talk · community record
Parallel Corpus Approach for Name Matching in Record Linkage
First-author paper presentation at ICDM 2014 (Shenzhen), with Leonid Zhukov and Alexandrin Popescul of Ancestry.com. This is the peer-reviewed version of the same work he later gave at Text By the Bay 2015: casting alternative name-spelling matching as a character-level statistical machine translation problem, trained on a crowd-sourced parallel corpus built from Ancestry's genealogy person records and user search query logs, and shown to beat standard phonetic and string-similarity baselines on precision/recall for record linkage.
01
Connections
1 relationship