talk · community record
Semantic Indexing of Four Million Docume...
Latent Semantic Analysis (LSA) is a technique in natural language processing and information retrieval that seeks to better understand the latent relationships and concepts in large corpuses. In this talk, we’ll walk through what it looks like to apply LSA to the full set of documents in English Wikipedia, using Apache Spark. Harnessing the Stanford CoreNLP library for lemmatization and MLlib’s scalable SVD implementation for uncovering a lower-dimensional representation of the data, we’ll undertake the modest task of enabling queries against the full extent of human knowledge, based on latent semantic relationships.
01
Connections
8 relationships
aboutApache Sparkproject ↗aboutMLlibproject ↗aboutStanford CoreNLPproject ↗affiliated withClouderacompany ↗documented byText By the Bay 2015photo ↗presented · incomingSandy Ryzaperson ↗presented atText by the Bayevent ↗recorded assfspark.org: Sandy Ryza, Semantic Indexing of Four Million Documents with Sparkvideo ↗