talk · community record
Unsupervised NLP Tutorial using Apache S...
Paraphrasing Tim O'Reilly, the person who has the most data wins. That's a neat slogan, but the more data one has, the more likely it is to be unlabeled. Unfortunately, there aren't that many unsupervised learning algorithms out there, for machine learning in general and for NLP in particular. Recent advances in deep learning provide new tools for text mining of large unsupervised datasets. In particular, I will talk about the math, intuition and implementation of the word2vec algorithm, its variants (skipgram and continuous bag of words), use cases, and extensions (e.g. paragraph2vec, doc2vec). I will wrap up with a simple demonstration at scale using Scala, Apache Spark, MLLib, and the Apache Zeppelin Notebook.
01
Connections
8 relationships
aboutApache Sparkproject ↗aboutApache Zeppelinproject ↗aboutword2vecproject ↗affiliated withNitrocompany ↗documented byText By the Bay 2015photo ↗presented · incomingMarek Kolodziejperson ↗presented atText by the Bayevent ↗recorded asText By the Bay 2015: Marek Kolodziej, Unsupervised NLP Tutorial using Apache Sparkvideo ↗