talk · community record
Classifying Text without (many) Labels
Supervised text classification is often hampered by the need to acquire relatively expensive labeled training sets. In some embodiments of the systems and methods disclosed herein, pre-existing Word2Vec or similar algorithms are leveraged to create vector representations of documents that enable a model to be successfully trained with a drastically reduced training set. By using this technique the implementer can now devote low investment to acquiring a small volume of labeled data examples in order to train proximity thresholds, without devoting significant resources using traditional text classification machine learning algorithms which typically require training volume examples that are orders of magnitude larger.