Adam Pingel, Q&A with Alexy Khrabrov of SF Text @VigLink 20150210
Recording: Adam Pingel, Q&A with Alexy Khrabrov of SF Text @VigLink 20150210
uh welcome everybody this is sf text the first inaugural meetup uh i'm the organizer alexis kravrov and we're on location at wiggling the company which works with a lot of text and uses very interesting technology uh we have uh two talks today from wiggling and from grant and gersol and we are going to talk with adam adam pingel who is our gracious host uh he will first tell us about uh wiggling itself so adam what is weak link what are you guys doing well um we help publishers monetize their content in a sentence and that comes in a couple forms we help help them affiliate links that they might have already in their content to product pages and then we also can identify product references and text and turn those into affiliated links and as we're technology people and we like search and nlp and tax money what are the underlying technologies they use what are the challenges in in doing this well um just to speak to the scale of it first we we work with tens of thousands of merchants we do many billions of page views every month so just scaling scaling the load is a challenge we run in three different amazon regions databases we use our cassandra mysql uh we use elasticsearch um so so there's that side of it and then on the algorithm side we have a whole data science team and a couple of nlp experts who uh who tune our name identity recognition and our matching technology um to make sure that we are linking to high quality products that are going to yield uh the highest possible expected uh revenue for the publishers uh and i know adam from sf scholar meet up where we've been colleagues in a long time so i wonder what is the rationale behind the choice of scholar and uh its uh associated technologies yes um we are historically a java shop um and so i think there have been several attempts over the years to use scala i think the gears really started to mesh a couple of years ago we started using it in the feed systems just to kind of get our our you know our feet wet um and it spread from there so now i think close to about a half of our ad hoc analytics uh queries are written in scala we do spark we're using akka to do a lot of our feed processing now we're slowly introducing akka or play-based uh web services into our infrastructure over time um so it the the fact that we're coming from a java background makes to some extent makes that a lot easier because it the the migration path is very can be very gradual and we can go as at a pace that feels comfortable to us um and then we also have a bunch of folks who are a little more well-versed in ruby or python but i think i think they also find that the functional paradigm is pretty pretty natural for them it doesn't require too much ramp up time to get them productive cool uh can you talk a lot a little bit about the nlp technologies you're using uh i know you're related to some of them i wonder what your experience with them you know uh can you recommend some of them so uh yeah we we use mallet um and i'll have to refer you to our nlp experts for for more on that um mallet we've we've looked at epic we've done some benchmarking with that i think katrina is going to talk about that tonight um just for the name that any recognition we got we did get some higher f1 scores but katrina's going to go into some depth in that later tonight um mallet's natural successor is factory we've we're aware of it we've looked at it we haven't yet benchmarked that but the epic and factory are the two um if we were to migrate beyond mallet uh those are the two most likely libraries we'll start using well and so what are the main challenges in in working with actual data um and what has been challenging seeing kind of organizing this information connecting it uh and a little bit about that wow that is a good question um gabor is going to speak a lot about creating uh training sets i think uh knowing when you're doing the right thing is challenging so getting all of the infrastructure in place to um to set up a b tests and make sure that that's all backed up with with solid statistics uh it's very challenging um the the traffic we see on any given day it varies it varies a lot the products that people are searching for that the seasonality of the business um it can be it can be difficult to know you know when a change you're making is really having desired effect so um i think that's that's one of the challenges i've observed um but again i think gabor is is the one who's really leading that charge here um uh let's see what else would i say this scaling search we do we do use search we've gone from using a single box running lucine embedded in a tomcat container to running a 30 node elastic search cluster over the last five years there have been a couple of points along the way um before before there was before there was shard sharding on solar we had to kind of do our own that was fairly brittle um so about a year ago we moved to elasticsearch um yeah i would say between the data science science and the and the scaling issues we definitely have our hands full dealing with this this data and also products products themselves as a rich um understanding a product is more than just its description there's a rich vocabulary of attributes that products might have a rich category a taxonomy that you might build uh to to classify um products so so that work is all ongoing um and still very much we're learning new things about that all the time great yeah uh so so it's interesting because you have essentially a team where you have engineers who have data scientists who have nlp experts how do you guys balance all of this how do you kind of make sure you have a production code in the end which runs a scale uh can you talk a little bit about uh inter play of these teams and uh how do you get to the actual working yes so we have we have a little over 20 engineers here um most of which are located in san francisco although we do have some remote offices we've actually just recently because we've been growing so fast split the engineering team into two halves and theoretically each team is capable of doing any kind of work in practice each team has has their affinities has their their specializations um so we we found it's beneficial to to keep everyone you know well appraised of what what's going on across all of engineering so we do have data science specific events we have something called the data science jam every friday at 11 o'clock but that's that's open not only to engineering but to anyone else in in the company who wants to to check it out it's also a good way for for the data scientists to kind of uh you know have a produce a deliverable and have a do a talk for a half hour so um we find that's really what are you guys doing these data science jams uh we'll talk about um we basically do two or three presentations over the course of an hour and uh the topics can be anything from restructuring the way we're doing click ids um and the math behind that um to to some of gabor's uh research on shallow parsing it it really runs a pretty wide range of topics statistics behind the kinds of a b testing we're doing a lot of stuff awesome do you see uh people learning and like engineers not yet familiar with these concepts wanting to actually pick it up and do something about it absolutely um yeah we actually we just had a hack day um which was a good chance for some of the engineers who don't get don't touch that stuff day to day to interact not only with the data scientists but but folks in marketing or operations we had a pan company hack day so that was a really great opportunity um to to give folks enough give folks a chance to work with stuff that they don't normally uh get to work with day to day excellent excellent uh so grant just join us and you will have to explain to him why you picked classic search over i know i'm gonna have to give him a heads up yeah so this is great so but since this is the first sf text method and uh and we are set in the direction and if we link has taxes at score i wonder what would you as the host and engineer uh running a team here what would you like to kind of get from these meetups and going to the future how do you see the most beneficial uh composition of these meet ups yeah well i think uh a lot of the meetups here in san francisco tend to be technology oriented this is something that katrine has comments to me about from from joining us almost a year ago she noticed this so i think having having a community built around a problem domain more than a technology would be really really useful to everyone here at big link and of course just having people over in person meeting them face to face a lot of us see these people in in get commits or in papers that they might write so i think just to establish uh more of a personal connection will be really helpful to everybody here um i think there's so much new there's so many problems uh so many there's so many unknowns in this in this business with this technology um that i think the upsides are clearly outweigh the any uh uh any i don't yeah maybe i'll allow that the the part but yeah i think i think the the the upside of sharing information uh is as far outweighs any uh any concern about uh you know ip or anything like that i think i think we can we can you know balance that and teach each other how to use the tools learn the algorithms true develop open source projects together potentially yeah i think open source is the key work here because i think the this group is coming together around shared open source and we have academics producing it we have practitioners taking it and hopefully by uh providing feedback yeah uh we'll improve uh our common open source and share as much as we can absolutely yeah well thanks adam it's great to be here and i'm looking forward to to your presentation all right thank you yeah we're happy to host thanks you