scala.bythebay.io: Adrian Mihai, Reactive Resumes
Recording: scala.bythebay.io: Adrian Mihai, Reactive Resumes
hi guys so delighted to meet you alien obviously from Dublin Ireland to try to give you guys a one-liner and so advance apologies for a bit of a goofy sort of definition ciasses kind of like silly for race you might so it basically built a knowledge graph in recruitment and we use that to train machines understand resumes we deploy that kind of stuff for job boards big recruitment agencies that kind of stuff so we're able to say powered monster comb and much older candidates with all open positions based only on resume file so what documents not kinda stuff right so concretely what is it that we do three areas in particular so ready my party at scale these symbols so obviously mostly I oh how are we doing that we are deploying a pipeline an application or an akka cluster application our workers are actually akka camel actors if you guys are familiar with camel to enable the integration we do Enterprise sort of data flows we ingest files from all kinds of sources so in our case know industry we can find rise you might so on hard drives on clouds you're pretty much everywhere so we actually over email by the way really important so we grab them from all over the place and yeah we post them second thing that we do information retrieval so identifying patterns within structure and phrasing of resumes so think linguistics right now not necessary words like TF you know write and match them with job descriptions automatically we of course tapping to external as we analyze the resume as we tap into external sources of signaling in in our case github which is very important so when we detect github links we go on fetch repositories actually non fort ones so really source code and then we grab all that kind of stuff sample it and expose it on the interface we also type in 30 tapping to LinkedIn all the cast off based on that we do real-time analytics so we do salarik recommendations for one so we mined 15 job boards currently as we each would I would say fifth position has a salary published right in terms of minimum maximum so by analyzing all that kind stuff and their training a couple of regressions and translating that on top of resumes we are forecasting because forecasting salaries for candidates the really cool part in in this area is the fact that we forecast also skills but we forecast skills that maximise can that the candidates set or candidates salary so if we realize that you kind of like have a background in Java that kind of stuff we probably will recommend you or closure or scholar or things like that because that will maximize your salary of course like we have lots of similarity between candidates search with job descriptions on the cast off right so first of all we don't actually have that many rights but when they arrive they actually arrive in the order of millions so see these in terms of a recruitment agency or having to ingest all their current database which are only fast really so what do we do in order not to have to go down when we hit the spike of say like lots of data we actually push data directly onto s3 from web program from everywhere and after that we start batching we have so I will actually talk about the whole pipe end-to-end pipeline which includes the front-end interactions backend which is no Jess we use that only to bridge requests and then obviously the arc are our application and I think get stored into Couchbase oh and by the way I think is immutable right so in our case I was saying I think is pretty much a stream so we start basically with documents we have an agnostic class photo on the front end reaching your back end which in our case currently it's a WebSocket or Ajax with each node from node we basically start querying all our internal pipelines and that's happening in our case we also are transport agnostic on that data and in a sense that we were to actually had your transpose like Kafka HTTP 0 and Q we were playing with and all that kind of stuff we enter on Augustine's pipeline which starts doing lots of flat maps and all that kind stuff and we output basically and the other define online the JSON data structure right so not to you know this is some sort of up pseudocode probably a bit smaller text but in big lines what he describes is is actually you know first step so as somebody uploads something of a cloud we basically only need to download that currently s3 we get back a battery output then we do combustion in our case conversion is performed by a master's cluster so docker some docker containers deployed over a masters cluster all the combustion is actually in memory so I will I will insist over that then we really start parsing you know in we start basically we convert I think of PDF we use PDF books we perform a bit of named entity extraction tokenization regex vectors all a kind stuff then we proceed with the lots of context so extracting context in our case education for instance that's elasticsearch so in our case currently we use percolation so we because universities we know which the world you world universities are we what we do basically we fed that to elasticsearch and we do an inverted sort of search we just pretty fast ok then although we also do one really razzuma we do lots of i/o so we if we detect the portfolio's we go and actually screenshot those ones if we detected I was saying github links and things like that we go on fetch source code all that kind of stuff it's it's managed by our customs I will actually talk about that and in particular we love the ARCA DSL the fact that we can actually broadcast the flows and we can schedule you know I all sort of flows one way or the logic on the other other way and Ola after that you we have a series of AI calls so mostly regressions maybe some interactions with some tensor flow models then we need to make sure that we can actually kill this stuff so it's basically passing through a shared or kill switch and in our case all the monitoring basically all the instrumentation all that kind of stuff we use Prometheus like it's an excellent thing we collect data from no deal on those carbs pretty much all over the place and the systems as well when I think is complete basically we return on next EP respawn so all these kind stuff is pretty much one click so it's one request really alright so as always he explained basically you know we actually have a pool of workers business logic all that kind of stuff traffic arrives on to s3 we start anarchist attacked us 4k orchestration and then I think is pretty much your single scene all right I would like surely explain right now how will do internal the workers how are how do they look basically so what we have in here we have a series of workers it's a word correction in our case means a producer and the consumer producer in the sense that we want to keep them agnostic as I was saying and that producer in our case that's the akka camel actor executor all have extras read access to hazel cast so just not to you know to have to tap into into data if we need it most often we don't but there is a hi hazel cast layer on top of it and your is the data stores basically we use the couch base it's immutable which Auto replicates exerc to elasticsearch we do some really really cool stuff or descent and I will insist over that at the end right so I on how do we do are you currently we deploy docker containers over what do they contain it's actually simple golang web servers written like in we use Ivy's framework I I believe and that's an example of you know a simple conversion so what do we do we actually mount around the schema in the docker container we use the LibreOffice and basically when a request arrives we just simply you know execute that stuff on to run disk and I'll return the answer back how we deploy this kind of stuff with scale we have another container in the system that subscribe to marathon or even seen and it's upon nor doing the hard checks it pushes although the correct all the instances all the IPS and the owner cast off into console and the from console we actually have a load balancer or HTTP load balancer which is Fabio written in gold that simply checks up to a console and the update the routing table so therefore we are actually able to tap into any of those containers we are simple HTTP requests I think he's basically route EDI nginx on some important the request really the aisle conversion is for us it's like something like this so based on request and Marshall the response on what you get back is even a byte array that happens with the both documents or Word documents or PDFs or dzifa's so older if we have many many files maybe in a zipper zipped we also perform more combustion in memory Oh deflection in the noise oh all right so how do we interactions work so we have obviously the front end in our case if you guys are familiar with the reactor all that guy stuff so we have a reactor view layer we also have a data layer or baobob currently so it's a state tree it contains all out front end States in one three lycée computation all that kind of stuff and all interactions are managed by cerebral see it as a state controller the beautiful thing about this guy is the fact that it actually allows you to decouple both the view under the moodle so in our case react can actually become angular really really quick we don't necessarily need to do that or you know high performance if you guys are familiar with react you also have a drop in like inferno the other the other side so the data port the model part it's actually we actually were towing with the immutable J so that I stopped so it's it's a simple data to write what goes on is always explaining it so we actually are performing your request it arrives by you know jsoc's of futures so in the and then responds back alright so in our case i was explaining that one worker in our case is composed by a producer and a consumer so the consumer so this is the Heysel cosplayer the memory memo Laird we also have a near cache I think h2k or something like that and this is the producer that we have a series of traits they could actually consume data from craft our currently it's mostly HTTP archive HTTP but yeah why do we do why do we have this kind of setup so we basically don't want to overload executors so that the smallest parts of work in our case this kind of stuff we don't want you to make them aware of you know cough context and communication all that kind of stuff and yeah basically it allows us do we know do super do maintain like really pure functions as small as less focused as possible scale of course like you were using a round tubing from round robin from more from our car we are using Yaka cluster to to scale horizontally isolated the octave system so pretty much no one I think is location transparent all the constant right so what's about Couchbase so if you guys I'm not sure if you guys are familiar with it or not in our case that's an immutable database so basically there's no update in place showing inserts what does this means for us is the fact that we can actually get free analytics of that because you also always have you know previous sort of data which you can tap into also Couchbase the difference between CouchDB for the guy for the ones of you that are familiar with it and Couchbase actually another layer which is man base like man cache sort of sort of thing we keep them all in the same same system why do we do that because if we actually want to tap into you know to interact with other systems we would like to be kind of like absolutely flexible enough to interact with the different other sessions from PHP for things like that which usually use memcache right-oh elasticsearch so obviously the classical case for search elastic searches you know such so we're not returning full documents we're only returning document keys and then we fetched up from the key value stores or from card place I was saying also elastic cell percolation so all inverted search we have the universities in university names we search with a row text and we get back really the matching matching universities in text right so why Kafka in a pipeline so we do that in order to do logging you know not to to do a synchronous logging obviously and job recovery so basically we don't know exactly we don't really want to lose candidates right so that's very important so if we fail we want to be able to know what kind of data fails so to be able to sort of reschedule he start I was explaining when we see for uploads we receive uploads from multiple sources so in our case batch we have cross-platform application as well allowing our ingestion or mac OS and windows mail web that can stop the flow I was explaining it's basically we send everything to s3 and then we bought batch afterwards and I'll get a bit into why I write now so in our case we do you know we write software for recruitment we're not recruiter sources so we actually need to adapt we we considered it a really wide wise decision to actually do define AI as you know parts industry specific data and parts sort of generic linguistics matching so for IT you know in in our case so we maintain basically two vectors describing each document write an industry-specific vector which is in case of IT we define that to ourselves and we we had no problem with that as you know software development web design data science stuff that people can actually relate to so it's an industry specific topic vector and your by the way we actually were looking we are assessing lots and lots of ways on how to generate those kind of vectors in our case we don't use classifiers at all so it doesn't really work you know doing skills on the abilities for that so we we actually future engineering is crucial in this in this kind of case and in our case the topic vector for IT it's like you know it's 11 all the topic vectors are 11 elements so it's like a really small vector but those 11 or elements in that vector are aggressions all of them also we also maintain on in the document vector so this is in the space of linguistics for the ones of you familiar with the work to back that kind of stuff so we map a thing onto onto a core knowledge sphere right and in our case those are you know higher dense vectors or in our case 300 lines within a semantic space so I have a couple of visualizations I have them live as well so we're actually when we're talking semantics we are for data science that's what you pick much you want to see not Islands nothing's like that so really nice neat area describing the document so in this case the right you mates leaning towards or data science that kind of stuff in this case we actually it's not really visible we are analyzing what kind of information meaning generated towards that the document vectors contributed to do towards doctor document vectors contribution how do we do similarity so in our case we have lots of similarities or in the six Pacific on use cases in case of IBM let's say they have one hundred thousand employees and they want you know for my next team or next best team around you know a particular set of job descriptions so those job descriptions of course are defined as well by vectors in that and in this case basically seen righty it's a pretty much cosine similarity or dot product in case of topic vectors yeah it's mostly supervised yes we maintain lotteries in case of document vectors obviously unsupervised how we do actually the document vectors I won't insist much but it's actually a process where we first of all we know what the topic the topic space is we do SVD just for a particular purpose so as we analyze the regiment we cluster all the information in that resume via K means so the problem in k-means as you may well be aware of is actually the K the number of clusters right so what do we do about that in order not to you know go blindly we basically do SVD figure out what the principal elements are we kind of like strict strip out the information and then we actually find the K that's the only purpose why we use is the SVD for once we have the number of clusters we we cluster things we employ our own glossaries to further filter out irrelevant clusters and then we have our own sort of bl2 computer document vectors from from those kind flusters so to be honest still we were looking currently on that that's certified it's you know lots of lots of voice we were looking at probability distributions ATP by the way like a hierarchical distribution so it worked for us we have a bit of a problem in the sense that we we it's still not perfect so our current area of research is not only defining a document by one vector we actually want to define a document by four vectors so in our case you know something like Java Scala that kind of area and I know business marketing which is a different sort of thing and then use something like words move or distance but it's a bit slow if you if your articles we could be declared you yeah so this is under constant research for us currently yeah right so all the similarities I was saying we are doing we are performing here cosine similarity if we actually have a really large number of documents we use our noise from Spotify and excellent project that has the currently has the binding so Python and Java from what I'm aware of so he fine may be wrong somehow so it depends from customer to customer so we we cater to recruitment agencies right so we don't when we when we're talking of a pool of candidates we of course we don't open up so everybody is like kind of isolating their own space so if we were powering jobs I you won't get to see us we were underneath right we don't even know I think it's pretty much bound you know confidence even you know all the data all the fun stuff right so I was saying again sorry sorry forecast I already explained that I believe so by mining that kind of stuff we realize how valuable our skills are we Dino is lots of stuff and then we translate that that's place into resumes and that's how we actually focus to our salaries this is a really interesting use case actually so mine we start mining with a simple simple Python - the process you know scraping basically running on Crone nothing really complicated we face data into Couchbase from college based in auto replicating to elasticsearch brought on the elasticsearch side we actually have a small plugin which tries to figure out right so is this document passed or not and if it's not it just triggers a small icing message or over Kafka or somewhere internally and a series of processes will actually pick up the document this document ID will do all the AI and then we'll be rewriting Couchbase which auto replication elasticsearch so for us what this translates to actually it's actually free sort of [Music] interaction if you will so we just have to write data into into code base and call it a name so we don't really need to do anything afterwards because I think that after its scheduled bye-bye that sort of round the long pipeline from time so as I was saying we are employer react just to full view we I have an example of you know an imposed form or Dropbox currently it's basically a simple WebSocket message or Ajax message breached by node arriving into into an arc esteem streams pipeline and once that's finished we send a message back with a confirmation that's pretty much it so it works exactly exactly the same in case of you know 1001 million files as it would for one single single file alright so how I was saying all the state is unified JSON state and font and controlled by your by cerebral view current really lacked Bob objects which has a cool source which are I would say cool I like elements within the three basically lazy state computation so whenever you're changing something in the state three I think as recomputes as you use it okay should all not cast off I was saying you know industry topic vector so this is where you actually can see one so in our case we have actually 11 categories on bottom and then you have a quick heat map over you know candidates abilities when we talk about IP we're actually talking in this case of you know data site where probably 200 percent and software development so for you to have a quick overview interactions so basically it's on the front and side it's kind of like the same thing as you do imagine on happening on acha acha seems i was explaining the whole information flow battery scales so we provide the cross-platform applications as i was saying that simply allows our customers to ingest folders a fast so in case of recruitment agencies everybody has files all over place hard drives all all the kind of stuff we just give them on application allowing them to browse for folder and ingest that container I also as I was saying we maintain a CLC's of docker containers so in our case you know for combustion go to PDF PDF to images so utilities really inside they they do have web server so written in go line or we also have a container which subscribes to the marathon or even stream whenever something happens to some of those containers this guy will update the console a console [Music] will have a console and then Fabio will actually ingest figure out which are the hata services and then expose them as a section p4 also monitoring instrumentations or prometheus graph and all that kind of stuff no than the Java clients for systems views are not exported from team prometheus for the couch with more to touch base we use the Telegraph plugins plugin alerting when something happens so we used olive manager for that so we use basically a couple of dashboards holding or anything together so both systems earned performance and you know isolate all all the kind stuff oh thank you so much I will try to pick up on a couple of questions if you bunch of tools usually don't need to because once you mine it you kind of like have an idea on the glossy we if we detect you know such things we definitely do but we didn't see the case actually the needful for that no it's actually in our case when we compute all the document vectors we know we know which what skills mean all that and stuff so we don't really need to actually with us and we'll be back here a second thanks a lot Adrian saying some [Applause] you