Devreal

SF Scala Nimbus interview 7 26 16

SF Scala Nimbus interview 7 26 16

Recording: SF Scala Nimbus interview 7 26 16

[Music] hello everybody I'm algan of Blomberg One Of Our venes to we usus Who is Spark engineer bl welc so given The Nature of this meet up and how we trying to do you know Kind of a Group project Night Interactive and stuff I Had this other Talk that done Before but it's not very Interactive so I just I work with an intern to prepare that is Interactive so we Gone and used dats Cloud prep notebooks can import and edit them and Just Try things Out And I think we got a Pretty Good starting Point we got like 20 or so data sets pared into data frames and we've got an example Where We're Going from One Of Those data frames transforming It featuring it and then Running naive base and linear regression Order to predict on One Of The columns And then Another ex dats Together and you know in this case we got tax data that by Zip code that has like you know the income levels and such and also some farmers market data Which has is Fairly granular and indicates whether certain types of vegetables are present Or Not So I think there is Some interesting hpo People Come With Them out one of the most popular hpo at the a little bit about how you got into this Field so you've been a regular you know host Park for very you know grateful and Kind of Tell us a little bit how you got into this Field why bloomberg why Spark Why It's All Together Now so Working with bloomberg for Just over a Year and was bally Because Of My Experience with Spark and Before that where I got my Experience was this startup called Radius Which basically does Intelligence on businesses for marketing and sales purposes and so for for Years I was basically evolving Kind of a similar Pipeline that does about the same thing that is resolving and D duplica large data sets of business locations and such and initi there solution with Elastic search and distributed Python stuff that you know we got to a certain point It Like takes too long to Run so we moved on to hup so I not having used doop Before Just seeing that is you know an appropriate tool To Use To solve The Problem you know I Spent about Two Years basically evolving A Pipeline initi doing native M produce then later into Just a complete cascading Pipeline and after that know Spark This Was Around Just Before Spark 1.0 When I was Kind Of seeing It as as a Real Thing That We Might practically switch to and and yeah I helped I helped basically make that happen at Radius basically writing the the existing thing and Spark and getting things done and even doing some Some things That We Wouldn't have Dreamed of doing in hadoop and so i' using Spark for a Year and half When I was in Do moreo I mean I do Enjoy applying particular data application but I would like the tools to be in a better place so that that work is more effective So You Know I've Had A few ideas about ways in Which The overall Experience of developing with Spark making These applications can be improved and That's what I'm trying to do Here at bloomberg andion There There's a Fair amount of Spark adoption happening across various teams bloomberg there you know There's like 4000 Engineers and you 15 teams or More That are trying to do something with Spark 4000 Engineers at bloomberg not 4000 do they have have util [Music] FR 4000 So basically You I find it very exciting that bloomberg like Years AG I to bloomberg Bet and Spark was New to them home mon mov quickly so I what you think bl Spark And how used Does It find Option Here Where do you see this Whole Company of in the tooling in so bloomberg Is basically You Know data Company It's information provider It Core product The bloomberg Terminal is a platform to use get act that Platform There spring amount of information that is coming in and Being Served up and There's several areas that Spark can be useful for that you know Among Which are Just directly supporting Analytics in the Terminal so There are certain things Kind of like calculations on of compute cared out and it's not something that you can necessarily anticipate before and Just Keep Around As An index Thing certainly some Of Those things can do that but some Of Those some Analytics work with that but There are some Where to be perfectly flexible It Would Be Nice If You Could do that Kind of computation On The Fly in a distributed Manner And that is one one of the M uses that People Do With It Here other uses so certainly There's a lot of internal data sets that people internal to bloomberg want to do machine Learning in order to you know build things and improve The product I recommendations and impr search capabilities and such and you know to do that effectively you need to be able to create These Models train Them Back test them and do it in a Time effent and also you know notebooks One Of The Things that know Kind of is is Kind of Being Being rolled Out is are These notebooks in the Jupiter notebooks in the bloomberg Terminal at the moment There they're accessing The bloomberg API so They Deploy on client machines People can Run code In The Notebook Access The bloomberg API and There's a lot of libraries provided that help that Out but One Of The eventual uses could be Out They with Access to the Blomberg dats There There are few exciting uses for Spark This is really [Music] for van