Devreal

Nick Elprin and Chris Yang of Domino Data Labs, Q&A with Alexy Khrabrov of SF Scala @Nitro 20150205

Nick Elprin and Chris Yang of Domino Data Labs, Q&A with Alexy Khrabrov of SF Scala @Nitro 20150205

Recording: Nick Elprin and Chris Yang of Domino Data Labs, Q&A with Alexy Khrabrov of SF Scala @Nitro 20150205

hello everybody this is SF Scola developer Series today we are doing a Meetup at Nitra um with Chris Young and Nick elrin of domino Labs hey alxi thanks for having us great to have you guys uh and they are actually using Scala to power their uh web Cloud hosted data science uh website where data scientists can upload the the code and run it in the cloud and share and interact around this and they're using Scola uh to actually make this technology and uh the topic of the today's me up is uh Scala notebook which is uh interactive for Apple in the browser where uh you can uh enter a code and uh look at it interactively much like Mathematica or IPython but powered by scholar uh and uh Chris uh was one of the original authors of scholar notbook uh so I want to ask you Chris can you tell us about the Genesis uh of that project how it came about what's the history behind it yeah sure so scholar notebook was a project that I worked on when Nick and I are both at Bridgewater Associates which is a big sort of financial institution in Connecticut and so what was going on there was that we were interested in developing an interactive environment where we could use um Scala to do analysis analytical type work and so we looked out across the current art and the best we sort of saw was IPython notebook there were some uh rather undeveloped sort of backends for Scola and so we wanted to see what we could do in sort of like a six weeks Skunk Works project to see if we could swap out the back end of IPython notebook and make it sort of Scala based all the way through and so we Fork the IPython notebook um front ends we replaced the back end with Scala and there were some features about interactivity uh that we'll be showing off and talking about tonight that were of particular interest for us so it wasn't just that you could run Scola code that's obviously very important but we were particularly interested in developing some more interactivity so the ability to um uh actuate things in the HTML and have that propagate back to the Scola backend and actually have schola code run as a result of that so was pretty exciting stuff so uh Bridgewater open sourced it I think a year and a half ago or something like that and uh you know in that time Nick and I both left Bridgewater uh sort of for our own reason um but you know we still contribute back to scholar notebook because it was sort of a fun project and sort of a neat thing and it's been cool to see how the notebook environment has developed uh over that time to uh create a lot of feature parody and a lot of the stuff that we were interested in so in particular IPython notebook has added a lot of interactivity I think in their version two uh and in the meantime a lot of other notebook projects have sprung up as well uh and so it's been a pretty exciting space to to to be a part of great uh so I'm curious about the U environment and Bridge Water which was conducive uh to make this kind of project because we usually associate Wall Street with you know frenzy and and stress and uh and uh you know high frequency trading and all this kind of things so but it's not the first project which I see coming out of Bridge Water and others uh which you know I think they also do some Stu with uh F fshp and uh you know data frames and R so I've seen a lot of interest in open S stores coming out of it and what is is it typical for Wall Street firms or was it you know what was it about bridgew and I guess what is kind of the motivation there to have these kind of tools in the first place sure yeah I think um I I I certainly can't speak for Bridgewater now it's been a little while since I've worked there but from what we've seen in sort of Finance broadly and in other sort of um other high technology places places that are really technology to sort of deliver value to whatever business they're in I think there's an increasing recognition of the power of Open Source both as a way to get more eyes on a problem potentially get some engineering Talent um as a way to contribute back to the community like so many places are powered by open source tools and so there's definitely a sense that um in these places they should be contributing back and you know hopefully it doesn't sound too cynical but as uh a recruiting or uh uh a way to sort of get your name out um that makes those places an attractive place to work there's a recognition that they're not Financial firms aren't only competing with other Financial firms they're also competing with all other firms for talent um and so developing making it clear to the world that you know this place is a place where exciting interesting technology happens um draws the best people in great so they actually uh piggy back on the same strateg startups are using to open source this tools which kind of segs me into the uh Domino so uh you guys are uh building Domino uh can you tell me a little bit about uh what it is what is the motivation what the potential users and how it differs from the current other offerings in the uh data science and I this um yeah so you know the central observation I think we had is that so everyone knows the world is becoming more analytical uh and I I think what we're seeing is it's no longer enough for firms to just point a visualization tool at a database it's not it's not just about querying data and visualizing it anymore it's about running code on top of that developing models whatever you want to call them these more sophisticated analytical uh projects and we basically saw from the the of the companies we talked to that there was a piece of the tech stack that was missing to really facilitate that at scale within an Enterprise that would do things like uh let you visualize the results that you produced from running this code be it in python or R scholar whatever um uh to Version Control properly for data science which meant Tracking not just your code but tracking the data you used and the results that you generated making all that available and accessible and sharable um so that's you know that's the Gap we're trying to fill with uh with Domino it's sort of a central hub for analytics within uh within the Enterprise that facilitates sharing tracking scaling collaborating and ultimately um deploying or or productionizing or operationalizing models so they can actually integrate with other business processes so that sort of end to-end analytical life cycle this is great so uh what um what kind of uh uh thinking have you been going with when you chose Scala like what what are other options you know how is it working out for you can you talk a little bit about that um in case it wasn't obvious uh Domino is written sort of from the ground up to be Scala a Play application sort of um you know 90% of the code there is is Scala and so the decision for us I think was pretty clear um Scala is a modern powerful language that we were familiar with uh from our past experience in a production context and so just to drive that point home a little bit you know it's one thing to be able to write hello world in Ruby it's quite another thing when you're up your third straight night with a Java debugger attached to a jvm trying to track down a crazy deadlocking thing and so the power of Scala as a language was as important to us as the power richness and flexibility of the whole jvm ecosystem that allows us to build like an actual system that can actually stay up that we can actually debug and so the maturity of the jvm combined with uh the power expressivity of skull itself I think just was a total SL dunk for us um and um and uh I I don't regret it at all I mean I think it's been a a great experience for us um in terms of Just The Scholar language itself like we make heavy use of um AA functional paradigms um and so skull has been a great match for how we we've thought about engineering our our stack which is I mean and uh as a result I think the code base has kept has been has remained um no bigger than it needed to be like it's uh so it's offici ly expressed and that's let us scale our development by bringing on bringing on contractors pretty easily people kind of everyone we've brought in has had a very easy time getting up to speed with the code base which is better for us because it means they can contribute faster so I think it's paid off I wonder um do we have any Reflections on what it's been like to try and hire Scala developers like what's what's the state of what's the state of the of the market in terms of like Scala talent I mean I think thinking back of like all the interviews we've had and rums that we've screened you know the scholar development Community is definitely smaller than say like the Java Community but um probably for the better too probably for the better there's some self- selection and so the the skill abilities maturity of the candidates that we do see that are sort of like people that are interested in working scull has been sort of much higher than than that there is there is one there is one sort of anti pattern of I think it's like Engineers who are attract Ed to skull and excited about it for a lot of the more esoteric features um and so they they get excited about a language for its features rather than and sort of like finding opportunities to use those features rather than using them as tools to write well-engineered code yeah that's true yeah because you know there's uh scholet gives you a lot of rope which can be used to hang yourself sort of and so that's one of the things we look for when we talk to candidates is just cuz that feature in language is there does that mean you have to use it when would you actually use it well uh we've all been there I think the first time you learn about implicit conversions you're like oh man everything should be implicitly converted from one thing to another uh and then you sort of within six weeks you sort of realize what a terrible terrible idea that is so um but yeah I think I definitely agree with that yeah that makes sense to me great so maybe I'll ask you to uh talk a bit about data science community and actually I I have I guess uh two uh questions because I think in in addition to scull we also share uh galvaniz as a location where you guys are working for and I recently joined the Advisory Board and uh I really believe that this kind of technosocial community where you have a lot of virtual community center around data science and kind of physical presence around meetups and Conference a great idea also education programs and uh we'll have uh three conference this year text by the bay Scala by the bay and Big Data Scola hosted at gvan I and big data scal is actually about uh end to end data pipelines in Scola and data science on on on Scala so I wonder uh what are you thinking in terms of community building through Domino which is in the cloud and kind of how do you uh connect to data scientists around you at galvaniz and like is there any kind of analogies you can draw between how real life collaboration happens versus this cloud-based collaborations which are uh envisioning and building um yeah let me think about that for one sec well one thing to clarify by the way is uh we have a CL I mean the demo we'll do is of kind of the cloud hosted version of our product but there's uh there's a parallel deployment option where the whole thing installs on premise and so a number of our customers who are more secure more concerned about data security and don't want things leaving their Network can sort of still use our stack behind their firewall prodct yeah yeah yeah um but in terms of the community I mean well one interesting thing we've seen is that uh well there's enthusiasm for an idea like GitHub for for analytics in the world GitHub for analytical projects we haven't seen that play out that much and I think part of the reason is that so much of it depends on um on the data which is often proprietary and inherently not sharable or I mean sharable within teams but the idea of something more analogous to GitHub where you have these massive projects that are done at scale across the world um I think I think you run into a tension where uh there's something so proprietary to somebody a company or a government agency or whatever that it's hard to actually effectively Implement that um so I think the opportunity for Community is much more around like you're saying uh meetups and cross-pollination of ideas and knowledge and um new technologies I'm not you know it would I think it' be a beautiful world to live in where um there was something like a worldwide GitHub you know massive collaboration on analytical projects but I think somebody will have to crack the nut of um how do you share the the data necessary to do that not and I don't mean in a technical way I mean what are around just the proprietary nature of it the legalities around it pracy yeah yeah um I mean we you know we know well obviously there's interesting we work with interest with companies doing interesting analytical work work and um and they wouldn't release their data but you know but then there are interesting projects that we know about that you'd think would be good candidates for it like people in Academia working on working on these analyses but what you find there is that a lot of the data they have is is proprietary data from certain vendors you know like neelen or some government agency that um that they're under by contract not allowed to share with anyone else so that'll be that'll be nice to figure that out uh I can actually uh propose one solution so we at Nitro are thinking uh hard about sharing a lot of Open Source Code around data problems we we have which is in the document space so we are build on Smart documents uh which are basically PDFs where we can understand both their internal structure and uh human intent behind them and the social graph where they're shared so one thing we've done we are securing uh scientific literature such as sit here or archive. or and this data which is already open can be a model right so we can uh test a lot of algorithms and run a lot of extraction tasks on the public documents and then apply techniques learn to the private documents so I think one way to share uh code around data would be to share some common data sets such as Network price right and what kagle is doing so hopefully that can be one way but another thing you have you may have a large organization which will be behind the firewall and then those data scientists may actually uh collaborate through the private Cloud right and so I think like in my mind that that may be one big Advantage so so do you think of any featur such as commenting sharing and annotation of uh the code which people upload Domino what is your kind of vision uh around around this well I think um one of the things one of the sort of core beliefs we've built into the product is you if you talk about Version Control and sharing for source code really all you need is a source code but if you talk about doing it for analytical work what you need is you need the code the data and the results kind of connected together and by the results I mean um you know charts that you produce or models that you generate or you know model performance scores because if you're looking at your progress and you want to say how how's this analysis evolving those are the metrics by which you evaluate your progress so um and uh so I mean the way we think about it is you got yeah you're kind of your coding your data together produce these results and then that is a unit that evolves over time and so when you talk about collaboration it's what's useful is a linkage between those things and commenting both the somewhat at the level of the code but also at the level of the results you know here's a here's a um a uh performance graph we generated showing you know prediction scores I want to I notice something that changed I want to Circle that and then I want that comment or question to be connected to what was the change in the code of the model that led to the difference in results so the analogy the analogy and source code would be something like tracking the you know the binaries associated with each build of your code but with analysis the results are much more visual interesting so that you know I'd like to see all that that will be a great uh uh a great tool for collaboration and uh um one question which is common to uh Scala data scientist is you know how good is Scala a platform for data science how good is jvm a platform from data science and you you actually provide a lot of functionality to python people right so you stradle the worlds so I wonder you know what are your insights around this can what can we tell python people uh you know about the state of Scola uh for data science how can we invite them to the scolar world what should we as a scholar Community do to make it easier for them to come to the scholar world for data science try F this one um I think projects like spark have been pretty exciting to see as this as they've developed and looking forward to um the demo I think you guys are going to be doing about how you're using it internally um I think there is something interesting at least from what we've seen the data scientists that are using scholar uh tend to be at the even highest end in terms of sophistication like I think everyone that we know that uses Scola that way like was an ex software engineer or like could have been in another life if that's you know how their interests uh went earlier on and so um what's interesting at least in the state of the market now is sort of this this conflation between language choice and sort of technical Acumen and I don't know if there's a like causal relationship like you need to be this amount of smart sure yeah yeah we can yeah we can try yeah yeah yeah um and so I think that'll be interesting I mean uh that's a hypothesis that I have uh about the space like it could be that for whatever reason scol is just a more difficult thing to go learn in which case it's not clear how much extra tooling you would need I'm not sure what my opinion about that is but that could explain some of the like gaps between python python data science and Scholar data science I would say that in terms of the actual literal tools and the actual literal libraries I think the like jvm ecosystem is probably a couple years behind there isn't like a standardized data frame type thing there isn't a standardized numpy scipi um and it's not that there aren't equivalent bits of functionality uh in the jvm scattered across like a thousand different like Maven repositories is that there hasn't been a consensus coales around harnessing everyone saying like hey this is the Matrix Library everyone's going to use like everyone contribute to it and so it makes it hard to get started it makes it hard to I think iterate and have those things improve release on release patch on patch and so I know a lot of people are trying to do this I think we were just looking at um what were what were we looking at something connected to deep learning I think like DL forj yeah oh yeah yeah yeah and so I think there's some evidence that that's getting some traction but um it would I I would love to know how the The Arc of adoption of um numpy sipai and pandis like actually started um and if there are lessons from the adoption and coalesence around those libraries if there's some lessons that we could apply to the jvm great I think there are some good news there I've seen the data bricks road map for this year and they standardize on scheme rdd actually the data frame format and they want to interface with pandas and in other libraries and I think they actually uh putting this forward not just for spark SQL but as a general interchange format and so that hopefully will have a distributed data frame which which is probably hard to have in in our python or other languages and uh in terms of uh linear algebra I think there are several good libraries there is Breeze uh there is n d forj there is LA forj from Twitter but it's really a good question uh you know for the community to basically uh rally behind one of them and create a lot of tools a lot of documentation and kind of socialize that uh that uh as as a as a good uh Library if anyone can do it it's probably data break Apache to just say like hey guys this is this is how it's going to go so yeah I hope that works that' be totally totally awesome if there's some standard interop especially some interop with uh Python and other framy type things that would be pretty cool actually so yeah let's hope so and of course uh we'll have you know uh several tracks at Big Data skull in August to to kind of help the community to uh to uh focus on on on these issues uh so I think in closing I'd like to ask you guys what is your vision for domino in one two three years and you know we'll come back and and chat with you then uh of course we'll do this before as well but it would be great to kind of uh compare predictions and uh in the future at some point so I'd like to see your vision as Founders you know how how it's going to go going to take a shot when don't you go first go first I'm going to draft off if your answer I think um well the you know the goal we're pursuing or the mission we have is to empower this new wave of more powerful analytical techniques and so be that machine learning or be that just much more sophisticated um you know computations to better do anything within an Enterprise it's this idea that more and more businesses are running their business not on data but on analysis um and so I think that means a bunch of different things I think that means better compounding of knowledge and uh being people being able to Leverage The ongoing work that other people within their organization are doing I think it means um reducing the time to test an idea once you have a thought for it so reducing the ratio of um uh thinking to doing time or something like that um I think it also means making it more fluid to take a model that you've built and deploy it so that it actually affects things in your business uh it's not just an analyst who runs something on their machine when somebody asks them for an answer but it is it is uh seamlessly integrated into the operations of the business itself so I think uh it's all those things that we're pursuing um and that's the way I think about you know how we prioritize the different features we develop so that's the level at which we hold a vision it's not a level of um you know the product will have these features and we look like X Y and Z but it's sort of it's whatever the most sophisticated companies you need in order to make these powerful analytical techniques really um uh Drive the the actual core work of the business they're doing sounds like a great plan Chris do have anything to this I think I may have successfully drafted off your answer by just saying Nick I think that was a great articulation of our vision so I think uh no more words are necessary all right great thanks guys uh for sharing these insights and we're looking forward to your talks uh tonight and uh great to have you in the community thanks Lei thanks so much thanks