Devreal

Panel II: Data Engineering and AI

Event: Scale by the Bay

Scale By The Bay 2018: Panel II: Data Engineering and AI

Recording: Scale By The Bay 2018: Panel II: Data Engineering and AI

you so it's my great pleasure to introduce the panel and I will introduce the moderator Lucas we walk and he will let the other folks and Tristan's also he will use them so Lucas is a long term friend of this conference he's being dated by the bay events and he found the company called CrowdFlower which when I was at clouds we were one of the earliest customers of CrowdFlower so they pioneered the whole notion of human in the loop and basically this kind of you know artificial artificial intelligence right when people realize actually we don't actually will have humans maybe they're even better than the AI and I think this is becoming even more and more and more important right so and and now Lucas actually went forward and he found different company called weights and biases and he blogged about going back to the fundamentals of deep learning and hacking his garage and like building deep learning like understanding how this works so I think it's super fitting that you know where the comfort which connects data engineering Nai building tools for engineers to Ronnie I understand that actually does what you want so and we have a lot series of folks who can actually explain how to do it best so it's all to you Lucas all right thank you so I'm not sure I've actually ever moderated a panel before and I was really excited to do it especially because four of you and I consider pretty good friends of mine I've know you well and one of you was actually really excited to meet so I actually was hoping I could do the introductions if you don't mind this is my personal opinion about all of you everything so Pete is uh-oh Pete I think I first met you as a customer of CrowdFlower great choice he was he was actually a data scientist at LinkedIn back when they had you know kind of one of the most forward-thinking data science groups and then he started a super cool startup called skip flag which kind of helped with knowledge discovery inside companies and they were very quickly acquired by workday which was a super cool super fast exit I'm sure your investors were thrilled and he's now actually writing a book supposedly writing a book I haven't seen it but I'm really excited to read it about AI and product and how you actually kind of make I work inside products he's also a social media celebrity he got us like 25 customers there's one tweet we didn't even pay him he'll get a sponsor so we have like a ML pundit and then so the Michelle is also I her as a crown flower customer great choice work for Cutler but I can't out myself so Michelle is actually a senior engineer on one of the coolest projects cube flow that's got an incredible amount of traction and just I think like nine months since the launch it's I think one of the most important projects happening right now in open source and AI and she has a lot to say about Scala that I'm excited to to learn about so I think really the one Scala expert on our panel today I think unless anybody else wants to claim expertise so Stepan runs hydrosphere he have our CTO of hydrosphere which is a super cool ml infrastructure company that helps you kind of store your models and deploy your models and make sure that everything is super reliable so it's not just the successful company that he funded himself which is incredibly impressive incredibly hard to do it's also a super popular open source project so incredible achievement there and so the one per today I don't think I've met you Chris but I know of you pleasure to meet you so Chris runs pipeline that AI which I was notable to me I see lots lots of companies websites but your website really stood out to me for its amazing number of high profile kind of fortune 500 customers and its aggressive use of Comic Sans which I just thought that combination was spectacular and intriguing also a super popular open-source project all right and Richard suture I mean you just saw him talk you know I I knew him as a start-up founder of a really exciting startup also one of those popular teachers Stanford has ever had apparently is retired from teaching now which is your loss but the videos remain still online yeah and you should watch them they're really really good I actually learned a lot of from your videos even though it took that class a long time before but you really updated it well and he's the chief scientist of Salesforce what you might not know is that Richard is a avid motorcyclist and paraglider and you should ask him about the time that he tried to do a wheelie on his motorcycle when he first don't ask and is his Instagram and Facebook feeds are truly they'll make you feel really bad about yourself so so alright so I was excited by this panel before you start new people have questions they want to ask I was kind of tweeting it out but I have a ton of questions let me just refer to my notes here for the top questions that I saw so I actually want that start with Michelle I'm sorry to do this but I can't help myself so I I was kind of curious about your thoughts on like Scala versus Python when it comes to machine learning and your experiences since I guess really relevant to the audience here I think I think you should start [Laughter] [Music] Scala I was introduced to Scala via spark so a little bit of background so I'm an engineer at Google now working on cue flow but prior to that I was at a few startups here in San Francisco building machine learning products on GCP and AWS and then prior to that for about a decade I was a data engineer in the enterprise so spark was really important for me that was kind of my bread and butter before I got into building machine learning products and so it was a very natural tool to reach to reach for whenever I started at any bond using CrowdFlower like sure spark belongs in this toolkit I don't think anyone at any bond was using Scala at all but enter Michelle and I was really excited about it and I think we we ended up releasing a platform idml which I don't think is alive anymore but that was all Scala base and one of the first talks that I ever gave publicly was that Chris's meetup and it was all about using Scala for machine learning it even was a like machine learning models as as a service company and so when I moved to another startup of course like follow the same pattern from before like of course I'm gonna use Scala to build these machine learning products and it works pretty well but I wasn't on AWS so I was using GCP and I tried to sort of hand that over to the engineering team I had docker eyes this really neat classifier tried to hand it over and they're like this is not a jar I don't I don't know what to do with this this is not gonna work and so that's how I discovered gke so it turns out that kubernetes works really really well for deploying machine learning models even if they're in Scala and we ended up moving all of the Scala code into kubernetes and it just really transformed everything to see a machine learning application alongside the rest of a sort of traditional application and so fast forward to joining Google in my second week like I'm still an orientation and I this team who had just launched this cute flow thing and it was essentially building machine learning building a framework for machine learning applications on kubernetes and that really spoke to me but Google is very heavy on Python and so I'm sort of bringing Scala into that world but there's a lot of Python around me so it's kind of strange place to be but there are there are many opinions assuming the Scala is very much a valid way of deploying machine learning models it's not necessarily a first-class citizen on coop low but pr's are accepted if anyone really wants to contribute I will absolutely support you with that any volunteers is this the part where we start fighting on the panel yeah so so beat I just call him out before the panel started I think he said that you thought that you know in ml software best practices are no longer relevant do you wanna okay cute Cuba tweet we were happening in the back channel so I think okay we're going there all right so I think I've heard that said in different ways so there was a great talk to the chief scientist or lead scientist at uber Carpathia gave a great talk where he's talking about software engineering 2.0 and always I'm sorry Tesla I so so many self-driving cars start I was just hard to keep them straight so anyway so the the interesting thing there is that a lot of the things that he was talking about were things that I think you know Facebook and Google and LinkedIn hit many years ago and and struggled with but it was such a small pocket of folks doing AI and I should obviously Microsoft and Amazon very early as well applying AI at scale and something you know AI and machine learning used interchangeably here and software engineering things like unit testing and you know agile and stand-ups and things like that really start breaking down I think when you're working on machine learning in AI Jaxson some of the friction could be attributed to okay data scientists or you know machine learning researchers are hard to work with and they want to write papers and things like that but I think that's only part of it the the other part of it is that the fundamental problems that you're working on are structured very differently it's not like you're an architecture astronaut where you can plan out what things are going to look like twelve months ahead and know in detail how it's going to work and then be right if you're a really good software engineering manager with this stuff you often don't know even for problems that look almost identical to ones you saw before when you try it on a different data set or a different subtly different problem it doesn't work at all and that's for the problems that you think you understand and so that makes all these things very difficult the data is changing underneath you there's a lot of people talking about data pipelines and streaming real-time data real-time data is great but guess what like how are you going to unit test that you know when when when you're your data is fundamentally shifting this is how for example one of the reasons not the only reason why you know Facebook and folks can be caught unaware of these new adversarial tactics and their machine learning algorithms can't catch you know bad actors right because people are changing underneath you while you're building a model I feel like Chris is step on your bone and jump in and disagree or you know you both offer sort of services that help with this right so back to the question about this cow and Python okay come on guys it is 2019 already like 2018 we have we we tend to build layers we tend to build micro services with enter the diversify our our stack between different languages and paradigms so it's it's absolutely absolutely okay to be able to to use Java Scala for middle layers for any orchestration stuff and use Python for some user specific dsls and if data scientists love python nothing bad with this so let's follow this this trend and provide more python-based libraries for data scientists to use it so it's nothing you've seen like the Salesforce has open sourced transform transform transfer is not my choice yeah yeah so it I believe it will find its niche and so it will find its users it's it's good like it's great product so but it's probably my not sure maybe 1% of the you of the of the data scientists would change would change would change they're stacked towards Scala even if it's great back to the question about the best practices the first like unit tests if you can think about how you need tests are being implemented in machine learning way so machine learning writes the code for itself and the way to you write a unit test is machine learning testing and train training basically the training procedure is reinvented the way we will write tests and if if as a software engineers we will think about this we don't need to write explicit tests right now they we should use just data to write the tests for us in in in a new way in any use of software 2.0 stack the same way how we how we write the tests but the different tests on a classical software once you wrote your tests and if they pass you kind of 100% on 90% aware of that will work and work in production the same way you tested it you push that production work forever in a machine-learning way you should write you should run those tests continuously even introduction because the date your data is being changed and your tests are being changed as well so it's kind of and this this is another way to writing these tests it's kind of probabilistic asserts yeah you can't you can't write explicit asserts you can't write like typesafe asserts there it's like probabilistic way to write those tests a new way of thinking about this so just that just a thought so I'm gonna say that's why I somewhat disagree with you Pete that like traditional software engineering practices don't work on machine learning so I feel like you just need the right framework for doing that so I agree that the tools don't exist yet but I think I think the way you build them is to take the take the concepts from traditional software engineering practices and apply those to the domain of machine learning so it's not the solution isn't going to look the same but I think we can sort of bootstrap what those frameworks look like by by using some of the existing tools and then building building out where it doesn't exist and that's that is kind of a concept behind coop flow is is taking a lot of those taking trying to distill out the things that do apply to machine learning and then build something very specific that that filled in that gap is your Lu yes I have a couple questions for you Lucas what's the relationship between figure eight like CrowdFlower and weights and biases because I get three different answers with like three different people so you start so it was first CrowdFlower then it was renamed to figure eight okay and then you went off and started wait some biases yeah exactly right okay thanks yes we could talk about that and you should buy all three of these services yeah we ran across CrowdFlower yes we have a way because like pipeline is the run time for like all predictions we're able to actually detect on confident predictions and send them off and then right now we send them to a slack channel and then yeah we've been opening up the slack channel to CrowdFlower folks or like right like Mechanical Turk so I didn't quite not Mechanical Turk yeah I don't quite know where this conversation was going so I'm just gonna talk about what I want to talk about but we thanks you really helping me with this moderated job yeah yeah I mean I think everything's been said over there right so we're seeing a lot it's funny because a few years ago I was at Netflix and the head of data science was very skeptical of spark and very skeptical of the JVM and like his name was chavier mechs days I don't know he's at he I think he just went to a healthcare company but yeah he went to court after word so he's famous for saying that like you don't even really need spark just rewrite everything and see take the existing models that you have because that's all that's happening anyway right Syfy you know like nom by all these are just going down to see tensorflow is going down to see so they were able to reduce their cluster from like a 15 note spark cluster to like one single node which was huge cost savings you know just just one big gigantic note so that's kind of stuff we're seeing cool do you want to just picture your company I'm not gonna go there - for the sake of this panel and our audience which I still I'd love to hear audience questions I think in terms of our stack we use a lot of PI torch in research and then PI Gorge 1.0 just came out which is great so now it has also serving capabilities and so maybe hopefully PI torch will be able to eventually be the one unifying framework for both the researchers and the engineers at this point does the audience have any questions I have a whole bunch more but [Music] is about so it's a pretty famous blog post yeah his name is chavier xav ier and then something just search that Netflix and I think I think it was like an ALS implementation or something so it took the developer you know four days to really get to that performance but four days of like developer time and then you save fifteen no it's you know through the rest of the right like the training life with that model so you know these are very specific cases right but I guess in the context of which languages this isn't a C++ conference but like we don't really see much Scala honestly I think last year I did a a Python talk here and I got yelled at by Alexi so I tried to squeeze in a little bit of Scala in there but is it yeah who's the guy out of Brussels that what's his name I can't think of it now famous guy created the spark notebook and eventually couldn't really get enough traction couldn't get people off of Scott or Python in the data science world yeah I was gonna chime in on that so I didn't answer the Python scholar question I didn't know if he's a popular yeah so I I don't remember there was some somebody on Twitter said something controversial a few weeks ago is that you normally do but it was something about so how's that Python versus our war going you know kind of jokingly in that realistically Python one right so if I thought is won that war and I think if you take a poll of most machine learning folks actually you don't have to do so the creator of chaos has a great chart where he mind hacker news job postings for example and he looked at you know what are the trends what are the keywords and a lot of them were you know hi torch Karis a lot more pipe torch actually the intense level of sorry google for for a lot of the startups and academics okay support yes cute below supports both but so that was a trend the other trend was if you looked at all the technologies almost all of them they were all python-based and so for a machine learning and AI both in academia and in startups and increasingly in large companies they're using you know if I torch intents are flow in Karis question I don't know if you've ever used Pais bark but like trying to debug that extra layer is I think this also comes from two very different worlds so if you take a tool and you try to make it do something that wasn't originally designed to do it can be a big problem so these other tools were both ground up by you know Yamla and crew and folks who were have been doing this for twenty years doing essentially you know neural networks versus this is my opinion spark I think worked really well for a lot of analytics workflows and I had we tried to use spark at our startup at one point and used by spark and the community seemed to be overloaded and burdened with all you know all these issues and bugs and when it came to like deep learning or anything related to that they were not being resolved right and and and so for us it just felt like they were really good at resolving things having to do with streaming data and analytics but not so good on how would that add so spark is not about machine learning at all so they have spark ml but I believe that machine learning stuff has to be decoupled from the data processing some solid students should be as separated data processing independent library to be used for machine learning so Richard I'm really genuinely curious kitty could you walk us through how you decided to go with PI torch over the other frameworks I'm not a big believer in research to force people until a tool because in the end at least in the very beginning of modeling research it's much more about that years and how quickly you can iterate on them and so at the same time when you have something that really worked and then you're trying to get into production and everybody has like a different framework that is kind of annoying and so I was happy to have seen a natural tendency so we started with chainer see some crazy people who tried to do that and were slow at iterating but fast once they were done iterating and chain er back then and in theano actually back in a day even but long story short it all merged we had some experts in pi torch in the group and they were able to convince everybody and once you look at some nice clean pi torch code it won also the war against tensorflow not really war but you know discussions and so on because in the end I think tensorflow and most other libraries other than pi torch really pulled you back into the convex hull of known models so if you want to just take an existing you know recurrent neural network sequence model and just run it on some data and go really large-scale then yeah you can use a lot of different standard tools but in research once you want to develop some new model that maybe adds reinforcement learning and has a memory component in your NLP model and this deep complex you know network of some crazy pointer mechanisms and whatnot then you just can't do that in any other framework other than pi torch and and and still iterate quickly through that so it's for research it's quite obvious what the right tool was for the group and then of course the problem is now that you once you've figured it out like trying to make it work intensive flow is then another group and because nothing ever ends right data changes but new models also keep improving quite rapidly it becomes a really annoying cat-and-mouse game between two groups of like oh we're finally figured out like we're finally able to replicate your research model from PI torch in this other framework and then they're like oh yeah but that was last last month and we already improved it and here are five more tweaks and they're like oh now we have to do that to get the higher accuracy and so long story short it would be great to push PI torch and have more pie tour serving infrastructure and not to you inside baseball that Chris are you seeing the same trend or yeah I mean we tend to see quite a bit of tensor flow just you know because that's dictum the group that we've been rolling with my question about pi torch the group so for like pi torch for the inference is it still onyx and all that stuff is that what they're pushing or we're just so unreal and we would try to use it and it didn't support like recurrent neural nets like I was like okay that's the end of that you could only use it for a computer vision in the beginning and they're making progress but it's still slow and I don't think it gets enough support so far so when we serve I torch at all it's usually in native PI George and or we just replicate the whole model in tensorflow manually so now this is very important right so running it in PI torch right so just it's like behind like a flask app then right then calling yeah so like that's where things break down and then to get the performance you have to convert it into tensor flow and then use tensile sir yeah because since our flows serving it just has all this nice infrastructure for engineering teams and pi torch did not have it until 1.0 came out quite recently so now things might change yeah so I think it's really important to to be iterating on your models in the same language that you're serving them in and so like any sort of integration points any sort of touch points that's where things break down and that is that is one of the biggest sources between for the conflict between data science and engineering and like if you want to solve that you provide a framework that gives you the workflow so that research is happy and engineering is happy and everyone's kind of all on the same team and there's no like throwing things over the fence you want data scientists to have access to be able to iterate on their models and then push it directly to production like you don't want them to have to ask permission to deploy models and and from the other side of the fence like you don't want DevOps having things thrown on their desk and like having an extra task you really want to automate as much of that so you can empower people at every stage of the development process well Sam did you have a question sir I think you said when you have to make real-time decisions with performance bounce how do you deal with that yeah and so that's why we like tensorflow and it's more than just the software layer so we're actually working with these two hardware manufacturers that are targeting tensorflow specifically right and so these are like TPU like chips basically right they're specific for deep learning and they've been basically delaying their hardware launch to make sure that they're compliance with the I think it's the excel a like layer right that's below tensorflow that's sort of the intermediate representation between the sea layer and the hardware the you know kind of last mile so once you start seeing these like hardware manufacturers paying attention to specific frameworks this is kind of a sign that you know this is gonna stick around so I think that's a really interesting question in that what often happens so you do have this this friction in this interface between research or prototype and then what's deployed in production actually the genesis back at LinkedIn of the data scientist title so Facebook and LinkedIn around the same time started calling people data scientists the original intent is not what it is now which means a lot of different things it's like computer scientist doing a lot of different things the original intent was actually just that what you said Michelle that we should be using the same language using the same code base machine learning folks and in research scientists should be pushing to production and so that was the way that we ran things over time now you have to get a scientist using Excel and you know whatever else but that was the original intent and we wanted it to be distinct from just research scientists and the ideas research scientists were more focused on writing papers and didn't have to worry about maintaining things in production in terms of scaling and like isn't that a concern if you're using you know some of these other tools I think we it's definitely an issue I think in any custom machine learning model you hit these edges pretty quickly so you said okay the persistence doesn't support wear you know recurrent networks you'll hit something like that another issue that we hit embeddings right so there's amazing things that you can do with like word embeddings and entity embeddings and salesforce has done some some really amazing stuff but then if you need to serve an answer in real time in production or on mobile right on edge computing or on mobile devices and I know you've look as you've done a bunch of things thinking about mobile and small devices then you hit these problems of memory constraints right and how do you squeeze you know terabytes of imbedding data and have it available in memory for these machine learning models and so one thing I'd call out there that's kind of interesting is a startup called Roxette so from X Facebook folks we use some similar technology and our startup for doing conversational AI and we squeezed embeddings into essentially rocks DB based system and that way you could you know store large amounts of data and access it in real time I just want to do a quick check because it's got super technical I'm trying to read the audience like how many folks actually work on like projects involving machine learning and they're their day jobs okay cool big fraction is awesome did you have a question I mean it's just obviously great and I think the the question is how do we feel about pi torch and gentle flow becoming you know coming closer and closer to one another and I torch having more serving capabilities tencel having you get execution and so on and I think it's just obviously great and if I was to work in programming languages I would probably think about the next generation of deep learning and the 2.0 kind of stack because there's still a lot of interesting research that can be done in creating programming frameworks like pi torch and tensorflow that make it easier and easier and lower the bar to for instance have connections to graphical models and have hidden variables and think about sampling and integrating RL and and having enough flexibility still like pi torch does to even come up with completely things like new things like you should also realize too that spark is becoming more like tensor flow and tensor flow is trying to become more like spark right so tons flow has the data set api which basically at that point you pretty much don't even need spark anymore right which I'm a big fan of trying to get rid of just like simplifying the stack right so it but yeah yeah so there's actually three frameworks that are coming together and actually it's a bit more than just frameworks it's politics behind every framework the company who drives that frameworks I'm not sure they they can converge into the single one Shadia thought of this I mean I think it's great that so the closer they come together the easier it is to support them both like are the last kilo release was to have API parity between tensorflow and PI torch and like as a platform developer like that makes a lot of sense that makes things easier it makes things easier for developing on top of a platform but ultimately I think so the concept behind building a platform that will support any framework you throw at it like that's most important because two fluent pie charts are really popular now but who knows what's coming tomorrow it may be the the platform of the future that combines the best of both worlds who knows but it's about building resilient applications that are like loosely coupled micro services that you can easily slot in something else so I think it's good that they're coming together but it's not something that I'm gonna Bank on so let's see how are we doing is this tutu inside baseball too technical or you like it I awesome more questions so I'm one of the things that we were talking about kind of you know before we we started this panel that I thought was pretty pretty interesting was you know kind of what tools like if you're you know like an average software company today and I'm sure you guys have thoughts here like what what tools should you be like buying or looking an open source for and like what should you be kind of building in-house like what part of your like should you even even be trying to train your own models if that's not the core thing that your company does it's a super interesting question I actually think when we have a lot of companies that reach out to us and they want to do something with AI and then I asked what's where's your data and the first thing there's they say is like well it's in a bunch of different places so that's one issue and so the data pipelines is you know and has been mentioned here before is a big thing but and once you have it all pipeline you're still often you to label it and I know you've worked on this a lot and I you've done a lot of that you've made a lot of progress in this and I see a lot of other startups working in that space too and you have you know big players like Amazon Amazon Mechanical Turk but it's still not really figured out how to really make data labeling as like part of a simple pipeline where it's low touch like you're just like here's the description of what I want just give me like ten thousand really cleanly example like cleanly labeled examples of that and like it'll just work like you still have to do like implement a bunch of checks and so on so making that and then there's so many special cases right maybe you have 3d lidar data and you want to connect that to 2d image data and then you have time series or your videos you have language and different languages you know different natural languages and so so there's still so much complexity that nobody has figured out that piece really well and I think as startups there's still space to make progress but then it's also a slow sort of slog to get enterprise customers to adopt your pipeline and so on and then the question is should large companies have that themselves like Google's spends a lot of money for really important search kinds of results on manual labeling themselves and so as a big company should you rely on an outside tool or company or should you try to build it yourself when it's a recurring thing that keeps happening and you keep getting more and more data that needs the label like new products and images of new products and you'll never be done with you know categorizing all products because there will always be new ones coming out and so when you think that there is a recurring kind of theta labeling question like how much can you rely on outside vendors and you know the cost and all that so that's one of the many thoughts I have another question yeah I think one decision that really makes or breaks a team is if you use air flow or some kind of like orchestration I guess that's orchestration or you know workflow kind of thing people a couple years ago I think airflow was the obvious choice I know coop flow uses Argo I don't know a single person that actually uses our go outside of when they're trying coop flow but like we can talk about that on fly but you know okay but I think that's an important one right because if you don't have clear visibility and and and if you're trying to write these things I actually have a like call this morning with a customer in Germany and they're using Izzy and like it's a question I didn't even ask like when we started working with them and then I was like I'll how are we can integrate with using it you know like does this thing even like work anymore and it's all XML and it's just horrible but yes I think like that's an important choice yeah about like tooling ml doing it's a completely wild space and I haven't even it's a melon the shelfs all you need you need to give a shell to a male engineer to be to be ready for to like iterate or develop and push things to production to label the label data quickly and reliably and all have all that like orchestration workflow managed all together so the tool the tooling spaces of just one new crazy crazy area to be to be to be like investing into the question is either you wanna build it build it internally as an engineers of course we will all have our own engineering a go to be to build something interesting yeah because building a systems system things are much more interesting than for engineers at least rather than implementing a business logic and and talking and to talking to business maybe so I'm and so tooling all kind of tooling in for my experience it's even hard to see the companies that go beyond having an inventory of your Jupiter notebooks with the pre-installed dependencies that that will satisfy all the all the needs of your data science teams even having this like a very simple stuff you do you have a docker containers with a superior notebooks that and you have to kind of inventory that you can run a droopier notebooks with tons of flow dependencies that your to do your research there you have a notebook or point Zeppelin notebook with the spark dependences and all that stuff is it's very basic thing but like it's missing in most of the companies I see so far my advice is always if there is a hosted managed version of something use it if you build your architecture in such a way that you can if it's these loosely coupled micro services you can change those components out later but the fastest way to build something useful is to use the managed version and the way you sort isolate from the lock-in that that can bring is to use managed versions of open-source products that way you can also swap out the underlying infrastructure I think like I've seen I've seen that beef successful for so many people and that's why I'm so involved in open-source because I think that that really benefits everyone vendors and engineers like yeah I think this is uh this is a really politically loaded question in many companies because everybody wants to work on cool problems and if you're lower in the stock you want to build your own infrastructure if you're an AI person you want to you know train the model to do the cool you know image recognition thing even if Facebook and Google have you know hunt orders of magnitude more data and can build a better model for sure than you can and so I think this is where it's on the leaders in the organization and and folks probably in this room to step back and make some of these tough calls right so the one way to think about it is what what is your [Music] competitive advantage there's a great talk actually by the CEO ServiceNow talking about this back at eBay years ago and then how he's thinking about it at ServiceNow and he was saying you know for the things we were building there were plenty of things it wasn't gonna be our competitive advantage that was not what people went to eBay for so you should focus on those things maybe it's like Richard was saying the specific data elements that you have that Google doesn't and you know double down on your models for that but it could also be you know like infrastructure maybe you should use Google hosted stuff use the Amazon right used to be nobody went broke broke buying IBM now it's nobody went broke buying GCP or AWS right so I think you have to make a strong case to do something different right now for many of those and then the other interesting area which is just emerging now is we definitely have these problems in the development flow and life cycle for building AI and m/l products and I think what's kind of emerge are these narrow companies like weights and biases like I just saw an interesting company that's rethinking testing for machine learning doing these probabilistic tasks on logs there's going to be this suite of startups that you can plug in so think of it like a new lamp stack right but this time I don't think it'll be open source it may be partially open source it'll be you know like a figure eight weight weights and biases and you know you take four or five of these things and plug them together and that'll get you 80% of the way there and guess what it's running on GCP or AWS yeah we're gonna get and integrate it with with weights and biases and figure eight and then PI torch for everything else but that's a really good question yeah and I think Richard and I probably have hit this I'm sorry the question was I hope I paraphrase this right is there like an inherent difference between sort of the b2b machine learning applications like a bank might do versus sort of like b2c applications like Facebook might do boy weird to start yeah there are a lot of things so it's happy hour starting yet yeah I think explain ability if you're in banking and you make some pretty fundamental decisions about people's lives should they get a loan to start a business you know should you pay for their you know insurance or their card damage and you know it could bankrupt them if you don't and so I think the bar is kind of higher to make some of those decisions correctly and often machine learning falls on sort of precision and recall trade-off curve you might not want to make all your decisions automatically have lower recall but you may want to the decisions that you do make and need to be really really precise whereas like in Google for instance it's okay if like not the first result is the right result but you know as long as you have the high recall precision can be a little bit lower if you don't click on every single ad which of course nobody does like that's okay but you know if as long as you click on some of the ads some of the time so precision versus recall is I think one way you can put this trade open of course an enterprise you know enterprise is a big space so there are lots of different things in it and and they fall into different parts of that curve and then I think explain ability is crucial when you want to change human behavior you can't just ask like salespeople who have done their job for like 20 years to like go and call this person instead of sending an email to that other organisation and they're like I've been doing my job for 20 years why is to say I telling me to something so for those areas where you want to change human behavior you need to give explain ability to a - and then of course the big an obvious one is you don't actually have the size that Google has for most enterprise applications most companies don't need to have like a thousands GPU machines each with like eight you know p1 hundreds in them to train on all of the you know images that come up on the internet right they have maybe 2000 different products and they have two pictures of each product and now you have this odd combination of low load high D low end kind of thing we actually have few training samples for each each category and now you still want to do something with that kind of data set and so I could go on forever but those are sort of the main main things you know the trust that comes with enterprise data is a pretty crucial piece you you can't really like you know Facebook lost millions of people's data if you do that as an enterprise you're dead Mike you just can't afford that so you know I've been kind of selling machine learning services to two companies for a long time and you know one of the big changes that I've seen I think I hadn't quite exactly thought about this until you asked the question but you know when we started like you know 10 15 years ago almost all the applications were b2c you know was sort of like a dranking and like search and it was like all around things that that sort of increased revenue for for b2c companies basically was where you saw a machine learning applications and I think I think b2b is taking longer for like some of the things that you're saying right the the precision recall trade-off is there Frank right it's like it's it's maybe more dangerous to make a mistake and maybe because in a b2b context you're often sort of saving money versus like making more money and I think companies tend to be less excited about that but that said I think that one of the big trends I've seen in the last six to three years that I think has been part of this explosion and in machine learning has been I think these is every b2b company is thinking like you know how do I apply especially Salesforce right thinking you know how do I apply machine learning to all the back office and all the business process stuff that that I do and I think there was something you could do in b2b that I think is kind of underused a kind of design pattern that you know we call human in the loop where you don't have to solve every single case right so like if you're serving up search results you can't just be like okay this one's too hard the machine learn is confused we're not gonna give you a result right but you know if you're trying to say you know you know should I like kind of like flag these sales records as like you know answered it correctly you can say you know what the algorithms not sure about these ones let's send these to a human operator to deal with and the rest the algorithm takes yeah so keep in mind too that like you guys have choices where you work and I you know so growing up in Chicago I do yes I was in like the whole FinTech like the insurance just hell right and then I come to Netflix and I'm like oh wait I actually had a choice this whole time and it's way more fun to work for a company that's not worrying so so much about right like the explain ability and each transaction going through and every penny being accounted for like that's just boring yeah well so I think there's something else that circles back to the beginning of the discussion we were talking about this you know intersection of data engineering and AI and there's a great Monica regatta who's somebody a few of us know worked with her back at LinkedIn data science she has this great data science hierarchy of needs of all the things you need before at the pinnacle of self-actualization you hit AI and a lot of it is is like plumbing right and so when I talk to a lot of enterprise companies that say hey we want to do AI like the people who are doing AI early like Google and other folks they were trying to solve certain problems and they said oh this is the right way to do it now what you have is a situation where people saw this great success and all the gains that they had at Facebook at Google etc and they say we want some of that right and then when you go in and talk to them a lot of these companies had lacked the infrastructure and if you think about it it makes sense so if your Google and you're optimizing page load time and they squeeze out every fraction for ads and everything and it matters everything is tracked in excruciating detail and that's kind of what you need for AI to work right you need a sensory system that's recording all the data so the robot arm can move around when you go into a lot of these environments that are more b2b they never needed to do that right so now I think the race is for people in this room you need to catch all these companies up and give them that nervous system of data that they need with tracking and search you need to track impressions and clicks and all these things and then you can apply the cool you know machine learning and deep learning yeah I would add so we are on scale by the way conference with scale means different things for AI for some people scale it's scale in terms of high load and low latency like Google for some companies like Salesforce scale is among a huge multi-tenancy scale across all the tenants and managing managing the AI applications across a highly compliant tenants for others like scale is a scale of training process for a scale of managing the multiple like hundreds of models in production and monitoring those models in production for for folks like coop flow team it's probably a scale of managing the workflow across and enterprise across companies how would you and people who both involved into this AI process and AI development process of the applications so it's a yep ai is very very different in terms of the big customers of this and maybe to point out the opposite of that which is yes that's true but it's also the opposite is true that a lot of companies are also you know they have their own end consumers and so now for us we don't just think as b2b as a one business to another business but we're thinking about the whole B to B to C so we also offer marketing solutions and we have to go through all twitter-like we have access to the full Twitter firehose and we need to deal with that every second - so it's like that's fully in the same world as in others we have IOT signals where we get robot arms to send us like millions and millions IOT signals every you know second or a minute and so we need to deal with those and find anomalies and in them and and all of that so I think that the worlds are also sometimes merging between the two but Enterprise certainly has some it's own unique challenges this is super cool oh one quick thing I think Pete said something really important about the the hierarchy that in order to do AI you need good data but don't mistake having good data with having good data infrastructure like there's a there's a large foundation that you need before it makes sense for you to do AI if anyone was in David's talk David from simple logic I think it was yesterday I I steal his slide in like every talk I give because he basically makes the point that if you can do it in a deterministic way you should do it like do not use ml if you don't have good data and good data in front well let's continue this conversation over drinks yeah thank you [Applause]