data.bythebay.io: Jay Kreps interview
Recording: data.bythebay.io: Jay Kreps interview
my name is Jake reps and I'm a founder of a company called confluent I thought speaking with a data by BAE was a pretty great opportunity to talk in this case to a bunch of people interested in data and data pipelines and particularly in the area that I've been working so it was a fun opportunity to get to talk to a bunch of people who are kind of deep in the thick of it I think you know that what's happening in the world is just maybe data in software I think it's getting a lot more central and what companies do in people's lives I think that makes it possible for you know programs to deal with data to do much more interesting things than they used to so whereas maybe data processing has always been used for kind of back-end optimization of you know big businesses and government applications now it's kind of really finding its way in a much more central manner into people's lives I think that's what's so exciting about it and then I think beyond that there's been big changes in technology just to keep up with the volume of data that makes a lot of things that would been harder to do before impractical or not cost-effective suddenly totally possible and so I think that that combination of how the world is changing and how the technology is changing is kind of happened together and it's left everybody pretty excited about the area and hence all the conferences you so my talk was about apache kafka and about processing streams that data thinking of data as a stream of something that's continually being updated or changing instead of kind of a static table or dump of data and that that I think is a paradigm that's just now coming to the forefront in how people deal with data how they process data how big you know companies design their software statics and so my top covered Apache Kafka which is a system I've worked on for a long time and to the new things in the kind of Kafka ecosystem framework called connect and framework called Kafka streams and so I kind of went into both of those and what they do and how they relate to copy on how they relate to this kind of overall vision for a streaming data I guess my advice for becoming a data scientist is a you know learn your math and then work with data yeah you know I think the the best data scientist I work with I can combined you know a strong theoretical background with a lot of like practical hands-on application I think like the problem you know actually addressing real problems and benchmarking yourself against those I think tends to develop the kind of skill sets that people need to be good at this stuff I think that's what kind of takes people out of a more theoretical mindset and to you know something which is actually quite hard which is to you know compete scores or test beds against real problems over and over and over again I think that's probably the the process I would fall unfortunately it's a easy time to get a job in this area so I think just starting starting low on the totem pole and working away up is I think a great with it but it's yes so the relationship between data science and data engineering is really interesting so I think in the short term maybe data engineering you know kind of building the pipelines and plumbing is the more important thing in a sense it's kind of like the difference between you know building a foundation and then decorating the house so at first building a foundation is really important once you have the foundation you really don't think about it that much but there's a lot of decoration that will continue happening over and over and over again I think the relationship is sort of similar in that right now the tools are pretty immature so there's a lot of them it's hard to use them it's hard to actually put them into production over time that kind of fades and it becomes you know much more kind of an off-the-shelf problem however i think the use of data is actually much harder to commoditize so as i think a lot of these pipeline technologies they reach a certain level of maturity they become kind of ubiquitous and then you just kind of can assume they're there and they become boring and the way that operating systems are boring meeting boring because I mostly work and when you kind of know someone they don't the problem of how to actually use data and apply it to the you know the challenge at hand for business that I think is actually much much harder to commoditize so it doesn't lend itself really to being turned into some piece of reusable infrastructure it's harder to turn it into an open source project people can download it actually requires people thinking about what the problem you're trying to solve is what data you actually have trying to model it and I think we're not that close to making that a kind of point and click process so I suspect that you know in the short term we kind of ignore the infrastructure side maybe too much but in the longer term you know the definition of success will be we can ignore it totally safely and you really don't think about that stuff you