Devreal

Bay.Area.AI: Interview with Jiang Chen, Zilliz

Bay.Area.AI: Interview with Jiang Chen, Zilliz

Recording: Bay.Area.AI: Interview with Jiang Chen, Zilliz

hello everybody my name is Alexi kro I'm the founder and organizer of B area AI which is the most established AI up in the Bay Area in the world running continuously for 10 years in Vera in the headquarters of the best technical companies in the world today we're at GitHub you see all the developer activity happening right behind us it's the live dashboard and uh tonight we have a meop uh which covers several key topics in AI uh programming uh it's DP PR from Stanford rigorous way to program LMS and we have serious about rag so we have uh rag Vector rag we have graph rag we have several speakers and uh uh today I have with me JN Chen who is the head of system and development relations at zillis and the database called Milos welcome John well thanks for having me yes thank you thanks a lot so you're one of the leading open source uh rare company so tell us a little bit you know uh it's fairly recent so how did you this you know project started how did you become the head of devil uh and how basically you see all this uh new area unfolding well yeah yeah sure um so everything started from um 2019 2020 when we started to build the first open source Vector database in the world which is called Milas um and the company behind this is zillis so I know we have two names so it created some confusion for the users um but yeah we are uh um we donated the mil product to Linux Foundation data and AI um but till today we're still the uh the main uh contributor and maintainer of the project M and it has been very popular among the developers it now has 29k stars um so yeah really really awesome numbers um we love to contribute and to create the value to the community so at the very early days of M us um you know probably most of the users of vector search were in the uh recommender system and you know machine learning um domain but however over the time as the you know AI is really democratizing the whole stack of um AI power search uh we have seen more and more users um like full like web developers mobile developers who didn't really have a lot of experience with machine learning start to use Vector search and start to embrace the power of semantic search uh into their uh technical uh into their tax stack so this is um really a a great time to you know um democratize the the very complex AI stack which was used which was um only available to large teams like with hundreds of people in the instructor team um now like with the uh EAS of use of iding models and Vector database um everyone can build a search everyone can embrace you know rag or image search into their applications MH now this is great so like you got an early start you were lucky when CH hit right suddenly you know this kind of whole opportunity happened and this is interesting because uh I am now ai commune architect at NE which is an established Dr database which created the segment of graph databases right it exists for decades and so it also is now the key vehicle to De hallucinate LMS right and so um it's very interesting so so when you guys started like what was the use case right it was it was not the ls right it was more traditional yeah we talk about history uh I mean I I I'll dive a bit deeper into that so actually the history about myself is that I had years of years of experience in Search and indexing in the traditional web search image and video search so in the old days like 5 years ago yeah not that long ago um there was two vehicles driving the innovation of search technology one is embedding AKA different networks and all the data generated from them and the other one is kg Knowledge Graph yes so I I really think that those are two um important Technologies Knowledge Graph the way I see is that it has a structured and systematic way to mining to to mine the dat the information out of the you know unstructured data and Vector embedding is a again um efficient systematic way to you know encode to extract the um you know representation to extract the the essence of the semantic out of the unstructured data so both of the Technologies are very important in terms of understanding the the like joint amount of unstructured data in the world because you know from the um even from the Y age um most of the information on the world in the world was like web pages and there wasn't really that much structure data yes yeah um so that in order to tackle this problem we really need some rather than you know having human labor to to you know label the data to do annotations we need a more scalable a more systematical way and de neural network and knowledge knowledge engineering are very important techniques to tackle that problem and that's kind of leads to why um I believe um both the vector database and uh graph database are popular and continue will continue to be pro pro Prosper yes you know it's interesting I when I basically started to follow the commun events you were guys were one of the first who started popularizing the r right like I think one of the first events I've seen where your Mups and what was super interesting to me majority of the people were not the people I've seen in developer commun before around developer m for 10 years and more right so I started the very first Spark M up with mat you know I run the biggest scull M up we have the you know apik kovka like all these distributed systems and like you kind of know what traditional developers do and now have this completely new generational developers some of them most of them were not developers before right they plunge into this like AI Engineers so like I wonder how do you go about like teaching uh right as Le like this new generation of people coming in right how do you kind of help them become a Engineers yes with with with ra you know talking about the community and um making friends with developers and creating value I think like empathy is the the first priority you need to know like what they are thinking about what's their um you know um the the challenges that they are facing um so that despite that um most of of us uh came from the database and infrastructure perspective we do know that you know as a fullstack mobile uh developer you probably don't know that much context in machine learning and data instructors and however in order to build a really awesome application you have to master that um so how how do you do that I think we are trying to provide two values um in in terms of this aspect one is that we're um kind of articulating the machine learning and Def NE Network Technologies from the perspective that uh like any like any U uh developer without deep experience in marchine learning can comprehend like you know some of the time you just need to tell a story that really makes sense like from their perspective like they are really familiar with say microservices right and then we talk about you know the machine learning stack not from the models but from a microservice service aceration perspective and I think that's also one of the reason why say longchain and Lama index are so popular among the community because they are really abstracting away those integrities in machine learning models and data stack they are really just um extracting the important pieces which are say uh in long chain you just need to have a few components and then you um streamline them into a chain L index do you know you provide abstraction out of the um the complex retrieval stack Maybe it will um you know touch quite a few pieces but with this simple and elegant abstraction you you know just abstract away all of the complexities and that's the similar thing uh we are doing here like we try to you know we choose to um partner with um all of the awesome R acction Frameworks and evaluation Frameworks out in their in the market in the open source Community um by integrating MERS and abstracting away um you know like vector uh distance Matrix and things like that uh extract away those complexities and only leave the um most frequently used features in those um obstructive API those interface however if you really have a complex um problem to T tackle and you need to like find through or like um um you know customize those knobs we still provide a um a way to do that like through a bench of um complex knobs for power users yeah for power user for the advanced use cases so that's one thing the other one is just like what we're doing today and also what have been doing a lot in the GitHub awesome GitHub menu U we um you know host the Meetup events we uh invite the speakers to talk about the new trends in this domain so that we um provide you know useful information and like um best practices to the developers so that they don't like 10 years of experience in data pipelines M they can still build a an awesome uh like production ready data pipeline by leveraging The Experience from all of the you know experts in this domain yes you know I really love it that like you really follow the spirit of Open Source right then you not only do your own right and then you invite others so I started joing this he days and what I see is amazing because multiple companies in the space they kind of bring each other and they together show the stack to developers right because Rising tide lifts all the boats if we help everybody do their piece in the stack better and we use open source you know we do it better so I've started this new commission called Dev real Dev real. it's basically Dev but keeping it real right and so so we kind of you know want to invite kind of the best Deval advocates in in it and my question to you is kind of one of the leading the real folks right doing this for a while how do you think we should kind of maximize impact what what works best for developers because it's really hard to like you said to find what should you do like you can just be overwhelmed with all this new information right and like there's so many possibilities you have five of these companies which should I pick how should I put them together like how do you help developers to think about this how can they start and how do we keep them happy by learning more and more and become more comfortable with this stuff yeah so um I think first of all I I I wish that we can you know collaborate even more so that we create those um integrated those well Illustrated dios and best practices through you know notebooks and and blogs and and even better video content so that uh we show developers like how to use those Technologies in action um by you know combining a suite of Technologies like NE 4J and M you know dat all those great data stack and in addition um I think we also need to build a lot of um uh uh we call it seamless integration and I know the word seamless is kind of overused already but really has to be because you know um when developers are using it you won't think of it as you know product a versus product B we think of as a a stack as a you know as a um a solution so that by you know uh doing Integrations between the um adjacent stack like upstream and downstream for example example um um embeding models and Vector database plus large language model which is kind of the three important pillars in R um by doing this kind of Integrations we can provide more like um easier way for developer to onboard to this new technology and as well as you know kind of plot plot through all of the challenges and you know obstacles that developer May came across um may come across when they are doing this integration and the last thing we want them to do is to feel frustrated when they are combining technology a b and c right yeah um and also you know make friends and and um you know host the community uh events so that we can talk we can communicate and collaborate fantastic this is what we're doing here and last question is you know maybe you can tell us something fun about yourself a fun fact or what do you like to do for fun okay um I guess yeah fun fact about uh my my role actually I I have never than DAV before all right this is uh this is probably probably my like third fourth month doing DAV I I really came from the engineering and product background so um back in the days uh well when I was at Google I was um you know building short video search um by you know doing semantic understanding of the short videos out there like from Tik Tok from YouTube shorts Instagram um so like at that time we didn't really have the sense of D real because we only have C customers you know partner teams like in the large organization so that you know it's really really different uh working in this uh open source and um you know this very open Community where you have unlimited amount of resources but at at the same time you also um you know con constantly face the challenge of connecting to to your users and developer friends so that yeah I I guess back to the you know the the the central topic of today is still like I think it's never um you know we can't really emphasize the importance of community um um too much it's really really important to foster a community and Foster the um the channel of communication and the free flow of information so that you know we can better build um Technologies but also products no I feel really happy that you know you said that because you know this is the first meet up in the 10 years year span we run it which I run as AI committee architect which is in the org so I've been doing it for many years but this is the third week you said a few months this third week I do it officially in the de or and right and so I really want to learn from Dev like you and all of us to teach developers how to tackle this a think so I think with folks like you and Roy and others like we're going to do it I think we're going to bring a eye to the people with open source so thank you Chun and looking forward to your talk thank you very much thank you all right looking forward all