Devreal

Graph Exchange: Alexy Khrabrov, OAKS ASKG: Agent-Server Knowledge Graph

Graph Exchange: Alexy Khrabrov, OAKS ASKG: Agent-Server Knowledge Graph

Recording: Graph Exchange: Alexy Khrabrov, OAKS ASKG: Agent-Server Knowledge Graph

I wear many hats. So, community all across. So, let's I'll just kind of, you know, do a few highlights. So, who knows what this is? This is a handbag made out of an H100 replica. It sells for a,000 bucks on eBay, right? That's a first small piece. It's much cheaper than the original. Uh but it just shows the tight guys, right? Like everybody is excited about AI, but it's hard to place where are you in AI? Like who knows everything about the AI landscape? Who knows where exactly you are, where you want to go, where you want to be. There are people like this

Please, please talk to me. I don't think there is anybody who knows even like where are you on the map, right? Like in any map on the city. So you are here then you know where you want to go. If you don't know where we are, how do we know we want to go? This is a problem. So, we propose a framework and we're going to talk about it here. We propose to talk about knowledge and AI, right? AI is only good and useful for business, not research if it solves something in the real world. And every real business knows something about the real world which is distinct. So, for instance, Coca-Cola is not a soda company

Coca-Cola is a distribution company, right? Coca-Cola knows logistics around the world. that knows how to make soda around the world, package it and ship it and retail it. And it happened so because it was a supplier to the US Army which was stationed around the world during the cold war. So thanks to the USSR, we now have Coca-Cola everywhere, right? And so this kind of knowledge is hard to capture. It's distinct business knowledge which a company has, your customer service has, you know, your crypto company has, your bank, right? So we need to capture this knowledge in the applications. And so currently we're talking about you know general purpose models. They're not capturing your business knowledge. They're capturing world knowledge

They're not capturing knowledge unique to your specific business. So we propose to look at the flow of knowledge through a fullstack open source application. And the way we do this we look at input transformation and output. And so input is where you know structure is extracted. You you shouldn't think of structure as syntax. You should think of structure as semantics. you extract a representation which has meaning. So an example is a recent project called Dockling

Who heard about Dockling? It's a number one PDF parser. It's at some point was number one project on GitHub. It's from my uh former colleagues at IBM. So it extracts semantic structure of a PDF uh in markdown for so basically you know sections become it's becomes a tree. So you don't think of this as textual elements. You think of this as a tree of meaning, right? And you should really reason about all kinds of data this way like the structure is actually meaning right subsumption containment means this is a detail right a caption or kind of what kind of table is this where is this table attached to name entities within sections you know pronoun resolution all of this you you should think of this as kind of knowledge which you extract and you never lose it you you keep enriching it and you never first of all lose the knowledge you never unbundle it and no you don't validate invalidate Right? It keeps being valid. So whatever record it is becomes in invariant and you pass it through and you transform it only in a way which does not lose knowledge. Right? So so I'm I'm a huge fan of pyantic

I'm very fortunate kind of think we're very fortunate that Python community picked pyantic for various often random reasons. But pyantic is a basically a type system right? So if you come from type languages like Java or C++ or Huskll you know this is extremely valuable right? If you use pidentic records then the knowledge will be bundled and it will keep going around and you will validate that it's still the same type. So it's very important that you transform things and this records can be basically stored in various ways right and we can put you know embeddings in them and we can store them in knowledge graphs and they can be uh building blocks of the memory and finally when the times come to do something in the world your your agents should take action do something provide some output actuate a robot right you do basically structured output which emits a sematically correct action which is you know for instance a receipt right you want to give a receipt which you learn infer the receipt should have sections they should have headings and it should also have a feature that all the numbers add up together right so this is both syntactically and semantically cohesive invariant and so you obey this invariant and if you do this knowledge flow through the application as I described then you will guarantee that your final output will be sematically correct so this is kind of the way we propose to think of knowledge and so is a social collaborator. We basically are very fortunate at Neo Forj that we integrate with dozens of open source partners. We are graph database and we do not sell open source products like we do not sell genai products. We provide this graph database right which is you know for almost 20 years has been the most reliable way to store knowledge graphs. So what happens is all our uh geni work is in open source and most of this is basically pair wise integrations. So we work with lench chain and llama index and wev8 and milvos and bine cone like it can be open source it can be clos source we have tons and tons of open source collaborations and and so like we're very fortunate that we're kind of in the center right of this knowledge ecosystem and we know a lot about people who store knowledge and what does it mean so um so we kind of have this building blocks which are both our partners and their projects we have production grade packages you know with lench chain with lamine is maintained by both companies update it regularly

An example is like when pyentic version is bumped then of course we have to do this because we use it and they use it. So um and uh this is basically the uh the flow and I added AI memory here as a distinct feature which is emerging now I think especially because agents need to store what they know the state of the world somewhere and so and this storage has to happen at many levels. It can be your operational storage like what does the agent know at this very moment? What do you know about your days of work? What do you know about your goal? What do you about know about your team? What do you know about your company's memory? What does the company know? Right? Because some people can know something about customers. What do you know about your you know human knowledge in general and in institutional knowledge your and furthermore like global knowledge right? So memory can exist at multiple levels and eventually I bet that agents will have to consult all levels of memory and you will have to be able to distinguish this and in in nearj you can give properties to nodes and you can have graph overlays and you can give features to relationships. Uh so so this is a project which we're specifically building right now in the design stage. It's called agent and service knowledge graph. So we want to take all them superours in the world and we want to put them in a big knowledge graph and uh we're going to use a standard for the data representation called Croson. Who heard about Croson? Uh Croson is a new format from Google for data sets in machine learning

It's basically a JSON schema for a file files in a data set. And there is an extension to this called semantic ran which is sounds weird maybe not tasty but uh that is basically an addition of ontologies to this data sets and so it's a global working group uh led by you know folks from Europe EU and and US and so we basically want to develop the anttology standard which is distributed so if you want an ontology in automotive industry you can solve the root kind of DNS of anttologies and you find out who maintains it maybe BMW and Tesla maybe some consortium maintains it and then you can say like what is the name of this spark plug and the spark plug should be have unique name which is identical for Borch repair center for you know Shenzhen manufacturing facility and for BMW factory which assembles the cars like this is very important that the terms you're using by all the agents in all the systems are resolving to actual unique objects in the world right like you will not do business unless you agree how you do this so so this is kind of what we want to do we want to first like align the anttologies and also figure out what MCP servers are doing in what domains they exist and basically what ontologies they're they're operating on. uh so we're designing this let me know if you are interested in participating this is you know sematic coran high level slides they exist on hugging face there are data set so it's a pretty advanced thing which you know came out of Google and friends so some more slides of this and so this is basically you know what we want to do here we want to have this community where folks do open source activities together they kind of figure out what to do with this MCP server ecosystem and agree on the common language And if we do this uh we're basically going to have network effects, right? We're all going to succeed by making all these pieces interoperable. So this is why we're here. We want to share this knowledge and crosspollinate it. And I have to mention in November we are running again longest running independent uh AI conference in the world called AI by the way which is the first AI by the way in deep learning in 2017. cars. Um, this is an awesome conference

Uh, early bird is still up, so check this out. Thousand people, you know, like this time 10, right? Uh, and this is like, uh, I don't think we have time for Q&A, so just make a screenshot if you want to follow up with me. Yeah.