Devreal

Slava Tykhonov and Alexy Khrabrov: Palefire, an advanced Context Graph AI Project

Slava Tykhonov and Alexy Khrabrov: Palefire, an advanced Context Graph AI Project

Recording: Slava Tykhonov and Alexy Khrabrov: Palefire, an advanced Context Graph AI Project

graphic exchange before? A few people. Who's been at the age of the self met up before? Maybe? Okay. So, for those of you who don't know me, I'm Alexy Krobov. I'm basically community organizer, um, meet up organizer, founder in the Bay Area for more than 15 years. I started Bay Area which is the oldest continuously running in person meet up in the world. Uh, running since 2014. Before, you know, deep learning and through all the waves of machine learning and data science and so forth. Uh, I also run the conference I founded called AI by the Bay which was scaled by the Bay, data by the Bay which is 13 year old

We just did one in November. Right? We did the first one in 2017 under AI by the Bay brand with Russell Norvig, the authors of the textbook. Who studied AI with Russell Norvig textbook? Do you guys remember the color it was? Red. See? Red was the first edition. So, you age yourself. I I I used the red. Now, apparently there's even like green or blue. So, depending on the edition, you know, people like ask people what color was the Russell Norvig, you will figure out roughly the age

So, uh, so this is, uh, a presentation of Pale Fire which is the project that, um, basically was incorporated into, um, AI stack. So, thank you Sumair for that. That's really really really smart. Right? Because it differentiated by placing it in, you know, the field where everybody can know what it's doing. Right? It's not some abstract telecom or like whatever AI in general. This is very good. So, so Slava Tikonov is my good friend, uh, and, uh, member of the knowledge group with which Adam leads. And Slava is an amazingly prolific individual

So, every month, every week he builds something new. He was doing it for decades. So, until recently he was the lead developer at Dutch Academy of Royal Dutch Academy of Arts and Sciences. And so, he built something called Dataverse. Uh, I don't know if who heard of Dataverse. Dataverse is the biggest public repository of data used for science. Uh, it started at Harvard. It's dataverse.org and then it spread basically all through the scientific community

EU is funding it. There are 800,000 data sets hosted at Dataverse and Slava built infrastructure and he built a lot of machinery which goes into that and a lot of different standards follow because scientists really need to interchange data, right? So, uh, so Pale Fire, uh, in case you didn't notice, our last names end in OV which means it's Russian or Slavic. And so Pale Fire is a Who Who knows what Pale Fire is outside of this talk? Pale Fire is a novel by Nabokov which is super weird and like I don't really like it. Uh, like Lolita is much better and you know, any any other novel almost, right? But apparently it's it's a masterpiece of abstraction. It has inside of it a nested novel, a nested poem with 999 lines which is, you know, the the novel talks about it as metadata. So So this is kind of very clever pawn to to, uh, see, we have we have a copy. So, yeah, so this is yeah, this is pretty interesting. I bought it, you know, but I didn't really I wasn't able to read it

It's not readable. In any case, the Pale Fire the computer program is much more fun. So, uh, so what we're doing here, right? Uh, so this is it's really, you know, another fantastic feature of Slava's work. It's just hard to nail. There is so much stuff going on. Knowledge has been organized, graphs and vector databases are being used, ontologies are being matched, right? Queries are being asked and LLMs are being basically used in a very specific way to disambiguate, uh, kind of drift to answer questions precisely using the graph structure. So, uh, if you don't get like all of this, don't despair. It will take you months or years to learn all of this

So, but you will have the link and you will basically kind of this is an entry point into this whole ecosystem, right? So, the the main thing is a kind of knowledge operating system. The if one thing you remember from this, it the whole goal is to basically take all these documents, build semantic knowledge graph, and do it very precisely. So, create, you know, ontologies, check them against a lot of existing knowledge, do retrieval properly, and uh and use a standard. So, Slava is also one of the founders of a standard called Croissant for a machine learning. So, Croissant is a metadata standard. So, that's another kind of, you know, pun with Pale Fire. Uh Croissant is used to describe data sets, which is important both for humans and now for agents, right? Because if you just have a data set with a bunch of columns, it's a CSV, right? It's a database. Uh you don't know like there are some truncated column names

You don't know what it really means, what even language it is in, and so on. coding. So, Croissant basically it's kind of obvious that surprisingly nobody did this before. Croissant adds a like big juicy JSON header to the whole thing which describes in excruciating detail what's in this data set. What is the language? What are the units of measure? What these numbers mean, right? This can be nested. So, so that's kind of if one thing you remember, there is this Croissant. And uh so, Slava was the, you know, member of the standards group which did Croissant 1.0, which is now MLCommons project. MLCommons is another nonprofit uh running the MLPerf uh benchmark for machine learning performance and some other, you know, community infrastructure

Um now, it's does narrative synthesis. So, it use uses LLMs to explain what different numbers in data means. And it's doing it using the semantic layer it extracted. So, it does it very precisely. And it's very well done. So, like there is a Docker version. So, really built something unusual. It's super easy to run it in a browser

I'll show you to you. So, this is Cross-on 1.0 for machine learning, right? These are the original authors. The first author is a lead in Google Paris. Slava now works for Codata. So, he's a head of AI and knowledge graphs at Codata, which is a European consortium funded by the EU to top data interchange for sciences around the world, right? Um and Omar is the lead at Google. So, this has always been an industry ready kind of industry centric standard. So, now you can basically get all of your data sets with Cross-on metadata from Google, Hugging Face, and others. It's a standard now

So, kind of you know, you have the editor that we you can edit your data sets and create these headers. It's a Hugging Face. Um you basically can, you know, crawl these repositories and you can use it in Colab. So, it's it's basically a production ready standard. And there is a lot of tools and this the ecosystem is growing which understand Cross-on and interchange. So, this is what it looks like. Uh Slava extended Cross-on in two ways now. So, he's leading a new working group for multilingual semantic Cross-on

Original Cross-on is basically flat. It doesn't try to build ontologies. It just describes what's in this set. Semantic Cross-on purpose to link you out if, let's say, it's a data set about automotive parts, about a BMW, it would be silly for you to maintain a BMW, you know, catalog of parts. It's better done by BMW. So, at the point of BMW takes ownership and publishes the authoritative ontology of BMW parts, you want to link out and say the full definitions are here at BMW. And why would you need to do this? because if you know BMW is being repaired by Bosch facility and like there is a self-driving BMW in a few years which will drive itself to Bosch, you know, center. It will order the part from China two days in advance which will be like flown in, right? So all of these three things, supplier, repair shop and BMW factory and machine they all should agree that this part is called exactly this, right? So you have to have ontology servers talking back to this ontology question, right? So so that that that will have to come

People will have to maintain authoritative ontologies for the business domains. So this is not the job of computer programmers. This is the job of business owners who own this ontologies of parts, ontologies of services, job names, roles and things like this, right? So so Croissant provides you a structure where you you can basically put this links in there, right? And here you can have like excruciating like detailed language annotation. Every word can have a language, you know, big data that we don't care. It's not much, right? But basically like you can call everything in the its original language as deep in the JSON as possible. So here's kind of a use case. So scientists would create a spreadsheet like this, right? And they need to share it. So they have some weird Windows device which emits some strange Excel, you know, lab like sequence and there are some variables in it

So what's another very interesting thing which happens here, scientists like variables and this is the thing which we don't hear about in Silicon Valley. And it really struck me because I work as a bunch of scientists. Scientists do not believe in vague, you know, digital transformation stuff. They they have numbers. And so also they like their experiments, their measurements, they come out of a bunch of numbers and now they have, you know, all these numbers need to be related, right? Like if you do experiments in chemistry, in in astronomy, like they end up with big tables and all these numbers without croissant annotation, they they it's not clear what they are. Once you annotated them, often experiments have related sets of numbers. So, if there is another thing to remember from this talk, is that we forgot that numbers are important, right? People who care about numbers are rich. Like here, they're not here, they're on Wall Street

Like people on Wall Street care a lot about numbers. They're very precise about numbers, right? Because numbers for them mean money. Numbers for scientists mean, you know, scientific results. And so, uh because this uh a tool was built in scientific environment, it's meant to be very careful, very precise around numbers. It can allow you to annotate and disambiguate numbers, and place numbers in context, which I found very refreshing. Right? Like we like we almost forgot it's all text now. Like it's actually not all text, right? Numbers are still underpinning a bunch of stuff. Uh like AI is just matrix multiplication, right? All the way down

So, it's all numbers in the end. So, we have a bunch of, you know, data. Uh now, what happens? We want to place it in this Harvard Dataverse Repository. So, the scientists, instead of little spreadsheet we're going to mail around, and anybody will receive it and ask, "What does this mean?" Like the scientist doesn't want to do this. They want to annotate it using croissant, place in the publicly available repository, which is can be at Harvard. Uh uh hopefully it keeps its funding, I don't know. Uh and basically it's going to stay up. But like this is replicated a bunch in the EU

So, if you find Harvard Dataverse is down, go check its mirror in the EU. Uh right. So, so basically what it does, right? Uh so, you you can basically create, uh you know, uh data. So, so now this is an example that you can download data in croissant. So, right? So, you you can get it, and it will be preceded by the croissant uh metadata. So, how do you do that? So, there is an agent data depositor, which takes this file, right? And basically goes through this and it it understands these numbers, it understands the the the headers and it looks them up and basically it's using them, you know, it's all the semantic machinery and it creates the data set from the spreadsheet and it creates the the header, uh right? And and and so this is what it does. The reason puts it analyzes web pages. If it's a web page, if it's published on the web, it analyzes the context, it extracts descriptions, metadata, and it creates a cross-on expert, drops it into a dataverse and it shares back the cross-on record, which describes this data set

Now it becomes shareable. So, uh the numbers in this uh data uh they actually become cross-checked to something called CDIF, which is cross-domain interoperability framework. That's something that CODATA supports. So, there is a lot of data which is, you know, represented in different ontologies, in different units of measure. So, CODATA uh maintains a correspondence between different ontologies scientists use already, right? And different num kind of numerical measurements. Uh so this is an authoritative way to translate some accepted scientific ontologies and numbers. Uh so CODATA maintains this, right? So, these numbers either are numbers extracted from the text that have been run through CDIF and are presented in the interoperable format with annotations. Uh so Bellfire can actually look at um uh web pages and actually um uh parse the page right in the browser and do all of this, create the the knowledge graph, uh extract the numbers and and understand it

It extracts metadata. I extract the variables. So, see, it's basically it extracted a bunch of variables. Uh it found units of measure. It found the context describing these numbers. From the whatever tables it found, right? So, instead of just like a sequence of numbers, now you have pretty rich numerical data. And then it creates a data set. And you can download it

Another thing Belfort does, so there is a bunch of plugins Lava built. One is is basically you can point it at any URL. If the URL is a video, the video has been transcribed using Gemini, and the transcript has been used as the source of building a knowledge graph, right? So, it does the same thing with the text it extracts from the video. It knows also the timings and everything else. And now you can talk to the video. You can ask different questions. What happens in the video? And it will answer it. Uh so, how it works

Um Again, so it does named entity extraction every time it kind of finds a piece of text it runs it through. It's using uses spaCy. Right? So, it's it's it's it's very fast. It doesn't feed like everything uh painstakingly through a lab. It uses pretty accept- kind of accepted existing stack uh which which works very fast. And it extracts right now 18 different kind of entities. And that all has been integrated with cross arm and data. Um It identifies different questions

Uh boost relevance. So, for the it's using both graph rag graph database Neo4j and it's using rag, which is quadrant, which is the industry leading vector database written in rust, extremely fast embeddable. uh and it has a variety of different ways to search through embeddings, right? So, so this this combination allows for very fast uh search for relevance. So, the repository is very much evolved right now. Like, when I tried Docker, Docker will depend on Nvidia GPUs, which only exist on Linux, and blah blah blah. So, this actually now you can run it in a browser. All right? And it will basically give you uh you know, meaning of the pages next to it. Uh So, Ghostwriter is a key part of Pale Fire

It's uh the engine which writes descriptions. Right? It basically, once you have the semantic layer extracted, it you can ask like, "What is this?" And it will write you it will generate text based not on whatever LLM thinks, but what your knowledge graph actually contains. Um there are skills. So, it can actually understand Dataverse. It can understand Excel spreadsheets. It knows, you know, interpret with the framework for very for variables. And it resolves uh that's another feature. Thank you

Yeah, I got some weird uh kind of cold or something. Which persists. Thank you very much. Uh yeah, so this is something that, you know, I cannot go into all the details. Like, again, Slav's um ecosystem is growing. Another very interesting thing I just mentioned very quickly he did. So, there is something called DID, decentralized identifiers, uh distributed decentralized identifiers. Who heard about them? Yes, Adam heard about them

So, this is something which Microsoft built uh 6 years ago. And like, it actually is a standard. Because people needed unique IDs. We We unique IDs, right? But we need globally unique resolvable IDs, which everybody in the world can use. Which unfortunately means something like blockchain, right? Because there is kind of an authoritative sequence of IDs, which everybody should verify that this is my ID. There is many IDs like this, but this one is mine. So, you need to have it. So, like we don't like the word blockchain around serious kind of communities, but what what this thing does, it's using kind of blockchain for authority, right? Not crypto mining speculation

And another beautiful feature is that you can stick any payload into the DID as it exists now. So, Slava puts all the prompts in the DID. He can version. So, he basically can snapshot the world using DID. And you can uh basically measure the drift because LLMs change every day. You can snapshot everything and know exactly what happened at every specific time. Right? And so, nobody's doing it properly. Every day we come and LLMs are new, right? Like DID framework lets you basically understand what's happening uh through time

And ODRL is another existing object description and rights language, which lets you define who has access to what. So, if you have this, you know, lots of this data extracted into your knowledge graph, you really don't want to give it to random people because what if it contains personal information? What's in an HR graph? Like it can be important. So, so you need something like DID plus ODRL to have very fine-grained control. So, we're kind of getting in the enterprise, you know, space. Again, that's multimodal. Um and that's kind of uh yeah, and so this is the, you know, ODRL demo. Uh kind of that's how it measure, you know, manages rights. A lot of a lot of things

You can look it up. Uh it's the And it has, of course, a few Everything now has an MCP. So, it does have an MCP as well. It's very modular. Uh you can get it from Docker, you can uh run it locally, you can build it from source, you can run it now in the browser, and it's built on open source. So, this is the page you can look at everything up. The most important thing is the GitHub URL for X tag. You can find it here

Obviously, so much as there is 46 more projects in it. So, check it out. Yes. Right? And footnotes is very interesting. This is the browser plugin which Slava built, which lets you run it in the browser. So, that's what I have. Uh I am not sure I have uh the demo. Actually, I do

Okay, cool. Now, this is what the demo looks like. And so, this is the connection to uh to X tag. So, you can basically put any page in here. So, Slava built an extension which opens a little tab. Right? It goes to Gemini. The magic of this is using free Gemini. You don't really need to configure API keys

You go say Gemini init. It sets up everything in your like basically there are two commands, and it's all in the docs. In your Python, you know, directory, and then you start the demon. That's all. You There are two things you do. It runs in the terminal, right? And so, let me see. Do I have my terminal? I have my terminal. So, see this is This is the chart like this is the video transcript

This is the page transcript, right? Like it converted everything to text, so you can see what it's doing. So, in this little tab, it basically It builds this knowledge graph, and it tells you, right, what's happening. So, this is um So, this is what I put like I put it on I on CNN page and like several hours ago, and CNN had There are numbers. So, look at this what it extracted. Because TSA is, you know, defunded, uh TSA agents are calling in droves to say they're sick, right? So, like so it extracted 33% as a number of TSA agents calling sick. And now, of course, the another fun part of the war that we have the gas prices going up. So, it found in the CNN page. You know, the number, right? It found the the the number, which is the gas price

So, like, it looked at this whole like mess, right? And it found some interesting numbers, and it explained economically that the country says sharp spike in energy costs. So, this understood it using LLM, and it built it, right? Like, this is actually much more meaningful than reading a bunch of garbage on the CNN page and trying to understand what's what's happening, right? Like, just point like point your browser and just read the little tab on the right. Like, your mind will be clear, and you will be serene. Uh now, like, this is the coffee. So, so Slava told me that coffee was one of the major driver. X tech, you know, took interest in this because coffee is a major multi-billion dollar economy. And if you know history, there are two things which led to kind of insurances, farmers trying to, you know, hedge against uncertainty in crops, and Venetians trying to hedge against uncertain pirates sinking their ships on the way to Alexandria, right? So, basically, and [clears throat] kind of, you know, the ships are gone away, but the farmers are still here. So, so like, agriculture is one of the most uncertain things

And strangely, it's a numbers business, right? It's futures, it's pork bellies. If you buy pork bellies futures and you don't unload them, they will send pork bellies to your house. All right, like, it will send you a train. Uh so, so this is the kind of one of the pages that and so Slava basically built this. So, you can look at history, right? So, you can page through the history of things you pointed Palantir at, and so it understands, right? I have mapped the extracted commodity data into the CFD, you know, framework. It shows you the table which it created, right? And it has a different table. It rebuilt the table from its understanding of this moving kind of blinking thing, right? So it it has supply projection, market status bearish, right? It extracts the sentiment of the market, right? So the apparently we have our supply of coffee, so the market is bearish as at least a few days ago. Uh so it's super fun, you know, and and so this is another one, right? So I pointed at the previous page

This page provides comprehensive real-time and historical data profile for Arabica So it extracted a very good summary. So you can point this to very keen financial pages. It will give you fairly reasonable explanations, right? And so I I can see how it leads to uh you know, uh basically very useful thing. So I will leave this slide up and maybe now maybe we can bring uh the mic to the people. So if you have questions, now I will bring the mic to you. Thank you very much. Also security is kicking us Security is chasing us out, but we'll do a couple questions. No problem

Until they physically remove us from here. We'll try to be fast, everyone. >> Yes. Hi, thank you for the presentation. So uh when you put when you pointed it to the CNN page, how did it figure out the relationships between the stock prices and the and the gas prices and the events that were going on with TSA? >> CNN it's just file numbers and context around them. So right, it basically there are no relationship between like the CNN page is just a bunch of disjoint news. So what it did, it found the TSA chunk, but then it did queries of LLMs and it LLMs now are up-to-date. They know the news, right? So it decorated the TSA number with some meaning which it extracted, right? From from LLMs

And then in a different chunk it talked about the gas prices because there was maybe a short kind of gas price up, you know, update and because again it went like in a query, what is this what is the gas price? It asked like, you know, current Frontier models which are updated every day, every hour, right? Now they keep in track. And so they they basically it showed you but then it put it in their own kind of semantic representation, generated the text which is meaningful. So it give you a very nice and clean explanation of this number. Not just from what it found in the page. It use it as an anchor and the query context to enrich it. Okay. Yeah. But if this was a chunk of text with multiple names talking about them having a relationship, right? Like, you know, you know, somebody buying coffee, then it would build a graph with entity buying as a relationship and coffee is another note, right? Like so that in that case it would extract relationships, right? There there was a richer page to to actually have more relationships between the entities it finds

Okay. Yeah. Maybe if it can just run keep running it would figure out that within CNN page there are these different relationships. It's a good question. I mean, it just shows that CNN page is really useless. But you know, or not very interesting. >> [laughter] >> It's interesting because it's a it's a an actual production example. Uh what were the dependencies in that example that you were using? Like what what model was that? Flash or Gemma or >> It's Gemini

It just goes to Gemini 3. >> Right, like it's basically using 3 version. It's a large model, yeah. And uh I just just here. You just hold it like >> Okay. Uh so other than that, what else did you use um for training sets? Um what led out of your say your GitHub repo, what else was was required as a prerequisite for performing that search? >> So again, like it it runs data through a lens, it runs uh basically name entity extractor and puts it into knowledge graph and it creates embeddings which it puts in a vector database. And there is a lot of glue. Yeah, so I cannot go really into details

>> What But what were the embeddings that in this case that was a prerequisite? >> I don't know exactly embeddings, right? Like Slava built this. I'm just presenting the overview of this. I am. I'm looking You can go to the repo. It's all open source. So I think it's basically using kind of standard open weights, Open weights. models and embeddings. Any time it gets, I mean, Gemini just you can get a hold of it for free, so why not, right? >> out of out of out of box embedding model

It should all be in the GitHub repo. Yeah, yeah, yeah. there are there are like you have configurations. You can configure what you want to use. Yeah. That's that's the trick. So kind of I I would really like the help of the community. You guys can kick the tires and ask Slava a bunch of questions

He's always online. All right, like cut issues against this repo. And if I may just add one comment to that is that the reason why this was put into the Linux Foundation is to make it irreversibly open source. Mhm. So there can be no bait and switch on that. So the community can really feel safe about contributing and >> [gasps] >> Yeah. That's the whole idea. Code it is not Open AI

It's a public European organization. It will not All right, like it will keep it open, for sure. And but the good thing is like now we all know about it. We should all use it. It's super fun. And I mean, Slava actually has a very interesting philosophical take that we're doing kind of AI wrong. Like he has much more precision, much more speed at a fraction of a cost. So he he basically he's not it's not just kind of a side project or a tool

He positions it as a