Devreal

Adam Pingel: Semiont at Graph Exchange

Adam Pingel: Semiont at Graph Exchange

Recording: Adam Pingel: Semiont at Graph Exchange

talk. Um uh so here's uh the rough outline. Um so there'll be less of a demo. It's it's possible. Um maybe as we do the other talks tonight, I could get something going. I'd love to show it to you live. Um like I've got some video I can show you, but it's it's a lot cooler doing it doing it live. Um okay, so I'm going to start in uh maybe since there's a little bit less of a demo emphasis, I'll I'll I'll expand on the story a little bit more

Um, so, uh, here's here's the intro slide. Make meaning. Um, before we get into what this all means for enterprise AI and, uh, architectures and all of that, I wanted to kind of I think it's really important to get the why here. Um, and this is actually, honest to God, how I got to this project. Uh, I don't always talk about this. I don't know if I will a lot going forward, but but I wanted to today because it really sets the stage. Um, one of the goals of this project is to uh create a a piece of technology that allows um its operator to create and curate a knowledge base suitable for use by AI tools um that is fully under the control of of that operator. It's open source

It can run on um in theory a really wide range of infrastructure. I've been I started um doing AWS but uh for the last six months or so really focusing mostly on local uh and code space uh uh environments. Um but um for all of that um complexity uh it's it's important to understand like what what are we talking about when we when we when we say research? um we could kind of instantly uh adopt the the deep research framing and and kind of lose track of of what we're trying to accomplish. Um as Alexi mentioned, I I have worked in legal tech for about seven years. I joined IBM four years ago and I was in legal tech for seven years prior to that. I was at a a startup here in San Francisco that got acquired by Lexus Nexus and then we um with Lexus's uh you know their resources and data we were able to kind of you know complete building out everything we had wanted to um just about everything. Um and then I joined IBM four years ago to to do drug discovery and and material science initially kind of with an AI twist. But then Chad GPT happened and like like many careers uh that kind of disrupted mine and kind of came back a little bit more towards AI and eventually knowledge management and uh and knowledge bases

Um but so this slide what is this slide? So um I uh I I want I want this to be really an inclusive slide. You know, families come in all kinds of shapes and sizes and colors and locations. Um, but the reason I I want to start here is it's not a domain like legal. Uh, if I talk start talking about legal, 90% of you might just immediately just just drift off. It's hard to keep people. Um, but uh we all have uh a group of people that we called the family when we were growing up. And there's a story that that group has. And that story, um, unless you really want to want to do otherwise, that story is kind of not no one else's damn business

Um, you want to control that. I've got, uh, I've got artifacts. Uh, I'm sort of really, you know, lucky in a lot of ways to have have artifacts that have been been passed down for for many generations. And for various reasons, I' I've wound up with a lot of that stuff. And, um, the story here is I uh had an uncle die a few years ago. and uh they didn't know what happened to the stuff that that people thought he had. Who has the Dirks family history. Uh turns out I had it

Um but in order to uh to figure out uh what I had, I had I wrote a book. That was the shortest path to being able to answer that question with any kind of um you know, with any kind of confidence. Um so I've got a few examples of things I learned as I was doing that. So that's my that's this is Fred Dirks on the upper left. That's my great great-grandfather. Um I grew up eating those these le cooking cookies from that recipe. Um and and I I didn't realize that that recipe had been published to a church cookbook in the 80s at some point. So found that made them once and then once when I got into the family stuff uh started wondering, okay, it's labeled Grandma Dirks, but that was my grandma that referred to it as grandma

So who's that grandma? after the the research I I it's probably that guy's either his wife or his or his uh his mother uh because uh his father was listed uh in a record as a as a master baker uh in the village where he was from but also his wife's father used to bake bread according to one story for for the local community. Um so there's just one little example. Um another one here is uh so I'm from Iowa. This point right here came into, you know, the the federal territory in I think 1832, 1833. Actually, all of this yellow did uh this wedge uh was marked by this rock. Now, I found a picture of this rock. This is my picture. I found another picture of this rock that had been taken in the 1920s in a in a small pamphlet that was published in the 1970s

that pamphlet or that short little book uh said that um said that that the rock had had disintegrated since the picture was taken. That didn't quite sound right because that that's solid granite. Um so I went looking for it and the my plan was to take my dad to a dive bar and to call it a dive bar is uh is uh really not not doing it justice. It's it's a hunting and fishing bar next to a river. Uh, and uh, turns out after a few minutes of conversation, the cook came out and said he knew exactly where it was and we found it. So, I've got the picture. Um, the last one is uh, this is my great another great great-grandfather. This is him in a picture next to his house

Uh, I um been spending a lot of time in that small town for various reasons. Um, I uh uh give you just a couple of details to and wrap this up. Um, but I got a Red Fin notification house for sale. Uh, turns out was this house. Long long story short, the reason the reason I got the notification is because I've been spending time in that area, I think. Um, the reason I recognized it was first I was curious. Oh, that's an interesting house. It's on a a big but weird plot of land

I wonder where it is. says, "Oh, it looks like it's on old property that used to belong to my family." Sure enough, and and this picture actually looks uh well, you can you can make out the house. The the Red Fin listing uh looks looks, you know, you can make it out. This is from the turn of the century or something like that. But then the crazy part um is that uh I I I found his grave site uh at one point and like I think it was the next day that that I got that red notification. So it was like that's my ghost story for uh uh for for for the rest of my life. But, you know, you can really get going crazy with the connections. And, you know, you can rationally explain that, well, you know, I was primed to be able to recognize that for what it was, and I had been spending time, uh, for reasons that caused me to be primed to recognize that, so, you know, it's probably not ghost

Um, all right. So, that's my extended motivation. Um, so this is this is what, uh, what I want to drive towards today. So I'm going to given that we're at a Neo Forj uh talk, I'm going to assume that people are familiar with graphs and graph rag and various techniques and just kind of get into to what I've built. Um so I I think these are a few of the questions that the community is answering. Uh this one, you know, often people refer to that that uh as I interpret it as the the cold start problem that um I was doing a lot of uh experimentation with graph rag about a year ago, but this kept coming up. How do you get the graph? And there were a lot of people uh kind of claiming they could automatically generate it. And I know since then we've brought ontologies into the picture

Uh but I I still think that there's there's some open questions about where these things come from, how we maintain them, and how we tie them to the primary material. Um and then also there's this question of like a lot of the diagrams starting with rag and a lot of things that have come come afterwards. Um the a lot of the you know the the voodoo the the AI magic is all just sort of happening inside some kind of backend data structure and it's not really user visible. um you know there might be some evidence of what's going on behind the scenes in in the the interface that you're using but uh uh you know it seems like there's a real bifurcation that there's the the stuff we're going to show the user and then there's all this fancy stuff that by the way I'm going to charge you a lot of money for um and control um and then lastly you know uh it it's hard to there are many situations where obtaining all of the data you might need by yourself paying for producing the graph or whatever it is is very challenging. I I think there are a number of um community efforts just beginning uh now that uh that kind of fit into this bigger vision of like having something that that looks like um an AI native knowledge base but not not having to do everything in that one being able to plug into others. Um, I'll mention, I don't think I'll get to the slide, but one of the things I did at the startup before the Lexus Nexus acquisition was digitize the Harvard Law Library. It was 40 million pages. They cut off all the spines and OCRD them

Um, there were waves of of uh processing. We were held to very very high standards for reason of of wanting to not infringe on copyright and so forth. Um, and and and several other reasons. Um but uh that to me I kind of look to as an example of how there there there was a business incentive. The venture capitalists wrote the check and helped fund that. Um it was it was well-managed became an asset that a lot of LLMs are trained on. Um the free law data set that the Harvard law uh data set is a big part of that. Um, so going back to to this project, um, what hit me, I showed you all that family stuff, and the reason I spent spent some time on that was, well, how do you how do you publish that? I mean, this is one way I I literally a couple summers ago printed out a bunch of copies of this, I think, for $17

I had passed out 15 copies at a at a at a family reunion. Uh, but that's clearly, so this represents about, you know, 1/8 of my great great-grandparents, right? have to do seven other books. I'm probably not going to do that. That's probably not the right way to go about that um in this this day and age. So, I started to think, well, how should I do it? Um and a wiki seemed like a good idea. I I and I went to I thought about well maybe I can use media wiki uh like Wikipedia does or maybe I can um use Google sites but it was clear that like they just they just weren't doing what what I'd like to I'd like to be able to kind of you know grow them in an organic way that feels like it's really leveraging the AI tools that we have available to us. So, for instance, get that little book from 1974 that had the picture of the rock, throw it in, and then start connecting it to, oh, it turns out somebody lived right next to that thing 100 years ago. Well, maybe maybe one of their grandchildren knows about it or something like that

You you just never know what connections you might be able to make when you throw some piece of let's say unstructured data into the pile. So I said, well uh I'm going to build a wiki. I'll make it uh you know feel like a good human wiki. That's the first task. And this is sort of a post hawk rationalization of what I've done. But um the next thing once you have that is factor out the event protocol. So I took a look at the few dozen event types uh that it that it took. Uh this was after I had kind of factored it away from restful interface and more towards an event bus

um let's look at those um that protocol and then wherever we can uh figure out how to make the the AI agents architectural equivalents. So um one of the things uh we'll talk about is uh the the automated AI generation. So given that you've got um a reference identified, click on that and you can automatically get the the um the thing you were referring to but also with the benefit of the context that the graph is is able to bring with it. So uh and that I call that yield. Um so humans can yield they can do the same thing given context write write uh an article uh and for for all of the other interactions and there are six of them. Um, the the goal is to be able to again make make the agents and the humans architectural equivalents. All right, I think I've covered most of that. All right, so this is where I would have um gone into a little bit more of a demo, but I'll start here

Um, this is this has been me doing the commits. I think depend dependabot has one or two but uh this has been going since since July. It's open source. It's uh released under Apache 2 license. Um there is a a code space uh that you can start up and and and get this um get this running. Um there is also a a peer repository called semi workflows. Uh this one uh is set up for kind of core semant development. That other one is more about kind of downloading the npms and um and just setting up a minimal install local to whatever the environment is

Uh but then it's also got a lot of um and you'll see some of this as we go forward. It's got a lot of domain specific light examples, some synthetic data, that kind of thing. All right. So when you start it up, it uh this is one of the first things you see. um uh you know should make you think a little bit of of a wiki I hope uh you can just go as a human just go kind of rooting around um see what you can find uh there's a really basic compose new document page um all right so that that covers yield uh that's what I was just so humans can compose and and uh AI agents can can compose as well we'll call that yield so there's these five other interactions that I've that I've defined um uh mark and I I've tried to chose recently I've chosen really simple names for these things. So um so yield uh we we'll let AI have the the longer word generate yield is what we'll we'll use internally to talk about that that interaction. Mark is uh you'll see a bunch of screenshots where we're going to support a variety of ways if you think of how do humans you know here's uh all these little red thingies here to mark errors. Uh humans mark up documents

Um so there's going to be uh five different motivations for for marking uh that that this system supports and and there's going to be this is pretty well developed. there's a a human interface to do this and the agents can do this automatically as well. That's really one of the first things I I built. Now what I'm calling gather um that's effectively um fetching context for it and it's generally done in ter in in terms of or relative to an annotation. So given an annotation what's the context around that that's relevant. Uh the next one is is bind. So that's really just applied to um one of the motivations. And so this these red red tabs would would be fall under probably the assessing

Um there's a W3C web annotation standard that I'm looking to for some of these concepts. Um so red underlines would uh are called assessing just creating a link which is the first thing you need to build a wiki. Uh that's the linking motivation that's broken in two halves. There's identifying the reference of course and then there's the the reference resolution. Uh so the the LA so the the former falls into the mark category but actually um the resolution uh I'm calling bind browse uh pretty pretty self-explanatory. It's a wiki you need to be able to to browse around. Now the other thing to keep in mind as we talk about these is that um this is set up to be collaborative too. So uh one of the the modes that I'm hoping to support here very shortly is like a presentation mode

So, uh, I could sit you down in front of a terminal and kind of give you a guided tour just by sending the right events to to your semi semiion interface. Uh, and then, uh, Becken, uh, this is kind of a weird one. I didn't plan on this one initially, but I thought it deserved its own category. I thinking back on my career, there's a lot of things, you know, anything that's advertising uh, kind of is kind of beckon. um that word actually applies uh in a lot of places in terms of the UI that that ends up being like the P like hey something's generating I'm going to pulse and then something is just generated I'll I'll have some little visual effect saying look here next or scrolling something into view. Um so those are the kind of the primitives that I'm I'm building with. Okay, there's just a to drive it home. Oh yeah, the I guess um yeah, I didn't mention so Oh yeah, the red underline got got taken away here, but that's you know misspelled word

You'd see that red red underline. That's the assessing motivation. Commenting we're we're really familiar with um dealing with uh you know word processors and so forth. Um linking we've covered highlighting yellow highlights and then uh tagging. That's kind of the the doorway to ontologies. Uh I'll show an example of uh uh marking up um analyzing you know long legal discourse according to the the Iraq uh the the issue rule application and conclusion. I'm using the tagging motivation to do that kind of document understanding. Okay

So, this is the uh the general form. And this is this is the one I really wish I could show you cuz this is the one I' I've just been working on in the last couple of weeks. Um so, I don't I've got a marker. Cool. Whiteboard style. Okay. Justin was in my uh my AI class in 2002. I was a TA and he was a student

Uh it's it's good to see you here. Thank you for coming. Let's do some whiteboarding. So, um we have some a document here. Uh we got green. Let me see if I can find a blue. I guess yeah, apologies to the folks all the way over there. Um, so if I want to, you know, there'll be some some references identified in the text

Each one of them will have an entry in the panel. Uh, if they're linked, they'll show a link. If they're not linked, they'll show a red question mark. I won't go get the red, but I'll show it as a question mark. If you click on that question mark, then this process begins. What how do you want to how do you want to how do you want to resolve that reference? There's a few different ways you can do it. Um so we've marked it already. We've created the blue

Uh so now we're always when we click this click this uh question mark, we're always going to gather the context that's going to inform whe whether this is an agent or a human doing this. Um that's you can imagine some some edge cases where you wouldn't need to gather context, but it's almost definitional that if if you're going to if you're going to resolve a reference, you you need some context. Um so gather that context, present it. Uh and then you have three choices. Give now you're looking at the context at that point. Uh you can say, "Oh, that that seems like something I probably already have. Let me try that route." So we're going to use the context to go kind of match with what's pre-existing. Uh, and then this one is just, uh, hey, um, I think I'm in the best position to to provide this

I might have a file lying around. I'll I'll upload the file. Um, and and then finally, uh, the fun one, we just click generate and given the context, we generate it. And then it'll appear as a document. It'll be linked, you know, from you could click on that, you could click on that, and you'll get to that new document. That's that's the end step. That's where we end up. Uh it's multimodal too

Uh this was something that I I didn't anticipate uh immediately. But of course you don't get get too far in building a wiki. You realize okay there's a there's a whole range of um and I'm trying to to use the worldwide web consortium terminology wherever I can. So internally uh these files uh get get uh their representation. So there's resources is more the conceptual uh meaning of this but the the file the JPEG is a representation. So we'll upload that representation store it and based on the the the media type of that representation will have different ways to annotate it. So here you can see um I' I can choose a shape. I'm going to create a reference annotation

So so this will show up initially as unresolved and then here that's that red question mark I was talking about earlier. Then we can click on it and resolve it. That's a that's a fake synthetic family. Then it's Yeah, it's kind there's definitely some gory things. I asked for something said in the 1890s, but you know, it's a little off. Uh uh and then um I where appropriate. I'm not going hog wild with the standards, but but where appropriate. And I'll explain uh a little bit more about why I I'm doing this in a bit, but um the the resources themselves and the annotations can be represented using JSON LD

And I also mentioned I'm using the the worldwide it's not 100% faithful um implementation of worldwide web web annotations. Uh but um I've been seriously investing in that standard for for a little while now. Um, okay. And here's the motivation. So, you build up one of these things. If you've got going back to instead of passing around 15, $17 books at a family reunion, if I could instead, uh, I know Jim has spoken about the land party, uh, it's one way to do it. Uh, what I did, uh, when I passed out books, I also to the younger folks, to the younger cousins, I passed out USB sticks. if I could pass out a USB stick that was basically self, you know, self-sufficient

It had the application uh they could be accumulated and to form their own knowledge bases. Um that that's the idea. So uh you know I I'm I'm not doing anything especially fancy here. It's a tarball uh at the end of the day, but uh I'm just putting putting the the underlying representations at the top level. And then in a semion file, we're going to have um the resources defined the JSON LD uh and also the annotations. So that's enough to reconstruct the all that that layer. So you can think there's a there's a set of primary material that you don't want to touch. I mean that's the that's the you know I got pictures of old documents

Uh, let me find one. Well, you know, yeah, you know what I'm talking about. Uh, here we go. An old picture. Um, this is this this is uh I just want to scan something like this once and uh and and you know, if there's a misspelling, uh sometimes there's a lot of information in that misspelling. So even like normalizations that look perfectly harmless, the fact that somebody threw in an extra H or something like that, that might have meaning, you just don't know. And so you really want to preserve that primary material as faithfully as you as you can um in in the knowledge base. Um so there's the there's a um the visual summary of of that gather step

uh and it's early days um in experimenting with this stuff but you can see I'm you know looking at the siblings of of the node in the graph and trying to understand what their types are um there's it's uh I don't want to go too in-depth because it's uh it's certainly not been stress tested but I am uh I am doing leveraging the graph in some it's fairly simple but um it it works well enough for a demo when I do have the demo running uh it it does do some smart things. Uh and then this is that next step after you've got the context if you if you decide you want to take the first option. So that's the you know the for um for resolving the reference you can either find something that exists supply something yourself or ask your AI collaborator to provide it for you based on the context. This is the first leg. You've you've got the the context, you've presented it to the knowledge base, and now it's come back with a list of candidates ranked by score, which is a composite of a bunch of signals from the graph and other things. Um, here's just a fun one in in Hindi. Uh, there's 29 languages supported. Um, this is all on the model

Uh, the I've I've got support for a lama now that's untested, but I generally use an anthropic API key when I'm developing. So the the annotation of this of this it's some fairy tale uh classic uh Hindi fairy tale uh I I asked it to uh identify you know I think danger so these are the parts in the story where some character is in danger this is the red underlines and these are I believe the locations and someone who speaks the language told me I think that's river um so you can see like this is the this clearly not me doing this this was uh this was some claude model uh that that figured out how to chop this up. Um here's another fun one. Uh this is, you know, obviously a little bit more kind of business enterprise friendly here, but um there's a little check box uh in the the reference detection menu. there's a little check box that asks you if you want to um include um descriptive references. And the reason that's the reason that's useful is for like these two here. You can see I asked it to identify locations. So we've got the facility at 1247 Oak Street

Uh this is a synthetic synthetic email. Uh and then I asked it to identify people. So, we got Sarah Chan, Michael Rodriguez, the the author is off screen here, but then I because I had that descriptive references box checked, it also picked up the owner and the client. We don't know who they are by name in this email, but uh we we probably have this one somewhere in the knowledge base. And if by clicking clicking on the red uh question mark, we will we will be able to resolve it. This one, however, uh you can tell from the context here, we don't know who that is actually. And so maybe what we might want to do at this point as somebody who's managing the process around this is let's go let's go write a a document uh with everything we know about the owner even though we don't know their their identity. Sometimes that comes in later

All right, I'll wrap up here shortly. So this is the Iraq analysis. I think this is um I don't know if this is citizens united but uh Iraq is a common way that legal writing is taught issue rule application and conclusion. So again just using the W3C tagging motivation uh we can chop this uh multi-page document up into um relevant um you know passages that match those those intents and it's in a in a legal opinion like this. It's a Supreme Court case from 15 years ago, something like that. Um, not everything is going to fall into one of those four categories. But, but those four when you when you find that a judge has written this law applies to the situation. We've established the fact pattern and now we're going to apply this law

That's a always a really key moment. And not only for a reader, but if you're doing some kind of analytics on top of this legal corpus, um, that's a really key uh key piece of processing. So, um, I'll wrap up here with just a few architecture slides. Um, I I'll go pretty quick. Um, this is the idealized uh architecture. It's not quite as clean as this, but conceptually, there's an event bus in the middle, and there's this this thing hanging off the bottom, this durable knowledge base, and I'll I'll zoom into this in a moment, but then there's a few actors. Um we've talked about what gathering is. The matching is the given the context what are the candidates and then rank them search basically

Uh and this is the thing that just makes it so this this actor is listening to the event bus and then making durable in the knowledge base uh all of the uh the resource creation events and the annotation events. And then you can see what I'm trying to depict up here is there's a um a mix of humans and AI agents collaborating through this protocol uh including you know you can think of a data feed as just doing a yield right it's not doing anything smart it's just getting a stream of of resources and throwing them into the bus and then this this apparatus does the right thing with it. All right. Uh, so this is zooming into that that bottom half. Um, and given that it's a Neo Forj event, I'll we'll dive into the graph a little bit. But the system of record here is effectively just an event log and then the content store, the blobs that came in. With that, you can reconstruct, you know, what else would you really need? there's a set of materialized views um that the UI is primarily hitting for most most of the navigation when you're browsing around. Um vectors are planned

Uh it's part of why I'm not going to go too deep into the the semantics of the gather and the match because you know they they probably will need need a vector database here pretty soon. Um and then then there's the graph hanging off which uh is very important for the the gathering and the matching but um you know it's it's just uh it's basically just a projection of the event log just like these materialized views are. However, this is actually just eventually consistent. We don't need anything that's uh you know a few milliseconds uh off is is perfectly tolerable. Um it's also worth noting that you know there's a lot more we can do with the semantics here. Um there's basically two at a high level two node types. There's the resources and the annotations on those resources. So there's a lot more there's information in the things in in the in the properties that they have

There's information in of course the the the full resources that those things point to. But um for like a graph rag setup there's uh the schema per se isn't especially useful. Um but that's something that's going to be a part of uh the work in the in the few months ahead. Um you know imbuing this with more semantic information and uh in in the right way um bringing bringing ontologies to bear so that we can do you know a lot of really high confidence trustworthy um navigation of this of this graph. Um, this is a summary of what the stower actor is doing. Here's the gatherer, >> the the matcher, and then last little bit just, you know, there is a fair amount of moving parts and this is what I was having trouble with this afternoon, a little bit right before the talk. Uh, just getting the the web interface up and running. uh you can see the the event bus you know what needs to happen is that this event bus will become a true standalone piece of infrastructure right now it's kind of run just running within the back end um and then the these HTTP routes that'll just be kind of an optional um transport protocol that you can configure but if you want to run it just purely locally you wouldn't need to have that running um won't get too far into that

All right. So, [clears throat] yeah, this is part of this is part of what I was doing instead of uh configuring, you know, get getting the demo lined up. But, um this is not well tested. Uh but if you go to this this uh URL, you can find a little bit more of an exposition on what's actually happening inside of that. So, if you debug, you can throw a verbose flag on it. Um uh here are the the the way that this gets configured out of the box uh depends on these um these settings and that software. And then you know this is a uh it's a work in progress. Um it is open source

Certainly welcome your ideas and and contributions. Let me go Oh, sorry. Yeah, I'll leave it here. But um yeah, any any questions? See what what time is it? We got >> raise hand. I'll bring you a tiny mic to record your question. >> Sure. >> All right. I see the question in the back

>> Yeah. Oh, yeah. Sure. >> Yeah. That's a good one to leave, huh? >> Like this. Yes. Like that. Yep

Um, >> yeah. >> So, I saw that you uh you described it on one of the pages as a kernel and then there were some places where you said it can identify the missing owner in your example and then kind of create a request for that to be resolved into this binding. >> So, could you start with some document and then it starts to kind of um build itself out both with the help of human and AI contributors? Is that kind of the idea self extending? Yeah. Um, one of the things, uh, the demos I like to show is, so you saw that old, you know, synthetic photo from 1890. Um, I've got a little bio for that guy. His his synthetic name is Elias Turner. You resolve that to his bio and it's got a bunch survived the grasshopper plague. It's kind of kind of fun

Like I I actually I really want to use it this way. In fact, uh, I've never heard of the grasshopper plague. Um, but it's pictured in it's in this synthetic bio of the synthetic man. Um, I I it's one thing you can hide. It shows up uh as an event if you if you want autodetect all the events. So, grasshopper plague generate that and sure enough there's it's this Rocky Mountain locust thing and uh it's it you know you it's real. Turns out it's real. Um, but I think what I'm trying to set up here is a virtuous cycle where the more of your own data you have set up modeled this way, the better detecting fakes you'll be

Uh, and the better you'll be able to fill in the gaps. Um it's not never going to be perfect, but um you know the there's another slide I have in here that just describes you know at the end of the day um you know there was this this article in the New Yorker three years ago but uh Chad GPT is a blurry JPEG of the web. So, you know, thought there was a kind of a fun analogy with this, here it is. If we really build it out, uh, this what I was trying to depict here is here's the Harvard law library, you know, inside this mountain of data that gets turned into an LLM that you then do what? And if if it's a blurry JPEG, um, what do we expect to get here? What's the process by which we just get this? Um, and it it clearly, you know, it's much more highly compressed. This is 10 to one. This is 100 100 to one. So just by that fact alone probably expect you're not going to get anything worth worth using. But with subject matter expert guidance and internal data you could you know at a macro level like a CTO level CEO level uh I think you know what you what you would hope to end up with is a knowledge base that's a critical asset for your business

>> How structured are you kind of pointing to how structured accurate and precise that kind of final knowledge graph can be? >> Yeah. Yeah. Yeah. Turn it into something that you can truly rely on and do do business with. But like, you know, don't don't handcode or or hand OCR like we did with the Harvard data. You don't have to go through that kind of process anymore. I do think we can bring to bear these tools. You just have to kind of couch them in the right kinds of monitoring uh processes and and automation

>> Other questions? >> Can we go back to the Oh, >> sorry. I need to bring you the mic. >> Just the slide. The slide that was verbose. It was really close. saw I don't know if it was forward or backwards. No, no, there was another. >> That's the code of Hamarabi, by the way

>> Priorities. >> Um, what's the acronym that you used? Iraq. >> Iraq. Yeah. IRA. >> The case. I think there's like legal case. >> Yeah

>> Was a screenshot from Simeon UI. >> There it is. >> Yep. Okay. >> Sorry. I need to bring in the mic. >> Okay. [clears throat] >> Was that was that the question? The Iraq

>> There you go. Um uh is it possible uh if uh in uh some source there is a link is it possible to open link take context from the link if so how many uh what could be the lay of links how many links it can be open one after another >> yeah uh are you ask would there be a limit was that your question >> oh then another. >> Oh, sure. I mean, I don't know. I I used to work at excite.com back in the 90s. We had our own crawler, and a lot of people have written crawlers uh since then. But, uh I I don't need I don't wouldn't necessarily need to get too opinionated about that. I mean, at the end of the day, as that data feed was showing, it's a yield

And so, yeah, you could you'd have to figure out how to how to loop in the mark the automated detection in an appropriate way. >> Right. Yeah. Yeah. >> Other questions? Anybody in the back? I guess >> just checking. >> Well, wait, wait. Where are you from in Iowa? >> Uh, little town south of Cedar Rapids. >> Yeah

>> Okay. >> Pretty [clears throat] near that point. That rock, actually. That wasn't hard. That wasn't a long drive. Yeah. So it's always hard to make predictions especially about the future but >> yes >> isn't a lot of this effort going to end up um in one of the foundation models in in at least in large chunks. >> Well uh >> so that as an organization I might say well why do am I doing this because >> this isn't going to end up in a large language

It better not uh I'll be pissed. Um I mean and that's really the point like yeah I will like you know to that that they've done this amazing work of compressing a bunch of knowledge which is this incredible tool that we can we can use but it's not everything I need I have my own stuff that that I need to I need to and and and I'm not willing to share the data. So just that setup alone I think just um I mean like you know I don't have any you know I haven't done a lot of measurement about like the effort required to produce a knowledge base this way versus other ways that that maybe is a part of what you're asking too but I I think part of it is definitional that I mean what I'm trying to do is kind of mimic the processes that humans use to do research and and that they trust that they can rely on and so it's it's more of a it's not a performance Yes, I think there are good performance implications, but it's almost more of a just a trust building like design um concern than than anything that's, you know, purely financial or or performance-based. Yeah. >> All right, let's do one more question and then we'll continue the program. >> Yeah. Not so much, >> Wait a sec. I'll bring you the mic

>> Yeah. >> Not so much a question, just a comment. um knowledge graphs store relationships whereas LLMs only store semantic similarity. So the advantage of a knowledge graph is that you can then traverse multiple relationships and so the value proposition is different for each and by combining the two systems together you actually have something that's greater than the >> yeah absolutely agree >> fantastic. So on that note uh thank you very much Adam again. Yeah. Thank you. [applause] >> So, our next speaker, Sum is here

So, we're going to switch speakers. You got couple of minutes if you want to grab a drink. So, Sum, please come over. Adam, let me grab your mic. So, you do the reverse, you >> so let's mic you up. So, basically what you do, you put it the holes, these holes facing up like this, >> and then you slide the magnet under your shirt. And it connects right. So like let's do it like this

>> It's good, right? >> I hope it's good. For some reason it's yellow if it stops. >> Just a question if you have time because I already did a lot of numbers. >> It's still yellow. Okay, let's swap it. Let's Can you just like reach out? >> Yeah. Yeah. Don't do it actually

Wait. H. So okay, that's fine. They both went yellow. I think I think what happens one sec.