Devreal

The World's Largest Open Source Entity Graph

Event: AI by the Bay

The World's Largest Open Source Entity Graph | Jose Plehn, AI By the Bay 2025.

Recording: The World's Largest Open Source Entity Graph | Jose Plehn, AI By the Bay 2025.

Thank you. Um, so yeah, uh, good to meet you all. Uh, Jose Plan. Uh, in terms of full disclosure, I wear three hats. Uh, so I'm founder CEO of Briqui or BQ as we like to call it. Uh, BQ is a data company. [clears throat] We're often compared to Dunham Street, Experian, Equifax companies of that sort. Uh I'm also founder of openendata.org

It's a brand new entity that you know uh hopefully that I'll be talking about and for you to visit. And I'm also a board member of the uh AI alliance and co-executive director. The AI Alliance is uh a nonprofit trade association and uh research lab that's u u in fact we have a booth u on this floor. So I would strongly encourage you to to go and uh talk to the folks there. it's uh free to join and uh part of this work is as a member of the AI alliance. So I'm going to give you some um background on BQ in general and a lot of work we do with government agencies because that actually uh is relevant to uh the open data sets that I'll be talking about and u basically we're introducing what we believe is the world's largest uh open u entity graph and by entities as we will learn I'm talking about companies people places of business addresses concepts of that sort but what we're hoping is that this be the launching pad for what we're calling the open data consortium which is for others to contribute entities that they like to track and care about and piece it all together which we'll be talking about. So um just by way of background I'm an economist by training. I'm not an engineer

So uh please uh you know uh uh you know take that into account if you ask any very deep technical questions. Uh but I was a professor. I was a economics professor then eventually I was an accounting professor at Cal very nearby at Berkeley and then I was a finance professor at UCLA. Uh I did one company and obviously it's my second one Briquery. So um Bryquery we service uh or we serve highly regulated industries. So we serve US government agencies, we serve Wall Street, we serve banks, insurance companies, um legal folks and so on. Whenever there's missionritical data um uh at stake for some kind of decision, whether it's a compliance decision, a financial decision, investment decision, lending decision, our data often feeds into that decision. We started off covering just the United States and uh that is still our core advantage and core competency but we do not have global coverage and u uh this will be relevant as we talk later on about the data sets that we are open sourcing uh for you all to benefit from

We um um we have a uh what I like to call as a symbiotic relationship with government agencies. So we we have contracts with them, we work with them and we often clean and analyze their data and build AI systems on top of it. And then all that work that goes into working with these agencies then you all can benefit from in terms of the data that we make available. Very often uh we get asked this uh especially at conferences and at events we're like well how come I haven't heard of you guys a brequery? So um uh for those of you old enough to remember when Intel was inside laptops and so on, we're kind of often described as the Intel inside or the Nvidia inside. So we power u a variety of platforms. We power platforms that service Wall Street power platforms that service banking uh insurance companies and so on. Through our platform clients, we service uh tens of millions of end users. And you've probably all been touched to some degree or another by data that we provide through platforms that um you use directly or indirectly

But yeah, we're kind of uh deep inside the the data bowels of of systems. Having said that, some at some point next year we are going to start servicing end users where we'll have a variety of uh user interfaces, APIs, MCP servers and such so that anyone who wants to build directly off of our data, you know, can do so rather than working through a partner. core of what we do is tracking entities and it starts often with legal entities and then oh let's see if it goes back on okay good so core of what we do is track legal entities and beyond that other types other entity types that that we'll be talking about uh we call them organizations companies call them people addresses locations and such but uh the tracking of legal entities and corporate structures is surprisingly complex, you know, because companies can have multiple children and grandchildren and subsidiaries and holding companies and the people related to them. Maybe it's related to an entity or a place and so on. They might have one office. They might have multiple offices. They could appeal boxes. They could try to hide their activities or could they be or they could be out in the open

And it's our job at Breckery to kind of track it all down. And then you can benefit from that with this entity graph. Um I point out kind of one of the things that we're famous for is the financial information that we sell. So a lot of folks who come to Briquery and license our data, our premium data, it's because in all this tracking that we do of the economy, we collect tax and financial information on every business in the world. So every every company, there's about 300 million companies in the world. We provide financial information on them all. But again, Wall Street and others rely on today. Um if there are any bankers in the audience uh so but we do have also have a deep partnership with Prosite

Prosite in case you're not familiar with it is u is the result of a merger between BAI and the RMA. BAI represents the banking industry uh the largest banks are all members and RMA is the risk management association. That's the trade association that uh represents risk professionals. I'm going to talk briefly about some government work that we do because again that will hopefully give you some confidence in the data that we are open sourcing and in and it will give you a better understanding of where it comes from. But some of the work that we have through a variety of contracts is um we're building out what's called the national secure data service for the US government and its allies. It's basically a centralized system that u kind of tracks, quantifies and centralizes data across all levels of government in one platform. It's actually the first time in history, United States history that such a system has been developed and um we're u we're building it. Here's a very u um uh you know little snapshot of kind of what what it entails but long story short it entails coordination across many many agencies and you know government it's difficult to get coordination um I should also point out that uh brequery of the company and the AI alliance uh which I help represent and brequery is a member of you I'm on the board of directors um we do have uh you know to be very open and transparent a um a western view of the world

So um we are contributing to America's AI action plan and u and its allies. So a lot of the work we do is in helping support uh the United States and its allies establish leadership globally in terms of open-source uh data and AI technologies. Now we have a lot of work in support of that through our collaboration with these agencies. a lot of the work we do and this entity graph that I've been talking about also uh the way it benefits government and those of us who reside in the United States and its allies uh revolves around AI readiness. So a lot of the work is trying to make the information collected by agencies uh AI ready and u any of you interested in some of this work be happy to uh to discuss in more uh detail. [clears throat] Last thing I will mention about the government work and again this is relevant to the entity graph because this entity graph that we're publishing uh we are um making uh we've already been using it in the work that we do with government and so uh it's sort of battle tested it's also been battle tested on Wall Street and and in a banking context but um one of the things we're doing is we're building a chatbot of the US for the US government and its allies and that is going to be available for the general public to use and consume. And this is a system whereby you can query um anything about uh the United States and they will provide a factual answer with a certain level of confidence and so we have a scoring system that uh kind of will tell you how confident we are in the answer being provided by the system. Having said that, the system does not rely on the open web

It relies on data that have been approved and trusted by the member uh government agencies. Here are just some examples of what this looks like. And built into this basically this graph that we're talking about powers this system. Okay, so every time you ask a question to the system about a country or a company or a person or an address or something like that, it's drawing from this open entity graph. Okay. Yeah. So here just some examples. This has not been released to the public yet

We expect it will be released next summer. Yes, this government uh AI system. So now to jump to what I believe you're all here for, although I kind of hopefully built a little bit of suspense. [laughter] So open data.org. Um what are we trying to do? The tagline that we like to use, let me There we go. The tagline we're trying to use is we're trying to make the world more factual one data set at a time. So I'm a firm believer that uh for us to build trust in AI systems, we need to trust in the information that they return back to us. And a core tenant to achieve that outcome is to have highquality entity resolution

And what do I mean by an entity? I'll go into details, but just to give you a preview, an entity to us and to me could mean a company, could be a human, could be an address, could be a place of business, could be a relationship between these things and extensions thereof. It could be universities, it could be academic publications, it could be a financial security, it could be a ship that's out at sea, it could be an island, it could be the people on the island. All of these so-called things are entities. They all deserve a unique ID. They all deserve to be tracked and monitored. And we must all agree as to what this ID is for this particular person, company, place, address, what have you. Okay. Now, again, confidentiality, privacy, and so on is important

We're not trying to let the world know everything about you as an individual. That is not the objective over here. The objective over here is to help ensure that AI systems are factual. And when they're referring to you as a human or your com the company you work for or some address you're associated with that but it's done with accuracy and critically provenence so we know where this came from. Okay. So I'm sure you're all familiar with open data but had one of our researchers kind of find some good reliable definition and here is one. Open data and content can be freely used modified and shared by anyone for any purpose. That seems like a very reasonable definition to me

We're all I think in general proponents of open source and open data and certainly that is at the heart of a mission of the AI alliance and myself personally and open data.org. Uh the US has led the way in some respects. The the EU has also led the way in other respects and Asia uh various countries in Asia are making way. others in Asia are not and there is inherent significant economic value associated with this uh you know with enhanced open data. Here are just some you know um little examples of that but hopefully you all agree that it's a good thing uh for there to be more open data and there is economic value associated with it. So what is this open entity graph? First of all, we are licensing it under Creative Commons Zero. That's the most permissive license there is. And the way to think about it is as building blocks

The most fundamental building building block is a global directory of every company in the world. There's about 300 million of them. It's a global directory of every place of business in the world. There's about 500 million of them. It's a global directory of every address in use in the world. global directory of every human, every adult, working adult associated with a company in the world. There's about one and a quarter billion. So it's identifying them all and not just identifying them and iding them but also the relationships between them

So I Jose plan reside in California. I'm associated with brequery there. Brequery has various offices and legal entities around the world. This is who they are and such. And so we can all agree on what those are and the relationships between them. By providing this directory, we can all have a golden source of truth. And as I mentioned, the way we're sourcing this information is through deep partnerships and work with government agencies. Okay? We have what we call the legal view and the business view, which I'll get into in a bit more detail

The legal view is basically the legal regulatory filing perspective. The business view is what you would see in the open web, public profiles on LinkedIn, corporate websites, news articles and such. Some people trust legal filings more. Some people trust the open web more. We make both of them available to you. Okay. Now because these are in many respects sort of nodes and edges between the nodes at heart this is a graph. I believe it's the largest graph in the world of entities

I don't know this for a fact but my instinct is it probably is. We're working very closely with all the major tech companies that track this kind of information. Google, Meta, Microsoft, Amazon, Neoforj and all the various providers out there. And as far as we know, this is the largest such graph the world has ever seen. And conceptually, the way I want you to think about it is as follows. Think of what Google did. What did Google do? It understood the indexing and linkages between all websites on the planet. at least the major ones, the ones that were being cited the most, right? Collected basic metadata about them, how they linked to each other, and then it ranked them with this page rank system so that when you run a query, it would find the websites that are most relevant to that query

The idea here is the same, but instead of linking websites and linking text and information on those websites, which is what LLMs and search engines do, instead this is at a more fundamental level, at the atomic level, entities. And if you think about it, us as humans when we consume information, generally speaking, it's information about what that we care about. At most fundamental level, it's information about companies and people. I believe now there are other related things obviously that we track but at its core those are the two building blocks of humanity I think and so all the world's information ideally through this effort will be linked to those two concepts okay now we do not know everything and it's probably a good idea that we don't know everything but if you look at a collective knowledge and wisdom of a world Collectively, we all know everything. So part of the objective here is to build out what we're calling the open data consortium, which is the opportunity for other companies or individuals to contribute data to this open data set. And there are various ways of contributing. One way is correcting errors. These errors happen

Anyone who's done data cleaning knows that even the most highquality data set from the most secure bank in the world will have errors and those cost but that's part of the cost of doing business. So we put a lot of work in curating our data. So do the government agencies we work with but we're not perfect. So part of the vision is and part of the benefit to us all in open sourcing this is to have some kind of crowdsourcing mechanism whereby the general population can help curate and fix errors that they know to be false. That's one objective. Another objective of the open data consortium is for you to contribute entities. Now I talked about before well you know we're tracking these so-called five pillars of economy which I'll show a visual nice pretty visual about that companies people legal entities and such but there are other entity concepts that you might care about in your line of work or in your company or in your industry. We've already been approached by academics who want to track publications in this graph

We've been approached by universities. We've been approached by construction companies that want to track physical assets around the world and link them to this graph also import export banks who want to track shipping. Environmentalists the other day I was presenting about this actually in Oakland as well and environmental lawyers said they wanted to help track environmental assets and link it all to this information. I think it's wonderful. I hope this helps solve at least a piece of a climate change problem. I mean often starts by tracking things, right? If you can't track them, how can you monitor progress? So that's the vision behind the open data consortium to have folks contribute entity concepts that they care about, but critically not just upload some data set because there's a bunch of open data sets out there, hugging faces full of them. I'm sure you've seen a bunch of them. No, for this to be useful, they all have to be linked

This is a graph. Okay? So any data set that you have about some ent that involves companies, people, ships, uh, beaches, I live in Huntington Beach, California, one of the most beautiful places right by the beach. I love it. I care a lot about that beach. My kids, we help clean it up once in a while. You know, things like that should be tracked and be part of this entity graph. So uh once we have this imagine a world in which every time you run a question you run a query a prompt of some system AI system that is imagine it was grounded in the truth of these entities. Okay that to me would would be ideal

So the five pillars of the economy are the building blocks of this graph. Again this is not this isn't it. This is the beginning and there will never be an end because I think this will be expanded upon probably hopefully beyond my lifetime. But these five things I think are the most fundamental aspects at least from an economics standpoint. And I am an economist. I apologize. But at least that's what I care about and that's what I think most businesses care about these concepts. organizations that's a company legal entities specific LLC's and corporations people of course locations aka places of business and addresses [clears throat] now as I mentioned we also we track both what we call the legal view and the business view right because some people trust regulatory and filings information and only those and some people trust the open web well sets of information are available in this system and here I give some you know kind of examples as to what exactly do we mean by the way this is critical all addresses are geocoded already so you will get latl long and we put a lot of work on that in fact most of the work we do is in cleanup and cleanup of what company names people names websites and addresses And addresses are very painful to clean

If anyone's been involved in that, hats off. Brutal work, especially globally. Imagine having to standardize addresses in Palestine. Imagine having to standardize addresses on the British Virgin Islands or mainland China or Hong Kong or Indonesia and doing it all in a standardized global way. No one had done this before in an open source fashion. we're doing. So, when I say the legal view, I mean something like this. This is an example of just a little chunk of a graph, right? We have companies, we have people, they have entities, they all relate to each other in some form or another

So, the legal view is kind of that official view that bankers and compliance officers and lawyers would care about. The business view is what you might see from corporate websites, public LinkedIn profiles, news articles, and such. Some folks care about that as well. So, that will be captured by this graph. That is captured by this graph. Now, we're obviously not the first to come up with an entity graph. We're not the first to come up with entity concepts. There are plenty of them out there

There's a company called Safecraft that invented a place key. There's the Gur's ID from Overture Maps Foundation of a Linux Foundation. There's Open Figgy from Bloomberg having to do with financial securities. There's leis from the Gly Foundation having to do with legal entities on and on and on and on. That's wonderful. All of these entity concepts should be brought into this graph. And if there's some entity concepts that you're familiar with in some country, in some vertical, in some industry, please bring them in. But keep in mind we want this to be CC0 creative commons open

How does this help? Well, obviously it helps in many ways, but this is an AI conference after all. So a core way is with entity resolution in the context of prompts and the responses. So whenever you get a chunk of text from an AI system, imagine getting a chunk of text and every entity in that text has been identified with a certain degree of accuracy. If there's the string apple, you know, okay, that's Apple, ticker APL, website apple.com, based in Certino, California, founded in 1970 or whatever it is, etc., etc. So in an ideal world, every entity, company, person, address, etc. refer to in a document or chunk of text has a unique ID associated with it. That would include brands. By the way, I didn't mention brands, but that's another complexity with companies

They don't just have subsidiaries, they also have brands. So if we think about the use cases of this, hopefully I've convinced you that this is useful and hopefully you will all be able to benefit from it, but critically hopefully contribute to it in your own fashion and spread the word. The world that my company lives in is a world listed various verticals of entity resolution, [clears throat] master data management, Wall Street, compliance checks, credit, sales and marketing. These are fairly fundamental, you know, verticals, all of which rely on foundational business, people, address type data. I should mention that one of the key partners in this effort, it was in one of the slides, but I'll mention it um also in another slide is Senzing. If you haven't heard of them, I strongly urge you to check them out. Sezing.com and Sensing is one of the leaders in entity resolution. You can learn more about the various fields and such that we have available and [clears throat] to encourage you all to make contributions

You will be rewarded. Okay, we're setting up a system of give to get, but if you add data to it, you will get data back. Okay, premium data, premium data from my company, Briquy, and some services from other alliance members. I'm a believer in rewarding those who who do good and this is one way we can do it by providing you with free additional data that you can use. Okay. Eventually we'll see hopefully there will be a plethora of services and benefits available uh to this. Here are some of the AI alliance members that are active in this effort. And uh let me summarize over here and then I think we'll have time for hopefully one or two questions

So we are publishing this in a variety of formats. If there's a format that you would like to see, let us know please. But we're are going to publish also in the sensing format. We're going to have APIs um also MCP server so you can obtain and query this information via natural language. We are first releasing US the end of this year and then summer next year we will release global. I'm happy to report that China data will be included. It's been a lot of work to be able to include China but we will include it. How to get involved? I strongly urge you all to join the AI alliance

We have a booth uh there. There should be Dave or Tim or both of them uh there on this floor. It's free to join. There's a nonprofit research lab aspect to it and there's also a trade association aspect to it. A nonprofit is free. A trade association is not. It cannot be free in fact by law. So I urge you to to go there and again please do contribute

I do believe in the power of crowdsourcing and again hopefully you can join us on this journey to make the world more factual one data set at a time. So thank you all. Hopefully we have time for a question or two. >> [applause] >> So next we'll have time for a few questions. So just raise your hand and I'll come and bring them back to you. >> Okay. Uh one question uh do you see it related to semantic web ideas and RDF formats what initially supposed to do exactly such entity databases? >> That's exactly right. Yeah, absolutely

In fact, in many ways that was the origin of this. Absolutely. And I should have mentioned it. I apologize. But absolutely, absolutely. Yes. I um my first thought was like, wow, this is scary. [snorts] Uh you probably been asked this question like so many times, and I I just wanted your comment on it

If this information gets into the hands of the bad actors, bad governments, even marketers, they're going to be marketing to my children very easily. Uh all of the CIA, whatever. Would you comment on that? Yes, absolutely. So, in fact, it's the privacy protections of people that were always our number one concern. And we've put a lot of uh thought and received a lot of legal advice around the proper way to do this. So first off, when it comes to the people data, no person's address is being divulged. Okay? So the people's privacy at least the way it's understood by various laws around the world are being maintained. The objective over here is for rare to be an ID

Now, let's keep in mind ad systems already have ids for us all, right? There's nothing we can do about that. What the main objective over here is to make sure that any information out there that's being claimed pertains to you is accurate and true. That is the main objective. But we are not going to help marketers target us. That is certainly not the objective over here. And there will not be contact information built into this open data set. So we're not going to be divulging your phone number or your email and such. Even though we have it, that will not be divulged

Yes. So I hope that makes you feel better [snorts] a little bit. [laughter] One step at a time. >> Yeah. Uh hi Jose. Quick question about the size of the graphs that you were talking about. So are you able to share roughly like in terms of rows or records how big these things get? >> Yeah, absolutely. So it's about 300 million companies globally, about 500 million locations or places of business globally and a little over 1 billion people >> and and this is before entity resolution, right? >> Sorry, >> this is before entity resolution is done

So once you dduplicate certain things >> Oh no, this is post dduplication. Yeah. >> Okay. >> Yeah. And then of course the relationships between them all the edges when you and that's of course I don't know trillions trillions trillions. Yeah. >> I think I saw a couple people at the front. So one here and then >> and there's a lady also in the back

>> Okay. So I'll go here and then here. Okay. >> Um first in response to this gentleman's question. I think the bad actors already have access to public data. In fact, they're actively curating non-public data. So, if you ever been potential target to pigs slaughtering, you know what I'm talking about. But I I honestly believe that a single version of truth and most importantly providence is the only weapon we have against AI slop and fake, you know, fill in the blank generation

>> Thank you. Here, here. No, no, but having said that, >> my biggest worry is that the people capable of maintaining that truth through crowdsourcing, will they have the wherewithal? Because the real power from these databases from I can tell are the edges that you can construct with absolute causality which only happens through build. So for example like Wikipedia these pages are kept true because >> somebody is an absolute nerd about that page. They love that topic with all their heart. >> Mhm. >> This involves many many many con connections that people going to try to introduce edges. How you going to enforce that schema so artificial edges that shouldn't be there and this is the the one bad actor can really skew toward the results and the you know replies that come up

what's going to be your guards you know safeguards against such things >> great question so first of all we are actually working with a wikime media foundation all this is being linked to wiki data and um Wikipedia articles and all that very actively involved in the development of this project uh that's one thing worth mentioning uh the other thing is at least when it comes to these fundamental entity types that's bquery's business that's what we do that's what we Do you know we day in day out we validate cross reference and check data provenence and sources and methods having to do with companies people places legal entities etc. So as far as the fundamental building blocks of society are concerned that's what we're going to be doing. We're going to be cleaning that up because we're already doing it anyways and that's part of our business model. That's part of our job. Having said that there will be areas that will be introduced that we're not experts in. So to use your analogy, we're nerds, experts in these five pillars of economy, but we're not nerds or experts in the other entity types. And that's where the AI alliance will come in. We're setting up a governance structure kind of like Wikipdia Foundation has done

And with the guidance in having teams of experts and companies that do have expertise in areas and having a series of controls, having said that, it hasn't all been figured out. But if you are interested and are passionate about this, we strongly encourage you to get involved. >> Yes, I think it was a lady. >> One final question. >> Okay. >> Um, thank you for talking about this. It's really interesting. Uh, something for me and my company that's been a problem is data staleness

So like for example, Zoom Info is another data third party data company where we can get entity information. Um, but for example, I leave my current company, I go somewhere else. How long does it take for that new node to appear in this solution? >> Yeah, great question. So, we update the data weekly fundamentally. Having said that, it's not practical to publish huge bulk data sets on a weekly basis. So, bulk data sets will be published on a quarterly basis, but there will be APIs available that have the real time well quasi real-time updates uh available uh to them. So we encourage folks to download the bulk data sets and do entity resolution and so on. But then when you want more current information to use the APIs

Yeah. Um I think the next presenter had to uh >> I'll be around. So happy to talk. >> So thank you so much for your talk and yeah we'll have a