Transcript: Jose Plehn, BrightQuery, on Reliable AI — Interview with Alexy
Hi, I'm Jose Plen. I uh wear three hats. I'm founder and CEO of Bryquery or BQ as we like to call it. I'm also the founder of open data.org. It's a very large significant open data initiative that I'll be talking about today. And uh finally, I'm on the board of directors of the AI alliance uh which is both a nonprofit uh uh research association representing the AI industry as well as a trade association also representing the industry to uh kind of help inform government and pursue policy objectives. So uh Brequery is uh first and foremost a data company. So what we do is we track what we call the five pillars of the economy. We track companies. We track legal entities. We track locations which are like places of business. Uh addresses, physical addresses and people. And in all respects, we track these entities using both what we call the legal view which involves uh tracking what's being reported in legal filings and regulatory filings as well as what we call the business view which is what's being kind of published in open web. that would include public LinkedIn profiles, corporate websites, news articles and such. So we parse all this information to track the entire global economy. There is over 300 million businesses, over 500 million places of business and over a billion people that we actively track and out of this we built the the world's largest entity graph kind of linking all these entity concepts. And so uh what we're exploring and have been doing now for some time uh working with US government, various government agencies, our clients on capital markets, Wall Street, our clients in banking, insurance and other areas is how that information can be used to make AI systems more factual and that's really kind of our our mission to make the world more factual one data set at a time. It's kind of a model that we like to use and um the one way of doing that is by grounding the systems in this entity graph as to you know who exactly is this chunk of text talking about or which company exactly is this chunk of text talking about. So basically it's an entity resolution problem. And to solve that entity resolution problem you need a gold standard directory of every entity on the planet whether it's a company a business legal entity an address or or a human. So we provide that directory and we're open sourcing that directory and then that can help power AI systems entity resolution uh systems. So that's when they are referring to a company a place a person etc. we can all be certain who exactly or what exactly we're talking about. So reliable AI has various facets to it. One aspect to it has to do with the entities that are being referred to in that uh text. And um what we what we're developing is a system that we're calling the factual shield which is basically a way of verifying that the uh claims that are being made in a chunk of text are actually factual. And one way of doing this is to look at okay which entities are being referred to whether it's companies, people uh places and such and then checking against trusted sources whether those claims are indeed facts are indeed correct. In the case of brequery how we solve that problem is by bringing to bear uh filings regulatory information trusted information from all over the world over 100,000 sources that we check these claims against. So effectively we're an automated AI fact checker and that is how we believe that we can make AI systems more reliable and hallucinate less. It's practically impossible to have zero hallucinations and to have 100% factual information. And therefore one of the features that we have of a system we're building is we provide confidence scores or like a percentage of how confident we feel about this particular piece of information being a so-called fact. So it might be 95% in this context, might be 70% in another context or it might be so much ambiguity that that claim cannot be assessed. And so that's what we strive to achieve. So to make it more reliable, uh one key feature is standardization you know. So um right now the MCP protocol has helped standardize things by having a system by which agents can and agents and others can communicate with one another. I do think that's critical. uh there's of course all kinds of web standards and there's also various data standards. I'm a big believer in standardization and there is still more work to be done. The area where we're helping in terms of standardization is with this entity resolution problem of having this global directory of entities so that every entity every person company and such has a unique ID associated with it and that ID should be permanent immutable. Moreover, the relationships between these virus ids has to be established and and published. So by having uh this dependable system of permanent ids, one can uh ensure that the that the claims and the so-called facts that are being uh reported in a uh in a chunk of text uh are grounded in some form of truth. Now having said that we do believe that going forward it's very important that uh others contribute their entity systems their ID systems you know so uh uh the AI alliance bquery we track these so-called five pillars of the economy companies people entities etc but there are all kinds of other concepts that matter all kinds of other entities that matter it could be financial securities it could be ships out at sea it could be um all kinds of assets and liabilities around the world. Uh it also could relate to the environment and and such. So all of these at heart should be identifiable with unique permanent IDs that can then be linked to the companies, the people, the governments and so on, you know, around the world and that is going to be a never- ending non-stop effort uh that we will all have to engage in hyperpersonalized. Um I do believe that eventually we will have an AI companion that knows us deeply inside out. um reads our emails, listens to our phone calls, reads our texts and such. Obviously, there are massive privacy concerns around that. Some people will want to opt out of the you know such a system, but I do believe that is the future of having an AI companion or AI avatar of us that really understands us very deeply. There will be in all likelihood a a personal version of this which is more for social relationships and there will perhaps be a professional version of this which is more for business relationships. So almost like a like a LinkedIn version of it and perhaps a you know Instagram or Facebook uh uh Tik Tok version you know of a social media version of it but um the system will likely reside on some kind of device. Um that device may or may not be a phone. It might be uh glasses. It might be some uh kind of small physical piece of technology, piece of hardware. And in an ideal world, that information about us is local local to that device. So it's not uploaded on some server and therefore accessible by bad actors or hackers and such. So in an ideal world, this personalized very confidential information about ourselves is held in this local device which others do not have access to and that we can readily and easily uh turn off if so desired.