Devreal

From Data 2.0 to Data 3.0: Making Data Autonomous for AI

Event: AI by the Bay

From Data 2.0 to Data 3.0: Making Data Autonomous for AI(Keynote) | Zhamak Dehghani, AI By the Bay25

Recording: From Data 2.0 to Data 3.0: Making Data Autonomous for AI(Keynote) | Zhamak Dehghani, AI By the Bay25

Thank you. [applause] Just want to give a quick shout out to Alexi and an all women team for putting this conference together and inviting me. Thank you. Um need to stand here. Okay, great. Well, um maybe super quick. uh who in the audience is working in data management, data processing, data engineering teams in organizations. Okay, so maybe about 20 20%

Um and who is depending on the work of people that are managing the data, building applications, agents, working with LLMs, data science, equal number. Uh so for the next 25 minutes I am going to challenge a lot of your assumptions and the way you've been working. Sorry I keep moving that way. I need to stay uh on this side. Um oh great I will take it. Yes I'm having a difficulty standing in one one spot. Um so bear with me suspend a lot of assumptions that you had uh to see where we might get to what I call uh is data 3.0. So if you imagine data management uh was big one stack terod data and formatica uh and then we as data 1.0 and data 2.0 was all about uh storage centric um kind of modern data technology um as data 2.0 data 3.0 is what I'm suggesting that it will come next uh why we need yet another approach to data management

There are times I think in in our history where the laws of physics the very foundational assumptions that underpins how we think about technology and who we build technology fundamentally changes. I'm using a dramatic example to make a point that for 200 years Newtonian physics just worked. We put um you know we projected the uh trajectory of the planets. We built bridges. uh we put satellites in orbits and they worked and they then it stopped working. It stopped working because we got closer to the speed of light to the quantum uh uh scale and then we needed a new set of laws of physics and assumptions. We needed Einstein theory theory of relativity to be able to build atomic clocks and uh GPS and particle accelerators. This is a similar moment in how we have assumed the data needs to be managed and the reason for it is that we are getting extremely challenged by the speed of uh innovation, the speed of change, uh the scale of impact of this new category of technologies based on uh agentic technologies or LLMs and the complexity they're introducing into our systems

It's been okay for teams to wait up to six months, weeks to build the to get the data to build the dashboard for a human to look at. It was okay. I mean, it wasn't really okay, but um the you know that that sometimes you know the data didn't match. The CEO look at a sales number and they get two sales numbers and they don't match. But the scale of uh mistakes is irreoverable when we have agents running around making bad decisions and taking actions ubiquitously across an organization. And the complexity that this new technology is introducing is combinatorial. And I'm going to make a kind of a bit of a point about the complexity and the the the scale why why it is taking so long why the the complexity is combinatorial and why we need to fundamentally challenge the way we have thought about data management. What I'm sharing here is off the back of working with a lot of organizations since about 2017 when I started kind of uh you know looking under the hood of data management and try to debug the situation

This is the process that large organizations go through to take data from some source. Maybe it's a purchase data, maybe data off the, you know, um, the the enterprise boundaries to get it to a point that it can be discovered, governed, and used. It's a waterfall. It's a broken waterfall process that hasn't changed really for decades. It starts with um as you see different roles handing o over different artifact trying to perform a fraction a step in that processing be able to get access to uh provisioning provision infrastructure be able to um model the data ingest the data move the data from one storage to another storage harmonizing across layers of uh data modeling um pass that over to governance to after the fact try to define semantic and business lexicon on top of the data that's already been produced. Define data quality again after the fact to decide if the data is good or bad way too late. So the steps that are a little bit out of step and that's what happens. You get friction along the way in each of these steps

And in architecture, what we end up is this hairball, the squishy middle between kind of your foundational technology, your storage and compute and security and the usage uh application layer, whether the dashboards and machine learning models or um LLMs. We end up with this herbball of data pipelines and duct tape, data cataloges, data semantic layers, a lot of different technologies that we need to kind of duct tape together to achieve that waterfall of data to value. And while the application is now evolve in minutes or seconds a new prompt, a new behavior, a new data required to satisfy the intention of that prompt, the data itself still takes weeks or months to become available. And that's a discrepancy that I think we need to have a very hard look at. And on the complexity, this is the reality. This is a very simple formula that describes the reality of the complexities as large organizations are dealing with. What I'm de demonstrating here is the access of diversity for managing data. So in a large organization, you have one of each

You have data stored in many different locations. You have the formats of that data in many shapes of form. park files, delta files, the you know the the the tables and vectors and so on. You have different modes of processing, different modes of access. People with very diverse skill sets. The people that are doing prompt engineering have a very different set of access skill sets and approach to people that are analyzing the data in a dashboard and so on and so forth. So any time that you add a new dimension just a you know a couple of extra numbers on any of those parameters you get a thousand new configuration to manage and that management the cost of managing this uh complexity is now on the consumer organizations that managing this stack. So hopefully by now you're slightly convinced that the situation that we're dealing with is not ideal

But this is not the first time that we have been faced by [snorts] complicated solutions or complexity. And architecturally we often applied three principles to tame that complexity. We encapsulate. We hide things in a box in manageable units. We abstract, we define APIs around that complexity, a clean interfaces we can talk to and then we automate it. Uh we remove a lot of manual um kind of interventions required to deal with that new abstraction, that new unit of processing. Can anybody think of prior evolutions of our technology where we had to apply these architectural principles? Some heads are shaking. Anybody wants to mention any data warehouses operating systems we defined we abstracted hardware we or virtualizations we abstracted hardware we defined Unix process we defined it streams and interfaces for that containerization kubernetes or docker again we abstracted diversity of the hardware and so on and so forth so I think we have seen this before and over the last seven years my kind of R&D has been focused on defining a new abstraction, a new unit that can manage that waterfall that all of those aspects of data processing in a more holistic way but also in a smaller scale applied to a particular domain

So what I'm offering is another approach to think about data from value to use as an autonomous unit that itself encapsulate all of these different facets that that really makes data usable. It's semantic, the data itself, the processing that is producing this new form or new meaning of the data and the conditions and rules that need to govern it. Uh it's a it has a mouthful name. Naming things is the hardest problem in computing. Uh this one is called autonomous data products. I named data products off the back of data mesh. The work that I did on data mesh and it got hijacked. So now I have a new name that is a little bit harder to pronounce

But but what is trying to say is that really we need to think about data and its process of generation as a product and then we need to kind of automate it. So again what I'm showing here is the not is is a concept that we created at next data. It's a running polycomputee process that it's been um instructed to define data or produce data for a particular domain semantic. It senses its environment in terms of where the source of the data for that it requires need to be processed and when that data need becomes available it process it. It guards itself with code as you know contracts and competition uh not governance to against bad data coming in and against bad data getting out. It's accessible through a global URL. So you can kind of I can email it to you and you can pass it to your friends and get to it and plug to it and ask questions about the data from it or read the data um uh get the to the access to the underlying data through it and be able to observe it. And most importantly that complexity we talked about that heterogeneous environment needs to m be managed and hidden somehow

So any one of these kind of processes uh data products it adapts itself through drivers just like an operating system to a different compute to different storage. So you can run it on a warehouse or a lakehouse or others. So then what would happen if we define this new unit of data as a product? What would happen in is that this herball of data management get replaced with essentially an autonomous set of interconnected each governing itself each running itself um set of data processing nodes that are serving data semantically aligned domain oriented data in different modes. So it can one might be producing um you know marketing information marketing promotion effectiveness KPIs uh and is producing that data as vector or as files or as tables and another one might be um processing the sales information or correlation of sales and news. And each of these are interplaying with each other because they're consuming probably data from one upstream node and serving data for to downstream nodes. and they each hide away that complexity of where their compute is running, where their storage is being stored or how they're running that compute as a stream or um or batch. And finally, when we think about the humans involved in generating data as a product, the process need to be much simpler. So if the humans are working with one artifact, one living breathing artifact, these data products, there are humans that are designing the kind of the modeling and the shape of that and the what business defines for example promotions um as a structured or storage agnostic way

uh and then there developers that are kind of turning that into the actual transformation that code that's generating the data according to that storage agnostic semantic and there are users that are finding the APIs to these data products and discovering them use them and in this case might be agentic users or um human users I'm just checking the time making sure we have time so there is no separate pipeline there is no separate catalog there is no separate observability. So like this all of these capabilities is now autonomously managed and orchestrated by this new um abstraction. Let's look at an example. Let's say this is a hypothetical retail world that we is uh you have an autonomous data product that is providing information about customer feedbacks and is calculating that information from upstream data products that are providing perhaps customer feedback at a point of sales through PDFs or documents and customer reviews from digital channels like Amazon and so forth. So this one data product is only has one job and its job is providing data around customer feedback across different channels. So it's it's self running and act retrieving the information that is relevant to it and it's listening when the upstream uh information changes. It uh enforces quality controls and contracts before and before processing that data from upstream and before promoting that data uh for availability to downstream. It has APIs for um human users or uh machine to be able to discover it and get access to the data that it needs

And most importantly, it provides that secure and governed data the the Amazon kind of sorry the cross channel reviews in different formats. So there might be applications that want to get the the same semantic of data in in vector um as vector embeddings and another application or a user that might see that as a tabular format and they should not worry about they have to go to a different technology to find it. that the producer doesn't need to worry about building five different pipelines for different modalities because all the producer needs to think about is that I'm just providing Amazon uh reviews right that that kind of business concept that they're focusing so what would the interaction look like with such a such a thing so once you have a mesh of these inter you know interconnected data products they're all emitting information about their intent about their objective about their semantic that can be captured by system level data products, right? So, we have like a system level data product running on this mesh that's providing the collection. So, if you're building an a LLM based let's say uh decision to be made and the decision is um you know what which which which product which of our products is causing the most complaint the process of that is really all automated. Um the from that evaluation of the prompt uh we can discover the intent of the prompt through that intent based on that intent based discovery we find the um uh the tooling all the tooling that are available and pick the tooling that is relevant and again the collection of the tooling uh available from all of these data products itself is provided by another autonomous data product which is a system level or a discovery data product and from there we can go to the right endpoint the right data product. Um and that endpoint will then evaluate and will give us access to what really matters which is here finding um what data product is uh what product is causing the complaint. So um it is a bold label data 3.0 So um it's a bit of a marketing label frankly but the uh intention that it's trying to convey is moving from imagining data as bits and bites in a lake hassle warehouse and accessorizing all of our uh um you know technology around accessorizing that the storage to get that data into storage as opposed to data being self-governing self-chestrating autonomous applications that serves their semantic ificant domain in different modes and once we make this shift in our thinking in our approach then we find out that there are shifts across different dimensions that need to happen that this dimension the very first dimension is moving from storage centric or um bits and bites on the disk centric to really semantic first and domain oriented and processing centric view of the data. the shift around modality to imagine data as really semantic first in any mode of action uh access as opposed to data as a table or data as a file and keep fighting about what what modality we need to have because frankly we don't really know what modality we might need tomorrow you know we were talking about vectors and rags as a modality of access and now we're talking about MCP endpoints and who knows what's going to be next tomorrow but what's going to be an unchanging truth is that your business logic and business semantics is still going to be the same thing

You're either selling shoes or you know solving um uh DNA mutation problems. Those semantics will remain to be relatively constant. Um the orchestration shift the who runs things there's no central brain or central orchestrator or central DAG anymore. Each of these data products are autonomous running and orchestrating their own process, their own objective, their own dependencies. And that's where really the friction of this data pipeline management will go away and we get the speed that we need. Reasoning it's built in. So there is no way that we can kind of after the fact try to reason and add lineage or other um annotations to the system to figure out how to reason about an answer. And that reasoning again is built into every data product because they know uh where the data is coming from, how they're calculating it, and h who is using it

And that emergence of the wider lineage is um that happens automatically. And finally, governance and controls can't be an after the-act thing that a tool externally with a group of people will come and enforce. The governance need to be computationally enforced as part of that [clears throat] autonomous uh data unit. And the reason that that that leads to again another excuse me leap in our thinking that these units are uh they have a kernel they have a running process they have a brain they have actual computation running inside them. The data is not just static on disk and computation outside. Computation and data live together in this small unit of um domain oriented data product. Bottom line um I've been trying to kind of solve the defic deficiencies of data management now for quite a while. I started with data mesh and data products and now with a technology but so what I think the so what of it is that sorry uh as organizations become more ambit ambitious around their data initiatives as that ex um access grows whether in terms of the number of data use cases ubiquitously within the organization applying in every aspect of the organization whether it's the number of data sources maybe the new modalities of the data with agents However you describe the ambitions of your um or data organization and increase of complexity

What we want to get that is to get out of this plateauing of the productivity where as speed at which we can use that data and innovate with it kind of plateaus and that would really require reimagination of data management and what I'm offering here is an approach that we've been researching and we have built with kind of data as autonomous units not data as bits and bytes. uh defenselessly sitting on disk. Um if you're interested to learn more, we have more resources on the website and we have a booth outside. Happy to give a demo. I think I'm giving a demo at 11:00 a.m. And um we will be um absolutely open sourcing and some of these buildtime and runtime specifications and interfaces and APIs that we have built and there will be some announcements coming up um early next year. Thank you. >> [applause] >> Um, I have some time for questions

>> Questions? Yes. >> Uh, hello, Zamach. Uh, I'm John. Nice to meet you. I love your talk. I've read the book. I drank the Kool-Aid. Uh, so super fan

I love where you're taking the conversation and I don't know if you heard Peter's talk uh earlier but it this my question is similar in the sense that you even had a slide that looked like his like the all of the roles in a data 2.0 team and when we think about agents and what's possible and interaction across the SDLC there's the hypothesis is that that's the way to think about what we might be forward generating autonomous so the question uh how do you think about uh the human interaction between people that are all part of that organization that has to be part of data 3.0 I'm just love to hear how you think about that. >> I I I thought about the humans I was the therapist almost for all of the data engineers for a while like talking to my teams and hearing their frustrations the data governance. So I think uh I'm not suggesting that humans would disappear and we just sprinkle a bunch of robots on those steps. I think what I'm proposing is that let's make humans more effective. If you look at for example the SDLC of data processing first of all I'm challenging that SDLC because the role of humans are after the fact and ineffective as in we have put for example the role of governance whom I have most empathy for after the data has been dumped to come and try to wrangle that data and put some meaning on top of it. So if we are going to make them more effective, their role needs to come at the beginning of that SDLC and then they most of them they're not actually programmers. So that's where kind of agentic technologies and co-pilots help them to bring that deep understanding of the business and the impact of that data in the business but be participating in how we're going to model that data early on. So I think I I what I'm proposing is that let's make the communication more effective

let's challenge that SDLC and come up with an STLC that's more um kind of doesn't have to inherit uh uh the legacy approach and the legacy thinking and then let's empower those people with the technology that you know gives them wings. Uh so as an example for us you can imagine we we had to first abstract this co this this messy you know wiring of technology into a declarative kind of abstraction where you can describe what is your semantic of your data you can describe what's the expectation from that data so once you do that then that human can be given a co-pilot we call it nextie to have a conversation both with the machine and the other humans to create an artifact that is more meaningful and closer to the human than closer to the machine. I feel like we have maybe abused the machine sympathy uh you know a little bit too too much in the data space. A lot of our practices are designed sympathetically to the machine which is how effectively can I store the data in a warehouse and how effectively can I index it. But we forgot the human that needed to actually understand that data. So, so I know I went a little bit too far, but um hopefully that gives you an indication of the direction of thinking. Um >> so just one question. Where is your demo and what time? >> Oh, demo 11 a.m

Level 4, I believe. >> I think so. Yes, level 4 11 a.m. >> Thank you so much. >> Thank you so much. [applause]