Devreal

AgStack: Open-Source for Feeding the World

AgStack: Open-Source for Feeding the World

Recording: AgStack: Open-Source for Feeding the World

Okay. So let me do a very quick introduction for simulator. Uh so uh the second talk is about palefire. Whalefire is a new project uh managing knowledge graphs uh and in a very precise way you will see. So I'll be giving the talk on behalf of Slava and uh Slava found a home in the link foundation with Simmer with the executive director of Xtech which is agriculture tech. I heard about it. I'm very I'm very excited because for instance I work with open source program offices opas and salon foundation founded a bunch of them one of them was founded at the nurse of vermont which has a patrick leaky school of agriculture and so there there are people doing open source for agriculture it's not all robots and whimos actually some people make food and you know what's the main thing which distinguishes humans and AI when it comes to food who who can guess when it comes to food what AI cannot do >> eat >> AI cannot eat so we eat and so this is something which we'll always cherish and have which AI cannot do and so this man is in charge of making sure we do it efficiently with open source so please take it away >> thank you Alexi can you uh thank you thank you can you hear me okay and can you all hear me all right well first of all thank you for inviting me to speak u and uh thank you for convening this wonderful uh gathering Um I'm only going to talk for about 10 15 minutes. Uh don't have a lot of slides but I'm happy to have uh question and answers and afterwards and networking happy to talk

So my name is Sumeir Johal. I'm the executive director of Astac. Uh quickly about me uh I grew up in an agriculture family in northern India. So I'm from the Punjab area of India. People who've been to India, it's a very agricultural rich uh area. We've been in farming and agriculture in my family for 10 generations. we've really uh kind of uh been in the involved in it and uh but I came to the US when I was a teenager and uh got an undergraduate and graduate degree in computer science and double e and so I'm tech trained and spent the last almost 30 years now building and running tech companies um both early stage companies startups saw the dotcom boom and bust and now it's AI so been around the block a little bit um and uh you know I really wasn't very in involved in open source for most of my career but right before co I started to look at uh a agriculture ecosystems um and I've been involved in a tech for about 15 years or so and we saw that there's a lot of um need for uh a pre-ompetitive stack that allowed innovation to happen faster and cheaper because agriculture does not have as they say perfect elasticity in terms of economics. it's very very you know resourced arena around around the world

So with that I started looking at what kinds of software and open- source work has been done around the world and was actually looking to poach some people from the Linux Foundation and in fact they poached me and and so one thing led to another and right before co actually joined to lead or to sort of start a project within the Linux Foundation called Astack. Um and so ADS stack is essentially a infrastructure stack that allows for anyone any stakeholder in agriculture who wants to build a digital you know product digital artifact to have some sort of pre-ompetitive underlying uh you know uh for lack of a better word stack uh to help them accelerate do it quickly not have to reinvent the wheel every time. And with all this information out there, weather systems and satellite systems and all kinds of information that's just around the globe and really exploded in the last 50 years. Um the work of the innovator should be made easy as they're trying to overcome some of the challenges in agriculture. That was the inspiration. Uh so you know fast forward through COVID now you know five six years uh we're now 47 repositories of different types of projects of which Palefire is is one of the later latest ones um and we're very excited about Palefire and Slava's work which I think Lexi is going to talk about um but it's essentially uh a whole bunch of uh projects that have been come together to essentially help that pre-ompetitive stack for innovation. Um, one of the other things that we did uh very early on and continue to do is get involved with institutions um that are really deeply invested in global agriculture. So the CGI, how many people here know the CGI? Anyone? So the CGI CGI is a um fairly large consortium

I would say they're in 30 countries. Uh hundreds of scientists who obsess over agriculture knowledges, knowledge bases. They're the ones who preserve the germoplasm of seeds and cultures across continents. I mean, they're really um quite invested with countries, the UN, big foundations like the Gates Foundation and others to manage this kind of global consortium and they've been around for I think over 100 years. So a lot of sort of old knowledge that has come together. A lot of the a extension services, anybody heard about a extension? Yes. So a extension services which come through universities in many countries uh including the United States are advised by the CGI which is a bunch of whole bunch of scientists. So it's a very uh incredible organization that very few people know about

Um and in fact they have something for fish and something for poultry and I mean the real ecosystem of the knowledge base of the world of food. Um so I got involved with them uh a few different governments um who were very interested in sort of building up on top of this and uh one thing led to another and that's where we are today. So that's a little bit about my background. Uh what I'm going to talk about today is just touch on a few things that intersect with Palefire um which is the graphbased new addition um uh you know repository and community that is now part of astack uh but it is one of several others and so I'll talk a little bit about a broader concept that pale fire sort of intersects um and so that's why one of the things that I thought use and I use gamma for generating this today so any gamma fans here you can critique it and and admire it at the same time. Um so you know the world's oldest knowledge systems really in agriculture and food uh really have no memory because we constantly develop no new research. A lot of times people you know research things that already been researched because previous research is not really visible. Most of the research is in PDF documents. scientists generate research for this sometimes for the sake of generating research and just gets out there and then nothing really happens

It's on the shelves. Um and you know what I think we're trying to do particularly with projects that uh Palefire and projects that intersect with Palefire is to really um unlock that knowledge and make it very accessible to innovators who want to then use it to drive change across the world. And let me say that you know before I go into it it's very very important work as somebody mentioned about something like food. I think Alexi, yeah, some of it has have the privilege of having food. We all need it. Not everybody gets it. So having uh easy way to make food more nutritious, cheaper, better and and less uh stressed uh with climate change on one side, labor shortages, geopolitical unrest, you know, you name it, the food ecosystems of the world are on very fragile ground. So this is very important work in my view and that's one of the reasons why we have this community that's come together

Um so essentially um the thesis here is that um agriculture really is the least digitized of all ecosystems. we really have a lot of knowledge that's locked up in at best PDF uh uh uh files and other things but even those are locked up they're not always public in fact I would say most of them are not public um they're not always published research a lot of the work in private enterprises for example very large monolithic uh organizations in the private sector uh big a companies that we know about have tons of research and information that is not public. So none of it has been used to train any LLM. Okay, that is essentially decades of field level observations, all kinds of internal information and it's sort of disconnected spreadsheets, databases, silos that exist in little pockets and the possibilities of being able to leverage that to improve the world's food supply are breathtakingly attractive. So that's the motivation behind it. Um the scale of the pro problem is huge. There's billions of of acres of farmland across the world. Um and you know CGR for example have 15 international centers and 30 different uh locations

Almost none of it is in a graph. So the discoverability of that knowledge is all siloed. Okay. Today which is why we're so excited about Palefire. uh the open source plumbing. This is what my my co project was. So the co project that I had for myself which is right after I started at Axtack was to create a uh a a geoid based or uh essentially a geographically indexed location um uh identifier a UU ID for every field parcel in the world uh and create that as an open layer on which knowledge could be tagged. And so that becomes kind of an underlying framework

The European Union has now adopted that standard for UDR regulations and it's coming soon to a US farm near you. It's US is kind of trailing behind the world the rest of the world. Not a big surprise. Um and of course Pailfire very excited about it because I think this can help with that indicy that that index unlock some of this information. And as I mentioned the UDR is a big forcing function. How many people know about the UDR here? Anybody? The EUDR stands for the European Deforestation Regulation, which means all food imported into Europe will have to define where it came from and will have to certify that it was not from a deforested polygon, land polygon through satellite data and others. So there's a huge amount of work that's happening. Coffee is the first one commodity being uh uh regulated

There's eight more within a year uh with the end of the decade. You know, it'll be a lot more. And it's a royal pain for people who are in the food industry, but if you've ever been to Europe, the food does taste better. So there is something to be said about that. Um so you know the scale of the problem, the years of research, the the you know the previous speaker was you know talking about how these the yield of information you know there's highgrade 10 meter pixel satellite coverage of the entire globe happening now. um and there's lots of information you can glean from that but how do we extract that and make it relevant in real time while we also have all this knowledge base in these big uh documents. So I think that is really where we are trying to sort of create this digital identity around information uh tie it to knowledge graphs and be able to make it easily usable by others. Uh here's a whole bunch of different uh projects

This is a subset. We have 47 projects. Now you can see there's a lot here. These are commits over time last 5 years. Uh this one was just recently added. Uh this is the u uh one of the open agri projects, but palefire is one of them. I think this red one is palefire. Uh it's been going on for a while even though it was just recently committed

Uh but you can see there's a whole bunch of communities that have come together and um open agri for example is an EU funded community out in in the Netherlands. Um they've bunch they've done a bunch of stuff. Ogri OS is an operating system for agriculture management was just donated literally I think two days ago. Um so there's a lot of stuff that's happening. Pancake is something that was geos uh spatial temporally uh indexed database uh built on top of Postgress that allows essentially for you to store information once it's extracted from knowledge graphs we put into it so that you can store it for a graph record. So there's a bunch of stuff that uh has happened around that. Um and you know the opportunity really is to um kind of use the identifier the GUID for the underlying information uh put the the graph knowledge on top of that index with that and create a digital public infrastructure. For example, the country of Honduras has signed up for a DPI layer that CGI and us are working together as a use case

Kenya is number two after that. So we're building these DPI infrastructures for the entire country to be able to see what can we do for unlocking uh agriculture. So I think that's the the work that's happening and um you know I think what we need from the community is uh how do we deal with agentic memory? You know memory is a big problem. Uh not just context management but long-term memory management. How do we deal with that? Um how do we optimize for temporal uh graphs? So not just knowledge graphs but over time how does that knowledge move um I think there's scaling issues and of course the latest and MCP so we can make it modular so that different agents can talk to each other so there's a lot going on uh happy to take questions and I think that's the end of my presentation >> thank you sir if you have a question Raise a hand. Okay. >> Um could you please uh do you have a slide to show visualization of you uh how you make this fields parcels cells >> how you build? >> Uh am I understand right? You make the whole agricultural fields like mapping all of them like in parcel in square kilometers maybe or meters. No, it's in uh the resolution is actually uh uh uh can be down to 3 cm

>> 3 cm. >> Yeah. The resolution of the pixel can be down to 3 cm. >> Do you have visualization of this uh table map? >> Uh I don't have it with me but essentially it's built on top of do people know about S2 libraries S2 H3. So there's existing open-source libraries that companies like Google Earth Engine, Google Maps or Uber use for visualization, underlying open source libraries. So the visualization is and it's hierarchical. So the pixels have other pixels and you know they're sort of >> you can go down to a a very small resolution >> and up to a very big resolution. And we're leveraging that open- source body of knowledge for indexing

So the resolution is really however you want it to be from >> hundreds of kilometers down to 3 cm. I hope I answered the question. I don't have a visualization for you right now, but >> if you can send us a link, we can visualize the folks. Yeah. >> Yeah, it can be visualized. Sorry. >> Yes. Yes, you're welcome

>> Are you all developing an ontology to guide the knowledge graph? It one ontology or multiple ontologies? M very very good question and important question. Uh we at at AGAC are not developing it but there are several communities that are part of our you know overall community several different repositories and and who are so CGI is one of them that has already developed over the last you know few decades detailed ontologies of different types of food systems. uh so we're going to start with that as a starting base but there are many many universities and professors and interest you know re who want to take that further and who want to define that I think ontological descriptors which are very relevant for graph graph-based uh knowledge discovery are absolutely at the heart of of this uh discovery project but we have a starting point I think the communities will once we unlock it our goal is to unlock it so the communities can just run with it and then we'll just have people opening up these ontologies and and may the best ones win kind of thing. >> For me it's something pale fire is going to help with this because as you will see pelar has a lot of machinery to work with ontologists. >> Yes. >> Okay. Last question. >> Thanks sir

>> Uh who uses today uh this system and for what kind of decisions? >> Yeah. So uh there's a lot of different types of users. Most of the users are either uh uh food companies or agriculture companies and they're developers that work there who are essentially using these to create innovations on top products. Uh and the types of decisions depend on the persona. So in that persona the decisions could be how do I irrigate a field using satellite data and weather? How much water should I put on my field? Or the decision could decision could be what should I plant? This is a how do I get soil information and weather information and market information and pricing information and labor information and figure out what crop should I plant here so that I have make some money as a farmer or when should I harvest? I have a tomato. If it's too late to harvest, it's going to basically be uh rotten. If it's too early, it's got too much water content and my processor wants a certain water content, certain bricks amount, when do I harvest? What's my harvest time? I mean, a a head of lettuce loses 12% of value per day. That is its weight, you know, in the on the shelves per day

So, the velocity of these decisions is >> absolutely unthinkably fast. It's very very quick at a globe scale. So I think those kinds of decisions are enabled by by different supply chain actors. I advise three different companies where they are doing supply chain optimization on fresh produce for example. Um you know some a hurricane happens in Missouri. What does that do to my lettuce supply that I'm basically buying? Now I have to flip to go from to Mexico and oh by the way there's just a terrorist attack in Mexico. So now the border is shut down. Oh I can't go to Mexico

what do I do for this is the kind of stuff that happens in the agriculture ecosystem behind the scenes. So when you show up at a grocery store it's all there. So those are the decisions being made by private sector, government sector policy, huge decisions on policy. What should be the discount rate? What should be the you know underwriting? The government of India for example provides all kinds of different um guarantees minimum support price to buy grain at a certain rate from farmers so that farmers have that assurance so they can get insurance. What should be that that rate? That's long-term planning. All that data right now comes from fragmented and unconnected sources. What should be the yield planning commission of Kenya? What should be the yield forecast for Kenya so Kenya can feed its people? All that data is changing dramatically day and day. And as climate change happens, as all the stuff is happening around the world, that volatility is actually growing

So this knowledge base is really really important to unlock the the decisions. I could talk for 2 hours here. There's unbelievable number of decisions. >> All right. Thank you. Let's thank Sume. >> Thank you. Are you going to hang out hang around a little bit? >> I'll be for a little

>> Okay. So, if you have other questions, please ask uh Sume after this. Now, I have to do a little bit of a dance because I am both the recorder and the presenter. So, just stay in place. Let me just double make sure everything is recording properly.