Devreal

Understanding World food economy with sa...

Event: Scale by the Bay

Scale By The Bay 2018: Aleksandra Kudriashova, Understanding World food economy with satellite

Recording: Scale By The Bay 2018: Aleksandra Kudriashova, Understanding World food economy with satellite

start with an example of my great friends the company called once world with their project of the regional map for the food production so for many countries they trained the modern classify the commodity crops production for several years it's I guess it's about three years for most region so that's primarily US and Europe so a really cool thing that it's possible to see how much food was produced and this includes we soybean maize cotton sugar cane so what's what's amazing I have a very fast give in there in this slide so then they analyzed the data on the field level so one field is like how many is like 12 to 6200 Hector's which is not very large so they used medium resolution data to see what's happening within the felt so they were talking about the data with a resolution of 10 to 20 to 30 meters not a 100 meters data because because of the crop rotation it's really important to see how much surface covered with this given felt to understand what was growing there was it was a successful was it harvested and did this field bring some input to the food production of the region so from there from the field level back to the regional and national level to assess how much for example potatoes were produced in us last year they they looked at they're very very deep today to the fuel panel and on the left hand there is a chart the chart is called NDVI and this is basically the way how did they analyze what was growing there so NDVI is one of the vegetation index is extracted from the multispectral data satellites all either the surface of earth in red means red green spectral bands and by this information it's possible to extract a metric of NDVI and use this metric to to see what was what was growing there so it's possible to plot this magic from day to day from week to week and see how they crop was developing this year from the previous comparing to previous here how this field is performing comparing to the other field and even do the cool things like crop classification and farming activity detection so India is just an example of the vegetation indexes there's like plenty of them and the reflectance of the green mass and the red and means for advance allow us to see to see the greenness the amount of green green mass over the field which finally the end of the season translates into how much sugars or how much stretch was produced for this given field NDVI is just an example of the of the index there are really plenty of them depending on the type of crop type of temperatures and most of them use different multispectral or very narrow hyper spectral bands and their main input for us from our practice that in order to do a successful machine learning or deploys a neural network with this data so the best way is to use all the synthesis as an input because they're pretty pretty deterministic and also their values of the row bands as their as the features so this was really the most helpful take away from us so there's not one index better than another it just used all over them some of them would help and for our company we realized that a successful project for I understand it what is growing sorry Google doesn't like me to have a successful project to do that crop classification or to do the prediction of how its graceful Fiddler's performance we identified three groups of challenges that a team should address so one of them is stream of the observation data so we cannot go to every feel and see what's going on there but we can to temporal basis assess the vegetation that says the temperature weather and get some statistical difficult information as frequently as possible so the second important thing which is in fact the most important bottleneck which provides us from having a real very high-resolution map of crops is the ground truth data which we use for training and validation because these data should be relevant it should represent this given type of crop in this given geography and for this given type of soil and this data unfortunately is pretty rare and for rural regions it's a challenge to extract it and the third part is the challenges associated with the methods of how these type tasks should be addressed and they includes both the algorithms and the methods and one of the most challenging for us is the data engineering part because we are talking about satellite images which are like a gigabyte per shot like 100 gigabytes per day and with this data there should be certain engineering routines to perform a to process and successfully and today I'll focus more on the data part of the problem and let's start from the observation data these data should be frequent and these data should be relevant so we are talking about weather optical radar imagery and also some of their location-based economical data from demographical from some statistics we identified it was also useful so for the most spectral data the most widely used are the I'll start with a LAN wrapper which is the data from our own constellation and also the data from NASA Landsat 8 and ISA Sentinel 2 which allow from 10 to 30 meters or a resolution and combined they give up to daily revisit frequency so the ones that it has the deepest archive those 2013 and for other applications is possible to look up into back to cemeteries from the Landsat one and all this data is like one way or another is also accessible but this data has a little bit lower resolution so there is the Sandhill to is really amazing it has a little bit smaller archive but it has more spectral bands higher GSD currently is there really industry standards of the data we should be used it is in Madagascar I can send you an address because I process it process it myself so the radar data which is also very important and sent no to provides accelerator data basically shows the waters its main application for agriculture and it's really possible to identify irrigation from non irrigation fields water resources artificial and natural and there homogeneities they are available which also helps to predict how successful is going to be the agriculture activity in this region this data is available at scale in two ways so one is obviously going to nursing lease and a two data sources which has its own challenges because they PI's are not that friendly there are limits bandwidth or the more cool ways to do the Google or Amazon cloud services where they have this data available in a pretty nice already processed products there are the challenges that it covers there since that their food scenes that cover like chance of square thousands of kilometers but overall this data is accessible and it's all open source ready to use obviously there is also a cool bunch of their commercial data and all these companies including the company I'm happy to be part of provided data with high frequency high resolution and has more dedicated Sparkle bands but in fact this data should be used only when it's really needed because for vast majority of applications lands and sent no is enough until you really care about some very narrow changes in certain bands so for example for cotton there should be a really very narrow spectrum of red band which shows their phonological phases and how the cotton is performing though the cotton is a large market itself or the data should be done more frequently because the multispectral data for there for advance the main challenge is the clouds because the agricultural activity active regions they are often cloudy and for example if you have a daily observations for a Four Seasons for high chance that half of the symmetry will be factored with the clouds and there is no use in it so this data is available and it should be only used when there's a real need in it yeah so our own data now company we build custom satellite missions and put different types of pillows and small sets and one of this type of the payload is the multispectral camera we use cameras dedicated to see the vegetation so these cameras who either earth in red and infrared spectral bands dextra their drone features and these data is consistent with lanceton sentinel so it's possible to build a continuous time series specifically to do the deal monitoring regarding the where the data source is there are two ways of doing it one is going to nor directly they have an API they have data says they have a historical data for up to 20 years back in history and the main problem is this data is pretty challenging to extract alternatively there are commercial API is like this read like weather on the ground error state and where the decision technologist we they provide a date and more agricultural analytics friendly way but it's pretty pricey and if it's not like a real regional analysis probably the state probably then nor should be used the last type of data we identified really you really helpful is their demand demographic statistical data information of how much food was produced how much food was sold consumed in previous years all these types of indexes these data sources provide large API is pretty nice and internally we identify that it's faster to deploy them and try to see some correlation between features then thinking that hmm do we need this or not so to really sometimes it's helpful and the second part of the power of our equation of our process is the ground truth data is how we put all the sub serrations and how we what do we use for training the model what do we use to identify that these type of features look like worn or this type of features look like a healthy cotton field so there are two type - two types of this of this data one is we need a lot of almost relevant data for training and the second part that we need a lot like specific local data for this region for this type of crop for this type of seeds for this soil to do their validation so for training I believe the vast majority of teams who do the crop classification they use these data set it is USDA it is amazing it has fifteen years of history and it has six years of high resolution history with the resolution of 30 meters in one pixels they do the crop classification at pretty high accuracy so they use D does the semi-automatic with like 85 to 90 percent of accuracy and this data's pretty good training data source so all the training should be done with these data and then the change should be done on the local date of given on the region you're looking at so yeah as I said it's a very high accuracy 30 meters resolution and it also has an API and the last set of challenges is how to use this data and how to how to do the engineering part of the problem so the largest challenge for us was in parallelization and processing all this data how to divide into chunks how to collect everything together and so sorry can I just go go there so in the semen that the deities processed on the personal computer take a crazy amount of years so these types of tasks for us were successfully deployed in the micro services implementation for Amazon like lambda functions and it turned out that given their modern limit of 15 minutes it's really possible to do their reading to do their processing to do calculation of index and use them as features all in 15 minutes for a given window of one by one kilometer and then assembly this one by one kilometer chunks into their whole global map was the way to go solution back to these slides so I collected the links to do the data sets and links to their both UI and programmatic methods of how to most effectively process their satellite it because the main challenge with it geometrically geometrically corrected and geometrical position so it's not only like just looking at a picture but it's also looking at the picture and understanding where it is and alignment of different data is a first just part of the problem is how to position it in the on the ground and collect different types of data within this ground so I'm trying to correct and my github links to to these types of resources so yeah love the functions was really saving for us and also I believe that Google for Google Cloud functions do pretty much the same thing as like reading the time series of images all together and and using them for further analysis so it allows us to paralyze to up to forty thousand of data calls with a pretty convenient cost model and it was possible to do them changes updates on the blog the most important part for us was traceability because before we moved to the server list for the instance type of implementation for us was really really nightmare how to how to manage this so this basically all I plan to say today and feel free to check out my github account they tried to pull different data sources with the links which may be interesting for the agricultural monitoring our different type of mapping we were trying to detect poppy seeds in Afghanistan but the problem is that the puppet seeds are planted together with cotton not together with wheat and they have exactly the same phonological phases so it starts in the same part of the season and it grows like exactly the same time so from the satellite pictures possible to classify so with trying hard but without the hyperspectral dated simply impossible so a really interesting so for that given application we use the high resolution image but it's only applicable when in the very end of the season when it's really visible what's going on there but in most cases we do they USDA dataset for validation though it's also somehow classified but they use of like a part of semi automatic classification for us the crop classification problem was really more rd because our for our team what we do we identify what kind of pillows to put our satellites and it turned out that for certain types of spectral bands we see they're more applicable sony infrared red age for specifically for like so range for cotton a certain range for wheat and certain range for corn and soybean so instead of having like a wide range imagery wishes has the same good pictures like Google Maps instead of that we do hyper spectrum so I give you answers as I said so if we would put in the production of the crop classification model and if we have to do it for the whole globe yes then we would need a GPU but specifically for our type of applications when we tell stories when we try different bands when our goal is not only provided map but also to make sure that our data is good for providing this map so for us till the price to value proportion of lambda functions was more relevant so yes in general terms they are more expensive but if you needs only like I don't know one melona for vacations then it's cheaper then they're having a GPU yes we did unfortunately I don't have like a visible results right now but it's a pretty straightforward problem once you utilize this data so for example USDA it has so for them they're the life cycle of that crop looks like a like waterfall so some of the crop is seeded some of the from this some of the crop is growing some of its harvested towards salt money get and for all of these types of the pipeline they have their monetary values so we have a very much here because we know how much food was seeded on how much did it grow and then we can you can predict the harvest and we can predict the supply and demand so this is basically only thing which happened and in this part of the of the analysis there are so many other things which are we don't have enough data to predict so we see it but we wouldn't have more more available input than anyone else so yeah say it's possible to see it so we have enough of the input the data but like god knows what happens here right any other questions thank you [Applause]