Devreal

From Bots to Brains: The Evolution of Web Scraping

Event: AI by the Bay

From Bots to Brains: The Evolution of Web Scraping | Lorenzo Padoan, AI By the Bay25

Recording: From Bots to Brains: The Evolution of Web Scraping | Lorenzo Padoan, AI By the Bay25

Hi everyone. Uh today I'm going to give a talk about a not sexy topic and usually extremely dry that is scraping. So I want to convey some coolness on this industry and I hope to do that. So I'm Lorenzo. I'm from Italy and I use it to work to the Italian palunteer. So I needed to scrape tons of data and do strange stuff. So um let's jump in. So imagine you you are in the 90s web 1.0 And basically uh you can scrape with a simple pearl script using some exoteric regax

And one day this guy called it Matthew Gray wanted to understand how big is currently worldwide web. And he was the guy that first discovered and first implemented the worldwide wonder the first crawler that we know. After that right now we are living in the AI gold rush. At the time there was the search engines gold rush in the 90s. To build a search engine you need an index and to do to have an index you need to to crawl websites and most of these companies were very po they coded poorly these crawlers. And so basically they overloaded the network of the websites. So they they blow up website here and there around the world. And for this reason this Martin Coer this this this guy saw these websites blow up and said we need to do something he proposed for the first time robotics

So what is robotics? You attach to your website this this this document and you say please scroller don't do this stuff don't do that instant cat is not mandatory but uh he was the guy that proposed that after that how many of you knows what what is back don't say what it is back is the name before Google so it was the research project at Stanford that after that uh became Google So they uh before to implement the page rank uh algorithm they need to crawl all these websites here and there and so for days they overloaded the Stanford network in order to crawl the websites all around the world and probably probably they didn't give a of robotics probably. So imagine now in the in the 2000s h web 2.0 to CSS interaction with the websites and also the internet industry started to grow and like big players with like eBay and there was an entire industry that was hungry of data like uh like beer's edge what they did with beer's edge they were an aggregator of uh eBay so basically they was a a price comparison of of the things inside eBay and one day eBay says no no guys you don't go you I I I don't want you do that. I'm going to sue you. What the court decided is okay. Okay, B, if you want, you can block them. What they did? They blocked over one 1,000 IP from the scraping machines of Beer's Edge. What did Beer's Edge for the first time? They used proxies in order to not be detected from them. So, it is the first way the first time that we know that proxies are used

And also in this in this period there was uh the development of these new technologies like beautiful soup in order to have this CSS selector in order to extract data from HTML. And also there was the era of meshup for example there was this guy a software engineer from Germany Pauler that said I need to find a house but Craigslist is a shitty way to find houses. So all this listing and so what he did was able he was able to scrape Craigslist put it together with Google maps at the time there wasn't any public available Google API so he was able to do that and he built housing maps so right now we are in the 2010s so the there is a new player inside the the the web JavaScript so with JavaScript you need to render the web page in order to collect all these kinds of data data and and also you need a a new set of tools to do that like selenium playright and also new uh uh libraries like scraping in order to do to build crawling frameworks at at this time the scraping industry started to be a an interesting industry at the same time also the security uh let's put it way this way the security industry that was like cloud that unfortunately today is down I'm Not so it's okay. Okay. Uh so uh and after that right now we are basically in in this era we are we are in the uh cut and mouse game. So there is this security uh company like Cloudflare that started to implement uh barriers like captures rate limits and also the attackers the industry that is eager of data with stealth browsers and also there was an important uh legal issue IQ versus linkering. linking suede HQ because it was scraping link in the court decided you can do that because if the data is public and it's not behind a paying wall you can do that you can scrape that. So it was for the scraping industry a pivotal moment that said okay you are not more in a gray area you can do this kind of things and also in these 2010s there was the development of these new technologies like computer vision and LP in order to perform better scraping but after that something happens AI so with AI you can use LLMs like like with scraper graph what we do with scrape graph we are there is right now a shift a shifting of paradigm from classic rules Python rules with CSS selector to using LMS

Why that? Why make sense? Because if you imagine a web page is like a content of semantic and the the the best semantic engines right now are LLMs because they can understand where are the data. They don't give a if they they are on top right and tomorrow they they change the order. they are able to scrape every time the data. So that way you can have resilient the ya changes and also less maintenance. So it it is what's going on right now. We are shifting from classic rules, programmatic rules to LLMs to extract data. And what's going on in the future? What we are working on? We are building this API that is this layer. But the in the future we need a new infrastructure like the guys that showed before

We need to some kinds of automation. Right now we are living in another subset of a gold rush. there is this gold rush of of the next layer of uh like search engines in the 90s. Right now there is this search layer for uh for um for agents and to do that right now there are two main approaches. The first one is using only brows automation like we have seen before but there is a problem you can't scrape at scale with that and the other problem is that there are another use case that using only to extract data also that is not working because sometimes you need to interact a very small interaction inside the websites and so what I'm thinking is going to we are going to have in the future is a a hybrid approach hybrid approach where we are able to uh build a a huge set of o of a swarm of agents that are able to collect data in hours, gigabytes of data for hedge funds for example for also small startup and uh yeah this is what I I believe and so what I we are building with scraper graph so if you want to check our website we are we have also open source project scrapraph bonus from open source project right now we have 21,000 stars on github where you can connect your LMS your local LM or your um cloud provide CL cloud code or whatever in order to scrape website and also if you want to scrape a scale million 10 million of website we provide an API so [snorts] we can able we can also afford this amount of data. Uh so thank you uh very much. Another thing we are hosting an araton at uh SPG. So if you want to if you want to uh build agents for production that can use a scrape graph and also use some kinds of evaluations on that this is the link to access to the hackathon

So if you have any question also technical feel free to ask. So thank you very much. [applause] >> Thank you Lorenzo. I think you came pretty close to another lightning talk. >> Yeah. >> You finished ahead of time. >> Yeah. >> So, with that, our sessions are complete, but there's one more session that's happening between 5 and 6 in the general auditorium, and the happy hour is downstairs on the first floor

Thank you so much for coming. Thank you. I really appreciate your patience. You're one of the best audiences. Um, look forward to seeing you tomorrow. [applause]