Devreal

Chief Scientist: Yann LeCun, Tapestry Kickoff, Paris

Chief Scientist: Yann LeCun, Tapestry Kickoff, Paris

Recording: Chief Scientist: Yann LeCun, Tapestry Kickoff, Paris

All right. Hello everybody. I'm Alexy from the founding committee and head of AI Alliance and now head of committee at Lexel and here we are at FPT in Paris, France with Yann LeCun who is a chief scientific advisor to the new Tapestry project of the AI Alliance. Please tell us what the Tapestry is, how it came about, and how you became basically scientific lead. So Tapestry project is a old idea I had that I was originally trying to push within Meta, but it didn't go very far. It's the idea that there's going to be a future where every one of us is going to get our entire information diet through AI assistants. And even though the AI assistants are produced by a handful of companies on the west coast of the US or in China, we're in trouble because we we're not going to have enough diversity of sources of information. So what I'm proposing is an open source platform uh which would constitute a foundation model on on top of which to train AI systems and those AI systems would be very high highly diverse

That would provide people with a high diversity of AI assistants with their language, their culture, their value system, their centers of interest, their political biases, whatever they want. And so the best way to do this is to have a coalition of multiple countries or regions or contributors collaborating in a some sort of federation to train a frontier model. And they don't have to exchange data. Uh they can retain sovereignty on their data. The way they exchange are parameter vectors. And uh and the system is trained collaboratively. Everybody will have access to it. Uh and each country or region may use their own local data and their local languages, the local library material, everything

So, I think it's the pretty much the only way that we're going to get access to a wide diversity of AI systems in the future. Right. And so, and basically, you know, you've been, you know, the head of AI at Meta and you led research at Meta and Llama was a big success, right? And so, essentially, it's not enough for the world, right? Because it basically we don't know exactly the data on on which it was trained, right? So, so how um like we've been here for a two-day workshop kicking off Tapestry. So, we've actually heard representatives of multiple countries, from multiple labs, and so there are kind of some of the architectural ideas. How would you summarize kind of the kind of current state of the idea? How should the world go about building these models, scientific and open source AI community? Right. So, first of all, I need I need to make it clear that I'm not at Meta anymore. I left Meta, you know, early 2026. I'm involved in a new company called Amina's, but this effort [clears throat] is separate

It's a it's really a collaborative, you know, non-commercial open open project. So, what I've seen in the past is a very strong desire from a lot of countries around the world um you know countries that are neither the US nor China to have some level of sovereignty in AI. And I see a lot of efforts in countries like Switzerland, Germany, France, UAE, Vietnam, India Japan, Korea efforts to build their own LLM. And and of a waste of resources, right? >> Yes. I mean all of those countries should get together and train a common LLM. I see uh some countries also uh that in the past have built on top of uh on top of Llama. Mhm. Uh but it's not clear where, you know, what role Meta is going to play in the open source world in the future

That's right. Cuz they're kind of, you know, coming, you know, becoming much much more uh much less open. Uh so so I think there is an important role to play for uh an organization like like this one. Uh where this is the kickoff meeting, you know, we hope to rally a lot of talented young people. Uh we want uh we want this to be a a bottom-up uh project where, you know, people who think they can contribute technically can contribute technically without having to go through like, you know, major uh bureaucracy or anything. But I think we're going to get a lot of support from uh governments around the world. >> Right. >> Uh you know, either money or political support or all kinds of stuff for resources in general

>> Right. Um because there is so so much of a demand. Right. You mentioned that, you know, uh so AI Alliance held uh several meetings uh at the AI Action Summits in Paris and New Delhi. And so you mentioned uh uh that basically those governments turned to you for advice. The heads of states were interested in this specifically. Right, so governments are uh understanding that this is important. And so there is demand

There is demand around the world and so this initiative is actually, you know, uh put together to answer the actual demand that people around the world have. That's right. There is a a huge demand and uh a lot of countries realize that their only path to AI sovereignty is through open source and collaboration. Right. Right. >> Right. Uh so so that that gives kind of a a big boost for projects like like like this one. Mhm

Um but I think what's important also is uh and pardon me to get a little technical but uh what we want is is build a software infrastructure so that you know various contributors can train their model locally on their own data preserve the sovereignty on their data they don't have to exchange their data the only information that circulates between the contributors are parameter vectors for models and we would train those models in distributed fashion but in such a way that they eventually arrive at a consensus model that is as good as if it had been trained on the totality of all the all the data accessible to everyone and that's an important aspect because that's probably a way to get those open model to get better than the proprietary models they will have access to more data more diverse data as well right using data that regional data that that the commercial entities don't have access to data So so this is a technical architectural challenge for the model builders so this is not something that yet exists No it's still we still have to figure out we know the techniques will exist we know that they kind of work they are prototypes It's feasible we know it's feasible but but the details have to be worked out Yes you know what I like about this that you are known you know to proclaim all the opinions that all LMS said that new scientists new researchers should not go in them and you your your company is building a world model but here you are advising a very practical today endeavor and you actually said that these things are useful there is a place for them there is a demand for them and you actually propose a way to make them better and make them useful for a world so it's probably it could be a PhD for somebody who will figure out how to do this a whole bunch of PhDs yeah [laughter] no but I mean I've I've never said a language is useless but I mean LMS are clearly very useful we all use them on a daily basis. That's right. Uh for code generation in particular, I mean, people use it for mathematics, but also for all kinds of stuff, right? Right. Uh for access to information, I mean, there's no question they're useful. Right. I mean, pretty much all of computer technology is useful, but pretty much pretty much all of computer technology is not a path towards human human-level intelligence. All right. >> What I've said is LLMs LLMs are not a path to human-level >> to AGI

Right. >> That's right. I don't like the word AGI, but you know, they're not a path towards human-like intelligence. >> Right. Uh but they're so useful. That's right. That's right. That's right

Uh so, uh now, obviously, you know, you're an academic and you're a leader and a startup founder and, you know, and you advise governments. So, you have multiple kind of uh kind of planes where you operate and so, how do and the AI alliance has a lot of brainpower and a lot of executive power. So, the question how do we go around the world and engage people because obviously we need, you know, smart uh PhD students to work in different leading universities on this problem. Uh and maybe at Sorbonne or maybe in one of you and maybe at Stanford or Berkeley or somewhere else. And also, we need, obviously, to align with government labs which already putting resources into the national uh LLMs and convince them that we'd rather all work together instead of them reinventing the wheel and being two steps behind, you know, ChatGPT or SAS or something like this. All right. And and so, and also we need to go, you know, and convince probably corporations in various countries that it's in their interest. So, so, uh how do you think we should kind of proceed uh as a young AI alliance to enlist people to work with us in different ways? So, there's uh a lot of really good reasons that will motivate people to contribute to this, right? Uh so, obviously, Here government level, it's just sovereignty

Yes. Right. Um and not be dependent upon technology from the US and China. There is obvious geopolitical reasons for this. Yes. Uh then there is motivation for uh companies in those various countries to not be again dependent upon a supplier that uh you know, can change their uh you know, their licensing term from one day to the next. Here, you would have access to a frontier open model, completely accessible. Uh you can do whatever you want with it

You can you can fine-tune it. You can you have the entire source code. You have part of the training data. You have the the training code. I mean, you know, it would be much more open than than than current open weight model. >> Yes. So, there is a motivation for uh for companies also to actually kind of rally to a project like this. But, what's most important is for individual contributors to uh actually be motivated by the mission

And the mission uh is basically to protect democracy. Yes. Right? To to ensure that uh people have access to wide variety of sources of information. Yes. And don't get the the information filtered by who you know, whatever someone in Silicon Valley or in Beijing decides. Right. Uh Or anywhere. Or anywhere for that matter, right? So, so there you can, you know, get the the foundation uh model produced by Tapestry and then fine-tune it with your own data, your own biases, your own political opinions if you have them

Mhm. Whatever your your value system. And uh and now you provide an assistant to people who are interested who have the similar interests to yours, right? Yes. Uh and and uh we need a high diversity of uh of assistants of this type, you know, for the same reason you need a high diversity of the press of, you know, newspapers, magazines, and information sources. Yes. Uh >> So, that I mean, the mission of, you know, preserving democracy, I think is and cultural diversity I think is quite motivating for a lot of people and there is really interesting technical problems to solve to get this to work, you know, how how do you uh train a a big uh uh model and an M or or or or whatever it is in a distributed fashion with with data centers spread around the world that you know cannot communicate information in a synchronous manner uh they're each trained on their own you know subset of the data and they still have to basically contribute to a a common model. That is a big technical challenge. And maybe using commodity hardware, right? Because maybe not using latest expensive Nvidia GPUs but maybe using smaller GPUs

That will be a very interesting question how do distributed training like this. Right, if you can do it this way I mean it's possible that you know even with like a few GPUs you can actually make significant contributions. >> Right. Right. Right, right. So uh so but that's you know remains to be uh invented or or fine tuned. There are techniques for this which we do not know if they work at that scale and so uh there's a project for a few PhD students there. Are you planning to advise any PhD students working on this? So this is not my expertise Right

and it's not the topic of my research. There is people here at that meeting and people working with them who are much more expert than me at this. But like I I'd be uh more than happy. Actually kind of wrote a white paper on this idea of distributed training but this was like 12 years ago. Right. Uh and uh um I I think it's a really interesting problem. I I'm uh happy to provide advice uh but I know there are a lot of people who are way more expert at this than I am. Well, the AI Alliance has multiple universities as members, right? Dozens around the world

So if you are a student PhD student watching this, you know, please connect with the Alliance. We'll put some links under this this video how to get in touch and contribute and we can route you to, you know, advisors who might be able to supervise this because this is an important scientific problem for the world. Right. And everything is to you know, still to be made. That's right. Uh we have a GitHub, but it's essentially empty at this point. Yeah. And so, if you have some good ideas, um come in

You can you can play any part on the wall. Uh we don't really have an organization with a technical leader yet. Uh so, you know, any idea you can you can provide, I think it'll be uh we'll be listen to. Thank you very much, Jan. Looking forward to the success of this endeavor under your scientific leadership. Thank you. All right.