Devreal

Bay Area AI: Hugging Face By the Bay

Bay Area AI: Hugging Face By the Bay

Recording: Bay Area AI: Hugging Face By the Bay

to be able to talk to you guys I wanted to start with a little bit of background about hugging face we started the company with more than four years ago now we based in blue clean and in Paris usually obviously these days everyone is remote so we have people all over the globe these days I'm personally in Connecticut to enjoy some more nature some more space than in my small Putin apartments we've been lucky to raise a bit more than twenty million dollars so far from really good investors some of them in New York and a lot of them actually on the west coast like Rodney Conway with a phone called a capital K V insurance when you were still on the on the west coast at the Warriors and the hats Richard searcher was like the chief scientist at Salesforce and many other angels LRE based on on the west coast we most popular for our open source library which is called hugging face transformers that I'll show it to you after we've been lucky to have been mentioned I think now it's more than a thousand research papers and have more than a thousand companies using our open source without do is that I would start by giving you a little bit of an overview of NLP if you've not been following it for the past few years and then I go in more details into what we're offering a taking face what is our opens was doing and how you can leverage it at your company and I want to start with camp like a statement that is bold or not depending on how much time is spent on and up to this past few months I strongly believe that today NLP is the most important field of machine learning my backgrounds before I studied her game face is in computer vision so I've had the chance to see the technology cycle of a lot of different technologies when I look at vision I look at time series I look at self-driving cars they're like it's been really kept like a strong movement for the development of an app in the past few years I try to explain kind of why and RP is becoming so big and so important in in the machining world now first thing like I wanted to define it's for people who are not familiar with it NLP is defined as a branch of artificial intelligence concerns with understanding natural language right so to make it simple you have computer vision for images or video subscribing car you have time series NLP is really the branch that has to do with any sort of text right and so usually new in club to that but only written text but also voice because most of the time to process voice what you do is you to speech to text a new process the text and then text to speech to reply with voice for example on Syria or LexA all the things like that if you can take a camp like a step back to understand the importance of NLP you realize that most of your day is spent in natural language is spent context obviously the way I'm talking to you now is true voice it's through text but then hopefully when you'll be able to hang out with your friends in person you're going to chat in text the emails that you're sending that you receive is in text the use that you read book that you read all of that is text even the activities which you wouldn't consider as mainly text wouldn't be understandable without text like it's the case for example for movies for T V shows if you remove the dialogue from it then it doesn't make any sense now if you think about companies you also realize that every single company is built around natural language and is built around text from customer support cells external communication internal communication all of that if you think is of it is is text if you think of the largest departments in most of the companies the departments where you have a lot of humans like customer support cells they are natural language processing department so I do take as input the text and they output any form of text so if you think of it as natural language is the interface of humans it is anywhere and everywhere and the interesting thing here is that if you look at what happened in terms of NLP science and the way technology can understand text have been a really big paradigm switch in the next in the last few years some of it is due to care fact longer technology trends obviously the ability to have more and more compute the ability to have access to large open text datasets which is basically the web and the last missing piece to the puzzle here is transferred by which I'll be introduced for ennopp te and studied to work with NLP around 3 years ago with the foundational paper attention is all you need and basically the combination of these three trends led to deep learning for an LP to finally work right finally now you manage to have NLP models that can be pre trained on very large data sets and then fine-tuned on any specific task what it led to is pretty insane progress in a lot of different end up data right so for example here I'm showing you the progress in text class education right so the task of taking it texts and being able to classify it these are results from the glue benchmark which is an aggregation of different evaluation metrics and you see that in just a matter of two years it went from 69 percent accuracy to 88 and what's even more interesting than the figures themselves is that if you compare it to a human baseline meaning that if you ask humans to perform the same tasks you see that we'd not pass the human baseline obviously I'm not saying that we are in a world where we can be human level and not be on every single task it's specific to this benchmark this evaluation and this specific context but at least the fact that we managed to reach the human baseline on this evaluation is really promising and exciting to me it doesn't just happen in text education and it happened in every single ineptly task so here I'm sure you showing you the progress for question answering question answering is the task where you give the model the text and then you formulate a question in natural language about this text and the model will show you where the answer is in in the document and if you see here in a matter of years we went from below 70 to now be over 90 and the interesting thing for this specific benchmark science benchmark that is called squad is that at some point the authors of the evaluation thought that maybe their variation wasn't good enough right that somehow the model could trump the evaluation and so they should work on more difficult benchmark for it to be more useful so they worked on squads 2.0 and what happened is that the exact same thing happens at the beginning it was a bit hard as you can see December 17 to do an 18 but really fast models were introduced that drastically improved the performance and the accuracy on these evaluations and we went back to 90 percent accuracy on this task of question-answering what it means is that basically thanks to the progress of transfer learning for an LP we're moving from a world were put it simply only humans could understand natural language as it was the case before which is why customer support cells communication are very human intensive departments that are using not so much technology at the moment we're moving from this world where only humans could understand natural language to a world work finally technology algorithms models are able to classify understand general generates natural language at some times as well as humans do what's even more fascinating on the topic is that pretty fast it went from science to prediction for multiple reasons one of the reasons being that most of the complexity of the science is hidden behind the weights it went in a matter of months from state-of-the-art to people applying it in production for real companies cases and here I'm going to give you a couple of examples of companies or they using transformers in production in real use cases today a very popular way to use these models is text classification as I talked about before you take text you classify it based on your own data so for example Monza which is a bank based in London I've been using transformers to classify the customer support emails right so they receive an email they basically train the model on hundred examples of email switch priority high and 100 examples of emails with priority low and then the model is able to predict the priority of a new incoming message and then they're able to reach it to the right team and the relevant they receive for example and you know I'm very unhappy about my credit card the model is going to be able to classify the priority the sentiments and for example the Department you can do that for any sort of classification right we can classify based on the sentiments it can classify based on the top tip it can classify based on any any sort of like topologies that you're interested in there was a really interesting use case is information extraction right so you take any form of text it can be here for this example for check it's like a homework or online classes could be news article can be email it can be comments it can be tweets community any form of text and thanks to these models you are able to extract information right so for example you're able to extract that it's talking about Washington as a person rather obviously Xander location extract dates or extract topics but you could also extract location you could extract companies any form of information to basically go from unstructured data right structured text to make it structured and create for example tags create categorization act on it use that in templates to follow up or things like that another really important task is question answering that is used for example by being they have for example the first page of results then they're going to apply this question answering tasks to give you an answer to your question in natural language right so if you ask what is the color of Kim Kardashian's hair then it's going to identify in the documents where the answer to your question is another one is text generation or text summarization so for example you've experienced that in Gmail right when you're typing an email and it's going to give you a recommendation of what to continue with with subbhu to complete it's using transform and models now the one is for example to go from a small piece of text to a longer piece of text right so Bloomberg for example can use Transformers to when the stock price of value of stock changes to generate a small paragraph about that can be used the other way around meaning that let's say you have a long news article using some models some transformer models you'll be able to create a summary of these these documents and last one - really the most camp-like mainstream and most like well known to the general public is conversational I write serial ixa and here in a good example is square that is using transformer models to power their customer support chatbots try to give the ability to customers to ask questions and get answers 24/7 right away and what's really interesting is that I started to talk about science and then use cases and this really kept like this very positive cycle where the more science progresses the more use cases in production you see and the more use cases are proving their value the more companies are investing in science and improving it so there's very much like this positive cycle that leads to exponential improvement and progress for the field of NLP and now after that I figured I would give you camp like a little bit of an overview of what we've been doing a tagging face we've been lucky to power a big chunk of this paradigm switch and this progress of NLP with our open source so if you go to github slash hugging face you will see all our most popular repositories most popular is obviously transformers that have had almost 30,000 github stores 7,000 Forks more than 400 contributors now that is used by more than 2,000 repositories now I'll give you calf like a deeper overview of it a bit later some other popular repositories from us or hugging faced organizers which is fast tokenizer tokenizer is what you use to basically go from text to what these models are able to process so tokenizer is a fast library to do that way faster than you would before and NLP is the most recent of our releases where you have a repository of datasets and evaluation matrix for NLP researchers and now I'm going to give you a camp like a tour of how you can use our technology and again face transformers which is now the most popular and ugly library out there so you can go from hugging face that seal and here you'll see that you have a search box for models for organizations here you can basically choose which framework you want to use by towards tensorflow you can choose the tasks that you want to use and obviously the language is that that you're interested in so maybe after we can take a model that is pretty popular maybe we can take Roberta large then on the Moodle page maybe some of an idea of the of the usage for the model and you'll have a snippet to use it directly into into transformers in addition to tools we create become because we became the standard for an RP there are a lot of ecosystem tools that come with it one of them is experts for example which is a explain ability tool to let you understand a little better how your new network behaves and kind of like take a look under the hood so now once you've kind of like identify identified one of the models that you're interested in one of the tasks that you're interested in I'm going to give you a sneak peek at the new feature that we're going to introduce in a few days it's not too big yet but that we believe is going to be really really cool for people to use our library so it's an inference API it's basically once you've identified models that you're interested in for the task you're interested in then we will provide you a demo and an endpoint to perform this and upitis right so for example we'll take text generation tasks right which is what you have in Gmail to basically autocomplete your text so here you'll see that all the models that we have in the library is available we have more than 2000 models now in the library and was in sturdy new models or added everyday to the library so by using it you get access to basically all the best of the NFP today so for example we'll take the model that has been very popular which is GPT - and here we'll take a digital version of GPT - if you're not familiar with distillation it's a process to reduce the size and so to make the models faster to run in in production right so here it's a text generation task so what it does is that it's going to complete your text right so if you say for example my name is Jack and I'm an engineer then you compute right and then it's going to suggest some auto completion completion right so here it completed and also the owner of the project and then it continues after that right you could try that for any of the models in our model hub let's try a different task from our model hub for example we can try translation right here we'll use for example this T 5 T 5 model right and here what it does all right it's still in beta so sometimes you still still see some problems so let me restore it here so I'm gonna translation again maybe take a smaller model all right I'll take 35 small if I do that right not all right it looks like there might be a problem with this specific model let me try again all right I'm gonna try to position so when the model is not in loaded in the hub already then it's going to automatically load it so let's now the model is loading we'll wait for two seconds and maybe in the meantime I'll take like other tasks so another tasks that is pre popular is text classification so here for example I'm gonna take the distal birds so birth is a another very popular model from our hub and if I - for example very happy today as I think this one is an emulsion classification model then obviously I'm very happy today you see that the label emotion negative is not probable and the positive one is very very probable then you have also tasks like token classification for example to find locations so if I pick example like this one friends equally in parish and for example here it dent if I'd an entity which is location with with friends and you can do that for every single model in our model hub so you have more than 2000 models that you can run the inference for them and the beauty of the fact that this is based on open-source is that you can start by using the pre-trained models or the fine-tuned models that other people and the community shared but if you want to go deeper if you want to train on your own data on your own data set for your specific task then you can obviously use our open-source train your own model and then add it to the hub and automatically you'll be able to run the inference for it so it gives you camp like a little bit of an overview of like how you can use our technology to do NLP in production there are many many different use cases as I said more than a thousand companies are using it in production today I'm sure if you think about your company you will find some some use cases so I that works for everyone now we can switch to camp like a more informal Q&A and maybe I'll start by taking a couple of questions I leave my screen share so that I can use it to to show some things so feel free to ask your questions in the chat I'll take the first question from modes which was how were you able to train so many models the GPU cost must have been sky-high so all our models are community based right so it's a whole community of NLP that I've been training models and then sharing it through our hub so obviously we train some models ourselves we did some models ourselves like this Dilbert for example or prune Birds which is a new model that we've released a few weeks ago but most of the models that you're seeing here in the community hub have been trained by other people other companies and by the community so for example you can see we organization page where you see some of the organizations that have trained their own models and share their own models so really grateful for them that's also what makes the NLP community grades which is that everyone collaborate sivan open sources everyone shares so for example here you have like Facebook that I've shared a lot of models you have Microsoft that I've shared a lot of models Google I shared a lot of models and enough like other companies have to in the same way if you train you on model we really encourage you to share it with the community rather than just keeping it for for yourself at working face we really believe that everyone has to come together to help push the field forward it's not going to be one company that is going to solve the field of NLP everyone will contribute to it so we really encourage this dynamic of collaboration of open source open science for everyone to come together and help accelerate the progress of the field all right I'm gonna take another question so Ahmed's is asking what labels is the classification based on it can be any any sort of labels really if you start from the pre train models or the models that are in the community hub you usually have the most popular types of classification right so like sentiments for example classification but then the same way Moodle is training a model on their own classification label in their case customer support priority right priority high priority low if you have a few hundred examples for a different classification then you can fine tune your own model really easily and have your own labels depending on your specific use case does it answer your question comment it's not feel free to add it to to the chat alright can you please share code get link steps for current example looks great so the best way to find that is obviously to go on github so on guitar then you go to our repository transformers right here you'll see examples and based on your use case you'll see a list of different examples so for example if you want text classification then you can see the example here in tensorflow and in PI torch and you'll even have the collab to replicate replicate the results so yeah the URL is hugging jitterbugging face - transformer / 3 masters examples so that's that's where you'll find all the examples alright Francisco is asking I noticed you support GF 2 and PI 2 watch is there feature parity between the two in hugging face and is one lagging behind the other so we started with by torch and added TF to support later so there are still some aspects where we need to catch up with tf2 but we aiming at having Q like a hundred percent feature parity very very soon so you really have the same capabilities in both frameworks obviously we're taking advantage of some aspects where one framework is stronger than the other there are still some specificities so we're taking advantage of that to make it easier for you and more powerful to you depending on the framework that you prefer but we aiming at having hundred person feature priority between between the two all right Greg is asking how do you and how do you do an LP inference getting and Eddings get results at scale since each request takes significant amount of time do you have any specific open source frameworks to manage the requests so it depends a lot on your use case obviously it depends a lot on your constraints right some companies will have more need for low latency others will have more flexibility on that some will want to optimize for inference costs others won't so it's hard to reply in a in a general way you have more and more resources online of companies who did it so for example roblox a couple of weeks ago shared a very nice blog post that I'm gonna paste on the chat now to explain how the scale to serve billion daily requests on on CPUs so you definitely have more and more companies that are running these models in production at scale we can help with some of its we helping a bunch of companies on that but it's it's very much still a case-by-case basis alright that was a question from Kevin yes could you give an example of NLP and one of these models in real time applications so I gave I gave a couple in in this slide basically these are all like real life examples right so for example if you go to being right what's the color of Kim Kardashian's hair here this is this this first paragraph this box is powered by transformer models so that's kept like a real-life examples that that you can see here mom zoo is using it for customer support text classification check for information extraction gmail is using it for autocomplete Square is using it for conversational AI these are all real life today's examples of use of transformer models in the context of a company Curtis asked you mentioned with which human legal success rates in text classification but what areas of NLP do you think will be harder to reach human level performance in the near future I believe the hardest NLP task today is open domain conversational AI and we know it 3-1 because we started the company working on this very topic because we were excited by it and we wanted to start with the hardest one of our penalty is basically the ability to have a conversation conversation with an agent in an open domain sitting like on a little kick not on a very specific topic it's really hard to do because it's a combination of a lot of different adopted tasks you need to be able to extract meaning you need to be able to extract information from the message from the user need to be able to classify with emotion you need to be able to understand like sarcasm for example and you need to be able to reply accurately to so you need to be good at text generation it's a really hard task open domain conversational AI that's probably going to be the last one to be solved at least in my mind our Julian is asking what techniques do you use to ensure low latency while evaluating and model response for example for text generation so I applied a little bit to this question before it were a little bit depends on your you case it will depend on your task you will depend on your training set and what do you brand in France for for example if you're doing text generation it's not the same thing to generate text on just a couple of characters and to run it on a long news news article so it's going going to be different depending on your use case really so Dan is asking can we fine-tune these models on the TP use absolutely you can we support TPU training we've been working in a really close fashion with the TPU team and we want to keep investing in in the topic so yeah there are some I think examples if I go back to what I was showing you before you have some some help especially in the trainer to train with with GPUs all right does hugging face focus more and NLP research or on democratizing and okay that's a good question we believe when you can't do one without the other anymore we believe there's no point in separating science from engineering and science from production anymore for resistor picking so we're doing both at the same time we both invest really heavily on science and trying to best build the best science for a nappy but at the same time making it as widely usable as possible so really blend the two we think it doesn't make sense to split the two you should be to be doing research with an engineering minds and with like thinking about how to democratize NLP at the same time when you want democratize and up key you have to understand the science and be able to move the science forward and it creates is very tough like a positive cycle again where the better you get into science the more companies can use it in introduction the more problems it solves for companies and and the more problems you solve for them the more they can invest into your science so yeah we want to be part of this cycle by doing both and if you look at the background of most of the team members at hugging face they have an interest both on the research side but also on the engineering side to democratize it so Dan is asking our most models available in tf2 yeah your steel models are only available in PI to watch but progressively we want to make every single model available both in PI to watch and tf2 and even further than that you see in the library we want to create more bridges I'm not on bigger huh yeah no sorry I think someone is uh is on I'll continue so even more than making models available in two frameworks we won't create more and more bridges with between the two because we realize that sometimes teams need to move from one to another even in the life cycle of building a new feature and building a new and ugly capability they want sometimes to be able to have the research team work on a model on the modeling side of things in pi torch and then for some reason maybe the production team is using more tensor school because using some production tools from from tensorflow so they want to be able to camp like do some modeling in pi torch and then do the production in tensorflow so even in the same project we want to encourage people to be able to go from tf2 to PI torch in a very smooth way alright another question from Raul openly I recently started providing the API do you see other companies following that or stop providing the models on platform like hugging face that's a good question obviously we have a very different approach than open the eye we're really investing on open source on community on being collaborative with with everyone openly I has done a fantastic job in in modeling with obesity tt-to and the new version of GPT too and we're happy because it validates the fact that NLP has become really important if you remember from a few years ago and you know Panera is still doing other things like robotics reinforcement learning computer vision but the fact that they would is a first monetization feature on something like NLP proves that there's very great potential for its we happy that they help pushing the field forwards they help kept like showing that companies can use an RP in production today you know not not tomorrow but really really today so we really happy about it we think it's it contributes to moving the field of NLP to to the mainstream Timothy is asking would you recommend using Apache arrow based and RP library to replace the F record writer based dataset preparation that's a good question unfortunately as a CEO of the company I can sometimes not go as deep as I would like to on specific topics I would recommend you to ask that on on github or maybe someone from the again faced team if they're listening now can jump in but I wouldn't be the right person to answer this question Alex is asking thank you for the great talk do you think there is a risk with reuse of the base language models which are it's so expensive to build without truly knowing that content totally yeah we think it's important to make these models not as black box so that's why for example we're encouraging if you go on the modal pages we're encouraging people that's a bad example because it's no card but we're encouraging people to add modal cards right which are like a way for people to explain what dataset they're using what are the potential biases what will work what will work to give more explanations and give people the tools to understand what is going to work and what won't work what kind of like yeah biases they need to be careful of so we want people to share that one more and that's one of the reasons why we've worked on adding datasets and matrix to the library so you see here you see like a list of all the datasets that can be used and then we even have a viewer so you can take a closer look at at the data sets you can change change like the keys it's a really good way to explore the datasets understand what the are about what's in there to potentially identify limitations in the identified biases and make sure it's not a problem for production setting oswald is asking what are your thoughts regarding gp3 and beyond so I answered a little bit before I think it's a really great thing I think it pushes the field forwards it makes it more mainstream than it is so I hope companies will keep investing on on these models to make end up even even better than it is today hamed is asking can you give a high-level description of the steps in order to leverage these pre trains models for domain-specific applications for example learning embeddings for particular domain or the mono example you gave also what do you do to work with languages that don't have a lot of pre train models so good resources on that is our blog so if you go you go here for example here you have a tutorial about how to train a new language model from scratch using transformers and tokenizer x' so it will give you like a very good overview on how how to how to do that and here it's to pre train on Esperanto but obviously can use that for for everything else and then in the examples then also you can you have a lot of things you can also go into the notebooks here where you have not only the hugging face notebooks but also community notebooks so that's what I was looking for before for example 25 on CPUs right so here the tutorial to to do that but also you have a lot of like different ones so feel free to check check them out they'll help you tremendously and obviously we're an open-source company so if you're running into any issue just open an issue on github submit a pull request if you want to make improvements we really appreciate that and we'll be really good at the replying really fast so Kennish Kanishka is asking what types of topics again face skin discover can it find if an entity is related to medical terms yeah depending on your training you know you can we have some companies who which pre-trained or fine tuned on Mindi codomain - quite quite successfully and then the type of photo picks it depends again on your use case a good way to test that also is to use one of our demo that I have here which is a zero shot demo so it gives you an idea of what it's capable of doing without any training right so for example here I have some sort of like a an article here about like Jupiter and then you specify some possible topic and then it's going to do the classification for you right so I can do a live example for you so I'll go to CNN right I'll take you know maybe I'll take that way I'll take this this article alright I'll copy the text here alright then I'll put here here take the custom one alright I'll paste it here and then I can specify a couple of like topics that I want classified right so here maybe I can do I can do entertainment politics health full rain first right and then I'll run it you and you see here that it classified with health and politics which counts like sound sounds quite right so you can find this I'll paste it here you can find this demo there to try to see how it does in terms of like topics classification in zero shot setting meaning like wizard in the training really out of the box alright so continue Greg is asking do you think the NLP models are moving to be here and inaccessible for an enthusiast in a nappy because for even fine-tuning cell Albert's there is a need for GPU TPU do you think there is more research required for less resource intensive methods for the same so that it is easy for anybody to get into an object I totally agree with the fact that we need to do more research on how to make these models more accessible we investing really heavily on that so in our model hub you'll see all the distilled version or and the print version of the models that we did ourselves to make them more accessible I think it's getting more and more accessible actually to NLP practitioners and beginners by the very nature of the technology which is transfer learning meaning that you can start using these models and fine-tune these models way more efficiently than you used to do training from scratch before right because you can transfer learn from the pre-training to the fine-tuning or even ultimately probably to the inference so we do think that it's actually getting easier and easier for newcomers in the field to get accurate and up here than than before the whole industry has to work on making these models more I creates more more accessible but they're getting more accessible faster than any other technology before if you think of any other technology when it's released as a research paper for example it usually takes like I don't know at least three four years before you can use in production now it's with Transformers specifically with our library it's really a matter of days after the research paper is published that companies are starting to use it in production which is pretty impressive and unprecedented in the grand scheme of technology before you need to wait really like a few years before being able to take advantage of the science progress now it's just a matter of days which is really impressive I think all right I got another question from Hammond for zhuzh interested in entrepreneurship in AI can you provide you input on viable AI business models and how they can compete very well financed industry giants it's a good question we haven't figured that out at working face yet we're not generating in your revenue right now it's not our main focus yet I believe that you'll be able to create a lot of value in AI so if you're getting usage if you're getting interest for what you're building ultimately you'll be able to build a viable business model I don't believe in the theory that the bigger players because they have arguably a bigger that I said that the only one who can do cool stuff on the topic I think you have different strengths as a start-up I think we kind of prove that by building the most popular open source and open library whereas like obviously a lot of the technology giants wanted to build that and didn't manage to beyond that I think as an entrepreneur and as a start-up you still have strengths on the topic of AI compared to larger companies that are moving slower almost by definition all right Stephen is asking can you remove CNN and see how it classifies alright I remove CNN here and I run that again all right I changed a little bit not not to too much right still 82% for health and 26% for politics and Stephen also asked and if you remove Texas Dixon replaced with Armenian I'll try to do that right Texas Armenian right there Nasser Texas all right here we go let's classify it like that all right yeah it adds a little bit of Foreign Affair so it's pretty interesting spool because the data set is a from the US right so if you add on mania then it considers themself like Foreign Affairs it's pretty interesting all right there's another question from Greg do you think OpenCL should be adopted by the deep learning frameworks since CUDA is the only choice of the industries I giving anybody adamantly even though AMD GPUs might be cheaper it's a good question again as the CEO face I don't get to go that deep sometimes on subjects like that so I'm not sure we'd be the best person to answer this question feel free to tweet it mentioning hugging face and I'm sure someone from the team would be able to answer something like that I put here my Twitter and all and also my email address so if anyone has any other answer or question or if you want to use an optional company and that you need help or yeah anything really feel free to feel free to reach out to me I think we covered most most of the questions unless there are more more than like we got got a good quick overview there let's see what do you think yeah thank you very much it was a great overview and I think that text questions work really well so I'll save this shot and posted alongside with the video recording but we have time for regular Q&A so if you guys want to ask questions just by saying them this is the time to do it so we can try that guys you will kept on mute yourself first because everybody is muted by default so Clement this is one see what inspired you to build this build hugging phase and this particular marvel of open-source yeah yeah I've been I've been obsessed with an LP for quite a while as I said I started with computer vision and I always figured that I NLP would be a much bigger field just because of the importance of text in every everyday life so yeah out of an obsession for now key we started at the beginning as I said with open domain conversational AI we wanted to build her from like from from the movie and we can fly yeah been lucky to underlying technology that we realized was useful for more companies than just ourselves so we started to open-source it and people started using it then it comes like makes sense with our vision that you know you not not one company we solve an option you know like we really need to take like a collaborative open source approach to these to solve such different problem and such a different challenge so open source was was the way to go and the way that we felt like we could more contribute to the ecosystem and there was a the beginning of the of the chat there was a question about why the huggy face name and logo we wanted to have an emoji as a name because when we go public we want to be the first company to go public with an emoji versus three identifier letters and companies usually have when they go to the stock market so we absolutely wanted an ode an emoji and then when we look the list of emojis we felt like the hugging face emoji was the best one to represent an LP because it's still an emoji right so it's a machine it's the technology it's not human but at the same time it's showing a very human emotion right the hug the same way NLP used to be something that is only human right only humans could understand and now with technology you manage to understand it so we felt like it was a good kept like metaphor for natural language processing is this hugging face emoji unfortunately right now with coronavirus we've tried to advocate for less hugging at some point we change our logo added a little mask to the logo to help spread the word of obviously social social distancing and everything but obviously hopefully in the next few months in the next few years when we will find a cure for chronic virus we'll be able to hug people again yeah can I ask yeah go ahead yeah Peyton wants to Stephen yeah thanks for your really good presentation so if you can probably see my question in chat and my first question it was easier once your zero shot model is it naturally us-centric because you wanted it to be or simply these sort of built-in bias that most new sources are in u.s. or currently are concentrating and us event and then B and B can we tweak that to give weight so as I say if I wanted a model that was focused in Central Asia where Kazakhstan was local and regional everything else was foreign yeah yeah yeah so here like for this you shot demo the model is board mnay right if you go back to the paper or the blog post and I mean I guess my sort of hidden intent is yeah pre train models are nice but especially if you're dealing with low resource languages if I dealt would say African language it's actually bad to learn the inbuilt bias of the dominant news sources and to be anxious so the word domestic if I say I want a domestic beer and all that important one that may mean something different and it's very complex I totally and to answer your second question first question at the same time it's all based on open source right so you can obviously find units or pre train it on different data sets for it to fit better your own context and your own tasks and and remove some of the biases that you you don't want buddy but if I take an off-the-shelf my intent is kind of more like if I take a standard off-the-shelf model I have to assume that sort of words like domestic and election and president are sort of polluted with regionally specific assumptions totally totally and that that's why we as I said we push people to give more and more information about biases like that in the model cards so that if you take their models then you can have a good view of what are the biases and see if it fits or if it's too much too much of a problem yeah like for example on this task so if you do if you wanted to [Music] classification which you want to text classification then you like all these models that you can you can look at and they're all out like different different biases but maybe you can find one that fits you use yes and if you can't then you can use our training scripts to basically train it on your own your own data sets yeah actually a fast and dirty trick I used to use with word factors this I would just look at say look at the word domestic and then compare it to every country so like whatever US Russia and then you can see what it aligns with and you know closest distance and that's good for expendability but is there an equivalent of that with transformer nearest your nearest nearest you know nearest cluster yeah a couple of so in the ecosystem there are couples like explain ability tools you know I should like at the beginning of the talk with experts so some people have been working on that I'm not super deep on the on the topic so I'm not sure I would be the right person to point you to except the exact right tool I can classify I can write a for loop I can classify the same sentence with 200 different country names and I can see which one domestic classifies closest to goes by being at the center of the ecosystem is to foster more and more explained in DBT tools to appear we think we still at the very beginning of it and we need to understand better what what happens in these models so if you if you end up working on a tool that works for you I would really encourage you to open source it try to integrate it with transformers because we still have demand for it and we believe it's very important to work on these topics sure absolutely it's the last last throwaway comment when I type to Mars Attacks it was classified as 99% foreign so we're all safe in Kazakhstan I can resume that's funny another way I wanted maybe at some point to show it do you see like diocese - is another demo that we have which is long form questions ring that is trained on Wikipedia and Ridgid with something called li5 so basically the the goal of the task is to explain complex questions in a very simple way so like for example what's the best way to treat sunburn and then it's going to take some information from different articles it's a really fun demo to play with it shows also a lot of like bias sometimes both from the model and from from the datasets but it's it's an interesting thing to check - the goal of most of our demos is to give people more tools to understand how these models work so that's usually the motivation for demos we have another one that is called right wrist transformer which is autocompletes so you basically I'm gonna post it in the chat at the same time that where you can test you to complete with like different like model size different topi different temperature and then to complete like that but the purpose of most our demos is to give people tools to look at what these models are producing to understand them better and hopefully take better choices when they want at some point to put these models in production I have a question can ask yeah go ahead I saw that you showed some content to your models via your website is there any chance to access it the REST API or programmatically yeah yeah that's something we're going to release soon that we've been working on pretty hard in the past few weeks so yeah yeah we if if you wanna if you want to preview feel free to send me an email and I can send you an early access to it but it's coming pretty soon did you post an email music yeah it's Clement so it's my my name at hugging face that lets you go thank you Oh any other questions got a question for ya and so you were saying the hardest the hardest topic topic area is general conversation do you have any general comments of an anal you because every time I see a new and people talking about speaker intent it's always like I want to buy somebody a dress for for a birthday present and it's never sort of deeper let's say multi-purpose and all you like that you can actually sustain a conversation with an artificial personality on multiple topics and you know it's got a sense of humor it's got a sense of context you know like a you know quote unquote real human conversation yeah yeah I mean I think it's really really hard to pick any you like it's addition of a lot of like different NRP tasks right and you have to be really good at each and every single one of them to give the impression in the context of a conversation to really understand and which human needle so its reorder if you want to take a look at what we believe is kept like state of the Arts I think we have Dino for that too it's pretty cool I think it's something like like that yeah what would you what would you say is missing what's you know the worst missing feature from conversation is currently the tough question it's a tough question it's a million-dollar question for if we're sure well I mean I could make it come like when you say context earlier I know these two say contact it's just like credit card or whatever it's it's it's early very shallow but if if this was like a bank would would we expect to see more context like this customer has been with us for twenty years they have this credit card with this limit the sister income the sister family this is their spending pattern is context and arbitrarily deep thing I don't know I mean I think to answer you question my guess but maybe it's not the right case and I'm not I don't have kept like but I think to me one of the areas where we working really hard at hugging face and that we think is really promising for NLP in general but specifically for conversational is the ability to add more structured data to the representations I feel like in like a conversation you can extract a lot of information you can use a lot of conversation that are more in a structured way like their bases and stuff like that and that's when you'll be able to really use that in addition to more Catholic and supervise models I think that that's when you'll see really could progress because that's when the conversation went Oni can't like make sense in the general way but also the accurate you know like the facts will be will be accurate which i think is still like a big thing missing for like a good conversation oh yeah so I would say yeah like ability to create more hybrids like [Music] knowledge representations in a way is something that with we're investing a lot on in terms of research that believe could bring conversationally I may be to to the next level sure but when I'm thinking we're about so that credit card example him so like say if it settle you like for suggesting financial product if you if you suggest like say college savings plan for somebody if you know that it in us if you know that they don't have kids that's one thing if you know their children age too that's another thing is they have children age 29 it's another thing so to what extent do you really want to incorporate like a very deep sort of context object in the model itself as opposed to just doing a sort of you know graph query on all of the information we have about the customer like maybe the model if we start over fitting the model to all sorts of users maybe that's the wrong place to stick the context yeah maybe it's a difficult question I I might have given you a better answer at the beginning of this talk now after an hour 15 and I'm not sure I'll be able to give you a good answer on that maybe I can give a very small answer I mean so I'm Thomas I'm walking on the science side a tiny face um because I think these models these are these are really interesting and definitely it's kind of strange right that we store all the knowledge in the model and then the model is this static object that you carry on for four years right like Bert will always think that the president of all the countries were the prison they were in 2018 when it was trained so we are actually actively working on that and we have a morality yep what if a morality ephemeral time changing yeah exactly yeah yeah we would like the model to evolve and to be able to ingest new knowledge and so we have a nice collaboration that we that we will publish that we will read is in a few weeks I hope with research teams who are working on models which include a database in the model we for retrieval components and so when you ask the question the model first queries its database and then retrieve the components and then digest it so it's on your net right it's so like this it's all deep learning but there is this evolving component in the models um we are very very excited about that at any phase and actually one of the reason we released the NLP the data set processing library which is actually a lot more than just a tacit processing but is basically just is this huge database processor library that makes it very efficient to query a big database and to use that model the one of the main reason we release this library is to allow people to use more this type of models because they are not really used well I mean people I don't think people working out with this type of evolving models right now and one of the main reasons that the tools are missing so as always we we set it apart as to fix the tool and to let people like use Maurice models so yeah just stay tuned and I'm happy to get your feedbacks on this release that when we do it should be it should be now in July pretty silly yeah I think I broke your convoy model I type I typed who should be shot and the answer is the police officer I work for I of a gun so the guy wants to shoot his boss or the guy believes his boss should be shot as any other questions I'm sorry if I showed you leave it off my gmail forgot I was sure sharing my screen is it is it so you guys are currently privately funded are you intending to be aborted offended no we're privately funded we think we can be like independent company that's what we're trying trying to do we've been lucky that investors allowed us so far to only focus on open source and open source adoption and and science at some point we'll have to work on our business plan grab a business model rather but but we're not there yet no I'm sure it's superb but I mean are that what are the KPI like even five years ago we probably wouldn't have been able to do this what are the KPI people want to see to sort of justify is it downloads usage API mean their source everybody investors anybody who wants to try to understand the valuation I mean I think you know usage really massive usage on our open source now so it's a good proxy to see that we're doing something useful for the community just because yeah a lot of researchers a lot of companies are using our open source that's a that's a good proxy for you to see that you're working on on something something good that's one of the beauty also of like you know releasing research and releasing open source which is at you you can see if it really makes sense or not you know like if we were like a research lab just working in a closed source not realizing things we could say we build the best thing but nobody could really say if it's true or not there's this camp like working in like a open set up an open source we can see if people are using our they're using our tools it means that it's probably useful if they don't it's probably not as useful as we as we stood and and we have to work harder sure no no sure I mean I'm saying that it's nobody's pressuring you to become like a SAS service so you're just a pure play open-source a toy company and yeah yeah you're not gonna become as any SAS start doing hosted or anything we don't know we don't know to be honest like right now we're still focusing on an open source it's always going to be the core thing for us we have many ideas for like more like page features that we might start experimenting with at some point but yeah we really don't know at the moment let's think and once again thinking whenever was awesome appreciate it we'll post the video soon on functional TV so you guys can share it with everybody else and I also will save the shot with the excellent questions and post them there thank you guys sounds good great thanks everyone thank you bye buddy safe