ai.bythebay.io: George Hotz, Self-Driving Lessons from Comma AI
Recording: ai.bythebay.io: George Hotz, Self-Driving Lessons from Comma AI
[Applause] good morning it's Tuesday it's true uh so my name is George I'm the CE of Kami and our slogan is ghost riding for the masses what we have a a Comm Neo um this you can download off of our GitHub you can download the plans for it uh and you can build one of these yourself um it goes in your car and we also have a dash cam app U it's a cloud-based dash cam app and the cool thing about it is you can download it right now uh it's available on both Google Play and the App Store and so right now it's just a dash but the dream the soon to be plan is to offer some ad ass features in the dash cam right so we'll offer some forward Collision warning some Lane departure warning good Lane departure warning not useless Lane departure warning like like most I mean seriously most people turn it off um but if you want those kind of features today if you build a common Neo you can install open pilot open pilot is our open source level two driving agent and I'll say this now it's better than autopilot too I mean it is uh it's not better than autopilot 8 autopilot 8 was pretty good on the old Hardware it's about on par with autopilot 7 autopilot 2 is terrible um ail's better you can download this and you can put this in your car uh we have a couple videos of it working that's the GitHub URL for it and it's open source um so if it doesn't work with your car you can add support for your car uh people have actually done this and hopefully we'll be uh merging the changes Upstream soon but yeah so I said level two what does that mean here the official self driving levels I think that's uh sa and that's Nisha let's zoom in a little on the sa one um the system takes over both steering and acceleration and deceleration in a defined use case that's level two level four the driver can hand over the entire driving task to the system in a defined use Cas wait a second steering acceleration deceleration what am I missing what's the difference they say the same thing right like like like steering acceleration deceleration the entire driving task that's the same um it's very confusing let maybe you know you got to employ people so they got make it confusing right but let's go through what self-driving levels actually are right so you have level zero everyone knows what level zero is no autonomy you have level one cruise control a ACC what you're seeing in a lot of the traditional auto manufacturers today and then there's really two level two systems out there there's Tesla autopilot and there's comma AI open pilot um and we'll get to like a better definition of sort of what this means but those things are capable of doing the steering the acceleration and the deceleration so why are there more levels have you guys seen this before who's seen the levels for self- driving cars good A lot of people have seen it right why are there more well the ones after two are liability levels right in 0o through two the automaker never takes liability in three the automaker takes liability sometimes and in four the automaker always takes liability the only difference between level two and level four is insurance right insurance is a real simple industry insurance is like reverse gambling right with they the house they they bet on things they know they have their actuary tables right so if if we build a system and we get it out there to a lot of people and we can show that like on one stretch of highway statistically it's actually safer than humans then we can sell that liability to an insurance company right and then for that stretch of highway it's it's level three right and then you do more and more and more and eventually you get to level four how anybody thinks that there's going to be some magical jump to level four baffles me right you want to prove something's a good driver do it statistically get a system out there as quickly as you possibly can iterate on it and this is I mean this is the Tesla strategy and this is also the Comm AI strategy if Tesla is the iOS of self-driving cars we want to be the Android um but okay so let's talk about the only number that actually matters right forget all this level stuff let's just talk about disengagement events right so this is Miles you can go between disengagements like kind of like meantime between failure kind of stuff right um I got a bunch of numbers it's a vaguely logarithmic scale down here in one you have like the Legacy automakers right if you get a Mercedes a Ford or a Honda you get on the highway you put on all of the self assist features you'll maybe go a mile between actually having to do something right um then you have a bunch of systems that are 10x better it's kind of where we are it's where Tesla is even Uber and Cruz fall around here they're kind of doing different parts of the problem but they still have a disengagement event about every 10 miles um this is where almost all the self-driving startups are right now uh there's one company that's 100x better and everyone should know what they are right the best technology by far um and then what's at 100,000 do anyone know that's what people are right uh people crash about once every 100,000 miles it's a rough number but um and then you have some people who are thinking autonomous vehicles are going to get you here this is impossible uh you can write the world's most perfect autonomy software what if a wheel falls off the car right like like there's there's there's so many failures never is impossible and it's absurd to try to talk about this this is the level of risk that we are willing to tolerate you know 30,000 people die every year in in car crashes in America right this is acceptable if we could get a system that's better than this that's it that's Victory right um so let's talk about how to actually do that I mean you can see there's maybe a 100x gap between all the other players okay so you have the automakers down there they're kind of hopeless um you see about 100x gap between the startups and Google and then you see another 100x gap between Google and the people uh yeah so let's talk about what it would actually what you'd actually have to do to close that Gap um so our mission this slide is a different color because it was in our pitch deck um it's been the same all along to build the world's first superhuman self-driving agent and I think from the last slide we now have a rigorous definition of what superum means right so I build a superhuman driving agent um let's talk about robotics right let's talk a little bit about it's kind of how I break down the self-driving car problem right you have localization where am I perception what's around me planning what do I do controls how do I do it uh one of these four blocks is already solved controls is solved um you have a model for the car uh also controls for like like highway driving and say it's so linear right it's such a simple problem you're not trying to do do some like fancy control system this isn't helicopters doing inverted hovers this is like I turn the wheel a tiny bit I step on the pedal a tiny bit right smooth driving is safe driving um let's talk about the other ones localization where am I down to like you know centimeters accuracy this is solved with money you can you can do this today um you have like Google and Cru what they're doing they're using a liar they have prepped regions and you can get really good localization if you throw money at the problem um perception is tricky on the long tail so remember Google's been at this for 10 years Google's probably the only uh company who you can kind of say has solved perception for self-driving cars I think almost all their disengagements are not perception failures they are planner failures um this is the hard problem with the top two you can get to 1,000 miles between disengagements I mean obviously you need some planning but not like like fancy sophisticated planning if you want to get to the 100,000 miles you need everything on the board um so yeah let's talk about localization and perception like I said uh perception is solved by Google using a lot a lot of debugging localization could build it for you in two weeks if you know you give me a lar and some maps um but how do you do it on commodity hardware and what's really interesting about commodity Hardware is commo you can get this today right but what's the advantage like you look at these you look at these uh these fleets right these companies that have fleets uh Google has a fleet Cruz has a fleet uh the automakers have fleets dely has a fleet uh they're tiny we already have the world's second largest Fleet obviously Tesla's the first right doing this on commodity Hardware enables you to actually ship this if you build a Neo today and put this on your car right now we're about 10 miles between disengagements you're going to get all the way up to a thousand just with software updates you're going to see the exact same thing happen with Tesla too camera only oh you need lar lar is the enabling technology for for self-driving cars I feel like a lot of auto manufacturers think that all you have to do is uh well we get a lar and we get maps and then we have our integration Engineers put the bolt between them and that's how you solve self-driving cars right um if you don't think camera only is possible let's do a little thought experiment 360 camera plus Oculus man he could drive the car right does anyone really doubt that does anyone doubt that that guy could drive the car with the Oculus and okay so uh the link is very hard assuming you have some magical high bandwidth low latency no network service issues link you drive the car no problem right and if that's possible it's clearly just a software problem right um the thing too about about about ldar is like a lot of uh what lar helps you with is like maybe the first five layers in your confet um you can do incredible things today with cameras uh like segets and depth Nets um you know this is this is a this is from coloring you can uh see where the Lan markings are where the road is where the cars are where the traffic lights are where the undrivable obstacles are you have depth Nets you have mono depth Nets um so you also don't eat two cameras two cameras is another one of these uh oh you need two cameras I think Bosch loves the two cameras Oh you don't need two cameras right um so you can infer tons of depth information even from just a static picture of a road because you're in a constrained environment right everyone kind of knows how big a car is um also we let people with one eye Drive uh to think that the two eyes are actually helping you with the driving problem at all is kind of absurd your Baseline is about 7even in as a human I think that gives you resolution out to maybe 10 or 20 mters right most of the things you're doing when you're driving is is way further away than than 10 or 20 meters actually people have no idea how far things actually are on the road or how long Lane lines are um but yeah you have you have seg Nets and depth Nets um so and obviously you don't just have to get depth in one picture you can also use structure for motion techniques right you can you can think about uh one camera moving as the same as two cameras here's a camera here's a camera I have good odometry on my car I know what the you know the the translation and rotation is between them there you go you got a you got depth right so you can do incredible things with cameras today um so that's kind of the perception problem localization this videos from our Dash can um the slam problem is much more constrained if you know that what you're in is a car right if you're trying to solve slam for like inside out VR right you're like waving a phone around and you want to recover both the path and the objects that's that's pretty hard right um especially when you start to deal with things like rolling shutter artifacts um but in a car this is all easy cuz the movement is so defined right you think about the six degrees of freedom of arbitrary camera Transformations and a car I mean there kind of only two right you got the this way this way and the this way this way right there's a little bit more to it but mostly with those two Dimensions you can specify a lot of the stuff about the camera transform and then your bundle adjustment does a phenomenal job kind of filling in the really small values for the other dimensions um so yeah I mean you want to build a you want to build a centimet accurate 10 cmet accurate local ization system from a cell phone gather pictures write the code slam the world um yeah I think we have a let me see what the next slide is let let's talk about the other problem right so that's localization and perception they're kind of solved with money we'll solve them on commodity Hardware now planning uh there's only one system in the world that solved planning that's people um AI hasn't solved this yet right and there's kind of uh a few tricky problems so you have long time Horizon learning right so you had like an lstm you could do like back propagation through time for maybe 50 steps um beyond that it gets kind of I mean even lstms right maybe in RNN you could go back five steps before your gradients explode or vanish lstms you could go back maybe 10x that but 50 steps back if you have a camera our our camera Loop runs at 20 HZ that's 2.5 seconds right that's not really long enough to have the kind of context you need you need to kind of get another 10x from somewhere um there's the other problem which is behavioral cloning doesn't really work to learn a task like this what you really want to do is reinforcement learning um but in order to do reinforcement learning you can't do this from fixed sets of data right so you need kind of two things you need a simulator uh which lets you deviate from what your data actually did and you need a reward function which how do you specify that for driving right I mean you could have this vague reward function of don't crash but certainly in your data sets crash events are very very rare and we're not going to let cars go out and actually explore on the roads right it's all about you ship something it better be all exploitation and not exploration um so you have like this thing called inverse reinforcement learning the main paper that you hear about when you hear about inverse reinforcement learning is the uh the Stanford helicopter paper really cool paper but they're solving fundamentally a different problem from uh self-driving cars the inverse uh the helicopter like getting a helicopter to do an inverted hover it's really easy to specify old school when a helicopter is doing an inverted hover it's just a question of what control inputs do I actually do to get it into that state and to get into that state and stable right so there inverse reinforcement learning works really well but the car driving problem isn't this it's actually really easy Once you know what state you want to be in to figure out how to get there right now of course you could say okay finally the state I want to be in is parked outside of this Taco Bell I mean yeah okay back propop that through your 30-minute commute it doesn't work right um so you need something else uh to solve PL planning and both of these things you're starting to see the writing on the wall uh so dilated causal convolutions from wavenet um I mean it's a cute trick they kind of turn n into log n and n into log n is well sort of sort of what you need to get that 10x um this is wavenet they did uh they did audio synthesis stuff um and the the convolutions are are causal the arrows all point forward through time because we make an assumption and it's a pretty good assumption that events in the future don't cause events in the past um so you kind of have the arrows all flowing this way it's dilated so you can get these really long time Horizons um yeah so you can kind of see that dealing with the time Horizon problem the other problem is kind of dealt with by Gans right you can learn a world model what's a world model it's a function given the state at time T and the action at time T that gives you the state at time t plus one you can even learn a reward function and you can learn a reward function by kind of assuming that the average is good um I mean not exactly the average is a little bit more sophisticated than that you have to learn to reject other kinds of noise it's clearly not just a beautiful gaussian um but you can learn a reward function using these uh probably using a again um we have a paper about this uh it's kind of like yeah we we learn a world model the problem with learning the world model is well it works for about 2.5 seconds so you need sort of both of these operating together to truly solve planning in a generic way right so you think about what's being done right now to try to solve planning uh you look at Google and Uber and they have this big well it's it's a hand engineered feature space right they talk about three-dimensional bounding boxes for cars they talk about threedimensional Bounty boxes for people they talk about the state of the police off officer's arm you ever say remember that in the Google TED video right like you're weirdly defining this this feature space and this is hand engineered feature spaces and if we've learned anything from the history of AI it's that hand engineered feature spaces always loose um so how do we not think about this like how many dimensions is the driving how many dimensions is the state when you're driving like it's I don't think it's thousands I don't think it's t I think it's probably like hundreds right I think you could succinctly describe almost any driving scenario in like a you know 512 dimensional Vector right so it's just the question of getting the right thing to take it down to that state for you and then in that state Vector is where you can do your RL you look at things like the Deep mine paper with Atari and it's cute but when you look at how much data really went in to uh to doing that it's it's it's it's a huge huge amount for a simple game like Atari what you really need is is to get into some abstract hopefully more smooth and more gaussian feature space and then do your reinforcement learning over there um and this is almost exactly what ganss do and it's it's pretty incredible um you know I think we're going to see huge uh advances in both of these this year um so yeah Kam is not a research lab um we're not trying to solve these problems I'm waiting for someone else to solve them right um what do we do well we gather really big dat cuz clearly for all of these things perception localization and planning Big Data helps so we have almost a million miles of data and you say oh that's nothing compared to Tesla's 100 million uh well a few things about Tesla's 100 million they don't have the videos back for that from the autopilot 2 Hardware they're getting the videos back but from the autopilot one Hardware they're not um and you can know they're not because Hardware wise the camera is only connected to the mobile ey uh and then the mobile ey is only connected over canvas right so you canvas one megabit the mobile ey doesn't have a video codak on board you could maybe store a few frames in the mobilized memory and then pull them out and you've seen this if you've seen people pull the crash data from Teslas right but they're not getting all the video back we have video for all of this um this is probably the most diverse uh driving data set in the world too right so you don't want data from the same geographic area you don't want data from the same phone mounting you don't want data from the same lens you don't want dat data from the same sensor you want all of these things to be uh like kind of marginalized right these are you should be invariant to all of these things right you should be invariant to the maybe even the exact uh placement of the thing right because humans kind of are um that's actually interesting if you ever try to if you drive on the right and you go to the UK and you try to drive on the left it takes a little bit to learn and then what's even scarier I come back from London and uh I almost sideswiped a truck I didn't realize like I just didn't realize that I didn't unlearn and I was way too far over in the lane um but you know you can learn this pretty fast so sideways sideways translations are maybe something that's a little bit harder to deal with but you know small variants of rotation uh different different sensors different Distortion models um it shouldn't it shouldn't matter that much for a driving agent it does matter if you want to do things like slam particularly there's a new uh from the LSD slam guy there's a new paper that came out that's that's that's it's a direct sparse odometry um he's using not just geometric models of cameras but photometric models of cameras and that stuff is super human I like in some weird ways like that computer vision is superum if you want to know like how far something is away I mean you like like getting your getting your features from a 16 megapixel camera over a few frames you can do this to some some some H superhuman level of accuracy um which is pretty cool but yeah uh so yeah get uh get big data um step two own the network so with data you can only kind of do the behavioral cloning kind of thing there's no exploration you want to actually take the models you've trained and try them and not try them in a simulator uh Chris armston uh gave me an awesome quote simulation is doomed to succeed right if you really want to deal with all the final edge cases of the world you're never going to code them into a simulator and people who think you are it's kind of like I mean it's the same things that kind of plagued computer vision for a long time right like okay maybe you're not specifying the you know you're not specifying the vision code but you're still specifying the graphics code you're still not dealing with all of these weird anomalies you're only sort of dealing with like you know your your your feature space is whatever you're putting into your your renderer right um so you want to own the network too and the really cool thing about owning the network is you deploy a model you see where the disengagement events are right this is all level two so people are driving these cars people are paying attention um they disengage the system you get the data back about that disengage and there you go it's reinforcement learning where the users are the agent it's a big reinforcement learning algorithm being run on the world right I mean your time your time for like Epoch is kind of slow right uh so you gotta you ship out the model in the morning you get a bunch of disengagement events in you back propop them and uh you know back propop them get a new model ship it out the next morning maybe we can even lower that model feedback time you don't start wanting to do learning on the device um learning on the device scares me because like you know you don't ever know if you can learn something bad right so you need you need in order to do this and in order to make this safe you need good you need good tests you need good tests that tests both you know static data and dynamic data simulators are good for testing they're just kind of not good for training um if your thing doesn't work in a simulator it won't work in reality probably um if your thing does work in a simulator it might tell you something about whether it works in reality but not that much um so yeah own the data own the network uh so people like when I have calls to action said this about uh Tech hunch talk when you call to action download a dash cam it's available in the uh Google Play Store and the App Store it's called shiffer it's pretty awesome uh we're going to be rolling out ads features to it as well um yeah download it and also we are hiring um W com AI reports of our demise have been greatly exaggerated uh we're here and I think we're going to win but don't take my word for it you know go build one it's pretty cool uh yeah take [Applause] questions no we we'll get it on [Music] video uh same question I asked uh the Nvidia guy what do you do with the bad data it looked like you're only keeping a fraction of the video um so we keep all the video uh there's kind of no such thing as bad data so all good drivers as the ganar rened a Quil I gave this to The Verge testing Vegas uh all all good drivers are good in the same way all bad drivers are bad in different ways um our Nets don't behave like the average driver they behave like the committee driver imagine you had 100 people in a room voting 100 times per second on what action the car should take even if 20 of them are bad drivers this doesn't matter even if 50 of them are bad drivers how are they all going to be bad in the same way the good really stands out so there is almost no such thing as like bad driving data now how do you filter out if you mean by bad data like people who turn the dash cam on when it's sitting on a desk I mean that's trivial to filter out right you have a we have a model that just kind of says is this driving does it look kind of like driving I mean just use a gan for that right it's easy uh do you know how many NEOS have been built and deployed um yes uh more than 50 less than 100 any other questions oh coming hi uh back to the data question so uh the good data is um automatically judged or somebody looks at them no one looks at the data so our our whole um pipeline is automatic ground TR thing uh so you can do the reason automatic ground truthing works and it works so well is there's so much redundancy kind of in the data right so say you build models that do prediction on single frame so you build models this a simple thing like Lanes right you build models that extract the lanes from single frames of data right you kind of know if the model is doing well or badly based on temporal continuity and now this isn't perfect there are still failures which can happen but fortunately the failures of neural Nets are usually not like these kind of failures they're usually not like persistently wrong unless it's persistently wrong in the ground truth so we do some amount of hand labeling um our seg nets come from hand labeled data sets um our depth Nets are automatically ground truthed using a combination of stereo and ldar um our model that we actually ship to the car is trained using those other models to assist in ground truth recovery um so if you know I mean what if you have a seget and you're run it on an image you know whether it's driving data or not right you expect to have some proportion of Sky some proportion of Road in the right sort of places you can clearly tell whether it's a desk or not right um You can also do this by some amount of uh 3D map recovery some amount of like I said the camera is generally confined to a uh two degree of Freedom space if it doesn't look like it's in a two degree of Freedom space it's not driving data you also have tons of correlation between the sensors right um even even things as bad as GPS speed uh work pretty well um the if you actually have a connection to the car you get even more sensors too you get the speedometer and you get the steering angle sensor and you get cars have separate every car that has uh traction control has a sensor on all four wheels and this is almost always accessible on the cam USS [Music] coming um what are your plans to get the cost down so more people can can buy the hardware cost is already pretty low um if you're if you're thinking about putting this thing on a car I'm not sure how many gains we would get by lowering the cost less than $1,000 right I'm not sure how many more people would build something and connect it to their car if the cost was say 100 instead of a th000 but I mean if you put things together in this talk right how different is the common Neo from shiffer what if you could make an app drive a [Applause] car um we we still use the radar uh we still use the radar for adaptive cruise control uh we can do Vision only adaptive cruise control but it's not as good um actually radar adaptive cruise control kind of superum because radar gives you speed as first order information right like humans are not very good at estimating 100 m 100 met out uh what the speed differential of a car is right it's kind of hard to tell because you're using like temporal difference and you're like okay I did did first frame was 100 was the next frame was it 99 or 101 you can't tell right um so yeah uh got to do that first but you can do it so as we uh move from left to right and have more miles per disengagement events I mean is there a concern around handling that trade-off between when the machine is in charge and when the human needs to take back over charge um so like the handoff problem yes yeah I mean there's some amount of concern right so what we do is we have a six-minute timer um the concern really only happens if like the car is driving for hours without uh any intervention and then the human needs to intervene suddenly um so we have a six-minute timer if the human hasn't interacted with the car in 6 minutes uh basically you know the the system stops accelerating and it beeps throws up All These Warnings this is the same thing Tesla's doing the my hope is this right so if you're building a level two system yeah you probably want to keep some kind of timer like that right or at least some better estimate of the system confidence my hope is that if you get enough of these level two systems out there even with the six-minute timeout you can start to classify what all of your disengagement events are and you can say look on these kinds of Roads there's not a disengagement and that's how you become level three um now you can't do this with a cell phone you know you're not going to start saying the cell phone can do any sort of level three stuff you're going to need you're going to need better Hardware to do it but the hardware certainly exists today you look at like nvidia's Drive px2 make sure that there's like an azal chip that can do some amount of safety critical stuff um our safety model basically works like this it works assuming that the human is paying attention right and six minute timeout warning when it pops up same as Tesla um to uh try our best to ensure that right and the truth is like if people don't pay attention they probably don't pay attention when they're driving that much anyway um the then there's there's even more you need than that to actually make a system safe um you need to make sure that whenever the human wants to take control of the car back it's very easy for them to do it and not all systems do this actually so most systems will disengage uh when you hit the Brak every every cruise control system disengages when you hit the Brak we disengage also when you hit the gas right because if a car is breaking when you don't want it to human's first reaction is hit the gas right and I'm surprised it Tesla autopilot doesn't do this right these systems let you hit the gas and stay engaged um then the really scary one and this is why we realized we needed a gas disengagement and this actually happens uh this happened on the stock Honda system um you step on the gas and okay the car wants to break and it sees that the human stepped on the gas what do you do right there's kind of two things you can do one thing is continue to step on the brake and hit both pedals at the same time not very good but the second one's actually worse which is if the human is hitting the gas suspend the break so what actually ends up happening here is the human steps on the gas The Brak stops they're like oh this is pretty nice they let off the gas the car still wants to break it hits The Brak that's how you get rear ended right um so this is why we disengage when you hit the gas so disengagements are very intuitive we don't disengage on steering though um some of these systems disengage on steering and like our our safety model for steering is this our systems incapable of putting more torque on the wheel so it's it's it's it's two things that are dangerous on steering um so yeah well this gets into like the second safety principle in general which is always make sure the car doesn't do something doesn't doesn't jerk doesn't do something quicker than a human can respond uh so in order to do this you put you put actuation limits on stuff you can have like ramp up on some actuation you know don't come out ever with one g of braking but if you slowly ramp up to oneg of braking and the user doesn't do anything that's probably the right choice right um and you can you know gather data for all of this and try it but uh yeah for steering we make sure that the amount of torque put on the steering wheel it's only like the torque you could put on with like a pinky right so imagine someone was trying to push the steering wheel with a pinky while you had your hand on it I mean you'd kind of feel them pushing a little bit but you know it doesn't do anything um so yeah the two safety principles are always make sure the user can take control intuitively and easily and make sure that the car will never do anything so quick that the user can't doesn't have time to respond assuming user is paying attention people do pay attention right I mean you you've seen this like okay there's the Tesla autopilot death it's it's all over the news but what about all the people who drove it and paid fine attention and then what about all the people who didn't pay attention without Tesla autopilot and also crash their cars right um so yeah I mean at the end of the day it's all about the numbers uh question back here this is just a little bit of like a a fanboyish question but do you ever host like Drive-In hackathons we thought about Drive in our cars and work with people who know a lot more about building these things than we do to build our own we thought about it we we we thought about kind of Hosting like a like a comma AI hackathon kind of as like a recruiting thing um yeah maybe we'll think more about it cool um one more question thank you um considering the bad data that we talked about I have to strongly disagree with you in one point and I think there's a lot of situations where a lot of people drive bad in the same way and that might even be because um human sensors are not good enough for example I could imagine invisible eyes that a human could not see but the camera sees it so all the humans they don't slow down because they don't see it and the camera or or your whole system system thinks okay there's ice but apparently that's not a reason to slow down because nobody did it so in that case I would definitely say that's bad data and we should somehow filter it out well okay so you can do you can do a few things um one you can look at you have the future having the future is a very very powerful thing you know what happened right so there is some amount of um you know regret that you can assign to something right um so yeah obviously if a human crashed that's clearly bad data you need a lot of data before you get a statistically significant amount of Crash events but I mean that specific Edge case yeah maybe you look for crashes right but the truth is humans driving badly if you define badly as like by the rules of the road you don't want to drive by the rules of the road you know how many times the Google's car Google car has been rear ended and you watch it drive around Mountain View and you see why so everybody working on this understands that you need to make your car predict the world around you you need to predict what's going to happen and what other people are going to do but there's another component as well you need to drive predictably you need to drive in the way humans expect you to drive anyone who is talking about vehicle to vehicle communication a future city where all the cars are these magical level four this is diluted in reality if you want to ship a self-driving car you're going to be interoperating with humans for a long time and in order to safely interoperate with humans you don't just have to predict what they are going to do you have to drive predictably yourself when is a driver bad you're on the road you look around why is that why is that that car driving badly because you're driving unpredictably swerving in between lanes turning without signaling right these are these are I mean these are obvious cases but there's so much more to it right driving is not some very rigid specified problem if it was then I could write up a reward function driving is a intuitive dance between humans and to solve it means kind of behaving like a human and I think the solution that solves this will also be the solution for an Untold number of Robotics applications um um I say this I shouldn't say it anymore but the real goal of Comm AI is to end all jobs just kidding chump just kidding just kidding all right all right thank you very much