Transcript: Julien Le Dem, Datadog, on Reliable AI — Interview with Alexy
Hi, I'm Julian. Uh, I work at Data Dog and I'm a principal engineer there. Oh, so recently I've been playing with uh Cloud Code to build a little uh visualizers for park metadata. Um, and I'm more of a back-end engineer, so I know very little about front end stuff. And I was able to make this little web app with it. Was pretty proud of myself of uh building new things out of my comfort zone using AI. Oh, that's a that's a vast question. Um, I think right now we don't really have fully reliable AI, right? We need to triple check everything. Um although we'll need to have a better definition of what that means in the future, right? Like that the result can be trusted uh that they accomplish what we were expecting them to do. So there'd be a lot of validation um in you know if they're using tools or building things and so on right like it's actually if the goal is the AI to produce a lot more like code say than we can build uh then it's going to be tricky to make sure it's not introducing uh flaws or security uh vulnerabilities and things like that. Well, we need to have uh good uh validation frameworks like verifications. How do we make sure the solutions what's happening is according to what we were expecting, right? If we're going to automate a bunch of things with AI. Um we'll need to have a much better validation framework. Um, I think right now, you know, if you use an LLM to do things, it it's pretty open-ended tool. So, it can do a lot of things you were not expecting it to do, including the wrong thing. It's been changing so fast in the next five years. I don't know. I think the next six months is tr difficult to guess. Um, I don't know. But I think it's all accelerating, right? So I guess the one of the dimension is all the problems for which it's easy to verify the solution we'd be solved very quickly because then that's where you know AI can try a lot of things quickly and as long as we can verify that it gives correct answers um we'll be able to use that very efficiently. for things that are harder to evaluate whether they're correct, it's going to be more difficult or potentially uh unsafe, right? Um so we see what happen. I don't know what this stack is going to look like, honestly.