Devreal

SBTB 2023: Karl Wehden, An introduction to Open Accelerated Discovery.

SBTB 2023: Karl Wehden, An introduction to Open Accelerated Discovery.

Recording: SBTB 2023: Karl Wehden, An introduction to Open Accelerated Discovery.

Hi everybody. Good afternoon. I hope you had a great lunch. Um you know, kind of continuing on the theme that we heard uh in the early keynote from uh my colleague Anthony. And if you got if you caught Dean's talk about OSS and open science, uh we're going to go into a little bit more detail today on one of the projects we're working on, something called the OpenADK framework. Um the goal of OpenADK framework is very simple. It's to solve some of the problems that we saw in terms of the exchange of data, the the reusability, um a lot of the the neat and fun things we know about open source and apply it to science uh overall. So, let's get started

So, um accelerated discovery is an interesting term, two interesting words together, but it means very something very specific within the context of the work we do at IBM Research. Um every year, we have something called the Global Technology Outlook. We put together a very fun glossy paper for um you know, the market and our internal teams. And a few years ago, we won on the idea of accelerating materials discovery. So, what does this mean? Um materials discovery can be anything from uh things like replacements for PFAS or forever forever plastics materials. It could be better photoresist materials so we can make higher fidelity um integrated circuits and gates. It could be um the idea of uh carbon capture uh tools and and facilities uh and carbon sequestration scenarios. Um all really interesting fundamental applications of how we're working through uh this idea of improving our life overall

This also happens to apply to um the identification of molecules for use in small molecule drugs. The uh materials that you might use to improve the delivery of drugs, things like higher temperature uh vaccines for COVID, for example, lots of other options in that category. Um and that's how it works. So, we did a paper talking about how we could apply a couple fundamentals of uh AI, quantum computing, a few other components to those areas with the objective of not just offering these technologies cuz we've been doing this for a long time. Uh part of the some of the people on that list uh from the paper built Summit, one of the early supercomputers. Uh we've got a lot of experience with high-performance computing and a few other pieces, but we wanted to take some of this technology and drive more availability, more accessibility to the general population overall cuz we really think the way to go, just like open source, uh with open science is to get as many people involved as possible so we can start to share and execute on targets. So, we looked at four fundamental areas. We looked at um the idea of what we call deep search or the analysis of the the corpus and body of technol- of of information available

Uh most molecules that you want to use for materials or drugs are actually packed into papers that are shoved into uh repositories like archive um and you know, it's really hard to tell who owns what. The other half of them are put into patents and patent databases uh as part of that story and you want to really understand A, what's available, B, what you can use, and C, how does it relate to the problem you're you're interested in. So, how do I relate properties to a specific molecule in terms of how that works. And that's what this DS4SD is. Uh ST4SD is a simulations technology which is uh fundamentally centered around the idea that um when you're screening whether a molecule works or not, you have to simulate it either um classically, Newtonian simulation, quant- um uh quantum mechanically or otherwise. So, you have a good sense of whether it will actually work in a situation. And that includes, you know, things that you would probably not initially think of, like um making sure that it's stable in uh 960° K water, meaning you um and other aspects of how things work in this category. Um and that's really a powerful way to do that

Um we also built something called GT4SD, which is eminently relevant um you know, now 5 years in. Uh GT4SD is a good example of early applications of generative uh models and the ability to essentially drive to a property driven outcome for this. So, what happens here is you input a set of properties. I would like something that is, you know, conductive or non-conductive and stable at this temperature target and a variety of other categories. And what it will do is it is cycle through and just like an LLM, generate the next token in terms of the overall molecular structure necessary to define a candidate molecule as part of that story. Now, this is where things get interesting. Um you know, a lot of people will take library like GT4SD and, you know, put in a property like toxicity. And all of a sudden, you're generating, you know, 40,000 candidates for a biological weapon

Reality is is most of them are not stable and probably are really hard to manufacture. Uh going back to the some of the comments earlier. But, you can, um you know, go after properties of both um benign and malicious in this category. And then finally, uh one of the key elements of testing, if you will, the manufacture of a molecule is a process called synthesis or uh in the case of how you want to if you want to figure out how to make the thing that you've just come up with, the idea of retrosynthesis uh as part of this. So, this is literally the path of a reaction, a state change within a set of molecules that combines the molecules to create a new one. Or the the reverse engineering of that process overall. Uh we also happen to to connect that to a robotic lab. So, in uh certain configurations and targets, you're able to essentially prototype the molecule once you've defined and validated it uh from that flow

So, whole idea here is all these things are driving acceleration through what is typically uh a bench-driven process from a chemistry perspective. So, chemistry, as I mentioned, is the study of properties and the behavior of of matter. Uh in many cases, our objective in application of technologies to this area is not to change any sort of viewpoint around that, but to open up new opportunities to uh that perspective and how it works. We're We really feel like this is a a worthwhile activity, and uh in many of these cases, we're driving in that direction. So, just to give you a little bit of a read of what I've just mentioned, uh we have the idea of a set of reaction conditions, which is the input set, if you will, that generates a molecular structure that has certain properties that you're interested in. And within that set, you also have a set of target properties that you want to express uh through that molecular structure. So, there's a constant feedback loop in terms of evaluation of both directions of this category. So, this is it synthesis retrosynthesis, um and then reaction prediction and property prediction are really the key to this game overall in terms of what that looks like

Now, that sounds like a great and opens ways to interpret how this works, but the reality is is that uh chemistry, like almost every other science, has its own biases built into it uh in many cases. So, um we are really comfortable with the idea uh of um you know, for lack of a better term, flat um molecules. Uh because we've had the language to describe them, they're easy to configure, and you know, certain times of um cross-coupling is very easy to do. So, most of the drugs we make are flat. Not because there are non-flat drugs There aren't many non-flat drugs. It's just we That's the area that we've been scoped in terms of our evaluation overall. Uh which means that there's a whole universe of non-flat things that could be really helpful, but we have to move past the prior art because the prior art has been, you know, essentially constraining the the search space that we've been using to define what we're targeting uh in this case. And this is just another example of where novel AI applications that are unconstrained in the sense that they will don't they don't care about dimensionality

In fact, they thrive in it. You've got a great opportunity to do some new things that otherwise you wouldn't be able to as part of that. Um so, we're really excited to push in that direction. But, there are some interesting barriers as well. Part of that bias in terms of understanding and interpreting the idea of uh uh representing what those molecules look like is also constrained by the data and language that we use, right? So, uh we use a bunch of text-based uh formats like SMILES, SELFIES, um InChI, a few other categories, which are very common ways to describe a molecule. But, in many of these cases, um you can describe a molecule using many different versions of SMILES or SELFIES itself. Um you could describe um different molecules in a similar fashion with this. And most of the the the structures or or natures of bonds that make up molecular structure are poorly described in, again, a two-dimensional uh interpretation, which is essentially where text comes into play here as part of it

Um you know, there's wonderful things like radial bonds in hexagonal structures that distribute uh charge state uh quantum mechanically across all of the the the the connected molecules that can't really write really well in a in a sequence of letters, unfortunately. So, the real question I think moving forward for us is how do we take what we know about the world and how it works both, you know, in a pure sense, probabilistically, quantum mechanically, and express it in a way that drives those those characters. Um not not too different than how do we express a specific form of logic or an operation in code? How do we describe the relationship of data in this category? We need to start evaluating and building out um some better understanding of what that looks like. Which I think, incidentally, is what we're really struggling with overall uh right now in this age of uh new AI. We've been working so hard for the last 40 or 50 years on encoding human consumable consumable data in a machine optimized format, you know, e.g. putting stuff into a database, uh for example, and now it's become possible for us to take encoded machine consumable data and store it in a human optimized format, meaning put it into text. I posit that the folks that take this kind of challenge for right now and skip the step in the middle, which is an encoded representation, uh will be the people who win this little battle, right? We have a great space uh in terms of a capacity to describe how things work, um and it is one of the most commonly used uh you know, essentially descriptive of tools we have, which the language I'm speaking right now, it's English. It's German, it's French, it's Spanish

It's a lot of the the fundamental formulas human languages or model languages that you can use to describe things. That is the embedding space. That is the con- the area with the by which we should be able to describe and effectively drive uh interaction and change as we continue to grow. We're going to need intermediate representations like uh programming languages in the short term to prove that we know what we're doing, but I think the ultimate goal is to push past that step to the point where what we describe is what we is what we create uh ultimately. So, how does this relate to the Open AD framework? The Open AD AD framework uh came to us because that just like everything else, there's never one way to do something. Uh we have folks who are bench chemists, meaning Erlenmeyer flasks and goggles and gla- and gloves. We have folks that are pure computational chemists uh that only work in a synthetic environment. Uh we have people in between that are trying to balance between these modalities in terms of how things work

So, we had a lot of different audiences to think about when we were putting together a framework to help with that acceleration overall. Um we think that uh by pulling together a single tool set that has many different faces, we're going to have an opportunity to get people to collaborate and drive things that we often take for granted uh in this situation, especially in the case of um you know, open source software, like repeatability. Repeatability for an experiment, you know, sounds like an odd combination of words to begin with, but it's a really important one ultimately as part of that. So, the Open AD framework uh gave us a couple of important uh you know, opportunities to correct for some of those initial lines of thinking and to move more toward a repeatability target in terms of the interpretation and gave us a great chance to test some of these new technologies in terms of the capacity to uh assistively and adaptively describe the work that people are doing. So, we put together a simple single API that feeds a CLI, um an API, and some notebook extended services because scientists love notebooks. I don't, but they do, so we use them. Um we integrated all the view of those pieces together, and we added tools that took things that were better expressed as visual context out of the 2D realm and into the 3D realm in terms of molecule viewers, runs, LLMs, etc. Um and then we consolidated all the illities like security, data access, uh collaboration, and put that into one single flow so that you don't have to worry about what it where it is or where you are and what it looks like

And then we made it extensible. So, if you come up with a novel form of simulation or a novel form of there candidate validation, you can use all these facilities to essentially extend and drive those pieces. And then we added the capacity to record, store, and manage uh experimental workflows, right? This is really important. You know, what's the best way if you don't have a predictive uh you know, path through a specific experiment? Record what you did and eliminate the parts later that aren't relevant or you believe aren't relevant to the overall experiment. That's part of that story. So, that looks essentially something like this. Um we have a notebook-based framework that lets you do that. We have the CLI side of the story

And then we have uh a strong DSL that we use to express uh workflow steps in a very clean and organized fashion. Like, I would like to get X from Y. Um rather than worrying about another pragma or or structure that you might need in this case. Now, this is the point where I wanted to go to a demo, but I am happily right now working on Dean's machine cuz mine decided not to not connect to the projector. So, I'm just going to walk through uh the demo backup I've got at this point. So, a great example of this is this idea of being able to pull out and have what I what I think of as self-describing or self-documenting set of commands that might be appropriate in this situation. So, for each of the frameworks that gets integrated, we essentially use an LLM to vector embed the base documentation uh into each of the segments that are that are here. So, if you wrote the docs, this system can answer questions straight from the docs in terms of what the function would be

See, any any you can combine a set of functions to do this. So, this is coming out of uh a Llama 2 Chroma DB combination um that we've put together that sits alongside the the standard software. We really like that idea because you can write something once and make it a usable user functional uh characteristic. Uh in this case, we're going to play with deep search, and we're going to look at and list uh some properties that are coming from some patents in terms of what those targets look like and how it can work. So, that usually constitutes searching for a set of paper through a set of papers for a specific molecules that might be relevant to you topic areas. So, in this case, we're going through some archive abstracts, and we're extracting um you know, some interesting information about power uh power uh conversion efficiency uh and a few other arguments, and we're pulling those out. This type of viewpoint comes back to the user either um in the notebook they're executing or uh through you know, basically a a Flask app that we kick up that goes straight into the web browser from the command line, which you can also invoke through the API. Um and all those tools and capacities are generalized so that um if you're a scientist and not a developer, you have a great opportunity to just use what's available to you give people immediate to how those tools tools work

So, say we find a property that we really like, and we think, "Hey, that's great. Uh let's evaluate what that looks like from a molecular structure perspective, um and let's find some similar molecules." What often happens in chemistry is that the great idea that you have, it's good chance, especially if you're working in 2D chemistry, someone else has come up with before. And not only have they come up with it, they've patented it and put it into a database at, you know, their their company just in case they need it later. Because that's how patents work. And puts you in a place where you might have a desired property or desired structure that needs to be expressed differently in order to essentially characterize a molecule that you can own if you will as part of that stream, which is a whole other problem in my opinion chemistry. We'll leave that one alone. Um the great way to do that is to do a molecule structure similarity search against this this database. Get a good look at what that looks like, how it works

This system itself isn't that complex. It's essentially leucine on top of elastic. But we've built some very very specific parsing capacity for for use in in this domain. Um a lot of that, to be really honest, is starting to be replaced by, you know, some of the core features in LLMs. But we're we're going to continue to mix that in in as it as we go, as we current we continue to develop it overall. So, once we have a set of candidates that make some sense, we've got a good view of what that looks like. Um you can go ahead and pull up a set of molecules, evaluate which ones are immediately interesting to you. How this works, you can kind of play with the representation and targets

Um and you can do that. So, we also built in not just the documentation view of how that works, but a capacity to combine commands through a little bit more of a complex embedding scenario here. So, we have this thing we call the the how-to assistant. So, you can ask Open AD how to do something specific like identify a property and then evaluate it in terms of a potential synthesis path. So, you can combine the tools that are that are in there, so you can quickly establish that. And if that one specific command doesn't work, no problem. Remember you're recording this. So you're going to be able to essentially extract the the bad value or augment it later if you need to

Uh but more importantly, you have a record of what you did, so you don't have to remember. In this case. Um so, this is a great example of just a direct output uh that gives you a clear visible understanding of what the tool and framework is. And this is nothing more than, you know, some chunking and uh vector embedding. This is a LangChain app uh that we added to this as well. So, um once we have a good sense of what that looks like and how it works, uh we can grab that molecule. Uh most chemists want to see a molecule cuz the the the structure is uh something that they remember like, you know, the texture of your favorite uh keyboard uh in many case. It's a very comfortable and familiar area and visual verification is pretty important, so we definitely show people what that looks like

We also happen to store it as a data frame behind the scenes, so if you want to play with it right away in terms of manipulation, it's it's pretty much directly available. And once you've got that going, you can uh essentially drive a prediction of the retrosynthesis process that's available from that perspective. So we take that molecule in. We show it to you in 3D cuz as I mentioned before, it's probably pretty important. Um and we evaluate the the output and push this in. So, right now, we um we do a couple things. One, uh we use a a generic property prediction to chemical output scenario, which is essentially nothing more than a next token problem, but we're starting to do what we call state change interaction graphs, which are starting to describe better ways to evaluate things using something called method method space chemistry uh to help do this. And we're doing that um combined with um you know, a framework called Ray, which you might have heard of several times today already, um, to be able to dynamically scale the cluster that executes it based on the scale of the molecular processing, uh, as part of that story

So, it really helps a a ton to be able to, uh, scale those functions. We off- often use, also, another interesting tool you that you might have heard about if you were downstairs, called Dapper, to handle all of the common security and, um, you know, secrets, uh, access management, uh, and other structures as part of that, uh, to simplify the the the service deployment in those cases. We think that kind of combination of things that, you know, have a lot of, uh, scale from zero uh, base properties really gives you a great opportunity to do some, uh, solid, uh, runs in terms of those characteristics. So, great. We like that molecule. We put it through, uh, the RXN retrosynthesis tool, or, sorry, synthesis in this case, um, and we come back with, and unsurprisingly, there are many different ways to perform synthesis for a specific molecule that you need. Now, um, why do we need so many of these? Some of these are practical, right? Some of these might require vacuum to be performed. Um, we've got to get get a get a good sense of what the constraints or targets in that case looks like and how it could work

These These are This is where you're moving more out of the scientific discovery phase and evaluating evaluating the overall cost or value engineering necessary to produce the molecule itself. Um, so, I imagine I would pick two or three of these. They They make sense to me, and I might say, "Hey, I'm going to send this synthesis synthesis path to RoboRXN." Or, you know, if you're if you have a lot of money and you're working in an industrial lab, a ChemSpeed or something like that, which is a a real a real, uh, industrial-grade lab, and perform some prototyping activity. Once you get that back, you have a good sense of what it does, what the constraints are, where it is, and what's in what's it going to cost to to build, and can you scale it? Which sounds to me like acceleration overall, and that's really where we're headed uh in this case. Now, I didn't say a lot about simulation because right now in many of these cases the advent of generative uh candidate identification I mentioned earlier, GT4SD, has done really interesting things to the profile of compute that we're applying to these uh these problems before. Um m- most chemists or uh medicinal chemists would evaluate a couple key molecules that they thought were sure bets, and they would go in and do things like docking simulations or uh in situ molecular dynamics, which are really expensive um high-performance computing operations. Uh but the way things are moving right now, you can generate 40,000 candidates and evaluate their their synthesis profile and quickly determine if there any of them are practical before you spend any of the money on high-grade uh simulation. Um and the high-grade simulation at the top end is really where some things things like quantum computing come into play uh as well, right? So, you may do something simple like a single operator validation of the polarity of a molecule with a quantum system, or you might look for small molecule observables, uh thermodynamic observables, which are really, you know, high-end problems to deal with just to be super clear uh in this case

But, you know, again, these are all things you can plug into this open data V framework and quickly and and and rapidly establish their their relevance to your overall molecular discovery workflow as part of the story. So, that's really where we want to go to. Um Oh, by the way, uh part of the synthesis path output is the recipe. So, stir this many times, this much in this situation uh as part of that, because hey, humans make this stuff too. We're not always We don't always have a robot around to do it for us. As sort of it from that perspective. And then, you know, we have forward prediction for batch and a variety of other scenarios. And all these tools, by the way, also work straight out of a CLI as well

Many HPC technology um aficionados, if you will, use things like Balsam and other shell-based workload tools and we need to stay compatible with this. So, there's nothing wrong with embedding a an Open AD call directly into your existing HPC workflow as part of that story. In fact, I think that's an ideal state. And, you know, the same fully integrated help works in in the in the command line view, as well as a direct API use as part of that story, too. Um as well as all the molecule viewers and inlines. You can kick those off straight from the command line, as well. As part of that story. So, the whole idea here is we want to take care of all the software-related um repeatability and accessibility concerns, so you can get down to to work and really drive the scientific output that you need using the fastest tools that you have available in any given time

As part of that and um pushing those directions. So, I wanted to share this with you because I think it's a really great intersection of open science and some of the key principles of AI and applications within this time frame. And you know, kind of embodies that spirit of exchange of information between open source software and open source science. As part of that story. So, any questions, folks? I think I'm on time, yeah? Right. Yeah, let's do one question after you. Uh you mentioned in passing that quantum computing is starting to be applied here. Maybe you could talk a little bit more about what people are doing with them in this context

Sure. Sure. So, um quantum computers computing is interesting, right? Because it's a probabilistic evaluation of certain states being uh driven in this category. Um most of those states are described uh in something called a Hamiltonian, which is a a perfect um definition of an energy state. Um and that's how, believe it or not, uh electron electrical qubits work. So, there's a very natural correlation between um articulating the energy state of a molecule and articu- articulate the articulating the energy state of a of a combined or a a set of entangled qubits. So, a lot of the early applications for quantum computing have been chemistry-related as a result. Um but, the problem with that is that we only have so many qubits entangled right now

So, and if you get more than like 10 or 20 or 30 together, you end up in this wonderful universe of distributed quantum computing, where you actually have to register the the superimposed qubit values across a classical computer and then perform the stochastic or essentially repetitive options to get an opinion as to what the state is. Uh so, there's lots of extra compute that goes into a hybrid quantum-classical calculation of an energy state as part of that story. Now, in a perfect world, we take all that mumbo jumbo I just said, we put it behind an API, and we we have the system do it for you, but you know, that's kind of where we're headed currently. I I think that's our time, but uh the Q&A uh table number two outside. Uh uh thank you very much. Thanks.