Is it forgotten junk, or a goldmine for tomorrow’s drug development? On this episode, we tackle the untapped potential of legacy data and what it really takes to turn that backlog into a competitive advantage.
The spotlight today is on Nathan Lewis—GRA Eminent Scholar at the Center for Molecular Medicine, Complex Carbohydrate Research Center, and Department of Biochemistry and Molecular Biology at the University of Georgia. A computational biology trailblazer, Nathan’s career journey runs from Sunday afternoons with PBS nature shows to the frontiers of genome-scale modeling, AI-powered analytics, and innovative bioprocessing strategies. His research bridges practical cell factory design and fundamental advances in how we make biologics.
Episode Highlights
- Rethinking the dogma: controlling protein glycosylation quality from the inside out, not just by bioprocess conditions [03:00]
- Nathan Lewis’s journey into science and bioprocessing, from unexpected college choices to pivotal advances in CHO cell engineering [05:12]
- The evolution of omics in bioprocessing: why actionable insights, not just big datasets, should be the goal [10:49]
- Strategic advice for structuring, annotating, and making old and new datasets ready for AI and LLM analysis [17:34]
- The balance between mechanistic and machine learning models—when each makes sense, and why hybrid modeling is gaining ground [22:20]
- The current and future role of digital twins in process development and why foundation models and data consortia matter for scalability [26:24]
In Their Words
This is where I think the huge innovations have been coming with a lot of agentic workflows. Now, we’ve been developing agentic workflows that can go through and go through old ELNs, spreadsheets, stuff like that. And based off of the notes that people have put on them, we can start to guess what the information is, what sort of structure it should have, and be able to take that, pull it all out of the legacy data structures.
Podcast Transcript
David Brühlmann [00:00:40]:
Somewhere in your lab, a hard drive holds omics data nobody has touched in years. Sequencing runs, proteomics, metabolomics, collected once and forgotten. Nathan Lewis thinks that’s a missed opportunity. Nathan Lewis is a GRA Eminent Scholar at the Center for Molecular Medicine, Complex Carbohydrate Research Center, and Department of Biochemistry and Molecular Biology at the University of Georgia.
He is an expert in computational biology with extensive experience in the analysis and design of cell factories and biologics. In this episode, he explains why that dormant data suddenly matters and what it actually takes to make it actionable at the bench. Let’s dive in.
Welcome, Nate. It’s good to have you on today.
Nathan Lewis [00:02:43]:
Thanks. It’s great being here.
David Brühlmann [00:02:44]:
So good to catch up again, and I’m very much looking forward to our conversation. And to start us off, Nate, share something that you believe about bioprocess development that most people disagree with.
Nathan Lewis [00:02:59]:
So we work a lot in different areas of bioprocess development, systems and synthetic biology, AI, genome editing, all that sort of stuff. And there’s a lot of really great tools out there and a lot of great things that can be done to be able to control the quality of a lot of these drugs that we’re developing. I would say one little niche area that we have been working on that I think actually has been kind of surprising to us that we were able to achieve, which kind of goes against the conventional wisdom, is in that one little niche area of product quality control that is looking at glycosylation, given that these glycans are on the majority of biologic drugs. Everywhere you read in a paper, it always says glycosylation is a non-template-driven process. That is, it’s all about the state of the cell that controls your product quality. We hold that that’s not totally true. We’ve been able to, through some innovations that we’ve been able to do, we’ve been able to identify a code in the sequence of any given protein that can control, at least provide a very pretty strong control on the product quality attributes themselves.
So it means that not only can we control the quality of our products outside in by controlling the bioprocess or even controlling the cells, but we can go in and we can actually control quality inside out, thinking about the protein itself. And so that’s one little niche area that I think that has been a little, been surprising to a number of people, but it seems to be working for us. So kind of changes the paradigm somewhat. I still think that the heart of what people are talking about when they say it’s non-template-driven and it’s in the sense that there is this external control on product quality. It just means that’s not the whole picture. We can control it from within. So that’s a small area amongst the many areas that we work in.
David Brühlmann [00:04:53]:
This is a great teaser for what’s coming up a bit later in our conversation. Let’s focus first on yourself, on your story. Draw us into your long career. What drew you into science and what were some interesting pit stops along your way?
Nathan Lewis [00:05:12]:
I mean, if you go back really far, I was always into science. I grew up in a home where on Sundays my mom would flip on PBS and with the nature shows, and we grew up really enjoying that stuff. But when I was in college, I was not a science major. I was not headed that direction until I had to take a chemistry class, loved it, and eventually found myself studying biochemistry and then going towards bioengineering for my PhD. What really drew me into where I am now started when I would start looking at scientific data and I realized how complex everything was, how much everything interacted. And I realized that to make sense of this, we’ve needed to be thinking about all of these biological processes as complex systems. And that’s what drove me to bioengineering was like, I can’t do this by hand. I need to have engineering capabilities and a lot of math behind it. So that drew me into systems biology.
And then as a PhD student, as I was studying the evolution of bacteria, thinking about all the metabolic pathways and stuff like that, I had to pay the bills. And so when somebody reached out to me and said, if I’d be willing to apply it to these hamster cells, Chinese hamster ovary cells, I was like, I don’t know what they do, but sure, I’ll try it. And I was absolutely fascinated when I started thinking about the theoretical stuff that we were working on during my PhD, with the idea of just being able to apply it to cell lines that are used for making biologic drugs. I was just fascinated. I was like, okay, this is no longer just a cute publication in a journal. This is a translatable technology with time that we’ll be able to get it to basically help people. And so that was a huge draw for me into the field where I am now. And so that was the start of it all.
And eventually we landed a really large grant to be able to do this for a long period of time, thanks to generous funding from the Novo Nordisk Foundation. When I was working at that startup, the standard procedure for cell line development and engineering, it was basically knock in a gene and put some sort of stress on the cells and try to select for a clone that produced something at high levels. It wasn’t thinking very much at all what’s going on inside of those cells. It was just like, can we pull something out of this, this heterogeneous mix of cells? And as an engineer, you start thinking about, well, how do we engineer better products? You have to have like an understanding of all the parts of your system. You have to have a wiring diagram of how all those parts are connected, and you have to have tools to engineer that system. And you have to have on top of that a unifying language, which is oftentimes mathematics, on how to control that system. We didn’t have very much of that at all for CHO cells. And so that was one thing we did at this startup that I was consulting for is we convinced BGI to sequence the genome of CHO K1 cells and then the Chinese hamster after that. And that gave us the parts list.
And so with that, we saw this, we have this whole parts list now. We could start building up the wiring diagrams of the cells. Be it the metabolic pathways that are controlling growth and protein production, the protein secretion pathway, glycosylation, apoptosis. We were able to go through and start looking at all these complex systems inside of the cells, and that gave us ideas on how do we engineer and control them. The problem was at the time, there was really no good genome editing technologies for CHO cells. The best we had was zinc finger nucleases, which were really difficult to use. You’d go one at a time. And then it just so happened that one day a friend of mine was in my postdoc, said to me he had this crazy idea to use this bacterial immune system to edit DNA. I was like, good luck on that one. Two months later, he was writing his paper for Science, and that was the introduction of CRISPR. And that paper came out literally a few days after I had started my faculty position. So that put us in a position where we had suddenly the parts list for CHO cells, we had the networks you needed to understand the complexity of the wiring diagram of the cell. And we had algorithms to work with that. And then now we had a tool to edit it. So that’s what brought me into this area where I am now. And so where people are doing this, a lot of people are going through and using these sorts of technologies to improve bioproduction.
David Brühlmann [00:09:37]:
Yeah, this is very exciting. And a little disclaimer, I’m a big fan of Nate’s work and your collaborators. I remember multiple ESACT conferences and cell culture engineering conferences where you presented the genome editing, you presented omics and glycosylation. And many, many years ago, I was doing my PhD in glycosylation, so I was very fascinated. Let’s address the elephant in the room because you’ve done a lot of modeling. And I remember a conversation during my PhD time when I was a PhD student when I had to justify doing omics because omics is great, you get a lot of data, but then the question is, what do you do with it? And that was the very conversation I had with my boss at that time. Like, does it really make sense to spend all that time and money to generate that much data if we can’t analyze it? So I would be curious, Nate, what is your perspective on that? Because there’s a lot of omics data sitting in labs untouched. Why do you believe today is it suddenly worth revisiting? And how have perhaps maybe the new technologies changed that?
Nathan Lewis [00:10:49]:
Yeah, exactly. I think that’s an excellent question. And it’s something I always ask myself too. I jokingly say, being in academia, you always want to do omics because that increases the impact factor of the journal you end up publishing your paper in.
And I think that’s when you’re— I jokingly say that, but there is some level of truth to it. At least in the past. However, what we’re doing here is we’re trying to develop technologies and trying to gain insights that will be useful and be able to benefit people’s lives, be able to hopefully cut down the costs in the manufacturing of these drugs and improving the quality of them. And so for a number of years now, I’ve been talking with a lot of people about, it’s not about doing omics, it’s doing actionable omics. The idea of where we can generate this data, but we do it in such a way that it guides us to make a decision to improve a process. Think of, for example, what many of the tech giants have done over the years. Google, Meta, all of these companies, what they’ve been doing is they’ve been generating tons of data. And what do you do with that data? Early on, it wasn’t immediately clear how they could really make it useful.
But through continued generation of the data, proper curation of it, and development of the right models to analyze it, they’ve been able to make all sorts of predictive frameworks. So now in bioprocessing, we— I have a brother who’s an econometrics expert, having worked for a number of these tech giants, and he laughs at me saying that we were working with small data in bioprocessing and in biology in general. And he jokingly, ever since we were both grad students, he would say, he’s like, I don’t believe anything you guys do because your number of samples is too small. However, we’re accumulating a lot of data. And now the key here is how do we turn all this data into actionable insights, actionable omics and stuff like that?
It comes down to a few key principles. And usually I lay them down into 2 or 3 different areas. One of them is identifying those data types that are most actionable inherently. We sequence the genome, incredibly important resource in my opinion. But if people are thinking about actionable omics, one of the easiest omics to generate is a genome sequence. You go through, you take your cell line, sequence the genome. But what do you do with that? You can find a few mutations, but does mutations really give you anything actionable? Oftentimes not. A couple of times we actually have found mutations that were causal and even in CHO cells, we’re able to tie that to a phenotype and engineer the cells. But oftentimes it’s pretty difficult.
Then you could go to something such as fluxomics where you are actually tracing the metabolite flows through the cell and connecting it to cell growth or other properties. And that is a lot more actionable because then you can figure out, well, then I can dial the metabolites up and down, but it’s more expensive. It’s harder to do those experiments. And so there’s this balance, this trade-off between how actionable they are inherently and how expensive and how difficult they are to generate.
Then you have the extreme where you have ones that are not only actionable, but interventional. And these are going to be, for example, pooled CRISPR screens or other things like that where you’re going to find a target and you’ve basically selected for that trait and you know exactly what to do with that cell line, potentially other cell lines too.
So when we’re thinking about the data that we’re going to be generating, you want to be thinking about the data that first is most actionable. It doesn’t mean we don’t do the other data types, but it means there has to be something more if you are. And that’s the second one. You find the right methods to analyze it.
And so there’s this whole spectrum of different types of models. You can go with detailed mechanistic models. You can think of series of ordinary differential equations that describe metabolic pathways. That’s going to be very, very— it’s going to take a long time to build those models, a lot of thought and put into it. But that’s going to be one class. Then you have all the way to the other spectrum where you have machine learning and AI models where you don’t really know much about what’s going on inside of the system, but you have an input and an output that you’re able to predict and, or at least connect. And then you have the middle where you can bring in models where they’re mechanistic, but you put on layers of machine learning and AI for some of these hybrid models. So you have the spectrum of models.
And so what you want to do to make your omics actionable or any of your data types actionable, you find the right types of models that allow you to gather that information. And I think going back to where I was talking about the big tech giants, they were able to develop modeling approaches. My brother helped develop some of these theories in econometrics that figured out how to take a lot of the data to be able to predict ad clicks. And now we could start thinking about, though, in bioprocessing, we can take all the different data that’s been generated over the years. If we get it into the right context, the right structure, we can figure out how to get it to improve productivity or to diagnose failures and so forth for root cause analysis, things like that.
And so that was the second thing. So the first thing was make sure you do the data types that are most actionable. Second one is going to be finding the right modeling approaches. The third one is kind of prospective study design. A lot of people, they’ll take their— a few samples, they’ll have maybe 3 bioreactors and they’ll run some stuff and they’ll compare a high and low producer. That on its own is not all that useful in the future. However, if you set it up so that you have proper bridging samples and then you not just randomize datasets, but if you— large panels of samples you’re running, but you properly balance for covariates and stuff like that, you can take these datasets and stack them on each other and so that you avoid batch effects, so you can eliminate a lot of that stuff. So that’s proper study design is the third one really.
And so if you take these 3 together, we find that we’ve been able to rescue bad experiments. We’ve been able to gain deep insights, mechanistic insights into the systems and so forth. And that allows us to get much more actionable insights into it.
David Brühlmann [00:16:52]:
You’re making an excellent point, Nate, because the purpose of the dataset finally is to derive some actions out of that, especially when you are developing a process or you’re going towards more commercial production. You want to use the insights to make some kind of decision, but that’s exactly where I see a lot of people struggling, even in quite established large companies where there’s tons and tons of data, but they’re either siloed or they’re not structured well. What advice would you give these people? Where should you start to actually make these datasets actionable? I’d say more on the probably very practical side to be able to use these datasets?
Nathan Lewis [00:17:34]:
That’s a very important question. And I think it’s something that a lot of people are struggling with, but there’s a lot of, I would say, over the past, especially this past year, there’s been so many advances that making it such that this will be less of a challenge. And that is given, so we, like many other people, have been jumping on the LLM bandwagon. These are incredibly powerful models that can be used. But what is absolutely critical is that the reason why they work so well is because there’s been structures put behind them to take unstructured data and give it structure.
How do we do that? How do we make our data AI-ready? There’s a number of things that need to be done along the way. One of them is you have to have controlled vocabularies. So what we need to do as a community, we need to go through and very clearly define that when we say a critical quality attribute, name your favorite one, whatever it is, that there’s a clear definition to what it is. And so that when we are doing our experiments, we are properly describing things.
You need to have minimal information standards that are put in place that when everyone’s doing an experiment, they’re going to provide information. What’s the experiment that’s being done? Who’s doing it? What was the temperature of the culture? What was the pH? Where you’re tracking as much as you can and then pulling that into a proper structure, into forms that can be easily machine-read.
The beauty of these LLMs is that they can take some not-so-great structured stuff and bring structure to it. But to do that, you need to have ontologies put in place. In other words, taking the controlled vocabularies, but finding connections between—if something is a cell line, is it connected to this term CHO-K1? Or is an impeller part of a bioreactor? Things like that where there’s different terms and there’s connections to them.
And then you have to have schema that gives you basically information about how to gather and where to put that, the data. As I’m saying this, you can already think like, oh my gosh, I do not want to be the person that has to go through and annotate all the datasets. Especially if you’re in an organization that has decades of data, nobody wants to go through and even touch the data from the last person that was in the organization, right?
This is where I think the huge innovations have been coming with a lot of agentic workflows. Now, we’ve been developing agentic workflows that can go through and go through old ELNs, spreadsheets, stuff like that. And based off of the notes that people have put on them, we can start to guess what the information is, what sort of structure it should have, and be able to take that, pull it all out of the legacy data structures, and populate in new structures, thus taking the unstructured data and getting it structured so that the LLMs can then go through and you can talk with your data.
And so that’s something we’ve done. We’re not the only ones doing it. There’s some companies that are doing stuff like this where you’re now able to take your bioprocess data and chat with it as if you were talking to Claude. And we’ve been absolutely surprised. We were making a tutorial video of our tool, for example, where we took a bunch of single-cell, or a bunch of, no, it was bulk RNA sequencing data from a number of bioprocess runs. And as on the fly while my research scientist was making the video, I was sitting there next to her and she said, I wonder if we could do like a GO enrichment analysis. She’s like, I haven’t coded anything like that into this. She types in like, do a GO enrichment analysis and rank the different pathways that are most associated with productivity and so forth.
And in real time, as we’re recording, I think it was within a minute or so, it’s sitting there saying, thinking, thinking, thinking. And suddenly a GO term enrichment analysis pops up and it was legitimate. And we looked at the code and it was really, it was real. Like, and so with this, you can now start to talk to your data and analyze it in real time without having to do the analysis that took me to code something up like that as a graduate student. It took me days to go through the theory, to understand everything and then to go through and code up all the statistics that would’ve then also scraped all the databases online to get the information on it. So now it can do it in real time.
David Brühlmann [00:21:47]:
Yeah, this is very powerful. I’d love to have your view on that because you have done all kinds of modeling. You have done mechanistic hybrids. You’re using LLMs now. Can you give us perhaps just some guidance on when each model makes more sense? And I think there’s a lot of hype also around AI and what AI can do and maybe cannot do. And perhaps also give us your vision about where you see this going in the next couple of years.
Nathan Lewis [00:22:20]:
There’s a lot of deep questions about and things to think about as a community. I think some of the challenges we face with LLMs, with a number of these AI models, is their opacity. We oftentimes don’t know what’s going on inside them. As I talk with my people in my group, I always tell them, like, we use these models, but we need to verify. We need to go and understand what is the code that is generating under the hood? What is the source of the data that it’s relying on? Is it just ignoring datasets that disagree with the statement that it’s going to make? And so that’s a challenge we face with AI. And of course, there’s a lot of great work being done on explainable AI and so forth to tease it apart.
I think, though, and this is where there’s been a push for hybrid models, we have decades of amazing work that scientists have done to tease apart the biochemical details of biological systems, metabolic networks, signaling pathways, all these sort of things. It would be ridiculous. It would be for us to throw that all away. Now, we wouldn’t be throwing it away because LLMs do read all those papers and structure things with them, but they may not obey the laws of thermodynamics, right?
When you’re thinking about mass balances of the substrates coming into a cell and the cell growth itself. And so I’ve been a huge proponent of hybrid models. We put in all the biochemical information we know and we trust into these models. We structure in such a way that it’s there, but then we have a lot of unknowns, parameters for enzyme kinetics, we have uptake rates, we can measure the uptake rates for metabolites into a cell, but we don’t necessarily know why it’s only taking up so much glucose. We’re only talking about taking up so much amino acid when you flooded the cells with a lot more. And these are a lot of unknowns, and that’s where the machine learning and the AI fits.
We put it there to fill in the gaps, but also to help us hypothesize about what those values should be, or maybe there’s a mechanism that’s missing. And so you can use it for discovery of the inner workings. The way I see it is that these tools, if you are working with a system that has been well characterized, you want to clearly have a mechanistic layer. They’re more interpretable. But you want to also use the best data-driven models on top of those to help manage the unknowns.
And so that’s going to be with the people that are really doing the complex modeling. Many other people may not have the years of experience working with ODE models or constraint-based models or agent-based models. All these different types that we can use. And so they are going to be more working with the LLMs to understand the bigger picture of the phenotypes they’re measuring and so forth.
And so the answer is we need to use a lot of them. Now, one way that we were looking at this, something we’re looking at doing is if you can set up these hybrid models that describe the system and then build the LLMs around them as a co-scientist to analyze those models and make tests, make up hypotheses, go check the literature to see what might be supported, and then present it and work back and forth with a scientist, then we can leverage these models even more efficiently, working with these stacks of models upon models as a co-scientist.
David Brühlmann [00:25:44]:
So the way you guys are working is less either/or, it’s more like and, right? You’re trying to find the best tool for the right application, and it’s in the combination of the various tools where you see the biggest difference.
Nathan Lewis [00:25:59]:
Exactly. Yeah. And it should make sense. You have— all tools are good for something, right? Or good for different questions, different— I mean, I’m not using my little laptop to do, like, supercomputer-level stuff, or I’m not using my laptop to cook my breakfast, right? So you have different tools for different tasks, and each model is like that too.
David Brühlmann [00:26:19]:
That makes a lot of sense. To what extent are you using digital twins?
Nathan Lewis [00:26:23]:
I think that we’ve been developing a number of digital twins, and this is again another area that we’ve been thinking a lot about because these can be used to really speed up stuff and cut down costs on experiments and so forth. So definitely we’re using them a lot. One of the issues I’ve had, though, is that with a number of these digital twin projects that people do in-house, they’ll take a small dataset that they have in-house. They’ll build and train this digital twin, but it may not be all that extendable to all the different deviations a bioprocess would go, or to a different cell line, or you change the media and suddenly it doesn’t work very well.
What we need to do is we need to be developing foundation models that allow people to take their small datasets and then have the information from all of these other experiments that has seen all sorts of different conditions to help inform it. So building a digital twin is not adequate. It’s really building a digital twin in the context of everybody else’s experiments.
And so that’s why one thing we’re thinking about, trying to get this funded, but we’re trying to build a community consortium through the iBioNetwork to bring in as much data as possible, from public data to train foundation models, that were then publicly available for people to go and bring into their company and say, okay, there’s all this information out there. We can now use that to transfer learning to our digital twins so that our digital twins are a lot more robust. And so that’s something we’re working on actively. With additional funding, we’d be able to move a lot faster. So as with everything, right?
David Brühlmann [00:28:02]:
At the end of the day, you do need some money to make things happen, right?
Nathan Lewis [00:28:06]:
For some reason, students and postdocs and staff want to eat, right?
David Brühlmann [00:28:11]:
The data you already have may be more valuable than the data you’re planning to generate next. That’s the thread running through this conversation with Nathan Lewis. In Part 2, we’ll keep exploring hybrid models, genome-scale predictions, and where AI actually earns its place in bioprocess development.
If you’re enjoying Smart Biotech Scientist episodes, please leave a review on Apple Podcasts or your favorite platform. Thank you for tuning in today, and I’ll see you in Part 2.
Disclaimer: This transcript was generated with the assistance of artificial intelligence. While efforts have been made to ensure accuracy, it may contain errors, omissions, or misinterpretations. The text has been lightly edited and optimized for readability and flow. Please do not rely on it as a verbatim record.
Next Step
If you found value in today’s episode, take a moment to like, follow, and leave a review on Apple Podcasts or your favorite platform—it helps us reach and support more scientists like you.
Thanks for tuning in to the Smart Biotech Scientist podcast and being part of this journey toward bioprocess mastery. For more insights and practical tips, visit
About Nathan Lewis
Nathan Lewis is a GRA Eminent Scholar at the University of Georgia, working across the Center for Molecular Medicine, Complex Carbohydrate Research Center, and Department of Biochemistry and Molecular Biology. His expertise spans biotechnology, computational biology, genomics, and systems biology, with a focus on designing and engineering cell factories and biologics. He has played a leading role in genome sequencing efforts for Chinese hamster and CHO cell lines and in developing AI-driven approaches to improve mammalian cell production.
Connect with Nathan Lewis on LinkedIn.
Further Listening
If this got you thinking about the data already sitting in your freezer — and what it would take to actually use it — start here. These four dig into AI-ready data, actionable omics, hybrid-model digital twins, and the cell-engineering biology underneath it all.
Episodes 263 - 264: Why AI and Automation Tools Won't Deliver Until Your Lab's Data Is Connected with David Hardy
Episodes 173 - 174: Mastering Hybrid Model Digital Twins: From Lab Scale to Commercial Bioprocessing with Krist Gernaey
Episodes 169 - 170: Why Your DNA Is a Terrible Disease Predictor (And How Multi-Omics Changes Everything) with Mo Jain
Episodes 77 - 78: Cell Factories Explained: How Synthetic Biology and AI Revolutionize Protein Production with Mauro Torres
If you'd rather follow the glycosylation thread, check Episodes 69 - 70: Glycoanalytics Explained with Róisín O'Flaherty
Want to Join?
Are you a CMC or biomanufacturing leader with hard-won lessons on process development, scale-up, or CDMO management? We’re always looking for practitioners with real execution stories to share on Smart Biotech Scientist.
Apply to be a guest or recommend someone:
David Brühlmann is a strategic advisor who helps C-level biotech leaders reduce development and manufacturing costs to make life-saving therapies accessible to more patients worldwide.
Hear It From The Horse’s Mouth
Want to listen to the full interview? Go to Smart Biotech Scientist Podcast.
Want to hear more? Do visit the podcast page and check out other episodes.
Do you wish to simplify your biologics drug development project? Contact Us