Bioprocess data: complex, messy, and absolutely essential. When real-world manufacturing collides with digital ambitions, how do you steer clear of hype and focus on results that truly move the needle? In this installment of the Smart Biotech Scientist Podcast, we’re peeling back the layers on digital twins—demystifying the models, the regulatory landscape, and the pitfalls, while laying out a pragmatic path for biotech innovators.
Joining the conversation is Ignasi Bofarull-Manzano, Senior Data Scientist and CMC Consultant at Körber Pharma. Ignasi brings hands-on experience from both development and manufacturing, advising biotechs on extracting value from their data and ensuring rigorous model validation—without losing sight of regulatory realities.
Episode Highlights
- Core differences—and surprising similarities—between modeling in development versus manufacturing [02:35]
- Regulatory requirements: credibility assessments, model risk, and validation steps for digital twins [05:07]
- Real-world example: How deploying an end-to-end process model led to 35% yield increase for Takeda, and considerations for ROI in manufacturing [08:34]
- Advice for startup leaders on when to invest in modeling and how to scale efforts case-by-case [11:42]
- Steps for scientists new to modeling: identifying bottlenecks, starting simple, and proving value offline before scaling up [12:26]
- The importance of understanding basic statistics before relying on AI-generated models [15:22]
- A stepwise summary for deploying digital modeling effectively in biotech [16:01]
In Their Words
First, identify what type of model you actually need. Usually, what I recommend is that even if you eventually need a complex model, you should start simple. Starting simple means building linear input-output models. For example, what is the final titer? What is the final mannosylation pattern in my fermentation? And which inputs influence those outputs? Even if you have temperature variations over time, you can initially use something as simple as the mean average temperature. So, start with simple linear input-output models.
Podcast Transcript
David Brühlmann [00:00:31]:
Welcome back to our conversation with Ignasi Bofarull-Manzano, who’s a Senior Data Scientist and CMC Consultant at Körber Pharma. In Part 1, we started unpacking what digital twins really are and how to bring order to messy bioprocess data. Now we push further into the practical side. We look at validation, the regulatory perspective, and what it takes to translate digital innovation into real CMC results.
Let’s continue the conversation. Using models in development is one thing—or in research and development—but when you want to use those models in manufacturing, you also have to address the regulatory aspects. How should we approach that?
Ignasi Bofarull-Manzano [00:02:35]:
Here, I would actually like to push back a little bit. It’s actually a bit philosophical, if I can say so. Right now, as products are being manufactured, the process already works like this.
Usually, you first perform process characterization and process validation. You generate data in the lab, then you fit a mathematical model. Based on that model, you define, for example, a Proven Acceptable Range (PAR) or a design space. Based on the results of this model, you prepare your regulatory filing, and these are the ranges within which you are allowed to operate during manufacturing.
So, already today, the manufacturing process is influenced by a mathematical model. The only difference is that this exercise is performed only once, based solely on small-scale data, which could actually be worse because, ideally, we scale up using dimensionless numbers, but we know there are always scale effects. Once the control strategy has been defined, we generally don’t modify it anymore. That’s how we’re used to manufacturing under GMP.
The concept of digital twins—especially the end-to-end digital twins that we’re implementing in manufacturing—is essentially based on the same type of models. Simply put, although there are some limitations that I can mention later, we take the models that were originally used during process validation for individual unit operations, where each unit operation is modeled separately, and first concatenate them into an end-to-end model because we’ve shown many times that this is very important.
Then we deploy the model in real time. That means the model receives real-time process data, makes a prediction, and—with a human still in the loop—that prediction is fed back into the manufacturing process. So it’s actually the same type of mathematical model.
I know that when it comes to GMP manufacturing, there’s a great deal of caution, and I think that’s a good thing. However, in my opinion, the worse approach is doing nothing.
In other words, we perform experiments at lab scale, fit a mathematical model, define operating ranges or a design space, and then, even if we observe process shifts during manufacturing because of scale or other factors, we continue using the original design space without adapting anything.
For me, that’s worse than using the same type of models but deploying them in real time to support risk-based decisions, while always keeping a subject matter expert in the loop. So, conceptually, there isn’t such a large difference between using models during process development and using them in manufacturing.
From a regulatory perspective, however, there are additional requirements—not because of the model itself. We need to perform a model credibility assessment, but because the model interacts with the manufacturing system, all the interfaces also need to be validated according to GAMP procedures. That’s mainly related to the documentation and validation activities.
We also need to manage things such as model updates, version control, and similar lifecycle activities. Apart from that regulatory documentation, I don’t see a major difference between using models in the lab and using them in manufacturing. As I said, the data connections and interfaces follow a separate validation pathway under the GAMP guidelines. Regarding the mathematical models themselves, it’s essential to perform a credibility assessment.
The standard approach is to follow ASME V&V 40, where ASME stands for the American Society of Mechanical Engineers. The framework requires you to clearly define the question of interest, the context of use, and then assess the model risk, which includes both the decision consequence and the model influence.
For example, if the model recommends releasing a batch, do you release it automatically, or do you first perform confirmatory testing? Based on these factors, you determine the model risk. Then, based on that risk, you establish a credibility plan, which consists of verification, validation, and applicability.
Verification essentially confirms that if you expect the model to compute “2 + 2,” it is actually computing “2 + 2” and not “2 + 3.” So verification is mainly concerned with the numerical, mathematical, and software implementation.
Validation then assesses model performance using approaches such as holding back independent runs for testing, evaluating challenging process runs, assessing prediction accuracy with respect to the context of use, performing uncertainty quantification, and similar activities.
With that credibility plan in place, it is generally sufficient to support deployment of the model in manufacturing.
David Brühlmann [00:07:24]:
Let’s zoom out a bit. We highlighted that the monetary return on investment (ROI) is great—or even much greater—in manufacturing. There’s definitely also a huge benefit during development. You can speed up development and find better process conditions. If we now look at the timing and the effort you want to invest in modeling, when should you start? What should you do at the beginning, and what can wait until later? Can you give some advice to startup leaders?
Ignasi Bofarull-Manzano [00:07:54]:
I would say that, in terms of return on investment, if you start during development, the return on investment begins on day one because applying models helps you perform the experiments that you need to do anyway in a more efficient way. So the return on investment is immediate.
In manufacturing, it’s much more use-case-based. The effort required is higher because it depends on how much equipment you have, the legacy of your systems, and how many interfaces need to be built and validated. The return on investment also depends very much on the individual case.
However, I’m allowed to mention one specific example. On September 3rd, we’re hosting an event here in Vienna called PAS-X User Day, where we invite customers to present real-world use cases. One of those customers will be Takeda, and I’d like to share some return-on-investment figures from that project.
Takeda approached us with manufacturing data from a biologics production process that had been running for quite some time, so they had a considerable amount of manufacturing data. The first thing we did was not start building a digital twin immediately. As I mentioned earlier, we always start offline. Using the historical manufacturing data, we trained an end-to-end process model.
What was particularly interesting was that, after training the model and running simulations—which is a multi-input, multi-output, multi-unit-operation optimization problem—the model suggested that changing the setpoints of six process parameters across four different unit operations could increase the yield by 35%, while also reducing the out-of-specification (OOS) rate significantly.
The OOS rate was already low, so the reduction of about 65% is somewhat less meaningful than the yield improvement. The key result is that we could increase yield by 35% while maintaining product quality within specification—or even improving it.
However, this was not yet a digital twin because everything was performed offline. It was an end-to-end process model. What became very clear was that this particular process already showed a strong business case.
Another aspect I found particularly interesting is that these six new operating setpoints were not based on exotic operating conditions. We never allowed the model to extrapolate beyond the available data. Instead, the recommendations were based entirely on process variations that had already occurred during previous years of manufacturing.
Some batches performed better than others. The model was able to identify, within all that historical data, that adjusting six process parameters could lead to a 35% increase in yield. In other words, within the existing design space, it was possible to improve the yield by 35%. That provided a very clear business case for using an end-to-end process model.
The next step was then to deploy it as a digital twin. At that point, we reached the digital shadow stage. The model now receives real-time data directly from the shop floor across all connected equipment.
Currently, we’re implementing the full end-to-end digital twin. Based on our simulations, we’re seeing that, on top of the initial 35% yield increase, the digital twin could provide an additional 4% yield improvement by adapting to the normal process variability observed during manufacturing. This is one of our most impressive real-world use cases, which is why we’ve invited Takeda to present it in Vienna.
By the way, you’re more than welcome to join us on September 3rd. That said, every manufacturing process is different. For example, if you produce only two batches per year, you might not actually need a digital twin. That’s simply the reality. Even in that situation, I would still recommend building an offline end-to-end process model because it may help you increase the yield—not necessarily by 35%, but by some meaningful amount—so that you obtain more product from those two annual batches.
However, because you’re manufacturing only twice per year, investing in a full digital twin—with all the required system interfaces and validation effort—may not provide sufficient return on investment. Ultimately, it’s a case-by-case decision.
David Brühlmann [00:12:14]:
These are impressive numbers, no doubt. And the business value is clear. Where should a smart biotech scientist start right now if they want to use modeling?
Ignasi Bofarull-Manzano [00:12:26]:
First, as I mentioned, identify what the bottleneck is. There are many problems and many business needs, but what is the biggest business need? Try to identify your bottleneck. For example, perhaps Protein A chromatography is your bottleneck, or you have a lot of variability in upstream processing, which leads to large variations in mannosylation patterns. Try to identify what is really the most critical issue because you may have several important challenges. Then identify what type of model you actually need.
This is important because the use case I mentioned—which will be presented in Vienna on September 3rd with Takeda—is based on linear mixed-effects models. In that particular project, we are not using fancy physics-informed AI or hybrid models. So first determine what type of model you need. My recommendation is that even if you ultimately need a complex model, start simple. Starting simple means building linear input-output models.
For example:
• What is the final titer?
• What is the final mannosylation pattern in my fermentation?
• Which process inputs influence those outputs?
Even if temperature varies over time, initially you can simply use the average temperature as an input. So start with simple input-output linear models.
Before moving on to more sophisticated models for individual unit operations, I always recommend concatenating the models—even for relatively simple unit operations such as virus inactivation.
Connect them into an end-to-end process model and evaluate how each process parameter affects not only the output of an individual unit operation, but the final drug substance quality.
That allows you to identify what is truly critical and where to focus your resources. Then, where necessary, you can begin building more sophisticated models for specific unit operations. And definitely start offline. Everyone talks about digital twins, but many of the examples people call digital twins are actually just digital models.
Start with offline end-to-end process models. Only after you’ve demonstrated the business value—for example, a 35% increase in yield—should you scale toward a digital twin.
Following these steps also helps secure support from C-level management, because you’ll already have evidence that the models deliver business value. Otherwise, you risk launching proof-of-concept projects that become expensive and ultimately go nowhere. So, start small and build from there.
David Brühlmann [00:14:52]:
This has been great. What additional questions should I have asked?
Ignasi Bofarull-Manzano [00:14:57]:
I think we’ve covered most of the important topics. We’ve discussed different types of models, regulatory considerations, and applications ranging from early-stage development to manufacturing.
Everything I’ve said today is, of course, my personal opinion. If I could add one more personal thought, it’s that with the current AI boom—including agentic AI—I’m seeing people delegate too much to AI. I also use AI every day for many different tasks, and it’s a fantastic tool. But when it comes to modeling, it’s very important to understand what you’re actually doing.
I know that’s not very appealing to some people because many people simply don’t like working with numbers. Still, I would recommend at least learning the basics of statistics so that you can understand what the model is telling you, whether the predictions have sufficient confidence, and how reliable the results are.
I think that’s extremely valuable so that you’re not simply relying on AI to evaluate the model for you.
David Brühlmann [00:15:56]:
What is the most important takeaway from our conversation?
Ignasi Bofarull-Manzano [00:16:01]:
I would summarize it like this. First, identify your business problem. Then determine what type of model is needed to solve that problem. Even if solving the problem ultimately requires a sophisticated model, don’t begin with one.
Start with simple linear models. Always look at the process end-to-end, taking a holistic or integrated process modeling approach. Identify only the data required to train that model.
Start with the minimum data you actually need. Demonstrate the business value offline—for example, through optimization studies. Only then evaluate whether it’s worthwhile to deploy and scale the model into a real-time implementation such as a digital twin.
David Brühlmann [00:16:44]:
Excellent. There you have it, smart biotech scientists. Thank you so much, Ignasi, for being on the show, for sharing your passion, and for explaining so clearly what digital twins are, what modeling is, and the benefits they can bring. Where can people get hold of you?
Ignasi Bofarull-Manzano [00:17:00]:
I’d say LinkedIn is the easiest way. People can send me a connection request, and we can connect there.
David Brühlmann [00:17:06]:
All right, smart biotech scientists, reach out to Ignasi, talk with him about modeling and much more. Thank you once again, Ignasi, for being on the show today.
Ignasi Bofarull-Manzano [00:17:17]:
Thanks a lot, David.
David Brühlmann [00:17:18]:
What Ignasi makes clear is that digital twins aren’t magic. They’re built on trust, validation, and physics grounded in real process understanding. Whether you’re just starting to untangle your data or already piloting hybrid models, the path forward is the same: start small, validate rigorously, and let the science lead.
Are you enjoying the show? Please leave us a review on Apple Podcasts or your favorite podcast platform. It helps other scientists like you discover the show. Thank you so much for tuning in, and I’ll see you next time.
Disclaimer: This transcript was generated with the assistance of artificial intelligence. While efforts have been made to ensure accuracy, it may contain errors, omissions, or misinterpretations. The text has been lightly edited and optimized for readability and flow. Please do not rely on it as a verbatim record.
Next Step
If you found value in today’s episode, take a moment to like, follow, and leave a review on Apple Podcasts or your favorite platform—it helps us reach and support more scientists like you.
Thanks for tuning in to the Smart Biotech Scientist podcast and being part of this journey toward bioprocess mastery. For more insights and practical tips, visit
About Ignasi Bofarull-Manzano
Ignasi Bofarull-Manzano is a Senior Data Scientist and CMC Consultant at Körber Pharma, specializing in the application of data science, process modeling, and digital technologies to biopharmaceutical manufacturing. With over seven years of industry experience, he works with biopharma companies to translate advanced modeling approaches into practical manufacturing and CMC strategies, including QbD-based development and regulatory support. Since 2024, Ignasi has also been pursuing an industrial PhD at Forschungszentrum Jülich and RWTH Aachen University under Professor Eric von Lieres. His research explores how physics-informed AI and scientific machine learning can enable end-to-end digital twins for regulated manufacturing.
Connect with Ignasi Bofarull-Manzano on LinkedIn.
Further Listening
If you enjoyed this episode you might also like listening to:
Episodes 215 - 216: From Data Silos to Autonomous Biomanufacturing: Digital Twins and AI-Driven Scale-Up with Ilya Burkov
Episodes 05 - 06: Hybrid Modeling: The Key to Smarter Bioprocessing with Michael Sokolov
Episodes 17 - 18: How Extracting Gold From Your Data Accelerates Process Development with Ioscani Jiménez del Val
Episodes 263 - 264: Why AI and Automation Tools Won't Deliver Until Your Lab's Data Is Connected with David Hardy
Want to Join?
Are you a CMC or biomanufacturing leader with hard-won lessons on process development, scale-up, or CDMO management? We’re always looking for practitioners with real execution stories to share on Smart Biotech Scientist.
Apply to be a guest or recommend someone:
David Brühlmann is a strategic advisor who helps C-level biotech leaders reduce development and manufacturing costs to make life-saving therapies accessible to more patients worldwide.
Hear It From The Horse’s Mouth
Want to listen to the full interview? Go to Smart Biotech Scientist Podcast.
Want to hear more? Do visit the podcast page and check out other episodes.
Do you wish to simplify your biologics drug development project? Contact Us