Digitalization is rapidly reshaping bioprocessing, enabling companies to accelerate product development, optimize manufacturing, and build a competitive advantage—even with limited data.
Ignasi Bofarull-Manzano, Senior Data Scientist and CMC Consultant at Körber Pharma, shares his expertise on digital twins, modeling, and making digital initiatives truly deliver results—not just slideware hype.
Episode Highlights
- Misconceptions about data requirements for digital twins—why quality and context of data matter more than sheer quantity [02:40]
- Ignasi’s journey from curiosity in biology to a career in data science, modeling, and digital twins [04:31]
- Clear distinctions between digital models, digital shadows, and digital twins, explained with real-world analogies [06:42]
- How to approach digital development when faced with legacy data silos and scattered analytics [09:56]
- The importance of starting with a focused business need instead of chasing trends or buzzwords [12:28]
- Insights into where modeling truly delivers value in the product lifecycle—development versus manufacturing [13:11]
- Strategies for small companies to leverage digitalization and data from the ground up [15:56]
- An accessible overview of physics-informed AI, physical AI, and hybrid modeling—and their application in bioprocessing [18:15]
- The comparative advantages of physics-informed AI versus hybrid models in different bioprocessing contexts [24:45]
In Their Words
The important difference here is that every effort made in manufacturing—even if it just leads to an increase of 1% yield—if this is an increase of 1% yield of every single batch throughout the full product lifecycle, this is a huge business benefit. So I would say that I’m a big fan of modeling, and I strongly believe modeling brings value throughout the full product lifecycle. But what we know is that the closer the model sits to manufacturing, the higher the return on investment is.
Podcast Transcript
David Brühlmann [00:00:30]:
You’ve got data everywhere—upstream, downstream, analytics—scattered across systems that never talk to each other. What if that mess is actually a gold mine you haven’t tapped yet? Today’s guest, Ignasi Bofarull-Manzano, who’s a Senior Data Scientist and CMC Consultant at Körber Pharma, shows how digital solutions and digital twins unlock that value: faster development, more robust processes, and higher productivity.
Let’s dig in. Welcome, Ignasi. It’s good to have you on today.
Ignasi Bofarull-Manzano [00:02:25]:
Thank you, David. Glad to be here. Very excited.
David Brühlmann [00:02:29]:
Let’s start our conversation with perhaps this controversial question. Share something that you believe about bioprocess development that most people disagree with.
Ignasi Bofarull-Manzano [00:02:40]:
I think many people overestimate what is required to build a digital twin in bioprocessing. They assume that usually a large amount of data is required, as well as very fancy models, but neither is necessarily true. So probably we will talk about modeling a lot throughout this podcast, but when it comes to the data requirements, in our industry we actually have a lot of data because it is spread throughout Excel files, historians, ELNs, MES, analytics systems, et cetera.
For me, as a data scientist, I’m always happy to have more data because basically I don’t want to miss any important source of variability. But when it comes to implementing digital twins that bring benefit, the essential question is: What decisions should the model support, and therefore which data are needed for that context of use?
What does this mean in plain language? For example, we have several published use cases where the typical data that you have from process validation, together with a handful of manufacturing runs, have been enough to build a useful end-to-end digital twin for manufacturing.
So what matters is the type of data. For example, many people focus a lot on product amount, but they don’t measure quality. It’s also very important that during process validation you investigate starting material variability. So it’s not so much about data quantity; it’s more about data quality.
David Brühlmann [00:04:09]:
You’re making several good points, and we’re going to unpack these various points during our conversation. I’d like to first talk about yourself. Draw us into your story. What was the spark that attracted you to science and then eventually to data science, modeling, and digital twins? And what were some interesting pit stops along the way?
Ignasi Bofarull-Manzano [00:04:31]:
So, since I was very young, I’ve always had a spark of curiosity in many diverse fields. Of course, biology has been the field that sparked the most curiosity in me because we’re all human, we all get sick, et cetera. So for me, it was clear that I wanted to study something related to biology.
At the same time, however, I’ve always been very interested in mathematics and numbers. Even though I have a huge curiosity to understand how things work, I also like to get things done. I’ve always been very interested in the application of things. So this is how I came across biotechnology.
I completed both my bachelor’s and master’s degrees in biotechnology. At the very beginning of my bachelor’s, there was actually a point where I was even hesitating to quit biotechnology. What happened was that, from the theoretical point of view, I really liked what I was studying. But, for example, I didn’t really like being in the lab. So it was a trade-off that made me hesitate about continuing.
Then, in the third year of my bachelor’s in Barcelona, I had a bioreactors course, and I realized that this was very applied. It was also a first step, somehow, into modeling. We could call it very simple modeling—just mass balances, for example: how much you have to inoculate your bioreactor, when you should stop the fermentation, et cetera. But it was already a motivation to delve deeper into this field.
When I really discovered that modeling was my thing was during my master’s thesis at a startup in Vienna, where I worked on hybrid modeling. That’s when I realized that I loved it and that it was really my place to be, let’s call it like this.
David Brühlmann [00:06:19]:
Fantastic. Let’s start with the buzzword. A lot of people talk about digital twins these days. And then, when you look into what they actually mean, you realize that they mean different things to different people. So what is actually a digital twin? Or what separates a true digital twin from, I’d say, a fancy dashboard or simulation?
Ignasi Bofarull-Manzano [00:06:42]:
A digital twin, as I understand it at least, is primarily a concept referring to connectivity to the physical process and the data decision loop, rather than a particular mathematical model.
The model running in the background can be statistical, mechanistic, or hybrid. One useful way to explain the different maturity levels is to distinguish between a digital model, a digital shadow, and a digital twin.
At the first stage, we have a digital model, which is basically a mathematical model running on a computer and fitted using historical process data. It can already be very useful for process development, simulation, and optimization, but it is not yet connected to the physical system.
The next step is to bring shop-floor data into the model. Once the model receives near-real-time data from the physical process, we move toward what people call a digital shadow. Now the model can follow the batch as it progresses, predict product quality, and issue early warnings, for example.
Finally, if we are able to close the data loop so that the model’s predictions or recommendations are communicated back into the manufacturing workflow and influence the physical process, this is what, according to the definition, is called a digital twin. So this bidirectional data flow is very important.
There is actually an example I like to use to make this easier to understand, which is Google Maps. Imagine that tomorrow you need to travel from point A to point B. A static route planner can calculate a route based on distance and typical traffic patterns. However, you have your route planned, but tomorrow, when you start the journey, traffic, roadworks, or accidents may change the best route.
Google Maps receives this live information, recalculates the prediction, and recommends a new action. Then you, as the driver, are still in the loop, so you have the decision to take the recommendation or not. But if you take the recommendation, it is basically influencing how you interact with the physical world. In that case, Google Maps would be acting as a digital twin.
The same principle applies to manufacturing. When we set up a control strategy, we define setpoints, which ideally represent the optimal or best route. But we know that, in reality, with normal process variability, the optimal route may change. A digital twin is responsible for interpreting this variability in real time and suggesting a new optimal route if the original one is no longer the best.
David Brühlmann [00:09:14]:
Now, to build a digital twin—or even just a model, I’d say, corresponding to the earlier stages you’ve outlined—we need a lot of data. You also mentioned at the beginning that it’s not only the amount of data that matters, but also the quality of the data.
What I’m observing is that a lot of biotech companies have huge data silos because of legacy systems, different site histories, different mentalities, and many other reasons. Where do you even start when you’re developing a more digital approach to process development and, hopefully, at the end of the day, an end-to-end digital approach?
Ignasi Bofarull-Manzano [00:09:56]:
Yes, very good question. We have a lot of data, as you said, and in an ideal world, we would use all this data to build a product lifecycle digital twin. However, for me as a data scientist, this is the dream, but it’s not actually the reality.
What happens, as you said, is that the bigger the company is, the more data chaos it has because there are many employees across many departments, sites, countries, and also systems.
So what is important is, first, instead of aiming for Mars, try to go to the Moon first. First define the decision and the context of use of the model that you want to build. Based on this, you need to identify which data are required for that model context of use and which interfaces you need to connect.
As I said at the very beginning, you don’t really need all available data, only the required data. What I also recommend is starting small. First of all, you always need a business need. What’s the bottleneck? Is it time to market? Is it manufacturing costs?
Then identify the type of model that you need for that problem and the required data to build a reliable model that fulfills that context of use. Then, don’t go fully into real time with a digital twin immediately, because that also comes with additional effort. First build a model.
It’s also very important, by the way, to build end-to-end process models. For example, if you think that upstream is a bottleneck in your production process, it’s still very important that you connect the model you built for upstream with all the subsequent unit operations to really identify or evaluate whether upstream is actually critical for drug substance quantity and quality.
So start small, always try to think end-to-end or holistically, and demonstrate the business value offline first. Only then, if there is a clear business value that these end-to-end process models offer, should you scale by connecting all the interfaces and systems to enable bidirectional data exchange.
David Brühlmann [00:12:04]:
So let me paraphrase this once again because you’ve made an important point. Start with the end in mind. What does the business need, or what are the bottlenecks of the business? Then focus on that. Don’t try to solve 10 or 20 problems at the same time. Focus on the most important one. Then, once you have a proof of concept, you can further expand on that.
Ignasi Bofarull-Manzano [00:12:28]:
Exactly. You summarized it very well. The reason I’m especially mentioning that you should first identify your problem is that, over the last two years, especially with the whole AI boom, I’ve seen some management teams mix digital twins with AI and other technologies, but they don’t really have a use case.
This is a common first mistake that I would recommend people avoid. Don’t start building models just for the sake of building models or because they sound sexy, if I can say that.
David Brühlmann [00:12:57]:
From what you’ve seen, is there a particular area in bioprocessing where people usually start because the business value is very clear, or is it more on a case-by-case basis?
Ignasi Bofarull-Manzano [00:13:11]:
It’s on a case-by-case basis. Perhaps this is also similar to your very first question about something I might disagree with compared with many people. Many people focus heavily on models for early- or late-stage development. We also work on those, and the numbers you obtain there are certainly more appealing.
For example, we’ve shown that you can reduce experimental effort by around 70% by using end-to-end process models during process characterization. That’s a big number, right? A 70% reduction.
However, you’re still in process development, and the product may still fail during clinical trials. So, of course, there is a benefit. If you reduce your experimental effort by more than half, that’s a significant business value, especially in terms of time. Usually, in process development, the business benefit is faster time to market rather than saving the cost of experimental runs, because those are relatively inexpensive. On the other hand, we have manufacturing. There, we’ve built models for root cause analysis as well as end-to-end process models.
The business benefit, expressed numerically, is smaller. For example, we’ve seen that an end-to-end process model in manufacturing can increase yield by an average of 4%. Now, if you compare 4% with 70%, 4% sounds much less impressive. But the important difference is that every improvement made in manufacturing—even if it leads to just a 1% increase in yield—applies to every single batch throughout the entire product lifecycle. That translates into a huge business benefit.
So I would say that I’m a big fan of modeling, and I strongly believe that modeling brings value throughout the entire product lifecycle. But what we know is that the closer the model sits to manufacturing, the higher the return on investment tends to be.
David Brühlmann [00:15:10]:
Yeah, I agree. Definitely, in terms of monetary return on investment, that’s where you’re going to see the biggest return on investment, considering that you’re going to produce your drug for the next 10 or 20 years. Perhaps I would say that in development, the ROI can also be non-monetary because companies are under increasing pressure to shorten development timelines. If you can save effort there and move faster, that’s also a tremendous return on investment. I’m curious—for a smaller company where you have limited data because you’re just starting out, what is your advice? What should they do?
Ignasi Bofarull-Manzano [00:15:56]:
If you’re a small company, you might actually have an advantage—a competitive advantage—because you’re starting from scratch. From a digitalization point of view, you can build that competitive advantage from the beginning. Some people talk about having a data-first mindset, although that’s more related to digitalization. When it comes to modeling, especially if you don’t have a lot of data, my advice is to start small.
For example, make sure that every experiment—which you have to perform anyway, whether for R&D or process validation—extracts as much knowledge as possible. You can do this by applying state-of-the-art Design of Experiments (DoE) methods, for example. Then, as the product progresses through its lifecycle, you can scale from there. I don’t know if this really answers your question, or if you were thinking in a different direction.
David Brühlmann [00:16:55]:
No, that’s good because, yes, as a small company, I would agree with you that you have a competitive advantage because you can start from scratch and get it right the first time instead of trying to rectify or correct past decisions.
Ignasi Bofarull-Manzano [00:17:11]:
Exactly. On the other hand, for very large companies, as we discussed earlier, there are so many different systems. Even when they use the same system, it may be used differently across different countries because different people are involved. It’s true that they have much more data, but that can also be counterproductive because you can get lost in the data, and not all of it is actually required.
That’s why it’s very important to focus on which data you need to fulfill the context of use of your model. And before defining the context of use of your model, you first need to identify the business need. What’s the problem you’re trying to solve? Then determine which model is appropriate for that context of use, and finally identify the data required to train that model.
David Brühlmann [00:17:51]:
Yes, absolutely. Now, you’re also working on physics-informed AI. What does that mean exactly? I’d also be curious to hear your perspective on the potential of AI because you already alluded to it. There’s a lot of hype out there. What can AI really deliver in this space, and what can it not?
Ignasi Bofarull-Manzano [00:18:15]:
Yeah. So first of all, I have to mention that the name physics-informed AI is a bit of a hype as well. I’m mentioning this because there is now another concept called Physical AI, which is, for example, a term that Jensen Huang from NVIDIA often mentions as one of his favorite topics. These are two very different technologies.
Physical AI is based on generative AI, and it’s essentially generative AI embedded into physical systems such as robots or autonomous vehicles. On the other hand, the more modest physics-informed AI, or physics-informed machine learning (PIML)—also known as scientific machine learning—is basically a machine learning model that, during the training process, is forced to obey physical principles, mainly first-principles differential equations.
The main purpose is that, once it’s trained, the model respects the underlying physics. What’s the advantage of this? This is twofold. On the one hand, it allows you to combine data with first principles. For example, in upstream processing, we cannot describe everything with differential equations. We know, for example, that it’s not entirely clear how pH influences cell growth. We don’t have a Monod equation for that relationship.
These types of models allow us to combine data with first principles, although this concept is not entirely new. There is already the concept of a hybrid model, which we can perhaps discuss later.
The second advantage is that, because the model has been trained to respect the physics, once training is complete, you no longer need to solve the physics explicitly. In other words, you no longer need to solve the differential equations during inference. This is a major advantage because some ordinary differential equations (ODEs) and especially partial differential equations (PDEs) are computationally expensive. They are relatively slow to solve.
This becomes a bottleneck, for example, when performing uncertainty quantification in real time or when solving the inverse problem, meaning estimating mechanistic parameters from high-throughput experimental data.
David Brühlmann [00:20:25]:
What is the difference with respect to hybrid modeling?
Ignasi Bofarull-Manzano [00:20:29]:
I would say that the philosophy is the same, which is combining data with first principles. The difference lies in how it is implemented. It may be a bit difficult to explain without a whiteboard, but imagine the typical cell growth differential equation, where X is the cell concentration.
We know that dX/dt, which represents the change in cell concentration over time, equals μ × X, where μ is the specific growth rate. We know this because it’s a logical equation. Everything happens through cells. If there are no cells, there cannot be any cell growth.
A hybrid model would, for example, train a neural network that predicts the specific growth rate, μ, as a function of pH. Once the neural network predicts μ, that value is passed into the ordinary differential equation:
dX/dt = μ × X
Then an ODE integrator, using the initial condition together with the predicted μ at each time point, computes the corresponding values of X. This is effectively solving the differential equation. So you’re still solving a differential equation during inference. Basically, a classical hybrid model is a system of differential equations where some parts of those equations are replaced by a black-box model, usually a neural network.
On the contrary, physics-informed machine learning, for example physics-informed neural networks (PINNs), directly learn not the specific growth rate μ, but rather how the cells grow over time and as a function of pH in this simplified example.
What’s the advantage? For upstream processing, you probably wouldn’t see a major difference between PINNs and hybrid models.
But imagine chromatography. With a hybrid model, you might have a system of partial differential equations describing convection, dispersion, diffusion, and binding, while the neural network predicts only one component—for example, the adsorption isotherm, describing the protein’s binding affinity.
Once the neural network predicts that isotherm, you still need to solve the full PDE system, including spatial and temporal discretization, which is computationally expensive. By contrast, a physics-informed neural network would directly predict the chromatogram over time. That is the major advantage.
I often use a simple example during conferences when introducing this technology. Imagine you bring a new dog home and have to train it. You don’t allow the dog to jump on the sofa because it could damage it or make it dirty. The dog doesn’t understand why it shouldn’t jump on the sofa. Instead, when it jumps on the sofa, you correct it, and when it behaves correctly, you reward it.
After enough repetitions, the dog learns what is acceptable and what is not, so eventually—even when you’re not at home—it shouldn’t jump onto the sofa. PINNs work in a similar way. During training, we use the physical equations, such as the partial differential equations, to guide the model.
The hope is that once the model is trained—and here the training procedure is extremely important—the model has learned the underlying physics well enough that it no longer needs to solve those equations explicitly during inference. As a result, predictions can be computed much faster.
David Brühlmann [00:24:16]:
From what I understood, it seems that, on the upstream side, physics-informed AI and hybrid models are fairly comparable. Obviously, there are some differences, and perhaps they’re complementary.
Whereas in downstream processing, hybrid modeling only captures part of the system, while physics-informed AI can create models that represent much more of the underlying physics and chemistry through the differential equations.
Ignasi Bofarull-Manzano [00:24:45]:
Based on the research we’ve done so far, we see the greatest value in downstream processing. For example, the father of these technologies, Professor George Karniadakis from Brown University—with whom we collaborate—worked with us on an upstream use case.
We asked them to model a bolus-fed CHO cell culture system using PINNs. In parallel, we used an advanced hybrid model called Neural ODE, which is essentially a hybrid model that is computationally more efficient because it uses automatic differentiation.
We found that the Neural ODE actually performed better. It was easier to train and generally gave better results than the physics-informed neural network. By contrast, in downstream applications, we’re seeing the opposite. There, physics-informed neural networks appear to provide greater benefits than hybrid models.
David Brühlmann [00:25:33]:
Ignasi is already reframing how we think about digital twins, and there’s much more ground to cover. In Part 2, we continue digging into the practical realities of building trust in these models and where the real value shows up for biotech teams.
If you’re enjoying the conversation, please leave a review on Apple Podcasts or your favorite podcast platform. Thank you so much for tuning in, and I’ll see you next time.
Disclaimer: This transcript was generated with the assistance of artificial intelligence. While efforts have been made to ensure accuracy, it may contain errors, omissions, or misinterpretations. The text has been lightly edited and optimized for readability and flow. Please do not rely on it as a verbatim record.
Next Step
If you found value in today’s episode, take a moment to like, follow, and leave a review on Apple Podcasts or your favorite platform—it helps us reach and support more scientists like you.
Thanks for tuning in to the Smart Biotech Scientist podcast and being part of this journey toward bioprocess mastery. For more insights and practical tips, visit
About Ignasi Bofarull-Manzano
Ignasi Bofarull-Manzano is a Senior Data Scientist and CMC Consultant at Körber Pharma, with more than seven years of experience in the biopharmaceutical industry. His expertise spans bioprocess engineering, data science, and advanced process modeling, helping biopharma companies leverage data and digital twin technologies to improve process understanding, support Quality by Design (QbD) development, and strengthen regulatory submissions. Ignasi is also an industrial PhD candidate at Forschungszentrum Jülich and RWTH Aachen University, where his research focuses on applying physics-informed AI and scientific machine learning to end-to-end digital twins for regulated manufacturing.
Connect with Ignasi Bofarull-Manzano on LinkedIn.
Further Listening
If you enjoyed this episode you might also like listening to:
Episodes 215 - 216: From Data Silos to Autonomous Biomanufacturing: Digital Twins and AI-Driven Scale-Up with Ilya Burkov
Episodes 05 - 06: Hybrid Modeling: The Key to Smarter Bioprocessing with Michael Sokolov
Episodes 17 - 18: How Extracting Gold From Your Data Accelerates Process Development with Ioscani Jiménez del Val
Episodes 263 - 264: Why AI and Automation Tools Won't Deliver Until Your Lab's Data Is Connected with David Hardy
Want to Join?
Are you a CMC or biomanufacturing leader with hard-won lessons on process development, scale-up, or CDMO management? We’re always looking for practitioners with real execution stories to share on Smart Biotech Scientist.
Apply to be a guest or recommend someone:
David Brühlmann is a strategic advisor who helps C-level biotech leaders reduce development and manufacturing costs to make life-saving therapies accessible to more patients worldwide.
Hear It From The Horse’s Mouth
Want to listen to the full interview? Go to Smart Biotech Scientist Podcast.
Want to hear more? Do visit the podcast page and check out other episodes.
Do you wish to simplify your biologics drug development project? Contact Us