Why the Lab Optimum Fails: Process Optimization for Scale-Up With Small Data

Key takeaways
- A lab optimum built on one peak yield often collapses under real noise at scale.
- Reliability for scale-up targets a stable operating window that tolerates real variation.
- The scientist sets the optimization goal: exploration, peak yield, or reliability.
- SuntheticsML campaigns commonly start from about five data points and improve from there.
- The model recommends the next experiment; the scientist decides what scales.
Process optimization for scale-up is the work of finding conditions that stay optimal after a reaction leaves the bench. For a bench chemist in pharmaceuticals, that reaction is one you developed and expect to defend through tech transfer. Many optimization efforts miss this. They chase the single highest yield in a small, clean dataset, then watch that result soften under the noise, variability, and constraints of a real plant. The number that looked like a win at 10 milliliters does not survive at 100 liters.
The right optimization goal is often reliability. A stable operating region matters more for scale-up than the single highest yield, and teams can reach it with very little data.
Why is a single best result risky?
A single best result is easy to over-trust. Run enough experiments, find the highest yield, and the job can look finished. The catch is that the highest point in a dataset can sit on a narrow peak. Shift any factor, add normal batch-to-batch variation, and performance falls away fast.
Bench chemists have seen this pattern. A promising bench result stalls in tech transfer. The conditions that looked optimal turn out to be sensitive to a degree of temperature or a small change in reagent quality. The team goes back to the lab, months into optimizing toward a point that was never going to hold.
Regulators build this concern into how they define development. ICH Q8(R2), the pharmaceutical development guideline, frames quality around a design space: the demonstrated multidimensional combination of input variables and process parameters that provides assurance of quality (ICH Q8(R2), 2009). The expectation is a proven operating region rather than one fixed set point. So the practical question is how wide and stable the good region around the optimum is.
Why does a fragile maximum break at scale?
Scale introduces variation that a clean bench study never shows. Mixing, heat transfer, addition rates, and raw material consistency all shift as volume grows. A process sitting on a sharp optimum has no margin to absorb those shifts.
Measurement noise starts to matter during development, well before manufacturing. If an optimization method treats every data point as exact truth, it will climb onto a spike created by noise. A method that models its own uncertainty behaves differently. It can settle on a slightly lower, far more stable region, because it accounts for how much confidence the data actually supports.
This is an established research area. Handling noisy experiments in Bayesian optimization is well documented (Letham et al., 2019), and work on robust Bayesian optimization shows that accounting for variation in the inputs steers the search toward stable optima instead of sharp, sensitive peaks (Fröhlich et al., 2020).
SuntheticsML is built around that uncertainty-aware view. In a published project with Pfizer, the platform matched the team's calibrated kinetic model at roughly 97.5% yield and reduced impurities relative to the kinetic model's optimized conditions. The result held even after 5% error noise was added to the training data. Experiment count barely moved, from 15 down to 14. The project mattered because a bench-friendly tool matched a demanding benchmark, stayed accurate under noise, and freed the modeling team for other work. Read the full Pfizer case study for the details.
How does choosing the optimization goal change the result?
Many optimization tools hide the optimization goal from the user. It should be a deliberate choice, and it belongs to the scientist. The chemist who developed the reaction sets the target, because they know which trade-offs the process can accept.
With the Lithium algorithm, a SuntheticsML campaign can target different objectives. A team can prioritize exploration when the space is poorly understood, pure optimization when speed to a peak matters, or reliability when the aim is a process robust enough to scale. The goal can shift between iterations as understanding grows, so early exploration can give way to a reliability focus once the promising region is clear. The Lithium release explains how this works.
That choice reframes the whole campaign. Optimizing for reliability tells the model to value a stable operating window rather than the single best coordinate. For a team headed toward manufacturing, that is usually the answer worth having.
The model recommends the next experiment. The scientist decides which trade-off the program can accept. The search belongs to the tool, and the judgment belongs to the chemist.
What does optimizing for reliability look like in practice?
Choosing reliability changes which experiments get run and which result gets accepted. The campaign spends runs mapping the width of the good region and testing how performance responds to small changes in the factors that will vary at scale, in place of confirming a single spike.
Lithium supports this with variable-specific modeling for mixed chemical and continuous factors, constraint handling for real formulation and process limits, and multi-objective tracking that lets a team watch yield, purity, and other outputs at once. Optimizing yield while ignoring impurities or operating limits is how a lab win becomes a manufacturing problem.
The outcome is a process description a scale-up team can trust: a defensible operating window with evidence behind it.
How much data is actually needed?
A common objection is that robustness demands far more data. It does not. The independent evidence for data-efficient optimization in chemistry is strong. In a widely cited study, Bayesian optimization found globally optimal conditions for a reaction within roughly 3% of a 1,728-point search space, and after a few rounds it matched or outperformed 50 expert chemists and engineers while producing more consistent worst-case results (Shields et al., Nature, 2021). Reviews of the field describe Bayesian optimization as a practical tool for sustainable reaction development because it reaches good answers with few runs (Nature Reviews Methods Primers, 2023).
SuntheticsML applies this in a bench-friendly way. Campaigns commonly begin with about five data points and improve from there, because the algorithms extract information from very small datasets and self-correct as new results arrive. The platform is also reaction-agnostic, spanning electrochemistry, catalysis, crystallization, and formulation. The technology page shows our approach.
None of this removes the need for good scientific setup. A clear goal, a quality metric tied to a real physical property, and reasonable measurement noise still decide the result. The algorithm serves that work; it does not replace it.
Conclusion
If your team has optimized toward a result that did not survive tech transfer, the optimization goal is worth revisiting. Optimizing for reliability produces a process built to hold at scale, and it works with the data you already have. Bring a system you care about and start a campaign from your existing dataset. Request a walkthrough of SuntheticsML. For related reading, see our guides on Bayesian optimization for chemical and pharmaceutical process development and frugal sampling strategies.
Frequently asked questions
What does process optimization for scale-up mean?
It means finding conditions that stay optimal when a process moves from the bench to larger equipment. The focus is a stable operating window that tolerates real variation across scale.
Why does a lab optimum often fail at scale?
Scale introduces variation in mixing, heat transfer, and material consistency. A process sitting on a narrow peak has no margin for that variation, so performance drops. A wider, more stable region holds up better.
How can scientists optimize for reliability instead of peak yield?
With SuntheticsML's Lithium algorithm, a campaign can be pointed at reliability for scale-up rather than pure optimization, and the goal can change between iterations. The model then values a stable operating window across the factors that vary at scale.
How much data is needed to start?
SuntheticsML campaigns commonly begin with around five data points. The algorithms are built to learn from very small datasets and to refine recommendations as new results come in.
Does adding noise break the model?
Not necessarily. Because the approach models its own uncertainty, it can stay accurate under noise. In the Pfizer project, results held even with 5% error added to the training data.
Sources
- ICH Q8(R2), Pharmaceutical Development (2009). International Council for Harmonisation, adopted by the EMA and FDA. Link
- Shields, B. J., Stevens, J., Li, J., et al. (2021). Bayesian reaction optimization as a tool for chemical synthesis. Nature, 590, 89-96. Link
- Letham, B., Karrer, B., Ottoni, G., & Bakshy, E. (2019). Constrained Bayesian Optimization with Noisy Experiments. Bayesian Analysis, 14(2), 495-519.
- Fröhlich, L. P., Klenske, E. D., Vinogradska, J., Daniel, C., & Zeilinger, M. N. (2020). Noisy-Input Entropy Search for Efficient Robust Bayesian Optimization. AISTATS, PMLR 108.
- Bayesian optimization as a valuable tool for sustainable chemical reaction development (2023). Nature Reviews Methods Primers. Link
- SuntheticsML technology overview. sunthetics.io/our-technology
- Introducing Lithium. sunthetics.io/blog-posts/introducing-lithium-ml-algorithm-update
- Sunthetics Small-Data ML Outperforms Established Kinetic Optimization Model (Pfizer). Case study

