Case Studies

Beating High-Throughput with ML: Optimizing Categorical Variables in Cross-Coupling Chemistry

94% Fewer Experiments, Same Breakthrough Result

Key Takeaways

  • 94% fewer experiments, same optimal results—machine learning outperformed expensive automation.

  • Categorical ML breakthrough: SuntheticsML optimizes directly from labels without complex parameterization or one-hot encoding.

  • New scientific insight: Base identity—not catalyst—was the key variable, guiding smarter future experimentation.

  • Lower barrier to entry: SuntheticsML achieves elite-level results without the infrastructure or cost of high-throughput robotics.

  • Repeatable & reliable: All five ML seeds converged to the correct optimum, confirming robustness.
Company Name

Category

Published Date

Highlighted features / use cases

Client

This study was conducted by a joint team from Ghent University and Sunthetics. The goal: optimize a Suzuki–Miyaura cross-coupling reaction using real-world, industrial-grade datasets—where categorical variables like solvent and catalyst dominate the design space​.


Challenge

  • Identify the best combination of catalyst, solvent, and base for a complex cross-coupling reaction.

  • Achieve this using minimal experimentation.

  • Tackle the high-dimensional complexity of categorical variables, which are often not easily optimized by standard ML or DoE tools.

  • Avoid the 768 experiments typically required by High-Throughput Experimentation (HTE) to explore the full design space.

Goal

  • Find the global optimum for reaction yield.

  • Cut down resource-intensive experimentation.

  • Prove that small-data ML can rival HTE in speed and precision—even for categorical variables.

Approach & Solution

  • SuntheticsML was deployed using proprietary:

    • Supervised Learning (SL) and

    • Active Learning (AL) algorithms

  • No variable encoding or parameterization was required—SuntheticsML directly optimized categorical inputs like catalyst identity.

  • Each optimization was initiated with 36 randomly selected experiments.

  • The platform iteratively recommended 6 experiments per cycle: 5 likely to yield maximum performance + 1 random for exploration.

  • Five independent "seeds" were used to evaluate repeatability and robustness.

Results & Metrics

  • HTE benchmark:

    • 768 experiments required to explore full combinatorial space and find optimal conditions.

  • SuntheticsML outcomes:

    • Average performance:

      • 6 iterations, 72 experiments, 91% experiment reduction

    • Best-case scenario:

      • 2 iterations, 48 experiments, 94% reduction

    • Worst-case scenario:

      • 9 iterations, 84 experiments, 89% reduction

  • Insight generated:

    • Variable importance analysis showed the base had the strongest effect on yield—surprising the research team and shifting future optimization focus away from catalysts and solvents​
      .

The Sunthetics Edge


“SuntheticsML matched the output of a fully combinatorial HTE campaign in a fraction of the time and cost—while revealing insights HTE didn’t, like the true variable driving reaction performance.”


Key Takeaways

The Sunthetics Edge


“SuntheticsML matched the output of a fully combinatorial HTE campaign in a fraction of the time and cost—while revealing insights HTE didn’t, like the true variable driving reaction performance.”

‍

‍

  • 94% fewer experiments, same optimal results—machine learning outperformed expensive automation.

  • Categorical ML breakthrough: SuntheticsML optimizes directly from labels without complex parameterization or one-hot encoding.

  • New scientific insight: Base identity—not catalyst—was the key variable, guiding smarter future experimentation.

  • Lower barrier to entry: SuntheticsML achieves elite-level results without the infrastructure or cost of high-throughput robotics.

  • Repeatable & reliable: All five ML seeds converged to the correct optimum, confirming robustness.

‍