Skip to content

Human and LLM Collaboration for Accelerated Materials Synthesis and Discovery

Source: arXiv:2607.07604 · Published 2026-07-08 · By Gregory Bassen, Wyatt Bunstine, Sarah Okandey, Sarah Cheung, Elaine Flowers, Ritwik Bose et al.

TL;DR

This paper investigates the complementary roles of human chemists and Large Language Models (LLMs) in predicting solid-state synthesis routes for known and novel materials within the Ruddlesden-Popper (RP) homologous series. While LLMs have accelerated prediction of candidate materials, experimental synthesis remains a bottleneck due to the complexity of designing reproducible reaction pathways and the incomplete literature reporting of critical synthesis variables. By experimentally testing synthesis recipes independently generated by humans and LLMs, then feeding outcomes back into a closed-loop refinement, the authors assess relative success rates and collaborative potential. Both humans and LLMs achieved similar success rates for known materials (approximately 75–83%) and lower but comparable success rates for unknown materials (14–22%). The collaboration culminated in the discovery of Ba3PtO5, a novel 1D perovskite-derived structural prototype, extending the RP family to a broader Rock-Salt Perovskite (RSP) homologous series. These results demonstrate that LLMs can approach human expert performance in synthesis planning and that iterative human-LLM collaboration can accelerate experimental materials discovery.

Key findings

  • Humans synthesized known RP materials with 83(±8)% success rate in round one vs LLMs 75(±9)%.
  • For unknown RP targets, humans achieved 17(±9)% success and LLMs 22(±10)% in round one.
  • After closed-loop feedback (round two), success rates slightly decreased: known material synthesis at 79(±8)% (human) vs 71(±9)% (LLM).
  • Unknown material discovery post-feedback was 22(±7)% (human) vs 14(±6)% (LLM).
  • LLM plans tended to specify more detailed reaction parameters (e.g., temperature ranges, apparatus type) than human recipes.
  • Humans outperformed LLMs in discovery of new phases after iterative experimental feedback.
  • Serendipitous synthesis of Ba3PtO5, a new 1D perovskite-derived phase, was enabled by human flux growth informed by PXRD feedback.
  • Variability observed between parallel experimental trials highlights sensitivity of synthesis outcomes to subtle procedural and environmental factors.

Threat model

The adversary is implicitly modeled as the originator of synthesis plans (human or LLM), with capabilities restricted to proposing reaction sequences. They cannot manipulate experimental equipment or conditions beyond recipe descriptions. The study focuses on efficacy of synthesis planning rather than active adversarial attacks or information leakage.

Methodology — deep read

The threat model involves an adversary limited to proposing synthesis recipes; success depends on recipe reproducibility and experimental validation. No direct adversarial scenario is discussed. The study used a set of target compositions primarily from the Ruddlesden-Popper homologous series An+1BnX3n+1 with chosen A (alkaline earth or rare earth), B (transition metals), and X = O to balance stability and exploration of new phases.

The experimental workflow involves:

  1. Selecting a target compound, either known or previously unreported.
  2. A human solid-state chemist independently writing a synthesis recipe based on domain expertise.
  3. Querying an LLM to generate an independent synthesis recipe for the same target.
  4. Executing both human and LLM recipes in two parallel experimental trials (A and B) via standard solid-state and flux methods.
  5. Characterizing resulting solids by powder X-ray diffraction (PXRD) to assess phase formation.
  6. Feeding PXRD results back to both human and LLM, who then revise synthesis plans iteratively (closed-loop).

All recipes are conducted under ambient atmosphere or specified oxygen conditions, using commercial precursors (e.g., BaCO3, CeO2) pre-weighed and mixed. PXRD data collected on Bruker D8 diffractometers with Cu Kα radiation, scanning 2θ=5°–60°, refined using GSAS software to determine phase purity. Single crystal X-ray diffraction (SCXRD) was performed for novel phases (e.g., Ba3PtO5) to solve crystal structure.

LLM outputs differed by providing explicit instructions for apparatus (e.g., Al2O3 vs Pt crucibles), temperature ranges, grinding styles, and heating/cooling rates. Humans used classical methods focused on flux synthesis for challenging oxidation states.

Evaluations considered synthesis success as formation of the known target phase by PXRD for established materials, and detection of new phases by unidentified PXRD peaks for unknown targets. Success rates and uncertainties were computed from batch trial results. Comparisons included multiple LLM models, with temperature parameter differences noted. Reproducibility was addressed by running dual trials per recipe.

The discovery of Ba3PtO5 was a concrete example where human-designed flux synthesis modified by PXRD feedback yielded single crystals characterized by SCXRD, revealing a new 1D connectivity motif extending the RP series to a broader Rock-Salt Perovskite homologous family (AX)m(ABX3)p. The workflow thus combines automated LLM planning with human intuition and experimental iteration to accelerate materials discovery.

Code for recipe generation is not public; experimental data and synthesized samples are curated at Johns Hopkins University with plans for future online data release.

Technical innovations

  • Use of closed-loop experimental feedback to iteratively refine both human- and LLM-generated solid-state synthesis recipes.
  • Direct prospective in-laboratory validation of LLM-predicted synthesis recipes for complex inorganic oxide materials.
  • Discovery and structural elucidation of a novel 1D perovskite prototype (Ba3PtO5) bridging known dimensionalities of RP series.
  • Characterization of synthesis recipe success rates comparing human expertise versus diverse LLM outputs on a broad homologous series.

Datasets

  • Ruddlesden-Popper homologous series targets — ~20 compounds — internal experimental dataset at Johns Hopkins University

Baselines vs proposed

  • Human round one known materials success rate = 83(8)% vs LLM round one known materials = 75(9)%
  • Human round one unknown materials discovery = 17(9)% vs LLM round one unknown = 22(10)%
  • Human round two known success = 79(8)% vs LLM round two known = 71(9)%
  • Human round two unknown discovery = 22(7)% vs LLM round two unknown = 14(6)%

Figures from the paper

Figures are reproduced from the source paper for academic discussion. Original copyright: the paper authors. See arXiv:2607.07604.

Fig 1

Fig 1: Overall workflow for the synthesis prediction experiment. a) A material candidate from the Ruddlesden-Popper series

Fig 2

Fig 2: Sample prompt and corresponding synthesis plans.

Fig 3

Fig 3: Experimental results of human vs LLM study. (a)

Fig 5

Fig 5: illustrates this progression, in which successive in-

Fig 4

Fig 4: Comparison of human synthesis plans to different

Limitations

  • The dataset is limited to a single homologous materials series (Ruddlesden-Popper oxides), so generalization is unclear.
  • No adversarial robustness testing of LLM-generated recipes was performed; failures were experimental.
  • Environmental and procedural variation between experimental trials caused inconsistencies, impacting reproducibility.
  • LLM recipes sometimes lacked full domain context, such as addressing unusual oxidation states requiring specialized atmospheres.
  • Comparison was limited to controlled human expert and LLM plans; other synthesis automation pipelines or AI models were not benchmarked.
  • Detailed closed-loop learning dynamics and improvements over multiple feedback rounds beyond two iterations remain unquantified.

Open questions / follow-ons

  • How do LLM-driven synthesis plans scale to more chemically diverse and complex materials classes beyond oxides?
  • Can automated feedback loops with real-time experimental data further improve synthesis success rates via machine optimization?
  • What is the role of environmental variables and minor procedural differences on LLM vs human reproducibility in synthesis?
  • How can domain knowledge gaps (e.g., handling high oxidation states) be encoded or augmented in LLM prompting to improve predictions?

Why it matters for bot defense

For bot-defense or CAPTCHA practitioners, this study demonstrates how closed-loop collaboration between human expertise and advanced AI (LLMs) can achieve task performance comparable to domain experts while accelerating experimental validation cycles. Similarly, in CAPTCHA contexts where iterative human-in-the-loop and AI approaches may be combined to improve challenge design or response synthesis, the principles of feedback-driven refinement and joint learning are applicable. The paper highlights that AI-generated procedural plans can match human performance given structured feedback, underscoring the importance of iterative validation and error correction in security-critical workflows. Additionally, the variability observed in parallel synthesis trials cautions about inherent stochasticity, which analogously might affect robustness in synthesizing adversarial or defensive CAPTCHA patterns from AI outputs. Overall, bot-defense engineers may find inspiration in this structured human-AI collaboration framework for evolving defenses or detection methods that integrate automated proposal with expert curation.

Cite

bibtex
@article{arxiv2607_07604,
  title={ Human and LLM Collaboration for Accelerated Materials Synthesis and Discovery },
  author={ Gregory Bassen and Wyatt Bunstine and Sarah Okandey and Sarah Cheung and Elaine Flowers and Ritwik Bose and Joshua Hummel and Christopher D. Stiles and Maxime A. Siegler and Tyrel M. McQueen },
  journal={arXiv preprint arXiv:2607.07604},
  year={ 2026 },
  url={https://arxiv.org/abs/2607.07604}
}

Read the full paper

Articles are CC BY 4.0 — feel free to quote with attribution