Benchmark Puts AI Agents to Work on Materials Discovery

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- AI models are benchmarked on a task to propose novel, dynamically stable, BEOL-compatible crystalline materials meeting targets for thermal conductivity, static dielectric constant, Young's modulus, and shear modulus, with every candidate required to come with a synthesis recipe an expert would attempt.
- The benchmark harness equips models with web search via Exa, a Python/bash sandbox with pymatgen, mp_api, and ASE, plus ML-based tools for dynamic stability, lattice thermal conductivity, static dielectric constant, and the compliance tensor.
- Models run without a stopping condition until they hit an error or exhaust a 100 million token budget, evaluated using the UK AI Security Institute's open-source Inspect framework.
- Property calculations lean on the PET-MAD universal machine learning interatomic potential, with Pheasy and Phonopy for phonon physics and a GMTNet model fitted to the JARVIS DFPT database for the static dielectric constant.
- Synthesis recipes are graded by GPT-5.6 Sol (worst of three runs) against a penalty-based rubric of critical and fixable items — in the article's worked example, Claude Opus 5's hexagonal diamond (lonsdaleite) proposal was graded WOULD NOT ATTEMPT because it lacked a credible pathway to the ordered P6₃/mmc phase.
Why it matters: The benchmark exposes a gap between AI's ability to hit computed materials-property targets and its ability to propose synthesis routes an experimentalist would actually attempt. Claude Opus 5's lonsdaleite recipe checked the computational boxes but lost on a critical phase-selection penalty — novel-material candidates still need credible experimental pathways, not just ML-friendly numbers.
Ask SkimNews


