Where are we?

Foundation models for materials, and where their labels come from

Sean Florez · Materials Science & Engineering
CSCI 7000, Fall 2026

The accuracy-efficiency spectrum: ab initio methods on the left, force fields on the right, machine learning in between, across system sizes from small molecules to a virion
Unke et al., Chem. Rev. 121, 10142 (2021), Fig. 1
Materials science

New technology relies on new materials

Materials science is how a new technology becomes possible at all. For most of its history it ran by hand. Synthesize a candidate, measure it, repeat. Discovery to deployment still takes 10–20 years. Computation changed the economics: predict the material before anyone makes it.

Given the positions of N atoms, what is the energy of that arrangement?

Differentiate E(r₁…rN) for the force on every atom, integrate the forces for dynamics and which structures are stable, take statistics over that, and nearly every property anyone wants comes out the other end.

A perovskite unit cell: a repeating three-dimensional arrangement of three atomic species
Five materials-science tasks that follow from an energy model: structure optimization, phonon prediction, mechanical property, phase diagram, molecular dynamics
Lattice: perovskite CaTiO₃, Solid State, Wikimedia Commons (2010), CC BY-SA 3.0. Tasks: Yang, Hu, Zhou, Liu, Shi, Li et al., MatterSim, arXiv:2405.04967 (2024), Fig. 1, left column; the five panels cropped and set in a row. One structure in. All of these out, and every one is the same calculation asked a different question.
State of practice

Three ways to answer that question

DFT

A benzene molecule inside a coloured electrostatic potential surface, red at the ring centre through green at the edges
Benzene, with the electrostatic potential surface computed around it. ChiralJon, Wikimedia Commons (2015), CC BY 2.0.

Solve for the electrons. No formula is assumed; you approximate the quantum mechanics directly. Accurate, and it scales as O(N³), so a few hundred atoms is the ceiling.

Force fields

Lennard-Jones potential curve, with the well depth epsilon and the zero crossing sigma marked
Lennard-Jones 12-6: V(r) = 4ε[(σ/r)12 − (σ/r)6]

A human writes the formula. This one has two numbers per element pair: ε, how deep the well is, and σ, where it crosses zero. Fit them to data. O(N) and fast, and valid only where it was fitted.

MLIPs

One atom neighbourhood drawn as a graph, feeding a network, returning that atom energy
One atom’s neighbourhood, straight to that atom’s energy

There is no formula. A network reads one atom’s neighbours and returns that atom’s energy; the total is the sum, and forces are its derivative. Millions of parameters, all fit to DFT labels. O(N).

DFT is the teacher. The force field is the old shortcut. The MLIP learns the shortcut from the teacher.
State of practice

Atoms get the same recipe as text

Train once on one broad corpus. Then run it zero-shot, or fine-tune.

Trained on Materials Project

An open database of over 150,000 inorganic materials, with computed structures, formation energies, band structures and elastic constants. Every number in it is calculated. MACE-MP-0 trains on MPtrj, 1.58M snapshots taken from those calculations.

Look at the periodic table at the top of that figure. That is the corpus, drawn as element counts, and it is the shape of everything the model knows. The elements it has nothing for stay dark.

MACE-MP-0 · GNoME · MatterSim · UMA, and a dozen more. Same graph, same message passing, same DFT labels. What changes is the data. Pt surfaces is one of the 22 domains on the right. My own system, and I didn’t train it.

MACE-MP-0 trained on Materials Project, deployed across 22 domains
Batatia, Benner, Chiang, Elena, Kovács, Riebesell et al., “A Foundation Model for Atomistic Materials Chemistry” arXiv:2401.00096 (2024), Fig. 1. Shown unmodified · CC BY-NC-ND.
State of practice

Data scaling

Out-of-domain MAE versus training-set size, log-log
Merchant et al., GNoME, Nature 624 (2023), Fig. 1e · CC BY 4.0. Out-of-domain error against training-set size.

Error falls as a power law in data.

Held constant: the architecture

CheckpointCorpusF1
EquiformerV3+DeNS-MP1.58M frames0.863
EquiformerV3+DeNS-OAM113M frames0.931

Same 30.3M parameters, 70× more labelled frames.

Two EquiformerV3+DeNS checkpoints from the Matbench Discovery leaderboard, August 2026. MP is trained on MPtrj; OAM is OMat24 pre-training then fine-tuned on MP and sAlex. Leaderboard: Riebesell, Goodall, Benner, Chiang, Deng, Ceder, Asta, Lee, Jain & Persson, Nat. Mach. Intell. 7, 836 (2025).
How that data is made

Propose, verify with DFT, retrain, repeat

GNoME discovery pipeline with active-learning loop
Merchant, Batzner, Schoenholz, Aykol, Cheon & Cubuk, “Scaling deep learning for materials discovery,” Nature 624, 80–85 (2023), Fig. 1a · CC BY 4.0.
Every arrow into that database is a DFT calculation.

The dashed line is the part that matters: verified results become training data, so the filter gets better each round. Six rounds took GNoME’s hit rate from under 6% to over 80%.

Turning point

Every label is manufactured

Text is self-supervised because it already exists. Atomic energies do not.

“Emphatically, unlike the case of language or vision, in materials science, we can continue to generate data.”
GNoME: Merchant et al., Nature 624, 80 (2023). OMol25 cost 6.6 billion CPU core-hours.
OMat24 dataset construction: Boltzmann sampling, ab-initio MD and rattled relaxation
Barroso-Luque, Shuaibi, Fu, Wood, Dzamba, Gao, Rizvi, Zitnick & Ulissi, OMat24, arXiv:2410.12771 (2024), Fig. 1(a), cropped. 100M+ structures, three sampling procedures.
Where the state of practice breaks, in three places
1
Where labels come from. Every one is manufactured
2
Atoms to parts. One model covers one level of five
3
What we optimize. A computed stand-in
Limitation 1 · where labels come from

DFT is the teacher, and the teacher is an approximation

Every training label in this field is a calculation. DFT approximates quantum mechanics, and its disagreements with measurement are systematic, known, and inherited.

Accurate against DFT, wrong against reality

A titanium potential trained to 6.0 meV/atom against DFT is 24% wrong on shear modulus against experiment.

The model is not bad. It is faithful to the wrong reference.

Where a measurement would come in

Autonomous synthesis and characterization. An instrument makes the candidate and measures it, and that measurement enters training as a label.

Even A-Lab’s perception layer is simulation-trained, deciding which phase it made from a network fitted to simulated diffraction patterns.

Close to $1 billion since 2023 is betting the fix is an autonomous lab. Slide 11 has what they published; S1b has what the literature did.

Shear modulus against temperature: the DFT-trained model sits well below the experimental reference across the whole range
Shear modulus against temperature. Röcken & Zavadlav, npj Comput. Mater. 10, 69 (2024), Fig. 3(b), cropped · CC BY 4.0. Black dashed is experiment. Red is the DFT-trained model, off across the whole range. The other two saw experimental data.
Train on DFT, test against DFT, optimize toward DFT. Reality is optional.
Limitation 2 · atoms to parts

Nothing connects the atom to the part

An MLIP answers one question at one level, and nobody wants energies. The properties people buy live several levels up, and every handoff across this row is fitted by hand.

Five levels from atomistic lattice structure through dislocation dynamics, subgrain structures and polycrystalline grain structure to macroscopic material behaviour, spanning nanometres and nanoseconds to millimetres and milliseconds
Multiscale modeling, Mechanics & Materials, ETH Zürich (mm.ethz.ch). Used for coursework with attribution. Foundation models live in the first circle only.
“There are no existing foundation model architectures specifically designed for multiscale modeling problems in materials science.”
Van, Verma, Zhao & Wu, arXiv:2506.20743 (2025)

Weather solved this with per-dataset encoders into one shared space. They had ERA5, a single self-consistent reanalysis. Materials has no ERA5.

Limitation 3 · what we optimize

Optimization is what pushes a model off its own map

We optimize energy above hull. A learned model computes it.

Reward hacking, as chemistry means it

The surrogate is queried on “inputs that deviate significantly from expected scenarios.”

Yoshizawa et al., DyRAMO, Nat. Commun. 16, 2409 (2025). No exploit. No adversary.
A search settles where the score reads highest.
Optimization manufactures the shift, then follows it.
Carries overDoes not
Goodhart. Over-optimization.No agent. Nothing here has goals.
RL against a learned scorer.Ground truth exists. Make it and measure it.
Energy above hull from a learned model plotted against DFT, R-squared 0.84
Xu et al., PLaID++, arXiv:2509.07150 (2025), Fig. 10 (left panel, eSEN). A verifier against the truth it stands in for, and this is their better one.

Chemeleon2, GRPO against exactly this reward: 80.3% elemental substitution, truly unmatched 18.9% → 11.4%.

The reward went up.
The goal stayed where it was.
Commercial state of practice

Nobody agrees whether simulation is enough

Close to $1 billion since 2023, and three incompatible bets on limitation 1. Everything peer-reviewed any of them has published is infrastructure: engines, datasets, benchmarks.

CompanyBet on experimental dataWhat is public
Orbital
London
Simulation is enough. Orb is trained on DFT only, and a stated commercial metric is experiments avoidedOrb potentials, Apache-2.0, independently ranked on Matbench Discovery. BASF licensed the software
Radical AI
New York
Both, and the join is unfinished. Their potential EGIP is DFT-trained; lab measurements currently tune a semi-empirical model and a vision modelTorchSim, MATRIX, LitXBench. Their own blog calls feeding lab data into the potential “the next problem… we’re currently working through this”
Lila Sciences
Cambridge, Mass.
The lab is the verifier. RL on a reasoning model, with the experiment as the rewardNothing technical
Periodic Labs
Menlo Park, Calif.
The lab manufactures the reward. RL on a frontier LLM, with the measurement as the reward signal. “Nature is the RL environment.”Four papers, none using their own experimental data. Their one MLIP paper benchmarks potentials against DFT. Lab still being commissioned
Machine learning works best on the training set distribution. In science and technology, we almost only care about out-of-domain generalization.
Ekin Doğuş Çubuk, senior author of GNoME, on why he left that paradigm to found Periodic Labs · Catalyst, 6 Nov 2025. Lila’s Chief Scientific Officer for Physical Sciences, tenured at MIT, says the same thing harder. Nobody in this table has published a material that its own lab made and its own model predicted.
Where I want to work
Making materials foundation models answer to reality.

The architectures are borrowed and they work. Data provenance, evaluation and the objective are where this field breaks, and all three are foundation-model problems.

Lab in the loop

Autonomous synthesis and characterization. An instrument makes the candidate and measures it without a person in the middle, and that measurement enters training as a label. The model then answers to something that was observed.
Where I would start: which observables can constrain a potential at all, and what a single measurement is worth counted in DFT labels.

One model, many levels

A single learned representation running from atoms through microstructure to process, so a prediction at the bottom reaches a property somebody buys.
Where I would start: two adjacent levels, contrastively aligned, on a system I already run.

Alignment mapping

Write the correspondence nobody has written between the LLM alignment stack and this one: reward model, proxy, over-optimization, verifier gaming.
Where I would start: reproduce Goodhart on open models and open code. That one is runnable this semester.

Questions

Thank you

Three limitations in the state of practice.

1
Where labels come from. Every one is manufactured
2
Atoms to parts. One model covers one level of five
3
What we optimize. A computed stand-in
And what is genuinely still open

Which experimental observable can constrain a potential, and what one measurement is worth in DFT-labels

What a materials ERA5 would even be

What over-optimization looks like, measured, in this domain

Supplementary · S0

Popular methods, by name

ModelYrWhoParamsPre-training corpus
MACE-MP-02023Cambridge4.7MMPtrj, 1.58M frames
GNoME2023Google DeepMind16.2MGNoME, 89M frames
MatterSim2024Microsoft4.6M17M structures, closed
UMA2025Meta FAIR1.4B~500M structures, 5 domains

Everything else is a variation on these. M3GNet, CHGNet, SevenNet, Orb, eSEN, EquiformerV3: same graph, same message passing, same DFT labels. They differ in corpus size, in how rotational symmetry is built in, and in how much compute went into training.

The lineage runs from Behler–Parrinello (2007), which fixed total energy as a sum of per-atom energies, a choice every model above still inherits.

Supplementary · S1

UMA’s architecture and MatterSim’s data engine

UMA Mixture of Linear Experts architecture
Wood, Dzamba, Fu, Gasteiger, Barroso-Luque, Shuaibi, Zitnick et al., “UMA: A Family of Universal Models for Atoms,” NeurIPS 2025, Fig. 2 (2025).

A router reads global metadata (charge, spin, which DFT setting produced the label) and mixes expert weight matrices once per structure. 1.4B parameters collapse to ~50M in the forward pass.

MatterSim active-learning data generation loop
Yang, Hu, Zhou, Liu, Shi, Li et al., MatterSim, arXiv:2405.04967 (2024), Fig. 2, panel (a), cropped at native resolution.

An ensemble flags configurations it is unsure about, molecular dynamics pushes them off equilibrium, and a “first-principles supervisor” (DFT) labels whatever survives. The database feeds back into training.

Supplementary · S1b

Every foundation-scale potential is trained on DFT alone

Experiment-informed potentials exist and they work. They are built one system at a time, and none of them is a foundation model. Meanwhile two reviews, two years apart, describe this field and neither mentions the other’s subject once.

Atomistic foundation-model review

TermOccurrences
self-driving0
autonomous lab0
closed-loop0
Bayesian optimization0
Yuan, Liu et al. & Head-Gordon, “Foundation models for atomistic simulation of chemistry and materials,” arXiv:2503.10538; published as Nat. Rev. Chem. 10, 212 (2026). Counts run on the preprint full text.

Self-driving-lab review

TermOccurrences
interatomic0
MLIP0
neural network potential0
MACE23
Tom, Schmid et al. & Aspuru-Guzik, Chem. Rev. 124, 9633 (2024), 476 references. All 23 “MACE” hits are inside the word pharmaceutical.

Its Table 4 tabulates 102 self-optimizing platforms. What they optimize with: local or heuristic search 48, Bayesian optimization 41, custom 7, deep RL 4, active learning 2. Not one is atomistic.

It has been done, on one system

Fine-tune a DFT-trained titanium potential on experimental elastic constants alone: energy RMSE goes 6.0 → 385 meV/atom, while forces (92.5→123.6 meV/Å) and virials (406.6→401.5) barely move. Those observables only see derivatives, so experiment alone leaves the energy scale free. Fuse the two objectives and it works: 7.9 meV/atom, and the elastic constants match. One element, hand-built.

Röcken & Zavadlav, npj Comput. Mater. 10, 69 (2024). The same paper: that 6.0 meV/atom DFT-trained model is 24% wrong on shear modulus. Neither reference is measuring the other.

The 2026 state of the art still keeps the weights DFT. PET-UAFD calibrates five universal potentials against 266 experimental systems, and only the ensemble weighting moves. The network weights never see an experiment. Kellner et al., arXiv:2604.24607  ·  The honest exception: Strieth-Kalthoff et al., Science 384 (2024) ran a pretrained GNN across five live labs, frozen as a featurizer with a GP head on 287 measurements, for 21 new state-of-the-art laser gain materials. It never touches a force, an energy or a configuration.

Supplementary · S1c

How close MatMCL actually gets

MatMCL: table, multimodal and vision encoders projected into one shared joint space
Wu, Ding, He, Wu, Jiang, Zhang & Ji, MatMCL, npj Comput. Mater. 11, 276 (2025), Fig. 1(b), cropped · CC BY 4.0.

What it does. Electrospun nanofibers, characterised in one lab. Three modalities off the same specimens: processing parameters (flow rate, concentration, voltage), SEM images of the microstructure, and measured mechanical properties. A table encoder, a vision encoder and a fused encoder project all three into one shared latent space, CLIP-style, so a missing modality can be recovered from the others.

It predicts properties without structural input at inference, and generates microstructures from processing parameters.

Which levels it actually bridges

Worth being precise, because it is easy to assume otherwise: MatMCL never touches atoms. It bridges process ↔ microstructure ↔ property, which is the top of the table on slide 10. There is no molecular dynamics in it and no energy anywhere.

So the mechanism is proven and the span is not. One material system, two adjacent levels, a lab-built dataset. The atomic level, where every foundation model in this talk lives, is the row that still connects to nothing.

Supplementary · S2

Where the architecture came from

Equivariant GNN architecture
Batzner et al., NequIP, Nat. Commun. 13, 2453 (2022), Fig. 1 · CC BY 4.0
  • 2007 Behler–Parrinello: total energy = sum of per-atom energies. Every modern MLIP inherits this.
  • Atoms as nodes, neighbours within a cutoff as edges; message passing over the graph.
  • Equivariance is built into the architecture: features are irreps of O(3), combined by Clebsch–Gordan tensor products.
  • 2026: EquiformerV3 is a graph attention transformer with SwiGLU ported to the sphere. The LLM playbook, applied to atoms.