Sean Florez · Materials Science & Engineering
CSCI 7000, Fall 2026
Materials science is how a new technology becomes possible at all. For most of its history it ran by hand. Synthesize a candidate, measure it, repeat. Discovery to deployment still takes 10–20 years. Computation changed the economics: predict the material before anyone makes it.
Differentiate E(r₁…rN) for the force on every atom, integrate the forces for dynamics and which structures are stable, take statistics over that, and nearly every property anyone wants comes out the other end.
Solve for the electrons. No formula is assumed; you approximate the quantum mechanics directly. Accurate, and it scales as O(N³), so a few hundred atoms is the ceiling.
A human writes the formula. This one has two numbers per element pair: ε, how deep the well is, and σ, where it crosses zero. Fit them to data. O(N) and fast, and valid only where it was fitted.
There is no formula. A network reads one atom’s neighbours and returns that atom’s energy; the total is the sum, and forces are its derivative. Millions of parameters, all fit to DFT labels. O(N).
Train once on one broad corpus. Then run it zero-shot, or fine-tune.
An open database of over 150,000 inorganic materials, with computed structures, formation energies, band structures and elastic constants. Every number in it is calculated. MACE-MP-0 trains on MPtrj, 1.58M snapshots taken from those calculations.
Look at the periodic table at the top of that figure. That is the corpus, drawn as element counts, and it is the shape of everything the model knows. The elements it has nothing for stay dark.
MACE-MP-0 · GNoME · MatterSim · UMA, and a dozen more. Same graph, same message passing, same DFT labels. What changes is the data. Pt surfaces is one of the 22 domains on the right. My own system, and I didn’t train it.
Error falls as a power law in data.
| Checkpoint | Corpus | F1 |
|---|---|---|
| EquiformerV3+DeNS-MP | 1.58M frames | 0.863 |
| EquiformerV3+DeNS-OAM | 113M frames | 0.931 |
Same 30.3M parameters, 70× more labelled frames.
The dashed line is the part that matters: verified results become training data, so the filter gets better each round. Six rounds took GNoME’s hit rate from under 6% to over 80%.
Text is self-supervised because it already exists. Atomic energies do not.
Every training label in this field is a calculation. DFT approximates quantum mechanics, and its disagreements with measurement are systematic, known, and inherited.
A titanium potential trained to 6.0 meV/atom against DFT is 24% wrong on shear modulus against experiment.
Autonomous synthesis and characterization. An instrument makes the candidate and measures it, and that measurement enters training as a label.
Even A-Lab’s perception layer is simulation-trained, deciding which phase it made from a network fitted to simulated diffraction patterns.
Close to $1 billion since 2023 is betting the fix is an autonomous lab. Slide 11 has what they published; S1b has what the literature did.
An MLIP answers one question at one level, and nobody wants energies. The properties people buy live several levels up, and every handoff across this row is fitted by hand.
Weather solved this with per-dataset encoders into one shared space. They had ERA5, a single self-consistent reanalysis. Materials has no ERA5.
We optimize energy above hull. A learned model computes it.
The surrogate is queried on “inputs that deviate significantly from expected scenarios.”
| Carries over | Does not |
|---|---|
| Goodhart. Over-optimization. | No agent. Nothing here has goals. |
| RL against a learned scorer. | Ground truth exists. Make it and measure it. |
Chemeleon2, GRPO against exactly this reward: 80.3% elemental substitution, truly unmatched 18.9% → 11.4%.
Close to $1 billion since 2023, and three incompatible bets on limitation 1. Everything peer-reviewed any of them has published is infrastructure: engines, datasets, benchmarks.
| Company | Bet on experimental data | What is public |
|---|---|---|
| Orbital London | Simulation is enough. Orb is trained on DFT only, and a stated commercial metric is experiments avoided | Orb potentials, Apache-2.0, independently ranked on Matbench Discovery. BASF licensed the software |
| Radical AI New York | Both, and the join is unfinished. Their potential EGIP is DFT-trained; lab measurements currently tune a semi-empirical model and a vision model | TorchSim, MATRIX, LitXBench. Their own blog calls feeding lab data into the potential “the next problem… we’re currently working through this” |
| Lila Sciences Cambridge, Mass. | The lab is the verifier. RL on a reasoning model, with the experiment as the reward | Nothing technical |
| Periodic Labs Menlo Park, Calif. | The lab manufactures the reward. RL on a frontier LLM, with the measurement as the reward signal. “Nature is the RL environment.” | Four papers, none using their own experimental data. Their one MLIP paper benchmarks potentials against DFT. Lab still being commissioned |
The architectures are borrowed and they work. Data provenance, evaluation and the objective are where this field breaks, and all three are foundation-model problems.
Autonomous synthesis and characterization. An instrument makes the candidate and measures it without a person in the middle, and that measurement enters training as a label. The model then answers to something that was observed.
Where I would start: which observables can constrain a potential at all, and what a single measurement is worth counted in DFT labels.
A single learned representation running from atoms through microstructure to process, so a prediction at the bottom reaches a property somebody buys.
Where I would start: two adjacent levels, contrastively aligned, on a system I already run.
Write the correspondence nobody has written between the LLM alignment stack and this one: reward model, proxy, over-optimization, verifier gaming.
Where I would start: reproduce Goodhart on open models and open code. That one is runnable this semester.
Three limitations in the state of practice.
Which experimental observable can constrain a potential, and what one measurement is worth in DFT-labels
What a materials ERA5 would even be
What over-optimization looks like, measured, in this domain
| Model | Yr | Who | Params | Pre-training corpus |
|---|---|---|---|---|
| MACE-MP-0 | 2023 | Cambridge | 4.7M | MPtrj, 1.58M frames |
| GNoME | 2023 | Google DeepMind | 16.2M | GNoME, 89M frames |
| MatterSim | 2024 | Microsoft | 4.6M | 17M structures, closed |
| UMA | 2025 | Meta FAIR | 1.4B | ~500M structures, 5 domains |
Everything else is a variation on these. M3GNet, CHGNet, SevenNet, Orb, eSEN, EquiformerV3: same graph, same message passing, same DFT labels. They differ in corpus size, in how rotational symmetry is built in, and in how much compute went into training.
The lineage runs from Behler–Parrinello (2007), which fixed total energy as a sum of per-atom energies, a choice every model above still inherits.
A router reads global metadata (charge, spin, which DFT setting produced the label) and mixes expert weight matrices once per structure. 1.4B parameters collapse to ~50M in the forward pass.
An ensemble flags configurations it is unsure about, molecular dynamics pushes them off equilibrium, and a “first-principles supervisor” (DFT) labels whatever survives. The database feeds back into training.
Experiment-informed potentials exist and they work. They are built one system at a time, and none of them is a foundation model. Meanwhile two reviews, two years apart, describe this field and neither mentions the other’s subject once.
| Term | Occurrences |
|---|---|
| self-driving | 0 |
| autonomous lab | 0 |
| closed-loop | 0 |
| Bayesian optimization | 0 |
| Term | Occurrences |
|---|---|
| interatomic | 0 |
| MLIP | 0 |
| neural network potential | 0 |
| MACE | 23 |
Its Table 4 tabulates 102 self-optimizing platforms. What they optimize with: local or heuristic search 48, Bayesian optimization 41, custom 7, deep RL 4, active learning 2. Not one is atomistic.
Fine-tune a DFT-trained titanium potential on experimental elastic constants alone: energy RMSE goes 6.0 → 385 meV/atom, while forces (92.5→123.6 meV/Å) and virials (406.6→401.5) barely move. Those observables only see derivatives, so experiment alone leaves the energy scale free. Fuse the two objectives and it works: 7.9 meV/atom, and the elastic constants match. One element, hand-built.
The 2026 state of the art still keeps the weights DFT. PET-UAFD calibrates five universal potentials against 266 experimental systems, and only the ensemble weighting moves. The network weights never see an experiment. Kellner et al., arXiv:2604.24607 · The honest exception: Strieth-Kalthoff et al., Science 384 (2024) ran a pretrained GNN across five live labs, frozen as a featurizer with a GP head on 287 measurements, for 21 new state-of-the-art laser gain materials. It never touches a force, an energy or a configuration.
What it does. Electrospun nanofibers, characterised in one lab. Three modalities off the same specimens: processing parameters (flow rate, concentration, voltage), SEM images of the microstructure, and measured mechanical properties. A table encoder, a vision encoder and a fused encoder project all three into one shared latent space, CLIP-style, so a missing modality can be recovered from the others.
It predicts properties without structural input at inference, and generates microstructures from processing parameters.
Worth being precise, because it is easy to assume otherwise: MatMCL never touches atoms. It bridges process ↔ microstructure ↔ property, which is the top of the table on slide 10. There is no molecular dynamics in it and no energy anywhere.
So the mechanism is proven and the span is not. One material system, two adjacent levels, a lab-built dataset. The atomic level, where every foundation model in this talk lives, is the row that still connects to nothing.