Act 8 · Seeing molecules

What chemistry cannot yet do

Prediction from first principles, catalysis by design, and the folding problem that turned out to be tractable after all. Three honest accounts of where the subject actually stands, and what the last of them did and did not settle.

1929 – now18 min
By the end you should be able to
  • State what Dirac claimed in 1929 and why it is both true and unhelpful
  • Explain why an error of a few kJ/mol in a calculated barrier is a large error
  • Say why catalysts are still found by screening rather than designed
  • Be precise about what CASP14 settled and what it left open

This is the last lesson of the chemistry path, so it should be an honest one. The subject started with a man weighing a sealed glass vessel and finding that the balance did not move. It arrives here able to photograph a single molecule and count the bonds in it. In between it acquired a conserved quantity, an atomic theory forced by whole-number ratios, a table whose gaps were filled by prediction, an explanation of that table from quantum mechanics, a theory of the bond, a thermodynamics that says which way a reaction goes, and a set of instruments that read structure off a chart. The way to respect that is not to end on a flourish. It is to say clearly what the method still cannot do — and to be equally careful not to mistake difficult for impossible, because one of the three problems below has just moved in a way almost nobody expected.

One: prediction from first principles

The underlying physical laws necessary for the mathematical theory of a large part of physics and the whole of chemistry are thus completely known, and the difficulty is only that the exact application of these laws leads to equations much too complicated to be soluble.

P. A. M. Dirac"Quantum Mechanics of Many-Electron Systems", Proceedings of the Royal Society A, 1929 — written three years after the Schrödinger equation, and still true in both halves.

Both halves of that sentence are worth taking seriously, because people usually quote only one of them. The first half is correct and is not a small claim. Chemistry is electrons and nuclei interacting electrostatically, obeying the Schrödinger equation. Nothing discovered in the century since has required an additional principle. Every bond in this course — the shared pair, the lattice, the aromatic ring, the enzyme’s transition state — is a consequence of that one equation and the Pauli principle. Chemistry is, in the strictest sense, reducible. The second half is also correct, and it is the reason reducible does not mean solved. A many-electron wavefunction is a function of three coordinates for every electron. Water has ten electrons, so its wavefunction is a function in thirty dimensions. Lay down a crude grid of ten points along each dimension and storing it takes 10³⁰ numbers, for a molecule with three atoms in it. Exact solution — full configuration interaction — scales exponentially with the size of the system. That is a qualitatively different obstacle from a hard calculation. Doubling your computer buys you a couple of extra electrons, not twice the molecule, and it will still be true in fifty years. Quantum computers are the one genuine prospect of changing the scaling here, and simulating chemistry is the application most often named for them; nothing yet built has done a calculation that a laptop could not.

What chemistry uses instead is density functional theory, and its foundation is one of the odder results in the subject. Hohenberg and Kohn proved in 1964 that the ground state is entirely determined by the electron density. Not by the wavefunction — by the density, ρ(x, y, z), a function of three coordinates however many electrons there are. The exponential disappears. Everything you want, including the total energy, is some functional of ρ. The proof is an existence proof. It establishes that the exact functional exists. It does not say what it is, and sixty years of work have not produced it. Kohn and Sham gave a practical scheme in 1965; Becke, Lee, Yang and Parr gave the approximation known as B3LYP in the early 1990s, which became one of the most-cited pieces of work in the history of science. Every one of those calculations approximates a function that is known to exist and known to nobody. This is not a complaint about a failed method. DFT works, it earned Walter Kohn and John Pople the Nobel Prize in Chemistry in 1998, and it is run a few hundred thousand times a year on real problems to real effect: reaction mechanisms, spectra, materials screening, catalysis. The honest difficulty is sharper than "it is approximate". The errors are not predictable in advance. A functional that handles organic thermochemistry to a few kJ/mol will fail on a transition-metal complex, or on dispersion interactions, or on anything strongly correlated, and nothing inside the calculation tells you which situation you are in. You find out by comparing with an experiment — which was the thing the calculation was meant to replace.

Problem

How accurate does a calculation have to be?

A calculated activation energy for a reaction comes out 8.0 kJ/mol too low. Rates follow the Arrhenius form, k = A·exp(−E_a/RT), and the pre-exponential factor is assumed correct. By what factor does the calculation overstate the rate at 298 K?

Error in barrier
ΔE_a = 8.0 kJ/mol, too low
Gas constant
R = 8.314 J mol⁻¹ K⁻¹
Temperature
T = 298.15 K
Rate law
k = A·exp(−E_a/RT)

Two: catalysis by design

The second thing chemistry cannot do is design a catalyst. The clearest way to see it is to look at how the most consequential catalyst in industrial history was found, and then to notice that it has not been replaced. Between 1909 and 1912, Alwin Mittasch and his group at BASF worked through roughly 2,500 different formulations in several thousand separate runs, searching for something that would combine nitrogen and hydrogen at a pressure a steel vessel could survive. The programme eventually ran to more than 20,000 tests. Haber’s bench catalyst had been osmium, and then uranium; there is not enough osmium in the world to run an industry on. What worked was iron — and specifically iron from a particular magnetite ore from Gällivare, in Swedish Lapland. Synthetic magnetite of nominally the same composition did not work. It took time to establish why: the impurities in the Swedish ore were doing the job. Alumina holds the iron surface open so it does not sinter shut; potassium makes the surface better at splitting the nitrogen triple bond. Two promoters, present by accident, and essential. That catalyst is essentially the catalyst in use today. Promoted iron, over a century later, in plants that fix more than a hundred million tonnes of nitrogen a year and feed roughly half the people alive. The one substantial alternative — ruthenium on graphite, commercialised in 1992 — occupies a small fraction of world capacity. The catalyst that underwrites modern agriculture was found by making a great many things and testing them, and nothing since has designed a better one.

There is a principle, and it is a good one, and it is not enough. Sabatier’s principle says the catalyst must bind the reagent neither too weakly nor too strongly: too weak and nothing sticks, too strong and the product will not leave. Plot activity against binding energy and you get a peak with falling sides — a volcano plot — and the best catalyst sits at the top. That tells you the shape of the answer. It does not tell you which material has the binding energy you want, and binding energy is exactly the quantity that first-principles calculation is least reliable about, because it involves a transition metal surface, which is the case where the functionals fail. Computational screening has produced real successes, and they should be counted honestly. A nickel–zinc alloy predicted from calculated binding energies and then made, for selective acetylene hydrogenation. A cobalt–molybdenum nitride for ammonia synthesis, designed by interpolating along the volcano between two metals that bind nitrogen too weakly and too strongly. These are genuine cases of a catalyst found by thinking rather than by trying, and there are perhaps a few dozen of them. And enzymes are the standing rebuke. Orotidine 5′-monophosphate decarboxylase accelerates its reaction by a factor of about 10¹⁷. Radzicka and Wolfenden measured the uncatalysed reaction in 1995 and found a half-life of 78 million years; with the enzyme it is a fraction of a second. The enzyme achieves this with no metal, no cofactor, and a handful of ordinary amino acid side chains held in a particular arrangement. Designed enzymes exist. They work. They also start out several orders of magnitude worse than natural ones, and they are improved mainly by directed evolution — random mutation, selection, repeat — which won Frances Arnold a share of the 2018 Nobel prize and is, in the end, screening at speed. It is an excellent method. It is not design.

Three: the folding problem

The third problem is the one that has actually moved, and the movement is recent enough that the story is usually told badly in both directions. The premise is Anfinsen’s. In experiments through the early 1960s, Christian Anfinsen unfolded ribonuclease completely — broke every disulphide bond, denatured it, reduced it to a floppy chain — and then removed the denaturant. It folded back into exactly the working enzyme, unaided, in a test tube containing nothing else. The conclusion, which took the 1972 Nobel Prize in Chemistry: the sequence determines the structure. All the information is in the order of the amino acids. And the paradox is Levinthal’s. In 1969 Cyrus Levinthal pointed out that this is nearly impossible. A chain of 100 residues with three accessible orientations per residue has about 3¹⁰⁰ ≈ 5 × 10⁴⁷ conformations. Sampling them at one every 10⁻¹³ seconds would take vastly longer than the age of the universe. Proteins fold in milliseconds to seconds. So folding cannot be a search — the chain must be running down some funnel in the energy landscape — and knowing that does not tell you where the bottom is. So: the information is in the sequence, provably. Getting it out was a fifty-year failure. And from 1994 there was a properly designed test of whether anyone was making progress.

The experiment

CASP14 — Critical Assessment of protein Structure Prediction, fourteenth round

Organised by John Moult and Krzysztof Fidelis; won by AlphaFold2, from DeepMind · 2020 — the assessment has run every two years since 1994 · Worldwide, blind

The question
Given only an amino acid sequence, can anyone predict the three-dimensional structure of the folded protein accurately enough to be useful — and can they do it on structures they have not seen?
The apparatus
The design of the test is the important part. Experimental groups supply structures they have solved but not yet published. The sequences alone are released to entrants; the coordinates are held back. Predictions are submitted to a deadline, then scored against the experimental answer by assessors who did not compete. Nobody can train on the answer, because the answer is not public. This has run for twenty-six years and is about as clean a blind trial as any field possesses.
Theory predicted

Steady, grinding improvement. Through the 2000s and 2010s the median GDT_TS score on the hardest targets — those with no known structure to use as a template — had crept from around 30 into the 50s and 60s. AlphaFold’s first version had won CASP13 in 2018 by a clear but ordinary margin. Nobody in the field expected a discontinuity.

They measured

A median GDT_TS of 92.4 across all targets, on a scale where a score around 90 is regarded as equivalent to an experimentally determined structure. In the hardest free-modelling category AlphaFold2 scored about 87, roughly 25 points clear of the next best group — the largest margin in the history of the assessment. Median backbone accuracy was under an ångström, which is about the width of an atom.

How sure could they be? GDT_TS is the percentage of residues falling within a set of distance cutoffs of the true position after optimal superposition, so it is a direct comparison against a measured structure rather than against another prediction. Roughly 90 is where the disagreement between prediction and experiment becomes comparable with the disagreement between two experimental determinations of the same protein.

Why it mattered

John Moult had been running this assessment since 1994 specifically to find out whether the problem was tractable, and said afterwards that in a real sense it had been solved. That judgement is from the person with the least incentive to overstate it. The 2021 paper released the method, and the resulting database now holds predicted structures for essentially every protein sequence that has been catalogued — over 200 million of them, against roughly 200,000 solved experimentally in sixty years of crystallography.

You might think

Protein folding has been solved.

Actually

Protein structure prediction has largely been solved, for a well-defined class of proteins, and that is a different sentence. The two get conflated because "the protein folding problem" was used loosely for both. What AlphaFold2 does is map sequence to final coordinates; it does not compute how the chain gets there, in what order, through what intermediates, or how fast — and it is not a physics simulation of anything, so it could not. The mechanism of folding, the reason Levinthal’s combinatorial disaster does not happen, misfolding and aggregation in amyloid disease, and the conformational changes that make proteins work are all open, and several of them matter more medically than the static structure does. The right summary is the narrow one: an unsolved problem in the field’s own benchmark was solved, at the level of that benchmark, and the benchmark was never the whole question.

  1. 1913The Oppau plant starts up on promoted iron, found by Mittasch after some 2,500 formulations were tried.
  2. 1929Dirac: the laws are completely known, and the equations are too complicated to solve.
  3. 1961Anfinsen refolds ribonuclease in a test tube: the sequence contains the structure.
  4. 1964Hohenberg and Kohn prove that an exact density functional exists. Nobody has found it since.
  5. 1969Levinthal: a chain cannot possibly search its conformations, and folds in milliseconds anyway.
  6. 1994Moult and Fidelis start CASP, a blind test nobody can train on.
  7. 1995Wolfenden measures an enzyme’s rate enhancement at 10¹⁷ — an uncatalysed half-life of 78 million years.
  8. 1998Kohn and Pople share the Nobel prize for computational chemistry that works despite an unknown functional.
  9. 2017Perdew and colleagues report that newer functionals fit energies better and densities worse.
  10. 2020CASP14: AlphaFold2 reaches a median GDT_TS of 92.4, and the assessment’s organiser calls the problem solved.
  11. 2024Baker, Hassabis and Jumper share the Nobel Prize in Chemistry — for design and for prediction.
  12. nowCatalysts are still found by making a great many of them and testing them.

One last thing, and it is the point of the whole path. Lavoisier had a balance and a sealed glass vessel. What he contributed was not oxygen — he did not discover it, and both men who did rejected what he made of it. It was a rule about how to ask: that the boundary of an experiment includes everything, that nothing may be left out of the sum, and that the answer is whatever the measurement says even when it contradicts a theory that has explained five other things beautifully. Every lesson here has been an instance of that rule. Two oxides of carbon in a ratio of exactly two. A gap in a table filled by an element with the density that had been written down for it years before anyone found it. X-ray frequencies giving every element an integer, and the integers correcting the table. Three peaks in the ratio 3:2:1. A single molecule photographed with its bonds visible. Those measurements will not be revised. The theories built on them will be, and this last act more than the rest. Whatever replaces the chemistry in this course will still have to account for every one of those numbers — that is what it means for a science to be cumulative, and it is why how do we know this? is the only question that has ever really mattered here. The questions that remain are not the ones Lavoisier could not answer. They are the ones his method opened.