A reconstruction that fails does not announce itself. It produces a full grain of atoms, a clean-looking density, and a respectable-looking recall score. This page is about how to catch it anyway — and what the 513 scored runs in this archive say about how often it happens and why.

1.What convergence looks like

Three things happen at once when a reconstruction succeeds, and the simultaneity is the point:

  1. The R-factor reaches the noise floor. The model's disagreement with the measurement drops until it equals the disagreement the true structure would have, which is set by photon shot noise. There is nothing structural left to explain.
  2. The atom count locks. The count stops changing — not because it hit a limit, but because adding or removing any atom makes the fit worse.
  3. The add/remove churn goes to zero. The solver runs out of moves it wants to make, and the run ends itself early rather than exhausting its iteration budget.

A failed run has none of these. Its R-factor flattens above the floor, its count drifts or stalls short, and it churns until it hits the cap. The best single number is the ratio of the two fits — χ²/floor — and it separates the archive cleanly, spanning 1.08 for the best run to 10,370 for the worst.

The archive's best and worst runs on identical panels: measured peak, density, atom overlay, count curve, R-factor curve, and a numeric scorecard for each.
Fig 1 The two extremes of the archive on identical axes. Top: the best-fitting run, χ²/floor 1.08, all 6,136 atoms found. Bottom: 25kD_s42, χ²/floor 10,370, which produced 34,125 atoms for a 24,800-atom truth. Read left to right and notice how late the difference becomes obvious — the measured peak and the density look comparable; it is the atom overlay and then the R-factor curve that give it away.

2.The archive

513 reconstructions with a complete figure suite and a full score. Converged means χ²/floor < 5; failed means > 50; the band between is real and populated, not a rounding artefact.

All 513 scored reconstructions, by grain size.
Grain sizeRunsConverged PartialFailedConverged
~5,000 atoms207146303170.5 %
~9,300 atoms1230925.0 %
~22,000–26,000 atoms21214531456.6 %
~50,000 atoms8207930.0 %
All51316316218831.8 %

Success falls off with grain size, and that is expected: the number of unknowns grows as three coordinates per atom while the number of usefully-measured intensity voxels grows much more slowly. Every run here used three Bragg peaks, which is the minimum that can constrain a 3D displacement field at all. Larger grains want more peaks, not more iterations.

The 50,000-atom row is work in progress, not a ceiling

79 of those 82 runs sit in the intermediate band — fitting better than a failure and worse than the floor. They have not stalled at a wrong answer; they have not finished. The three genuine failures there are a much smaller fraction than at 25,000 atoms.

There is almost no partial credit

One property of the archive is worth stating on its own, because it shapes everything that follows. Reconstruction accuracy is not spread smoothly from good to bad — it is bimodal. Bin all 513 runs by their final mean position error and the middle is nearly empty:

Final mean position error over all 513 scored runs. A tungsten bond is 2.7411 Å.
Position errorRunsShareInterpretation
below 0.05 Å21441.7 %atomic resolution — the lattice is resolved
0.05 – 0.5 Å265.1 %the empty middle
0.5 Å and above27353.2 %off by a substantial fraction of a bond

In the tightly controlled 144-run 5,000-atom library the gap is starker still: 129 runs land below 0.05 Å, 14 land at or above 0.5 Å, and exactly one falls in between.

Why an empty middle is a clue, not a curiosity

Continuous error sources — photon noise, an imperfect potential, insufficient iterations — produce continuous degradation. A gap means the outcome is being decided by something discrete: the model either latches onto the correct lattice registry or latches onto a wrong one, and there is no stable state part-way between two lattice sites. This is the single strongest piece of evidence for the mechanism in §4, and it is also good news, because a discrete failure is a failure you can detect and retry rather than an accuracy limit you have to live with.

Three panels: 144 R-factor trajectories against iteration coloured by cascade dose, the same runs' reduced chi-squared, and final count error plotted against final position error showing two separated clusters.
Fig 2 All 144 library runs at once. Left and middle: every trajectory, coloured by cascade energy. Two bundles are visible — one that descends to the floor and one that flattens high and stays there — along with a handful that break out late, past iteration 150. Right: final count error against final position error, with tungsten's bond length marked. The two clusters and the gap between them are the bimodality above, drawn.

3.It is not the radiation damage

The obvious hypothesis is that damaged grains are harder to reconstruct, since defects are exactly what the ideal starting lattice does not contain. The 144-run 5,000-atom library was built to test that, with a pristine control and four cascade energies on sixteen different grains. The result runs the other way.

The 5,000-atom library by damage level. L0 is the pristine control — a relaxed but undamaged grain. Energies are the primary knock-on atom energy of the molecular-dynamics cascade.
DamageRunsConvergedPartialFailedConverged
L0 — pristine, no cascade16102462.5 %
L1 — 150 eV32214765.6 %
L2 — 400 eV32254378.1 %
L3 — 1,000 eV32283187.5 %
L3deep — 1,000 eV, deep32239071.9 %

The pristine control is the worst row in the table, and the most violent cascade is the best. Not one of the 32 deep-cascade runs failed outright. This is a monotone trend over four of the five levels and it reverses the naive expectation completely.

Why damage helps

Because the failure mode is a registry problem, and defects break the symmetry that causes it. A perfect crystal looks the same after sliding it by one lattice site, so nothing in the data prefers one alignment over another and the solver can settle into the wrong one. Put a recognisable cascade in the middle of the grain and that degeneracy is gone — there is now exactly one alignment that explains the measurement.

The gallery's lib_5k_g03_L0 is this argument in a single run: an undamaged 4,828-atom grain, the easiest target in the archive, stalled at 1.48 Å with χ²/floor 1,478.

Grain shape matters far more than dose. Among the 22,000–26,000-atom runs, sorted by which grain they were reconstructing: grain g09 converged in 5 of 9 attempts (55.6 %), g04 in 5 of 19 (26.3 %), g01 in 4 of 48 (8.3 %), and g02 in 0 of 25. Same code, same damage recipes, same peak set.

4.Two distinct failure modes

Failures are not all alike. Fit one rigid-body transform — a single translation and rotation applied to the whole reconstruction — and ask whether that fixes it. The answer splits the failures cleanly in two.

Two scatter plots: recall before versus after a single rigid transform, and the size of the best rigid translation against chi-squared over floor, with converged, registry-failure and warp groups distinguished.
Fig 3 Every failure in the measured cache, classified by whether one rigid transform repairs it. The horizontal dashed line in the right panel is tungsten's Wigner–Seitz radius, 1.371 Å — the natural scale of one lattice site.

Registry failure — the crystal is right, its position is wrong

In this group the reconstruction is internally correct. Every bond length, every defect, the whole lattice — right. It is simply sitting in the wrong place, shifted bodily by roughly one lattice site. Apply one translation and it snaps into agreement.

Measured over this group: median recall at 0.5 Å goes from 0.0000 to 0.9551 under a single rigid transform. The median translation needed is 1.419 Å — essentially tungsten's Wigner–Seitz radius of 1.371 Å — and the median rotation is 0.0176°, i.e. none. It is a pure shift, not a misorientation.

Four panels: stuck-registry rate falling with cascade dose, a bimodal position-error distribution, the rotation of the repairing shift plotted against its magnitude showing it is a pure translation, and a text verdict.
Fig 4 Registry lock-in across the 144-run library. Left: the fraction of runs stuck above 1 Å falls with cascade energy — 25.0 % pristine, 18.8 % at 150 eV, 9.4 % at 400 eV, 1.6 % at 1 keV (Spearman −0.287, p = 0.0005). Middle: position error is bimodal, clustering at ~0.02 Å or ~1.5 Å with almost nothing between. Right: the repairing transform carries essentially no rotation, so the reconstruction is a single misplaced domain — a mosaic or multi-domain explanation was tested and rejected.
Where the shift comes from

The seed lattice is built by laying an ideal BCC grid over the support and keeping the sites that fall inside it, with the grid's origin taken from the support's centroid. Nothing in that construction ties the lattice's phase to the actual atoms — the support knows the grain's outline, not where its lattice planes sit. So the seed can start up to one Wigner–Seitz radius out of registry, and the refinement loop then has to walk the entire crystal into place using single-atom moves. Sometimes it does. Sometimes it cannot, because moving one atom at a time out of a globally-shifted lattice makes the fit worse before it makes it better.

Non-rigid warp — a distorted field

The other group is not fixable by any rigid transform. Its best single translation lifts median recall at 0.5 Å only from 0.0335 to 0.1939. The error is a smoothly varying field across the grain rather than a uniform offset: different regions have drifted different ways. 25kD_s42 in the gallery is an extreme member — and it is also the one that overpopulated the grain by 9,325 atoms, which is what a warped model needs in order to keep explaining the measurement.

In the classification drawn on Fig. 2 the split is 18 registry failures against 30 warps; the medians quoted above were recomputed independently from the measurement cache and reproduce the figure's values (1.419 against 1.43 Å translation, 0.0176 against 0.020° rotation). The exact group sizes shift by a couple of runs depending on where the “repaired” threshold is drawn; the two-mode structure does not.

5.Why a stuck run cannot stop itself

A converged run ends early: it stops proposing changes, and the loop notices and exits. The rule is deliberately strict — a long unbroken run of iterations with no atom added and none removed. A run in a limit cycle never gets there. It keeps making changes, each one undone a few iterations later, and burns its entire iteration budget without moving.

Panels showing the distribution of consecutive zero-change iteration streaks, the rate at which runs hit the iteration cap by damage level, and a site-occupancy raster revealing a repeating add-remove cycle.
Fig 5 The early-stop rule and what defeats it. The occupancy raster on this figure tracks individual lattice sites across iterations and shows a genuine limit cycle: a small population of sites being filled and emptied over and over. The cap-hit rate falls with damage level, mirroring the table in §3.

This is why the poster says a failed run “churns indefinitely”. It is not a figure of speech and it is not slow convergence — it is a cycle, and waiting longer does not help. The archive contains runs given 3,600 iterations that ended no better than siblings given 200.

6.The metric that hides all of this

The scoring code's default matching tolerance is half a lattice parameter, read from each run's own metadata. In this campaign that metadata field carried aluminium's lattice parameter of 4.046 Å, giving a default tolerance of 2.023 Å — while tungsten's first-neighbour bond is 2.7411 Å.

The tolerance is therefore 74 % of a bond: wider than the registry offset it is supposed to detect. An atom displaced onto a neighbouring lattice site scores as correctly found. The failure mode of §4 is invisible to the metric by construction.

Recall at the honest 0.5 angstrom tolerance plotted against recall at the loose 2.023 angstrom tolerance, showing failed runs clustered at high loose recall and near-zero tight recall, plus a histogram of the gap.
Fig 6 Every point below the diagonal is a run flattered by the loose tolerance. The cluster at the bottom right — loose recall near 1.0, honest recall near 0 — is the failure population.
Recall at the two tolerances, median over 195 runs with a cached re-measurement. Both columns use exclusive one-to-one matching and are twin-resolved.
GroupRunsRecall @2.023 ÅRecall @0.5 ÅGap
Converged1211.00001.00000.0000
Failed500.96920.00130.9678
All 50 failed runs score above 0.90 on the loose ruler

Not most of them — all of them. Meanwhile 30 of those 50 score below 0.01 at an honest half-ångström. Quoted alone, the loose recall cannot distinguish a single failed reconstruction in this archive from a converged one.

This is why every figure on this site prints both numbers, and why χ²/floor rather than recall is the headline metric. Note what the loose ruler did not corrupt: it is a scoring default, not something the solver optimises, so no reconstruction was ever trained toward it. The conclusions change; the runs do not.

7.Picking the good run without knowing the answer

Everything above is diagnosis. The practical question is what to do at a real beamline, where there is no truth to score against and no way to know whether the reconstruction in front of you is one of the good ones.

The answer is the one on the poster: run several reconstructions and keep the best-fitting. What makes it work is that χ²/floor requires no truth — the noise floor follows from photon statistics, which are a property of the measurement.

Chi-squared over floor plotted against true position error across the scored library, showing that the truth-free metric separates good runs from bad with an overlap region drawn explicitly.
Fig 7 The truth-free selector against the truth. χ²/floor is computable from the measurement alone; position error is not. The decision boundary and the region where the two populations overlap are drawn explicitly rather than hidden — the separation is very good, not perfect.
Measured: 99.3 % of an oracle's benefit, for free

Compare three strategies across the archive's seed pairs. Pick blind and you get mean recall 0.9065 with 8.6 % of runs broken. Pick the lower χ²/floor of the pair — no truth used — and you get 0.9548 with 4.7 % broken. Let an all-knowing oracle pick the genuinely better run and you get 0.9552 with 4.7 % broken.

The truth-free rule captures 99.3 % of what perfect knowledge would buy you, and roughly halves the broken-run rate. That is the whole justification for launching reconstructions in parallel rather than serially refining one.

The main page's eight-repeat table is this in miniature: eight independent repeats on one grain, χ²/floor of 1.15, 1.16 and 1.53 against 151, 275, 938, 1,046 and 1,204. Keep the lowest and you get a reconstruction accurate to 0.0017 Å without ever having seen the answer.

8.What would actually fix it

Three things follow from the diagnosis above, in rough order of expected value:

Where to go next

Back to the gallery →

Eight reconstructions in detail — four that reached the noise floor and four that did not, with every diagnostic figure.

How to read these figures →

Panel by panel through the diagnostic sheets, including which single panel decides whether a run worked.

How BCDI works →

The phase problem, why three Bragg peaks, the twin ambiguity, and what the noise floor actually is.