Cookie Consent by Free Privacy Policy Generator Appendix B3 — How to Grade a Model Honestly | Igor Moiseev
Lab › Geometry of the Cosmic Web › Appendix B3

Appendix B3 — How to Grade a Model Honestly

Every evaluation failure the research program hit, told technically: the matched-length artifact and its sign flip, reference circularity, held-out seed discipline, jackknife versus pair bootstrap, physically-null controls, oracle bounds as information limits — and two byte-identical datasets caught by hashing.

By Igor Moiseev · 8 July 2026
Geometry of the Cosmic Web
  1. The Geometry of the Cosmic Web: A Research Program
  2. Two Ways to See a Cosmic Filament
  3. From Cosmic Filaments to Curved Spacetime
Appendices — Theory Background
  1. B1. How the Universe Moves Its Matter: Transport Models
  2. B2. The Tidal Frame: How Collapse Chooses Directions
  3. B3. How to Grade a Model Honestly ← you are here
  4. B4. The Transverse-Damping Model, in Full
  5. B5. Reading the Sky's Hot Gas
What this appendix covers
The research program's most transferable products are not cosmological. They are the evaluation failures it committed, caught, and repaired — each one a general lesson in how a plausible protocol manufactures a discovery. This appendix tells them technically: the matched-length artifact (a p ≈ 10⁻¹⁴ result with the wrong sign), reference circularity in skeleton scoring, seed discipline and competitor-favouring calibration, why pair bootstrap overstated a detection by a factor of 1.6, physically-null controls that exposed a 4.7σ instrumental ghost, oracle bounds that separate physics from estimation error, and a dataset-provenance check that caught two "independent" simulations being byte-identical. Numbers cite the experiment reports in research/cosmic-web/docs/.

Match on the quantity that buys the score

The program’s benchmark compared two filament detectors by completeness: the fraction of the true filament network lying within 2 voxels of the recovered skeleton. Completeness has a mechanical confounder — a longer skeleton covers more of the truth regardless of quality — so both methods were required to output skeletons of equal total length. The first implementation enforced that match on a proxy: the volume of the hysteresis mask from which each method’s skeleton is then extracted, on the assumption that realised skeleton lengths would follow.

They did not. At sparse sampling the orientation-lift’s realised skeletons came out about 19% longer than the Hessian baseline’s — 1,468 versus 1,239 voxels — and the resulting “+0.06 completeness, 48/50 seeds, p ≈ 10⁻¹⁴” advantage was pure length. Under the corrected extractor, which iterates until the realised skeleton itself hits the target length, the sign flips: the Hessian wins completeness at every sampling level from 2,500 galaxies up (Δ = −0.045 to −0.057, 50/50 seeds, p ≈ 10⁻¹⁵), with a statistical tie only at ultra-sparse sampling (SYNTHESIS; PAPER-DRAFT §5.1). The artifact was detected only because a later instrument change retroactively shifted archived scores, and the archived length fields confirmed the mismatch.

Two statistical debts an honest checklist also names. Multiplicity: the sky-side estimator went through a redesign sequence (E3→E3e) before the final pre-registered gate (E3c’s 3σ); results from the exploratory stages are reported as exploration, and only the gated, null-subtracted numbers are carried forward — but a reader should know the garden of forking paths was walked. Tiny held-out samples: transfer rows rest on n = 3 seeds, so the quoted ±’s are jackknife estimates with large variance-of-variance; treat them as scale indicators, not precision intervals.

The lesson generalises beyond skeletons: a matched comparison must enforce the match on the quantity that determines the score, not on a proxy upstream of it — and statistical strength (p ≈ 10⁻¹⁴) certifies nothing about whether the comparison was constructed correctly.

The same race, two ways to "match length" — experiment E0. Completeness advantage of the orientation lift over the Hessian (Δ = lift − Hessian; above zero the lift wins, below the Hessian wins), versus galaxy count. Both skeletons were pruned to equal length before scoring — but equal on what? Matched on a proxy (the hysteresis-mask volume), the lift's realised skeleton ran ~19% longer at sparse sampling, and length buys completeness: the lift appears to win by up to +0.06 (p ≈ 10⁻¹⁴). Flip the match to the realised skeleton length — the quantity that actually determines the score — and the sparse "win" inverts to a clean Hessian lead at every level. The phantom was in the matching, not the method. Proxy numbers from e0_results.json; corrected from the 50-seed reruns (E0 report and its E0b sweeps in research/cosmic-web/docs/).

Reference circularity

When no analytic truth exists, the temptation is to score each method’s sparse-data skeleton against a reference skeleton extracted from the clean, noise-free field. The program’s E1 ran the full 2×2 design — each method’s sparse skeleton against each method’s clean reference — and found each method scores 0.07–0.19 higher against the reference produced by its own family; the two clean references agree with each other only at completeness ≈ 0.60 (E1, 12 seeds). Any single-reference comparison is therefore circular: the choice of reference decides the winner before the data are consulted.

The repair is a method-neutral criterion that needs no reference at all: which skeleton’s spines capture more mass (total, and in the filament-density band) at matched length (E1b). Under that criterion the verdicts stabilised across field types and resolutions.

Held-out seeds and competitor-favouring calibration

Every tunable choice is an opportunity to overfit the benchmark. The transport-model campaign used a three-tier seed protocol: the crossing threshold \(\rho_c\) was calibrated on seed 1 by optimising the isotropic competitor, not the proposed model; configuration selection used seeds {2, 3}; every reported number comes from held-out seeds ({4, 5, 6}, or the five held-out seeds of E5), with jackknife-over-seed error bars (E5–E5d). Calibrating on the competitor is the cheap trick that makes a win conservative: any residual bias in the shared knob favours the opponent. When the frozen model then beats the ZA by −9% on seeds it never saw, the number means something.

Jackknife versus pair bootstrap

The observational campaign stacked Compton-y maps at 876,639 galaxy pairs. Bootstrap over pairs treats each pair as independent — but pairs share galaxies and overlap on the sky, so their map noise is correlated, and the bootstrap error is too small. The validity battery (E3e) replaced it with a jackknife over 50 right-ascension patches: delete one sky patch at a time, remeasure, and read the variance from the spread. The deflation was a factor 1.6–2.4 across instruments: the headline bridge significance on ACT dropped from 5.24σ (pair bootstrap) to 3.30σ (jackknife), and Planck’s from 19.3σ to 7.99σ. Rule: resample units that are actually exchangeable — sky patches, not overlapping pairs.

Physically-null controls

An estimator can pass every statistical test and still measure an instrumental artifact. The decisive check is a control sample in which the physics is absent by construction but every systematic is present. For the pair-bridge measurement: pairs with the same transverse separation window but line-of-sight separation 25–40 h⁻¹Mpc — same sky geometry, same halos, no possible physical bridge between them. On ACT the null pairs are consistent with zero (1.65σ). On Planck they show a 4.7σ “bridge” (1.80×10⁻⁸) — direct, quantitative confirmation of second-halo beam leakage at Planck’s 10′ resolution (E3e; Appendix B5 explains the mechanism). Without the null, that leakage would have been booked as astrophysics. The physically meaningful number is the null-subtracted one: ~1.2–1.4×10⁻⁸ at ≈2σ per instrument (2.4σ Planck, 1.6σ ACT).

The same discipline applied on the sky in an earlier form: the spine-stack controls were rejection-matched to the spine points’ galactic-latitude histogram, so a latitude-dependent foreground could not inflate the detection (E3b) — and the stricter control strengthened the result, from 4.8σ to 9.0σ. Stricter controls do not always shrink signals; they shrink lies.

Oracle bounds as information limits

When a model underperforms, is the physics wrong or the estimation noisy? The program answered with an oracle: rerun the transverse-damping model with tidal frames computed from the true final N-body field instead of the model’s own evolving estimate (E5b). With oracle frames the model beats the Zel’dovich baseline in every distance bin (−20% at the web; 4.47 versus 5.00 overall), so the damping physics is uniformly correct, and the practical model’s small outskirts deficit was frame estimation, not wrong physics. The oracle also prices the remaining headroom — 0.46 voxels of recoverable error, more than any other refinement lever — and grades the final model: the frozen recipe’s 4.52 closes 89% of the gap to the 4.47 bound (E5d). An oracle bound converts “could do better” into a number.

Trust the bytes: the CAMELS twins

The referee battery required validating the program’s own N-body truth against two independent production codes. The second dataset chosen, SIMBA_DM CV_0 from the public CAMELS suite, was checked before use with a range-request hash comparison — and found byte-identical to the first, IllustrisTNG_DM CV_0: the dark-matter-only “CV_0” runs of those two suites are one shared simulation on the server (E12). Astrid_DM CV_0 (run with MP-Gadget) was genuinely distinct and became the second code; the two real codes agree with each other to 0.029 h⁻¹Mpc median per particle, and the program’s particle-mesh truth sits 0.41 h⁻¹Mpc from that consensus — an order of magnitude below the 4–5 h⁻¹Mpc effects measured (E12). The lesson: “independent dataset” is a hypothesis, and it is testable for pennies — hash before you validate.

The checklist

Back to the series

Back to the series: The Geometry of the Cosmic Web: A Research Program · Two Ways to See a Cosmic Filament · From Cosmic Filaments to Curved Spacetime.

References

  1. F. Wilcoxon (1945). "Individual comparisons by ranking methods." Biometrics Bulletin 1, 80–83. The paired signed-rank test used throughout the benchmark.
  2. J. W. Tukey (1958). "Bias and confidence in not-quite large samples." Ann. Math. Statist. 29, 614; B. Efron (1979). "Bootstrap methods: another look at the jackknife." Ann. Statist. 7, 1–26.
  3. F. Villaescusa-Navarro et al. (2021). "The CAMELS project." ApJ 915, 71. Source of the external-validation simulations, including the byte-identical CV_0 twins.
  4. Experiment reports SYNTHESIS, E1, E1b, E5–E5d, E3b, E3e, E12 and the paper draft, in research/cosmic-web/docs/ of the repository.