Calibration, Uncertainty, and Cross-Model Inference
When several holographic models reproduce the same observables, calibration is not identification. A defensible comparison separates training data from held-out tests, propagates numerical and effective-theory errors, includes model discrepancy, and asks which predictions survive changes of priors and observable choice. The result may discriminate models, establish non-identifiability, or support a shared mechanism; it cannot silently promote a phenomenological fit to a microscopic duality.
Required background. Holographic Models of Quantum Critical Matter and Strange Metals supplies a case with degenerate scaling explanations. QCD-Like Holography and Phenomenological Limits supplies top-down and bottom-up model classes.
Helpful background. Claim–Evidence Records, Replication, and Retraction Handling supplies evidence-status distinctions. Evidence Triangulation and Reproducibility supplies independent-probe design.
Evidence cutoff: 25 July 2026. Method and phenomenology references below are bounded by this date.
A generative comparison
Section titled “A generative comparison”For model class with parameters , write the prediction for common, dimensionless observables as . A Gaussian likelihood has
with
represents structured model discrepancy: missing operators, incorrect microscopic content, or an inadequate crossover form. It is conceptually different from numerical error. Treating it as an unconstrained independent error on every point can make any model fit; setting it to zero assumes the model class is exact.
The posterior is
Priors must respect physical constraints such as positive kinetic terms, regular horizons, causal propagation, and the regime of the derivative expansion. Calibration methodology with explicit model discrepancy was developed by Kennedy and O’Hagan 2001.
A held-out three-observable application
Section titled “A held-out three-observable application”Consider two two-parameter models with normalized outputs
Calibrate both to and with negligible training error. Each gives and therefore fits perfectly. Their held-out predictions are
If the independent measurement is , the standardized residuals are and . The held-out datum discriminates the models even though their calibration scores are identical. If instead , it does not: the honest result is non-identifiability at the stated precision.
For a holographic comparison, the three outputs might be equilibrium pressure, dc conductivity, and a frequency-resolved spectral moment. They must use identical units, ensembles, current normalizations, Kubo limit order, and covariance. The third observable should probe a bulk sector not already fixed by the first two.
Uncertainty that belongs in the answer
Section titled “Uncertainty that belongs in the answer”For a posterior prediction report the decomposition, or samples sufficient to reconstruct it,
Parameter uncertainty alone is too small whenever alternative bulk potentials or omitted derivative terms give comparable fits. Numerical convergence must vary radial resolution, fitting window, and horizon cutoff. Truncation uncertainty must scale with the first omitted , inverse-coupling, or derivative order rather than with solver tolerance.
Bayesian holographic QCD studies now propagate parameter posteriors into multiple transport predictions Chen et al. 2025. Their credible intervals are conditional on the chosen action, calibration inputs, priors, emulator, and discrepancy treatment; they are not probabilities that a gravity dual of real QCD is true.
Leave-one-out and prior adversaries
Section titled “Leave-one-out and prior adversaries”Three controls are especially revealing.
- Leave one observable out. Refit and predict the omitted datum. A large shift reveals that the original “prediction” was calibration-dominated.
- Change physically admissible priors. If the model ranking changes, report prior sensitivity rather than a decisive preference.
- Enlarge the model class. Add the leading allowed coupling or an alternative potential. If its posterior is broad but its predictions spread, structural uncertainty dominates.
Also compare posterior predictive residuals across temperature, density, and frequency rather than compressing them into one score. Correlated residuals indicate a missing mechanism even when the total is acceptable.
The strongest warranted outcome is conditional: among the specified models, data, priors, and error model, one predicts independent observables better over the declared range. A stable shared scaling relation can support a universality class. Neither conclusion identifies a real material or plasma without the missing microscopic and experimental evidence.
Exercises
Section titled “Exercises”In the toy application, suppose the held-out uncertainty contains independent data error and numerical error . What is the combined standard deviation, and what are the two residuals?
Solution
The standard deviation is . For , the residuals are and .
The chapter overview contains the structure diagram and validity and failure diagram. They are embedded there once so that their shared chapter-level context is not repeated on every article.
For the chapter-wide comparison of assumptions, counterevidence, falsifiers, and claim ceilings, see the claim-domain table.
References
Section titled “References”- Chen, Bing, Liqiang Zhu, Xun Chen, Defu Hou, and Xurong Chen. “Transport Properties of QGP within a Bayesian Holographic QCD Model.” arXiv:2508.16167 [hep-ph] (2025). arXiv.
- Kennedy, Marc C., and Anthony O’Hagan. “Bayesian Calibration of Computer Models.” Journal of the Royal Statistical Society: Series B 63, 425–464 (2001). DOI.