Validation and Theory Uncertainties
Validation asks whether a perturbative prediction implements its stated mathematics and physics. Uncertainty analysis asks how incomplete information propagates after those checks have passed. These are different tasks: a bug is not a theory uncertainty, a Monte Carlo standard error is not a missing-order estimate, and a scale envelope is not automatically a confidence interval. A defensible result keeps analytic, numerical, perturbative, parametric, and modeling components separate until a justified correlation model combines them.
Required background. Fixed-Order Organization and Scale Dependence supplies perturbative order labels and scale diagnostics. Phase-Space Integration and Monte Carlo Estimators supplies sampling errors and convergence hypotheses.
Helpful background. Resummation and Fixed-Order Matching distinguishes fixed-order, profile, and matching-scheme variations.
Validation before uncertainty
Section titled “Validation before uncertainty”Let a prediction be represented as
where is the measurement definition, physical inputs, unphysical scales, a collection of schemes and approximations, and the numerical sample. Validation holds all of these fixed and tests identities that the implementation must satisfy: pole cancellation, Ward identities, factorization limits, normalization, map invertibility, or agreement with a trusted benchmark.
Only after those identities pass should variations of , , , or be interpreted. Inflating an error band does not cure a failed identity.
Validation matrix
Section titled “Validation matrix”The table is a reusable minimum set for fixed-order, loop, subtraction, matching, and resonance predictions. Not every row applies to every process, but every omitted row should be marked inapplicable with a reason rather than silently skipped.
| Check class | Concrete procedure | Pass condition | What it does not establish |
|---|---|---|---|
| Analytic poles and dimensions | Cancel every regulator pole; verify mass dimension, symmetry factors, and crossing conventions. | Coefficients vanish symbolically or to a declared precision before phase-space integration. | Finite constants and numerical mappings can still be wrong. |
| Gauge or BRST identities | Replace external polarizations by momenta and vary gauge parameters where available. | Physical results obey the stated Ward or Slavnov–Taylor identity and are gauge-parameter independent. | A gauge-invariant but incorrectly normalized answer can pass. |
| Infrared limits and subtraction | Approach every soft, collinear, and overlap limit along independent trajectories. | The real-to-counterterm ratio tends to one and mapped measurements approach the lower multiplicity. | Integrated finite remainders and global coverage need separate tests. |
| Numerical convergence | Use independent batches, alternative mappings, tail diagnostics, and analytic benchmark integrands. | Batch estimates agree statistically and errors scale as predicted under the estimator's hypotheses. | Convergence does not exclude a common bias in the integrand. |
| Scale and matching behavior | Repeat declared scale, profile, and matching variations and expand the matched result. | Known fixed-order coefficients are reproduced and residual scale dependence starts beyond claimed order. | The variation band has no universal confidence level. |
| Independent benchmark | Compare selected phase-space points, integrated limits, or coefficients with a separate implementation or analytic result. | Agreement falls within a predeclared tolerance using matched inputs and conventions. | Shared code, formulas, or hidden conventions reduce independence. |
| Physical limits and symmetries | Take threshold, high-energy, soft, collinear, decoupling, and crossing limits that apply. | The prediction approaches the independently known power, sign, and normalization. | Passing a few limits does not prove correctness in the full domain. |
| Inputs and model choices | Propagate parameter covariance and compare explicitly defined approximation or scheme alternatives. | Sources, correlations, domains, and combination rules are published separately. | The spread of chosen models is not guaranteed to cover all missing physics. |
The matrix is intentionally redundant. Dimensional analysis, a Ward identity, a local subtraction test, and a benchmark fail for different reasons; agreement among them is stronger than repeating one test at more points. In particular, the distinction between pointwise real-to-counterterm limits and cancellation of integrated poles is explicit in Catani and Seymour 1997, § 2.1, pp. 297–298; §§ 7.1–7.2, pp. 343–346.
Numerical uncertainty
Section titled “Numerical uncertainty”For independent importance-sampled weights , the standard error is
This estimate assumes finite variance and the declared sampling structure. Correlated bins require the covariance matrix of the estimated bin means,
Publishing only per-bin errors discards correlations needed for normalized shapes, integrals, and fits. If several separately sampled contributions cancel, their covariances vanish only when the samples are actually independent; common random numbers can produce useful nonzero covariance and must be retained.
Numerical error is reduced by more samples or better mappings. A stable estimate with a biased phase-space map, missing channel, or wrong counterterm is precisely estimated but wrong.
Parametric uncertainty
Section titled “Parametric uncertainty”Let the physical inputs have covariance . Linear propagation gives
Finite differences must use steps large compared with integration noise and small enough to remain in the linear regime. Nonlinear propagation can use replicas or direct sampling from a documented joint input distribution. Correlated parameters must not be varied one at a time and then added in quadrature unless that procedure reproduces their covariance.
An input fitted using the same data as the target observable creates additional correlations and possible circularity. That inference problem lies beyond this theory-level page, but the dependence should be identified.
A two-bin covariance calculation
Section titled “A two-bin covariance calculation”Take a two-bin prediction
Suppose an independently calibrated dimensionless parameter has standard deviation and local response
Linear propagation gives
Independence of the parameter calibration and the Monte Carlo sample permits addition of these two covariance matrices:
For the sum ,
Discarding the off-diagonal entries instead gives . The discrepancy is not an additional uncertainty; it is the result of using the wrong covariance model. A scale envelope remains a separately labeled perturbative diagnostic and should not be inserted into as if its variations were random draws.
Perturbative truncation diagnostics
Section titled “Perturbative truncation diagnostics”Residual renormalization and factorization scale dependence probes selected higher-order logarithms. For resummed predictions, profile-scale and matching-scheme variations probe additional structures. Keep these components labeled:
where is a declared combination rule, not an automatic quadrature sum. Validate the prescription on processes or orders where the next coefficient is known. Conventional factor-two scale variation has no process-independent confidence-level interpretation; empirical surveys quantify this limitation in Bagnaschi et al. 2015, § 3, pp. 8–12; § 5, p. 25.
Alternative probabilistic models for missing coefficients add explicit prior and coefficient-distribution assumptions. Their probabilities are conditional on those assumptions and require calibration; they should not be presented as consequences of perturbation theory alone.
Approximation and model dependence
Section titled “Approximation and model dependence”Examples include a narrow-width approximation, a chosen power-correction model, a factorization ansatz, or a treatment of non-global effects. A difference between two schemes that agree through the claimed order is a higher-order diagnostic. A difference between models is conditional on the models considered and need not bracket the truth.
For each approximation, state:
- the expansion parameter and claimed domain;
- terms kept and dropped;
- a lower-order comparison with a more complete calculation when available;
- how cuts or endpoints can enhance the omitted terms;
- whether the variation is correlated across bins and processes.
Do not hide these effects in a generic “systematic” category. Their physical origins determine how they should be propagated and correlated.
Reproducibility and independent checks
Section titled “Reproducibility and independent checks”A benchmark should freeze input values, scales, schemes, cuts, phase-space points, precision, and conventions. Compare complex amplitudes or coefficient components before squaring when possible; agreement only after integration can hide compensating errors. Independent code paths should avoid sharing the same algebraic simplifier, phase-space generator, or copied formula for the quantity under test.
For numerical results, retain source revisions, run configuration, random-seed policy, sample counts, compiler or precision choices when consequential, and machine-readable benchmark values. This provenance supports reproduction; it does not replace scientific checks.
Common pitfalls
Section titled “Common pitfalls”Treating a failed check as an uncertainty source. A residual regulator pole or Ward-identity violation is a defect. Fix it before interpreting variation bands.
Adding unlike errors in quadrature by reflex. Quadrature follows from a covariance model, not typography. Preserve source labels and correlations first.
Quoting excessive digits. Numerical precision should reflect the dominant relevant uncertainty and benchmark tolerance, not the raw floating-point output.
Using two calculations with shared ingredients as independent validation. Identify common amplitudes, mappings, libraries, and conventions; test at a level where the implementations genuinely differ.
Exercises
Section titled “Exercises”For two bins with equal standard deviations and correlation coefficient , derive the uncertainties of their sum and difference .
Solution
With
linear propagation gives
Positive correlation increases the uncertainty of the sum and decreases that of the difference; negative correlation does the reverse. Setting can therefore understate either observable depending on the actual sign.
Where to continue
Section titled “Where to continue”- Fixed-Order Organization and Scale Dependence develops the scale derivative and order labels used in the matrix.
- Phase-Space Integration and Monte Carlo Estimators develops sampling variance and mapping tests.
- Unstable-Particle Observables and Controlled Resonance Approximations applies gauge, pole, interference, and approximation checks near resonances.
References
Section titled “References”- Bagnaschi, Emanuele, Matteo Cacciari, Alberto Guffanti, and Laura Jenniches. “An Extensive Survey of the Estimation of Uncertainties from Missing Higher Orders in Perturbative Calculations.” Journal of High Energy Physics 02 (2015): 133. doi:10.1007/JHEP02(2015)133. Open PDF.
- Catani, Stefano, and Michael H. Seymour. “A General Algorithm for Calculating Jet Cross Sections in NLO QCD.” Nuclear Physics B 485 (1997): 291–419; erratum 510 (1998): 503–504. doi:10.1016/S0550-3213(96)00589-5. Open PDF.