Skip to content

Correlated Evidence, Independence, and Triangulation

Apparently different nonperturbative methods are independent only to the extent that they have distinct failure opportunities. Shared actions, renormalization conditions, scale setting, ensembles, ansätze, matching coefficients, calibration observables, perturbative inputs, or analytic-continuation assumptions correlate their outputs. Validation should therefore combine methods through their dependency graph and covariance, count shared inputs once, and seek held-out checks that can distinguish the branches—not multiply agreement as though every result were a fresh experiment.

Required background. Claim status across methods supplies proposition-level evidence labels, and observable translation ensures that the outputs being compared are genuinely commensurate.

Helpful background. Functional-method validation provides concrete truncation, branch, regulator, and benchmark axes.

Shared framework. The evidence-triangulation graph is the canonical dependency picture. The claim-and-control comparison, confinement claim-and-evidence comparison, functional-method comparison, and exact and rigorous status comparison expose different shared inputs and failure tests.

Independence means distinct failure opportunities

Section titled “Independence means distinct failure opportunities”

Method names do not determine independence. Two calculations are correlated if the same latent error can move both conclusions in the same direction. That latent source may be visible—a common scale-setting observable—or structural, such as a shared vertex ansatz, operator basis, symmetry restriction, or perturbative matching coefficient.

For each result, list its ancestors in five groups:

  1. Theory data: action, global form, state, boundary conditions, and parameters.
  2. Method structure: regulator, truncation, basis, contour, closure, or exactness hypotheses.
  3. External inputs: calibration data, phenomenological parameters, ensembles, priors, and matching coefficients.
  4. Translation: renormalization, analytic continuation, scale setting, and continuum or volume limits.
  5. Validation: identities, benchmarks, alternative observables, and theorem consequences.

An edge means that changing the parent can change the child. Common ancestors merge before the evidence nodes. A result used to tune a model cannot then serve unchanged as a held-out validation of that model, and a Ward identity imposed in two closures is a common consistency condition rather than two independent confirmations.

The relevant question is counterfactual: Could this check fail while the other methods still pass? If not, it supplies little discrimination even when it is valuable for internal consistency.

Correlation changes the combined uncertainty

Section titled “Correlation changes the combined uncertainty”

The simplest calculation already shows why agreement cannot be counted naively. Let two unbiased estimates x1x_1 and x2x_2 of the same quantity have equal variance σ2\sigma^2 and correlation coefficient ρ\rho. Their average has

Var ⁣(x1+x22)=σ22(1+ρ).\operatorname{Var}\!\left(\frac{x_1+x_2}{2}\right) =\frac{\sigma^2}{2}(1+\rho).

For ρ=0\rho=0, averaging halves the variance. For ρ=1\rho=1, it produces no reduction at all. The corresponding effective number of independent equal-precision estimates is

Neff=21+ρ,N_{\rm eff}=\frac{2}{1+\rho},

which ranges from two at zero correlation to one at complete positive correlation. Negative correlation can improve a particular combination, but only when it is established rather than assumed.

More generally, write a vector of method outputs as

x=θ1+Au+ϵ,\mathbf x =\theta\mathbf 1+A\mathbf u+\boldsymbol\epsilon,

where u\mathbf u contains shared uncertain inputs, AA records each method’s sensitivity to them, and ϵ\boldsymbol\epsilon contains method-specific errors. Then

C=AΣuAT+ΣϵC =A\,\Sigma_u A^{\mathsf T}+\Sigma_\epsilon

is the covariance that belongs in a combination. For known positive-definite CC, the minimum-variance linear estimator is

θ^=1TC1x1TC11,Var(θ^)=11TC11.\widehat\theta =\frac{\mathbf 1^{\mathsf T}C^{-1}\mathbf x} {\mathbf 1^{\mathsf T}C^{-1}\mathbf 1}, \qquad \operatorname{Var}(\widehat\theta) =\frac1{\mathbf 1^{\mathsf T}C^{-1}\mathbf 1}.

These formulas do not solve unknown model discrepancy. If an ansatz omits the same tensor structure in two methods, assigning it a zero-mean Gaussian variance without evidence can merely hide the problem inside CC. Vary the shared structure, add an adversarial model, or report a qualified envelope. Lyons, Gibaut, and Clifford discuss the nonintuitive behavior that can arise even when combining correlated estimates with a stated covariance Lyons, Gibaut, and Clifford 1988, pp. 110–117.

Suppose Dyson–Schwinger equations (DSE), functional renormalization group (FRG), and a regulated lattice calculation report a scalar-channel mass. Before combining them, translate all three to the same target: for example, the lowest stable infinite-volume pole in a specified renormalized theory.

The dependencies might be:

  • DSE and FRG start from the same continuum action and renormalization conditions;
  • both use a closely related dressed-vertex ansatz constrained by the same perturbative ultraviolet behavior;
  • their numerical solvers and hierarchy closures differ, but a missing tensor structure can bias both;
  • the lattice calculation uses independent gauge or scalar configurations and a different regulator, yet shares the physical scale-setting input;
  • all three identify the mass through Euclidean information, so a common spectral or pole interpretation may remain; and
  • one lattice observable was used to tune the functional ansatz.

Agreement among the tuned DSE result, FRG result using the same vertex family, and the calibration observable is then weaker than three independent determinations. A better design reserves a second channel, momentum range, or parameter point not used in tuning. It also varies the shared vertex basis and asks whether DSE and FRG move together while the regulated result does not.

One possible dependency table is:

Input or checkDSEFRGRegulated computationInterpretation
Continuum action and global dataUsesUsesMatched targetNecessary shared theory input
Renormalization and scale settingUsesUsesSupplies or sharesCorrelated dimensionful uncertainty
Vertex or operator ansatzExplicit closureProjection may reuse itNot used in the same formCommon functional-method discrepancy
Configurations or raw samplesIndependent solverIndependent solverIndependent ensembleAlgorithmic independence, not full inferential independence
Imposed symmetry identityEnforcedEnforcedChecked in regulated formShared consistency condition
Held-out second channelPredictedPredictedMeasured independentlyGenuinely discriminating if not calibrated
Continuum and volume variationNo lattice limitNo lattice limitDirectly variedTests a failure mode unique to the regulator route
Closure or basis enlargementDirectly variedDirectly variedOperator basis varied separatelyDistinguishes truncation from extraction error

The first QFT application is complete only after the last two columns are acted on: identify a check that can fail all methods and one that distinguishes among them.

A strong validation program uses several benchmark types without pretending they are interchangeable.

Structural checks. Dimensions, symmetries, Ward identities, anomaly relations, unitarity, and normalization can reject an inconsistent calculation. If imposed during construction, they test implementation but provide limited evidence about omitted dynamics.

Solvable limits. Free theory, weak coupling, large NN, low dimension, or an integrable point can test normalization and asymptotic behavior. A method may still fail away from that limit, so vary the distance from it.

Synthetic recovery. Generate data from a known model, hide its parameters, and test whether the full inference procedure recovers them with calibrated uncertainty. This checks the workflow but can favor discrepancies resembling the generator.

Held-out observables. Calibrate on one channel or kinematic region and predict another. This supplies a new failure opportunity if the second observable was not used indirectly through priors, tuning, or scale setting.

Independent formulations. Change variables, regulator, contour, basis, or experimental input so that the dominant omitted effects differ. Mere code independence is insufficient when the mathematical approximation is shared.

Adversarial cases. Seek a regime where candidate assumptions make different predictions: a threshold, symmetry-breaking source, alternative volume, complex momentum, or parameter sign. A discriminating test is more informative than another point where every approximation is expected to agree.

Calibration and validation must remain separated. Kennedy and O’Hagan’s treatment of computer models makes model discrepancy an explicit part of inference rather than silently absorbing it into parameter fitting Kennedy and O’Hagan 2001, §§ 2–4, pp. 425–446. In QFT, the analogue is to distinguish parameter uncertainty from the error caused by a missing operator, saddle, branch, or continuum limit.

For any three-method claim, design at least these two tests.

Choose an exact relation or controlled limit that every method must satisfy after observable translation. Examples include a Ward identity with nontrivial momentum dependence, a spectral sum rule, a known weak-coupling coefficient, or a finite-volume shift tied to scattering data. Failure of any branch blocks the combined claim.

This test checks correctness but may remain correlated. If every method enforces the relation by construction, success should be labeled consistency, not confirmation.

Choose an observable or regime on which the leading shared approximations differ. For the DSE–FRG–regulated example, enlarge the vertex basis and predict a held-out bound-state channel across several volumes. A closure-driven shift common to DSE and FRG but absent from the regulated extrapolation points toward shared truncation error; volume dependence isolated to the regulated branch points toward extraction or limit error.

The decision should be specified before inspecting the held-out result. Otherwise repeated choices of fit range, regulator, basis, and benchmark can turn validation into post hoc selection.

A concise synthesis states:

  • the common renormalized observable and agreed regime;
  • the dependency graph and which inputs are shared;
  • the covariance treatment for quantified shared errors;
  • method-specific systematic variations;
  • checks used in calibration versus checks held out;
  • the weakest unresolved edge; and
  • the bounded conclusion that survives removing any single correlated branch.

Do not compress this into a vote count. “Three methods agree” is less informative than “two closures sharing a vertex family and one independent regulator agree after matching; the held-out channel is reproduced, while real-time continuation remains unresolved.” The latter tells a reader what changed and what would falsify the conclusion.

Counting implementations instead of assumptions. Independent codes can implement the same truncation and inherit the same bias. Trace mathematical and data dependencies, not software ancestry alone.

Using calibration data twice. A datum used to tune an ansatz cannot be reused as an independent check without conditioning the interpretation on that tuning.

Forcing every discrepancy into a covariance matrix. Unknown model incompleteness may be directional, regime dependent, or discontinuous. Stress tests and qualified conclusions are safer than invented Gaussian errors.

Derive the variance of the average of two equal-variance correlated estimates and check the limits ρ=0\rho=0 and ρ=1\rho=1.

Solution

With Var(xi)=σ2\operatorname{Var}(x_i)=\sigma^2 and Cov(x1,x2)=ρσ2\operatorname{Cov}(x_1,x_2)=\rho\sigma^2,

Var ⁣(x1+x22)=14[operatornameVar(x1)+Var(x2)+2Cov(x1,x2)]=σ22(1+ρ).\begin{aligned} \operatorname{Var}\!\left(\frac{x_1+x_2}{2}\right) &=\frac14\left[operatorname{Var}(x_1)+\operatorname{Var}(x_2) +2\operatorname{Cov}(x_1,x_2)\right]\\ &=\frac{\sigma^2}{2}(1+\rho). \end{aligned}

It is σ2/2\sigma^2/2 for independent estimates and σ2\sigma^2 for perfectly correlated estimates.

A functional calculation is tuned to a lattice propagator at spacelike momenta and then predicts a bound-state mass. Name one shared check and one discriminating check.

Solution

A shared check is a symmetry identity linking the propagator, vertex, and bound-state kernel; all routes must satisfy it after matching, though it may be imposed in the functional construction. A discriminating check is a held-out bound-state channel or parameter point not used in tuning, evaluated with an enlarged vertex basis and an independently varied lattice volume and operator basis. The correlated scale-setting uncertainty remains common in either comparison.

  • Kennedy, Marc C., and Anthony O’Hagan. “Bayesian Calibration of Computer Models.” Journal of the Royal Statistical Society: Series B 63, no. 3 (2001): 425–464. DOI.
  • Lyons, Louis, Duncan Gibaut, and Peter Clifford. “How to Combine Correlated Estimates of a Single Physical Quantity.” Nuclear Instruments and Methods in Physics Research Section A 270, no. 1 (1988): 110–117. DOI.