Uncertainty, disagreement, and negative results
Two results disagree only after they refer to the same object under compatible conditions. Before interpreting a tension, align the observable, kinematics, conventions, approximation order, inputs, and inferential question. Then keep shared uncertainties and genuinely independent checks visible.
This guide ends with a comparison another researcher can inspect: what agrees, what differs, how uncertainties are related, which conclusion is supported, and which new result would change the comparison.
Helpful background. Bring the claim sheet from Scope, conventions, and status and the protocol from Reproduce and validate a result. If covariance, effective sample size, or conditional inference is unfamiliar, use the statistics preparation check first.
Match the objects before comparing the numbers
Section titled “Match the objects before comparing the numbers”Write one row for each result:
| Field | Result A | Result B |
|---|---|---|
| Quantity and units | ||
| External state, geometry, or dataset | ||
| Kinematic or parameter domain | ||
| Conventions and scheme | ||
| Approximation or truncation order | ||
| Shared inputs | ||
| Method-specific inputs | ||
| Central result | ||
| Uncertainty components and correlations | ||
| Strongest supported conclusion |
Do not translate the central values without translating their uncertainty and scope. A coefficient quoted at two renormalization scales, a Euclidean correlator and a real-time response, or a partonic and a fiducial cross section are different objects until the required evolution or forward map is supplied.
Three lists make the comparison legible:
- Common ground: identities, inputs, limits, or qualitative behavior both calculations share.
- True differences: assumptions, data, approximation order, algorithm, prior, observable, or physical interpretation that remains different after translation.
- Apparent differences: notation, units, basis, parameterization, or scale choices that disappear under a checked transformation.
Only a true difference about a common target is a scientific disagreement.
Running case: two truncations are not a contradiction
Section titled “Running case: two truncations are not a contradiction”Continue the heavy-exchange example from the reproduction step. Normalize one exchange channel as
The local expansion through order is
Relative to the exact result,
Suppose one group uses leading order, , while another uses next order, . At , their values are and , whereas the exact value is . Their absolute difference is , or ten percent relative to , but they do not make contradictory claims: their relative truncation errors are ten percent and one percent, respectively.
The correct comparison is
| Approximation | Value at | Relative truncation error | Supported statement |
|---|---|---|---|
| Leading contact term captures the scale and sign, not percent precision. | |||
| The first momentum correction reaches percent accuracy at this point. | |||
| — | Reference value for this one-channel toy calculation. |
If both groups instead claim to have computed at the same and obtain different values after convention translation, there is a genuine calculation-level disagreement. The one-channel series illustrates controlled truncation; it does not establish the full decoupling theorem, loop matching, crossing relations, or the behavior of nondecoupling couplings.
Decompose uncertainty where it enters
Section titled “Decompose uncertainty where it enters”Do not begin with a single error bar. Record each contribution at the stage where it arises:
- definition: ambiguity in the observable, event class, operator, or phase criterion;
- input: calibration, external parameters, boundary data, ensemble, or dataset selection;
- theory: omitted perturbative orders, EFT powers, nonperturbative matrix elements, or model discrepancy;
- numerical: discretization, finite volume, sampling, solver tolerance, or unstable inverse problems;
- translation: matching, basis conversion, scale evolution, units, or detector response; and
- inference: likelihood, priors, nuisance parameters, regularization, selection, and model comparison.
For every contribution, state how it was estimated, whether it is signed or bounded, what it correlates with, and which check supports it. The JCGM 100:2008, §§4–5 gives a general framework for propagating input uncertainties and covariance for a well-defined quantity. An EFT truncation estimate or model discrepancy still requires subject-specific reasoning; it should not be relabeled as sampling variation.
In the running case, the first omitted geometric-series term gives a useful truncation scale because the remainder is known exactly. In a real EFT, unknown coefficients, logarithms, thresholds, and multiple expansion parameters can weaken that estimate. Order variation is evidence only when the power counting and observed convergence support it.
Shared inputs change the uncertainty of a difference
Section titled “Shared inputs change the uncertainty of a difference”Let estimates and have standard uncertainties and correlation . The variance of their difference is
Positive correlation can make a difference more precise because a common shift cancels. It can also make agreement less independent: two pipelines using the same data, calibration, perturbative coefficients, code library, or fit model may share the dominant failure.
Record dependencies explicitly. Agreement between correlated implementations is a valuable verification, but it is not the same as confirmation by a method with different failure modes. The distinction between computational reproducibility and broader scientific replication is developed in the National Academies 2019 report, chs. 3–4.
Use the weakest accurate relation
Section titled “Use the weakest accurate relation”After matching objects and propagating shared uncertainty, classify the relation conservatively:
- compatible: both claims can hold in their recorded scopes;
- tension: their preferred values or regions pull apart, but the evidence does not justify contradiction;
- contradiction: they make incompatible statements about the same target under the same assumptions;
- underdetermined: available evidence does not discriminate among the relevant explanations;
- non-comparable: the target objects or scopes remain different; or
- correction: one calculation contains an identified error and the revised result replaces it for the present comparison.
A quoted “number of sigma” does not choose the category on its own. The sampling model, nuisance treatment, selection procedure, look-elsewhere effects, parameter boundaries, and asymptotic regime are part of its meaning. Likelihood-based asymptotic formulae and their conditions are presented in Cowan et al. 2011, §§2–3.
Read a negative result within its sensitivity
Section titled “Read a negative result within its sensitivity”A negative result says what a test did not find in the region where it had power. Record:
- the target signal and operational signature;
- the parameter or function space tested;
- expected outcomes under the relevant alternatives;
- observed result and uncertainty;
- inference procedure and coverage or calibration checks;
- blind spots and failed assumptions; and
- the broader question left open.
For the heavy-scalar example, define the leading local operator by
so that its coefficient in this normalization is
An interval consistent with does not separately determine and , exclude heavy fields with other couplings, or test energies near the heavy pole. It constrains the specified operator combination within the measurement, matching, and EFT domain. A failed numerical solver is even narrower: it is a negative result about that method and setup, not evidence that the physical solution does not exist.
End with a discriminating test
Section titled “End with a discriminating test”A useful comparison points toward a result that could change it:
common target:decision-relevant difference:candidate explanations:observable or theorem that separates them:expected outcome under each explanation:required precision or domain:method with a different dominant failure:shared dependencies that remain:blind spots:proceed / revise / stop rule:In the running example, evaluating at several values separates a wrong normalization from a missing higher-order term. A normalization error remains roughly constant in ; an omitted first correction produces a residual proportional to at leading order. Near , neither pattern validates the local expansion because the scale hierarchy itself is failing.
Exercises
Section titled “Exercises”1. Translate two error statements
Section titled “1. Translate two error statements”One paper writes the leading relative EFT error as . Another writes with . Do the statements disagree?
Solution
No. Substituting gives
They use different variables for the same power suppression. A disagreement would remain only if their and referred to different kinematics or if one statement included a coefficient or logarithm the other omitted.
2. Include a shared uncertainty
Section titled “2. Include a shared uncertainty”Two estimates are and . Compute first for and then for .
Solution
For independent uncertainties,
For ,
so the standardized difference is about . The positive shared component cancels in . This calculation does not by itself establish a scientific tension; the Gaussian model, correlation estimate, selection procedure, and common target still need checking.
3. Interpret a null contact-interaction result
Section titled “3. Interpret a null contact-interaction result”A fit finds a contact coefficient compatible with zero within its stated interval. Which conclusions are supported?
Solution
The fit constrains that coefficient—or a stated combination of coefficients—under the chosen data, likelihood, operator basis, matching, truncation, and kinematic cuts. If the interval has validated coverage, one may report the corresponding allowed region.
It does not prove that no heavy particle exists, determine and separately when only enters, exclude interactions outside the chosen basis, or establish validity near . Those are the nearest stronger nonclaims.
References
Section titled “References”- Cowan, Glen, Kyle Cranmer, Eilam Gross, and Ofer Vitells. 2011. “Asymptotic Formulae for Likelihood-Based Tests of New Physics.” European Physical Journal C 71: 1554. DOI and open article.
- Joint Committee for Guides in Metrology. 2008. Evaluation of Measurement Data—Guide to the Expression of Uncertainty in Measurement. JCGM 100:2008. DOI. Official PDF.
- National Academies of Sciences, Engineering, and Medicine. 2019. Reproducibility and Replicability in Science. Washington, DC: National Academies Press. DOI and open book.
Carry the matched comparison, shared dependencies, and discriminating test into Scope a first research project.