Cross-Method Benchmarks and Adversarial Null Tests
Cross-method agreement is persuasive only to the extent that the methods have different failure channels. Correlation-matrix, replica, randomized, and tomography estimates can share state preparation, region maps, normalization, calibration, or continuum fits. A robust comparison therefore freezes a benchmark, maps shared dependencies, includes adversarial nulls, and reserves held-out observations before examining agreement.
Required background. Continuum Extrapolation, Bias, and Uncertainty supplies joint error propagation.
Helpful background. Entropy and Rényi Estimation Protocols, Randomized Measurements and Classical Shadows, Replica, Swap, and Multi-Copy Measurement Protocols, and Correlation-Matrix and Gaussian-State Reconstruction supply the four comparison methods.
Independence map
Section titled “Independence map”Build a table before analysis:
| Stage | Method A | Method B | Shared? | Targeted check |
|---|---|---|---|---|
| state preparation | batch and controls | copies or randomized runs | often | interleaved reference states |
| region/mode map | same boundary and basis | independently reconstructed map | maybe | deliberate offset scan |
| calibration | detector matrix | alternate detector or | often | blind gain injection |
| estimator | covariance model | swap/shadow statistic | no | known positive/negative states |
| normalization | common reference partition function | independent ratio | maybe | unit and product-state identities |
| continuum fit | same code and ansatz | independent implementation | often | frozen synthetic scaling family |
Agreement is independent only below the first shared unchallenged stage.
Frozen free-field benchmark
Section titled “Frozen free-field benchmark”Generate one checksum-frozen regulated free-field dataset with known covariance, exact purity, replica ratio, and continuum target. Provide raw records rather than only summaries. Teams or isolated code paths implement:
- covariance entropy;
- swap or replica purity;
- randomized purity or selected shadows;
- full finite-mode likelihood where feasible.
Freeze conventions, target regions, and acceptance tests, but do not share estimator internals. Compare central values, interval coverage over repeated datasets, residuals, and computational failures.
Two seeded failures
Section titled “Two seeded failures”Shared normalization error. Multiply all raw field amplitudes or partition-function ratios by a common hidden factor. Methods using the same calibration should shift together; a truly independent reference should expose the agreement as shared bias.
Method-specific branch error. Introduce a wrong replica-sheet branch or a sign in the permutation estimator. Replica results should fail exact small-system and product-state tests while covariance results remain correct.
Additional nulls include non-Gaussian states matching covariance, copy drift, ill-conditioned random ensembles, finite-volume recurrences, and unresolved symmetry sectors. Each null has a prespecified expected signature.
Decision rule
Section titled “Decision rule”Do not declare success from pairwise overlap of intervals alone. Evaluate a discrepancy statistic with the full cross-method covariance and prespecified tolerance. If methods disagree, trace the dependency map and test one failure channel at a time. If they agree but share an untested stage, the claim remains conditional on that stage. Transparent preregistration and open-research practices help distinguish tests fixed in advance from choices made after seeing the answer Nosek et al. 2015, pp. 1422–1425.
As of 10 August 2026, method diversity is a validation strategy, not proof of correctness. Independent data, calibration, code, and analytic assumptions determine its strength.
Exercises
Section titled “Exercises”False independence. Two codes give the same answer from the same binned file. What independence has been demonstrated?
Solution
Only implementation differences after binning. Acquisition, calibration, selection, normalization, and binning biases are shared. Preserve and independently process earlier-stage records to test them.
Blind injection. Why should a seeded failure be hidden from the analysis team?
Solution
Knowing the failure can encourage tuning the test to the answer. A blind injection measures whether the prespecified validation actually detects the class of error at realistic magnitude.
Inference and failure-control maps
Section titled “Inference and failure-control maps”The first diagram traces the complete path from raw records to a bounded information claim; inspect the assumption attached to every arrow. The second maps shared and method-specific failure channels to held-out tests, regulator variation, replication, and correction.
Entropy, tomography, witness, and recovery methods enter at the estimator stage, but all share calibration, uncertainty, continuum, and alternative-model tests. The final statement is no stronger than the least validated arrow. The diagram is schematic and not to scale.
Different estimators can share the same calibration or normalization bias, so numerical agreement is not automatically independent replication. Adversarial nulls, held-out observables, regulator variation, and genuinely independent implementations set the claim ceiling and trigger correction when needed. The diagram is schematic.
References
Section titled “References”- Nosek, Brian A., George Alter, George C. Banks, et al. “Promoting an Open Research Culture.” Science 348 (2015): 1422–1425. DOI.