Correlated Fits, Model Selection, and Stability Tests
A lattice fit usually acts on a highly correlated vector whose covariance is itself estimated. A convincing parameter estimate therefore needs more than a plausible curve and acceptable : the model must be identifiable on the chosen window, covariance treatment must be stable and calibrated, residuals must be structured as expected, and competing descriptions must be tested on synthetic or held-out data.
Required background. Estimators, covariance, and resampling define the joint data uncertainty; operator bases and excited states define the spectral models.
Helpful background. Continuum extrapolation shows the same correlated-fit principles across ensembles.
Correlated objectives
Section titled “Correlated objectives”Local fit convention and regime. The fitted vector, time or ensemble window, covariance estimator, effective number of independent blocks, model family, and any priors are fixed as one procedure. is the covariance of the estimated data vector, not of individual raw measurements, and every synthetic or resampled repetition re-estimates and regularizes it exactly as production does. Reported intervals cover the complete declared selection procedure, not only a fit conditional on the chosen model and window.
For data vector , model , and covariance , the Gaussian objective is
This form is appropriate only when the estimated mean is sufficiently Gaussian and the covariance procedure is part of the calibration. If is estimated from the same finite sample, the nominal distribution is altered. Resampling or synthetic repetitions should include covariance re-estimation; Michael 1994, pp. 2616–2619 treats the finite-sample correlated-fit issue in a lattice setting.
Whitened residuals
should show no systematic time, momentum, or ensemble structure. A small total can coexist with coherent residuals when covariance eigenmodes are poorly estimated or overregularized.
Spectral-window example
Section titled “Spectral-window example”Consider
A one-state fit at late time reduces excited-state bias but loses signal; a two-state fit at earlier time adds weakly identified parameters. Vary and , operator basis, and model order. Inspect parameter correlations and the singular values of the model Jacobian
If a direction of is nearly null, the data do not identify that parameter without a prior or external constraint. A stable caused by a strong prior is prior-conditioned evidence and should be labeled as such; Lepage et al. 2002, pp. 12–20 gives the constrained-fitting framework rather than a license to erase prior sensitivity.
Synthetic closure should draw correlated two-state data with known parameters, rerun the complete window and model-selection procedure, and measure bias and interval coverage. A plateau-only adversary with an unresolved second state is particularly important.
Covariance regularization
Section titled “Covariance regularization”If is noisy, hard singular-value cuts replace by a pseudoinverse on retained modes; shrinkage uses
for a declared target . Both alter weights and goodness-of-fit calibration. Vary the cutoff or shrinkage strength, show which residual modes are removed, and repeat coverage tests. Choosing the setting that minimizes the final error is an analysis choice, not validation.
Model comparison and averaging
Section titled “Model comparison and averaging”Information criteria, marginal likelihoods, cross-validation, or predeclared fit-quality rules compare models under different assumptions. None proves that the candidate set contains the truth. Akaike’s information-criterion construction is derived in Akaike 1974, pp. 716–723. If estimates are combined with weights , a useful variance decomposition is
State whether are fixed, data dependent, Bayesian, or heuristic, and calibrate the complete selection or averaging procedure. Treating the spread of only accepted models as a frequentist confidence interval has no automatic coverage guarantee.
The graph places model choices upstream of scale, matching, and continuum fits, so their covariance and alternatives must propagate.
Correlated fitting is one node in a larger dependency graph. Window, covariance, and model alternatives must be repeated through scale, renormalization, and continuum stages rather than appended as an uncorrelated percentage. The diagram is schematic.
Stability tests that answer different questions
Section titled “Stability tests that answer different questions”| Variation | Diagnoses | Does not prove |
|---|---|---|
| Move fit window | contamination and signal loss | completeness of model set |
| Change operator basis | overlap and rank sensitivity | absence of all missing states |
| Vary covariance conditioning | noisy eigenmode dependence | correct likelihood calibration |
| Add/remove state or cutoff term | model-order sensitivity | that the larger model is identifiable |
| Hold out times or ensembles | local predictive adequacy | continuum universality |
| Synthetic closure | bias and coverage under fixture | correctness outside fixture family |
Adversarial failure: a discarded covariance mode carries the discrepancy
Section titled “Adversarial failure: a discarded covariance mode carries the discrepancy”Take a trusted positive-definite fixture with eigenpair and inject
relative to the fitted model. With the full covariance, this one direction contributes . A hard singular-value cut that discards assigns it zero weight and can turn the same visibly constrained discrepancy into an acceptable reported fit. The production test must expose the removed mode, repeat the complete cut rule on synthetic data, and fail if nominal intervals lose coverage or the target shifts materially as that mode is restored.
Observable-level validation checklist
Section titled “Observable-level validation checklist”- Record the complete data vector, fit window, covariance source, conditioning rule, candidate models, priors, and selection rule before interpreting the target parameter.
- Inspect whitened residuals by time, momentum, channel, and ensemble rather than relying on a scalar .
- Diagnose identifiability with parameter correlations and singular values of ; vary any influential prior.
- Track every removed or shrunk covariance eigenmode and repeat goodness-of-fit calibration after covariance re-estimation.
- Run the full window, conditioning, and model-selection procedure on synthetic correlators and measure bias plus interval coverage.
- Inject the five-sigma covariance-mode adversary above and require the stability report to reveal its effect.
- Hold out at least one informative time range, channel, or ensemble when the data volume permits, and state what that prediction does not test.
Common pitfalls
Section titled “Common pitfalls”Using visual smoothness as fit validation. Strong correlations make many points move together; a curve can look excellent while missing a constrained eigenmode.
Selecting a window after viewing the desired value. Freeze the rule, blind the target, or propagate the selection procedure through alternatives.
Reporting only the preferred model. Show the plausible set, weights or acceptance rule, and the between-model component.
Learning outcomes
Section titled “Learning outcomes”- Fit a correlated synthetic correlator with at least two plausible spectral models and quantify parameter, window, covariance-conditioning, and prior sensitivity with bias and coverage results.
- Given a smooth-looking fit or discarded-mode adversary, decide whether the target is identifiable and report a model-dependent interval without treating visual stability or nominal goodness of fit as proof of adequacy.
Exercises
Section titled “Exercises”- Why can dropping a small covariance eigenvalue reduce dramatically?
Solution
The inverse weights residuals by . A small eigenvalue gives a large penalty to its residual direction; dropping it removes that constraint entirely, whether it was noisy or informative.
- What does a successful two-exponential synthetic closure fail to establish?
Solution
It establishes behavior for the tested noise, covariance, amplitudes, gaps, and candidate models. It does not cover oscillating states, continuum spectra, different gaps, non-Gaussian tails, or other omitted structures.
References
Section titled “References”- Akaike, H. (1974). A new look at the statistical model identification. IEEE Transactions on Automatic Control, 19, 716–723. DOI.
- Lepage, G. P. et al. (2002). Constrained curve fitting. Nuclear Physics B Proceedings Supplements, 106–107, 12–20. DOI.
- Michael, C. (1994). Fitting correlated data. Physical Review D, 49, 2616–2619. DOI.