Statistical ensembles and probability repair
Statistical reasoning in field theory becomes unreliable when three different questions are answered with the same formula. A probability law determines expectations once the law is given. A physical ensemble explains why a particular law models the system. An estimator and its uncertainty describe what finite data can reveal about a target quantity. This lesson separates those layers through one finite canonical ensemble, one dependence example, and one correlated-sample calculation.
Required background. Linear and tensor methods supplies finite sums, quadratic forms, and covariance matrices. If those operations are comfortable, no measure theory or thermodynamic limit is needed to begin.
Probability laws, ensembles, and samples answer different questions
Section titled “Probability laws, ensembles, and samples answer different questions”Suppose a finite set of configurations is labeled by , a normalized law is , and an observable has value . Then
are probability identities. They are true for any declared normalized law; they do not explain where came from.
A canonical ensemble adds physical content:
This law represents equilibrium with a heat bath at fixed inverse temperature under specified Hamiltonian, volume, conserved quantities, and boundary conditions. The algebra that follows from is exact, but the claim that a laboratory system or simulation is described by that law is a model assumption requiring separate support.
Finally, observations introduce an inference problem. The sample mean
is an estimator of an ensemble mean only after its sampling process is specified. Independent equilibrium draws, a correlated stationary Markov chain, and a chain that has not equilibrated can have the same list of recorded values but justify different uncertainty statements.
Keep this classification visible:
| Layer | Question | Typical statement |
|---|---|---|
| Probability | What follows from the declared law? | |
| Physical ensemble | Why should this law describe the system? | The system is in canonical equilibrium at fixed and volume. |
| Inference | What does finite, possibly dependent data establish? | estimates with a covariance-aware standard error. |
The first statement can be proved algebraically, the second is part of the physical model, and the third needs sampling assumptions. None can substitute for the others.
A finite canonical ensemble gives an exact response identity
Section titled “A finite canonical ensemble gives an exact response identity”Consider three nondegenerate configurations with energies
Assume canonical equilibrium at inverse temperature , fixed volume, and . These are the physical assumptions. No samples or estimators have entered. Writing gives
The mean and second moment are
so
For any finite canonical law,
Multiplying by and summing gives
Because in the chosen units, the heat capacity at fixed external parameters is
This response–fluctuation identity is exact for the declared finite canonical ensemble. It is not evidence that an arbitrary dataset has equilibrated, and it does not hold unchanged after replacing the canonical law by a different ensemble. The general partition-function method is developed in Partition functions and thermodynamic response and Kardar 2007, chs. 1–3.
Two limits check the arithmetic. As , and the ground state dominates:
As , all three configurations become equally weighted:
These are temperature limits of a fixed finite system. Since is a finite sum of positive analytic functions for real , it has no finite-system thermodynamic nonanalyticity. A phase transition, spontaneous symmetry breaking, or equivalence of ensembles requires an appropriate infinite-system limit and its hypotheses.
Zero covariance does not imply independence
Section titled “Zero covariance does not imply independence”Let be uniform on and set . Then
so
Nevertheless, and are dependent: is determined by , and
Covariance tests one bilinear moment. Independence requires the entire joint law to factorize, equivalently for all relevant events. In a Gaussian model, zero covariance has stronger consequences, but that is an additional distributional assumption, not the definition of independence. See Durrett 2019, § 2.1, pp. 43–49, PDF for the standard independence criteria.
This example is purely probabilistic. It makes no physical ensemble claim and uses no finite-data inference.
Correlated samples change the uncertainty of the mean
Section titled “Correlated samples change the uncertainty of the mean”Let be a stationary sequence with common mean , variance , and autocorrelation
Stationarity and the common mean imply , so correlation alone does not bias this estimator. It does change its variance:
Define the finite- integrated autocorrelation time and a variance-equivalent effective sample size by
Then
For the covariance model with ,
at fixed and large . With and , the exact finite sum gives . The correlation-aware standard error is therefore about , almost three times the independent-sample value .
Effective sample size is a summary of the variance for this estimator and observable; it is not a literal count of independent configurations. Negative correlations can even produce . When and the autocorrelation tail are estimated from the same finite series, window or block choices add uncertainty. Stationarity, equilibration, and a suitable central limit theorem must be justified separately before turning a standard error into a confidence interval. Autocorrelation sums, window selection, and error-estimate stability are treated in Wolff 2004/2006, §§ 2 and 3.1–3.3, pp. 4–13.
Do not confuse these quantities:
- is the ensemble fluctuation of one draw;
- is the sampling variance of an estimator;
- a confidence interval needs a sampling-distribution argument as well as an estimated standard error; and
- discretization, finite-volume, model, and equilibration errors are distinct possible biases, not contributions automatically captured by the standard error.
Finite-data and thermodynamic limits are different
Section titled “Finite-data and thermodynamic limits are different”Several limits can appear in one field-theory calculation. State which one is being taken and what remains fixed.
| Limit | Question answered | What it does not establish by itself |
|---|---|---|
| samples | Does an estimator converge under the sampling law? | That the physical ensemble or simulation algorithm is correct |
| Volume | Does a thermodynamic state, phase, or nonanalyticity emerge? | A central limit theorem for the recorded samples |
| Lattice spacing | Does a continuum observable exist after tuning and renormalization? | Ensemble equivalence or equilibration |
| Source | Does a response or order parameter persist? | That the order of and is interchangeable |
For the finite three-state example, a microcanonical law at exact energy places all probability on the middle configuration and gives zero energy variance. The canonical law mixes all three energies and generally has nonzero variance. Relating those descriptions requires additional degrees of freedom, a declared thermodynamic limit, and appropriate regularity or concavity conditions. Use Thermodynamic limits, phases, and ensemble equivalence before making an equivalence or phase-transition claim.
Exercises
Section titled “Exercises”Normalize first, then condition
Section titled “Normalize first, then condition”Three outcomes have weights , respectively. Let the observable take values on those outcomes. Normalize the law, compute and , and find . Classify the calculation by layer.
Solution
The total weight is , so
The first two moments are
and
Therefore
The conditioning event has probability , so
Every step is a probability identity once the weights and observable are declared. No physical reason for those weights and no sample-based estimator has been supplied.
Generate a mean, variance, and covariance
Section titled “Generate a mean, variance, and covariance”For finite configurations with energy and another observable , set
Show that derivatives of generate , , and . Identify the physical assumption and the probability identities.
Solution
The normalized law is
Differentiating once gives
A second derivative differentiates both the numerator and the normalization:
Similarly,
Choosing the Boltzmann form as a model of canonical equilibrium is the physical ensemble assumption. Once that normalized law is accepted, the derivative relations are probability identities. No inference choice appears because no finite sample has been introduced.
Retain every covariance term
Section titled “Retain every covariance term”Let , , and . Compute , , and the ratio of the correct standard error to the independent-sample standard error.
Solution
There are three diagonal terms, four ordered pairs at lag one, and two ordered pairs at lag two. Hence
Matching this to gives
well below the three recorded observations. The independent formula would give variance , so the standard-error ratio is
Thus the correct standard error is about larger than the independent-sample value for this finite covariance model. This conclusion assumes the stated stationary covariance law is known; estimating it from three data points would not be credible.
Return when the three layers stay separate
Section titled “Return when the three layers stay separate”You are ready to return when you can do all three of the following without looking back:
- normalize a finite law, compute moments or conditional probabilities, and test independence from the joint law rather than covariance alone;
- state the physical assumptions selecting an ensemble and derive one response–fluctuation identity without treating it as evidence of equilibration; and
- name an estimand, retain covariance terms in its uncertainty, and state which theorem or sampling assumptions would be needed for an interval or large- claim.
Re-run the statistical and probability diagnostic with a different finite spectrum and correlation parameter. If all three tasks are Demonstrated, return to the page that sent you here. If correlated uncertainty is now clear but precision, convergence, testing, or reproducibility is still blocking the work, continue to Numerical and reproducibility repair and then repeat the computational and evidence diagnostic.
References
Section titled “References”- Durrett, Rick. Probability: Theory and Examples. 5th ed. Cambridge University Press, 2019. Official open PDF.
- Kardar, Mehran. Statistical Physics of Particles. Cambridge University Press, 2007. DOI.
- Wolff, Ulli. “Monte Carlo Errors with Less Errors.” Computer Physics Communications 156, no. 2 (2004): 143–153. DOI; arXiv:hep-lat/0306017.