Probabilistic Convergence, Laws of Large Numbers, and Central Limit Theorems
An estimator is justified only after its limiting claim has been matched to the right mode of convergence and to a theorem whose hypotheses actually hold. Consistency is usually a statement such as , justified by a law of large numbers. Pathwise stabilization calls for almost-sure convergence, while control of a mean error calls for convergence. A Gaussian fluctuation approximation is different: it is a convergence-in-distribution statement for a centered and rescaled error, usually supplied by a central limit theorem. Gallager 2011, slides 9–31 compares these convergence modes and the law-of-large-numbers-to-central-limit progression in a common set of examples.
For independent identically distributed observations, an integrable first moment is enough for the strong law of large numbers, and a finite positive variance is enough for the classical central limit theorem. Those statements do not by themselves give a finite-sample error bound, convergence of moments, or a central limit theorem for correlated Monte Carlo output. This page makes those boundaries explicit and ends with one finite-regulator Euclidean QFT calculation where every probabilistic assumption can be checked.
Required background. Probability Spaces, Random Variables, and Conditional Expectation, supplies probability spaces, expectation, integrability, and independence.
Helpful background. Limits, Completeness, and Modes of Convergence provides the general convergence language used to compare almost-sure, probability, , and distributional limits.
Modes of probabilistic convergence
Section titled “Modes of probabilistic convergence”Readers who need a shorter review can use the statistical ensembles and probability repair.
Unless stated otherwise, are random vectors on one probability space , taking values in a finite-dimensional Euclidean space. The norm is , and is understood in every limit. We write
The first three notions compare random variables on a common probability space. Convergence in distribution compares their laws, so the variables can even be realized on different probability spaces. In the QFT application, the sampling limit is taken at fixed finite regulator and fixed finite volume. No continuum or thermodynamic limit is interchanged with the sample size limit.
Four notions of probabilistic convergence
Section titled “Four notions of probabilistic convergence”Almost-sure convergence
Section titled “Almost-sure convergence”The sequence converges to almost surely when
Outside one fixed null event, every sample path eventually follows the ordinary pointwise limit. This is the natural conclusion when a theorem is supposed to describe the long-run behavior of almost every repeated experiment.
Convergence in probability
Section titled “Convergence in probability”The sequence converges to in probability when, for every ,
This is the usual mathematical form of consistency. For a parameter , the claim says that every fixed error tolerance is violated with a probability tending to zero. It does not give a rate unless the proof supplies one.
Convergence in mean
Section titled “Convergence in mean”For and , the sequence converges to in when
Thus convergence controls mean absolute error, and convergence controls mean-square error. Membership of both the sequence and its limit in is part of the definition; requiring only an integrable difference would fail to exclude a shared non- component.
Convergence in distribution
Section titled “Convergence in distribution”The sequence converges to in distribution when
for every bounded continuous function . Equivalently, the probability laws of converge weakly to the law of . For real-valued variables, this is equivalent to
at every continuity point of the distribution function of . Distributional convergence is therefore the right language for a central limit theorem, but it is deliberately insensitive to unbounded test functions such as or .
What implies what
Section titled “What implies what”On a probability space, for ,
and separately
If the proposed limit is a constant , there is also the useful equivalence
Here the symbol on the left denotes the degenerate random variable equal to .
Proof of the moment and probability arrows. Since , Hölder’s inequality gives
For every , Markov’s inequality then gives
Thus convergence implies convergence, and convergence implies convergence in probability.
Proof of the almost-sure arrow. If almost surely, then almost surely. The indicators are bounded by , so dominated convergence shows that their expectations, which are the relevant probabilities, tend to zero.
The implication from probability to distribution and the constant-limit converse are standard weak-convergence results. One route to the converse is to apply the Portmanteau theorem to the closed complement of every ball about . Durrett 2019, § 3.2, pp. 116–124, PDF proves these weak-convergence tools.
The missing arrows are genuinely missing
Section titled “The missing arrows are genuinely missing”The following examples are not pathologies to be ignored; they identify the extra hypotheses needed in later theorems.
Probability need not imply almost surely. On , write each uniquely as with , and define
Then
so and even for every finite . Yet each lies in one selected interval at every level , so infinitely often. There is no almost-sure convergence to zero.
Almost surely need not imply . On with Lebesgue measure, set
For every , eventually , so . Nevertheless,
for every . The shrinking exceptional set is offset by a growing spike.
Distribution need not imply probability. Let take the values and with equal probability, and define for every . Then and have the same law, so , but
Weak convergence sees the marginal laws, not the coupling between the two variables.
Weak convergence need not preserve expectations. Taking in the spike example gives almost surely, hence in distribution, while
Bounded continuous test functions cannot detect the missing mass in the unbounded first moment.
Moving limits through functions and expectations
Section titled “Moving limits through functions and expectations”Limit theorems become useful for estimators only after one knows which operations preserve their conclusions.
Continuous mapping
Section titled “Continuous mapping”Continuous mapping theorem. Suppose , and let be measurable. If the set of discontinuity points of satisfies
then
This is a cited theorem, not proved here; Durrett 2019, § 3.2, pp. 116–124, PDF gives the Portmanteau and continuous-mapping arguments. The condition on cannot simply be deleted. For example, converges to , but with one has and .
Slutsky’s theorem
Section titled “Slutsky’s theorem”Slutsky’s theorem. For real-valued variables, if and for a constant , then
and, when ,
Independence of and is not required. The key fact is the joint convergence ; continuous mapping then supplies the displayed conclusions. Vector and matrix versions follow by choosing the corresponding continuous map. This is what makes replacement of an unknown standard deviation by a consistent estimate legitimate.
Uniform integrability
Section titled “Uniform integrability”Convergence in probability controls the probability of a fixed-size error, not its expectation. A family of random vectors is uniformly integrable when
The spike example is bounded in but is not uniformly integrable, so an bound alone is insufficient. A convenient sufficient condition is
because the tail in the definition is then bounded by a constant times .
Vitali convergence theorem. If and is uniformly integrable, then and
Conversely, convergence implies uniform integrability of the sequence. More generally, if , , and is uniformly integrable, then . If instead almost surely and for a common , dominated convergence gives directly.
For real random variables, weak convergence plus uniform integrability of is enough to recover convergence of first moments. Without such a tail condition, a CLT or any other weak limit cannot be used to pass an expectation through the limit. Durrett 2019, § 4.6, pp. 244–247, PDF develops uniform integrability and its connection with convergence.
Laws of large numbers
Section titled “Laws of large numbers”Let be iid real random variables, and define
The independence, identical-distribution, and moment assumptions below are part of the theorem. More general laws exist, but none follows merely from the notation .
Weak law
Section titled “Weak law”Weak law of large numbers. If
then
The integrable iid theorem is cited from Durrett 2019, § 2.2, pp. 56–65, PDF. Here is the complete elementary proof under the stronger assumption . Independence gives
Chebyshev’s inequality therefore yields, for every ,
This proof supplies a conservative probability bound as well as consistency. Finite variance is sufficient for the argument, but it is not necessary for the weak law.
Strong law
Section titled “Strong law”Strong law of large numbers. Under the same iid first-moment hypothesis,
one has
This full-strength result is cited from Durrett 2019, §§ 2.3–2.5, pp. 67–87, PDF. A proof roadmap is as follows: truncate the variables to control rare large values, use summability and the Borel–Cantelli lemmas to show that truncation changes only finitely many terms almost surely, control the centered truncated sums, and then remove the truncation. Each step is substantive; the finite-variance Chebyshev calculation above is not a proof of this almost-sure statement.
The strong law answers a pathwise question and implies the weak law. It still does not specify how rapidly the error decreases or what its rescaled distribution looks like.
Why the assumptions matter
Section titled “Why the assumptions matter”If has a standard Cauchy law, then every average has the same Cauchy law. There is no finite mean for the averages to estimate, and they do not converge in probability to a constant.
Independence also does real work. If for all , with a nondegenerate integrable random variable, then
for every . Identical marginal distributions alone do not create averaging. Dependent sequences require an ergodic theorem or another law of large numbers with its own hypotheses.
Central limit theorems
Section titled “Central limit theorems”A law of large numbers describes the unscaled error . A central limit theorem describes its leading random fluctuation after multiplication by .
The iid finite-variance theorem
Section titled “The iid finite-variance theorem”Classical central limit theorem. Let be iid with
Then
No third moment and no moment-generating function are required. The conclusion concerns the standardized sum, not the law of an individual , and it is asymptotic rather than an assertion that the finite- distribution is exactly Gaussian.
Proof sketch by characteristic functions. Put . Finite second moment gives the expansion
Independence then gives
The limit is the characteristic function of a standard normal law, and the Lévy continuity theorem completes the proof. The expansion and continuity theorem are the key lemmas; their systematic development belongs to Characteristic Functions, Moments, Cumulants, and Generating Functionals. Durrett 2019, § 3.4.1, pp. 143–146, PDF gives the complete theorem and characteristic-function proof.
Independent but non-identical summands
Section titled “Independent but non-identical summands”The plural “central limit theorems” matters because the iid theorem is not the only useful form. Let
be a triangular array whose variables are independent within each row, centered, and have finite variances. Define
Lindeberg–Feller central limit theorem. If, for every ,
then
This is a cited theorem; see Durrett 2019, § 3.4.2, pp. 147–152, PDF. The Lindeberg condition says that no asymptotically significant fraction of the total variance comes from rare terms comparable to the full row scale. It replaces identical distribution, but not row-wise independence. A Lyapunov moment condition is a common stronger sufficient condition.
What controls finite-sample normal accuracy
Section titled “What controls finite-sample normal accuracy”The classical CLT alone gives no universal value of at which a Gaussian approximation becomes accurate. Under the additional third-moment assumption
the Berry–Esseen theorem gives a quantitative bound. If is the standard normal distribution function, then a universal constant exists such that
This controls absolute error between distribution functions. It is not a relative tail bound and does not guarantee that densities are close. Durrett 2019, § 3.4.4, pp. 157–158, PDF states and proves a Berry–Esseen estimate.
Vector observables and studentized errors
Section titled “Vector observables and studentized errors”Let be iid random vectors in with
Then the multivariate central limit theorem states
The covariance matrix may be singular. A proof sketch follows from the scalar theorem: for every fixed ,
and the Cramér–Wold device reconstructs the joint weak limit.
Durrett 2019, § 3.10, pp. 199–201, PDF gives the multivariate convergence criterion and central limit theorem. Moreover,
exactly for iid samples.
Return to iid real under the classical CLT assumptions. For , define the sample variance
The identity
and the strong law applied to and show that almost surely. Let be the nonnegative square root. Slutsky’s theorem therefore gives the studentized asymptotic limit
On the event , the ratio may be defined arbitrarily; because , the probability of that event tends to zero. Thus the result justifies replacing by a consistent sample estimate in the limit. It does not assert a finite- Student law unless the additional normal-sample assumptions for that exact result hold.
Sampling estimates in a regulated Euclidean theory
Section titled “Sampling estimates in a regulated Euclidean theory”Consider a finite regulator with real field coordinates, a measurable action , and a positive, normalizable Euclidean weight
For independent exact draws and a measurable scalar observable , set
The theorem to use is now selected by the integrability that the regulated measure actually provides:
- If , the strong law gives almost surely.
- If also , the classical CLT gives .
- For a vector of observables with finite second moments, the multivariate theorem carries their full covariance matrix into the Gaussian limit.
These are sampling statements. They do not prove that the regulator may be removed, that the volume may be sent to infinity, or that a formal continuum functional integral exists.
An exact two-mode check
Section titled “An exact two-mode check”Continue the finite Gaussian model from the prerequisite page,
Under its normalized Euclidean measure, the first coordinate is Gaussian with mean zero and variance
For independent exact configurations, let denote that coordinate and estimate its two-point moment by
The Gaussian moments give
Consequently, the strong law and CLT yield
This example also exposes the difference between exact and asymptotic normality. Since are independent standard normals,
Thus is not Gaussian at finite ; its centered and scaled law only approaches a Gaussian. By contrast, is exactly standard normal for every because a linear combination of Gaussian variables is Gaussian. The comparison is a useful check that a CLT conclusion should not be stronger than the finite-sample law permits.
Why Markov-chain output needs a different theorem
Section titled “Why Markov-chain output needs a different theorem”Practical lattice calculations usually record successive states of a Markov chain rather than independent exact draws. For a real, second-order stationary observable sequence , sample the terms and define
Counting all pairs at each lag gives the exact identity
If every nonzero-lag covariance vanishes, this expression reduces to ; vanishing covariance alone does not establish independence. If the covariance series is absolutely summable, the exact identity and dominated convergence give
If, in addition, theorem-specific moment and dependence hypotheses establish a CLT with this covariance-sum variance, then
The phrase “if a suitable CLT has been proved” is essential. Stationarity and ergodicity alone do not imply a central limit theorem, and covariance summability by itself is not a universal substitute for mixing, regeneration, spectral, or other theorem-specific hypotheses. Wolff 2004/2006, §§ 1–2, pp. 2–6 starts from equilibrated chain output, makes its normal approximation and finite correlation-scale assumptions explicit, and shows how autocorrelations enter the variance.
The mathematical hypotheses for Markov-chain laws and CLTs belong to Markov Generators, Semigroups, Ergodicity, and Correlated-Sample Error. The developed physical treatment of estimator covariance, blocking, jackknife/bootstrap, ratios, and autocorrelation-aware error analysis belongs to Estimators, Covariance, and Resampling.
Finally, the probability model itself must exist. A positive normalized finite-dimensional Euclidean weight can support the claims above. A Lorentzian factor , a sign-changing weight, or a merely formal continuum path integral is not a probability measure to which these theorems can be applied directly.
Choosing the conclusion before choosing the theorem
Section titled “Choosing the conclusion before choosing the theorem”| Intended claim | Mathematical target | Typical iid hypothesis | What still needs separate work |
|---|---|---|---|
| Consistency of a sample mean | rate and finite- coverage | ||
| Pathwise long-run average | iid and | convergence rate | |
| Mean-square accuracy | estimator-specific bounds | tail control and bias | |
| Gaussian root- fluctuations | iid and | finite- accuracy | |
| Studentized Gaussian limit | replace by | CLT plus consistent | exact finite- law |
| Correlated-sample Gaussian limit | dependent-sequence CLT | theorem-specific dependence and moment conditions | autocorrelation estimation and diagnostics |
This table gives sufficient hypotheses for standard routes, not necessary conditions for every possible model. The decisive habit is to write the desired convergence statement first and then verify the theorem’s assumptions against the actual random variables.
Common pitfalls
Section titled “Common pitfalls”Calling every limiting statement “convergence.” Almost-sure, probability, , and distributional convergence answer different questions. State the mode explicitly and use only the implications that have been proved.
Reading moment convergence from a CLT. Weak convergence tests bounded continuous functions, whereas moments are unbounded. Add uniform integrability or another justified tail bound before passing expectations through a weak limit.
Treating a CLT as a finite-sample Gaussian identity. A CLT describes a centered, scaled sequence as . Berry–Esseen needs an additional third moment, and even that theorem controls distribution functions rather than making the data exactly normal.
Applying iid formulas to a Markov chain. Correlation changes the variance of a mean and can invalidate the iid proof entirely. Establish an appropriate ergodic theorem and dependent-sequence CLT before interpreting a Monte Carlo standard error.
Using a formal QFT weight as a probability law. Positivity, normalization, measurability, and the relevant moments are hypotheses, not notation. Check them at the regulated level and keep regulator limits separate from sampling limits.
Check your understanding
Section titled “Check your understanding”1. Retrieve the implication structure
Section titled “1. Retrieve the implication structure”Suppose . Which of , probability, and distributional convergence follow automatically? Does almost-sure convergence follow?
Solution
Because the underlying measure has total mass one,
so convergence follows. Markov’s inequality then gives convergence in probability, which implies convergence in distribution. Almost-sure convergence does not follow in general; the moving-interval example above converges in every finite while taking the value infinitely often on every sample path.
2. Diagnose missing hypotheses
Section titled “2. Diagnose missing hypotheses”For each construction, identify why the usual sample-mean conclusion fails:
- for every , where is integrable and nondegenerate.
- are iid standard Cauchy variables.
Solution
In the first construction the variables are identically distributed but not independent, and for all . It cannot converge in probability to the constant unless is already degenerate.
In the second construction independence holds, but the first absolute moment does not exist. Stability of the Cauchy law gives for every , so there is no deterministic law-of-large-numbers limit and no finite-variance classical CLT.
3. Extract a finite-sample guarantee
Section titled “3. Extract a finite-sample guarantee”Let be iid with mean and variance . For and , find a sufficient sample size for
using only the information stated.
Solution
Chebyshev’s inequality gives
Therefore it is sufficient to choose
This bound can be conservative, but unlike the bare CLT it is a valid finite- statement under only a finite-variance assumption.
4. Transfer the result to a correlated sequence
Section titled “4. Transfer the result to a correlated sequence”Suppose a stationary scalar sequence has autocovariance
and suppose separately that a suitable CLT for the sequence has been established. Compute its asymptotic variance for the sample mean.
Solution
The geometric series gives
Positive inflates the iid variance, while negative reduces it. The covariance calculation does not itself prove the assumed CLT; that requires separate dependence hypotheses.
Synthesis and continuations
Section titled “Synthesis and continuations”Use convergence in probability to state ordinary consistency, almost-sure convergence for pathwise stabilization, convergence for mean-error control, and convergence in distribution for a rescaled fluctuation law. For iid data, integrability supplies both weak and strong laws of large numbers, while finite positive variance supplies the classical central limit theorem. Uniform integrability is the extra bridge needed when unbounded moments must pass through a limit, and dependence requires a different LLN or CLT rather than an iid formula with new notation.
The next page develops the transform language used in the CLT proof: Characteristic Functions, Moments, Cumulants, and Generating Functionals. For correlated stochastic dynamics, continue later to Markov Generators, Semigroups, Ergodicity, and Correlated-Sample Error. For production lattice estimators and uncertainty analysis, continue to Estimators, Covariance, and Resampling.
References
Section titled “References”- Rick Durrett, Probability: Theory and Examples, fifth edition, PDF, Cambridge University Press, 2019. §§ 2.2–2.5, pp. 56–87, support the convergence modes and weak and strong laws; § 3.2, pp. 116–124, supports weak convergence, Portmanteau, and continuous mapping; §§ 3.4.1–3.4.2, pp. 143–152, support the iid and Lindeberg–Feller central limit theorems; § 3.4.4, pp. 157–158, supports the Berry–Esseen bound; § 3.10, pp. 199–201, supports multivariate weak convergence and the multivariate CLT; and § 4.6, pp. 244–247, supports uniform integrability and convergence.
- Robert Gallager, “Lecture 3: Laws of Large Numbers, Convergence”, MIT OpenCourseWare 6.262, Spring 2011, slides 9–21. This is the teaching source for the distinction among mean-square, probability, distributional, and almost-sure convergence and for the LLN-to-CLT progression.
- Ulli Wolff, “Monte Carlo Errors with Less Errors”, arXiv:hep-lat/0306017v4, 2006 revision, §§ 1–2, pp. 2–6. These sections derive error estimates for equilibrated Markov-chain output, autocovariance contributions, integrated autocorrelation time, and the explicit assumption behind a Gaussian error approximation.