Markov Generators, Semigroups, Ergodicity, and Correlated-Sample Error
A Markov transition kernel answers a local question: given the present state, what is the probability law of the next state? Acting on observables, the kernel takes a conditional expectation; acting on probability measures, it propagates a law. Repeated kernels form a discrete semigroup , while a time-homogeneous continuous process gives a one-parameter semigroup . On a function space where is strongly continuous, its infinitesimal generator is
The domain is part of this definition. A formal differential expression without a function space, boundary conditions, and a domain is not yet a generator.
An invariant probability measure is a fixed point of the forward evolution. It does not by itself imply convergence to equilibrium, an ergodic theorem, a mixing rate, or a central limit theorem. Once suitable hypotheses do give a stationary Markov-chain CLT, serial covariance changes the uncertainty of the sample mean from the iid expression. With the convention used here,
Both quantities depend on the observable . The page develops these statements from kernels through uncertainty and then checks them in a regulated Gaussian field mode whose stationary evolution can be reversible or nonreversible.
Required background. Stochastic Processes and Correlation Functions for conditional laws, stationarity, covariance, and correlation functions; Probabilistic Convergence, Laws of Large Numbers, and Central Limit Theorems for modes of convergence, iid limit theorems, and the distinction between a limit law and a finite-sample bound.
Markov kernels, semigroups, and sampling scope
Section titled “Markov kernels, semigroups, and sampling scope”Work first on a measurable state space . Bounded measurable observables are denoted by , probability laws by , and the expectation of under an invariant law by
The operator convention is backward on observables: is the conditional expectation of the future observable given the present state . The corresponding forward action propagates measures. This is the same convention used for the local diffusion operator on Brownian Motion, Stochastic Calculus, Langevin Equations, and Fokker–Planck Dynamics, but no stochastic-calculus result from that page is needed below.
The probability statements apply to genuine, normalized, nonnegative measures. A regulated Euclidean or thermal field mode can supply such a measure. A formal continuum functional integral, a sign-changing weight, or the oscillatory factor does not automatically do so; Markov-chain probability theorems cannot be transferred to those expressions without a separate construction.
Pavliotis develops transition functions, Markov semigroups, generators, and invariant measures in Pavliotis 2014, §§ 2.2–2.4, pp. 30–39, PDF. The discussion below makes the operator domains and the additional hypotheses behind long-time conclusions explicit.
A transition kernel has two actions
Section titled “A transition kernel has two actions”A time-homogeneous one-step Markov kernel satisfies two requirements:
- for fixed , the map is a probability measure on ;
- for fixed , the map is measurable.
For a Markov chain ,
The kernel acts on observables by conditional expectation,
and on laws by
These are not two unrelated conventions. They are dual:
Thus pulls a future observable back to the present, whereas pushes the present law forward. If the initial law is , then the law of is .
The Markov properties visible at operator level are
If is invariant, Jensen’s inequality also gives contraction on for :
These positivity and normalization properties are fundamental; symmetry is not. A Markov operator need not be self-adjoint.
Chapman–Kolmogorov composition and semigroups
Section titled “Chapman–Kolmogorov composition and semigroups”Define the composition of two kernels by
Conditional expectation and the Markov property give the Chapman–Kolmogorov law. For a time-homogeneous chain,
Indeed, on observables,
For a continuous-time, time-homogeneous Markov process, write
Then
This is a one-parameter Markov semigroup. The word semigroup records composition for nonnegative times; it does not assert that is invertible. Random evolution generally loses information, so negative-time operators need not exist.
Discrete and continuous notation should remain distinct. means applications of a one-step kernel. is indexed by a continuous time parameter and needs regularity in before it has an infinitesimal generator.
The generator includes its domain
Section titled “The generator includes its domain”Choose a Banach space on which every acts and suppose the semigroup is strongly continuous:
For a Feller process on a locally compact state space, a common choice is , the continuous functions that vanish at infinity. Other problems require other spaces. The generator is the generally unbounded operator
For , semigroup theory gives
The shorthand is useful only with this semigroup meaning. It does not license a term-by-term power series for an arbitrary unbounded operator.
For a finite-state continuous-time chain, the generator is a matrix with for and , and is an ordinary matrix exponential. For a diffusion, the same abstract definition can lead on a suitable core: a manageable test-function subspace on which the operator determines its closure. On that core the differential expression can be
Closing the operator means adjoining limits of test functions and their images in the chosen function space. Boundary conditions affect that closure, so they help determine which generator—and therefore which Markov evolution—is meant.
The forward law satisfies, weakly for ,
Writing this as is an adjoint notation. A density-level partial differential equation requires additional regularity and justified integrations by parts.
Invariance, stationarity, and reversibility
Section titled “Invariance, stationarity, and reversibility”A probability measure is invariant for a one-step kernel when
or for a continuous semigroup when for every . Equivalently,
for every bounded measurable . If , invariance makes the whole time-indexed process stationary. If , the dynamics can have an invariant law without the realized process being stationary.
For a continuous semigroup, differentiating the invariant identity at gives
This is the weak content of . The converse is not a purely formal calculation: a candidate solution of must be a normalized probability measure, must lie in the appropriate adjoint setting, and must determine an invariant law for the well-posed semigroup.
Detailed balance is the stronger symmetry condition
Integrating over proves invariance. Moreover, for ,
so is self-adjoint on . Such a chain is reversible: a stationary path has the same finite-dimensional law when time is reversed. The converse relationship is understood in this stationary setting. Neither detailed balance nor self-adjointness is part of the definition of a Markov process. Roberts and Rosenthal state detailed balance and its implication for stationarity on Roberts and Rosenthal 2004, pp. 4–5, PDF.
Long-time statements are logically distinct
Section titled “Long-time statements are logically distinct”Several claims that are colloquially called “ergodic” require different hypotheses:
- Existence and uniqueness of an invariant law identify a possible equilibrium but do not say that a given initial law approaches it.
- Convergence to equilibrium asserts, in a named mode such as total variation, that or approaches .
- An ergodic theorem controls time averages along one trajectory.
- Mixing quantifies the loss of dependence between separated times; a geometric rate is stronger information than convergence without a rate.
- A central limit theorem controls fluctuations of a centered time average on the scale.
A finite-state theorem gives a clean baseline. If a finite Markov chain is irreducible, then it has a unique invariant law and, from any initial state,
If it is also aperiodic, then in total variation; in the finite setting the convergence is geometric. Under these irreducible and aperiodic finite-state hypotheses, additive observables also satisfy a possibly degenerate central limit theorem.
A two-state counterexample isolates the role of aperiodicity:
The chain is irreducible and is its unique invariant law. Time averages still converge to , but the one-time law alternates between the two states and has no limit from a point mass. Uniqueness therefore does not imply convergence of .
Outside finite state spaces, the theorem must name the recurrence and irreducibility conditions. For an irreducible countable-state chain, existence of an invariant probability is tied to positive recurrence; the simple random walk on is recurrent but null recurrent and has no invariant probability. Norris proves this equivalence in Norris 1997, § 1.7, PDF pp. 4–5. On general state spaces, -irreducibility, positive Harris recurrence, aperiodicity, and the exceptional set of starting states all matter. Roberts and Rosenthal separate these issues in Roberts and Rosenthal 2004, § 3.2, pp. 14–18, PDF.
A safe route to a Markov-chain CLT
Section titled “A safe route to a Markov-chain CLT”Let be a stationary chain with invariant law , and let be real-valued. Center the observable:
Stationarity and an ergodic theorem can give , but they do not by themselves imply that has a Gaussian limit. One standard sufficient route on a general state space is geometric ergodicity: for some and finite ,
Together with this rate, require
For a reversible geometrically ergodic chain, the condition is enough. These are sufficient conditions, not definitions and not necessary conditions. Roberts and Rosenthal give the hypotheses, counterexamples, and extensions beyond stationary starts in Roberts and Rosenthal 2004, § 5, pp. 42–47, PDF.
The Poisson equation exposes why the fluctuation problem is an operator problem:
Suppose a suitable exists and the stationary chain is ergodic. Define
Then , and the Poisson equation gives the martingale decomposition
The boundary term divided by vanishes in probability. Under the corresponding stationary-ergodic martingale CLT,
where
This derivation does not require reversibility. The real work is proving that the Poisson equation has an adequate solution and that the martingale CLT hypotheses hold; stationarity alone supplies neither.
Covariance determines the sample-mean variance
Section titled “Covariance determines the sample-mean variance”For the stationary scalar observable, define
The variance of
has the exact finite- identity
It follows by grouping the covariances in the double sum according to their lag. No CLT is needed for this identity.
If
then
Absolute covariance summability gives this variance limit. A Gaussian limit still requires a valid Markov-chain CLT. When both statements hold,
If the series converges in the required sense, it solves the Poisson equation and
linking the operator and covariance descriptions.
Integrated autocorrelation time and effective sample size
Section titled “Integrated autocorrelation time and effective sample size”Assume , and assume the correlation series below converges. Define
This page uses the convention standard in Wolff’s lattice-error analysis:
Here is the asymptotic variance ratio relative to independent samples from . It may be larger or smaller than one. Consequently,
When , the effective sample size is the variance-equivalent quantity
If , the asymptotic variance on the scale vanishes. It is clearer to report that degeneracy than to describe the formal value as a literal sample count.
Some authors call , rather than , the integrated autocorrelation time. A numerical value without its defining convention is therefore ambiguous by a factor of two.
Neither nor is generally a property of the chain alone. Different observables couple to different slow modes. Negative or oscillating correlations can make and hence . This does not create extra stored configurations: only compares the asymptotic variance of this estimator with the variance of an iid sample mean. Wolff derives the covariance sum, this autocorrelation convention, and the effective-count interpretation in Wolff 2004/2006, § 2, pp. 4–6.
Controlled example: a rotating regulated Gaussian mode
Section titled “Controlled example: a rotating regulated Gaussian mode”Consider two real components of a regulated Gaussian field mode with equal stiffness . Introduce
and the antisymmetric matrix
A continuous-time Ornstein–Uhlenbeck evolution with an additional rotational drift is
Its observable generator, on a suitable core of smooth functions, is
The calculation can be read entirely at kernel level. With ,
and
The relations
verify the semigroup law for the Gaussian transition kernels. The centered Gaussian law
is invariant because covariance propagation gives
In density language, its stationary probability current is
This current is nonzero when , but it is divergence-free because and . Thus the Gaussian law is invariant even though the continuous evolution is nonreversible. The antisymmetric part of the drift preserves the density while circulating probability around its level sets.
The sampled chain
Section titled “The sampled chain”Sample at spacing and set
The sampled process obeys the exact recursion
with . It has the same invariant law . For its one-step matrix , the stationary pair is jointly Gaussian. Detailed balance holds exactly when its cross-covariance is symmetric, equivalently when
Thus gives the usual reversible relaxation, while generic gives a nonreversible chain with the same invariant Gaussian. At the sampled kernel is again reversible and has alternating correlations. This sampling-time symmetry does not make the underlying continuous evolution reversible when .
A linear observable
Section titled “A linear observable”For , stationarity gives
Taking the real part of a geometric series yields
Therefore
Three checks show what the formula means:
-
At ,
so slow positive correlation reduces the effective sample size.
-
At ,
because successive values are anticorrelated.
-
In the dense-sampling limit,
The chain is Gaussian, so is Gaussian for every finite . The exact variance identity above and the convergent geometric series imply
Here the CLT can be verified directly; it is not being inferred from stationarity alone.
A radial observable sees a different time scale
Section titled “A radial observable sees a different time scale”For
the Gaussian fourth-moment identity gives
The rotation angle has disappeared. The same Markov chain therefore gives a rotation-sensitive autocorrelation time for and a different, rotation-insensitive autocorrelation time for . This is why one cannot attach a single effective sample size to a stream of configurations without naming the observable.
The regulator condition is essential. At , diverges and the displayed stationary Gaussian ceases to be a probability law. The calculation is a finite-dimensional benchmark, not a construction of an interacting continuum field measure or a production lattice algorithm.
Estimating error from finite correlated output
Section titled “Estimating error from finite correlated output”Suppose are equilibrated observations of one scalar observable. A sample autocovariance convention is
A truncated estimator of the asymptotic variance is
When this finite-window estimate is nonnegative, the corresponding Monte Carlo standard-error estimate is
A raw rectangular covariance sum need not be nonnegative at finite . A negative value is a warning that this window and estimator do not provide a usable error estimate; it must not be hidden by taking an absolute value.
The cutoff is not cosmetic. A small window misses a slow positive tail and biases the uncertainty downward; a large window accumulates noisy sample autocovariances. Consistency requires dependence and moment hypotheses together with an increasing window that remains small relative to , not a fixed universal cutoff. Wolff analyzes this bias–noise balance, automatic windows, and the uncertainty of the error estimate in Wolff 2004/2006, §§ 3.1–3.3, pp. 7–13. Those procedures assume an equilibrated run, a resolvable decay scale, and a run long compared with that scale; an apparent plateau is evidence to inspect, not a theorem.
Three error mechanisms must remain separate:
- an unequilibrated initial distribution produces transient bias;
- serial dependence changes the variance of an otherwise stationary mean;
- a finite run makes the estimated autocorrelation tail itself uncertain.
Discarding a prefix addresses only the first mechanism, and an arbitrary “burn-in” does not prove equilibration. Using addresses none of the off-diagonal covariances.
Window selection, replicas, long-chain checks, observable-specific slow tails, nonlinear estimators, and production lattice diagnostics require a larger workflow. They belong to Autocorrelation Times and Effective Sample Size.
Common pitfalls
Section titled “Common pitfalls”Treating invariance as convergence. The equation says what happens if the chain already has law . It does not show that another initial law approaches ; irreducibility, recurrence, periodicity, and a chosen convergence mode still matter.
Defining a Markov process by detailed balance. Detailed balance is a useful sufficient condition for invariance and makes self-adjoint in . Nonreversible kernels and generators remain fully Markovian and can be invariant, ergodic, and asymptotically normal.
Using a formal differential operator as a complete generator. The same differential expression can represent different boundary behavior. The function space, domain, and closed semigroup determine the process.
Promoting an ergodic theorem to a CLT. Convergence of does not determine the fluctuation law. A Markov-chain CLT needs its own mixing, moment, spectral, martingale, or Poisson-equation hypotheses.
Using the iid error after observing correlation. The exact variance contains every lag covariance. Positive tails usually enlarge error bars; negative or oscillating correlations can reduce them.
Quoting an autocorrelation time without a convention or observable. The factor-of-two convention varies across fields, and different observables couple to different dynamical modes. State both the defining sum and the observable.
Confusing transient bias with serial-correlation error. Removing early states cannot remove covariance among the retained states. Equilibration and uncertainty require distinct evidence.
Applying probability language to an unconstructed field measure. A formal functional weight is not automatically a normalized probability law. Regulator removal and continuum existence require separate analysis.
Check your understanding
Section titled “Check your understanding”1. Recall the two kernel actions
Section titled “1. Recall the two kernel actions”State how a one-step kernel acts on an observable and on a probability law . Then state the identity that relates the two actions.
Solution
The observable action is conditional expectation,
and the forward action on a law is
They are dual:
Thus pulls a future observable back to the present, whereas pushes the present law forward.
2. Separate time averages from marginal convergence
Section titled “2. Separate time averages from marginal convergence”For
find the invariant law and compare the large- time average with the large- one-time law from state .
Solution
Solving gives . Starting at , the path is , so for any ,
However, the one-time law is at even times and at odd times, so it has no limit. Finite-state irreducibility supplies the time-average result here; aperiodicity is the missing condition for ordinary marginal convergence.
3. Solve the Poisson equation for an eigen-observable
Section titled “3. Solve the Poisson equation for an eigen-observable”Suppose , where , , and . Find a Poisson solution and compute .
Solution
Since
a solution is
Stationarity gives
Therefore
For , correlations inflate the variance. For , alternating correlations give and hence an effective sample size larger than .
4. Check the rotating Gaussian benchmark
Section titled “4. Check the rotating Gaussian benchmark”For the sampled Gaussian chain, sum and recover . Then compare the linear observable with the radial observable .
Solution
Using a complex geometric series,
Hence
For the radial observable, the stationary Gaussian fourth-moment identity gives and therefore
Thus the rotation can shorten or oscillate the linear-observable correlations while leaving the radial autocorrelation factor unchanged. The effective sample size is observable-specific.
Synthesis and continuation
Section titled “Synthesis and continuation”A Markov kernel propagates both observables and laws, and the Chapman–Kolmogorov identity organizes repeated evolution into a semigroup. A continuous-time generator is the derivative of that semigroup on a declared domain. Invariance fixes a law; stationary initialization realizes it; reversibility adds a symmetry; and ergodic, mixing, and CLT statements each require further hypotheses.
For a stationary scalar observable, the exact sample-mean variance contains all lag covariances. When covariance summability and a valid Markov-chain CLT hold, their infinite sum is the asymptotic variance. Integrated autocorrelation time and effective sample size repackage that variance, but only after the convention and observable are named. The rotating Gaussian mode shows explicitly that invariance does not require reversibility and that two observables of one chain can have different error inflation.
For production lattice analysis—including window and tail choices, replicas, long-chain checks, and observable-specific diagnostics—continue to Autocorrelation Times and Effective Sample Size.
References
Section titled “References”-
J. R. Norris, Markov Chains, § 1.7: Invariant Distributions, PDF, Cambridge University Press, 1997. The linked § 1.7 PDF, pp. 4–5, supports the relation between invariant probabilities and positive recurrence for irreducible countable-state chains, including the null-recurrent random walk contrast.
-
Grigorios A. Pavliotis, Stochastic Processes and Applications, PDF, Springer, 2014; linked author manuscript dated November 11, 2015. §§ 2.2–2.4, pp. 30–39, support transition functions, the Chapman–Kolmogorov identity, Markov semigroups and generators, forward law evolution, and invariant measures.
-
Gareth O. Roberts and Jeffrey S. Rosenthal, “General State Space Markov Chains and MCMC Algorithms”, PDF, Probability Surveys 1 (2004), 20–71, with the linked author version corrected through 2023. pp. 4–5 support detailed balance; § 3.2, pp. 14–18, supports the distinctions among invariance, irreducibility, aperiodicity, convergence, and laws of large numbers; and § 5, pp. 42–47, supports the Markov-chain CLT conditions, covariance-sum variance, and Poisson-equation route.
-
Ulli Wolff, “Monte Carlo Errors with Less Errors”, Computer Physics Communications 156 (2004), 143–153; arXiv:hep-lat/0306017v4, 2006 revision. § 2, pp. 4–6, supports the covariance-sum error, the convention , and ; §§ 3.1–3.3, pp. 7–13, support the finite-window bias–noise tradeoff and error-estimation cautions.