Skip to content

Markov Generators, Semigroups, Ergodicity, and Correlated-Sample Error

A Markov transition kernel answers a local question: given the present state, what is the probability law of the next state? Acting on observables, the kernel takes a conditional expectation; acting on probability measures, it propagates a law. Repeated kernels form a discrete semigroup PnP^n, while a time-homogeneous continuous process gives a one-parameter semigroup PtP_t. On a function space where PtP_t is strongly continuous, its infinitesimal generator is

Lf=limt0Ptfft,fD(L).Lf=\lim_{t\downarrow0}\frac{P_tf-f}{t}, \qquad f\in D(L).

The domain D(L)D(L) is part of this definition. A formal differential expression without a function space, boundary conditions, and a domain is not yet a generator.

An invariant probability measure is a fixed point of the forward evolution. It does not by itself imply convergence to equilibrium, an ergodic theorem, a mixing rate, or a central limit theorem. Once suitable hypotheses do give a stationary Markov-chain CLT, serial covariance changes the uncertainty of the sample mean from the iid expression. With the convention used here,

τint,f=12+k=1ρf(k),Neff,f=N2τint,f.\tau_{\mathrm{int},f} =\frac12+\sum_{k=1}^{\infty}\rho_f(k), \qquad N_{\mathrm{eff},f} =\frac{N}{2\tau_{\mathrm{int},f}}.

Both quantities depend on the observable ff. The page develops these statements from kernels through uncertainty and then checks them in a regulated Gaussian field mode whose stationary evolution can be reversible or nonreversible.

Required background. Stochastic Processes and Correlation Functions for conditional laws, stationarity, covariance, and correlation functions; Probabilistic Convergence, Laws of Large Numbers, and Central Limit Theorems for modes of convergence, iid limit theorems, and the distinction between a limit law and a finite-sample bound.

Markov kernels, semigroups, and sampling scope

Section titled “Markov kernels, semigroups, and sampling scope”

Work first on a measurable state space (E,E)(E,\mathcal E). Bounded measurable observables are denoted by ff, probability laws by μ\mu, and the expectation of ff under an invariant law π\pi by

π(f)=Ef(x)π(dx).\pi(f)=\int_E f(x)\,\pi(\mathrm dx).

The operator convention is backward on observables: Ptf(x)P_tf(x) is the conditional expectation of the future observable given the present state xx. The corresponding forward action propagates measures. This is the same convention used for the local diffusion operator on Brownian Motion, Stochastic Calculus, Langevin Equations, and Fokker–Planck Dynamics, but no stochastic-calculus result from that page is needed below.

The probability statements apply to genuine, normalized, nonnegative measures. A regulated Euclidean or thermal field mode can supply such a measure. A formal continuum functional integral, a sign-changing weight, or the oscillatory factor eiSe^{iS} does not automatically do so; Markov-chain probability theorems cannot be transferred to those expressions without a separate construction.

Pavliotis develops transition functions, Markov semigroups, generators, and invariant measures in Pavliotis 2014, §§ 2.2–2.4, pp. 30–39, PDF. The discussion below makes the operator domains and the additional hypotheses behind long-time conclusions explicit.

A time-homogeneous one-step Markov kernel PP satisfies two requirements:

  1. for fixed xEx\in E, the map AP(x,A)A\mapsto P(x,A) is a probability measure on (E,E)(E,\mathcal E);
  2. for fixed AEA\in\mathcal E, the map xP(x,A)x\mapsto P(x,A) is measurable.

For a Markov chain (Xn)n0(X_n)_{n\geq0},

P(x,A)=P(Xn+1AXn=x).P(x,A)=\mathbb P(X_{n+1}\in A\mid X_n=x).

The kernel acts on observables by conditional expectation,

(Pf)(x)=Ef(y)P(x,dy)=Ex[f(X1)],(Pf)(x) =\int_E f(y)P(x,\mathrm dy) =\mathbb E_x[f(X_1)],

and on laws by

(μP)(A)=EP(x,A)μ(dx).(\mu P)(A) =\int_E P(x,A)\,\mu(\mathrm dx).

These are not two unrelated conventions. They are dual:

E(Pf)(x)μ(dx)=Ef(y)(μP)(dy).\int_E(Pf)(x)\,\mu(\mathrm dx) =\int_E f(y)\,(\mu P)(\mathrm dy).

Thus PP pulls a future observable back to the present, whereas μP\mu P pushes the present law forward. If the initial law is μ\mu, then the law of XnX_n is μPn\mu P^n.

The Markov properties visible at operator level are

f0Pf0,P1=1,Pff.f\geq0\Longrightarrow Pf\geq0, \qquad P\mathbf1=\mathbf1, \qquad \lVert Pf\rVert_\infty\leq\lVert f\rVert_\infty.

If π\pi is invariant, Jensen’s inequality also gives contraction on Lp(π)L^p(\pi) for 1p<1\leq p<\infty:

PfLp(π)pEP(fp)dπ=fLp(π)p.\lVert Pf\rVert_{L^p(\pi)}^p \leq\int_E P(|f|^p)\,\mathrm d\pi =\lVert f\rVert_{L^p(\pi)}^p.

These positivity and normalization properties are fundamental; symmetry is not. A Markov operator need not be self-adjoint.

Chapman–Kolmogorov composition and semigroups

Section titled “Chapman–Kolmogorov composition and semigroups”

Define the composition of two kernels by

(PQ)(x,A)=EQ(y,A)P(x,dy).(PQ)(x,A)=\int_E Q(y,A)P(x,\mathrm dy).

Conditional expectation and the Markov property give the Chapman–Kolmogorov law. For a time-homogeneous chain,

Pn+m=PnPm,P0=I.P^{n+m}=P^nP^m, \qquad P^0=I.

Indeed, on observables,

(Pn+mf)(x)=Ex[f(Xn+m)]=Ex ⁣[E[f(Xn+m)Fn]]=Ex[(Pmf)(Xn)]=(PnPmf)(x).\begin{aligned} (P^{n+m}f)(x) &=\mathbb E_x[f(X_{n+m})]\\ &=\mathbb E_x\!\left[ \mathbb E[f(X_{n+m})\mid\mathcal F_n] \right]\\ &=\mathbb E_x[(P^mf)(X_n)] =(P^nP^mf)(x). \end{aligned}

For a continuous-time, time-homogeneous Markov process, write

(Ptf)(x)=Ex[f(Xt)].(P_tf)(x)=\mathbb E_x[f(X_t)].

Then

P0=I,Ps+t=PsPt,s,t0.P_0=I, \qquad P_{s+t}=P_sP_t, \qquad s,t\geq0.

This is a one-parameter Markov semigroup. The word semigroup records composition for nonnegative times; it does not assert that PtP_t is invertible. Random evolution generally loses information, so negative-time operators need not exist.

Discrete and continuous notation should remain distinct. PnP^n means nn applications of a one-step kernel. PtP_t is indexed by a continuous time parameter and needs regularity in tt before it has an infinitesimal generator.

Choose a Banach space B\mathcal B on which every PtP_t acts and suppose the semigroup is strongly continuous:

limt0PtffB=0,fB.\lim_{t\downarrow0}\lVert P_tf-f\rVert_{\mathcal B}=0, \qquad f\in\mathcal B.

For a Feller process on a locally compact state space, a common choice is B=C0(E)\mathcal B=C_0(E), the continuous functions that vanish at infinity. Other problems require other spaces. The generator is the generally unbounded operator

D(L)={fB:limt0Ptfft exists in B},Lf=limt0Ptfft,fD(L).\begin{aligned} D(L) &=\left\{ f\in\mathcal B: \lim_{t\downarrow0}\frac{P_tf-f}{t} \text{ exists in }\mathcal B \right\},\\ Lf &=\lim_{t\downarrow0}\frac{P_tf-f}{t}, \qquad f\in D(L). \end{aligned}

For fD(L)f\in D(L), semigroup theory gives

ddtPtf=PtLf=LPtf.\frac{\mathrm d}{\mathrm dt}P_tf =P_tLf =LP_tf.

The shorthand Pt=etLP_t=e^{tL} is useful only with this semigroup meaning. It does not license a term-by-term power series for an arbitrary unbounded operator.

For a finite-state continuous-time chain, the generator is a matrix QQ with qij0q_{ij}\geq0 for iji\ne j and jqij=0\sum_jq_{ij}=0, and Pt=etQP_t=e^{tQ} is an ordinary matrix exponential. For a diffusion, the same abstract definition can lead on a suitable core: a manageable test-function subspace on which the operator determines its closure. On that core the differential expression can be

Lf=biif+12aijijf.Lf=b^i\partial_i f+\frac12a^{ij}\partial_i\partial_jf.

Closing the operator means adjoining limits of test functions and their images in the chosen function space. Boundary conditions affect that closure, so they help determine which generator—and therefore which Markov evolution—is meant.

The forward law μt=μ0Pt\mu_t=\mu_0P_t satisfies, weakly for fD(L)f\in D(L),

ddtEfdμt=ELfdμt.\frac{\mathrm d}{\mathrm dt}\int_Ef\,\mathrm d\mu_t =\int_E Lf\,\mathrm d\mu_t.

Writing this as tμt=Lμt\partial_t\mu_t=L^\dagger\mu_t is an adjoint notation. A density-level partial differential equation requires additional regularity and justified integrations by parts.

Invariance, stationarity, and reversibility

Section titled “Invariance, stationarity, and reversibility”

A probability measure π\pi is invariant for a one-step kernel when

πP=π,\pi P=\pi,

or for a continuous semigroup when πPt=π\pi P_t=\pi for every t0t\geq0. Equivalently,

EPfdπ=Efdπ\int_E Pf\,\mathrm d\pi=\int_Ef\,\mathrm d\pi

for every bounded measurable ff. If X0πX_0\sim\pi, invariance makes the whole time-indexed process stationary. If X0≁πX_0\not\sim\pi, the dynamics can have an invariant law without the realized process being stationary.

For a continuous semigroup, differentiating the invariant identity at t=0t=0 gives

ELfdπ=0,fD(L).\int_E Lf\,\mathrm d\pi=0, \qquad f\in D(L).

This is the weak content of Lπ=0L^\dagger\pi=0. The converse is not a purely formal calculation: a candidate solution of Lπ=0L^\dagger\pi=0 must be a normalized probability measure, must lie in the appropriate adjoint setting, and must determine an invariant law for the well-posed semigroup.

Detailed balance is the stronger symmetry condition

π(dx)P(x,dy)=π(dy)P(y,dx).\pi(\mathrm dx)P(x,\mathrm dy) =\pi(\mathrm dy)P(y,\mathrm dx).

Integrating over xx proves invariance. Moreover, for f,gL2(π)f,g\in L^2(\pi),

f,Pgπ=Pf,gπ,\langle f,Pg\rangle_\pi =\langle Pf,g\rangle_\pi,

so PP is self-adjoint on L2(π)L^2(\pi). Such a chain is reversible: a stationary path has the same finite-dimensional law when time is reversed. The converse relationship is understood in this stationary L2(π)L^2(\pi) setting. Neither detailed balance nor self-adjointness is part of the definition of a Markov process. Roberts and Rosenthal state detailed balance and its implication for stationarity on Roberts and Rosenthal 2004, pp. 4–5, PDF.

Long-time statements are logically distinct

Section titled “Long-time statements are logically distinct”

Several claims that are colloquially called “ergodic” require different hypotheses:

  • Existence and uniqueness of an invariant law identify a possible equilibrium but do not say that a given initial law approaches it.
  • Convergence to equilibrium asserts, in a named mode such as total variation, that Pn(x,)P^n(x,\cdot) or Pt(x,)P_t(x,\cdot) approaches π\pi.
  • An ergodic theorem controls time averages along one trajectory.
  • Mixing quantifies the loss of dependence between separated times; a geometric rate is stronger information than convergence without a rate.
  • A central limit theorem controls fluctuations of a centered time average on the N1/2N^{-1/2} scale.

A finite-state theorem gives a clean baseline. If a finite Markov chain is irreducible, then it has a unique invariant law and, from any initial state,

1Nn=0N1f(Xn)π(f)almost surely.\frac1N\sum_{n=0}^{N-1}f(X_n) \longrightarrow\pi(f) \qquad\text{almost surely}.

If it is also aperiodic, then Pn(x,)πP^n(x,\cdot)\to\pi in total variation; in the finite setting the convergence is geometric. Under these irreducible and aperiodic finite-state hypotheses, additive observables also satisfy a possibly degenerate central limit theorem.

A two-state counterexample isolates the role of aperiodicity:

P=(0110),π=(12,12).P= \begin{pmatrix} 0&1\\ 1&0 \end{pmatrix}, \qquad \pi=\left(\frac12,\frac12\right).

The chain is irreducible and π\pi is its unique invariant law. Time averages still converge to π(f)\pi(f), but the one-time law alternates between the two states and has no limit from a point mass. Uniqueness therefore does not imply convergence of Pn(x,)P^n(x,\cdot).

Outside finite state spaces, the theorem must name the recurrence and irreducibility conditions. For an irreducible countable-state chain, existence of an invariant probability is tied to positive recurrence; the simple random walk on Z\mathbb Z is recurrent but null recurrent and has no invariant probability. Norris proves this equivalence in Norris 1997, § 1.7, PDF pp. 4–5. On general state spaces, φ\varphi-irreducibility, positive Harris recurrence, aperiodicity, and the exceptional set of starting states all matter. Roberts and Rosenthal separate these issues in Roberts and Rosenthal 2004, § 3.2, pp. 14–18, PDF.

Let (Xn)(X_n) be a stationary chain with invariant law π\pi, and let fL2(π)f\in L^2(\pi) be real-valued. Center the observable:

f~=fπ(f),SN=n=0N1f~(Xn).\widetilde f=f-\pi(f), \qquad S_N=\sum_{n=0}^{N-1}\widetilde f(X_n).

Stationarity and an ergodic theorem can give SN/N0S_N/N\to0, but they do not by themselves imply that SN/NS_N/\sqrt N has a Gaussian limit. One standard sufficient route on a general state space is geometric ergodicity: for some 0<ρ<10<\rho<1 and finite M(x)M(x),

Pn(x,)πTVM(x)ρn.\lVert P^n(x,\cdot)-\pi\rVert_{\mathrm{TV}} \leq M(x)\rho^n.

Together with this rate, require

π(f2+δ)<for some δ>0.\pi(|f|^{2+\delta})<\infty \qquad\text{for some }\delta>0.

For a reversible geometrically ergodic chain, the L2(π)L^2(\pi) condition is enough. These are sufficient conditions, not definitions and not necessary conditions. Roberts and Rosenthal give the hypotheses, counterexamples, and extensions beyond stationary starts in Roberts and Rosenthal 2004, § 5, pp. 42–47, PDF.

The Poisson equation exposes why the fluctuation problem is an operator problem:

(IP)h=f~.(I-P)h=\widetilde f.

Suppose a suitable hL2(π)h\in L^2(\pi) exists and the stationary chain is ergodic. Define

Dn+1=h(Xn+1)(Ph)(Xn).D_{n+1}=h(X_{n+1})-(Ph)(X_n).

Then E[Dn+1Fn]=0\mathbb E[D_{n+1}\mid\mathcal F_n]=0, and the Poisson equation gives the martingale decomposition

SN=h(X0)h(XN)+n=0N1Dn+1.S_N =h(X_0)-h(X_N) +\sum_{n=0}^{N-1}D_{n+1}.

The boundary term divided by N\sqrt N vanishes in probability. Under the corresponding stationary-ergodic martingale CLT,

SNNN(0,σas2),\frac{S_N}{\sqrt N} \Longrightarrow \mathcal N(0,\sigma_{\mathrm{as}}^2),

where

σas2=Eπ[D12]=π(h2)π((Ph)2)=2f~,hπf~,f~π.\begin{aligned} \sigma_{\mathrm{as}}^2 &=\mathbb E_\pi[D_1^2]\\ &=\pi(h^2)-\pi((Ph)^2)\\ &=2\langle \widetilde f,h\rangle_\pi -\langle \widetilde f,\widetilde f\rangle_\pi. \end{aligned}

This derivation does not require reversibility. The real work is proving that the Poisson equation has an adequate solution and that the martingale CLT hypotheses hold; stationarity alone supplies neither.

Covariance determines the sample-mean variance

Section titled “Covariance determines the sample-mean variance”

For the stationary scalar observable, define

γf(k)=Covπ(f(X0),f(Xk))=f~,Pkf~π,k0.\gamma_f(k) =\operatorname{Cov}_\pi(f(X_0),f(X_k)) =\langle \widetilde f,P^k\widetilde f\rangle_\pi, \qquad k\geq0.

The variance of

fN=1Nn=0N1f(Xn)\overline f_N=\frac1N\sum_{n=0}^{N-1}f(X_n)

has the exact finite-NN identity

Varπ(fN)=1N[γf(0)+2k=1N1(1kN)γf(k)].\operatorname{Var}_\pi(\overline f_N) =\frac1N\left[ \gamma_f(0) +2\sum_{k=1}^{N-1} \left(1-\frac{k}{N}\right)\gamma_f(k) \right].

It follows by grouping the N2N^2 covariances in the double sum according to their lag. No CLT is needed for this identity.

If

k=1γf(k)<,\sum_{k=1}^{\infty}|\gamma_f(k)|<\infty,

then

NVarπ(fN)σas2=γf(0)+2k=1γf(k).N\operatorname{Var}_\pi(\overline f_N) \longrightarrow \sigma_{\mathrm{as}}^2 =\gamma_f(0)+2\sum_{k=1}^{\infty}\gamma_f(k).

Absolute covariance summability gives this variance limit. A Gaussian limit still requires a valid Markov-chain CLT. When both statements hold,

N[fNπ(f)]N(0,σas2).\sqrt N\,[\overline f_N-\pi(f)] \Longrightarrow \mathcal N(0,\sigma_{\mathrm{as}}^2).

If the series h=k0Pkf~h=\sum_{k\geq0}P^k\widetilde f converges in the required sense, it solves the Poisson equation and

2f~,hπf~,f~π=γf(0)+2k=1γf(k),2\langle \widetilde f,h\rangle_\pi -\langle \widetilde f,\widetilde f\rangle_\pi =\gamma_f(0)+2\sum_{k=1}^{\infty}\gamma_f(k),

linking the operator and covariance descriptions.

Integrated autocorrelation time and effective sample size

Section titled “Integrated autocorrelation time and effective sample size”

Assume γf(0)>0\gamma_f(0)>0, and assume the correlation series below converges. Define

ρf(k)=γf(k)γf(0).\rho_f(k)=\frac{\gamma_f(k)}{\gamma_f(0)}.

This page uses the convention standard in Wolff’s lattice-error analysis:

τint,f=12+k=1ρf(k),Cf=1+2k=1ρf(k)=2τint,f.\begin{aligned} \tau_{\mathrm{int},f} &=\frac12+\sum_{k=1}^{\infty}\rho_f(k),\\ \mathcal C_f &=1+2\sum_{k=1}^{\infty}\rho_f(k) =2\tau_{\mathrm{int},f}. \end{aligned}

Here Cf\mathcal C_f is the asymptotic variance ratio relative to independent samples from π\pi. It may be larger or smaller than one. Consequently,

σas2=γf(0)Cf=2γf(0)τint,f,\sigma_{\mathrm{as}}^2 =\gamma_f(0)\mathcal C_f =2\gamma_f(0)\tau_{\mathrm{int},f},

When Cf>0\mathcal C_f>0, the effective sample size is the variance-equivalent quantity

Neff,f=NCf=N2τint,f.N_{\mathrm{eff},f} =\frac{N}{\mathcal C_f} =\frac{N}{2\tau_{\mathrm{int},f}}.

If Cf=0\mathcal C_f=0, the asymptotic variance on the N1N^{-1} scale vanishes. It is clearer to report that degeneracy than to describe the formal value Neff,f=N_{\mathrm{eff},f}=\infty as a literal sample count.

Some authors call Cf\mathcal C_f, rather than τint,f\tau_{\mathrm{int},f}, the integrated autocorrelation time. A numerical value without its defining convention is therefore ambiguous by a factor of two.

Neither τint,f\tau_{\mathrm{int},f} nor Neff,fN_{\mathrm{eff},f} is generally a property of the chain alone. Different observables couple to different slow modes. Negative or oscillating correlations can make Cf<1\mathcal C_f<1 and hence Neff,f>NN_{\mathrm{eff},f}>N. This does not create extra stored configurations: Neff,fN_{\mathrm{eff},f} only compares the asymptotic variance of this estimator with the variance of an iid sample mean. Wolff derives the covariance sum, this autocorrelation convention, and the effective-count interpretation in Wolff 2004/2006, § 2, pp. 4–6.

Controlled example: a rotating regulated Gaussian mode

Section titled “Controlled example: a rotating regulated Gaussian mode”

Consider two real components of a regulated Gaussian field mode with equal stiffness κq>0\kappa_{\mathbf q}>0. Introduce

λ=Γκq>0,D=ΓT,v=Dλ=Tκq,\lambda=\Gamma\kappa_{\mathbf q}>0, \qquad D=\Gamma T, \qquad v=\frac{D}{\lambda}=\frac{T}{\kappa_{\mathbf q}},

and the antisymmetric matrix

J=(0110).\mathsf J= \begin{pmatrix} 0&-1\\ 1&0 \end{pmatrix}.

A continuous-time Ornstein–Uhlenbeck evolution with an additional rotational drift is

dΦs=(λI+ωJ)Φsds+2DdWs.\mathrm d\Phi_s =(-\lambda I+\omega\mathsf J)\Phi_s\,\mathrm ds +\sqrt{2D}\,\mathrm dW_s.

Its observable generator, on a suitable core of smooth functions, is

Lf(ϕ)=(λϕ+ωJϕ)f(ϕ)+DΔf(ϕ).Lf(\phi) =(-\lambda\phi+\omega\mathsf J\phi)\mathbin{\cdot}\nabla f(\phi) +D\Delta f(\phi).

The calculation can be read entirely at kernel level. With Rα=eαJ\mathsf R_\alpha=e^{\alpha\mathsf J},

Mt=e(λI+ωJ)t=eλtRωt,Qt=v(1e2λt)I,M_t=e^{(-\lambda I+\omega\mathsf J)t} =e^{-\lambda t}\mathsf R_{\omega t}, \qquad Q_t=v(1-e^{-2\lambda t})I,

and

Pt(ϕ,dϕ)=N(Mtϕ,Qt)(dϕ).P_t(\phi,\mathrm d\phi') =\mathcal N(M_t\phi,Q_t)(\mathrm d\phi').

The relations

Ms+t=MsMt,Qs+t=MsQtMsT+QsM_{s+t}=M_sM_t, \qquad Q_{s+t}=M_sQ_tM_s^{\mathsf T}+Q_s

verify the semigroup law for the Gaussian transition kernels. The centered Gaussian law

π(dϕ)=12πvexp ⁣[ϕ22v]d2ϕ\pi(\mathrm d\phi) =\frac1{2\pi v} \exp\!\left[-\frac{|\phi|^2}{2v}\right]\mathrm d^2\phi

is invariant because covariance propagation gives

Mt(vI)MtT+Qt=vI.M_t(vI)M_t^{\mathsf T}+Q_t=vI.

In density language, its stationary probability current is

J(ϕ)=ωJϕπ(ϕ).\mathcal J_*(\phi) =\omega\mathsf J\phi\,\pi(\phi).

This current is nonzero when ω0\omega\ne0, but it is divergence-free because trJ=0\operatorname{tr}\mathsf J=0 and ϕJϕ=0\phi\mathbin{\cdot}\mathsf J\phi=0. Thus the Gaussian law is invariant even though the continuous evolution is nonreversible. The antisymmetric part of the drift preserves the density while circulating probability around its level sets.

Sample at spacing Δ>0\Delta>0 and set

r=eλΔ,θ=ωΔ.r=e^{-\lambda\Delta}, \qquad \theta=\omega\Delta.

The sampled process obeys the exact recursion

Φn+1=rRθΦn+v(1r2)ξn,ξniidN(0,I2),\Phi_{n+1} =r\mathsf R_\theta\Phi_n +\sqrt{v(1-r^2)}\,\xi_n, \qquad \xi_n\stackrel{\mathrm{iid}}{\sim}\mathcal N(0,I_2),

with 0<r<10<r<1. It has the same invariant law π\pi. For its one-step matrix A=rRθA=r\mathsf R_\theta, the stationary pair (Φn,Φn+1)(\Phi_n,\Phi_{n+1}) is jointly Gaussian. Detailed balance holds exactly when its cross-covariance is symmetric, equivalently when

A=ATsinθ=0.A=A^{\mathsf T} \quad\Longleftrightarrow\quad \sin\theta=0.

Thus θ=0\theta=0 gives the usual reversible relaxation, while generic θ\theta gives a nonreversible chain with the same invariant Gaussian. At θ=π\theta=\pi the sampled kernel is again reversible and has alternating correlations. This sampling-time symmetry does not make the underlying continuous evolution reversible when ω0\omega\ne0.

For f(Φ)=ϕ1f(\Phi)=\phi_1, stationarity gives

γf(k)=vrkcos(kθ),ρf(k)=rkcos(kθ).\gamma_f(k) =v r^k\cos(k\theta), \qquad \rho_f(k)=r^k\cos(k\theta).

Taking the real part of a geometric series yields

Cf=1+2k=1rkcos(kθ)=1r212rcosθ+r2.\begin{aligned} \mathcal C_f &=1+2\sum_{k=1}^{\infty}r^k\cos(k\theta)\\ &=\frac{1-r^2}{1-2r\cos\theta+r^2}. \end{aligned}

Therefore

τint,f=1r22(12rcosθ+r2),Neff,f=N12rcosθ+r21r2.\tau_{\mathrm{int},f} =\frac{1-r^2} {2(1-2r\cos\theta+r^2)}, \qquad N_{\mathrm{eff},f} =N\frac{1-2r\cos\theta+r^2}{1-r^2}.

Three checks show what the formula means:

  • At θ=0\theta=0,

    Cf=1+r1r,τint,f=12coth ⁣(λΔ2),\mathcal C_f=\frac{1+r}{1-r}, \qquad \tau_{\mathrm{int},f} =\frac12\coth\!\left(\frac{\lambda\Delta}{2}\right),

    so slow positive correlation reduces the effective sample size.

  • At θ=π\theta=\pi,

    Cf=1r1+r<1,Neff,f>N,\mathcal C_f=\frac{1-r}{1+r}<1, \qquad N_{\mathrm{eff},f}>N,

    because successive values are anticorrelated.

  • In the dense-sampling limit,

    Δτint,f0eλtcos(ωt)dt=λλ2+ω2.\Delta\tau_{\mathrm{int},f} \longrightarrow \int_0^\infty e^{-\lambda t}\cos(\omega t)\,\mathrm dt =\frac{\lambda}{\lambda^2+\omega^2}.

The chain is Gaussian, so fN\overline f_N is Gaussian for every finite NN. The exact variance identity above and the convergent geometric series imply

NfNN(0,vCf).\sqrt N\,\overline f_N \Longrightarrow \mathcal N(0,v\mathcal C_f).

Here the CLT can be verified directly; it is not being inferred from stationarity alone.

A radial observable sees a different time scale

Section titled “A radial observable sees a different time scale”

For

q(Φ)=Φ22v,q(\Phi)=|\Phi|^2-2v,

the Gaussian fourth-moment identity gives

ρq(k)=r2k,Cq=1+r21r2.\rho_q(k)=r^{2k}, \qquad \mathcal C_q=\frac{1+r^2}{1-r^2}.

The rotation angle has disappeared. The same Markov chain therefore gives a rotation-sensitive autocorrelation time for ϕ1\phi_1 and a different, rotation-insensitive autocorrelation time for Φ2|\Phi|^2. This is why one cannot attach a single effective sample size to a stream of configurations without naming the observable.

The regulator condition κq>0\kappa_{\mathbf q}>0 is essential. At κq=0\kappa_{\mathbf q}=0, v=T/κqv=T/\kappa_{\mathbf q} diverges and the displayed stationary Gaussian ceases to be a probability law. The calculation is a finite-dimensional benchmark, not a construction of an interacting continuum field measure or a production lattice algorithm.

Estimating error from finite correlated output

Section titled “Estimating error from finite correlated output”

Suppose f0,,fN1f_0,\ldots,f_{N-1} are equilibrated observations of one scalar observable. A sample autocovariance convention is

γ^N(k)=1Nn=0Nk1(fnfN)(fn+kfN),0k<N.\widehat\gamma_N(k) =\frac1N\sum_{n=0}^{N-k-1} (f_n-\overline f_N)(f_{n+k}-\overline f_N), \qquad 0\leq k<N.

A truncated estimator of the asymptotic variance is

σ^as2(M)=γ^N(0)+2k=1Mγ^N(k),\widehat\sigma_{\mathrm{as}}^2(M) =\widehat\gamma_N(0) +2\sum_{k=1}^{M}\widehat\gamma_N(k),

When this finite-window estimate is nonnegative, the corresponding Monte Carlo standard-error estimate is

MCSE^(fN)=σ^as2(M)N.\widehat{\operatorname{MCSE}}(\overline f_N) =\sqrt{\frac{\widehat\sigma_{\mathrm{as}}^2(M)}{N}}.

A raw rectangular covariance sum need not be nonnegative at finite NN. A negative value is a warning that this window and estimator do not provide a usable error estimate; it must not be hidden by taking an absolute value.

The cutoff MM is not cosmetic. A small window misses a slow positive tail and biases the uncertainty downward; a large window accumulates noisy sample autocovariances. Consistency requires dependence and moment hypotheses together with an increasing window that remains small relative to NN, not a fixed universal cutoff. Wolff analyzes this bias–noise balance, automatic windows, and the uncertainty of the error estimate in Wolff 2004/2006, §§ 3.1–3.3, pp. 7–13. Those procedures assume an equilibrated run, a resolvable decay scale, and a run long compared with that scale; an apparent plateau is evidence to inspect, not a theorem.

Three error mechanisms must remain separate:

  1. an unequilibrated initial distribution produces transient bias;
  2. serial dependence changes the variance of an otherwise stationary mean;
  3. a finite run makes the estimated autocorrelation tail itself uncertain.

Discarding a prefix addresses only the first mechanism, and an arbitrary “burn-in” does not prove equilibration. Using γ^N(0)/N\widehat\gamma_N(0)/N addresses none of the off-diagonal covariances.

Window selection, replicas, long-chain checks, observable-specific slow tails, nonlinear estimators, and production lattice diagnostics require a larger workflow. They belong to Autocorrelation Times and Effective Sample Size.

Treating invariance as convergence. The equation πP=π\pi P=\pi says what happens if the chain already has law π\pi. It does not show that another initial law approaches π\pi; irreducibility, recurrence, periodicity, and a chosen convergence mode still matter.

Defining a Markov process by detailed balance. Detailed balance is a useful sufficient condition for invariance and makes PP self-adjoint in L2(π)L^2(\pi). Nonreversible kernels and generators remain fully Markovian and can be invariant, ergodic, and asymptotically normal.

Using a formal differential operator as a complete generator. The same differential expression can represent different boundary behavior. The function space, domain, and closed semigroup determine the process.

Promoting an ergodic theorem to a CLT. Convergence of fN\overline f_N does not determine the N\sqrt N fluctuation law. A Markov-chain CLT needs its own mixing, moment, spectral, martingale, or Poisson-equation hypotheses.

Using the iid error after observing correlation. The exact variance contains every lag covariance. Positive tails usually enlarge error bars; negative or oscillating correlations can reduce them.

Quoting an autocorrelation time without a convention or observable. The factor-of-two convention varies across fields, and different observables couple to different dynamical modes. State both the defining sum and the observable.

Confusing transient bias with serial-correlation error. Removing early states cannot remove covariance among the retained states. Equilibration and uncertainty require distinct evidence.

Applying probability language to an unconstructed field measure. A formal functional weight is not automatically a normalized probability law. Regulator removal and continuum existence require separate analysis.

State how a one-step kernel PP acts on an observable ff and on a probability law μ\mu. Then state the identity that relates the two actions.

Solution

The observable action is conditional expectation,

(Pf)(x)=Ef(y)P(x,dy),(Pf)(x)=\int_Ef(y)P(x,\mathrm dy),

and the forward action on a law is

(μP)(A)=EP(x,A)μ(dx).(\mu P)(A)=\int_EP(x,A)\,\mu(\mathrm dx).

They are dual:

EPfdμ=Efd(μP).\int_EPf\,\mathrm d\mu =\int_Ef\,\mathrm d(\mu P).

Thus PP pulls a future observable back to the present, whereas μP\mu P pushes the present law forward.

2. Separate time averages from marginal convergence

Section titled “2. Separate time averages from marginal convergence”

For

P=(0110),P= \begin{pmatrix} 0&1\\ 1&0 \end{pmatrix},

find the invariant law and compare the large-NN time average with the large-nn one-time law from state 00.

Solution

Solving πP=π\pi P=\pi gives π=(1/2,1/2)\pi=(1/2,1/2). Starting at 00, the path is 0,1,0,1,0,1,0,1,\ldots, so for any ff,

1Nn=0N1f(Xn)f(0)+f(1)2=π(f).\frac1N\sum_{n=0}^{N-1}f(X_n) \longrightarrow\frac{f(0)+f(1)}2 =\pi(f).

However, the one-time law is δ0\delta_0 at even times and δ1\delta_1 at odd times, so it has no limit. Finite-state irreducibility supplies the time-average result here; aperiodicity is the missing condition for ordinary marginal convergence.

3. Solve the Poisson equation for an eigen-observable

Section titled “3. Solve the Poisson equation for an eigen-observable”

Suppose Pu=αuPu=\alpha u, where π(u)=0\pi(u)=0, π(u2)>0\pi(u^2)>0, and α<1|\alpha|<1. Find a Poisson solution and compute Cu=2τint,u\mathcal C_u=2\tau_{\mathrm{int},u}.

Solution

Since

(IP)u=(1α)u,(I-P)u=(1-\alpha)u,

a solution is

h=u1α.h=\frac{u}{1-\alpha}.

Stationarity gives

ρu(k)=αk.\rho_u(k)=\alpha^k.

Therefore

Cu=1+2k=1αk=1+α1α.\mathcal C_u =1+2\sum_{k=1}^{\infty}\alpha^k =\frac{1+\alpha}{1-\alpha}.

For 0<α<10<\alpha<1, correlations inflate the variance. For 1<α<0-1<\alpha<0, alternating correlations give Cu<1\mathcal C_u<1 and hence an effective sample size larger than NN.

For the sampled Gaussian chain, sum k1rkcos(kθ)\sum_{k\geq1}r^k\cos(k\theta) and recover Cf\mathcal C_f. Then compare the linear observable f(Φ)=ϕ1f(\Phi)=\phi_1 with the radial observable q(Φ)=Φ22vq(\Phi)=|\Phi|^2-2v.

Solution

Using a complex geometric series,

k=1rkcos(kθ)=Re ⁣(reiθ1reiθ)=rcosθr212rcosθ+r2.\begin{aligned} \sum_{k=1}^{\infty}r^k\cos(k\theta) &=\operatorname{Re}\!\left( \frac{re^{i\theta}}{1-re^{i\theta}} \right)\\ &=\frac{r\cos\theta-r^2} {1-2r\cos\theta+r^2}. \end{aligned}

Hence

Cf=1+2k=1rkcos(kθ)=1r212rcosθ+r2.\mathcal C_f =1+2\sum_{k=1}^{\infty}r^k\cos(k\theta) =\frac{1-r^2}{1-2r\cos\theta+r^2}.

For the radial observable, the stationary Gaussian fourth-moment identity gives ρq(k)=r2k\rho_q(k)=r^{2k} and therefore

Cq=1+2k=1r2k=1+r21r2.\mathcal C_q=1+2\sum_{k=1}^{\infty}r^{2k} =\frac{1+r^2}{1-r^2}.

Thus the rotation can shorten or oscillate the linear-observable correlations while leaving the radial autocorrelation factor unchanged. The effective sample size is observable-specific.

A Markov kernel propagates both observables and laws, and the Chapman–Kolmogorov identity organizes repeated evolution into a semigroup. A continuous-time generator is the derivative of that semigroup on a declared domain. Invariance fixes a law; stationary initialization realizes it; reversibility adds a symmetry; and ergodic, mixing, and CLT statements each require further hypotheses.

For a stationary scalar observable, the exact sample-mean variance contains all lag covariances. When covariance summability and a valid Markov-chain CLT hold, their infinite sum is the asymptotic variance. Integrated autocorrelation time and effective sample size repackage that variance, but only after the convention and observable are named. The rotating Gaussian mode shows explicitly that invariance does not require reversibility and that two observables of one chain can have different error inflation.

For production lattice analysis—including window and tail choices, replicas, long-chain checks, and observable-specific diagnostics—continue to Autocorrelation Times and Effective Sample Size.

  • J. R. Norris, Markov Chains, § 1.7: Invariant Distributions, PDF, Cambridge University Press, 1997. The linked § 1.7 PDF, pp. 4–5, supports the relation between invariant probabilities and positive recurrence for irreducible countable-state chains, including the null-recurrent random walk contrast.

  • Grigorios A. Pavliotis, Stochastic Processes and Applications, PDF, Springer, 2014; linked author manuscript dated November 11, 2015. §§ 2.2–2.4, pp. 30–39, support transition functions, the Chapman–Kolmogorov identity, Markov semigroups and generators, forward law evolution, and invariant measures.

  • Gareth O. Roberts and Jeffrey S. Rosenthal, “General State Space Markov Chains and MCMC Algorithms”, PDF, Probability Surveys 1 (2004), 20–71, with the linked author version corrected through 2023. pp. 4–5 support detailed balance; § 3.2, pp. 14–18, supports the distinctions among invariance, irreducibility, aperiodicity, convergence, and laws of large numbers; and § 5, pp. 42–47, supports the Markov-chain CLT conditions, covariance-sum variance, and Poisson-equation route.

  • Ulli Wolff, “Monte Carlo Errors with Less Errors”, Computer Physics Communications 156 (2004), 143–153; arXiv:hep-lat/0306017v4, 2006 revision. § 2, pp. 4–6, supports the covariance-sum error, the convention τint=12+k1ρ(k)\tau_{\mathrm{int}}=\tfrac12+\sum_{k\geq1}\rho(k), and N/(2τint)N/(2\tau_{\mathrm{int}}); §§ 3.1–3.3, pp. 7–13, support the finite-window bias–noise tradeoff and error-estimation cautions.