Skip to content

Markov-Chain Sampling of Lattice Fields

A Markov chain samples a regulated lattice field theory correctly only when its transition kernel preserves the declared target measure, communicates across every target-supported sector relevant to the observable, and is run under conditions for which a law of large numbers applies. Detailed balance is a convenient sufficient route to stationarity, not the definition of correctness; acceptance rate and short autocorrelation are efficiency diagnostics, not substitutes for invariance or coverage. This page makes those distinctions at a finite regulator, where every object can be defined and, on a small lattice, checked exactly.

Required background. Markov generators, ergodicity, and sampling error supplies transition kernels, invariant measures, and ergodic averages. Scalar actions and difference operators supplies the finite Euclidean action whose Boltzmann weight is sampled.

Helpful background. Probabilistic convergence and limit theorems supplies laws of large numbers and central-limit hypotheses.

Finite-regulator target and transition kernel

Section titled “Finite-regulator target and transition kernel”

On a finite lattice, let Ω\Omega be the configuration space and let the Euclidean action S(ϕ)S(\phi) define

π(dϕ)=1ZeS(ϕ)μ0(dϕ),Z=ΩeS(ϕ)μ0(dϕ).\pi(d\phi)=\frac{1}{Z}e^{-S(\phi)}\,\mu_0(d\phi), \qquad Z=\int_\Omega e^{-S(\phi)}\,\mu_0(d\phi).

The reference measure μ0\mu_0 is counting measure for discrete spins, Lebesgue measure for noncompact scalar variables, or the product Haar measure for compact links. The normalization ZZ need not be known to form ratios. This chapter assumes a nonnegative, normalizable weight. Complex or sign-indefinite weights require the methods and qualifications of sign problems and finite density.

Local regulator and convention card. The main benchmark is a periodic four-site Ising ring, si{1,+1}s_i\in\{-1,+1\} and s5=s1s_5=s_1, with dimensionless coupling K=βJ0K=\beta J\ge0 and action S(s)=Ki=14sisi+1S(s)=-K\sum_{i=1}^4s_i s_{i+1}. General formulas use a finite or standard Borel state space, a time-homogeneous kernel K(x,dy)K(x,dy), and row-vector convention νn+1=νnK\nu_{n+1}=\nu_nK. All generated-chain expectations are distinguished from exact target expectations π\langle\cdot\rangle_\pi.

A Markov kernel assigns a probability measure K(x,)K(x,\cdot) to the next configuration given xx. The target is invariant when

Ωπ(dx)K(x,A)=π(A)\int_\Omega \pi(dx)K(x,A)=\pi(A)

for every measurable set AA. For a finite state space this is the matrix equation πK=π\boldsymbol\pi K=\boldsymbol\pi. The stronger reversible, or detailed-balance, condition is

π(dx)K(x,dy)=π(dy)K(y,dx).\pi(dx)K(x,dy)=\pi(dy)K(y,dx).

Integrating either side over xx proves invariance. The converse is false: a directed probability current can circulate while leaving π\pi invariant. Therefore a nonreversible kernel is not suspect merely because detailed balance fails; it must instead come with a direct stationarity proof and the same support and convergence checks.

For importance sampling from an independent proposal qq, the identity

Oπ=Eq[O(X)w(X)]Eq[w(X)],w(x)=eS(x)q(x),\langle O\rangle_\pi =\frac{\mathbb E_q[O(X)w(X)]}{\mathbb E_q[w(X)]}, \qquad w(x)=\frac{e^{-S(x)}}{q(x)},

requires q(x)>0q(x)>0 wherever eS(x)O(x)e^{-S(x)}O(x) contributes and sufficient moments for the estimator used. Markov-chain sampling instead changes the distribution of the draws through KK; it does not remove the support condition.

The figure below places the two Markov-chain obligations in their necessary order. Inspect the separate invariant-kernel and coverage gates: neither acceptance nor autocorrelation appears until both have passed.

A declared lattice target must pass stationarity and support-coverage gates before autocorrelation or cost can be used to judge a Markov chain.

For a Markov chain, stationarity establishes the invariant distribution while support communication establishes the scope of ergodic averages. Only a kernel that passes both gates may be compared by observable autocorrelation and cost. Original schematic, not to scale.

The canonical sampler correctness and performance matrix gives the corresponding structured checks for every update family in this chapter.

From stationarity to a valid field estimator

Section titled “From stationarity to a valid field estimator”

Suppose X0,X1,X_0,X_1,\ldots is generated by KK. The usual estimator is

O^N=1Nn=0N1O(Xn).\widehat O_N=\frac{1}{N}\sum_{n=0}^{N-1}O(X_n).

If X0πX_0\sim\pi, then every marginal is already π\pi, so E[O^N]=Oπ\mathbb E[\widehat O_N]=\langle O\rangle_\pi. This statement does not imply that one trajectory explores the target. If a kernel merely flips all spins, sss\mapsto-s, the Ising measure is invariant and the chain is reversible, yet it remains in a two-state orbit. An ergodic theorem additionally needs the chain to communicate across the relevant support and avoid a purely periodic obstruction. On a finite state space, irreducibility plus aperiodicity is a standard sufficient pair; on general state spaces the corresponding recurrence and irreducibility conditions require a measure-theoretic formulation Tierney 1994, §§2–3.

When X0≁πX_0\not\sim\pi, the expectation at step nn is ν0KnO\nu_0K^nO, not πO\pi O. Discarding an arbitrary number of early states does not prove the remaining bias negligible. Convergence must follow from a kernel-level argument and be challenged with starts in separated regions. After convergence, correlated draws still change the variance. For a stationary scalar observable,

Var(O^N)=ΓO(0)N[1+2t=1N1(1tN)ρO(t)],\operatorname{Var}(\widehat O_N) =\frac{\Gamma_O(0)}{N} \left[1+2\sum_{t=1}^{N-1}\left(1-\frac{t}{N}\right)\rho_O(t)\right],

where ΓO(t)=Covπ[O(X0),O(Xt)]\Gamma_O(t)=\operatorname{Cov}_\pi[O(X_0),O(X_t)] and ρO(t)=ΓO(t)/ΓO(0)\rho_O(t)=\Gamma_O(t)/\Gamma_O(0). The limit of the bracket, when it exists, is twice the integrated autocorrelation time in the convention of Chapter 6. Estimating it from realized chains, choosing windows, and propagating it into uncertainties belong to statistical inference and error budgets. Here it appears only to separate correctness from efficiency.

Metropolis–Hastings as a stationarity construction

Section titled “Metropolis–Hastings as a stationarity construction”

Let q(x,dy)q(x,dy) propose yy from xx. When forward and reverse proposal densities exist on the same support, accept with

α(x,y)=min ⁣{1,π(y)q(y,x)π(x)q(x,y)}.\alpha(x,y)=\min\!\left\{1, \frac{\pi(y)q(y,x)}{\pi(x)q(x,y)}\right\}.

The off-diagonal transition density is K(x,y)=q(x,y)α(x,y)K(x,y)=q(x,y)\alpha(x,y); rejected proposals supply the diagonal probability. Multiplication by π(x)\pi(x) gives

π(x)q(x,y)α(x,y)=min{π(x)q(x,y),π(y)q(y,x)},\pi(x)q(x,y)\alpha(x,y) =\min\{\pi(x)q(x,y),\pi(y)q(y,x)\},

which is symmetric under xyx\leftrightarrow y. Thus the accepted moves obey detailed balance, and the rejection term is automatically symmetric on the diagonal. This is the Metropolis–Hastings construction introduced in its general form by Hastings 1970, pp. 97–109, extending the symmetric-proposal method of Metropolis et al. 1953, pp. 1087–1092.

Two qualifications matter in field theory. First, if q(x,y)>0q(x,y)>0 but q(y,x)=0q(y,x)=0, the proposed move cannot be accepted by this reversible construction. Second, a correct ratio does not establish irreducibility: the proposal may conserve a hidden parity, topological sector, or boundary value.

The sixteen configurations of the periodic ring fall into three bond-energy classes. Their exact partition function is

Z4(K)=2e4K+12+2e4K.Z_4(K)=2e^{4K}+12+2e^{-4K}.

Global spin-flip symmetry gives m=0\langle m\rangle=0 for m=14isim=\frac14\sum_i s_i. Differentiating logZ4\log Z_4 gives the nearest-neighbor correlation

C114i=14sisi+1=14logZ4K=4sinh(4K)Z4(K).C_1\equiv\frac14\sum_{i=1}^4\langle s_i s_{i+1}\rangle =\frac14\frac{\partial\log Z_4}{\partial K} =\frac{4\sinh(4K)}{Z_4(K)}.

For a single-spin-flip proposal, choosing one site uniformly and setting si=sis_i'=-s_i changes the action by

ΔS=2Ksi(si1+si+1),\Delta S=2K s_i(s_{i-1}+s_{i+1}),

so the symmetric-proposal acceptance is min(1,eΔS)\min(1,e^{-\Delta S}). Enumerate all sixteen rows of the resulting transition matrix. Each row must sum to one, every entry must be nonnegative, and the exact weights π(s)=eS(s)/Z4\pi(s)=e^{-S(s)}/Z_4 must satisfy both

ry=xπ(x)K(x,y)π(y)=0,Bxy=π(x)K(x,y)π(y)K(y,x)=0.r_y=\sum_x\pi(x)K(x,y)-\pi(y)=0, \qquad B_{xy}=\pi(x)K(x,y)-\pi(y)K(y,x)=0.

These are exact algebraic checks, not noisy Monte Carlo comparisons. The single-spin proposal also connects the full hypercube and has a positive rejection probability at K>0K>0, establishing irreducibility and aperiodicity. Only after these generator checks should sampled estimates of C1C_1 be compared with the closed form above.

A stationary but nonergodic chain. Replace the update by a global spin flip. The exact stationarity and detailed-balance residuals still vanish, but every bond product is conserved. Chains started in the ferromagnetic and alternating configurations never agree on C1C_1. This negative control proves why invariance alone is insufficient.

An asymmetric proposal treated as symmetric. Bias the proposed site or direction in a state-dependent way, but accept with min(1,π(y)/π(x))\min(1,\pi(y)/\pi(x)). The missing q(y,x)/q(x,y)q(y,x)/q(x,y) factor produces a nonzero balance residual and generally a wrong stationary distribution even if the acceptance rate looks healthy.

A correct kernel with a misleading short run. At large KK, a local chain can remain near one magnetized phase over the recorded interval. Symmetry-related starts and exact enumeration expose the finite-time failure; neither deleting a convenient prefix nor reporting m±1\langle m\rangle\approx\pm1 repairs it.

  • Declare SS, μ0\mu_0, boundary conditions, parameter values, and the support on which π\pi is positive.
  • Prove invariance—by detailed balance or a direct calculation—for the implemented transition including rejection and exceptional branches.
  • Test row normalization, nonnegativity, support communication, conserved quantities, and periodicity on the smallest exactly enumerable lattice.
  • Compare at least two symmetry-separated starts and one independent kernel against exact Z4Z_4, C1C_1, and another observable sensitive to sector trapping.
  • Inject a missing proposal ratio or a support-preserving global flip and require the planned residual or cross-start test to fail.
  • Record efficiency separately: proposal cost, acceptance, slow observable, autocorrelation convention, and cost per effective estimate. Chapter 6 determines the uncertainty from the realized chain.

After completing this page, you should be able to:

  • prove invariance of a finite-lattice transition kernel, or identify the exact correction factor that restores it; and
  • distinguish stationarity, ergodicity, finite-time initialization bias, autocorrelation, and computational efficiency by tests that can fail independently.

1. Prove the Ising normalization. Count the four-spin configurations by the number of domain walls and recover Z4(K)Z_4(K).

Solution

Periodic spins have an even number of domain walls. There are two configurations with zero walls, weight e4Ke^{4K} each; choose two of four bonds for the walls and one initial spin, giving 2(42)=122\binom42=12 configurations of weight one; and there are two alternating configurations with four walls, weight e4Ke^{-4K} each. Their sum is 2e4K+12+2e4K2e^{4K}+12+2e^{-4K}.

2. Separate invariance from ergodicity. Write the transition matrix for the deterministic global flip on the sixteen spin configurations. Show that πK=π\pi K=\pi but exhibit a target observable whose time average depends on the starting orbit.

Solution

The matrix has K(s,s)=1K(s,-s)=1 and all other entries zero. Because S(s)=S(s)S(-s)=S(s), π(s)=π(s)\pi(-s)=\pi(s), so xπ(x)K(x,y)=π(y)=π(y)\sum_x\pi(x)K(x,y)=\pi(-y)=\pi(y). Each orbit is only {s,s}\{s,-s\}. The bond correlation C1(s)=14isisi+1C_1(s)=\frac14\sum_i s_i s_{i+1} is constant on such an orbit: it equals 11 from an all-aligned start and 1-1 from an alternating start, while the target value lies strictly between them at finite KK.

  • Hastings, W. K. (1970). “Monte Carlo sampling methods using Markov chains and their applications.” Biometrika 57(1), 97–109. doi:10.1093/biomet/57.1.97.
  • Metropolis, N., Rosenbluth, A. W., Rosenbluth, M. N., Teller, A. H., and Teller, E. (1953). “Equation of state calculations by fast computing machines.” The Journal of Chemical Physics 21(6), 1087–1092. doi:10.1063/1.1699114.
  • Tierney, L. (1994). “Markov chains for exploring posterior distributions.” The Annals of Statistics 22(4), 1701–1762. doi:10.1214/aos/1176325750.