For every fixed inverse temperature $0\lt \beta\lt 1$, we prove that the unscaled spectral gap of single-site heat-bath dynamics for the zero-field Gaussian Sherrington–Kirkpatrick model is bounded away from zero with probability tending to one over the disorder. Equivalently, the Gibbs law satisfies a dimension-free Poincaré inequality for all functions.
How quickly does local resampling remove fluctuations in a densely coupled spin system? For the Sherrington–Kirkpatrick model [19], the signs of the interactions matter: each spin interacts weakly with every other spin, but the sum of the absolute interaction strengths grows with the dimension. We study heat-bath dynamics, which updates a spin from its conditional Gibbs law, in the tradition of Glauber’s local stochastic dynamics [17]. Our result is a dimension-independent Poincaré inequality throughout the high-temperature phase, at each fixed inverse temperature below one. It controls every real observable on a single event of disorder probability tending to one.
Fix \(0<\beta<1\) and put \(j=\beta^2\). Let \(J\) be a real symmetric \(n\times n\) matrix with zero diagonal and independent off-diagonal entries \(J_{ij}\sim N(0,j/n)\), \(i<j\). For a field \(h\in\mathbb R^n\), write \[\mu_h(x)=\frac{1}{Z_h}\exp\left\{\frac12x^{\mathsf T}Jx+h^{\mathsf T}x\right\},
\qquad x\in\{-1,1\}^n.\] Let \(P_i^h\) denote conditional expectation under \(\mu_h\) given all spins except the \(i\)th. Define the unscaled Dirichlet form \[
\mathcal D_h(f)=\sum_{i=1}^n\mathbb E_{\mu_h}(f-P_i^h f)^2.
\tag{1}\]
We use a rate-one clock at each site in continuous time, giving the generator \(\mathcal L_h=\sum_i(P_i^h-I)\). A discrete attempt selects one site uniformly and has transition operator \(P_h=n^{-1}\sum_iP_i^h\). Thus \(\mathcal L_h=n(P_h-I)\), and the discrete gap is the continuous gap divided by \(n\). The gap of a reversible generator is the infimum of its Dirichlet form divided by the variance, over all nonconstant functions.
Theorem 1. There is a finite constant \(C_\beta\), depending only on \(\beta\), such that \[\mathbb P_J\left\{\mathop{\mathrm{Var}}_{\mu_0}(f)\le C_\beta\mathcal D_0(f)
\text{ for every }f:\{-1,1\}^n\longrightarrow\mathbb R\right\}
\longrightarrow1.\] Consequently, the discrete-time heat-bath chain that chooses a uniform site has spectral gap at least \(1/(C_\beta n)\) with probability tending to one.
The parameter \(\beta\) is fixed before \(n\) tends to infinity; the constant \(C_\beta\) may depend on that fixed choice. External fields enter the proof through posterior Gibbs laws, while the theorem concerns \(\mu_0\). The diagonal convention is immaterial to the Gibbs law: because \(x_i^2=1\), adding any diagonal matrix to \(J\) adds a constant to the Hamiltonian. This observation will let us introduce a Gaussian diagonal for matrix calculations and remove it again.
Earlier results and methods
The interval \(\beta<1\) is the classical zero-field high-temperature range. Aizenman, Lebowitz, and Ruelle [2] proved that \(n^{-1}\log Z_0\) converges in probability and in mean to \(\log2+\beta^2/4\), the annealed free-energy value, throughout this interval. The dynamical question is whether local resampling also removes fluctuations at a rate independent of the system size.
A basic route to such a bound controls conditional influences. Wu [24] proved a Poincaré inequality under a Dobrushin spectral-radius condition on a matrix of nonnegative conditional influences. In dense SK systems, stronger fixed-temperature estimates become possible by exploiting the signed random interactions. Bauerschmidt and Bodineau [6] used a Gaussian field decomposition to prove a high-temperature logarithmic Sobolev inequality; its full spin-flip energy differs from the heat-bath form (1). Eldan, Koehler, and Zeitouni [16] established a spectral condition for the heat-bath gap, yielding the SK range \(\beta<1/4\). Adhikari, Brennecke, Xu, and Yau [1] obtained constant unscaled gaps for mixed \(p\)-spin models under a sufficiently small weighted condition on their interaction coefficients.
For mixing in single-site discrete attempts, Anari, Jain, Koehler, Pham, and Vuong [3] used entropic independence to prove an \(O(n\log n)\) bound in the SK range \(\beta<1/4\). Anari, Koehler, and Vuong [4] extended this order to an inverse-temperature threshold approximately equal to \(0.295\). More recently, Wang [23] proved worst-start mixing in \(O_\beta(n\log(n/\varepsilon))\) attempts for every fixed \(\beta<1/2\), simultaneously over external fields. Boban, Li, and Oveis Gharan [7] obtained a constant unscaled gap and \(O(n^2)\) zero-field mixing for \(\beta<1/2+\varepsilon_0\), with an absolute \(\varepsilon_0\ge5\cdot10^{-5}\). Theorem 1 reaches every fixed \(\beta<1\) at zero field.
Equilibrium covariance estimates already cover this entire range. El Alaoui and Gaitonde [12] proved that the expected operator norm of the zero-field spin covariance matrix stays bounded for every fixed \(\beta<1\). Brennecke, Schertzer, Xu, and Yau [9] proved the operator-norm approximation \[\left\|\mathop{\mathrm{Cov}}_{\mu_0}(x)-\bigl((1+\beta^2)I-J\bigr)^{-1}\right\|
\longrightarrow0
\qquad\text{in probability}.\] For \(a\in\mathbb R^n\), the variance of the linear observable \(a^{\mathsf T}x\) is \(a^{\mathsf T}\mathop{\mathrm{Cov}}_{\mu_0}(x)a\). These covariance results therefore control linear fluctuations. The all-functions heat-bath inequality requires additional estimates; our proof obtains them at the random posterior fields generated by Gaussian observations.
Our argument transfers a functional inequality along stochastic localization, introduced by Eldan [15]. We use its Gaussian-observation representation developed by El Alaoui and Montanari [13]. The variance-retention and Dirichlet-form framework of Chen and Eldan [11] explains how posterior estimates can imply an inequality for the initial law. Here an observation event of exponentially small probability can still carry substantial variance for a particular test function. We control its probability after weighting paths by the square of that function. The entropy calculation is a stopped finite-mixture version of the drift-energy representation; see Lehec [18] and the work of Föllmer discussed there.
Gaussian observations, planted laws, and approximate message passing (AMP) are also combined in the SK sampling method of El Alaoui, Montanari, and Sellke [14]. Celentano [10] proved local convexity of the TAP free energy used by that sampler near sufficiently late AMP iterates. Our argument needs invertibility at every sufficiently accurate solution of the field equation along the observation path. A scalar entropy inequality and a local-volume estimate give this uniformity in Section 5. The field recursions retain the Onsager correction of the Thouless–Anderson–Palmer equations [21] and are related to Bolthausen’s iterative construction [8]. We initialize them at a Gibbs spin sample and use exact conditional-spin identities through a fixed finite depth, without requiring convergence of the recursion. Finally, the terminal gap adapts Wang’s signed two-spin method [23]; Section 8 proves the variant with an entrywise-square cost and a quartic row error.
Notation and elementary identities
For a vector \(v\), set \(\langle v\rangle=n^{-1}\sum_i v_i\) and let \(D_v\) be the diagonal matrix with diagonal \(v\). Vector products and scalar functions of vectors are understood coordinatewise. Unless another subscript is displayed, \(\lVert \cdot\rVert\) is Euclidean norm for vectors and operator norm for matrices; \(\lVert \cdot\rVert_{\mathrm F}\) is Frobenius norm. A centered Gaussian orthogonal ensemble (GOE) matrix of scale \(j\) has variances \(j/n\) off diagonal and \(2j/n\) on the diagonal.
If \(x^{(i)}\) is obtained by flipping the \(i\)th spin, put \[d_i f(x)=\frac{f(x)-f(x^{(i)})}{2x_i},\qquad
(\nabla u)_{ij}=d_j u_i.\] Thus \(d_i f\) is independent of \(x_i\) and \[
d_i(fg)=f\,d_i g+g\,d_i f-2x_i(d_i f)(d_i g).
\tag{2}\] At a fixed field \(h_0\) set \[h^1=h_0+Jx,\qquad m^1=\tanh h^1,\qquad v^1=\mathop{\mathrm{sech}}^2 h^1.\] Since \(J_{ii}=0\), \(m_i^1\) and \(v_i^1\) are respectively the conditional mean and variance of \(x_i\). In particular, \[
\mathcal D_{h_0}(f)=\mathbb E_{\mu_{h_0}}\sum_i v_i^1(d_i f)^2,\qquad
\lVert G\rVert_*^2=\mathbb E_{\mu_{h_0}}G^2+\mathcal D_{h_0}(G).
\tag{3}\] For probability laws \(\nu,\mu\), their relative entropy is \(D(\nu\Vert\mu)=\int\log(\mathrm d\nu/\mathrm d\mu)\,\mathrm d\nu\) when \(\nu\ll\mu\), and \(+\infty\) otherwise.
Constants are independent of \(n\); dependence on a fixed temperature, construction depth, or positive stability margin is recorded at its use. Their values may change between displays. The prerequisites are Gaussian concentration and comparison, Gaussian integration by parts, and stochastic calculus. We prove the model-specific estimates in the sections below.
The proof mechanism
There are two parts to the argument: prove useful inequalities for posterior Gibbs laws, then transfer them to the unobserved law. A long but fixed observation makes most individual spins predictable. The corresponding small conditional variances, combined with a two-spin curvature inequality, give a terminal gap. To recover the initial variance, we need a directional covariance estimate along the path and a mean estimate under square reweighting. The former bounds the loss of variance; the latter prevents that variance from accumulating on exceptional paths.
The posterior estimates come from finite Onsager-corrected recursions. Exact conditional-spin identities control their residuals. A telescoping identity for the squared increments then selects two consecutive small increments without requiring convergence of the recursion. Uniform matrix estimates handle the resulting adaptive diagonal coefficients under an explicit inverse margin. To compare these approximate fields with an exact root, stability must hold throughout the small-residual region and persist when the diagonal variance matrix is replaced by the nearby response coefficients used in the comparison. A scalar entropy estimate controls a Gaussian integral, and a local-volume lower bound turns that integral control into exclusion of unstable approximate roots. All tolerances and finite depths are chosen before the large-\(n\) limit.
The three inputs
We observe a spin sample in independent Gaussian noise: \[
X\sim\mu_0,\qquad Y(t)=tX+B(t),\quad 0\le t\le T,
\tag{4}\] where \(B\) is standard \(n\)-dimensional Brownian motion independent of \(X\). Conditional on the observations up to \(t\), the spin law is \(\mu_{Y(t)}\). The proof uses the following three inputs, established in the later sections.
Stability along ordinary observation paths. Proposition 9 gives a deterministic notion of a good field such that, for typical \(J\), all fields \(Y(t)\) are good except on an event of observation probability at most \(e^{-a n}\), where \(a>0\) is fixed.
Two estimates at a good field. For every \(G\), Proposition 23 proves \[
\lVert \mathop{\mathrm{Cov}}_{\mu_h}(G,x)\rVert^2
\le C\bigl(\mathop{\mathrm{Var}}_{\mu_h}(G)+\mathcal D_h(G)\bigr).
\tag{5}\] For \(N=\mathbb E_{\mu_h}f^2>0\) and \(\nu=f^2\mu_h/N\), Proposition 24 proves, for each sufficiently small fixed \(\delta>0\), \[
\frac{\lVert \mathbb E_\nu x-t_*(h)\rVert}{\sqrt n}
\le C\delta+\frac{C_\delta}{\sqrt n}
\left(1+\frac{\mathcal D_h(f)}{N}\right).
\tag{6}\] Here \(t_*(h)=\tanh r(h)\), where \(r(h)\) is the unique solution of \[r-h+\{j(1-n^{-1}\|\tanh r\|^2)I-J\}\tanh r=0.\] Existence and uniqueness at the fields under consideration are part of Lemma 15. The constant \(C\) is independent of \(\delta\); \(C_\delta\) may depend on it. Taking \(f=1\) gives the same comparison for \(\mathbb E_{\mu_h}x\).
A terminal gap. For a sufficiently large fixed \(T\), Proposition 27 gives \(\mathop{\mathrm{Var}}_{\mu_{Y(T)}}f\le2\mathcal D_{Y(T)}(f)\) outside an event of ordinary observation probability at most \(e^{-a n}\).
The transfer uses the two estimates for different purposes. Directional covariance controls posterior variance loss, while square-reweighted means bound the entropy of the changed observation law. Their combination controls the contribution of the bad paths before the terminal gap is applied. Section 2 proves this implication first, so the later technical sections have a precise target.
Organization
After the observation transfer in Section 2, Section 3 constructs the matrix diagnostics used by the adaptive field calculations. Sections 4 and 5 identify the scalar equality law and establish stability along observations. Section 6 develops the finite recursions and their weak residual rules; Section 7 uses them to prove the two posterior estimates. Section 8 closes the argument with the fixed-time terminal gap.
From posterior estimates to a spectral gap
We now prove the transfer announced in the introduction. Its inputs are a covariance bound, a square-reweighted mean bound, and a terminal gap. The stopping rule and good-disorder event must be independent of the test function, so we keep the exceptional-path estimate explicit. The argument uses the Gaussian-channel formulation of localization [13] and the variance/Dirichlet-form framework of [11]; the finite-spin identities are derived here.
Posterior martingales and Dirichlet forms
Let \(\mathcal F_t=\sigma(Y(s):s\le t)\) and write \(\mathbb E_t\) for expectation under \(\mu_{Y(t)}\). Bayes’ formula gives this posterior: the likelihood of an observation path given \(X=x\) is proportional to \(\exp(x^{\mathsf T}Y(t)-tn/2)\). The process \[I(t)=Y(t)-\int_0^t\mathbb E_s x\,\mathrm ds\] is the innovation Brownian motion in \(\mathcal F_t\): conditional expectation makes it a continuous martingale, and its quadratic covariation is \(tI_n\). The Brownian martingale characterization therefore applies. Applying Itô’s formula to the finitely many posterior weights gives, for every spin function \(g\), \[
\mathrm d(\mathbb E_t g)=\mathop{\mathrm{Cov}}_t(g,x)^{\mathsf T}\mathrm dI(t).
\tag{7}\] To verify the cancellation, differentiate the posterior expectation with respect to \(Y\); its gradient is \(\mathop{\mathrm{Cov}}_t(g,x)\). Each likelihood weight \(\exp(x^{\mathsf T}Y(t)-tn/2)\) satisfies a stochastic exponential equation. The quotient rule for their weighted sums then gives (7): subtracting \(\mathbb E_t x\,\mathrm dt\) from \(\mathrm dY\) removes the finite-variation term.
At every deterministic time, \[
\mathbb E\mathcal D_{Y(t)}(f)\le\mathcal D_0(f).
\tag{8}\] This is conditional-variance contraction: the \(i\)th summand on the left is \(\mathbb E\mathop{\mathrm{Var}}(f(X)\mid X_{-i},\mathcal F_t)\), and the one on the right is \(\mathbb E\mathop{\mathrm{Var}}(f(X)\mid X_{-i})\). The conditional variance decomposition gives the inequality after averaging over observations. In this section, an expectation without a field subscript refers to that experiment, with \(J\) held fixed.
Stopping before stability fails
Stop before the first possible loss of the field estimates. More precisely, fix the terminal horizon \(T\), take the observation stopping time \(\tau_0\) from Corollary 10, and set \(\tau=T\wedge\tau_0\). This rule does not use the test function. Before \(\tau\), stability holds through residual threshold \(\rho_0\sqrt n\), so both posterior estimates apply. On the specified matrix event, \(\mathbb P(\tau_0\le T)\le e^{-an}\).
For a fixed \(J\), choose the canonical terminal good set \[\mathcal T_n(J)=\{h:\mathop{\mathrm{Var}}_{\mu_h}(g)\le2\mathcal D_h(g)\text{ for every }g\}.\] Membership means positivity of a finite-dimensional symmetric quadratic form with coefficients continuous in \(h\), so this set is Borel. Proposition 27 controls the probability that \(Y(T)\) misses it. Declare the stopped path bad if \(\tau_0\le T\), or if \(\tau_0>T\) and \(Y(T)\notin\mathcal T_n(J)\). The event is \(\mathcal F_\tau\)-measurable: before \(T\) only stopping can make it bad, and at \(T\) the remaining condition is determined by \(Y(T)\). The two failure probabilities give, on the typical matrix events, \[
\mathbb P(\mathrm{bad})\le 2e^{-a n}
\tag{9}\] for a fixed \(a>0\), after decreasing the rate if necessary.
Suppose \(\mathbb E_{\mu_0}f=0\) and \(\mathbb E_{\mu_0}f^2=1\). Set \[M_t=\mathbb E_t f,\quad N_t=\mathbb E_t f^2,\quad V_t=N_t-M_t^2,
\quad d_0=\mathcal D_0(f).\] By (7) and optional stopping, \(u(t)=\mathbb EV_{t\wedge\tau}\) is absolutely continuous and satisfies, almost everywhere, \[u'(t)=-\mathbb E\bigl[\mathbf 1_{\{t<\tau\}}\lVert \mathop{\mathrm{Cov}}_t(f,x)\rVert^2\bigr].\] The directional estimate bounds the loss term before stopping. Together with (8), applied at the deterministic time \(t\), it gives \(u'(t)\ge-C_{\mathrm d}(u(t)+d_0)\), with fixed \(C_{\mathrm d}\). Solving this scalar inequality from \(u(0)=1\) yields \[
\mathbb EV_\tau\ge v-(1-v)d_0,\qquad v=e^{-C_{\mathrm d}T}>0.
\tag{10}\] All uses of optional stopping here are at bounded times and for bounded random variables at each fixed \(n\).
Changing the law of the observations
The variance lower bound alone does not control how much variance lies on bad paths. For that purpose change the initial spin law to \(f^2\mu_0\), which is a probability law by normalization, and call the resulting observation law \(\mathbb P^f\). Its density on \(\mathcal F_t\) is \(N_t\), and its spin posterior is \(\nu_t=f^2\mu_{Y(t)}/N_t\). Since \(\mu_{Y(t)}\) has full support and \(f\) is nonzero, \(N_t>0\). Bounded stopping gives the density \(N_\tau\) on \(\mathcal F_\tau\). Let \[a_t=\mathbb E_{\nu_t}x-\mathbb E_t x.\] Equation (7) with \(g=f^2\) gives \(\mathrm dN_t/N_t=a_t^{\mathsf T}\mathrm dI(t)\). Under \(\mathbb P^f\), \(I\) has drift \(a_t\). The drift-energy entropy identity, in the form discussed by Lehec [18], has a direct finite-mixture proof here. Itô’s formula for \(\log N_t\) yields \[
D\bigl(\mathbb P^f|_{\mathcal F_\tau}\,\Vert\,
\mathbb P|_{\mathcal F_\tau}\bigr)
=\mathbb E^f\log N_\tau
=\frac12\mathbb E^f\int_0^\tau\lVert a_t\rVert^2\,\mathrm dt.
\tag{11}\] The stopped identity needs no limiting integrability argument. At fixed \(n\), the drift satisfies \(\lVert a_t\rVert\le2\sqrt n\) and \(\tau\le T\). Novikov’s condition therefore holds, and the stochastic integral in the log-density formula has zero mean.
Before stopping, applying (6) to \(f\) and to \(1\) gives \[b_t:=\frac{\lVert a_t\rVert}{\sqrt n}
\le2C_{\mathrm w}\delta+\frac{C_\delta}{\sqrt n}
\left(2+\frac{\mathcal D_{Y(t)}(f)}{N_t}\right).\] Since \(b_t\le2\), we may use \(b_t^2\le2b_t\). At deterministic \(t\), the density identity and (8) imply \[
\mathbb E^f\frac{\mathcal D_{Y(t)}(f)}{N_t}
=\mathbb E\mathcal D_{Y(t)}(f)\le d_0.
\tag{12}\] Write \(\mathcal H_\tau\) for the entropy in (11). The preceding estimates give the quantitative bound \[
\frac{\mathcal H_\tau}{n}
\le 2C_{\mathrm w}T\delta
+\frac{C_\delta T}{\sqrt n}(2+d_0).
\tag{13}\] The separation of constants is essential: \(C_{\mathrm w}\) is independent of \(\delta\), whereas \(C_\delta\) may grow when \(\delta\) decreases. We next convert this entropy bound into a bound on the bad-path probability. If an event has probabilities \(p\) and \(q\) under two laws, the log-sum inequality for the event and its complement bounds relative entropy below by \(p\log(1/q)-\log2\). With \(p=\mathbb P^f(\mathrm{bad})\), (9) therefore gives, for large \(n\), \[p\le\frac{\mathcal H_\tau+\log2}{an-\log2}.\] If the ordinary bad probability is zero, absolute continuity gives \(p=0\). For \(d_0\le1\), (13) therefore implies \[
p\le\frac{4C_{\mathrm w}T}{a}\delta
+\frac{6C_\delta T}{a\sqrt n}
+\frac{2\log2}{an}.
\tag{14}\] Choose \(\delta>0\) once so that the first term is at most \(v/8\). With this choice fixed, enlarge \(n\) until the last two terms sum to at most \(v/8\). The resulting bound holds simultaneously for all normalized \(f\) with \(d_0\le1\): \[
\mathbb P^f(\mathrm{bad})\le v/4.
\tag{15}\]
Figure 1 records how the two posterior estimates will control the stopped variance and its exceptional contribution.
Variance retention and control of exceptional paths, for a fixed typical matrix. Normalize \(f\) by \(\mathbb E_{\mu_0}f=0\) and \(\mathbb E_{\mu_0}f^2=1\), and put \(d_0=\mathcal D_0(f)\le1\). Along \(Y(t)=tX+B(t)\), let \(V_t=\mathop{\mathrm{Var}}_{\mu_{Y(t)}}f\), \(N_t=\mathbb E_{\mu_{Y(t)}}f^2\), and \(\tau=T\wedge\tau_0\). Here \(\mathcal H_\tau\) is the relative entropy of the stopped observation law induced by the spin law \(f^2\mu_0\), relative to the ordinary law, and \(v=e^{-C_{\mathrm d}T}\). The density \(N_\tau\) converts its bad-path probability into an upper bound on exceptional variance. If \(d_0>1\), the final lower bound is automatic.
Completion of the theorem
We can now close the argument, keeping the order of choices explicit. First choose \(T\) from the terminal estimate. Next choose the stable-field constants, the ordinary failure rate \(a\), and the finite diagnostic length for the directional estimate; these choices determine \(C_{\mathrm d}\) and \(v\). Root stability and the operator bound on \(J\) determine the leading weighted-mean constant \(C_{\mathrm w}\) independently of selection depth. We may therefore select \(\delta\) in (14), then its diagnostic length and \(C_\delta\), and finally \(n_0\). Every depth and tolerance is fixed before \(n\) grows.
With the choices made, the proof needs a finite intersection of matrix events. Include the directional word event and the word events for the selected weighted-mean construction, each with its fixed length and coefficient bounds. Lemma 2 accommodates each fixed \(M_0\), and Lemma 18 supplies an actual inverse margin independent of depth. Add the events of Propositions 9 and 27. Their intersection still has probability tending to one, and all constants and the threshold \(n_0\) are uniform on it and over the test functions.
For such a matrix and a normalized \(f\) with \(d_0\le1\), use \(V_\tau\le N_\tau\) and the stopped density identity to obtain \[\mathbb E[\mathbf 1_{\mathrm{bad}}V_\tau]
\le\mathbb E[\mathbf 1_{\mathrm{bad}}N_\tau]
=\mathbb P^f(\mathrm{bad})\le v/4.\] On a good path \(\tau=T\) and the terminal gap applies, so \[\mathbb E[\mathbf 1_{\mathrm{good}}V_\tau]
\le2\mathbb E\mathcal D_{Y(T)}(f)\le2d_0.\] Combining these estimates with (10) gives \[v-(1-v)d_0\le2d_0+v/4,
\qquad\text{hence}\qquad
d_0\ge\frac{3v}{4(3-v)}\ge\frac v4.\] The case \(d_0>1\) already satisfies this bound. Center and rescale any nonconstant function to obtain Theorem 1 once the three inputs have been proved, with the finite constant \(C_\beta=4e^{C_{\mathrm d}T}\). We do not optimize this constant. For the clock conversion, the self-adjoint operator \(H=\sum_i(I-P_i^0)\) has form \(\mathcal D_0\), and one discrete attempt is \(I-H/n\). Its gap is therefore the asserted unscaled gap divided by \(n\).
Uniform random-matrix estimates
The field analysis uses random-matrix estimates in two forms. Finite field recursions produce products whose diagonal coefficients may be chosen after the disorder is observed. Lemma 2 controls such products of bounded length, subject to a separate positivity margin whenever an inverse occurs. We also obtain fixed-vector bilinear estimates and a simultaneous determinant bound for the stability argument in Section 5. Their expectation calculations use the same inverse identity. We begin with fixed deterministic coefficients and then use concentration and finite meshes to obtain the required simultaneous estimates.
Throughout this section, \(j\in(0,1)\) is fixed. A centered GOE matrix \(W\) has independent entries on and above the diagonal, with \[\mathbb EW_{ab}^{2}=j/n\quad(a<b),\qquad
\mathbb EW_{aa}^{2}=2j/n.\] We couple \(J\) to \(W\) by putting \(J=W-D_{\mathop{\mathrm{diag}}W}\). In particular, the off-diagonal law of \(J\) is the law used in the theorem. Fix \(\eta>0\) so that, with \(a_+=1+\eta\), \[
j a_+^2<1.
\tag{16}\] For a nonnegative diagonal matrix \(A\le a_+I\), put \[
\begin{gathered}
\chi_A=\frac1n\mathop{\mathrm{Tr}}A,\qquad B_A=j\chi_A,\qquad
K_A(M)=K_A(1;M),\\
K_A(z;M)=(I-zAM+z^2B_AA)^{-1}.
\end{gathered}
\tag{17}\] The associated symmetric matrix is \[S_A(z;M)=I+z^2B_AA-zA^{1/2}MA^{1/2}.\] The symmetric matrix lets us control the inverse without imposing a positive lower bound on the individual diagonal entries of \(A\). Whenever \(S_A(z;M)\) is invertible, the following identities hold: \[\begin{align*}
K_A(z;M)
&=I+A^{1/2}S_A(z;M)^{-1}A^{1/2}(zM-z^2B_AI),
\tag{18}\\
K_A(z;M)A&=A^{1/2}S_A(z;M)^{-1}A^{1/2}.
\tag{19}\end{align*}\] To verify the first identity, apply \((I-UV)^{-1}=I+U(I-VU)^{-1}V\). For the second, multiply the two sides using \(S_A=I-A^{1/2}(zM-z^2B_AI)A^{1/2}\). These formulas will also let us truncate the inverse while retaining bounded matrix factors.
Words and their predictions
A word is an ordered product of diagonal matrices, copies of a symmetric matrix \(M\), and at most one factor \(K_A(M)\). Its length is the number of factors. Consecutive diagonal factors may always be combined, and an absent diagonal factor is interpreted as \(I\). To each word \(Q\) we associate a diagonal matrix \(\mathsf P(Q)\) by the contraction rule below. The estimates will compare the diagonal of \(Q\) with this explicitly defined matrix.
Begin with a word without \(K_A\). Consider every complete noncrossing pairing of its occurrences of \(M\), in their linear order. For each pairing, remove an innermost pair by the rule \[
MDM\ \longmapsto\ j\frac{\mathop{\mathrm{Tr}}D}{n}I,
\tag{20}\] where \(D\) is diagonal after previous removals. Combine adjacent diagonals and continue. Sum the resulting diagonal matrices over all complete noncrossing pairings. If there is no complete pairing, the prediction is zero; if there is no occurrence of \(M\), it is the word itself. The result for a fixed pairing does not depend on the order of removal: two available innermost pairs are disjoint, and their removals commute. For words without an inverse factor, this is the operator-valued semicircular moment rule with covariance map \(D\mapsto jn^{-1}\mathop{\mathrm{Tr}}(D)I\); see Speicher [20].
For example, four occurrences of \(M\) have two complete noncrossing pairings, giving \[\begin{split}
\mathsf P(MD_1MD_2MD_3M)
={}&j^2\frac{\mathop{\mathrm{Tr}}D_1}{n}\frac{\mathop{\mathrm{Tr}}D_3}{n}D_2\\
&+j^2\frac{\mathop{\mathrm{Tr}}D_2}{n}\frac{\mathop{\mathrm{Tr}}(D_1D_3)}n I.
\end{split}\] The first term pairs the first two and last two occurrences; the second pairs the outside occurrences around the middle pair.
For a word containing \(K_A\), replace that factor by the formal series \[
K_A(z;M)=\sum_{r\ge0}(zAM-z^2B_AA)^r,
\tag{21}\] and apply the preceding rule coefficient by coefficient. Ordinary occurrences of \(M\) outside this factor are left unscaled. We prove below that the resulting series \(\mathsf P(Q;z)\) is a polynomial; we define \(\mathsf P(Q)=\mathsf P(Q;1)\).
The corresponding trace prediction is \(\mathsf p(Q)=n^{-1}\mathop{\mathrm{Tr}}\mathsf P(Q)\). To see its cyclic invariance, draw an ordinary trace word on a circle, with a chord for each pair in a noncrossing pairing. Its contribution is \(j\) for each chord, multiplied by, for each face, the average of the product of the diagonal matrices on that face’s boundary arcs. Removing an innermost chord gives exactly (20), and the average over the last face is the final normalized trace. This description is unchanged by rotating the circle. For words with \(K_A(z;M)\), apply this argument coefficient by coefficient in (21).
Lemma 2 (Simultaneous word estimates). Fix \(L\in\mathbb N\), \(M_0<\infty\), and \(c>0\). There are constants \(C,c_1>0\) and \(n_0\), depending only on \(j,a_+,L,M_0,c\), such that the following event has probability at least \(1-Ce^{-c_1n}\) for \(n\ge n_0\). Simultaneously for all words \(Q\) of length at most \(L\), all their real diagonal factors of norm at most \(M_0\), and all \(0\le A\le a_+I\), one has \[
\lVert Q\rVert\le C,\qquad
\left(\sum_{a\ne b}|Q_{ab}|^4\right)^{1/4}\le C,\qquad
\lVert \mathop{\mathrm{diag}}(Q-\mathsf P(Q))\rVert_2\le C.
\tag{22}\] For a word containing \(K_A(J)\), the assertion is restricted to parameters satisfying \(S_A(1;J)\succeq cI\). The same conclusion holds with \(W\) in place of \(J\), with the corresponding stability restriction. Thus all diagonal parameters in this assertion may be chosen after the random matrix has been observed.
If a word has \(r\) ordinary occurrences of \(M\) outside its one \(K_A(z;M)\) factor, then \(\mathsf P(Q;z)\) has degree at most \(r\).
We prove the simultaneous estimate in three steps. For a fixed deterministic diagonal, Gaussian comparison gives stability along \(0\le z\le1\). We then calculate the expected diagonal of a truncated word. Finally, concentration makes the word estimates simultaneous in the diagonal parameters. The first step supplies stability for each deterministic choice; it does not remove the separate stability restriction for adaptive \(A\) in Lemma 2.
Lemma 3 (Stability for a deterministic diagonal). There is \(c_0>0\), depending only on \(j,a_+\), with the following property. For every fixed deterministic \(0\le A\le a_+I\), and every fixed \(R>2\sqrt j\), there are constants \(C,c_2>0\) and \(n_0\), uniform in \(A\), such that \[
\mathbb P\left\{\lVert W\rVert\le R\ \hbox{and}\quad
S_A(z;W)\succeq c_0I\text{ for all }0\le z\le1\right\}
\ge1-Ce^{-c_2n}\qquad(n\ge n_0).
\tag{23}\] In addition, for every fixed \(\varepsilon>0\), \(\mathbb P\{\lVert W\rVert>2\sqrt j+\varepsilon\}\le
C_\varepsilon e^{-c_\varepsilon n}\) for sufficiently large \(n\).
Proof. We first bound the expected quadratic process that defines the smallest eigenvalue of \(S_A\). Concentration will give a margin at each fixed \(z\), and a finite mesh will extend it along the path.
We use two standard Gaussian facts. A real-valued \(L\)-Lipschitz function of a standard Gaussian vector satisfies \(\mathbb P\{|F-\mathbb EF|>u\}\le2\exp(-u^2/(2L^2))\). Also, if two centered Gaussian processes have ordered squared increments, the process with larger squared increments has the larger expected supremum (Sudakov–Fernique). The latter comparison remains valid after the same deterministic function is added to both processes; see the deterministic-shift comparison in Vitale [22]. Both facts apply by finite nets and continuity to the compact index sets below. The linear map from the independent standard Gaussian coordinates to \(W\), in Frobenius norm, has norm \(\sqrt{2j/n}\).
For the comparison, write \(w=A^{1/2}p\) for \(\lVert p\rVert=1\) and consider \[X_p=w^TWw,\qquad
Y_p=2\sqrt{j/n}\,\lVert w\rVert\,g^Tw,\] where \(g\) is standard Gaussian in \(\mathbb R^n\). If \(w'=A^{1/2}p'\), direct calculation gives \[\begin{align*}
&\mathbb E(Y_p-Y_{p'})^2-\mathbb E(X_p-X_{p'})^2\\
&\qquad=\frac{2j}{n}
\left[(\lVert w\rVert^2-\lVert w'\rVert^2)^2
+2(\lVert w\rVert\lVert w'\rVert-w^Tw')^2\right]\ge0.
\end{align*}\] Consequently, for each \(z\in[0,1]\), \[\mathbb E\sup_{\lVert p\rVert=1}\{zX_p-z^2j\chi_A\lVert w\rVert^2\}
\le \mathbb E\sup_{\lVert p\rVert=1}\{zY_p-z^2j\chi_A\lVert w\rVert^2\}.\] Set \(T_g=\sqrt{j/n}\lVert A^{1/2}g\rVert\). Since \(g^Tw=p^TA^{1/2}g\le\lVert A^{1/2}g\rVert\), the last supremum is at most \[\sup_{0\le v\le\sqrt{a_+}}
(2zT_gv-z^2j\chi_Av^2).\] The remaining supremum depends on the Gaussian vector only through \(T_g\). The variable \(T_g^2\) has mean \(j\chi_A\) and variance \(2j^2\mathop{\mathrm{Tr}}(A^2)/n^2\le C/n\). The inequality \(|\sqrt x-\sqrt y|\le\sqrt{|x-y|}\) therefore shows, uniformly in \(A\), that \(\mathbb E|T_g-\sqrt{j\chi_A}|=o(1)\). On replacing \(T_g\) by \(\sqrt{j\chi_A}\) in the preceding display, the expression becomes \(2d-d^2\), with \(0\le d=z\sqrt{j\chi_A}v\le\sqrt j\,a_+<1\). Thus, with \(\delta=(1-\sqrt j\,a_+)^2>0\), \[
\mathbb E\sup_{\lVert p\rVert=1}
\{zp^TA^{1/2}WA^{1/2}p-z^2j\chi_Ap^TAp\}
\le1-\delta+o(1),
\tag{24}\] uniformly in \(A,z\). The supremum in (24) is \(a_+\)-Lipschitz in \(W\) for Frobenius norm. Gaussian concentration gives an exponentially small probability that it exceeds \(1-\delta/2\), uniformly for each fixed \(A,z\).
To pass from a fixed \(z\) to the entire interval, we also need a bound on \(\lVert W\rVert\). Taking \(A=I\) in the comparison without deterministic offsets gives \[\mathbb E\lambda_{\max}(W)\le2\sqrt{j/n}\mathbb E\lVert g\rVert\le2\sqrt j.\] The same holds for \(-W\). Concentration of these two eigenvalues proves the stated norm tail. On \(\lVert W\rVert\le R\), the function of \(z\) inside the supremum in (24) is uniformly Lipschitz, with constant at most \(a_+R+2ja_+^2\). A sufficiently fine constant mesh of \([0,1]\), followed by a union bound, therefore gives (23), for example with \(c_0=\delta/4\). ◻
Truncation and the loop recursion
The inverse in a word need not exist for every realization of the matrix. We therefore replace each word by a globally defined bounded function of \(W\), agreeing with the original word whenever the required norm and stability conditions hold. This permits concentration and integration by parts before those conditions are imposed.
Fix a sufficiently large constant \(R\). Let \(\Pi_R\) be metric projection, in the Hilbert space of real symmetric matrices with Frobenius norm, onto \(\{M:\lVert M\rVert\le R\}\). This is a nonexpansive map and is equivariant under orthogonal conjugation. Put \(\widehat W=\Pi_R(W)\). Choose a smooth compactly supported real function \(f\) which equals \(x^{-1}\) on the interval \[\left[\tfrac12\min(c,c_0),\,2+c+c_0+ja_+^2+a_+R\right].\] This interval contains, with slack, every stable spectrum needed both for deterministic-parameter estimates and for the restriction \(S_A(1;M)\succeq cI\) in Lemma 2. Replace \(K_A(z;W)\) by \[
\widehat K_A(z)
=I+A^{1/2}f(S_A(z;\widehat W))A^{1/2}
(z\widehat W-z^2B_AI).
\tag{25}\] Replace every ordinary \(W\) in a word by \(\widehat W\) as well; denote the resulting word by \(\widehat Q_\theta(z)\), where \(\theta\) lists all its diagonal parameters. These truncated factors equal the original factors on the stability events where they will be used.
For each fixed length and fixed bounds on the diagonal parameters, there is a constant \(C\) such that \[\begin{align*}
\lVert \widehat Q_\theta(z)\rVert&\le C,
&\lVert \widehat Q_\theta(z;W)-\widehat Q_\theta(z;W')\rVert_{\mathrm F}
&\le C\lVert W-W'\rVert_{\mathrm F},
\tag{26}\\
\mathop{\mathrm{Lip}}_{\mathrm F}\bigl(\widehat Q_\theta(z;\cdot)
-\widehat Q_{\theta'}(z;\cdot)\bigr)
&\le C\sqrt\Delta
&&\text{if }\lVert \theta-\theta'\rVert_\infty\le\Delta\le1.
\tag{27}\end{align*}\] The constants are uniform for \(z\in[0,1]\) and for all matrix sizes. Here the Lipschitz constants concern maps into matrices with Frobenius norm.
We verify both estimates with constants independent of dimension. On the operator-norm ball, all factors are bounded, and their differentials in a perturbation \(E\) are bounded in Frobenius norm by \(C\lVert E\rVert_{\mathrm F}\). The only functional-calculus factor is handled by writing \(f\) as a Fourier integral and using \[\mathrm d(e^{itS})[E]
=it\int_0^1 e^{it(1-u)S}E e^{ituS}\,\mathrm du.\] The Fourier transform of \(f\) decays faster than any power, so its first two weighted absolute moments are finite. This formula and the corresponding difference formula give \[\lVert \mathrm df(S)[E]\rVert_{\mathrm F}\le C\lVert E\rVert_{\mathrm F},\qquad
\lVert (\mathrm df(S)-\mathrm df(S'))[E]\rVert_{\mathrm F}
\le C\lVert S-S'\rVert\lVert E\rVert_{\mathrm F}.\] Changing the diagonal parameters by at most \(\Delta\) changes their square roots in norm by at most \(\sqrt\Delta\). Consequently the factors, and their matrix differentials, change by respectively \(C\sqrt\Delta\) and \(C\sqrt\Delta\lVert E\rVert_{\mathrm F}\). The product rule proves the same differential assertions for a word. Integrating along the line segment between two matrices in the convex operator-norm ball proves (26) and (27) there. Composition with the nonexpansive projection \(\Pi_R\) proves them everywhere.
These Lipschitz estimates first give the moment bounds needed for the expectation calculation. For a fixed parameter tuple, \[
\mathop{\mathrm{Var}}(\widehat Q_{ab})\le C/n,\qquad
\mathop{\mathrm{Var}}\left(\frac1n\mathop{\mathrm{Tr}}\widehat Q\right)\le C/n^2.
\tag{28}\] The normalized trace gains a factor \(n^{-1/2}\) in its Frobenius Lipschitz constant. The same estimates for a difference of two parameter tuples gain a factor \(\Delta\) on the right-hand sides. Fourth centered entry moments are bounded by \(C/n^2\), or \(C\Delta^2/n^2\) for such a difference, by integrating the Gaussian concentration tail.
Lemma 4 (Fixed-parameter diagonal expectations). For the truncated words just defined, uniformly in deterministic diagonal parameters in the indicated bounded ranges, in \(z\in[0,1]\), and in the diagonal index \(i\), \[
\mathbb E\widehat Q_\theta(z)_{ii}
=\mathsf P(Q_\theta;z)_{ii}+O(n^{-1}).
\tag{29}\] For words without \(K_A\), the prediction has no \(z\) dependence. The formal prediction for a word with \(r\) ordinary \(W\) letters is a polynomial of degree at most \(r\).
Proof. The calculation has two parts. We first justify integration by parts for the truncated words and compute the inverse factor alone. We then obtain a recursion in the number of ordinary \(W\) factors; the same recursion determines the formal prediction exactly and the expectation up to \(O(n^{-1})\).
All constants in this proof depend only on the fixed length and parameter bounds. Lemma 3 shows that, for every fixed deterministic parameter tuple, the truncations agree with the actual matrices for all \(z\in[0,1]\), except on an event of probability \(Ce^{-c_2n}\). Gaussian integration by parts can be applied to the globally Lipschitz truncated factors, using weak derivatives. More explicitly, for a scalar Lipschitz function \(H\), \[
\mathbb E[W_{ab}H(W)]
=\frac jn\mathbb E\bigl[\mathrm dH(W)[E_{ab}+E_{ba}]\bigr].
\tag{30}\] When \(a=b\), the elementary direction on the right is doubled, as required by the diagonal variance \(2j/n\).
For (30), remove the projection from the distinguished ordinary \(W\) factor. The resulting expectation error is exponentially small, since the other factors are bounded and Gaussian norm tails control the removed factor. On the event of agreement, the differential of the inverse is \[
\mathrm dK_A(z;W)[E]=zK_A(z;W)AEK_A(z;W).
\tag{31}\] We may use the expressions on the right, with truncated factors, on the whole probability space. The error is exponentially small even after the index sums: all factors and their true truncated differentials are bounded, there are at most a fixed polynomial number of indices, and the disagreement event has exponentially small probability. More precisely, the elementary direction \(E_{ab}+E_{ba}\) has Frobenius norm at most \(2\). Both the true weak differential and its formal replacement are therefore bounded by \(C\) in each entry; their difference vanishes on the agreement event. Each such differential error has expected absolute value at most \(Ce^{-c_2n}\). Expanding the diagonal entry of a length-\(L\) word and applying the product rule uses at most \((L-1)n^{L-1}\) scalar contributions before any index identifications; the factor \(j/n\) in integration by parts only improves this bound. Thus their total error is at most a fixed polynomial in \(n\) times \(e^{-c_2n}\), hence at most \(C'e^{-c'n}\). Subsequent words have bounded lengths depending only on the original length. We therefore suppress hats during the loop calculation, with this convention understood.
We now compute the base case, in which the only nondiagonal factor is \(K=K_A(z;W)\). Write \(a_i=A_{ii}\), \(k_i(z)=\mathbb EK_{ii}\), and \(r(z)=n^{-1}\sum_i a_i k_i(z)\). Differentiating \(K_{bi}\) in (30) gives \[\begin{align*}
\mathbb E(WK)_{ii}
&=\frac{zj}{n}\mathbb E\left[(\mathop{\mathrm{Tr}}KA)K_{ii}\right]
+\frac{zj}{n}\sum_b\mathbb E[(KA)_{bi}K_{bi}]+O(e^{-c'n})\\
&=zjr(z)k_i(z)+O(n^{-1}).
\end{align*}\] For the second sum, the absolute value of the sum over \(b\) is at most \(\lVert KAe_i\rVert\lVert Ke_i\rVert\le C\). The factorization in the first term follows from (28); its covariance is at most \(C n^{-3/2}\), which is more than enough. The inverse identity \(K=I+zAWK-z^2B_AAK\) now yields \[
\mathbb E(WK)_{ii}=zjr k_i+O(n^{-1}),\qquad
k_i=1+z^2ja_i(r-\chi_A)k_i+O(n^{-1}).
\tag{32}\]
The loop equations must select the solution near \(k_i=1\). We establish this by continuation from \(z=0\), where the inverse is the identity. Put \(h(z)=r(z)-\chi_A\). The truncated \(k_i\) are uniformly bounded and continuous in \(z\). If \(|h(z)|\le\delta_1\), the second equation in (32) gives \(\max_i|k_i(z)-1|\le C\delta_1+C/n\). Choose \(\delta_1>0\) small, using (16), so that throughout this region and for large \(n\), \[\left|z^2j\frac1n\sum_i a_i^2 k_i(z)\right|
\le1-\delta_2\] for some fixed \(\delta_2>0\). Averaging the second equation in (32) with weights \(a_i\) gives \[h(z)=z^2j\left(\frac1n\sum_i a_i^2k_i(z)\right)h(z)+O(n^{-1}),\] and hence \(|h(z)|\le C/n\) while \(|h(z)|\le\delta_1\). Since \(h(0)=0\), continuity rules out a first exit from that region for large \(n\). It follows uniformly in \(z\) that \[
\mathbb EK_{ii}=1+O(n^{-1}),\qquad
\mathbb E(WK)_{ii}=zB_A+O(n^{-1}).
\tag{33}\]
To complete the base case, we verify the formal identity \(\mathsf P(K_A(z;W))=I\) independently of the expectation calculation. The cancellation takes place at each coefficient of the series. In the expansion (21), a block \(-z^2B_AA\) may be regarded as the negative contraction of two adjacent blocks \(zAW,zAW\), since their contraction is \(z^2j\chi_AA=z^2B_AA\). Fix a completely noncrossing-paired chain \((AW)^{2m}\) with \(m>0\). Let \(\mathcal A\) be the nonempty set of its pairs which join two originally adjacent occurrences of \(W\). It is nonempty because every nonempty noncrossing pairing of a linear order has an adjacent pair. All pairs in \(\mathcal A\) are disjoint. The terms of (21) which give this paired chain are in bijection with subsets \(T\subseteq\mathcal A\): contract the pairs in \(T\) into the negative deterministic blocks, and pair the remaining occurrences normally. The magnitude of the final diagonal contribution is unchanged, and its sign is \((-1)^{|T|}\). Thus the total multiplier is \(\sum_{T\subseteq\mathcal A}(-1)^{|T|}=0\). This argument is finite at each power of \(z\) and proves the asserted formal identity. Multiplication by surrounding diagonal matrices proves both claims of the lemma for words with no ordinary \(W\) letters.
For the induction step, select an ordinary occurrence of \(W\) in a general word and apply integration by parts at that occurrence. We describe the resulting finite recursion directly, without requiring convergence of the formal series. In each product-rule term, mark the partner of the selected occurrence: for a derivative falling on an ordinary factor, this is the \(W\) being differentiated. If the derivative falls on \(K_A\), replace that factor by \(zK_AA\,\boxed W\,K_A\) and take the boxed occurrence as the partner. In either case, list the two marked positions in their linear order, so the marked product reads \[U\,\boxed W\,V\,\boxed W\,Q.\] The two orientations in (30) contribute respectively \[
\frac jn\mathbb E\bigl[(UQ)_{ii}\mathop{\mathrm{Tr}}V\bigr],\qquad
\frac jn\mathbb E(UV^TQ)_{ii}.
\tag{34}\] There is an additional factor \(z\) when the derivative fell on \(K_A\). The second expression in (34) is \(O(n^{-1})\) by operator-norm bounds, including when it contains two copies of \(K_A\). The first expression is \[
j\left(\frac1n\mathbb E\mathop{\mathrm{Tr}}V\right)\mathbb E(UQ)_{ii}+O(n^{-1}),
\tag{35}\] again by (28). The index sums explain the two terms. The opposite orientation identifies the two interior indices and gives \(\mathop{\mathrm{Tr}}V\); the equal orientation reverses the interior path and gives \(V^T\).
Consequently \(\mathbb E\widehat Q_{ii}\) is the finite sum of the leading products (35) over all ordinary partners, plus \(z\) times the corresponding leading product for the one \(K_A\) factor if present, with an \(O(n^{-1})\) remainder. Each word \(V\) and \(UQ\) in this recursion has fewer ordinary \(W\) letters than the original word. Each contains at most one \(K_A\): if the derivative fell on \(K_A\), the splitting places one of its two copies in each word. Extra diagonal factors \(A\) cause no problem, since at most one is added per decrease of the induction index.
The decrease in the number of ordinary letters holds whether the selected occurrence lies before or after the inverse. The two possible orders can be checked explicitly. If the original word is \(P K_A R W T\), the contraction with \(K_A\) contributes \[zj\,\mathbb E\left[(P K_A A T)_{ii}\frac{\mathop{\mathrm{Tr}}(K_A R)}n\right].\] If it is \(P W R K_A T\), that contribution is \[zj\,\mathbb E\left[(P K_A T)_{ii}\frac{\mathop{\mathrm{Tr}}(R K_A A)}n\right].\] Here \(P,R,T\) contain no inverse factor. The displayed outer and inner words each have one inverse factor, even though the original distinguished occurrence may lie anywhere among the ordinary letters.
It remains to compare this expectation recursion with the formal prediction. Sort the noncrossing pairings by the partner of the selected ordinary occurrence. The inside and outside may then be paired independently; their predictions are respectively \(\mathsf p(V)\) and \(\mathsf P(UQ)\). If the partner occurs inside the series for \(K_A(z;W)\), marking that occurrence in the formal series is exactly the formal differentiation identity \(\mathrm dK_A[E]=zK_AAEK_A\). Thus the entire sum over positions in that series is the single \(z\)-weighted product just described.
We can now induct on the number \(r\) of ordinary \(W\) letters. An ordinary partner leaves a total of \(r-2\) such letters in the two smaller words; a partner in \(K_A\) leaves \(r-1\). The base prediction is constant. The recursion therefore proves that the formal prediction is a polynomial of degree at most \(r\). It also proves (29): use the induction hypothesis for the two expectations in each leading product, and average the diagonal induction hypothesis to handle its normalized trace. There are only finitely many products and all factors are uniformly bounded. The same induction, with no inverse factor, proves the statement for ordinary words. ◻
Making the word estimates simultaneous
Proof of Lemma 2. We first make the centered estimates uniform over the diagonal parameters, then add the expectation estimate from Lemma 4. Fix a letter pattern and let \(\theta\) range over its parameter box, which has at most \(C_Ln\) real coordinates. Work initially with the truncated words at \(z=1\). Conjugation by any diagonal sign matrix preserves the law of \(W\) and conjugates \(\widehat Q_\theta\) by that same sign matrix. The diagonal parameters commute with the conjugation, and both the projection and the functional calculus respect it. Thus \(\mathbb E(\widehat Q_\theta)_{ab}=0\) whenever \(a\ne b\).
Define the two matrix seminorms \[N_2(M)=\lVert \mathop{\mathrm{diag}}M\rVert_2,\qquad
N_4(M)=\left(\sum_{a\ne b}|M_{ab}|^4\right)^{1/4}.\] Both are \(1\)-Lipschitz with respect to Frobenius norm. The entry moment bounds following (28) give \[
\mathbb EN_k(\widehat Q_\theta-\mathbb E\widehat Q_\theta)\le C
\quad(k=2,4).
\tag{36}\] Indeed, sum the \(n\) second moments for \(k=2\), and the at most \(n^2\) fourth moments for \(k=4\), then apply Jensen’s inequality. For two parameter tuples at distance at most \(\Delta\), the same argument, using (27), gives \[\mathbb EN_k\bigl((\widehat Q_\theta-\mathbb E\widehat Q_\theta)
-(\widehat Q_{\theta'}-\mathbb E\widehat Q_{\theta'})\bigr)
\le C\sqrt\Delta.\] The random norms in this last display concentrate around their means at Gaussian scale \(C\sqrt{\Delta/n}\), again by (27). The centered expectations are deterministic and do not change Lipschitz constants.
To control every parameter tuple, approximate it successively on finer meshes. For each \(\ell\ge0\), choose a sup-norm mesh \(\mathcal N_\ell\) of spacing \(2^{-\ell}\) in the parameter box. It may be chosen with \[|\mathcal N_\ell|\le\exp(C_L(\ell+1)n).\] For a tuple \(\theta\), choose nearest mesh points \(\theta_\ell\in\mathcal N_\ell\). Then \(\lVert \theta_\ell-\theta_{\ell-1}\rVert_\infty\le C2^{-\ell}\). For all pairs of mesh points at this distance, concentration and a union bound show that their centered increments have both \(N_2\) and \(N_4\) norm at most \[C'\sqrt{\ell+1}\,2^{-\ell/2},\] except with probability \(Ce^{-c(\ell+1)n}\), provided \(C'\) is chosen sufficiently large. Specifically, the squared ratio of this threshold to the Gaussian scale is a constant multiple of \((\ell+1)n\), and its constant can be chosen to dominate the logarithm of the number of mesh pairs. The initial mesh is bounded in the same way using (36). Summing the failure probabilities over \(\ell\), and summing the convergent increment bounds along each chain, proves, with probability \(1-Ce^{-c_1n}\), \[
\sup_\theta N_k(\widehat Q_\theta-\mathbb E\widehat Q_\theta)\le C
\quad(k=2,4).
\tag{37}\] At each fixed \(n\), continuity of the truncated word and its expectation in \(\theta\) justifies taking the limit of the mesh chain. There are only finitely many letter patterns of length at most \(L\), so their events may be intersected.
The centered words are now controlled on one event for every parameter tuple. Lemma 4 supplies the remaining bias estimate, \(N_2(\mathbb E\widehat Q_\theta-\mathsf P(Q_\theta))\le C/\sqrt n\), uniformly in \(\theta\). Sign symmetry makes the off-diagonal of \(\mathbb E\widehat Q_\theta\) identically zero. Combining these facts with (37) proves the desired diagonal and off-diagonal assertions for the truncated words. Their operator norms are already bounded by construction. On \(\lVert W\rVert\le R\) and at parameters with \(S_A(1;W)\succeq cI\), the truncations equal the exact words, proving the GOE version.
The GOE estimate transfers to the zero-diagonal model through \(J=W-D_{\mathop{\mathrm{diag}}W}\). This last step uses a Frobenius bound on the removed diagonal. Since \[\lVert D_{\mathop{\mathrm{diag}}W}\rVert_{\mathrm F}^{2}
\ \stackrel{\mathrm{law}}=\ \frac{2j}{n}\chi_n^2,\] it is bounded by a fixed constant except with exponentially small probability. The norm tail for \(W\), and a sufficiently large choice of \(R\), ensure also that both \(W\) and \(J\) lie in the operator-norm ball except with exponentially small probability. For every parameter tuple, (26) then bounds the Frobenius norm of the difference of the truncated word evaluated at \(W\) and at \(J\) by a fixed constant. Both \(N_2\) and \(N_4\) of this difference are bounded by its Frobenius norm. The prediction is unchanged by this substitution. At parameters with \(S_A(1;J)\succeq cI\), the truncated factors evaluated at \(J\) equal the exact factors. This proves all of (22). The polynomial assertion was proved in Lemma 4. ◻
Bilinear forms and regularized determinants
We next extract two forms of the preceding estimates. The first controls bilinear forms for a prescribed finite family of deterministic vectors. The second controls a regularized determinant simultaneously over its diagonal parameters and bounded finite-rank perturbations. Their quantifiers differ, and we keep the two statements separate.
Lemma 5 (Bilinear deterministic equivalents). Fix an integer \(m\ge1\), a vector bound \(M_1<\infty\), a tolerance \(\varepsilon>0\), and \(R>2\sqrt j\). For every deterministic \(0\le A\le a_+I\) and deterministic vectors \(v_1,\ldots,v_m\in\mathbb R^n\) with \(\lVert v_r\rVert\le M_1\), there is an event of probability at least \(1-Ce^{-c_3n}\) on which (23) holds and, with \[G_A=A^{1/2}S_A(1;W)^{-1}A^{1/2}=K_A(W)A,\] one has, for every \(r,s\), \[\begin{align*}
|v_r^T(G_A-A)v_s|&\le\varepsilon,
\tag{38}\\
|v_r^T(G_AW-B_AA)v_s|&\le\varepsilon,
\tag{39}\\
|v_r^T(WG_AW-B_AI-B_A^2A)v_s|&\le\varepsilon.
\tag{40}\end{align*}\] Here \(C,c_3,n_0\) may depend on \(j,a_+,m,M_1,\varepsilon,R\) but are uniform in \(A\), the vectors, and \(n\ge n_0\). The assertion is for each deterministic choice of these parameters; it does not claim simultaneity over all vector choices.
Proof. For each of the specified bilinear forms, we calculate its expectation and then apply concentration. Use the truncations above with a cutoff agreeing with the inverse on the path-stability event. Sign conjugation makes the off-diagonal expectations of each relevant matrix vanish. Equation (33) at \(z=1\) gives \[\mathbb E\widehat K_{ii}=1+O(n^{-1}),\qquad
\mathbb E(\widehat W\widehat K)_{ii}=B_A+O(n^{-1}).\] The identities, valid on the event of agreement, \[G_AW=K_A(W)(I+B_AA)-I,
\qquad
WG_AW=WK_A(W)(I+B_AA)-W\] therefore give the three diagonal expectation predictions \(A\), \(B_AA\), and \(B_AI+B_A^2A\), with entrywise errors \(O(n^{-1})\). Their use with truncated factors changes expectations only exponentially little, by the same argument as in Lemma 4. In particular, their errors in any of the specified bilinear forms are \(O(n^{-1})\): for a diagonal error matrix \(E\), one has \(|v_r^TEv_s|\le\lVert E\rVert\lVert v_r\rVert\lVert v_s\rVert\).
The expectation calculation has identified the three required predictions. To control fluctuations, each truncated bilinear form is \(C M_1^2\)-Lipschitz in \(W\) for Frobenius norm, by (26). Gaussian concentration gives failure probability \(Ce^{-c_3n}\) at any fixed positive tolerance. Intersect these events for the \(3m^2\) forms and with Lemma 3. For large \(n\) the expectation errors are less than half the specified tolerance, proving the result. ◻
Lemma 6 (A simultaneous regularized determinant bound). Fix \(\varepsilon>0\), an integer \(r_0\ge0\), and \(M_2<\infty\). There is \(\gamma_0>0\), depending only on \(j,\varepsilon\), such that for every fixed \(0<\gamma\le\gamma_0\) and all sufficiently large \(n\), an event of probability at least \(1-Ce^{-c_4n}\) has the following property. Simultaneously for all diagonal \(0\le V\le I\) and all real matrices \(E\) with \(\mathop{\mathrm{rank}}E\le r_0\) and \(\lVert E\rVert\le M_2\), put \[b=\frac1n\mathop{\mathrm{Tr}}V,\qquad B=jb,\qquad
L=I+(BI-W)V+E.\] Then \[
\frac1n\log\det(L^TL+\gamma I)\le jb^2+\varepsilon.
\tag{41}\] The constants \(C,c_4\) and the lower threshold on \(n\) may depend on \(j,\varepsilon,\gamma,r_0,M_2\). In particular, \(V\) and \(E\) may depend on \(W\).
Proof. We first compute the mean log determinant with \(E=0\) and a fixed deterministic \(V\). Regularization will then provide a global Lipschitz bound, allowing a finite mesh in \(V\) and finally the finite-rank perturbation. Set \[L_z=I+z^2BV-zWV,
\qquad K(z)=(I-zVW+z^2BV)^{-1}.\] Since \(V\le I\), we use the proof of Lemma 3 with the diagonal upper bound equal to \(1\); its stability and inverse-norm constants here depend only on \(j\). On the path-stability event from Lemma 3 with \(A=V\), the determinant of \(L_z\) equals the positive determinant of \(S_V(z;W)\) by \(\det(I-UV)=\det(I-VU)\), and both \(L_z^{-1}\) and \(K(z)\) have uniformly bounded operator norm. Indeed \(L_z=K(z)^{-T}\), and (18) applies. Differentiating the normalized log determinant on this event gives \[\begin{align*}
\frac{\mathrm d}{\mathrm dz}\frac1n\log\det L_z
&=\frac1n\mathop{\mathrm{Tr}}\bigl[L_z^{-1}(2zBV-WV)\bigr]\\
&=2zB\frac1n\mathop{\mathrm{Tr}}(K(z)V)
-\frac1n\mathop{\mathrm{Tr}}(WK(z)V).
\tag{42}\end{align*}\] Here \(VL_z^{-1}=K(z)V\), which follows by multiplying \((I-zVW+z^2BV)V=V L_z\), and cyclicity of trace was used.
The derivative is expressed in the two matrices whose diagonal expectations were computed in (33). Let \(\mathcal E_V\) be the path-stability event intersected with \(\lVert W\rVert\le R\), for a fixed \(R>2\sqrt j\). This one event applies to the entire interval of \(z\). Equations (33), and replacement by the truncated factors off \(\mathcal E_V\), show uniformly in \(z,V\) that the expectation of the right-hand side of (42), multiplied by \(\mathbf 1_{\mathcal E_V}\), is \[2zBb-zBb+O(n^{-1})=zjb^2+O(n^{-1}).\] All differentiated quantities are bounded on \(\mathcal E_V\), so integration and expectation may be interchanged. Since \(L_0=I\), it follows that \[
\mathbb E\left[\mathbf 1_{\mathcal E_V}\frac2n\log\det L_1\right]
=jb^2+O(n^{-1}).
\tag{43}\] On \(\mathcal E_V\), the singular values of \(L_1\) are bounded below by a positive constant. Therefore \[0\le\frac1n\log\det(L_1^TL_1+\gamma I)
-\frac2n\log\det L_1
\le C\gamma.\] Off \(\mathcal E_V\), the absolute value of the regularized normalized log determinant is bounded by \[|\log\gamma|+C+2\log(1+\lVert W\rVert),\] for \(0<\gamma\le1\). The norm tails of \(W\) and (23) make its expected contribution exponentially small, for each fixed \(\gamma\). Consequently, uniformly in deterministic \(V\), \[
\mathbb E\left[\frac1n\log\det(L_1^TL_1+\gamma I)\right]
\le jb^2+C\gamma+O(n^{-1})+O_\gamma(e^{-c n}).
\tag{44}\]
The mean bound is uniform in deterministic \(V\). To obtain a single event for all \(V\), we use the concentration supplied by regularization. Define on all real matrices \[F_\gamma(M)=\frac1n\log\det(M^TM+\gamma I).\] Its Frobenius gradient is \(2n^{-1}M(M^TM+\gamma I)^{-1}\). Each singular value of this gradient is at most \(1/(n\sqrt\gamma)\), and hence \[
\mathop{\mathrm{Lip}}_{\mathrm F}(F_\gamma)\le\frac1{\sqrt{n\gamma}}.
\tag{45}\] For fixed \(V\), the map \(W\mapsto L_1\) has Frobenius Lipschitz constant at most one. Gaussian concentration therefore gives \[
\mathbb P\{|F_\gamma(L_1)-\mathbb EF_\gamma(L_1)|>u\}
\le2\exp\left(-\frac{\gamma n^2u^2}{4j}\right).
\tag{46}\]
Choose \(\gamma_0\le1\) small enough that \(C\gamma_0\) in (44) is at most \(\varepsilon/8\). For fixed \(\gamma\le\gamma_0\) and sufficiently large \(n\), the expectation bound and (46) bound all points of any fixed sup-norm mesh of the cube \([0,1]^n\) by \(jb^2+\varepsilon/2\), with failure \(\exp(-c_\gamma n^2)\). Indeed a mesh of spacing \(\Delta>0\) has at most \((C/\Delta)^n\) points, whose logarithm is only of order \(n\). On \(\lVert W\rVert\le R\), if \(\lVert V-V'\rVert\le\Delta\), then \[|b-b'|\le\Delta,\qquad
\lVert (I+(BI-W)V)-(I+(B'I-W)V')\rVert_{\mathrm F}
\le\sqrt n(2j+R)\Delta.\] By (45), the corresponding \(F_\gamma\) values differ by at most \((2j+R)\Delta/\sqrt\gamma\), whereas \(|jb^2-j(b')^2|\le2j\Delta\). Choose \(\Delta\), depending only on \(j,R,\gamma,\varepsilon\), so that these errors sum to at most \(\varepsilon/4\). Including the exponentially likely norm event proves, simultaneously for all \(V\), \[F_\gamma(I+(BI-W)V)\le jb^2+3\varepsilon/4.\]
We have obtained the simultaneous bound for \(E=0\) with room for the perturbation. The same Lipschitz estimate (45) gives, for every allowed \(E\), \[|F_\gamma(M+E)-F_\gamma(M)|
\le\frac{\lVert E\rVert_{\mathrm F}}{\sqrt{n\gamma}}
\le M_2\sqrt{\frac{r_0}{n\gamma}}.\] This is at most \(\varepsilon/4\) for large \(n\), uniformly even when \(E\) is chosen from the observed matrix. It proves (41) and the lemma. ◻
A scalar entropy inequality
The stability argument needs a scalar criterion that forces an empirical field law toward a Gaussian law. We first bound relative entropy below by a squared discrepancy involving the field and its hyperbolic tangent. We then identify the Gaussian law compatible with the scalar consistency equations. Together with the equality case of the first estimate, this characterizes the consistent equality laws.
Throughout this section, \(N(d,s)\) denotes the Gaussian law with mean \(d\) and variance\(s\). All scalar expectations are taken with respect to the indicated probability law.
Lemma 7 (Scalar entropy inequality). Let \(0<j<1\), \(d\in\mathbb R\), and \(s>0\). Let \(P\) be a probability law on \(\mathbb R\) with finite second moment, and set \[m(y)=\tanh y,\qquad q=\mathbb E_P m^2,\qquad b=1-q,\qquad
a=\mathbb E_P[(y-d)m],\qquad S=s+jq.\] If \(s\ge jq\), then \[
D(P\Vert N(d,s))\ge \frac{j(a-sb)^2}{2sS}.
\tag{47}\] Equality holds if and only if \(P=N(d,s)\).
Proof. The second-moment assumption makes the right-hand side finite, so it suffices to consider the case \(D:=D(P\Vert N(d,s))<\infty\). In this case \(P\) is absolutely continuous, and hence \(q>0\): otherwise \(P\) would be the point mass at zero. Also \(q<1\), because \(\tanh^2 y<1\) for every finite \(y\). We use the following dimensionless parameters: \[
k=\frac{jq}{s}\in(0,1],\qquad C=\frac{ja}{s},\qquad
B=jb\in(0,1),\qquad a_0=1-B\in(0,1).
\tag{48}\] In these parameters, multiplying the desired inequality by \(j\) gives the target \[
jD\ge \frac{(C-B)^2}{2(1+k)}.
\tag{49}\]
We will combine two entropy bounds. A monotone change of variables handles one range of the discrepancy \(C-B\); comparison with a Gaussian of a different variance handles the remaining range. We retain strictness in both arguments to identify all equality cases.
A change-of-variables bound.
For the change of variables, fix \(u<1\). The map \[T_u(y)=y-u\tanh y\] is an increasing \(C^1\) diffeomorphism of \(\mathbb R\) onto itself, with derivative \(T_u'(y)=1-u(1-m(y)^2)>0\). Change of variables in relative entropy yields \[\begin{align*}
D
&=D((T_u)_\#P\Vert N(d,s))
+\frac{2ua-u^2q}{2s}
+\mathbb E_P\log\bigl(1-u(1-m^2)\bigr)\\
&\ge \frac{2ua-u^2q}{2s}
+\mathbb E_P\log\bigl(1-u(1-m^2)\bigr).
\end{align*}\] The entropy identity is well defined: for fixed \(u<1\), \(\log T_u'\) is bounded, and the second-moment assumption makes the difference of the Gaussian log densities integrable. To estimate the logarithmic term, apply concavity on the segment joining \(1\) and \(1-u\). For \(0\le v\le1\), this gives \[\log(1-uv)\ge v\log(1-u).\] Consequently, with \(\ell(u)=-\log(1-u)-u\), \[
jD\ge u(C-B)-\frac{k u^2}{2}-B\ell(u),\qquad u<1.
\tag{50}\] Completing the square isolates the amount by which this bound can exceed the target in (49): \[
jD-\frac{(C-B)^2}{2(1+k)}
\ge h(u)-\frac{\bigl(C-B-(1+k)u\bigr)^2}{2(1+k)},
\qquad h(u)=\frac{u^2}{2}-B\ell(u).
\tag{51}\]
We first locate a range in which the extra term \(h(u)\) is strictly positive: \[
h(u)>0\quad\text{for }u\le a_0,\ u\ne0,
\qquad h(a_0)\ge\frac{a_0^3}{6}.
\tag{52}\] For the negative range, integrate \(v/(1+v)\le v\) over \(0\le v\le -u\) to obtain \(\ell(u)\le u^2/2\). Since \(B<1\), this proves strict positivity for \(u<0\). For the positive range, use \(\ell(u)=\sum_{r\ge2}u^r/r\): it shows that \(\ell(u)/u^2\) is increasing on \(0<u<1\). Moreover, using \(B=1-a_0\), \[h(a_0)=\frac{a_0^2}{2}-(1-a_0)\ell(a_0)
=\sum_{r\ge3}\frac{a_0^r}{r(r-1)}
\ge\frac{a_0^3}{6}.\] It follows that \(h(u)/u^2\ge h(a_0)/a_0^2>0\) for \(0<u\le a_0\), proving (52).
We now choose the transport parameter. If \(u_*=(C-B)/(1+k)\le a_0\) and \(C\ne B\), set \(u=u_*\) in (51). The square vanishes and \(h(u_*)>0\), so (49) is strict. If \(u_*>a_0\), put \[C_0=1+ka_0,
\qquad C_1=C_0+\sqrt{\frac{1+k}{3}}\,a_0^{3/2}.\] In this case \(C>C_0\). The endpoint choice \(u=a_0\) in (51) still gives \[jD-\frac{(C-B)^2}{2(1+k)}
\ge \frac{a_0^3}{6}-\frac{(C-C_0)^2}{2(1+k)}.\] This proves strictness throughout \(C_0<C<C_1\) and leaves only the range beginning at \(C_1\).
The remaining range.
For \(C\ge C_1\), we use the second moment to obtain the remaining entropy bound. Put \(v=\mathbb E_P[(y-d)^2]/s>0\). Comparison with the Gaussian \(N(d,sv)\) gives the exact identity \[D=D(P\Vert N(d,sv))+\frac{v-1-\log v}{2},\] and hence \(2D\ge v-1-\log v\). Cauchy–Schwarz gives \[v\ge\frac{a^2}{sq}=\frac{C^2}{jk}>1,\] where the last inequality follows from \(C\ge C_1>C_0>1\) and \(jk<1\). The function \(v\mapsto v-1-\log v\) is increasing for \(v\ge1\), so \[2jD\ge \frac{C^2}{k}-j-j\log\frac{C^2}{jk}
\ge \frac{C^2}{k}-1-\log\frac{C^2}{k}.\] The second inequality removes \(j\) from the lower bound. Indeed, with \(C,k\) held fixed, the derivative of \(C^2/k-\lambda-\lambda\log(C^2/(\lambda k))\) is \(-\log(C^2/(\lambda k))\le0\) for \(j\le\lambda\le1\). Thus the remaining comparison with the target reduces to \[f(C):=\frac{C^2}{k}-1-\log\frac{C^2}{k}
-\frac{(C-B)^2}{1+k}>0
\quad(C\ge C_1).\] We estimate \(f\) at \(C_0\) and control its increase from \(C_0\) to \(C_1\). Direct simplification at \(C_0\) gives \[\begin{align*}
f(C_0)
&=\frac1k-1+\log k
+2a_0-a_0^2-2\log(1+ka_0)\\
&\ge -\frac{2a_0^3}{3}.
\end{align*}\] Indeed \(1/k-1+\log k\ge0\) for \(0<k\le1\), and \[\log(1+ka_0)\le\log(1+a_0)
\le a_0-\frac{a_0^2}{2}+\frac{a_0^3}{3}.\] The last bound follows by integrating \((1+x)^{-1}\le1-x+x^2\) from zero to \(a_0\). Differentiation also gives \[\begin{align*}
f'(C_0)&=\frac2k-\frac{2}{1+ka_0}
\ge\frac{2a_0}{1+a_0},\\
f''(C)&=\frac{2}{k(1+k)}+\frac{2}{C^2}\ge1
\qquad(C>0).
\end{align*}\] For completeness, the first inequality follows from \[\frac1k-\frac1{1+ka_0}-\frac{a_0}{1+a_0}
=\frac{(1-k)(1+a_0+ka_0^2)}{k(1+ka_0)(1+a_0)}\ge0.\] The derivative bounds compensate for the possible negative value at \(C_0\). Write \(\Delta=C_1-C_0\ge a_0^{3/2}/\sqrt3\) and integrate them to obtain \[\begin{align*}
f(C_1)
&\ge -\frac{2a_0^3}{3}
+\frac{2a_0}{1+a_0}\Delta+\frac{\Delta^2}{2}\\
&\ge a_0^3\left(
-\frac23+\frac{2}{\sqrt3(1+a_0)\sqrt{a_0}}+\frac16
\right)>0.
\end{align*}\] The strict inequality uses \(0<a_0<1\) and \(1/\sqrt3+1/6>2/3\). Since \(f'\) is positive at \(C_0\) and increasing thereafter, \(f(C)>0\) for every \(C\ge C_1\).
The two ranges together prove strictness in (49) whenever \(C\ne B\). To finish the equality case, suppose \(C=B\). The target then has right-hand side zero, so equality is equivalent to \(D=0\), or \(P=N(d,s)\). Conversely, Gaussian integration by parts gives \(a=s\mathbb E_P\mathop{\mathrm{sech}}^2 y=sb\) for this law, which verifies equality. ◻
We next impose the scalar consistency equations used in the application. Here the Gaussian parameters themselves depend on \(P\): for given \(t\ge0\) and \(\sigma\ge0\), set \[
M=\mathbb E_P\tanh y,\qquad q=\mathbb E_P\tanh^2 y,\qquad
d=t+jM,\qquad s=t+\sigma^2+jq.
\tag{53}\] These equations give \(s\ge jq\), the comparison required by Lemma 7. When \(s>0\), that lemma restricts equality to Gaussian laws. The next result identifies which Gaussian laws satisfy the consistency equations when \(\sigma=0\), and also treats the degenerate case at \(t=0\).
Lemma 8 (The consistent Gaussian law). Fix \(0<j<1\). For each \(t\ge0\), the equation \[
s_*(t)=t+j\mathbb E_{N(s_*(t),s_*(t))}\tanh y
\tag{54}\] has a unique nonnegative solution, where \(N(0,0)\) means the point mass at zero. Its associated Gaussian law \(P_*(t)=N(s_*(t),s_*(t))\) is the unique Gaussian law satisfying (53) with \(\sigma=0\). Thus, among consistent laws with \(s>0\), equality in (47) holds precisely at \(P_*(t)\).
The function \(t\mapsto s_*(t)\) is Lipschitz with constant \((1-j)^{-1}\), and \[
t\le s_*(t)\le\frac{t}{1-j},\qquad
q_*(t):=\mathbb E_{P_*(t)}\tanh^2 y
=\frac{s_*(t)-t}{j}\le\frac{t}{1-j}.
\tag{55}\] In particular \(s_*(0)=q_*(0)=0\), and \(s_*(t),q_*(t)>0\) for \(t>0\).
Proof. We first show that a consistent Gaussian must have equal mean and variance. We then solve the resulting one-variable fixed-point equation and prove the bounds in the statement.
Two Gaussian symmetry identities provide the first step. For \(s>0\), the density \(\rho_s\) of \(N(s,s)\) satisfies \[\frac{\rho_s(y)-\rho_s(-y)}{\rho_s(y)+\rho_s(-y)}=\tanh y.\] Pairing the integrals over \(y\) and \(-y\) therefore gives \[
\mathbb E_{N(s,s)}\tanh y=\mathbb E_{N(s,s)}\tanh^2 y.
\tag{56}\] Second, if \(d\ge0\) and \(s>0\), pairing the positive and negative half-lines shows that \[
\mathbb E_{N(d,s)}[\tanh y\,\mathop{\mathrm{sech}}^2 y]\ge0.
\tag{57}\] Here the integrand is odd and nonnegative for \(y\ge0\), while the density at \(y\) is at least its value at \(-y\) when \(d\ge0\).
Now suppose a Gaussian law \(N(d,s)\) with \(s>0\) satisfies (53) with \(\sigma=0\). We begin by locating its mean. For fixed \(s\), the map \[d\longmapsto t+j\mathbb E_{N(d,s)}\tanh y\] is a contraction with Lipschitz constant at most \(j\), because its derivative is \(j\mathbb E_{N(d,s)}\mathop{\mathrm{sech}}^2 y\in[0,j]\). Iteration from zero stays nonnegative, so its unique fixed point \(d\) is nonnegative.
To compare this nonnegative mean with the variance, keep \(s\) fixed and define \[H_s(d)=\mathbb E_{N(d,s)}[\tanh y-\tanh^2 y].\] For \(d\ge0\), Gaussian differentiation and (57) give \[H_s'(d)=\mathbb E_{N(d,s)}[(1-m^2)(1-2m)]\in[-1,1].\] Indeed the integrand is bounded below by \(-1\), and its expectation is at most \(\mathbb E(1-m^2)\le1\) after subtracting \(2\mathbb E[m(1-m^2)]\ge0\). Thus \(H_s\) is 1-Lipschitz on \([0,\infty)\). By (56), \(H_s(s)=0\). On the other hand, consistency says \(d-s=j(M-q)=jH_s(d)\). Consequently \[|H_s(d)|\le |d-s|=j|H_s(d)|,\] forcing \(H_s(d)=0\) and \(d=s\). The consistent Gaussian law must therefore satisfy (54).
It remains to solve (54). Set \[F(s)=\mathbb E_{N(s,s)}\tanh y\quad(s>0),\qquad F(0)=0.\] It is continuous at zero by bounded convergence. For \(s>0\), differentiating both the Gaussian mean and variance gives \[\begin{align*}
F'(s)
&=\mathbb E_{N(s,s)}\left[(\tanh y)'
+\tfrac12(\tanh y)''\right]\\
&=\mathbb E_{N(s,s)}[(1-m^2)(1-m)]\in[0,1].
\end{align*}\] The lower bound is pointwise; the upper bound follows from (57). Hence \(F\) is nonnegative and 1-Lipschitz on \([0,\infty)\), with \(F(s)\le s\). Consequently, \(s\mapsto t+jF(s)\) is a contraction of \([0,\infty)\) into itself and has exactly one fixed point \(s_*(t)\). The identity (56) then verifies both consistency equations for \(N(s_*(t),s_*(t))\). If a consistent law has \(s=0\), then \(0=t+jq\) forces \(t=q=0\), and hence the law is the point mass at zero and \(d=0\); this agrees with the solution just constructed.
Existence and uniqueness are now established, including the degenerate case. To obtain the dependence on time, compare the fixed-point equations at \(t,t'\): \[|s_*(t)-s_*(t')|
\le |t-t'|+j|s_*(t)-s_*(t')|,\] which proves the asserted Lipschitz bound. The size estimates follow from \(0\le F(s)\le s\), which gives \(t\le s_*(t)\le t/(1-j)\), and (56) gives \(q_*(t)=F(s_*(t))=(s_*(t)-t)/j\le t/(1-j)\). For \(t>0\), the variance \(s_*(t)\) is positive, so its Gaussian law is nondegenerate and \(q_*(t)>0\). The equality assertion follows from Lemma 7. ◻
Stable fields along Gaussian observations
The observation argument requires control of every approximate solution of the field equation at the fields visited by the observation process. We first prove this stability statement, including perturbations of the diagonal variance matrix, and then derive a deterministic existence, uniqueness, and inverse estimate for the field equation.
For a symmetric matrix \(J\), a field \(h_0\in\mathbb R^n\), and \(y\in\mathbb R^n\), define \[
\begin{split}
m(y)&=\tanh y,\qquad q(y)=\langle m(y)^2\rangle,\qquad
b(y)=1-q(y),\\
V_y&=D_{1-m(y)^2},\qquad B(y)=jb(y),\\
F_{J,h_0}(y)&=y-h_0+(B(y)I-J)m(y).
\end{split}
\tag{58}\] For fixed \(J\), abbreviate \(F_h=F_{J,h}\); we also suppress \(h\) when there is no ambiguity. Hyperbolic functions act coordinatewise. To measure stability while allowing the diagonal factor to vary, define, for each nonnegative diagonal matrix \(A\), \[
\mathsf H_J(y,A)
=I+A^{1/2}\left(B(y)I-J-\frac{2j}{n}m(y)m(y)^{\mathsf T}\right)A^{1/2}.
\tag{59}\] The matrix \(\mathsf H_J(y,V_y)\) is the symmetric version of the possibly nonsymmetric derivative \(F'_{J,h_0}(y)\). The external field \(h_0\) enters the residual \(F_{J,h_0}(y)\) but not \(\mathsf H_J(y,A)\). This separation will let us use the same matrix witness when the observation field moves.
Proposition 9 (Stability along observations). Fix \(j\in(0,1)\) and \(0<T<\infty\). There are constants \(\eta,c,\epsilon_A,\rho_0,a>0\), depending only on \(j,T\), and events \(\mathcal J_n\) for the zero-diagonal Gaussian matrix \(J\), such that \(\mathbb P(J\in\mathcal J_n)\longrightarrow1\) and the following assertion holds for every \(J\in\mathcal J_n\). Let \(X\sim\mu_0\), let \(B_0\) be an independent standard \(n\)-dimensional Brownian motion, and put \(Y(t)=tX+B_0(t)\). With conditional probability at least \(1-e^{-an}\), simultaneously for \(0\le t\le T\), \(y\in\mathbb R^n\), and all diagonal \(A\), \[
\left.
\begin{gathered}
\lVert F_{J,Y(t)}(y)\rVert\le2\rho_0\sqrt n,\\
0\le A\le(1+\eta)I,\qquad
\lVert A-V_y\rVert_{\mathrm F}\le\epsilon_A\sqrt n
\end{gathered}
\right\}
\quad\Longrightarrow\quad
\mathsf H_J(y,A)\succ cI.
\tag{60}\] We may require \(j(1+\eta)^2<1\) and \(\lVert J\rVert\le2\sqrt j+\theta\) on \(\mathcal J_n\), for any sufficiently small fixed \(\theta>0\) chosen in the proof.
Corollary 10 (Stopping formulation). For every \(J\in\mathcal J_n\) there is an observation stopping time \(\tau_0\) such that \[\mathbb P(\tau_0\le T\mid J)\le e^{-an},\] and, for every \(t\in[0,T]\) with \(t<\tau_0\), the implication (60) holds at \(Y(t)\) with residual threshold \(\rho_0\sqrt n\) in place of \(2\rho_0\sqrt n\). The stopping rule depends on \(J\) and the observations, and is independent of any test function.
Proof. We express failure of stability through a witness set that is fixed before observing the path. Let \(\mathcal W_J\) consist of pairs \((y,A)\) satisfying \(0\le A\le(1+\eta)I\) and \(\lVert A-V_y\rVert_{\mathrm F}\le\epsilon_A\sqrt n\) for which \(\mathsf H_J(y,A)\not\succ cI\). Define the least normalized residual among these witnesses by \[d_J(h)=\inf_{(y,A)\in\mathcal W_J}
\frac{\lVert F_{J,h}(y)\rVert}{\sqrt n},
\qquad \inf\varnothing=+\infty.\] If \(\mathcal W_J\) is nonempty, \(d_J\) is finite and \(n^{-1/2}\)-Lipschitz, because \(F_{J,h}(y)=F_{J,0}(y)-h\). Let \[\tau_0=\inf\{t\in[0,T]:d_J(Y(t))\le\rho_0\},\] with value \(+\infty\) when the set is empty. Continuity makes the first hitting time an observation stopping time. Before the hit, every violating witness has residual greater than \(\rho_0\sqrt n\). At a hit, the definition of the infimum supplies a violating witness with residual strictly below \(2\rho_0\sqrt n\); attainment of the infimum is unnecessary. Proposition 9 therefore bounds the hitting probability. ◻
The proof of Proposition 9 has three parts. A Gaussian integral bounds the weight of unstable configurations; its scalar rate is controlled by the entropy inequality, and its conditional matrix factor is controlled near the equality laws. A local-volume estimate then shows that any unstable approximate zero would contribute too much to that integral. Finally we transfer the resulting estimate to the original observation law, uniformly over the time interval.
A Gaussian integral and its scalar rate
We carry out the integral calculation first under a planted Gaussian law, whose relation to the original law will be proved when we assemble the path estimate. Under this law, \[
\widetilde J=W+\frac jn\mathbf 1\mathbf 1^{\mathsf T},\qquad
h_0=t\mathbf 1+\sqrt t\,g,
\tag{61}\] where \(W\) is centered GOE of scale \(j\) and \(g\) is independent standard Gaussian. We retain the diagonal of \(\widetilde J\) during this calculation. For a Borel set \(\mathcal E\subseteq\mathbb R^n\times\mathrm{Sym}_n\), define \[
I_{t,\sigma}(\mathcal E)
=(2\pi\sigma^2)^{-n/2}\int
\mathbf 1_{\{(y,\widetilde J)\in\mathcal E\}}
\exp\left\{-\frac{\lVert F_{\widetilde J,h_0}(y)\rVert^2}{2\sigma^2}
+\frac{njb(y)^2}{2}\right\}\,\mathrm dy,
\qquad \sigma>0.
\tag{62}\] The weight \(e^{njb(y)^2/2}\) compensates for the exponential scale of the local determinant: Lemma 14 will use the regularized determinant bound to turn an approximate zero into a lower bound for this integral. The sets to which we apply the integral depend on \(y,\widetilde J\) and are independent of \(h_0\). Thus the observation can be averaged while retaining the restriction to \(\mathcal E\).
For fixed \(y,t,\sigma\), introduce the scalar quantities \[
\begin{gathered}
m=\tanh y,\quad q=\langle m^2\rangle,\quad b=1-q,\quad B=jb,
\quad M=\langle m\rangle,\\
d=t+jM,\qquad s=t+\sigma^2+jq,\qquad S=s+jq,\\
a_y=\langle (y-d\mathbf 1)m\rangle,\qquad z=y-d\mathbf 1+Bm,\qquad
\psi=\frac{j(a_y-sb)^2}{2sS}.
\end{gathered}
\tag{63}\] Write \(\varphi_{d,s}\) for the density of \(N(d,s)\), with \(s\) denoting the variance, and define \[
G_{t,\sigma,n}(y)
=\prod_{i=1}^n\varphi_{d,s}(y_i)\,e^{n\psi}.
\tag{64}\] The Gaussian parameters in this expression depend on \(y\) through the preceding averages; the product is therefore not a fixed product density.
Lemma 11 (Conditional Gaussian reduction). With the preceding definitions, and with \(g'\) an independent standard Gaussian vector, \[
\mathbb EI_{t,\sigma}(\mathcal E)
=\int \sqrt{\frac sS}\,G_{t,\sigma,n}(y)
\mathbb P\left((y,\widetilde J)\in\mathcal E\,\middle|\,
Wm+\sqrt{t+\sigma^2}\,g'=z\right)\,\mathrm dy.
\tag{65}\] The conditional distribution in this identity is the usual Gaussian conditional distribution and is defined for every \(y\), since \(\sigma>0\).
Proof. We first evaluate the integral without its indicator. For fixed \(m\), the GOE covariance is \[\mathop{\mathrm{Cov}}(Wm)=jqI+\frac jnmm^{\mathsf T}.\] Convolution with the heat kernel of variance \(\sigma^2\) and with \(\sqrt t\,g\) therefore produces the Gaussian vector \(Wm+\sqrt{t+\sigma^2}\,g'\), of covariance \(sI+jmm^{\mathsf T}/n\). Its determinant is \(s^nS/s\), and its inverse is \[\frac1sI-\frac{j}{nsS}mm^{\mathsf T}.\] Moreover, \(F_{\widetilde J,h_0}(y)=z-Wm-\sqrt t\,g\). The Gaussian density at \(z\), multiplied by \(e^{njb^2/2}\), differs from \(\prod_i\varphi_{d,s}(y_i)\) by the prefactor \(\sqrt{s/S}\) and the exponential whose logarithm divided by \(n\) is \[-\frac{Ba_y}{s}-\frac{B^2q}{2s}
+\frac{j(a_y+Bq)^2}{2sS}+\frac{jb^2}{2}
=\frac{j(a_y-sb)^2}{2sS}.\] This calculation gives the identity for the unrestricted integral. To retain the indicator, condition the joint Gaussian pair formed by \(W\) and the convolved observation on the value \(z\) of the latter. The resulting conditional probability is exactly the factor in (65). All integrands are nonnegative, so Tonelli’s theorem justifies the interchange of integrations. ◻
The scalar entropy inequality controls the remaining integral in (65). To apply it, replace empirical averages in (63) by integration against a probability measure \(P\) with finite second moment, keeping the notation \(M,q,b,d,s,S,a_y,\psi\). Where needed, write \(a(P)\) instead of \(a_y\). Lemma 7 then gives \[
\psi(P,t,\sigma)-D(P\Vert N(d,s))\le0
\qquad (s>0).
\tag{66}\] At \(\sigma=0\), Lemma 8 identifies all equality cases: they form the curve \[P_t=N(s_*(t),s_*(t)),\qquad
s_*(t)=t+j\mathbb E_{P_t}\tanh y,\] with the point-mass convention \(P_0=\delta_0\). In particular, \(s_*(t)\le t/(1-j)\) and \(q(P_t)=(s_*(t)-t)/j\le t/(1-j)\).
Lemma 12 (Uniform scalar rate bounds). Fix \(j\in(0,1)\) and \(T<\infty\).
For every \(s_{\min}>0\) and \(L<\infty\), one can choose \(R<\infty\) such that, for all sufficiently large \(n\), uniformly over \(0\le t\le T\) and \(0<\sigma\le1\), \[\int_{\{s\ge s_{\min},\,\langle y^2\rangle>R^2\}}
G_{t,\sigma,n}(y)\,\mathrm dy\le e^{-Ln}.\] For every \(\varepsilon>0\), the unrestricted integral over \(\{s\ge s_{\min}\}\) is at most \(e^{\varepsilon n}\) for all sufficiently large \(n\).
Fix \(t_0>0\). Outside any prescribed open neighborhood of the equality curve \(\{(P_t,t):t_0\le t\le T\}\), there are \(a,\sigma_*>0\) such that the integral of \(G_{t,\sigma,n}\) is at most \(e^{-an}\), uniformly over \(t_0\le t\le T\) and \(0<\sigma\le\sigma_*\), for all sufficiently large \(n\). The neighborhood can impose weak closeness of the empirical law and closeness of any finite list of continuous, at most linearly growing averages, including all parameters in (63).
For every \(q_0>0\), there are \(t_0,\sigma_*,a>0\) such that \[\int_{\{q\ge q_0\}}G_{t,\sigma,n}(y)\,\mathrm dy\le e^{-an}\] uniformly over \(0\le t\le t_0\) and \(0<\sigma\le\sigma_*\), for all sufficiently large \(n\).
Fix \(t_0>0\) and \(\nu>0\). For all sufficiently large \(R_1\) there is \(a>0\) such that \[\int_{\{\langle y^2\mathbf 1_{\{|y|>R_1\}}\rangle>\nu\}}
G_{t,\sigma,n}(y)\,\mathrm dy\le e^{-an},\] uniformly over \(t_0\le t\le T\) and \(0<\sigma\le1\), for all sufficiently large \(n\).
All constants are independent of \(n,t,\sigma\) in the indicated ranges. In (ii), neighborhoods are interpreted on each fixed second-moment ball; the complement of a sufficiently large such ball is covered by (i).
Proof. We first obtain a Gaussian tail envelope, then prove a compact-set bound by a finite cover. Strict negativity away from the equality laws will give the second and third assertions; an exponential tilt will give the fourth.
Throughout, \(|d|\le T+j\) and \(s\le T+1+j\). Also \(a(P)^2\le q\int(y-d)^2\,\mathrm dP\) and \(jq/S\le1/2\). The elementary inequality \((r-u)^2\le\frac32r^2+3u^2\) consequently gives \[n\psi\le\frac{3jq}{4sS}\sum_i(y_i-d)^2+\frac{3nj}{2}
\le\frac{3}{8s}\sum_i(y_i-d)^2+\frac{3nj}{2}.\] On \(s\ge s_{\min}\), it follows that \[
G_{t,\sigma,n}(y)\le
\exp\left(Cn-c\sum_i y_i^2\right),
\tag{67}\] where \(c>0\) and \(C<\infty\) depend only on \(j,T,s_{\min}\). On \(\sum_i y_i^2>nR^2\), use half of the remaining Gaussian square for integration and the other half for the required exponential decay. This proves the first assertion of (i).
The estimates on bounded second moments follow from one compactness argument, which we now give in full. Put \[\mathcal P_R=\left\{P:\int y^2\,\mathrm dP\le R^2\right\}.\] Tightness and preservation of the second-moment bound under weak limits make this set weakly compact. Integration of a continuous function of at most linear growth is continuous on it: the second-moment bound makes the contribution of \(|y|>K\) uniformly \(O(K^{-1})\), while bounded truncations are weakly continuous. In particular, \(d,s,a(P),\psi\) are continuous on \(\mathcal P_R\times[0,T]\times[0,1]\) where \(s\ge s_{\min}\).
For every compact subset \(\mathcal C\) of that region, \[
\limsup_{n\to\infty}\frac1n\log
\sup_{t,\sigma}\int_{\{(P_y,t,\sigma)\in\mathcal C\}}
G_{t,\sigma,n}(y)\,\mathrm dy
\le\sup_{\mathcal C}\bigl\{\psi-D(P\Vert N(d,s))\bigr\},
\tag{68}\] where \(P_y=n^{-1}\sum_i\delta_{y_i}\); empty integrals are zero. To establish this bound, fix a strict finite upper bound \(r\) for its right-hand side. At each \(\xi=(P,t,\sigma)\in\mathcal C\), choose a bounded continuous test function \(\varphi_\xi\) from the variational formula for relative entropy so that \[\psi(\xi)-\int\varphi_\xi\,\mathrm dP
+\log\int e^{\varphi_\xi}\,\mathrm dN(d,s)<r.\] Such a test also exists when the entropy is infinite. Retain positive slack in the inequality and then choose a neighborhood of \(\xi\) on which continuity controls \(\psi\), the empirical average of \(\varphi_\xi\), and \(d,s\). Moreover, on \(\sum_i y_i^2\le nR^2\), the logarithmic density ratio between \(N(d',s')^{\otimes n}\) and the fixed law \(N(d,s)^{\otimes n}\) is at most \(n\) times a quantity tending uniformly to zero as \((d',s')\to(d,s)\); this follows directly by expanding their Gaussian log densities. For \(y\) in that neighborhood, insert \(e^{\sum_i\varphi_\xi(y_i)}\) and bound its expectation under the fixed product Gaussian. The integral on the neighborhood is then at most \(e^{nr}\), after using less than the reserved slack for the three continuity errors. Cover \(\mathcal C\) by finitely many such neighborhoods. An arbitrarily small enlargement of \(r\) absorbs their number and proves (68). Combining this bound with (66) on the ball and (67) outside it proves the second assertion of (i).
To obtain a strictly negative rate, observe that the same variational formula expresses relative entropy as a supremum of continuous functions. Thus it is jointly lower semicontinuous in \(P,d,s\) when \(s\) is bounded away from zero, and \(\psi-D\) is upper semicontinuous. At \(\sigma=0\), its zero set is precisely the equality curve identified above. On a compact set disjoint from any open neighborhood of that curve its supremum is strictly negative. The same remains true for all sufficiently small \(\sigma\): otherwise a convergent sequence with \(\sigma\downarrow0\) and rates approaching zero would give an equality point outside that neighborhood. Apply (68), with \(s_{\min}=t_0\), and then choose the second-moment tail rate in (i) strictly negative. This proves (ii).
Near time zero, the restriction on \(q\) supplies the positive variance needed for the same argument. For (iii), choose \(t_0\) so small that \(t_0/(1-j)<q_0/2\). The set \(q\ge q_0\) contains no equality point at \(\sigma=0\) and \(0\le t\le t_0\). On this set \(s\ge jq_0>0\), so the preceding compactness argument and the tail envelope apply without change, including at \(t=0\).
It remains to control the contribution of large coordinates in (iv). Fix \(t_0,\nu>0\), and take \(0<\lambda<1/[8(T+1+j)]\). Let \(h_{R_1}(y)=y^2\mathbf 1_{\{|y|>R_1\}}\). On a fixed second-moment ball, repeat the proof of (68), this time inserting \(e^{\lambda\sum_i h_{R_1}(y_i)}\). Begin with any small positive rate bound \(\varepsilon\) in that proof and its finite collection of tests \(\varphi_\xi\). Uniformly for \(|d|\le T+j\) and \(t_0\le s\le T+1+j\), \[\int\bigl(e^{\lambda h_{R_1}(y)}-1\bigr)\,\mathrm dN(d,s)(y)
\longrightarrow0\qquad(R_1\to\infty),\] by Gaussian tails; the choice of \(\lambda\) leaves a uniformly negative quadratic exponent. The same holds with the integrand multiplied by any of the finitely many bounded \(e^{\varphi_\xi}\). The additional logarithmic moment can therefore be made arbitrarily small in every member of the finite cover. For sufficiently large \(R_1\), the tilted integral on the ball is at most \(e^{2\varepsilon n}\). On \(\langle h_{R_1}(y)\rangle>\nu\), remove the tilt at cost \(e^{-\lambda\nu n}\). Choosing \(2\varepsilon<\lambda\nu/2\) gives a strictly negative rate there. The integral off the ball is estimated without the tilt, using (i) with a larger negative rate. This proves (iv). ◻
The conditional Hessian
The scalar bounds concentrate the integral near the equality curve. We now show that, there, the conditional Hessian is positive with an exponential probability bound. The conclusion is simultaneous over all admissible nearby diagonal matrices, which will be needed to exclude every unstable approximate zero.
Lemma 13 (Conditional stability near the scalar equality curve). Fix \(0<t_0\le T\), \(R<\infty\), and \(\eta>0\) with \(j(1+\eta)^2<1\). There are \(c_*,\epsilon_*,a_*,\sigma_*>0\), a neighborhood \(\mathcal U\) of \(\{(P_t,t):t_0\le t\le T\}\), a tail tolerance \(\nu>0\), and a finite \(R_1\) such that the following holds. Suppose \[\langle y^2\rangle\le R^2,\qquad (P_y,t)\in\mathcal U,\qquad
\langle y^2\mathbf 1_{\{|y|>R_1\}}\rangle\le\nu,
\qquad t_0\le t\le T,\quad0<\sigma\le\sigma_*.\] Under the Gaussian conditioning in (65), with probability at least \(1-e^{-a_*n}\), simultaneously for all diagonal \(A\) satisfying \[0\le A\le(1+\eta)I,\qquad
\lVert A-V_y\rVert_{\mathrm F}\le\epsilon_*\sqrt n,\] one has \(\mathsf H_{\widetilde J}(y,A)\succeq c_*I\). All constants are uniform over the displayed ranges and sufficiently large \(n\). The tail cutoff \(R_1\) may be required to be arbitrarily large after choosing \(\nu\); \(\epsilon_*\) is then decreased if necessary.
Proof. We first compute the Gaussian regression and the finite-rank correction at an equality law. We then show that the empirical-law and diagonal hypotheses preserve its positive margin, uniformly over all admissible diagonals.
Gaussian regression and whitening.
In a sufficiently small neighborhood of the equality curve, \(s\) and \(q\) have positive lower bounds depending on \(j,t_0,T\). Indeed \(s_*(t)\ge t_0\), and \(q(P_t)=\mathbb E_{P_t}\tanh^2 y\) is positive and continuous there. Set \[u=\frac{m}{\sqrt{nq}},\qquad
Z_1=\frac{y-d\mathbf 1}{\sqrt{ns}},\qquad P_u=I-uu^{\mathsf T}.\] The norms of \(\mathbf 1/\sqrt n,u,Z_1\) are bounded uniformly by the assumed second-moment bound. Write \[W=P_uWP_u+uw^{\mathsf T}+wu^{\mathsf T}+\alpha uu^{\mathsf T}.\] The three blocks in this decomposition are independent before conditioning, with \(\mathop{\mathrm{Cov}}(w)=(j/n)P_u\) and \(\mathop{\mathrm{Var}}(\alpha)=2j/n\). The projection of the observation in (65) perpendicular to \(u\) observes \(\sqrt{nq}\,w\) with independent noise of covariance \((t+\sigma^2)P_u\); projecting along \(u\) observes \(\sqrt{nq}\,\alpha\) with noise variance \(t+\sigma^2\). The elementary Gaussian conditional-mean and conditional-variance formulas therefore give the following representation, using a fresh centered GOE \(W_0\): \[
\begin{split}
P_uWP_u&=P_uW_0P_u,\\
w&=\bar w+\kappa P_uW_0u,\qquad
\bar w=j\sqrt{q/s}\,Z_1-\frac{ja_y}{s}u,\qquad
\kappa=\sqrt{\frac{t+\sigma^2}{s}},\\
\alpha&=\frac{2j(a_y+Bq)}S
+\sqrt{\frac{t+\sigma^2}S}\,u^{\mathsf T}W_0u.
\end{split}
\tag{69}\] For example, the perpendicular conditional mean is \(j\sqrt{q/s}\,P_uZ_1\), which equals the displayed \(\bar w\) because \(u^{\mathsf T}Z_1=a_y/\sqrt{qs}\). The parallel conditional mean follows from \(u^{\mathsf T}z=\sqrt{n/q}(a_y+Bq)\). The parallel residual variance is \((2j/n)(t+\sigma^2)/S\). These conditional means and residual variances verify all three blocks of (69).
For the moment, fix a deterministic admissible \(A\). We compare the conditional Hessian with the positive matrix built from the fresh GOE by setting \[\chi=\frac1n\mathop{\mathrm{Tr}}A,\qquad B_A=j\chi,\qquad
S_A=I+B_AA-A^{1/2}W_0A^{1/2},\qquad
\mathcal Y(p)=S_A^{-1/2}A^{1/2}p.\] Lemma 3 gives \(S_A\succeq c_0I\) except with exponentially small probability, where \(c_0>0\) is uniform in \(A\). Let \(G_A=A^{1/2}S_A^{-1}A^{1/2}\). Lemma 5, applied to the three deterministic vectors \(\mathbf 1/\sqrt n,u,Z_1\), says that their Gram entries after \(\mathcal Y\) are as close as desired to their \(A\)-weighted Gram entries, and that the extra vector \(\mathcal Y(W_0u)\) has Gram entries as close as desired to those of \[
B_A\mathcal Y(u)+\sqrt{B_A}\,N,
\tag{70}\] where \(N\) is a unit vector orthogonal to the three deterministic images in this comparison. Specifically, its mixed entries are \(B_Ap^{\mathsf T}Au\) and its squared norm is \(B_A+B_A^2u^{\mathsf T}Au\). Here \(N\) describes only Gram entries; this notation does not assert a coupling with an additional random or deterministic vector. The comparison concerns finite Gram matrices; the associated symmetric finite-rank eigenvalues depend continuously on their entries, as we verify when returning to the empirical parameters. Each fixed Gram tolerance has failure probability at most \(Ce^{-a n}\) uniformly over the bounded deterministic vectors and over \(A\). Also \(u^{\mathsf T}W_0u\) is smaller than any fixed tolerance with exponentially high probability, since its variance is \(2j/n\).
The correction at an equality law.
We now compute the finite-rank perturbation in the ideal Gram comparison. Set \(\sigma=0\), take \(s=s_*(t)\), and replace each weighted average \(n^{-1}\sum_i A_{ii}f(y_i)\) by \(\mathbb E_{N(s,s)}[\mathop{\mathrm{sech}}^2 y\,f(y)]\). Evaluate all other empirical averages under the same law. Thus \(a_y=sb\), \(B_A=B=jb\), and \(\kappa^2=1-jq/s\). Write \[X_1=\mathcal Y(\mathbf 1/\sqrt n),\qquad
U=\mathcal Y(u),\qquad Z_2=\mathcal Y(Z_1),\] using the Gram interpretation just described. Subtracting \(W_0\) from (69) gives exactly \[
\begin{split}
W-W_0={}&u\bar w^{\mathsf T}+\bar w u^{\mathsf T}
-(1-\kappa)\{u(W_0u)^{\mathsf T}+(W_0u)u^{\mathsf T}\}\\
&+\{\alpha+(1-2\kappa)u^{\mathsf T}W_0u\}uu^{\mathsf T}.
\end{split}
\tag{71}\] At equality, \(\bar w=j\sqrt{q/s}\,Z_1-Bu\) and \(\alpha=2B\) after discarding the vanishing Gaussian scalar. Whiten by \(S_A\), and use (70) with \(B_A=B\). The perturbation to subtract from the identity consists of the planted term \(jX_1X_1^{\mathsf T}\), the term \(2jqUU^{\mathsf T}\), and the whitened version of (71). The terms \(-2BUU^{\mathsf T}\) and \(+2BUU^{\mathsf T}\) cancel. Eliminating the orthogonal direction \(N\) by Schur complement adds \((1-\kappa)^2BUU^{\mathsf T}\) to the perturbation. Its total \(UU^{\mathsf T}\) coefficient becomes \[2jq-2(1-\kappa)B+(1-\kappa)^2B
=jq(2-B/s).\] Consequently the Schur complement is the identity minus \[
j\{X_1X_1^{\mathsf T}+(2-B/s)M_1M_1^{\mathsf T}
+M_1G_1^{\mathsf T}+G_1M_1^{\mathsf T}\},
\qquad M_1=\sqrt q\,U,\quad G_1=Z_2/\sqrt s.
\tag{72}\]
To prove positivity of this Schur complement, it suffices to examine its three-vector Gram matrix. Put \(e=\mathbb E_{N(s,s)}[m^2(1-m^2)]\). The Gram matrix of \((X_1,M_1,G_1)\) equals \[
H=\begin{pmatrix}
b&e&-2e\\
e&e&b-3e\\
-2e&b-3e&b/s-2b+6e
\end{pmatrix}.
\tag{73}\] We derive the entries of this matrix explicitly. The density of \(N(s,s)\) satisfies \(p(y)/p(-y)=e^{2y}\). Pairing \(y\) and \(-y\) shows that every integrable odd function \(f\) satisfies \(\mathbb Ef(y)=\mathbb E[f(y)\tanh y]\). This gives \(\mathbb E[m(1-m^2)]=e\). Gaussian integration by parts gives \[\frac1s\mathbb E[(y-s)(1-m^2)]=-2e,\qquad
\frac1s\mathbb E[(y-s)m(1-m^2)]=b-3e.\] Applying it twice, and using \((1-m^2)''=4m^2(1-m^2)-2(1-m^2)^2\), gives the final diagonal entry \(b/s-2b+6e\). The remaining entries are immediate, proving (73).
The coefficient matrix in (72) is \[C=j\begin{pmatrix}1&0&0\\0&2-B/s&1\\0&1&0\end{pmatrix}.\] Set \(d_0=1-jb\) and \(g_1=je\). The first and third columns of \(I-HC\) are \((d_0,-g_1,2g_1)^{\mathsf T}\) and \((-g_1,-g_1,d_0+3g_1)^{\mathsf T}\). Adding \(B/s\) times the third column to the second makes the second column \((0,d_0+g_1,0)^{\mathsf T}\). Hence \[
\det(I-HC)=(d_0+2g_1)(d_0+g_1)^2>0.
\tag{74}\] The same determinant is obtained from the identity minus (72) on any ambient space containing its three vectors. To determine its sign as a quadratic form, first subtract only \(jX_1X_1^{\mathsf T}\); the result is positive definite because \(j\lVert X_1\rVert^2=jb<1\). The remaining two-vector perturbation has at most one positive eigenvalue, since its coefficient block \(j\left(\begin{smallmatrix}2-B/s&1\\1&0\end{smallmatrix}\right)\) has one positive and one negative eigenvalue. The resulting matrix can consequently have at most one nonpositive direction. The positive determinant in (74) excludes both a zero eigenvalue and a single negative eigenvalue. Hence the matrix is positive definite. The Schur complement then also gives positivity before \(N\) was eliminated.
Perturbing the scalar law and the diagonal.
The equality calculation supplies the positive margin we need. We first explain why the hypotheses put the actual weighted Gram entries near these equality values.
Each deterministic vector has coordinates \(n^{-1/2}\) times one of \(1,m/\sqrt q,(y-d)/\sqrt s\). Products of two such coordinate functions are bounded by \(C(1+y^2)\). For any such product \(f\), the tail and Frobenius bounds imply \[
\left|\frac1n\sum_i(A_{ii}-(V_y)_{ii})f(y_i)\right|
\le C_{R_1}\frac{\lVert A-V_y\rVert_{\mathrm F}}{\sqrt n}+C\nu.
\tag{75}\] On \(|y_i|\le R_1\), use Cauchy–Schwarz and boundedness of \(f\); outside, use \(|A_{ii}-(V_y)_{ii}|\le2+\eta\) and the tail hypothesis, with \(R_1\ge1\). The corresponding \(V_y\)-weighted functions are bounded and continuous, even when multiplied by \(y^2\), because \(\mathop{\mathrm{sech}}^2 y\) decays exponentially. Their averages are therefore close to the equality-law averages in a sufficiently small weak neighborhood. The parameters \(d,s,q,a_y\) are continuous on the second-moment ball, so their denominators and coefficients are close as well. Finally \(|B_A-B|\le j\lVert A-V_y\rVert_{\mathrm F}/\sqrt n\).
We retain the positivity uniformly for \(t\in[t_0,T]\) and under small Gram errors. If \(Q\) has finitely many columns, the nonzero eigenvalues of \(QCQ^{\mathsf T}\) agree with those of \((Q^{\mathsf T}Q)^{1/2}C(Q^{\mathsf T}Q)^{1/2}\), including multiplicity. The largest eigenvalue, with zero adjoined, is therefore a continuous function of the Gram matrix and the coefficients. All equality parameters vary continuously on the compact interval \([t_0,T]\), and \(s,q\) stay positive there. The exact calculation thus gives a uniform positive whitened margin. Choose a fixed tolerance \(\xi>0\) on coefficients and Gram entries small enough to preserve, say, half of it. Lemmas 3 and 5, the Gaussian scalar estimate, and (75) show that this tolerance is achieved for each fixed admissible \(A\) with failure probability at most \(Ce^{-a n}\), after choosing the scalar neighborhood, tail tolerance, \(\sigma_*\), and Frobenius distance sufficiently small. Returning from whitening multiplies the margin by at least \(c_0\). The omitted mass \((B-B_A)A\) has norm at most \(j(1+\eta)\lVert A-V_y\rVert_{\mathrm F}/\sqrt n\) and can be absorbed by decreasing that distance again.
At this stage the margin is proved for each fixed \(A\). To make it simultaneous, we use that its failure exponent is uniform over the bounded deterministic vector norms. Those bounds were fixed before choosing the Frobenius distance. Also (69) implies \(\lVert W\rVert\le K_1\) with failure \(Ce^{-a_1n}\) for some uniform \(K_1\): the norms of the deterministic means are bounded, and the remaining terms are controlled by \(\lVert W_0\rVert\). On this event, \(A\mapsto\mathsf H_{\widetilde J}(y,A)\) is uniformly \(1/2\)-Hölder in diagonal sup norm. Indeed \(\lVert A^{1/2}-\widehat A^{1/2}\rVert\le
\lVert A-\widehat A\rVert^{1/2}\), while the middle factor in (59) has bounded norm.
Choose a fixed mesh size \(\zeta>0\) small enough that the Hölder error is below one quarter of the margin. Given \(A\) with normalized Frobenius distance at most \(\epsilon\) from \(V_y\), let \(\mathcal S_A=\{i:|A_{ii}-(V_y)_{ii}|>\zeta\}\). It has at most \(\alpha n\) elements, where \(\alpha=\epsilon^2/\zeta^2\). Outside \(\mathcal S_A\) replace \(A_{ii}\) by \((V_y)_{ii}\); on \(\mathcal S_A\), round it within \(\zeta\) to a grid of \([0,1+\eta]\), including the endpoints. The resulting \(\widehat A\) obeys \[\lVert A-\widehat A\rVert\le\zeta,\qquad
\lVert \widehat A-V_y\rVert_{\mathrm F}\le2\epsilon\sqrt n.\] For fixed \(y\), these approximants are deterministic and number at most \[\sum_{k\le\alpha n}\binom nk N_\zeta^k,
\qquad N_\zeta\le C(1+1/\zeta).\] Its logarithm divided by \(n\) tends to zero as \(\alpha\downarrow0\); for instance it is at most \(-\alpha\log\alpha-(1-\alpha)\log(1-\alpha)
+\alpha\log N_\zeta+o(1)\) when \(\alpha<1/2\). Choose \(\epsilon_*\) so small that \(2\epsilon_*\) meets all moment tolerances and this counting rate is less than half the fixed-\(A\) failure exponent. A union bound, followed by the Hölder comparison, proves the desired uniform margin.
For later use, record the order that preserves the concentration exponents. First fix \(R,t_0,T,\eta\), the limiting margin, and the Gram tolerance \(\xi\). This fixes the vector norm bounds, conditional concentration exponents, \(K_1\), and the mesh \(\zeta\). Next take the equality neighborhood and tail tolerance \(\nu\) small. Choose any sufficiently large finite \(R_1\), then take \(\epsilon_*\) small enough for (75) and the union bound. Finally decrease \(\sigma_*\) so all scalar coefficients lie within the required tolerance. No concentration exponent is lost when \(\epsilon_*\) is decreased. This also proves the last assertion of the lemma. ◻
From a small integral to exclusion of approximate zeros
The conditional estimate controls all nearby diagonals for each fixed \(y\) satisfying the scalar hypotheses. Combined with the scalar rate bounds, it will make the expected integral over unstable configurations exponentially small. To exclude approximate zeros at any \(y\), the next deterministic lemma gives the complementary lower bound: an approximate zero satisfying a regularized determinant bound contributes a definite exponential amount to the integral over a surrounding ball. The regularization makes a sign assumption on the Jacobian unnecessary.
Lemma 14 (Local volume at an approximate zero). Fix \(j,K<\infty\), \(r>0\), and \(\varepsilon>0\). Suppose \(\lVert \widetilde J\rVert\le K\), and a set \(E\subseteq\mathbb R^n\) contains the ball \(\{y':\lVert y'-y\rVert\le r\sqrt n\}\). Put \[L=F'_{\widetilde J,h_0}(y)
=I+(BI-\widetilde J)V_y-\frac{2j}{n}mm^{\mathsf T}V_y,\qquad
Q=L^{\mathsf T}L+\gamma I.\] If \[
\frac1n\log\det Q\le jb(y)^2+\varepsilon_1,\qquad
\lVert F_{\widetilde J,h_0}(y)\rVert\le\rho\sqrt n,
\tag{76}\] then \[
(2\pi\sigma^2)^{-n/2}\int_E
e^{-\lVert F_{\widetilde J,h_0}(y')\rVert^2/(2\sigma^2)+njb(y')^2/2}\,\mathrm dy'
\ge e^{-\varepsilon n}
\tag{77}\] for all sufficiently large \(n\), provided the constants are chosen in this order: first \(\varepsilon_1>0\) sufficiently small in terms of \(\varepsilon\); then any fixed \(\gamma>0\); then \(\sigma>0\) sufficiently small in terms of \(j,K,r,\varepsilon,\gamma\); and finally \(\rho>0\) sufficiently small in terms of those constants and \(\sigma\). The bound is uniform in \(y,h_0,n\) under the stated hypotheses.
Proof. We integrate over a Gaussian ellipsoid centered at the approximate zero. Its covariance is chosen to control the linearized residual. Let \(g''\) be standard Gaussian and put \(\Delta=\sigma Q^{-1/2}g''\), and restrict to \(\mathcal B=\{\lVert g''\rVert\le2\sqrt n\}\). For \(2\sigma/\sqrt\gamma\le r\), all resulting points \(y+\Delta\) belong to \(E\). Changing variables in the integral gives the lower bound \[
(\det Q)^{-1/2}\mathbb P(\mathcal B)
\mathbb E\left[\exp\left\{
\frac{\lVert g''\rVert^2}{2}
-\frac{\lVert F(y+\Delta)\rVert^2}{2\sigma^2}
+\frac{njb(y+\Delta)^2}{2}\right\}\,\middle|\,\mathcal B\right].
\tag{78}\] Here and for the remainder of this proof \(F=F_{\widetilde J,h_0}\). The conditioning has probability tending to one, and all estimates below remain valid under it with a uniform factor, since \(\mathbb P(\mathcal B)\ge1/2\) for large \(n\).
We must control the nonlinear residual on this ellipsoid uniformly in its center. Coordinatewise Taylor expansion gives \[m(y+\Delta)=m+V_y\Delta+r_m,\qquad
|(r_m)_i|\le C\Delta_i^2.\] Writing \(b(y+\Delta)=b+\ell_b+r_b\), its linear part is \(\ell_b=-2\langle m(1-m^2)\Delta\rangle\), and \(|r_b|\le C\lVert \Delta\rVert^2/n\). Thus, when the product \(jb,m\) in \(F\) is expanded, its remainder is bounded in norm by \[C\left(\lVert r_m\rVert+\frac{\lVert \Delta\rVert^2}{\sqrt n}\right).\] For the product of first differences use \(|b(y+\Delta)-b(y)|\le C\lVert \Delta\rVert/\sqrt n\) and \(\lVert m(y+\Delta)-m(y)\rVert\le\lVert \Delta\rVert\). The remaining term \(-\widetilde Jm\) has remainder bounded by \(K\lVert r_m\rVert\). Since the covariance of \(\Delta\) is bounded above by \(\sigma^2I/\gamma\), Gaussian fourth moments give \[\mathbb E\sum_i\Delta_i^4\le Cn\sigma^4/\gamma^2,\qquad
\mathbb E\lVert \Delta\rVert^4\le Cn^2\sigma^4/\gamma^2.\] Combining these Taylor estimates with the Gaussian moments, and writing \(R_F=F(y+\Delta)-F(y)-L\Delta\), yields \[
\mathbb E[\lVert R_F\rVert^2\mid\mathcal B]
\le Cn\sigma^4/\gamma^2,\qquad
\mathbb E[|b(y+\Delta)^2-b(y)^2|\mid\mathcal B]
\le C\sigma/\sqrt\gamma.
\tag{79}\] All constants depend only on \(j,K\).
The choice of ellipsoid now controls the linear term: the singular values of \(LQ^{-1/2}\) are at most one. Hence \(\lVert L\Delta\rVert\le\sigma\lVert g''\rVert\) pointwise, including on \(\mathcal B\). Expanding \(F(y+\Delta)=F(y)+L\Delta+R_F\) and applying Cauchy–Schwarz to its cross terms yields \[\begin{align*}
&\frac1n\mathbb E\left[
\frac{\lVert g''\rVert^2}{2}
-\frac{\lVert F(y+\Delta)\rVert^2}{2\sigma^2}
\,\middle|\,\mathcal B\right]\\
&\hspace{1cm}\ge
-C\left\{
\frac\rho\sigma+\frac\sigma\gamma
+\left(\frac\rho\sigma\right)^2
+\left(\frac\sigma\gamma\right)^2
+\frac\rho\gamma\right\}.
\end{align*}\] The last mixed term may also be absorbed into the two squares. Apply Jensen’s inequality to the conditional expectation in (78). Its logarithm divided by \(n\) is at least \[-\frac{1}{2n}\log\det Q+\frac{jb(y)^2}{2}
-C\left\{\frac\rho\sigma+\frac\sigma\gamma
+\left(\frac\rho\sigma\right)^2
+\left(\frac\sigma\gamma\right)^2
+\frac\sigma{\sqrt\gamma}\right\}+o(1).\] The determinant hypothesis cancels \(jb(y)^2/2\), with a loss of at most \(\varepsilon_1/2\). Choose the constants in the order in the statement; the other errors then have total size smaller than \(\varepsilon\). This proves the lower bound. ◻
Proof of the observation-path statement
Proof of Proposition 9. We assemble the scalar, matrix, and local-volume estimates under the planted law (61). After excluding unstable approximate zeros at a fixed time, we remove the auxiliary diagonal, control the whole path by a fixed mesh, and undo the planting.
Let \(J\) be \(\widetilde J\) with its diagonal zeroed. Choose \(\eta,\theta,q_0>0\) sufficiently small that \(j(1+\eta)^2<1\) and, whenever \(q\le2q_0\), \[
I+(j-3jq-2\sqrt j-2\theta)A\succeq c_{\mathrm{sm}}I
\qquad(0\le A\le(1+\eta)I)
\tag{80}\] for some \(c_{\mathrm{sm}}>0\). These choices are possible because at \(\eta=\theta=q=0\) the lower scalar bound is \(1+j-2\sqrt j=(1-\sqrt j)^2>0\). On \[
\lVert J\rVert\le2\sqrt j+\theta,\qquad
\max_i|\widetilde J_{ii}|\le\theta,
\tag{81}\] we have \(\lVert \widetilde J\rVert\le2\sqrt j+2\theta\). Since \(mm^{\mathsf T}/n\preceq qI\), (80) implies \(\mathsf H_{\widetilde J}(y,A)\succeq c_{\mathrm{sm}}I\) for all \(q\le2q_0\). Choose \(t_0\in(0,T]\) small enough for Lemma 12(iii).
For the remaining times \(t\in[t_0,T]\), the conditional Hessian lemma provides the margin near the equality curve. Choose a second-moment ball whose complement has a fixed negative rate, using Lemma 12(i) with \(s_{\min}=t_0\). Apply Lemma 13 on this ball. Take its equality neighborhood and tail tolerance, choose its tail cutoff \(R_1\) sufficiently large for Lemma 12(iv), and then choose its diagonal-distance radius. Decrease \(\sigma_*>0\) so that Lemma 12(ii) excludes the complement of that equality neighborhood, and so that the small-time scalar bound also holds. The scalar regions excluded by these choices all have fixed strictly negative rates, uniformly for \(0<\sigma\le\sigma_*\). On the region left over, the conditional matrix margin fails with probability at most \(e^{-a_*n}\).
Let \(\epsilon_*>0\) denote the full diagonal-distance radius just chosen. Choose \(c_E>0\) smaller than both \(c_*/2\) and \(c_{\mathrm{sm}}/2\). For \(t\ge t_0\), define \[E_t(W)=\left\{y:\exists A,\quad
0\le A\le(1+\eta)I,\quad
\lVert A-V_y\rVert_{\mathrm F}\le\epsilon_*\sqrt n,\quad
\lambda_{\min}(\mathsf H_{\widetilde J}(y,A))\le c_E\right\}.\] For \(0\le t<t_0\), instead put \(E_t(W)=\{y:q(y)\ge q_0\}\). There is \(a_0>0\) such that \[
\mathbb EI_{t,\sigma}(E_t)\le e^{-a_0n}
\quad(0\le t\le T, 0<\sigma\le\sigma_*)
\tag{82}\] for all sufficiently large \(n\), uniformly in \(t,\sigma\). For small times, this follows from the scalar bound and Lemma 11. For larger times, split the conditional integral into the excluded scalar regions and their complement. The excluded regions have negative rates. On the complement multiply the conditional probability \(e^{-a_*n}\) by the unrestricted scalar integral, whose rate is at most any prescribed positive number, by Lemma 12(i). Taking that number below \(a_*/2\) proves (82). Finitely many pieces can be absorbed by decreasing \(a_0\).
We now use local volume to exclude every unstable approximate zero. An individual witness will force a ball of the excluded set to contribute too much to the integral. Suppose \[
\lVert F_{\widetilde J,h_0}(y)\rVert\le\rho\sqrt n,\qquad
\lVert A-V_y\rVert_{\mathrm F}\le\tfrac12\epsilon_*\sqrt n,\qquad
\lambda_{\min}(\mathsf H_{\widetilde J}(y,A))\le\tfrac12c_E
\tag{83}\] for some \(0\le A\le(1+\eta)I\), and assume (81). There is a constant \(r>0\), independent of this witness and of \(n,t\), such that \(\lVert y'-y\rVert\le r\sqrt n\) implies \(y'\in E_t(W)\). For \(t\ge t_0\), the same \(A\) witnesses instability throughout a sufficiently small ball. Indeed, \[\lVert V_{y'}-V_y\rVert_{\mathrm F}\le2\lVert y'-y\rVert,\qquad
|b(y')-b(y)|\le2\lVert y'-y\rVert/\sqrt n,\] and \[\lVert m(y')m(y')^{\mathsf T}/n-m(y)m(y)^{\mathsf T}/n\rVert
\le2\lVert y'-y\rVert/\sqrt n\] preserve the full diagonal-distance constraint and the bound \(\lambda_{\min}\le c_E\) if \(r\) is small enough. For \(t<t_0\), (80) and the chosen \(c_E\) imply \(q(y)>2q_0\); the same Lipschitz bound for \(q\) then gives \(q(y')\ge q_0\) for a smaller \(r\).
Set \(L=F'_{\widetilde J,h_0}(y)\). It is a bounded-rank, bounded-norm perturbation of \[I+(BI-W)V_y:
\quad L=I+(BI-W)V_y
-\frac jn\mathbf 1\mathbf 1^{\mathsf T}V_y
-\frac{2j}{n}mm^{\mathsf T}V_y.\] Here \(n^{-1}\mathop{\mathrm{Tr}}V_y=b(y)\), so the determinant rate in Lemma 6 matches the weight in (62). That lemma says that, for any chosen \(\varepsilon_1>0\), one may choose a fixed \(\gamma>0\) sufficiently small so that, except on an event of probability \(Ce^{-a_1n}\), simultaneously at every \(y\), \[\frac1n\log\det(L^{\mathsf T}L+\gamma I)
\le jb(y)^2+\varepsilon_1.\] The lemma’s rank-perturbation allowance applies here with two rank-one terms of norms at most \(j\) and \(2j\). Choose \(\varepsilon<a_0/2\), choose \(\varepsilon_1\) for Lemma 14, then choose \(\gamma\) by the log-determinant lemma. Next fix \(\sigma\in(0,\sigma_*]\) sufficiently small for the local-volume lemma, and finally fix \(\rho>0\) sufficiently small for it. A witness (83) on the determinant event forces \(I_{t,\sigma}(E_t)\ge e^{-\varepsilon n}\). Markov’s inequality and (82) show that the probability of such a witness, under (81), is at most \(Ce^{-a_2n}\), uniformly in \(t\in[0,T]\). Here the uniformity of the negative rates permits the later choice of \(\sigma\) required by the local-volume estimate.
The integral argument used a full GOE matrix. We now pass to its zero-diagonal version, retaining slack in both residual and Hessian bounds. Choose a fixed \(\delta_D>0\) so small that \[\delta_D\le\theta,\qquad
\delta_D\le\rho/2,\qquad
(1+\eta)\delta_D\le c_E/4.\] The diagonal entries of \(\widetilde J\) are independent \(N(j/n,2j/n)\), so \(\mathbb P(\max_i|\widetilde J_{ii}|>\delta_D)\le Ce^{-a_Dn}\) for large \(n\). On its complement, \[\frac{\lVert F_{\widetilde J,h_0}(y)-F_{J,h_0}(y)\rVert}{\sqrt n}
\le\delta_D,\qquad
\lVert \mathsf H_{\widetilde J}(y,A)-\mathsf H_J(y,A)\rVert
\le(1+\eta)\delta_D.\] Thus a zero-diagonal witness with residual at most \(\rho\sqrt n/2\), radius \(\epsilon_*\sqrt n/2\), and Hessian minimum at most \(c_E/4\) would give a witness (83). We have proved an exponentially small fixed-time failure probability for the zero-diagonal matrix, restricted only to \(\lVert J\rVert\le2\sqrt j+\theta\).
To pass from a fixed time to the whole interval, we control observation increments on a deterministic mesh. Take a mesh of size \(\Delta>0\) that includes both endpoints. For a scalar Brownian motion, the reflection bound implies that \(\sup_{0\le u\le\Delta}|B(u)|^2/\Delta\) has an exponential moment at a fixed positive argument. The coordinate Brownian motions are independent. Exponential Markov applied to the sum of their squared suprema consequently gives, for each desired \(\kappa_0>0\), a mesh \(\Delta\) small enough that \[\mathbb P\left(\max_k\sup_{t_k\le t\le t_{k+1}}
\frac{\lVert B_0(t)-B_0(t_k)\rVert}{\sqrt n}>\kappa_0\right)
\le Ce^{-a_3n}.\] The deterministic drift contributes at most \(\Delta\sqrt n\) to each increment. Choose the mesh so the full normalized increment is at most \(\rho/4\). Set \(\rho_0=\rho/8\), \(\epsilon_A=\epsilon_*/2\), and \(c=c_E/4\). A pathwise violation of (60) then gives a fixed-time violation at a mesh point: the Hessian and its witness do not use the field, and the residual increases by at most \(\rho\sqrt n/4\). A union bound over the fixed finite mesh proves an annealed path failure bound \(Ce^{-a_4n}\) under the planted law, restricted to \(\lVert J\rVert\le2\sqrt j+\theta\).
We have obtained a path estimate under planting. To recover the original disorder law, introduce the partition-function density \[Z(J)=\sum_x e^{x^{\mathsf T}Jx/2},\qquad
L(J)=Z(J)/\mathbb EZ(J).\] Size-bias the original joint law of \(J,X\), with \(X\sim\mu_0\), by \(L(J)\). Its density in \(J,X\) is proportional to \(e^{x^{\mathsf T}Jx/2}\) times the original Gaussian density. Gaussian tilting makes \(X\) uniform and, conditionally on \(X\), makes the off-diagonal entries of \(J\) independent with means \(jX_iX_j/n\) and variance \(j/n\). Add independent diagonal entries with law \(N(j/n,2j/n)\)—a centered GOE diagonal plus \(jI/n\)—and conjugate by \(D_X\); the resulting matrix and observation have exactly the law (61). All witness conditions are equivariant under this conjugation, because \(m(D_Xy)=D_Xm(y)\), \(V_{D_Xy}=V_y\), and diagonal \(A\) commutes with \(D_X\). Conditional observation paths given \(J\) are unchanged by size bias.
Undoing this change of law requires a lower bound for \(L(J)\) on typical original matrices. For independent uniform signs \(\varepsilon_i\), Gaussian averaging gives \[\frac{\mathbb EZ^2}{(\mathbb EZ)^2}
=e^{-j/2}\mathbb E\exp\left\{\frac j{2n}
\left(\sum_i\varepsilon_i\right)^2\right\}
\le(1-j)^{-1/2}.\] To obtain the inequality, introduce an independent standard Gaussian \(g\) and use \(\mathbb E_\varepsilon e^{\sqrt{j/n}g\sum_i\varepsilon_i}
=\cosh(\sqrt{j/n}g)^n\le e^{jg^2/2}\). Paley–Zygmund then yields a constant \(p_j>0\) with \(\mathbb P(Z\ge\mathbb EZ/2)\ge p_j\). As a function of the independent standard off-diagonal Gaussians, \(\log Z\) has gradient squared at most \((j/n)\binom n2\le jn/2\). Gaussian concentration about a median therefore has scale \(O(\sqrt n)\). The preceding positive-probability lower bound implies that this median is at least \(\log\mathbb EZ-O_j(\sqrt n)\): otherwise its upper concentration tail would be smaller than \(p_j\). It follows that, for every fixed \(\delta>0\), \[
\mathbb P\{L(J)<e^{-\delta n}\}\longrightarrow0.
\tag{84}\]
We can now turn the annealed planted estimate into the required conditional observation bound. Let \(p_n(J)\) denote the conditional observation-path failure probability for (60). The planted estimate is \[\mathbb E\bigl[L(J)p_n(J)
\mathbf 1_{\{\lVert J\rVert\le2\sqrt j+\theta\}}\bigr]
\le Ce^{-a_4n}.\] Choose \(\delta<a_4/2\). On \(L(J)\ge e^{-\delta n}\), this gives an original expectation bound \(Ce^{-(a_4-\delta)n}\) for the restricted \(p_n(J)\). Markov’s inequality makes \(p_n(J)\le e^{-an}\) outside an event of probability tending to zero, for a fixed sufficiently small \(a>0\). Equation (84) and the usual GOE norm bound, including diagonal removal, show that the restrictions on \(L(J)\) and \(\lVert J\rVert\) hold with probability tending to one. Intersect these events to obtain \(\mathcal J_n\) and finish the proof.
We finish by checking measurability of the witness events used in the argument. At fixed \(y\), the matrix-witness sets are projections over the compact diagonal range. For existence of \(y,A,t\) along a continuous path, additionally restrict \(\lVert y\rVert\), \(\lVert J\rVert\), and \(\sup_{t\le T}\lVert Y(t)\rVert\) by integers. The witness inequalities are closed and the witness parameter space is compact, so these restricted projections are closed. Taking a countable union removes the bounds. This also justifies the conditional probabilities and the restricted integrals above. ◻
A unique root and quantitative inversion
The preceding proposition supplies stability along random observations at fixed subcritical \(j\). The following consequence is deterministic: its hypothesis requires positivity throughout the region of small residual. From that whole-region assumption we obtain a unique zero and a bound on the distance to it. This separates the probabilistic construction of stable fields from the inversion used later.
Lemma 15 (Root and inverse bound at a stable field). Fix \(j,K,c,r_0>0\). Let \(J\) be symmetric with \(\lVert J\rVert\le K\), and let \(h_0\in\mathbb R^n\). Suppose \[
\mathsf H_J(y,V_y)\succeq cI
\quad\text{whenever}\quad
\lVert F_{J,h_0}(y)\rVert\le r_0\sqrt n.
\tag{85}\] Then \(F_{J,h_0}\) has exactly one zero \(r\in\mathbb R^n\), and \[
\lVert y-r\rVert\le C\lVert F_{J,h_0}(y)\rVert
\quad\text{if}\quad\lVert F_{J,h_0}(y)\rVert<r_0\sqrt n,
\qquad C=1+(K+3j)/c.
\tag{86}\] In particular this conclusion applies to every field satisfying (60), even if that implication is required only through residual threshold \(\rho_0\sqrt n\).
Proof. We first bound the derivative inverse on the region in the hypothesis. We then prove existence and uniqueness of a zero, and integrate the inverse bound along a path to that zero.
Put \(L_1=BI-J-2jmm^{\mathsf T}/n\). Since \(F'(y)=I+L_1V_y\), the elementary inverse identity gives \[F'(y)^{-1}
=I-L_1V_y^{1/2}
(I+V_y^{1/2}L_1V_y^{1/2})^{-1}V_y^{1/2}.\] Under (85), this inverse exists and has norm at most \(1+(K+3j)/c\), since \(0<V_y\le I\) and \(\lVert L_1\rVert\le K+3j\).
For existence and uniqueness, pass to the coordinates \(m\in(-1,1)^n\) and define the potential \[\Phi(m)=\sum_i\int_0^{m_i}\mathop{\mathrm{atanh}}z\,\mathrm dz-h_0^{\mathsf T}m
-\frac12m^{\mathsf T}Jm+\frac j2\lVert m\rVert^2
-\frac j{4n}\lVert m\rVert^4.\] Its critical-point equation is \(F_{J,h_0}(\mathop{\mathrm{atanh}}m)=0\). At any critical point, \[V_y^{1/2}\Phi''(m)V_y^{1/2}
=\mathsf H_J(y,V_y)\succeq cI,\qquad y=\mathop{\mathrm{atanh}}m,\] so the critical point has positive definite Hessian and is isolated. Choose \(a\in(0,1)\) close enough to one that \[\mathop{\mathrm{atanh}}a>1+\max_i\left(|(h_0)_i|+\sum_j|J_{ij}|+j\right).\] Every critical point lies in the interior of \([-a,a]^n\), and on each face of this box the negative gradient \(-\nabla\Phi\) points strictly inward. The minimum of \(\Phi\) on the box is therefore attained at an interior critical point. This proves existence.
To prove uniqueness, consider the negative gradient flow from any point of the box. The inward-pointing condition keeps the trajectory in the box for all positive times. Its potential decreases and is bounded below, so \(\int_0^\infty\lVert \nabla\Phi(m(t))\rVert^2\,\mathrm dt<\infty\). The squared gradient is uniformly continuous in time, because the gradient and Hessian are bounded on the box. It follows that the gradient tends to zero: a nonzero limsup would, by uniform continuity, force disjoint intervals of fixed positive integral. Every limit point of a trajectory is therefore critical. The critical set is finite, since it is a compact set of isolated zeros. The limit set of a trajectory is connected: it is the intersection of the nested compact connected closures of its tails. Each trajectory therefore converges to one critical point.
It remains to show that all trajectories converge to the same point. The positive Hessian makes every critical point a strict local attractor. For clarity, in a sufficiently small ball about it one has \((m-m_*)^{\mathsf T}\nabla\Phi(m)\ge c_*\lVert m-m_*\rVert^2\), which gives exponential contraction of the distance along the flow and invariance of the ball. Continuous dependence of the flow on its starting point, up to the time it enters that ball, shows that each attraction basin is relatively open in \([-a,a]^n\). These disjoint nonempty basins cover the connected box. Hence there is only one basin and one critical point, proving uniqueness of \(r\).
We now derive the distance estimate. Suppose \(\lVert F(y)\rVert<r_0\sqrt n\). Starting at \(y(0)=y\), solve on the open region \(\lVert F\rVert<r_0\sqrt n\) the ordinary differential equation \[\dot y(s)=-F'(y(s))^{-1}F(y).\] The chain rule gives \(F(y(s))=(1-s)F(y)\) while the solution exists. Its velocity has norm at most \(C\lVert F(y)\rVert\). If its maximal existence interval ended at or before time one, this velocity bound would give a finite limiting point, whose residual is still strictly less than \(r_0\sqrt n\) and whose derivative is invertible. Local existence would extend the solution, a contradiction. The solution consequently reaches time one, at which its value is the unique zero \(r\). Integrating the velocity bound along this path gives (86). ◻
Weak residual identities and discrete differentiation
We derive the residual identities needed for the two posterior estimates. Moments of a fixed finite sequence of primary iterates will locate two consecutive small increments. Ordinary test-vector recipes then control means under square reweighting; one additional implicit recipe supplies the linear-response test for the directional covariance estimate. The proof proceeds from discrete derivative bounds to a weak residual rule, and ends with the terminal identity for that implicit test.
This section is deterministic. We fix \(J\) satisfying the word estimates of Lemma 2, fix an external field \(h_0\in\mathbb R^n\), and take all expectations with respect to \(\mu_{h_0}\). We fix the depth, the finite construction, and its coefficient bounds independently of \(n\). These choices determine a finite required word length, which can increase when the choices change. Constants in derivative estimates can depend on these choices and on the corresponding word bounds; they are uniform in \(n\), \(h_0\), the spin configuration, and the permitted seeds. The size and inverse bounds for the initial implicit construction below have the stronger, depth-independent uniformity explicitly stated there.
We use the discrete derivatives, product rule (2), and norm \(\lVert G\rVert_*\) introduced above, with \(dg=(d_i g)_{i=1}^n\) for scalar \(g\). All vector functions such as \(\tanh h^1\) are understood coordinatewise.
Primary fields
The following recursion has the Onsager-corrected form of iterative TAP constructions, as in Bolthausen [8]. Here it starts from a Gibbs spin sample and is used only through a fixed finite depth. Define \[
\begin{gathered}
m^0=x,\quad b_0=0,\qquad
m^l=\tanh h^l,\quad b_l=\langle 1-(m^l)^2\rangle,\quad R_l=m^{l-1}-m^l,\\
h^{l+1}=h_0+Jm^l-jb_lm^{l-1}\quad(l\geq1).
\end{gathered}
\tag{87}\] The finite-depth objective is visible in the scalar pairings \(S_l=R_l^{\mathsf T}m^l/\sqrt n\). Expanding the squared increment gives \[
\frac{\lVert R_l\rVert^2}{n}=b_l-b_{l-1}-\frac{2S_l}{\sqrt n}.
\tag{88}\] Since \(0\le b_l\le1\), control of \(S_l/\sqrt n\) through a sufficiently large fixed depth forces two consecutive residuals to be small. We will prove the moment bounds \[\mathbb ES_l^2\le C_l,\qquad
\mathbb E_\nu\left(\frac{S_l}{\sqrt n}\right)^2
\le\frac{C_l}{\sqrt n}\left(1+\frac{\mathcal D_{h_0}(f)}{N}\right),
\quad N=\mathbb Ef^2>0,\quad \nu=\frac{f^2\mu_{h_0}}{N}.\] These bounds will supply the ordinary and square-reweighted selection estimates in Section 7; convergence of the primary recursion is not required.
The first residual explains what is needed from a test vector \(U\). Conditional covariance in one spin gives the exact identity \[\mathbb EG R_1^{\mathsf T}U=\mathbb E\sum_i v_i^1d_i(GU_i).\] The product rule bounds its right side by \(C\lVert G\rVert_*\) whenever \(\lVert U\rVert\) and \(\sum_i|d_iU_i|\) are bounded pointwise. Thus the essential property of our tests is a bound on the sum of their diagonal derivatives, in addition to a bound on their Euclidean norm. The Onsager corrections will preserve this property through the finite constructions below. We begin with the primary fields, whose derivatives enter those tests. We use \(\lVert Q\rVert_{\ell^4}=(\sum_{i,j}|Q_{ij}|^4)^{1/4}\) for the unnormalized entrywise norm.
Lemma 16 (Primary differentiation). Let \[
H_1=J,\qquad M_0=I,\qquad
M_l=D_{1-(m^l)^2}H_l,\qquad
H_{l+1}=JM_l-jb_lM_{l-1}.
\tag{89}\] For each fixed \(l\geq1\), \[\begin{align*}
&\lVert \nabla h^l-H_l\rVert_{\mathrm F}
+\lVert \nabla m^l-M_l\rVert_{\mathrm F}\leq C,\tag{90}\\
&\lVert \nabla h^l\rVert+\lVert \nabla m^l\rVert
+\lVert \nabla h^l\rVert_{\ell^4}
+\lVert \mathop{\mathrm{diag}}\nabla h^l\rVert_2\leq C,
\qquad \lVert db_l\rVert\leq Cn^{-1/2}.
\tag{91}\end{align*}\] More generally, let \(g_i\) be a smooth function of a fixed list of coordinates \(h_i^l\) and averages \(b_l\), with uniformly bounded derivatives through the orders being used. Deterministic site labels are permitted, provided the same bounds hold uniformly in \(i\). Its derivative matrix has bounded operator norm, bounded diagonal \(\ell^2\) norm, and bounded entrywise \(\ell^4\) norm.
Proof. We first bound the formal matrices in (89), then compare them with the actual discrete derivatives. Expanding \(H_{l+1}\) produces the chain \[J D_{1-(m^l)^2}J\cdots D_{1-(m^1)^2}J\] and the terms obtained by contracting disjoint pairs of originally adjacent \(J\)’s. Contracting \(J D_{1-(m^s)^2}J\) contributes \(-j b_s I\). This description follows inductively from the two terms defining \(H_{l+1}\): either keep its first link, or contract that link with the next one. Multiplicities are retained when different contractions happen to give the same matrix.
The contraction terms cancel the diagonal predictions. To see this, fix a complete noncrossing pairing of the original chain. Reinserting each contracted adjacent pair identifies its contribution with the prediction of this full pairing, multiplied by the signs of the contractions. The originally adjacent pairs belonging to a pairing are mutually disjoint, and every nonempty complete noncrossing pairing contains at least one such pair. Summing over their subsets gives \((1-1)^a=0\), where \(a\geq1\) is their number. An odd chain has no complete pairing. Lemma 2 therefore bounds the diagonal \(\ell^2\) norm of \(H_l\), and bounds its operator and off-diagonal \(\ell^4\) norms term by term. Its full \(\ell^4\) norm is bounded as well, since \(\lVert \mathop{\mathrm{diag}}H_l\rVert_4\leq\lVert \mathop{\mathrm{diag}}H_l\rVert_2\).
We now compare these formal bounds with the actual derivatives. For a scalar function \(\psi\) with bounded second derivative, Taylor’s theorem gives \[\lvert d_j\psi(h_i^l)-\psi'(h_i^l)d_jh_i^l\rvert
\leq C|d_jh_i^l|^2.\] The Frobenius norm of the remainder is at most \(C\lVert \nabla h^l\rVert_{\ell^4}^2\). This proves the magnetization part of (90) once the field part is known. The same argument for \(\psi(z)=\mathop{\mathrm{sech}}^2z\) gives an operator bound for its site derivative matrix \(Q_l\). Averaging its rows yields \[db_l=n^{-1}Q_l^{\mathsf T}\mathbf 1,\qquad
\lVert db_l\rVert\leq n^{-1/2}\lVert Q_l\rVert.\] In differentiating \(b_lm^{l-1}\), the term from \(db_l\) is a rank-one matrix of Frobenius norm at most \(\lVert m^{l-1}\rVert\lVert db_l\rVert\leq C\). The discrete product error is \(\nabla m^{l-1}\) times a diagonal whose diagonal has norm at most \(2\lVert db_l\rVert\). Its Frobenius norm is bounded, since \(\lVert \nabla m^{l-1}\rVert\leq C\). Starting at \(\nabla h^1=J\), these estimates close the induction. The resulting bounded Frobenius perturbation preserves the operator, diagonal \(\ell^2\), and entrywise \(\ell^4\) bounds of the formal matrices.
For the final assertion, apply the same comparison to a smooth function of several arguments. Expand along the segment between the arguments at \(x\) and \(x^{(j)}\). The preceding bounds control the primary linear terms. A linear term involving \(db_s\) is a rank-one matrix with a bounded-entry left vector, hence has bounded Frobenius norm. The squared-average remainder is controlled by \[n\sum_j |d_jb_s|^4\leq n\lVert db_s\rVert^4\leq C.\] Mixed remainders are bounded by sums of pure squares, and primary quadratic remainders are controlled by the entrywise \(\ell^4\) bounds. This proves the assertion. ◻
We have in particular bounded the Euclidean change of every primary vector under a single spin flip, at each fixed depth. This is the estimate that will keep the local constructions inside their buffers.
Finite recipes for small vectors
The residuals will be tested against vectors built by finitely many operations. We call these vectors “small” when their Euclidean norm is bounded; the word does not refer to coordinatewise magnitude. The following definition specifies the operations and the meaning of their site partials.
Definition 17 (Ordinary admissible recipe). Choose finitely many deterministic seeds \(s\in\mathbb R^n\) of bounded Euclidean norm. Seeds may depend on \(J\) and \(h_0\), but are fixed as functions of \(x\). Choose a finite tuple of scalar functions \(\theta\) satisfying \[\sup_x\bigl(|\theta_\alpha(x)|+\lVert d\theta_\alpha(x)\rVert\bigr)\leq C\] for each component. Such scalar functions will be called parameters; they may include the averages \(b_l\).
Given auxiliary fields \(y^1,\ldots,y^{a-1}\) already constructed, a source has the form \[
w_i^a=\sum_{s}\phi_{a,s,i}\bigl((h_i^l)_{l\in I_a},\theta\bigr)s_i
+\sum_{b<a}\phi_{a,b,i}\bigl((h_i^l)_{l\in I_a},\theta\bigr)y_i^b.
\tag{92}\] All sums and index sets are finite. Coefficients are smooth and bounded, with bounded derivatives of every fixed order needed in the construction, uniformly on their argument ranges and the connecting line segments. Define the new auxiliary field by \[
y^a=Jw^a
-j\sum_l\langle \partial_lw^a\rangle\,m^{l-1}
-j\sum_{b<a}\langle \partial_{y^b}w^a\rangle\,w^b.
\tag{93}\] The partial \(\partial_l\) differentiates the explicitly displayed local argument \(h_i^l\), holding every parameter and every auxiliary argument fixed. In particular it holds \(b_l\) fixed even when that average is among the parameters. The partial \(\partial_{y^b}\) differentiates the indicated linear auxiliary argument, with its coefficient held fixed. A partial in an absent argument is zero. Thus \(\partial_{y^b}w^a\) is simply the coefficient vector \(\phi_{a,b}\).
A terminal vector is another source of the form (92), using any already constructed auxiliary fields. A source \(w^b\) used again as a terminal vector or in a new source is replaced by its expression (92); auxiliary fields themselves are not expanded. The recipe is admissible from level \(p\) if every explicit primary argument appearing in its source formulas or terminal formula has index at least \(p\). This condition does not restrict the parameters, or the vectors \(m^{l-1}\) in the correction terms of (93).
The definition permits two operations that we will use repeatedly. Multiplication by a bounded smooth coefficient of the same allowed arguments gives another source. At any finite stage we may also add parameters satisfying the stated size and difference bounds. We will use the deterministic seed \(\mathbf 1/\sqrt n\) freely; thus \(m^l/\sqrt n\) is a source with the single primary argument \(h^l\). Linearity in the auxiliary arguments is preserved in every source. It will be used in both the derivative estimates and the trace cancellation.
For a terminal vector \(U\) admissible from level \(p\), the estimate we seek is \[\left|\mathbb EG R_p^{\mathsf T}U\right|\le C\lVert G\rVert_*.\] Its square-density counterpart bounds \(|\mathbb Ef^2H_0R_p^{\mathsf T}U|\) by \(C(\mathbb Ef^2+\mathcal D_{h_0}(f))\) when \(H_0\) has bounded size and difference norm. With \(U=m^p/\sqrt n\), the choices \(G=S_p\) and \(H_0=S_p/\sqrt n\) will give the moment estimates above. Lemma 20 proves these rules after we establish the required diagonal derivative bounds.
A scalar identity for comparing fields
To pass from \(R_{p+1}=\tanh h^p-\tanh h^{p+1}\) to earlier residuals, we use a scalar identity valid for arbitrary real fields. For \(z,r_0,a\in\mathbb R\), define \[
\Phi(z;r_0,a)=\int_0^1 e^{a(1-u^2)/2}
\frac{\cosh(u(z-r_0))}{\cosh z\cosh r_0}\,\mathrm du.
\tag{94}\] Differentiating \(e^{a(1-u^2)/2}\sinh(u(z-r_0))\) in \(u\) gives \[
\bigl[(z-r_0-a\tanh z)-a\partial_z\bigr]\Phi(z;r_0,a)
=\tanh z-\tanh r_0.
\tag{95}\] The boundary at \(u=0\) is zero; the boundary at \(u=1\), after division by \(\cosh z\cosh r_0\), is the right side. The identity holds for either sign of \(a\). At \(a=0\), \(\Phi\) is the secant slope of \(\tanh\), with continuous value \(\mathop{\mathrm{sech}}^2r_0\) at \(z=r_0\).
These formulas also give the uniform coefficient bounds needed for the recipes. Positivity of the integrand and the secant-slope formula imply \(0<\Phi\le e^{|a|/2}\). For \(a\) in a fixed bounded interval, every fixed-order derivative in \(z,r_0,a\) is bounded by a constant times \(\Phi\): derivatives of the hyperbolic factors introduce bounded ratios, and those of the exponential introduce bounded powers of \((1-u^2)/2\). The quotient rule then bounds every fixed-order derivative of each ratio \((\partial^\alpha\Phi)/\Phi\). These assertions are uniform in the two real field arguments and require no stability assumption.
In the ordinary residual induction we will substitute \((z,r_0,a)=(h_i^p,h_i^{p+1},j(b_p-b_{p-1}))\). The same identity also provides the two tests comparing an approximate field with the stable root, which we construct next.
An initial implicit auxiliary field
The directional estimate needs one auxiliary field defined by a linear implicit equation. We construct it near the stable root, establish its inverse margin, and then control its discrete derivatives on a smaller region.
Assume the stable-field implication (60) holds at \(h_0\) through the fixed residual threshold required in Lemma 15. Let \(r\) be its unique root, and put \(t_*=\tanh r\) and \(q_*=\langle t_*^2\rangle\). Fix \(k\geq2\) and consider \[
\Omega_\rho=
\{x:\max(\lVert R_k(x)\rVert,\lVert R_{k-1}(x)\rVert)\leq\rho\sqrt n\}.
\tag{96}\] All choices of \(\rho\) below depend only on the stable-field constants and the operator bound on \(J\), until derivative estimates are considered. For sufficiently small \(\rho\), the primary recursion and Lemma 15 give on \(\Omega_{2\rho}\)\[
\lVert h^k-r\rVert+\lVert m^{k-1}-t_*\rVert\leq C\rho\sqrt n,
\qquad |1-b_{k-1}-q_*|\leq C\rho.
\tag{97}\] The comparison is independent of \(k\): the primary recursion gives \[F(h^k)=JR_k-jb_{k-1}(R_{k-1}+R_k)
+j(b_k-b_{k-1})m^k,
\qquad
|b_k-b_{k-1}|\leq2\lVert R_k\rVert/\sqrt n,\] so the residual is at most \(C\rho\sqrt n\). Applying the root estimate therefore gives the asserted comparison with the same depth independence.
Use the scalar kernel with \[
a=j(1-b_{k-1}-q_*),\qquad A_i=\Phi(h_i^k;r_i,a).
\tag{98}\] Write \(A=D_{(A_i)}\) and \(\chi=\langle A_i\rangle\). The initial source \(w_0\) is one of \[w_{0,i}=A_i e_i\quad(\lVert e\rVert\leq1),\qquad
w_{0,i}=n^{-1/2}\partial_{r_0}\Phi(h_i^k;r_i,a).\] The partial in the second formula holds \(a\) fixed. The two choices have complementary roles. Holding \(r_i,a\) fixed in the site partial \(\partial_k\), the scalar identity and its derivative in \(r_0\) give \[\begin{align*}
\sum_i[(h_i^k-r_i-am_i^k)-a\partial_k](A_i e_i)
&=(m^k-t_*)^{\mathsf T}e,\\
\sum_i[(h_i^k-r_i-am_i^k)-a\partial_k]
\frac{\partial_{r_0}\Phi(h_i^k;r_i,a)}{\sqrt n}
&=\sqrt n\{\chi-(1-q_*)\}.
\tag{99}\end{align*}\] Indeed differentiating (95) in \(r_0\) gives \([(z-r_0-a\tanh z)-a\partial_z]\partial_{r_0}\Phi
=\Phi-\mathop{\mathrm{sech}}^2r_0\). Thus the first source tests a direction of the magnetization error, while the second tests the discrepancy of the averaged coefficient \(\chi\). Both quantities enter the posterior comparison. To put these scalar identities into a form controlled by the weak residual rule, introduce the coupled source and field \[
\begin{gathered}
w=w_0+A(y^{\rm imp}-jc_1t_*),\qquad
y^{\rm imp}=Jw-j\chi w-jc_1m^{k-1},\\
c_1=\langle \partial_k w\rangle.
\end{gathered}
\tag{100}\] In this partial, \(y^{\rm imp}\), \(c_1\), and \(a\) are independent arguments held fixed. Consequently \(\partial_{y^{\rm imp}}w=A\), and the field equation in (100) is exactly (93) with one allowed self-dependence: \(\chi\) is its auxiliary-partial average, and \(c_1\) its primary-partial average. The correction \(jc_1m^{k-1}\) will supply the scalar cancellation in Lemma 21. Its only primary local argument is \(h^k\). Once constructed, \(y^{\rm imp}\) is available as the first auxiliary field in any ordinary finite recipe. The scalar \(\sqrt n c_1\) is then an additional parameter, and \(t_*/\sqrt n\) is an additional seed.
Lemma 18 (Implicit construction). For sufficiently small \(\rho>0\), (100) has a unique solution on \(\Omega_{2\rho}\). There are constants independent of the primary depth \(k\) such that, throughout this region, \[
\lVert w\rVert+\lVert y^{\rm imp}\rVert+\sqrt n|c_1|\leq C,
\qquad
I+j\chi A-A^{1/2}JA^{1/2}\succeq cI.
\tag{101}\] On \(\Omega_{3\rho/2}\), for all sufficiently large \(n\) depending on the fixed depth, the solution also satisfies \[
\lVert \nabla w\rVert_{\mathrm F}+\lVert \nabla y^{\rm imp}\rVert_{\mathrm F}
+\sqrt n\lVert dc_1\rVert\leq C_k.
\tag{102}\]
Proof. We first control the coefficients and solve the implicit linear system. Only then do we differentiate it, using the larger region to cover spin flips from the smaller one. The uniform kernel and quotient bounds proved above will control the coefficients even when some \(A_i\) are small.
Put \(\ell_i=(\partial_z\Phi/\Phi)(h_i^k;r_i,a)\). At \(z=r_0\) this ratio equals \(-\tanh r_0\), for every \(a\). Comparison with \(z=r_0,a=0\), followed by (97), gives \[
\lVert A-D_{1-(m^k)^2}\rVert_{\mathrm F}\leq C\rho\sqrt n,
\qquad \lVert \ell+t_*\rVert\leq C\rho\sqrt n,
\qquad |\chi-b_k|\leq C\rho.
\tag{103}\] In particular, the comparisons put \(A\) among the diagonals allowed by the stable-field implication, once \(\rho\) is decreased.
Let \(u_0=w_0/A\) coordinatewise and \(p_0=A\partial_k u_0\). The quotient bounds just proved give \(\lVert u_0\rVert+\lVert p_0\rVert\leq C\); their local derivatives have coordinatewise bounds proportional to \(|e_i|\) in the first case and to \(n^{-1/2}\) in the second. Direct differentiation with the specified partial convention gives \[\partial_k w=D_\ell w+p_0,\qquad
c_1=\langle \ell w+p_0\rangle.\] We can now eliminate \(y^{\rm imp}\) and \(c_1\) and solve for the source alone. The resulting linear equation is \[
\left[I-A(J-j\chi I)
+\frac jn A(m^{k-1}+t_*)\ell^{\mathsf T}\right]w
=A\left[u_0-j(m^{k-1}+t_*)\langle p_0\rangle\right].
\tag{104}\] Write the matrix on the left as \(I+AL_2\). Its symmetrically normalized version \(I+A^{1/2}L_2A^{1/2}\) differs in operator norm by \(O(\rho)\) from \[I+jb_k A-A^{1/2}JA^{1/2}
-\frac{2j}{n}A^{1/2}m^k(m^k)^{\mathsf T}A^{1/2}.\] For the rank term, use (97) and (103) in the elementary inequality \(\lVert uv^{\mathsf T}-u'v'^{\mathsf T}\rVert
\leq\lVert u-u'\rVert\lVert v\rVert+\lVert u'\rVert\lVert v-v'\rVert\). The displayed symmetric matrix has the fixed positive margin supplied by (60). Decrease \(\rho\) until the perturbation is less than half that margin. The normalized matrix then has bounded inverse, even though it need not itself be symmetric. The identity \[(I+AL_2)^{-1}
=I-A^{1/2}(I+A^{1/2}L_2A^{1/2})^{-1}A^{1/2}L_2\] gives the same bound for the unnormalized inverse, since \(\lVert L_2\rVert\) is bounded. The right side of (104) is bounded in norm, because \(|\langle p_0\rangle|\leq\lVert p_0\rVert/\sqrt n\). This establishes the size bounds. To obtain the last assertion of (101), remove the negative rank term from the stable-field matrix and change \(b_k\) to \(\chi\). The constants used to solve the system and obtain these bounds depend only on the fixed stability and operator margins.
It remains to bound the discrete derivatives. Compare the equations at \(x\) and \(x^{(i)}\). On the stated inner region both configurations belong to \(\Omega_{2\rho}\) for large \(n\), by Lemma 16. In every coefficient-difference term, retain the solution vector at the common point \(x\). In particular, \[\begin{align*}
d_i c_1
&=\langle (d_i\ell)w(x)+d_i p_0\rangle
+n^{-1}\ell(x^{(i)})^{\mathsf T}d_iw,\\
d_i\{c_1(m^{k-1}+t_*)\}
&=c_1(x)d_i m^{k-1}
+(m^{k-1}(x^{(i)})+t_*)d_ic_1.
\end{align*}\] Use the same product arrangement for \(A\) and \(\chi\) in \[w=A\{u_0+(J-j\chi I)w-jc_1(m^{k-1}+t_*)\}.\] For clarity, put \(v=m^{k-1}+t_*\) and \(\mathcal B=u_0+(J-j\chi I)w-jc_1v\), and use primes for evaluation at \(x^{(i)}\) in the next display. With \(\zeta_i=n^{-1}\{(d_i\ell)^{\mathsf T}w+\mathbf 1^{\mathsf T}d_ip_0\}\), the exact equation is \[\begin{align*}
&\left[I-A'(J-j\chi'I)+\frac jn A'v'(\ell')^{\mathsf T}\right]d_iw
\\
&\qquad=(d_iA)\mathcal B+A'd_iu_0
-jA'w\,d_i\chi-jc_1A'd_im^{k-1}-jA'v'\zeta_i.
\tag{105}\end{align*}\] Thus the left matrix is precisely the matrix in (104) evaluated at \(x^{(i)}\). The bracket \(\mathcal B\) on the right is consequently one common-point vector for all columns. This allows the following estimates to control the Frobenius norm of the entire right side.
We bound its terms in turn. A site coefficient difference times a common vector \(s\) of bounded norm has bounded Frobenius norm, since \[\sum_{i,j}s_i^2|d_j\psi_i|^2
\leq \lVert s\rVert^2\max_i\sum_j|d_j\psi_i|^2\leq C.\] This applies to \(A,\ell,u_0,p_0\), using their coefficient bounds and Lemma 16; the seed factors are retained for \(u_0,p_0\). The bracket multiplying \(d_iA\) is bounded in norm by the already established size estimates, so no bound on \(1/A_i\) is needed. Also \(\lVert d\chi\rVert\leq C/\sqrt n\), because \(A_i\) is a primary site function with an average parameter of that derivative scale. If \(\zeta_i=\langle (d_i\ell)w+d_i p_0\rangle\), then \[\lVert \zeta\rVert\leq n^{-1/2}
\bigl(\lVert D_w\nabla\ell\rVert_{\mathrm F}
+\lVert \nabla p_0\rVert_{\mathrm F}\bigr)
\leq C/\sqrt n.\] Its multipliers \(m^{k-1}(x^{(i)})+t_*\) have norm at most \(2\sqrt n\), which still gives a bounded Frobenius norm after summing the squared column norms. Finally \(|c_1|\lVert \nabla m^{k-1}\rVert_{\mathrm F}\leq C\), by its operator bound and \(|c_1|\leq C/\sqrt n\). Applying the uniformly bounded flipped inverses column by column proves \(\lVert \nabla w\rVert_{\mathrm F}\leq C_k\). Returning to the two displayed difference identities gives \(\lVert dc_1\rVert\leq C_k/\sqrt n\) and the required bound for \(\nabla y^{\rm imp}\). The operator in (104), the regions \(\Omega_\rho\), and the required buffer inclusions do not depend on the directional seed \(e\). All bounds for its right side and its derivatives use only \(\lVert e\rVert\leq1\); indeed \(u_0=e\) and \(p_0=0\) for that source. The inverse margins, the threshold on \(n\), and all these estimates are therefore uniform over the full Euclidean unit ball of seeds. ◻
Local recipes and their buffers.
A local recipe starts with this implicit field and adds only ordinary auxiliaries. Such recipes are evaluated on the inner region \(\Omega_\rho\), with all estimates also available on its fixed finite flip-neighborhoods. To define them as functions on the whole cube, keep the actual solution through \(\Omega_{2\rho}\) and assign finite values to \(y^{\rm imp},c_1\) outside that region, defining sources by their stated expressions. The implicit equation is required only inside the buffer. Every test in a local statement is supported on \(\Omega_\rho\). For a fixed recipe, once \(n\) is sufficiently large, every difference used in the proof stays strictly inside the buffer. The arbitrary extension therefore does not enter the estimates.
The buffer sizes are chosen before the construction depth; the threshold on dimension is chosen afterward. To verify this order, choose a finite recipe and let \(q\) be the largest number of successive spin flips used by any difference in its proof. The induction below adds only finitely many operations at each smaller residual index, so \(q\) is finite. The primary bounds give, for every involved level \(l\), \[\|R_l(x^{(i)})-R_l(x)\|
\le 2\|\nabla R_l(x)e_i\|\le C_l.\] With \(C=\max_l C_l\), a path of at most \(q\) flips increases either residual defining \(\Omega_\rho\) by at most \(qC\). Thus \(n>(2qC/\rho)^2\) puts every required neighborhood of \(\Omega_{3\rho/2}\) strictly inside \(\Omega_{2\rho}\). Here \(q,C\), and the threshold on \(n\) may depend on the already fixed recipe and depth; \(\rho\) still depends only on the earlier stability and size margins.
Trace control for small derivatives
The primary derivative bounds do not yet give the diagonal trace control needed for integration by parts. For small vectors, the source structure and its correction terms supply this additional cancellation. We prove it together with the size and derivative bounds used to construct later auxiliary fields.
Lemma 19 (Small-vector differentiation). Every source, auxiliary field, and terminal vector \(U\) in a fixed ordinary admissible recipe satisfies \[
\lVert U\rVert+\lVert \nabla U\rVert_{\mathrm F}
+\lVert \mathop{\mathrm{diag}}\nabla U\rVert_1\leq C.
\tag{106}\] The same holds locally for recipes with the single initial implicit field above. In these recipes the initial partial \(\partial_k w\) is also a small source. For every primary coefficient and auxiliary coefficient in (93), respectively, \[
\begin{split}
\alpha_l=\langle \partial_lw^a\rangle:&\quad
|\alpha_l|+\lVert d\alpha_l\rVert\leq C/\sqrt n,\\
\gamma_b=\langle \partial_{y^b}w^a\rangle:&\quad
|\gamma_b|+\lVert d\gamma_b\rVert\leq C.
\end{split}
\tag{107}\] All local statements include the required fixed flip buffers. Constants can depend on the entire finite recipe and its primary depth.
Proof. The proof has two parts. We bound the errors introduced by discrete differentiation, then cancel the trace predictions of the retained formal terms. Write \(\lVert E\rVert_{\rm nuc}\) for the sum of the singular values of \(E\). The error estimates repeatedly use \[
|\mathop{\mathrm{Tr}}(PE)|\leq\lVert P\rVert\lVert E\rVert_{\rm nuc},\qquad
\lVert BC\rVert_{\rm nuc}\leq\lVert B\rVert_{\mathrm F}\lVert C\rVert_{\mathrm F},\qquad
\lVert E\rVert_{\rm nuc}\leq\sum_j\lVert E_{\cdot j}\rVert.
\tag{108}\] Every bounded-length diagnostic word \(P\) has bounded operator norm and bounded off-diagonal entrywise \(\ell^4\) norm. For the errors arising before we solve the differentiated implicit equations, \(P\) may contain one inverse \(K\) from (17). Its required margin is supplied by the last assertion of (101).
Source errors.
Consider one source term \(\phi_i s_i\). Retain the derivative \[
D_\phi\nabla s+\sum_lD_{s\partial_l\phi}\nabla h^l,
\tag{109}\] and put parameter derivatives into the error rather than into the formal expression. Set \[Q_{ij}=\sum_l|d_jh_i^l|,\qquad
\Theta_j=\sum_\alpha|d_j\theta_\alpha|.\] The finite sums have the bounds \[\max_i\sum_jQ_{ij}^2+\lVert Q\rVert_{\ell^4}^2
+\lVert \mathop{\mathrm{diag}}Q\rVert_2^2\leq C,
\qquad \lVert \Theta\rVert\leq C.\] Taylor expansion of \(d_j\phi_i\) has linear parameter terms and a remainder bounded entrywise by \[
C\{Q_{ij}^2+Q_{ij}\Theta_j+\Theta_j^2\}.
\tag{110}\] After multiplication by \(s_i\), the linear parameter terms have bounded nuclear norm: each is a rank-one matrix with a bounded-norm left vector and a bounded-norm right vector. For the \(Q^2\) part, the matrix before multiplication by \(D_s\) has bounded Frobenius norm. Its product with \(D_s\) has bounded nuclear norm by (108) and \(\lVert D_s\rVert_{\mathrm F}=\lVert s\rVert\leq C\). For the mixed part the sum of column norms is at most \[C\sum_j\Theta_j\Bigl(\sum_i s_i^2Q_{ij}^2\Bigr)^{1/2}
\leq C\lVert \Theta\rVert
\Bigl(\sum_i s_i^2\sum_jQ_{ij}^2\Bigr)^{1/2}\leq C.\] For the last part it is at most \(C\lVert s\rVert\sum_j\Theta_j^2\leq C\). These separate estimates also apply to the combined remainder in (110): entry by entry, assign each summand in the bound its proportional share of the actual remainder.
The product rule also has a remainder with entries bounded by \[C|d_js_i|\{Q_{ij}+\Theta_j\}.\] Assume \(\lVert \nabla s\rVert_{\mathrm F}\leq C\), which is zero for a seed, holds for the initial implicit field by Lemma 18, and is inductive for later fields. The \(\Theta\) part has bounded sum of column norms by Cauchy–Schwarz. For the \(Q\) part, test against \(P\) and split the trace into diagonal and off-diagonal terms. Hölder’s inequality gives \[\begin{align*}
\sum_{i\ne j}|P_{ji}|\,|d_js_i|Q_{ij}
&\leq \lVert P_{\rm off}\rVert_{\ell^4}
\lVert \nabla s\rVert_{\mathrm F}\lVert Q\rVert_{\ell^4}\leq C,\\
\sum_i|P_{ii}|\,|d_is_i|Q_{ii}
&\leq\lVert P\rVert\lVert \mathop{\mathrm{diag}}\nabla s\rVert_2\lVert \mathop{\mathrm{diag}}Q\rVert_2\leq C.
\end{align*}\] The first line uses exponents \(4,2,4\). In particular, the argument does not assume the diagonal \(\ell^1\) bound for \(\nabla s\) that we are still proving. Bounded entries of \(Q\) also give a bounded Frobenius norm for this error. We have therefore controlled each source error both in Frobenius norm and in trace against the indicated words.
Average corrections and size bounds.
Next we control the averaged coefficients in the correction terms. Differentiating a coefficient function in a primary argument gives another source with the same inputs and bounds. Apply the size and Frobenius parts of the preceding estimates to that source, then average: \[|\langle u\rangle|\leq\lVert u\rVert/\sqrt n,\qquad
\lVert d\langle u\rangle\rVert\leq\lVert \nabla u\rVert_{\mathrm F}/\sqrt n,\] which proves the first line of (107). The implicit coefficient \(c_1\) already has these bounds from its direct construction. For an auxiliary partial the scale is different: it is a bounded coefficient vector, rather than a small vector. Averaging its difference estimate \(|d_j\phi_i|\leq C(Q_{ij}+\Theta_j)\) and using the row bounds above gives \(|\langle \phi\rangle|+\lVert d\langle \phi\rangle\rVert\leq C\). This proves the second line, including the self-coefficient \(\chi\) when it occurs.
For a primary correction \(\alpha_lm^{l-1}\) retain \(\alpha_l\nabla m^{l-1}\). The dropped rank-one matrix has nuclear norm at most \(\lVert m^{l-1}\rVert\lVert d\alpha_l\rVert\leq C\). The dropped product error is \(\nabla m^{l-1}\) times a diagonal of signed entries \(2d_i\alpha_l\). Its nuclear norm is bounded by \[2\lVert \nabla m^{l-1}\rVert_{\mathrm F}\lVert d\alpha_l\rVert
\leq C\sqrt n\,n^{-1/2}\leq C.\] This also holds for \(m^0=x\). For an auxiliary correction \(\gamma_bw^b\), the analogous rank-one and product errors have bounded nuclear norm using \(\lVert w^b\rVert+\lVert \nabla w^b\rVert_{\mathrm F}\leq C\) and \(|\gamma_b|+\lVert d\gamma_b\rVert\leq C\). When primary derivatives in retained terms are replaced by \(H_l,M_l\), their Frobenius errors from Lemma 16 are multiplied either by a small-vector diagonal or by an \(O(n^{-1/2})\) scalar. In the former case (108) applies; in the latter use \(\lVert E\rVert_{\rm nuc}\leq\sqrt n\lVert E\rVert_{\mathrm F}\). Both give bounded nuclear norm.
We can now close the induction for the size and Frobenius parts of (106). In (93), the terms are a bounded-norm leading term, primary corrections bounded by \(C\sqrt n\,n^{-1/2}\), and bounded-norm auxiliary corrections. The retained source derivative (109) is Frobenius-bounded: use the preceding auxiliary derivative bounds and \(\lVert D_{s\partial_l\phi}\nabla h^l\rVert_{\mathrm F}
\leq\lVert s\partial_l\phi\rVert\lVert \nabla h^l\rVert\). The retained primary correction is bounded in Frobenius norm by \(|\alpha_l|\lVert \nabla m^{l-1}\rVert_{\mathrm F}\leq C\). Thus the size and Frobenius induction closes before we use or prove the diagonal \(\ell^1\) assertion.
Solving the differentiated self-dependence.
The initial implicit source needs one additional step before the trace cancellation. Retain its site partials, freeze the parameters and averages, and replace the primary derivatives by their formal matrices. This gives \[
\begin{split}
\nabla w&=A\nabla y^{\rm imp}+D_{\partial_k w}H_k+E_1,\\
\nabla y^{\rm imp}&=(J-j\chi I)\nabla w-jc_1M_{k-1}+E_2.
\end{split}
\tag{111}\] The errors have bounded Frobenius norm and bounded trace against words containing at most one \(K\). The parameter \(c_1\) causes no exception: write the corresponding source term as \(-j(\sqrt n c_1)A(t_*/\sqrt n)\), so that it is exactly of the source form above. Its difference bound follows from (102). Solving (111) yields \[
\nabla w
=K\{D_{\partial_k w}H_k-jc_1AM_{k-1}+E_1+AE_2\},
\qquad K=(I-AJ+j\chi A)^{-1}.
\tag{112}\] Testing a propagated error against an ordinary word now puts at most one \(K\) into the test word. The preceding error estimates therefore give both bounded trace and bounded Frobenius norm. Substitution gives the same conclusion for \(\nabla y^{\rm imp}\). Subsequent fields are ordinary: propagating these errors multiplies them only by ordinary words. Induction therefore bounds the trace of every final error against an ordinary word. Since the recipe is fixed and finite, the required increases in word length are finite as well.
Cancellation of the formal traces.
All differentiated errors are now controlled. To finish, we must bound the diagonal \(\ell^1\) norm of the retained formal derivative. A single ordinary field already illustrates the cancellation. From \(w_i=\phi(h_i^1)s_i\) we obtain \(y=Jw-j\langle \phi'(h^1)s\rangle x\). If \(P=D_{\phi'(h^1)s}\), its retained derivative is \(JPJ-j\langle \phi'(h^1)s\rangle I\), whose two diagonal predictions cancel. We now show that the correction at each node gives this same cancellation along every dependency path, including paths through the implicit field.
At the point \(x\), freeze the local partial vectors and all their averages. Index each primary or auxiliary field by a node \(a\), and attach its source to the same node: for example, \(h^l\) has source \(m^{l-1}\), and \(y^a\) has source \(w^a\). Let \(p_{ad}\) be the partial of the source at node \(a\) with respect to the field at node \(d\). Introduce a formal variable \(z\), and denote the formal source and field derivatives by \(\mathfrak S_a\) and \(\mathfrak H_a\), respectively. They obey \[
\mathfrak S_a=\sum_dD_{p_{ad}}\mathfrak H_d,
\qquad
\mathfrak H_a=zJ\mathfrak S_a
-z^2j\sum_d\langle p_{ad}\rangle\mathfrak S_d.
\tag{113}\] At the spin leaf, the source derivative for \(h^1\) is \(I\) and its field derivative is \(zJ\). A terminal source uses just the first equation. At \(z=1\) these are precisely the retained formal derivatives. In the implicit case the only cyclic node is \(y^{\rm imp}\) and its source \(w\); closing it gives \[
\mathfrak S_{\rm imp}
=K(z)\{D_{\partial_k w}\mathfrak H_{h^k}
-z^2jc_1A\mathfrak S_{h^k}\},
\quad K(z)=(I-zAJ+z^2j\chi A)^{-1}.
\tag{114}\] Every other dependency is acyclic and linear. Solving the single cyclic node therefore leaves at most one inverse \(K(z)\) in each product of the closed formulas.
The averaged term in (114) is essential even for \(k=2\). Write \(D=D_{1-(m^1)^2}\) and \(P=D_{\partial_2w}\), so that \(b_1=n^{-1}\mathop{\mathrm{Tr}}D\) and \(c_1=n^{-1}\mathop{\mathrm{Tr}}P\). Then \[\mathfrak H_{h^2}=z^2(JDJ-jb_1I),\qquad
\mathfrak S_{h^2}=zDJ,\] and the implicit derivative is \(K(z)\{z^2P(JDJ-jb_1I)-z^3jc_1ADJ\}\). At degree four, the two noncrossing pairings of \(AJAJPJDJ\) have predictions \[j^2\chi b_1 AP,
\qquad j^2c_1\frac{\mathop{\mathrm{Tr}}(AD)}n A.\] The first is canceled by the terms from \(-j\chi A\) in the quadratic coefficient of \(K(z)\) and \(-jb_1I\) in \(\mathfrak H_{h^2}\). The second is canceled by the degree-four term \(-jc_1AJADJ\) from the averaged correction. This calculation exhibits the first cancellation involving both the self-dependent node and its primary input. We now account for every degree by following dependency paths.
We work coefficient by coefficient. An uncontracted dependency path, followed backwards from a field to the spin leaf, has successive links \(JD_{p_{ad}}\) and ends with the \(J\) link of \(h^1\). A terminal source contributes an initial partial diagonal and then such a path. The choices of input nodes index these paths separately, preserving all multiplicities. Expanding (113) is inclusion–exclusion over disjoint originally adjacent link pairs: the first term keeps the first link, while the second contracts it with the next link, with factor \(-j\langle p_{ad}\rangle\), and continues from the input node’s source. That continuation skips exactly the link just contracted. Every contraction has degree two in \(z\), equal to the degree of the replaced links. Consequently a coefficient of degree \(L\) involves only uncontracted paths of length \(L\), of which there are finitely many even with the self-dependence.
To compute the trace prediction, test the formal derivative by a bounded diagonal \(D\). Fix a nonempty path and a full noncrossing pairing of its links. Reinserting contracted adjacent pairs restores that same full-pairing prediction. Each contraction contributes its minus sign. Conversely, all subsets of the originally adjacent pairs of that full pairing occur, with precisely these signs. Such a full pairing has at least one originally adjacent pair, so its total coefficient is zero. Paths of odd length have no full pairing. There is no empty path contributing a derivative of a small seed, because all seeds are fixed with respect to \(x\); parameter derivatives were already put into the error. Thus the trace prediction vanishes as a formal power series. The polynomial-prediction assertion for single-inverse words in Lemma 2 then makes it zero at \(z=1\) in the closed formulas.
To pass from this zero prediction to a bounded trace, we use the small-vector structure. Every product in a closed formal derivative of a small vector contains either a small-vector diagonal or a scalar of size \(O(n^{-1/2})\): a path giving a nonzero derivative must enter the primary nodes before reaching the spin leaf. Entry from a small source uses either its primary partial, a small vector, or the average of that partial in a correction term. The latter has the indicated scalar size. For the implicit starting node these are exactly \(\partial_k w\) and \(c_1\) in (114).
For a term containing a small diagonal \(D_s\), rotate the tested trace to put that factor first. The word estimate for the remaining factors and Cauchy–Schwarz give \[\left|\mathop{\mathrm{Tr}}(D_sQ)-\sum_i s_i\,[\mathsf P(Q)]_{ii}\right|
\leq\lVert s\rVert\lVert \mathop{\mathrm{diag}}(Q-\mathsf P(Q))\rVert_2\leq C.\] The normalized-trace prediction is cyclically invariant by Lemma 2, so the rotation does not change the prediction being cancelled. For a term with a scalar of size \(O(n^{-1/2})\), its unscaled trace error is at most \(C\sqrt n\) by the diagonal \(\ell^2\) word estimate, and multiplication gives the same \(O(1)\) bound. The closed formula has finitely many terms. Combining their errors with the zero prediction and the already controlled differentiated errors gives \[|\mathop{\mathrm{Tr}}(D\nabla U)|\leq C\qquad(\lVert D\rVert\leq1,\ D\text{ diagonal}).\] Taking \(D_{ii}\) to be the sign of \(d_iU_i\) proves the remaining assertion \(\lVert \mathop{\mathrm{diag}}\nabla U\rVert_1\leq C\). All arguments were pointwise inside the prescribed buffers, which proves the local version as well. ◻
The weak residual rule
We now apply the diagonal \(\ell^1\) estimate to the exact conditional integration-by-parts identity at the first primary field. The scalar integral identity then reduces each higher residual to earlier ones. We carry out this induction for ordinary tests and square-density tests together.
Lemma 20 (Weak residual rule). Let \(U\) be a terminal vector in a fixed recipe admissible from level \(p\geq1\). For an ordinary recipe and every scalar \(G\), \[
\left|\mathbb EG R_p^{\mathsf T}U\right|\leq C\lVert G\rVert_*.
\tag{115}\] For a local recipe with implicit starting index \(k\), the same conclusion holds when \(p\leq k\) and \(G\) is supported on \(\Omega_\rho\).
There is also an ordinary square-density version. If \(\sup_x(|H_0(x)|+\lVert dH_0(x)\rVert)\leq C_0\), then for every \(f\), \[
\left|\mathbb Ef^2H_0 R_p^{\mathsf T}U\right|
\leq C\{\mathbb Ef^2+\mathcal D_{h_0}(f)\}.
\tag{116}\] Constants are uniform in \(f,G,H_0\) subject to the stated bound, and may depend on \(p\), the recipe, its coefficient bounds, and the finite set of word diagnostics needed in the proof.
Proof. We induct strongly on \(p\) over all finite admissible recipes, proving the two assertions together. Each step may enlarge the original recipe, but it adds only finitely many auxiliary operations and invokes only smaller residual indices. Consequently every particular instance of the induction uses finitely many operations and a finite word-length bound.
At the base level \(p=1\), condition on all spins except the one being differentiated. Its conditional mean is \(m_i^1\) and its conditional variance is \(v_i^1\). A function of this single spin is affine, with linear coefficient \(d_i\) of the function. Taking its conditional covariance with the spin gives the exact identity \[
\mathbb E\sum_i(x_i-m_i^1)GU_i
=\mathbb E\sum_i v_i^1d_i(GU_i).
\tag{117}\] By (2), its right side is \[\mathbb E\sum_i v_i^1Gd_iU_i
+\mathbb E\sum_i v_i^1(d_iG)\{U_i-2x_id_iU_i\}.\] The diagonal derivative bound makes the first term at most \(C\mathbb E|G|\). For the second term, the coefficient vector has pointwise norm at most \(\lVert U\rVert+2\lVert \mathop{\mathrm{diag}}\nabla U\rVert_2\leq C\). Weighted Cauchy–Schwarz and \(0<v_i^1\leq1\) therefore bound that term by \(C\mathcal D_{h_0}(G)^{1/2}\). This proves (115). For a supported local test, its derivatives vanish outside the one-flip neighborhood of its support; all bounds used here are available there.
For (116), absorb the bounded multiplier by replacing \(U\) in the base identity with \(\widetilde U=H_0U\). This vector still has bounded norm and bounded diagonal derivative \(\ell^1\) norm: the additional term obeys \(\sum_i|U_i d_iH_0|\leq\lVert U\rVert\lVert dH_0\rVert\), and the product-error term is bounded by the same argument using \(\lVert \mathop{\mathrm{diag}}\nabla U\rVert_2\). Apply (117) with \(G=f^2\) and use \[d_i(f^2)=2f\,d_i f-2x_i(d_i f)^2.\] The diagonal-derivative term is bounded by \(C\mathbb Ef^2\); the remaining terms are bounded by \(C\sqrt{\mathbb Ef^2\,\mathcal D_{h_0}(f)}+C\mathcal D_{h_0}(f)\). The base case is now proved in both forms. Before proceeding to higher residuals, we record how bounded scalar multipliers affect the tests. If \(|\vartheta|+\lVert d\vartheta\rVert\leq C\), then \[
\lVert \vartheta G\rVert_*\leq C'\lVert G\rVert_*.
\tag{118}\] Indeed arrange the product rule as \(d_i(\vartheta G)=\vartheta(x^{(i)})d_iG+G(x)d_i\vartheta\), square, and sum with weights \(v_i^1\leq1\). In the square-density case such a multiplier can be absorbed into \(H_0\); the product retains its bounded size and difference norm. For local tests only the buffer bounds are needed, since multiplication does not enlarge their support.
For the induction step, suppose both assertions hold through level \(p\) and take \(U\) admissible from level \(p+1\). The scalar integral identity suggests the following new source. Set \[a_p=j(b_p-b_{p-1}),\qquad
U'_i=\Phi(h_i^p;h_i^{p+1},a_p)U_i.\] This is a source admissible from level \(p\). Its coefficient and all needed derivatives are bounded because \(|a_p|\leq j\) and the bounds for \(\Phi\) proved with (95) hold uniformly in its two real field arguments. The level condition on \(U\) makes its site partial at level \(p\) vanish: \(\partial_pU\) differentiates only explicit primary arguments. Parameters and auxiliary arguments remain fixed, including when their values depend on earlier primary fields. Applying (95) coordinatewise therefore yields \[R_{p+1}^{\mathsf T}U
=(h^p-h^{p+1}-a_pm^p)^{\mathsf T}U'
-a_p n\langle \partial_pU'\rangle.\] The primary recursion gives \[h^p-h^{p+1}-a_pm^p
=JR_p+a_pR_p-jb_{p-1}R_{p-1},\] where the last term is absent for \(p=1\). Add a last auxiliary field \(y'\) by (93) with source \(U'\), and write \(\alpha_l=\langle \partial_lU'\rangle\) and \(\gamma_b=\langle \partial_{y^b}U'\rangle\). Symmetry of \(J\) now gives \[\begin{align*}
R_{p+1}^{\mathsf T}U
={}&R_p^{\mathsf T}y'+a_pR_p^{\mathsf T}U'
-jb_{p-1}R_{p-1}^{\mathsf T}U'
+j\sum_b\gamma_bR_p^{\mathsf T}w^b\\
&+j\sum_{l\geq p}\alpha_lR_p^{\mathsf T}m^{l-1}
-a_pn\alpha_p.
\tag{119}\end{align*}\] The first line now consists of earlier residuals with admissible sources and bounded multipliers. After testing, (118), (107), and the level constraints allow us to apply the induction hypothesis to each term. In particular the enlarged recipe is admissible from level \(p\), which is also sufficient for its \(R_{p-1}\) term.
On the second line, the terms with \(l>p\) use the terminal source \(m^{l-1}/\sqrt n\), admissible from level \(p\), and multiplier \(\sqrt n\alpha_l\). The multiplier bounds follow from Lemma 19. The remaining term is handled by the identity \[
R_p^{\mathsf T}m^{p-1}
=n(b_p-b_{p-1})-R_p^{\mathsf T}m^p.
\tag{120}\] This follows from \((m^{p-1}-m^p)^{\mathsf T}(m^{p-1}+m^p)
=\lVert m^{p-1}\rVert^2-\lVert m^p\rVert^2\). The first term on the right of (120) cancels \(a_pn\alpha_p\) in (119); the remaining term is handled with the admissible source \(m^p/\sqrt n\) and multiplier \(\sqrt n\alpha_p\). This completes the induction for both (115) and (116). In the local version it stops at level \(k\); each added source respects the support and buffer requirements of the initial implicit construction. ◻
The last step puts the implicit source into the form needed for the posterior comparison. The following identity is a direct application of the weak residual rule, with the averaged correction retained.
Lemma 21 (Terminal residual identity). For the implicit construction at index \(k\) and every \(G\) supported on \(\Omega_\rho\), \[
\left|\mathbb EG\left[
\{h^k-h_0-(J-jb_{k-1}I)m^k\}^{\mathsf T}w
-j(b_k-b_{k-1})n c_1\right]\right|
\leq C\lVert G\rVert_*.
\tag{121}\]
Proof. The vector in braces is \(JR_k-jb_{k-1}(R_{k-1}+R_k)\). To express its pairing with the implicit source in terms of earlier residuals, use symmetry of \(J\), the equation \(Jw=y^{\rm imp}+j\chi w+jc_1m^{k-1}\), and (120). The bracketed expression becomes \[R_k^{\mathsf T}y^{\rm imp}
+j(\chi-b_{k-1})R_k^{\mathsf T}w
-jb_{k-1}R_{k-1}^{\mathsf T}w
-jc_1R_k^{\mathsf T}m^k.\] Apply Lemma 20 to the first three terms, absorbing their bounded parameters into the test. For the last term, use source \(m^k/\sqrt n\) and multiplier \(\sqrt n c_1\). The bounds needed for this multiplier are precisely (101) and (102). ◻
Posterior estimates at a stable field
The residual rules now yield the two posterior estimates used in Section 2. We first select a finite collection of small-residual regions, then compare the primary iterates with the exact root. Throughout, \(J\) satisfies the required fixed-length word diagnostics, and \(h_0\) satisfies the whole stable-field implication through residual threshold \(\rho_0\sqrt n\). Denote the unique zero of \(F_{h_0}\) by \(r\) and set \(t_*=\tanh r\), \(q_*=\langle t_*^2\rangle\). All expectations and forms here are at \(\mu_{h_0}\), and \(R_l=m^{l-1}-m^l\) denotes the residual from Section 6.
Selecting a small-residual region
The primary recursion need not converge. The finite-depth fact we need is that two consecutive residuals are small on most of the Gibbs mass, with a separate bound after square reweighting. Smooth cutoffs make that selection compatible with discrete differentiation.
Lemma 22. For each sufficiently small fixed \(\rho>0\), there is a fixed depth \(d\) and functions \(C_k\), \(2\le k\le d\), with the following properties:
\(0\le C_k\le1\), and \(C_k\) is supported where \(\max(\lVert R_k\rVert,\lVert R_{k-1}\rVert)\le\rho\sqrt n\). It equals one where both norms are at most \(\rho\sqrt n/2\), and \(\lVert dC_k\rVert\le C_\rho/\sqrt n\) pointwise.
Set \(\widehat C_k=C_k\prod_{i=2}^{k-1}(1-C_i)\) and \(D=1-\sum_{k=2}^d\widehat C_k\). Then \(0\le D\le1\) and \[\mathbb P_{\mu_{h_0}}(D>0)\le\frac{C_\rho}{n}.\] For \(N=\mathbb Ef^2>0\) and \(\nu=f^2\mu_{h_0}/N\), \[\mathbb P_\nu(D>0)\le\frac{C_\rho}{\sqrt n}
\left(1+\frac{\mathcal D_{h_0}(f)}{N}\right).\]
All constants are uniform over the field and matrix under the stated diagnostics. The depth and the required diagnostic length may depend on \(\rho\).
Proof. We use the scalar residual pairing \(S_l=R_l^{\mathsf T}m^l/\sqrt n\) to control the total squared increment. First, the primary derivative bounds give \(\lVert dS_l\rVert\le C_l\) pointwise. In the product rule, the linear terms are controlled by the operator norms of \(\nabla R_l\) and \(\nabla m^l\), multiplied by \(\lVert m^l\rVert/\sqrt n\) and \(\lVert R_l\rVert/\sqrt n\). The cross term for coordinate \(i\) has size at most \(2\lVert d_i R_l\rVert\lVert d_i m^l\rVert/\sqrt n\); summing its square uses the bounded column norms and gives the same scale.
Apply Lemma 20 with target \(m^l/\sqrt n\) and test \(G=S_l\). It gives \[\mathbb ES_l^2\le C_l\bigl(\mathbb ES_l^2+C_l\bigr)^{1/2},
\qquad\text{hence}\qquad \mathbb ES_l^2\le C_l'.\] For the square-density version use \(H_0=S_l/\sqrt n\), which is bounded with bounded difference norm. Dividing the resulting estimate by \(N\sqrt n\) gives \[
\mathbb E_\nu\left(\frac{S_l}{\sqrt n}\right)^2
\le\frac{C_l}{\sqrt n}\left(1+\frac{\mathcal D_{h_0}(f)}{N}\right).
\tag{122}\] We now use the squared-increment identity (88). Choose a fixed \(d\) with \(\lfloor d/2\rfloor\rho^2/4>2\), followed by \(\varepsilon>0\) with \(2d\varepsilon<1\). Whenever \(\max_{l\le d}|S_l|/\sqrt n\le\varepsilon\), telescoping gives \(\sum_{l=1}^d\lVert R_l\rVert^2/n<2\). If every consecutive pair contained a residual larger than \(\rho\sqrt n/2\), the disjoint pairs \((1,2),(3,4),\ldots\) would force the sum to exceed this bound. At least one pair therefore meets the required threshold.
To implement the selection smoothly, apply a scalar cutoff equal to one on \([0,1/4]\) and zero on \([1,\infty)\) to each of \(\lVert R_k\rVert^2/(\rho^2n)\) and \(\lVert R_{k-1}\rVert^2/(\rho^2n)\), and multiply the results to define \(C_k\). The primary operator and column bounds, with the product rule, give \(\lVert d(\lVert R_l\rVert^2/n)\rVert\le C_l/\sqrt n\), hence the claimed derivative bound. Finally \(D=\prod_{k=2}^d(1-C_k)\): positive deficit implies \(\max_{l\le d}|S_l|/\sqrt n>\varepsilon\). The preceding two moment bounds and a finite union bound give the ordinary and square-reweighted estimates. ◻
A directional covariance bound
On each selected region, the implicit test vector converts the residual identity into a comparison with \(t_*\). Its auxiliary scalar error is estimated first; only then do we test in an arbitrary Euclidean direction.
Proposition 23. There is a constant \(C\) such that, at every field satisfying the above stable-field implication, \[
\lVert \mathop{\mathrm{Cov}}_{\mu_{h_0}}(G,x)\rVert
\le C\bigl(\mathop{\mathrm{Var}}_{\mu_{h_0}}(G)+\mathcal D_{h_0}(G)\bigr)^{1/2}
\tag{123}\] for every \(G\). The constant is uniform in \(n,h_0\), and \(G\) on the specified matrix events; only a fixed finite diagnostic length is required.
Proof. Fix an index \(k\ge2\) and work on the small-residual region (96), with the buffered implicit construction (100). Abbreviate \(m=m^k\), \(p=m-t_*\), and put \[a=j(1-b_{k-1}-q_*),\qquad
L_* =\sqrt n\bigl(\chi-(1-q_*)\bigr).\] Here \(\chi=\langle A_i\rangle\) and \(A_i=\Phi(h_i^k;r_i,a)\) are defined globally by (98), even though the implicit vector is only needed locally. Thus \(|L_*|\le C\sqrt n\) and \(\lVert dL_*\rVert\le C_k\) pointwise, by the primary derivative bounds and \(\lVert db_{k-1}\rVert=O(n^{-1/2})\).
For either of the two allowed sources \(w_0\), the following pointwise identity holds on the buffered region: \[\begin{align*}
&\bigl(h^k-h_0-(J-jb_{k-1}I)m\bigr)^{\mathsf T}w
-j(b_k-b_{k-1})nc_1 \\
&\quad=\sum_i\bigl[(h_i^k-r_i-am_i)-a\partial_k\bigr]w_{0,i}
+j(1-q_*-\chi)p^{\mathsf T}w-jc_1R_k^{\mathsf T}p.
\tag{124}\end{align*}\] Here the zeroth-order term in the first sum acts by multiplication on \(w_{0,i}\). For the verification, subtract the root equation to write the vector on the left as \(h^k-r-am-Jp+j(1-q_*)p\), then substitute \(Jw=y^{\rm imp}+j\chi w+jc_1m^{k-1}\). Apply (95) to the contribution \(A(y^{\rm imp}-jc_1t_*)\), holding its auxiliary arguments and parameters fixed, and use \(nc_1=\sum_i\partial_k w_i\). The remaining scalar terms cancel by \[p^{\mathsf T}(m^{k-1}+t_*)+n(b_k-b_{k-1})
=n(1-b_{k-1}-q_*)+R_k^{\mathsf T}p,\] which follows by expanding \(p=m-t_*\) and \(R_k=m^{k-1}-m\).
Lemma 21 controls the tested left-hand side of [eq:implicit-comparison]. The last term on the right also has tested expectation at most \(C\lVert G\rVert_*\): apply Lemma 20 to target \(p/\sqrt n\) and multiplier \(\sqrt n c_1\). The latter is bounded with bounded difference norm on the buffer. Meanwhile the size estimates for the implicit solution and Lemma 15 give \[
\lvert j(1-q_*-\chi)p^{\mathsf T}w\rvert\le C_0\rho|L_*|.
\tag{125}\] The source size and stability margins determine \(C_0\) independently of \(k\) and the selection depth. This lets us choose \(\rho\) first so that \(C_0\rho<1/2\) and all stability and flip-buffer conditions hold. The selection depth in Lemma 22 is chosen afterward. This order will permit absorption of the scalar error without a self-dependent choice of depth.
First control the averaged coefficient by taking \(w_{0,i}=\partial_{r_0}\Phi(h_i^k;r_i,a)/\sqrt n\). The second identity in (99) makes the first sum in [eq:implicit-comparison] equal to \(L_*\). Test with \(G=C_k^2L_*\). The cutoff and difference bounds give \[\lVert C_k^2L_*\rVert_*\le C+\lVert C_kL_*\rVert_{L^2(\mu_{h_0})}.\] Indeed its difference vector is bounded pointwise: \(|L_*|=O(\sqrt n)\) multiplies \(\lVert d(C_k^2)\rVert=O(n^{-1/2})\), and \(\lVert dL_*\rVert=O(1)\). Absorbing (125) now yields \[
\lVert C_kL_*\rVert_{L^2(\mu_{h_0})}\le C_k'.
\tag{126}\]
Now test a direction by taking \(w_0=Ae\), where \(e\) is any deterministic vector with \(\lVert e\rVert\le1\). By the first identity in (99), the first sum in [eq:implicit-comparison] is \(p^{\mathsf T}e\). Test with \(\widehat C_kG\). The residual estimates control the left side and the last term by \(C\lVert G\rVert_*\); the middle term is bounded by (126) and Cauchy–Schwarz, since \(0\le\widehat C_k\le C_k\). Thus \[\lvert \mathbb E[\widehat C_kG(m^k-t_*)^{\mathsf T}e]\rvert\le C\lVert G\rVert_*.\] Summing Lemma 20 with the constant target \(e\) over \(l=1,\ldots,k\) gives the same bound for \(\mathbb E[\widehat C_kG(x-m^k)^{\mathsf T}e]\). Sum over \(k\). The deficit has contribution at most \[2\sqrt n\,\lVert G\rVert_{L^2(\mu_{h_0})}
\mathbb P_{\mu_{h_0}}(D>0)^{1/2}\le C\lVert G\rVert_*.\] Thus \(|\mathbb E[G(x-t_*)^{\mathsf T}e]|\le C\lVert G\rVert_*\) uniformly over \(\lVert e\rVert\le1\). Centering \(G\) and taking the supremum over this seed class gives (123). The uncentered estimate also records the ordinary-mean consequence: take \(G=1\) to obtain \(\lVert \mathbb E_{\mu_{h_0}}x-t_*\rVert\le C\). This norm is the unnormalized Euclidean norm. It compares the mean itself with \(t_*\); no approximation of the derivative of the mean map is used here. The local construction required \(n\) above a fixed threshold. In the remaining dimensions, the elementary bounds \(\lVert \mathop{\mathrm{Cov}}(G,x)\rVert^2\le n\mathop{\mathrm{Var}}(G)\) and \(\lVert \mathbb E_{\mu_{h_0}}x-t_*\rVert\le2\sqrt n\) extend both conclusions after enlarging \(C\). ◻
Means under square reweighting
For the change of observation law we need a different estimate, valid for arbitrary square densities. It uses only ordinary recipes and trades a small fixed leading error for a constant depending on that tolerance.
Proposition 24. For every sufficiently small fixed \(\delta>0\), there is a constant \(C_\delta\) such that, for \(N=\mathbb E_{\mu_{h_0}}f^2>0\) and \(\nu=f^2\mu_{h_0}/N\), \[
\frac{\lVert \mathbb E_\nu x-t_*\rVert}{\sqrt n}
\le C\delta+\frac{C_\delta}{\sqrt n}
\left(1+\frac{\mathcal D_{h_0}(f)}{N}\right).
\tag{127}\] The leading constant \(C\) is independent of \(\delta\) and of the selection depth. This estimate requires only ordinary word diagnostics, of a fixed length that may depend on \(\delta\).
Proof. Select the cutoffs with a fixed threshold \(\rho\le\delta\) small enough for Lemma 15. On each cutoff support, the residual-to-root comparison gives \[\lVert m^k-t_*\rVert\le C\delta\sqrt n.\] Here the constant comes from the recurrence and the inverse bound in Lemma 15, and is independent of \(k\). For \(\lVert e\rVert\le1\), sum the square-density version of Lemma 20 with target \(e\) through \(k\), using multiplier \(\widehat C_k\). After division by \(N\), this bounds \[\lvert \mathbb E_\nu[\widehat C_k(x-m^k)^{\mathsf T}e]\rvert
\le C_\delta\left(1+\frac{\mathcal D_{h_0}(f)}{N}\right).\] The cutoff weights sum to at most one, so all root-comparison terms together cost at most \(C\delta\sqrt n\). The leftover mass contributes at most \(2\sqrt n\mathbb P_\nu(D>0)\), which is controlled by the second part of Lemma 22. Add the terms, divide by \(\sqrt n\), and take the supremum over \(e\) to obtain (127). The leading \(C\) comes from root stability; only \(C_\delta\) and the finite diagnostic length depend on the selected tolerance. Any dimensions below a fixed threshold required by this construction are covered by the elementary bound \(\lVert \mathbb E_\nu x-t_*\rVert/\sqrt n\le2\) after increasing \(C_\delta\); this leaves the leading constant unchanged. The argument has used no implicit test vector. ◻
The spectral gap at a large observation time
We complete the observation argument by proving a gap after a sufficiently large fixed observation time. The deterministic two-spin curvature estimate is uniform over external fields. The final probabilistic statement instead concerns the specified observation law, at fixed \(j<1\). We adapt the signed two-spin method of Wang [23], retaining the signed first-order interaction and estimating the remaining terms by an entrywise-square matrix and a quartic row error. The complete calculation below proves the precise variant used here.
For a symmetric matrix \(J\) with zero diagonal, define the symmetric matrices \[U_{ij}=\tanh J_{ij},\qquad Q_{ij}=U_{ij}^2,\qquad
r_4(U)=\max_i\sum_j U_{ij}^4.\] In particular, \(U_{ii}=Q_{ii}=0\). The square defining \(Q\) is entrywise. At a field \(h_0\), put \[Hf=\sum_i(I-P_i)f
=\sum_i(x_i-m_i^1)d_i f,\qquad
g_i=d_i f,\qquad a_i=v_i^1g_i.\] Each \(P_i\) is an orthogonal projection in \(L^2(\mu_{h_0})\), and therefore \(H\) is self-adjoint and positive, with \[
\mathbb E_{\mu_{h_0}}[fHf]=\mathcal D_{h_0}(f).
\tag{128}\]
Lemma 25 (Signed curvature inequality). There are universal constants \(u_0>0\) and \(C<\infty\) such that, for every \(n\), every symmetric zero-diagonal \(J\) satisfying \(\max_{i,j}|U_{ij}|\le u_0\), every \(h_0\in\mathbb R^n\), and every real function \(f\) on the discrete cube, \[
\mathbb E_{\mu_{h_0}}[(Hf)^2]
\ge (1-Cr_4(U))\mathcal D_{h_0}(f)
-\mathbb E_{\mu_{h_0}}\bigl[a^{\mathsf T}Ua
+C|a|^{\mathsf T}Q|a|\bigr].
\tag{129}\] Here \(|a|\) denotes coordinatewise absolute value.
Proof. The curvature bound is obtained by summing an exact two-spin calculation. Fix \(i\ne j\), condition on the other spins, and call the remaining spins \(x,y\). Their conditional density is proportional to \(\exp(\alpha x+\gamma y+J_{ij}xy)\). Introduce the scalar parameters \[a=\tanh\alpha,\quad b=\tanh\gamma,\quad
A=1-a^2,\quad B=1-b^2,\quad u=\tanh J_{ij}.\] The letters \(a,b,A,B\) have these scalar meanings only during the pair calculation. Use \(\mathbb E_0\) for the product law of means \(a,b\), and \(\mathbb E_{ij}\) for the interacting conditional law. The identity \(e^{J_{ij}xy}=\cosh(J_{ij})(1+uxy)\) gives their exact density ratio: \[
\frac{\mathrm d\mu_{ij}}{\mathrm d\mu_{0,ij}}(x,y)
=\frac{1+uxy}{1+uab}.
\tag{130}\] Expand \(f\) in the four-dimensional space of functions of \(x,y\). Its two gradients then have the unique representation \[g_i=p+r(y-b),\qquad g_j=s+r(x-a)\] with real \(p,s,r\). The shared coefficient \(r\) is the mixed-spin term of this expansion. The conditional means and variances satisfy \[\begin{align*}
x-m_i^1&=\frac{(x-a)(1-uxy)}{1+uay},&
v_i^1&=\frac{A(1-u^2)}{(1+uay)^2},\\
y-m_j^1&=\frac{(y-b)(1-uxy)}{1+ubx},&
v_j^1&=\frac{B(1-u^2)}{(1+ubx)^2}.
\tag{131}\end{align*}\]
For a function \(F\) of one spin, write \(F(z)=\overline F+z\,dF\), and set \[F_1(y)=\frac{g_i(y)}{1+uay},\qquad
F_2(x)=\frac{g_j(x)}{1+ubx}.\] The identity \((1-uxy)^2(1+uxy)=(1-u^2)(1-uxy)\) and independence under \(\mathbb E_0\) give the following exact expression for the cross term: \[
\frac{\mathbb E_{ij}[(x-m_i^1)g_i\,(y-m_j^1)g_j]}{AB}
=\frac{1-u^2}{1+uab}
\bigl(dF_1dF_2-u\overline F_1\overline F_2\bigr).
\tag{132}\] Both terms follow from the one-spin identities \(\mathbb E_0[(x-a)F_2(x)]=A\,dF_2\) and \(\mathbb E_0[(x-a)xF_2(x)]=A\overline F_2\), together with their \(y\) counterparts. Thus the cross term has been reduced to four scalar coefficients.
The four coefficients appearing here are \[\begin{align*}
dF_1&=\frac{r(1+uab)-uap}{1-u^2a^2},&
\overline F_1&=\frac{p-(b+ua)r}{1-u^2a^2},\\
dF_2&=\frac{r(1+uab)-ubs}{1-u^2b^2},&
\overline F_2&=\frac{s-(a+ub)r}{1-u^2b^2}.
\end{align*}\] Consequently the right-hand side of (132) is \[
K\bigl[(-u+u^2ab)ps+u^2bA\,rp+u^2aB\,rs+Rr^2\bigr],
\tag{133}\] where \[\begin{align*}
K&=\frac{1-u^2}{(1+uab)(1-u^2a^2)(1-u^2b^2)},\\
R&=(1+uab)^2-u(b+ua)(a+ub)\\
&=1+uab+u^2(a^2b^2-a^2-b^2)-u^3ab.
\end{align*}\] The mixed coefficients \(rp\) and \(rs\) begin at order \(u^2\): their first-order terms have canceled. Keeping this cancellation is what will leave only a quartic row cost after Young’s inequality.
For \(|u|\le1/4\) and \(|a|,|b|\le1\), all denominators above are bounded away from zero. There is thus a universal constant \(C_1\) such that \[|K-1|+|R-1|\le C_1|u|.\] Consequently the \(ps\) coefficient in (133) is \(-u+O(u^2)\), both mixed coefficients are \(O(u^2)\), and the \(r^2\) coefficient is \(1+O(|u|)\). These constants are uniform in the conditional means \(a,b\): all the displayed rational functions and their required derivatives are bounded on the compact parameter range.
Return temporarily to the notation \(a_i=v_i^1g_i\) and \(a_j=v_j^1g_j\) for the two gradient weights. Equations (130) and (131) imply \[
\frac{\mathbb E_{ij}[a_i a_j]}{AB}
=\mathbb E_0\bigl[w_u(x,y)g_i(y)g_j(x)\bigr],\qquad
w_u(x,y)=\frac{(1-u^2)^2(1+uxy)}
{(1+uab)(1+uay)^2(1+ubx)^2}.
\tag{134}\] Uniformly in \(a,b,x,y\), one has \[c_1\le w_u(x,y)\le C_2,\qquad
|w_u(x,y)-1|\le C_2|u|\qquad (|u|\le1/4)\] for universal positive constants \(c_1,C_2\). Expanding \(g_i(y)g_j(x)\) and using that the centered factors have zero product-law mean gives \[
\frac{\mathbb E_{ij}[a_i a_j]}{AB}
=ps+e_0ps+e_1rp+e_2rs+e_3r^2,
\qquad |e_\ell|\le C_3|u|.
\tag{135}\] To obtain the four error bounds, expand the product into its \(ps\), \(rp\), \(rs\), and \(r^2\) terms. Each centered spin is bounded by two, and the uniform estimate for \(w_u-1\) controls every resulting coefficient.
Adding \(u\) times (135) to (133), and decreasing the universal \(u_0>0\) if necessary, yields \[\begin{align*}
&\frac{\mathbb E_{ij}[(x-m_i^1)g_i\,(y-m_j^1)g_j]
+u\mathbb E_{ij}[a_i a_j]}{AB}\\
&\hspace{15mm}\ge
\frac12r^2-C_4u^2|ps|
-C_4u^2|r|(|p|+|s|)\\
&\hspace{15mm}\ge -C_5u^2|ps|-C_5u^4(p^2+s^2).
\end{align*}\] The last inequality is Young’s inequality, absorbing the mixed terms into \(r^2/2\).
We must now express the two scalar costs in the heat-bath weights that can be summed over pairs. The lower bound for \(w_u\) in (134), followed by product independence, gives \[\begin{align*}
\mathbb E_{ij}|a_i a_j|
&\ge c_1AB\,\mathbb E_0|g_i(y)g_j(x)|\\
&=c_1AB\,\mathbb E_0|g_i(y)|\,\mathbb E_0|g_j(x)|
\ge c_1AB|ps|.
\end{align*}\] Similarly, the density in (130) and the ratios \(v_i^1/A\), \(v_j^1/B\) are bounded above and below by universal positive constants for \(|u|\le u_0\). Hence \[\mathbb E_{ij}[v_i^1g_i^2]\ge c_2A\mathbb E_0g_i^2
\ge c_2Ap^2\ge c_2ABp^2,
\qquad
\mathbb E_{ij}[v_j^1g_j^2]\ge c_2ABs^2.\] The comparison constants do not deteriorate when a conditional variance \(A\) or \(B\) is small. Substituting these bounds into the preceding scalar inequality proves the pair estimate \[\begin{align*}
\mathbb E_{ij}[(x-m_i^1)g_i\,(y-m_j^1)g_j]
\ge{}&-u\mathbb E_{ij}[a_i a_j]-C_6u^2\mathbb E_{ij}|a_i a_j|\\
&-C_6u^4\mathbb E_{ij}[v_i^1g_i^2+v_j^1g_j^2].
\tag{136}\end{align*}\]
To assemble the curvature form, first condition on all but spin \(i\) in the diagonal term of \(\mathbb E(Hf)^2\). Their sum is \(\sum_i\mathbb E[v_i^1g_i^2]=\mathcal D_{h_0}(f)\). Next average (136) over the outside spins and sum over ordered pairs \(i\ne j\). The signed and absolute pair costs become the two quadratic forms in (129). The quartic cost is bounded by \[2C_6\,r_4(U)\sum_i\mathbb E[v_i^1g_i^2].\] Increasing \(C\) proves the lemma. ◻
The curvature inequality is useful when the few coordinates with appreciable conditional variance occupy a small principal submatrix. Bandeira, El Alaoui, and Rödder [5] combine uniform bounds on small principal submatrices with a sufficiently large constant field to prove polynomial mixing. Here the matrix bounds enter the signed curvature form directly. We next obtain matrix control uniform over every such coordinate set; the set will later depend on the spin configuration.
Lemma 26 (Small principal submatrices). Fix \(j\in(0,1)\). There is a finite constant \(M\), depending only on \(j\), with the following property. For every \(\varepsilon>0\) there is \(\alpha\in(0,1/2)\) such that, with probability tending to one over \(J\), \[\begin{gather*}
\max_{i,j}|J_{ij}|\le n^{-1/4},\qquad
\|J\|,\|U\|,\|Q\|\le M,\qquad
r_4(U)\le Mn^{-1/2},
\displaybreak[1]\tag{137}\\
\|U_{SS}\|\le\varepsilon,\qquad
\|Q_{SS}\|\le\varepsilon
\quad\text{for every }S\subseteq\{1,\ldots,n\}
\text{ with }|S|\le\alpha n.
\tag{138}\end{gather*}\] The same \(M\) works for all \(\varepsilon\).
Proof. For any deterministic unit vector \(z\), the centered Gaussian variable \(z^{\mathsf T}Jz\) has variance \[\frac{4j}{n}\sum_{i<k}z_i^2z_k^2\le\frac{2j}{n}.\] Thus \[
\mathbb P\{|z^{\mathsf T}Jz|>t\}\le
2\exp\{-nt^2/(4j)\}.
\tag{139}\] To pass from fixed quadratic forms to norms, choose a \(1/4\)-net of the unit sphere in \(\mathbb R^k\) with at most \(9^k\) points. Every symmetric \(k\times k\) matrix \(B\) then satisfies \[\|B\|\le2\max_{z\text{ in the net}}|z^{\mathsf T}Bz|;\] to see this, approximate a unit vector maximizing the absolute quadratic form and bound the change by \(2\|B\|/4\). Taking \(k=n\) in (139) proves \(\|J\|\le M_0\) with probability at least \(1-e^{-c n}\) for a sufficiently large constant \(M_0=M_0(j)\).
For a fixed row \(i\) and a fixed set \(S\) of at most \(k\) indices, Gaussian independence within that row gives \[
\mathbb P\left\{\sum_{\ell\in S}J_{i\ell}^2>t\right\}
\le \exp\left\{-\frac{nt}{4j}+\frac{k}{2}\log2\right\}.
\tag{140}\] Use exponential Markov with parameter \(n/(4j)\) to obtain this estimate: each nonzero Gaussian square contributes the factor \(\sqrt2\) to the moment generating function. Choosing the full index set and taking a union bound over rows bounds every row square sum by a fixed \(M_1=M_1(j)\), with failure at most \(e^{-cn}\). Separately, the Gaussian entry tail and a union bound give \[\mathbb P\{\max_{i,\ell}|J_{i\ell}|>n^{-1/4}\}
\le 2n^2\exp\{-\sqrt n/(2j)\}=o(1).\] Since \(|\tanh z|\le|z|\), these bounds imply \(\|Q\|\le M_1\) and \(r_4(U)\le M_1n^{-1/2}\), using the maximal absolute row-sum bound for the operator norm of a symmetric matrix. Moreover, \[|\tanh z-z|\le |z|^3/3,
\qquad
\|U-J\|\le \frac13\max_i\sum_\ell|J_{i\ell}|^3
\le\frac{M_1}{3}n^{-1/4}.\] The scalar inequality follows by integrating \(|1-\mathop{\mathrm{sech}}^2 z|=\tanh^2 z\le z^2\). This proves (137), with a fixed \(M\).
The global bounds have fixed \(M\); we now select the small-set fraction \(\alpha\). Put \(k=\lfloor\alpha n\rfloor\). Bounding sets of size exactly \(k\) suffices, because a smaller set can be enlarged and restriction cannot increase a principal-submatrix norm. The \(\binom nk\) possible sets have logarithmic count at most \(n\mathfrak h(k/n)\), where \[\mathfrak h(t)=-t\log t-(1-t)\log(1-t),\] and \(\mathfrak h(t)\to0\) as \(t\downarrow0\). Applying the sphere-net argument on each support and using (139) gives \[\mathbb P\left\{\max_{|S|=k}\|J_{SS}\|>\varepsilon/2\right\}
\le 2\binom nk9^k
\exp\{-n\varepsilon^2/(64j)\}.\] Choose \(\alpha>0\) sufficiently small that this tends to zero exponentially in \(n\). Then \(\|U-J\|=o(1)\) proves the assertion for \(U\).
Finally, (140), with \(t=\varepsilon\), and a union bound over all rows and all \(k\)-element sets give \[\mathbb P\left\{\max_i\max_{|S|=k}\sum_{\ell\in S}J_{i\ell}^2
>\varepsilon\right\}
\le n\binom nk
\exp\left\{-\frac{n\varepsilon}{4j}+\frac{k}{2}\log2\right\}.\] Decreasing \(\alpha\) further makes this probability exponentially small. On its complement, the maximal row sum of every \(Q_{SS}\) is at most \(\varepsilon\), which proves the remaining assertion. ◻
We apply these matrix bounds to a noisy observation. The following statement allows any initial spin law, with the interaction still drawn at a fixed \(j\in(0,1)\); the exceptional probability concerns the observation at that fixed matrix.
Proposition 27 (Terminal observation gap). Fix \(j\in(0,1)\). There are finite \(T>0\), \(c>0\), and events \(\mathcal E_n^{\mathrm{term}}\) depending only on \(J\), with \(\mathbb P_J(\mathcal E_n^{\mathrm{term}})\to1\), such that the following holds for every sufficiently large \(n\) and every \(J\in\mathcal E_n^{\mathrm{term}}\). Let \(X\) have any probability law on \(\{-1,1\}^n\), possibly depending on \(J\), and let \(Z\) be an independent standard Gaussian vector. Set \(Y=TX+\sqrt T\,Z\). Then, conditionally on \(J\), except on an event of probability at most \(e^{-cn}\), the inequality \[
\mathop{\mathrm{Var}}_{\mu_Y}(f)\le 2\mathcal D_Y(f)
\tag{141}\] holds simultaneously for every real function \(f\) on the cube. All constants and matrix events are independent of the law of \(X\).
Proof. Choose the tolerances in the order needed to absorb the curvature costs. Let \(C\) be the constant in Lemma 25, and choose \(\varepsilon>0\) with \((1+C)\varepsilon\le1/8\). Lemma 26 supplies \(\alpha\) and the fixed global bound \(M\). Finally choose the fixed observation time \(T\) large enough that \[
\frac{16M^2}{T^2}\le\frac\alpha2,
\qquad
(1+C)M\bigl(2\mathop{\mathrm{sech}}(T/4)+\mathop{\mathrm{sech}}^2(T/4)\bigr)\le\frac18,
\qquad
e^{-T/8}\le\frac{\alpha}{4(e-1)}.
\tag{142}\] Let \(\mathcal E_n^{\mathrm{term}}\) be the events supplied by that lemma. For all sufficiently large \(n\), their bounds also ensure \(\max|U_{ij}|\le u_0\) and \(Cr_4(U)\le1/8\).
First count coordinates whose observation field remains small. Given \(X\), the indicators \(\mathbf 1_{\{|Y_i|<T/2\}}\) are independent; symmetry of the Gaussian noise gives the same success probability \(p_T\) for either sign of \(X_i\). The Gaussian tail bound yields \[p_T=\mathbb P\{|T+\sqrt T Z_1|<T/2\}
\le\mathbb P\{Z_1<-\sqrt T/2\}\le e^{-T/8}.\] For \(B_Y=\#\{i:|Y_i|<T/2\}\), exponential Markov inequality yields \[\mathbb P\{B_Y>\alpha n/2\mid X,J\}
\le e^{-\alpha n/2}(1+p_T(e-1))^n
\le e^{-\alpha n/4}.\] The same bound holds after averaging over \(X\).
Suppose now that \(B_Y\le\alpha n/2\). For every spin configuration \(x\), \(\|Jx\|^2\le M^2n\), and consequently \[\#\{i:|(Jx)_i|>T/4\}\le\frac{16M^2n}{T^2}\le\frac\alpha2 n.\] It follows, simultaneously for all \(x\), that \[
S(x)=\{i:|Y_i+(Jx)_i|<T/4\}
\quad\text{satisfies}\quad |S(x)|\le\alpha n.
\tag{143}\] At this fixed field \(Y\), set \(q_f(x)=\sum_i v_i^1(x)g_i(x)^2\) and \(\eta=\mathop{\mathrm{sech}}^2(T/4)\). Since \(0<v_i^1\le1\), one has pointwise \[\|a\|^2\le q_f(x),\qquad
\|a_{S(x)^c}\|^2\le\eta q_f(x).\] We may now use \(S(x)\) even though it depends on the configuration, because (138) holds for every small set on one matrix event. Split the quadratic form between \(S(x)\) and its complement. The small-set bound handles its first block, and the global norm handles the cross and complementary terms, giving \[a^{\mathsf T}Ua
\le\bigl[\varepsilon+M(2\sqrt\eta+\eta)\bigr]q_f(x).\] The identical argument, with \(|a|\) and \(Q\), gives the same upper bound for \(|a|^{\mathsf T}Q|a|\). Lemma 25 and (142) therefore imply \[\begin{align*}
\mathbb E_{\mu_Y}[(Hf)^2]
&\ge\left[1-Cr_4(U)
-(1+C)\{\varepsilon+M(2\sqrt\eta+\eta)\}\right]\mathcal D_Y(f)\\
&\ge\frac12\mathcal D_Y(f).
\end{align*}\]
It remains to translate curvature into a gap. The kernel of the self-adjoint operator \(H\) contains exactly the constants: vanishing \(\mathbb E[fHf]\) forces every one-spin conditional variance to vanish, and the strict positivity of the Gibbs law propagates equality across every cube edge. For an eigenfunction with positive eigenvalue \(\lambda\), the curvature estimate and (128) imply \(\lambda^2\ge\lambda/2\), hence \(\lambda\ge1/2\). Expand in an orthonormal eigenbasis to obtain (141). The observation failure bound was \(e^{-\alpha n/4}\), uniformly in the initial spin law, so we may set \(c=\alpha/4\). ◻
A. Adhikari, C. Brennecke, C. Xu, and H.-T. Yau, Spectral gap estimates for mixed \(p\)-spin models at high temperature, Probability Theory and Related Fields 189 (2024), nos. 3–4, 879–907. doi:10.1007/s00440-024-01261-9.
M. Aizenman, J. L. Lebowitz, and D. Ruelle, Some rigorous results on the Sherrington–Kirkpatrick spin glass model, Communications in Mathematical Physics 112 (1987), 3–20. doi:10.1007/BF01217677.
N. Anari, V. Jain, F. Koehler, H. T. Pham, and T.-D. Vuong, Entropic independence: optimal mixing of down-up random walks, Proceedings of the 54th Annual ACM SIGACT Symposium on Theory of Computing (STOC 2022), 1418–1430. doi:10.1145/3519935.3520048. Full Ising proof: arXiv:2106.04105v2.
N. Anari, F. Koehler, and T.-D. Vuong, Trickle-down in localization schemes and applications, Proceedings of the 56th Annual ACM Symposium on Theory of Computing (STOC 2024), 1094–1105. doi:10.1145/3618260.3649622.
A. S. Bandeira, A. El Alaoui, and A. Rödder, Mixing of Glauber Dynamics on High Overlap Gibbs Measures, arXiv:2607.06813v1, July 7, 2026.
R. Bauerschmidt and T. Bodineau, A very simple proof of the LSI for high temperature spin systems, Journal of Functional Analysis 276 (2019), no. 8, 2582–2588. doi:10.1016/j.jfa.2019.01.007.
M. Boban, A. Li, and S. Oveis Gharan, Rank-1-perturbed trickledown theorems: Mixing time of Glauber dynamics for the Sherrington–Kirkpatrick model up to \(\beta\leq\frac12+\varepsilon\), arXiv:2609.13138v1, September 11, 2026.
E. Bolthausen, An iterative construction of solutions of the TAP equations for the Sherrington–Kirkpatrick model, Communications in Mathematical Physics 325 (2014), no. 1, 333–366. doi:10.1007/s00220-013-1862-3.
C. Brennecke, A. Schertzer, C. Xu, and H.-T. Yau, The two point function of the SK model without external field at high temperature, Probability and Mathematical Physics 5 (2024), 131–175. doi:10.2140/pmp.2024.5.131.
M. Celentano, Sudakov–Fernique post-AMP, and a new proof of the local convexity of the TAP free energy, Annals of Probability 52 (2024), no. 3, 923–954. doi:10.1214/23-AOP1675.
Y. Chen and R. Eldan, Localization schemes: A framework for proving mixing bounds for Markov chains, Duke Mathematical Journal 174 (2025), no. 8, 1431–1510. doi:10.1215/00127094-2024-0063. Preprint: arXiv:2203.04163v2.
A. El Alaoui and J. Gaitonde, Bounds on the covariance matrix of the Sherrington–Kirkpatrick model, arXiv:2212.02445v3, November 12, 2024.
A. El Alaoui and A. Montanari, An information-theoretic view of stochastic localization, arXiv:2109.00709v2, September 9, 2021.
A. El Alaoui, A. Montanari, and M. Sellke, Sampling from the Sherrington–Kirkpatrick Gibbs measure via algorithmic stochastic localization, Proceedings of the 63rd IEEE Annual Symposium on Foundations of Computer Science (FOCS 2022), 323–334. Full version: arXiv:2203.05093v2.
R. Eldan, Thin shell implies spectral gap up to polylog via a stochastic localization scheme, Geometric and Functional Analysis 23 (2013), no. 2, 532–569. doi:10.1007/s00039-013-0214-y.
R. Eldan, F. Koehler, and O. Zeitouni, A spectral condition for spectral gap: fast mixing in high-temperature Ising models, Probability Theory and Related Fields 182 (2022), nos. 3–4, 1035–1051. doi:10.1007/s00440-021-01085-x.
R. J. Glauber, Time-dependent statistics of the Ising model, Journal of Mathematical Physics 4 (1963), no. 2, 294–307. doi:10.1063/1.1703954.
J. Lehec, Representation formula for the entropy and functional inequalities, Annales de l’Institut Henri Poincaré, Probabilités et Statistiques 49 (2013), no. 3, 885–899. doi:10.1214/11-AIHP464.
D. Sherrington and S. Kirkpatrick, Solvable model of a spin-glass, Physical Review Letters 35 (1975), no. 26, 1792–1796. doi:10.1103/PhysRevLett.35.1792.
R. Speicher, Combinatorial theory of the free product with amalgamation and operator-valued free probability theory, Memoirs of the American Mathematical Society 132 (1998), no. 627. doi:10.1090/memo/0627.
D. J. Thouless, P. W. Anderson, and R. G. Palmer, Solution of ‘Solvable model of a spin glass’, Philosophical Magazine 35 (1977), no. 3, 593–601. doi:10.1080/14786437708235992.
R. A. Vitale, Some comparisons for Gaussian processes, Proceedings of the American Mathematical Society 128 (2000), no. 10, 3043–3046. doi:10.1090/S0002-9939-00-05367-3.
S. Wang, Optimal Mixing of Glauber Dynamics for the Sherrington–Kirkpatrick Model at \(\beta<1/2\), arXiv:2608.22159v2, August 28, 2026.
L. Wu, Poincaré and transportation inequalities for Gibbs measures under the Dobrushin uniqueness condition, Annals of Probability 34 (2006), no. 5, 1960–1989. doi:10.1214/009117906000000368.
LEVEL 3 COMPLETE!
You read 26,827 words and 2,323 formulas. Your math teacher would be proud. Converted from the LaTeX source. Something look off? The original PDF is the real thing.