A
D
V
E
R
T
I
S
E
M
E
N
T
ADVERTISEMENT
The spherical perceptron with bi-orthogonally invariant disorder
expertly designed by an internal OpenAI model  ·  released 2026-09-24  ·  original PDF
Theorems: 3 Lemmas: 4 Proofs: 7
Formulas: 1,255 Words: 18,511 Play time: ~2 hours

>>> How to Play <<<
We determine the limiting free energy of a spherical perceptron with a bounded continuous activation and a bi-orthogonally invariant disorder matrix. The singular values may have any compact limiting distribution, provided there are no outliers. We give an explicit variational formula for this limit.

>>> Level Map <<<
  1. Introduction
  2. The variational formula
  3. Angular pressure and the spherical theorem
  4. A Gaussian model with outer and inner columns
  5. Paths with a shared coordinate
  6. Explicit evaluation of the angular pressure
  7. The Gaussian theorem and its probabilistic tools
  8. Marked cascades, zero restrictions, and visit variables
  9. Joint overlap identities and ordered matrix paths
  10. Fresh fields and the one-pattern variation
  11. The Gaussian canonical calculation
  12. A canonical multiplier
  13. From canonical rows to Stiefel columns
  14. The Gaussian upper bound
  15. GG at the deterministic contact point
  16. The contradiction from derivatives
  17. The Gaussian cavity lower bound
  18. Generic parameters and two-pattern identification
  19. Splitting a Stiefel frame
  20. Expansion and fixed-size integrability
  21. Matching physical and canonical overlaps
  22. The expected logarithmic cost of the shell
  23. Cancellation and telescoping
  24. Quadratic tails and global soft constraints
  25. Extension of the terminal class
  26. Differentiability and the powered restriction ratio
  27. Soft duality
  28. Haar flags and weighted shells
  29. Gaussian spans and shell entropy
  30. A majorization principle without concavity
  31. The coefficient integral and triangular entropy
  32. Upper and lower identification of the shell
  33. Full output dimension and bounded continuous activations
  34. Radial decomposition and proof of the main theorem
  35. Spectral and activation scope
  36. General dimension sequences
  37. Negligible spectral groups inside the edge interval
  38. Activation extensions
  39. Examples delimiting the hypotheses

Introduction

A spherical perceptron assigns a weight to each point on a sphere by evaluating a scalar potential on its linear measurements. Let \(A\) be a real \(M\times N\) matrix, and let \(\sigma_N\) be uniform probability measure on the sphere \(\{x\in\mathbb R^N:\lVert x\rVert^2=N\}\). For a bounded continuous function \(f:\mathbb R\to\mathbb R\), the normalized log-partition pressure is \[ F_N(A,f)=\frac1N\log\int_{\lVert x\rVert^2=N} \exp\left\{\sum_{r=1}^M f((Ax)_r)\right\}\,d\sigma_N(x). \tag{1}\] A temperature parameter is included by writing \(f=\beta\phi\). We study the limit as \(M/N\to\alpha>0\) when the law of \(A\) is invariant under left and right orthogonal multiplication. Such a matrix has Haar singular spaces, and its singular values describe the strength and rank of the measurements. Our objective is a variational formula that retains this spectral information for a general bounded continuous potential.

A central question in the statistical mechanics of perceptrons is the volume of weight vectors that satisfy prescribed random pattern constraints. Gardner and Gardner–Derrida developed this viewpoint for spherical and binary coupling constraints [6, 7]. Gardner and Derrida also studied an error-counting partition function and identified regimes in which its replica-symmetric saddle point becomes unstable. For nonnegative-margin hard constraints, Shcherbina and Tirozzi rigorously established the replica-symmetric formula for the typical spherical feasible volume [28].

The finite-temperature overlap-path formalism also has substantial antecedents. Györgyi and Reimann derived a nested Gaussian recursion and spherical entropy within a continuous replica-symmetry-breaking ansatz, retaining a general scalar potential in their formulation and specializing their calculations to error counting beyond storage capacity [9]. For independent Gaussian patterns, Montanari and Zhou give a replica derivation of a finite-temperature variational functional for general projection potentials [15]; their rigorous theorems concern the performance of approximate-message-passing algorithms. These variational predictions motivate an overlap hierarchy beyond the regimes described by one scalar order parameter.

Prescribed singular spectra and independent Haar singular bases already appear in Kabashima’s framework for inference from correlated patterns [11] and in Shinzato and Kabashima’s analysis of perceptron capacity [29]. Shinzato and Kabashima derive replica-symmetric and one-step replica-symmetry-breaking expressions, and apply the replica-symmetric calculation to spherical weights. Abbara, Aubin, Krzakala, and Zdeborová subsequently used such finite-temperature free-energy calculations for rotationally invariant perceptrons in their study of Rademacher complexity [1]. These works are direct predecessors of the present spectral model and its variational description. The present theorem establishes a limiting pressure through a full matrix-path variational problem for bounded continuous potentials, with compact limiting singular spectrum and no outliers.

There are two obstacles to extending a Gaussian perceptron calculation. First, fixing the singular values makes the matrix entries dependent; scalar covariances of individual measurements no longer determine the model. Second, replacing orthogonal subspaces by Gaussian spans introduces auxiliary vectors whose entropy would change the desired pressure. We address the first issue by conditioning on the norm allocated to each singular-value block and computing the remaining angular pressure. We address the second by imposing one-sided distance tests to a nested sequence of Haar subspaces and integrating their witnesses with a small positive exponent. The exponent tends to zero only after the large-dimension limit; it is never a continuation in replica number.

The geometry is supported by a weighted shell comparison. After sorting the target squared norm per block dimension, a redistribution of squared norms toward earlier blocks cannot increase the limiting weighted shell volume. Random subdivision of Haar subspaces proves this comparison by transferring weighted volume, even when the activation concentrates that weight. Consequently the one-sided projection tests recover the angular pressure without requiring concavity. Gaussian orthogonalization supplies the underlying Haar frames, as in the classical constructions described by Mezzadri [14].

The remaining Gaussian problem involves several Stiefel columns, with one column shared at an intermediate level of a cascade. Its upper bound adapts Mourrat’s enriched-free-energy and supersolution contact-point method [16, 18]. The hierarchical Gaussian fields used in this method are part of the Hamilton–Jacobi framework developed by Mourrat [17] and Mourrat and Panchenko [19]. These works provide the comparison framework; their theorems do not directly cover the present nonlinear pattern Hamiltonian. Here the interaction parameter is the number of patterns, and the field parameters are matrix covariance increments on the active Stiefel coordinates. The proof establishes the corresponding contact identities while retaining the prescribed sharing of one coordinate.

The matrix overlap argument draws on Panchenko’s vector-spin synchronization theory [24] and its enriched formulation in [18]. Related matrix spherical entropy formulas were developed by Ko [13], and the scalar spherical cascade transform appears in [17]. The entropy calculation here includes the change of active spin space at the sharing level. For the lower bound, restoring two independent patterns identifies the covariance of the field acting on newly added spin coordinates. A Gaussian calculation predicts one overlap path for those coordinates, while spatial rotation invariance predicts the overlap path of the original system. Comparing the two laws on a thick shell forces these paths to agree and identifies the entropy in the lower bound. The companion scalar paper [20] gives the corresponding scalar comparison and cavity approach; no theorem from it is needed below.

The probabilistic framework is the cascade and overlap theory of spin glasses, including the marked-cascade identities of Panchenko and Talagrand [26], Panchenko’s scalar ultrametricity theorem [23], and the vector-spin synchronization framework [24]. We retain independent residual Gaussian noise on every visit to a cascade leaf and prove the marking rules for the fractional moments and zero restrictions used here. Section 2 states the angular and spherical theorems, then gives the explicit functional. Sections 3–7 establish the Gaussian formula and global soft penalties. Section 8 proves the angular reduction, and Section 9 performs the radial and spectral limits. Appendix 10 allows arbitrary dimension sequences, weakens spectral support control to control by the interval between its edges, and treats continuous subquadratic and bounded Borel activations under their stated additional hypotheses. Its examples explain why the spectral and activation qualifications matter.

The variational formula

The formula separates the output direction from the allocation of input norm among singular-value blocks. We first state this decomposition using an angular pressure with a direct probabilistic meaning. We then compute that pressure through a Gaussian model with projection witnesses.

All sphere and Stiefel measures are probability measures. A matrix inequality is in the Loewner order. Determinants and inverses on the zero-dimensional space are, respectively, \(1\) and the zero operator.

Angular pressure and the spherical theorem

Let \(E_1,\ldots,E_h\subset\mathbb R^M\) be mutually orthogonal Haar subspaces with \(\dim E_j=n_j\) and \(n_j/M\to p_j>0\). Thus \(p:=\sum_jp_j\le1\). For \(a_j\ge0\), let \(\sigma_{E_j,Ma_j}\) be uniform probability measure on the sphere in \(E_j\) of squared radius \(Ma_j\), with the point mass at zero when \(a_j=0\). Define the random angular pressure \[ P_M^f(a)=\frac1M\log\int\exp\left\{\sum_{r=1}^M f\left(\sum_jy_{j,r}\right)\right\} \prod_jd\sigma_{E_j,Ma_j}(y_j). \tag{2}\] The subspaces are held fixed in this integral; the vectors \(y_j\) are integrated independently. The subscript \(p\) on its limit records the entire list of proportions \((p_j)_j\).

Theorem 1 (Angular formula). For \(\sum_jp_j\le1\) and \(f\in C_b(\mathbb R)\), \(P_M^f(a)\) converges in probability, uniformly on compact sets of \(a\in[0,\infty)^h\), to a deterministic continuous function \(P_p^f(a)\). Its explicit evaluation is given by (13) and (14) below. For \(p<1\) and positive \(a\), (13) holds for every sufficiently large fixed coefficient radius \(\kappa\).

With this angular limit, the remaining optimization describes how a spherical input divides its norm among the singular-value blocks. For the pressure (1), let \(r_N=\min(M,N)\) and let \(s_{N,1},\ldots,s_{N,r_N}\) be the singular values of \(A\), including zero values in these slots. Structural right or left null coordinates outside these slots are kept separate.

Theorem 2. Fix \(\alpha>0\), \(M=\lfloor\alpha N\rfloor\), and \(f\in C_b(\mathbb R)\). Suppose the law of \(A_N\) is invariant under left and right orthogonal multiplication and, in probability, \[ \frac1{r_N}\sum_{r=1}^{r_N}\delta_{s_{N,r}}\Longrightarrow\nu, \qquad \max_{r\le r_N}\operatorname{dist}(s_{N,r},\operatorname{supp}\nu)\longrightarrow0, \tag{3}\] where \(\nu\) is a deterministic compactly supported probability measure on \([0,\infty)\). Then \(F_N(A_N,f)\) converges in probability and in mean to a deterministic value \(F_{\alpha,\nu}(f)\).

Here is its formula. Approximate the singular slots in operator norm by finitely many positive values \(s_1,\ldots,s_h\) whose multiplicities satisfy \(n_j/M\to p_j>0\). Thus \(\sum_jp_j=\min(1,1/\alpha)\). Put \(v_j=\alpha p_j\), and, if \(\alpha<1\), add an input null block with \(v_0=1-\alpha\). Let \(\mathcal I=\{1,\ldots,h\}\) when \(\alpha\ge1\) and \(\mathcal I=\{0,1,\ldots,h\}\) otherwise. The variable \(\rho_i\) is the fraction of the input squared norm assigned to block \(i\). The value for this finite approximation is \[ \sup_{\substack{\rho_i>0\\\sum_{i\in\mathcal I}\rho_i=1}} \left\{ \frac12\sum_{i\in\mathcal I}v_i\log\frac{\rho_i}{v_i} +\alpha P_p^f\left(\left(\frac{s_j^2\rho_j}{\alpha}\right)_{j=1}^h\right) \right\}. \tag{4}\] The limit of (4) as the spectral mesh tends to zero is \(F_{\alpha,\nu}(f)\). Blocks with zero limiting proportion are not used in (4); the no-outlier assumption ensures the required approximations with positive limiting proportions.

All formulas in this section apply directly to bounded continuous \(f\). They may equivalently be evaluated through a final bounded Lipschitz approximation, uniform on compact intervals; the proof shows that the two prescriptions agree. Conditional deterministic singular values are allowed; the theorem does not require a density, a connected support, or a free-probability regularity condition on \(\nu\).

A Gaussian model with outer and inner columns

The auxiliary Gaussian model distinguishes the vector whose weight we seek from the vectors that witness its projection constraints. For now fix arbitrary column counts \(d_j\ge1\) and a positive exponent \(0<\theta<1\); the application will specify \(d_j\). Here we allow arbitrary positive proportions \(p_j\), without restricting their sum.

Let \(n_j/M\to p_j>0\), and let \[X_j\in\operatorname{St}(n_j,d_j):=\{X\in\mathbb R^{n_j\times d_j}:X^{\mathsf T}X=n_jI_{d_j}\}\] have uniform law. The first columns, jointly across species, are outer variables; the remaining columns have the conditional Stiefel law and are inner variables. For independent standard Gaussian vectors \(g_{r,j}\in\mathbb R^{n_j}\), define the row fields \[\xi_{r,j}(X)=\frac{X_j^{\mathsf T}g_{r,j}}{\sqrt{n_j}},\qquad 1\le r\le M.\] For a Hamiltonian \(H\) set \[ \mathcal Z_{M,\theta}(H)= \left[\int\left(\int e^H\,d\mu_{\rm in}\right)^\theta d\mu_{\rm out}\right]^{1/\theta}. \tag{5}\] The natural tilted path law first samples the outer variables with density proportional to the inner partition function to power \(\theta\), and then samples the tilted conditional inner law.

The first column is therefore drawn once before the inner columns are integrated. In the cascade representation it is shared by visits whose common level is at least \(\theta\); the other columns remain visit variables. The matrix paths below record the covariances of Gaussian row fields seen by such visits. Theorem 3 will identify the limit of the normalized log integral with their variational functional.

Paths with a shared coordinate

Keep the block proportions, column counts, and exponent from Section 2.2. The first coordinate in each block is shared at level \(\theta\).

Choose a common finite grid \[ 0=\zeta_0<\zeta_1<\cdots<\zeta_k=1, \qquad w_i=\zeta_{i+1}-\zeta_i\quad(0\le i<k), \tag{6}\] containing \(\theta\). A trial in block \(j\) consists of symmetric matrices \[0\preceq q_{j,0}\preceq\cdots\preceq q_{j,k-1}\preceq I_{d_j}, \qquad q_{j,-1}=0,\quad q_{j,k}=I_{d_j},\] with \(q_{j,i}e_1=e_1\) whenever \(\zeta_i\ge\theta\). Regard \(q_j(u)=q_{j,i}\) on \([\zeta_i,\zeta_{i+1})\). Let \(F_{j,i}=\mathbb R^{d_j}\) before the sharing level and \(F_{j,i}=e_1^\perp\) thereafter; put \(F_{j,k}=\{0\}\). Define the susceptibility and precision \[ \mathsf K_{j,i}=\sum_{r=i+1}^k\zeta_r(q_{j,r}-q_{j,r-1}), \qquad \mathsf P_{j,i}=(\mathsf K_{j,i}|_{F_{j,i}})^{-1}. \tag{7}\] An eligible trial has strictly positive susceptibility on each nonempty active space. Extend its precision by zero on the orthogonal complement. Let \(q_j^0\) be zero on active coordinates and the identity on shared coordinates, with \(q_{j,k}^0=I\), and define \(\mathsf K_{j,i}^0\) in the same way. The block entropy is \[ 2S_\theta(q_j)=\operatorname{tr}(\mathsf P_{j,0}q_{j,0})+ \sum_{i=1}^k\frac1{\zeta_i} \log\left( \frac{\det\mathsf K_{j,i-1}}{\det\mathsf K_{j,i}} \frac{\det\mathsf K_{j,i}^0}{\det\mathsf K_{j,i-1}^0} \right). \tag{8}\] Each determinant is taken on the active space at its own index. In particular the normalization compensates for the dimension change at \(\theta\).

For a continuous terminal \(D:\prod_j\mathbb R^{d_j}\to\mathbb R\) satisfying \(-C(1+|z|^2)\le D(z)\le C\) for some \(C<\infty\), let \(Z_0,\ldots,Z_k\) be independent centered Gaussian vectors with block-diagonal covariances \(\operatorname{diag}_j(q_{j,i}-q_{j,i-1})\). Set \(U_k=D\) and define backwards, for \(i=k,k-1,\ldots,1\), \[U_{i-1}(z)=\frac1{\zeta_i}\log\mathbb E\exp\{\zeta_i U_i(z+Z_i)\}.\] The single-pattern value is \(V_D(q)=\mathbb EU_0(Z_0)\); this last ordinary expectation averages the root Gaussian data. Define \[ \mathcal G_{\theta,p}(D)=\inf_q\left\{\sum_{j=1}^h p_jS_\theta(q_j)+V_D(q)\right\}, \tag{9}\] where all eligible finite grids and trials are allowed. Refinements always use a common grid for all blocks. The diagonal increment \(I-q_{j,k-1}\) is integrated independently on every visit in the cascade interpretation, including repeated visits to the same leaf.

Explicit evaluation of the angular pressure

Return to \(p=\sum_jp_j<1\) and positive radii \(a_j\). Generate the Haar flag from consecutive column blocks \(G_j\) of a matrix with independent \(N(0,1/M)\) entries, as proved in Section 8. For outer coefficients \(x_j\in\mathbb R^{n_j}\) and inner coefficients \(z_{i,j}\in\mathbb R^{n_j}\), \(j\le i<h\), put \[y=\sum_jG_jx_j,\qquad u_i=\sum_{j\le i}G_jz_{i,j}.\] Thus \(y\) is the physical field and \(u_i\) is a witness in the flag subspace \(E_1\oplus\cdots\oplus E_i\). Requiring \(\lVert y\rVert^2/M=\sum_ja_j\) and \[\lVert y-u_i\rVert^2/M\le\sum_{j>i}a_j\] forces the squared norm in each tail of the orthogonal decomposition of \(y\) to be no larger than the prescribed tail. Equivalently, the initial blocks carry at least their prescribed total mass. After sorting \(a_j/p_j\), Lemma 6 shows that these one-sided tests recover the desired weighted shell. The witnesses are integrated as the inner variables of (5); sending \(\theta\) to zero removes their entropy cost. Section 8 proves the construction and these limits.

Permute the pairs \((p_j,a_j)\) so that \(a_j/p_j\) is nonincreasing; ties may be ordered arbitrarily. Set \[ s_a=\sum_j a_j,\qquad H_p(a)=\frac12\sum_jp_j\log\frac{2\pi e a_j}{p_j},\qquad J_p=\frac12\int_0^p\log(1-u)\,du. \tag{10}\] Here \(H_p\) is the limiting entropy per output coordinate of the product of Lebesgue shells, and \(J_p\) is the limiting logarithmic volume Jacobian of the Gaussian span map. Orthogonalizing the coefficient columns \(x_j,z_{j,j},\ldots,z_{h-1,j}\) after division by \(\sqrt M\) gives \(d_j=h-j+1\) columns in block \(j\). Choose a coefficient radius \(\kappa>0\), penalty strength \(\lambda>0\), and truncation threshold \(B>0\). In each block choose an upper triangular \(d_j\times d_j\) matrix \(R_j\) with positive diagonal and every column norm at most \(\kappa\). For one row \(\xi_j\in\mathbb R^{d_j}\), label the entries of \(\xi_j^{\mathsf T}R_j\) by \(x,j\) and \(z_i,j\), \(j\le i<h\). Write \[y=\sum_j (\xi_j^{\mathsf T}R_j)_{x,j}, \qquad u_i=\sum_{j\le i}(\xi_j^{\mathsf T}R_j)_{z_i,j}.\] The symbols \(y,u_i\) in these row formulas denote single coordinates of the fields above. Let \(g\) be the vector of row functions and \(c\) the corresponding constants in \[ (y^2,s_a),\qquad(-\min\{y^2,B\},-s_a),\qquad \left((y-u_i)^2,\sum_{j>i}a_j\right),\quad 1\le i<h. \tag{11}\] The soft penalties below enforce, in the limit, the row-average inequalities \[M^{-1}\sum_r g_\ell(y_r,u_{1,r},\ldots,u_{h-1,r})\le c_\ell.\] Define \[\begin{align*} \mathcal B_{p,a}(\kappa,\lambda,B,\theta) =\sup_{R_1,\ldots,R_h}\Bigg[& \sum_jp_j\left\{ \log\!\left(\kappa\sqrt{\frac{2\pi e}{p_j}}\right) +\log\frac{R_{j,11}}\kappa +\theta\sum_{\ell=2}^{d_j}\log\frac{R_{j,\ell\ell}}\kappa \right\}\\ &+\theta\min_{0\le t_\ell\le\lambda/\theta} \left\{t\cdot c+\mathcal G_{\theta,p}\big(f(y)/\theta-t\cdot g\big)\right\}\Bigg]. \tag{12}\end{align*}\] The functions \(y,u_i,g\) in the second line depend on the chosen triangular matrices. The terminal is bounded above and has at most quadratic negative growth. The explicit angular formula is \[ P_p^f(a)=J_p-H_p(a)+ \lim_{\lambda\to\infty}\lim_{B\to\infty}\lim_{\theta\downarrow0} \mathcal B_{p,a}(\kappa,\lambda,B,\theta). \tag{13}\] Theorem 1 asserts that these successive limits exist and give the same angular pressure for every sufficiently large fixed \(\kappa\). Values at zero coordinates of \(a\) are obtained by continuity. When \(\sum_jp_j=1\), the evaluation is \[ P_p^f(a)=\lim_{\eta\downarrow0}P_{(1-\eta)p}^f(a), \tag{14}\] with \(a\) fixed. No analytic continuation in \(\theta\) is used: \(\theta\) is a positive exponent throughout.

The Gaussian theorem and its probabilistic tools

We use the Stiefel model and powered integral (5) from Section 2.2, with fixed \(0<\theta<1\), fixed column counts, and \(n_j/M\to p_j>0\). The Gaussian formula holds without a restriction on \(\sum_jp_j\); its soft-penalty extension will supply the one-sided angular tests.

Theorem 3 (Gaussian formula and global penalties). For fixed \(0<\theta<1\), fixed column counts \(d_j\), and \(n_j/M\to p_j>0\), let \(D\) be bounded continuous, or the sum of a bounded continuous function and a nonpositive quadratic form. Then \[ \frac1M\log\mathcal Z_{M,\theta}\left(\sum_{r=1}^M D(\xi_r)\right) \longrightarrow\mathcal G_{\theta,p}(D) \tag{15}\] in probability and in expectation. More generally, let each \(g_\ell\) be bounded continuous or a nonnegative quadratic form, let \(\bar g=M^{-1}\sum_r g(\xi_r)\), and fix \(d>0\) and \(c\in\mathbb R^s\). The pressure for \[H=\sum_rD(\xi_r)-Md\sum_{\ell=1}^s(\bar g_\ell-c_\ell)_+\] converges in probability to \[ \min_{t\in[0,d]^s}\left\{t\cdot c+\mathcal G_{\theta,p}(D-t\cdot g)\right\}. \tag{16}\] No restriction on \(\sum_jp_j\) is needed.

The proof occupies Sections 3–7. We begin with \(D\) smooth and compactly supported. In this case every derivative used below is bounded.

Marked cascades, zero restrictions, and visit variables

We use Ruelle’s probability cascades [27] and the marked-cascade framework of Panchenko and Talagrand [26]. Their sampling and cavity structure were developed by Bolthausen and Sznitman [4]. The direct calculation below permits nonnegative marks with a positive finite fractional moment; its separate logarithmic estimates cover the restrictions used later. A finite cascade on (6) has Poisson branching exponents \(\zeta_1,\ldots,\zeta_{k-1}\), all strictly below one. At a vertex with exponent \(z\), take independent Poisson points of intensity \(z x^{-1-z}\,dx\). Products along paths are normalized at the leaves. Common-level probabilities for two sampled leaves are \(w_i\), \(0\le i<k\). The step at \(\zeta_k=1\) is an ordinary integral over a fresh visit variable, never a Poisson branching level. Root Gaussian data have coefficient zero and are averaged at the end.

We use the following precise change-of-measure rule. If a positive edge weight \(x\) is multiplied by a nonnegative mark \(F\) satisfying \[ 0<\mathbb E(F^z\mid\text{parent data})<\infty, \tag{17}\] the image of its positive points is again Poisson, with intensity multiplied by that conditional moment, and its marks have density proportional to \(F^z\). This follows by changing variables in the marked Poisson intensity. Points for which \(F=0\) are deleted. Applying this rule backwards with \(F=\exp(U_i-U_{i-1})\) proves both the log-exponential recursion and the tilted child kernel \[ \exp\{\zeta_i(U_i-U_{i-1})\}\,d\mu_i. \tag{18}\] Conditional on root data, the transformed unmarked cascade has its original law and is independent of the path marks with these kernels. Restrictions by an indicator are therefore permitted whenever surviving fractional moments are positive and finite. A subtree with zero normalizer is deleted before the next backward step; the condition is checked at surviving parents, and the final root normalizer must be positive. Restrictions require separate logarithmic integrability estimates; these will be supplied where they are used.

The independence assertion follows through the same backward induction. At each child, include its already transformed unmarked descendant tree in the mark. By induction that tree has its ordinary law independently of the path marks. The parent multiplier depends only on the child and ancestor marks, so its Poisson image intensity factors into the ordinary edge intensity, the tilted child kernel, and the unchanged descendant-tree law. For the normalized multiplier \(F=\exp(U_i-U_{i-1})\), its conditional \(\zeta_i\)-moment is one, including after deletion of zero marks. Thus no ancestor-dependent scale remains, and the transformed unmarked tree has its ordinary law independently of the original root data.

For completeness, the Poisson Laplace formula says that a sum of \(z\)-Poisson points multiplied by independent positive factors with finite \(z\)th moment is a scaled positive stable variable of index \(z\). Such a variable has finite logarithmic second moment and finite positive moments of every order below \(z\). These facts follow by integrating its Laplace transform \(e^{-ct^z}\); its lower tail is bounded using \(\mathbb P(S\le u)\le e^{tu}\mathbb Ee^{-tS}\) and optimizing \(t\). Successive exponents are strictly increasing, so these fractional moments justify all backwards normalizations. For these recursively normalized multipliers, the original and transformed leaf totals have the same law conditional on the root data. Their logarithms are integrable by the preceding stable-moment estimates, so the conditional mean of their difference is zero. Telescoping the multipliers gives the expected logarithm identity.

If the original marks have independent components and their terminal energy is additive, the backward values and tilted kernels factor over those components. The preceding induction shows that conditioning on a finite sampled shape of the transformed unmarked tree preserves this factorization. Averaging the original independent root data therefore leaves the canonical spatial rows independent, with their common averaged row law. This statement concerns the transformed sampled shape; conditioning on the original numerical labels or on realized root fields would be a different assertion.

Merged masses at a branching exponent \(z\) have the law of normalized stable jumps at that exponent. Indeed, after summing descendants, mark the points at that level by the descendant totals; the Poisson mapping and superposition give the assertion. We will use this at the fixed exponent \(\theta\), independently of how finely the other levels are subdivided.

Sampling first-column marks at exponent \(\theta\) and conditional inner variables at each visit implements (5). For the cascade consisting only of the split at \(\theta\), the expected cascade logarithm, conditional on the disorder, is exactly \[\frac1\theta\log\int\left(\int e^H d\mu_{\rm in}\right)^\theta d\mu_{\rm out}.\] Adding redundant cascade levels leaves this representation unchanged when the Hamiltonian has no additional label dependence.

Joint overlap identities and ordered matrix paths

Matrix synchronization is a central feature of vector-spin theory [24]. We follow the joint scalarization and label-overlap formulation of [18], and establish the version needed for the present shared-coordinate construction. For replicas \(a,b\), write \[\mathcal R_j^{ab}=(X_j^a)^{\mathsf T}X_j^b/n_j.\] The array indexed jointly by replica and column is Gram, and the self blocks are fixed. Include also the common-level label overlap. Gaussian perturbations below impose the Ghirlanda–Guerra identities (GG) [8]: for every bounded test \(F\) of the complete arrays on the first \(r\) replicas and every continuous test \(\psi\) of the generated scalarization-and-label vector on one edge, \[ \mathbb E\langle F\psi(T^{1,r+1})\rangle =\frac1r\mathbb E\langle F\rangle\,\mathbb E\langle\psi(T^{12})\rangle +\frac1r\sum_{a=2}^r\mathbb E\langle F\psi(T^{1a})\rangle. \tag{19}\] Here brackets denote Gibbs expectation and \(\mathbb E\) averages all randomness. Initially \(T^{ab}\) denotes the scalarization-and-label vector specified next; after symmetry is proved it determines the full matrix-and-label edge value.

The perturbing scalarizations are \[ v^{\mathsf T}(\mathcal R_j^{ab})^{\circ \ell}v, \qquad \ell=1,2, \tag{20}\] for a countable dense set of vectors \(v\), together with the numerical label overlap. Each is a Gram kernel: use tensor products of the column vectors for the entrywise powers. Products, nonnegative integral powers, and positive sums remain Gram kernels. After individual rescaling their absolute values are at most one and their diagonals are constant. Monomial GG identities give all continuous tests of this generated vector by polynomial approximation on its compact range. At this stage they do not yet identify a nonsymmetric matrix from its scalarizations; the test \(F\) may nevertheless depend on the complete finite matrix arrays.

We recall why the resulting paths have the structure used in (9). A scalar symmetric weakly exchangeable positive semidefinite array with deterministic bounded diagonal and full off-diagonal GG is ultrametric. To check the exact applicability of the established result, normalize a positive diagonal to one (a zero diagonal gives the zero array). The Dovbysh–Sudakov representation [22] writes it as \[T^{ab}=h^a\cdot h^b+a^a\mathbf 1_{\{a=b\}},\qquad a^a\ge0,\] with conditionally iid pairs and \(\lVert h^a\rVert\le1\). The sphere-support theorem [21], whose hypotheses use only off-diagonal GG tests, puts the directing Hilbert-space measure on the sphere of squared radius \(q_*=\sup\operatorname{supp}\mathcal L(T^{12})\). The represented inner-product array therefore has a deterministic diagonal. Its full-overlap GG identities follow from the off-diagonal identities, and [23] applies.

The identities for continuous edge tests extend to bounded Borel tests, since they identify finite Borel measures on the compact edge space. We also need a simple all-pairs extension consequence of GG. On an event requiring every old edge to lie in a fixed marginal bin of mass \(b>0\), GG gives conditional probability \((1-b)/r\) that a new edge from each fixed old vertex misses the bin. A union bound leaves extension probability at least \(b\). Hence positive-mass bins support arbitrarily large all-pairs samples. A bin uniformly below zero is impossible for a Gram array, by testing its large all-ones quadratic form. Thus scalar overlaps are nonnegative. These facts apply to every scalarization, every positive sum of scalarizations, and the label overlap.

For matrix symmetry, we use Panchenko’s argument from the proof of [25], in the scalarized form recalled in [18]. The first- and second-power scalarizations determine each transposed pair of entries up to interchange. If an antisymmetric part persisted with positive probability, choose a small positive-mass joint bin fixing these unordered pairs while keeping a fixed asymmetry gap. The all-pairs extension just used produces arbitrarily large samples in this bin. Finite color thinning gives arbitrarily many ordered vertices whose forward cross matrices choose the same ordering of every pair, to the prescribed tolerance. This elementary thinning follows by successively keeping a color class of the remaining vertices and then taking a common outgoing color. For a homogeneous set of \(2m\) replicas, average each normalized column over its first and second halves, writing these averages as \(a_c,b_c\) for column \(c\). If the bin width is \(\varepsilon\) and \(t_c\) is its same-column overlap center, the fixed self overlaps and the cross overlaps give \[\|a_c-b_c\|^2=\frac{2(1-t_c)}m+O(\varepsilon).\] The averages have norm at most one. Thus \(|a_c\cdot b_d-b_c\cdot a_d|=O(\sqrt\varepsilon)+O(m^{-1/2})\), whereas the homogeneous forward matrices make this difference approach their fixed antisymmetric entry gap. Taking arbitrarily large \(m\) and then arbitrarily small bins contradicts that gap. Thus \(\mathcal R_j^{12}\) is symmetric in the limit. Dense first-power quadratic forms now determine each matrix, so the identities extend to all joint matrix-and-label edge tests.

Finally, the first-power scalarizations and label are comonotone on an edge, by the ultrametric-sum argument in [16]. Otherwise joint GG gives crossed values on edges \(12\) and \(13\). Ultrametricity of each of two crossed coordinates forces edge \(23\) to have the lower value in each coordinate, contradicting ultrametricity of their sum. A positive weighted sum of a countable separating set orders all coordinates simultaneously. Dense quadratic forms therefore give deterministic matrix quantiles \(q_j(u)\) nondecreasing in Loewner order. Gram positivity and self overlap \(I\) imply \(0\preceq q_j\preceq I\). The label coordinate retains its prescribed atom masses, so its sharing threshold lies at quantile \(\theta\). When the first column is shared, \(\mathcal R_j^{ab}e_1=e_1\); hence the matrix quantiles satisfy the pinning condition at that same threshold.

Every finite monotone rounding of this joint order has the cascade law. At any threshold, ultrametricity makes the above-threshold relation an equivalence relation; GG specifies the probability that the next sample joins each existing class. These nested joining probabilities determine the rounded array. To verify that a cascade has them, at a parent of exponent \(z_0\) its child masses of exponent \(z\) are stable jumps tilted by their total to power \(z_0\). The Gamma integral and the Poisson factorial formula show that an existing child with \(n_a\) of \(n\) visits is selected next with probability \((n_a-z)/(n-z_0)\). Multiplying conditional probabilities up the tree gives the GG joining rule. Thus cascade approximations preserve label atoms and matrix pinning.

Fresh fields and the one-pattern variation

Conditionally on finitely many base replicas, independent fresh Gaussian fields have exactly the covariance matrices given by their overlaps. Finite-array convergence therefore passes bounded continuous field tests. Bounded weights bounded away from zero also permit logs and reciprocal partition functions: approximate these uniformly by polynomials on a compact interval, and represent powers by additional replicas. This proves passage of bounded reweightings to finite cascade approximations. Independent residual visit noises are essential when self variance exceeds the largest shared covariance.

For smooth \(D\) the single-pattern functional is Lipschitz in the sum of the \(L^1\) path distances. More precisely, along a step variation, \[ \frac{d}{dt}V_D(q+t\dot q)\Big|_{t=0} =-\frac12\sum_{j,i}w_i\operatorname{tr}\left(\dot q_{j,i}\mathbb E[M_{j,i}M_{j,i}^{\mathsf T}]\right), \tag{21}\] where \(M_{j,i}\) is the conditional mean of the terminal \(j\)-gradient under the successive tilted kernels. Gaussian covariance differentiation inserts half the Hessian and half the step coefficient times the gradient square. Differentiating all increments and propagating Hessians makes these terms telescope; consecutive coefficients differ by \(w_i\), giving (21). The conditional gradients form a martingale, so their second-moment matrices \(G_{j,i}=\mathbb E[M_{j,i}M_{j,i}^{\mathsf T}]\) are nondecreasing and uniformly bounded. To see the resulting order explicitly, put \(T_{j,i}=\sum_{r=i}^{k-1}w_r(\widetilde q_{j,r}-q_{j,r})\) and \(T_{j,k}=0\) on a common grid. Summation by parts gives \[\sum_iw_i\operatorname{tr}((\widetilde q_{j,i}-q_{j,i})G_{j,i}) =\operatorname{tr}(T_{j,0}G_{j,0})+ \sum_{i=1}^{k-1}\operatorname{tr}(T_{j,i}(G_{j,i}-G_{j,i-1})).\] Every term is nonnegative when the tail differences are positive semidefinite. Apply this along the segment between the two paths in (21) to obtain \[ \int_t^1\widetilde q_j(u)\,du\succeq\int_t^1q_j(u)\,du\quad\hbox{for every }j,t \quad\Longrightarrow\quad V_D(\widetilde q)\le V_D(q). \tag{22}\] Approximation by step averages extends these statements to bounded monotone paths. By the fresh-field argument this extended \(V_D\) is the limiting expected log increment of one independent pattern at a GG limit.

The Gaussian canonical calculation

The scalar spherical cascade transform is computed in [17]; matrix multipliers and susceptibility determinants also occur in the vector spherical Crisanti–Sommers formula [13]. Here the active spin space changes at the sharing level. The calculation below derives the entropy with that dimension change and the strict Gaussian precision conditions it requires. We work first in one block, suppressing its index. On a fixed grid let \(0\preceq b_0\preceq\cdots\preceq b_{k-1}\) be linear-field covariance levels. These are covariances of fields coupled to spin coordinates, not the row-field covariances \(q\). Each spatial coordinate has an independent field copy. The first spin column is sampled at exponent \(\theta\) and shared thereafter; the remaining columns are sampled independently at each visit.

Let \(l_n(b)\) be the expected cascade log integral per spatial coordinate for this linear field on \(\operatorname{St}(n,d)\), minus \(\operatorname{tr}b_{k-1}/2\). We shall prove \[ \limsup_{n\to\infty}l_n(b)\le S_\theta(q)-\frac12\sum_iw_i\operatorname{tr}(q_i b_i) \tag{23}\] for every eligible \(q\), locally uniformly in the field levels.

A canonical multiplier

Replace each spatial spin row by a standard Gaussian vector \(s\in\mathbb R^d\), with the same sharing rule. Write \(m\) for the index satisfying \(\zeta_m=\theta\). Take independent field increments \[z_0\sim N(0,b_0),\qquad z_i\sim N(0,b_i-b_{i-1})\quad(1\le i<k).\] The first spin coordinate \(s_1\sim N(0,1)\) is sampled at level \(m\); the remaining coordinates \(s_F\sim N(0,I_{d-1})\) are sampled at the terminal visit. All these reference variables are independent.

Here is the one-row recursion defining the canonical integral. For a symmetric multiplier \(C\), start with \[U_k(x,s)=x^{\mathsf T}s+\tfrac12s^{\mathsf T}Cs, \qquad U_{k-1}(x,s_1)=\log\mathbb E_{s_F}e^{U_k(x,s)}.\] For \(i=k-1,\ldots,1\), replace \(U_i\) by \[U_{i-1}(x)=\frac1{\zeta_i}\log \mathbb E_{z_i,\,s_1\text{ if }i=m} \exp\{\zeta_i U_i(x+z_i)\}.\] The argument \(s_1\) is retained until the step \(i=m\) and integrated at that step; its integration and the \(z_m\) integration have the same exponent and may be performed in either order. Finally put \(\Psi(C,b)=\mathbb E_{z_0}U_0(z_0)\). This extended-real function is finite exactly when its successive Gaussian precisions are strictly positive. A zero precision is insufficient: the corresponding Gaussian exponential integral is infinite.

Lemma 4 (Canonical duality). There is a minimizing \(C=C(b)\) for \(\Psi(C,b)-\operatorname{tr}C/2\). It lies in the strict finite domain, its tilted spin satisfies \(\mathbb Ess^{\mathsf T}=I\), and the cross covariances \(q^*\) of two tilted visits form an eligible path. Moreover \[ \begin{split} J(b)&:=\Psi(C(b),b)-\tfrac12\operatorname{tr}(C(b)+b_{k-1})\\ &=S_\theta(q^*)-\frac12\sum_iw_i\operatorname{tr}(q_i^*b_i) =\inf_q\left\{S_\theta(q)-\frac12\sum_iw_i\operatorname{tr}(q_i b_i)\right\}. \end{split} \tag{24}\] For fixed \(\theta,d\) and bounded \(b_{k-1}\), all such multipliers are bounded uniformly over the grids. The tilted terminal field has bounded second and fourth moments under the same uniformity.

Proof. We first construct the multiplier and its covariance path. We then use the inverse identity across the sharing level and the entropy gradient to identify that path as the variational minimizer. Comparing derivatives in \(b\) identifies the values; the final moment estimate supplies uniformity for the cavity calculation.

The recursion is lower semicontinuous, with value \(+\infty\) outside its finite domain. This follows either from the Gaussian formulas or from Fatou at successive integrations, using Jensen lower bounds at the root. There is also uniform coercivity. For any covariance \(Q\) with \(\frac12I\preceq Q\preceq\frac32I\), use the centered Gaussian spin law \(N(0,Q)\) in the variational inequality at the spin integrations, leaving the field kernels unchanged. Writing \(\gamma_r=N(0,I_r)\) and \(Q_1,Q_{F\mid1}\) for the marginal and conditional laws of this trial, their costs satisfy \[\theta^{-1}\operatorname{Ent}(Q_1\mid\gamma_1) +\mathbb E_{Q_1}\operatorname{Ent}(Q_{F\mid1}\mid\gamma_{d-1}) \le\theta^{-1}\operatorname{Ent}(N(0,Q)\mid\gamma_d).\] This is the chain rule for relative entropy, and allows correlated trial coordinates. Hence \[\Psi(C,b)-\tfrac12\operatorname{tr}C \ge\tfrac12\operatorname{tr}(C(Q-I))- \frac1{2\theta}(\operatorname{tr}Q-\log\det Q-d).\] Choosing \(Q=I+\frac12\operatorname{sign}(C)\) gives \[ \Psi(C,b)-\tfrac12\operatorname{tr}C\ge\tfrac14\lVert C\rVert_*-c_d/\theta, \tag{25}\] where \(\lVert \cdot\rVert_*\) is the trace norm. If \(\lVert b_{k-1}\rVert\le B_0\), the comparison \(C=-(B_0+1)I\) has \(\Psi(C,b)\le0\): nested exponents at most one and Jensen bound it by the ordinary Gaussian exponential integral. Thus minimizers lie in a compact set depending only on \(d,\theta,B_0\). Lower semicontinuity gives existence, and finiteness forces strict precision inequalities. Differentiating in \(C\) then gives \(\mathbb Ess^{\mathsf T}=I\).

The tilted kernels are Gaussian with linear conditional means and deterministic conditional covariances. Under their joint path law, for \(0\le i<k\) let \(\mathcal H_i\) be generated by \(z_0,\ldots,z_i\) and also \(s_1\) when \(i\ge m\); set \(\mathcal H_k=\sigma(\mathcal H_{k-1},s_F)\). Define \(M_i=\mathbb E[s\mid\mathcal H_i]\) and let \(\mathsf K_i^*\) be the Hessian of \(U_i\) with respect to its field argument. The frozen components of \(M_i\) are the already sampled spin coordinates, while the Hessian is supported on \(F_i\). Differentiating one log-exponential integration gives \(\mathsf K_{i-1}^*=\mathsf K_i^*+\zeta_i\operatorname{Cov}(M_i\mid\mathcal H_{i-1})\); these conditional covariance matrices are deterministic. Averaging and using the martingale property gives \[ q_i^*=\mathbb EM_iM_i^{\mathsf T},\qquad \mathsf K_k^*=0,\qquad \mathsf K_{i-1}^*=\mathsf K_i^*+\zeta_i(q_i^*-q_{i-1}^*),\qquad q_0^*=\mathsf K_0^*b_0\mathsf K_0^*. \tag{26}\] The self moment is \(q_k^*=I\). Unsampled coordinates retain nondegenerate conditional variation, so \(\mathsf K_i^*\) is strictly positive on the active space. Shared coordinates have the same cross and self entries, giving pinning. Thus (26) identifies \(\mathsf K_i^*\) with (7).

The inverse identity needed below is \[ (\mathsf P_i-\mathsf P_{i-1})|_{F_i} =\zeta_i(b_i-b_{i-1})|_{F_i},\qquad 1\le i<k. \tag{27}\] At a step without freezing this is the ordinary quadratic Gaussian identity. At a freezing step it involves the compression of the inverse, and we spell it out without a commutativity assumption. Decompose the previous active space into the newly frozen space \(E\) and the remaining space \(F\). Holding the new spin \(s_E\) fixed, the backwards quadratic value has the form \[s_E^{\mathsf T}x_E+\tfrac12x_F^{\mathsf T}H x_F +(B_1s_E+a)^{\mathsf T}x_F+c(s_E),\qquad H\succ0.\] Write the field increment as \(\Delta b=\left(\begin{smallmatrix}U&V\\V^{\mathsf T}&W\end{smallmatrix}\right)\) and its exponent as \(t\). Integrating this increment at fixed \(s_E\) changes the active Hessian and the coefficient of \(s_E\) to \[\Gamma=(H^{-1}-tW)^{-1},\qquad D_1=(I-tHW)^{-1}(B_1+tHV^{\mathsf T}).\] The strict finite canonical domain guarantees \(H^{-1}-tW\succ0\). The remaining affine vector changes similarly. Now integrate \(s_E\); its tilted precision is also strictly positive in that domain, so its covariance \(T\) is positive definite. The full Hessian before freezing is \[ \begin{pmatrix} tT&tTD_1^{\mathsf T}\\ tD_1T&\Gamma+tD_1TD_1^{\mathsf T} \end{pmatrix}. \tag{28}\] Its Schur complement after eliminating \(E\) is exactly \(\Gamma\). Consequently the \(F\)-compression of its inverse is \(\Gamma^{-1}=H^{-1}-tW\). Subtracting it from the post-freezing inverse \(H^{-1}\) proves (27). The mixed covariance \(V\) affects \(D_1\) and cancels through the actual Schur complement.

We next verify convexity of the entropy. For a line variation \(e_i=\dot q_i\) preserving pinning, put \[A_i=-\dot{\mathsf K}_i=\zeta_{i+1}e_i+\sum_{r>i}^{k-1}w_r e_r, \qquad \lVert A\rVert_P^2=\operatorname{tr}(PAPA).\] Matrix differentiation of (8) gives the active gradient \[ \frac2{w_i}\nabla_{q_i}S_\theta =\left(\mathsf P_0q_0\mathsf P_0+ \sum_{r=1}^i\frac{\mathsf P_r-\mathsf P_{r-1}}{\zeta_r}\right)\Big|_{F_i}. \tag{29}\] The second derivative of \(2S_\theta\) equals \[\begin{align*} &2\operatorname{tr}(\mathsf P_0e_0\mathsf P_0A_0) +2\operatorname{tr}(q_0\mathsf P_0A_0\mathsf P_0A_0\mathsf P_0) -\zeta_1^{-1}\lVert A_0\rVert_{\mathsf P_0}^2\\ &\hspace{1cm}+\sum_{i=1}^{k-1}(\zeta_i^{-1}-\zeta_{i+1}^{-1}) \lVert A_i\rVert_{\mathsf P_i}^2. \tag{30}\end{align*}\] The second term is nonnegative. Since \(\mathsf K_i\preceq\mathsf K_0\), Schur complementation implies \(\mathsf P_i\succeq\mathsf P_0|_{F_i}\). For symmetric \(A\) the corresponding squared norm increases with the positive precision (expand the difference with a positive increment). Replace every positive term’s precision by \(\mathsf P_0\). Extend each \(e_i\) by zero off \(F_i\). The remaining terms satisfy the exact identity \[2\operatorname{tr}(\mathsf P_0e_0\mathsf P_0A_0) -\zeta_1^{-1}\|A_0\|_{\mathsf P_0}^2 +\sum_{i=1}^{k-1}(\zeta_i^{-1}-\zeta_{i+1}^{-1}) \|A_i\|_{\mathsf P_0}^2 =\sum_iw_i\|e_i\|_{\mathsf P_0}^2.\] Indeed, \(A_{i-1}=A_i+\zeta_i(e_{i-1}-e_i)\), so expansion cancels the cross terms successively. Thus \(S_\theta\) is strictly convex. Its value and active gradient vanish at \(q^0\), so \(S_\theta\ge0\).

At the canonical path, (26) gives \(\mathsf P_0q_0^*\mathsf P_0=b_0\). Compress (27) further to each nested \(F_i\) and telescope in (29). The scaled gradient in (29) becomes \(b_i|_{F_i}\). Hence \(q^*\) is the unique minimizer of the rightmost expression in (24).

To identify its value, vary the fields on the segment from zero to \(b\). Covariance differentiation of \(\Psi\) before subtracting the diagonal shift inserts \(\frac12(\mathsf K_i^*+\zeta_i\mathbb EM_iM_i^{\mathsf T})\) for an increment at level \(i\). Telescoping and \(\mathbb Ess^{\mathsf T}=I\) give \[\dot J=-\tfrac12\sum_iw_i\operatorname{tr}(q_i^*\dot b_i).\] The entropy infimum has the same derivative by its unique minimizer. Envelope differentiation is justified by the compact multiplier bound: limit optimizers remain in the strict finite domain, where the Gaussian formulas and their directional derivatives are continuous, including degenerate field increments. Both values are zero at \(b=0\); the minimizer then is \(C=0\) and \(q=q^0\). This proves (24).

Finally the field moments are uniform in the grid. Write \(z=\sum_{\ell=0}^{k-1}z_\ell\), \(\Delta_0=b_0\), \(\Delta_\ell=b_\ell-b_{\ell-1}\) for \(\ell\ge1\), and let \(\rho\) be the full tilted path density relative to the independent Gaussian reference variables. Differentiating the product of the kernels (18) makes all intermediate terms telescope: \[L_\ell:=\nabla_{z_\ell}\log\rho =s-\sum_{i=\ell}^{k-1}w_iM_i, \qquad 0\le\ell<k.\] Since \(\mathbb Ess^{\mathsf T}=I\) and \(M_i\) are its conditional means, \(\|v^{\mathsf T}L_\ell\|_{L^2}\le2|v|\) for every deterministic \(v\). Gaussian integration by parts in the increments therefore gives, with \(X=\mathbb E|z|^2\) and \(B=\operatorname{tr}b_{k-1}\), \[X=B+\sum_\ell\mathbb E[z^{\mathsf T}\Delta_\ell L_\ell] \le B+2B\sqrt X.\] This bounds \(X\) independently of the grid. The full tilted path law is centered Gaussian in the strict precision domain, so \(\mathbb E|z|^4\le3X^2\) is bounded as well. ◻

From canonical rows to Stiefel columns

Use \(n\) independent canonical spatial rows. Their unrestricted expected log integral is \(n\Psi(C,b)\), since the recursion factors over rows. Decompose the Gaussian spin columns by Gram–Schmidt, drawing the first column at its prescribed vertex and the remaining columns at visits. Restrict all normalized triangular lengths and cross readings to a fixed small neighborhood of the identity. The first restriction and each conditional inner restriction have probability tending to one, independently of the preceding frame; their recursion normalization costs are \(o(n)\).

On the restriction the directions have the exact Stiefel sharing law. Replacing the quadratic energy by \(n\operatorname{tr}C/2\) costs \(o_\epsilon(n)\). Replacing the linear-field covariance by that of exact Stiefel columns also costs \(o_\epsilon(n)\): its normalized covariance discrepancy is uniformly small, and Gaussian interpolation bounds the expected log change by half the diagonal-minus-pair covariance difference. Let \(n\to\infty\) and then shrink the neighborhood. Restriction gives a lower bound on the canonical integral, so \[\limsup_n l_n(b)\le\Psi(C,b)-\tfrac12\operatorname{tr}(C+b_{k-1})=J(b).\] Lemma 4 proves (23). The same covariance comparison and compact multiplier bound give local uniformity in \(b\). Gaussian differentiation for random reference measures can be justified throughout by finite conditional iid discretizations, followed by truncation: at fixed system size bounded covariances give finite moments of exponential integrals and of Jensen lower bounds for their inverses.

The Gaussian upper bound

The comparison enriches the pressure by hierarchical linear fields and tests an affine bound at a deterministic contact point. We adapt the supersolution method of [16] and its matrix-valued extension [18]. The changes required here concern Poisson differentiation in the pattern count, the changing active field spaces, and the shared-coordinate entropy bound. All contact identities for this pattern Hamiltonian are proved below. Fix an eligible trial \(q\) and its label cascade. Use \(\operatorname{Pois}(Mr)\) patterns, \(0\le r\le1\), and independent linear fields in block \(j\) with covariance levels \(b_{j,i}\). Include the shift \(-\frac12\sum_jn_j\operatorname{tr}b_{j,k-1}\). The increment \(\Delta b_{j,i}\) is required to be positive semidefinite and supported on \(F_{j,i}\); field levels are cumulative sums of these increments. A fixed cap will bound the sum of their traces.

To impose joint GG, enumerate the normalized monomial kernels described in (20), including label factors, and add \[ \sqrt M\,e_M\sum_{a\le M}c_a u_aY_a, \qquad e_M=M^{-1/16},\quad 1\le u_a\le2. \tag{31}\] The fields \(Y_a\) are independent with those kernels as covariances. Choose \(c_a>0\) decreasing rapidly enough that their sums, and the sums of squared coefficients times kernel Lipschitz constants, are finite. Tensor Gaussian fields in the spin columns realize these kernels, with label increments on the tree; kernels without a label factor are common root data. Interpolation bounds the expected pressure change by \(O(e_M^2)\). Under reweighting, common-level probabilities remain \(w_i\).

Let \(P_M(r,b,u)\) be the random cascade log integral divided by \(M\), including the shifts and perturbation. Uniformly for bounded fields, \[ \mathbb E(P_M-\mathbb EP_M)^2=O(M^{-1}). \tag{32}\] Indeed replacing a pattern changes the root log by at most \(2\lVert D\rVert_\infty\), and adding one by at most \(\lVert D\rVert_\infty\). Independent replacements and the Poisson variance bound handle patterns and their number. Common Gaussian field and perturbation data have log-Lipschitz constant \(O(\sqrt M)\) in standard coordinates, since spin norms are fixed. Recursion preserves this bound. Gaussian concentration handles that part. The transformed-versus-original log-total noise has bounded second moment for this fixed grid, as in Section 3.

Suppose the desired upper bound failed for this trial by a positive gap along a subsequence. For a small fixed \(\delta>0\), minimize \[\begin{align*} \mathcal A_M(r,b,u)={}&-\mathbb EP_M(r,b,u) -\frac12\sum_j\frac{n_j}{M}\sum_iw_i\operatorname{tr}(q_{j,i}b_{j,i}) +\sum_j\frac{n_j}{M}S_\theta(q_j)\\ &+r(V_D(q)+\delta)+\sum_{a\le M}c_a(u_a-3/2)^2 \tag{33}\end{align*}\] over these domains. At \(r=1,b=0,u_a=3/2\), failure makes the value strictly negative for large \(M\); Poissonization has \(o(1)\) cost. At \(r=0\), (23) and block independence give a lower bound \(-o(1)\), locally uniformly in \(b\).

The field cap can be chosen not to bind, before using any overlap identities. For fixed \(0<\eta<1\), the trial \(q'_i=q_i+\eta(I-q_i)\) is eligible. The weighted difference against cumulative fields is \[\sum_iw_i\operatorname{tr}((q'_i-q_i)b_i) =\eta\sum_r\operatorname{tr}\left(\Delta b_r\sum_{i\ge r}w_i(I-q_i)\right) \ge c_q\sum_r\operatorname{tr}\Delta b_r,\] where \[c_q=\eta\min_{r:F_r\ne\{0\}} \{w_r\lambda_{\min}((I-q_r)|_{F_r})\}>0.\] Strict positivity holds because \(I-q_r\succeq\mathsf K_r\succ0\) on \(F_r\). Apply (23) with \(q'\); bounded pattern energy and the \(O(e_M^2)\) perturbation cost make [eq:contact] positive at a sufficiently large fixed trace cap. Thus a negative minimizing point has \(r>0\) and permits small increases of every allowed field increment.

GG at the deterministic contact point

The quadratic penalty in the perturbation parameters implements the contact-point selection used in [16] and [18]. It yields GG at the parameters selected by the comparison, rather than only after averaging over those parameters. The minimizing parameters are deterministic, being selected from an expected pressure. Comparing with all \(u_a=3/2\) shows that their penalty is \(O(e_M^2)\). Every fixed \(u_a\) is consequently interior for large \(M\). Hold all other minimizing coordinates fixed. Convexity in \(u_a\) and minimality give the explicit local remainder bound \[ 0\le\mathbb EP_M(u+s e_a)-\mathbb EP_M(u)-s\partial_{u_a}\mathbb EP_M(u) \le c_a s^2. \tag{34}\] This estimate precedes and does not assume GG.

Differentiation gives \[\partial_{u_a}P_M=e_Mc_a\langle Y_a\rangle/\sqrt M, \qquad \partial_{u_a}^2\mathbb EP_M=e_M^2c_a^2\mathbb E\langle(Y_a-\langle Y_a\rangle)^2\rangle.\] Use convex difference quotients at \(s=M^{-1/4}\), the variance bound (32), and (34). The expected absolute disorder fluctuation of the first derivative is \(O_a(s+1/(s\sqrt M))\). Dividing (34) by \(s^2\) and letting \(s\to0\) also gives \[\mathbb E\langle(Y_a-\langle Y_a\rangle)^2\rangle \le\frac{2}{e_M^2c_a}.\] Consequently the disorder and thermal contributions, respectively, obey \[\frac{\mathbb E\langle|Y_a-\mathbb E\langle Y_a\rangle|\rangle} {\sqrt M e_M} \le C_a\left(\frac{M^{-1/4}}{e_M^2} +\frac1{\sqrt M e_M^2}\right) =O_a(M^{-1/8}+M^{-3/8}).\] In particular, \[ \mathbb E\langle|Y_a-\mathbb E\langle Y_a\rangle|\rangle =o(\sqrt M e_M) \tag{35}\] for every fixed \(a\). Gaussian integration by parts against a bounded test \(F\) on \(r\) replicas yields \[\mathbb E\langle Y_a(1)F\rangle =\sqrt M e_Mc_au_a\,\mathbb E\left\langle F\left(\sum_{b=1}^r\kappa_a(1,b)-r\kappa_a(1,r+1)\right)\right\rangle.\] Replacing \(Y_a\) by its expected mean using (35), and subtracting the one-replica identity, proves GG in every subsequential array limit. Compactness supplies such limits, and Section 3 describes them by jointly ordered paths \(\widetilde q_j\) and the fixed label path.

The contradiction from derivatives

Increase a field increment at index \(i\) by a positive matrix supported on \(F_{j,i}\). The derivative of \(-\mathbb EP_M\) is \(n_j/(2M)\) times the Gibbs mean overlap contracted with that cumulative field increment. At the minimum, the limiting right derivative of [eq:contact] therefore gives \[ \left.\int_{\zeta_i}^1(\widetilde q_j(u)-q_j(u))\,du\right|_{F_{j,i}} \succeq0. \tag{36}\] Outside the active space the pinned-coordinate differences vanish. The tail difference is zero at \(t=1\). On each grid interval, the trial is constant and \(\widetilde q\) is monotone; testing quadratic forms shows the tail inequality between endpoints as well (the tail difference is concave there). Thus (22) implies \(V_D(\widetilde q)\le V_D(q)\).

Poisson differentiation in \(r\) gives minus the expected log integral of one fresh pattern. Its subsequential limit is \(-V_D(\widetilde q)\) by fresh-field convergence. Consequently the limiting \(r\)-derivative of [eq:contact] is at least \(\delta>0\). A minimizing point with \(r>0\) permits a left variation and must have nonpositive left derivative. This contradiction proves \[ \limsup_M\frac1M\mathbb E\log\mathcal Z_{M,\theta}\left(\sum_rD(\xi_r)\right) \le\mathcal G_{\theta,p}(D). \tag{37}\]

The Gaussian cavity lower bound

We prove the reverse inequality to (37) for smooth compactly supported \(D\). The proof follows the variational cavity approach of Aizenman, Sims, and Starr [2]. Its central comparison is between two descriptions of the new spin coordinates. Spatial rotation invariance makes their empirical overlaps approximate the physical overlaps of the enlarged system. A Gaussian calculation on a thick shell makes the same empirical overlaps approximate the canonical path \(q^*\) of Lemma 4. Matching these descriptions selects the entropy in the lower bound. This comparison is related to the critical-point method of Chen and Mourrat [5]; here two-pattern restoration and the shared-coordinate shell calculation establish the required consistency relation.

Use fixed pattern count \(M\) and the cascade with only the outer split at \(\theta\). Add the same monomial perturbations, now with amplitude \(s_M=M^{3/8}\) in place of \(\sqrt M e_M\). Average expected logarithms over a single product probability law \(\pi\) under which every \(u_a\) is uniform on \([1,2]\); use this same law at every system size. The perturbation changes a log partition function by \(O(s_M^2)=o(M)\).

Let \(L_{M,\boldsymbol n}(u)\) be the expected perturbed cascade log partition function with block dimensions \(\boldsymbol n=(n_j)_j\). For a fixed cavity size \(\ell\), increase \(M\) to \(M+\ell\) and \(n_j\) to \(n_j+r_j\), where \(r_j=p_j\ell+O(1)\) are integers, positive when \(\ell\) is sufficiently large. Define \[\Delta_{M,\ell}(u)=L_{M+\ell,\boldsymbol n+\boldsymbol r}(u)-L_{M,\boldsymbol n}(u).\] We will bound its parameter average from below by \(\ell\mathcal G_{\theta,p}(D)-o(\ell)\). The order of limits is \(M\to\infty\) with \(\ell\) fixed, followed by \(\ell\to\infty\); at each fixed \(\ell\), pass to a subsequence on which the integer increments are fixed. The same parameter law \(\pi\) at the two sizes will make these increments telescope.

Generic parameters and two-pattern identification

There are sets of perturbation parameters of \(\pi\)-measure tending to one on which every subsequential overlap limit of the base system with two patterns omitted satisfies joint GG. To see this, for each fixed perturbation coordinate the pressure variance is \(O(M^{-1})\) as before. Integration by parts bounds its expected first derivative by \(O_a(s_M^2/M)\). Integrating the second derivative controls thermal variance. Integrated convex difference quotients with step \(s\) give the bound \[C_a s\,s_M^2/M+C/(s\sqrt M)\] on the integrated absolute fluctuation of the first derivative. At \(s=M^{-1/8}\) this implies \[\int\mathbb E\langle|Y_a-\mathbb E\langle Y_a\rangle|\rangle/s_M\,d\pi\longrightarrow0.\] The estimates hold on a slightly larger coupling interval, so endpoint difference quotients cause no difficulty. Markov’s inequality and a diagonal selection give the stated good sets for every fixed coordinate. Integration by parts then proves GG exactly as in the upper bound. No separate GG assertion at size \(M+\ell\) is needed.

Choose an arbitrary sequence of good parameters, and pass to joint finite-array subsequential limits of the two-omitted system and the full size-\(M\) system. Write \(\xi^r(X)\) for the base row fields and \(D'_j,D''_{jj}\) for block derivatives. In the full system retain the averages \[\begin{align*} B_j(X,\widetilde X)&= \frac1{n_j}\sum_{r\le M}D'_j(\xi^r(X))D'_j(\xi^r(\widetilde X))^{\mathsf T},\\ c_j(X)&=\frac1{n_j}\sum_{r\le M} \left\{\operatorname{tr}D''_{jj}(\xi^r(X))-(\xi_j^r(X))^{\mathsf T}D'_j(\xi^r(X))\right\}. \tag{38}\end{align*}\] Also retain separately the entries of the one-path Hessian and \(\xi_j^r(D'_j)^{\mathsf T}\) averages. They are bounded because \(D\) is smooth and compactly supported.

Lemma 5 (Bulk covariance identification). The full-system overlap array has the same GG law as the two-omitted limit. Along its joint matrix-and-label order, the off-diagonal limits of \(B_j\) are deterministic bounded nondecreasing positive semidefinite paths \(b_j(u)\). Their diagonal is a deterministic matrix \(B_{j,\mathrm d}\succeq b_j(u)\). The indicated one-path averages have deterministic limits, including \(c_j\), and \[ c_j+\operatorname{tr}B_{j,\mathrm d} =\int_0^1\operatorname{tr}(q_j(u)b_j(u))\,du. \tag{39}\] The paths \(b_j\) are functions of the actual joint overlap and label value; there is no extra variation on a flat portion of that value.

Proof. Restore the two omitted patterns by bounded reweighting. In finite cascade roundings their marks have independent components conditional on the transformed sampled shape. Reweighting preserves the unmarked overlap law. For one pattern and two visits at common level \(i\), branching independence gives its conditional terminal-gradient cross moment as the second moment of the conditional gradient mean through level \(i\). These matrices are positive semidefinite, nondecreasing, and bounded by the terminal-gradient second moment. Two independent restored patterns give products of these conditional moments. One-path observables similarly have constant means and products of means for the two patterns.

If the joint matrix-and-label path is flat, every restored-pattern Gaussian increment is zero and there is no change of outer sharing. The corresponding recursion is the identity, its tilted kernel is a point mass, and its conditional gradient mean does not change. Thus \(b_j\) is constant there. This conclusion survives rounding: bounded smooth gradients and Hessians bound the \(L^2\) gradient-martingale variation over a collection of intervals by a constant times the total covariance-trace increment. Collapse equal joint values and keep whole atoms as bins. Vanishing matrix increments then cannot leave a hidden gradient-covariance increment. A sharing transition is retained by the label coordinate even if the matrices themselves are already pinned. The independent diagonal residual is not part of this off-diagonal assertion.

Here is the conditional-variance step. The conditional terminal-gradient moment matrices on finite roundings are bounded and monotone. Select their limits as the rounding is refined, retaining whole atoms of the joint order. For one matrix entry, let \(\beta(T)\) be \(1/p_j\) times this limiting conditional moment at the joint overlap-and-label value \(T\) of the sampled pair, and let \(A\) be the subsequential limit of the corresponding entry of \(B_j^{12}\). Exchangeability gives the following identities after passing through the finite roundings, for every bounded continuous \(\varphi\): \[\mathbb E[A\varphi(T)]=\mathbb E[\beta(T)\varphi(T)].\] Indeed the empirical sum contributes the factor \(M/n_j\to1/p_j\). Its square contains two distinct restored patterns with coefficient \(M(M-1)/n_j^2\to1/p_j^2\); repeated indices contribute \(o(1)\). Conditional factorization for the distinct patterns consequently gives \[\mathbb EA^2=\mathbb E\beta(T)^2.\] The first identity extends to bounded measurable \(\varphi\); taking \(\varphi=\beta\) and using the second gives \(\mathbb E(A-\beta(T))^2=0\). Thus the empirical covariance is a deterministic function of \(T\). For diagonal and one-path observables the same calculation has a constant conditional mean, giving their deterministic limits.

Finally apply Gaussian integration by parts to one full pattern. The expected \(\xi_j^{\mathsf T}D'_j\) equals the Hessian trace plus the diagonal gradient square minus the two-replica contraction of their gradients with \(\mathcal R_j^{12}\). Sum over patterns, divide by \(n_j\), and pass to the identified limits. This gives (39) with the displayed sign. ◻

Splitting a Stiefel frame

Write both the base and enlarged Stiefel frames with unit-length columns. The enlarged frame admits the decomposition \[ \begin{pmatrix} X_jT_j/\sqrt{n_j}\\ Z_j/\sqrt{n_j+r_j} \end{pmatrix},\qquad T_j^{\mathsf T}T_j=I-\frac{Z_j^{\mathsf T}Z_j}{n_j+r_j}, \tag{40}\] where \(X_j^{\mathsf T}X_j=n_jI\), the bottom block has \(r_j\) rows, and \(T_j\) is upper triangular with positive diagonal. The base frame \(X_j\) has its uniform law and is independent of \(Z_j\). The first column is the outer variable in (5), so its coupling must be fixed before drawing the conditional inner columns. The decomposition has exactly this order: the first column of \(Z_j\) is drawn at the outer vertex independently of the base outer column; its remaining columns are drawn at visits conditionally on that first column and independently of the drawn base frame. These facts follow by splitting the first column and then using rotational invariance and triangular orthogonalization for the conditional remaining columns.

For fixed \(r_j\), the entries of \(Z_j\) converge to independent standard Gaussians. The conditional inner statement holds for every converging bounded outer column: the projection of its orthocomplement onto the last \(r_j\) rows has Gram matrix tending to identity. Representing a uniform frame by orthonormalized standard Gaussian columns proves the assertion and its conditional version.

Fix a small \(\epsilon>0\) and restrict the larger partition function to the thick shell \[ \mathcal I=\left\{\lVert Z_j^{\mathsf T}Z_j/r_j-I\rVert\le\epsilon \text{ for every }j\right\}. \tag{41}\] This only lowers its logarithm. On the restriction replace the larger perturbation by the base perturbation at the same \(u\). For fixed \(\ell\), overlaps differ by \(O_\ell(M^{-1})\). Summable kernel Lipschitz constants, the amplitude \(s_M\), and the exponentially small coefficient tail imply an \(O_\ell(M^{-1/4})+o(1)\) covariance error in unnormalized logarithms. Gaussian interpolation gives a vanishing cost both for expected logs and for bounded expected replica tests. In the latter case differentiation inserts a finite linear combination of covariance differences. The estimates are uniform in \(u\).

There is a second, physical prediction for the new-coordinate overlaps. Under the original larger Gibbs law, \[ (Z_j^1)^{\mathsf T}Z_j^2/r_j\simeq\mathcal R_{j,\mathrm{whole}}^{12}, \qquad (Z_j^1)^{\mathsf T}Z_j^1/r_j\simeq I, \tag{42}\] with exceptional expected probability tending to zero, first as \(M\to\infty\) and then as \(\ell\to\infty\), at each fixed tolerance, uniformly in \(u\). Indeed the disorder-averaged sampled frames are invariant under a common spatial rotation in each species: this is true of the Gaussian patterns, perturbation kernels, and reference marks. Projection of a fixed finite collection of vectors onto \(r_j\) random coordinates approximates all its normalized inner products, by polarization and the spherical squared-projection variance. This uses rotation invariance after disorder averaging, not independence of coordinates under fixed disorder.

The self part of (42) makes the discarded Gibbs mass in (41) small for large \(\ell\). Markov’s inequality transfers (42) to the restricted Gibbs law. On the restriction, whole and base overlaps differ uniformly by \(o(1)\) at fixed \(\ell\). Perturbation replacement therefore preserves the same prediction with the base overlap on the right.

Expansion and fixed-size integrability

Divide the restricted enlarged cascade integral, after the preceding perturbation replacement, by the unrestricted base cascade integral, and call this random ratio \(R_{M,\ell}\). Restriction lowers the numerator and perturbation replacement costs \(o_M(1)\) at fixed \(\ell\), so \(\Delta_{M,\ell}(u)\ge\mathbb E\log R_{M,\ell}-o_M(1)\). After division, the reference measure is the base Gibbs law augmented by fresh outer \(Z\) marks at its labels and the conditional inner \(Z\) variables from (40). It is this ratio that we now expand. In an old pattern the field in block \(j\) is \[T_j^{\mathsf T}\xi_j^r(X)+Z_j^{\mathsf T}h_j^r/\sqrt{n_j+r_j},\] where \(h_j^r\) are independent standard Gaussian vectors of length \(r_j\).

Taylor expansion gives a linear Gaussian field coupled to each row of \(Z_j\). Up to a vanishing error at fixed \(\ell\), it is \(n_j^{-1/2}\sum_{r\le M}h_j^r(D'_j(\xi^r))^{\mathsf T}\); the exact denominator \(\sqrt{n_j+r_j}\) may be replaced by \(\sqrt{n_j}\) because \(r_j\) is fixed. Its conditional cross covariance is therefore \(B_j(X,\widetilde X)\) in the limit. The deterministic second-order correction is a one-path quadratic function \(W(X,Z)\) with coefficients given by the retained one-path averages. Uniformly on the shell, \[ W(X,Z)=\frac12\sum_jr_jc_j(X)+O(\epsilon\ell). \tag{43}\] To check the normalization term, the first-order upper triangular part of \(T_j-I\) has symmetric sum \(-Z_j^{\mathsf T}Z_j/n_j\); at \(Z_j^{\mathsf T}Z_j=r_jI\) it is \(-r_jI/(2n_j)\). The mean fresh quadratic term is the Hessian contracted with \(Z_j^{\mathsf T}Z_j/(2n_j)\). The diagonal specialization of their sum is \(r_jc_j(X)/2\); the retained matrix coefficients are bounded, so moving within the shell costs \(O(\epsilon r_j)\). Summing over \(j\) gives (43). Cross-species fresh quadratic terms are centered. The \(\ell\) added patterns can use their independent base fields at this accuracy.

The expansion requires two estimates: errors in the positive integral must vanish, and logarithms of the resulting ratios must be uniformly integrable. These are separate issues because the restricted insertion may approach zero; its reciprocal also enters normalized Gibbs tests. For fixed \(\ell\), the expected supremum of the dilation and third-order remainders on the shell tends to zero. Compact support controls powers of the base fields occurring in near-identity dilations, while bounded third derivatives and Gaussian moments give an \(O_\ell(M^{-3/2})\) error per old pattern for the fresh shifts. Centered fresh quadratic terms have conditional second moments \(O_\ell(M^{-1})\) at fixed spins and bounded \(Z\). Their exponential absolute-value moments at every fixed power are uniformly bounded for large \(M\), since their supremum is bounded by \[C_\ell\left(1+M^{-1}\sum_{r,j}|h_j^r|^2\right).\] The linear insertion has uniformly bounded conditional exponential moments on the shell. Integration over the independent reference variables thus makes the absolute-integral error from dropping centered terms tend to zero in probability.

Small denominators require a further estimate. At the fixed sharing exponent \(\theta\), merged base Gibbs cluster masses, in transformed order, are normalized \(\theta\)-stable jumps independent of the fresh outer marks. Retain outer \(Z\) norm bands strictly inside the shell. At fixed sufficiently large \(\ell\) their probabilities are positive; conditional inner shell probabilities are uniformly positive on these bands, by the Gaussian frame limit. Hence the reference shell mass dominates a positive constant times an independent thinning of those stable masses. Its negative logarithm has bounded second moment: if \(q>0\) is the retention probability, the retained mass is \[ W_q=\frac{S_q}{S_q+S_{1-q}}, \tag{44}\] with independent positive \(\theta\)-stable totals, and \(S_q\stackrel d=q^{1/\theta}S_1\). Stable logarithmic second moments give the bound. Jensen on the normalized shell handles the Gaussian linear insertion; its conditional variance is bounded for fixed \(\ell\). Other terms are bounded there. These estimates control negative log tails, while the preceding exponential moments control positive tails. Therefore the ratios are tight away from zero in probability and their logs are uniformly integrable.

One can now first truncate exponentials and denominators and then use the finite-array polynomial approximation of Section 3. Shared \(Z\) marks are included through the exact label split; their conditional frame limits hold, and the shell boundary is Gaussian-null. Removing the truncations proves convergence of both logs and bounded Gibbs tests through the expansion. The same estimates give, for each fixed \(\ell\), a uniform lower bound on \(\Delta_{M,\ell}(u)\) even at parameters outside the good set.

Matching physical and canonical overlaps

Lemma 5 and the preceding passage express a lower bound on the limiting increment through arbitrarily fine finite cascade roundings of the joint paths \(q,b\). In such a rounding, draw standard Gaussian \(Z\) rows with the sharing rule, couple them linearly to fields of shared covariances \(b_{j,i}\), retain the shell, add (43) with its deterministic limiting coefficients, and add \(\ell\) independent fresh-pattern factors with covariances \(q\). The linear field also has an independent visit residual with covariance \(B_{j,\mathrm d}-b_{j,k-1}\). Integrating it adds \[\frac12\sum_j\operatorname{tr}\big((B_{j,\mathrm d}-b_{j,k-1})Z_j^{\mathsf T}Z_j\big)\] to the correction. The fresh patterns likewise use their true diagonal covariance. These rounded \(q\) paths describe overlap arrays and need not yet have strict active susceptibility.

We now compare two measures on this same rounded cascade. The first is the restricted cavity measure just obtained from \(R_{M,\ell}\). The second is unrestricted and canonical: keep the new-pattern factors and the Gaussian \(Z\) reference marks, but replace the correction by \(\frac12\sum_j\operatorname{tr}(C_jZ_j^{\mathsf T}Z_j)\), where \(C_j=C(b_j)\) is the multiplier from Lemma 4; do not include a linear visit residual in the canonical comparison. On the shell the difference of log weights is \[ \frac12\sum_jr_j\left\{ c_j+\operatorname{tr}(B_{j,\mathrm d}-b_{j,k-1}-C_j)\right\}+O(\epsilon\ell). \tag{45}\] All constants here are uniform in the rounding grid, because the multipliers and bulk coefficients are uniformly bounded. The constant term in (45) changes the log partition function but cancels when normalizing either measure on the shell. Only the \(O(\epsilon\ell)\) variation affects the comparison of their Gibbs probabilities.

Conditional on the transformed unmarked sampled shape, the canonical spatial rows are iid jointly Gaussian across replicas. Their self covariance is \(I\) and their cross covariance at common level \(i\) is \(q_{j,i}^*\) from Lemma 4. The independent root rows are averaged in this assertion. Additive row energies and new-pattern energies factor the tilted kernels, so new patterns do not bias these row laws. This assertion is used only in the unrestricted canonical measure; the joint shell restriction may couple rows.

Gaussian concentration therefore gives, uniformly over the grids, exponentially small expected canonical Gibbs probability of a shell failure at fixed \(\epsilon\), and of \[ \lVert (Z_j^1)^{\mathsf T}Z_j^2/r_j-q^*_{j,\mathrm{level}}\rVert>\eta \tag{46}\] at fixed \(\eta>0\). Denote these bounds by \(Ce^{-c_\epsilon\ell}\) and \(Ce^{-c_\eta\ell}\). If the canonical shell mass is at least \(1/2\), (45) bounds the actual restricted pair law by \[ 4e^{4C\epsilon\ell} \tag{47}\] times the unrestricted canonical pair law. The probability that this shell mass is below \(1/2\) is at most \(2Ce^{-c_\epsilon\ell}\). Choose \(\epsilon<c_\eta/(8C)\) and then \(\ell\) large. The canonical prediction (46) therefore transfers to the restricted cavity measure with arbitrarily small exceptional probability. This uses row independence only under the unrestricted canonical law; the shell itself can couple the rows.

The restricted cavity measure also retains the physical prediction (42), because the expansion passed bounded Gibbs tests as well as logarithms. It predicts the rounded overlap \(q_{j,i}\) at the sampled common level. After averaging the cascade weights and marks, marked Poisson reweighting, including the shell restriction, keeps the sampled common-level probabilities exactly \(w_i\). Since all matrix overlaps have bounded norm, agreement of both empirical predictions outside small exceptional events gives \[ \sum_{j,i}w_i\lVert q_{j,i}-q^*_{j,i}\rVert\longrightarrow0 \tag{48}\] in the following precise sense. First choose the desired overlap tolerance, then \(\epsilon\) as above, then a sufficiently large fixed \(\ell\). At this fixed \(\ell\), take the finite-array subsequential limit as \(M\to\infty\) and choose a sufficiently fine rounding. Both predictions hold in the resulting joint sampled calculation, and their failure probabilities bound the weighted path distance by twice the chosen tolerance plus a bounded multiple of their sum. No finite-size convergence uniform over all grids is asserted or needed.

For the marking assertion, an outer Gaussian column in a narrower norm band leaves an open set of inner columns satisfying the shell constraint. Both events have positive Gaussian probability, so the surviving root normalizer is positive. Infeasible outer columns simply give zero marks. The canonical fractional moments are bounded by the unrestricted strict-domain recursion; in the actual shell calculation bounded spin norms give finite Gaussian exponential moments. Thus pruning and backward marking apply and preserve the unmarked level law. Independence of canonical spatial rows was used before restriction; after restriction only the comparison (47) is used.

The expected logarithmic cost of the shell

The preceding comparison identifies the overlap path. To bound \(\mathbb E\log R_{M,\ell}\), we also need the log partition function on the canonical shell. Its high Gibbs probability alone does not control its expected logarithmic cost. Write \(Z_c=\int e^{H_c}d\mu\) and \(N_c=\int_{\mathcal I}e^{H_c}d\mu\) for the canonical denominator and shell numerator on the ordinary cascade/spin reference measure.

Let \(A=\mu(\mathcal I)\). At the fixed exponent \(\theta\), choose narrower bands for the outer columns. For all large \(\ell\) their joint probability \(q_\ell\) is bounded below by \(q_0>0\); conditional inner Gaussian Gram bands have probability at least \(c_0>0\). Thus \(A\ge c_0W_{q_\ell}\) with the thinning mass (44). In particular \[ \mathbb E(-\log A)^2\le C(\theta,q_0,c_0), \tag{49}\] independently of every other cascade exponent and of the depth of the grid.

Jensen on the normalized reference shell gives \(\log N_c\ge\log A+\int H_c\,d\mu_{\mathcal I}\). On the shell the quadratic and bounded-pattern terms have magnitude at most \(C\ell\). Conditional on reference spin marks and weights, the averaged linear insertion is centered Gaussian with variance at most \(C\ell\), using the bounded total field covariance and total shell spin norm \(O(\ell)\). Therefore \[ \mathbb E[(\log N_c)_-^2]\le C\ell^2. \tag{50}\] For the denominator, relative entropy gives \(\log Z_c\le\langle H_c\rangle_c\). The canonical spin self covariance is \(I\), its fourth moments are bounded, and the terminal-field fourth moments are uniformly bounded by Lemma 4. Cauchy–Schwarz over \(O(\ell)\) row energies and \(\ell\) bounded patterns yields \[ \mathbb E[(\log Z_c)_+^2]\le\mathbb E\langle H_c^2\rangle_c\le C\ell^2. \tag{51}\] These bounds are uniform in the grid. Since \(\mathbb E(1-N_c/Z_c)\le Ce^{-c_\epsilon\ell}\), the event \(N_c/Z_c<1/2\) has exponentially small probability. On its complement \(-\log(N_c/Z_c)\le\log2\); on the event use (50)–(51) and Cauchy–Schwarz. We obtain \[ \mathbb E[-\log(N_c/Z_c)]=o(\ell). \tag{52}\] This argument depends on the fixed exponent \(\theta\), not on a lower bound for the intermediate grid exponents.

Cancellation and telescoping

We can now evaluate the increment. Additive independent mark recursion gives the unrestricted canonical log value \[\ell V_D(q)+\sum_jr_j\Psi(C_j,b_j).\] Add (45) and subtract the \(o(\ell)\) shell cost. For block \(j\), the coefficient of \(r_j\) is \[\Psi(C_j,b_j)+\tfrac12\{c_j+\operatorname{tr}(B_{j,\mathrm d}-b_{j,k-1}-C_j)\} =J(b_j)+\tfrac12\{c_j+\operatorname{tr}B_{j,\mathrm d}\}.\] Thus the multiplier and the terminal field shift cancel, including the independent visit residual. Lemma 4 and (39) turn the lower bound into \[ \ell V_D(q)+\sum_jr_j\left\{ S_\theta(q_j^*)+\frac12\sum_iw_i \operatorname{tr}((q_{j,i}-q^*_{j,i})b_{j,i})\right\}-o_{\rm tol}(\ell). \tag{53}\] Rounding errors in (39) tend to zero at fixed \(\ell\). The paths \(q_j^*\) are eligible, and \(S_\theta(q_j^*)\) is uniformly bounded: use (24), comparison with \(q^0\), bounded \(b\), and \(S_\theta\ge0\). With (48), the Lipschitz estimate (21), and \(r_j=p_j\ell+O(1)\), division of (53) by \(\ell\) gives a lower bound \(\mathcal G_{\theta,p}(D)-o_{\rm tol}(1)\).

Here is the uniform sequential statement used to sum increments. For every \(\delta>0\), there is \(\ell_0\) such that for each fixed integer \(\ell\ge\ell_0\), every sequence of base sizes and dimensions with \(n_j/M\to p_j\) and increments \(r_j=p_j\ell+O(1)\) satisfies \[ \liminf_{M\to\infty}\frac1\ell\int\Delta_{M,\ell}(u)\,d\pi(u) \ge\mathcal G_{\theta,p}(D)-\delta. \tag{54}\] The constant in \(O(1)\) is fixed across these choices. To prove the statement, begin with any good-parameter sequence and any further compact array limit. The tolerance, shell, and rounding choices above apply to it with uniform canonical constants. Failure of the eventual bound would give such a sequence contradicting (53). The bad parameter sets have vanishing measure and use the fixed-\(\ell\) uniform lower increment bound proved earlier. This establishes (54) after averaging.

For arbitrary target dimensions, telescope along pattern sizes \(M'\) in the residue class of the target \(M\) modulo \(\ell\), with block dimensions \(\lfloor M'n_j/M\rfloor\). The last point is exactly \((M,\boldsymbol n)\). A block increment differs from \(\ell n_j/M\) by at most one; for fixed \(\ell\) and large target \(M\), it therefore differs from \(p_j\ell\) by a fixed bounded amount. The sequential criterion applies uniformly to all large base sizes along these triangular chains; otherwise a failing sequence would contradict (54). Finitely many small initial sizes, including those too small to support the frames, may be skipped with bounded starting cost. The same parameter law \(\pi\) at both endpoints makes the log increments telescope exactly. Divide by the target \(M\), let \(M\to\infty\), then \(\ell\to\infty\) and \(\delta\downarrow0\), and remove the \(o(M)\) perturbation. This proves the matching lower bound in (15). For smooth compactly supported \(D\), convergence in probability follows from bounded differences in the independent patterns, and expectation convergence follows as well.

Quadratic tails and global soft constraints

The Gaussian formula has so far been proved for smooth compactly supported terminals. The Haar reduction needs negative quadratic terms and penalties on empirical row averages. We first control Gaussian row tails uniformly in the trial path. We then show that an additive multiplier makes each constrained average concentrate at a deterministic value; the box optimality conditions select the multiplier for the soft penalty.

Extension of the terminal class

The smooth proof extends to continuous potentials satisfying \[ -C(1+|\xi|^2)\le D(\xi)\le C \tag{55}\] by bounded smooth approximation on compact sets. We justify a uniform tail estimate, including its use in the variational infimum.

Omit one row and view its restoration in the cascade representation. Its partition ratio is an integral of \(e^{D(\xi)}\) against an omitted-row Gibbs reference independent of that row. Jensen bounds this ratio from below by \[\exp\{-C(1+\langle|\xi|^2\rangle_{\rm omit})\}.\] If \(R_D\) denotes this ratio, another Jensen inequality gives, for a sufficiently small fixed \(a>0\), \[\mathbb ER_D^{-a}\le e^{aC}\mathbb E\langle e^{aC|\xi|^2}\rangle_{\rm omit}<\infty.\] The bound is uniform because every fresh self field is a standard Gaussian vector of the fixed total dimension \(\sum_jd_j\). The numerator of a tail event is bounded above by \(e^C\) times its omitted-row probability. Split according to whether the denominator is smaller than \(e^{-cR^2}\) and use its negative moment together with Gaussian tails in the numerator. This gives constants \(c,C'>0\), independent of system size and of a trial grid, such that under the expected natural tilted path law \[ \mathbb E\langle\mathbf 1_{\{\lVert \xi\rVert>R\}}\rangle\le C'e^{-cR^2}. \tag{56}\] The same proof applies to one-pattern trial recursions via their cascade representation. It remains valid uniformly over bounded families in (55).

Without perturbations, the expected cascade path law conditional on the common disorder is precisely the natural powered outer/inner law from (5); this follows by the marked change of measure. Interpolation between compact approximations and (56) therefore bounds the \(L^1\) difference of normalized log partition functions by a quantity tending to zero. It also bounds the difference of \(V_D\) uniformly over all eligible paths. Consequently the infima converge and both probability and expectation convergence extend to (55). Bounded continuous functions plus nonpositive quadratic forms are included. Every differentiation used here can first be done with truncated functions; at finite size fields on the compact Stiefel domain are bounded for each fixed disorder.

Differentiability and the powered restriction ratio

Write \(p(t)=\mathcal G_{\theta,p}(D-t\cdot g)\). It is convex as a limit of convex finite pressures. For any bounded continuous row observable \(h\), the limiting pressure with added coupling \(s h\) is differentiable in \(s\). To prove this without assuming uniqueness of a minimizing path, differentiate one finite recursion. Its first derivative is the tilted conditional mean of \(h\). Differentiating the recursion once more sums the conditional variance increments of that mean martingale, each multiplied by its step coefficient \(\zeta_i\in[0,1]\). Without these coefficients the nonnegative increments telescope to at most the terminal variance of \(h\), which is bounded by \(\lVert h\rVert_\infty^2\). Hence \[ 0\le\frac{d^2}{ds^2}V_{D+s h}(q)\le\lVert h\rVert_\infty^2, \tag{57}\] uniformly over grids, paths, and \(s\). Subtract \(\lVert h\rVert_\infty^2s^2/2\) from every entropy-plus-recursion objective. Each becomes concave, and their infimum is concave. The limiting pressure is therefore semiconcave as well as convex, hence differentiable. Infinite-entropy paths do not affect this argument; the baseline trial and the terminal bounds keep the value finite.

Differentiability and exponential tilting imply that the empirical average of \(h\) concentrates at its deterministic derivative under the natural tilted path law, in probability including disorder. The ratio needed for this argument is not an ordinary Gibbs mass. If \(u(x)\) is the conditional inner probability of an event under an additive tilt, then exactly \[ \frac{\mathcal Z(H;\text{event})}{\mathcal Z(H)} =\left(\mathbb E_{\rm out}^{H}u^\theta\right)^{1/\theta}. \tag{58}\] For a one-sided empirical-average event, the usual pointwise exponential bound survives inner integration, the positive outer power, and the final root. Probability convergence of the nearby coupled pressures and differentiability make the ratio exponentially small in probability. More explicitly, if \(v=\mathbb E_{\rm out}^{H}u\) is its natural probability and \(R\) is the restricted ratio, then \(u\le u^\theta\) and Jensen’s inequality give \[v^{1/\theta}\le R\le v.\] Thus an exponentially small \(R\) forces \(v\to0\), while \(v\to1\) forces \(R\to1\) at fixed \(\theta\). These two implications will be used for concentration and for a cost-free restriction, respectively.

Apply this to bounded clippings of every \(g_\ell\). Estimate (56) is uniform for \(t\) in the compact box \([0,d]^s\): all unbounded quadratic coefficients remain nonpositive. Thus \(\bar g\) concentrates at a deterministic vector \(v(t)\), obtained as the limit of the clipped centers. For every direction \(a\) into the box, the one-sided derivative is \[ \partial_a p(t)=-a\cdot v(t). \tag{59}\] To justify passage of slopes through clipping, interpolate only the increment in the coupling. Increasing a nonnegative quadratic coupling adds a nonpositive term. If it decreases, write its contribution, for \(a_\ell<0\) and \(u>0\), as \[-(t_\ell-u|a_\ell|)g_\ell -u|a_\ell|(1-\gamma)(g_\ell-g_{\ell,R}),\qquad 0\le\gamma\le1,\] where \(g_{\ell,R}\) is a nonnegative bounded clipping. Both unbounded coefficients are nonpositive while the direction remains in the box. Uniform tails then make the directional slope error at most \(\sum_\ell|a_\ell|\epsilon_R\), with \(\epsilon_R\to0\). This proves (59), including boundary directions.

Soft duality

Choose a minimizer \(t\in[0,d]^s\) of \(p(t)+t\cdot c\), which exists by continuity. Equation (59) gives the box conditions \[t_\ell=0\Rightarrow v_\ell(t)\le c_\ell,\qquad 0<t_\ell<d\Rightarrow v_\ell(t)=c_\ell,\qquad t_\ell=d\Rightarrow v_\ell(t)\ge c_\ell.\] Equivalently, \[ t_\ell(v_\ell(t)-c_\ell)=d(v_\ell(t)-c_\ell)_+. \tag{60}\] For every configuration the soft Hamiltonian is bounded above by the additive Hamiltonian \[H_t=\sum_rD(\xi_r)-Mt\cdot\bar g+Mt\cdot c,\] because \(d(x-c)_+\ge t(x-c)\) for \(t\in[0,d]\). This proves the upper bound in (16). On \(\lVert \bar g-v(t)\rVert_1\le\delta\), (60) gives the converse inequality up to \(2dM\delta\). This set has natural tilted probability tending to one; by (58) its restriction costs \(o(M)\) in the log partition function. Letting \(\delta\downarrow0\) gives the matching lower bound in probability. Theorem 3 is proved without any interchange of a minimum and an integral.

Haar flags and weighted shells

We prove Theorem 1 for the Haar angular pressure (2). Keep the mutually orthogonal subspaces \(E_j\), dimensions \(n_j/M\to p_j>0\), and radii \(Ma_j\) from its definition. The Gaussian formula will evaluate a coefficient integral with projection witnesses; a weighted comparison of Haar shells then identifies that value with the desired angular pressure.

We prove the theorem first for bounded Lipschitz \(f\) and \(p<1\). Coupling block directions gives the useful uniform estimate \[ |P_M^f(a)-P_M^f(a')| \le\operatorname{Lip}(f)\left(\sum_j(\sqrt{a_j}-\sqrt{a'_j})^2\right)^{1/2}. \tag{61}\]

Gaussian spans and shell entropy

Generate the flag \(E_1\subset E_1\oplus E_2\subset\cdots\) from consecutive blocks of an \(M\times n\) matrix \(G\) with independent \(N(0,1/M)\) entries, \(n=\sum_jn_j\). Gaussian orthogonalization gives its Haar law; see [14] for this construction. On an event of probability at least \(1-e^{-cM}\), its norm and the norm of its inverse on its image are bounded by fixed constants \(C_{\rm op},C_{\rm inv}\). An elementary net proof suffices: an \(r\)-dimensional unit sphere has a \(\delta\)-net of size at most \((1+2/\delta)^r\); fixed-vector Gaussian norms have arbitrarily strong exponential upper tails at a large constant and small-ball probability at most \((C\epsilon)^M\). First obtain the upper bound at a fixed mesh and then take a mesh of order \(\epsilon\) for the lower bound. The strict inequality \(n/M\to p<1\) makes the small-ball exponent dominate the net size for small enough \(\epsilon\).

The volume Jacobian of \(x\mapsto Gx\) is \(\sqrt{\det(G^{\mathsf T}G)}\). Its logarithm divided by \(M\) converges in probability to \(J_p\): successive orthogonal residual lengths are independent with laws \((\chi^2_{M-r+1}/M)^{1/2}\), and chi-square concentration, followed by a Riemann sum, gives \[ \frac1{2M}\log\det(G^{\mathsf T}G)\longrightarrow \frac12\int_0^p\log(1-u)\,du. \tag{62}\]

The angular pressures concentrate. On the good norm and inverse event, the flag projectors are Lipschitz in the Frobenius distance of \(G\), with a fixed operator-norm constant: represent unit vectors by bounded coefficients, compare their distances to the other image, and apply this to each initial block. Differences of consecutive projectors give the individual blocks. Nearby frames may be aligned by projecting and multiplying by the inverse Gram square root; their operator-norm change is bounded by the projector change. Coupling directions and using (61) bounds the pressure change. Boundedness handles pairs farther apart. A Lipschitz extension from the good event has Lipschitz constant \(O(M^{-1/2})\) in standard Gaussian entries, hence exponential concentration. For reference, the Gaussian exponential bound follows by centering with an independent copy, joining the copies by a circular Gaussian interpolation, and applying Jensen to the integrated directional derivative; the independent Gaussian velocity gives \(\mathbb Ee^{t(F-\mathbb EF)}\le e^{Ct^2\operatorname{Lip}(F)^2}\). Smoothing removes differentiability assumptions.

By boundedness, equicontinuity, and concentration, every subsequence has a further subsequence on which \(P_M^f\) converges in probability uniformly on compact sets to a deterministic bounded continuous \(P_\infty\). We identify this function. Radial integration and Stirling’s formula show that the weighted Lebesgue volume in \(E_1\oplus\cdots\oplus E_h\) of a small shell with \(\lVert y_j\rVert^2/M\) near a positive \(a_j\) has logarithm per \(M\), after shrinking the shell, equal to \[ \Phi(a)=H_p(a)+P_\infty(a). \tag{63}\] The same radial calculation gives the upper Laplace bound on compact sets. Set \(\Phi=-\infty\) at a zero coordinate; the divergent radial entropy and bounded \(P_\infty\) justify this convention in upper bounds.

A majorization principle without concavity

Lemma 6 (Weighted Haar-shell majorization). Sort \(a_j/p_j\) in nonincreasing order. If \(b_j\ge0\) satisfy \[ \sum_jb_j=\sum_ja_j,\qquad \sum_{j\le i}b_j\ge\sum_{j\le i}a_j\quad(1\le i<h), \tag{64}\] then \(\Phi(b)\le\Phi(a)\).

Proof. If some \(b_j=0\), then \(\Phi(b)=-\infty\) by convention, so assume all \(b_j>0\). On an interval of length \(p\), let \(A\) be the decreasing density with values \(a_j/p_j\) on successive intervals of lengths \(p_j\), and let \(B^*\) be the decreasing rearrangement of the density list \(b_j/p_j\) with those same lengths. At each original cutpoint, the initial integral of \(B^*\) dominates the original \(b\)-prefix and hence the \(a\)-prefix. Between cutpoints the former cumulative integral is concave and the latter affine. Thus initial integrals of \(B^*\) dominate those of \(A\) everywhere, with equal totals.

The allocation step is the finite-weight Hardy–Littlewood–Pólya convex-order principle. For classical rearrangement background see [10]; we give the required weight-preserving allocation proof here. This credit concerns the allocation, not the Haar subdivision and volume transfer that follow.

Consider allocation tables \(\lambda_{ij}\ge0\) with column sums \(p_j\) and row sums \(p_i\). The entry \(\lambda_{ij}\) is the dimension proportion sent from old block \(j\) to new block \(i\); its squared-mass contribution is \(\lambda_{ij}b_j/p_j\). To show that some table attains \(a\), fix arbitrary prices \(c_i\), and let \(C\) equal \(c_i\) on the interval of length \(p_i\). Write \(C^*\) for its decreasing rearrangement. The support function of the attainable mass vectors assigns the largest source densities to the largest prices. Rearrangement and then cumulative dominance give \[\sum_i c_i a_i=\int_0^p CA \le\int_0^p C^*A \le\int_0^p C^*B^* =\max_{\lambda}\sum_{i,j}c_i\lambda_{ij}b_j/p_j.\] The second inequality is summation by parts; equal total mass permits arbitrary, including negative, prices. Finite-dimensional separation now gives a table attaining \(a\). Set \(T_{ij}=\lambda_{ij}/p_j\); then \[ T\ge0,\qquad \sum_iT_{ij}=1,\qquad Tp=p,\qquad Tb=a. \tag{65}\]

The relation (65) is majorization relative to the block proportions \(p\); see [30]. We realize it geometrically as follows. Randomly subdivide each old \(E_j\) into orthogonal pieces of relative dimensions approaching \(T_{ij}\), and form the new block \(i\) by their sum. The identities \(Tp=p\) and \(Tb=a\) give, respectively, the new dimension proportions and the target squared masses. Dimensions with the correct two margins are obtained by \(o(M)\) adjustments and bounded rounding corrections. The new decomposition has the same Haar marginal as the old one. For each fixed vector in the old \(b\)-shell, squared norms in the pieces concentrate about their prescribed fractions. The failure bound is uniform in the fixed vector; simultaneous concentration for all shell vectors is unnecessary. Thus each such vector lies in a slightly enlarged new \(a\)-shell with failure probability \(\epsilon_M\to0\).

This transfers weighted volume even if the activation weight is concentrated. Conditional on the old decomposition, let \(W_b\) be its weighted shell volume and \(W_{\rm tr}\) the part transferred. Fubini gives \[\mathbb E[W_b-W_{\rm tr}\mid E]\le\epsilon_M W_b.\] With conditional probability at least \(1-\sqrt{\epsilon_M}\), at least \((1-\sqrt{\epsilon_M})W_b\) transfers. The weight and ambient image do not change. Intersect the high-probability shell bounds for the two Haar marginals; they need not be independent. Taking logarithms and shrinking shell widths gives \(\Phi(b)\le\Phi(a)\). ◻

The coefficient integral and triangular entropy

The outer coefficient blocks \(x_j\in\mathbb R^{n_j}\) have Lebesgue measure on balls of radius \(\kappa\sqrt M\). For each \(j\le i<h\), let \(z_{i,j}\) independently have the uniform probability law on the same ball. Put \[y=Gx,\qquad u_i=\sum_{j\le i}G_jz_{i,j},\qquad \overline g=\frac1M\sum_{r=1}^M g(y_r,u_{1,r},\ldots,u_{h-1,r}),\] with the tests and constants from (11). Define the soft coefficient pressure \[ Q_M=\frac1M\log\int dx\left[ \int\exp\left\{\frac{\sum_rf(y_r)-M\lambda\sum_\ell(\bar g_\ell-c_\ell)_+}{\theta}\right\}\,d\mu(z) \right]^\theta. \tag{66}\]

In species \(j\), orthogonalize the coefficient columns divided by \(\sqrt M\) in the order \(x_j,z_{j,j},\ldots,z_{h-1,j}\). They are a Haar orthonormal frame times an upper triangular \(R_j\). Conditional on the first column, the remaining frame has precisely the inner Stiefel law, independently of the triangular parameters. Rotational invariance of the independent coefficient balls makes the inner parameters independent of the first length and direction.

For ball probability measure, the normalized log density of these triangular parameters contributes \(p_j\log(R_{j,\ell\ell}/\kappa)\) for each column. Indeed the density of its readings in previous directions and its new perpendicular radius is proportional to that radius to power \(n_j-\ell\) on a fixed-dimensional half ball. Converting the first ball from probability to Lebesgue measure adds \(p_j\log(\kappa\sqrt{2\pi e/p_j})\) by Stirling. Only inner-column entropy is multiplied by \(\theta\) in (66).

For the normalization, write the coefficient columns as \(\sqrt M\,O_jR_j\), where \(O_j^{\mathsf T}O_j=I\), and set \(X_j=\sqrt{n_j}O_j\). Writing \(G_j=M^{-1/2}\widehat G_j\) gives \[G_j(\sqrt M\,O_jR_j) =\frac{\widehat G_jX_j}{\sqrt{n_j}}R_j.\] Thus for fixed \(R_j\) the row readings are exactly \(\xi_j^{\mathsf T}R_j\) in Theorem 3. Apply that theorem with terminal \(f(y)/\theta\), penalty strength \(d=\lambda/\theta\), and the tests in (11). Multiplication by the outer factor \(\theta\) gives the second line of [eq:Bfunctional]. Ordinary finite-mesh Laplace bounds then give \[ Q_M\longrightarrow\mathcal B_{p,a}(\kappa,\lambda,B,\theta) \quad\hbox{in probability}. \tag{67}\] The finite-mesh Laplace bounds remain valid after taking the \(\theta\)th power of the inner integral. On a bounded operator-norm event all normalized energies are uniformly continuous in the triangular parameters, uniformly in frames; the squared and truncated-squared tests use the uniform \(\ell^2\) bounds on readings. Upper and lower finite-mesh sums are therefore valid. For a finite sum and \(0<\theta<1\), its \(\theta\)th power is between the maximum powered summand and the sum of powered summands; the number of cells contributes no extensive cost. Fixed trial neighborhoods give the lower bound. At fixed \(\theta>0\), zero diagonal faces have arbitrarily negative entropy and may be discarded before refining the mesh.

On the bounded-norm event, the normalized log in (66) is Lipschitz in the standard Gaussian entries with constant \(C(\lambda,\kappa)M^{-1/2}\). This constant is independent of \(B\) and \(\theta\). The derivative of a truncated square has the same \(\ell^2\) bound as that of a square, and the inner division by \(\theta\) cancels the outer power in a supremum change bound. Extend from the good event as above. For every fixed deviation, (67) consequently has an exponential concentration rate independent of \(B,\theta\); the size at which its limiting center is accurate may depend on them.

Upper and lower identification of the shell

The upper bound first removes the witnesses. Write \(b_j=\lVert P_{E_j}y\rVert^2/M\) and \(T=\sum_jb_j\). Since \(u_i\) belongs to the \(i\)th flag subspace, \[\frac{\lVert y-u_i\rVert^2}{M}\ge\sum_{j>i}b_j, \qquad \frac1M\sum_r\min\{y_r^2,B\}\le T.\] The first two test penalties therefore sum to at least \((T-s_a)_++(s_a-T)_+=|T-s_a|\), and each remaining penalty is at least the corresponding tail-norm excess. This bound is independent of the witnesses, so it survives their probability integral and the outer power. Change variables from coefficients to the image and use (62) and the radial upper Laplace bound. Along the chosen subsequence, \[ J_p+\mathcal B_{p,a}\le \sup_{\substack{b_j\ge0\\\sum_jb_j\le C(\kappa)}} \left\{\Phi(b)-\lambda\left( \left|\sum_jb_j-s_a\right|+ \sum_{i<h}\left(\sum_{j>i}b_j-\sum_{j>i}a_j\right)_+ \right)\right\}. \tag{68}\] One may take \(C(\kappa)=C_{\rm op}^2h\kappa^2\). Boundary mass splits with a zero coordinate have divergent negative radial entropy. As \(\lambda\to\infty\), compactness forces the total and prefix constraints of Lemma 6; thus the upper limit of (68) is at most \(\Phi(a)\).

For the reverse bound fix \[ \kappa>2C_{\rm inv}\sqrt{s_a+1}. \tag{69}\] Every \(y\) in a sufficiently narrow desired shell and every one of its flag projections then has a coefficient representative strictly inside the appropriate balls, with slack proportional to \(\sqrt M\). A neighborhood of normalized radius \(\delta\) around each witness makes the projection penalty small. Its inner probability is at least \(\exp\{-C(1+|\log\delta|)M\}\), so its cost in (66) is only \(\theta C(1+|\log\delta|)\).

It remains to control loss of squared norm to the bounded lower-norm test. Fix a small \(\eta>0\). Every vector of normalized squared norm at most \(C(\kappa)\) has all coordinates of absolute value larger than \(b_0(\eta)\) within some set of \(\lfloor\eta M\rfloor\) coordinates. The number of such sets is at most \(\exp(Mh(\eta))\), with \(h(\eta)\to0\). One set therefore captures at least that inverse fraction of the weighted desired shell volume. Rotate those coordinates by an independent Haar matrix. The activation energy changes by at most \(2\lVert f\rVert_\infty\eta M\). To quantify the truncation loss, write \(m=\lfloor\eta M\rfloor\) and let \(v\) be the restriction of a captured vector to the selected coordinates. For a Haar rotation \(U\) on these coordinates, \[\mathbb E_U\frac1M\sum_{r=1}^{m}(Uv)_r^4 =\frac{3\lVert v\rVert^4}{M(m+2)} \le\frac{C_1C(\kappa)^2}{\eta}.\] Choose \(B>b_0(\eta)^2\), so the unrotated entries lose no squared norm. Let \(\mu_{\rm cap}\) be the probability measure obtained by normalizing the old activation weight on the captured shell, whose weighted volume is \(W_{\rm cap}\). Extend \(U\) by the identity outside the selected coordinates and put \[L_B(Uy)=\frac1M\sum_r\big((Uy)_r^2-\min\{(Uy)_r^2,B\}\big).\] The inequality \(z^2-\min\{z^2,B\}\le z^4/B\) and weighted Fubini give, for every \(\tau>0\), \[\mathbb E_U\mu_{\rm cap}\{y:L_B(Uy)>\tau\} \le\frac{C_1C(\kappa)^2}{\eta B\tau}.\] Choose \(B\) so the right side is at most \(1/4\). Markov’s inequality shows that with probability at least \(1/2\), the surviving vectors carry at least \(W_{\rm cap}/2\) of the old weighted volume. Their rotated weighted volume is at least \(e^{-2\lVert f\rVert_\infty\eta M}W_{\rm cap}/2\). Rotating the matrix with the vectors preserves the shell norms, flag distances, and Jacobian.

The useful set can depend on the old matrix. To make the probability argument legitimate, supply independent rotations for every candidate coordinate set. Each rotated matrix has the original Gaussian marginal. With probability bounded away from zero, at least one candidate has coefficient pressure satisfying \[ J_p+Q_M\ge\Phi(a)-\epsilon -2\lVert f\rVert_\infty\eta-h(\eta) -\theta C(1+|\log\delta|), \tag{70}\] where at fixed \(\lambda\) the shell, truncation, Jacobian, and witness errors can be made smaller than \(\epsilon\).

The choices respect exactly the order in (13). Fix \(\kappa\) by (69); then fix any finite \(\lambda\) and a desired error \(\epsilon\). The concentration rate for \(Q_M\) at that error is independent of \(B,\theta\). Choose \(\eta\) so that \(h(\eta)\) is below half this rate and \(2\lVert f\rVert_\infty\eta+h(\eta)<\epsilon\). Choose shell widths and \(\delta\) so their errors times \(\lambda\) are small, then choose a sufficiently large \(B\), independently of \(\theta\). Finally take \(\theta\) small enough that the witness entropy is below \(\epsilon\). A union bound over the candidate rotations and the concentration estimate turns (70) into the corresponding deterministic bound for \(J_p+\mathcal B_{p,a}\), with one additional \(\epsilon\) loss. Thresholds may depend on the fixed \(\lambda\); no uniformity in its later limit is required.

For fixed \(M\), the soft integral increases as \(\theta\downarrow0\) because the inner expression is an increasing \(L^{1/\theta}\) norm. It increases with \(B\) and decreases with \(\lambda\). These monotonicities pass to the deterministic limits. The upper bound and the witness lower bound thus establish the successive limits and \[\Phi(a)=J_p+ \lim_{\lambda\to\infty}\lim_{B\to\infty}\lim_{\theta\downarrow0} \mathcal B_{p,a}(\kappa,\lambda,B,\theta).\] Using (63) proves (13) and identifies every subsequential \(P_\infty\). Equicontinuity extends convergence to compact sets and to zero radii.

Full output dimension and bounded continuous activations

If \(\sum_jp_j=1\), retain, independently in each block, a random subspace of dimension \((1-\eta)n_j+o(M)\). Under the untilted angular law, the fraction of squared norm lost in each block is at most \(2\eta\) with probability tending to one. The same is true of the old Gibbs mass with probability tending to one over the chosen retained subspaces: for each fixed old vector the projection failure probability is uniformly small, so Fubini and Markov apply after arbitrary bounded activation reweighting. Ignore zero-radius blocks.

On this good set, rescale each retained projection to its old radius. Orthogonality gives a field change of norm at most \(C\sqrt{\eta M}\), and hence an energy change at most \(C\sqrt\eta M\). Under the untilted angular law the retained directions are uniform and independent of all split lengths. For an upper bound on the old partition function, restrict to its Gibbs-good mass and integrate the transformed weight; this is at most the new partition function times \(e^{C\sqrt\eta M}\). For the reverse bound, restrict the untilted split lengths to the same typical set and use the pointwise opposite energy comparison; its probability tends to one independently of the retained directions. Therefore \[|P_M^f(a)-P_{M,\eta}^f(a)|\le C\sqrt\eta+o_{\mathbb P}(1).\] The retained system has proportions \((1-\eta)p\) and has already been treated. The bound, followed by \(\eta\downarrow0\), proves (14) and convergence, uniformly on compact radius sets. Only the angular dimensions are retained; no input Dirichlet entropy is changed by this prescription.

For \(f\in C_b(\mathbb R)\) choose bounded Lipschitz approximants with a common bound and uniform convergence on compact intervals. Every angular configuration satisfies \(M^{-1}\sum_ry_r^2=\sum_ja_j\). On a compact radius set, the fraction of entries outside \([-R,R]\) is uniformly \(O(R^{-2})\). Thus compact approximation controls the normalized energy uniformly and transfers all angular conclusions. The same estimate applies to (66) on a bounded operator-norm event, independently of \(\lambda,B,\theta\), so its limiting formula has the same continuity. This completes the proof of Theorem 1.

Radial decomposition and proof of the main theorem

For fixed singular values, bi-orthogonal invariance permits independent Haar left and right singular spaces: insert independent Haar rotations into any singular-value decomposition and use invariance. The input sphere absorbs the right rotation. For a matrix with positive singular values \(s_j\) of multiplicities \(n_j\), let \(\rho_j\) be the input squared norm fraction in block \(j\), and add \(\rho_0\) for a positive-proportion structural right null block if present. Their joint law is Dirichlet with parameters half the corresponding dimensions. An input block of squared norm \(N\rho_j\) produces an output block of squared norm \[Ns_j^2\rho_j=M a_j,\qquad a_j=s_j^2\rho_j/(M/N).\] The logarithmic Dirichlet density rate per input coordinate is \[ I_v(\rho)=\frac12\sum_i v_i\log\frac{\rho_i}{v_i}, \qquad \rho_i>0,\quad \sum_i\rho_i=1. \tag{71}\] This follows directly from the Dirichlet density and Stirling’s formula. At a boundary with a positive \(v_i\) the rate tends to \(-\infty\), while the angular pressure is bounded. On compact interiors, Theorem 1 and continuity give ordinary upper and lower radial Laplace bounds. They yield exactly (4), including the structural null block’s entropy. The argument uses only the limiting positive proportions.

For bounded Lipschitz \(f\), two matrices of the same shape satisfy \[ |F_N(A,f)-F_N(\widetilde A,f)| \le\operatorname{Lip}(f)\sqrt{M/N}\,\lVert A-\widetilde A\rVert_{\rm op}. \tag{72}\] Indeed \(\sum_r|(A-\widetilde A)x|_r\le\sqrt M\lVert A-\widetilde A\rVert_{\rm op}\sqrt N\) for every spherical input. Partition the compact spectral support into finitely many small sets with positive limiting masses, choosing boundaries away from atoms, and positive representatives within the desired mesh error. A finite cover of the support by such positive-mass neighborhoods, together with (3), assigns every singular slot within mesh error plus \(o(1)\); representatives for zero values may be lifted a small positive amount. The associated multiplicities have positive limiting proportions. Equation (72) now compares the original matrix to the finite-block matrix, so its finite-block limits are Cauchy as the mesh shrinks and are independent of the choices.

For \(f\in C_b\), bounded Lipschitz compact approximation again suffices. On a good spectral event \(\lVert A\rVert_{\rm op}\le C\), every input satisfies \[\frac1M\sum_r(Ax)_r^2\le\frac NM C^2.\] Large field entries are uniformly sparse, so the normalized energy difference between \(f\) and its approximant is uniformly small. This proves the claimed formula and convergence for \(C_b\).

If the spectral hypotheses hold only in probability, take an arbitrary subsequence and a further one on which the empirical law and no-outlier error converge almost surely. Condition on its realized singular values. The deterministic-sequence proof applies to every such realization, including the varying integer block dimensions. Bounded conditional probabilities and integration give convergence in probability on this further subsequence. The subsequence criterion proves it on the original sequence. Finally \(|F_N(A,f)|\le(M/N)\lVert f\rVert_\infty\) makes the pressures uniformly bounded, giving convergence in mean. This proves Theorem 2.

Spectral and activation scope

The main theorem has a compact rescaled spectrum, controls outliers on the rectangular singular slots, and assumes a bounded continuous activation. We give several extensions and then show by examples why unrestricted interpretations fail. All limits in this appendix refer to the same formula, with a final approximation limit when stated.

General dimension sequences

Proposition 7. Theorem 2 remains valid for arbitrary deterministic dimensions with \(M/N\to\alpha\in(0,\infty)\).

Proof. The angular and Dirichlet arguments only use limiting proportions. The additional case is \(\alpha=1,N>M\), with \(k=N-M=o(N)\) structural right-null coordinates. Let \(B\) be the restriction of the matrix to the \(M\) visible input coordinates, and let \(Z_M(r)\) be their angular partition function at radius \(r\). Assume first that \(f\) is bounded Lipschitz, and put \(b=\lVert f\rVert_\infty\).

Delete two output terms. The remaining readings have a common kernel of dimension at least two, so the integrand factors through the orthogonal complement of a chosen two-dimensional subspace. The projection of a uniform sphere in \(\mathbb R^M\) onto that \((M-2)\)-dimensional complement is uniform on the ball. Nested balls and restoration of the deleted terms give \[ Z_M(r')(r'/r)^{M-2}\le e^{4b}Z_M(r),\qquad 0<r'\le r. \tag{73}\] The visible squared fraction is \(V\sim\operatorname{Beta}(M/2,k/2)\), independently of direction. For \(k>0\), \[\mathbb EV^{-(M-2)/2}=\frac{\Gamma(N/2)}{\Gamma(M/2)\Gamma(1+k/2)}, \qquad \log\mathbb EV^{-(M-2)/2}=O\big(k\log(eN/k)+\log N\big)=o(N).\] Thus (73) gives \(\mathbb EZ_M(\sqrt{NV})\le e^{4b+o(N)}Z_M(\sqrt N)\). Conversely \(V\to1\) in probability and \[|\log Z_M(\sqrt{Nv})-\log Z_M(\sqrt N)| \le\operatorname{Lip}(f)\sqrt{MN}\lVert B\rVert_{\rm op}(1-\sqrt v).\] Restrict to \(V\ge1-\delta\), then let \(N\to\infty\) and \(\delta\downarrow0\). This gives the opposite comparison up to \(o(N)\) on bounded-norm events. Finally the same Lipschitz estimate compares radii \(\sqrt N\) and \(\sqrt M\), with logarithmic error \(o(N)\) because \(M/N\to1\). Division by \(N\), instead of \(M\), also changes the normalized pressure by \(o(1)\). The square model for \(B\) is therefore the required comparison. If \(k=0\) no comparison is needed. Compact approximation extends the result to \(C_b\), spectral good events handle random hypotheses, and boundedness gives mean convergence. ◻

Negligible spectral groups inside the edge interval

A block of \(o(N)\) directions can still carry a positive fraction of the input norm under the Gibbs weight. The next lemma records its entropy cost. We will reproduce that cost by reserving one or two directions in macroscopic spectral blocks.

Lemma 8 (Entropy with a negligible block). Partition the input coordinates into finitely many blocks with \(n_i/N\to v_i>0\) and an exceptional block of dimension \(r=o(N)\), so \(\sum_i v_i=1\). For macroscopic fractions \(\rho_i>0\) with \(\sum_i\rho_i\le1\), the upper logarithmic shell rate is \[ I_v(\rho)=\frac12\sum_i v_i\log(\rho_i/v_i). \tag{74}\] It holds also for shells touching zero exceptional mass; when \(r=0\), shells with \(\sum_i\rho_i<1\) are empty. If the exceptional coordinates are replaced by a fixed positive number of reserved directions, every strictly positive prescribed split has this same rate, with boundary splits obtained by approximation.

Proof. For \(r>0\) the total macroscopic fraction \(q\) is \(\operatorname{Beta}((N-r)/2,r/2)\). The logarithm of its normalizing constant is \(o(N)\), and its extensive density factor away from \(q=0\) is \(q^{N/2+o(N)}\). Near \(q=1\), integrate \((1-q)^{r/2-1}\) rather than bounding it pointwise; its integral is \(2/r\), with subextensive logarithm. The upper rate is therefore \(\frac12\log q\), with the matching lower rate on interior intervals and at \(q=1\) by \(q\to1\) in probability. Conditional Dirichlet fractions \(\rho_i/q\) give the remaining rate \(\frac12\sum_i v_i\log((\rho_i/q)/v_i)\); the sum is (74). When \(r=0\) only \(q=1\) is feasible, and the ordinary Dirichlet statement suffices for the upper bound.

For finitely many reserved directions, each parameter is \(1/2\). Their density factors and additional gamma normalizations contribute only \(o(N)\) on fixed positive shells. Their extensive cost is already in (74): writing \(q=\sum_i\rho_i\) exposes \(\frac12\log q\). There is no further extensive term. ◻

Proposition 9 (Edge-interval control). In Theorem 2, or Proposition 7, let \(s_-=\min\operatorname{supp}\nu\) and \(s_+=\max\operatorname{supp}\nu\). The no-outlier assumption may be weakened to \[ \max_{r\le\min(M,N)}\operatorname{dist}(s_{N,r},[s_-,s_+])\longrightarrow0 \quad\hbox{in probability}. \tag{75}\] If \(\alpha<1\), the interval may instead be \([0,s_+]\), using the macroscopic structural right null space at its lower endpoint. Forced left zeros do not provide such an endpoint.

Proof. Work first with bounded Lipschitz \(f\) and a fixed finite spectral discretization. If both edges are zero, the operator norm tends to zero and the limit is \(\alpha f(0)\). Otherwise take positive representatives \(s_j\) for the non-null macroscopic groups, with multiplicities \(n_j/M\to p_j>0\), and add the structural input-null group when extensive. Representatives near the edges can be chosen arbitrarily close to them. Up to a further arbitrarily small operator error, all other slots lie between the smallest and largest available macroscopic representatives (including the allowed structural zero); their total number is \(o(M)\). At square aspect first remove a subextensive structural right null space by the radius comparison in the proof of Proposition 7; that comparison requires only a bounded operator norm.

The lower bound: small exceptional mass. Call the \(o(M)\) slots exceptional and compare to the model assigning them all to the top macroscopic group. The latter has the main-theorem limit. Restricting the original model to an exceptional squared input fraction at most \(\epsilon\) has untilted probability tending to one; discarding that field changes pressure by \(O(\sqrt\epsilon)\). The remaining proportions and radial law give the main value as a lower bound, by the angular theorem and Dirichlet Laplace bounds, followed by \(\epsilon\downarrow0\).

The upper bound: one additional Haar direction. For the upper comparison, fix \(\eta>0\) and retain \(l_j=\lfloor(1-\eta)n_j\rfloor\) random dimensions in every non-null macroscopic output block. For every old configuration the probability that any block loses more than \(2\eta\) of its squared norm tends to zero uniformly. Fubini under the old Gibbs measure and Markov show that this restriction retains Gibbs mass tending to one for a typical choice of subspaces. Rescale retained components to their original radii. The field changes by \(O(\sqrt{\eta M})\), hence pressure by \(O(\sqrt\eta)\). Retained directions are uniform independently of split lengths. The reverse comparison used below follows by restricting untilted split lengths to that same typical set.

For macroscopic fractions \(\rho_i>0\) with \(\sum_i\rho_i\le1\), set \(a_j=s_j^2(N/M)\rho_j\). Let \(E_j^\eta\) be the retained Haar blocks, and let \(e\) be a uniform unit vector orthogonal to them, jointly Haar with them. Define \[\begin{align*} \Psi_M^\eta(a,t;E^\eta,e) &=\frac1M\log\int \exp\left\{\sum_{r=1}^M f\left(\sum_jy_{j,r}+t\sqrt M\,e_r\right)\right\} \prod_jd\sigma_{E_j^\eta,Ma_j}(y_j),\\ D_M^\eta(a,t)&=\mathbb E\Psi_M^\eta(a,t;E^\eta,e). \end{align*}\] The companion direction \(e\) is part of the quenched Haar data; only the \(y_j\) are integrated in the partition function. The total dimension of the retained spaces plus one stays below \(M\) by a positive fraction. The Gaussian-span concentration proof therefore gives exponential concentration about \(D_M^\eta\), uniformly at fixed deviations for bounded \(a,t\). Coupling gives uniform Lipschitz continuity in \((\sqrt{a_j})_j,t\). No limit of \(D_M^\eta\) is assumed.

For a fixed normalized exceptional coefficient vector its output has exactly this companion law at its specified length. A fixed-mesh net of the exceptional ball has size \((1+2/\delta)^{o(M)}=\exp(o(M))\). A union bound, finite radius nets, and the Lipschitz estimate give simultaneous concentration for all exceptional coefficients and bounded radii. Their output length has the form \[t^2=(N/M)(1-\textstyle\sum_i\rho_i)\tau, \qquad s_{\rm lo}^2\le\tau\le s_{\rm hi}^2.\] Lemma 8 and a radial upper bound consequently yield \[\begin{align*} F_N(A,f)\le C\sqrt\eta+o_{\mathbb P}(1)+ \sup\bigg\{&\frac12\sum_i v_i\log\frac{\rho_i}{v_i} +\frac MN D_M^\eta(a,t):\\ &\rho_i>0,\quad\sum_i\rho_i\le1,\quad a_j=s_j^2(N/M)\rho_j,\\ &t^2=(N/M)(1-\textstyle\sum_i\rho_i)\tau,\quad s_{\rm lo}^2\le\tau\le s_{\rm hi}^2\bigg\}. \end{align*}\] Fractions tending to zero in a macroscopic block can be excluded because their entropy tends to \(-\infty\) and the angular pressure is bounded. On the remaining compact domain, equicontinuity justifies the upper mesh bound with this size-dependent objective.

Matching the upper bound by endpoint reservations. The reassigned model has a matching lower bound for this supremum, up to \(C\sqrt\eta+o_{\mathbb P}(1)\). Fix a deterministic near-maximizing trial and reserve one input direction from each endpoint group. Write \(\gamma=1-\sum_i\rho_i\) and choose \(\gamma_{\rm lo},\gamma_{\rm hi}\ge0\) with \[\gamma_{\rm lo}+\gamma_{\rm hi}=\gamma,\qquad \gamma_{\rm lo}s_{\rm lo}^2+\gamma_{\rm hi}s_{\rm hi}^2=\gamma\tau.\] Allocate these squared fractions to the two reserved directions and \(\rho_i\) to the remaining macroscopic blocks. Their combined output is a single Haar companion field of normalized squared length \(t^2\) relative to the other blocks. Prescribing signs costs only a constant factor. Lemma 8 supplies precisely the entropy in [eq:edgeupper], including the extensive cost of concentrating norm in two directions.

The reassignment changes \(o(M)\) dimensions, and reservation removes at most two; the extensive slack \(\eta n_j\) lets us retain the same \(l_j\) dimensions in each remaining block. Restricting to typical projected lengths supplies the same \(D_M^\eta(a,t)\) within \(O(\sqrt\eta)\). Here concentration is needed only for the chosen trial. If endpoints coincide, use one direction or two from the same group. A structural zero endpoint uses a direction in the extensive input-null group, which consumes input norm without contributing output. Boundary reserved fractions follow by positive-shell approximation and continuity. Comparing the two models and sending \(\eta\downarrow0\) proves equality of their limits. Finally shrink the spectral mesh, approximate \(C_b\) on compacts, and condition on good spectral sequences as in the main proof. Boundedness gives mean convergence. ◻

Activation extensions

Proposition 10. Under either compact spectral condition above, the following extensions hold.

  1. If \(f\) is continuous and \(|f(z)|/(1+z^2)\to0\) as \(|z|\to\infty\), the pressure converges in probability to the final bounded-cutoff limit of the main formula. Convergence in mean also holds, for example, if the spectra are uniformly bounded by a deterministic constant.

  2. If \(f\) is bounded Borel and \(\nu\ne\delta_0\), the pressure converges in probability and in mean. Its value is a final uniformly bounded continuous approximation limit, with approximation in Lebesgue measure on compact intervals as in the proof.

Proof. For (i), let \(\chi_R\) be a continuous cutoff equal to one on \([-R,R]\) and zero outside \([-2R,2R]\), and put \(f_R=\chi_Rf\). With \(\epsilon_R=\sup_{|z|>R}|f(z)|/(1+z^2)\to0\), every input gives \[\frac1N\sum_r|f((Ax)_r)-f_R((Ax)_r)| \le\epsilon_R(M/N+\lVert A\rVert_{\rm op}^2).\] This also bounds the difference of the pressures. Good spectral events make the cutoff error uniformly small, so the bounded-activation limits are Cauchy and imply probability convergence. A deterministic uniform operator bound also bounds the pressures uniformly, proving the stated mean conclusion. Spectral hypotheses only in probability do not imply this uniform integrability for an unbounded \(f\).

For (ii), put \(b=\lVert f\rVert_\infty\), taking \(b>0\). Since \(\nu\ne\delta_0\), good spectral events have \(\lVert A\rVert_{\rm op}\le C\) and at least \(\kappa N\) input singular directions with singular value at least \(s_*>0\), where \(0<\kappa<1/2\). Write \(\lambda_x=\lVert Ax\rVert/\sqrt M\). It has a uniform upper bound \(\Lambda\), while its small-length input mass has arbitrarily large exponential cost: for every \(L>0\), sufficiently small \(\lambda_0>0\) gives \[ \sigma_N\{\lambda_x<\lambda_0\}\le e^{-LM} \tag{76}\] for all large good spectra. Indeed the event forces the squared projection onto the indicated \(\kappa N\) directions below \(C'\lambda_0^2N\), and its beta density has rate tending to \(-\infty\) as \(\lambda_0\downarrow0\).

Conditional on input coefficients and singular values, Haar left rotation makes the field uniform on the sphere of radius \(\lambda_x\sqrt M\). The density of any \(k<M-2\) entries is \[\frac{\Gamma(M/2)}{\Gamma((M-k)/2)(\pi M\lambda_x^2)^{k/2}} \left(1-\frac{|z|^2}{M\lambda_x^2}\right)_+^{(M-k-2)/2}.\] For \(k\le M/2\) and \(\lambda_x\ge\lambda_0\) this is at most \((C_0/\lambda_0)^k\). If a measurable set \(B\subset\mathbb R\) has Lebesgue measure at most \(\delta\), then with \(k=\lceil\gamma M\rceil\), \(0<\gamma<1/2\), \[ \mathbb P\{\#\{r:(Ax)_r\in B\}\ge k\} \le\binom{M}{k}(C_0\delta/\lambda_0)^k. \tag{77}\] For fixed \(\gamma,\lambda_0\), this decay rate is arbitrarily large when \(\delta\) is small.

Choose \(R\) so that \(\Lambda^2/R^2\le\gamma\). Every input then has at most \(\gamma M\) field entries outside \([-R,R]\). Approximation in Lebesgue measure gives a continuous \(g\), \(|g|\le b\), for which \(B=\{z\in[-R,R]:|f(z)-g(z)|>\epsilon\}\) has arbitrarily small measure. One can obtain it by simple functions, regularity of Lebesgue measure, continuous approximations of their finitely many indicator sets, and clipping to \([-b,b]\). Choose \(\lambda_0\) and then the approximation accuracy so that (76)–(77) have any prescribed exponential rate. Fubini and Markov show, with probability tending to one over the Haar orientation, that the untilted mass of exceptional inputs is at most \(e^{-B_0M}\) for a fixed \(B_0>2b\) as large as desired.

On the other inputs the average absolute energy difference is at most \(\epsilon+4b\gamma\). The exceptional partition mass is at most \(e^{(b-B_0)M}\), whereas each full partition function is at least \(e^{-bM}\); it is negligible. Thus \[|F_N(A,f)-F_N(A,g)|\le(M/N)(\epsilon+4b\gamma)+o_{\mathbb P}(1).\] A sequence with \(\epsilon,\gamma\downarrow0\) makes the continuous-activation formula values Cauchy and independent of the approximants, proving probability convergence. The uniform bound \((M/N)b\) gives mean convergence. ◻

Examples delimiting the hypotheses

Forced output zeros can hide a small singular outlier.

Take \(M=1000N\), \(\alpha=1000\), and independent Haar singular spaces. On alternate sizes let all \(N\) singular slots equal \(\sqrt\alpha\); on the other sizes replace one by zero. Take \(f(z)=-c\min(z^2,1)\) with \(c\ge1\).

On full-rank sizes write \(A=\sqrt\alpha U\) with \(U\) a Haar orthonormal \(M\times N\) frame, and for a unit input \(u\) put \[Q(u)=\frac1M\sum_{r=1}^M\min\{M(Uu)_r^2,1\}.\] The uniform lower bound below is a Kashin-type random-subspace spreading estimate; see [12] and the norm-comparison restatement [3]. The explicit clipped-square constant and probability estimate below are proved directly here.

For fixed \(u\), \(\sqrt M Uu\) has law \(G/\sqrt S\), where \(G_r\) are iid standard normals and \(S=M^{-1}\sum_rG_r^2\). Since \(\mathbb P(1\le|G_r|\le2)>1/10\), Chernoff bounds give fewer than \(M/20\) such entries with probability at most \(e^{-M/80}\), and \(\mathbb P(S>2)\le e^{-(1-\log2)M/2}\). Thus \(\mathbb P(Q(u)<1/40)\le2e^{-M/80}\). The function \(Q\) is \(2\)-Lipschitz on the input unit sphere. A \(1/160\)-net has at most \(321^N\) points, so \[\mathbb P\{\inf_uQ(u)<1/80\}\le2e^{-N(\alpha/80-\log321)}\longrightarrow0.\] The full-rank pressure is consequently at most \(-c\alpha/80=-12.5c\) with high probability.

With one right-null direction, the squared fraction \(V\) outside the kernel is \(\operatorname{Beta}((N-1)/2,1/2)\), and for every orientation the average of \(\min((Ax)_r^2,1)\) is at most \(\lVert Ax\rVert^2/M=V\). For fixed \(v\in(0,1)\) the event \(V\le v\) has log probability per \(N\) tending to \(\frac12\log v\). Taking \(v=1/160\) gives \[\liminf F_N(A,f)\ge-\tfrac12\log160-6.25c>-12.5c.\] The two subsequences are separated. Nevertheless the output eigenvalue measure tends to \((1-\alpha^{-1})\delta_0+\alpha^{-1}\delta_\alpha\), and every output eigenvalue lies in this support exactly. The rectangular-slot singular measure tends to \(\delta_{\sqrt\alpha}\) with no upper outlier. Output-side support control including forced left zeros, or upper-only outlier control, is therefore insufficient.

Unbounded limiting support under weak convergence.

Let \(M=N\), \(d_k=\exp(8^k)\), and \(\nu=\sum_{k\ge1}2^{-k}\delta_{d_k}\). For large \(N\) choose \(K=\lfloor\log_2\log N\rfloor\), multiplicities \(n_j=\lfloor N2^{-j}\rfloor\) for \(j<K\), and \(r=n_K=N-\sum_{j<K}n_j\). Use these singular values with Haar orientations. Their measures converge weakly to \(\nu\), every value is in its support, and eventually \(2^{-K}N/2\le r\le N/2\). Both parities of \(K\) occur infinitely often.

The intervals \([d_k^{1/2},d_k^2]\) are disjoint because \(d_{k+1}^{1/2}=d_k^4\). Choose a continuous \(f:\mathbb R\to[0,1]\) equal to \(c_k\in\{0,1\}\) on their absolute-value versions, alternating with \(k\), and interpolate in the gaps. For a uniform unit vector \(\omega\in\mathbb R^N\) and \(r\le N-2\), integration of its projection density gives \[ \mathbb P(\lVert P_r\omega\rVert\le t) \le\frac{\Gamma(N/2)t^r}{\Gamma(r/2+1)\Gamma((N-r)/2)} \le(C\sqrt{N/r}\,t)^r. \tag{78}\] Put \(d=d_K\) and \(\lambda_x=\lVert Ax\rVert/\sqrt N\). Always \(\lambda_x\le d\). If \(\lambda_x<d^{3/4}\), the normalized input projection norm in the top block is below \(d^{-1/4}\) (its squared fraction is below \(d^{-1/2}\)), so (78) gives \[\sigma_N\{\lambda_x<d^{3/4}\} \le(C\sqrt{N/r}\,d^{-1/4})^r\le e^{-NL_K},\qquad L_K\longrightarrow\infty.\] For example \(L_K=4^K/8-\log C-1\) is valid for all large \(K\).

Conditionally on input coefficients, the output direction after Haar rotation is uniform. Fix \(0<\epsilon<1/4\) and \(m_N=\lceil\epsilon N\rceil\). At \(\lambda_x\ge d^{3/4}\), requiring a specified set of \(m_N\) entries to have absolute value below \(d^{1/2}\) forces its unit-direction projection norm below \(\sqrt{m_N/N}\,d^{-1/4}\). Equation (78) and a union bound give \[\mathbb P\{\#\{r:|(Ax)_r|<d^{1/2}\}\ge m_N\} \le2^N(Cd^{-1/4})^{m_N}\le e^{-NH_K},\qquad H_K\longrightarrow\infty.\] Deterministically at most \(Nd^{-2}\) entries exceed \(d^2\), by the total norm bound. Therefore the expected untilted mass of inputs for which the average \(f\) differs from \(c_K\) by more than \(\epsilon+d^{-2}\) is at most \(2e^{-NT_K}\) with \(T_K\to\infty\). Fubini and Markov make its quenched mass at most \(e^{-NT_K/2}\) with probability tending to one. It is negligible even under the maximal tilt \(e^N\). Letting \(\epsilon\downarrow0\), the pressures tend to zero on one parity and one on the other, in probability and in mean. Mere weak convergence to an unbounded law, even with every singular value in its support, cannot replace compact spectral size control.

A bounded Borel activation at a zero limiting spectrum.

Take \(M=N\), with \(A_N=0\) on even sizes and \(A_N=s_NO_N\) on odd sizes, where \(s_N>0\), \(s_N\to0\), and \(O_N\) is Haar orthogonal. For \(f(z)=\mathbf 1_{\{z\ne0\}}\) the pressure is zero on even sizes and one on odd sizes. On odd sizes every coordinate is nonzero for almost every input, since the exceptional inputs form a finite union of great subspheres. Both spectral subsequences converge to \(\delta_0\) without outliers. This explains the exclusion in Proposition 10(ii).

Continuous activation without a growth bound.

Again take \(M=N\), with \(A_N=0\) on even sizes and \(A_N=N^{-1/4}O_N\) on odd sizes, but now take \(f(z)=z^8\). The singular measure still converges to \(\delta_0\) without outliers. On odd sizes rotational invariance of the input gives, for \(x=\sqrt N\,\omega\), \[\sum_r(A_Nx)_r^8=N^2\sum_r\omega_r^8.\] On the cap \(\omega_1^2\in[1/4,1/2]\), this is at least \(N^2/256\). Its beta density gives cap probability at least \(\frac18\,2^{-(N-3)/2}\) for \(N\ge3\). Thus \[F_N(A_N,f)\ge N/256-\tfrac12\log2-\frac{3\log2}{2N}\longrightarrow+\infty\] on odd sizes, while it is zero on even sizes. Continuity alone does not suffice in the compact no-outlier regime.

  1. 99 A. Abbara, B. Aubin, F. Krzakala, and L. Zdeborová, Rademacher complexity and spin glasses: a link between the replica and statistical theories of learning, in Proceedings of the First Mathematical and Scientific Machine Learning Conference, Proc. Mach. Learn. Res. 107 (2020). PMLR article page.
  2. M. Aizenman, R. Sims, and S. L. Starr, Extended variational principle for the Sherrington–Kirkpatrick spin-glass model, Phys. Rev. B 68 (2003), 214403. doi:10.1103/PhysRevB.68.214403.
  3. M. Bayati and A. Montanari, The LASSO risk for gaussian matrices. Author manuscript, Appendix F, Theorem F.1. https://web.stanford.edu/~bayati/papers/lasso.pdf.
  4. E. Bolthausen and A.-S. Sznitman, On Ruelle’s probability cascades and an abstract cavity method, Comm. Math. Phys. 197 (1998), no. 2, 247–276. doi:10.1007/s002200050450.
  5. H.-B. Chen and J.-C. Mourrat, On the free energy of vector spin glasses with nonconvex interactions, Probab. Math. Phys. 6 (2025), no. 1, 1–80. doi:10.2140/pmp.2025.6.1. arXiv:2311.08980v3.
  6. E. Gardner, The space of interactions in neural network models, J. Phys. A: Math. Gen. 21 (1988), no. 1, 257–270. doi:10.1088/0305-4470/21/1/030.
  7. E. Gardner and B. Derrida, Optimal storage properties of neural network models, J. Phys. A: Math. Gen. 21 (1988), no. 1, 271–284. doi:10.1088/0305-4470/21/1/031.
  8. S. Ghirlanda and F. Guerra, General properties of overlap probability distributions in disordered spin systems. Towards Parisi ultrametricity, J. Phys. A: Math. Gen. 31 (1998), no. 46, 9149–9155. doi:10.1088/0305-4470/31/46/006.
  9. G. Györgyi and P. Reimann, Beyond storage capacity in a single model neuron: continuous replica symmetry breaking, J. Stat. Phys. 101 (2000), 679–702. arXiv:cond-mat/0003360.
  10. G. H. Hardy, J. E. Littlewood, and G. Pólya, Inequalities, second edition, Cambridge University Press, 1952.
  11. Y. Kabashima, Inference from correlated patterns: a unified theory for perceptron learning and linear vector channels, J. Phys.: Conf. Ser. 95 (2008), 012001. doi:10.1088/1742-6596/95/1/012001.
  12. B. S. Kashin, Diameters of some finite-dimensional sets and classes of smooth functions, Math. USSR-Izv. 11 (1977), no. 2, 317–333. doi:10.1070/IM1977v011n02ABEH001719.
  13. J. Ko, The Crisanti–Sommers formula for spherical spin glasses with vector spins, preprint, 2019. arXiv:1911.04355.
  14. F. Mezzadri, How to generate random matrices from the classical compact groups, Notices Amer. Math. Soc. 54 (2007), 592–604. arXiv:math-ph/0609050.
  15. A. Montanari and K. Zhou, Which exceptional low-dimensional projections of a Gaussian point cloud can be found in polynomial time?, preprint, 2024, version 3, October 17, 2025. arXiv:2406.02970v3.
  16. J.-C. Mourrat, Nonconvex interactions in mean-field spin glasses, Probab. Math. Phys. 2 (2021), no. 2, 281–339. doi:10.2140/pmp.2021.2.281.
  17. J.-C. Mourrat, The Parisi formula is a Hamilton–Jacobi equation in Wasserstein space, Canad. J. Math. 74 (2022), no. 3, 607–629. doi:10.4153/S0008414X21000031.
  18. J.-C. Mourrat, Free energy upper bound for mean-field vector spin glasses, Ann. Inst. Henri Poincaré Probab. Stat. 59 (2023), no. 3, 1143–1182. doi:10.1214/22-AIHP1292.
  19. J.-C. Mourrat and D. Panchenko, Extending the Parisi formula along a Hamilton–Jacobi equation, Electron. J. Probab. 25 (2020), no. 23, 1–17. doi:10.1214/20-EJP432.
  20. OpenAI, The free energy of the spherical random perceptron, OpenAI Math Release preprint OAI:The-free-energy-of-the-spherical-random-perceptron-September-24-2026, 2026.
  21. D. Panchenko, A connection between the Ghirlanda–Guerra identities and ultrametricity, Ann. Probab. 38 (2010), no. 1, 327–347. doi:10.1214/09-AOP484.
  22. D. Panchenko, On the Dovbysh–Sudakov representation result, Electron. Commun. Probab. 15 (2010), no. 31, 330–338. doi:10.1214/ECP.v15-1562.
  23. D. Panchenko, The Parisi ultrametricity conjecture, Ann. of Math. (2) 177 (2013), no. 1, 383–393. doi:10.4007/annals.2013.177.1.8.
  24. D. Panchenko, Free energy in the mixed \(p\)-spin models with vector spins, Ann. Probab. 46 (2018), no. 2, 865–896. doi:10.1214/17-AOP1194.
  25. D. Panchenko, Free energy in the Potts spin glass, Ann. Probab. 46 (2018), no. 2, 829–864. doi:10.1214/17-AOP1193. arXiv:1512.00370v3.
  26. D. Panchenko and M. Talagrand, On one property of Derrida–Ruelle cascades, C. R. Math. Acad. Sci. Paris 345 (2007), no. 11, 653–656. doi:10.1016/j.crma.2007.10.035.
  27. D. Ruelle, A mathematical reformulation of Derrida’s REM and GREM, Comm. Math. Phys. 108 (1987), no. 2, 225–239. doi:10.1007/BF01210613.
  28. M. Shcherbina and B. Tirozzi, Rigorous solution of the Gardner problem, Comm. Math. Phys. 234 (2003), no. 3, 383–422. doi:10.1007/s00220-002-0783-3.
  29. T. Shinzato and Y. Kabashima, Perceptron capacity revisited: classification ability for correlated patterns, J. Phys. A: Math. Theor. 41 (2008), 324013. doi:10.1088/1751-8113/41/32/324013.
  30. F. vom Ende and G. Dirr, The \(d\)-majorization polytope, Linear Algebra Appl. 649 (2022), 152–185. doi:10.1016/j.laa.2022.05.005. arXiv:1911.01061v5.
LEVEL 3 COMPLETE!
You read 18,511 words and 1,255 formulas. Your math teacher would be proud.
Converted from the LaTeX source. Something look off? The original PDF is the real thing.

Cool Links: openai/math   Lean   Mathlib   arXiv   the real Coolmath Games