A D V E R T |
I S E M E N T |
| Math Sites: lean ages 13-∞ readme referees parents | >>> MAITH GAMES <<< | all 372 compute stand |
|
LEVEL 2 OF 6 · Memory–sample lower bounds for noiseless Gaussian regression
Posterior replicas and conditional information in Gaussian regression
expertly designed by an internal OpenAI model · released 2026-09-27
· original PDF
IntroductionConsider an unknown unit vector \(s\in S^{d-1}\) and a stream of independent Gaussian rows \(X_j\sim N(0,I_d)\) with exact labels \(Y_j=\langle X_j,s\rangle\). If all raw observations can be retained, \(d\) rows determine \(s\) almost surely, because their matrix is invertible. A learner with finite persistent memory must instead replace each observed pair by an update of its state. The question is how this restriction changes the number of observations needed for a prescribed angular accuracy. Exact labels make the question particularly delicate: an individual real label has no prescribed bit precision. We study this problem through a conditional information bound. A finite message is formed from one block of exact observations, and a separate, independent projection of the signal is then revealed to the analyst. We bound the information that remains in the message after this revelation. The projection is an analysis variable throughout; it is not available to the learner or to the rule forming the message. Conditional information from one blockLet \(\sigma_d\) denote uniform probability on \(S^{d-1}\). In a block experiment, the signal \(S\) has law \(p=f\sigma_d\), where \(f\) is a bounded nonnegative Borel density. Independently draw a standard Gaussian matrix \(A\), and observe \((A,AS)\). A finite-valued message \(W\) is drawn from a measurable kernel of these data, conditionally independently of \(S\). Finally draw a standard Gaussian matrix \(B\) independently of the entire experiment producing \((S,A,AS,W)\). Our principal quantity is \[I(S;W\mid B,BS).\] All logarithms and entropies below are in nats. Theorem 1 (Gaussian replica block bound). For all sufficiently large \(d\), set \[k=2\lfloor d/16\rfloor,\qquad m=k,\qquad t=m+1,\qquad \ell=4k.\] Suppose \(p=f\sigma_d\) is a probability with \(0\le f\le L<\infty\). In the block experiment above, let \(A\) have \(k\) rows and let \(B\) have \(\ell\) rows. There is an absolute constant \(C\) such that \[ I(S;W\mid B,BS) \le \frac{H(W)}{t}+Cd+C\log(2+\log L). \tag{1}\] The bound is uniform over the prior, the density bound \(L\), and the message kernel. The labels used to produce \(W\) are the exact values \(AS\). The entropy cost is divided by a number \(t\) proportional to \(d\). The weak dependence on \(L\) matters when the theorem is applied to a streaming state: conditioning a uniform signal on a state value \(v\) of probability \(p_v>0\) gives a density at most \(1/p_v\). Rare states may therefore have large density bounds, even when their average information cost is small. The proof couples several possible signals with the same observed data. Given \((A,AS)\), draw \(S_1,\ldots,S_t\) independently from the posterior, independently also of \(W\) given the data. Each pair \((S_i,W)\) has the original signal–message law. The replicas are dependent after the data are forgotten: they satisfy \(AS_1=\cdots=AS_t\) for the common hidden matrix. Write \(K\) for the relative entropy of their joint law from \(p^{\otimes t}\), and \(K'\) for the corresponding relative entropy of \((BS_1,\ldots,BS_t)\) from the product of its marginals, conditional on \(B\). These are their total correlations before and after projection. Lemma 13 gives \[tI(S;W\mid B,BS)\le H(W)+K-K'.\] Thus the proof must bound the loss of correlation. Bounding \(K\) alone would discard the geometric dependence still visible to the analyst. The exact-label geometry supplies the comparison. For the difference matrix \(D=[S_2-S_1,\ldots,S_t-S_1]\), the replica likelihood contains \(\det(D^{\mathsf T}D)^{-k/2}\). Theorem 15 identifies the underlying finite measure, including its normalization and density values on the constraint. An observable test of the projected replicas supplies a matching inverse-volume term. Section 4 then bounds the remaining projection-density ratio by a dyadic decomposition, proving Theorem 1. A streaming consequenceThe learner receives the fresh pairs \((X_j,Y_j)\) in order, cannot choose the rows, and cannot revisit them. At each index it has at most \(2^M\) persistent states. Its transition and stopping rules may be arbitrary measurable randomized functions of the current state and current pair; all information retained for later indices must be in the state. It stops at some \(\tau\in\{0,\ldots,T\}\) for a prescribed deterministic horizon \(T\). Its unit-vector output uses only the terminal state, the rule at index \(\tau\), and fresh randomness. Thus the final observation affects the output only through that state. Rules may depend on the index, dimension, and target accuracy. Shared randomness and initialization are jointly independent of the signal and samples, and fresh random choices are independent before use. The entire experiment is required to be jointly measurable. Definition 3 records the full model and the completed-measurability convention. Theorem 2 (Uniform-prior sample lower bound). Let \(M(d)=o(d^2)\) be nonnegative integer-valued, and let \(0<\epsilon(d)\le1/10\). Suppose a family of learners in Definition 3 satisfies \[\mathbb P\{\arccos\langle\widehat S,S\rangle\le\epsilon(d)\}\ge\frac23 \qquad\text{when }S\sim\sigma_d,\] where the probability includes the signal, samples, and learner randomness. Then there is an absolute \(c>0\) such that \[T(d)\ge c\,d\log\frac1{\epsilon(d)}\] for all sufficiently large \(d\). The eventual threshold may depend on the memory sequence. The theorem assumes success averaged over the uniform sphere. A guarantee of success at least \(2/3\) for every fixed signal implies that hypothesis by integration, so the same lower bound holds for such learners. Section 6 proves the result by conditioning on the current state, applying a block estimate, and averaging its prior cost. The independent projection leaves a residual sphere of dimension proportional to \(d\); accurate recovery still requires order \(d\log(1/\epsilon)\) conditional information on that sphere. Information and geometric comparisonsThe total-correlation argument uses Watanabe’s measure of multivariate dependence (Watanabe 1960) and its general probability-space chain rules (Austin 2020, sec. 3.3 and 6). Posterior resampling preserves the joint law with the observed data, as in the general Nishimori identity recorded by Lelarge and Miolane (Lelarge and Miolane 2019, Proposition 16). Here its role is to preserve the signal–message pair law while introducing dependence among replicas. The projected tests use the relative-entropy variational principle; a convenient statement is (Dupuis and Mao 2022, Equations (1.1)–(1.2)). The equal-label measure is related to coarea (Federer 1959, Theorem 3.1) and to affine simplex-volume changes of variables (Drury 1984, Lemma 1). We prove the Gaussian probability normalization and the required evaluations at constrained labels directly. Thus these classical tools provide the background for the exact-fiber calculation, rather than an unproved formula for its particular probability law. The subsequent comparisons retain different information from that geometry. For a prior density at most \(e^b\), Theorem 21 uses \(q=\lfloor d/8\rfloor\) replicas and gives the alternative cost \(H(W)/q+C(d+b/d)\). Its test records successive simplex heights in distance bins and compares the observed bin with a mixture over fixed bins. The direct linear dependence on \(b\) makes averaging over state values immediate. The other routes use different comparison laws: Table 1 records their geometric inputs and the state information fixed during iteration. Each route verifies its own normalization and exceptional sets. They offer distinct tests and ways to average the prior cost, rather than a succession of stronger asymptotic sample bounds.
One complete route occupies Sections 2–6. Section 2 supplies the learner reductions and the residual-sphere information requirement. Section 3 establishes the replica inequality, exact equal-label measure, and logarithmic integrability. Sections 4 and 5 prove the two block estimates, and Section 6 gives their current-state iterations. The final three sections develop the other geometric comparisons and their streaming applications; the synthetic and incidence iterations condition on full state histories. All applications use the same exact-row, finite-state, deterministic-horizon learner model. Finite states and the information required at the endpointThe block estimates will bound information in a state after an independent projection is known. After specifying the learner model, we show why an accurate state must still carry order \(d\log(1/\epsilon)\) information in that experiment. We also record the reductions for randomized rules, rare states, and bounded stopping, so that later applications need only specify their block dimensions and their way of charging the clock. Finite-state learnersDefinition 3 (Finite-state noiseless regression). An unknown unit vector \(s\in S^{d-1}\) is fixed before sampling. At index \(j\ge1\) the learner receives \[X_j\sim N(0,I_d),\qquad Y_j=\langle X_j,s\rangle,\] where the \(X_j\) are independent. It cannot choose a row or revisit an earlier one. At each index its persistent state has at most \(2^M\) values. Its transition and stopping rule may be any measurable randomized kernel of the current state and the entire current pair \((X_j,Y_j)\). Computation within a transition is unrestricted, and every piece of information retained for later indices must be in the finite state. The learner stops at an index \(\tau\in\{0,\ldots,T\}\), where the integer \(T\) is a prescribed deterministic horizon. Its unit-vector output is drawn using only the terminal state, the rule at index \(\tau\), and fresh randomness. In particular, the last observation can affect the output only through the terminal state. Rules may depend on \(d\), the target accuracy, and the index, and may be selected by a shared random seed. The initial state and that seed are jointly independent of the signal and the entire sample sequence. Fresh transition and output random variables may also be used; each is independent of that pair, the signal, the sample sequence, and all earlier fresh random variables before it is used. Memory is measured in bits, not real registers. The construction from the signal, pre-generated sample rows, shared seed, initialization, and fresh transition and output randomness through the state path, stopping index, and output must form a jointly measurable probability experiment. Per-seed measurability of the sample kernels alone is not a definition of the mixture over seeds. For each fixed seed, the transition and stopping coordinates may be measurable in the completion of the Borel sigma field under \(\lambda(dx,dy)=\gamma_{1,d}(dx)\,dy\) on \(\mathbb R^d\times\mathbb R\), where \(\gamma_{1,d}\) is standard Gaussian probability on \(\mathbb R^d\). Lemma 8 supplies Borel versions for the fixed-seed uniform-prior experiment. The original randomized experiment must still be jointly measurable as stated in the definition. Spherical caps and residual fibersLemma 4 (Spherical balls). Let \(\sigma_q\) be uniform probability on \(S^{q-1}\), where \(q\ge2\). For every \(z\in S^{q-1}\) and \(r>0\), \[ \sigma_q\{u:\lVert u-z\rVert\le r\}\le r^{q-1}. \tag{2}\] In particular an angular cap of radius \(0<\epsilon\le\pi\) has mass at most \(\epsilon^{q-1}\). Proof. For \(r\le1\), rotate \(z\) to the north pole. A point in the chordal ball has last coordinate at least \(1-r^2/2\ge1/2\). As a graph over its first \(q-1\) coordinates, the surface has area element at most twice Lebesgue measure, and its projection lies in the \((q-1)\)-ball of radius \(r\). The whole sphere has area at least twice the volume of the unit \((q-1)\)-ball, by projecting both hemispheres. This proves (2). For \(r>1\), total mass one suffices. Angular distance at most \(\epsilon\) implies chordal distance \(2\sin(\epsilon/2)\le\epsilon\), proving the last assertion. ◻ Lemma 5 (Residual sphere and radius bounds). Let \(S\sim\sigma_d\) and let \(G\) be independent of \(S\), with \(b\) rows, where \(0\le b\le d-2\). For \(b\ge1\), let \(G\) be either a standard Gaussian matrix or a uniform Haar frame of \(b\) orthonormal rows in \(\mathbb R^d\). For \(b=0\), interpret \(G\) and \(GS\) as empty. Conditional on \((G,GS)\), the signal is uniform on an affine sphere parallel to \(\ker G\), with ambient kernel dimension \(d-b\) and radius \[R=\lVert P_{\ker G}S\rVert.\] For \(b=0\), this is the original unit sphere and \(R=1\) deterministically. For \(1\le b\le d-2\), the squared radius has distribution \[ R^2\sim\operatorname{Beta}\bigl((d-b)/2,b/2\bigr). \tag{3}\] For \(0<r_0<1\) one has \[ \mathbb P\{R<r_0\}\le \frac{b}{d(1-r_0^2)}. \tag{4}\] When \(0\le b\le d/2\), two useful bounds at radius \(1/2\) are \[ \mathbb P\{R<1/2\}\le\frac{8}{d+2}, \qquad \mathbb P\{R<1/2\}\le\frac{90}{d}. \tag{5}\] The second inequality also follows directly from two Gaussian norm estimates, without using the beta variance. Proof. The zero-row case is immediate: conditioning on empty data changes nothing, and the residual radius is one. It does not involve a beta law with a zero shape parameter. Suppose \(b\ge1\). The matrix is full row rank almost surely in the Gaussian case and by definition in the Haar case. For fixed \(G\), write \(S=Z/\lVert Z\rVert\) with \(Z\) standard Gaussian and split \(Z\) into the row space of \(G\) and its kernel. The squared lengths are independent chi-squares with \(b\) and \(d-b\) degrees of freedom. Their normalized ratio gives (3); the direction in the kernel is independent and uniform. Given the projection, its row-space component is fixed, so the remaining signal is uniform on the asserted affine sphere. Its radius has \[\mathbb ER^2=\frac{d-b}{d},\qquad \operatorname{Var}(R^2)=\frac{2b(d-b)}{d^2(d+2)}.\] These formulas also hold for the deterministic zero-row radius. Markov’s inequality applied to \(1-R^2\), whose mean is \(b/d\), proves (4). If \(b\le d/2\), the mean of \(R^2\) is at least \(1/2\) and its variance is at most \(1/(2(d+2))\). Chebyshev’s inequality at distance \(1/4\) from the mean gives the first bound in (5). For the second bound, still with \(b\le d/2\), let \(Z_{\ker}\) be the kernel component of \(Z\). Its squared norm has mean \(d-b\ge d/2\) and variance at most \(2d\), while \(\lVert Z\rVert^2\) has mean \(d\) and variance \(2d\). Chebyshev gives \[\mathbb P\{\lVert Z_{\ker}\rVert^2<d/3\}\le\frac{72}{d},\qquad \mathbb P\{\lVert Z\rVert^2>4d/3\}\le\frac{18}{d}.\] Outside these events \(R^2\ge1/4\). Their union gives \(90/d\). For \(b=0\) the probability is zero, as already established. ◻ The next estimate tests success against a law in which the signal and output are conditionally independent. This is the binary-relative-entropy event method; see Bassily et al. (Bassily et al. 2018, Appendix A.1, Lemmas 14–15) for finite-space information bounds. Here the reference signal lies on a residual sphere. Lemma 6 (Information on a residual sphere). Let \(S\sim\sigma_d\), let \(V\) be a random variable with values in a standard Borel space, jointly distributed with \(S\), and let \(G\) be independent of \((S,V)\), with \(0\le b\le d-2\) rows. For \(b\ge1\), \(G\) may be either a standard Gaussian matrix or a uniform Haar frame of \(b\) orthonormal rows; at \(b=0\) it is empty. Let \(\widehat S\in S^{d-1}\) be generated from \(V\) by a kernel conditionally independent of \((S,G)\). Write \[p=\mathbb P\{\arccos\langle\widehat S,S\rangle\le\epsilon\}, \qquad R=\lVert P_{\ker G}S\rVert.\] For \(2\epsilon<r_0\le1\), if \(\alpha:=p-\mathbb P\{R<r_0\}>0\), then \[ I(S;V\mid G,GS)\ge \alpha(d-b-1)\log\frac{r_0}{2\epsilon}-\log2. \tag{6}\] Proof. Put \(Z=(G,GS)\). Compare the true law of \((S,V,Z)\) with \(P_ZP_{S\mid Z}P_{V\mid Z}\), and in both laws generate \(\widehat S\) using the same kernel from \(V\). Adding that kernel preserves the divergence, which is \(I(S;V\mid Z)\). Fix side data with residual radius \(R\ge r_0\) and fix an output. If the set of successful points on the fiber is nonempty, choose one such point. Angular success implies Euclidean distance at most \(\epsilon\) from the output, so any two successful fiber points are at Euclidean distance at most \(2\epsilon\). After translating and rescaling the fiber to its unit sphere, the successful set therefore lies in a chordal ball of radius \(2\epsilon/R\le2\epsilon/r_0<1\) centered on that sphere. By Lemma 4, its conditional mass is at most \((2\epsilon/r_0)^{d-b-1}\). Under the comparison law, the output and signal are conditionally independent given \(Z\), so the event of success and \(R\ge r_0\) has probability at most this quantity. Under the true law it has probability at least \(\alpha\). For Bernoulli probabilities \(a,c\), relative entropy obeys \[D_{\mathrm{KL}}(\operatorname{Bern}(a)\Vert\operatorname{Bern}(c)) \ge a\log(1/c)-\log2;\] this follows by dropping the nonnegative term \((1-a)\log(1/(1-c))\) and using binary entropy at most \(\log2\). Apply data processing to the event just described. If the comparison probability is zero the conclusion is immediate; otherwise its upper bound and the true lower bound give (6). ◻ We record three numerical forms used below. The first uses an angular cap bound: a chordal radius \(4\epsilon\) corresponds to angular radius \(2\arcsin(2\epsilon)\le5\epsilon\) for \(\epsilon\le1/10\). Corollary 7 (Three endpoint substitutions). Assume \(0<\epsilon\le1/10\) in Lemma 6. For sufficiently large \(d\) the following bounds hold.
Each right side is at least \(c\,d\log(1/\epsilon)\) for an absolute \(c>0\) and all sufficiently large \(d\). Proof. Use \(r_0=1/2\). In (i), the failure probability is at most \(8/(d+2)\), so the true event has mass at least \(1/2\) for large \(d\). The angular cap bound is \((5\epsilon)^{d-b-1}\), which gives the displayed version of the binary entropy calculation. In (ii), Markov’s estimate leaves event mass at least \(1/3\), and the chordal cap is \((4\epsilon)^{d-b-1}\). In (iii), the stated \(90/d\) estimate leaves mass at least \(1/2\). The last assertion follows from \(d-b-1\ge d/2-1\) and \[\log\frac1{c_0\epsilon}\ge \left(1-\frac{\log c_0}{\log10}\right)\log\frac1\epsilon \qquad(c_0=4\text{ or }5).\] The positive multiple of \(d\log(1/\epsilon)\) absorbs \(\log2\) eventually. ◻ Borel rules, randomness, and bounded stoppingLemma 8 (Borel versions for the averaged experiment). Fix \(d\ge2\), a finite horizon \(T\), and finite state sets. For fixed data-independent rule parameters, suppose each transition and stopping coordinate is measurable in the completion of the Borel sigma field on \(\mathbb R^d\times\mathbb R\) under \(\lambda(dx,dy)=\gamma_{1,d}(dx)\,dy\), where \(\gamma_{1,d}\) is standard Gaussian probability on \(\mathbb R^d\). The output law depends only on the terminal state and stopping index. There are Borel transition and stopping kernels, with the same state sets and output laws, having the same joint law of the signal, state path, stopping index, and output when \(S\sim\sigma_d\). Proof. For each state and index combine the next-state and stopping alternatives into a finite destination set. Choose a Borel version of every coordinate of this finite probability vector. Such versions exist for completed- measurable real functions: approximate by simple functions and replace their countably many level sets by Borel sets modulo null sets. On the Borel null set where the resulting vector is negative in a coordinate or fails to sum to one, replace it by a fixed point mass. The finite union over all states and indices is contained in one Borel \(\lambda\)-null set \(N\). For a nonzero fixed \(x\), the scalar \(\langle x,S\rangle\) has a Lebesgue density when \(S\sim\sigma_d\) and \(d\ge2\). Rotation and the one-coordinate spherical law show this directly. The event \(x=0\) is Gaussian-null. Consequently each pre-generated sample pair \((X_j,\langle X_j,S\rangle)\) has law absolutely continuous with respect to \(\lambda\), and it avoids \(N\) almost surely. This remains true on the event that any particular state is reached, because that event intersected with the pair lying in \(N\) is contained in the marginal null event. The finite union over sample indices is still null. Coupling each original finite kernel and its Borel version by the same fresh uniform random variable now makes their state paths and stopping decisions agree by induction. Their output laws agree at the identical terminal pair. ◻ This lemma preserves the experiment averaged over the uniform signal; it does not claim pointwise preservation at every fixed signal. We may therefore apply it after passing to the uniform-prior hypothesis of Theorem 2. When a shared seed is present, apply the replacement separately in its fixed-seed experiments. The original jointly measurable law supplies well-defined conditional success probabilities and their average. No jointly measurable selection of the replacement Borel versions is needed. Lemma 9 (Fixing randomness and padding the clock). Consider a learner in Definition 3 under the uniform prior. If its average success is at least \(p\), some fixed shared-seed conditional experiment has success at least \(p\). Its initialization remains independent of the signal and samples, and its fixed-seed Borel replacement has the same success. In this Borel experiment the initialization and all remaining randomized rules can be represented by data-independent variables chosen before sampling, followed by deterministic state-based rules. For every \(p'<p\), some realization of those variables gives deterministic rules with average success at least \(p'\). Padding a learner stopped by \(T\) so that its halted states are absorbing, while retaining the stopping index and postponing the output, uses at most \((T+2)2^M\) states at every layer. Hence every padded state has entropy at most \[ h_T=M\log2+\log(T+2). \tag{10}\] Extra ignored samples may be generated to complete a last block. Proof. The original jointly measurable law gives a measurable conditional success function of the shared seed whose average is the original success. Work on a common full-measure set of seed values where the conditional law is the stated fixed-seed experiment and the conditional initialization remains independent of the signal and rows. The latter property follows from the joint independence of the initial state and shared seed from the signal and all rows. Some seed in this set has conditional success at least \(p\): otherwise the strictly positive difference between \(p\) and that conditional success would have a positive integral. Apply Lemma 8 to this one fixed-seed experiment. Its uniform-prior success is unchanged. This step does not require a jointly measurable choice of Borel replacements across seeds. In the resulting fixed-seed Borel experiment, for a finite destination kernel use an independent uniform variable and the cumulative coordinate probabilities to choose the destination. There are only finitely many state/index rules; unused choices can be generated in advance as well. Similarly pre-generate, for each terminal state/index, an output with its prescribed distribution, independently of the data. Together with the conditional initial state, these variables are jointly independent of the signal and samples. Conditional on their realization, the initialization and rules are deterministic and state based. The finite Borel construction is jointly measurable, and averaging its success gives that of the selected fixed-seed experiment. Some realization therefore has success at least \(p'<p\). At a layer the active state set has at most \(2^M\) values. Each earlier stopping index contributes a disjoint frozen copy of a state set of that size, and there are at most \(T+1\) stopping indices, including zero. The stated count follows. The frozen pair retains exactly the information allowed to determine the postponed output, and ignored fresh samples add no information to it. ◻ Lemma 10 (Conditional priors at finite events). Let \(S\sim\sigma_d\) and let \(H\) be a finite-valued random variable jointly distributed with \(S\). Every positive-probability value \(h\) gives a conditional signal density bounded by \(L_h=1/\mathbb P\{H=h\}\) relative to \(\sigma_d\), and \[ \mathbb E\log L_H=H(H). \tag{11}\] Proof. Bayes’ formula gives the Borel density \[\frac{dP_{S\mid H=h}}{d\sigma_d}(s) =\frac{\mathbb P\{H=h\mid S=s\}}{\mathbb P\{H=h\}}\le L_h.\] Averaging \(\log L_h\) gives (11). In particular, \(H\) may be one padded state or a history of padded states; these have different entropy costs. ◻ Lemma 11 (Output capacity and the early clock bound). Consider a family of learners indexed by \(d\), with uniform-prior success at least a fixed \(p_0>0\), and set \(\beta=\log(1/\epsilon)\) for \(0<\epsilon\le1/10\). At each dimension, the terminal output capacity satisfies \[ p_0\le (T+1)2^M\epsilon^{d-1}. \tag{12}\] On the branch \(T\le d\beta\), if \(M=o(d^2)\), then \[ \beta=o(d),\qquad T=o(d^2),\qquad h_T=o(d^2). \tag{13}\] Proof. First fix a shared seed in the original conditional experiment. There are at most \((T+1)2^M\) terminal state/index pairs. For a fixed pair, write \(\nu_{j,v}\) for its prescribed output law and \(r_{j,v}(s)\) for the conditional probability of reaching it given \(S=s\). Its contribution to success is \[\int r_{j,v}(s)\int \mathbf 1_{\{\arccos\langle u,s\rangle\le\epsilon\}} \,\nu_{j,v}(du)\,\sigma_d(ds) \le\epsilon^{d-1},\] because \(r_{j,v}\le1\) and each angular cap has mass at most \(\epsilon^{d-1}\) by Lemma 4. Summing gives (12) for its conditional success. The same numerical bound holds for almost every seed, so averaging the original conditional success probabilities gives the stated bound for the jointly measurable mixture. On \(T\le d\beta\), the inequalities \(\beta\ge\log10\) and \(T+1\le2d e^\beta\) imply \[(d-2)\beta\le M\log2+\log(2d)+\log(1/p_0).\] Dividing by \(d\) proves \(\beta=o(d)\) for a subquadratic memory sequence. The branch then gives \(T=o(d^2)\), and (10) gives \(h_T=o(d^2)\). ◻ Lemma 12 (A raw-sample endpoint). Let \(0<\epsilon\le1/10\), and let an estimator use all exact measurements \((G,GS)\) from \(b\le\lfloor d/2\rfloor\) independent Gaussian rows, together with independent randomness, when \(S\sim\sigma_d\). Its angular success probability is at most \[ \frac{8}{d+2}+(4\epsilon)^{d-b-1}. \tag{14}\] The separate Gaussian-norm estimate gives the also valid bound \(90/d+(4\epsilon)^{d-b-1}\). In particular, uniform-prior success at least \(3/5\) requires \(T>\lfloor d/2\rfloor\) for all sufficiently large \(d\), even if all raw data up to the horizon are retained. Proof. For \(d\le2\), both numerical bounds exceed one. We may therefore assume \(d\ge3\), when \(b\le\lfloor d/2\rfloor\le d-2\). Condition on all the data. The estimator and the signal are then conditionally independent, and the conditional signal is the residual sphere in Lemma 5. On radius at least \(1/2\), the successful points have relative chordal cap mass at most \((4\epsilon)^{d-b-1}\), by the argument in Lemma 6. Add the corresponding failure probability from (5). The zero-row case uses the deterministic radius one. If a learner stops by \(T\le\lfloor d/2\rfloor\), grant it all of the first \(\lfloor d/2\rfloor\) rows and labels, including unused ones. Its output is a kernel of those data and independent randomness. The displayed bound tends to zero uniformly for \(\epsilon\le1/10\), contradicting success \(3/5\). This proves the last assertion. The angular-cap version used in the first replica route follows as well by replacing \(4\epsilon\) by the weaker \(5\epsilon\). ◻ Posterior replicas and exact equal-label measuresThe information estimate for a block has two parts. First, a general identity converts information in one finite message into the loss of dependence among several signals. Second, exact Gaussian observations give an explicit law for signals sampled from the same posterior. This section proves both statements. The subsequent sections estimate how much of that dependence remains visible after another projection. All probability spaces used below are standard Borel. Conditional laws mean measurable regular conditional probabilities, and conditional relative entropies include integration over the conditioning variable. We write \(D_{\mathrm{KL}}(P\Vert Q)\) for relative entropy, with value \(+\infty\) when \(P\) is not absolutely continuous with respect to \(Q\). Logarithms and entropies are in nats. Information in a message and loss of correlationFor a joint law \(Q\) of \((S_1,\ldots,S_t)\) with common marginal \(p\), its total correlation is \[K=D_{\mathrm{KL}}(Q\Vert p^{\otimes t}).\] This quantity was introduced by Watanabe (Watanabe 1960); see Austin (Austin 2020, sec. 3.3 and 6) for its measure-theoretic form and information identities. Let \(B\) be independent of the tuple and let \(U_i=\phi_B(S_i)\), where \((B,s)\mapsto\phi_B(s)\) is measurable. Writing \(p_B=(\phi_B)_\#p\), define the correlation visible after projection by \[K'=\mathbb E_BD_{\mathrm{KL}}\bigl(Q_{U_1,\ldots,U_t\mid B}\Vert p_B^{\otimes t}\bigr).\] Lemma 13 (Replica information inequality). Let \(t\ge1\) and suppose \(K<\infty\). Let \(W\) be finite-valued and suppose that every pair \((S_i,W)\) has the same law as \((S_1,W)\). In the conclusion write \(S=S_1\). If \(B\) is independent of \((S_1,\ldots,S_t,W)\), then \[ t I(S;W\mid B,\phi_B(S))\le H(W)+K-K',\qquad 0\le K'\le K. \tag{15}\] The same statement holds in a conditional experiment when all of these hypotheses hold there, including the matching pair laws and the independence of \(B\). Equivalently, the left side may use a separate representative of the common signal–message pair law if it is coupled independently with \(B\). Proof. Let \(K_W\) be the average total correlation conditional on \(W\), computed relative to the product of the conditional one-coordinate marginals. Decomposing the relative entropy of \((S_1,\ldots,S_t,W)\) from \(p^{\otimes t}\otimes Q_W\) in the two orders gives \[ tI(S;W)=I((S_i)_{i=1}^t;W)+K-K_W. \tag{16}\] The relative entropy being decomposed is at most \(K+H(W)\), so every nonnegative chain-rule term in this display is finite. Define \(K'_W\) in the same way after projection, conditionally on \((B,W)\). The corresponding identity after projection is \[ tI(\phi_B(S);W\mid B) =I((U_i)_{i=1}^t;W\mid B)+K'-K'_W. \tag{17}\] Coordinatewise data processing gives \(K'\le K\) and \(K'_W\le K_W\). Independence of \(B\) and the fact that \(\phi_B(S)\) is determined by \((B,S)\) also give \[I(S;W\mid B,\phi_B(S))=I(S;W)-I(\phi_B(S);W\mid B).\] Subtracting (17) from (16) therefore gives the exact identity \[tI(S;W\mid B,\phi_B(S)) =I((S_i)_{i=1}^t;W\mid B,(U_i)_{i=1}^t) +(K-K')-(K_W-K'_W).\] The conditional tuple-information is at most \(H(W)\), and \(K_W-K'_W\ge0\). This proves the inequality. The remaining assertion is nonnegativity of relative entropy. ◻ The coupling used in this paper has precisely the pair-law property in the lemma. Suppose that data \((A,Y)\) are generated from \(S\), and that \(W\) is drawn from a kernel of the data, conditionally independently of \(S\). Given the same \((A,Y)\), draw \(S_1,\ldots,S_t\) independently from the posterior law of \(S\), independently also of \(W\). Conditioning on the data shows that \((S_i,W)\) has the original signal–message law for every \(i\). We call these signals posterior replicas. Equal one-signal marginals alone would not be sufficient for Lemma 13. The invariance of the joint data–signal law under posterior resampling is the Nishimori identity; see Lelarge and Miolane (Lelarge and Miolane 2019, Proposition 16) for a general formulation. When an independent \(B\) is added to both experiments, equality of the pair laws gives equality of the corresponding \((S_i,W,B)\) laws as well, so the lemma’s left side is the information in the original signal–message pair. We shall use the following entropy test. If \(Q\ll R\) are probabilities and \(F\ge0\) is measurable with \(\mathbb E_R F\le1\), then \[ D_{\mathrm{KL}}(Q\Vert R)\ge\mathbb E_Q\log F \tag{18}\] whenever the right side is defined. For a bounded positive test this follows by tilting \(R\) by \(F/\mathbb E_R F\) and using nonnegativity of relative entropy; bounded truncations give the stated form. This is the entropy variational principle in the form used here; see Dupuis and Mao (Dupuis and Mao 2022, Equations (1.1)–(1.2)). In particular, if the relative entropy is finite, applying the bounded form to truncations of \((1+F)/2\) and then monotone convergence yields \[ \mathbb E_Q\log(1+F)\le D_{\mathrm{KL}}(Q\Vert R)+\log2. \tag{19}\] This supplies positive-log integrability for the likelihood tests below. Projection fibers and inverse affine volumesLet \(\sigma_d\) be uniform probability on \(S^{d-1}\) and let \(\gamma_{k,d}\) be the law of a \(k\times d\) matrix with independent standard Gaussian entries. For a subspace \(F\subset\mathbb R^d\), let \(\gamma_{k,d}^{F}\) denote the law of a matrix whose rows are independent standard Gaussians in \(F^\perp\). For a tuple \(\mathbf u=(u_1,\ldots,u_t)\) with \(m=t-1\), set \[ D_{\mathbf u}=[u_2-u_1,\ldots,u_t-u_1],\qquad F_{\mathbf u}=\mathop{\mathrm{col}}(D_{\mathbf u}),\qquad J(\mathbf u)=\det(D_{\mathbf u}^{\mathsf T}D_{\mathbf u})^{1/2}. \tag{20}\] At full column rank \(J\) is the product of the successive distances to predecessor affine spans. For two distinct points, the difference \(v=u_2-u_1\) has \(Av\sim N(0,\lVert v\rVert^2I_k)\), whose density at zero is \((2\pi)^{-k/2}\lVert v\rVert^{-k}\). This is the inverse-volume factor for two points. With \(m\) differences, each row of \(AD_{\mathbf u}\) has covariance \(D_{\mathbf u}^{\mathsf T}D_{\mathbf u}\), giving the factor \((2\pi)^{-km/2}J(\mathbf u)^{-k}\). The next lemma provides the integrability needed to identify an exact measure on the constraint \(Au_1=\cdots=Au_t\). Lemma 14 (Projection fibers and inverse affine volumes). Let \(1\le k\le d-2\) and let \(P^0\) be either \(\sigma_d\) or standard Gaussian probability on \(\mathbb R^d\). Every full-row-rank \(k\times d\) matrix \(A\) projects \(P^0\) to a Lebesgue density \(p_A^0\). With \(A^\dagger=A^{\mathsf T}(AA^{\mathsf T})^{-1}\), the spherical density is \[ p_A^0(y)=\frac{\Gamma(d/2)}{\pi^{k/2}\Gamma((d-k)/2)} \frac{(1-\lVert A^\dagger y\rVert^2)^{(d-k-2)/2}} {\sqrt{\det(AA^{\mathsf T})}} \mathbf 1_{\{\lVert A^\dagger y\rVert<1\}}. \tag{21}\] For \(\lVert A^\dagger y\rVert<1\), its conditional law \(P^0_{A,y}\) is uniform on the sphere centered at \(A^\dagger y\), of radius \(\sqrt{1-\lVert A^\dagger y\rVert^2}\), in the affine space parallel to \(\ker A\). For the Gaussian base, \(p_A^0\) is the density of \(N(0,AA^{\mathsf T})\), and the conditional point is \(A^\dagger y\) plus a standard Gaussian in \(\ker A\). These conditional laws have jointly measurable versions; unused fibers may be assigned any fixed probability law. If \(m\ge1\), \(t=m+1\), and \(k+m\le d-2\), then \(J>0\) almost surely under \((P^0)^{\otimes t}\) and \[ \int J(\mathbf u)^{-k}\,(P^0)^{\otimes t}(d\mathbf u)<\infty. \tag{22}\] If \(k+m\le d-3\), the same assertion holds with exponent \(k+1\). Proof. Split a standard Gaussian vector into the row space and kernel of \(A\). Its squared lengths in those spaces are independent chi-squares with \(k\) and \(d-k\) degrees of freedom. After dividing by the total length, their ratio has the beta density. Polar integration in orthonormal row coordinates gives (21); changing those coordinates to \(y\) supplies the determinant factor. The remaining direction is independent and uniform in the kernel. This is the spherical beta projection law, discussed geometrically by Frankl and Maehara (Frankl and Maehara 1990). Without normalization, orthogonal Gaussian independence gives the Gaussian assertion. For measurability, generate the kernel component by projecting an extra independent Gaussian with the Borel matrix \(I-A^\dagger A\), and normalize its length in the spherical case. Assign fixed values on rank failures and on the null event of a zero projected Gaussian. For the inverse moment, first note that an affine \(j\)-plane \(V\) satisfies \[ \sigma_d\{u:\mathop{\mathrm{dist}}(u,V)\le r\}\le C^d r^{d-1-j},\qquad 0<r\le1. \tag{23}\] Indeed, cover the part of \(V\) within distance one of the unit sphere by at most \((C/r)^j\) balls of radius \(r\). Each nonempty intersection with the sphere is contained in a radius-\(4r\) ball centered on the sphere. The spherical ball bound in Lemma 4, with the trivial bound one when necessary, proves (23). For a standard Gaussian point, its projection onto the orthogonal complement of the direction space of \(V\) has a translated standard Gaussian density. The density is bounded uniformly over the translation, so the analogous estimate is \(C^d r^{d-j}\). Gram–Schmidt gives \[J(\mathbf u)=\prod_{i=2}^t \mathop{\mathrm{dist}}\bigl(u_i,\mathop{\mathrm{aff}}(u_1,\ldots,u_{i-1})\bigr).\] The predecessor affine span has dimension at most \(i-2\). Thus the smallest spherical small-distance exponent in (23) is \(d-1-(t-2)=d-m>k\). Integrating its tail bounds the inverse \(k\)th moment of each distance uniformly in the preceding points: for a nonnegative distance \(X\), \[\mathbb EX^{-k}\le1+k\int_0^1 r^{-k-1}\mathbb P\{X<r\}\,dr.\] Integrate the last point first and repeat. This proves the finite product moment and, since every lower-dimensional affine span has zero base mass, full affine rank almost surely. The Gaussian exponents are larger. Replacing \(k\) by \(k+1\) gives the last assertion under \(k+m\le d-3\). ◻ The exact finite measure on equal labelsLet \(f\) be a bounded nonnegative Borel function and set \(\nu=fP^0\). The measure \(\nu\) is finite and need not be a probability. For a full-rank \(A\), choose the projected density and the normalized conditional probability by \[ p_A(y)=p_A^0(y)h_A(y),\qquad h_A(y)=\int f\,dP^0_{A,y},\qquad \nu_{A,y}(du)=\frac{f(u)}{h_A(y)}P^0_{A,y}(du) \quad\text{when }h_A(y)>0. \tag{24}\] The conditional probability on an unused fiber may be arbitrary. In the spherical case the displayed density is zero outside the open projection ellipsoid. The constructions in Lemma 14 make these versions jointly measurable. The equality below distinguishes two natural descriptions. The left first selects a matrix and a common label, with the label weighted by \(p_A^t\), then selects points independently on that fiber. The right first selects independent points and then Gaussian rows annihilating their differences, with an inverse affine-volume weight. These are finite measures, not the two unweighted probability experiments. When \(\nu\) is a probability, its actual common label has density \(p_A\), so recovering posterior replicas will require multiplying this identity by \(p_A^{1-t}\) on positive fibers. We first identify the measures and their null sets, before that division. The general background is coarea (Federer 1959, Theorem 3.1); related affine simplex-volume Jacobians occur in Drury (Drury 1984, Lemma 1). The following proof fixes the normalization and the versions used at the exact constrained labels. Theorem 15 (Finite measure for equal Gaussian labels). Let \(k,m\ge1\), \(t=m+1\), and \(k+m\le d-2\). Let \(P^0\) be either uniform probability on \(S^{d-1}\) or standard Gaussian probability on \(\mathbb R^d\), and let \(\nu=fP^0\) for a bounded nonnegative Borel function \(f\). With the notation in (20) and (24), for every nonnegative Borel function \(H(A,y,\mathbf u)\), \[\begin{align*} &\int\gamma_{k,d}(dA)\int_{\mathbb R^k}p_A(y)^t \int H(A,y,\mathbf u)\,\nu_{A,y}^{\otimes t}(d\mathbf u)\,dy \\ &\quad=(2\pi)^{-km/2}\int J(\mathbf u)^{-k}\,\nu^{\otimes t}(d\mathbf u) \int H(A,Au_1,\mathbf u)\,\gamma_{k,d}^{F_{\mathbf u}}(dA). \tag{25}\end{align*}\] Both sides are finite measures before the test \(H\) is inserted. They are supported on \(Au_i=y\) for every \(i\), full-column-rank \(D_{\mathbf u}\), and full-row-rank \(A\). In the spherical case the common fiber has positive radius. Thus (25) is an identity of finite measures on \((A,y,\mathbf u)\); the right side places \(y\) at \(Au_1\). For a full-rank \(A\), alternatively define \(p_A(y)\) by the lower limit of its radius-\(1/n\) Lebesgue ball averages or by the lower limit of its Gaussian convolutions with covariance \(n^{-2}I_k\). Each choice is jointly measurable and agrees with (24) for Lebesgue-almost every \(y\). All three versions agree at \(y=Au_1\), with a finite positive value, almost everywhere for the finite reference measure \[ R(d\mathbf u,dA)=\nu^{\otimes t}(d\mathbf u) \gamma_{k,d}^{F_{\mathbf u}}(dA). \tag{26}\] When \(\nu=0\) the last assertion is vacuous. The equality is obtained using either ball or Gaussian approximate identities in both integration orders for the base law \(P^0\), then multiplying the resulting measure identity by the bounded density factors. Proof. We first take \(\nu=P^0\). Let \(\varphi_\eta\) be either the uniform density on the radius-\(\eta\) \(k\)-ball or the density of \(N(0,\eta^2 I_k)\). For a continuous compactly supported test \(G(\mathbf u,A)\) whose matrix support lies in the open full-row-rank set, integrate \[ G(\mathbf u,A)\prod_{i=2}^t\varphi_\eta(A(u_i-u_1)) \tag{27}\] against \((P^0)^{\otimes t}(d\mathbf u)\gamma_{k,d}(dA)\). Fix \(A\) first and disintegrate each point by its projected label. Given the first label \(y\), the other labels approach \(y\) under either probability approximate identity. The base density is continuous in the interior of its support, and the conditional fiber law varies weakly continuously there. The spherical boundary has zero projected probability. Thus the integral converges to \[\int\gamma_{k,d}(dA)\int (p_A^0(y))^t \int G(\mathbf u,A)(P^0_{A,y})^{\otimes t}(d\mathbf u)\,dy.\] To justify the passage, condition first also on \(u_1\). The absolute value of the remaining integral is at most \(\lVert G\rVert_\infty(\sup p_A^0)^m\). The density suprema are uniformly bounded on the compact matrix support: this follows from (21), whose exponent is nonnegative, and from the Gaussian density formula. The bound is global in all the other labels, including labels near or outside the spherical support, because each \(\varphi_\eta\) is a probability density. Dominated convergence applies for both kernels. In the other order, fix a tuple of full affine rank. Each row of \(AD_{\mathbf u}\) is Gaussian with covariance \(D_{\mathbf u}^{\mathsf T}D_{\mathbf u}\). Its joint density has maximum at zero, with value \[w_{k,m}(\mathbf u)=(2\pi)^{-km/2}J(\mathbf u)^{-k}.\] Conditional on \(AD_{\mathbf u}=Z\), the matrix \(A\) is a draw from \(\gamma_{k,d}^{F_{\mathbf u}}\) plus \(Z(D_{\mathbf u}^{\mathsf T}D_{\mathbf u})^{-1}D_{\mathbf u}^{\mathsf T}\). This translation tends to zero with \(Z\). The integral in (27), conditioned on the tuple, therefore converges to \[w_{k,m}(\mathbf u)\int G(\mathbf u,A) \gamma_{k,d}^{F_{\mathbf u}}(dA).\] For either kernel its absolute value is at most \(\lVert G\rVert_\infty w_{k,m}(\mathbf u)\). Indeed, the joint density of the whole matrix \(Z=AD_{\mathbf u}\) is bounded by \(w_{k,m}\), while \(\prod_{i=2}^t\varphi_\eta(Z_{:,i-1})\) has Lebesgue integral one on \(\mathbb R^{k\times m}\). This argument does not require the columns of \(Z\) to be independent. The bound is integrable by Lemma 14. Dominated convergence gives the second integration order and equates these two limits. For clarity, the left measure is locally finite on the open full-row-rank matrix set before we identify it. If a compact set in that space has matrix projection \(K_A\), its left mass is at most \[\int_{K_A}\!\int (p_A^0(y))^t\,dy\,\gamma_{k,d}(dA) \le \sup_{A\in K_A}(\sup_y p_A^0(y))^m<\infty,\] because \(\int p_A^0(y)\,dy=1\). Thus both sides are locally finite Borel measures on this locally compact second-countable space; the right one is globally finite by (22). Continuous compactly supported tests identify them on that open set. Exhausting it shows that the left total mass there equals the finite right total mass. The right measure gives no mass to remaining matrices: its rows live in a space of dimension \(d-m\ge k+2\). The left measure also has no such mass because \(\gamma_{k,d}\) is full rank almost surely. Equality therefore extends to all nonnegative Borel tests. The right measure is supported on full affine rank by Lemma 14, so the same holds on the left. In the spherical case its label is in the open ellipsoid, and its fiber radius is positive. Equivalently, a boundary fiber could contain only one unit vector and hence could not contain a full-affine- rank tuple with \(t\ge2\). Now multiply the base equality by \(\prod_{i=1}^t f(u_i)\). The formula (24) converts the fiber side to \(p_A(y)^t\nu_{A,y}^{\otimes t}\); the other side becomes \(w_{k,m}\nu^{\otimes t}\gamma_{k,d}^{F_{\mathbf u}}\). On a fiber with \(h_A(y)=0\), the measure \(fP^0_{A,y}\) is zero, so both expressions on the fiber side are zero regardless of the unused conditional probability. This conversion never divides on a zero fiber. Boundedness of \(f\) preserves finiteness. Inserting a test \(H(A,Au_1,\mathbf u)\) gives (25), since both sides are supported on equal labels. No normalization by \(\nu(\mathbb R^d)\) has been made. In particular, the two equal-label limit calculations were made only for the base law; no continuity of \(f\) or of the reweighted fiber kernel is used in them. It remains to justify density evaluation on that constraint. For either choice of \(\varphi\), set \[\widetilde p_A(y)=\liminf_{n\to\infty} \int\varphi_{1/n}(Au-y)\,\nu(du).\] The integral and lower limit are jointly measurable. Lebesgue differentiation for the ball and Gaussian approximate identities makes each version agree with (24) at almost every \(y\) for each full-rank \(A\). On the left of (25), the common label has a measure with Lebesgue density \(p_A(y)^t\) and hence avoids the null sets where the versions disagree or are not finite and positive. Equality of the undivided finite measures transfers this fact to the right. There the weight \(w_{k,m}\) is finite and strictly positive at full affine rank. The reference \(R\) has finite mass \(\nu(\mathbb R^d)^t\) and full affine rank almost everywhere, because \(\nu^{\otimes t}\ll(P^0)^{\otimes t}\). Restricting successively to \(\{1/j\le w_{k,m}\le j\}\) therefore transfers the same null statement to the unweighted reference (26). The labels where \(p_A=0\) have zero left mass because their weight is \(p_A^t\), so this argument does not divide on them. If \(\nu=0\), both finite measures and \(R\) are zero. The only additional limiting assertion for a possibly discontinuous \(f\) is the ordinary almost-everywhere differentiation of its projected \(L^1\) density used in this paragraph. This proves the claimed versions before any division by a projection density. ◻ Corollary 16 (Posterior replica likelihood). Under the hypotheses of Theorem 15, suppose \(p=fP^0\) is a probability. Draw \(A\sim\gamma_{k,d}\) independently of \(S\sim p\), put \(Y=AS\), and conditionally on \((A,Y)\) draw \(t\) independent replicas from \(p_{A,Y}\). Let \(Q\) be the law of their tuple together with \(A\), and let \(R=p^{\otimes t}(d\mathbf u)\gamma_{k,d}^{F_{\mathbf u}}(dA)\). Then \(R\) is a probability and \[ \frac{dQ}{dR}(\mathbf u,A) =(2\pi)^{-km/2}J(\mathbf u)^{-k}p_A(Au_1)^{-m}. \tag{28}\] In particular, the tuple marginal of \(Q\) is absolutely continuous relative to \(p^{\otimes t}\). Proof. Multiply the equal finite measures by \(p_A(Au_1)^{-m}\) on their common finite positive set, first with bounded truncations and then by monotone convergence. The exponent \(t=m+1\) on the left becomes one, so that side is exactly the probability experiment defining \(Q\). The right side is the displayed density relative to \(R\). Marginalization gives the last assertion. ◻ Remark 17 (The uniform instance used with a further inverse moment). Suppose the replica count is \(t\ge2\) and \(k+t+1<d\). Then the uniform spherical instance of Theorem 15 applies with \(m=t-1\), and Lemma 14 also gives \(\int J^{-(k+1)}d\sigma_d^{\otimes t}<\infty\). This stronger inverse moment controls a logarithmic volume under a reference weighted by \(J^{-k}\). Extending the likelihood argument to an unbounded finite-entropy density requires an additional proof; it is not included in the bounded-reweighting theorem. Finite correlation for bounded spherical priorsThe finite-measure identity applies to both bases. The additional logarithmic assertion needed for the first two block comparisons concerns a spherical probability prior. It does not assert an entropy bound for every Gaussian reweighting. Lemma 18 (Logarithmic integrability for spherical replicas). In Corollary 16, let \(p=f\sigma_d\) with \(0\le f\le L<\infty\). Under the replica law \(Q\), both \(\log p_A(Y)\) and \(\log J(S_1,\ldots,S_t)\) are absolutely integrable. Consequently the tuple total correlation is finite and satisfies \[ K\le -\frac{km}{2}\log(2\pi)-k\mathbb E_Q\log J-m\mathbb E_Q\log p_A(Y). \tag{29}\] Proof. Write \(u_A\) for the uniform-sphere projection density. Since \(p_A\le Lu_A\), (21) bounds the positive logarithm of \(p_A(Y)\) by a constant depending on \(d,k,L\) plus \(\tfrac12|\log\det(AA^{\mathsf T})|\). Gaussian row Gram–Schmidt expresses that determinant as a product of independent chi-squares with degrees \(d,d-1,\ldots,d-k+1\). Their absolute logarithms are integrable, directly from their densities at zero and infinity. For the negative part, the support of \(p_A\) lies in the ball of radius \(\lVert A\rVert_{\mathrm{op}}\) and \(-z\log z\le1/e\) for \(0<z\le1\). The expected volume of that ball is finite because a finite Gaussian matrix has all positive norm moments. For a simplex height, condition on \((A,Y)\) and let \(\rho\) be the fiber radius. Set \(q=d-k\). After translating and rescaling to the unit sphere in \(\ker A\), the posterior has density at most \[L' =\frac{L u_A(Y)}{p_A(Y)}\ge1.\] For \(a_i=\mathop{\mathrm{dist}}(S_i,\mathop{\mathrm{aff}}(S_1,\ldots,S_{i-1}))\), the tube estimate (23), also conditional on the predecessors, gives \[\mathbb P\{a_i\le r\rho\mid A,Y,S_1,\ldots,S_{i-1}\} \le\min\{1,L'C^q r^{q-1-(i-2)}\},\qquad 0<r\le1.\] The exponent is positive. With \(\log^-x=\max\{0,-\log x\}\), tail integration yields \[\mathbb E[\log^-(a_i/\rho)\mid A,Y,S_1,\ldots,S_{i-1}] \le\log L'+Cq+1.\] Moreover the bounded density ratio makes its relative entropy finite and \[ 0\le\mathbb E\log L'=\log L-\mathbb E_AD_{\mathrm{KL}}(p_A\Vert u_A)\le\log L. \tag{30}\] For completeness, the negative part of the density-ratio logarithm is integrable because \(-z\log z\le1/e\) against the probability density \(u_A\); the positive part is at most \(\log L\). Each pair \((A,S_i)\) has the original independent matrix–signal law. For every fixed unit signal, rotational invariance of \(A\) makes \(\rho^2\) a beta variable with parameters \((d-k)/2,k/2\), whose logarithm is integrable. Thus \(\mathbb E[-\log\rho]<\infty\). Each unscaled height is at most two, so its positive logarithm is bounded. Summing its two logarithmic parts over the Gram–Schmidt product proves absolute integrability of \(\log J\). We may now integrate the logarithm of (28); this is the finite divergence \(D_{\mathrm{KL}}(Q\Vert R)\). Forgetting \(A\) sends \(R\) to \(p^{\otimes t}\), so data processing proves (29). ◻ A Gaussian test that cancels the affine volumeWe prove Theorem 1. Fix its dimensions \(k=m=2\lfloor d/16\rfloor\), \(t=m+1\), and \(\ell=4k\), and let \(p=f\sigma_d\) with \(0\le f\le L\). Draw the block data \((A,Y)\), where \(Y=AS\), and then draw \(t\) independent posterior replicas conditional on those data. Write \(Q\) for the law of the replicas together with \(A\), and use \(D,F,J\) for their difference matrix, difference span, and affine volume as in (20). For large \(d\), \(k+m\le d-2\). Corollary 16 applies with \[c=(2\pi)^{-km/2},\qquad R(d\mathbf s,dA)=p^{\otimes t}(d\mathbf s)\gamma_{k,d}^{F}(dA).\] The evaluated density is the common exact version of Theorem 15. If \(K\) is the total correlation of the replicas, Lemma 18 gives \[ K\le\log c-k\mathbb E_Q\log J-m\mathbb E_Q\log p_A(Y)<\infty. \tag{31}\] The potentially large term is the negative logarithm of \(J\). We recover it from a test determined by the projected replicas. After that cancellation, only a comparison between two equal-label projection densities remains. The observable projected testLet \(B\) be the independent \(\ell\)-row matrix in Theorem 1, and set \(U_i=BS_i\). The projected tuple determines \[BD=[U_2-U_1,\ldots,U_t-U_1].\] At its almost sure full rank, choose a \(k\times\ell\) matrix \(C_0\) with orthonormal rows in the orthogonal complement of the columns of \(BD\), measurably as a function of \(BD\). Projecting the standard basis and taking the first successful Gram–Schmidt choices supplies such a rule. There is room for these rows because \(\ell-m=3k\). Define the selector arbitrarily on rank failures. Put \[\widetilde A=C_0B.\] This matrix is known from \((B,U_1,\ldots,U_t)\) and satisfies \(\widetilde A S_i=C_0U_1\) for every replica. It supplies an observable equal-label likelihood to compare with the actual likelihood in (31). Lemma 19 (Projected likelihood test). Let \(G\) be an \(\ell\times m\) standard Gaussian matrix and set \[Z_{k,m}=\mathbb E\det(G^{\mathsf T}G)^{-k/2}.\] This constant is finite and positive. The nonnegative function \[ \Lambda(B,U_1,\ldots,U_t) =\frac{c}{Z_{k,m}} \det((BD)^{\mathsf T}(BD))^{-k/2} p_{C_0B}(C_0U_1)^{-m}, \tag{32}\] with value zero on the exceptional null sets, has expectation one under the law that draws \(B\) and projects independent \(p\)-distributed signals. Conditional on the replica tuple and its original matrix \(A\), \(\widetilde A\) has law \(\gamma_{k,d}^{F}\); this law depends only on the tuple. If \(K'\) is the projected total correlation of the replicas, then \[ K-K'\le Ckm+ m\mathbb E\!\left[\log p_{\widetilde A}(\widetilde A S_1)-\log p_A(Y)\right]. \tag{33}\] Every logarithm in this display is absolutely integrable. Proof. For a fixed full-rank tuple write \(D=ET\), where the columns of \(E\) form an orthonormal basis of \(F\) and \(T\) is invertible. Then \(G=BE\) is standard Gaussian and \[\det((BD)^{\mathsf T}(BD)) =J^2\det(G^{\mathsf T}G).\] The action of \(B\) on \(F^\perp\) is independent of \(G\). Conditional on \(G\) and the tuple, \(C_0\) is fixed with orthonormal rows and \(C_0BE=0\). Thus \(\widetilde A=C_0B\) has law \(\gamma_{k,d}^{F}\), independently of \(G\) given the tuple. Since \(B\) is independent of the replica experiment, this remains its law after conditioning also on the original \(A\). The test is observable from \((B,U_1,\ldots,U_t)\): \(C_0\), \(C_0B\), and its evaluated label \(C_0U_1\) are all determined there. The common density versions of Theorem 15 make the last factor measurable, finite, and positive almost surely under the independent-tuple reference. The replica tuple is absolutely continuous with respect to that reference’s tuple marginal, so the same assertion holds for replicas. Under the independent-signal reference law, integrate the factor \(\det(G^{\mathsf T}G)^{-k/2}\) first. It contributes \(Z_{k,m}\). The remaining expectation is \[\int cJ^{-k}p_{\widetilde A}(\widetilde A S_1)^{-m} p^{\otimes t}(d\mathbf s)\gamma_{k,d}^{F}(d\widetilde A)=1\] by Corollary 16. This proves the normalization of (32). Here are the needed Gaussian determinant bounds. Gram–Schmidt expresses \(\det(G^{\mathsf T}G)\) as a product of independent chi-squares with degrees \(\nu=\ell,\ell-1,\ldots,\ell-m+1\); this standard Gaussian reduction is also described by Edelman and Rao (Edelman and Rao 2005, sec. 5). The smallest degree is \(3k+1>k\). Since \(k\) is even, direct integration of the chi-square density gives \[\mathbb E(\chi_\nu^2)^{-k/2} =\prod_{j=1}^{k/2}(\nu-2j)^{-1}\le(2k)^{-k/2}.\] In particular \(Z_{k,m}<\infty\). The same densities give absolute logarithmic integrability, and Jensen gives \(\mathbb E\log\det(G^{\mathsf T}G)\le m\log\ell\). Consequently \[ \log Z_{k,m}+\frac{k}{2}\mathbb E\log\det(G^{\mathsf T}G)\le Ckm. \tag{34}\] Before taking logarithms of the test, we control the positive logarithm of its projection density. For a fixed tuple put \[\delta(s')=\lVert P_{F^\perp}(s'-S_1)\rVert.\] The ball-average density version in Theorem 15 holds at \(\widetilde A S_1\) under the independent-tuple reference and hence under the replica law. The coordinates of \(\widetilde A(s'-S_1)\) are independent centered Gaussians with variance \(\delta(s')^2\). A ball has probability at most its volume times the maximum Gaussian density. Apply this bound to the ball averages and then use Fatou’s lemma: \[ \mathbb E_{\widetilde A}p_{\widetilde A}(\widetilde A S_1) \le (2\pi)^{-k/2}\int\delta(s')^{-k}\,p(ds'). \tag{35}\] Here and below such conditional estimates hold for almost every tuple. The affine space \(S_1+F\) has dimension \(m\), so (23) gives \[p\{\delta\le r\}\le LC^d r^{d-1-m},\qquad 0<r\le1.\] Since \(d-1-m>k\), tail integration bounds the inverse moment uniformly over tuples for fixed \(d,L\). Thus \(\log^+p_{\widetilde A}(\widetilde A S_1)\) is integrable. The projected law has divergence \(K'\le K<\infty\) from the independent-signal reference. Normalization of the test and (19) imply that \((\log\Lambda)^+\) is integrable. Its defining identity is \[ \begin{aligned} \log\Lambda &=\log c-\log Z_{k,m}-k\log J -\frac{k}{2}\log\det(G^{\mathsf T}G)\\ &\qquad-m\log p_{\widetilde A}(\widetilde A S_1). \end{aligned} \tag{36}\] The negative part of this expression is bounded by a constant and the positive logarithms of \(J\), the Gaussian determinant, and the last density. Those are integrable by Lemma 18, the chi-square calculation, and (35). Thus \(\log\Lambda\) is absolutely integrable. Solving (36) for the last logarithm proves its absolute integrability as well. We may now apply the entropy test (18) to obtain \(K'\ge\mathbb E\log\Lambda\). Subtract this bound from (31). Equation (36) displays the exact cancellation of \(-k\mathbb E\log J\), and (34) bounds the remaining determinant terms by \(Ckm\). This gives (33). ◻ Comparing the remaining projection densitiesBoth the actual \(A\) and the observable \(\widetilde A\) annihilate \(F\). The preceding lemma has removed the affine volume from the information bound. It remains to compare the two densities in (33). We first bound the conditional mean of the new density by a sum of averages of the actual density. The radii in that sum will depend only on \(A\), allowing us to average the actual label with its original density \(p_A\). Lemma 20 (Dyadic density comparison). In the replica experiment with the independent side matrix and \(\widetilde A=C_0B\) defined above, \[ \mathbb E\!\left[\log p_{\widetilde A}(\widetilde A S_1)-\log p_A(Y)\right] \le Ck+C\log(2+\log L). \tag{37}\] Proof. Fix the tuple and its actual \((A,Y)\), and again write \(\delta(s')=\lVert P_{F^\perp}(s'-S_1)\rVert\). Conditional on these data, \(\widetilde A\) has law \(\gamma_{k,d}^{F}\). We refine the inverse-moment bound (35) by using \(A\). Choose the deterministic integer \[ J_0=1+\left\lceil \frac{\log_2 L+C_1d}{d-1-m-k}\right\rceil \tag{38}\] with a sufficiently large absolute \(C_1\). The denominator is bounded below by a positive absolute multiple of \(d\). The contribution of \(\delta\le2^{-J_0}\) to the inverse moment, including the Gaussian constant, is at most \[(2\pi)^{-k/2} \sum_{j\ge J_0}2^{(j+1)k}LC^d2^{-j(d-1-m)}\le1.\] For the remaining distances use the bins \(\delta\in(r/2,r]\), where \(r=2,1,\ldots,2^{-J_0+1}\). These cover the remaining values because \(\delta(s')\le\lVert s'-S_1\rVert\le2\). Since \(A\) annihilates \(F\), \[\lVert As'-Y\rVert\le\lVert A\rVert_{\mathrm{op}}\,\delta(s').\] Let \(v_k\) be the unit \(k\)-ball volume, and let \(\overline p_A^{\,a}(y)\) be the average of \(p_A\) over the radius-\(a\) ball centered at \(y\). The contribution of one bin is at most \[(2/r)^k p\{\delta\le r\} \le 2^k v_k\lVert A\rVert_{\mathrm{op}}^{k} \overline p_A^{\,\lVert A\rVert_{\mathrm{op}}r}(Y).\] Integrating \(e^{-\lVert z\rVert^2/2}\) on the radius-\(\sqrt{k}\) ball gives \(v_k\le(2\pi e/k)^{k/2}\). Hence \[ \mathbb E_{\widetilde A}p_{\widetilde A}(\widetilde A S_1) \le1+\left(\frac{C\lVert A\rVert_{\mathrm{op}}}{\sqrt{k}}\right)^k \sum_r\overline p_A^{\,\lVert A\rVert_{\mathrm{op}}r}(Y). \tag{39}\] The right side depends on the tuple only through \((A,Y)\). Divide by the finite positive density \(p_A(Y)\) and apply Jensen first to \(\widetilde A\); all logarithms are integrable by Lemma 19. Only after obtaining (39) do we average over \(Y\) given \(A\). Its law is then the original \(p_A(y)\,dy\), and for every radius fixed by \(A\), \[\mathbb E_{Y\mid A}\frac{\overline p_A^{\,a}(Y)}{p_A(Y)}\le1,\qquad \mathbb E_{Y\mid A}\frac1{p_A(Y)} \le v_k\lVert A\rVert_{\mathrm{op}}^{k}.\] The first inequality integrates an averaged probability density over \(\{p_A>0\}\), and the second is the volume of that set, contained in the radius-\(\lVert A\rVert_{\mathrm{op}}\) ball. A second application of Jensen thus bounds the expected log ratio by \[ k\mathbb E\log\!\left(C\left(1+\frac{\lVert A\rVert_{\mathrm{op}}}{\sqrt{k}}\right)\right) +\log(J_0+2). \tag{40}\] This order of averaging is essential: before the tuple has disappeared, its conditional label need not have density \(p_A\) under further conditioning. For completeness, choose \(1/4\)-nets of the unit spheres in \(\mathbb R^d\) and \(\mathbb R^k\), with at most \(9^d\) and \(9^k\) points. The operator norm is at most twice the maximum absolute Gaussian bilinear form on these nets. The expected maximum of \(N\) absolute standard normals is at most \(\sqrt{2\log(2N)}\), by the exponential-moment bound; their independence is not required. Thus \(\mathbb E\lVert A\rVert_{\mathrm{op}}\le C\sqrt d\). Since \(k\) is a fixed positive fraction of \(d\), Jensen bounds the first term of (40) by \(Ck\). The cutoff (38) gives \(\log(J_0+2)\le C\log(2+\log L)\), proving the lemma. ◻ Proof of Theorem 1. Generate \(W\) conditionally on \((A,Y)\) independently of the replicas. Each pair \((S_i,W)\) has the original pair law, and the side matrix \(B\) is independent of the whole construction. Lemma 13, followed by Lemmas 19 and 20, gives \[tI(S;W\mid B,BS) \le H(W)+Ckm+m\bigl(Ck+C\log(2+\log L)\bigr).\] Here the representative signal is coupled independently with \(B\), exactly as in the original block experiment. Since \(t=m+1\) and \(k\) is proportional to \(d\), division by \(t\) proves (1). ◻ Shared data and distance-bin comparisonsThe determinant test in the preceding section pays a double logarithm of the largest prior density. We instead encode the dyadic scale of each projected simplex height. Under a prior bounded by \(e^b\), these scales have entropy \(O(1+b/d)\). The projected correlation recovers the height and label-density terms, leaving this scale cost. Since the chosen scale depends on the side matrix, the central step compares it with a mixture over fixed scales before averaging the matrix. Throughout this section, for sufficiently large \(d\), set \[ k=\lfloor d/16\rfloor,\qquad q=\lfloor d/8\rfloor,\qquad p=\lfloor d/2\rfloor. \tag{41}\] Here \(k\) is the number of actual rows, \(q\) is the number of replicas, and \(p\) is the number of side rows. In particular \(k\ge d/32\), \(q\ge d/16\), \(q\ge2\), and \(p-i+2\ge2k\) for \(2\le i\le q\). Theorem 21 (Information from shared exact data). Let \(P=f\sigma_d\) be a probability with a nonnegative Borel density \(f\le e^b\), where \(b\ge0\). Let \(S\sim P\), let \(A\) be an independent \(k\times d\) standard Gaussian matrix, and put \(Y=AS\). Let a finite message \(W\) be drawn from a kernel of \((A,Y)\), conditionally independently of \(S\). Let \(B\) be a \(p\times d\) standard Gaussian matrix independent of the entire joint experiment producing \((S,A,Y,W)\). Then \[ I(S;W\mid B,BS)\le\frac{H(W)}q+C\left(d+\frac bd\right) \tag{42}\] for an absolute \(C\), uniformly over the prior and message kernel. The order of (42) also follows from Theorem 1: append \(k\) ignored fresh rows to \(A\) and first use the initial \(8k\) rows of \(B\). The Gaussian theorem gives the entropy denominator \(2k+1\ge q\), and \(8k\le p\). Adding the remaining independent side rows can only decrease the conditional information, because independence of the side matrix gives \[I(S;W\mid B,BS)=I(S;W)-I(W;B,BS).\] Finally \(\log(2+b)\le\log(2+d)+b/d\le Cd+b/d\). The direct proof below identifies the entropy of the chosen distance scales as the prior cost. Projected height binsDraw the independent side matrix \(B\) and put \(C_i=BS_i\). Conditional on the tuple, the image of the direction space of \(F_i\) has rank \(i-2\) almost surely. The orthogonal residual of \(S_i\) from \(F_i\) has length \(\rho_i\). Its image under \(B\) is an \(N(0,\rho_i^2 I_p)\) vector, independent of the restriction of \(B\) to that direction space. Projecting off the image of the direction space shows that \[ T_i:=\mathop{\mathrm{dist}}(C_i,BF_i)\ \stackrel{\mathrm{law}}=\ \rho_i\chi_{m_i}, \qquad m_i=p-i+2, \tag{48}\] conditional on the tuple. Here \(\chi_{m_i}\) is a chi variable. In particular \(m_i\ge2k\). Define the integer bin \[K_i=\left\lceil\log_2(T_i/\sqrt{m_i})\right\rceil,\] with any fixed value at \(T_i=0\), a null event. Use the reference probabilities \(\pi_j=2^{-|j|}/3\) for \(j\in\mathbb Z\); they sum to one. The chi-square second moment and the Gaussian small-ball bound show uniformly in \(m\ge1\) that \[\mathbb E\log(\chi_m/\sqrt m)\le0,\qquad \mathbb E|\log(\chi_m/\sqrt m)|\le C.\] For the first inequality use Jensen on the square. For the second, its positive part follows from the second moment, while the negative part follows by integrating \(\mathbb P\{\chi_m/\sqrt m<e^{-t}\}\le\min\{1,(Ce^{-t})^m\}\). Together with Lemma 22, this gives \[ (\log2)\mathbb EK_i\le\mathbb E\log\rho_i+\log2,\qquad H(K_i)\le\mathbb E[-\log\pi_{K_i}] \le C\left(1+\frac bd\right). \tag{49}\] The entropy bound is the elementary cross-entropy inequality, since \(-\log\pi_j=\log3+|j|\log2\) and the expected absolute bin index is bounded. For fixed \(B\), let \(P_B\) be the law of \(BS'\) for \(S'\sim P\). Given \((B,C_1,\ldots,C_{i-1})\), the affine set \(BF_i\) is known. Let \(E_j\) be its tube of radius \(\sqrt{m_i}\,2^j\). The bin event \(K_i=j\) is contained in \(E_j\). Since the marginal of \(C_i\) given \(B\) is \(P_B\), conditional relative entropy followed by the bin map yields \[ I(C_i;C_1,\ldots,C_{i-1}\mid B) \ge-H(K_i\mid B,C_1,\ldots,C_{i-1})-\mathbb E\log P_B(E_{K_i}). \tag{50}\] To justify the logarithms, coordinatewise data processing makes the projected total correlation finite by (47). Each conditional divergence in its chain rule is therefore finite. Its image under the bin map is finite as well. The expected negative logarithm of the reference bin probability is that image divergence plus the finite conditional bin entropy. Since the tube probability is at least the bin probability, \(-\log P_B(E_{K_i})\) is integrable. This also proves (50) without an undefined entropy difference. A fixed mixture before averaging the side matrixWe will prove the one-height estimate \[ \mathbb E\log P_B(E_{K_i}) \le k\mathbb E\log\rho_i-h_k+C\left(k+1+\frac bd\right). \tag{51}\] Together with (50), this recovers the terms \(h_k-k\mathbb E\log\rho_i\) in the unprojected correlation bound. The selected bin \(K_i\) depends on \(B\). We first bound every fixed bin using a reference density that depends only on \(A\), so that the eventual label average uses the original density \(g_A\). Fix the shared-data tuple and its \((A,Y)\), and fix \(i\). For a candidate \(s'\), put \(\rho'=\mathop{\mathrm{dist}}(s',F_i)\). Its distance to \(BF_i\) after projection has law \(\rho'\chi_{m_i}\), by the same calculation as (48). The Gaussian small-ball bound gives \[\mathbb E_B P_B(E_j) \le\int\min\left\{1,\left(\frac{4\cdot2^j}{\rho'}\right)^{m_i}\right\}P(ds').\] The minimum is one at zero denominator. Since \(A\) is constant with value \(Y\) on \(F_i\), \(\lVert As'-Y\rVert\le\lVert A\rVert_{\mathrm{op}}\rho'\). Pushing \(P\) forward by \(A\) therefore gives the further bound \[ \mathbb E_B P_B(E_j) \le\int\min\left\{1, \left(\frac{4\cdot2^j\lVert A\rVert_{\mathrm{op}}}{\lVert y'-Y\rVert}\right)^{m_i}\right\} g_A(y')\,dy'. \tag{52}\] For \(m>k\) and \(\beta>0\), radial integration in \(\mathbb R^k\) gives the exact kernel mass \[ \int_{\mathbb R^k}\min\{1,(\beta/\lVert u\rVert)^m\}\,du =v_k\beta^k\frac{m}{m-k}. \tag{53}\] Indeed the inside ball contributes \(v_k\beta^k\), and the exterior contributes \(kv_k\beta^k/(m-k)\). Let \(a_{A,j}\) be the probability density obtained by convolving \(g_A\) with this normalized kernel, using \(m=m_i\) and \(\beta=4\cdot2^j\lVert A\rVert_{\mathrm{op}}\). Since \(m_i\ge2k\) and \(v_k\le(2\pi e/k)^{k/2}\), (52) becomes \[ 2^{-kj}\mathbb E_B P_B(E_j) \le\left(\frac{C\lVert A\rVert_{\mathrm{op}}}{\sqrt k}\right)^k a_{A,j}(Y). \tag{54}\] For each fixed \(i\), this density depends on \(A\) and \(j\), but not on the replica tuple and not on \(Y\) as a parameter of the density. Define \[a_A=\sum_{j\in\mathbb Z}\pi_j a_{A,j},\qquad T_B=\sum_{j\in\mathbb Z}\pi_j2^{-kj}P_B(E_j).\] The first is a probability density depending only on \(A\) and the fixed index \(i\). Conditional on the shared-data outcome, (54) gives \[\mathbb E_B T_B\le \left(\frac{C\lVert A\rVert_{\mathrm{op}}}{\sqrt k}\right)^k a_A(Y).\] The \(j=K_i\) summand gives the pointwise inequality \[ \log P_B(E_{K_i}) \le kK_i\log2-\log\pi_{K_i} +\log T_B. \tag{55}\] Before averaging logarithms, we verify their integrability. If \(Z_{A,j}=v_k(4\cdot2^j\lVert A\rVert_{\mathrm{op}})^k m_i/(m_i-k)\), then for \(y\) in the support of \(g_A\), \[\frac{\pi_0}{Z_{A,0}}\le a_A(y)\le\sup g_A.\] The lower bound holds because the support has diameter at most \(2\lVert A\rVert_{\mathrm{op}}\), so the unnormalized kernel at \(j=0\) equals one throughout that support. The upper bound follows by convolving a bounded density with a probability density. The positive log of \(\sup g_A\) is integrable by the uniform projected-density formula and the Gaussian log-determinant bound. The absolute log of \(Z_{A,0}\) is integrable: \(\log^+\lVert A\rVert_{\mathrm{op}}\) is controlled by its finite mean, and \(\log^-\lVert A\rVert_{\mathrm{op}}\le\log^-\lVert Au_0\rVert\) for any fixed unit \(u_0\), whose norm is chi-distributed. Thus \(\log a_A(Y)\) is absolutely integrable. The pointwise inequality (55) bounds \(\log^-T_B\) by the sum of \(-\log P_B(E_{K_i})\), \(k|K_i|\log2\), and \(-\log\pi_{K_i}\), all integrable. For the positive part, \(\mathbb E_B\log^+T_B\le\log(1+\mathbb E_BT_B)\), and (54) bounds its right side using \(a_A(Y)\) and \(\lVert A\rVert_{\mathrm{op}}\) as above. Thus \(\log T_B\) is absolutely integrable as well. Conditional Jensen now gives \[\mathbb E_B\log T_B\le k\log\left(\frac{C\lVert A\rVert_{\mathrm{op}}}{\sqrt k}\right)+\log a_A(Y).\] The right side depends only on \((A,Y)\). In the full experiment \(Y\mid A\) has density \(g_A\), and nonnegativity of \(D_{\mathrm{KL}}(g_A\Vert a_A)\) therefore gives \[ \mathbb E\log a_A(Y)\le\mathbb E\log g_A(Y)=-h_k. \tag{56}\] This is where the fact that \(a_A\) depends only on \(A\) is needed. Also \(\mathbb E\log(C\lVert A\rVert_{\mathrm{op}}/\sqrt k)\le C\) by the Gaussian norm estimate in Section 4, since \(k\ge d/32\). Consequently \(\mathbb E\log T_B\le Ck-h_k\). Combining this with (55) and (49) proves (51). Insert (51) in (50), use the entropy bound in (49), and sum the conditional divergences over \(i=2,\ldots,q\). The projected correlation satisfies \[ \mathop{\mathrm{tc}}(C_1,\ldots,C_q\mid B) \ge(q-1)h_k-k\sum_{i=2}^q\mathbb E\log\rho_i -C(q-1)\left(k+1+\frac bd\right). \tag{57}\] Proof of Theorem 21. Generate \(W\) from the shared data independently of the replicas conditional on \((A,Y)\). Each replica has the original signal–message pair law, and \(B\) is independent of this coupling. The finite correlation bound (47) therefore permits Lemma 13 with \(t=q\): \[qI(S;W\mid B,BS) \le H(W)+\mathop{\mathrm{tc}}(S_1,\ldots,S_q)-\mathop{\mathrm{tc}}(C_1,\ldots,C_q\mid B).\] Subtract (57) from (47). The measurement entropy and all height logarithms cancel. Since \(k\le d/16\), division by \(q\) gives (42). Section 6 applies this bound while conditioning only on the current state and retaining its separate late elimination of the clock cost. ◻ Current-state recurrences and the streaming boundWe now complete the streaming applications of the Gaussian determinant test and the distance-bin test through two current-state iterations. Both use one independent side matrix throughout the stream. They condition on the present padded state, apply a complete bounded-prior block bound in that conditional experiment, and only then average over state values. Thus their information upper bounds never subtract two possibly infinite averages of conditional total correlations. The Gaussian replica clockProposition 23 (Current-state iteration of Gaussian replicas). Let \(M(d)=o(d^2)\) be nonnegative and integer-valued, and let \(0<\epsilon(d)\le1/10\). A learner in Definition 3 whose uniform-prior angular success is at least \(3/5\) satisfies \[T(d)\ge c\,d\log(1/\epsilon(d))\] for an absolute \(c>0\) and all sufficiently large \(d\). The threshold may depend on the memory sequence. Proof. Use the first step of Lemma 9 to choose a fixed shared-seed conditional experiment with success at least \(3/5\) and then take its Borel replacement. Its initialization is independent of the signal and all rows; fresh transition and output randomness remains in the kernels. This selection uses the original jointly measurable experiment and does not require a jointly measurable family of Borel replacements. Put \(\beta=\log(1/\epsilon)\). If \(T>d\beta\), the conclusion is immediate. On the remaining branch, Lemmas 9 and 11 give padded states of entropy at most \[ h_T=M\log2+\log(T+2)=o(d^2), \qquad \beta=o(d),\qquad T=o(d^2). \tag{58}\] Let \[k=2\lfloor d/16\rfloor,\qquad t=k+1,\qquad \ell=4k,\qquad n=\lceil T/k\rceil.\] Append ignored fresh samples so that the padded stream has \(n\) complete blocks. Let \(V_i\) be the padded state at the end of block \(i\), with \(V_0\) the data-independent initial padded state. The latter includes a possible stopping decision at index zero. Draw a single standard Gaussian \(\ell\times d\) matrix \(B\) independently of the entire selected learner experiment and all appended rows, and put \(Z=BS\). Consider a positive-probability value \(V_i=v\), and put \(p_v=\mathbb P\{V_i=v\}\). By Lemma 10, its conditional signal law has a Borel density \(f_v\) satisfying \[f_v\le L_v:=1/p_v \quad\hbox{relative to }\sigma_d,\qquad \sum_v p_v\log L_v=H(V_i).\] The next block matrix is independent of the pair \((S,V_i)\): the state uses only the signal, completed rows, and learner randomness, and the new rows are fresh. Given \(v\), the next state is a finite kernel of the new matrix \(A\) and its exact label \(AS\). The side matrix \(B\) remains independent of this entire conditional signal, block, and message experiment. Therefore Theorem 1 applies with \(L=L_v\) and with message \(V_{i+1}\). Apply that numerical bound at each \(v\) before averaging. Concavity of \(x\mapsto\log(2+x)\) and the entropy bound give \[\begin{align*} I(S;V_{i+1}\mid V_i,B,Z) &\le \frac{H(V_{i+1}\mid V_i)}{t} +Cd+C\sum_vp_v\log(2+\log L_v) \\ &\le \frac{h_T}{t}+Cd+C\log(2+H(V_i)) \\ &\le \frac{h_T}{t}+Cd+C\log(2+h_T). \tag{59}\end{align*}\] All conditional uses have bounded priors and finite conditional replica correlation, as established in the block theorem. Since \(t\asymp d\) and \(h_T=o(d^2)\), the last line is at most \(C_1d\) for all sufficiently large \(d\), with one absolute \(C_1\). Write \(\mathcal I(V)=I(S;V\mid B,Z)\). The chain rule and data processing give \[\mathcal I(V_{i+1}) \le \mathcal I(V_i)+I(S;V_{i+1}\mid V_i,B,Z).\] Here \(\mathcal I(V_0)=0\) by the independence of initialization and the index-zero decision from \((S,B,Z)\). Iterating (59) yields \[ \mathcal I(V_n)\le C_1dn. \tag{60}\] The side row count satisfies \(\ell=8\lfloor d/16\rfloor\le d/2\). The postponed output is a kernel of \(V_n\) with fresh randomness, so (7) gives \[ \mathcal I(V_n)\ge \tfrac12(d-\ell-1)\log\frac1{5\epsilon}-\log2 \ge c_1d\beta \tag{61}\] for large \(d\). Lemma 12 also gives \(T>\lfloor d/2\rfloor\). In particular \(T>0\), and for \(d\ge32\) the inequality \(k\ge d/16\) gives the explicit clock comparison \[dn\le d(T/k+1)\le16T+d\le18T.\] Combining this with (60) and (61) proves the proposition. ◻ The distance-bin clockProposition 24 (Current-state iteration of distance bins). Under the hypotheses of Theorem 2, the distance-bin block estimate (42) also gives \(T(d)\ge c\,d\log(1/\epsilon(d))\) for an absolute \(c>0\) and all sufficiently large \(d\). Proof. Apply Lemma 9 with \(p=2/3\) and \(p'=3/5\). It first fixes a shared seed using the original jointly measurable experiment, then chooses Borel versions in that conditional experiment and fixes its independent initialization and remaining fresh randomness. We obtain deterministic Borel state-based rules and an output with uniform-prior success at least \(3/5\). Pad stopping and postpone this output. The final output is a function of the final tagged state. For this route keep the convenient clock bound \[ \mathfrak m_T=M\log2+\log(T+3). \tag{62}\] Indeed, the exact count in Lemma 9 is \((T+2)2^M\), so the entropy of every padded state is at most \(h_T\le\mathfrak m_T\). Use \[k=\lfloor d/16\rfloor,\qquad q=\lfloor d/8\rfloor,\qquad p=\lfloor d/2\rfloor,\qquad n=\lceil T/k\rceil,\] and append ignored fresh rows through the last block. Let \(V_0,\ldots,V_n\) be the block-end states, starting with the data-independent tagged initial state. Independently of the entire deterministic learner experiment and all rows, draw one \(p\times d\) standard Gaussian matrix \(B\). It is used only for analysis. Put \(Z=BS\) and \(\mathcal D(V)=I(S;V\mid B,Z)\). For a block from \(V\) to \(W\), condition on a value \(v\) with probability \(p_v>0\). By Lemma 10, the conditional signal density is bounded by \(\exp(b_v)\), where \(b_v=\log(1/p_v)\). The next standard Gaussian matrix \(A\) is independent of \((S,V)\), the ending state is a function of \((v,A,AS)\), and \(B\) remains independent of the entire conditional experiment. Apply Theorem 21 at this fixed \(v\) and then average. Since \(\sum_vp_vb_v=H(V)\), \[\begin{align*} I(S;W\mid V,B,Z) &\le \frac{H(W\mid V)}q+C\left(d+\frac{H(V)}d\right) \\ &\le C_2\left(d+\frac{\mathfrak m_T}{d}\right). \tag{63}\end{align*}\] The last inequality uses \(q\ge d/16\) for large \(d\). In this calculation the prior cost \(b_v/d\) is averaged linearly; no entire state history is retained. The chain rule gives \(\mathcal D(W)-\mathcal D(V)\le I(S;W\mid V,B,Z)\), and \(\mathcal D(V_0)=0\). Consequently \[ \mathcal D(V_n)\le C_2 n\left(d+\frac{\mathfrak m_T}{d}\right). \tag{64}\] For the exact side row count \(p=\lfloor d/2\rfloor\), the independent residual-sphere endpoint (9) gives \[ \mathcal D(V_n)\ge \tfrac12(d-p-1)\log\frac1{4\epsilon}-\log2 \ge c_2d\beta,\qquad \beta=\log(1/\epsilon), \tag{65}\] for sufficiently large \(d\). This is the direct \(90/d\) radius estimate and the event of mass at least \(1/2\) used in the distance-bin route. The corresponding raw-data bound in Lemma 12 also yields \(T>\lfloor d/2\rfloor\), equivalently \(T>d/2\) for integer \(T\). The zero-row experiment in that lemma has residual radius one. If \(T\ge d\beta\), the asserted lower bound is immediate. Otherwise \(\mathcal D(V_n)\le H(V_n)\le\mathfrak m_T\) and (65) give the route’s clock inequality \[c_2d\beta\le M\log2+\log(T+3) \le M\log2+\log(d\beta+3).\] Since \(d\beta+3\le(d+3)(\beta+1)\) and \(\log(1+\beta)\le\beta\), this implies, for large \(d\), \[ \beta\le \frac{C(M+\log d+1)}d=o(d). \tag{66}\] Thus \(T<d\beta=o(d^2)\) and \(\mathfrak m_T=o(d^2)\). The right side of (64) is consequently at most \(C_3nd\). For \(d\ge32\), \(k\ge d/32\), and \(T>d/2\) gives \[dn\le32T+d\le34T.\] Comparing the resulting upper bound \(\mathcal D(V_n)\le C_4T\) with (65) proves the proposition. ◻ Proof of Theorem 2. Its uniform-prior success \(2/3\) is at least the \(3/5\) required by Proposition 23, which proves the theorem. Proposition 24 gives a second proof using the linear entropy cost. The later sections give three further proofs: the Haar current-state argument in Proposition 31, and the synthetic and incidence full-history arguments in Proposition 35 and Theorem 44. Each uses the same learner model, while retaining its own side law, conditional experiment, and clock estimate. ◻ Haar projections and posterior replica informationThis comparison uses Haar frames rather than the Gaussian side matrices of Sections 4 and 5. We express the replica measure by successive simplex heights. The normalization changes at each insertion because the remaining ambient dimension shrinks. A projected-distance distribution-function test, valid even with atoms, recovers a fraction of the height cost. We choose its exponent below one as a function of \(d\) and keep the resulting entropy remainder explicit. Proposition 29 gives that remainder in terms of the state entropies. In particular, the block cost is \(O(d)\) when all state alphabets have logarithmic width \(o(d^2)\). We first prove the fixed-prior height identity and cancellation, then average over state values. A Gaussian observation block is converted to its orthonormal row frame only in that final step; the learner’s observations are still the original exact Gaussian pairs. Write \(\sigma=\sigma_d\) and \(\log^+x=\max\{0,\log x\}\) for \(x>0\). A \(j\)-frame is a \(j\times d\) matrix with orthonormal rows; its Haar law is the probability law invariant under orthogonal changes of ambient coordinates. Throughout the section set \[ d\ge64,\qquad k=r=\lfloor d/16\rfloor,\qquad m=r+1, \qquad \ell=\lfloor d/4\rfloor. \tag{67}\] Here \(k\) is the actual block row count, \(r\) is the number of height insertions, \(m=r+1\) is the replica count, and \(\ell\) is the side-frame row count. At the \(i\)th insertion, \(1\le i\le r\), the two dimensions that will matter are \[ n_i=d-i+1,\qquad q_i=\ell-i+1. \tag{68}\] They satisfy \[ n_i>k+2,\qquad q_i\ge\frac{3d}{16}>k, \qquad \frac{q_i}{q_i-k}\le\frac32, \qquad \frac{n_i}{q_i}\le\frac{16}{3},\qquad \frac d k\le32. \tag{69}\] Indeed, \(q_i\ge\ell-r+1\ge3d/16\) and \(k\le d/16\); also \(k\ge d/32\) for \(d\ge64\). The other inequalities follow at once. In particular, none of the denominators in (69) vanishes. An exact identity for the replica measureWe begin with the normalization of spherical projection. For integers \(1\le j<n-2\), put \[\alpha_{n,j}=\frac{\Gamma(n/2)}{\pi^{j/2}\Gamma((n-j)/2)}, \qquad v_j=\frac{\pi^{j/2}}{\Gamma(j/2+1)}.\] Here \(v_j\) is the volume of the unit ball in \(\mathbb R^j\). The first \(j\) coordinates of a uniform point of \(S^{n-1}\) have density \[ p_{n,j}(u)=\alpha_{n,j}(1-\lVert u\rVert^2)^{(n-j-2)/2} \mathbf 1_{\{\lVert u\rVert<1\}}. \tag{70}\] Conditional on these coordinates, the remaining direction is uniform, with radius \(\sqrt{1-\lVert u\rVert^2}\). These are the density and conditional law in Lemma 14, specialized to ambient dimension \(n\) and a matrix with orthonormal rows. We keep the ambient dimension in the notation because it changes at each insertion. We will also use the elementary estimates \[ \alpha_{n,j}\le n^{j/2},\qquad v_j\alpha_{n,j}\le (Cn/j)^{j/2}, \tag{71}\] where \(C\) is absolute. For completeness, gamma recursion and \(\Gamma(z+1/2)\le\sqrt z\,\Gamma(z)\) imply \(\Gamma(n/2)/\Gamma((n-j)/2)\le(n/2)^{j/2}\); the half-step inequality is Cauchy–Schwarz in the gamma integral. Integrating that integral over \([j/2,j/2+1]\) gives \(\Gamma(j/2+1)\ge(j/2)^{j/2}e^{-j/2-1}\). These two bounds prove (71). Let \(P=g\sigma\) be a probability measure with a bounded Borel density \(0\le g\le K<\infty\). Necessarily \(K\ge1\). Draw an independent Haar \(k\)-frame \(A\) and \(S\sim P\), and write \(Y=AS\). Under \(\sigma\), the density of \(Y\) is \[p^0(y)=\alpha_{d,k}(1-\lVert y\rVert^2)^{(d-k-2)/2} \mathbf 1_{\{\lVert y\rVert<1\}}.\] For \(\lVert y\rVert<1\), let \(U_{A,y}\) be uniform probability on the sphere \(\{s\in S^{d-1}:As=y\}\). On \(\lVert y\rVert\ge1\), choose any fixed measurable probability kernel for \(U_{A,y}\); it will always be multiplied by \(p^0(y)=0\). Fix the following density and posterior versions: \[ \begin{split} p_A(y)&=p^0(y)\int g\,\mathrm dU_{A,y} \quad (\lVert y\rVert<1),\qquad p_A(y)=0\quad (\lVert y\rVert\ge1),\\ P_{A,y}(\mathrm ds)&= \frac{g(s)}{\int g\,\mathrm dU_{A,y}}\,U_{A,y}(\mathrm ds) \quad\text{when }p_A(y)>0. \end{split} \tag{72}\] The posterior on the unused set \(\{p_A=0\}\) may be any fixed measurable probability kernel. These are measurable versions. For example, a fiber direction can be generated by projecting a standard Gaussian onto \(\ker A\) and normalizing it. If \(u\) is that unit direction, the fiber point is \(A^{\mathsf T}y+\sqrt{1-\lVert y\rVert^2}\,u\). A fixed coordinate convention handles the null normalization failures. This construction also makes \(p_A(y)\) jointly measurable in \((A,y)\). We have \(p_A\le Kp^0\) and, for every \(A\), \(p_A\) is a probability density. For a tuple \(\boldsymbol s=(s_0,\ldots,s_r)\), define \[D_i=\operatorname{span}\{s_1-s_0,\ldots,s_i-s_0\},\qquad D_0=\{0\}, \qquad \delta_i=\operatorname{dist}(s_i-s_0,D_{i-1}).\] Thus \(\delta_i\) is the height added by the \(i\)th point. Write \(\mathrm dA_D\) for Haar probability on the \(k\)-frames in \(D^\perp\). All subspaces arising here have \(\dim D^\perp\ge d-r>k\). A measurable version of this frame kernel is obtained by projecting independent standard Gaussians onto \(D^\perp\) and applying a fixed Gram–Schmidt rule. We use \(\mathrm dA\) when \(D=\{0\}\). Lemma 25 (Haar simplex-height identity). For the dimensions in (67) and every probability \(P=g\sigma\) with \(0\le g\le K<\infty\), the following finite measures on the frame and the whole tuple are equal: \[ \mathrm dA\int_{\mathbb R^k}p_A(y)^m\,\mathrm dy \prod_{i=0}^r P_{A,y}(\mathrm ds_i) =\left(\prod_{i=0}^rP(\mathrm ds_i)\right) \left(\prod_{i=1}^r\alpha_{d-i+1,k}\delta_i^{-k}\right) \mathrm dA_{D_r}. \tag{73}\] Equality means equality after integration against every bounded Borel function of \((A,\boldsymbol s)\). Define the displayed product density to be zero on degenerate tuples. The measure on the right is supported on tuples with \(\dim D_i=i\) for all \(1\le i\le r\). Proof. We first take \(P=\sigma\). Fix \(s_0,\ldots,s_{i-1}\) with \(\dim D_{i-1}=i-1\), and put \(H=D_{i-1}^\perp\), so \(\dim H=n_i=d-i+1\). For a bounded continuous function \(\varphi(A,s)\), consider \[ I_t=\int\mathrm dA_{D_{i-1}}\int\sigma(\mathrm ds)\, \varphi(A,s) \frac{\mathbf 1_{\{\lVert A(s-s_0)\rVert\le t\}}}{v_kt^k}, \qquad t>0. \tag{74}\] We compute its limit as \(t\downarrow0\) in both orders of integration. Fix \(A\) first. Disintegration of \(\sigma\) under \(s\mapsto As\) rewrites the inner integral as the average over the ball of radius \(t\) about \(As_0\) of \[y\longmapsto p^0(y)\int\varphi(A,s)\,U_{A,y}(\mathrm ds).\] This function is continuous in the open unit ball. Indeed, use the fiber representation following (72) with a common uniform direction and apply bounded convergence. Because \(d>k+2\), the factor \(p^0\) tends to zero at the boundary, so the product extends continuously by zero to \(\mathbb R^k\). The average therefore tends to \[ p^0(As_0)\int\varphi(A,s)\,U_{A,As_0}(\mathrm ds). \tag{75}\] Its absolute value, and that of every preceding average, is at most \(\lVert\varphi\rVert_\infty\alpha_{d,k}\). Dominated convergence thus permits the frame integration as well. For the other order, fix \(s\) and put \(\delta=\lVert\operatorname{proj}_H(s-s_0)\rVert\). When \(\delta>0\), the vector \(A(s-s_0)\) under \(\mathrm dA_{D_{i-1}}\) has density \[ z\longmapsto\delta^{-k}\alpha_{n_i,k} \left(1-\frac{\lVert z\rVert^2}{\delta^2}\right)^{(n_i-k-2)/2} \mathbf 1_{\{\lVert z\rVert<\delta\}}. \tag{76}\] This is (70) in the current ambient space \(H\), not in the original \(d\)-dimensional space. Since \(n_i>k+2\), the density is bounded by its value \(\alpha_{n_i,k}\delta^{-k}\) at zero. The conditional frame law also has a concrete limit. Put \(e=\operatorname{proj}_H(s-s_0)/\delta\) and \(w=z/\delta\). If \(C\) is a Haar \(k\)-frame in \(H\cap e^\perp\), then the conditional law of \(A\) given \(A(s-s_0)=z\), for \(\lVert w\rVert<1\), is the law of \[ A=w e^{\mathsf T}+(I_k-ww^{\mathsf T})^{1/2}C. \tag{77}\] Its rows are orthonormal and \(Ae=w\). Invariance under rotations fixing \(e\) identifies the remaining factor as Haar; conversely the displayed formula describes every such frame. The formula is continuous at \(w=0\). At zero the frame law is \(\mathrm dA_{D_i}\), since \(H\cap e^\perp=D_i^\perp\). It follows that the frame integral in (74) tends, for this \(s\), to \[ \alpha_{n_i,k}\delta^{-k} \int\varphi(A,s)\,\mathrm dA_{D_i}. \tag{78}\] The bound \(\lVert\varphi\rVert_\infty\alpha_{n_i,k}\delta^{-k}\) dominates the frame averages for every \(t\). It is integrable in \(s\). Choose a \((k+1)\)-dimensional subspace \(L\subset(D_{i-1}+\operatorname{span}\{s_0\})^\perp\); the available dimension is at least \(d-i\ge d-r>k+1\). Then \(\delta\ge\lVert\operatorname{proj}_L s\rVert\). The latter vector has a bounded density by (70), because \(d>k+3\). In \(k+1\) dimensions its inverse \(k\)th power is integrable at zero: the radial integral is a constant times \(\int_0^1 u^k u^{-k}\,\mathrm du<\infty\). This also shows that \(\delta>0\) for \(\sigma\)-almost every \(s\). Dominated convergence proves the second limit after integration in \(s\). Equating the two integrated limits in (75) and (78) gives the insertion identity. Bounded continuous functions determine finite Borel measures on the compact frame–sphere space, so the identity also holds for bounded Borel tests. The frame and fiber constructions above make the kernels measurable in the preceding tuple, allowing successive integration. Start with \(\sigma(\mathrm ds_0)\mathrm dA\) and insert, for \(i=1,\ldots,r\), the kernels \(p^0(As_0)U_{A,As_0}(\mathrm ds_i)\). Each insertion changes the frame law from \(\mathrm dA_{D_{i-1}}\) to \(\mathrm dA_{D_i}\) and contributes \(\alpha_{d-i+1,k}\delta_i^{-k}\). On the original side, disintegration of \(s_0\) contributes one more factor \(p^0(y)\), so its total power is \(r+1=m\). These measures are finite at every step because \(p^0\le\alpha_{d,k}\); in particular their final mass is \(\int p^0(y)^m\,\mathrm dy\le\alpha_{d,k}^{m-1}\). The preceding almost-sure positivity of every inserted height gives the full-rank support. This proves (73) for \(\sigma\). Finally multiply this equality of measures by the bounded Borel function \(\prod_{i=0}^r g(s_i)\). Formula (72) changes the left side into the stated measure for \(P\), while the right side becomes the product of \(P\) measures in (73). The bounded tilt preserves finiteness. No continuity of \(g\) or of \(p_A\) was used. ◻ We now pass from the finite measure to the actual posterior experiment. Draw \((A,Y)\) with law \(\mathrm dA\,p_A(y)\mathrm dy\), and, conditionally on \((A,Y)\), draw \(S_0,\ldots,S_r\) independently from \(P_{A,Y}\). Write \(Q\) for the law including the frame and tuple, and \(Q_S\) for its tuple marginal. Each \((A,S_i)\) has law \(\mathrm dA\,P(\mathrm ds)\), and all the replicas satisfy \(AS_i=Y\) almost surely. Define the probability measure \(\Lambda(\mathrm dA,\mathrm d\boldsymbol s) =P^{\otimes m}(\mathrm d\boldsymbol s)\mathrm dA_{D_r}\). The degenerate tuples have zero \(P^{\otimes m}\)-measure, by the positivity proved during insertion; fix any measurable convention there. Multiply (73) by \(p_A(As_0)^{-r}\) on its positive set, and by zero on its zero set. This operation is justified first for the minimum with a finite constant and then by monotone convergence. On the left it changes \(p_A^m\) to \(p_A\), since \(m=r+1\) and the used posterior is supported on \(As_0=y\). Consequently \[ \frac{\mathrm dQ}{\mathrm d\Lambda}(A,\boldsymbol s) =\left(\prod_{i=1}^r\alpha_{d-i+1,k}\delta_i^{-k}\right) p_A(As_0)^{-r} \tag{79}\] on the positive-density, full-rank set, with value zero elsewhere. For a fresh Haar \(\ell\)-frame \(B\), independent of this whole experiment, put \(Z_i=BS_i\) and let \(P_B\) be the image of \(P\) under \(B\). Define the two total correlations \[ J_S=D_{\mathrm{KL}}(Q_S\Vert P^{\otimes m}),\qquad J_Z=\mathbb E_B D_{\mathrm{KL}} (Q_{Z_0,\ldots,Z_r\mid B}\Vert P_B^{\otimes m}). \tag{80}\] Lemma 26 (Finite upper correlation). The correlations satisfy \(0\le J_Z\le J_S\le U<\infty\), where \[\begin{align*} U&=\sum_{i=1}^r\left[ \log\alpha_{d-i+1,k}+k\mathbb E_Q\log(1/\delta_i) -\mathbb E_{A,S\sim P}\log p_A(AS)\right], \tag{81}\\ 0\le U&\le Crd(\log K+\log d). \tag{82}\end{align*}\] Here \(\mathbb E_{A,S\sim P}\) draws the Haar frame independently of a vector with law \(P\). Every logarithm in (81) is integrable. Proof. The positive part of \(\log p_A\) under \(p_A\) is bounded uniformly in \(A\) by \(\log^+(K\alpha_{d,k})\). Its negative part is also integrable uniformly: the density is supported in a ball of finite volume, and \(-x\log x\le1/e\) for \(0<x\le1\). Comparison with uniform probability on the unit ball now gives \[ -\mathbb E_{A,S\sim P}\log p_A(AS)\le\log v_k. \tag{83}\] To control a height, condition on \((A,Y,S_0,\ldots,S_{i-1})\). Put \(h=\sqrt{1-\lVert Y\rVert^2}\) and \(n=d-k\). Almost surely \(h>0\). The conditional residual direction of \(S_i\) in \(\ker A\) has density at most \[K'=\frac{Kp^0(Y)}{p_A(Y)}\ge1\] relative to uniform probability on \(S^{n-1}\). Choose a unit vector \(e\in\ker A\) perpendicular to the preceding residual directions, including that of \(S_0\). There are at most \(i\le r<n\) such directions; a fixed Gram–Schmidt rule makes the choice measurable. Then \(e\perp D_{i-1}\). If \(U_i\) is the new unit residual direction, then \[\delta_i\ge |\langle e,S_i-S_0\rangle| =h|\langle e,U_i\rangle|.\] The one-coordinate case of (70) and (71) yield \[\mathbb P\{|\langle e,U_i\rangle|\le t \mid A,Y,S_0,\ldots,S_{i-1}\} \le\min(1,CK'\sqrt d\,t).\] Integrating this tail with \(t=e^{-u}\) shows that the conditional expectation of \(\log(1/|\langle e,U_i\rangle|)\) is at most \(1+\log(CK'\sqrt d)\). For fixed \(A\), the expected density tilt is \[\mathbb E[\log K'\mid A] =\log K-D_{\mathrm{KL}}(p_A\Vert p^0)\le\log K.\] The divergence is finite: \(p_A/p^0\le K\) controls its positive part, and \(-x\log x\le1/e\) controls its negative part under \(p^0\). Also the individual marginal \((A,S_0)\) consists of an independent Haar frame and a unit vector. Thus \(\lVert AS_0\rVert^2\) has beta parameters \(k/2,(d-k)/2\) for every fixed target, with both parameters positive. The beta integral gives \[ \mathbb E h^{-2}=\frac{d-2}{d-k-2}\le C, \qquad \mathbb E\log(1/h)\le\tfrac12\log\mathbb E h^{-2}\le C. \tag{84}\] Combining these estimates gives \(\mathbb E\log(1/\delta_i)\le C+\log K+C\log d\). The negative part is bounded because \(\delta_i\le2\). Both signs of every height logarithm are therefore integrable. We may now take the expected logarithm of (79). It is exactly the finite relative entropy \(D_{\mathrm{KL}}(Q\Vert\Lambda)=U\). Forgetting \(A\) by data processing yields \(J_S\le U\). Adding the independent \(B\) and then projecting coordinatewise yields \(J_Z\le J_S\). Finally (71), (83), and the height bound give (82). ◻ The inverse-height terms in the upper bound can grow with the density bound \(K\). We will keep those terms until the projected correlation has supplied matching lower bounds. This order is what permits averaging over rare states at the end. The correlation retained by the auxiliary projectionWe will bound the probability that an independent projected signal lies close to the affine span of preceding projected replicas. Averaging over the side frame turns this tube probability into ball averages of \(p_A\) around the actual label \(AS_0\). Their radii vary with the simplex height; the next lemma controls the logarithmic cost of taking a supremum over those radii. For a probability density \(f\) on \(\mathbb R^k\), define its centered maximal average by \[\mathcal M f(u)=\sup_{t>0}\frac{1}{v_kt^k} \int_{\lVert u'-u\rVert\le t} f(u')\,\mathrm du'.\] For fixed \(f\in L^1\), the ball integral is continuous in its center and positive radius; the supremum may therefore be restricted to positive rational radii. In particular \(\mathcal M f\) is measurable. For the jointly measurable family \(p_A\), the same rational supremum and Tonelli’s theorem give joint measurability in \((A,u)\). For historical background on one-dimensional maximal averages, see Hardy and Littlewood (Hardy and Littlewood 1930, sec. III). Stein and Strömberg (Stein and Strömberg 1983, 259) discuss the exponential dimension cost of customary covering proofs. We prove the centered weak estimate needed here, including its constant. Lemma 27 (Logarithmic maximal estimate). Suppose \(f\) is a probability density supported in the unit ball of \(\mathbb R^k\) and \(f\le K\alpha_{d,k}\). Then \[ \int f(u)\log\frac{\mathcal M f(u)}{f(u)}\,\mathrm du \le Cd+\log(1+Cd+\log K). \tag{85}\] Both logarithms are integrable under \(f(u)\mathrm du\). The integrand is taken as zero on \(\{f=0\}\). Proof. For every \(\lambda>0\), \[ \bigl|\{u:\mathcal M f(u)>\lambda\}\bigr| \le\frac{3^k}{\lambda}. \tag{86}\] To prove this, take a compact subset of the displayed superlevel set. Open and closed balls have the same \(f\)-mass, since their boundaries have Lebesgue measure zero. At each point choose an open ball centered there whose \(f\)-mass exceeds \(\lambda\) times its volume. The strict inequality persists under a sufficiently small enlargement, so these balls give an open cover. Take a finite subcover and greedily keep disjoint balls in decreasing order of radius. Every discarded ball meets a retained ball of at least its radius, and is contained in that ball’s threefold dilation. Hence the compact set has volume at most \(3^k\) times the sum of the retained volumes. The retained balls are disjoint and have total \(f\)-mass at most one, so their volumes sum to less than \(1/\lambda\). Inner regularity proves (86). Set \(b=3^k\) and \(L_f=K\alpha_{d,k}\). Since \(\mathcal M f\le L_f\), layer integration over the unit ball gives \[\begin{align*} \int_{\lVert u\rVert\le1}\mathcal M f(u)\,\mathrm du &\le\int_0^{L_f}\min(v_k,b/\lambda)\,\mathrm d\lambda\\ &\le b\bigl(1+\log^+(L_fv_k/b)\bigr). \end{align*}\] On the support of \(f\), a radius-two ball contains the whole unit ball, so \(\mathcal M f\ge(v_k2^k)^{-1}\) there. Its logarithm is bounded on that support, while \(f\log f\) is integrable by the bounded-density argument in Lemma 26. Jensen’s inequality under the probability density \(f\) now yields \[\int f\log(\mathcal M f/f) \le\log\int_{\{f>0\}}\mathcal M f \le k\log3+\log\bigl(1+\log^+(K\alpha_{d,k}v_k/3^k)\bigr).\] Use (71) and \(d/k\le32\) to obtain (85). ◻ The next test uses only the projected points, including when their conditional distance distribution has atoms. Lemma 28 (Lower projected correlation). For the replica law above and every \(0<a<1\), \[ J_Z\ge a\sum_{i=1}^r\left[ k\mathbb E_Q\log(1/\delta_i)-\log(C_1^kv_k) -\mathbb E_{A,S\sim P}\log\mathcal M p_A(AS)\right] +r\log(1-a), \tag{87}\] where \(C_1\) is an absolute constant, independent of \(P,K,d,a\). Proof. Given \((B,Z_0,\ldots,Z_{i-1})\), define for \(z\in\mathbb R^\ell\) \[R_i(z)=\operatorname{dist}\bigl(z-Z_0, \operatorname{span}\{Z_j-Z_0:1\le j<i\}\bigr).\] Let \(F_i(t)=P\{s:R_i(Bs)\le t\}\) be its distribution function under an independent \(s\sim P\). The distance, and hence this conditional distribution function, is measurable in the projected history: the distance is the infimum of the corresponding norms over rational coefficient vectors. Let \(\rho_i=R_i(Z_i)\) be the actual replica distance. If a random variable \(R\) has distribution function \(F\), then \[ \mathbb P\{F(R)\le u\}\le u\qquad(0\le u\le1). \tag{88}\] This includes distributions with atoms. In fact, the set of real \(x\) with \(F(x)\le u\) is an initial interval, whose probability is the supremum of \(F\) on that interval and is at most \(u\). Tail integration of (88) gives, for \(0<a<1\), \[ \mathbb E F(R)^{-a} \le1+\int_1^\infty t^{-1/a}\,\mathrm dt=\frac1{1-a}. \tag{89}\] In particular \(F(R)>0\) almost surely under its own law. Apply the relative-entropy variational inequality to the test \(z\mapsto F_i(R_i(z))^{-a}\) under the conditional one-vector reference law \(P_B\). For an unbounded test, use its minimum with a finite constant and pass monotonically to the limit; (89) bounds its reference expectation. The chain rule for \(J_Z\) has one term for each conditional law of \(Z_i\) given \((B,Z_0,\ldots,Z_{i-1})\), and its \(i=0\) term is zero. It follows that \[ J_Z\ge a\sum_{i=1}^r\mathbb E_{Q,B}\log\frac1{F_i(\rho_i)} +r\log(1-a). \tag{90}\] The logarithmic expectations are finite: the variational bound after averaging controls their nonnegative parts by the finite correlation \(J_Z\le J_S\) from Lemma 26. We next estimate \(F_i(\delta_i)\) while holding the entire replica outcome \((A,S_0,\ldots,S_r)\) fixed. Put \(D=D_{i-1}\) and \(E=\operatorname{row}(B)\). The isometry \(B^{\mathsf T}:\mathbb R^\ell \to E\) gives the exact identity \[ \operatorname{dist}(Bu,BD) =\lVert\operatorname{proj}_{E\cap D^\perp}u\rVert. \tag{91}\] Indeed, the orthogonal complement within \(E\) of \(\operatorname{proj}_E D\) is precisely \(E\cap D^\perp\). The projection of the fixed \((i-1)\)-space \(D\) into a Haar \(\ell\)-space \(E\) has full rank almost surely: generate \(E\) with independent standard Gaussians and restrict them to an orthonormal basis of \(D\), obtaining an \(\ell\times(i-1)\) Gaussian matrix of full column rank. Thus \(\dim(E\cap D^\perp)=q_i=\ell-i+1\). Rotations preserving \(D\) show that this intersection is a uniform \(q_i\)-space in the \(n_i\)-space \(D^\perp\). For fixed \(u\), let \(v=\operatorname{proj}_{D^\perp}u\). Spherical projection in that \(n_i\)-space, followed by (71), gives \[ \mathbb P_B\{\operatorname{dist}(Bu,BD)\le t\} \le\min\left(1,\left(\frac{C_0t}{\lVert v\rVert}\right)^{q_i}\right), \tag{92}\] with value one if \(v=0\). The constant \(C_0\) is absolute because \(n_i/q_i\le16/3\); the density formula is applicable because \(n_i-q_i=d-\ell>2\). In the replica experiment \(A\) annihilates \(D\). Therefore for \(u=s-S_0\), \(\lVert v\rVert\ge\lVert A(s-S_0)\rVert\). Also \(\rho_i\le\delta_i\), since \(B\) is a contraction and the distance minimization permits all vectors in \(D\). Set \(u_0=AS_0\) and \(q=q_i\). The identity \[\min(1,(b/x)^q)=\int_1^\infty qt^{-q-1} \mathbf 1_{\{x\le bt\}}\,\mathrm dt \qquad(b>0,\ x\ge0)\] also holds at \(x=0\). Integrating (92) against \(s\sim P\) and using Tonelli’s theorem yields \[\begin{align*} \mathbb E_B F_i(\delta_i) &\le\int_1^\infty qt^{-q-1} P\{\lVert As-u_0\rVert\le C_0\delta_i t\}\,\mathrm dt\\ &\le v_k(C_0\delta_i)^k\mathcal M p_A(u_0) \int_1^\infty qt^{k-q-1}\,\mathrm dt\\ &=\frac{q_i}{q_i-k}v_k(C_0\delta_i)^k\mathcal M p_A(u_0) \le C_1^kv_k\delta_i^k\mathcal M p_A(u_0). \tag{93}\end{align*}\] The integral converges because \(q_i>k\), and the last step uses the exact bound \(q_i/(q_i-k)\le3/2\). One absolute choice of \(C_1\) works for every \(k\ge1\). Now \(F_i(\rho_i)\le F_i(\delta_i)\). Jensen’s inequality for \(-\log\) and (93), first for each fixed replica outcome, give \[\mathbb E_{Q,B}\log\frac1{F_i(\rho_i)} \ge k\mathbb E_Q\log(1/\delta_i)-\log(C_1^kv_k) -\mathbb E_{A,S\sim P}\log\mathcal M p_A(AS).\] Here the last expectation uses the individual marginal \((A,S_0)\). The right side is finite by Lemmas 26 and 27; Jensen may also be read in the extended sense before averaging. Substitution into (90) proves (87). ◻ Cancellation and the block information boundThe preceding estimates can now be subtracted. Retaining every term of \(U\) gives, for every \(0<a<1\), \[\begin{align*} J_S-J_Z\le{}&(1-a)U-r\log(1-a)\\ &+a\sum_{i=1}^r\left[ \log(\alpha_{d-i+1,k}C_1^kv_k) +\mathbb E_{A,S\sim P} \log\frac{\mathcal M p_A(AS)}{p_A(AS)}\right]. \tag{94}\end{align*}\] Thus only the fraction \(a\) of the height contribution has been canceled; the residual \((1-a)U\) remains explicit. Take \[ a=1-d^{-3}. \tag{95}\] Using (82), (71), and Lemma 27 in (94) yields \[ \frac{J_S-J_Z}{m} \le Cd+\log(1+Cd+\log K) +C\frac{\log K+\log d}{d^2}. \tag{96}\] In detail, \((1-a)U/m\le C(\log K+\log d)/d^2\) because \(r/m\le1\), and \(-r\log(1-a)/m\le3\log d\). Also \(\log(\alpha_{n_i,k}C_1^kv_k)\le Cd\) by (71) and \(n_i/k\le32\). These are all the terms in (94); the \(3\log d\) term is absorbed into \(Cd\). Proposition 29 (Haar block information bound). Let \(S\sim\sigma\), let \(V\) be a finite-valued state, and let a block consist of \(k=\lfloor d/16\rfloor\) fresh independent standard Gaussian designs, independent of \((S,V)\) and the past, with their exact labels. Suppose a finite-valued state \(W\) is generated by a measurable randomized kernel from \(V\) and the full block, with no other retained information from the past. Let \(B\) be a Haar \(\ell=\lfloor d/4\rfloor\) frame drawn independently of the entire joint experiment producing \(S,V\), the block, and \(W\). Then \[\begin{align*} I(S;W\mid V,B,BS)\le{}&\frac{H(W\mid V)}{m}+Cd +\log(1+Cd+H(V))\\ &+C\frac{H(V)+\log d}{d^2}, \qquad m=\lfloor d/16\rfloor+1. \tag{97}\end{align*}\] The constant \(C\) is absolute. If every state has logarithmic width at most \(H_*=o(d^2)\), the left side is at most \(Cd\) for all sufficiently large \(d\). Proof. Fix a state value \(v\) with probability \(p_v>0\). Its conditional signal law has the bounded density \[g_v(s)=\frac{\mathbb P\{V=v\mid S=s\}}{p_v}\le K_v:=\frac1{p_v}\] relative to \(\sigma\). The new Gaussian design matrix is independent of \((S,V)\). Its row Gram–Schmidt factorization is \(RA\) with an invertible triangular \(R\) and a Haar \(k\)-frame \(A\), almost surely. The exact labels are \(R(AS)\), so, conditional on the design, observing them is equivalent to observing \(Y=AS\). Because the whole design is independent of \(S\) conditional on \(V=v\), its coefficients \(R\) add no information about \(S\) once \((A,Y)\) is specified. Generate the replicas conditionally independently from this posterior, and independently of \(W\) given \(v\) and the complete block data. A single replica then has the same joint law with the block data and \(W\) as the original signal: both use the posterior given the block, and \(W\) has no further dependence on the signal once those data and \(v\) are fixed. In particular each pair \((S_i,W)\) has the same conditional law as \((S,W)\), the auxiliary \(B\) is independent of the entire tuple and \(W\), and the original pair used on the left is also coupled independently with \(B\). The common map \((B,s)\mapsto Bs\) is continuous, and the unprojected total correlation is finite by Lemma 26. These are precisely the application hypotheses of Lemma 13. Applied in the conditional experiment \(V=v\), it gives \[mI(S;W\mid V=v,B,BS) \le H(W\mid V=v)+J_S(v)-J_Z(v).\] No conditional independence of the replicas given \(W\) is asserted or needed. Insert (96) with \(K=K_v\) and average over \(v\). The identity \[\sum_v p_v\log K_v=\sum_vp_v\log(1/p_v)=H(V)\] accounts exactly for rare state values. Jensen’s inequality bounds the average of \(\log(1+Cd+\log K_v)\) by \(\log(1+Cd+H(V))\). This proves (97). Under the width hypothesis, \(H(W\mid V)/m=o(d)\), the last term is \(o(1)\), and the logarithmic term is \(O(\log d)\) for all sufficiently large \(d\). The stated \(Cd\) bound follows. ◻ The terminal fiber and the streaming conclusionThe block comparison is complete. We apply the residual-sphere endpoint of Section 2 with the present Haar side frame, then compare the resulting information with the number of blocks. Lemma 30 (The terminal Haar fiber). Let \(S\sim\sigma\), let \(V_f\) be finite-valued, and let the unit-valued output \(\widehat S\) be generated from \(V_f\) by a measurable kernel using fresh randomness, conditionally independently of \(S\). Let \(B\) be a Haar \(\ell=\lfloor d/4\rfloor\) frame independent of \((S,V_f,\widehat S)\). If \(\mathbb P\{\arccos\langle\widehat S,S\rangle\le\epsilon\}\ge2/3\) for \(0<\epsilon\le1/10\), then \[ I(S;V_f\mid B,BS) \ge\frac{d-\ell-1}{3}\log\frac1{4\epsilon}-\log2. \tag{98}\] Proof. Apply Lemma 6 with \(G=B\), \(V=V_f\), \(b=\ell\), and \(r_0=1/2\). Its hypotheses hold because \(B\) is independent of \((S,V_f,\widehat S)\). Explicitly, the Markov bound for the residual radius \(h\) is \(\mathbb P\{h<1/2\}\le4\ell/(3d)\le1/3\). Success together with \(h\ge1/2\) therefore has probability at least \(1/3\), while the conditional prior cap probability is at most \((4\epsilon)^{d-\ell-1}\). These are exactly the constants in (98). ◻ Proposition 31 (Streaming consequence). For each dimension \(d\), let \(S\sim\sigma\) and consider a learner in Definition 3, with deterministic horizon \(T=T(d)\) and nonnegative integer memory \(M=M(d)\). If \(M(d)=o(d^2)\), \(0<\epsilon(d)\le1/10\), and its angular success probability averaged over the signal, samples, and learner randomness is at least \(2/3\), then \[T(d)=\Omega\bigl(d\log(1/\epsilon(d))\bigr).\] The implicit constant is absolute; the dimension threshold may depend on the memory sequence. Proof. Apply the fixed-shared-seed part of Lemma 9 to the learner in Definition 3. The resulting Borel experiment has data-independent initialization and uniform-prior success at least \(2/3\), with fresh transition and output randomness retained. This is the success threshold in Lemma 30. First, success forces \(T=\Omega(d)\) without any memory restriction. Suppose \(T\le d/8\), and grant an estimator all \(T\) raw pairs, filling in ignored samples after any early stopping. Lemma 5 with \(b=T\) gives a uniform residual sphere in kernel dimension \(d-T\); (4) bounds its probability of radius below \(1/2\) by \(4T/(3d)\). At \(T=0\) the radius is one deterministically. On radius at least \(1/2\), the cap argument of Lemma 12 bounds conditional success by \((4\epsilon)^{d-T-1}\). Thus \[ \mathbb P\{\text{angular error}\le\epsilon\} \le\frac{4T}{3d}+(4\epsilon)^{d-T-1} \qquad(T\le d/8). \tag{99}\] The output has no further dependence on \(S\) conditional on all the raw data, so the bound also covers randomized estimators. Its right side is at most \(1/6+(0.4)^{7d/8-1}<2/3\) for all sufficiently large \(d\). Hence \(T>d/8\) on those dimensions. Put \(L=\log(1/\epsilon)\ge\log10\). If \(T\ge dL\), the conclusion already holds. On the remaining branch \(T<dL\), apply the output-capacity bound of Lemma 11 to this uniform-prior experiment with \(p_0=2/3\). It gives \[ (d-1)L\le M\log2+\log(T+1)+\log(3/2). \tag{100}\] For large \(d\), \(\log(T+1)\le\log(dL+1)\le dL/2\): here \(dL\ge d\log10\) and \(\log(x+1)\le x/2\) for \(x\ge3\). Thus (100) implies \[ dL\le C(M+1),\qquad L=o(d),\qquad T=o(d^2). \tag{101}\] For example, after subtracting \(dL/2\), the left side is \((d/2-1)L\ge dL/4\) for \(d\ge4\). The second and third conclusions follow from \(M=o(d^2)\) and \(T<dL\). Use Lemma 9 to retain the stopping index, postpone the fresh randomized output, and append ignored rows through the last block. Its state count gives every padded state entropy at most \[ H_*=M\log2+\log(T+2) \le M\log2+\log(C(M+1)+2)=o(d^2). \tag{102}\] This count restricts the branch where the block argument is needed; it does not replace that argument in the representable-precision regime. Use a single independent Haar \(\ell\)-frame \(B\) for the entire analysis, put \(Z=BS\), and divide the padded stream into \(N=\lceil T/k\rceil\) blocks. For consecutive states \(V,W\), the chain rule and data processing give \[I(S;W\mid B,Z) \le I(S;V,W\mid B,Z) =I(S;V\mid B,Z)+I(S;W\mid V,B,Z).\] The initial state is independent of \((S,B,Z)\). Proposition 29 and (102) therefore imply \[ I(S;V_f\mid B,Z)\le Cd\lceil T/k\rceil. \tag{103}\] The terminal lower bound (98) is at least \(c dL\) for all sufficiently large \(d\): indeed \(d-\ell-1\ge d/2\) and \[\log\frac1{4\epsilon} \ge\left(1-\frac{\log4}{\log10}\right)L,\] while the subtractive \(\log2\) is absorbed because \(dL\to\infty\). Finally \(k\ge d/32\) and the already proved \(T>d/8\) give \[d\lceil T/k\rceil\le32T+d\le40T.\] Comparison with (103) proves the proposition. ◻ A learner that succeeds with probability at least \(2/3\) for every fixed signal also has uniform-prior success at least \(2/3\) by integration. Thus Proposition 31 separately gives the corresponding worst-signal lower bound, with the same exact observations and finite-state model. Synthetic Gaussian projectionsThe Gaussian test in Section 4 uses an orthonormal selector of the projected difference space’s complement. Here we instead draw additional Gaussian rows in that complement. Their composition with the side matrix has a random covariance, which must be averaged in the likelihood comparison. To normalize that comparison, we apply the Gaussian-base equal-label identity to the projected prior, rather than applying only its spherical instance to the original prior. The inverse-determinant scale of this random covariance cancels the scale of the projected simplex volume. A critical inverse-distance kernel then gives a double-logarithmic density cost. Although the cost has the same order as in Section 4, the comparison law and cancellation are different. We use this bound while conditioning on an entire state history. We use the spherical and Gaussian notation of Section 3. Fix \(d\ge16\) and put \[ k=\ell=\left\lfloor\frac d8\right\rfloor,\qquad u=\ell-1,\qquad m=\left\lfloor\frac d2\right\rfloor,\qquad q=m-u. \tag{104}\] In particular, \[ k+u\le m-2,\qquad q-k-1=m-2k>0,\qquad d-1-u-k=d-2k\ge\frac{3d}{4},\qquad \frac{m}{q-k-1}\le2. \tag{105}\] The letter \(m\) in this section is the dimension of the side projection. The difference count in Theorem 15 will be \(u\). Proposition 32 (Information beyond an independent Gaussian projection). Let \(P=f\sigma_d\) be a probability, where \(f\) is a nonnegative Borel function satisfying \(f\le\Lambda<\infty\), with \(\Lambda\ge1\). Draw \(X\sim P\) and \(A\sim\gamma_{k,d}\) independently. Let \(z=z(A,AX)\) be a measurable function taking at most \(K\ge1\) values. Independently of this entire experiment draw \(B\sim\gamma_{m,d}\), and put \(Y=BX\). With the parameters in (104), \[ I(X;z\mid B,Y) \le \frac{\log K}{\ell} +C\bigl[d+\log(1+\log\Lambda)\bigr], \tag{106}\] where \(C\) is absolute and independent of \(P,\Lambda,K\), and the message rule. The proof couples \(\ell\) signals by giving them the same actual measurements. Corollary 16 describes their joint law. We compare its total correlation before and after applying \(B\). For the comparison after \(B\), a Gaussian matrix is constrained to annihilate the projected differences. We first construct the resulting probability law, then estimate its density while keeping its dependence on \(B\) explicit. The actual and projected replica lawsFor a tuple \(\mathbf x=(x_1,\ldots,x_\ell)\) in \(\mathbb R^h\), where \(h\) will be \(d\) or \(m\), write \[D_{\mathbf x}=[x_2-x_1,\ldots,x_\ell-x_1],\qquad L_{\mathbf x}=\mathop{\mathrm{col}}(D_{\mathbf x}),\qquad \Delta(\mathbf x)=\det(D_{\mathbf x}^{\mathsf T}D_{\mathbf x})^{1/2}.\] Thus \(L_{\mathbf x}\) and \(\Delta(\mathbf x)\) are the objects \(F_{\mathbf x}\) and \(J(\mathbf x)\) of (20), with difference count \(u\). At full affine rank define \[ w(\mathbf x)=(2\pi)^{-ku/2}\Delta(\mathbf x)^{-k}. \tag{107}\] We need the same version of a projection density before and after a composition of matrices. If \(\nu\) is a finite measure and \(C\) has \(r\) rows, write \[ g_C^\nu(v)=\liminf_{j\to\infty} \int\varphi_{1/j,r}(Cs-v)\,\nu(ds),\qquad \varphi_{\eta,r}(v)=(2\pi\eta^2)^{-r/2} \exp\!\left(-\frac{\lVert v\rVert^2}{2\eta^2}\right). \tag{108}\] This formula defines a jointly measurable function even for matrices of deficient rank. In every use below the matrix has full row rank almost surely, and Theorem 15 identifies this version with the projection density and transfers its finite positive values to the constrained evaluations. We abbreviate \(g_C^P\) to \(g_C\). Fix the data of Proposition 32. In Corollary 16, take ambient dimension \(d\), \(k\) rows, difference count \(u\), and prior \(P\). The condition \(k+u\le d-2\) follows from (105). Draw \(A\sim\gamma_{k,d}\), then a common label \(b\) with density \(g_A\), and finally draw \(X_1,\ldots,X_\ell\) independently from \(P\) conditional on \(AX=b\), and set \(z=z(A,b)\). Denote this law by \(Q\). Its tuple–matrix likelihood is \[ Q(d\mathbf x,dA) =P^{\otimes\ell}(d\mathbf x)\gamma_{k,d}^{L_{\mathbf x}}(dA) w(\mathbf x)g_A(Ax_1)^{-u}. \tag{109}\] Every pair \((X_i,z)\) has the original signal–message law. Independently draw \(B\) and put \(Y_i=BX_i\). For fixed \(B\), let \(P^B=B_\#P\), and define \[\mathcal C_X=D_{\mathrm{KL}}(Q_{\mathbf X}\Vert P^{\otimes\ell}),\qquad \mathcal C_Y=\mathbb E_BD_{\mathrm{KL}}(Q_{\mathbf Y\mid B}\Vert(P^B)^{\otimes\ell}).\] Lemma 18, with the same substitution and density bound \(L=\Lambda\), proves absolute integrability of \(\log g_A(b)\) and \(\log w(\mathbf X)\). It gives \[ \mathcal C_X \le \mathbb E\log w(\mathbf X)-u\,\mathbb E\log g_A(b)<\infty. \tag{110}\] Data processing under the coordinatewise map \(x_i\mapsto Bx_i\) gives \(0\le\mathcal C_Y\le\mathcal C_X\). We next construct a probability density against which to test \(\mathcal C_Y\). Almost every \(B\) has full row rank. For such a fixed \(B\), the spherical projection formula (21), with \(m\) rows, and \(P\le\Lambda\sigma_d\) show that \(P^B\) has bounded Lebesgue density on a compact ellipsoid. Here \(d-m-2\ge0\). The standard Gaussian density has a positive minimum on that ellipsoid, so \(P^B\) has bounded density relative to standard Gaussian probability on \(\mathbb R^m\). This bound is finite for each such \(B\) and is allowed to depend on \(B\). Apply Corollary 16 in ambient dimension \(m\), again with \(k\) rows and difference count \(u\). Its dimension condition is \(k+u\le m-2\), as in (105). The resulting same-label tuple law has density \[ R_B(\mathbf y) =w(\mathbf y)\, \mathbb E_{F\sim\gamma_{k,m}^{L_{\mathbf y}}}g_{FB}(Fy_1)^{-u} \tag{111}\] relative to \((P^B)^{\otimes\ell}\). Indeed, the density version agrees pointwise under composition: \[ g_F^{P^B}(v) =\liminf_j\int\varphi_{1/j,k}(Fy-v)\,P^B(dy) =g_{FB}^{P}(v). \tag{112}\] The finite-measure theorem transfers version, zero, and infinity exceptions before the negative power is taken. Thus \(R_B\) is finite and positive almost everywhere for the product law and integrates to one. It can be chosen jointly measurable in \(B,\mathbf y\): on the full-affine-rank set, realize the projected Gaussian rows with the Borel orthogonal projection onto \(L_{\mathbf y}^{\perp}\), and assign fixed values on the null complement. The actual law \(Q_{\mathbf Y\mid B}\) is absolutely continuous relative to \((P^B)^{\otimes\ell}\), because \(Q_{\mathbf X}\ll P^{\otimes\ell}\) and \(B\) is independent of the tuple. At the actual \((B,\mathbf Y)\), draw \(F\) with the conditional law in (111) and put \[\Xi=g_{FB}(FY_1).\] The same null-set transfer makes \(\Xi\) finite and positive almost surely in this enlarged experiment. The synthetic density and its critical kernelThe comparison density \(R_B\) contains the inverse \(u\)th moment of \(\Xi\). To compare its logarithm with the actual likelihood in (110), we will control \(\Xi\) using the actual matrix \(A\) and label \(b\). The next lemma first averages the synthetic Gaussian covariance at a fixed tuple. It then bounds the resulting inverse-distance integral by a function \(U_A(b)\) whose total mass can be estimated independently of that tuple. Lemma 33 (Synthetic likelihood and critical kernel). Let \(P=f\sigma_d\) be as in Proposition 32. Fix a tuple \(\mathbf x\in(S^{d-1})^\ell\) of affine rank \(u\), and put \[x=x_1,\qquad L=L_{\mathbf x},\qquad \delta(s)=\mathop{\mathrm{dist}}(s-x,L).\] Draw \(B\sim\gamma_{m,d}\). Given \(B\), draw \(F\sim\gamma_{k,m}^{BL}\), so each row of \(F\) is a standard Gaussian in \((BL)^\perp\). Then \[ \mathbb E_{B,F}g_{FB}(FBx) \le (2\pi)^{-k/2}a_q\int\delta(s)^{-k}\,P(ds),\qquad a_q=\mathbb E\det(HH^{\mathsf T})^{-1/2}\le(q-k-1)^{-k/2}, \tag{113}\] where \(H\) is a standard Gaussian \(k\times q\) matrix. There is an absolute \(C_0>0\) such that the following holds with \[ r_*=\exp\!\left[-\frac{C_0(d+\log\Lambda)}{d-1-u-k}\right]\in(0,1). \tag{114}\] For every full-row-rank \(k\times d\) matrix \(A\), put \(a=\lVert A\rVert_{\mathrm{op}}\) and define \[\begin{align*} U_A(b)={}&(2\pi)^{-k/2}\mathbf 1_{\{\lVert b\rVert\le a\}} \left[1+a^k\int g_A(v)\mathbf 1_{\{\lVert v-b\rVert\le2a\}}\right. \\[-2mm] &\hspace{42mm}\left. {}\cdot\min\{(ar_*)^{-k},\lVert v-b\rVert^{-k}\}\,dv\right]. \tag{115}\end{align*}\] At \(v=b\), the second entry in the minimum is interpreted as \(+\infty\). If \(Ax=b\) and \(L\subset\ker A\), then \[ (2\pi)^{-k/2}\int\delta(s)^{-k}\,P(ds)\le U_A(b). \tag{116}\] Writing \(V_k\) for the volume of the Euclidean unit ball in \(\mathbb R^k\), the same function satisfies \[\begin{align*} \int U_A(b)\,db &\le (2\pi)^{-k/2}V_k a^k[2+k\log(2/r_*)], \tag{117}\\ \mathbb E_{A\sim\gamma_{k,d}}\log\int U_A(b)\,db &\le C\bigl[d+\log(1+\log\Lambda)\bigr]. \tag{118}\end{align*}\] The logarithm in the last display is absolutely integrable, and \(\mathbb E_A\log(1+\int U_A)<\infty\). Proof. At the fixed tuple, the restriction of \(B\) to \(L\) has rank \(u\) almost surely. Choose an orthonormal \(m\times q\) matrix \(N\) whose columns span \((BL)^\perp\), using only \(BD_{\mathbf x}\). Gram–Schmidt with a fixed ordering gives a measurable choice on the full-rank set. With an independent standard Gaussian \(k\times q\) matrix \(H\), the required conditional law of \(F\) is realized by \[ F=HN^{\mathsf T}. \tag{119}\] The choice of \(N\) uses only \(B|_L\). The Gaussian restriction \(B|_{L^\perp}\) is independent of \(B|_L\). Write \(s-x=v+w\) with \(v\in L\) and \(w\in L^\perp\). Then \(N^{\mathsf T}Bv=0\), while, conditionally on \(B|_L\), the vector \(N^{\mathsf T}Bw\) has law \(N(0,\lVert w\rVert^2I_q)\). Therefore, conditionally on \(H\) and \(B|_L\), \[ FB(s-x)\sim N(0,\delta(s)^2HH^{\mathsf T}). \tag{120}\] This conditional calculation accounts for the dependence of \(F\) on \(B\). For \(\eta>0\), Tonelli’s theorem and (120) give \[\mathbb E_{B,F}\int\varphi_{\eta,k}(FB(s-x))\,P(ds) =(2\pi)^{-k/2}\int \mathbb E_H\det(\delta(s)^2HH^{\mathsf T}+\eta^2I_k)^{-1/2}\,P(ds).\] The integrand on the right increases as \(\eta\downarrow0\). The tube bound (23) gives \(\delta(s)>0\) for \(P\)-almost every \(s\); it also gives a finite inverse \(k\)th moment, since \(d-1-u>k\). The matrix \(H\) has full row rank almost surely. Monotone convergence on the right and Fatou’s lemma on the left, along \(\eta=1/j\), prove the first inequality in (113). For completeness, successive orthogonal residuals of the rows of \(H\) have independent squared lengths \[V_i\sim\chi^2_{q-i+1},\qquad \det(HH^{\mathsf T})=\prod_{i=1}^kV_i.\] Indeed, conditionally on the preceding rows, the new row is standard Gaussian on a complement of dimension \(q-i+1\); this conditional law is constant, which also proves independence of the residual lengths. This is the elementary Gaussian orthogonal reduction described in Edelman and Rao (Edelman and Rao 2005, sec. 5); the rectangular product used here follows from the displayed conditioning. Direct integration of the chi-squared density gives \(\mathbb EV^{-1}=1/(p-2)\) for \(V\sim\chi_p^2\), \(p>2\). Cauchy–Schwarz hence gives \[a_q=\prod_{i=1}^k\mathbb EV_i^{-1/2} \le\prod_{i=1}^k(q-i-1)^{-1/2} \le(q-k-1)^{-k/2}.\] We next use the actual observations. If \(Ax=b\) and \(L\subset\ker A\), then, for \(v=P_L(s-x)\), \[ \lVert As-b\rVert=\lVert A(s-x-v)\rVert\le a\,\delta(s). \tag{121}\] Let \(\alpha=d-1-u\). The tube bound, now applied to \(x+L\), gives \[P\{\delta\le t\}\le\Lambda C^d t^\alpha,\qquad 0<t\le1.\] Layer-cake integration therefore yields, for \(0<r\le1\), \[\int_{\{\delta\le r\}}\delta^{-k}\,dP \le \Lambda C^d \left(1+\frac{k}{\alpha-k}\right)r^{\alpha-k}.\] Since \(\alpha-k=d-2k\ge3d/4\), an absolute choice of \(C_0\) makes this quantity at most one at \(r=r_*\). On \(\{\delta>r_*\}\), (121) gives both \(\delta^{-k}\le r_*^{-k}\) and \(\delta^{-k}\le a^k\lVert As-b\rVert^{-k}\). The pushforward of \(P\) under \(A\) has density \(g_A\), and both \(As\) and \(b\) have norm at most \(a\). Integration of these two bounds gives (116) with the two support cutoffs in (115). The remaining singularity has the critical exponent \(k\) in dimension \(k\). Radial integration gives the exact mass \[ \int_{\{\lVert w\rVert\le2a\}} \min\{(ar_*)^{-k},\lVert w\rVert^{-k}\}\,dw =V_k[1+k\log(2/r_*)]. \tag{122}\] The ball of radius \(ar_*\) contributes \(V_k\), and the annulus contributes \(kV_k\int_{ar_*}^{2a}dt/t\). Dropping the outer cutoff in (115) only for its convolution term, and using \(\int g_A=1\), proves (117). We record the scale in this bound before taking its logarithm. Integrating the standard Gaussian density over the ball of radius \(\sqrt k\) gives \((2\pi)^{-k/2}V_k\le(e/k)^{k/2}\). Also \(\mathbb E\lVert A\rVert_{\mathrm{op}}\le C\sqrt d\): choose \(1/4\)-nets on the two unit spheres of cardinalities at most \(9^k\) and \(9^d\), bound the operator norm by twice the maximum of the scalar bilinear forms on the nets, and integrate the union bound for their standard Gaussian tails. Jensen’s inequality gives \[\mathbb E\log\bigl[(2\pi)^{-k/2}V_k\lVert A\rVert_{\mathrm{op}}^k\bigr] \le\frac k2(1-\log k)+k\log(C\sqrt d)\le Cd.\] The cutoff (114) and (105) give \[\log[2+k\log(2/r_*)] \le C+\log d+\log(1+\log\Lambda).\] These estimates prove (118). Finally, \(\log\lVert A\rVert_{\mathrm{op}}\) is integrable below because the norm dominates the absolute value of one standard Gaussian entry. Retaining the first term of \(U_A\) bounds its total mass below by \((2\pi)^{-k/2}V_k\lVert A\rVert_{\mathrm{op}}^k\). Together with (117), these observations prove the two remaining integrability assertions. ◻ Integrability and cancellation of the correlation costsThe majorant now controls the synthetic density at the actual label. Before subtracting the two correlation bounds, we verify integrability of the comparison logarithms. This uses both the finite actual correlation and the normalization of \(R_B\). Lemma 34 (Integrability of the projected comparison). In the replica experiment \(Q\), enlarged by the independent \(B\) and the conditional \(F\) above, \[\mathbb E|\log w(\mathbf Y)|+\mathbb E|\log\Xi| +\mathbb E|\log R_B(\mathbf Y)|<\infty.\] If \(G\) is a standard Gaussian \(m\times u\) matrix, then \[\begin{align*} \mathbb E\log w(\mathbf Y) &=\mathbb E\log w(\mathbf X)-\frac{k}{2}\mathbb E\log\det(G^{\mathsf T}G), \tag{123}\\ \frac12\mathbb E\log\det(G^{\mathsf T}G)&\le\frac u2\log m. \tag{124}\end{align*}\] Moreover, \[ \mathcal C_Y \ge\mathbb E\log R_B(\mathbf Y) \ge\mathbb E\log w(\mathbf Y)-u\,\mathbb E\log\Xi. \tag{125}\] Proof. At a fixed full-affine-rank \(\mathbf X\), write \(D_{\mathbf X}=UR\), where \(U\) has orthonormal columns. Since \(B\) is independent, \(BU\) is a standard Gaussian \(m\times u\) matrix. Consequently \[\Delta(\mathbf Y)\stackrel{\mathrm{law}}{=} \Delta(\mathbf X)\det(G^{\mathsf T}G)^{1/2} \quad\text{conditionally on }\mathbf X.\] The logarithm of \(\Delta(\mathbf X)\) is integrable by Lemma 18. Gaussian column residuals factor \(\det(G^{\mathsf T}G)\) into independent chi-squared variables of degrees \(m,m-1,\ldots,m-u+1\). Their absolute logarithms are integrable from their densities at zero and infinity. This proves (123) and the asserted integrability of \(\log w(\mathbf Y)\). Jensen applied to each factor gives (124). The actual equal-label experiment has \(L_{\mathbf X}\subset\ker A\). Conditionally on \((A,b,\mathbf X)\), the matrix \(B\) remains independent, and the conditional law of \(F\) is \(\gamma_{k,m}^{B L_{\mathbf X}}\). Lemma 33 therefore gives \[ \mathbb E[\Xi\mid A,b,\mathbf X]\le a_qU_A(b). \tag{126}\] It follows that \(\mathbb E\log(1+\Xi)<\infty\). Here are the details needed for the possible small values of \(g_A(b)\). Conditional Jensen reduces the claim to integrability of \(\log(1+U_A(b))\). On the actual finite positive set for \(g_A\), \[\log(1+U_A(b)) \le\log(1+g_A(b))+\log(1+U_A(b)/g_A(b)).\] The first term is integrable by Lemma 18. For the second, average over \(b\) with density \(g_A\) and apply Jensen to obtain at most \(\log(1+\int U_A)\). Lemma 33 makes its average finite. Thus \((\log\Xi)^+\) is integrable. Let the product reference first draw \(B\) and then draw \(\ell\) independent points from \(P^B\). The divergence of the actual joint law from this reference is \(\mathcal C_Y<\infty\), and \(R_B\) integrates to one under this reference. Equation (19) gives \[ \mathbb E(\log R_B(\mathbf Y))^+ \le\mathbb E\log(1+R_B(\mathbf Y))\le\mathcal C_Y+\log2<\infty. \tag{127}\] Conditionally on \(B,\mathbf Y\), Equation (111) has a finite positive right side, so \(\mathbb E_F\Xi^{-u}<\infty\). Hence \(\mathbb E_F(\log\Xi)^-<\infty\) on almost every such fiber; also \(\mathbb E_F(\log\Xi)^+<\infty\) there by the preceding joint bound. Conditional Jensen is now legitimate and gives \[\log R_B(\mathbf Y) \ge\log w(\mathbf Y)-u\,\mathbb E_F\log\Xi.\] Rearranging this inequality, \[u\,\mathbb E_F(\log\Xi)^- \le(\log R_B(\mathbf Y))^+ +|\log w(\mathbf Y)|+u\,\mathbb E_F(\log\Xi)^+.\] Its right side has finite expectation, proving joint integrability of \((\log\Xi)^-\). The same Jensen lower bound gives \[(\log R_B(\mathbf Y))^- \le|\log w(\mathbf Y)|+u\,\mathbb E_F|\log\Xi|,\] so the negative comparison logarithm is integrable as well. Finally use the normalized test \(R_B\) in (18). All logarithms are now integrable, and that test together with the conditional Jensen bound proves (125). ◻ Proof of Proposition 32. The preceding lemma permits subtraction of all likelihood and correlation terms. Conditional Jensen in (126) gives \[\mathbb E[\log\Xi\mid A,b,\mathbf X]\le\log a_q+\log U_A(b).\] The last logarithm is integrable under the actual label law: \(U_A(b)\) is bounded below there by \((2\pi)^{-k/2}\), and its positive logarithm was controlled in the proof of Lemma 34. At fixed \(A\), Jensen with \(b\sim g_A\) gives \[\int g_A(b)\log\frac{U_A(b)}{g_A(b)}\,db \le\log\int U_A(b)\,db.\] Together with (113) and (118), this yields \[ \mathbb E\log\Xi-\mathbb E\log g_A(b) \le-\frac k2\log(q-k-1) +C[d+\log(1+\log\Lambda)]. \tag{128}\] Combine (110), (123), (124), (125), and (128). The result is \[\begin{align*} \mathcal C_X-\mathcal C_Y &\le\frac{ku}{2}\log\frac{m}{q-k-1} +Cu[d+\log(1+\log\Lambda)] \\ &\le C\ell[d+\log(1+\log\Lambda)]. \tag{129}\end{align*}\] The first term is \(O(k u)\) by (105). Thus the scale from the projected Gaussian determinant cancels the inverse-determinant scale in the synthetic likelihood. The critical kernel was integrated before taking a logarithm, leaving only \(\log(1+\log\Lambda)\). Apply Lemma 13 with \(t=\ell\), the representative \(S=X_1\), message \(W=z\), and map \(\phi_B(x)=Bx\). Every replica has the required signal–message pair law, \(B\) is independent of the whole tuple and message, and \(\mathcal C_X<\infty\). Its conclusion and \(H(z)\le\log K\) give \[\ell I(X_1;z\mid B,BX_1) \le\log K+\mathcal C_X-\mathcal C_Y.\] The joint law of \((B,X_1,z)\) is the product of the \(B\)-law and the original \((X,z)\)-law. Substitution of (129) therefore proves (106). ◻ Iteration along the full state historyProposition 35 (The synthetic route to the sample bound). For every nonnegative integer-valued \(M(d)=o(d^2)\) and every sequence \(0<\epsilon(d)\le1/10\), consider learners in Definition 3 with deterministic horizon \(T(d)\) and signal \(X\sim\sigma_d\). If their angular success probability, averaged over the uniform signal, rows, and learner randomness, is at least \(2/3\), then \[T(d)\ge c\,d\log(1/\epsilon(d))\] for all sufficiently large \(d\), with an absolute \(c>0\). The threshold may depend on the memory sequence. Proof. Apply Lemma 9 with \(p=2/3\) and \(p'=s_0=16/25=0.64\) to the learner in Definition 3. This gives deterministic Borel state-based rules and data-independent initialization, with uniform-prior success at least \(s_0\); all remaining randomness is fixed. Pre-generate all \(T\) rows, including rows after an early stop. Put \(L_\epsilon=\log(1/\epsilon)\). It suffices to consider dimensions for which \(T\le dL_\epsilon\). At other dimensions the claimed lower bound already holds. Lemma 11, applied with \(p_0=s_0\), gives \[ s_0\le(T+1)2^M\epsilon^{d-1},\qquad L_\epsilon=o(d),\qquad T=o(d^2). \tag{130}\] Pad first to time \(T\) and then to the least multiple of \(k\) at least \(T\), using the stopping-index tags and ignored fresh pairs of Lemma 9. Its exact per-index state count gives \[ K=(T+2)2^M,\qquad \log K=o(d^2). \tag{131}\] Since \(s_0>3/5\), Lemma 12 gives \(T>\lfloor d/2\rfloor\) for all sufficiently large \(d\). We use the weaker consequence \[ T>d/4. \tag{132}\] Let \(p=\lceil T/k\rceil\). Write \(H_i\) for the tuple of states at the ends of the first \(i\) blocks, with \(H_0\) constant. Use one independent \(B\sim\gamma_{m,d}\) and \(Y=BX\) throughout. If a history \(h\) has probability \(p_h>0\), Lemma 10 gives its conditional signal density \[\frac{\mathbb P\{H_{i-1}=h\mid X=x\}}{p_h}\le\frac1{p_h} \quad\text{relative to }\sigma_d.\] The next block matrix is independent of \((X,H_{i-1})\). Given \(h\), the initial state is fixed, and the next state is a measurable \(K\)-valued function of the new block and its exact labels. The side matrix remains independent of this conditional experiment. Proposition 32 therefore applies with \(\Lambda=1/p_h\). Averaging it over histories gives \[\begin{align*} I(X;H_i\mid H_{i-1},B,Y) &\le\frac{\log K}{\ell} +C\left[d+\sum_h p_h\log(1+\log(1/p_h))\right] \\ &\le\frac{\log K}{\ell} +C[d+\log(1+(i-1)\log K)]. \tag{133}\end{align*}\] The last line uses concavity of \(t\mapsto\log(1+t)\) and \(H(H_{i-1})\le(i-1)\log K\). In particular it applies the bounded-prior block theorem before averaging its density cost. Since \(p=o(d)\), \(\ell\ge d/16\), and \(\log K=o(d^2)\), the right side is at most \(Cd\), with one absolute \(C\), for all sufficiently large \(d\). The chain rule then gives \[ I(X;H_p\mid B,Y)\le Cdp\le C(T+d). \tag{134}\] The postponed deterministic output is a function of the final tagged state, and therefore of \(H_p\). To obtain the matching endpoint, use Lemma 6 with \(S=X\), \(V=H_p\), \(G=B\), and \(r_0=1/4\). The output is a function of \(H_p\), and \(B\) is independent of the entire signal–history experiment. Moreover, \(2\epsilon<1/4\) and \(m\le d/2\le d-2\). For the residual radius \(\rho_B=\lVert P_{\ker B}X\rVert\), (4) gives \[\mathbb P\{\rho_B<1/4\} \le\frac{m/d}{1-1/16}\le\frac8{15}.\] Thus the lemma applies with success minus radius-failure probability at least \(s_0-8/15=8/75\), and yields \[\begin{align*} I(X;H_p\mid B,Y) &\ge\frac8{75}(d-m-1)\log\frac1{8\epsilon}-\log2 \\ &\ge c_1dL_\epsilon \tag{135}\end{align*}\] for all sufficiently large \(d\), with an absolute \(c_1>0\). Indeed, \[\log\frac1{8\epsilon} \ge\left(1-\frac{\log8}{\log10}\right)L_\epsilon,\] and \(d-m-1\ge d/2-1\). By (132), \(T+d<5T\). Combining (134) and (135) therefore gives \(T\ge c_2dL_\epsilon\) with an absolute \(c_2>0\). The constants are absolute; the eventual threshold comes only from \(M(d)=o(d^2)\). All replicas and auxiliary projections have served solely to bound information. The learner’s observations remain the prescribed fresh Gaussian rows and exact labels. ◻ Stiefel incidence and conditional block informationThis comparison bounds conditional information directly by the relative entropy of two laws on tuples of signals. Both laws make the signals share an exact projection. The first uses the actual observation block; the second obtains its shared projection from an orthonormal frame inside the coordinates of the independent side projection. Given the separate side projections of all the signals, the second law leaves them conditionally independent with their original fiber laws. This product structure makes the divergence of the two laws an upper bound for the information in the message, multiplied by the number of signals. The geometric work is to express both laws relative to one incidence measure. We first construct the laws and prove their information comparison. We then compute their exact Jacobians, transfer density exceptions before division, and estimate the resulting likelihood ratio. Write \(\sigma=\sigma_d\), and let \(v_q\) denote the volume of the unit ball in \(\mathbb R^q\). For the conditional chain rule on standard Borel spaces, see Austin (Austin 2020, secs. 3.3–3.4). We use, for all sufficiently large \(d\), \[ n=k=\lfloor d/16\rfloor,\qquad \ell=k-1,\qquad p=\lfloor d/4\rfloor . \tag{136}\] Thus \(\ell\ge2\), while \(p-n-\ell\) and \(d-n-\ell\) are at least an absolute positive multiple of \(d\). Proposition 36 (Information in one block). Let \(\mu=f\sigma\) be a probability measure with a Borel density \(0\le f\le F<\infty\), where \(F\ge1\). Draw \(S\sim\mu\) and standard Gaussian matrices \(A\in\mathbb R^{n\times d}\) and \(B\in\mathbb R^{p\times d}\), all independently. For an integer \(w\ge1\), a message \(J\) takes values in a fixed \(w\)-point alphabet (unused symbols are allowed) and is drawn by an arbitrary measurable probability kernel from \((A,AS)\); conditional on \((A,AS)\), its randomness is independent of \((S,B)\). There is an absolute constant \(C\) such that \[ I_\mu(S;J\mid B,BS) \le \frac{\log w}{k}+Cd+C\log(1+\log F). \tag{137}\] The dependence on \(F\) is doubly logarithmic. This permits us later to condition on an entire finite history: a history of probability \(a\) changes the uniform prior to a density bounded by \(1/a\). Two probability laws on compatible signalsWe use the spherical projection versions from (24), with prior \(\mu=f\sigma\). Recall them locally: for a full-row-rank \(q\times d\) matrix \(L\), with \(1\le q\le d-2\), and an interior label \(t\), let \(\sigma_{L,t}\) be uniform probability on the spherical fiber \(Ls=t\), and let \(p_L^0\) be the uniform-sphere projection density in (21). Then \[ p_L(t)=p_L^0(t)\int f\,d\sigma_{L,t},\qquad \mu_{L,t}=\frac{f\,\sigma_{L,t}}{\int f\,d\sigma_{L,t}} \quad\text{when }p_L(t)>0. \tag{138}\] Set \(p_L=0\) outside the interior projection ellipsoid and fix one measurable probability kernel on unused fibers. These jointly measurable versions specify all the conditional draws below. Their geometric finite-measure form will enter only when we compute the two likelihoods. Write \(\phi_j(A,y)\) for the probabilities of the message kernel. For a tuple \(\mathbf s=(s_1,\ldots,s_k)\), put \[U=[s_2-s_1,\ldots,s_k-s_1]\in\mathbb R^{d\times\ell}.\] The notation \(U^\perp\) means the orthogonal complement of the column space of \(U\), and \(P_{U^\perp}\) is its orthogonal projector. Let \(\mathcal V_{n,h}\) be the manifold of \(n\times h\) matrices with orthonormal rows, with the metric induced by the Euclidean matrix norm. Write \(\upsilon_{n,h}\) for normalized Riemannian volume on it. This is the Euclidean Stiefel metric discussed by Edelman, Arias, and Smith (Edelman et al. 1998, sec. 2.2.1); we compute its volume factor below. We use an auxiliary \(C\in\mathcal V_{n,p}\) and set \(D=CB\). For a full-rank \(p\times\ell\) matrix \(W\), write \(\upsilon_W\) for uniform frames whose rows lie in \((\operatorname{col}W)^\perp\). Write \(\gamma_{r,d}\) for standard Gaussian measure on \(\mathbb R^{r\times d}\), and \(\gamma_U\) for the law of \(n\) independent \(N(0,P_{U^\perp})\) rows. All these kernels are measurable on their rank domains. For example, projecting a Gaussian matrix onto \((\operatorname{col}W)^\perp\) and taking its orthonormal row polar factor gives \(\upsilon_W\); orthogonal invariance identifies this with normalized induced volume. We may assign arbitrary values on rank exceptions. Define \(Q\) by the following experiment.
Then \(AU=0\). Conditional on \((A,y)\), all \(k\) signals are independent with law \(\mu_{A,y}\), which is absolutely continuous on a positive-radius sphere of dimension \(d-n-1\). Each new point therefore avoids the affine span of its predecessors; hence \(U\) has rank \(\ell\) almost surely. The independent Gaussian \(B\) makes \(BU\) have rank \(\ell\) almost surely, so \(\upsilon_{BU}\) is available because \(p-\ell\ge n\). Each marginal \[ \mathcal L_Q(s_i,J,B,Bs_i)=\mathcal L_\mu(S,J,B,BS) \qquad(1\le i\le k) \tag{139}\] is the law in Proposition 36. The matrix \(D=CB\) has a useful conditional law under \(Q\): \[ \mathcal L_Q(D\mid \mathbf s,A,y,J) =N(0,P_{U^\perp})^{\otimes n}. \tag{140}\] To verify this, split each row of \(B\) between \(\operatorname{col}U\) and \(U^\perp\). The second component is independent of \(BU\). Given the tuple, the kernel for \(C\) depends on \(B\) only through \(BU\); its orthonormal rows send the independent Gaussian second component to \(n\) independent standard Gaussian rows in \(U^\perp\). The first component is annihilated by \(C\). Define \(R\) in the other order. First draw independent \(B\sim\gamma_{p,d}\) and \(C\sim\upsilon_{n,p}\); then \(D=CB\) is standard Gaussian. Draw \(s_1\sim\mu\) independently and, conditional on \((D,Ds_1)\), draw \(s_2,\ldots,s_k\) independently from \(\mu_{D,Ds_1}\). Draw \(J\) uniformly from its \(w\) possible values, independently, and finally draw \(A\sim\gamma_U\). Again \(U\) has rank \(\ell\). Conditional on \((C,D)\), the other \(p-n\) rows in an orthogonal row rotation of \(B\) are independent standard Gaussian rows and were not used to draw the tuple. Since \(p-n\ge\ell\), they make \(BU\) have rank \(\ell\) almost surely. In these experiments \(A\) and \(D\) also have row rank \(n\) almost surely: each is either standard Gaussian or has conditional Gaussian support \(U^\perp\), of dimension \(d-\ell\ge n\). Lemma 37 (Product conditioning). Put \(Z_i=Bs_i\). After \(A,C\) are discarded under \(R\), \[ \mathcal L_R(\mathbf s\mid B,Z_1,\ldots,Z_k,J) =\bigotimes_{i=1}^k\mu_{B,Z_i}. \tag{141}\] Consequently, \[ \operatorname{KL}(Q\|R)\ge kI_\mu(S;J\mid B,BS). \tag{142}\] Proof. Retain \(B,C\) temporarily and put \(t=Ds_1\). Conditional on \((B,C,t)\), the signals are independent with law \(\mu_{D,t}\). Disintegrating this law further according to \(z=Bs\) gives \(\mu_{B,z}\): the earlier equation \(Ds=t\) is just \(Cz=t\), already determined by \(z\). Formally, disintegrate the law of \(z=Bs\) according to \(Cz\), and compose its conditional kernel with \(\mu_{B,z}\); the resulting kernel is a conditional law given \(Ds=t\). These identities hold at the needed values because each individual signal has marginal law \(\mu\) given \((B,C)\). Conditioning the independent signals on their separate \(Z_i\)’s therefore yields the product in (141). That product does not depend on \(C,t\), so discarding them leaves it unchanged. The independent uniform \(J\) and the later kernel for \(A\) do not change the marginal under consideration. Data processing and the conditional chain rule give \[\operatorname{KL}(Q\|R)\ge \mathbb E_Q\operatorname{KL}\!\left( Q(\mathbf s\mid B,Z_1,\ldots,Z_k,J) \,\middle\|\,\bigotimes_{i=1}^k\mu_{B,Z_i}\right).\] This is valid as an extended inequality; absolute continuity will be established below. Relative entropy to a product is at least the sum of the relative entropies of the marginals. For each such summand, convexity of relative entropy permits us, after averaging, to discard the other \(Z_j\)’s from its first conditional law. The result is at least \[\mathbb E_Q\operatorname{KL}\bigl( Q(s_i\mid B,Z_i,J)\,\|\,\mu_{B,Z_i}\bigr) =I_\mu(S;J\mid B,BS)\] by (139). Sum over \(i\). ◻ The product conditional law has reduced the block information bound to one task: bound \(\operatorname{KL}(Q\|R)\) by \(\log w\) plus the geometric and prior-density costs. We next put both laws on a common incidence measure. Its exact densities will make that divergence computable, after exceptional sets have been removed before division. The common measure and its density versions
To compare the two laws, we express their conditional probabilities as normalized finite measures on the same incidence set. The underlying geometric formula is the regular-submersion case of coarea; see Federer (Federer 1959, Definition 2.10 and Theorem 3.1, pp. 423 and 426–427). The following calculation specifies the induced fiber volume and the probability normalization used here. Lemma 38 (Spherical projection measure). Suppose \(L\in\mathbb R^{q\times d}\) has full row rank and \(1\le q\le d-2\). In the interior \[E_L=\{t:t^{\mathsf T}(LL^{\mathsf T})^{-1}t<1\}\] define a finite measure on \(\{s\in\mathbb S^{d-1}:Ls=t\}\) by \[ \lambda_{L,t}(ds)= \frac{f(s)}{|\mathbb S^{d-1}|\,j_L(s)} \,d\operatorname{vol}_{d-1-q}(s),\qquad j_L(s)=\det\!\bigl(L(I-ss^{\mathsf T})L^{\mathsf T}\bigr)^{1/2}. \tag{143}\] Set \(\lambda_{L,t}=0\) outside \(E_L\). Its mass and normalization are exactly the versions recalled in (138): \[p_L(t)=\lambda_{L,t}(\mathbb S^{d-1}),\qquad \mu_{L,t}=\lambda_{L,t}/p_L(t)\quad\text{where }p_L(t)>0.\] The finite-measure kernel is jointly measurable in \((L,t)\). Moreover, \[ V_L\|p_L\|_\infty\le Fd^d,\qquad V_L=|E_L|=v_q\det(LL^{\mathsf T})^{1/2}. \tag{144}\] Proof. The tangent space at \(s\) is \(s^\perp\). The Gram matrix of the differential \(L:s^\perp\to\mathbb R^q\) is \(L(I-ss^{\mathsf T})L^{\mathsf T}\), proving the formula for its normal Jacobian. With \(P_L=L^{\mathsf T}(LL^{\mathsf T})^{-1}L\), the determinant lemma also gives \[j_L(s)=\det(LL^{\mathsf T})^{1/2}\sqrt{1-\|P_Ls\|^2}.\] The differential is onto except when \(s\) belongs to the row space of \(L\); that sphere has zero \((d-1)\)-dimensional surface measure. Applying change of variables on coordinate patches of its complement, and then exhausting the complement, gives \[\int h(s)\,d\mu(s) =\int_{\mathbb R^q}\int h(s)\,\lambda_{L,t}(ds)\,dt \quad(h\ge0).\] To identify these measures with the versions already used, put \(s_0=L^{\mathsf T}(LL^{\mathsf T})^{-1}t\) and \(r=\sqrt{1-\|s_0\|^2}\). An interior fiber is \(s_0+r\mathbb S(\ker L)\), where \(\mathbb S(\ker L)\) is the unit sphere of \(\ker L\), and \(j_L\) is constant there. Its induced volume is \(|\mathbb S^{d-q-1}|r^{d-q-1}\). Dividing this volume by \(|\mathbb S^{d-1}|j_L\) gives precisely the uniform projection density \(p_L^0(t)\) in (21). Consequently, as finite measures on every interior fiber, \[\lambda_{L,t}=p_L^0(t)f\,\sigma_{L,t}.\] Taking mass and then normalizing gives (138) pointwise on its positive set. A uniform direction in \(\ker L\) can be generated by projecting a standard Gaussian vector onto that subspace and normalizing. The projection matrix is a measurable function of full-rank \(L\); integrating the bounded Borel function \(f(s_0+r\theta)\) against this kernel therefore gives measurable \(\lambda_{L,t}\), \(p_L(t)\), and \(\mu_{L,t}\), with the fixed convention on unused fibers from (138). For orthonormal rows and \(f=1\), the displayed fiber formula gives a density proportional to \((1-\|t\|^2)^{(d-q-2)/2}\) on the unit ball. Its normalizing integral is at least \(\tfrac12v_qd^{-q/2}\): restrict to the ball of radius \(d^{-1/2}\) and use \((1-1/d)^{(d-q-2)/2}\ge1/2\). The maximum of the density times \(v_q\) is consequently at most \(2d^{q/2}\le d^d\). The bound \(f\le F\) multiplies this estimate by at most \(F\). Finally, an invertible change from orthonormal row coordinates to \(L\) divides the density by \(\det(LL^{\mathsf T})^{1/2}\) and multiplies the ellipsoid volume by the same factor. This proves (144); in particular our chosen \(p_L\) is finite everywhere. ◻ Lemma 39 (Interchange on a regular incidence set). Let \(X,Z\) be second-countable Riemannian manifolds, and let \(G:X\times Z\to\mathbb R^a\) be smooth. On a relatively open subset \(\Omega\) of \(G^{-1}(0)\), suppose both partial differentials \(d_xG\) and \(d_zG\) are onto. Write \(J_xG,J_zG\) for their normal Jacobians. For every nonnegative measurable \(h\) on \(\Omega\), \[\begin{align*} &\int_Z\int_{\{x:(x,z)\in\Omega\}} \frac{h(x,z)}{J_xG(x,z)} \,d\operatorname{vol}_{\dim X-a}(x)\,d\operatorname{vol}_Z(z) \\ &\qquad= \int_X\int_{\{z:(x,z)\in\Omega\}} \frac{h(x,z)}{J_zG(x,z)} \,d\operatorname{vol}_{\dim Z-a}(z)\,d\operatorname{vol}_X(x). \tag{145}\end{align*}\] The fiber volumes are induced volumes, and the equality holds for nonnegative integrals even when their common value is infinite. Proof. Near a point of \(\Omega\), choose coordinates \(x=(x',x'')\), \(z=(z',z'')\) so that \(x''\) and \(z''\) each have \(a\) coordinates and both \(G_{x''}\) and \(G_{z''}\) are invertible. Let \(u_X(x)\,dx\), \(u_Z(z)\,dz\) be the volume forms in these coordinates. Changing variables from \((x',x'')\) to \((x',G)\) shows that the induced \(x\)-fiber volume divided by \(J_xG\) is \[\frac{u_X(x)}{|\det G_{x''}|}\,dx'.\] Indeed, the tangent-volume factor and the normal-volume factor of this change of variables are respectively the induced fiber volume and the reciprocal normal Jacobian. Thus the left integral has local density \(u_Xu_Z/|\det G_{x''}|\) in \((x',z',z'')\); the right integral has density \(u_Xu_Z/|\det G_{z''}|\) in \((x',z',x'')\). At fixed \((x',z')\), implicit differentiation gives \[\frac{dx''}{dz''}=-G_{x''}^{-1}G_{z''},\qquad \left|\det\frac{dx''}{dz''}\right| =\frac{|\det G_{z''}|}{|\det G_{x''}|}.\] The two local integrals agree. Split \(\Omega\) into measurable pieces subordinate to a countable coordinate cover and sum using Tonelli’s theorem. This also proves the assertion for infinite integrals. ◻ Let \(\mathsf H\) be the probability law that draws \(\mathbf s\sim\mu^{\otimes k}\), then independent \(B\sim\gamma_{p,d}\), and, conditional on these, independently draws \(A\sim\gamma_U\) and \(C\sim\upsilon_{BU}\). Include an independent uniform \(J\). Independent points from \(\mu\) give \(\operatorname{rank}U=\ell\) almost surely. The same Gaussian decomposition as in (140) then gives full ranks for \(BU,A,D\) under \(\mathsf H\). Lemma 40 (Exact undivided identities and probability densities). Define the following marginal probability measures of \(\mathsf H\): \[\begin{split} \mathsf H_A(d\mathbf s,dA) &=\mu^{\otimes k}(d\mathbf s)\,\gamma_U(dA),\\ \mathsf H_C(d\mathbf s,dB,dC) &=\mu^{\otimes k}(d\mathbf s)\,\gamma_{p,d}(dB)\,\upsilon_{BU}(dC). \end{split}\] Using the unnormalized measures of (143), define nonnegative measures on the corresponding exact incidence sets by \[\begin{align*} \mathsf M_A(d\mathbf s,dA) &=\mu(ds_1)\gamma_{n,d}(dA) \prod_{i=2}^k\lambda_{A,As_1}(ds_i),\tag{146}\\ \mathsf M_C(d\mathbf s,dB,dC) &=\mu(ds_1)\gamma_{p,d}(dB)\upsilon_{n,p}(dC) \prod_{i=2}^k\lambda_{D,Ds_1}(ds_i),\qquad D=CB. \tag{147}\end{align*}\] Values at row-rank failures may be set to zero. The weights on the right below are assigned arbitrary finite values on their Gram-rank null sets. These measures satisfy the exact identities \[\begin{align*} \mathsf M_A &=(2\pi)^{-n\ell/2}\det(U^{\mathsf T}U)^{-n/2}\,\mathsf H_A, \tag{148}\\ \mathsf M_C &=\kappa\det(U^{\mathsf T}B^{\mathsf T}BU)^{-n/2}\,\mathsf H_C, \qquad \kappa=\frac{|\mathcal V_{n,p-\ell}|}{|\mathcal V_{n,p}|}. \tag{149}\end{align*}\] These are equalities of nonnegative measures. For \(L=A,D\), the values \(p_L(Ls_1)\) are positive and finite \(\mathsf H\)-almost surely and equal the shrinking-ball limits \[p_L(Ls_1)=\lim_{j\to\infty} \frac{\mu\{z:\|Lz-Ls_1\|<2^{-j}\}}{v_n2^{-jn}}.\] On this full-measure set, the densities of \(Q\) and \(R\) with respect to \(\mathsf H\) are \[\begin{align*} q_*&=w\phi_J(A,As_1)(2\pi)^{-n\ell/2} \det(U^{\mathsf T}U)^{-n/2}p_A(As_1)^{-\ell}, \tag{150}\\ r_*&=\kappa\det(U^{\mathsf T}B^{\mathsf T}BU)^{-n/2} p_D(Ds_1)^{-\ell}. \tag{151}\end{align*}\] In particular \(Q\ll R\). Proof. We first establish the identities without dividing by any projection density. In \(\mathsf M_A\), integrating the tuple variables gives the exact \((s_1,A)\)-marginal \[ \mu(ds_1)\gamma_{n,d}(dA)\,p_A(As_1)^\ell. \tag{152}\] For fixed full-rank \(A\), this weight is finite by (144). Where it is positive, \(\lambda_{A,As_1}^{\otimes\ell} =p_A(As_1)^\ell\mu_{A,As_1}^{\otimes\ell}\). The rank argument for the \(Q\) experiment therefore shows that \(\operatorname{rank}U=\ell\) outside an \(\mathsf M_A\)-null set. At zero weight the fiber product has zero mass. The same reasoning in \(\mathsf M_C\) gives the exact \((s_1,B,C)\)-marginal \[ \mu(ds_1)\gamma_{p,d}(dB)\upsilon_{n,p}(dC)\,p_D(Ds_1)^\ell, \tag{153}\] and the rank argument for \(R\) removes the loci where \(U\) or \(BU\) has rank less than \(\ell\). Multiplication by the finite weight in (153) preserves these native null sets. On either incidence set, the spherical partial differentials are regular on the remaining locus. Indeed, if some \(s_i\) belonged to the row space of \(A\), then its row-space projection would have norm one. The equations \(As_j=As_i\) would give every \(s_j\) that same row-space projection, and their unit norms would force \(s_j=s_i\) for all \(j\), contradicting \(\operatorname{rank}U=\ell\). The same argument applies to \(D\). The Gaussian rank assertions above also remove row-rank failures. Thus removing the complements of the full-rank regular loci changes neither \(\mathsf M_A,\mathsf M_C\) nor the corresponding \(\mathsf H_A,\mathsf H_C\). Fix \(s_1\) and apply Lemma 39 to \[x=(s_2,\ldots,s_k),\qquad z=A,\qquad G(x,z)=AU.\] For fixed \(A\), the target has \(\ell\) independent \(n\)-coordinate blocks, so the partial normal Jacobian is \(\prod_{i=2}^k j_A(s_i)\). These are precisely the denominators in \(\prod_{i=2}^k\lambda_{A,As_1}(ds_i)\). For fixed tuple, the differential is \(\Delta A\mapsto\Delta A\,U\). It consists of \(n\) row maps with Gram matrix \(U^{\mathsf T}U\); its normal Jacobian is \(\det(U^{\mathsf T}U)^{n/2}\). The \(A\)-fiber has rows in \(U^\perp\). Restricting the ambient standard Gaussian density to this subspace and then expressing it relative to the normalized Gaussian \(\gamma_U\) contributes \((2\pi)^{-n\ell/2}\). This proves (148) on the regular locus, and hence exactly after the null loci just proved are removed. In particular, integrating over a restricted regular tuple fiber gives at most \(p_A(As_1)^\ell\), and gives equality for native almost every \((s_1,A)\). For the second identity, fix \((s_1,B)\) and use \(G(x,C)=CBU\). Put \(W=BU\). The tuple partial normal Jacobian is \(\prod_{i=2}^k j_D(s_i)\). At \(CW=0\), \[T_C\mathcal V_{n,p} =\{\Delta:\Delta C^{\mathsf T}+C\Delta^{\mathsf T}=0\}.\] Variations whose rows belong to \(\operatorname{col}W\) lie in this tangent space, since \(C W=0\). They form the orthogonal complement of the tangent space to the constrained frame fiber. On this normal space the differential is \(\Delta\mapsto\Delta W\), with normal Jacobian \(\det(W^{\mathsf T}W)^{n/2}\). There is no additional metric factor in these normal directions. The constrained fiber, with its induced metric, is isometric to \(\mathcal V_{n,p-\ell}\). Replacing its unnormalized volume by \(\upsilon_W\) leaves the volume ratio \(\kappa\). Lemma 39 now proves (149), again exactly after removal of native rank-null sets. The corresponding restricted tuple mass is at most \(p_D(Ds_1)^\ell\) and equals it for native almost every \((s_1,B,C)\). We next transfer the density exceptions using these undivided measures. For a full-rank \(n\times d\) matrix \(L\), define \[b_j(L,t)= \frac{\mu\{z:\|Lz-t\|<2^{-j}\}}{v_n2^{-jn}},\] and let \(\mathcal E(L,s)\) be the condition that either \(p_L(Ls)\notin(0,\infty)\), or the limit of \(b_j(L,Ls)\) does not exist, or that limit differs from \(p_L(Ls)\). This is a measurable condition. Set \[\mathcal E_A=\{(\mathbf s,A):\mathcal E(A,s_1)\},\qquad \mathcal E_D=\{(\mathbf s,B,C):\mathcal E(CB,s_1)\}.\] For every fixed full-rank \(L\), Lebesgue differentiation and the density formula imply \(\mu\{s:\mathcal E(L,s)\}=0\). The exact native marginals (152) and (153), whose density weights are finite for each fixed native projection, therefore give \[\mathsf M_A(\mathcal E_A)=0,\qquad \mathsf M_C(\mathcal E_D)=0.\] On the right sides of (148) and (149), the multiplying Gram weights are strictly positive and finite almost surely. If \(h\) denotes either weight, then on \(\{1/m\le h\le m\}\) the equality \(\int_{\mathcal E}h\,d\mathsf H_\bullet=0\) gives \(\mathsf H_\bullet(\mathcal E\cap\{1/m\le h\le m\})=0\). Letting \(m\to\infty\) proves that \(\mathcal E_A\) is \(\mathsf H_A\)-null and \(\mathcal E_D\) is \(\mathsf H_C\)-null. This argument is valid even if an undivided measure has infinite total mass. Adding the unused probability kernels gives both assertions under \(\mathsf H\). Only now replace each unnormalized fiber measure by its normalized conditional law. The \(Q\) experiment divides (148) by \(p_A(As_1)^\ell\), then attaches the same kernels for \(B,C\) and the message density \(w\phi_J(A,As_1)\) relative to uniform \(J\). This gives (150). The \(R\) experiment divides (149) by \(p_D(Ds_1)^\ell\) and attaches the unchanged \(\gamma_U\) kernel for \(A\) and uniform \(J\), giving (151). The divisions are legitimate on the full-measure set just proved. Since \(r_*>0\) and is finite there, \(Q\ll R\). ◻ Lemma 41 (The induced Stiefel volume). For \(1\le n\le h\), with the induced Euclidean matrix metric, \[ |\mathcal V_{n,h}|= 2^{n(n-1)/4}\prod_{j=1}^n|\mathbb S^{h-j}|. \tag{154}\] Consequently, for (136), \[ \kappa=\pi^{-n\ell/2}\prod_{j=1}^n \frac{\Gamma((p-j+1)/2)}{\Gamma((p-j+1-\ell)/2)} \ge\left(\frac{p-n+1-\ell}{2\pi}\right)^{n\ell/2}. \tag{155}\] Proof. For \(n=1\), the frame manifold is \(\mathbb S^{h-1}\), giving the base case of (154). For \(n\ge2\), the map from \(\mathcal V_{n,h}\) to its first row in \(\mathbb S^{h-1}\) has constant normal Jacobian \(2^{-(n-1)/2}\). To compute it, move the frame by an orthogonal transformation to the first \(n\) coordinate rows. Tangent matrices then consist of a skew-symmetric block in the first \(n\) columns and an arbitrary remaining block. Each of the \(n-1\) skew coordinates affecting the first row appears twice in the squared matrix norm, so the corresponding singular value of the first-row map is \(2^{-1/2}\). The other first-row coordinates have singular value one. The fiber is isometric to \(\mathcal V_{n-1,h-1}\). Slicing thus gives \[|\mathcal V_{n,h}|= 2^{(n-1)/2}|\mathbb S^{h-1}|\,|\mathcal V_{n-1,h-1}|.\] Iteration proves (154). The metric factor depends only on \(n\), so it cancels in the ratio defining \(\kappa\). The sphere-area formula \(|\mathbb S^{h-1}|=2\pi^{h/2}/\Gamma(h/2)\) gives the product in (155). For \(a>0\) and \(\ell/2\ge1\), a gamma random variable \(X\) of shape \(a\) and scale one satisfies \(\Gamma(a+\ell/2)/\Gamma(a)=\mathbb E X^{\ell/2} \ge(\mathbb E X)^{\ell/2}=a^{\ell/2}\) by Jensen’s inequality. Apply this with \(a=(p-j+1-\ell)/2\) and then use \(a\ge(p-n+1-\ell)/2\). ◻ An inverse kernel at two scalesThe density ratio in Lemma 40 involves \(p_D(Ds_1)/p_A(As_1)\). It has no useful pointwise bound. We control its positive logarithm instead, using the fact that \(A\) annihilates the same difference space that governs the conditional Gaussian law of \(D\). For \(u>0\), write \(\log^+u=\max\{0,\log u\}\), and set \(\log^+0=0\). Lemma 42 (Two-scale inverse kernel). Under \(Q\), define \[\begin{split} K&=\int \operatorname{dist}(z-s_1,\operatorname{col}U)^{-n}\,d\mu(z),\\ p_A^*(t)&=\sup_{\rho>0} \frac{\mu\{z:\|Az-t\|<\rho\}}{v_n\rho^n},\qquad K_0=v_n\|A\|_{\mathrm{op}}^n p_A^*(As_1). \end{split}\] Then \(K_0\ge2^{-n}\), and \[ K\le K_0L_F,\qquad L_F=1+\frac{n}{d-\ell-n}+n\log(2/\tau),\qquad \tau=\left(\frac{2^{-n}}{Fd^d}\right)^{1/(d-\ell-n)}. \tag{156}\] In particular, \(\log L_F\le C\log d+C\log(1+\log F)\), and \(K<\infty\). Moreover, \[ \mathbb E_Q[p_D(Ds_1)\mid\mathbf s,A,y,J] \le (2\pi)^{-n/2}K . \tag{157}\] Proof. For fixed \(z\), (140) makes the coordinates of \(D(z-s_1)\) independent centered Gaussians with variance \(\operatorname{dist}(z-s_1,\operatorname{col}U)^2\). Hence \[\frac{\Pr_Q\{\|D(z-s_1)\|<\rho\mid\mathbf s,A,y,J\}}{v_n\rho^n} \le (2\pi)^{-n/2} \operatorname{dist}(z-s_1,\operatorname{col}U)^{-n}.\] The right side is interpreted as \(+\infty\) when the distance is zero. Integrate in \(z\) using Tonelli’s theorem and then apply conditional Fatou along \(\rho=2^{-j}\). Lemma 40 and \(Q\ll\mathsf H\) identify the almost-sure limit of the ball averages with \(p_D(Ds_1)\). This proves (157) without evaluating an arbitrary density version at a dependent point. Let \(H(t)\) be the \(\mu\)-probability that the distance in \(K\) is at most \(t\). Since \(AU=0\), \[\operatorname{dist}(z-s_1,\operatorname{col}U)\le t \quad\Longrightarrow\quad \|A(z-s_1)\|\le\|A\|_{\mathrm{op}}t.\] The projection law has a Lebesgue density, so projection-ball boundaries have zero mass. The definition of \(p_A^*\) therefore gives \(H(t)\le K_0t^n\). All these distances are at most two; taking \(t\downarrow2\) from above also gives \(K_0\ge2^{-n}\). A second estimate uses the projection onto \(U^\perp\), of dimension \(d-\ell\le d-2\). Applying Lemma 38 to a matrix with orthonormal rows for this projection bounds its density by \(Fd^d/v_{d-\ell}\). A ball of radius \(t\) centered at the projection of \(s_1\) consequently has probability at most \(Fd^dt^{d-\ell}\). Thus \[ H(t)\le\min\{K_0t^n,\;Fd^dt^{d-\ell}\}. \tag{158}\] The exponent \(d-\ell\) uses the bounded projection density and the condition \(\ell\ge2\). The layer-cake formula, including the possibility of an initially infinite integral, is \[K=2^{-n}+n\int_0^2H(t)t^{-n-1}\,dt.\] Here \(d-\ell-n>0\) and \(\tau\le1\). Use the second estimate in (158) up to \(\tau\) and the first thereafter: \[\begin{split} K&\le2^{-n} +\frac{nFd^d\tau^{d-\ell-n}}{d-\ell-n} +nK_0\log(2/\tau)\\ &\le K_0\left(1+\frac{n}{d-\ell-n}+n\log(2/\tau)\right). \end{split}\] The last inequality uses \(Fd^d\tau^{d-\ell-n}=2^{-n}\le K_0\). Finally, \[\log(1/\tau) =\frac{n\log2+\log F+d\log d}{d-\ell-n}.\] The dimension ratios in (136) give \(L_F\le C(1+d\log d+\log F)\). Taking its logarithm gives the stated bound and finiteness of the inverse-distance integral \(K\). ◻ Lemma 43 (Maximal average at an actual label). For every full-rank \(A\), \[ \int\log\frac{p_A^*(As)}{p_A(As)}\,d\mu(s) \le Cd+C\log(1+\log F). \tag{159}\] The integrand is nonnegative almost surely. Proof. Lebesgue differentiation gives \(p_A^*(t)\ge p_A(t)\) for \(A_\#\mu\)-almost every label. We prove the particular centered maximal inequality we need. Its one-dimensional ancestry is the maximal-function theory of Hardy and Littlewood (Hardy and Littlewood 1930, sec. III); for the dimension cost of customary covering arguments, see Stein and Strömberg (Stein and Strömberg 1983, 259). For \(a>0\), a compact subset of \(\{t:p_A^*(t)>a\}\) has a finite cover by balls centered in that set, each of projection mass greater than \(a\) times its volume. Select disjoint balls in decreasing order of radius. Every discarded ball is contained in the triple of a selected ball, since it meets a selected ball of at least its radius. The selected masses sum to at most one, so the compact set has volume at most \(3^n/a\). Inner regularity gives \[ |\{t:p_A^*(t)>a\}|\le3^n/a\qquad(a>0). \tag{160}\] The function is measurable: the supremum may be restricted to positive rational radii, because a density gives zero mass to ball boundaries. Let \(E_A\) be the projection ellipsoid and \(V_A=|E_A|\). Since \(p_A^*\le\|p_A\|_\infty\), integration of level sets, using the bound \(V_A\) up to level \(1/V_A\) and (160) above that level, gives \[\int_{E_A}p_A^*(t)\,dt \le1+3^n\log^+(V_A\|p_A\|_\infty) \le1+3^n(\log F+d\log d).\] Jensen’s inequality with respect to the probability density \(p_A\) then yields \[\int p_A(t)\log\frac{p_A^*(t)}{p_A(t)}\,dt \le\log\int_{\{p_A>0\}}p_A^*(t)\,dt \le\log\int_{E_A}p_A^*(t)\,dt.\] Values on \(\{p_A=0\}\) do not enter the integral. The preceding estimate and \(n\asymp d\) prove (159). ◻ Proof of Proposition 36. On the full-measure set in Lemma 40, the log likelihood ratio is the sum of the message term \(\log(w\phi_J(A,As_1))\), the term \[ -\frac{n\ell}{2}\log(2\pi)-\log\kappa+ \frac n2\log \frac{\det(U^{\mathsf T}B^{\mathsf T}BU)}{\det(U^{\mathsf T}U)}, \tag{161}\] and \(\ell\log(p_D(Ds_1)/p_A(As_1))\). The message term is at most \(\log w\). Under \(Q\), \(B\) is independent of \(U\). Conditional on \(U\), put \(V=U(U^{\mathsf T}U)^{-1/2}\). Then \(G=BV\) is a standard Gaussian \(p\times\ell\) matrix and \[\frac{\det(U^{\mathsf T}B^{\mathsf T}BU)}{\det(U^{\mathsf T}U)} =\det(G^{\mathsf T}G).\] The lower bound in (155) and the determinant–trace inequality bound (161) above by \[\frac{n\ell}{2}\log \frac{\operatorname{tr}(G^{\mathsf T}G)} {\ell(p-n+1-\ell)}.\] Its positive part has expectation \(O(n\ell)\). Indeed, \(\mathbb E\operatorname{tr}(G^{\mathsf T}G)=p\ell\), the ratio \(p/(p-n+1-\ell)\) is bounded, and \(\mathbb E\log^+X\le\log(1+\mathbb E X)\) for \(X\ge0\). For the remaining ratio, use \(\log^+u\le\log(1+u)\), conditional Jensen, and (157) and (156). Since \(p_A^*(As_1)/p_A(As_1)\ge1\) almost surely, this gives \[\begin{align*} &\mathbb E_Q\!\left[ \log^+\frac{p_D(Ds_1)}{p_A(As_1)} \,\middle|\,\mathbf s,A,y,J\right]\\ &\quad\le \log2+ \log^+\!\bigl((2\pi)^{-n/2}v_n\|A\|_{\mathrm{op}}^n\bigr) +\log L_F+\log\frac{p_A^*(As_1)}{p_A(As_1)}. \tag{162}\end{align*}\] The Gaussian operator term has expectation \(O(n)\). Here are the elementary estimates. A \(1/4\)-net of the unit sphere in \(\mathbb R^h\) has size at most \(9^h\), by packing disjoint radius \(1/8\) balls. Nets in the two dimensions bound \(\|A\|_{\mathrm{op}}\) by twice the maximum of at most \(9^{n+d}\) absolute standard Gaussian bilinear forms. A Gaussian tail bound gives \(\mathbb E\|A\|_{\mathrm{op}}\le C\sqrt d\). Also, integrating the Gaussian density over the ball of radius \(\sqrt n\) gives \(v_n\le(2\pi e/n)^{n/2}\). Therefore \[\log^+\!\bigl((2\pi)^{-n/2}v_n\|A\|_{\mathrm{op}}^n\bigr) \le n\log(1+\sqrt{e/n}\,\|A\|_{\mathrm{op}}),\] whose expectation is \(O(n)\) by Jensen and \(n\asymp d\). The \((A,s_1)\)-marginal under \(Q\) is the original independent Gaussian/prior draw. Averaging (162) and using Lemmas 42 and 43 consequently gives \[ \mathbb E_Q\log^+\frac{p_D(Ds_1)}{p_A(As_1)} \le Cd+C\log(1+\log F). \tag{163}\] All positive parts in the log likelihood ratio are now integrable. Since \(Q\ll R\), its negative part is integrable as well: with \(h=dQ/dR\), \(\int_{\{h<1\}}-h\log h\,dR\le1/e\). Thus no undefined subtraction occurs, and \[\operatorname{KL}(Q\|R) \le\log w+Ckd+Ck\log(1+\log F).\] Combine this with (142) and divide by \(k\). ◻ Finite histories and the precision endpointWe state the consequence first for average success under the uniform spherical prior, using the exact finite-state randomized model in Definition 3. Theorem 44 (Average-prior precision bound). There is a universal \(c>0\) with the following property. Let \(M=M(d)\) be a nonnegative integer with \(M(d)=o(d^2)\). For all sufficiently large \(d\), with the threshold allowed to depend on the rate \(M(d)/d^2\to0\), suppose \(S\sim\sigma\) and a learner in Definition 3 stops by a deterministic integer horizon \(T\ge0\). If \[\Pr\{\angle(\widehat S,S)\le\epsilon\}\ge2/3, \qquad 0<\epsilon\le1/10,\] then \[ T\ge c\,d\log(1/\epsilon). \tag{164}\] The bound permits arbitrary computation within a transition. For the proof of Theorem 44 and its auxiliary lemmas below, apply the fixed-shared-seed part of Lemma 9 to the stated learner in Definition 3. We work in the resulting Borel experiment with data-independent initialization and uniform-prior success at least \(2/3\), retaining fresh transition and output randomness. This is the standing reduction for this subsection. Lemma 45 (Short-horizon reduction). Put \(L_\epsilon=\log(1/\epsilon)\). Pad stopping to \(T\) by retaining the stopped state and its stopping index. At most \[ w=(T+2)2^M \tag{165}\] padded states suffice. Under the hypotheses of Theorem 44, if \(T<dL_\epsilon\), then \[ L_\epsilon=O(1+M/d)=o(d),\qquad T=O(d^2),\qquad \log w=o(d^2). \tag{166}\] Proof. The padding count in Lemma 9, including stopping at index zero, gives (165). Its stopped records remain fixed while fresh unused observations are generated. The output-capacity bound in Lemma 11, with success at least \(2/3\) and \((T+1)2^M\le w\), gives \[\frac23\le w\epsilon^{d-1},\qquad (d-1)L_\epsilon\le M\log2+\log(T+2)+\log(3/2).\] Because \(L_\epsilon\ge\log10\) and \(T<dL_\epsilon\), \(\log(T+2)\le\log d+\log L_\epsilon+C\). Using \(\log L_\epsilon\le L_\epsilon\) yields \[(d-2)L_\epsilon\le M\log2+\log d+C.\] This proves \(L_\epsilon=O(1+M/d)=o(d)\). The remaining assertions follow from \(T<dL_\epsilon\) and \(M=o(d^2)\). ◻ Lemma 46 (Information accumulated over histories). In the short-horizon case of Lemma 45, put \(b=\lceil T/n\rceil\) and pad with fresh unused samples to \(bn\) observations. Let \(V_i\) be the transcript of the first \(i\) block-end padded states, with \(V_0\) trivial. Let \(B\sim\gamma_{p,d}\) be independent of the entire learner experiment and used only for analysis. Then, for all sufficiently large \(d\), \[ I(S;V_b\mid B,BS)\le Cbd \tag{167}\] for an absolute \(C\). The dimension threshold may depend on the rate \(M/d^2\to0\). Proof. For a positive-probability history \(V_{i-1}=v\), Lemma 10 gives the Bayes density \[f_v(s)=\frac{\Pr\{V_{i-1}=v\mid S=s\}}{\Pr\{V_{i-1}=v\}} \le F_v:=\frac1{\Pr\{V_{i-1}=v\}}\] relative to \(\sigma\). A Borel version is available from the Borel learner. The history fixes the previous padded state. Its update over the next fresh block is a kernel of \((A,AS)\) to at most \(w\) next states. For the first block, independent initialization can be integrated into the same kernel. The matrix \(B\) is absent from the learner and independent of its entire experiment; it remains independent of the entire conditional signal, fresh-block, and message experiment after conditioning on \(v\). Proposition 36 therefore applies at each such history. The finite history has entropy \(H(V_{i-1})\le b\log w\). Hence Jensen’s inequality gives \[\mathbb E_v\log(1+\log F_v) \le\log(1+H(V_{i-1})) \le\log(1+b\log w)=O(\log d).\] The last bound uses \(b=O(d)\) and \(\log w=O(d^2)\) eventually, from (166) and \(n\asymp d\). The same bounds give \(\log w/k=o(d)\). Averaging the one-block estimate thus bounds each increment \(I(S;V_i\mid V_{i-1},B,BS)\) by \(Cd\) eventually. The conditional chain rule, with the same independent \(B\) throughout, proves (167). ◻ Lemma 47 (Precision inside an independent fiber). Let \(S\sim\sigma\), and let a unit-vector output \(\widehat S\) have angular error at most \(\epsilon\le1/10\) with probability at least \(2/3\). Let \(B\sim\gamma_{p,d}\), \(p=\lfloor d/4\rfloor\), be independent of the entire experiment producing \((S,\widehat S)\). Then, for all sufficiently large \(d\), \[ I(S;\widehat S\mid B,BS) \ge \frac{d-p-1}{3}\log\frac1{4\epsilon}-\log2 \ge c_1dL_\epsilon \tag{168}\] with an absolute \(c_1>0\). In the learner model above, success probability at least \(2/3\) is also impossible for \(T\le d/4\), even if all observations are retained. Proof. Apply Lemma 6 with \(V=\widehat S\), the identity output kernel, \(b=p\), and \(r_0=1/2\). For the residual radius \(R=\|P_{\ker B}S\|\), Lemma 5 gives \[\Pr\{R<1/2\}\le\frac{4p}{3d}\le\frac13.\] Success at least \(2/3\) therefore leaves event mass at least \(1/3\), and (6) gives the first inequality in (168). The second follows from \[\log\frac1{4\epsilon} \ge\left(1-\frac{\log4}{\log10}\right)L_\epsilon,\] because \(d-p-1\) is an absolute positive fraction of \(d\); the subtractive \(\log2\) is absorbed for large \(d\). For the final assertion, grant the learner all pre-generated observations through its horizon, including ignored rows after stopping. Its output is a kernel of those data and independent randomness. Lemma 12 shows that success at least \(3/5\), and hence at least \(2/3\), already requires \(T>\lfloor d/2\rfloor\) for large \(d\). In particular, \(T>d/4\). That lemma includes the deterministic zero-row residual sphere, so the assertion covers \(T=0\) as well. ◻ Proof of Theorem 44. If \(T\ge dL_\epsilon\), the desired bound is immediate. Otherwise apply Lemmas 45 and 46. The terminal output is a kernel of the final padded state, hence of \(V_b\); conditional data processing and Lemma 47 give \[Cbd\ge I(S;V_b\mid B,BS) \ge I(S;\widehat S\mid B,BS) \ge c_1dL_\epsilon.\] Thus \(b\ge c_0L_\epsilon\) for an absolute \(c_0>0\). If \(L_\epsilon\ge2/c_0\), then \[T>n(b-1)\ge(c_0/2)nL_\epsilon\ge c\,dL_\epsilon\] after adjusting an absolute constant. If \(L_\epsilon<2/c_0\), the separate conclusion \(T>d/4\) of Lemma 47 gives the same order after another absolute adjustment. This proves (164) uniformly over the allowed \(\epsilon\). ◻ A guarantee of success probability at least \(2/3\) for every fixed unit signal implies the uniform-prior average guarantee, so it has the same consequence. The argument uses fresh exact labels, a finite persistent state alphabet, and a deterministic horizon; its additional projections and tuple resampling belong only to the analysis.
Austin, Tim. 2020. “Multi-Variate Correlation and Mixtures of Product Measures.” Kybernetika 56 (3): 459–99. https://doi.org/10.14736/kyb-2020-3-0459.
Bassily, Raef, Shay Moran, Ido Nachum, Jonathan Shafer, and Amir Yehudayoff. 2018. “Learners That Use Little Information.” Proceedings of Algorithmic Learning Theory, Proceedings of machine learning research, vol. 83: 25–55. https://proceedings.mlr.press/v83/bassily18a.html.
Dagan, Yuval, Gil Kur, and Ohad Shamir. 2019. “Space Lower Bounds for Linear Prediction in the Streaming Model.” Proceedings of the Thirty-Second Conference on Learning Theory, Proceedings of machine learning research, vol. 99: 929–54. https://proceedings.mlr.press/v99/dagan19b.html.
Drury, S. W. 1984. “Generalizations of Riesz Potentials and \(L^p\) Estimates for Certain \(k\)-Plane Transforms.” Illinois Journal of Mathematics 28 (3): 495–512.
Dupuis, Paul, and Yixiang Mao. 2022. “Formulation and Properties of a Divergence Used to Compare Probability Measures Without Absolute Continuity.” ESAIM: Control, Optimisation and Calculus of Variations 28. https://doi.org/10.1051/cocv/2022002.
Edelman, Alan, Tomás A. Arias, and Steven T. Smith. 1998. “The Geometry of Algorithms with Orthogonality Constraints.” SIAM Journal on Matrix Analysis and Applications 20 (2): 303–53. https://doi.org/10.1137/S0895479895290954.
Edelman, Alan, and N. Raj Rao. 2005. “Random Matrix Theory.” Acta Numerica 14: 233–97. https://doi.org/10.1017/S0962492904000236.
Federer, Herbert. 1959. “Curvature Measures.” Transactions of the American Mathematical Society 93 (3): 418–91. https://doi.org/10.1090/S0002-9947-1959-0110078-1.
Frankl, Peter, and Hiroshi Maehara. 1990. “Some Geometric Applications of the Beta Distribution.” Annals of the Institute of Statistical Mathematics 42 (3): 463–74. https://doi.org/10.1007/BF00049302.
Hardy, G. H., and J. E. Littlewood. 1930. “A Maximal Theorem with Function-Theoretic Applications.” Acta Mathematica 54: 81–116. https://doi.org/10.1007/BF02547518.
Lelarge, Marc, and Léo Miolane. 2019. “Fundamental Limits of Symmetric Low-Rank Matrix Estimation.” Probability Theory and Related Fields 173: 859–929. https://doi.org/10.1007/s00440-018-0845-x.
OpenAI. 2026a. Memory and precision in noiseless Gaussian regression. OpenAI Math Release preprint OAI:Memory-and-precision-in-noiseless-Gaussian-regression-September-27-2026.
OpenAI. 2026b. Replacing Gaussian observations in memory-constrained inference. OpenAI Math Release preprint OAI:Replacing-Gaussian-observations-in-memory-constrained-inference-September-27-2026.
Raz, Ran. 2016. Fast Learning Requires Good Memory: A Time-Space Lower Bound for Parity Learning. https://arxiv.org/abs/1602.05161v1.
Raz, Ran. 2017. “A Time-Space Lower Bound for a Large Class of Learning Problems.” Proceedings of the 58th Annual IEEE Symposium on Foundations of Computer Science, 732–42. https://doi.org/10.1109/FOCS.2017.73.
Sharan, Vatsal, Aaron Sidford, and Gregory Valiant. 2019. “Memory-Sample Tradeoffs for Linear Regression with Small Error.” Proceedings of the 51st Annual ACM SIGACT Symposium on Theory of Computing, 890–901. https://doi.org/10.1145/3313276.3316403.
Stein, E. M., and J.-O. Strömberg. 1983. “Behavior of Maximal Functions in \(\mathbb{R}^n\) for Large \(n\).” Arkiv för Matematik 21: 259–69. https://doi.org/10.1007/BF02384314.
Steinhardt, Jacob, and John Duchi. 2015. “Minimax Rates for Memory-Bounded Sparse Linear Regression.” Proceedings of the 28th Conference on Learning Theory, Proceedings of machine learning research, vol. 40: 1564–87. https://proceedings.mlr.press/v40/Steinhardt15.html.
Watanabe, Satosi. 1960. “Information Theoretical Analysis of Multivariate Correlation.” IBM Journal of Research and Development 4 (1): 66–82. https://doi.org/10.1147/rd.41.0066.
|
| ||||||||
|