A
D
V
E
R
T
I
S
E
M
E
N
T
ADVERTISEMENT
Pointwise Multiple Ergodic Averages for Mixing Transformations
expertly designed by an internal OpenAI model  ·  released 2026-10-04  ·  original PDF
Theorems: 3 Lemmas: 12 Proofs: 18
Formulas: 907 Words: 11,038 Play time: ~1 hour

>>> How to Play <<<
Let T be an invertible mixing probability-preserving transformation. For every integer n ≥ 2 and every fixed tuple of bounded measurable functions, we prove that the consecutive multiple ergodic averages of length n converge almost everywhere to the product of the integrals, as the averaging length tends to infinity through all positive integers. The probability space need not be standard, and no rate of mixing is required.

>>> Level Map <<<
  1. Introduction
  2. Context and prior work
  3. The arbitrary-length step
  4. Dynamical inputs and standard models
  5. Standard models and the two companion inputs
  6. Relative weak mixing and distal factors
  7. Triple convergence along all lengths
  8. Selected lengths and their joining laws
  9. A positive witness at selected lengths
  10. Selected laws on all extensions
  11. Distal parameters and small marginals
  12. Channels and closure under selected laws
  13. Minimal correlations and saturated extensions
  14. Spatial roles and minimum support
  15. Saturating the parameter and single-slot data
  16. Regrouping the other slots as outputs
  17. The square argument
  18. A common corner and two conditional expectations
  19. Independence of the two parameter descriptions
  20. Why the pairing cannot vanish
  21. Increasing the number of output roles
  22. Tensor packing

Introduction

Let \((X,\mathcal F,\mu)\) be a probability space and let \(T:X\to X\) be invertible, bimeasurable, and measure preserving. We say that \(T\) is mixing if \[ \mu(A\cap T^{-r}B)\longrightarrow\mu(A)\mu(B) \qquad (|r|\to\infty) \tag{1}\] for every \(A,B\in\mathcal F\). For a fixed tuple of bounded measurable functions, the consecutive multiple ergodic averages are \[A_N(f_1,\ldots,f_n)(x) =\frac1N\sum_{k=1}^N\prod_{j=1}^n f_j(T^{jk}x).\] The problem is to determine their almost-everywhere behavior, not merely their limit in norm. We prove the following result.

Theorem 1. Suppose that \(T\) satisfies (1). For every integer \(n\ge2\) and every fixed \(f_1,\ldots,f_n\in L^\infty(\mu)\), \[A_N(f_1,\ldots,f_n)(x) \longrightarrow\prod_{j=1}^n\int_X f_j\,d\mu \qquad\text{for $\mu$-almost every }x,\] as \(N\) tends to infinity through all positive integers. The probability space need not be standard, and no rate of mixing is required.

The exceptional null set may depend on the system and the fixed tuple. The theorem concerns the times \(k,2k,\ldots,nk\) and does not require a common exceptional set for all bounded functions.

Context and prior work

Arithmetic-progression averages became central to ergodic Ramsey theory through Furstenberg’s proof of Szemerédi’s theorem (Furstenberg 1977). The distinction between norm and pointwise convergence is substantial. Host and Kra proved \(L^2\) convergence of these averages for every finite length on arbitrary probability-preserving systems (Host and Kra 2005, Theorem 1.1). Ziegler obtained another proof through universal characteristic factors (Ziegler 2007, Corollary 1.8). For pointwise convergence, Bourgain’s double recurrence theorem treats two factors (Bourgain 1990). Higher-length results have been obtained under additional structural assumptions: K-systems (Derrien and Lesigne 1996, Theorem 4.2), weakly mixing systems with singular spectrum on their Pinsker factor (Assani 1998), and weakly mixing systems whose pairwise independent self-joinings are independent (Gutman et al. 2018, Theorem 3.4). Huang, Shao, and Ye proved pointwise convergence for every finite length on ergodic measure-distal systems (Huang et al. 2019, Theorem C).

The present proof extends the fourfold mixing theorem of (OpenAI 2026a, Theorem 1.1). It uses two substantial companion results. The first proves that ordinary mixing implies mixing of every finite order, the assertion of Rokhlin’s multiple-mixing problem (OpenAI 2026b, Theorem 1.1). The second gives an oscillation estimate for triple averages at distinct integer slopes on arbitrary invertible probability-preserving systems (OpenAI 2026c, Proposition 3.5); together with its profile approximation, this yields pointwise triple convergence at distinct positive integer slopes without a mixing assumption. Section 2 states their exact forms and derives the required triple pointwise theorem along all lengths. These companion proofs are not reproduced here. The selected-law construction, arbitrary-length extension argument, and tensor estimate used below are proved in full.

The arbitrary-length step

We briefly describe what must change when the number of factors increases. A failure of pointwise convergence allows an averaging length to be chosen measurably at each starting point so that a centered average stays positive. After passing to a suspension flow, these choices define a probability law on \(n\) output coordinates. The law has the correct individual marginals and is invariant under different speeds in the different coordinates. The same construction works on extensions of the flow.

The method of (OpenAI 2026a, secs. 3–6) organizes additional outputs into three groups. Their spatial coordinates are independent, and the original spatial coordinate is independent of any two groups; independence from all three groups is not imposed. We retain this three-group structure even when there are \(n\) slots. Higher-order mixing and a tensor-packing estimate show that these independence properties survive the selected joining construction.

Two changes make this approach work for arbitrary \(n\). First, we minimize the number of slots supporting a nonzero centered correlation. This gives independence for every smaller collection of the specific spatial lists being tested. A slot may read both the original spatial coordinate and the outputs in any proper subset of the three groups. That flexibility permits the outputs to be regrouped without imposing independence on arbitrary collections of whole states.

Second, an extension is chosen so that conditioning on all distal parameters together with a single state cannot improve under any further extension. This is a conditional-expectation energy construction, related to the sated-extension method (Austin 2015) and adapted from (OpenAI 2026a, sec. 6). Applying the joining construction again gives a square array. The single-state conditioning identity controls its distinguished row and column. Their parameter laws and a Hilbert-space argument then produce a new nonzero correlation with the same number of slots but one more output-only slot. Maximizing that number gives the contradiction. Only the pointwise theorem for at most three slots on arbitrary systems is used.

Section 2 records the dynamical inputs. Section 3 constructs the selected laws and their conditional marginals. Section 4 proves preservation of the three-group structure. Sections 5 and 6 carry out the minimal-correlation and square arguments. Appendix 7 proves the tensor estimate.

Dynamical inputs and standard models

Fix an integer \(n\ge2\) and put \(I=\{1,\ldots,n\}\). We first state the dynamical results used below. The distinction between norm and pointwise convergence matters: norm convergence is available for any finite number of slots, whereas the pointwise input on arbitrary systems will be used only for at most three slots.

Standard models and the two companion inputs

For the fixed tuple in Theorem 1, record all integer translates of bounded measurable representatives of the functions. This gives a measurable map into a countable product of compact scalar discs, intertwining \(T\) with the invertible shift. The pushforward measure defines a standard Borel probability system, and mixing passes to this factor. The original functions are its coordinate functions. Changing representatives affects at most a countable union of translated null sets, so almost-everywhere convergence on the factor pulls back. We may therefore work on standard probability spaces.

Throughout the proof, factors and conditional laws are understood modulo null sets. We use Borel versions of measurable functions. Sub-\(\sigma\)-algebras on a standard probability space have standard models modulo null sets. A flow is a jointly Borel action \((S_Y^t)_{t\in\mathbb R}\) on a standard Borel probability space \((Y,m_Y)\), preserving \(m_Y\). Factor maps commute with every fixed time modulo null sets. Joint measurability permits Fubini in all time integrals. For a preserved factor, induced transformations at fixed times can be modeled as invertible maps; we take time integrals on the original flow.

We use the following statement from (OpenAI 2026b, Theorem 1.1).

Theorem 2 (Higher-order mixing, companion input). Let \(T\) be an invertible mixing probability-preserving transformation. For any fixed \(s\ge2\), \(g_1,\ldots,g_s\in L^\infty(\mu)\), and integer times \(t_1,\ldots,t_s\), \[\int_X\prod_{j=1}^s g_j(T^{t_j}x)\,d\mu(x) \longrightarrow\prod_{j=1}^s\int_Xg_j\,d\mu \quad\text{as}\quad \min_{i\ne j}|t_i-t_j|\longrightarrow\infty.\] For \(s=1\) the equality holds for every time.

The cited theorem is stated for sets and ordered times. Simple-function approximation, invariance, and the finitely many orderings of the times give this formulation. No quantitative mixing rate is asserted or used.

Our second companion input is the smooth-average convergence consequence of (OpenAI 2026c, Proposition 3.5 and Lemma 3.6). An averaging profile there is a real compactly supported smooth function \(\varphi\) of mass one with moments of orders \(1,2,3\) equal to zero, and, for some \(G\ge1\), \[\lVert\varphi^{(r)}\rVert_\infty \le \bigl(G(r+2)\bigr)^{G(r+2)}\qquad(r\ge0).\] The constant \(G\) may depend on the profile. We need the following two precise conclusions.

Proposition 3 (Smooth triple averages, companion input). Let \(R\) be any invertible probability-preserving transformation, and let \(b_1,b_2,b_3\) be distinct nonzero integers with greatest common divisor one. For bounded \(g_1,g_2,g_3\) and every averaging profile, \[V_k^\varphi(y)=\sum_{l\in\mathbb Z}2^{-k}\varphi(l/2^k) \prod_{j=1}^3g_j(R^{b_jl}y)\] converges almost everywhere as \(k\to\infty\). Moreover, for each \(c\in[1,2]\) and \(\epsilon>0\) there is such a profile with \[\lVert\varphi-c^{-1}\mathbf 1_{(0,c]}\rVert_{L^1(\mathbb R)}<\epsilon.\]

The approximation is (OpenAI 2026c, Lemma 3.9). Neither part requires mixing or ergodicity. We will not assume pointwise convergence of longer averages on arbitrary systems.

Relative weak mixing and distal factors

We recall the structure facts in the form needed for the auxiliary systems. An extension of an invertible system \(R\) over a factor \(B\) is relatively weakly mixing if, for bounded \(f,g\) with \(\mathbb E[f\mid B]=0\), \[ \frac1H\sum_{h=1}^H \lVert\mathbb E[\overline f(g\circ R^h)\mid B]\rVert_2^2\longrightarrow0. \tag{2}\] An extension is relatively compact if the invariant finitely generated \(L^\infty(B)\)-modules span a dense subspace of its \(L^2\) space. A module consists of finite sums \(\sum_j a_jF_j\), with \(a_j\in L^\infty(B)\) and fixed \(F_j\in L^2\); invariance is under composition with both \(R\) and \(R^{-1}\). Individual modules need not be closed; see the corrected module characterization in (Jamneshan 2026, Theorem 2.5(ii)\('\)). A relatively distal extension is a tower of relatively compact extensions and inverse limits. A system is distal if it is distal over the trivial factor. A joining is an invariant probability law on a product with the prescribed marginals.

Proposition 4 (Relative structure). The following statements hold for standard probability-preserving systems, without an ergodicity assumption.

  1. Every system \(Y\) has a maximal measure-distal factor \(W_Y\), and \(Y\to W_Y\) is relatively weakly mixing. Every distal factor pulls into the maximal distal factor of an extension. Automorphisms commuting with the transformation preserve \(W_Y\).

  2. Factors and joinings of distal systems are distal. A distal system is relatively distal over each of its factors. Distality is equivalent for a transformation and any of its nonzero integer powers.

  3. Relative weak mixing passes to nonzero integer powers. The conditional product of relatively weakly mixing fibre laws over an invariant joining of their base laws is relatively weakly mixing over that joining.

  4. A relatively weakly mixing extension and a relatively distal extension of the same base are relatively disjoint: every invariant joining identifying their bases is the conditional product over that base.

  5. For a distal transformation \(R\), every \(d\ge1\), and bounded \(g_1,\ldots,g_d\), the averages \[\frac1N\sum_{l=1}^N\prod_{i=1}^d g_i\circ R^{il}\] converge almost everywhere.

For a flow, the maximal distal factor of time one is preserved by the flow, and its action at every nonzero rational time is distal.

These are the conclusions of (OpenAI 2026a, Proposition 2.2). They include the nonergodic form of the Furstenberg–Zimmer structure theorem (Zimmer 1976; Jamneshan 2023) and the pointwise theorem for distal systems (Huang et al. 2019, Theorem C). The cited proposition gives the component-disintegration argument needed for the latter in the nonergodic setting. In the flow assertion, invariance of \(W_Y\) follows because each time commutes with time one. If \(r=p/q\ne0\) is rational, the \(q\)th power of its action is the distal action at integer time \(p\), so the power assertion applies.

The following short argument, from (OpenAI 2026a, Lemma 2.3), will also fix the norm-convergence input explicitly.

Lemma 5 (Relative norm convergence). Suppose \(Y\to B\) is relatively weakly mixing for \(R\). For distinct nonzero integers \(a_1,\ldots,a_d\) and bounded \(G_1,\ldots,G_d\), \[ \left\|\frac1N\sum_{l=1}^N \left(\prod_{i=1}^dG_i\circ R^{a_il} -\prod_{i=1}^d\mathbb E[G_i\mid B]\circ R^{a_il}\right)\right\|_2 \longrightarrow0. \tag{3}\]

Proof. By telescoping and scaling, it suffices to show that the average of \(v_l=\prod_i G_i\circ R^{a_il}\) tends to zero when all functions are bounded by one and one of them, \(G_s\), is centered over \(B\). Induct on \(d\). For fixed \(h\ge1\) put \(F_{i,h}=\overline{G_i}(G_i\circ R^{a_ih})\). Invariance gives \[\langle v_l,v_{l+h}\rangle =\int_Y F_{s,h}\prod_{i\ne s} F_{i,h}\circ R^{(a_i-a_s)l}\,dm_Y.\] The induction hypothesis replaces the factors indexed by \(i\ne s\) in their average by their conditional expectations over \(B\). The remaining product is \(B\)-measurable and bounded by one; hence \[\limsup_{N\to\infty} \left|\frac1N\sum_{l=1}^N\langle v_l,v_{l+h}\rangle\right| \le\lVert\mathbb E[F_{s,h}\mid B]\rVert_1.\] For \(d=1\) this bound holds directly. The Hilbert-space van der Corput inequality now bounds the limiting squared norm by a constant times \[\frac1H+\frac1H\sum_{h=1}^H\lVert\mathbb E[F_{s,h}\mid B]\rVert_1.\] Relative weak mixing for \(R^{a_s}\) and Cauchy–Schwarz make this expression tend to zero as \(H\to\infty\). ◻

For positive slopes, inserting functions \(1\) in omitted positions and using Proposition 4(5) shows that the distal-factor averages have an \(L^2\) limit. Thus Lemma 5 gives an \(L^2\) limit for arbitrary bounded tuples with distinct positive slopes on every standard system. On the original mixing system, the trivial factor already satisfies (2); the limit is the product of the means.

Triple convergence along all lengths

Lemma 6 (Triple pointwise convergence). On any standard invertible probability-preserving system, averages of at most three bounded functions at distinct positive integer slopes converge almost everywhere along all positive integer lengths.

Proof. Add constant-one functions if necessary to have three distinct positive slopes, write them as \(gb_1,gb_2,gb_3\) with \(\gcd(b_1,b_2,b_3)=1\), and apply Proposition 3 to \(S=R^g\). By scaling, suppose the three functions are bounded by one, and put \(F_l=\prod_{j=1}^3g_j\circ S^{b_jl}\).

For each rational \(c\in[1,2]\), choose a sequence of profiles tending in \(L^1\) to \(h_c=c^{-1}\mathbf1_{(0,c]}\). On a common conull set all the corresponding smooth dyadic averages converge. For \(N_k=\lfloor c2^k\rfloor\), the difference between the ordinary average at \(N_k\) and \(V_k^\varphi\) is at most \[\frac1{c2^k} +2^{-k}\sum_{l\in\mathbb Z}|h_c(l/2^k)-\varphi(l/2^k)|.\] The last sum tends to \(\lVert h_c-\varphi\rVert_1\) by Riemann sums. Arbitrarily close profiles therefore give the Cauchy property on each of these scales. The \(L^2\) limit established above identifies all their almost-everywhere limits with the same function.

For any bounded sequence \(|F_l|\le1\), changing the averaging length from \(N\) to \(M\) changes the average by at most \[ \frac{2|M-N|}{\max(M,N)}. \tag{4}\] For \(2^k\le N<2^{k+1}\), compare \(N\) with \(\lfloor c2^k\rfloor\) on a finite rational grid in \([1,2]\). Equation (4) makes the limiting error arbitrarily small as the grid mesh tends to zero. This proves convergence along all integer lengths. ◻

Selected lengths and their joining laws

We assume that the asserted pointwise convergence fails for the fixed set of slopes \(I=\{1,\ldots,n\}\). This section records that failure in a probability law and constructs compatible laws on all flow extensions. The construction retains a positive correlation while forcing every one-, two-, or three-slot marginal to have its ordinary conditional law. We follow the selector and joining construction in (OpenAI 2026a, secs. 3–6), keeping the extension and conditioning properties explicit for the arbitrary number of slots.

A positive witness at selected lengths

Expanding each function into its mean and centered part, and then into real and imaginary parts, gives a nonempty \(J_0\subseteq I\) and real functions \(f_i\), \(i\in J_0\), such that \[\int_X f_i\,\,d\mu=0,\qquad \lVert f_i\rVert_\infty\le1, \qquad B_L(x)=\frac1L\sum_{l=1}^L\prod_{i\in J_0}f_i(T^{il}x)\] does not tend to zero on a set of positive measure. Here and below \(B_L\) has an integer subscript. Lemma 5, applied over the trivial factor, gives \(B_L\to0\) in \(L^2(\mu)\). After changing the sign of one \(f_i\), there are \(\delta>0\) and a measurable set \(E_*\subseteq X\) of positive measure such that \[ B_L(x)>2\delta\quad\text{for arbitrarily large integers }L, \qquad x\in E_*. \tag{5}\]

We use the roof-one suspension \[ C=X\times[0,1),\qquad m_C=\mu\otimes\,du, \qquad S_C^t(x,u)=\bigl(T^{\lfloor u+t\rfloor}x,\{u+t\}\bigr). \tag{6}\] Write \(\xi(x,u)=x\) and \(\upsilon(x,u)=u\), regarding \(\upsilon\) also as a coordinate in the circle \(\mathbb R/\mathbb Z\).

Lemma 7 (Selected positive witness). There are a finite set \(F\subseteq[1,2]\) containing \(1\) and measurable functions \(N_a:C\to[2^a,\infty)\), \(a\in\mathbb N\), independent of the phase, with \[ N_a(x,u)\in\{c2^k:c\in F,\ k\ge a\},\qquad N_a(x,u)=2^a\quad(x\notin E_*), \tag{7}\] such that \(B_{\lfloor N_a(x,u)\rfloor}(x)>\delta\) on \(E_*\). Moreover, there is a continuous nonnegative function \(\Psi\) on \((\mathbb R/\mathbb Z)^I\) for which \[ \liminf_{a\to\infty}\int_C\frac1{N_a(z)} \int_0^{N_a(z)} \Psi\bigl((\upsilon(S_C^{it}z))_{i\in I}\bigr) \prod_{i\in J_0}f_i\bigl(\xi(S_C^{it}z)\bigr) \,\,dt\,\,dm_C(z)>0. \tag{8}\]

Proof. By (4), changing the length from \(L\) to \(M\) changes an average of a sequence bounded by one by at most \(2\lvert M-L\rvert/\max(M,L)\). Choose \(F\) to be a sufficiently fine finite grid in \([1,2]\). For every sufficiently large integer \(L\), some \(c\in F\) and \(k=\lfloor\log_2L\rfloor\) satisfy \[\lvert B_L(x)-B_{\lfloor c2^k\rfloor}(x)\rvert<\delta \qquad\text{for every }x.\] Equation (5) therefore supplies, for each \(x\in E_*\) and \(a\), a pair \((c,k)\) with \(k\ge a\) at which the latter average exceeds \(\delta\). Choose the first such pair in a fixed enumeration. This is a measurable choice from a countable set and defines \(N_a\) on \(E_*\); use \(2^a\) elsewhere.

Choose continuous nonnegative circle functions \(\alpha,\beta\), neither identically zero, supported respectively inside \((1/4,1/2)\) and \((0,1/(4n))\). Set \[ \Psi(v)=\alpha(2v_1-v_2)\beta(v_2-v_1), \qquad v=(v_i)_{i\in I}\in(\mathbb R/\mathbb Z)^I. \tag{9}\] For \(z=(x,u)\) and \(t=l+s\), where \(l\in\mathbb Z\) and \(0\le s<1\), the phase factor is \(\alpha(u)\beta(s)\). If it is nonzero, then \(0<u+is<1\) for every \(i\in I\), so \(\xi(S_C^{it}z)=T^{il}x\). Summing complete unit intervals gives, uniformly in \(z\) and real \(L\ge2\), \[\begin{align*} &\frac1L\int_0^L \Psi\bigl((\upsilon(S_C^{it}z))_i\bigr) \prod_{i\in J_0}f_i\bigl(\xi(S_C^{it}z)\bigr)\,\,dt \\ &\hspace{18mm} =\alpha(u)\left(\int_0^1\beta(s)\,\,ds\right) B_{\lfloor L\rfloor}(x)+O(L^{-1}). \tag{10}\end{align*}\] The error includes the incomplete last interval, the change from indices \(0,\ldots,\lfloor L\rfloor-1\) to \(1,\ldots,\lfloor L\rfloor\), and the normalization. On \(E_*\) the main term at \(L=N_a(z)\) has integral at least \[\delta\mu(E_*) \left(\int_0^1\alpha(u)\,\,du\right) \left(\int_0^1\beta(s)\,\,ds\right)>0.\] On \(X\setminus E_*\) its integral tends to zero because \(N_a=2^a\) there and \(B_{2^a}\to0\) in \(L^2(\mu)\). The uniform error is \(O(2^{-a})\), proving (8). ◻

Selected laws on all extensions

Fix these selectors and a free ultrafilter \(\mathcal U\) on \(\mathbb N\). A flow over \(C\) is a standard probability space \((Y,m_Y)\) with a jointly Borel probability-preserving action \(S_Y^t\), \(t\in\mathbb R\), and a probability-preserving factor map \(\pi_Y:Y\to C\) satisfying \(\pi_Y S_Y^t=S_C^t\pi_Y\) modulo null sets for each fixed \(t\). An extension \(\rho:Y'\to Y\) of such flows is a probability-preserving factor map, equivariant in the same sense, with \(\pi_{Y'}=\pi_Y\rho\) modulo null sets. We do not require a common conull set on which every real-time factor identity holds. Joint measurability and Fubini justify their uses inside time integrals.

Proposition 8 (Selected joining laws). Every flow \(Y\) over \(C\) has a unique probability law \(\lambda_Y\) on \(Y^I\) satisfying, for all bounded measurable functions \(G_i\) on \(Y\), \[ \int_{Y^I}\prod_{i\in I}G_i(y_i)\,\,d\lambda_Y =\lim_{a\to\mathcal U}\int_Y\frac1{N_a(\pi_Yy)} \int_0^{N_a(\pi_Yy)}\prod_{i\in I}G_i(S_Y^{it}y) \,\,dt\,\,dm_Y(y). \tag{11}\] Its marginals are \(m_Y\), and it is invariant under \((y_i)_i\mapsto(S_Y^{it}y_i)_i\), \(t\in\mathbb R\). If \(\rho:Y'\to Y\) is an extension over \(C\), then \((\rho^I)_*\lambda_{Y'}=\lambda_Y\).

Countable inverse limits of extensions over \(C\) exist in this category. The selected law at such a limit projects to the selected law at every stage. Pullbacks of the stage selected-law spaces have dense union in \(L^2\) of the limit selected law.

Proof. For each \(a\) the expression before the ultralimit defines a genuine probability measure \(\lambda_{Y,a}\) on \(Y^I\), by integrating the measurable orbit map over the indicated finite time interval. We first show that each of its marginals converges against every bounded measurable test \(G\) to \(m_Y\). The one-slot continuous-time averages \[\frac1L\int_0^L G(S_Y^{it}y)\,\,dt\] converge almost surely as \(L\to\infty\). For example, apply Birkhoff’s theorem for \(S_Y^i\) to the function \(y\mapsto\int_0^1G(S_Y^{is}y)\,\,ds\) and discard the incomplete unit interval. Their almost-sure limit has integral \(\int G\,\,dm_Y\), since every deterministic average has that integral. As \(N_a(\pi_Yy)\ge2^a\), bounded convergence gives the same integral limit at the selected lengths.

Embed \(Y\) as a Borel subset of a compact metrizable space \(K\). Ultralimits of the \(\lambda_{Y,a}\) against continuous functions on \(K^I\) give a positive normalized functional, hence a probability measure on \(K^I\). Its marginals are \(m_Y\), by the preceding paragraph, so it is concentrated on \(Y^I\). Call the resulting law \(\lambda_Y\). To obtain (11) for arbitrary bounded slot functions, normalize their bounds to one and approximate each in \(L^1(m_Y)\) by continuous functions on \(K\) with the same bound. The difference between the two product integrals is bounded by the sum of the single-slot errors. Under \(\lambda_Y\) these errors are the \(L^1(m_Y)\) errors, while under \(\lambda_{Y,a}\) they converge to those errors by the marginal convergence already proved. Letting the errors tend to zero proves the formula. Product tests determine a probability measure on \(Y^I\), giving uniqueness. In particular, all these identities depend only on the almost-everywhere classes of their slot functions.

Applying the product action at a fixed time \(t_0\) shifts the time interval in (11) by \(t_0\) without changing its length or its initial-point selector. For slot functions bounded by one the error is at most \(2\lvert t_0\rvert 2^{-a}\), proving invariance. For an extension \(\rho\), substitute \(G_i\circ\rho\) in the formula. Equivariance, the equality of the selectors, and Fubini identify its right-hand side with that for \(Y\), proving functoriality.

For a tower \(Y_0\leftarrow Y_1\leftarrow\cdots\), take the full countable product of the underlying standard spaces, with the consistent inverse-limit probability measure and coordinatewise flow. This is a jointly Borel action. A cylinder through stage \(r\) has probability computed from \(m_{Y_r}\) by the finitely many bonding maps. At any fixed time their equivariance and preservation of \(m_{Y_r}\) leave that probability unchanged. Cylinders determine the measure, so every fixed time preserves it. The measure is concentrated on compatible sequences; using the full product avoids any need for one compatibility set invariant at all real times. The stage projections are factor maps modulo null sets, and projection through \(Y_0\) gives the factor onto \(C\). Functoriality now gives the asserted selected-law projections. Finally, the stage coordinates generate the inverse-limit sigma-algebra. The compatibility relations hold under the selected law slot by slot because its marginals are the inverse-limit probability measure. Thus the pulled-back stage sigma-algebras on the selected-law space are increasing and generate it modulo null sets. The usual \(L^2\) approximation by an increasing sequence of sigma-algebras proves density. ◻

Distal parameters and small marginals

For a flow \(Y\) over \(C\), let \(\omega_Y:Y\to W_Y\) be its maximal measure-distal factor for time one, and let \(\nu_Y=(\omega_Y)_*m_Y\). Write \(m_{Y,w}\) for its conditional probability measures, so that \[m_Y=\int_{W_Y}m_{Y,w}\,\,d\nu_Y(w).\] Under \(\lambda_Y\) put \[ P_i=\omega_Y(y_i),\qquad P=(P_i)_{i\in I}. \tag{12}\] Proposition 4 ensures that \(W_Y\) is preserved by every flow time. Such statements about factor sigma-algebras are understood modulo null sets; all time integrals below are evaluated on the original jointly Borel flow. If \(\rho:Y'\to Y\) is an extension, the pullback of \(W_Y\) is distal for time one and hence is contained in \(W_{Y'}\). Consequently the old parameter array is a measurable function of the new parameter array on selected-law spaces.

Proposition 9 (Parameter and small-marginal laws). For every flow \(Y\) over \(C\) and bounded measurable \(H_i\) on \(W_Y\), \[ \mathbb E_{\lambda_Y}\prod_{i\in I}H_i(P_i) =\lim_{L\to\infty}\int_Y\frac1L\int_0^L \prod_{i\in I}H_i\bigl(\omega_Y(S_Y^{it}y)\bigr) \,\,dt\,\,dm_Y(y). \tag{13}\] For every nonempty \(J\subseteq I\) with \(\lvert J\rvert\le3\), \[ \mathcal L_{\lambda_Y}\bigl((y_i)_{i\in J}\mid P\bigr) =\bigotimes_{i\in J}m_{Y,P_i}. \tag{14}\]

Proof. We first explain the passage from discrete to continuous time used in both assertions. Write \(t=l+s\), with \(l\ge0\) an integer and \(0\le s<1\). For each fixed \(s\), a continuous-time product becomes the discrete product at slopes \(i\) for the shifted slot functions \(G_i\circ S_Y^{is}\). Whenever the corresponding discrete averages converge almost surely, Fubini and bounded convergence show that their integrals in \(s\) converge almost surely. Starting the discrete average at \(l=0\) rather than \(l=1\), and dropping an incomplete final unit interval, each has an error tending to zero. The resulting continuous-time convergence holds along all real lengths \(L\).

Apply this observation to functions from \(W_Y\). Their shifts still belong to \(W_Y\), and pointwise convergence on the distal factor is part of Proposition 4. Thus the time averages in (13) converge almost surely along all lengths, and therefore at the selected lengths as well. Bounded convergence and (11) prove that equation.

For \(\lvert J\rvert\le3\), Lemma 6 and the same continuous-time argument show that the selected \(J\)-marginal equals its deterministic-length integrated law. The extension \(Y\to W_Y\) is relatively weakly mixing. Apply Lemma 5 to the distinct integer slopes \(i\in J\) and the shifted functions \(G_i\circ S_Y^{is}\), for each fixed \(s\). Conditional expectation commutes with these shifts because \(W_Y\) is preserved. Integration in \(s\), using the uniform norm bound and bounded convergence, gives \[ \mathbb E_{\lambda_Y}\prod_{i\in J}G_i(y_i) =\mathbb E_{\lambda_Y}\prod_{i\in J} \mathbb E_{m_Y}[G_i\mid W_Y](y_i). \tag{15}\] Indeed the difference of the deterministic averages tends to zero in \(L^2(m_Y)\), so also in their integrated values. Applying this identity to \(G_i\) multiplied by arbitrary bounded functions of \(P_i\) yields \[ \mathcal L_{\lambda_Y}\bigl((y_i)_{i\in J}\mid(P_i)_{i\in J}\bigr) =\bigotimes_{i\in J}m_{Y,P_i}. \tag{16}\]

Conditioning on the full parameter array requires one more step; it does not follow from (16) alone. Use the invariant transformation with speed \(i\) in slot \(i\). By the stability under nonzero powers and conditional products in Proposition 4, the \(J\)-marginal system in (16) is relatively weakly mixing over its parameter law. The full \(P\)-system is a joining of the distal systems \(W_Y\) at these integer powers. It is distal, and is therefore relatively distal over the subarray \((P_i)_{i\in J}\). Their joint law under \(\lambda_Y\) is an invariant joining identifying that common subarray. Relative disjointness in Proposition 4 makes this joining conditionally independent over the subarray, which is precisely (14). ◻

The full parameter law is thus unaffected by the selectors, although the full state law need not be. We now turn that state law into a new flow while preserving one chosen slot as the original system.

Lemma 10 (Normalized joining flow). For \(b\in I\), define \[ \mathsf M_b(Y)= \left(Y^I,\lambda_Y, S_{\mathsf M_b(Y)}^t(y_i)_i=(S_Y^{it/b}y_i)_i\right). \tag{17}\] This is a standard probability-preserving flow over \(C\) via \((y_i)_i\mapsto\pi_Y(y_b)\), and projection onto slot \(b\) is an extension onto \(Y\). The array \(P\) belongs to the maximal measure-distal factor of \(\mathsf M_b(Y)\) for time one.

Proof. Invariance follows from Proposition 8, after rescaling the time parameter by \(1/b\). Slot \(b\) has speed one and marginal \(m_Y\), giving the stated factor maps. In slot \(i\) the \(W_Y\)-factor is distal for time \(i/b\), by Proposition 4. The law of \(P\) is an invariant joining of these rational-time distal systems, hence is distal. Maximality of the distal factor gives the last assertion. ◻

Lemma 11 (The retained positive correlation). With the functions and phase test of Lemma 7, \[ \int_{C^I}\Psi\bigl((\upsilon(z_i))_i\bigr) \prod_{i\in J_0}f_i\bigl(\xi(z_i)\bigr)\,\,d\lambda_C>0. \tag{18}\] The phase array in this integral is measurable in the distal parameter array of \(C^I\).

Proof. Continuous functions on the compact product of circles are uniformly approximable by finite sums of products of continuous one-coordinate functions. Apply this to \(\Psi\), use (11) for each resulting product, and then use the strictly positive lower limit in (8). Finally, time one fixes the phase on \(C\), so the phase factor is distal and lies in \(W_C\). ◻

Channels and closure under selected laws

The selected length is determined by the original point in the suspension. It can therefore change the joint law of other coordinates, even when their unconditioned law is a product. We organize those coordinates in three groups. Independence of the original point from any two groups will control the effect of selecting a length on all three groups. This channel construction and its closure argument adapt (OpenAI 2026a, sec. 5); the proof below permits every fixed finite set of slots \(I=\{1,\ldots,n\}\).

Definition 12 (Channel). A channel is a flow \(Y\) over the suspension \(C\), with factor map \(\pi_Y:Y\to C\), together with a finite list of probability-preserving maps \(D_\gamma:Y\to C\), called outputs, and positive rational numbers \(q_\gamma\), such that \[D_\gamma(S_Y^t y)=S_C^{q_\gamma t}D_\gamma(y)\] for each fixed \(t\), modulo null sets. Partition the output indices into three groups, numbered \(1,2,3\); empty groups are allowed. Set \[x_0=\xi\circ\pi_Y,\qquad x_\gamma=\xi\circ D_\gamma, \qquad d_Y=(x_\gamma)_\gamma,\] and let \(d_{Y,H}\) be the sublist from groups in \(H\subseteq\{1,2,3\}\). Let \(V\) collect the phase coordinates of \(\pi_Y\) and all the \(D_\gamma\). Under the measure \(m_Y\) require \[ \begin{split} \mathcal L(d_Y\mid V)&=\bigotimes_\gamma\mu,\\ \mathcal L(x_0,d_{Y,H}\mid V)&= \mu\otimes\bigotimes_{\gamma\text{ in groups }H}\mu, \qquad H\subsetneq\{1,2,3\}. \end{split} \tag{19}\] An empty spatial list has the one-point law.

Thus the spatial outputs have a product law independent of the phases. The original spatial coordinate has a product law with the outputs in any proper collection of groups, but no independence from the complete output list is required. The suspension \(C\), with no outputs, is a channel.

Lemma 13 (Conditioning and extensions). For a channel \(Y\), both identities in (19) hold with conditioning on its maximal measure-distal factor \(W_Y\) in place of \(V\). Every flow extension of \(Y\) over \(C\) is a channel with the pulled back output data, and the same strengthening holds there.

Proof. Choose a positive integer \(h\) such that \(hq_\gamma\) is an integer for each output. The phase vector \(V\) is a factor fixed by time \(h\), and its time-one action is distal, so \(V\) belongs to \(W_Y\). For any spatial tuple appearing in (19), its joint factor with \(V\) has product law over \(V\). At time \(h\) the spatial coordinates evolve by a product of positive integer powers of \(T\). This product is mixing: the assertion follows first for product tests from mixing of each power, and then for all bounded tests by approximation. Consequently this factor is relatively weakly mixing over \(V\) for time \(h\).

On the other hand, \(W_Y\) is distal for time \(h\) and relatively distal over \(V\). Relative disjointness, in the form of (OpenAI 2026a, Proposition 2.2), makes the spatial factor and \(W_Y\) relatively independent over \(V\). Its conditional spatial law given \(W_Y\) is therefore the product law in (19). Pullback to an extension preserves the joint laws of the original data, proving the channel identities there; the same argument gives their stronger conditioning. ◻

We now prove the closure property that will permit repeated extensions. Recall from Lemma 10 that \(\mathsf M_b(Y)\) has state space \(Y^I\), measure \(\lambda_Y\), and flow \((y_i)_i\mapsto(S_Y^{it/b}y_i)_i\).

Proposition 14 (Channel closure). Let \(Y\) be a channel and \(b\in I\). On \(\mathsf M_b(Y)\) take \(\pi_Y(y_b)\) as original factor and take every \(D_\gamma(y_i)\), \(i\in I\), as an output, in the old group of \(\gamma\). These data make \(\mathsf M_b(Y)\) a channel.

Proof. Each new output is probability-preserving and has positive rational speed \(iq_\gamma/b\). It remains to prove the two conditional laws. Test them with products of bounded spatial functions and products of bounded phase functions. Such tests determine the laws. Expanding each spatial function into its mean and centered part, and scaling, reduces the proof to real spatial functions that are either \(1\) or centered and bounded in absolute value by \(1\). We may also bound every phase function by \(1\). Whenever at least one spatial factor is centered, the tested integral must be zero. A test using the new original coordinate uses outputs from only a proper set of groups.

Time intervals and offsets.

Use the selected-law formula (11) and condition on the old phase vector \(V\) under \(m_Y\). Fix an integer \(h>0\) clearing all output-speed denominators and write \(t=hl+s\), where \(l\ge0\) is an integer and \(0\le s<h\). For fixed \(V,s\), all phase weights in the test are independent of \(l\). The spatial function for output \(\gamma\) at slot \(i\) is evaluated at \[ T^{e_\gamma i l+a_{\gamma,i}}x_\gamma, \qquad e_\gamma=hq_\gamma\in\mathbb N. \tag{20}\] The offsets \(a_{\gamma,i}\) depend on \(V,s\) and range over a fixed finite set, by the suspension formula. If the original-factor test is present, its argument is \(T^{hbl+a_0}x_0\), with the same finite-offset property. Equivariance and Fubini justify these formulas in the time integrals.

On an event where the selected length is \(N=c2^k\), \(c\in F\) and \(k\ge a\), put \(m=\lfloor c2^k/h\rfloor\). Replacing the normalized time integral by the average over the \(m\) complete intervals, followed by integration against \(ds/h\), has error at most \(2h2^{-a}\). For all sufficiently large \(a\) we have \(m\ge1\).

Tests using a proper set of groups.

Given \(V\), the relevant old output coordinates have product law and are independent of \(x_0\), by (19). The selector is determined by \(V,x_0\). Suppose first that some output occurrence is centered. For fixed \(V,s,x_0\), integrate the spatial output product at index \(l\). The expectation factors over the old output indices \(\gamma\). For an index having a centered occurrence, the times (20) in distinct slots separate as \(l\to\infty\). Mixing of all orders makes its integral tend to the product of the means, hence to zero. A single centered occurrence already has zero integral. The convergence is uniform over the finitely many offset patterns. All other factors have absolute value at most \(1\). The Cesàro averages of these absolute expectations consequently tend to zero uniformly over all sufficiently large lengths. This applies to the selected lengths as well.

If every output test is \(1\), the centered factor must be the original one. Birkhoff’s theorem for the ergodic power \(T^{hb}\) gives convergence to zero for \(\mu\)-almost every \(x_0\), simultaneously for all the finitely many offsets \(a_0\). Since the selected lengths tend to infinity pointwise, they have the same limit. The conditional law of \(x_0\) given \(V\) is \(\mu\), so bounded convergence completes this case.

Three groups and almost orthogonal vectors.

It remains to treat the full output list, with no original-factor spatial test. If one group has no centered occurrence, every spatial test in that group is \(1\), and the preceding argument applies; the phase tests of that group were allowed throughout. Thus suppose each group contains a centered occurrence. For fixed \(V,s\), write its spatial product at index \(l\) as \[u_l\in L^2(\Omega_1),\qquad v_l\in L^2(\Omega_2),\qquad w_l\in L^2(\Omega_3),\] where \(\Omega_j\) carries the product measure with one copy of \(\mu\) for every old output in group \(j\). These are real unit-ball vectors, indeed functions bounded by \(1\). Given \(V\), the old spatial output law is the product of these three spaces. The vectors depend on \(V,s\) only through the finite offset patterns.

For every \(\eta>0\), after omitting finitely many initial indices, the families \((u_l)\) and \((v_l)\) each have absolute pair inner products at most \(\eta\) outside a graph of bounded maximum degree \(L_0\). All these bounds are uniform in the offset patterns. To verify this, in each of the two groups choose a centered occurrence at some output index \(\gamma\). Its coordinate integral in a pair inner product is small by mixing of all orders whenever the times in the two copies of (20) are sufficiently separated. The times within each copy separate once \(l,l'\) exceed a fixed threshold. Failure of separation between copies then requires one of finitely many inequalities \[\lvert e_\gamma i l-e_\gamma i'l'+a'\rvert\le R.\] For fixed \(l\), each inequality permits only a bounded number of integers \(l'\), because \(e_\gamma,i'>0\); the same holds with \(l,l'\) interchanged. Their union is the required bounded-degree graph. The coordinate integrals for other output indices are bounded by \(1\). The separation threshold and \(L_0\) may depend on \(\eta\), but no rate of mixing is needed.

Conditioning on a selected length.

For fixed selector index \(a\), partition into the events \(N_a=c2^k\), \(c\in F\), \(k\ge a\), choosing one representation when a length has more than one. Given \(V\), write their probabilities as \(p_{k,c}(V)\). On one such event of probability \(p>0\), the conditional spatial law of any two groups, and of each single group, is unchanged. Indeed the event is determined by \(V,x_0\), which is independent of any proper set of groups by (19). The full three-group law may change. Its density \(\rho\) relative to the product law satisfies \[ 0\le p\rho\le1,\qquad \int p\rho=p,\qquad \lVert p\rho\rVert_2\le\sqrt p. \tag{21}\] The reason is that the joint measure of the outputs and the event is dominated by the unconditioned product measure. These statements hold for almost every \(V\); the length events are countable.

We choose the rank bound to meet two requirements: the packing bound must be \(o(m)\), and the reciprocal bounds must be summable over dyadic length scales. For \(m=\lfloor c2^k/h\rfloor\) set \[ d_m=\left\lceil(1+\log m)^{3/2}\right\rceil, \qquad z_l=u_l\otimes v_l\quad(0\le l<m). \tag{22}\] Let \(\Pi_m\) be the orthogonal projection onto up to \(d_m\) leading eigenvectors of the positive covariance operator \(\sum_{l<m}z_lz_l^*\), taking its full range if its rank is smaller. Fix \(0<\varepsilon\le1\) and then fix \(0<\eta\le\varepsilon^2/10^4\) in the preceding graph bound. Theorem 21, proved in Appendix 7 from (OpenAI 2026a, Theorem 4.1), gives \[ \frac1m\sum_{l<m}\lVert\Pi_m z_l\rVert \le\varepsilon+o(1). \tag{23}\] Here the error tends to zero uniformly over the offset patterns. Indeed, after the fixed initial indices are removed, the number of projection norms at least \(\varepsilon\) is at most \[C(1+L_0)\varepsilon^{-2} \exp\!\bigl(C\varepsilon^{-2}\sqrt{d_m}\log(2+d_m)\bigr)=o(m).\] The removed indices contribute \(o(1)\) to the average. Importantly, \(\varepsilon,\eta\), and the resulting exceptional degree are fixed before \(m\) tends to infinity.

For the projected terms, Cauchy–Schwarz under the conditional event law uses its unchanged two-group and one-group marginals. Term by term, including the event probability, it gives \[\left|p\int (\Pi_mz_l)(\omega_1,\omega_2) w_l(\omega_3)\rho(\omega)\,d\omega\right| \le p\lVert\Pi_mz_l\rVert_2\lVert w_l\rVert_2 \le p\lVert\Pi_mz_l\rVert_2.\] Thus their average contributes at most \(p(\varepsilon+o(1))\).

For the residual vectors \(z_l'=(1-\Pi_m)z_l\), let \(A=(\langle z_l',z_{l'}'\rangle)_{l,l'<m}\) and \(B=(\langle w_l,w_{l'}\rangle)_{l,l'<m}\). The covariance trace is at most \(m\), so its remaining eigenvalues, and hence the nonzero eigenvalues of \(A\), are at most \(m/d_m\). If its rank was at most \(d_m\), the residual is zero. Also \(B\) is positive semidefinite with trace at most \(m\). Consequently \[\left\|\frac1m\sum_{l<m}z_l'\otimes w_l\right\|^2 =\frac{\operatorname{tr}(AB)}{m^2}\le\frac1{d_m}.\] By (21), the residual contribution, including the event probability, is at most \(\sqrt{p/d_m}\). Multiplication by the fixed phase weight does not increase either bound.

Summing over lengths.

Write \(m(k,c)=\lfloor c2^k/h\rfloor\). For fixed \(h\) this is comparable to \(2^k\) for all sufficiently large \(k\), uniformly in the finite set \(F\). Since \(\sum_{k,c}p_{k,c}(V)=1\), Cauchy–Schwarz yields \[ \sum_{k\ge a,\ c\in F}\sqrt{\frac{p_{k,c}(V)}{d_{m(k,c)}}} \le \left(\sum_{k\ge a,\ c\in F}\frac1{d_{m(k,c)}}\right)^{1/2} \longrightarrow0. \tag{24}\] The series converges because its summands are \(O(k^{-3/2})\). The projected contributions have total limit superior at most \(\varepsilon\), by (23) and the same probability identity. These estimates are uniform in \(V,s\) and may be integrated. There is no measurable-choice issue for the projections: choose one for each of the countably many sizes and finitely many offset patterns. Letting \(a\) tend to infinity in the selected-law formula and then letting \(\varepsilon\) tend to zero proves the required vanishing. Both laws in (19) follow. ◻

Corollary 15. For every channel \(Y\), the copied spatial list \((d_Y(y_i))_{i\in I}\) under \(\lambda_Y\) has its full product law conditional on \(P=(\omega_Y(y_i))_{i\in I}\).

Proof. Apply Proposition 14 and Lemma 13 to \(\mathsf M_b(Y)\). By Lemma 10, the old parameter array \(P\) belongs to its maximal distal factor. Conditioning the resulting constant product law down to \(P\) proves the assertion. ◻

Minimal correlations and saturated extensions

The positive correlation in (18) records the assumed pointwise failure. We continue the channel-and-saturation argument of (OpenAI 2026a, secs. 5–6), with an additional minimum-support choice. We choose such a correlation with the fewest possible spatial slots. This choice supplies the product laws for proper subcollections that a pointwise theorem of greater length would otherwise have to provide. Among correlations of this minimum size, we maximize the number of slots that use only outputs. The final section will increase that number and obtain a contradiction.

Spatial roles and minimum support

Let \(Y\) be any channel. At a slot \(i\in I\), we allow either of the following spatial lists: \[\begin{align*} L_i(y_i)&=d_Y(y_i) &&\text{(an output role)},\\ L_i(y_i)&=(x_0(y_i),d_{Y,H_i}(y_i)),\qquad H_i\subsetneq\{1,2,3\} &&\text{(an input role)}. \end{align*}\] An input role may therefore read the original coordinate jointly with outputs from any proper subset of the three groups. This allowance is important: later, regrouping outputs will produce precisely these joint lists. A test may ignore some entries of its designated list.

For a nonempty set \(J\subseteq I\), designate one role at each \(i\in J\). Choose bounded real functions \(h_i\) of the corresponding lists, centered for their product spatial laws. By Lemma 13, \(h_i(L_i)\) also has conditional mean zero over \(W_Y\). We consider correlations \[ \mathbb E_{\lambda_Y}\left[\Phi(P) \prod_{i\in J}h_i(L_i(y_i))\right], \tag{25}\] where \(\Phi\) is any bounded real function of the full distal parameter array \(P=(\omega_Y(y_i))_{i\in I}\). The channel \(C\), with no outputs, has a nonzero correlation of this form by Lemma 11.

Let \(k\) be the minimum of \(|J|\) among all nonzero correlations (25), over all channels and all such data. The small-slot law (14) gives \(k\geq4\): when \(|J|\leq3\), the conditional expectation of the product given \(P\) is the product of the conditional means, all zero. Thus \(n<4\) already contradicts the assumed failure. In what follows \(4\leq k\leq n\).

Lemma 16 (Product laws below the minimum support). On any channel, choose fewer than \(k\) distinct slots and designate one allowed spatial list at each. Under its selected law, these lists, jointly, have their product spatial law conditional on the full distal parameter array.

Proof. Test the conditional law by a product of bounded real functions, one for each designated list, and by a bounded real function of the full parameter array. Write each spatial test as its product-law mean plus a centered function. Every nonconstant term in the expansion is a correlation of the form (25) with fewer than \(k\) slots, and hence vanishes. Only the product of the means remains. Product tests determine the asserted conditional law on these finite products of standard spaces. ◻

Among the nonzero correlations with support size \(k\), let \(r\) be the largest number of output roles. Corollary 15 implies \(r<k\), because the full copied output collection has product spatial law given \(P\). Choose a channel and a correlation attaining these two extremal values. We next pass to an extension on which conditioning on the parameters, or on the parameters and one slot, cannot be improved by any further extension.

Saturating the parameter and single-slot data

For any flow \(Y\) over \(C\), define sub-\(\sigma\)-algebras of the selected law space \((Y^I,\lambda_Y)\) by \[ \mathcal D_0(Y)=\sigma(P),\qquad \mathcal D_i(Y)=\sigma(P,y_i)\quad(i\in I). \tag{26}\] If \(\rho:Y'\to Y\) is a flow extension over \(C\), functoriality gives \((\rho^I)_*\lambda_{Y'}=\lambda_Y\). The pullback of \(W_Y\) is a distal factor of \(Y'\), so is contained in \(W_{Y'}\). Consequently the pullback of each \(\mathcal D_l(Y)\) is contained in \(\mathcal D_l(Y')\).

Lemma 17 (Simultaneous saturation). Every standard flow \(Y\) over \(C\) has a standard flow extension \(Y_*\to Y\) over \(C\) with the following property. For every further standard flow extension \(\rho:Y'\to Y_*\) over \(C\), every \(f\in L^2(\lambda_{Y_*})\), and every \(l\in\{0\}\cup I\), \[ \mathbb E_{\lambda_{Y'}}[f\circ\rho^I\mid\mathcal D_l(Y')] =\mathbb E_{\lambda_{Y_*}}[f\mid\mathcal D_l(Y_*)]\circ\rho^I. \tag{27}\]

Proof. We use a countable bounded-energy construction of the kind underlying Austin’s sated extensions (Austin 2015, Theorem 3.11); the argument below proves the assertion for the particular data (26). In a tower of flow extensions, for a fixed test \(f\) introduced at one stage, the quantity \[\left\|\mathbb E[f\mid\mathcal D_l]\right\|_2^2\] at later stages is nondecreasing, by inclusion of the pulled-back data, and is bounded by \(\|f\|_2^2\). At each stage introduce a countable dense set in the \(L^2\) space of its selected law, retaining all previously introduced tests. Schedule all pairs consisting of a retained test and a datum index so that each pair is treated infinitely often, with positive tolerances tending to zero. A diagonal enumeration gives such a schedule even though new tests are introduced at every stage.

At a scheduled step, extend the current flow so that the corresponding energy is within the prescribed tolerance of its supremum over all standard flow extensions of that stage over \(C\). The supremum is finite. It can be taken over a set: standard Borel spaces, their probability laws, Borel actions, and factor maps admit codes on fixed standard spaces. The identity extension is among the competitors. Let \(Y_*\) be the countable inverse limit of the resulting tower, which exists in our category by Proposition 8.

Fix a retained test, an index \(l\), and a further extension \(Y'\to Y_*\). At every scheduled optimization of this pair, \(Y'\) is also an extension of the stage just before the optimization. Its energy is therefore at most the energy achieved immediately after that optimization plus its tolerance. The latter achieved energy is at most the energy at \(Y_*\). As the tolerances tend to zero, the energy at \(Y'\) is at most that at \(Y_*\). The reverse inequality follows from monotonicity, so the two energies are equal.

Identify \(L^2(\lambda_{Y_*})\) with its pullback in \(L^2(\lambda_{Y'})\). The two conditional expectations are orthogonal projections onto nested closed subspaces. Equality of their squared norms forces equality of the projections, proving (27) for each retained test. Stage functions have dense span in \(L^2(\lambda_{Y_*})\) by Proposition 8; the retained dense sets and the contraction property of conditional expectation extend the identity to every \(f\) in that space. ◻

Apply Lemma 17 to the chosen channel. Pulling back its data preserves the channel laws by Lemma 13, and preserves its nonzero correlation by functoriality. Its old parameter test is a function of the enlarged distal parameters. We henceforth call this saturated channel \(Y\). Write \(J\) for its support, \(A\subset J\) for its output roles, and choose \[b\in J\setminus A, \qquad |J|=k,\quad |A|=r.\] Thus the role at \(b\) is an input role. The purpose of the remaining construction is to turn it into an output role while keeping a nonzero correlation on the same \(k\) slots.

Regrouping the other slots as outputs

Set \(Z=\mathsf M_b(Y)\). Its state is \((y_j)_{j\in I}\), its law is \(m_Z=\lambda_Y\), and projection to \(y_b\) is a flow factor \(Z\to Y\). We give \(Z\) a channel structure different from the full-copy structure of Proposition 14. Retain all old outputs from slot \(b\) with their original group numbers. For each \(j\in J\setminus\{b\}\), also include the \(C\)-valued factors whose spatial coordinates form \[Q_j=L_j(y_j).\] These are old output factors and, when the role at \(j\) is an input role, the original factor \(\pi_Y(y_j)\). Choose a surjection \[\tau:J\setminus\{b\}\longrightarrow\{1,2,3\},\] possible because \(k\geq4\), and put the entire list \(Q_j\) in new group \(\tau(j)\). Every new output has the correct marginal and a positive rational speed: the old speed at slot \(j\) is multiplied by \(j/b\).

Lemma 18 (The regrouped channel). The preceding output lists and groups make \(Z\) a channel over its original factor \(\pi_Y(y_b)\).

Proof. We prove both channel laws conditional on the old full parameter array \(P\) on \((Y^I,\lambda_Y)\). Every phase in the new channel is a function of \(P\), so conditioning down then proves (19) for \(Z\).

For the full output collection, its spatial lists are \[d_Y(y_b),\qquad Q_j\quad(j\in J\setminus\{b\}).\] These are allowed lists at \(k\) distinct old slots. Expand a product of tests on these lists into constant and centered parts. Every term with fewer than \(k\) centered factors has the product-law value by Lemma 16. The all-centered term also vanishes against every bounded parameter test. Otherwise it would give a nonzero correlation with support \(k\) and \(r+1\) output roles: slot \(b\) now has an output role, and the other roles are unchanged. This contradicts the definition of \(r\). If the output list at \(b\) is empty, its centered test is identically zero, with the same conclusion. Thus the full output collection has product spatial law given \(P\).

For the second channel law, choose a proper subset \(H\subsetneq\{1,2,3\}\) of the new groups and include \(x_0(y_b)\). At slot \(b\) the resulting list is \((x_0(y_b),d_{Y,H}(y_b))\), an allowed input list. At any other retained slot \(j\) it is exactly \(Q_j\), and such a slot occurs only when \(\tau(j)\in H\). Surjectivity of \(\tau\) omits at least one of the \(k-1\) other slots. Fewer than \(k\) old slots therefore occur, so Lemma 16 gives the required product law conditional on \(P\). This proves the second channel law as well. ◻

The square argument

We now use the saturated channel \(Y\), its nonzero correlation (25), and the regrouped channel \(Z=\mathsf M_b(Y)\) from Section 5. The selected law \(\lambda_Z\) is a law on an ordered square of \(Y\)-states. Its \(b\)th row and \(b\)th column both have law \(\lambda_Y\), but for different reasons. Saturation will relate their correlations without requiring any symmetry of the square or any pointwise theorem for many slots on an arbitrary extension.

A common corner and two conditional expectations

On \((Y^I,\lambda_Y)\), put \[ q((y_i)_{i\in I})= \prod_{j\in J\setminus\{b\}}h_j(L_j(y_j)), \qquad g(P,y_b)=\mathbb E_{\lambda_Y}[q\mid P,y_b]. \tag{28}\] The function \(g\) is not zero in \(L^2(\lambda_Y)\): otherwise the nonzero correlation would vanish after conditioning on \((P,y_b)\). On the same state space regarded as \(Z\), the function \(q\) reads only the newly added outputs. Write \(e=\omega_Z(z)\) for the new distal parameter and define \[ u(e,y_b)=\mathbb E_{m_Z}[q\mid e,y_b]. \tag{29}\] All these functions have bounded real Borel versions.

By Lemma 10, the old array \(P\) is a factor of the new distal parameter. Write \[R(e)=P, \qquad \kappa(e)=R(e)_b.\] Thus \(\kappa\circ\omega_Z\) is the old distal parameter of the slot-\(b\) factor \(Z\to Y\).

A point of the selected law space \((Z^I,\lambda_Z)\) has the form \[\mathbf z_i=(y_{ij})_{j\in I},\qquad E_i=\omega_Z(\mathbf z_i),\qquad E=(E_i)_{i\in I}.\] The first index specifies the row. Define the column states and their parameters by \[ y_i=y_{ib},\qquad p=(\kappa(E_i))_{i\in I},\qquad w=\kappa(E_b). \tag{30}\] In particular \(y_b=y_{bb}\) is the common corner. The column \((y_i)_{i\in I}\) has law \(\lambda_Y\), by functoriality for \(Z\to Y\). The row \(\mathbf z_b\) also has law \(\lambda_Y\), because the one-slot marginal of \(\lambda_Z\) is \(m_Z=\lambda_Y\). Figure 1 records these distinct identifications.

\[\begin{array}{c|ccccc|l} &1&\cdots&b&\cdots&n&\text{row parameter}\\ \hline 1&y_{11}&\cdots&\colorbox{blue!10}{$y_{1b}$}&\cdots&y_{1n}&E_1\\ \vdots&\vdots&&\colorbox{blue!10}{$\vdots$}&&\vdots&\vdots\\ b&\colorbox{green!12}{$y_{b1}$}& \colorbox{green!12}{$\cdots$}& \fcolorbox{black}{yellow!25}{$y_{bb}$}& \colorbox{green!12}{$\cdots$}& \colorbox{green!12}{$y_{bn}$}&E_b\\ \vdots&\vdots&&\colorbox{blue!10}{$\vdots$}&&\vdots&\vdots\\ n&y_{n1}&\cdots&\colorbox{blue!10}{$y_{nb}$}&\cdots&y_{nn}&E_n \end{array}\] \[\underbrace{\mathbf z_b=(y_{bj})_{j\in I}}_{\text{row }b: \text{ old parameters }R(E_b)} \qquad \underbrace{(y_{ib})_{i\in I}}_{\text{column }b: \text{ old parameters }p} \qquad w=p_b=R(E_b)_b.\]

The ordered square under \(\lambda_Z\). The highlighted row and column meet at \(y_{bb}\), whose old distal parameter is \(w\). Both have law \(\lambda_Y\): the row by the single-slot marginal, the column by the factor map \(Z\to Y\). No invariance under transposition is asserted.

Lemma 19 (The row–column identity). Under \(\lambda_Z\), \[ \mathbb E\big[q((y_i)_i)q(\mathbf z_b)\mid E\big] =\int_Y g(p,y)u(E_b,y)\,\,dm_{Y,w}(y). \tag{31}\] Moreover, in a single \(Z\)-state with law \(m_Z\), the conditional law of its slot-\(b\) state given \(e\) is \(m_{Y,\kappa(e)}\).

Proof. Apply saturation (27) with datum index \(0\) to the extension \(Z\to Y\). It says that the conditional law of the entire column given \(E\) is the law of \((Y^I,\lambda_Y)\) given \(p\). This assertion follows first for \(L^2\) tests, and then as an identity of conditional kernels by a countable determining class. By the singleton case of (14), the conditional law of \(y_b\) given \(E\) is consequently \(m_{Y,w}\). Conditioning down to \(E_b\) gives the same law for the slot-\(b\) state of row \(b\) given \(E_b\). Since row \(b\) has marginal \(m_Z\), this proves the last assertion of the lemma.

Saturation with datum index \(b\) gives the stronger identity \[ \mathbb E_{\lambda_Z}[q((y_i)_i)\mid E,\mathbf z_b]=g(p,y_b). \tag{32}\] Indeed the old datum is \((P,y_b)\) and the new datum is \((E,\mathbf z_b)\). Multiply (32) by \(q(\mathbf z_b)\) and condition on \(E\). The singleton case of (14), now for \(Z\), states that the law of \(\mathbf z_b\) given \(E\) is the \(m_Z\)-fibre law over \(E_b\). Within that fibre, (29) conditions \(q(\mathbf z_b)\) on its slot-\(b\) state, whose law is \(m_{Y,w}\) by the first part of the proof. This gives (31). ◻

Independence of the two parameter descriptions

The integral in (31) pairs a column function with a row function. Its nonvanishing will follow from the fact that, apart from the parameter at their common corner, the two parameter descriptions are independent.

Lemma 20 (Parameter independence). Under \(\lambda_Z\), the random variables \(p\) and \(E_b\) are conditionally independent given \(w\).

Proof. Let \(F_0\) be a bounded real test on \(W_Z\). In the single-state parameter law \((\omega_Z)_*m_Z\), let \(\overline F_0\) satisfy \[\overline F_0(\kappa(e))=\mathbb E[F_0(e)\mid\kappa(e)].\] We prove that \(F_0(E_b)\) can be replaced by \(\overline F_0(w)\) when tested against any product \(\prod_i H_i(p_i)\) of bounded real coordinate tests. All the variables are functions of \(E\), so the deterministic parameter-law identity (13) on \(Z\) computes their joint integral. At each fixed averaging time \(t\), its state integral is \[\int_Z F_0(\omega_Z(S_Z^{bt}z)) \prod_{i\in I}H_i\big(\kappa(\omega_Z(S_Z^{it}z))\big) \,\,dm_Z(z).\] Make the measure-preserving change of variable \(z'=S_Z^{bt}z\). The product of the \(H_i\) becomes \[\prod_{i\in I}H_i\big( \kappa(\omega_Z(S_Z^{(i-b)t}z'))\big).\] For each fixed \(t\), this is measurable with respect to \(\kappa(\omega_Z(z'))\). To see this, the slot-\(b\) map \(Z\to Y\) is a flow factor, and the old factor \(W_Y\) is preserved by every flow time. Hence the old parameter at time \((i-b)t\) is a measurable function of the old parameter at time zero.

Conditional expectation in this single-state integral therefore replaces the first factor by \[\overline F_0\bigl(\kappa(\omega_Z(z'))\bigr).\] Reversing the change of variable and taking the deterministic averaging limit gives \[\mathbb E_{\lambda_Z}\left[F_0(E_b)\prod_i H_i(p_i)\right] =\mathbb E_{\lambda_Z}\left[\overline F_0(w)\prod_i H_i(p_i)\right].\] The fixed-time identities suffice under the time integral by measurability and Fubini; no simultaneous pointwise model for the factor action is needed. The \(E_b\)-marginal is the single-state parameter law, so \(\overline F_0(w)=\mathbb E[F_0(E_b)\mid w]\). Products of the \(H_i\) determine the \(p\)-law, and \(w=p_b\) is itself a coordinate of \(p\). The displayed identity is exactly conditional independence given \(w\). ◻

Why the pairing cannot vanish

Let \(\alpha_w\) be the conditional law of the column parameter array \(p\) over its \(b\)th parameter \(w\). Let \(\beta_w\) be the conditional law of \(E_b\) over \(w\). These are regular conditional laws on standard spaces. Lemma 20 gives \[ \mathcal L(p,E_b\mid w)=\alpha_w\otimes\beta_w, \qquad R_*\beta_w=\alpha_w. \tag{33}\] The second identity uses the row marginal \(m_Z=\lambda_Y\): its old full parameter array \(R(e)\) has the same law as the column parameter array, with the same distinguished \(b\)th coordinate.

For almost every \(w\), work in the separable real Hilbert space \[\mathcal H_w=L^2(m_{Y,w};\mathbb R).\] The bounded functions \(g(p,\cdot)\) and \(u(e,\cdot)\) define square-integrable random vectors in this space, with laws of their parameters respectively \(\alpha_w\) and \(\beta_w\). There is no measurability ambiguity in this fibrewise assertion. Bounded Borel versions of \(g\) and \(u\) are fixed; disintegration and Fubini yield their sections for almost every \(w\). A countable generating algebra of the standard Borel space \(Y\) has rational simple span dense in every \(L^2(m_{Y,w})\). Tests against that span give weak measurability of the sections, and separability gives strong measurability.

The conditional mean of the second random vector recovers the first: \[ \mathbb E_{\beta_w}[u(e,\cdot)\mid R(e)] =g(R(e),\cdot) \quad\text{in }\mathcal H_w. \tag{34}\] Here is the conditional-law justification. In a single \(Z\)-state, write \(y\) for its slot-\(b\) state. Lemma 19 gives, conditional on \(w\), \[\mathcal L(e,y\mid w)=\beta_w(\,de)m_{Y,w}(\,dy).\] Conditioning (29) down on \((R(e),y)\) gives \(g(R(e),y)\) by (28), because \(m_Z=\lambda_Y\). Under the preceding product law, conditioning \(e\) on \((R(e),y)\) is the same as conditioning it on \(R(e)\). This proves (34). One can impose the scalar conditional identities first for countably many generating tests and then use Fubini and density to obtain the stated Hilbert-space identity for almost every \(w\).

Suppose that the right side of (31) vanished almost surely. By (33), for almost every \(w\) this would give \[\langle g(p,\cdot),u(e,\cdot)\rangle_{\mathcal H_w}=0 \quad(\alpha_w\otimes\beta_w)\text{-almost surely}.\] Condition on \(R(e)=p'\) and use (34). We obtain \[ \langle g(p,\cdot),g(p',\cdot)\rangle_{\mathcal H_w}=0 \quad(\alpha_w\otimes\alpha_w)\text{-almost surely}. \tag{35}\] A square-integrable random vector in a separable real Hilbert space cannot be orthogonal almost surely to an independent copy unless it vanishes almost surely. Indeed, otherwise its norm is at least some \(\epsilon>0\) with positive probability. By a countable cover with balls of radius \(\epsilon/4\), this event has positive probability in one such ball. Two independent copies then have positive probability of both lying in that ball and having norm at least \(\epsilon\). Their distance is less than \(\epsilon/2\), so \[\langle v,v'\rangle =\tfrac12(\|v\|^2+\|v'\|^2-\|v-v'\|^2)>0,\] a contradiction. Apply this observation to (35). It forces \(g(p,\cdot)=0\) for \(\alpha_w\)-almost every \(p\), for almost every \(w\). The singleton conditional law for the column then implies \(g(P,y_b)=0\) in \(L^2(\lambda_Y)\), contrary to (28). Thus (31) does not vanish almost surely.

Increasing the number of output roles

Choose a bounded real function \(\Theta(E)\) that detects the nonzero conditional expectation in (31); its sign is one such choice. Then \[ \mathbb E_{\lambda_Z}\left[ \Theta(E)q((y_i)_i)q(\mathbf z_b)\right]\ne0. \tag{36}\] We interpret this as a correlation on the regrouped channel \(Z\). For \(i\in J\setminus\{b\}\), the factor \(h_i(L_i(y_{ib}))\) retains its old role. An old output role uses only the retained outputs from slot \(b\). An old input role uses the new original coordinate and retained outputs from its old proper subset \(H_i\) of group numbers. It may ignore the additional outputs in those groups. Each such test is centered for its new product spatial law, since its marginal on the coordinates it actually uses is the old product law.

At row \(b\), the factor \(q(\mathbf z_b)\) reads only the added outputs with spatial lists \(Q_j\), \(j\in J\setminus\{b\}\), and thus has an output role. It is centered: under the product spatial law of all new outputs these lists are independent product lists, so the mean of this product of centered tests is zero. Hence (36) is a nonzero correlation of support size \(k\) with \(r+1\) output roles. This contradicts the maximal choice of \(r\).

Completion of the proof of Theorem 1. If the required pointwise limit failed for a fixed tuple and a fixed \(n\geq2\), the reduction and selectors of Section 3 would give the positive correlation (18). The minimum-support argument rules this out immediately when \(n<4\), and the saturated, regrouped square argument rules it out when \(n\geq4\). Therefore the averages converge almost everywhere to the product of the means, along all positive integer lengths. Finally, the countable-coordinate coding reduction pulls this conclusion back to the original, possibly nonstandard, probability space. The tuple and \(n\) were fixed throughout, as allowed in the statement of the theorem. ◻

Tensor packing

The closure proof uses a deterministic fact: a subspace of moderately growing dimension captures a substantial part of only a small fraction of an almost orthogonal tensor family. We reproduce the argument of (OpenAI 2026a, Theorem 4.1), with the same quantitative estimate. Throughout this appendix, Hilbert spaces are real, tensor products are Hilbert tensor products, and logarithms are natural.

Theorem 21 (Tensor packing). There is a universal constant \(C\) with the following property. Let \(0<\varepsilon\le1\), and let \(d\ge1\) and \(L\ge0\) be integers. Let \(x_1,\ldots,x_m\) and \(y_1,\ldots,y_m\) lie in the unit balls of real Hilbert spaces \(H_1,H_2\). For each family suppose that a graph on \(\{1,\ldots,m\}\) of maximum degree at most \(L\) contains every off-diagonal pair whose absolute inner product exceeds \(\eta\), where \[0\le\eta\le\varepsilon^2/10^4.\] The two exceptional graphs may differ. For every orthogonal projection \(P\) on \(H_1\otimes H_2\) of rank at most \(d\), \[ \#\{j:\lVert P(x_j\otimes y_j)\rVert\ge\varepsilon\} \le C(1+L)\varepsilon^{-2} \exp\!\left(C\varepsilon^{-2}\sqrt d\log(2+d)\right). \tag{37}\] For \(d=d_m\) as defined in (22), this is \(o_{\varepsilon,L}(m)\), uniformly over the Hilbert spaces, vectors, projections, and allowed values of \(\eta\).

The proof separates directions on which the testing space has large variance and applies Gaussian comparison to the remaining directions. The comparison is the finite Sudakov–Fernique inequality; see (Fernique 1974, sec. 1, Equation (1.2)) and (Sudakov 1971). The following matrix estimate uses the Gaussian integration-by-parts moment argument behind matrix Khintchine inequalities; compare (Tropp 2018, Lemma 6.1 and Proposition 7.3). We include its proof to retain dependence on the covariance trace, rather than on the ambient dimensions.

Lemma 22 (Gaussian matrix estimate). Let \(B_1,\ldots,B_r\) be matrices between two finite-dimensional real Hilbert spaces, and put \[H_j=\begin{pmatrix}0&B_j\\B_j^*&0\end{pmatrix}, \qquad S=\sum_{j=1}^rH_j^2.\] If \(\lVert S\rVert_{\rm op}\le\sqrt d\) and \(\mathop{\mathrm{tr}}S\le2d\) for some \(d\ge1\), then for independent standard normal variables \(g_j\), \[ \mathbb E\left\|\sum_{j=1}^rg_jB_j\right\|_{\rm op} \le C d^{1/4}\sqrt{\log(2+d)}. \tag{38}\] The constant is independent of the dimensions of the two spaces.

Proof. Write \(H=\sum_jg_jH_j\). Gaussian integration by parts, followed by differentiation of the matrix power, gives for every integer \(q\ge2\) \[\mathbb E\mathop{\mathrm{tr}}H^{2q} =\sum_{j=1}^r\sum_{a=0}^{2q-2} \mathbb E\mathop{\mathrm{tr}}(H_jH^aH_jH^{2q-2-a}).\] For a fixed self-adjoint \(H\) with eigenvalues \(\lambda_u\) and a self-adjoint matrix \(K\), the absolute value of such a trace is at most \[\sum_{u,v}|K_{uv}|^2 |\lambda_v|^a|\lambda_u|^{2q-2-a} \le\mathop{\mathrm{tr}}(K^2|H|^{2q-2}).\] The inequality follows from weighted arithmetic–geometric mean and the symmetry \(|K_{uv}|=|K_{vu}|\). Summing over \(j\) gives \[ \mathbb E\mathop{\mathrm{tr}}H^{2q} \le(2q-1)\lVert S\rVert_{\rm op}\mathbb E\mathop{\mathrm{tr}}|H|^{2q-2}. \tag{39}\] Since \(\mathbb E\mathop{\mathrm{tr}}H^2=\mathop{\mathrm{tr}}S\), iteration yields \[\mathbb E\mathop{\mathrm{tr}}H^{2q} \le(2q-1)!!\,\lVert S\rVert_{\rm op}^{q-1}\mathop{\mathrm{tr}}S.\] The operator norm of \(H\) equals that of \(\sum_jg_jB_j\). Taking a \(2q\)-th root and using the hypotheses bounds its expectation by \[C\sqrt q\,d^{1/4}(2\sqrt d)^{1/(2q)}.\] Choose an integer \(q\ge2\) comparable to \(\log(2+d)\) to obtain (38). ◻

Proof of Theorem 21. We first record an elementary consequence of the exceptional-graph hypothesis. If unit-ball vectors \(v_j\) have absolute inner products at most \(\eta\) outside a graph of degree \(L\), then for every nonempty set \(A\) of indices, \[ \left\|\frac1{|A|}\sum_{j\in A}v_j\right\|^2 \le\frac{1+L}{|A|}+\eta. \tag{40}\] Indeed the expansion contains at most \((1+L)|A|\) ordered diagonal or exceptional pairs. A lower bound on the norm of an average therefore bounds the number of vectors in that average.

It suffices to work in the finite spans of the \(x_j\) and \(y_j\). To see this, let \(Q\) project onto their tensor product and replace the testing subspace \(\operatorname{Ran}P\) by \(Q(\operatorname{Ran}P)\). For any vector \(z\) in that tensor product, the norm of its projection onto the new subspace is at least \(\lVert Pz\rVert\): in the supremum defining \(\lVert Pz\rVert\), replace each unit testing vector by its \(Q\)-image, whose norm is at most one and whose pairing with \(z\) is unchanged. The dimension does not increase.

Identify tensors with matrices from \(H_2\) to \(H_1\). Choose a Hilbert–Schmidt orthonormal basis \(A_1,\ldots,A_r\) of the testing space, where \(r\le d\). The matrices \[R_1=\sum_jA_jA_j^*,\qquad R_2=\sum_jA_j^*A_j\] have trace \(r\). Let \(E_1,E_2\) be their spectral subspaces for eigenvalues strictly greater than \(\sqrt d\); each has dimension at most \(\sqrt d\).

Discard indices for which the projection of \(x_j\) onto \(E_1\) has norm greater than \(\varepsilon/10\), and do the same for \(y_j,E_2\). Cover either exceptional subspace’s unit ball by at most \((1+80/\varepsilon)^{\sqrt d}\) balls of radius \(\varepsilon/40\). If one ball contains \(s\) of the discarded projections, their average has norm at least \(\varepsilon/20\). Applying (40) to the original vectors, with \(\eta\le\varepsilon^2/10^4\), gives \(s\le C(1+L)\varepsilon^{-2}\). The number of discarded indices is therefore at most \[ C(1+L)\varepsilon^{-2} \exp\!\bigl(\sqrt d\log(1+80/\varepsilon)\bigr). \tag{41}\]

Let \(Q_1,Q_2\) project onto \(E_1^\perp,E_2^\perp\), and put \(B_j=Q_1A_jQ_2\). The coefficient map \(z\mapsto(\langle A_j,z\rangle_{\rm HS})_{j=1}^r\) is a contraction. For a remaining index, \[\lVert x_j\otimes y_j-Q_1x_j\otimes Q_2y_j\rVert \le\lVert(1-Q_1)x_j\rVert+\lVert(1-Q_2)y_j\rVert \le\varepsilon/5.\] Thus each remaining index counted in (37) has coefficient vector \[v_j=(\langle x_j,B_ay_j\rangle)_{a=1}^r, \qquad\lVert v_j\rVert\ge\varepsilon/2.\] The map from \(x_j\otimes y_j\) to \(v_j\) remains contractive. Outside the union of the two exceptional graphs, the tensor inner products have absolute value at most \(\eta^2\), and that union has degree at most \(2L\). A collection of these coefficient vectors in a ball of radius \(\varepsilon/8\) centered at one of them has average norm at least \(3\varepsilon/8\). Applying (40) to its preimage tensors bounds its size by \(C(1+L)\varepsilon^{-2}\). A maximal \(\varepsilon/8\)-separated subset therefore has cardinality \(l\) such that the number of indices still to be counted is at most \(C(1+L)\varepsilon^{-2}l\).

We now bound this separated set. For \(G=\sum_ag_aB_a\), the centered Gaussian variables \(Z_j=\langle x_j,Gy_j\rangle\) satisfy \[\mathbb E|Z_j-Z_{j'}|^2=\lVert v_j-v_{j'}\rVert^2.\] Gaussian comparison gives, when \(l\ge2\), \[ c\varepsilon\sqrt{\log l} \le\mathbb E\max_j Z_j\le\mathbb E\lVert G\rVert_{\rm op}. \tag{42}\] For completeness, the comparison of finite centered Gaussian families follows by applying Gaussian integration by parts to \(\beta^{-1}\log\sum_j e^{\beta z_j}\) along their independent Gaussian interpolation. Its derivative is \[\frac\beta4\mathbb E\sum_{j,j'}p_jp_{j'} \left(\mathbb E|Z_j-Z_{j'}|^2-\mathbb E|Z'_j-Z'_{j'}|^2\right), \qquad p_j=\frac{e^{\beta z_j}}{\sum_a e^{\beta z_a}}.\] It is nonnegative when the first increment variances dominate. Letting \(\beta\to\infty\) compares the expected maxima. For (42), take the \(Z'_j\) to be independent normals of variance \(\varepsilon^2/128\) and use \(\mathbb E\max_{j\le l}g_j\ge c\sqrt{\log l}\).

Compression gives \[\sum_jB_jB_j^*\le Q_1R_1Q_1,\qquad \sum_jB_j^*B_j\le Q_2R_2Q_2.\] These operators have norms at most \(\sqrt d\) and traces at most \(d\). Lemma 22 therefore applies. Together with (42) it yields \[\log l\le C\varepsilon^{-2}\sqrt d\log(2+d).\] The resulting bound holds for \(l\le1\) as well. Combine it with (41) and absorb \(\log(1+80/\varepsilon)\le C\varepsilon^{-2}\) to prove (37).

For \(d=d_m\), the exponent is \(O_\varepsilon((\log m)^{3/4}\log\log m)=o(\log m)\), which proves the asserted \(o(m)\) bound. ◻

Since projection norms are at most one, the estimate also gives \[ \frac1m\sum_{j=1}^m\lVert P(x_j\otimes y_j)\rVert \le\varepsilon+ \frac{C(1+L)}{m\varepsilon^2} \exp\!\left(C\varepsilon^{-2}\sqrt{d_m}\log(2+d_m)\right) \tag{43}\] when \(\operatorname{rank}P\le d_m\). Here \(\varepsilon\) and then \(\eta\) are fixed before \(m\) increases; \(L\) may depend on that chosen \(\eta\). This is exactly the order of quantifiers supplied by mixing in Proposition 14.

Assani, I. 1998. “Multiple Recurrence and Almost Sure Convergence for Weakly Mixing Dynamical Systems.” Israel J. Math. 103: 111–24. https://doi.org/10.1007/BF02762270.
Austin, T. 2015. “Pleasant Extensions Retaining Algebraic Structure, I.” J. Anal. Math. 125: 1–36. https://doi.org/10.1007/s11854-015-0001-9.
Bourgain, J. 1990. “Double Recurrence and Almost Sure Convergence.” J. Reine Angew. Math. 404: 140–61.
Derrien, J.-M., and E. Lesigne. 1996. “Un Théorème Ergodique Polynômial Ponctuel Pour Les Endomorphismes Exacts Et Les \(K\)-Systèmes.” Ann. Inst. H. Poincaré Probab. Statist. 32 (6): 765–78.
Fernique, X. 1974. “Minorations Des Fonctions Aléatoires Gaussiennes.” Ann. Inst. Fourier (Grenoble) 24 (2): 61–66. https://doi.org/10.5802/aif.506.
Furstenberg, H. 1977. “Ergodic Behavior of Diagonal Measures and a Theorem of Szemerédi on Arithmetic Progressions.” J. Analyse Math. 31: 204–56. https://doi.org/10.1007/BF02813304.
Gutman, Y., W. Huang, S. Shao, and X. Ye. 2018. “Almost Sure Convergence of the Multiple Ergodic Average for Certain Weakly Mixing Systems.” Acta Math. Sin. (Engl. Ser.) 34 (1): 79–90. https://doi.org/10.1007/s10114-017-6366-1.
Host, B., and B. Kra. 2005. “Nonconventional Ergodic Averages and Nilmanifolds.” Ann. Of Math. (2) 161 (1): 397–488. https://doi.org/10.4007/annals.2005.161.397.
Huang, W., S. Shao, and X. Ye. 2019. “Pointwise Convergence of Multiple Ergodic Averages and Strictly Ergodic Models.” J. Anal. Math. 139 (1): 265–305. https://doi.org/10.1007/s11854-019-0061-3.
Jamneshan, A. 2023. “An Uncountable Furstenberg–Zimmer Structure Theory.” Ergodic Theory Dynam. Systems 43 (7): 2404–36. https://doi.org/10.1017/etds.2022.43.
Jamneshan, A. 2026. “An Uncountable Furstenberg–Zimmer Structure Theory—Corrigendum.” Ergodic Theory Dynam. Systems 46 (1): 211–17. https://doi.org/10.1017/etds.2025.10214.
OpenAI. 2026a. Pointwise convergence of fourfold ergodic averages for mixing transformations. OpenAI Math Release preprint OAI:Pointwise-convergence-of-fourfold-ergodic-averages-for-mixing-transformations-October-4-2026.
OpenAI. 2026b. Rokhlin’s multiple-mixing problem for one transformation. OpenAI Math Release preprint OAI:Rokhlins-multiple-mixing-problem-for-one-transformation-September-23-2026.
OpenAI. 2026c. Triple ergodic averages with distinct integer slopes. OpenAI Math Release preprint OAI:Triple-ergodic-averages-with-distinct-integer-slopes-October-4-2026.
Sudakov, V. N. 1971. “Gaussian Random Processes and Measures of Solid Angles in Hilbert Space.” Dokl. Akad. Nauk SSSR 197 (1): 43–45.
Tropp, J. A. 2018. “Second-Order Matrix Concentration Inequalities.” Appl. Comput. Harmon. Anal. 44 (3): 700–736. https://doi.org/10.1016/j.acha.2016.07.005.
Ziegler, T. 2007. “Universal Characteristic Factors and Furstenberg Averages.” J. Amer. Math. Soc. 20 (1): 53–97. https://doi.org/10.1090/S0894-0347-06-00532-7.
Zimmer, R. J. 1976. “Extensions of Ergodic Group Actions.” Illinois J. Math. 20 (3): 373–409. https://doi.org/10.1215/ijm/1256049780.
LEVEL 1 COMPLETE!
You read 11,038 words and 907 formulas. Your math teacher would be proud.
Converted from the LaTeX source. Something look off? The original PDF is the real thing.

Cool Links: openai/math   Lean   Mathlib   arXiv   the real Coolmath Games