A D V E R T |
I S E M E N T |
| Math Sites: lean ages 13-∞ readme referees parents | >>> MAITH GAMES <<< | all 372 compute stand |
|
A zero-entropy system without a smooth positive-volume model
expertly designed by an internal OpenAI model · released 2026-09-25
· original PDF
IntroductionThe smooth realization problem asks which measure-preserving systems can be represented by smooth dynamics on a compact manifold. Here a smooth positive-volume model of an invertible probability-preserving transformation \((X,\mu,T)\) consists of a compact smooth manifold \(M\) of finite dimension, a smooth positive probability volume \(\nu\), and a \(C^\infty\) diffeomorphism \(S\) preserving \(\nu\), together with a measure-preserving isomorphism \[F:X_0\longrightarrow M_0,\qquad F\circ T=S\circ F,\] on invariant conull measurable sets. Both \(F\) and \(F^{-1}\) are measurable; no continuity is required. Here smooth positive volume means a probability measure with strictly positive smooth local densities, \(d\nu=f(x)|dx_1\cdots dx_d|\), where \(f>0\) is smooth. The manifold may have a smooth boundary, with these densities smooth up to the boundary. No orientation of \(M\) or orientation preservation by \(S\) is required. The manifold and its dimension may depend on the original system. A standard nonatomic probability space means a probability space isomorphic, modulo null sets, to the unit interval with Lebesgue probability. Finite entropy is a necessary restriction in the smooth realization problem. For example, Ruelle’s differentiable entropy inequality (Ruelle 1978, Theorem 2) gives finite entropy from compact derivative bounds. The question addressed here is whether this entropy restriction suffices when the invariant measure must be a smooth positive density. For invertible ergodic transformations of standard nonatomic probability spaces, two positive realization results help locate this restriction. Lind and Thouvenot realized finite-entropy ergodic transformations by Lebesgue-preserving homeomorphisms of the two-torus (Lind and Thouvenot 1977). They also proved that a hyperbolic two-torus automorphism represents every such system with measure entropy below the automorphism’s topological entropy, using a suitable invariant Borel probability of full support; Quas and Soo recall this earlier result in (Quas and Soo 2016, sec. 1). Quas and Soo’s Theorem 2 extends the result to toral automorphisms with no root-of-unity eigenvalue. Choosing a sufficiently high iterate of a hyperbolic automorphism therefore gives a smooth model for any prescribed finite-entropy system, with an invariant Borel probability of full support. Full support does not require a smooth positive density. The requirement on the invariant measure is essential. A realization with an arbitrary invariant Borel probability, which may be singular, does not by itself give a smooth positive-volume model. This distinction is part of the classical formulation discussed by Katok and Thouvenot (Katok and Thouvenot 1997, 324–26). They developed slow entropy invariants and constructed ergodic \(\mathbb Z^2\) actions having no Lipschitz realization on a compact space of finite box dimension, even when the model may carry an arbitrary invariant Borel probability (Katok and Thouvenot 1997, Proposition 4 and Corollary 2). The commuting generators must be realized together in that problem; its obstruction does not by itself answer the question for one transformation. Foreman and Weiss (Foreman and Weiss 2022, 2607) also discuss finite-entropy realization, in a formulation allowing an invariant measure merely equivalent to volume. Our conclusion concerns the smooth positive-density version; we do not assert an obstruction for every equivalent-to-volume invariant measure. Their anti-classification theorems concern the complexity of the measure-isomorphism relation among smooth transformations, including an ergodic area-preserving torus subclass (Foreman and Weiss 2022, Theorems 1–2 and Corollary 4). They do not assert that a particular abstract finite-entropy system lacks such a model. Subsequent anti-classification results treat weakly mixing smooth transformations (Kunde 2024, Theorem B) and mixing zero-entropy smooth transformations (Gerber and Kunde 2025, Theorem 2); these too study equivalence inside families of smooth models. We give a negative answer to the smooth-volume realization question, already among systems of zero entropy. For a finite probability vector \(p\), its Shannon entropy is \[H(p)=-\sum_a p(a)\log_2 p(a),\qquad 0\log_2 0=0.\] The entropy \(H_\mu(\mathcal P)\) of a finite measurable partition is the entropy of its atom probabilities; the entropy of a finite random variable is likewise the entropy of its law. The Kolmogorov–Sinai entropy of \(T\) is \[h_\mu(T)=\sup_{\mathcal P}\lim_{n\to\infty} \frac1n H_\mu\!\left(\bigvee_{j=0}^{n-1}T^{-j}\mathcal P\right),\] where \(\mathcal P\) ranges over finite measurable partitions. Theorem 1. There exists an ergodic invertible measure-preserving transformation \((X,\mu,T)\) of a standard nonatomic probability space such that \(h_\mu(T)=0\) and \((X,\mu,T)\) is not measurably conjugate, on invariant conull sets, to any \(C^\infty\) diffeomorphism of a compact finite-dimensional manifold preserving a strictly positive smooth probability density. The manifold may have smooth boundary and need not be orientable. The proof isolates a property of smooth positive volume that remains useful after an arbitrary measurable change of coordinates. At a dyadic time \(n=2^k\), partition the manifold into small cells, choose a cell with its volume as probability, and choose two independent points \(x,y\) from its normalized volume. On most such pairs, one can find a point \(z\) whose future shadows that of \(x\) and whose past shadows that of \(y\). More than the existence of this point is needed: the law of \(z\) on the retained pairs is bounded by twice the original volume. The number of cells is at most \(\exp(CB_kn)\) for a sequence with \(\sum_k B_k^{16}<\infty\). The precise statement is 7. This switch estimate is the main analytic ingredient. Its proof begins with a square-summable dyadic estimate for stationary diagonal derivative logs, rather than an assumed rate in an ergodic theorem. Finite-interval triangular changes of basis give the linear estimates needed to solve a boundary problem joining the two orbit halves. A relative exterior-power estimate controls the change in small endpoint minors, and a determinant identity and global injectivity control the law of the switch. The quantitative summability and the bound on the output measure are the features that will obstruct realization. We construct the opposing symbolic system from finite sets \(\mathcal W_i\) of words of lengths \(m_i\). Each word at the next level is a concatenation of words at the preceding level, and the system records their compatible alignment phases. Word labels are nonzero vectors over \(\mathbb F_2\). Certain pairs of labels separated by a power-of-two number of blocks must be orthogonal for a fixed alternating form. These constraints coexist with high entropy of short label lists and with a positive symbol-distance separation between distinct or misaligned words. The entropy of the label lists makes two independent lists unlikely to satisfy most of the orthogonality tests. A smooth switch would force their phases and most of their labels to agree with the corresponding past and future lists of a single legal symbolic point. Those lists do satisfy the tests. Summability of the smooth partition cost, together with the growth of the word scales, supplies times at which conditioning on a cell loses too little entropy to reconcile the two conclusions. The order of choices is important: the symbolic system is fixed first. Only then is a hypothetical smooth model, with any one finite dimension, considered. The scale argument absorbs all constants depending on that model. Thus the construction does not select a different counterexample for each dimension. Organization.[sec:stationary,sec:linear,sec:switching] prove the switching estimate. 5 constructs the finite word families. 6 supplies an ergodic probability on their phase extension and establishes the window entropy and graph constraints. 7 compares the two laws of one graph-passing statistic and proves 1. Conventions.Logarithms in entropies and combinatorial estimates are to base \(2\); natural logarithms are written \(\ln\), and derivative growth is written with \(\exp\) or \(e\). Total variation distance between finite laws \(p,q\) is \(\mathop{\mathrm{TV}}(p,q)=\tfrac12\sum_a|p(a)-q(a)|\). For finite random variables, conditional entropy means \[H(Y\mid Z)=\sum_{z:\Pr(Z=z)>0}\Pr(Z=z)\, H\bigl(\mathcal L(Y\mid Z=z)\bigr).\] Conditional partition entropy uses the same definition for atom labels. We write \(H_{\mathrm{bin}}(s)=-s\log_2 s-(1-s)\log_2(1-s)\) for binary entropy, with the same zero-mass convention at the endpoints. Constants denoted by \(C\) may change from line to line. In the smooth part they depend only on the fixed smooth data, dimension, and a fixed exceptional-mass tolerance, unless otherwise specified. All asymptotic estimates there are uniform over the points, bins, and finite-interval frames satisfying the displayed hypotheses. The empty determinant and empty product equal one. Stationary estimates for derivative growthThe switching construction requires control of derivatives on a long finite orbit segment. The accuracy may depend on the length of the segment, but the squared errors must be summable over dyadic lengths. We first obtain this summability for an arbitrary bounded stationary process. Applying the result to a triangular representation of the derivative then gives large sets of orbit segments with the required control. Deviation from a dyadic chordLet \((Y_t)_{t\in\mathbb Z}\) be a real stationary process with \(\mathbb EY_0^2<\infty\). Stationarity means that translating all indices in any finite list does not change its joint distribution. For \(h\geq 0\) define the block average and normalized maximal chord deviation by \[A_h=2^{-h}\sum_{t=0}^{2^h-1}Y_t, \qquad Z_h=2^{-h}\max_{0\leq u\leq 2^h} \left|\sum_{t=0}^{u-1}Y_t-uA_h\right|.\] The maximum is over integer \(u\), and an empty sum is zero. In particular \(Z_0=0\). The chord here uses the average on the particular block under consideration; it is not a line whose slope is the mean of the stationary process. Lemma 2 (Dyadic chord estimate). For the process above, put \[v_l=\mathbb EA_{l-1}^2-\mathbb EA_l^2,\qquad l\geq 1.\] Then \(v_l\geq 0\), \(\sum_{l\geq 1}v_l\leq\mathbb EY_0^2\), and \[ \|Z_h\|_{L^2} \leq \frac12\sum_{a=0}^{h-1}2^{-a/2}\sqrt{v_{h-a}} \qquad(h\geq 1). \tag{1}\] Consequently \[ \sum_{h\geq 1}\|Z_h\|_{L^2}^2 \leq C\,\mathbb EY_0^2 \tag{2}\] for an absolute constant \(C\). The same estimates hold for a block beginning at any deterministic integer time. Proof. Consider a block of length \(2^l\), and let \(U\) and \(V\) be the averages on its left and right halves. Its own average is \(W=(U+V)/2\). The identity \[\frac{(U-W)^2+(V-W)^2}{2} =\frac{U^2+V^2}{2}-W^2\] and stationarity give \[v_l =\mathbb E\frac{(U-W)^2+(V-W)^2}{2} =\mathbb E(U-W)^2.\] The last equality also follows pointwise from \(V-W=-(U-W)\). Thus \(v_l\geq0\), and telescoping the differences of second moments proves the asserted bound on their sum. No independence of the two halves is used. We next compare the cumulative-sum polygon on \([0,2^h]\) with its chord. For \(a=0,\ldots,h\), let \(F_a\) be the piecewise-linear interpolant that agrees with the cumulative sum at the multiples of \(2^{h-a}\). Then \(F_0\) is the chord and \(F_h\) agrees with every integer-time cumulative sum. Passing from \(F_a\) to \(F_{a+1}\) inserts the midpoint of each of \(2^a\) intervals of length \(2^{h-a}\). If \(U_j,W_j\) are respectively the left-child and parent averages on the \(j\)th such interval, the displacement at its midpoint is \(2^{h-a-1}(U_j-W_j)\). The difference of the interpolants is linear between the old and new vertices, so \[2^{-h}\|F_{a+1}-F_a\|_\infty \leq 2^{-a-1}\max_{1\leq j\leq 2^a}|U_j-W_j|.\] Bounding a squared maximum by the sum of squares and using stationarity of each interval gives \[\left\|2^{-h}\|F_{a+1}-F_a\|_\infty\right\|_{L^2} \leq 2^{-a-1}\bigl(2^a v_{h-a}\bigr)^{1/2}.\] The triangle inequality, first for the polygonal sup norm and then in \(L^2\), proves (1). Extend the sequence \((\sqrt{v_l})_{l\geq1}\) by zero to nonpositive indices. Its squared \(\ell^2\) norm is at most \(\mathbb EY_0^2\). The right side of (1) is its convolution with the summable sequence \((2^{-a/2-1})_{a\geq0}\). The triangle inequality for the \(\ell^2\) norms of its translates therefore proves (2). Finally, stationarity identifies the distribution of each translated block with that of the block just considered. ◻ A stationary triangular derivative representationFix a compact smooth manifold \(M\) of dimension \(d\geq1\), a smooth Riemannian metric, a smooth diffeomorphism \(S:M\to M\), and an \(S\)-invariant probability measure \(\nu\). In the switching argument \(\nu\) will have a smooth positive density. That extra assumption is not needed in this section. The constructions also apply to a compact manifold with smooth boundary. The compact frame extension and triangular cocycle have a classical precedent in Oseledets’s proof of the multiplicative ergodic theorem (Oseledets 1968, sec. 4, pp. 208–209, equations (6)–(10)). We use the full orthogonal frame bundle and a positive-diagonal upper triangular convention. The finite-time estimates below follow from the stationary chord lemma; we do not invoke a limiting Lyapunov splitting or a convergence rate in the multiplicative ergodic theorem. The orthonormal frame bundle \(\mathcal F M\) consists of the linear isometries \(e:\mathbb R^d\to T_xM\), with bundle projection \(\pi(e)=x\). For each frame there is a unique positive-diagonal QR factorization \[ D S_x\circ e=\widetilde S(e)\,R(e), \tag{3}\] where \(\widetilde S(e)\) is an orthonormal frame at \(Sx\) and \(R(e)\) is upper triangular with positive diagonal. This is ordinary Gram–Schmidt in the indicated order, with the length of each new perpendicular component chosen positive. It depends continuously, indeed smoothly, on the frame and the base point. The map \(\widetilde S\) is invertible. If \(DS_x e=e'R\), then \((DS_x)^{-1}e'=D(S^{-1})_{Sx}e'=eR^{-1}\), where the inverse linear map has domain \(T_{Sx}M\) and codomain \(T_xM\). The matrix \(R^{-1}\) is again upper triangular with positive diagonal. Taking positive-diagonal QR after applying the inverse derivative therefore recovers \(e\). This also proves continuity of the inverse. The construction uses the full orthonormal frame bundle and does not require a global frame or an orientation. There is a \(\widetilde S\)-invariant probability \(\widetilde\nu\) whose projection is \(\nu\). To see this, first put the normalized orthogonal-group Haar measure on every fiber over \(\nu\). Orthogonal changes of local trivialization preserve this fiber measure, so the resulting probability is defined without choosing a global section. Each of its iterates under \(\widetilde S\) projects to \(\nu\). Cesaro averages have a weakly convergent subsequence, because \(\mathcal F M\) is compact. The difference between the pushforward of an average and the average itself tends to zero against every continuous function, so its limit is invariant. Continuity of \(\pi\) preserves the marginal \(\nu\) in the limit. Compactness and invertibility provide a constant \(A_*>1\) such that \[ \|R(e)\|,\ \|R(e)^{-1}\|\leq A_* \qquad(e\in\mathcal F M). \tag{4}\] In particular \(A_*^{-1}\leq R(e)_{jj}\leq A_*\). The functions \[g_j(e)=\ln R(e)_{jj},\qquad 1\leq j\leq d,\] are continuous and bounded, and \((g_j(\widetilde S^t e))_{t\in\mathbb Z}\) is stationary under \(\widetilde\nu\). Proposition 3 (Good finite orbit segments). For each \(\eta>0\) there is a positive sequence \((b_k)_{k\geq1}\) with \[ b_k\geq 2^{-k/10}, \qquad \sum_{k\geq1}b_k^2<\infty, \tag{5}\] and compact sets \(\widetilde{\mathcal G}_k\subset\mathcal F M\) and \(\mathcal G_k=\pi(\widetilde{\mathcal G}_k)\subset M\) such that \[\widetilde\nu(\widetilde{\mathcal G}_k)\geq1-\eta/100, \qquad \nu(\mathcal G_k)\geq1-\eta/100.\] For \(n=2^k\), \(b=b_k\), and \(e\in\widetilde{\mathcal G}_k\), define \[A_t=R(\widetilde S^t e)\quad(-n\leq t<n), \qquad \lambda_j=\frac1{2n}\sum_{t=-n}^{n-1}\ln(A_t)_{jj}.\] Then for every \(j\) and every integer \(-n\leq u\leq n\), \[ \left|\sum_{t=-n}^{u-1}\ln(A_t)_{jj} -(u+n)\lambda_j\right|\leq bn. \tag{6}\] Consequently, on every subinterval \(-n\leq u\leq v\leq n\), \[ \left|\sum_{t=u}^{v-1}\ln(A_t)_{jj} -(v-u)\lambda_j\right|\leq2bn. \tag{7}\] Proof. Apply 2 to each \(g_j\). Let \(c_{j,h}\) be the right side of (1) for that process, and put \(\alpha_h^2=\sum_{j=1}^d c_{j,h}^2\). The lemma gives \(\sum_h\alpha_h^2<\infty\). For example, take \[b_k=40\eta^{-1/2}\alpha_{k+1}+2^{-k/10}.\] This sequence has the properties in (5). The interval \([-n,n)\) has length \(2^{k+1}\). Stationarity and the squared Markov inequality give \[\widetilde\nu\!\left( \max_{\substack{1\leq j\leq d\\-n\leq u\leq n}} \left|\sum_{t=-n}^{u-1}g_j(\widetilde S^t e) -(u+n)\lambda_j\right|>b_kn \right) \leq \frac{4\alpha_{k+1}^2}{b_k^2} \leq\frac{\eta}{400}.\] Define \(\widetilde{\mathcal G}_k\) by the complementary non-strict inequalities. These are finitely many continuous inequalities on a compact bundle, so this set and its projection are compact. Moreover \[\pi^{-1}(M\setminus\mathcal G_k) \subset\mathcal F M\setminus\widetilde{\mathcal G}_k.\] The pushforward identity \(\pi_*\widetilde\nu=\nu\) gives the asserted base-measure estimate. No measurable choice of a frame over each point of \(\mathcal G_k\) is needed. Finally, subtracting the chord error at \(u\) from the chord error at \(v\) gives (7), with the same rates \(\lambda_j\). ◻ We call a frame in \(\widetilde{\mathcal G}_k\) good at scale \(k\), and a base point in \(\mathcal G_k\) good at that scale. Every good base point has a good witness frame. The estimates below hold for every such witness; witnesses at different points or scales need not be chosen coherently. We do not require an individual point to be good at all sufficiently large scales. The stationary argument has now supplied the summable quantity that will control the cost of switching. Fix \(\eta\) and its sequence \((b_k)\), and use the following parameters at scale \(n=2^k\): \[ \begin{aligned} b&=b_k,& B&=b^{1/8},& g&=b^{1/2},\\ L&=\lceil b^{1/4}n\rceil,& h_0&=e^{-Bn},&r&=e^{-2Bn},&\delta&=e^{-3Bn/2}. \end{aligned} \tag{8}\] The parameter \(b\) controls the diagonal chord error \(bn\); \(g\) will be the separation margin for the finite-interval rates, \(L\) is the short comparison time, and \(B\) sets the spatial scales and the upper bound \(CBn\) for the logarithm of the number of partition cells. Here \(h_0\) and \(r\) will be spatial lengths; \(\delta\) will be a neighborhood radius for an orbit construction. Since \(b_k\to0\) and \(bn\geq n^{9/10}\), \[ \ln n=o(bn),\quad bn=o(gL),\quad gL=o(gn),\quad gn=o(L),\quad L=o(Bn). \tag{9}\] The ceiling in \(L\) has no effect on these comparisons. We also have \[\sum_k B_k^{16}=\sum_k b_k^2<\infty, \qquad Bn\geq n^{79/80}.\] All subsequent asymptotic estimates are for \(k\to\infty\). Constants may depend on the fixed smooth data, dimension, and \(\eta\), but not on the point, its good witness, or the scale. Linear estimates on finite orbit segmentsWe now turn the diagonal control of 3 into estimates for the full derivative. There are three steps. A finite-interval change of basis separates groups of diagonal rates, short derivatives identify nearly the same spaces at nearby base points, and a product-expansion estimate controls small perturbations relative to the possibly small exterior volumes. These statements will supply the linear data for switching orbit segments. Triangular products and separation of ratesFor invertible matrices \(A_t\) indexed by \(-n\leq t<n\), write \[A(t,s)=A_{t-1}\cdots A_s\quad(t>s),\qquad A(s,s)=\mathrm{Id},\qquad A(s,t)=A(t,s)^{-1}.\] All vector norms and exterior-power norms in the next proposition are Euclidean operator norms. If \(E\) is a \(q\)-dimensional subspace and \(A\) is injective, let \[J_q(A|E)=\|(\mathord{\bigwedge}^q A)\xi\|,\] where \(\xi\) is either unit simple \(q\)-vector spanning \(E\). This is the intrinsic \(q\)-dimensional volume expansion. We set \(J_0=1\), and the zeroth exterior power is the identity on \(\mathbb R\). Proposition 4 (Finite-interval splitting). Fix \(d\geq1\) and \(A_*>1\). There is a constant \(C\), depending only on these two quantities, with the following properties. Let \(n\geq2\), \(b>0\), and \(\ln(2n+1)\leq bn\). Suppose that \(A_t\), \(-n\leq t<n\), are upper triangular with positive diagonal and \[\|A_t\|,\ \|A_t^{-1}\|\leq A_*, \qquad \left|\sum_{t=-n}^{u-1}\ln(A_t)_{jj} -(u+n)\lambda_j\right|\leq bn\] for \(-n\leq u\leq n\) and \(1\leq j\leq d\), where \(\lambda_j=(2n)^{-1}\sum_{t=-n}^{n-1}\ln(A_t)_{jj}\). First, order the rates as \(\lambda_{(1)}\geq\cdots\geq\lambda_{(d)}\). For every \(0\leq q\leq d\) and every \(-n\leq s\leq t\leq n\), \[ \left|\ln\|\mathord{\bigwedge}^q A(t,s)\| -(t-s)\sum_{j=1}^q\lambda_{(j)}\right|\leq Cbn. \tag{10}\] The analogous estimate for \(A(s,t)\) uses the \(q\) largest members of \(\{-\lambda_1,\ldots,-\lambda_d\}\). Next, suppose that real \(\gamma\) and \(g>0\) separate all the rates: each \(\lambda_j\) is at most \(\gamma-g\) or at least \(\gamma+g\). Let \(F\) be the set of indices in the second group, let \(p=|F|\), and put \[\Lambda=\sum_{j\in F}\lambda_j,\qquad \Omega=-\sum_{j\notin F}\lambda_j.\] There are decompositions \(\mathbb R^d=E_t^s\oplus E_t^f\), \(-n\leq t\leq n\), invariant under each \(A_t\), with \(\dim E_t^f=p\), whose two projections have norm at most \(e^{Cbn}\). For \(s\leq t\) and \(l=t-s\) they satisfy \[\begin{align*} \|A(t,s)|E_s^s\|&\leq e^{(\gamma-g)l+Cbn}, & \|A(s,t)|E_t^f\|&\leq e^{-(\gamma+g)l+Cbn}, \tag{11}\\ \|\mathord{\bigwedge}^p A(t,s)\|&\leq e^{\Lambda l+Cbn}, & e^{\Lambda l-Cbn}\leq J_p(A(t,s)|E_s^f)&\leq e^{\Lambda l+Cbn}, \tag{12}\\ \|\mathord{\bigwedge}^{d-p} A(s,t)\|&\leq e^{\Omega l+Cbn}, & e^{\Omega l-Cbn}\leq J_{d-p}(A(s,t)|E_t^s)&\leq e^{\Omega l+Cbn}. \tag{13}\end{align*}\] If \(0<p<d\) and at least one of \(v_1,\ldots,v_p\) belongs to \(E_s^s\), then \[ \|A(t,s)v_1\wedge\cdots\wedge A(t,s)v_p\| \leq e^{(\Lambda-2g)l+Cbn}\prod_{j=1}^p\|v_j\|. \tag{14}\] The backward analogue has exterior degree \(d-p\), exponent \((\Omega-2g)l+Cbn\), and at least one factor in \(E_t^f\). When a group is empty its space is \(\{0\}\), all statements on that space are vacuous, and the exterior-degree-zero conventions apply. Proof. Diagonal normalization. Write \[e_j(u)=\sum_{t=-n}^{u-1}\bigl(\ln(A_t)_{jj}-\lambda_j\bigr), \qquad D_u=\mathop{\mathrm{diag}}(e^{e_1(u)},\ldots,e^{e_d(u)}).\] The chord hypothesis gives \(\|D_u\|,\|D_u^{-1}\|\leq e^{bn}\). The matrices \[\widehat A_t=D_{t+1}^{-1}A_tD_t\] are upper triangular with constant diagonal \((e^{\lambda_1},\ldots,e^{\lambda_d})\) and entries bounded by \(e^{Cbn}\). Since the rates are averages of logarithms of diagonal entries, \[ |\lambda_j|\leq\ln A_*. \tag{15}\] Their inverse matrices also have entries bounded by \(e^{Cbn}\). We record how triangular products are estimated. If \(m\) is fixed and upper-triangular \(m\)-by-\(m\) matrices have constant diagonal \(e^{\alpha_1},\ldots,e^{\alpha_m}\), bounded \(|\alpha_i|\), and off-diagonal entries at most \(e^{Cbn}\), a term in any product entry can change its coordinate index strictly at most \(m-1\) times. There are at most a fixed constant times \((2n+1)^{m-1}\) choices for those times and the intermediate indices. If the product length is \(l\) and a term has \(a\leq m-1\) strict changes, its diagonal factors contribute at most \(e^{\alpha_{\max}(l-a)}\). The omitted \(a\) diagonal factors cost at most a fixed factor \(e^{a|\alpha_{\max}|}\) after \(e^{\alpha_{\max}l}\) is factored out. The off-diagonal factors contribute \(e^{C_mbn}\). The polynomial number of terms is absorbed by \(\ln(2n+1)\leq bn\). Thus the product norm is at most \(e^{\alpha_{\max}l+C_mbn}\). Apply this argument to \(\mathord{\bigwedge}^q\widehat A_t\) in the coordinate-wedge basis, ordered lexicographically. This matrix is upper triangular of fixed dimension \(\binom dq\), and its diagonal entries are \(\exp(\sum_{j\in I}\lambda_j)\) for \(q\)-element index sets \(I\). The largest such sum is \(\sum_{j=1}^q\lambda_{(j)}\). For the reverse inequality, the diagonal of the product indexed by a maximizing set \(I\) is exactly \(\exp(l\sum_{j\in I}\lambda_j)\); that matrix entry is no larger than the operator norm. Converting back multiplies the exterior norms only by the two endpoint gauge bounds. This proves (10). Applying the same argument to inverse normalized steps in reverse time proves its backward version. The cases \(q=0,d\) follow as well: the former is the identity, and the latter is the determinant line. Removing entries between the groups. We next conjugate \(\widehat A_t\) by time-dependent upper unitriangular matrices to make the two coordinate index groups invariant. Process the superdiagonals in increasing order of their distance \(j-i\) from the diagonal. At a position \(i<j\) belonging to different groups, suppose the current entry is \(h_t\). Use an elementary change \(Q_t=\mathrm{Id}+u_tE_{ij}\), where \(E_{ij}\) has its only nonzero entry, equal to one, in position \((i,j)\). In \(Q_{t+1}^{-1}\widehat A_tQ_t\) the new entry at this position is \[h_t+e^{\lambda_i}u_t-e^{\lambda_j}u_{t+1}.\] To set it to zero, solve \[ u_{t+1}=q u_t+f_t,\qquad q=e^{\lambda_i-\lambda_j},\qquad f_t=e^{-\lambda_j}h_t. \tag{16}\] If \(q<1\), prescribe \(u_{-n}=0\) and solve forward: \[u_t=\sum_{a=-n}^{t-1}q^{\,t-1-a}f_a.\] If \(q>1\), prescribe \(u_n=0\) and solve backward: \[u_t=-\sum_{a=t}^{n-1}q^{\,t-1-a}f_a.\] In each formula the coefficients have modulus at most one, so \(|u_t|\leq2n\max_a|f_a|\). Different groups cannot have \(q=1\). This bound avoids dividing by the possibly small separation of rates. Right multiplication by \(Q_t\) changes column \(j\) by a multiple of column \(i\), and left multiplication by \(Q_{t+1}^{-1}\) changes row \(i\) by a multiple of row \(j\). Upper triangularity shows that every affected off-diagonal position has distance at least \(j-i\); at that distance only \((i,j)\) changes. The possible quadratic cross term is zero because \(E_{ij}\widehat A_tE_{ij}=(\widehat A_t)_{ji}E_{ij}=0\). Thus a later operation cannot restore an entry on an earlier superdiagonal. There are at most \(d(d-1)/2\) operations. At each operation the bound above, the current entry bounds, and (15) bound the shear and its inverse by \(e^{Cbn}\). Multiplying or adding a fixed number of such quantities only changes \(C\); the factor \(2n\) is absorbed by \(\ln(2n+1)\leq bn\). Induction over this fixed number of operations therefore gives matrices \(C_t\) with \[ \|C_t\|,\ \|C_t^{-1}\|\leq e^{Cbn}, \qquad B_t=C_{t+1}^{-1}A_tC_t, \tag{17}\] where \(B_t\) is upper triangular, has diagonal \(e^{\lambda_j}\), and has no entries between the two groups. Its entries and those of \(B_t^{-1}\) are bounded by \(e^{Cbn}\). All constants are independent of the length of the interval. Let \(\mathbb R^F\) be the coordinate span with indices in \(F\) and \(\mathbb R^{F^c}\) its coordinate complement. Set \[E_t^f=C_t\mathbb R^F,\qquad E_t^s=C_t\mathbb R^{F^c}.\] These spaces are invariant under the original matrices. Their projections are conjugates by \(C_t\) of coordinate projections, and hence have norm at most \(e^{Cbn}\) after changing \(C\). The two groups may be interlaced in the original coordinate order. Each induced block is nevertheless upper triangular in its inherited order; a fixed permutation makes the two blocks consecutive if desired. The triangular-product estimate within these blocks proves (11). Exterior volumes and mixed factors. The fast rates are exactly the \(p\) largest rates. The general exterior bound already proves the first inequality of (12). The normalized fast block has determinant \(e^{\Lambda l}\) on an interval of length \(l\). To account explicitly for its possibly oblique basis, define \[V_F(t)= \sqrt{\det\bigl((C_t|_{\mathbb R^F})^*(C_t|_{\mathbb R^F})\bigr)}.\] For \(p=0\) this Gram determinant is one. The endpoint bounds in (17) imply \(e^{-Cbn}\leq V_F(t)\leq e^{Cbn}\). The intrinsic volume expansion is exactly \[J_p(A(t,s)|E_s^f) =e^{\Lambda l}\frac{V_F(t)}{V_F(s)}.\] This proves both volume bounds in (12). Applying the same argument to the inverse slow block gives (13), including either sign of \(\Omega\). The normalized block decomposition induces an invariant decomposition of each exterior power according to the number of fast and slow factors. In degree \(p\), every sector with a slow factor has diagonal rate sum at most \(\Lambda-2g\): at least one fast rate has been replaced by a slow rate. The triangular-product estimate therefore bounds this entire sector by \(e^{(\Lambda-2g)l+Cbn}\). If one original vector is slow, split each of the other vectors into its fast and slow parts. There are only a fixed number of terms, each still having a slow factor. The projection norms and endpoint gauges are absorbed into \(e^{Cbn}\), proving (14). Reversing time and exchanging the groups gives the backward statement. The empty-group cases require none of these mixed-factor assertions and follow from the stated conventions. ◻ The estimates concern every subinterval of one good finite trajectory. The change of basis is made once on that trajectory; only its values at the two ends of a subinterval enter a product estimate. In particular, a gauge loss is not multiplied once for each time step. For a good frame in 3, the matrices of the derivative in the moving orthonormal frames satisfy the hypotheses above for all sufficiently large \(k\). The proposition therefore gives tangent spaces \(E_t^s,E_t^f\subset T_{S^t x}M\) and the corresponding derivative estimates. We use this notation also after expressing the spaces in chart coordinates. Coordinate maps from a fixed uniformly regular finite atlas conjugate every product only at its endpoints, so their bounded norms and inverse norms change the constants in the estimates, but not their form. Spaces that can be chosen uniformly in a spatial binWe need spaces at nearby starting points to agree much more accurately than their possible small angle. The comparison uses short derivative products and the parameters in (8). Fix a finite family of coordinate charts with smaller relatively compact subcharts covering \(M\). Their coordinate maps, inverse maps, and the local expressions of \(S\) and \(S^{-1}\) have uniform bounds through order two on the relevant compact sets. Each spatial bin below is contained in one of the smaller subcharts, has coordinate diameter at most \(C h_0\), and uses that specified chart to identify its tangent spaces with \(\mathbb R^d\). At a boundary of \(M\) use nested half-box charts. For equal-dimensional subspaces \(E,F\subset\mathbb R^d\), measure their distance by \(\|\Pi_E-\Pi_F\|\), where \(\Pi_E,\Pi_F\) are Euclidean orthogonal projections. This distance concerns the subspaces and does not require a choice of bases in them. Lemma 5 (Common spaces in a bin). For all sufficiently large \(k\), each spatial bin \(U\) meeting \(\mathcal G_k\) has a number \[\gamma_U\in[g,(8d+1)g]\] such that every good witness frame over \(U\) has all its rates at most \(\gamma_U-g\) or at least \(\gamma_U+g\), with the same number \(p_U\) in the second group. Apply 4 using this threshold. There is \(c>0\), independent of the bin and the witnesses, such that:
In particular, fix one representative space for each nonempty combination of bins in (ii). Orthogonal projection onto it, restricted to any corresponding actual endpoint space, is an isomorphism with inverse norm at most \(2\). If \(p_U=0\) or \(p_U=d\), the relevant zero or full spaces agree exactly. Proof. Short products in common charts. Choose one reference point in a starting bin. Uniform Lipschitz bounds for \(S\) and \(S^{-1}\) give, for another point in that bin, \[d_M(S^t x,S^t y)\leq C A_0^{|t|}h_0 \qquad(|t|\leq L)\] with a fixed \(A_0>1\). Since \(L=o(Bn)\), this tends to zero uniformly. At each time on either short reference orbit choose a chart whose smaller subchart contains the reference point. All nearby orbits then remain in the same larger chart at that time. Use the prescribed bin chart at the initial time. The one-step coordinate derivatives and their inverses are bounded by a fixed constant, and their differences are at most \(C A_0^{|t|}h_0\). The telescoping product identity consequently gives \[ \|P_x^\pm-P_y^\pm\|\leq e^{-Bn+CL}=:\zeta_k,\qquad \sigma_{\min}(P_x^\pm)\geq e^{-CL}=:m_k. \tag{18}\] Here \(P_x^+\) and \(P_x^-\) denote \(DS^L_x\) and \(DS^{-L}_x\), respectively, in these common initial and terminal charts. For example, summing the \(L\) terms in the product difference bounds it by \(C L A_1^{2L}h_0\), which has the stated form. The inverse product bound gives the minimum stretch. On a manifold with boundary the same proof uses the nested half-box charts and uniform one-sided smooth bounds; equivalently the finitely many local expressions may be extended to their surrounding boxes. The singular values are Lipschitz with respect to operator norm. Since \(\zeta_k/m_k\to0\), (18) implies, after increasing the constants if necessary, \[ |\ln\sigma_j(P_x^+)-\ln\sigma_j(P_y^+)| \leq 2\zeta_k/m_k . \tag{19}\] For completeness, the singular-value Lipschitz bound follows from the min–max characterization, because \(\|Pv\|\) and \(\|Qv\|\) differ by at most \(\|P-Q\|\) on unit vectors. The logarithmic inequality then follows from the lower bound \(m_k\). A threshold for every witness. Let \(\lambda_{(j)}(e)\) be the sorted rates of a good witness \(e\). Subtracting (10) for exterior degrees \(j\) and \(j-1\) gives \[ \left|L^{-1}\ln\sigma_j(P_{\pi(e)}^+) -\lambda_{(j)}(e)\right| \leq Cbn/L. \tag{20}\] This argument uses the no-gap part of 4; a threshold has not yet been chosen. Fix one good witness \(e_*\) in \(U\). Equations (19) and (20) show that every good witness in \(U\), including every other witness over the same point, satisfies \[\max_j|\lambda_{(j)}(e)-\lambda_{(j)}(e_*)| \leq Cbn/L+2\zeta_k/(m_kL)=o(g).\] Consider the \(d+1\) candidate thresholds \((1+8j)g\), \(0\leq j\leq d\). A real number lies within distance \(3g\) of at most one candidate. The \(d\) representative rates therefore leave a candidate whose distance from all of them is greater than \(3g\). For large \(k\), all the other sorted rate lists differ by less than \(g\), so this candidate separates every list with a margin greater than \(2g\). It has the weaker required margin \(g\), and the number of rates above it is constant. Fix this candidate as \(\gamma_U\). Comparison of subspaces using one short map. Assume first that \(0<p_U<d\) and put \(q=d-p_U\). Use the single forward short map \(P_*=P_{\pi(e_*)}^+\). Let \(W\) be its bottom \(q\) right-singular subspace. Its top \(p_U\) singular values are at least \(\exp((\gamma_U+g)L-Cbn)\) by (20). If \(v\) is a unit vector in the slow space of any good witness \(e\), then (11) and (18) give \[\|P_*v\|\leq e^{(\gamma_U-g)L+Cbn}+\zeta_k.\] Orthogonality of the images of right-singular vectors implies \[\|\Pi_{W^\perp}v\| \leq e^{-2gL+Cbn} +\zeta_k e^{-(\gamma_U+g)L+Cbn} \leq e^{-2gL+Cbn}+e^{-Bn+CL}.\] In the last bound, constants absorb \(|\gamma_U|=O_d(g)\) and the other terms bounded by a constant times \(L\). This is at most \(e^{-c_1gL}\) for some \(c_1>0\) and large \(k\). To pass from this directed estimate to a subspace estimate, observe that if \(\|\Pi_{W^\perp}|E\|\leq\epsilon<1\) and \(\dim E=\dim W\), then \(\Pi_W|E\) is an isomorphism. The space \(E\) is a graph over \(W\) of norm at most \(\epsilon/\sqrt{1-\epsilon^2}\). The corresponding orthogonal projectors differ by at most a fixed multiple of \(\epsilon\) when \(\epsilon\leq1/2\). Thus every slow space is within \(C e^{-c_1gL}\) of this same \(W\); the triangle inequality compares any two slow spaces. No continuity assertion about singular vectors or chosen good frames has been used. For fast spaces use the single backward short map \(P_{\pi(e_*)}^-\). Its bottom \(p_U\) right-singular space is the comparison space. The top \(q\) singular values are at least \(\exp(-(\gamma_U-g)L-Cbn)\), whereas a unit fast vector has backward stretch at most \(\exp(-(\gamma_U+g)L+Cbn)\). The identical calculation gives the error bound \[e^{-2gL+Cbn} +\zeta_k e^{(\gamma_U-g)L+Cbn} \leq e^{-2gL+Cbn}+e^{-Bn+CL}.\] This proves (i) after reducing \(c\) to absorb fixed factors. In particular, the error is \(o(e^{-C_0bn})\) for every fixed \(C_0\), by \(bn=o(gL)\). Endpoint comparisons. Keep the threshold of \(U\). For paths starting in \(U\) and ending at time \(n\) in a common bin \(V\), choose one representative path for this pair of bins. Apply the preceding short-product comparison from their time-\(n\) points backward for length \(L\). Their initial closeness for this comparison is supplied by \(V\), not by propagating the starting-bin closeness for \(n\) steps. All intervals \([n-L,n]\) lie inside the good interval, and all paths use the same group dimension and threshold. In particular, (10) applies on each \([n-L,n]\) with the same full-segment rates, so the singular-value comparison and the backward argument just proved compare their fast spaces at \(n\). At \(-n\), the identical all-subinterval estimate for forward short maps on \([-n,-n+L]\) gives the slow-space assertion. Finally, if \(E\) and its chosen representative \(F\) have equal dimension and \(\|\Pi_E-\Pi_F\|\leq\epsilon\), then for \(v\in E\), \[\|\Pi_F v\|\geq\sqrt{1-\epsilon^2}\,\|v\|.\] The restriction is onto \(F\) by equality of dimensions, and its inverse has norm at most \((1-\epsilon^2)^{-1/2}\leq2\) for large \(k\). If one group is empty, the associated spaces are identically zero or the full coordinate space, so all comparisons and projection conclusions follow directly. ◻ For later use, choose one good witness in each nonempty starting bin and orthonormal bases within each of its two spaces at time zero. The matrix having these two bases as columns has norm bounded by a dimension-only constant and inverse norm at most \(e^{Cbn}\), by the projection bounds of 4. It gives fixed product coordinates for that bin. Likewise 5 supplies fixed endpoint projections for each relevant pair of bins. All these choices are finite at a fixed scale. Their constancy within the indicated bins, rather than any continuity of choices between bins, is what the switching construction will use. Relative stability of exterior productsThe bounds on every unperturbed subinterval also control a perturbed long product. This must be a relative estimate: a leading volume can be exponentially small, so continuity with an absolute error alone is insufficient. Lemma 6 (Exterior-product perturbation). Fix \(d\geq1\) and \(K\geq1\). Let \(0\leq p\leq d\), let \(m\geq1\), and let \(A_0,\ldots,A_{m-1}\) and \(\widehat A_0,\ldots,\widehat A_{m-1}\) be real \(d\)-by-\(d\) matrices such that \[\|A_t\|\leq K,\qquad \|\widehat A_t-A_t\|\leq\epsilon\leq1.\] Use the same product notation as above for nonnegative indices. Suppose \(a\geq0\), \(\sigma\in\mathbb R\), and \[\|\mathord{\bigwedge}^p A(t,s)\| \leq e^{\sigma(t-s)+a} \qquad(0\leq s\leq t\leq m).\] There is \(C=C(d,K)\) such that \[ \left\|\mathord{\bigwedge}^p\widehat A(m,0) -\mathord{\bigwedge}^p A(m,0)\right\| \leq e^{\sigma m+a} \left[(1+C\epsilon e^{a+|\sigma|})^m-1\right]. \tag{21}\] For \(p=0\) the left side is zero. In particular, use the scales (8), take \(m\leq2n\), \(a\leq C_1bn\), \(|\sigma|\leq C_2\), and \(\epsilon\leq C_3e^{-3Bn/2}\), with fixed constants. Then for every fixed \(C_0>0\) the left side of (21) is \[ o\!\left(e^{\sigma m-C_0bn}\right), \tag{22}\] uniformly for all matrices satisfying these hypotheses. Proof. The exterior power is multiplicative. Multilinearity in its \(p\) columns, together with the one-step bound \(K+1\) for the perturbed matrices, gives \[E_t:=\mathord{\bigwedge}^p\widehat A_t -\mathord{\bigwedge}^p A_t, \qquad \|E_t\|\leq C\epsilon.\] Expand the product using either the unperturbed exterior matrix or \(E_t\) at each time. A term with \(q\geq1\) insertions has \(q+1\) unperturbed intervals whose total length is \(m-q\). Allowing empty intervals, the hypothesis bounds that term by \[e^{\sigma(m-q)+(q+1)a}(C\epsilon)^q =e^{\sigma m+a}(C\epsilon e^{a-\sigma})^q.\] There are \(\binom mq\) choices of the inserted positions. Sum over \(q=1,\ldots,m\) and use \(-\sigma\leq|\sigma|\) to prove (21). This calculation explains why negative \(\sigma\), as may occur for backward exterior volumes, is harmless: each removed time step contributes only the displayed factor \(e^{-\sigma}\). For the final assertion put \(u=C\epsilon e^{a+|\sigma|}\) and use \((1+u)^m-1\leq mu e^{mu}\). After division by \(e^{\sigma m-C_0bn}\) the bound is at most \[C m\epsilon e^{(2C_1+C_0)bn+C_2} \exp\!\bigl(Cm\epsilon e^{C_1bn+C_2}\bigr).\] The logarithm of its decisive first factor is at most \(-3Bn/2+O(bn)+\ln(2n)\), which tends to \(-\infty\) by (9). The quantity in the second exponential tends to zero for the same reason. All constants are fixed, proving the uniform little-oh assertion. For exterior degree zero the two products are both the identity, so the result holds without an expansion. ◻ The lemma applies to forward degree-\(p\) products with \(\sigma=\Lambda\) and to backward degree-\((d-p)\) products with \(\sigma=\Omega\) from 4. Together with 5, it supplies uniform linear and relative-volume estimates for orbit segments whose starting points occupy one small spatial bin. Switching a past and a future with controlled volumeThe finite-interval estimates now give a geometric operation on pairs of nearby points. It joins the future of one point to the past of the other. We must control the distribution of the resulting point as well as its orbit: this requires constructing both complementary joins and estimating their joint Jacobian. The joining geometry is related to the canonical local-product coordinates of hyperbolic dynamics (Smale 1967, sec. I.7). Solving a linear difference equation and then a nonlinear contraction also has a standard shadowing precedent (Chow et al. 1989, Lemma 3.2 and Proposition 4.1). Here the slow group may include zero or positive rates; we prove the finite-interval boundary inverse directly. The measure inequality requires the further determinant and injectivity arguments below. Theorem 7 (Smooth switching). Let \(M\) be a compact smooth manifold of positive finite dimension, let \(\nu\) be a smooth positive probability volume in the density sense of 1, and let \(S:M\to M\) be a smooth diffeomorphism preserving \(\nu\). Fix a Riemannian distance \(d_M\) and \(\eta>0\). There exist positive numbers \(B_k\) with \(\sum_k B_k^{16}<\infty\) and a constant \(C\) such that the following holds for all sufficiently large \(k\), with \(n=2^k\). There is a finite measurable partition \(\mathcal D_k\) of \(M\) with \(\#\mathcal D_k\le \exp(CB_kn)\). Define the probability measure \[ \rho_k=\sum_{\substack{Q\in\mathcal D_k\\\nu(Q)>0}} \frac{(\nu|_Q)\otimes(\nu|_Q)}{\nu(Q)}. \tag{23}\] Thus one first samples a cell with its \(\nu\)-weight and then samples two independent points in that cell. There are a measurable set \(G_k\subset M\times M\), with \(\rho_k(G_k)\ge1-\eta\), and a measurable map \(z_k:G_k\to M\) such that \[ (z_k)_*(\rho_k|_{G_k})\le 2\nu \tag{24}\] and \[ \sup_{(x,y)\in G_k}\left[ \max_{0\le t\le n}d_M(S^t z_k(x,y),S^t x) +\max_{-n\le t\le0}d_M(S^t z_k(x,y),S^t y) \right]\longrightarrow0. \tag{25}\] The measure on the left of (24) is not normalized after restriction to \(G_k\). The same conclusions hold if \(M\) has a smooth boundary. The proof occupies the rest of this section. Its first two steps produce two complementary switches \((z,w)\). The last two steps show that their pair map is injective and has volume Jacobian \(1+o(1)\). This is what will give (24). Grids, retained pairs, and fixed endpoint projectionsWe may assume \(0<\eta<1\). Apply 3 with parameter \(\eta/2\). It supplies \(b=b_k\ge n^{-1/10}\), \(\sum_k b_k^2<\infty\), and a compact set of good base points, denoted here by \(\mathcal A_k\) to distinguish it from the retained pair set, with suitable frame paths on \([-n,n]\), such that \(\nu(M\setminus\mathcal A_k)<\eta/100\). Set \[ B=b^{1/8},\quad h_0=e^{-Bn},\quad r=e^{-2Bn},\quad L=\lceil b^{1/4}n\rceil,\quad g=b^{1/2},\quad \delta=e^{-3Bn/2}. \tag{26}\] Here and below constants depend only on the fixed smooth data, dimension, and \(\eta\). They never depend on a point, bin, frame witness, or \(k\). In particular, \[ \log n=o(bn),\quad bn=o(gL),\quad gL=o(gn),\quad gn=o(L),\quad L=o(Bn). \tag{27}\] These relations absorb every fixed power of a quantity bounded by \(e^{Cbn}\) or \(e^{Cgn}\) at the places where it is used below. Choose a finite system of coordinate patches with smaller coordinate boxes covering \(M\), each smaller box having closure in its patch. Assign points to the first smaller box containing them. This gives disjoint chart regions whose boundaries lie in finitely many smooth coordinate faces. The coordinate maps, their inverses, and the chart density of \(\nu\) have uniform bounds on slightly larger compact patches; the density has a uniform positive lower bound. Within each region use the grid of coordinate cubes of side \(h_0\), keeping only cubes whose closures lie in that region. We call these cubes coarse bins. The discarded part has volume \(O(h_0)\): it is contained in a fixed multiple of the \(h_0\)-neighborhood of the finitely many region boundaries. All comparisons at a common starting or endpoint bin use its one fixed coordinate map. For a starting bin \(Q\) meeting \(\mathcal A_k\), 5 gives a common threshold \(\gamma_Q\in[g,Cg]\), a fast dimension \(p_Q\), and reference spaces \(E_Q^s,E_Q^f\) at time zero. The spaces of every good frame path beginning in \(Q\) are within \(e^{-cgL}\) of these spaces. Their complementary projections and the inverse of a basis orthonormal within each space have norm at most \(e^{Cbn}\). The threshold and the reference spaces are chosen once for this bin. This involves finitely many choices at each scale, not a choice of a frame as a measurable function on \(\mathcal A_k\). In the starting chart let \(E_Q:\mathbb R^{d-p_Q}\times\mathbb R^{p_Q}\to\mathbb R^d\) be the matrix with these orthonormal slow and fast columns, respectively. Tile the chart by the parallelepipeds \(E_Q([0,r)^d+rj)\), \(j\in\mathbb Z^d\), and keep the ones whose closures lie in \(Q\). Translations of the grid are immaterial. These are the product cells. Although \(\|E_Q^{-1}\|\le e^{Cbn}\), the forward norm satisfies \(\|E_Q\|\le\sqrt2\): for \((a,b)\), \(|E_Q(a,b)|\le |a|+|b|\). Thus every cell has physical diameter at most \(C_dr\) and volume at least \(e^{-Cbn}r^d\) up to a fixed factor. The total number of full cells is consequently at most \[ C e^{Cbn}r^{-d}\le e^{CBn}. \tag{28}\] Every point of a coarse bin at distance greater than \(C_dr\) from its walls belongs to a full product cell. Hence the partial-cell loss, summed over all coarse bins, is \(O(r/h_0)=o(1)\). Use one additional cell for all omitted points and for bins not meeting \(\mathcal A_k\). Assign grid boundaries arbitrarily; they are null. This completes the partition \(\mathcal D_k\) without changing (28). For two points \(x=(x_s,x_f)\) and \(y=(y_s,y_f)\) in the product coordinates of one cell, the ideal exchange \[z_0=(y_s,x_f),\qquad w_0=(x_s,y_f)\] stays in that cell. These points combine the slow coordinate of one input with the fast coordinate of the other. For the pairs retained below, we will construct actual orbit joins \(z,w\) whose product coordinates differ from \(z_0,w_0\) by \(o(r)\). Removing thin coordinate strips will therefore keep those joins in their original cell. This is why the cells follow the two reference spaces, which need not be orthogonal; see 1. Fix a small number \(\tau>0\), independent of \(k\). The interior of a product cell retained at this stage consists of points whose every product coordinate is at least \(\tau r\) from its two cell faces. The discarded coordinate strips have at most \(2d\tau\) of its Euclidean volume. The constant determinant of \(E_Q\) cancels in this fraction. Moreover, the ratio between the largest and smallest smooth density on the physical cell is \(1+O(r)\), uniformly in \(Q\). Thus the total smooth volume of these strips is at most \(2d\tau+o(1)\). The corresponding strips of width \(\tau h_0\) in coarse bins have the same bound. These are estimates for the original measure \(\nu\), without conditioning on membership in \(\mathcal A_k\). If \(M\) has boundary, also exclude points having any iterate at an integer time in \([-n,n]\) at distance at most \(c_0h_0\) from \(\partial M\), where \(c_0>0\) is fixed. A smooth boundary collar has volume \(O(h_0)\), so invariance bounds this loss by \(O(nh_0)=o(1)\). The latter limit follows from \(Bn\ge n^{79/80}\). Local coordinate extensions across the boundary may be used for uniform derivative bounds; all orbit tubes constructed below lie in the original manifold. Let \(R_k\) be the Borel set of points in \(\mathcal A_k\) that lie in the retained interiors of full product cells, whose two endpoints \(S^{\pm n}u\) lie in coarse bins at coordinate distance at least \(\tau h_0\) from their faces, and that satisfy the boundary provision when needed. Bins not meeting \(\mathcal A_k\) contain no such point. The preceding estimates and invariance give \[ \nu(M\setminus R_k)\le \eta/100+C_d\tau+o(1). \tag{29}\] Choose \(\tau\) so that \(C_d\tau<\eta/100\), and then increase \(k\) so that the right-hand side is less than \(\eta/4\). Both marginals of (23) are exactly \(\nu\). Consequently, for \[ G_k=(R_k\times R_k)\cap \bigcup_{Q'\text{ a full product cell}}(Q'\times Q'), \qquad \rho_k(G_k)\ge1-\eta/2\ge1-\eta. \tag{30}\] The measure \(\rho_k\) is supported on pairs in the same partition cell, and \(R_k\) meets only full product cells. Hence \(\rho_k(G_k)=\rho_k(R_k\times R_k)\), and the two marginal bounds give the displayed estimate. No uniform distribution of good points inside individual cells is required. We next fix the maps that will specify the switch. For each starting bin \(Q\) and endpoint bin \(U\) reached at time \(n\) by a point of \(\mathcal A_k\cap Q\), choose a representative endpoint fast space. Let \(P^+_{Q,U}:\mathbb R^d\to\mathbb R^{p_Q}\) be its orthogonal projection, followed by a fixed orthonormal identification with \(\mathbb R^{p_Q}\). Similarly, for an endpoint bin \(V\) reached at time \(-n\), define \(P^-_{Q,V}:\mathbb R^d\to\mathbb R^{d-p_Q}\) from a representative slow space. By the endpoint conclusion of 5, these projections restrict to isomorphisms on the actual fast and slow endpoint spaces of every corresponding good path, with inverse norm at most \(2\) for large \(k\). The maps depend only on the displayed bins and the starting threshold. They do not depend on a later choice of input in those bins. Zero-dimensional projections and determinants have their usual empty space meanings. For a retained pair \((x,y)\), write \(Q\) for its coarse starting bin and \(U_x,V_x,U_y,V_y\) for the coarse bins of \(S^nx,S^{-n}x,S^ny,S^{-n}y\), respectively. If \(\chi_U\) denotes the coordinate map attached to a bin, define the following smooth maps on the open sets where the coordinates make sense: \[ \begin{aligned} F_x(u)&=P^+_{Q,U_x}\chi_{U_x}(S^nu),& P_x(u)&=P^-_{Q,V_x}\chi_{V_x}(S^{-n}u),\\ F_y(u)&=P^+_{Q,U_y}\chi_{U_y}(S^nu),& P_y(u)&=P^-_{Q,V_y}\chi_{V_y}(S^{-n}u). \end{aligned} \tag{31}\] The subscripts denote fixed bin labels. We seek \(z,w\) satisfying \[ F_x(z)=F_x(x),\quad P_y(z)=P_y(y),\qquad F_y(w)=F_y(y),\quad P_x(w)=P_x(x). \tag{32}\] The next step both solves these equations near the prescribed halves and gives a larger neighborhood in which each solution is unique. 2 displays the required exchange of endpoint data. A boundary inverse and a nonlinear orbitLemma 8. For every sufficiently large \(k\) and every \((x,y)\in G_k\), the first two equations of (32) have a solution \(z\) with \[ \max_{0\le t\le n}d_M(S^tz,S^tx) +\max_{-n\le t\le0}d_M(S^tz,S^ty) \le e^{Cgn}r. \tag{33}\] In fixed regular coordinates along these halves the solution is unique among displacement arrays of sup norm at most \(\delta\). The complementary solution \(w\) has the analogous properties with \(x,y\) exchanged. In the product coordinates of the common cell, if \(x=(x_s,x_f)\) and \(y=(y_s,y_f)\), then \[ z=(y_s,x_f)+o(r),\qquad w=(x_s,y_f)+o(r), \tag{34}\] uniformly over all retained pairs. Both outputs remain in the original product cell, and their relevant endpoints remain in the respective four original endpoint bins. Proof. Uniform coordinates. At each reference location choose one of finitely many slightly larger regular patches containing a fixed-radius neighborhood of that location. At time zero both strings use \(\chi_Q\); at the terminal times they use the endpoint charts already specified. Shrinking the fixed radius once, the one-step coordinate maps \(\chi_{t+1}\circ S\circ\chi_t^{-1}\) and their first two derivatives are uniformly bounded wherever they will be used. The inverse one-step maps also have uniform first and second derivative bounds. For example, differentiating the inverse identity bounds their second derivatives by a fixed multiple of the forward second-derivative bound and the cube of the inverse first-derivative bound. This follows from a finite atlas with nested compact patches and smoothness on compact \(M\); it does not depend on the number of steps. A patch can be chosen separately at each internal time. No iterate \(S^n\) is assigned a uniform second derivative bound. Use displacement variables \(v_t^+\) about \(S^tx\), \(0\le t\le n\), and \(v_t^-\) about \(S^ty\), \(-n\le t\le0\). Write \(A_t^+,A_t^-\) for the derivative steps in these charts. Matching the two locations at zero and imposing the terminal equations gives the linear system \[ \begin{aligned} v_{t+1}^{\pm}-A_t^{\pm}v_t^{\pm}&=f_t^{\pm},\\ v_0^+-v_0^-&=m,\\ P^+_{Q,U_x}v_n^+&=h_+,\qquad P^-_{Q,V_y}v_{-n}^-=h_-. \end{aligned} \tag{35}\] The matching datum for the actual problem is \(m=\chi_Q(y)-\chi_Q(x)\); its norm is \(O(r)\). The actual terminal data are zero. The inverse in weighted coordinates. Let \(\gamma=\gamma_Q\) and put \(\widetilde v_t=e^{-\gamma t}v_t\). Use the same weighting on equation data, with the step datum at its output time \(t+1\). By 4, the weighted propagators on slow spaces forward and fast spaces backward have norm at most \(e^{Cbn}e^{-g|t-s|}\). The associated projections have norm at most \(e^{Cbn}\). Geometric sums cost at most \(C/g\le e^{Cbn}\), by (27) and the floor on \(b\). On the future half, prescribe a slow vector \(a\in E_0^s(x)\). Variation of constants gives its slow part at every time by propagating \(a\) forward and adding the projected step data. If \(s_n\) is the resulting slow value at time \(n\), the fast value there is uniquely \[c_n=(P^+_{Q,U_x}|_{E_n^f(x)})^{-1} (\widetilde h_+-P^+_{Q,U_x}s_n).\] Propagate \(c_n\) and the fast step data backward. All forced terms have norm at most \(e^{Cbn}\) times the sup norm of the weighted data. The dependence of the resulting fast value at zero on \(a\) is the map \[ K_+=-\widetilde A_x(0,n)|_{E_n^f(x)} (P^+_{Q,U_x}|_{E_n^f(x)})^{-1} P^+_{Q,U_x}\widetilde A_x(n,0)|_{E_0^s(x)}, \qquad \|K_+\|\le e^{-2gn+Cbn}. \tag{36}\] Here \(\widetilde A(t,s)=e^{-\gamma(t-s)}A(t,s)\). Both decaying factors are present in (36); there is no assumption that the endpoint projection annihilates the slow space. Thus \(\widetilde v_0^+=a+K_+a+q_+\), with \(q_+\) bounded by the weighted data as above. On the past half prescribe \(b'\in E_0^f(y)\), propagate fast backward, use the slow endpoint projection to fix the slow value at \(-n\), and then propagate slow forward. This gives \(\widetilde v_0^-=b'+K_-b'+q_-\), where \[ \|K_-\|\le e^{-2gn+Cbn},\qquad \|q_-\|\le e^{Cbn}\|\text{weighted data}\|_{\infty}. \tag{37}\] The map \((a,b')\mapsto a-b'\) from \(E_0^s(x)\oplus E_0^f(y)\) to the starting coordinate space has inverse norm at most \(e^{Cbn}\). Indeed the corresponding reference spaces are complementary with this bound, and their perturbation is \(e^{-cgL}\); multiplication by the reference inverse still gives a perturbation tending to zero. The terms \(K_+a-K_-b'\) are smaller still. A Neumann series therefore solves the matching equation with inverse cost \(e^{Cbn}\). This proves an \(e^{Cbn}\) inverse bound for the full weighted system. Removing the weights costs at most \(e^{2\gamma(n+1)}\), so its unweighted inverse \(\mathcal L^{-1}\) satisfies \[ \|\mathcal L^{-1}\|_{\infty\to\infty}\le Q_k, \qquad Q_k=e^{Cgn}. \tag{38}\] If a group is empty, omit its variables, propagation formula, and terminal datum. The matching map on the remaining spaces and the same bounds apply, including the cases \(p_Q=0,d\). Contraction with a larger uniqueness radius. Taylor expansion of each one-step map gives a residual \(\mathcal R(v)\) with \[ \|\mathcal R(v)\|_\infty\le C\|v\|_\infty^2, \qquad \|\mathcal R(v)-\mathcal R(u)\|_\infty \le C\delta\|v-u\|_\infty \tag{39}\] on the sup norm ball of radius \(\delta\). Every residual coordinate depends on just its one step, so the constant has no factor \(n\). The matching and terminal equations are affine and have zero nonlinear residual. The nonlinear problem is \(v=\mathcal L^{-1}(a_0+\mathcal R(v))\), where \(\|a_0\|_\infty\le Cr\). The scale relations give \[CQ_k\delta\longrightarrow0,\qquad Q_kr/\delta\longrightarrow0.\] For large \(k\) the displayed map sends the \(\delta\)-ball into its \(\delta/2\)-ball and has Lipschitz constant at most \(1/4\). It has a unique fixed point in the larger ball, with \(\|v\|_\infty\le 2CQ_kr\). The exact one-step equations and the matching equation make this array one genuine orbit; call its point at zero \(z\). Chart bounds give (33). In the boundary case its distance from each reference orbit is \(o(h_0)\), so it stays inside \(M\). Time-zero coordinates and wall margins. For the linear solution with zero step and terminal data, the two initial displacement spaces are the graphs in (36) and (37). In the reference product coordinates they differ from the slow and fast coordinate spaces by at most \(e^{Cbn}(e^{-cgL}+e^{-2gn+Cbn})\). Solving their matching equation therefore differs from ideal exchange of the fast and slow coordinates by at most \[ e^{Cbn}\bigl(e^{-cgL}+e^{-2gn+Cbn}\bigr)r=o(r). \tag{40}\] All basis conversion and matching inverse factors are included in the constant \(C\). The nonlinear solution differs from the linear one by at most \(CQ_k(2CQ_kr)^2=e^{Cgn}r^2\) in ordinary coordinates, and still by \(e^{Cgn}r^2=o(r)\) in product coordinates. This proves (34). The proof for \(w\) is identical with the two inputs exchanged. Increase \(k\) so that the error in each product coordinate is less than \(\tau r/4\). The ideal exchanged point has distance at least \(\tau r\) from every product face because both inputs do. Thus \(z,w\) have coordinate margins at least \(3\tau r/4\) in their original cell. Similarly, (33) in the fixed endpoint charts is less than \(\tau h_0/4\), so each corresponding endpoint retains margin at least \(3\tau h_0/4\) in its original coarse bin. These strict margins will recover all labels from the outputs. The point obtained does not depend on the auxiliary internal charts or on the frame witnesses used to justify the estimates. Two such constructions have the same endpoint equations and sharp radius \(e^{Cgn}r=o(\delta)\). In either one’s charts both are therefore in its uniqueness ball, and hence coincide. ◻ Relative minors and the joint determinantWe have constructed complementary orbit joins. A single join need not have a convenient Jacobian, so we now keep both of them. Their endpoint coordinates are permuted, which will cancel the large expansion and contraction factors. For the fixed labels of a retained pair define \[ H_{ab}(u)=(F_a(u),P_b(u)),\qquad a,b\in\{x,y\}. \tag{41}\] Each target has dimension \(p_Q+(d-p_Q)=d\) and uses the fixed orthonormal coordinates in the projection ranges. Lemma 9. All four derivatives \(DH_{xx}(x)\), \(DH_{yy}(y)\), \(DH_{xy}(z)\), and \(DH_{yx}(w)\) are nonsingular. Locally the equations (32) define a smooth pair map \(\Phi(x,y)=(z,w)\) with \[ J_{\nu\otimes\nu}\Phi(x,y)=1+o(1), \tag{42}\] uniformly on the retained pairs. The conclusion includes empty slow or fast groups. Proof. Put \(p=p_Q\). For the future map at \(x\), use the actual fast and slow spaces first, with orthonormal bases within each. The fast \(p\)-minor has absolute value at least \(e^{\Lambda_x n-Cbn}\), where \(\Lambda_x\) is the sum of its fast rates. This follows from the fast volume lower bound in 4 and from the uniformly invertible restriction of the endpoint projection. Every other maximal minor contains a slow starting vector and is at most \(e^{(\Lambda_x-2g)n+Cbn}\) when both groups occur. Replacing the actual bases with the common reference basis changes their spaces by \(e^{-cgL}\). After a change of orthonormal basis within each space, multilinearity and the full exterior norm bound show that the leading minor stays nonzero and that all other maximal minors are \(o(1)\) times its absolute value. Factors from the inverse of the possibly oblique starting basis are at most \(e^{Cbn}\). Call the resulting leading absolute minor \(A_x\). More explicitly, when \(0<p<d\), every nonleading maximal future minor \(\Delta\) in the common reference basis satisfies \[\frac{|\Delta|}{A_x} \le C\left(e^{-2gn+Cbn}+e^{-cgL+Cbn}\right)=o(1).\] The first term comes from the mixed-exterior estimate. The second comes from the change of spaces, using the full exterior upper bound and the leading-minor lower bound. Increasing \(C\) includes the oblique-basis factors. If \(p=0\) or \(p=d\), there are no nonleading maximal future minors, so this comparison is unnecessary. Define \(A_y\) in the same way for the future of \(y\). For the past maps the leading columns are slow, and the backward exterior rate is \[\Omega_x=-\sum_{\text{slow at }x}\lambda_j(x).\] Denote their leading absolute minors by \(B_x,B_y\). They have lower bounds \(e^{\Omega_x n-Cbn}\) and \(e^{\Omega_y n-Cbn}\), respectively, and all other maximal minors are relatively negligible. No sign is prescribed for \(\Lambda_x\) or \(\Omega_x\) in a perturbation estimate. We must transfer the future minors from \(x\) to \(z\) relative to \(A_x\), even if this leading quantity is small. Use the reference charts along the future of \(x\) for both nearby orbits. By (33), their one-step derivatives differ by at most \(C\delta\) for large \(k\). On every unperturbed subinterval the \(p\)-th exterior product has norm at most \(e^{\Lambda_x\ell+Cbn}\). Applying 6 gives, for every fixed \(C_0\), \[ \left\|\bigwedge^pD(S^n)_z- \bigwedge^pD(S^n)_x\right\| =o\bigl(e^{\Lambda_xn-C_0bn}\bigr), \tag{43}\] where the derivatives are represented in these common starting and endpoint coordinates. To see the scale in this application, an expansion term with \(q\ge1\) altered factors is bounded by \[e^{\Lambda_x(n-q)+(q+1)C_1bn}(C\delta)^q.\] The binomial sum, divided by \(e^{\Lambda_xn-C_0bn}\), tends to zero because \(n\delta e^{Cbn}\to0\). The bounded factor \(e^{q|\Lambda_x|}\) covers either sign of the rate. The same reasoning backward uses \(\Omega_y\) and \(e^{q|\Omega_y|}\). Fixed endpoint projections and the common starting basis preserve (43), after increasing \(C_0\) if an inverse basis factor is needed. Hence at \(z\) the leading future and past minors are \(A_x(1+o(1))\) and \(B_y(1+o(1))\), with the same relative smallness of all the other minors. At \(w\) they are \(A_y\) and \(B_x\). Expand each full determinant along its \(p\) future rows. There is one leading product, from fast future columns and slow past columns. Every other term has a relatively negligible maximal minor. In the common starting basis the four absolute determinants are therefore \[ \begin{array}{c|cccc} \text{map and point}&H_{xx}(x)&H_{yy}(y)&H_{xy}(z)&H_{yx}(w)\\ \hline \text{absolute determinant}& A_xB_x(1+o(1))&A_yB_y(1+o(1))& A_xB_y(1+o(1))&A_yB_x(1+o(1)). \end{array} \tag{44}\] All leading factors are positive. In particular these maps are locally invertible. If \(p=0\), future targets and minors are empty and \(A_x=A_y=1\); if \(p=d\), the past minors are \(B_x=B_y=1\). The same determinants and conclusions apply, without mixed minors of the missing kind. The defining equations say that \((H_{xy}(z),H_{yx}(w))\) is a fixed permutation of \((H_{xx}(x),H_{yy}(y))\). Differentiation gives the exact identity \[ |\det D\Phi(x,y)|= \frac{|\det DH_{xx}(x)|\,|\det DH_{yy}(y)|} {|\det DH_{xy}(z)|\,|\det DH_{yx}(w)|}. \tag{45}\] Each full derivative represented in the common product basis differs from its coordinate derivative by the factor \(\det E_Q\). There are two such factors in the numerator and two in the denominator, so they cancel exactly. Substituting (44) cancels \(A_x,B_x,A_y,B_y\) individually, with no comparison between the rates of \(x\) and \(y\). Thus the coordinate Jacobian is \(1+o(1)\). All four starting points lie in the same cell of diameter \(O(r)\). The ratio of the two output smooth densities to the two input densities is therefore \(1+o(1)\) as well. This proves (42). ◻ Measurability, recovered labels, and global injectivityThe Jacobian calculation is local. To convert it into the desired measure inequality we must ensure that no multiplicity arises from different choices of bins or local branches. Lemma 10. The map \(\Phi:G_k\to M\times M\), \(\Phi(x,y)=(z,w)\), is measurable and injective. Its restriction to countably many measurable pieces agrees with smooth local diffeomorphisms having the Jacobian in (42). Proof. First fix a retained pair. Its cell and four endpoint labels are locally constant among retained inputs near it, because of the positive wall margins. By 9, the implicit equations define a smooth local pair branch near its outputs. This branch agrees with the constructed switch on nearby retained inputs. Indeed, for this fixed \(k\), continuity of the finitely many maps \(S^t\), \(|t|\le n\), makes their reference halves arbitrarily close when their inputs are close. The constructed solutions have sharp radius \(e^{Cgn}r=o(\delta)\) about those halves. Both the old local implicit branch and the construction for the nearby pair therefore lie in the old \(\delta\)-tube, and satisfy the same fixed endpoint equations with the nearby right-hand sides. The contraction uniqueness argument identifies them even for these changed data: two solutions \(u,v\) in the old ball satisfy \(\mathcal L(v-u)=\mathcal R(v)-\mathcal R(u)\), so \(\|v-u\|_\infty\le CQ_k\delta\|v-u\|_\infty\) forces equality. Both candidate arrays here use the old reference halves. Their matching datum is therefore the same old coordinate difference; only the prescribed endpoint values change with the nearby inputs. The charts along the old pair may be kept fixed in this neighborhood. In particular no regularity of a choice of good frame witnesses is needed. The retained set is Borel, and \(M\times M\) is second countable. A countable subcollection of these relative neighborhoods covers it. Subtracting earlier members partitions it into countably many Borel pieces. On each piece \(\Phi\) is the restriction of the indicated smooth local diffeomorphism; this proves measurability and supplies the pieces for change of variables. For injectivity suppose \(\Phi(x,y)=\Phi(x',y')=(z,w)\). The interiors of product cells are disjoint, and both outputs stayed in the original cell. Thus the two input pairs have the same product cell and hence the same starting bin \(Q\), threshold, and dimensions. The four endpoint-bin labels are recovered from the ordered output pair by \[ (U_x,V_x,U_y,V_y)= \bigl(\operatorname{bin}(S^nz),\operatorname{bin}(S^{-n}w), \operatorname{bin}(S^nw),\operatorname{bin}(S^{-n}z)\bigr). \tag{46}\] All four endpoints have positive margins in their bins. Therefore these labels, and exactly the four projections and coordinate maps in (31), are the same for the competing inputs. On the future half, both \(x\) and \(x'\) are within \(e^{Cgn}r\) of \(z\); on the past half both are within that distance of \(w\). Consequently their full orbits on \([-n,n]\) are within \(2e^{Cgn}r=o(\delta)\) of each other. Moreover the recovered equations give \[F_x(x')=F_x(z)=F_x(x),\qquad P_x(x')=P_x(w)=P_x(x).\] Apply the boundary uniqueness argument of 8 with both reference halves equal to the true good orbit of \(x\). It uses \(x\)’s endpoint projections, zero matching datum, and zero terminal data. The orbit of \(x\) is one solution. In the same charts the orbit of \(x'\) is another solution in its \(\delta\)-ball, since its distance is \(o(\delta)\) even after fixed chart conversions. Uniqueness gives \(x'=x\). The same argument gives \(y'=y\). Only the true input orbits were required to be good; neither \(z\) nor \(w\) was required to belong to \(\mathcal A_k\) or to \(R_k\). ◻ Proof of 7. The construction above applies for all sufficiently large \(k\). Equation (26) gives \(\sum_kB_k^{16}=\sum_kb_k^2<\infty\); any finitely many unused terms can be chosen positively. The partition bound follows from (28), retention from (30), and uniform shadowing from (33). It remains to prove the output measure bound. On the square of a positive-volume cell \(Q'\), the original law \(\rho_k\) has density \(1/\nu(Q')\) with respect to \(\nu\otimes\nu\). The map \(\Phi\) sends retained pairs in that square into the same square. By 9, its inverse volume Jacobian is at most \(2\) for all sufficiently large \(k\). Apply change of variables on the countably many pieces of 10. Their images are disjoint by global injectivity. Summing therefore gives \[\Phi_*(\rho_k|_{G_k})\le 2\rho_k.\] The factor \(1/\nu(Q')\) is unchanged. Restriction to endpoint-label pieces or to \(G_k\) is never followed by normalization, and zero-volume cells contribute nothing. Taking first marginals proves (24), because the first marginal of \(\rho_k\) is \(\nu\). The construction already handled the optional smooth boundary. ◻ Words with prescribed relations and recognizable boundariesWe now construct finite families of words that will define the symbolic system. At each level, short lists of constituent labels have large entropy, although specified pairs of labels satisfy a bilinear relation. We also require that a word can be recognized, together with its alignment, from a sufficiently accurate copy. The construction in this section is entirely finite at each step; no invariant probability is chosen yet. Nested-word constructions and probabilistic control of frequencies and shifted comparisons have precedents in (Foreman et al. 2011, sec. 4.1 and Proposition 43). Related word and phase frameworks appear in (Foreman and Weiss 2022, secs. 4.2.3–4.2.4). Our graph constraints and the per-word condition below are proved separately; in particular, we do not require every lower-level word to occur in every parent or import unique ergodicity from those constructions. All positions in a finite word are numbered from zero. For a word \(u\), write \(u[t,t+q)\) for its consecutive subword of length \(q\) starting at \(t\). For two words of the same length, \(d_{\mathrm H}(u,v)\) denotes the number of coordinates at which they differ. Thus this distance counts symbols or constituent labels according to the alphabet being used. For probabilities \(p,q\) on a finite set, we use \[\mathop{\mathrm{TV}}(p,q)=\frac12\sum_a |p(a)-q(a)|.\] Write \(\mathcal L(Y)\) for the distribution of a random variable \(Y\). Parameters and the four word propertiesEach new word will concatenate words from the preceding family. We control two sampling experiments: choosing a word uniformly from its family, and fixing a parent at the next level and choosing one of its constituents uniformly. We compare the laws of windows at the same internal offset in these two chosen words. Requiring this approximation for every parent will allow arbitrary parent weights in the eventual invariant measure. The construction first samples constituent lists satisfying the relations inside independent groups. It then keeps lists satisfying the per-parent approximation, and finally selects a family with the required window distributions and distances. The parameters below provide enough groups and candidate words for these selections. Fix \[ \varepsilon_i=2^{-i-2},\qquad L_i=2^{2s_i}-1,\qquad \ell_i=\log L_i,\qquad K_i=\lfloor\sqrt{\ell_i}\rfloor, \tag{47}\] and put \[ N_i=2^{K_i},\qquad H_i=2^{\lfloor K_i/2\rfloor},\qquad D_i^*=2^{\lfloor K_i/8\rfloor},\qquad m_1=1,\qquad m_{i+1}=m_iN_i. \tag{48}\] Starting from the integer \(s_1\), define \(2s_{i+1}\) to be the greatest even integer not exceeding \[ X_i:=\frac{\varepsilon_i}{100}N_i\ell_i. \tag{49}\] The proof below shows that this recursion is well defined at every level if \(K_1\ge256\). There is no need to optimize this initial size. Let \(\mathcal A_i=\mathbb F_2^{2s_i}\setminus\{0\}\). On the full vector space use the nondegenerate alternating form \[[u,v]_i=\sum_{a=1}^{s_i} \bigl(u_a v_{s_i+a}+u_{s_i+a}v_a\bigr)\in\mathbb F_2.\] The fixed symbol alphabet is \(\mathcal A=\mathcal A_1\). Each family \(\mathcal W_i\subset\mathcal A^{m_i}\) will have \(L_i\) members. Once it has been constructed, choose a bijection with \(\mathcal A_i\); this equips its word labels with the form \([\cdot,\cdot]_i\). In particular, a word label and its vector label are identified only through this fixed bijection. Whenever \(w\) is a concatenation of \(N_i\) members of \(\mathcal W_i\), let \[c_i(w)=(u_0,\ldots,u_{N_i-1})\in\mathcal W_i^{N_i}\] be its list of constituents. The list is unambiguous because all members of \(\mathcal W_i\) have the same length and are distinct. Here \(m_i\) is the symbol length, \(L_i\) counts distinct level-\(i\) words, and \(N_i\) counts constituent occurrences in a level-\((i+1)\) word, with repetitions allowed. In the properties below, \(H_i\) is the group length and \(D_i^*\) bounds the window lengths and prescribed power-of-two lags; these lengths and lags are measured in level-\(i\) constituent positions. We fix the approximation tolerance \[ \vartheta=10^{-4}. \tag{50}\] Proposition 11 (Finite word families). Choose \(s_1\) so that \(K_1\ge256\), and define the parameters by (47)–(49). There are families \(\mathcal W_i\subset\mathcal A^{m_i}\) of cardinality \(L_i\), with \(\mathcal W_1=\mathcal A\), such that every \(w\in\mathcal W_{i+1}\) is a concatenation of \(N_i\) members of \(\mathcal W_i\) and the following properties hold.
Property ([word:frequency]) compares the two sampling experiments described above. The following elementary estimate controls empirical laws both within each parent and across the selected family. Lemma 12 (Averages of probability vectors). Let \(Y_1,\ldots,Y_q\) be independent random probability vectors on a set of cardinality \(A\), all with mean \(p\). Then \[\mathbb E\,\mathop{\mathrm{TV}}\left(\frac1q\sum_{b=1}^qY_b,p\right) \le\frac12\sqrt{\frac Aq}.\] Coordinates within an individual vector need not be independent. Proof. For \(\widehat p=q^{-1}\sum_bY_b\), independence between vectors gives \[\sum_a\operatorname{Var}(\widehat p(a)) =\frac1{q^2}\sum_{b,a}\operatorname{Var}(Y_b(a)) \le\frac1{q^2}\sum_b\mathbb E\sum_aY_b(a)^2 \le\frac1q.\] Apply the inequality \(\mathbb E|Z|\le(\mathbb E|Z|^2)^{1/2}\) coordinatewise and then Cauchy–Schwarz to the sum over the \(A\) coordinates. ◻ Proof of 11. Growth of the fixed parameters. Whenever \(X_i\ge6\), rounding in (49) gives \[ X_i-3<\ell_{i+1}\le X_i,\qquad \ell_{i+1}\ge\frac{X_i}{2} =\frac{\varepsilon_i}{200}2^{K_i}\ell_i. \tag{52}\] Indeed, \(X_i-2<2s_{i+1}\le X_i\), and \(2s_{i+1}-1<\log(2^{2s_{i+1}}-1)<2s_{i+1}\) when \(2s_{i+1}\ge2\). We claim inductively that \[ K_i\ge256,\qquad \varepsilon_iK_i\ge32\,2^{i-1},\qquad K_{i+1}\ge2^{K_i/3}\ge K_i^2\ge4K_i. \tag{53}\] The first two inequalities hold at \(i=1\). If they hold at level \(i\), then \(\ell_i\ge K_i^2\) implies \[X_i\ge\frac8{25}K_i2^{K_i}\ge6, \qquad \ell_{i+1}\ge\frac4{25}K_i2^{K_i}\ge2^{K_i}.\] Consequently \(K_{i+1}\ge2^{K_i/2}-1\ge2^{K_i/3}\ge K_i^2\ge4K_i\). These elementary inequalities hold for every \(K_i\ge256\); for example, \(K/3-2\log K\) is increasing there and is positive at \(256\). Since \(\varepsilon_{i+1}=\varepsilon_i/2\), the second bound in (53) propagates as well. In particular, all parameters are positive and defined, \(H_i\) divides \(N_i\), and \(D_i^*\le H_i\). Sampling the relations. Suppose that \(\mathcal W_i\) and its vector labeling have been fixed. Generate a list \((U_0,\ldots,U_{N_i-1})\) in independent consecutive groups of \(H_i\) positions. Within each group, proceed from left to right. Choose each new label uniformly among the nonzero vectors orthogonal to all earlier labels at the lags \(2^k\), \(0\le k\le\lfloor K_i/8\rfloor\), that stay inside the group. There are at most \(\lfloor K_i/8\rfloor+1\le K_i\) linear equations. Their common kernel has codimension at most \(K_i\), whether or not the equations are independent. As \(2s_i\ge\ell_i\ge K_i^2\), the number of nonzero choices is at least \[ 2^{2s_i-K_i}-1 \ge2^{2s_i-K_i-1} \ge\frac{L_i}{2^{2K_i}}. \tag{54}\] Denote this unconditioned law on \(\mathcal W_i^{N_i}\) by \(P_i^0\). Every sampled list satisfies the relations in Property ([word:graph]). For later use, set \(\zeta_i=2^{2K_i}/L_i\) within this proof. The conditional probability of any specified label at any position, given all preceding positions, is at most \(\zeta_i\). If \(r_1<\cdots<r_{d'}\) are specified positions, conditioning first at \(r_{d'}\) and then successively at the earlier specified positions gives \[ P_i^0\{U_{r_j}=v_j\text{ for all }j\}\le\zeta_i^{d'}. \tag{55}\] All unspecified labels are integrated out. This argument does not require independence within a group. Each individual \(U_r\) nevertheless has a uniform marginal on \(\mathcal W_i\). To see this, the group of linear maps preserving the alternating form acts transitively on its nonzero vectors. Given such a vector \(v\), nondegeneracy provides \(w\) with \([v,w]_i=1\); the plane spanned by \(v,w\) is nondegenerate, and its orthogonal complement carries a nondegenerate alternating form of dimension two less. Induction extends \(v,w\) to a symplectic basis, proving transitivity by mapping one such basis to another. The sequential rule is equivariant under these maps, including its uniform choices, so each marginal is invariant under a transitive action and hence uniform. Imposing the approximation in each word. For \(i\ge2\) and an offset \(0\le t\le N_{i-1}-D_{i-1}^*\), let \[f_t(u)=c_{i-1}(u)[t,t+D_{i-1}^*),\qquad u\in\mathcal W_i,\] and let \(p_t\) be its law under uniform choice of \(u\). For each of the \(q_i=N_i/H_i\) independently sampled groups, form the empirical probability vector of the \(H_i\) values \(f_t(U_r)\) in that group. Its mean is \(p_t\) by the uniform one-label marginals just proved. The average of these independent vectors is exactly the law of \(f_t(U_J)\) when \(J\) is uniform over the entire sampled list. The alphabet of \(f_t\) has cardinality at most \(A_i=L_{i-1}^{D_{i-1}^*}\). Thus 12, Markov’s inequality, and a union over at most \(N_{i-1}\) offsets bound the probability of violating any requirement in Property ([word:frequency]) by \[ F_i:=\vartheta^{-1}N_{i-1} \sqrt{\frac{L_{i-1}^{D_{i-1}^*}}{N_i/H_i}}. \tag{56}\] Below we verify \(F_i<1/2\) at every level \(i\ge2\). Let \(\mathcal E_i\) be the event that all these requirements hold, and let \(P_i=P_i^0(\,\cdot\mid\mathcal E_i)\). For \(i=1\), take \(\mathcal E_1\) to be the whole sample space and \(P_1=P_1^0\). In all cases \(P_i^0(\mathcal E_i)\ge1/2\), so (55) yields \[ P_i\{U_{r_j}=v_j\text{ for all }j\}\le2\zeta_i^{d'}. \tag{57}\] The factor two is paid once for the whole conditioning event. No uniform marginal or independence between groups is asserted for \(P_i\). The relations hold on its support, as does Property ([word:frequency]) for each individual list. Selecting the finite family. Put \(M_i=L_{i+1}\) and sample \(M_i\) independent complete lists with law \(P_i\). For each offset \(0\le t\le N_i-D_i^*\), let \(Q_{i,t}\) be the law under \(P_i\) of the length-\(D_i^*\) window at \(t\). It satisfies (51) by (57). The empirical window law of the \(M_i\) samples is an average of independent point-mass probability vectors on an alphabet of size \(L_i^{D_i^*}\). A second application of 12 and Markov’s inequality, now with a union over at most \(N_i\) offsets, bounds the probability of violating any required window approximation by \[ E_i:=\vartheta^{-1}N_i \sqrt{\frac{L_i^{D_i^*}}{L_{i+1}}}. \tag{58}\] We will have \(E_i<1/4\) at every level. It remains to impose the distances, including comparisons in which some of the sampled-word indices coincide. First use the reference law in which the \(N_i\) labels in each distinct indexed list are independent and uniform. Consider one pairwise comparison between distinct indices, or one comparison of a target list with a shifted window from two source lists. In the latter case write \(t\) for its offset, with \(1\le t<N_i\); any of the three indices may coincide. For a specified set of matching positions, form a multigraph whose vertices are variables (sample index, position) and whose edges are the asserted equalities. Every target position occurs once. The source positions also occur once: they use \(t,\ldots,N_i-1\) in the first source and \(0,\ldots,t-1\) in the second, disjoint ranges even if the source indices agree. The degree is therefore at most two, counting repeated edges. There is no loop. In a pairwise comparison the indices are distinct, and in a shifted comparison the source position differs from the target position by \(t\) modulo \(N_i\). A connected component with \(v\) vertices requires \(v-1\) independent equalities and has probability \(L_i^{-(v-1)}\) under the reference law. A path has one fewer edge than vertices; a cycle has equally many, with at least two edges. Thus every component imposes at least half as many independent equalities as edges, including a pair of parallel edges. If \(e\) positions are specified to match, their probability is at most \(L_i^{-e/2}\). Failure of a distance requirement entails at least \(\lceil\varepsilon_iN_i\rceil\) matching positions. Summing over their possible sets, of which there are at most \(2^{N_i}\), gives the bound \[ 2^{N_i}L_i^{-\varepsilon_iN_i/2}. \tag{59}\] By (54), a complete list under \(P_i^0\) has likelihood ratio at most \(2^{2K_iN_i}\) relative to the reference law, and under \(P_i\) at most \(2\,2^{2K_iN_i}\). Each comparison uses at most three distinct sampled lists, so its probability under the actual sampling law is at most \((2\,2^{2K_iN_i})^3\) times (59). Repeated indices do not introduce an additional independent list. There are fewer than \(4N_iM_i^3\) pairwise and shifted comparisons in total. Their union failure probability is consequently at most \[ T_i:=32N_iL_{i+1}^3\, 2^{(6K_i+1)N_i}L_i^{-\varepsilon_iN_i/2}. \tag{60}\] We will verify \(T_i<1/4\). These estimates address separate requirements: the first sampling law enforces the relations, conditioning enforces the approximation in each individual word, and the final independent samples provide the window laws and distances. We now check that the fixed initial size makes all three probability bounds hold at every step. Uniform bounds from one initial choice. For \(K\ge256\), set \[B_0(K)=14+K+\tfrac12\,2^{K/8}(K+1)^2.\] The elementary inequalities \(14+K\le2^{K/8}\) and \((K+1)^2\le2^{K/8}\) hold at \(256\) and persist thereafter, as the ratios of their left sides to their right sides decrease. Hence \[ B_0(K)\le2^{K/4}\le\tfrac18\,2^{K/3}. \tag{61}\] We also have \(\log(\vartheta^{-1})<14\) and \(K_i^2\le\ell_i<(K_i+1)^2\). For the first probability bound, put \(K=K_{i-1}\) and use \(\log(N_i/H_i)\ge K_i/2\) to obtain \[\log F_i \le B_0(K)-\frac{K_i}{4} \le-\frac18\,2^{K/3}.\] For the second, put \(K=K_i\). Since \(\ell_{i+1}\ge2^K\), \[\log E_i \le B_0(K)-\frac{\ell_{i+1}}2 \le2^{K/4}-\frac{2^K}{2} \le-\frac{2^K}{4}.\] Finally (52) and \(\varepsilon_i\ell_i\ge\varepsilon_iK_i^2\ge32K_i\) give \[\begin{align*} \log T_i &\le5+K_i+N_i \left(6K_i+1-\frac{47}{100}\varepsilon_i\ell_i\right)\\ &\le5+K_i-9K_i2^{K_i} \le-K_i2^{K_i}. \end{align*}\] It follows that \[ F_i\le2^{-2^{K_{i-1}/3}/8}<\tfrac12\quad(i\ge2),\qquad E_i\le2^{-2^{K_i}/4}<\tfrac14,\qquad T_i\le2^{-K_i2^{K_i}}<\tfrac14. \tag{62}\] All right sides are decreasing functions of their \(K\) argument. They apply simultaneously at every level by (53); no later adjustment of \(s_1\) is made. The conditioning event thus has positive probability. Under the conditioned law the two subsequent failures have total probability less than \(1/2\), so a tuple satisfying every requirement exists. Its pairwise distance requirement is imposed on distinct sample indices and has a positive lower bound. Equal sampled words would have distance zero, so the tuple has exactly \(L_{i+1}\) distinct members. Flatten their constituent lists to form \(\mathcal W_{i+1}\). The empirical distribution is now exactly the uniform distribution on this set, without deduplication or a change of weights. It therefore has all four asserted properties. Choose its bijection with \(\mathcal A_{i+1}\) and continue the induction. ◻ Recognition at every symbol offsetThe distances in 11 concern constituent labels. The next consequence converts them into distances in the fixed alphabet \(\mathcal A\), including shifts that do not respect any lower-level boundary. In coding terminology the two conclusions below are a positive minimum distance and a positive comma-free index: the latter measures distance from every nontrivial splice of two legal words. This is the synchronization notion described in (Tonchev 2003, sec. 1); the recursive lower bound here is proved directly. Lemma 13 (Recognition of words and alignment). For the families in 11, put \[d_i=\prod_{j=1}^{i-1}(1-\varepsilon_j),\qquad d_*=\prod_{j=1}^{\infty}(1-\varepsilon_j)>0.\] Distinct \(v,w\in\mathcal W_i\) satisfy \[d_{\mathrm H}(v,w)\ge d_i m_i\ge d_*m_i.\] For all \(v,w,z\in\mathcal W_i\) and every \(1\le t<m_i\), \[d_{\mathrm H}\bigl((vw)[t,t+m_i),z\bigr) \ge d_i m_i\ge d_*m_i.\] These distances count symbols in \(\mathcal A\), and repeated source words are allowed. Proof. The finite products are bounded below by \(1-\sum_{j\ge1}\varepsilon_j=3/4\), proving that \(d_*>0\). At level one, distinct one-symbol words have distance one and the offset assertion is vacuous. Suppose both assertions hold at level \(i\). Distinct level-\((i+1)\) words differ in at least \((1-\varepsilon_i)N_i\) constituent labels by 11, Property ([word:distance]). Each unequal constituent pair contributes at least \(d_i m_i\) symbol mismatches, giving the bound \(d_{i+1}m_{i+1}\). Now consider a nonzero offset \(t<m_{i+1}\) in two concatenated level-\((i+1)\) words, compared with a third such word. If \(m_i\) divides \(t\), the offset \(t/m_i\) lies between \(1\) and \(N_i-1\). The constituent-label offset requirement gives at least \((1-\varepsilon_i)N_i\) unequal pairs, and the preceding distinct-word bound again gives \(d_{i+1}m_{i+1}\) symbol mismatches. If \(m_i\) does not divide \(t\), each of the \(N_i\) aligned target constituents faces a length-\(m_i\) source window with the same nonzero offset modulo \(m_i\). That source window crosses exactly two consecutive legal level-\(i\) words. The inductive offset assertion therefore gives at least \(d_i m_i\) mismatches on each target constituent. It applies also when the two source constituents lie on opposite sides of the join of the two larger source words. Adding over the disjoint target constituents gives \(d_i m_{i+1}\ge d_{i+1}m_{i+1}\), completing the induction. ◻ The number and size of the available scalesThe word construction must retain enough label information across levels to compete with the summable smooth-switching costs. The following estimates are the quantitative input for that comparison. Lemma 14 (Word rates). Let \(R_i=\ell_i/m_i\). Every \(m_i\) is a power of two, with \[\log m_i=\sum_{j<i}K_j.\] The rates satisfy \[ R_{i+1}\le\frac{\varepsilon_i}{100}R_i, \qquad R_i\longrightarrow0. \tag{63}\] They also have the lower bound \[ R_i\ge R_1\,200^{-(i-1)} 2^{-(i-1)(i+4)/2} \ge R_1\,2^{-(i-1)(i+20)/2}, \tag{64}\] and \[ \sum_{i=1}^{\infty}K_iR_i^{16}=\infty. \tag{65}\] Proof. The formula for \(m_i\) follows from \(m_1=1\) and \(N_i=2^{K_i}\). Dividing (52) by \(m_{i+1}=N_i m_i\) gives \[\frac{\varepsilon_i}{200}R_i \le R_{i+1}\le\frac{\varepsilon_i}{100}R_i.\] Since \(\varepsilon_i\le1/8\), the upper bound implies \(R_i\le800^{-(i-1)}R_1\to0\), proving (63). Iterating the lower bound and using \(\sum_{j<i}(j+2)=(i-1)(i+4)/2\) prove the first inequality in (64); the second uses \(\log200<8\). Also (53) implies \(K_i\ge K_1^{2^{i-1}}\). Hence \[\log(K_iR_i^{16}) \ge 2^{i-1}\log K_1+16\log R_1-8(i-1)(i+20) \longrightarrow\infty.\] In particular the terms in (65) do not tend to zero, and the series diverges. ◻ The families are now fixed. Their per-word approximation transfers short-window entropy to any invariant measure on the resulting sequence space, while 13 recovers block labels and phases from small symbol errors. We next construct that space and prove the required measure-theoretic statements. An invariant process and its window lawsWe now turn the finite word families of 11 into one ergodic system. The probability on that system need not assign uniform frequencies to the different words. The per-word frequency condition in the construction is what will make the window estimates valid for every invariant probability. The compact phase extensionLet \(\mathcal A=\mathcal W_1\), regarded as an alphabet, and let \[\Omega=\mathcal A^{\mathbb Z}\times\prod_{i\ge1}\mathbb Z/m_i\mathbb Z\] with its compact product topology. Write a point as \((\omega,a)\), where \(\omega=(\omega_t)_{t\in\mathbb Z}\) and \(a=(a_i)_{i\ge1}\). Define \(X\subset\Omega\) by the following conditions:
The coordinate \(a_i\in\{0,\ldots,m_i-1\}\) records the index of the origin in its level-\(i\) word. Let \[T(\omega,a)=(\omega',a'),\qquad \omega'_t=\omega_{t+1},\quad a'_i=a_i+1\pmod{m_i}.\] The defining conditions are closed, and \(T\) is a homeomorphism of \(X\). For each \(j\), repeat a word of \(\mathcal W_j\) bi-infinitely with its level-\(j\) boundary at zero. Since higher-level words are concatenations of lower-level ones, this sequence with zero phases satisfies all conditions through level \(j\). The finite-level closed sets are nested and nonempty. Compactness proves that \(X\) is nonempty. The invariant-measure facts used here are the standard compact-system ones; see (Dajani and Dirksin 2008, Theorems 6.1.4–6.1.6). We give the construction together with the checks specific to the phase extension. Proposition 15. There is a \(T\)-invariant ergodic Borel probability \(\mu\) on \(X\). Every \(T\)-invariant Borel probability on \(X\) is nonatomic, is standard modulo null sets, and has zero Kolmogorov–Sinai entropy. Proof. For any \(x\in X\), weak limits of \(n^{-1}\sum_{j=0}^{n-1}\delta_{T^jx}\) are invariant probabilities. Indeed, the difference between the integrals of \(f\circ T\) and \(f\) against these averages tends to zero for each continuous \(f\). The set \(\mathcal M_T(X)\) of invariant probabilities is nonempty, compact, and convex. Here is an explicit way to select an extreme point. Choose a countable uniformly dense family \((f_j)\) in the continuous real functions on \(X\). Starting from \(\mathcal M_T(X)\), successively restrict to the compact face on which \(\int f_j\,d\mu\) is maximal. The nested intersection is nonempty and contains only one measure, because the integrals of every \(f_j\) are then fixed and determine a probability. An intersection of faces is a face, so this measure is extreme. If it had a \(T\)-invariant measurable set of mass strictly between zero and one, normalized restriction to that set and its complement would express it as a nontrivial convex combination of invariant probabilities. Thus it is ergodic. The same argument applies to invariant sets modulo null sets after taking Borel representatives. For any invariant probability, \(a_i\) is uniform on \(\mathbb Z/m_i\mathbb Z\), because \(T\) cyclically permutes its level sets. If a measurable atom had mass \(c>0\), then for each \(i\) it would be contained, modulo null sets, in one of these level sets. Hence \(c\le1/m_i\) for every \(i\), which is impossible since \(m_i\to\infty\). The space \(X\) is a compact metrizable subspace of a countable product of finite discrete sets. For completeness, its nonatomic probability is standard in the precise sense needed here. Encode the finite coordinates into successive disjoint blocks of ternary digits \(0,2\); this gives a continuous injective map of \(X\) into a compact subset of \([0,1]\). Let \(\lambda\) be the image law and \(F(t)=\lambda([0,t])\) its continuous distribution function. Define the measurable quantile \[Q(u)=\inf\{t:F(t)\ge u\},\qquad 0<u<1.\] Continuity gives \(F(Q(u))=u\), and \(\{u:Q(u)\le t\}\) has Lebesgue measure \(F(t)\), so \(Q_*\operatorname{Leb}=\lambda\) and \(F_*\lambda=\operatorname{Leb}\). Moreover \(Q(F(x))\le x\) for \(\lambda\)-almost every \(x\); both sides have law \(\lambda\). Their nonnegative difference has zero integral, and therefore \(Q(F(x))=x\) almost surely. After restricting to measurable conull sets, \(F\) and \(Q\) are inverse isomorphisms. Composing with the coordinate embedding establishes the assertion. Passing to completed sigma fields does not change it. It remains to check the entropy of the entire phase extension. A partition depending on finitely many coordinates is refined by one that records \(\omega_{-r},\ldots,\omega_r\) and \(a_I\) for some fixed \(r,I\). Compatibility determines all the lower phases from \(a_I\), and the phase evolves deterministically. Fix any \(j\ge I\). The initial phase \(a_j\) determines \(a_I\) and the boundaries of all level-\(j\) words. The symbol interval \([-r,n-1+r]\) used in an \(n\)-step name meets at most \(\lceil(n+2r)/m_j\rceil+1\) such words, each in \(\mathcal W_j\). Consequently the number of names is at most \[m_j L_j^{\lceil(n+2r)/m_j\rceil+1}.\] Keeping \(j\) fixed and sending \(n\to\infty\) bounds the entropy rate by \(R_j=(\log L_j)/m_j\). By 14, \(R_j\to0\). Only now let \(j\to\infty\): every finite-coordinate partition has entropy rate zero. This count uses no assumption about word frequencies. Let \(\mathcal P\) be any finite measurable partition with \(A\) atoms. Cylinder sets generate the Borel sigma field, and their algebra is dense in measure. Thus the atom label of \(\mathcal P\) can be approximated by an \(A\)-valued finite-coordinate label \(\mathcal Q\) with mismatch probability \(s\) arbitrarily small. Completion of the measure only adds null-set changes. The chain rule and conditioning give \[H(\mathcal P^n) \le H(\mathcal Q^n) +\sum_{t=0}^{n-1}H(T^{-t}\mathcal P\mid T^{-t}\mathcal Q) =H(\mathcal Q^n)+nH(\mathcal P\mid\mathcal Q),\] where \(\mathcal P^n=\bigvee_{t=0}^{n-1}T^{-t}\mathcal P\). Recording first whether the two labels differ shows that \[H(\mathcal P\mid\mathcal Q)\le H_{\rm bin}(s)+s\log A.\] Entropy rates exist by subadditivity. Divide by \(n\), take the limit, and then let \(s\downarrow0\). This proves \(h_\mu(T)=0\). ◻ Fix henceforth an ergodic probability supplied by 15. The next estimates use only invariance, so their validity is unaffected by this choice. What alignment does to the sampling lawFor \(x=(\omega,a)\in X\), let \(V_i(x)\in\mathcal W_i\) be the word containing the origin, namely the symbol string beginning at \(-a_i(x)\). Set \[A_i=\{x:a_i(x)=0\},\qquad \mu_i=m_i\,\mu|_{A_i}.\] The law \(\mu_i\) is a probability. It describes a point observed at a level-\(i\) boundary; it is not asserted to give uniform word frequencies. For an integer \(t\), let \[r_i(x,t)=(-a_i(T^t x))\bmod m_i\in\{0,\ldots,m_i-1\}.\] The first level-\(i\) boundary at or after \(t\) is \(t+r_i(x,t)\). For an integer \(D\ge2\), define the vector \[ \mathcal L_{i,D}(x,t)= \bigl(V_i(T^{\,t+r_i(x,t)+jm_i}x)\bigr)_{j=0}^{D-2}. \tag{66}\] All its words lie in \([t,t+Dm_i)\), since \(r_i+(D-1)m_i\le Dm_i\). Lemma 16. The law of (66) under \(\mu\) is the law under \(\mu_i\) of the first \(D-1\) level-\(i\) words beginning at zero. Under \(\mu_i\), the index \(J=a_{i+2}/m_i\) is uniform on \(\{0,\ldots,N_iN_{i+1}-1\}\) and is independent of \(V_{i+2}\). More precisely, for each such \(j\) and \(v\in\mathcal W_{i+2}\), \[ \mu_i(J=j,V_{i+2}=v) =\frac{\mu_{i+2}(V_{i+2}=v)}{N_iN_{i+1}}. \tag{67}\] Proof. Invariance first allows us to take \(t=0\). On the phase class \(a_i=r\), the map \(T^{(-r)\bmod m_i}\) sends the normalized restriction of \(\mu\) bijectively and measure-preservingly onto \(\mu_i\). The first assertion follows by applying this separately on all \(m_i\) phase classes. For the second assertion, \(T^{jm_i}\) maps \(\{a_{i+2}=0,V_{i+2}=v\}\) onto \(\{a_{i+2}=jm_i,V_{i+2}=v\}\). It does not change the containing parent word because \(0\le jm_i<m_{i+2}\). These sets have the same \(\mu\)-mass. Multiply by \(m_i\), and use \(m_{i+2}/m_i=N_iN_{i+1}\), to obtain (67). ◻ The identity is the reason that a deterministic condition on every parent word suffices. At a fixed internal offset in a parent, the slot being sampled is uniform even when the parent itself is not uniformly distributed. Entropy of short label listsWe use two elementary entropy bounds. The second is the standard maximal-coupling and mismatch-flag entropy argument; compare (Sason 2013, Theorems 1–3). We include the weaker form needed here. If a law on a finite set has every atom at most \(M\), its entropy is at least \(\log(1/M)\). If two laws \(P,Q\) on an \(A\)-point set have \(\mathop{\mathrm{TV}}(P,Q)\le s\), then \[ H(P)\ge H(Q)-1-s\log A. \tag{68}\] Indeed, couple their labels with mismatch probability at most \(s\). The mismatch flag and the chain rule bound either conditional entropy by \(1+s\log A\). We use this simple bound rather than an optimized continuity estimate. Proposition 17. For all sufficiently large \(i\), for every power of two \(2\le D\le D_i^*\), and every integer \(t\), \[ H_\mu(\mathcal L_{i,D}(\,\cdot\,,t)) \ge .996(D-1)\ell_i,\qquad \ell_i=\log L_i. \tag{69}\] The large-\(i\) threshold is independent of \(D,t\) and of the invariant probability on \(X\). Proof. Write \(q=D-1\). By 16 we work under \(\mu_i\). Decompose its uniform parent-slot index as \[J=hN_i+u,\qquad 0\le h<N_{i+1},\quad 0\le u<N_i.\] Conditional on the parent \(v=V_{i+2}\) and the internal offset \(u\), the slot \(h\) is uniform among the \(N_{i+1}\) constituent \(\mathcal W_{i+1}\) words of \(v\). Call \(u\) admissible if \(u+D_i^*\le N_i\). Its failure probability is \[\beta_i=\frac{D_i^*-1}{N_i}\le\frac{D_i^*}{N_i}.\] For every fixed \(v\) and admissible \(u\), 11, Property ([word:frequency]), applied at level \(i+1\), puts the law of the \(D_i^*\) labels at this offset within \(10^{-4}\) of their law in a uniform word of \(\mathcal W_{i+1}\). 11, Property ([word:window]), applied at level \(i\), puts that law within a further \(10^{-4}\) of a law whose every specified \(q\)-coordinate atom is at most \[M_i=2\left(\frac{2^{2K_i}}{L_i}\right)^q.\] Restriction to the first \(q\) entries cannot increase total variation. The comparison law may depend on \(u\), but this atom bound does not. It follows from (68), with \(s=2\cdot10^{-4}\) and alphabet size \(L_i^q\), that the actual \(q\)-list conditional on \(v,u\) has entropy at least \[q\ell_i-2K_iq-1-(1+sq\ell_i) =q\bigl((1-s)\ell_i-2K_i\bigr)-2.\] Entropy is concave. Mixing over all \(v\) and admissible \(u\) preserves this lower bound; adding the remaining offsets loses at most their fraction, since entropy is nonnegative. Thus \[ H(\mathcal L_{i,D}) \ge(1-\beta_i) \left\{q\bigl((1-s)\ell_i-2K_i\bigr)-2\right\}. \tag{70}\] Here \(\beta_i\to0\), \(2K_i/\ell_i\to0\), and \(2/\ell_i\to0\). Once each of these three quantities is at most \(.0005\), division by \(q\ell_i\), with \(q\ge1\), gives a coefficient at least \[(1-.0005)(1-.0002-.0005-.0005)>.996.\] This proves the uniform assertion. ◻ Only the two indicated levels of the construction enter this argument. There is no accumulation of total variation errors over all ancestors, and no independence assumption on the sequence of adjacent words. The graph relation across two windowsTwo level-\(i\) labels \(u,v\in\mathcal W_i\) pass the graph test when \([u,v]_i=0\), using their fixed vector identification. The same alignment law gives a quantitative bound on the one exception to the graph test: crossing a group boundary. Lemma 18. Let \(2\le D\le D_i^*\) be a power of two and set \(n=m_iD\). For all sufficiently large \(i\), with \(\mu\)-probability at least \[1-\alpha_{i,D},\qquad \alpha_{i,D}=\frac{2D-2}{H_i},\] all corresponding entries of \(\mathcal L_{i,D}(x,-n)\) and \(\mathcal L_{i,D}(x,0)\) pass the level-\(i\) graph test simultaneously. In particular \(\sup_{D\le D_i^*}\alpha_{i,D}\to0\). Proof. Because \(n\) is a multiple of \(m_i\), the advance to alignment is the same in the two windows. Denote it by \(q_0\). The entry starts are \[ -n+q_0+jm_i \quad\hbox{and}\quad q_0+jm_i,\qquad 0\le j\le D-2. \tag{71}\] Corresponding entries are separated by exactly \(D\) level-\(i\) positions. The first start through the last start span \(2D-1\) such positions, including the unused intervening position. 3 illustrates these positions and their common lag when \(D=4\). By 16, the first boundary has the boundary-start law. Its position modulo \(H_i\) is uniform: its position in a \(\mathcal W_{i+1}\) parent is uniform modulo \(N_i\), and \(H_i\) divides \(N_i\). Once \(2D-1\le H_i\), exactly \(2D-2\) residues have a span that crosses an \(H_i\)-group boundary. All parent boundaries are already group boundaries. On every other residue all the labels in (71) belong to one of the prescribed groups. 11, Property ([word:graph]), applies because \(D\) is an allowed power-of-two lag. Finally \(2D_i^*/H_i\to0\) by the definitions. ◻ We now have a fixed ergodic system with zero entropy, uniformly large list entropy, and a graph relation across adjacent windows. These properties are compatible: the information per symbol in a level-\(i\) list is of order \(R_i\), which tends to zero. The many available test scales will compensate for this decreasing rate. The remaining task is to compare that relation with the conditional independence in a smooth switch. The obstruction to a smooth volume modelFix the ergodic system constructed in 6. The final argument concerns a single statistic of a pair of points: the fraction of corresponding past and future labels that pass the graph test. We define that statistic before restricting the pair experiment. Its upper and lower bounds will therefore refer to the same probability law. The quantitative interfaceThe smooth and symbolic constructions meet through the quantities in 1. At \(n=2^k\), write \(\mathcal D_k\) for the smooth partition, \(\rho\) for its cell-weighted pair law, \(G\) for the retained pairs, and \(z\) for the switch. The exceptional-mass tolerance \(\eta>0\) is fixed before these objects are chosen. On the symbolic side, \(L_i=\#\mathcal W_i\), \(R_i=(\log L_i)/m_i\), and \(K_i=\lfloor\sqrt{\log L_i}\rfloor\). Every \(m_i\) is a power of two. The symbolic tests have length \(n=m_iD\), where \(D\) is a power of two and \(2^{\lceil K_i/16\rceil}\le D\le2^{\lfloor K_i/8\rfloor}\). In a window of \(Dm_i\) symbols we use the \(D-1\) full words beginning at the first level-\(i\) boundary. Each input initially uses its own alignment. The compact system recording the symbol sequence and all compatible alignment phases has zero entropy for every invariant probability, by 15.
Choosing times with a negligible partition costLemma 19. Let \((B_k)_{k\ge1}\) be any positive sequence with \(\sum_k B_k^{16}<\infty\). There exist levels \(i\to\infty\) and powers of two \(D\) satisfying \[2^{\lceil K_i/16\rceil}\le D\le2^{\lfloor K_i/8\rfloor}\] such that, for \(n=m_iD\), \[ \frac{B_{\log n}}{R_i}\longrightarrow0, \qquad R_i=\frac{\ell_i}{m_i}. \tag{72}\] In particular any partitions with \(\log\#\mathcal D_k\le C B_k2^k\) obey \[ \log\#\mathcal D_{\log n}=o((D-1)\ell_i) \tag{73}\] on this sequence, for every fixed \(C\). Proof. The lengths \(m_i\) are powers of two. Put \(b_i'=\log m_i\), so \(b_{i+1}'=b_i'+K_i\). The integer sets \[J_i=\{b_i'+j:\lceil K_i/16\rceil\le j\le \lfloor K_i/8\rfloor\}\] are disjoint and, for large \(i\), contain at least \(K_i/32\) integers. If the ratios in (72) were bounded below by some \(c>0\) on all sufficiently late tests, then \[\sum_kB_k^{16} \ge c^{16}\sum_{i\ {\rm large}}|J_i|R_i^{16} \ge \frac{c^{16}}{32} \sum_{i\ {\rm large}}K_iR_i^{16}=\infty,\] contrary to the hypothesis and 14. Thus one can select successively later tests with ratio tending to zero. Since \(B_{\log n}n=(B_{\log n}/R_i)D\ell_i\) and \(D/(D-1)\le2\), (73) follows. ◻ Suppose for a contradiction that a smooth volume model \((M,\nu,S)\) exists. Its dimension is positive because a zero-dimensional compact manifold is finite and cannot carry a nonatomic probability. Apply 7 with \(\eta=.01\). The model is now fixed, so all its constants and its sequence \((B_k)\) are fixed. Use 19 for this sequence. In what follows, \(i,D,n\) run through the selected tests and are taken sufficiently large for every stated eventual estimate. We use the measurable isomorphism to regard every finite symbolic observation as a measurable observation on \(M\). The original domain can be replaced by the intersection of all its integer translates, so all orbit identities hold on an invariant conull set. Extend the observations arbitrarily on its null complement. For each test let \(C\) be the cell sampled in the switching experiment, and let \(\rho\) be the original law of \((x,y)\). Its two marginals are exactly \(\nu\), and its conditional law given \(C\) is the product of normalized volume on that cell. The independent upper boundLet \[U=\mathcal L_{i,D}(y,-n),\qquad V=\mathcal L_{i,D}(x,0),\qquad q=D-1.\] Each list uses its own input’s alignment advance. Both are defined on every pair before a switch is selected. Under the identification of \(\mathcal W_i\) with the nonzero vectors of \(\mathbb F_2^{2s_i}\), write \([\,\cdot\,,\,\cdot\,]_i\) for the alternating form from the word construction, and set \[ Z(x,y)=\frac1q\sum_{j=0}^{q-1} \mathbf 1_{\{[U_j,V_j]_i=0\}}. \tag{74}\] For each fixed cell and each fixed \(j\), the laws of \(U_j\) and \(V_j\) are independent. Apart from the diagonal pairs, the orthogonality relation gives the classical binary symplectic graph; see (Seidel 1973, Definition 3.3 and Theorem 3.4) for its finite-geometric description and spectral scale. We include the character calculation, keeping the diagonal pairs, and pass from average Shannon entropy to the required graph bound by a heavy–light split of each law. Lemma 20. Let \(L=2^{2s}-1\) and equip the nonzero vectors of \(\mathbb F_2^{2s}\) with a nondegenerate alternating form. Suppose a family of pairs of laws \((p_\xi,q_\xi)\) on these vectors is indexed by a probability space, and \[\mathbb E_\xi H(p_\xi)\ge .99\log L,\qquad \mathbb E_\xi H(q_\xi)\ge .99\log L.\] For all sufficiently large \(L\), independent samples from the two laws, followed by averaging over \(\xi\), pass the test \([u,v]=0\) with probability at most \(.62\). Proof. For a law \(p\), put \(A_p=\{u:p(u)>L^{-.8}\}\) and \(\theta_p=p(A_p)\). The set \(A_p\) has at most \(L^{.8}\) elements. Splitting entropy according to this set gives \[ H(p)\le1+.8\theta_p\log L+(1-\theta_p)\log L =1+(1-.2\theta_p)\log L. \tag{75}\] The same formula, with the usual zero-mass conventions, covers empty sets. The entropy hypotheses imply \[\mathbb E_\xi\theta_{p_\xi},\ \mathbb E_\xi\theta_{q_\xi} \le .05+\frac5{\log L}\le .055\] once \(\log L\ge1000\). Let \(A\) be the matrix \(A_{uv}=(-1)^{[u,v]}\) on nonzero vectors. On the full vector space, the rows of the same matrix are orthogonal: for \(u\ne w\), nondegeneracy makes \(v\mapsto[u+w,v]\) a nonzero linear functional, so \(\sum_v(-1)^{[u+w,v]}=0\). Every row has squared norm \(L+1\). The full operator norm is therefore \(\sqrt{L+1}\), and restriction to the nonzero rows and columns has no larger norm. The light subprobability \(p^{\rm l}=p\,\mathbf 1_{A_p^c}\) satisfies \[\|p^{\rm l}\|_2^2 \le L^{-.8}\sum_u p^{\rm l}(u)\le L^{-.8}.\] Hence \[\left|\sum_{u,v}p^{\rm l}(u)A_{uv}q^{\rm l}(v)\right| \le\sqrt{L+1}\,L^{-.8}=:\beta_L.\] The passing mass of the light product is at most \((1+\beta_L)/2\). All pairs with at least one heavy coordinate have total mass at most \(\theta_p+\theta_q\). After averaging, and taking \(L\) large enough that \(\beta_L\le .02\), the passing probability is at most \[.055+.055+\frac{1+.02}{2}=.62.\] The matrix includes its diagonal entries, so the automatic relations \([u,u]=0\) require no separate exclusion. ◻ By the two original marginals and 17, \[H_\rho(U),H_\rho(V)\ge .996q\ell_i.\] Conditioning each vector on the sampled cell loses at most \(H(C)\le\log\#\mathcal D_{\log n}=o(q\ell_i)\). Conditional subadditivity consequently gives \[ \frac1q\sum_{j=0}^{q-1}H(U_j\mid C)\ge .99\ell_i, \qquad \frac1q\sum_{j=0}^{q-1}H(V_j\mid C)\ge .99\ell_i \tag{76}\] for late tests. Apply 20 with index \(\xi=(C,j)\), assigning \(j\) uniform weight independently of the cell. Independence is used conditional on this full index: equivalently, first fix \(j\), use the product law conditional on \(C\), and then average. No product assertion conditional only on \(C\) after mixing over a shared random \(j\) is needed. We obtain \[ \mathbb E_\rho Z\le .62. \tag{77}\] From geometric shadowing to symbol agreementLet \(d_*>0\) be the symbol-distance constant in 13. It is fixed by the word construction. Let \(f:M\to\mathcal A\) be the transported symbol at time zero. We first choose compact cores for this one fixed finite partition, rather than for the growing level-\(i\) label sets. Lemma 21. For all sufficiently large dyadic times in the smooth switching estimate, the retained pairs for which \[\begin{align*} \#\{0\le t<n:f(S^tz)\ne f(S^tx)\} &<d_*n/100,\tag{78}\\ \#\{-n\le t<0:f(S^tz)\ne f(S^ty)\} &<d_*n/100 \tag{79}\end{align*}\] have original \(\rho\)-mass at least \(.96\). Proof. Choose \(\kappa<d_*/20000\). Inner regularity of the smooth probability gives disjoint compact sets \(K_a\subset f^{-1}(a)\), for \(a\in\mathcal A\), with \[\nu\left(M\setminus\bigcup_{a\in\mathcal A}K_a\right)<\kappa.\] Measurable atoms may first be replaced by Borel representatives modulo null sets. There are finitely many cores, and the distance between any two nonempty distinct cores is positive. If there is only one nonempty core, the following equal-symbol implication is automatic. Otherwise take the positive minimum of these distances. Uniform shadowing is smaller than this fixed separation at all large scales. Therefore two compared iterates both in the cores have the same symbol. Write \(G\) for the retained set and \(\sigma=z_*(\rho|_G)\). The switching theorem gives \(\sigma\le2\nu\). For every integer \(t\), \[(S^t)_*\sigma\le2(S^t)_*\nu=2\nu.\] The \(x\) and \(y\) time marginals under the original law are \(\nu\). At each future comparison time, the \(\rho\)-mass within \(G\) where either \(x\) or \(z\) misses the cores is at most \(\kappa+2\kappa=3\kappa\). The same holds for the past comparison with \(y\). The integrals over \(G\) of the fractions of mismatches on either half are therefore at most \(3\kappa\). Markov’s inequality at threshold \(d_*/100\), on each half, bounds their combined failure mass by \(600\kappa/d_*<.03\). Since \(\rho(G)\ge .99\), the simultaneous success mass is at least \(.96\). This is an estimate in the original law, without dividing by \(\rho(G)\). The input marginals and \(\sigma\le2\nu\) give zero mass to the complement of the invariant conjugacy domain. Thus the symbol strings and all orbit identities used in the argument are defined outside a \(\rho\)-null set. ◻ Recognition preserves the original entry pairingConsider a pair in the success event of 21. If the phases of \(z\) and \(x\) modulo \(m_i\) differed, take the \(D-1\) full level-\(i\) blocks of \(x\) beginning at its first aligned boundary in \([0,n)\). At each such interval, the symbol string of \(z\) is a window of length \(m_i\) cut at a nonzero offset across two consecutive legal level-\(i\) words. This remains true if those words are equal or if one extends beyond the observed window. By 13, each of the \(D-1\) disjoint intervals contributes at least \(d_*m_i\) mismatches. Their total is at least \[d_*(D-1)m_i\ge d_*n/2,\] contradicting (78). The same reasoning in the past gives \[ a_i(z)=a_i(x)=a_i(y). \tag{80}\] Here \(a_i(S^{-n}u)=a_i(u)\) for every \(u\) in the conjugacy domain, since \(n\) is a multiple of \(m_i\). Let \(q_0\) be the common advance to alignment. The already defined entries \(U_j,V_j\) occupy, in the symbolic orbit of \(z\), the positions \[-n+q_0+jm_i,\qquad q_0+jm_i,\qquad 0\le j\le D-2.\] Their difference is exactly \(n=Dm_i\). The index \(j\) has not been changed after selecting the switch. When one of these entries differs in label from the corresponding word of \(z\), distinct-word separation in 13 costs at least \(d_*m_i\) mismatches. Consequently fewer than \(D/100\) entries on each side can differ. Let \(E_{i,D}\subset X\) be the group-boundary exception from 18, transported to \(M\) and extended arbitrarily on its null complement. Its original measure is at most \(\alpha_{i,D}\le2D_i^*/H_i\to0\), uniformly over the chosen tests. The retained switch has \[\rho\{(x,y)\in G:z(x,y)\in E_{i,D}\} \le2\alpha_{i,D}.\] Take a late test for which \(2\alpha_{i,D}\le.01\). On a subset of the original pair space of mass at least \(.96-.01\), the symbol-matching event holds and the two \(z\)-lists satisfy every graph relation. Removing the fewer than \(D/100\) wrong entries on each side gives \[Z(x,y)\ge 1-\frac{2D}{100(D-1)}\ge.96.\] The statistic is nonnegative elsewhere, so \[ \mathbb E_\rho Z\ge(.96-.01)\cdot.96=.912. \tag{81}\] This contradicts (77). Completion of the proof of 1. 15 supplies a standard nonatomic ergodic invertible probability-preserving system of zero entropy. If it had a smooth volume model, 7 and 19 would supply the tests above. The original pair law at each late test would satisfy both (77) and (81), which is impossible. The dimension and all smooth constants were arbitrary but fixed after construction of the system. Thus no model in any finite dimension exists. ◻ Remark 22. The role of smooth positive volume is quantitative. It gives a switch whose unnormalized output law is controlled by the original invariant measure. That control allows a measurable symbol observation to be compared along shadowing orbits, despite the absence of regularity in the conjugacy. The proof makes no analogous assertion for models carrying only a singular measure or a nonsmooth density equivalent to volume.
Chow, S.-N., X.-B. Lin, and K. J. Palmer. 1989. “A Shadowing Lemma with Applications to Semilinear Parabolic Equations.” SIAM Journal on Mathematical Analysis 20: 547–57.
Dajani, Karma, and Sjoerd Dirksin. 2008. A Simple Introduction to Ergodic Theory. Lecture notes, Utrecht University, December 18.
Foreman, Matthew, Daniel J. Rudolph, and Benjamin Weiss. 2011. “The Conjugacy Problem in Ergodic Theory.” Annals of Mathematics 173 (3): 1529–86.
Foreman, Matthew, and Benjamin Weiss. 2022. “Measure Preserving Diffeomorphisms of the Torus Are Unclassifiable.” Journal of the European Mathematical Society 24 (8): 2605–90.
Gerber, Marlies, and Philipp Kunde. 2025. Non-Classifiability of Mixing Zero-Entropy Diffeomorphisms up to Isomorphism.
Katok, Anatole, and Jean-Paul Thouvenot. 1997. “Slow Entropy Type Invariants and Smooth Realization of Commuting Measure-Preserving Transformations.” Annales de l’Institut Henri Poincaré, Probabilités Et Statistiques 33 (3): 323–38.
Kunde, Philipp. 2024. “Anti-Classification Results for Weakly Mixing Diffeomorphisms.” Mathematische Annalen 390: 5607–68.
Lind, D. A., and J.-P. Thouvenot. 1977. “Measure-Preserving Homeomorphisms of the Torus Represent All Finite Entropy Ergodic Transformations.” Mathematical Systems Theory 11: 275–82.
Oseledets, V. I. 1968. “A Multiplicative Ergodic Theorem. Characteristic Ljapunov Exponents of Dynamical Systems.” Trudy Moskovskogo Matematicheskogo Obshchestva 19: 179–210.
Quas, Anthony, and Terry Soo. 2016. “Ergodic Universality of Some Topological Dynamical Systems.” Transactions of the American Mathematical Society 368 (6): 4137–70.
Ruelle, David. 1978. “An Inequality for the Entropy of Differentiable Maps.” Boletim Da Sociedade Brasileira de Matemática 9 (1): 83–87.
Sason, Igal. 2013. “Entropy Bounds for Discrete Random Variables via Maximal Coupling.” IEEE Transactions on Information Theory 59 (11): 7118–31.
Seidel, J. J. 1973. On Two-Graphs, and Shult’s Characterization of Symplectic and Orthogonal Geometries over GF(2). T.H.-Report 73-WSK-02. Technische Hogeschool Eindhoven.
Smale, Stephen. 1967. “Differentiable Dynamical Systems.” Bulletin of the American Mathematical Society 73: 747–817.
Tonchev, Vladimir D. 2003. “Difference Systems of Sets and Code Synchronization.” Rendiconti Del Seminario Matematico Di Messina, Series II 9: 217–26.
|
| ||||||||
|