A
D
V
E
R
T
I
S
E
M
E
N
T
ADVERTISEMENT
Exact derandomization of logarithmic space: L = RL = BPL
expertly designed by an internal OpenAI model  ·  released 2026-09-23  ·  original PDF
Theorems: 7 Lemmas: 53 Proofs: 67
Formulas: 2,976 Words: 55,488 Play time: ~6 hours

>>> How to Play <<<
We prove $\mathsf L=\mathsf{RL}=\mathsf{BPL}$, resolving the derandomization problem for polynomial-time randomized logarithmic space.

>>> Level Map <<<
  1. Introduction
  2. Historical context
  3. Proof strategy
  4. The transition model and resource parameters
  5. Rank channels for retained vertex sets
  6. The exact correction and copy hierarchy
  7. Correction with a fixed column multiplier
  8. A finite pilot chooses the column multiplier
  9. The finite residual
  10. Copies reduce density while controlling active vertices
  11. Reward transport and the final deterministic target
  12. Stochastic atoms and their interfaces
  13. Table domains and the normalization principle
  14. Dense bins and preparatory normalization
  15. Corrections with an incoming gate
  16. Detector scores
  17. Actual ports and estimated assembly
  18. Integrated losses and assembly bounds
  19. Conditional averaging in the shared environment
  20. Generating paths without retaining their endpoints
  21. Additive estimation and conditional prefix bounds
  22. Tasks and dependency order
  23. Raw sample terms
  24. Mean comparisons and the cost of releasing a source
  25. Choosing precision to pay for scale and domain changes
  26. The additive schedule
  27. Absolute difference bounds and averaging noise
  28. Compression of additive row estimates
  29. The row interface
  30. The update map
  31. Deterministic comparison arguments
  32. Charging overflow events
  33. Bounding numerical ranges
  34. The error reserve with exact endpoint grouping
  35. Adaptive fingerprints within budget blocks
  36. Uniform salts and their access rule
  37. The update rule and the clean calculation
  38. Witness envelopes and escaped paths
  39. Bucket algebra and the collision charge
  40. Closing the coupled induction
  41. Catalytic key equality
  42. Finite numerical queries with bounded suspended storage
  43. Streams, records, and uniformity
  44. Truncations, comparisons, and binary arithmetic
  45. Multiplication by minimum-index bands
  46. Reciprocals and the early-word convention
  47. Indexed ranks, sums, scales, and structural leaves
  48. The numerical interface
  49. Uniform controller implementation
  50. Parameters and interfaces
  51. Program states and record bounds
  52. Contours and incoming operations
  53. Certificates for the estimator recipes
  54. The uniform space theorem
  55. Derandomization and quantitative consequences
  56. Choosing constants and finite ranges
  57. A dyadic approximation at the specified start
  58. The class equality and promise separation
  59. Accuracy supplied as part of the input
  60. Certified accepting computations
  61. The effective compiler
  62. One fixed interpreter with enforced resource bounds
  63. Fixing one deterministic library
  64. Padding and exact preservation of acceptance probability
  65. Compiling the virtual input operations
  66. Explicit simultaneous bounds

Introduction

A logarithmic-space machine can remember a bounded number of positions in its input, but cannot keep a polynomial-length computation or random tape. Can fresh random bits increase what such a machine can decide? We answer this question by giving a deterministic logarithmic-space simulation of every polynomial-time randomized logarithmic-space computation with bounded error.

We use machines with finite control, an endmarked read-only input, finitely many work tapes, and fresh independent fair bits. Input heads are confined between the endmarkers. All randomized running-time bounds hold on every random tape. Space counts traversed writable cells in bits, including every simultaneously live counter, numerical register and suspended call; visiting a blank work cell counts toward space. A transducer has a one-way write-only output tape. Throughout, \(n\) is the input length and logarithms have base two.

The class \(\mathsf L\) consists of languages decidable by deterministic machines in \(O(\log(n+2))\) work space. Such a halting deterministic machine has only polynomially many configurations and hence runs in polynomial time. The class \(\mathsf{RL}\) permits polynomial-time randomized machines in this space, with acceptance probability zero on no inputs and at least \(1/2\) on yes inputs. For \(\mathsf{BPL}\), the corresponding probabilities are at most \(1/3\) and at least \(2/3\).

Theorem 1 (Exact logarithmic-space derandomization). \[\mathsf L=\mathsf{RL}=\mathsf{BPL}.\]

The proof approximates arbitrary acceptance probabilities. This gives the two-sided statement directly, and also applies to computations whose acceptance probabilities have no decision gap. Its quantitative form separates the memory needed for the input from the memory needed for the requested accuracy.

Theorem 2 (Acceptance-probability approximation). Fix a randomized machine \(\mathcal M\) with polynomial worst-case running time and \(O(\log(n+2))\) work space. Let \(p_{\mathcal M}(x)\) be its acceptance probability on an input \(x\) of length \(n\). There is a uniform deterministic algorithm which, on input \((x,1^q)\) with \(q\ge1\), outputs an integer \(0\le a\le2^{q+2}\) satisfying \[\left|a2^{-(q+2)}-p_{\mathcal M}(x)\right|\le2^{-q}.\] It uses \(O(\log(n+2)+q)\) work space and \((n+2)^{O(1)}2^{O(q)}\) time. The constants depend only on \(\mathcal M\).

In particular, every fixed inverse-polynomial accuracy is attained in logarithmic space and polynomial time. The precision parameter is unary; the running-time bound is exponential in general \(q\). We first prove the fixed inverse-polynomial result and then obtain the single variable-precision algorithm by padding. Section 12 also derives total deterministic separators for bounded-error promise problems and constructs accepting runs when the machine’s acceptance probability has a specified inverse-polynomial lower bound.

The simulation is also effective at the level of machine descriptions. Given a source program and supplied polynomial-time and logarithmic-space bounds, one terminating compiler produces a deterministic decider and explicit numerical bounds for its space and time (Theorem 74). The supplied bounds are hypotheses on the source, rather than properties the compiler must decide. Section 13 constructs this compiler by specializing one fixed deterministic library.

Historical context

The complexity of probabilistic Turing machines was developed by Gill (Gill 1977), including measures of both time and tape usage. Random walks made the logarithmic-space question concrete: Aleliunas, Karp, Lipton, Lovász and Rackoff obtained polynomial-time randomized algorithms for undirected connectivity and explicitly asked whether polynomial-time randomized logarithmic space could be simulated in deterministic logarithmic space (Aleliunas et al. 1979). Early work also studied probabilistic machines without a polynomial-time restriction. Borodin, Cook and Pippenger obtained a quadratic-space deterministic simulation in this broader setting (Borodin et al. 1983). The distinction between a space bound and simultaneous time and space bounds is essential throughout this history.

Pseudorandom generators attack the problem by replacing a long sequence of fresh random bits with the output generated from a short seed. Ajtai, Komlós and Szemerédi gave early logarithmic-space constructions that hit small-space tests using short random strings (Ajtai et al. 1987). Babai, Nisan and Szegedy subsequently obtained pseudorandom generators through multiparty communication bounds (Babai et al. 1992). Nisan’s generator uses \(O(S\log R)\) seed bits for a space-\(S\) computation reading \(R\) random bits (Nisan 1992). For logarithmic space and polynomially many bits this gives \(O(\log^2 n)\) seed length; direct seed enumeration has a quasipolynomial time bound. Nisan’s later simultaneous simulation achieves polynomial time and \(O(\log^2 n)\) space (Nisan 1994). Although its title uses \(\mathrm{RL}\), that result includes two-sided bounded error. For \(S(n)\ge\log n\), Nisan and Zuckerman showed that a space-\(S\) computation using \(\operatorname{poly}(S)\) random bits can be simulated using only \(O(S)\) random bits and \(O(S)\) space (Nisan and Zuckerman 1996). The restriction on the original number of random bits is substantive; for logarithmic space it permits polylogarithmically many bits.

Saks and Zhou reduced the deterministic space needed for general bounded-error logarithmic-space computation to \(O(\log^{3/2}n)\) (Saks and Zhou 1999). Cai, Chakaravarthy and van Melkebeek established a simultaneous tradeoff: for each fixed rational \(0\le\alpha\le1/2\), the simulation uses time \(n^{O(\log^{1/2-\alpha}n)}\) and space \(O(\log^{3/2+\alpha}n)\) (Cai et al. 2006). The polynomial-time endpoint has squared-logarithmic space. Hoza subsequently improved the general space bound to \(O(\log^{3/2}n/\sqrt{\log\log n})\) (Hoza 2021, Theorem 1.6 and Corollary A.5). These smaller space containments do not assert the same polynomial-time endpoint. Hoza’s work also develops weighted pseudorandomness, in which signed or real weights replace an ordinary probability distribution; the space-containment improvement and the new weighted-generator construction are distinct results of that paper.

Acceptance-probability estimation also connects derandomization to stochastic matrix products and inverses. Ahmadinejad, Kelner, Murtagh, Peebles, Sidford and Vadhan developed high-precision small-space algorithms for random-walk probabilities, with stronger bounds for Eulerian graphs (Ahmadinejad et al. 2020). Cohen, Doron, Sberlo and Ta-Shma later used varying precision across recursion to improve the space complexity of long stochastic matrix products (Cohen et al. 2023). The construction below works with a finite inverse on a time-layered configuration graph; its correction and copy hierarchy controls the information retained while individual entries are evaluated.

Structural restrictions on a computation provide another route. For undirected connectivity, Nisan, Szemerédi and Wigderson obtained deterministic space \(O(\log^{3/2}n)\) (Nisan et al. 1992), subsequently improved to \(O(\log^{4/3}n)\) by Armoni, Ta-Shma, Wigderson and Zhou (Armoni et al. 2000). Reingold’s logarithmic-space algorithm, whose conference version appeared at STOC 2005, settles undirected connectivity (Reingold 2005, 2008). Reingold, Trevisan and Vadhan studied regular directed graphs and gave logarithmic-space path search for Eulerian directed graphs (Reingold et al. 2006). Their oblivious walk constructions impose additional labeling conditions. General directed configuration graphs need not have these properties, so those results did not settle the unrestricted randomized computation problem. A read-once branching program is a layered deterministic computation that reads the input symbols once in a fixed order (bits in the binary case). Its length is the number of steps, and its width is the largest number of states in a layer. A \(1/2\)-hitting set for a class of programs contains an accepting input for every member that accepts at least half of all input strings. The relation between one-sided and two-sided derandomization also requires care. Cheng and Hoza proved that one uniformly logarithmic-space-enumerable \(1/2\)-hitting-set family for all width-\(n\), length-\(n\) read-once branching programs would yield \(\mathsf L=\mathsf{BPL}\) (Cheng and Hoza 2022). The premise is a hitting-set construction; it is not merely the language-class equality \(\mathsf{RL}=\mathsf L\).

Recent results refine both general and structured models. Chen, Cohen, Doron, Khaskelberg and Ta-Shma improve error reduction for weighted generators; in the polynomial-width regime their construction has seed length \(\log|\Sigma|+O(\log^2 T+\log(1/\delta))\), where \(T\) is branching-program length, \(\Sigma\) its alphabet and \(\delta\) the error (Chen et al. 2026). Cheng and Wu obtain simultaneous polynomial-time and polylogarithmic-space simulations for regular branching programs, with extensions to polynomial-time randomized logarithmic-space machines whose two-way read-only random tape visits each cell at most constantly often, and to such machines with read-once randomness and an auxiliary stack whose push, pop and idle schedule is independent of the random bits. Consecutive stays on a random-tape cell count as one visit (Cheng and Wu 2026, Theorems 5–6). Hardness-versus-randomness approaches additionally give conditional, certified and per-instance alternatives; Pyne and Tell survey this line of work (Pyne and Tell 2026).

Effectivity has its own substantial antecedents. Pyne, Raz and Zhan construct a universal deterministic estimator whose \(O(S(n))\) space guarantee follows from \(\mathsf{prBPL}\subseteq\mathsf{SPACE}[O(S(n))]\), for space-constructible \(S(n)\ge\log n\) (Pyne et al. 2023, full version, Theorem 5.1). Here \(\mathsf{prBPL}\) denotes promise problems solvable in polynomial time and randomized logarithmic space with bounded error. The compiler below uses a fixed separator established in this paper and supplies explicit source-specific numerical resource bounds. Its construction should be read in this context of universal simulation, with the source bounds supplied as part of the compiler input.

Proof strategy

Fix a randomized machine and an input of length \(n\), and write \(B=\Theta(\log(n+2))\). The proof constructs a finite family of estimates of its acceptance probability, indexed by \(O(B)\)-bit descriptions. More than three quarters of these estimates are accurate; every member of the family can be evaluated by a halting deterministic computation using \(O(B)\) space. Exhaustive enumeration then gives a deterministic answer. The second guarantee is essential: enumeration must finish even on choices for which the estimate is inaccurate.

The exact target.

The configuration reduction of Section 2 gives a strictly forward substochastic matrix \(S\) on a polynomial-size set \(V_0\) and a vector \(e\in\{0,1\}^{V_0}\) indicating accepting final configurations. The vector \[p_0=(I-S)^{-1}e\] records acceptance probabilities from all configurations. The inverse is a finite sum, because every transition advances time. Section 4 replaces \(S\) successively by transitions \(C_l\) on sets \(V_l\) of configuration copies, carrying accumulated rewards \(W_l\in[0,1]^{V_l}\). Each copy step has a designated primary copy. If \(z_l\) is obtained from the original start \(x_0\) by always choosing that copy, then \[0\le p_0(x_0)-W_l(z_l) \le |V_0|\,2^{-(H-1)l},\] where \(H\ge4\) is fixed. Thus \(L=O(B)\) stages make \(W_L(z_L)\) accurate to any fixed inverse-polynomial error. Every stage has at most \(N=2^{C_NB}\) vertices for a fixed constant \(C_N\).

From the estimator to a deterministic decision.

The algorithm estimates \(W_L\) in one shared finite random environment \(\sigma\). Its law is uniform on a specified finite set of matrix-and-bit-string encodings of length \(O(B)\). On each channel a retention test selects vertices; write \(X_f(v)\) for the indicator that \(v\) is retained at the final stage rate \(f=2^{-HL}\), and \(\pi_f\) for its probability. Section 3 gives \(\pi_f\asymp f\). Theorem 46 proves, for the actual estimator \(\bar W_L\), \[\frac1N\mathbb E_\sigma\sum_{v\in V_L}\frac{X_f(v)}f |\bar W_L(v)-W_L(v)|^8\le C^8 2^{-8p}.\] Here \(p\) is the statistical accuracy exponent. Keeping only the summand at \(z_L\) gives \[\mathbb E\!\left[|\bar W_L(z_L)-W_L(z_L)|^8 \mid X_f(z_L)=1\right] \le N\frac f{\pi_f}C^8 2^{-8p}.\] The factor \(N\) is the price of passing from the integrated bound to one specified start. Since \(\log N=O(B)\), taking the estimator’s statistical budget to be a sufficiently large fixed multiple of \(B\) makes \(p\) large enough to pay this price and gives an accurate value on more than three quarters of the conditional environments.

The implementation computes dyadic approximations to the same fixed value \(\bar W_L(z_L)\) on every environment. Numerical precision changes the requested digits, not the lists and ports defining that value. Choosing the stage count, statistical budget and digit accuracy makes the total error less than \(1/8\) on more than three quarters of the encodings conditioned on \(X_f(z_L)=1\). Their median has the same error bound. We enumerate the original encodings with their original multiplicities, so the multiset reproduces exactly this conditional law.

Repeated enumeration finds the median while retaining only the current encoding and \(O(B)\)-bit numerical and counting registers. Every digit query used here requests one of the first \(O(B)\) binary places, terminates in \(O(B)\) space, and confines its physical input heads between the real endmarkers. It therefore has only \(2^{O(B)}\) complete configurations and takes at most that many steps. There are also \(2^{O(B)}\) encodings, so the entire computation takes polynomial time. Comparing its answer with \(1/2\) distinguishes acceptance probabilities at most \(1/3\) from those at least \(2/3\). This proves the nontrivial containment in Theorem 1.

Why the hierarchy can be estimated.

The hierarchy reduces transition entries more rapidly than it increases the number of vertices incident to positive entries. A correction \(E\) replaces \(C\) by \(D=C-E+EC\), with \[I-D=(I+E)(I-C).\] This identity transfers the removed transition contribution into the reward. Detectors identify columns with many substantial entries and suppress correction there. The copy step distributes the remaining large entries while controlling how many vertices are incident to them. An ordered-support argument gives a factorial truncation bound, so only constantly many correction steps are needed at each stage.

The exact matrices and conditional expectations define the targets; they are not stored. Section 5 represents their contributions by sample-dependent tables with bounded row support. Corrections also have bounded incoming degree. Their true conditional means recover the required entries. The algorithm substitutes earlier estimates into these formulas, so the analysis must bound the resulting bias without assuming that the estimated factors are independent. Conditional averaging preserves the specified source and endpoint retention tests. Its uniform mixing bound is derived from classical property (T) results of Shalom (Shalom 1999); Section 6 proves the finite-action and sampling consequences used here.

At parent budget \(M\), Section 7 averages differences between child estimates at consecutive budgets \(m\) and \(m-1\). These differences have small norm because both approximate the same true sample term. Their contributions at the same endpoint are combined before measuring the difference. The slack \(M-m\) has two roles: it determines the length of conditional averaging available to reduce the noise, and it bounds the description of an immediate sample path. Grouping comparable slacks lets Section 8 replace each row array by a short list of exceptions to a common default. The proof charges deterministic conditional-mean error separately from residual noise.

A product follows a correction edge to an intermediate vertex and then queries a right factor there. Releasing the old source changes the weighted error domain. Bounded incoming degree controls that change, and the task’s accuracy assignment pays the resulting factor. Thus the same construction that permits a source address to be discarded also appears explicitly in the error analysis.

Why the lists fit in logarithmic space.

A list entry is found by a path. Retaining a full endpoint at every recursive level would cost too much. Section 9 therefore compares endpoints by short fingerprint keys at small slacks, combining polynomial evaluation and affine hashing (Schwartz 1980; Carter and Wegman 1979). Candidate endpoints can depend on earlier comparisons, so a fixed-pair collision estimate alone is insufficient. A comparison calculation using exact equality within one budget block produces candidate sets independent of that block’s fingerprint bits. Smooth support functions charge paths that leave these sets to lower-call errors; collisions within the fixed sets supply the remaining error. This coupled induction proves the moment bound for the actual estimator.

Sections 10 and 11 establish its total space bound. Numerical operations request one bounded digit at a time. The controller keeps one vertex cursor, uses bounded inverse ports for correction moves, and recovers incoming accesses by traversing finite program trees, using the contour idea of Cook and McKenzie (Cook and McKenzie 1987) and the space-preserving traversal perspective of Lange, McKenzie and Tapp (Lange et al. 2000). A shared catalytic bit vector permits fingerprint comparisons without keeping a key index at every recursive level. The controller assigns each call a decreasing recursion allowance \(w\), with top allowance \(O(B)\), and uses the requested digit place \(t\) as its numerical clock. If the child has allowance \(w'\) and digit place \(j\), all local data retained during that call occupy at most \[A(w-w')+D(t-j)\] bits, including the controller copies used by a traversal. The child clock may increase only in proportion to the decrease of allowance. These differences telescope along the active call chain; the environment, cursor, catalyst and scratch are shared once. The bound and termination do not depend on statistical accuracy.

Section 12 carries out the median construction for fixed inverse-polynomial error and derives the class equality and promise separation. It then obtains one algorithm for all unary precision requests by padding, and constructs certified accepting computations when the machine’s acceptance probability has an inverse-polynomial lower bound. Section 13 is a separate application of fixed-precision approximation: one resource-capped universal interpreter yields a fixed deterministic library, whose syntactic specialization supplies explicit resource bounds for each source machine.

The transition model and resource parameters

We first express acceptance probability as a finite matrix inverse. Time labels make the transition matrix nilpotent; keeping the complete configuration at each vertex gives local access in both directions. The construction will use the resulting probability vector at every configuration, as well as its value at the designated start.

A matrix indexed by a finite set \(V\) is strictly forward if there is an integer time attached to every vertex and every nonzero entry goes to a strictly later time. It is substochastic if its entries are nonnegative and each row sums to at most one. An edge is specified by a vertex and a port from a fixed finite set. A paired edge operation returns the other endpoint together with its inverse port; applying the inverse operation restores the first endpoint. Absent ports are explicitly recognized.

Lemma 3 (Configuration reduction). Fix a polynomial-time randomized machine using \(O(\log(n+2))\) work space. On an input of length \(n\) one can specify, uniformly in logarithmic space, a polynomial-size vertex set \(V_0\), a strictly forward substochastic matrix \(S\), a vector \(e\in\{0,1\}^{V_0}\), and a start vertex \(x_0\) with the following properties.

  1. The incoming and outgoing port degrees of \(S\) are bounded by constants depending only on the machine. Port validity, edge weights, and pairing are computable in logarithmic space.

  2. The vector \(p_0=(I-S)^{-1}e\) is well defined and belongs to \([0,1]^{V_0}\).

  3. The coordinate \(p_0(x_0)\) is the acceptance probability of the machine.

In particular, an \(\mathsf{RL}\) computation gives the promise \(p_0(x_0)=0\) or \(p_0(x_0)\ge1/2\).

Proof. Let \(T(n)\) be a fixed polynomial upper bound on the running time. A vertex records a time in \(\{0,\ldots,T(n)\}\), the finite control, all input and work head positions, and the contents of a sufficiently large box of \(O(\log(n+2))\) work cells. Take elementary machine steps with fixed dyadic coin probabilities. A transition advances the time by one. Invalid configurations or moves leaving the box have no corresponding transition. The box is chosen large enough that this never affects a computation from the specified start.

Each elementary step has constantly many successor choices. To find predecessors, guess the old finite control, locally overwritten symbols, head movements, and coin choice, then check that the forward step gives the specified configuration. There are constantly many such guesses. Combine transitions with the same ordered pair of endpoints by adding their probabilities. Order the distinct successors and predecessors by their configuration encodings; their positions give paired ports for the resulting matrix entries. These lists have constant length, so duplicate removal, summed weights, and the inverse port are computed in logarithmic space by local enumeration. This base computation makes no recursive graph queries. All records use \(O(\log(n+2))\) bits, and the read-only input can be consulted during these checks.

All input-head coordinates in the vertex universe are restricted to a fixed polynomial interval containing the positions reachable within \(T(n)\) local steps from their specified initial positions. Configurations and predecessor candidates outside that interval are invalid. To obtain a simulated head’s scanned symbol, compare its stored coordinate with the input boundaries, supply any required endmarker or off-data blank according to the source tape convention, and otherwise scan the real endmarked input to the designated position using a logarithmic-space counter. A position forbidden by the source tape convention remains invalid. The physical scan stays between the real endmarkers, even when a valid simulated coordinate is outside the data. This routine makes no recursive graph query.

An already halted computation keeps its complete tape and head data while its time advances. In particular, halted configurations are not merged into a single vertex. Let \(e\) be one precisely at accepting configurations at time \(T(n)\), and zero elsewhere. Final-time rows have no outgoing transitions. Thus \(S\) is strictly forward and substochastic, and it is nilpotent. Its inverse expression is the finite sum \[(I-S)^{-1}e=\sum_{j=0}^{T(n)}S^je.\] Starting from any valid configuration, at most one reward can be collected, at final time. The sum is therefore a probability in \([0,1]\). At \(x_0\) it is exactly the machine’s acceptance probability. ◻

The coordinatewise bound on \(p_0\) will be used after weighted projections to larger vertex sets. It holds at every configuration, independently of the acceptance probability at the designated start. Thus the matrix construction and the probability approximation below apply to arbitrary randomized computations satisfying the resource bounds.

Size convention.

Fix an integer parameter \(B\), a power of two, comparable with \(\log_2(n+2)\) and at least a sufficiently large fixed constant. For example, take the least power of two at least the maximum of that constant and \(\lceil\log_2(n+2)\rceil\). This choice is computable in logarithmic space and includes all bounded input lengths in the same construction. All constants throughout the construction are independent of the input; they may depend on the fixed machine and, for fixed inverse-polynomial accuracy, its fixed exponent. The uniform precision parameter of Theorem 2 is handled separately by padding. The number of stages is \(L=O(B)\). At stage \(l\), a vertex has a base identifier followed by \(l\) symbols from a fixed radix \(D_0\). The radix will be selected by the hierarchy. Encode each base vertex by a fixed-width radix-\(D_0\) word and prepend the digit \(1\); unused base words are invalid. Appending one copy digit at each stage gives injective numerical identifiers even across different stage lengths. Indeed, a base width \(w\) places stage-\(l\) identifiers in \([D_0^{w+l},2D_0^{w+l})\), and these intervals are disjoint. Appending a fixed suffix is an affine map on identifiers.

Every stage has size at most a common bound \[N=2^{C_NB}.\] We may increase the fixed constant \(C_N\) as needed. This is a bound on one stage, not a claim that there are only \(N\) pairs of vertices. Integrated entry errors below carry a factor \(1/N\) while summing over both endpoints.

Fix integers \(H\ge4\) and \(2\le h_0<H\), and put \[b_k=2^{-Hk}\qquad(k\ge0).\] These are the nominal sampling rates. Stage \(l\) will have density \(f=b_l\). The stochastic tables will use rates from this same grid for their source and terminal retention tests. The error power in the later analysis is always \(s=8\).

Uniform constants and finite ranges.

All thresholds needed by the algorithm can be chosen rational, with fixed gaps, and all discrete constants can be chosen integer. We first prove estimates with constants uniform in the ambient finite universe. The sampling construction applies to every finite range of identifiers and rates once its prime field is sufficiently large. The hierarchy uses \(O(B)\) copy digits and rates with \(k=O(B)\); Section 12.1 will verify this rate bound also for the recursive estimator and give a compatible order of constant choices.

We can then take the prime field to have size \(2^{O(B)}\), with its exponent chosen after these constants. Its elements and all identifiers occupy \(O(B)\) bits. Increasing that exponent does not change the probabilistic constants proved in the next section.

Rank channels for retained vertex sets

The hierarchy will use random retained sets to detect columns with many substantial entries. The atom construction will use the same sets to represent matrix entries. We define the sampling channels and prove the marginal, conditional-pair, and count estimates needed for these two constructions. Section 6 later constructs averaging walks that preserve specified retention conditions.

The environment consists of the independent rank channels below together with a finite string of independent uniform auxiliary bits. The length and uses of that string will be specified when needed.

We use a fixed number of independent rank channels, including channels denoted \(X,Y,Z\). Extra channels, when needed, are also a fixed finite number. A channel is a uniformly random matrix \[A=(a_0,F)\in\mathbb F_P^{H\times h_0}\] conditioned on the last \(r=h_0-1\) columns \(F\) being linearly independent. We call such an ordered collection of columns a frame. The integer \(P\) is prime. It is chosen larger than every identifier and sufficiently large for all rate comparisons below. For a numerical identifier \(y\), set \[v(y)=(1,y,\ldots,y^{h_0-1})^{\mathsf T}, \qquad a(y)=Av(y)\in\mathbb F_P^H.\] Field elements in a retention test use their representatives in \(\{0,\ldots,P-1\}\). At nominal rate \(b_k\), let \[L_k=\lceil P2^{-k}\rceil,\qquad \pi_{b_k}=(L_k/P)^H.\] The identifier is retained if all coordinates of \(a(y)\) lie in \(\{0,\ldots,L_k-1\}\). Write \(X_a(y)\) for the indicator on channel \(X\) at rate \(a\), and similarly for other channels. The same notation permits a specified fixed suffix in the channel argument. Retention boxes are nested. All rates in this section belong to the grid \(b_k=2^{-Hk}\), and the admissible range satisfies \(P2^{-k}\ge1\). The choices of \(H,h_0\) and the identifier range are those of Section 2.

Lemma 4 (Rank-channel estimates). There are positive constants \(c_\pi,C_\pi,C_{\rm pair},C_{\rm var}\), depending only on \(H,h_0\), with the following properties in every admissible finite range.

  1. A specified identifier is retained at rate \(a\) with probability \(\pi_a\), where \(c_\pi a\le\pi_a\le C_\pi a\).

  2. For distinct identifiers \(x,y\) and any admissible rates \(u,a\), \[\Pr\{X_a(y)=1\mid X_u(x)=1\}\le C_{\rm pair}a.\]

  3. If \(A_0\) is a deterministic set of \(q\) distinct identifiers and \(T_a=\sum_{y\in A_0}X_a(y)\), then \[\Pr\{T_a<q\pi_a/2\}\le \frac{C_{\rm var}}{q\pi_a} \quad(q>0).\]

  4. Replacing every channel argument \(y\) by the same injective affine map \(\alpha y+\beta\) leaves the joint law of all evaluations unchanged, provided \(\alpha\ne0\) in \(\mathbb F_P\). In particular, a common fixed padding suffix leaves their law unchanged.

Proof. The probability that an unrestricted \(H\)-by-\(r\) matrix is a frame is \[q_{\rm fr}=\prod_{j=0}^{r-1}(1-P^{j-H})\ge c_{\rm fr}>0.\] For fixed positive-degree coefficients, \(a_0\) is uniform, so any one evaluation is uniform in \(\mathbb F_P^H\) and independent of frame success. This proves the exact marginal retention probability. In the ranges used, take \(P2^{-k}\ge1\). Then \(P2^{-k}\le L_k\le2P2^{-k}\), proving comparability. The prime exponent can in fact make all these sides much larger than one.

Without conditioning on a frame, two evaluations at distinct field points are independent uniform vectors: the two evaluation functionals have rank two since \(h_0\ge2\). Consequently \[\Pr\{X_a(y)=1,X_u(x)=1\mid F\text{ is a frame}\} \le \frac{\pi_a\pi_u}{q_{\rm fr}}.\] The first marginal after conditioning is still \(\pi_u\), which proves the second assertion.

For the third assertion, in the unrestricted distribution the hit count has mean \(q\pi_a\) and variance at most \(q\pi_a\), by pairwise independence. Chebyshev’s inequality bounds the stated lower-tail event by \(4/(q\pi_a)\). Conditioning on frame success increases this probability by at most \(1/q_{\rm fr}\).

Finally, substituting \(\alpha y+\beta\) in a polynomial acts on its positive-degree coefficient columns by an invertible triangular matrix, with diagonal \(\alpha,\ldots,\alpha^{h_0-1}\). It therefore preserves the frame condition. Its action on the full coefficient array is an invertible linear change of variables, so it preserves uniform measure conditional on that condition. Appending digits in radix \(D_0\) has \(\alpha\) a power of \(D_0\); taking \(P>D_0\) ensures the required invertibility. ◻

The exact correction and copy hierarchy

At each stage we reduce the transition entries and transfer their contribution to a reward. The algebra behind this transfer is simple. Given a transition matrix \(C\) and a nonnegative correction \(E\), put \[D=C-E+EC.\] Then \(I-D=(I+E)(I-C)\). Thus a relation \(W=(I-C)p\) becomes \((I+E)W=(I-D)p\): the same vector \(p\) is represented by the new transition and the updated reward. The correction must make \(D\) nonnegative and keep its row sums at most one. It must also reduce entries sufficiently fast that, after a controlled increase in the number of active vertices, the remaining transition contribution becomes small. The construction uses only the model reduction, with no one-sided acceptance promise.

At one stage the input is a finite ordered set \(V\), a rate \(0<f\le1\) on the grid of Section 3, and a nonnegative matrix \(C\) with \[ C(x,y)=0\quad\text{unless }x<y,\qquad C\mathbf1\le\mathbf1,\qquad C(x,y)\le f. \tag{1}\] All matrix and vector inequalities are entrywise. Write \[\mathcal R(C)=\{x:\text{some }C(x,y)>0\},\quad \mathcal T(C)=\{y:\text{some }C(x,y)>0\},\quad \mathcal A(C)=\mathcal R(C)\cup\mathcal T(C).\] The stage will reduce the entry bound from \(f\) to \(f/2^H\) and increase \(|\mathcal A(C)|\) by at most a factor of two. We first analyze the correction with an arbitrary fixed column multiplier. We then choose that multiplier by a finite detector calculation, and use copies to divide the entries in columns where correction is not fully enabled.

Correction with a fixed column multiplier

Put \(q=2^H\). Fix rational numbers \(0<\theta_1<\theta_2<q^{-1}/32\) and a nondecreasing Lipschitz function \(Q:[0,\infty)\to[0,\infty)\) satisfying \[ (t-\theta_2)_+\leq Q(t)\leq(t-\theta_1)_+. \tag{2}\] Its eighth root is required to be Lipschitz on bounded intervals. Here and below the scalar cutoffs can be chosen piecewise polynomial with rational coefficients. To see this explicitly, let \(\beta\) be the normalized integral of \(t^7(1-t)^7\) on \([0,1]\), extended by zero to the left and one to the right. Both \(\beta^{1/8}\) and \((1-\beta)^{1/8}\) are Lipschitz: near their vanishing endpoints the corresponding functions vanish to order eight, and away from those endpoints the assertion follows by differentiation. For example, \[Q(t)=(t-\theta_1)_+ \beta\left(\frac{t-\theta_1}{\theta_2-\theta_1}\right)\] has all the stated properties. Near \(\theta_1\) its eighth root is of order \((t-\theta_1)_+^{9/8}\); elsewhere its root has bounded derivative on every bounded interval. Let \(L_Q\) be a fixed Lipschitz constant for \(Q\). For real arguments, extend \(Q\) by \(Q(t)=0\) when \(t<0\). This extension preserves its Lipschitz and root-Lipschitz properties and permits evaluation on real approximations. The exact hierarchy itself uses only nonnegative arguments.

Fix a column multiplier \(g:V\to[0,1]\). Starting from zero, define \[ E_0=0,\qquad G'_i=C+E_{i-1}C,\qquad E_i(x,y)=fQ(G'_i(x,y)/f)g(y)\quad(i\ge1). \tag{3}\] The cutoff leaves a positive residual at each corrected entry. Keeping \(g\) fixed makes the update map monotone. We next show that a fixed number of iterations suffices for every such multiplier, independently of the size of \(V\). At a nonnegative fixed point \(E_*\) of this recurrence, put \(T=C-E_*+E_*C\). In a column with \(g(y)=1\), the lower bound on \(Q\) gives \(T(x,y)\le\theta_2f\). The truncation estimate below will transfer this entry bound to a finite correction, with a small additional error. The fixed point is only a comparison object; the hierarchy will use a finite iteration.

Lemma 5 (Fixed point and factorial truncation). For every fixed column multiplier \(0\leq g\leq1\), there is a unique strictly forward nonnegative solution of \[ E_*(x,y)=fQ((C+E_*C)(x,y)/f)g(y). \tag{4}\] For the recurrence in Equation (3), \(0\leq E_i\leq E_*\), and for every \(k\geq0\), \[ \max_{x,y}\frac{E_*(x,y)-E_k(x,y)}f \leq e^{1/\theta_1}\frac{(L_Q/\theta_1)^k}{k!}. \tag{5}\] In particular \(k_*\) can be chosen as a fixed integer so that the right side of Equation (5) at \(k=k_*\) is at most \(q^{-1}/32\), uniformly in \(C,f,V\), and \(g\).

Proof. Order the vertices compatibly with the strictly forward relation. For a fixed row \(x\), the entry \((E_*C)(x,y)\) involves only \(E_*(x,z)\) with \(z<y\). Equation (4) therefore defines the row successively, in increasing column order. Columns at or before \(x\) receive zero. This proves existence and uniqueness in the stated class, without a contraction hypothesis. Monotonicity of the update map proves \(E_i\leq E_*\) and the nonnegativity of the differences.

Set \(T=C-E_*+E_*C\). If \(a=(C+E_*C)(x,y)\), then Equation (2) implies \[0\leq E_*(x,y)\leq(a-\theta_1f)_+.\] Thus \(T\geq0\), and at every positive position of \(E_*\) one has \(T(x,y)\geq\theta_1f\). Also \[ T\mathbf 1=C\mathbf 1-E_*(\mathbf 1-C\mathbf 1) \leq C\mathbf 1\leq\mathbf 1. \tag{6}\] Consequently, for each fixed row \(x\), its set \(A_x=\{y:E_*(x,y)>0\}\) has cardinality \(m_x\leq1/(\theta_1f)\).

Unroll the inequality \(E_*\leq C+E_*C\) along this row’s support. Each resulting term for \(E_*(x,y)\) is a product of entries of \(C\) along a sequence \[x<z_1<\cdots<z_j<y, \qquad z_1,\ldots,z_j\in A_x.\] Every substitution keeps the source coordinate equal to \(x\): the recursive factor is \(E_*(x,z)\), while only its terminal coordinate changes. All intermediate vertices therefore belong to the same set \(A_x\). Strict ordering makes a choice of intermediate vertices specify at most one such sequence. There are at most \(\binom{m_x}{j}\) such sequences, and each product is at most \(f^{j+1}\). The expansion terminates because every sequence is strictly increasing. Therefore \[ E_*(x,y)\leq f\sum_{j=0}^{m_x}\binom{m_x}{j}f^j =f(1+f)^{m_x}\leq fe^{1/\theta_1}. \tag{7}\] Only the support of the single fixed row \(x\) was used here; the number of vertices in the whole graph does not enter the exponent.

Let \(\Delta_i=E_*-E_i\). These matrices are nonnegative, and their row-\(x\) supports are contained in \(A_x\). The Lipschitz bound gives \[\Delta_i(x,y)\leq L_Q\sum_{z<y}\Delta_{i-1}(x,z)C(z,y).\] Outside \(A_x\) the left side is zero, so when this inequality is unrolled it suffices to use intermediate vertices from \(A_x\) at every step. After \(k\) steps there are at most \(\binom{m_x}{k}\) strictly increasing sequences of intermediate vertices. Their initial differences are at most \(fe^{1/\theta_1}\) by Equation (7), and their \(k\) transition factors are at most \(f^k\). Thus \[\Delta_k(x,y) \leq fe^{1/\theta_1}L_Q^kf^k\binom{m_x}{k} \leq fe^{1/\theta_1}\frac{(L_Q/\theta_1)^k}{k!}.\] For \(k>m_x\) the first upper bound is zero. This proves Equation (5). The factorial eventually dominates the fixed exponential factor, giving the stated fixed choice of \(k_*\). ◻

The iteration count \(k_*\) is now fixed independently of the detectors. We next choose \(g\) so that correction is supported on columns with only boundedly many substantial entries. A pilot selects this multiplier; the actual recurrence then uses that one multiplier for all \(k_*\) steps.

A finite pilot chooses the column multiplier

All expectations defining the detectors are true expectations; an estimate is never used to define its own target.

For the detectors we use independent channels \(X,Y\) from Section 3. Denote their common retention probability at rate \(f\) by \(\pi_f\). The consequences of Lemma 4 needed here can be written, with fixed positive constants, as \[\begin{align*} c_\pi f\leq\pi_f\leq C_\pi f,\qquad &\Pr\{Y_f(w)=1\mid Y_f(y)=1\}\leq C_{\rm pair} f &&(w\ne y), \tag{8}\\ &\Pr\left\{\sum_{x\in A}X_f(x)<\tfrac12|A|\pi_f\right\} \leq\frac{C_{\rm var}}{|A|\pi_f} &&(\varnothing\ne A\subseteq V\text{ deterministic}). \tag{9}\end{align*}\] These constants are uniform in the stage, the admissible moduli, and common fixed padding suffixes. No independence between different rows of a detector trial will be assumed.

All threshold choices and ranks below are fixed constants. The pilot uses the iteration count from Lemma 5. Choose \[0<\delta_0<\delta_1<\theta_1/100, \qquad 0<\eta_-<\eta_+<\delta_0/3.\] If other row scores require still smaller detector thresholds, make those choices before fixing the ranks. Let \(\phi\) be a nondecreasing cutoff which is zero through \(\delta_0\) and one from \(\delta_1\) on. Let \(\rho\) be a nonincreasing cutoff which is one through \(\eta_-\) and zero from \(\eta_+\) on. Finally, let \(\gamma:[0,1]\to[0,1]\) be nonincreasing, one on \([0,1/2]\), and zero on \([4/5,1]\). These functions and their complements, when used, have the root regularity just described.

For a finite multiset of nonnegative numbers, \(\operatorname{rank}_j\) means its \(j\)-th largest member, with enough additional zeros that this is always defined. Starting from \(\widehat E_0=0\), perform \(k_*\) pilot steps. At step \(i\), first set \[ G_i=C+\widehat E_{i-1}C. \tag{10}\] For a draw of the terminal channel, set \[ R_i(x)=\rho\left(\operatorname{rank}_K \{G_i(x,w)/f:Y_f(w)=1\}\right). \tag{11}\] Define the deterministic column number \[ U_i(y)=\mathbb E\left[ \operatorname{rank}_J \{\phi(G_i(x,y)/f)R_i(x):X_f(x)=1\} \,\middle|\,Y_f(y)=1\right]. \tag{12}\] Then put \[ \widehat g_i(y)=\gamma\left(\max_{1\leq r\leq i}U_r(y)\right), \qquad \widehat E_i(x,y)=fQ(G_i(x,y)/f)\widehat g_i(y). \tag{13}\] The order of these definitions is essential: \(G_i\) is already deterministic when its detector is formed, and \(U_i\) is defined before \(\widehat E_i\). The running maximum makes \(\widehat g_i\) nonincreasing. Its final value is therefore no larger than any earlier pilot multiplier; monotonicity of \(Q\) will let us bound the actual corrections by the pilot corrections.

After the pilot, set \(g=\widehat g_{k_*}\) and specialize the fixed-multiplier recurrence of Equation (3) to the actual correction \[ E_0=0,\qquad G'_i=C+E_{i-1}C,\qquad E_i(x,y)=fQ(G'_i(x,y)/f)g(y)\quad(1\leq i\leq k_*). \tag{14}\] Write \[ E=E_{k_*},\qquad D=C-E+EC. \tag{15}\]

Lemma 6 (Elementary pilot bounds). All pilot and actual inputs and corrections are nonnegative and strictly forward. For \(1\leq i\leq k_*\), \[\begin{align*} G_i\mathbf 1,\ \widehat E_i\mathbf 1, \ G'_i\mathbf 1,\ E_i\mathbf 1&\leq i\mathbf 1,\\ G_i(x,y),\ \widehat E_i(x,y), \ G'_i(x,y),\ E_i(x,y)&\leq if. \end{align*}\] Moreover, \[E_{i-1}\leq E_i\leq\widehat E_i, \qquad G'_i\leq G_i.\] The row and column supports of every matrix in this statement are contained in \(\mathcal R(C)\) and \(\mathcal T(C)\), respectively.

Proof. Equation (2) gives \(Q(0)=0\) and \(0\leq Q(t)\leq t\) for \(t\ge0\). Consequently \(\widehat E_i\leq G_i\). If \(\widehat E_{i-1}\mathbf 1\leq(i-1)\mathbf 1\), then \[G_i\mathbf 1=C\mathbf 1+\widehat E_{i-1}C\mathbf 1 \leq\mathbf 1+\widehat E_{i-1}\mathbf 1 \leq i\mathbf 1.\] For an individual entry, \[G_i(x,y)\leq f+f\sum_z\widehat E_{i-1}(x,z)\leq if.\] This proves the pilot bounds by induction.

The sequence \(\widehat g_i\) is nonincreasing, so \(g\leq\widehat g_i\) for every \(i\). Induction in Equations (10)–(14), using monotonicity of \(Q\), proves \(E_i\leq\widehat E_i\) and \(G'_i\leq G_i\). The map \(F\mapsto fQ((C+FC)/f)g\), applied entrywise after forming the matrix product, is monotone. Its iteration from zero is therefore nondecreasing. The actual bounds follow from the pilot bounds.

A product of strictly forward matrices is strictly forward. A zero row of \(C\) remains a zero row in all these formulas, by induction from the zero correction. A zero column of \(C\) is a zero column of every product of the form \(FC\) and hence of every input and correction. This proves the remaining assertions. ◻

Lemma 7 (The detector bounds overloaded columns). Let \[D_{\rm row}=1+C_{\rm pair}k_*/\eta_-.\] Choose the integer \(K\geq160D_{\rm row}\), increasing it if necessary for the other finitely many row-gate estimates. For any fixed positive integer \(J\), choose \[ R_*\geq c_\pi^{-1}\max(4J,40C_{\rm var}). \tag{16}\] Then the pilot has the following properties.

  1. If \(\#\{x:G_i(x,y)\geq\delta_1f\}>R_*/f\), then \(U_i(y)\geq19/20>9/10\).

  2. If \(U_i(y)>1/4\), then \[\#\{x:G_i(x,y)>\delta_0f\} >\frac{J}{4C_\pi f}.\]

  3. For any pilot correction with multiplier \(\widehat g_i(y)>0\), \[\#\{x:G_i(x,y)\geq\delta_1f\}\leq R_*/f.\] For every actual correction, \(g(y)>0\) implies \[\#\{x:G'_i(x,y)\geq\delta_1f\}\leq R_*/f.\]

Proof. Fix \(i,y\) and condition on \(Y_f(y)=1\). In a fixed row, the set \(\{w:G_i(x,w)>\eta_-f\}\) has cardinality at most \(k_*/(\eta_-f)\) by Lemma 6. Its expected number of retained members is at most \(D_{\rm row}\): the specified terminal \(y\), if in this set, contributes at most one, and Equation (8) applies to every other member. If the row gate is not full, at least \(K\) such members have been retained. Markov’s inequality therefore gives \[ \Pr\{R_i(x)<1\mid Y_f(y)=1\}\leq D_{\rm row}/K. \tag{17}\]

Let \(A=\{x:G_i(x,y)\geq\delta_1f\}\) and \(n=|A|\). Let \(N_A\) count its retained starts, and let \(L_A\) count those retained starts whose row gate is not full. Independence of the two channels and Equation (17) imply \[\mathbb E[L_A\mid Y_f(y)=1]\leq n\pi_fD_{\rm row}/K.\] Thus \[\Pr\{L_A>n\pi_f/4\mid Y_f(y)=1\}\leq1/40.\] The law of the start channel is unchanged by the conditioning, so Equation (9) gives \[\Pr\{N_A<n\pi_f/2\mid Y_f(y)=1\} \leq C_{\rm var}/(n\pi_f).\] If \(n>R_*/f\), Equation (16) ensures \(n\pi_f\geq\max(4J,40C_{\rm var})\). Outside an event of probability at most \(1/20\), there are at least \(J\) retained high rows with full gates. Each has score exactly one, so the trial’s \(J\)-th largest score is one. All trial scores lie in \([0,1]\), proving the first assertion. This argument used no independence among the row-gate failures.

For the second assertion let \(A_0=\{x:G_i(x,y)>\delta_0f\}\), and let \(N_0\) be its number of retained starts. A positive score must come from \(A_0\). Pointwise the trial is at most \(\mathbf 1_{\{N_0\geq J\}}\), and hence at most \(N_0/J\). Its conditional expectation is at most \(|A_0|\pi_f/J\leq |A_0|C_\pi f/J\). The asserted lower bound on \(|A_0|\) follows.

Finally, a positive \(\widehat g_i(y)\) implies \(U_i(y)<4/5\), so the first assertion excludes a high-entry count greater than \(R_*/f\). If \(g(y)>0\), all pilot multipliers at \(y\) are positive. Apply the same bound to every \(G_i\) and then use \(G'_i\leq G_i\) from Lemma 6. ◻

The finite residual

The factorial estimate gives a fixed iteration count. We now return to the finite correction \(E=E_{k_*}\) and show that its residual is substochastic and has small entries on columns where correction is fully enabled. We also bound the unsigned constituents of its defining expression by the residual itself, as needed when that expression is estimated term by term.

Lemma 8 (Residual bounds and unsigned-term domination). For the \(k_*\) chosen in Lemma 5, put \[K_*=k_*+1,\qquad \varepsilon_{\rm good}=\theta_2+q^{-1}/32<q^{-1}/16, \qquad c_{\rm dom}=\theta_1/(2k_*).\] The matrix \(D\) in Equation (15) is nonnegative, strictly forward, and substochastic. Its entries satisfy \[ D(x,y)\leq K_*f, \qquad g(y)=1\ \Longrightarrow\ D(x,y)\leq\varepsilon_{\rm good}f. \tag{18}\] Furthermore, \[ D\geq c_{\rm dom}(C+E+EC). \tag{19}\] Thus every unsigned term in \(C-E+EC\) is bounded by a fixed multiple of \(D\). In the formulas \(G_i=C+\widehat E_{i-1}C\) and \(G'_i=C+E_{i-1}C\), each unsigned term is bounded by the corresponding input itself. Moreover, \(\mathcal R(D)\subseteq\mathcal R(C)\) and \(\mathcal T(D)\subseteq\mathcal T(C)\).

Proof. Set \(A=G'_{k_*}\) and \(\delta=(E-E_{k_*-1})C\geq0\). Then \[ D=A-E+\delta, \qquad A-E\geq\min(A,\theta_1f). \tag{20}\] Here and in the minimum the scalar \(\theta_1f\) denotes the constant entrywise bound. The second inequality follows from \(E\leq(A-\theta_1f)_+\). It proves \(D\geq0\). Since \(C\mathbf 1\leq\mathbf 1\) and \(E\geq0\), \[D\mathbf 1=C\mathbf 1-E(\mathbf 1-C\mathbf 1) \leq C\mathbf 1\leq\mathbf 1.\] Also \(D\leq C+EC\), whose entries are at most \(f+f\sum_z E(x,z)\leq(k_*+1)f\) by Lemma 6. Strict forwardness and the support claims follow from the same formulas.

Because \(A(x,y)\leq k_*f\), Equation (20) gives \(A-E\geq(\theta_1/k_*)A\). On the other hand, \[C+E+EC=A+E+\delta\leq2A+\delta.\] Since \(\theta_1/k_*<1\), these two inequalities prove Equation (19). The remaining unsigned-term assertions are immediate from nonnegativity of the summands in the two input formulas.

For the good-column bound use the fixed point with this same multiplier \(g\), and set \(\Delta=E_*-E\). With \(T=C-E_*+E_*C\), \[D=T+\Delta-\Delta C\leq T+\Delta.\] If \(g(y)=1\), the lower bound in Equation (2) gives \(T(x,y)\leq\theta_2f\). Equation (5) gives \(\Delta(x,y)\leq q^{-1}f/32\). This proves Equation (18). ◻

The domination in Equation (19) is also an estimation bound. It lets us compare the three unsigned contributions \(C\), \(E\), and \(EC\) separately against their estimates, measuring each error relative to the same nonnegative target \(D\). Their signs are then combined by a triangle inequality in Lemma 26. The estimated raw arrays may have negative entries; positivity here is a property of the exact target.

Copies reduce density while controlling active vertices

The residual already has small entries in columns with \(g(y)=1\). We distribute the remaining large entries among copies of their destinations. The detector incidence bound will limit the number of parents with active extra copies, and a weighted projection will relate the new transition to the residual.

Choose an integer \(D_0>K_*q\), and let \(\tau_0=q^{-1}/8\), \(\tau_1=q^{-1}/4\). Choose a rising cutoff \(\chi\) which is zero through \(\tau_0\) and one from \(\tau_1\) on. For \(0\leq j<D_0\) set \[ s_0(t)=1-(1-D_0^{-1})\chi(t),\qquad s_j(t)=D_0^{-1}\chi(t)\quad(j>0). \tag{21}\] They are nonnegative, sum to one, and have Lipschitz eighth roots on bounded ranges. Define \(U^{\max}(x)=\max_{1\leq i\leq k_*}U_i(x)\). Let \(\kappa\) be a rising cutoff, zero through \(1/4\) and one from \(1/2\) on, and put \[ r_0(x)=1,\qquad r_j(x)=\kappa(U^{\max}(x))\quad(j>0). \tag{22}\] On \(V^+=V\times\{0,\ldots,D_0-1\}\) write \(x_i=(x,i)\) and set \[ C^+(x_i,y_j)=r_i(x)D(x,y)s_j(D(x,y)/f). \tag{23}\] Order the copies first by their parent and then by their copy symbol.

Proposition 9 (One-stage density and active-count reduction). Choose the detector rank \(J\) large enough that \[ J\geq\frac{2C_\pi(D_0-1)k_*(k_*+1)}{\delta_0}, \tag{24}\] and then choose \(R_*\) as in Equation (16). The matrix \(C^+\) is strictly forward and substochastic, and \[C^+(x_i,y_j)\leq f/q, \qquad |\mathcal A(C^+)|\leq2|\mathcal A(C)|.\] If \(P^r\) is the matrix from \(V^+\) to \(V\) defined by \[ P^r(x_i,x)=r_i(x),\qquad P^r(x_i,z)=0\quad(z\ne x), \tag{25}\] then \[ C^+P^r=P^rD. \tag{26}\]

Proof. The row sum at \(x_i\) is \(r_i(x)\sum_yD(x,y)\leq1\). Strict forwardness follows from that of \(D\). Put \(t=D(x,y)/f\). If \(t<\tau_1\), then every routed entry is at most \(ft<f/q\). If \(t\geq\tau_1\), all fractions equal \(1/D_0\), so every entry is at most \(K_*f/D_0<f/q\). This proves the density bound.

If a nonprimary fraction \(s_j(D(x,y)/f)\) is positive on a positive entry, then \(D(x,y)/f>\tau_0>\varepsilon_{\rm good}\). Lemma 8 therefore implies \(g(y)<1\). Since \(\gamma\) equals one through \(1/2\), this gives \(U^{\max}(y)>1/2\), and hence \(r_j(y)=1\). The primary factor is always one. Thus for every positive entry of \(D\), \[ \sum_{j=0}^{D_0-1}s_j(D(x,y)/f)r_j(y)=1. \tag{27}\] For a zero entry no identity of its fractions is needed. Multiplying Equation (23) by \(P^r\) and using Equation (27) proves Equation (26) entrywise.

It remains to count active vertices. Put \[\mathcal O=\{y:U^{\max}(y)>1/4\},\qquad n'=|\mathcal A(C)|.\] Every \(y\in\mathcal O\) has, for some pilot step \(i\), more than \(J/(4C_\pi f)\) entries with \(G_i(x,y)>\delta_0f\), by Lemma 7. Counting these incidences over all steps and all active rows, and using Lemma 6, gives \[\begin{align*} |\mathcal O|\frac{J}{4C_\pi f} &\leq\sum_{i=1}^{k_*}\#\{(x,y):G_i(x,y)>\delta_0f\}\\ &\leq\frac{n'}{\delta_0f}\sum_{i=1}^{k_*}i =\frac{n'k_*(k_*+1)}{2\delta_0f}. \end{align*}\] The assertion is also valid when \(n'=0\), since then all matrices and detectors vanish. Thus \[|\mathcal O|\leq \frac{2C_\pi n'k_*(k_*+1)}{\delta_0J}.\]

Every active child has an active parent, by the support conclusion of Lemma 8. A nonprimary child active as a row has \(r_i(x)>0\), so its parent belongs to \(\mathcal O\). A nonprimary child active as a column has a positive extra routing fraction and, as shown above, its parent has \(U^{\max}>1/2\), so again belongs to \(\mathcal O\). Therefore \[|\mathcal A(C^+)|\leq n'+(D_0-1)|\mathcal O|\leq2n',\] where the last inequality is Equation (24). ◻

Remark 10 (Order of the one-stage constants). The bounds used to select \(k_*\) in Lemma 5 hold for every multiplier in \([0,1]\). Thus \(\theta_1,\theta_2,Q,k_*,K_*\) and \(D_0\) can be fixed before any detector rank. Next fix all finitely many correction, detector, and access-score threshold ladders, taking \(\delta_0,\delta_1\) below the correction support thresholds as required. Choose \(K\) large enough for the corresponding row estimates. Choose \(J\) by Equation (24), and then \(R_*\) by Equation (16). An incoming rank used later can be chosen after \(R_*\): its full-gate probability is controlled by the column bound in Lemma 7, independently of the auxiliary row gates. There is therefore no requirement to enlarge \(K\) again in response to that incoming rank. Subsequent estimator constants may depend on these already fixed ranks and thresholds.

Reward transport and the final deterministic target

The copy step controls transition entries and active vertices. We now iterate it and transport the terminal reward through the same weighted projections. The resulting reward will approximate the original acceptance probability, including probabilities strictly between zero and one.

Starting with \(C_0=S\) on \(V_0\), repeat the one-stage construction with \[f_l=q^{-l},\qquad V_{l+1}=V_l\times\{0,\ldots,D_0-1\},\qquad C_{l+1}=C_l^+.\] Let \(E_l,D_l,P_l^r\) denote the correction, residual, and projection at stage \(l\). Define matrices from \(V_l\) to \(V_0\) by \[ P_0^c=I,\qquad P_{l+1}^c=P_l^rP_l^c, \qquad B_0=I,\qquad B_{l+1}=P_l^r(I+E_l)B_l. \tag{28}\] The superscript \(c\) here denotes the cumulative projection, not a transpose. Let \(e\) be the base final-success reward and \(p_0=(I-S)^{-1}e\), so \(0\leq p_0\leq1\) by the model reduction.

Proposition 11 (Reward identity and a shallow target). For every stage, \[ C_lP_l^c=P_l^c-B_l(I-S). \tag{29}\] The vectors \(W_l=B_le\) satisfy \[ 0\leq W_l\leq1,\qquad W_l=P_l^cp_0-C_lP_l^cp_0, \tag{30}\] and their updates are \[ W_{l+1}(x_i)=r_i(x)\bigl(W_l(x)+(E_lW_l)(x)\bigr). \tag{31}\] If \(\bar x_l\) is the all-primary copy of a base vertex \(x\), then \[ 0\leq p_0(x)-W_l(\bar x_l) \leq |V_0|\,2^{-(H-1)l}. \tag{32}\] Consequently there is \(L=O(B)\) for which the last upper bound is strictly less than \(1/16\) for every base vertex \(x\). In particular, \(p_0(x)=0\) implies \(W_L(\bar x_L)=0\), and \(p_0(x)\ge1/2\) implies \(W_L(\bar x_L)>7/16\). Every full stage domain has size \(2^{O(B)}\) and admits \(O(B)\)-bit addresses.

Proof. The assertion at \(l=0\) is \(S=I-(I-S)\). The residual identity can be written as \[D_l=I-(I+E_l)(I-C_l).\] Using Equation (26) and then the inductive hypothesis, \[\begin{align*} C_{l+1}P_{l+1}^c &=P_l^rD_lP_l^c\\ &=P_l^r\bigl(P_l^c-(I+E_l)(P_l^c-C_lP_l^c)\bigr)\\ &=P_{l+1}^c-P_l^r(I+E_l)B_l(I-S)\\ &=P_{l+1}^c-B_{l+1}(I-S). \end{align*}\] This proves Equation (29).

Every \(B_l\) is nonnegative. Every \(P_l^r\) has at most one nonzero entry per row, of value in \([0,1]\), and thus \(P_l^c\) is substochastic. Apply Equation (29) to \(p_0\), using \((I-S)p_0=e\). This gives the identity in Equation (30). Its lower bound follows from \(B_le\geq0\); its upper bound follows from \(W_l\leq P_l^cp_0\leq1\). Equation (31) is Equation (28) applied to \(e\).

By Proposition 9, induction gives \[|\mathcal A(C_l)|\leq2^l|\mathcal A(S)|\leq2^l|V_0|, \qquad C_l(x,y)\leq q^{-l}.\] There are at most \(|\mathcal A(C_l)|\) columns containing any positive entry, so each row sum of \(C_l\) is at most \(|V_0|2^lq^{-l}=|V_0|2^{-(H-1)l}\). The all-primary projection has every row factor equal to one and hence \((P_l^cp_0)(\bar x_l)=p_0(x)\). The identity in Equation (30), together with \(0\leq P_l^cp_0\leq1\), now proves Equation (32).

Choose an integer \(L\) such that \((H-1)L>\log_2(16|V_0|)\). Since \(\log_2|V_0|=O(B)\), this is \(O(B)\). If \(p_0(x)=0\), nonnegativity forces \(W_L(\bar x_L)=0\). If \(p_0(x)\geq1/2\), the strict error bound gives \(W_L(\bar x_L)>7/16\). Finally, \[|V_l|=|V_0|D_0^l=2^{O(B)}\qquad(0\leq l\leq L),\] because \(D_0\) is fixed. Parent projection deletes one radix symbol, and appending a specified symbol is affine in the numerical ID. Thus these domain changes preserve the address convention required by the sampler and by the later implementation. ◻

The remainder of the proof will estimate the finite targets just constructed. In particular, it need only implement the \(k_*\) pilot and \(k_*\) actual updates at a stage and the reward update in Equation (31); it does not evaluate \(E_*\). The same error bound in Equation (32) gives any fixed inverse-polynomial accuracy by increasing the constant in \(L=O(B)\). Section 12 makes that choice together with the statistical and numerical precisions.

Stochastic atoms and their interfaces

The exact hierarchy uses matrix entries, correction products, and column detector expectations. We represent each of these by sample-dependent tables: bins represent entries of the transitions and residuals, correction tables represent the pilot and actual corrections, and score tables represent the detector trials. Each normalized table has an earlier preparatory table that supplies its denominator.

We first prove the exact conditional-mean identities. We then extend the same row formulas to computed raw arrays and prove the support and access bounds that hold even when those arrays are inaccurate. Finally we bound the effect of errors in the incoming gates and column multipliers. The later implementation supplies enumeration and pairing of the resulting ports.

Exact conditional-mean identities refer only to the true tables defined here. Applying the same formulas to approximate raw signals gives approximate tables; their normalization need not be unbiased.

Weighted pseudorandomness provides a related way to represent expectations with few random bits; see Braverman, Cohen and Garg (Braverman et al. 2020) and the error-reduction work of Cohen, Doron, Renard, Sberlo and Ta-Shma (Cohen et al. 2021). The tables below have exact conditional-mean identities. Their estimates use a common environment, so products and subsequent normalization are analyzed through joint error bounds rather than independent sampling.

Table domains and the normalization principle

Throughout this section \(s=8\). Fix a stage, write \(f=b_l\), and let \(G\) be one of its nonnegative raw matrix inputs. The hierarchy gives absolute constants \(M_G,K_G\) such that \[ \sum_y G(x,y)\le M_G,\qquad G(x,y)\le K_G f. \tag{33}\] These bounds include pilot and actual inputs and the residual \(D\); see Lemmas 6 and 8. For a correction \(F=fQ(G/f)g_0\), the additional input is \[ g_0(y)>0\quad\Longrightarrow\quad \#\{x:G(x,y)\ge\delta_1 f\}\le R_*/f, \tag{34}\] from Lemma 7.

A row task has start rate \(u\le f\), end scale \(a\), and, when needed, a usage rate \(t\le a\), all on the rate grid. Its domain is \(X_u(x)=Y_a(y)=1\), with distinct start and terminal channels. We use \(a\le K_f f\) for an absolute \(K_f\), and use \(a=f\) for corrections and detector scores. Here and in the estimation and compression sections, \(t\) denotes a sampling rate; the numerical clock \(t\) is introduced separately in Section 10. The true first raw signal is \[ d_y=G(x,y)/a. \tag{35}\] An optional second raw signal is \(\zeta\mathbf1_{\{Y_t(y)=1\}}q_y\), where \(q_y=q(x,y)\in[0,1]\) is a deterministic conditional mean and \(\zeta>0\) is a fixed sufficiently small constant. That mean will be specified separately for each normalized task.

All parameters of a table are common to the whole table: its stage and iteration, task type, channel roles, channel padding suffixes, rates \(u,a,t\), and subsequently its budget and numerical convention. They do not depend on the row or column at which the table is queried.

Write \(\operatorname{rk}_k\) for the \(k\)th largest value, with unlimited zero padding. For two arrays on the same padded index set, \[ |\operatorname{rk}_k(v)-\operatorname{rk}_k(w)| \le\sup_z|v_z-w_z|. \tag{36}\] Indeed, if every coordinate changes by at most \(\epsilon\), at least \(k\) coordinates at or above the old \(k\)th value remain at or above that value minus \(\epsilon\); apply this in both directions.

The normalization can be seen before choosing the row functions. Fix a pair \((x,y)\) on its conditioning slice, a deterministic nonnegative value \(k(d_y)\), and a random row gate \(R\in[0,1]\). Choose a preparatory factor \(h\) that is one wherever \(k\) is positive. If \[q(x,y)=\mathbb E[h(d_y)R\mid X_u(x)=Y_t(y)=1]\ge\tfrac12\] on that support, then \(k(d_y)R/q(x,y)\) has conditional mean \(k(d_y)\) there. A clipped denominator keeps the formula bounded at every other position and for computed inputs. The constructions below arrange a uniform positive lower bound on this mean while the row gate limits the number of nonzero retained entries. Corrections use an additional incoming gate, with the same normalization principle.

For real \(\alpha<\beta\), let \(\chi_{\alpha,\beta}\) be a fixed increasing switch, zero on \((-\infty,\alpha]\) and one on \([\beta,\infty)\). Choose it by rescaling the normalized integral of \(z^{s-1}(1-z)^{s-1}\) on \([0,1]\). Both \(\chi_{\alpha,\beta}^{1/s}\) and \((1-\chi_{\alpha,\beta})^{1/s}\) are Lipschitz: at either endpoint the vanishing function has a zero of order \(s\), and on the interior ordinary differentiation applies. Products of bounded functions with Lipschitz \(s\)th roots again have Lipschitz \(s\)th roots.

Every scalar saturation below uses fixed min/max clipping, \[\operatorname{clip}_{[\alpha,\beta]}(v) =\min\{\beta,\max\{\alpha,v\}\},\] which is \(1\)-Lipschitz. We specify its bounds when defining the function. Functions initially defined for nonnegative first entries are extended by zero below zero; each already vanishes below a positive support threshold. These conventions apply only to evaluation of the row functions. Ranks and access tests always use the full prescribed raw arrays before scalar clipping.

Here is one explicit set of margins used below. Let \(\vartheta>0\) be a fixed lower support threshold of a main row function, and put \(\rho=\operatorname{rk}_K(v)\) for its first raw row \(v\). Define \[\begin{align*} h_{\vartheta}(v) &=\chi_{\vartheta/2,\,3\vartheta/4}(v),\\ R_{\vartheta}(\rho) &=1-\chi_{\vartheta/1000,\,\vartheta/500}(\rho),\\ I_{\vartheta}(v,\rho) &=\chi_{\vartheta/8,\,\vartheta/4}(v) \bigl(1-\chi_{\vartheta/100,\,\vartheta/50}(\rho)\bigr). \tag{37}\end{align*}\] We take \(\vartheta=\theta_1\) for corrections, \(\vartheta=1/2\) for dense bins, and \(\vartheta=\delta_0\) for the detector score. For that detector we make the allowed hierarchy choices \[\eta_-=\delta_0/1000,\qquad \eta_+=\delta_0/500,\qquad \rho(t)=1-\chi_{\delta_0/1000,\,\delta_0/500}(t) =R_{\delta_0}(t).\] Here \(\rho\) denotes the cutoff in the hierarchy, and the same switch \(\chi\) is used in both occurrences. These choices satisfy \(0<\eta_-<\eta_+<\delta_0/3\). Thus the detector row gate here is exactly \(R_i\) in Equation (11). The detector does not require \(h_{\vartheta}\). The hierarchy’s choices may and do satisfy \(\delta_1<\theta_1/100\) and \(0<\delta_0<\delta_1\). Thus the correction auxiliary score \(I_{\theta_1}\) can be positive only when \(G(x,y)>\delta_1 f\).

Dense bins and preparatory normalization

We first split a raw entry into contributions at the rate-grid scales. At each scale a row gate bounds the number of retained entries, and an exact preparatory conditional mean corrects the loss caused by that gate.

Put \(q_*=2^H\), and let \(\chi=\chi_{1/2,1}\). For \(w\ge0\) define \[\begin{align*} \omega_0(w)&=\chi(w),\\ \omega_k(w)&=\chi(w/b_k)-\chi(w/b_{k-1})\qquad(k\ge1). \tag{38}\end{align*}\] The top bin has no preceding switch. The weights are nonnegative and sum to one for \(w>0\), since their partial sums telescope to \(\chi(w/b_k)\to1\). Hence \[ w=\sum_{k\ge0}w\omega_k(w). \tag{39}\] At scale \(a=b_k\), the bin in atom units is \[ k_a(d)=d\omega_k(ad). \tag{40}\] For \(k\ge1\) its support is contained in \(1/2<d<q_*\). Its two switch transition intervals are disjoint, and \[\chi(d)-\chi(d/q_*)=\chi(d)(1-\chi(d/q_*)),\] which proves directly that its \(s\)th root is Lipschitz. For \(k=0\), fix \(B_{\rm top}>\max\{K_G,1\}\) and evaluate the defining expression at \(\operatorname{clip}_{[0,B_{\rm top}]}(d)\). This makes the top-bin function bounded with Lipschitz root. It preserves every true value, since \(a=b_0=1\) and \(G(x,y)\le K_G f\le K_G\). For \(k\ge1\), retain the zero extension outside the support \((1/2,q_*)\); no upper clipping changes a zero bin value. Bins with \(a>2K_G f\) vanish on the true input, so a fixed \(K_f\ge2K_G\) suffices for the permitted scale range.

For the routed residual bin use instead \[ k_a^{(j)}(d)=k_a(d)s_j(ad/f). \tag{41}\] The routing fractions have bounded Lipschitz roots. Since \(a/f\le K_f\) and the bin is zero above a fixed argument range (or is top-bin clipped), these functions have uniformly bounded Lipschitz roots as well. The same zero-extension and clipping conventions apply to routed bins. Summing Equation (41) over scales gives the routed residual contribution exactly on every true input.

Choose \(R_{1/2}\) and \(h=h_{1/2}\) from Equation (37). For the true row set \[\rho_x=\operatorname{rk}_K \{G(x,z)/a:Y_a(z)=1\},\qquad R=R_{1/2}(\rho_x).\] Write \[ q_y=\mathbb E\bigl[h(d_y)R\mid X_u(x)=Y_t(y)=1\bigr], \qquad c(q)=\min\{1,\max\{1/2,q\}\}. \tag{42}\] The number \(q_y\) is the deterministic conditional mean at the indicated pair. The second raw signal carries the \(Y_t\) membership mask. In addition, \(q_y=0\) if \(d_y\le1/4\). The true preparatory table on the broad terminal domain \(Y_a\) is the unnormalized value \(h(d_y)R\). It uses only the first raw signal. Its conditional mean on the narrower slice is exactly the \(q_y\) in Equation (42). The usage rate \(t\) is part of the fixed task tuple: changing the slice on which this mean is taken would in general change the denominator. The broad preparatory domain and the narrower normalization slice must therefore be retained as separate parts of the interface.

Lemma 12 (Dense atom identity). For sufficiently large fixed \(K\), the table \[ \mathcal B_a(x,y)= \mathbf1_{\{Y_t(y)=1\}}\frac{k_a(d_y)R}{c(q_y)} \tag{43}\] on retained starts has bounded values and row support. On the slice \(X_u(x)=Y_t(y)=1\) its conditional expectation times \(a\) equals \(G(x,y)\omega_k(G(x,y))\). The same assertion holds for \(\mathcal B_a^{(j)}\) with the routing fraction included. Apart from its start membership, its true value uses only the terminal channel.

Proof. The row gate can fail to be full only if at least \(K\) sampled columns have \(G(x,z)/a>r\), where \(r=(1/2)/1000\) is fixed. There are at most \(M_G/(ar)\) such deterministic columns. Conditional on \(Y_t(y)=1\), the specified column contributes at most one hit and each other column has \(Y_a\)-hit probability at most \(C a\), by Lemma 4. Thus the expected count is at most \(1+CM_G/r\). Conditioning also on the independent source channel does not alter this bound. Markov’s inequality makes the failure probability at most \(1/8\) by a sufficiently large \(K\).

Where \(k_a(d_y)>0\), we have \(h(d_y)=1\). Therefore \(q_y=\mathbb E[R\mid Y_t(y)=1]\in[7/8,1]\) there, and the clipping does nothing. The conditional mean of Equation (43) is consequently \(k_a(d_y)\). Outside that support both sides vanish. If a row output is positive, its first entry exceeds \(1/2\) and its \(K\)th first rank is below \(1/1000\). Hence there are at most \(K-1\) positive entries. Boundedness follows from the clipped denominator. The routed fraction is deterministic at truth and obeys the same argument. ◻

For products of a correction with a deterministic bin contribution, we will use the following bounds.

Lemma 13 (Bounds for a fine product contribution). Let \(C\) be the transition matrix at the current stage, and let \(F\) be any finite pilot or actual correction there. If a matrix \(A_b\geq0\) on the same vertex set satisfies \(A_b\mathbf 1\leq\mathbf 1\) and \(A_b(x,y)\leq c_b b\) for some \(b>0\), then \[(FA_b)\mathbf 1\leq k_*\mathbf 1, \qquad (FA_b)(x,y)\leq k_*c_b b.\] If also \(A_b\leq C\), then \(FA_b\leq FC\). These estimates have constants independent of the bin scale \(b\).

Proof. Use \(F\mathbf 1\leq k_*\mathbf 1\) from Lemma 6: \((FA_b)\mathbf 1\leq F\mathbf 1\), and \((FA_b)(x,y)\leq c_b b\sum_zF(x,z)\). The last assertion follows by multiplying \(A_b\leq C\) on the left by the nonnegative matrix \(F\). ◻

Corrections with an incoming gate

Corrections also need a bound on incoming degree. Their deterministic column-support bound permits an additional incoming gate while keeping the probability that both gates are full bounded away from zero.

Consider \(F=fQ(G/f)g_0(y)\) and put \(a=f\), \(\vartheta=\theta_1\). Fix \(B_Q>\max\{K_G,\theta_2\}\) and define the atom cutoff \[Q_{\rm at}(d)=Q\bigl(\operatorname{clip}_{[0,B_Q]}(d)\bigr).\] It is bounded, has a Lipschitz \(s\)th root on the real line, and agrees with \(Q(d)\) for every true argument \(d=G(x,y)/f\). Its extension to negative arguments is zero. The exact hierarchy continues to use the uncapped \(Q\). Let \(R=R_{\theta_1}(\rho_x)\), \(h=h_{\theta_1}\), and \(I=I_{\theta_1}\) be the functions in Equation (37), where the true row rank is \(\rho_x=\operatorname{rk}_K\{G(x,z)/f:Y_f(z)=1\}\). On a true or computed raw table, form \(I(x,y)\) from that table’s own row entry and own rank. With \(\chi_{1/4,1/2}\) as above, define \[ T_y=1-\chi_{1/4,1/2}\left( \operatorname{rk}_{K_c}\{I(x,y):X_u(x)=1\}\right). \tag{44}\] Thus \(T_y=1\) through rank \(1/4\) and \(T_y=0\) from rank \(1/2\). This rank is taken over the fixed retained source set before any incoming port filtering; there is no gate-dependent change of its index set. The true correction preparation is \(h(d_y)R T_y\) on \(Y_a\). Define its conditional mean by \[ q_y=\mathbb E\bigl[h(d_y)R T_y\mid X_u(x)=Y_t(y)=1\bigr]. \tag{45}\] The number \(q_y\) is deterministic; as before, its second raw signal carries the \(Y_t\) membership mask. Define the normalized true correction using the clipped expression everywhere: \[ \mathcal F(x,y)=\mathbf1_{\{Y_t(y)=1\}} \frac{Q_{\rm at}(d_y)R T_y}{c(q_y)}\,g_0(y). \tag{46}\] Here \(q_y=0\) whenever \(d_y\le\theta_1/2\). Thus both kinds of second signal are supported above a fixed positive first-signal threshold. Choose \(\zeta\le10^{-6}\min\{\theta_1,\delta_0,1/2\}\), decreasing it later if a fixed compression margin requires this. The second signals are then uniformly tiny on the first-signal threshold scale.

Lemma 14 (Correction atom identity, including zero multipliers). For sufficiently large fixed \(K\) and then \(K_c\), the conditional mean of \(f\mathcal F(x,y)\) on \(X_u(x)=Y_t(y)=1\) is exactly \(F(x,y)\). On a true position with \(Q(d_y)>0\) and \(g_0(y)>0\), \[ q_y=\alpha(x,y):= \mathbb E[R T_y\mid X_u(x)=Y_t(y)=1]\ge3/4. \tag{47}\] No lower bound on \(q_y\) or \(\alpha(x,y)\) is required when \(g_0(y)=0\). The true gates and normalization use only the two role channels \(X,Y\); \(g_0\) is deterministic.

Proof. The row count argument in Lemma 12, with the correction thresholds, gives \(\Pr(R\ne1\mid\text{the two hits})\le1/8\). If \(g_0(y)>0\), every source with \(I(x',y)>0\) belongs to the deterministic set in Equation (34). Conditional on \(X_u(x)=1\), the number of sampled members of that set has expectation at most \[1+C uR_*/f\le1+C R_*.\] The bound remains valid after conditioning on the independent terminal hit; it ignores the auxiliary row gate and therefore needs no additional independence between rows. If fewer than \(K_c\) such sources are sampled, the \(K_c\)th auxiliary rank is zero and \(T_y=1\). Choose \(K_c\ge8(1+C R_*)\) to make \(\Pr(T_y\ne1\mid\text{the two hits})\le1/8\). The union bound, without independence of the gates, gives \(\mathbb E[RT_y\mid\text{the two hits}]\ge3/4\).

Since \(h(d_y)=1\) on \(Q\)’s support, Equation (47) holds and \(c(q_y)=q_y\) there whenever \(g_0(y)>0\). The claimed conditional identity follows by division by this exact conditional mean. If \(Q(d_y)=0\) or \(g_0(y)=0\), both sides of the identity are zero. In the latter case the pre-incoming quantity \(Q_{\rm at}(d_y)R/c(q_y)\) is still a well-defined bounded comparison target, even if the unclipped mean vanishes. ◻

Detector scores

The unnormalized detector table is \[ \mathcal S_i(x,y)=\phi(G_i(x,y)/f) R_{\delta_0}\!\left(\operatorname{rk}_K \{G_i(x,w)/f:Y_f(w)=1\}\right) \quad\text{on }X_f(x)=Y_f(y)=1. \tag{48}\] Its column trial is the \(J\)th largest such score over retained starts. By the identical cutoff choice, its conditional mean given \(Y_f(y)=1\) is exactly \(U_i(y)\) from Equation (12). The trial is an inverse rank access to a row-bounded table and does not require an incoming degree bound.

Actual ports and estimated assembly

The conditional-mean identities now give the true contributions. A product estimate will query a correction at \((x,z)\) and then a bin at \((z,y)\). Bounded outgoing degree limits the possible midpoints; bounded incoming degree is needed when errors at the new source \(z\) are summed over old sources \(x\). We therefore specify access relations as well as nonzero table supports. The access relations may contain entries of weight zero.

Passing through one copy projection is partitioned by its fixed source and destination symbols. Within each part the inherited hash argument appends that symbol before the inherited suffix. Lemma 4 then gives the same distribution as the unpadded channel. The task tuple records these fixed symbols and suffixes. By definition, neither a right factor’s task parameters nor a column vector’s task parameters retain an earlier source address. The environment supplied to the table is common to all its entries. The vertices \(x,y\) are arguments of that table, rather than additional parameters. Thus two paths that discover the same entry in that environment query the same value; the path used to discover it is not part of its mathematical definition.

We will apply these row functions to arrays represented by a finite list of exceptional values and one common value at every omitted endpoint. Call that common value the default. Adjoin arbitrarily many dummy coordinates with the same default, and include them in every rank; they are never matrix entries or ports. For the true row arrays, dummy values are zero. Section 8 constructs the finite-list representations. The elementary rank convention below lets us prove support and access bounds before choosing that construction.

Remark 15 (Ranks of arrays with a common default). Suppose a row array equals a common default outside a finite list, and arbitrarily many dummy coordinates have that same default. Its \(k\)th rank can be computed from the listed coordinates and \(k\) copies of the default, with the stipulated zero padding. Indeed, every omitted coordinate merely supplies another copy of a value already available with multiplicity at least \(k\). Additional listed real coordinates equal to the default therefore do not change the rank. The same observation applies to a rank of entry sizes when each omitted entry has the same default vector.

This is an equality of ranks for one prescribed full array. The dummy coordinates are never matrix entries or ports. A common componentwise clamp of the entries and their default preserves this representation: an omitted position still equals the clamped default. No assertion that different calculations select identical lists or ports is needed.

Lemma 16 (Support margins and exposed row classes). Suppose a bounded nonnegative main function \(k(v)\) vanishes for \(v\le\vartheta\). Its product \(k(v)R_{\vartheta}(\rho)\) and the preparatory product \(h_{\vartheta}(v)R_{\vartheta}(\rho)\) have at most \(K-1\) nonzero real entries in a row. They admit a common exposed class \(A\), of at most \(K-1\) real entries, containing their supports and satisfying \(I_{\vartheta}=1\) throughout \(A\). These assertions hold for arbitrary real raw estimates, without a sign or accuracy hypothesis.

The auxiliary score \(I_{\vartheta}\) also has at most \(K-1\) nonzero real entries and has an exposed row class with a bounded-support smooth witness equal to one on that class.

Proof. For the main and preparatory products a nonzero entry has \(v>\vartheta/2\) and \(\rho<\vartheta/500\). There cannot be \(K\) such entries, since then the \(K\)th rank would exceed \(\vartheta/2\).

For an explicit class choose consistently computed approximations \(\widetilde v,\widetilde\rho\) with absolute errors at most \(\eta=\vartheta/10000\), and set \[ A=\{y:Y_a(y)=1,\quad \widetilde v_y>3\vartheta/8,\quad \widetilde\rho<\vartheta/200\}. \tag{49}\] Every positive main or preparatory entry belongs to \(A\). An entry in \(A\) satisfies \(v_y>(3/8-1/10000)\vartheta>\vartheta/4\) and \(\rho<(1/200+1/10000)\vartheta<\vartheta/100\), so its witness \(I_{\vartheta}\) equals one. Its entry threshold exceeds its rank threshold, proving \(|A|\le K-1\).

The support of \(I_{\vartheta}\) requires \(v>\vartheta/8\) and \(\rho<\vartheta/50\), again giving the same row bound. To access this score itself, the class \[A_I=\{y:Y_a(y)=1,\quad \widetilde v_y>\vartheta/10,\quad \widetilde\rho<\vartheta/40\}\] contains its support and has at most \(K-1\) entries. A smooth unit witness on \(A_I\) is \[J_{\vartheta}(v,\rho)= \chi_{82\vartheta/1000,\,90\vartheta/1000}(v) \bigl(1-\chi_{26\vartheta/1000,\,27\vartheta/1000}(\rho)\bigr).\] The approximation margins imply \(A_I\subseteq\{J_{\vartheta}=1\}\). Moreover \(82/1000>3(27/1000)\), so this witness also has bounded row support and the required entry versus rank separation. The witness is an analytical output; its support need not be broadened recursively.

These arguments use the rank of the prescribed entire raw array. In particular, if at least \(K\) dummy positions have a common virtual default \(\beta\), then \(\rho\ge\beta\). A real position still equal to that default cannot meet either exposed class: the required lower bound on \(v=\beta\) exceeds the required upper bound on \(\rho\). Thus omitted default positions do not create unlisted ports. ◻

The fixed-accuracy approximations in this lemma are part of the access interface, not tests for equality of exact reals. The implementation must use the same approximation and tie conventions whenever it queries the same table. The strict margins prove correctness for every such choice. Auxiliary-only access classes require a row bound but no incoming bound. In particular, different computed versions may expose different zero-weight entries. The assertions are applied to each version’s own final raw array and repeatable approximation convention; they do not require those classes to coincide. Accuracy of the raw array is used later in the error estimates, not in this support argument.

An estimated table applies the preceding bounded row functions to its final raw row arrays. In a normalized task its second input is a computed value approximating \(\zeta\mathbf1_{\{Y_t(y)=1\}}q_y\): the calculator divides that input by \(\zeta\) and applies \(c\). It does not launch another averaging operation after selecting an endpoint. Its auxiliary incoming scores are formed from its own final first-signal rows, and its incoming gate is Equation (44) for those scores. A preparatory task is an earlier, separate task; its raw arrays, scores, and gates are its own, and its target is the unnormalized preparatory table specified above. Corrections additionally multiply by a clipped vector estimate of the prescribed \(U\)-cutoff \(g_0(y)\), queried with a common broad \(Y_f\)-domain tuple. The unnormalized preparations \(hR\) and \(hRT\) remain on their broad \(Y_a\) domain; their usage parameter \(t\) specifies the conditional slice on which they are averaged. The resulting second raw signal and the normalized main output carry the \(Y_t\) mask. Dense bins and their preparations and detector tables have no additional incoming multiplier. Because \(c\) takes values in \([1/2,1]\), every such pre-incoming row function has a globally Lipschitz \(s\)th root as a function of its saturated first entry, first rank, and scaled second entry. The constants may depend on the fixed \(\zeta\), but not on the stage, rates, or budget.

For clarity, Table 1 records the thresholds needed by the compression argument. For a pre-incoming row function \(H\), its entry factor vanishes through \(e_H\), and its row gate vanishes from rank \(g_H\) onward. Every listed pair satisfies \(3g_H<e_H\), including the broadest witness, for which \(82\vartheta/1000-3(27\vartheta/1000)=\vartheta/1000>0\). All gates are full in a neighborhood of rank zero. The correction preparation in this table is its pre-incoming factor \(hR\); its incoming gate is assembled afterward.

Entry and row-gate thresholds. The entry factor vanishes through \(e_H\), and the row gate vanishes from rank \(g_H\) onward. Here \(\vartheta\in\{1/2,\theta_1,\delta_0\}\) as appropriate.
Pre-incoming row function \(e_H\) \(g_H\)
Dense bin, including routed bins \(1/2\) \(1/1000\)
Dense preparation \(hR\) \(1/4\) \(1/1000\)
Correction clipped-ratio function \(\theta_1\) \(\theta_1/500\)
Correction preparation \(hR\) \(\theta_1/2\) \(\theta_1/500\)
Detector score \(\delta_0\) \(\delta_0/500\)
Auxiliary score \(I_{\vartheta}\) \(\vartheta/8\) \(\vartheta/50\)
Witness \(J_{\vartheta}\) \(82\vartheta/1000\) \(27\vartheta/1000\)

The bounded first-entry ranges and gate transition intervals already specified give fixed sensitivity ranges for these functions. If \(w\) is a computed second raw entry, its only use is through \(c(w/\zeta)\); clipping \(w\) to \([\zeta/2,\zeta]\) leaves this quantity unchanged. Its denominator stays in \([1/2,1]\) even for negative or inaccurate inputs. Usage masks only remove entries and do not change any of these support or continuity bounds.

Lemma 17 (Port contract). All main and preparatory row tables, auxiliary scores, and the smooth witnesses just constructed have bounded values and bounded row support. Every row access class used for a main or preparatory output, or for access to its auxiliary score, has bounded size and a smooth witness equal to one on every exposed entry, including exposed entries of weight zero. A witness itself is needed as an analytical row function, without an additional requirement to expose its support through a further witness. For a correction or its preparatory table let \(A\) be its pre-incoming class from Equation (49). Retain an exposed entry \((x,y)\in A\) only when \[ \#\{x':X_u(x')=1,\ (x',y)\in A\}\le K_c. \tag{50}\] This deletion preserves all table values and gives at most \(K-1\) outgoing and at most \(K_c\) incoming exposed entries.

These assertions hold separately for every computed version, provided that its values, ranks, scores, and access classes use the same final raw arrays and common task tuple. Access ports denote distinct entries: the implementation must expose each pair at most once and pair its incoming and outgoing ports consistently. Access to auxiliary scores alone has only the asserted row bound.

Proof. The row and witness assertions are Lemma 16; clipping and multiplication by bounded factors do not enlarge supports. If a column has more than \(K_c\) entries in \(A\), its auxiliary score is one at all those entries, so its \(K_c\)th auxiliary rank is one. Equation (44) makes \(T_y=0\) there. Consequently every value deleted by Equation (50) was already zero. The surviving relation has the stated two degree bounds even when zero weights are retained. The argument applies to arbitrary raw arrays, so it applies equally to the estimated and true versions.

The cap removes every precursor port at an overfull column; it does not choose \(K_c\) of them. It is applied even when the particular entry has weight zero. Every surviving port still lies in \(A\), so its pre-incoming witness \(I\) equals one regardless of a zero usage mask, incoming gate, or column multiplier. The broader auxiliary class \(A_I\) is used to enumerate scores for an incoming rank. It is not the capped correction relation used as a paired left factor in a product.

The classes are relations on row and column IDs, rather than multisets of paths. Choosing one port per relation element gives the stated combinatorial interface. Establishing its single-address enumeration and port pairing within the space bound is an implementation obligation; the degree argument does not itself grant random access to both IDs. ◻

For later reference, the full interface supplied here is the following. Each task has fixed common parameters; first ranks include the prescribed dummy coordinates; exposed ports are real distinct entries and lie in a unit witness; sparse main and preparatory ports obey the actual incoming cap; and values, auxiliary scores, incoming ranks, and column multipliers are all queried at the same supplied environment and tuple. The path by which an entry is discovered must not change that environment or tuple for its assembly. Only completed earlier output types, or stated fixed local functions and aggregations of those types, may be used as raw inputs of a later task. No prefix still waiting for assembly is such an output. These requirements distinguish finite-support analytical bounds from the access properties used in the subsequent implementation. Completing a raw array here means defining its entries from the allowed earlier outputs, rather than storing the whole array at once. Its auxiliary scores and precursor classes are defined before its incoming gate and cap, so their definitions do not call that same gate or cap.

These contracts apply directly to each computed array after compression and clamping. The later comparison in Lemma 37 preserves relevant row-function values under its stated common-increment hypothesis. It need not preserve the finite-accuracy access descriptors of two calculations. Each calculation uses its own consistent classes, and Lemma 17 gives their support and degree bounds separately.

Integrated losses and assembly bounds

It remains to control how errors in the row functions, incoming scores, and detector vectors affect an assembled table. The incoming degree bound lets us sum these errors by columns without multiplying by the number of possible source vertices.

For a real row array \(Z\) define \[ \|Z\|_{u,a;s}^s= \frac1N\mathbb E_\sigma\sum_x\frac{X_u(x)}u \sum_{y:Y_a(y)=1}|Z(x,y)|^s, \qquad \mathfrak d_{u,a}(H,H')= \|H^{1/s}-(H')^{1/s}\|_{u,a;s}. \tag{51}\] Here \(\sigma\) has the product law of the rank channels and independent auxiliary bits specified in Section 3. All table values are extended by zero outside their membership domain. The denominator \(u\) is the nominal rate, and actual-rate comparability in Lemma 4 absorbs the resulting absolute constants. For a vertex array on a channel with rate \(u\), write \[ \|v\|_{u;s}^s= \frac1N\mathbb E_\sigma\sum_x\frac{X_u(x)}u|v(x)|^s; \tag{52}\] the indicated role channel replaces \(X\) when necessary. These are also the comparison norms for two sample-dependent versions. All required pre-incoming functions and smooth support witnesses are included among the row outputs whose discrepancies are to be bounded.

Bounded row support and range imply a uniform bound on these row-output norms, since \(|V_l|\le N\) and \(\mathbb E X_u/u=O(1)\). Bounded vertex arrays have the analogous uniform bound. Ordinary error is controlled by root error on bounded nonnegative outputs, since \[ |z-w|\le s C^{(s-1)/s}|z^{1/s}-w^{1/s}| \quad(0\le z,w\le C). \tag{53}\]

Lemma 18 (Assembly comparisons). Compare two versions with the same environment, rates, and task tuple. For a sparse normalized output write \(H=P T g\) and \(H'=P'T'g'\), where \(P,P'\) are the pre-incoming clipped-ratio row functions (including their usage masks), \(I,I'\) are their auxiliary scores, and \(g,g'\) are their column multipliers. If \(g,g'\) are the prescribed smooth cutoffs of bounded vertex arrays \(U_i,U_i'\), then \[ \mathfrak d_{u,f}(H,H')^s \le C\mathfrak d_{u,f}(P,P')^s +C\mathfrak d_{u,f}(I,I')^s +C\frac fu\sum_i\|U_i-U_i'\|_{f;s}^s. \tag{54}\] For a sparse preparation omit the last term and use its unnormalized pre-incoming row function. The corresponding dense and detector row functions have no incoming or column-multiplier assembly cost.

If \(V_y,V_y'\) are the detector trials formed from \(\mathcal S_i,\mathcal S_i'\) by their column order statistic, then \[ \|V-V'\|_{f;s}^s \le C\mathfrak d_{f,f}(\mathcal S_i,\mathcal S_i')^s. \tag{55}\]

Proof. All factors are bounded. On the union \(E\) of the two final nonzero supports, the triangle inequality for the products of \(s\)th roots gives \[|H^{1/s}-(H')^{1/s}| \le C|P^{1/s}-(P')^{1/s}| +C|T^{1/s}-(T')^{1/s}| +C|g^{1/s}-(g')^{1/s}|.\] Outside \(E\) the left side is zero. The pre-incoming discrepancy may be charged on its entire row domain; it needs no incoming cap. Each column has at most \(2K_c\) edges in \(E\), by Lemma 17. By Equation (36), the root-Lipschitz incoming cutoff, and Equation (53), \[|T_y^{1/s}-(T_y')^{1/s}|^s \le C\sup_{x:X_u(x)=1} |I(x,y)^{1/s}-I'(x,y)^{1/s}|^s.\] Consequently its integrated contribution is at most \[\frac C{Nu}\mathbb E\sum_{y:Y_f(y)=1} \sum_{x:X_u(x)=1} |I(x,y)^{1/s}-I'(x,y)^{1/s}|^s,\] which is \(C\mathfrak d_{u,f}(I,I')^s\).

The maximum over a fixed number of \(U_i\) values is Lipschitz, and the multiplier cutoff has Lipschitz root. Thus the last discrepancy is bounded in \(s\)th power by \(C\sum_i|U_i(y)-U_i'(y)|^s\). Summing over \(E\) costs at most \(2K_c/u\) per retained column. The vertex metric has weight \(1/f\) instead, giving precisely the factor \(f/u\) in Equation (54). All columns in this comparison are in the common \(Y_f\) domain, including those whose final values carry a narrower usage mask. The common-tuple condition ensures that the vertex discrepancy at \(y\) is the same quantity for every old start. Explicitly, if \(m_E(y)\) counts the edges of \(E\) entering \(y\), then for each \(i\) its contribution obeys \[\frac1{Nu}\mathbb E\sum_{y:Y_f(y)=1} m_E(y)|U_i(y)-U_i'(y)|^s \le 2K_c\frac fu\|U_i-U_i'\|_{f;s}^s.\] This is the column-counting step; it uses no independence between the two row tables and the vector estimates.

For a detector trial, Equation (36) bounds the trial discrepancy by the supremum of the input score discrepancies over retained starts. Its \(s\)th power is at most the sum of their \(s\)th powers. The factor \(Y_f(y)/f\) in the vertex metric is exactly the factor obtained by reversing the two sums in the row metric with \(u=a=f\). Equation (53) gives Equation (55). ◻

The required task types are therefore the dense bins \(\mathcal B_b\) of \(C_l\), the routed bins \(\mathcal B_b^{(j)}\) of \(D_l\), detector scores \(\mathcal S_i\), and the pilot and actual correction tables \(\mathcal F\), together with one earlier unnormalized preparation for each normalized type. Their true conditional-mean identities are Lemmas 12 and 14. Their computed versions use the same bounded row functions and the assembly contract above. The next sections bound their discrepancies through the additive schedule and implement their accesses without storing two unrestricted vertex addresses.

Conditional averaging in the shared environment

The atom identities use conditional expectations on slices that retain specified source and terminal vertices. To approximate these expectations, we need a walk that mixes on each slice. We construct such a walk using the channels of Section 3, then express its steps by transformations that can be specified before the retained identifiers are known. Validation tests those identifiers when their vertices are visited.

The identifiers are fixed when the conditional space is defined. We condition on one retained identifier per conditioned channel; additional independent bit coordinates may be varied or frozen as specified below.

The full environment law already includes the frame restriction on each channel. A conditional slice additionally imposes the indicated retention events and fixes any designated frozen bits. Its averaging operator and \(L_s\) norms use the normalized probability law on that slice.

The only external group-theoretic inputs are property (T) of \(\mathrm{SL}_H(\mathbb Z)\) for \(H\ge3\) and relative property (T) of \[(\mathrm{SL}_2(\mathbb Z)\ltimes\mathbb Z^2,\mathbb Z^2).\] We use Shalom’s quantitative statements: (Shalom 1999, Theorem 2.6, p. 156) for the linear group and (Shalom 1999, Theorem 2.1, p. 152) for the relative property. We derive the affine finite-action gap and the conditional sampler from these inputs below.

Lemma 19 (A uniform affine spectral gap). Let \(G_H=\mathrm{SL}_H(\mathbb Z)\ltimes\mathbb Z^H\), with fixed \(H\ge3\). There is a fixed finite symmetric generating multiset and a lazy rational random walk on it whose action on every finite transitive \(G_H\)-set has a spectral gap bounded below by a positive constant depending only on \(H\).

Proof. Let \(\mathcal S\) consist of the elementary matrices with a single \(\pm1\) off the diagonal, acting linearly, and the translations by \(\pm e_i\), where \(e_i\) are the standard basis vectors of \(\mathbb Z^H\). This is a fixed finite symmetric generating set of \(G_H\). For a unitary representation \(\rho\) and a vector \(\xi\), put \[\delta=\max_{g\in\mathcal S}\|\rho(g)\xi-\xi\|.\] We first bound the distance from \(\xi\) to the invariant subspace.

For each coordinate pair \(\{i,j\}\), embed the two-dimensional affine group on those coordinates, and let \(P_{ij}\) be the orthogonal projection onto the subspace fixed by its translations. This subspace and its orthogonal complement are invariant under that pair group, because its translation subgroup is normal. The eight generators in (Shalom 1999, Theorem 2.1, p. 152) belong to \(\mathcal S\). Applying their relative Kazhdan constant \(1/10\) on the orthogonal complement gives \[\|\xi-P_{ij}\xi\|\le10\delta.\] This is the projection argument in the proof of (Shalom 1999, Corollary 2.3, p. 154).

The projections \(P_{ij}\) commute: they are limits of averages of commuting coordinate translations. Their product \(P_{\rm tr}\) projects onto the subspace fixed by all translations. With \(q=\binom H2\), telescoping the product gives \[\|\xi-P_{\rm tr}\xi\| \le\sum_{i<j}\|\xi-P_{ij}\xi\| \le10q\delta.\] Since the full translation subgroup is normal, \(P_{\rm tr}\) commutes with every \(\rho(g)\). Its range carries a representation of \(G_H/\mathbb Z^H=\mathrm{SL}_H(\mathbb Z)\). Let \(P_{\rm inv}\) project onto the full \(G_H\)-invariant subspace, and let \(\kappa_H>0\) be the Kazhdan constant for the elementary generators from (Shalom 1999, Theorem 2.6, p. 156); adjoining inverse generators preserves the bound. Applying that constant on the range of \(P_{\rm tr}\) gives \[\|P_{\rm tr}\xi-P_{\rm inv}\xi\|\le\kappa_H^{-1}\delta, \qquad \|\xi-P_{\rm inv}\xi\|\le C_H\delta, \quad C_H=10q+\kappa_H^{-1}.\] Thus the same generating set controls distance to invariants in every unitary representation, with a constant depending only on \(H\).

Now let \(\rho\) be the permutation representation of a finite transitive action, with the uniform probability measure, and put \(m=|\mathcal S|\). The rational Markov operator \[K_H=\frac12I+\frac1{2m}\sum_{g\in\mathcal S}\rho(g)\] is self-adjoint, and its spectrum lies in \([0,1]\). The invariant functions are exactly the constants. For \(f\) orthogonal to constants, the preceding inequality and the Dirichlet identity give \[\langle f,(I-K_H)f\rangle =\frac1{4m}\sum_{g\in\mathcal S}\|\rho(g)f-f\|_2^2 \ge\frac1{4mC_H^2}\|f\|_2^2.\] This proves the uniform spectral gap. ◻

Lemma 20 (Conditional averaging). Consider the fixed product of rank channels, conditioned on retention of one specified identifier per conditioned channel. Include any designated set of independent uniform bits, with length padded to a multiple of \(H\), and freeze any remaining independent bits. There is a rational Markov averaging operator \(K\) on this conditional space such that its stationary distribution is the indicated uniform conditional law and, for each fixed \(2\le s<\infty\), \[\|K^d-\Pi\|_{L_s\to L_s}\le C_s2^{-c_sd} \qquad(d\ge0).\] Here \(\Pi\) is conditional expectation over the variables being averaged. The constants are independent of the field size, retention side lengths, bit lengths, indicated identifiers, and frozen bits.

Proof. On a channel conditioned at \(y\) with side \(L\), use coordinates \((F,w)\), where \(F\) is its field frame and \(w=Av(y)\in\{0,\ldots,L-1\}^H\). The constant column is uniquely determined by \(F,w,y\). As observed above, \(F\) is uniform over frames and \(w\) is independent uniform in its box. Identify this box with \((\mathbb Z/L\mathbb Z)^H\). For \((M,c)\in G_H\) act by \[(F,w)\longmapsto(MF,Mw+c),\] using reduction modulo \(P\) on the first coordinate and modulo \(L\) on the second.

This action is transitive. Reduction of \(\mathrm{SL}_H(\mathbb Z)\) contains the elementary generators of \(\mathrm{SL}_H(\mathbb F_P)\). That finite group acts transitively on ordered frames of length \(r<H\): extend either frame to a basis and adjust a complementary column to make the change-of-basis determinant one. Once the frame has been moved, a translation adjusts \(w\) to any specified value. The two moduli need not agree. On an unconditioned channel use a public identifier and side \(P\), giving the same description.

If the free bit length is \(Hq\), identify its strings with \((\mathbb Z/2^q\mathbb Z)^H\) and use the affine action on that coordinate alone. It is transitive. All currently free bit blocks can be concatenated into this one component; frozen bits are not transformed. There are a bounded number of components. Apply the lazy walk from Lemma 19 to a uniformly selected component. If there are \(J_{\rm comp}\) components and the single-component gap is at least \(\gamma_H\), the product walk has gap at least \(\gamma_H/J_{\rm comp}\). Here \(J_{\rm comp}\) is bounded independently of the field and bit lengths. The walk preserves the desired product measure.

Writing its gap as \(\gamma>0\), we have \(\|K^d-\Pi\|_{2\to2}\le(1-\gamma)^d\) and \(\|K^d-\Pi\|_{\infty\to\infty}\le2\). The diagonal case of the Riesz–Thorin interpolation theorem, already covered by Riesz (Riesz 1927, Theorems II and II\('\), p. 472), gives \[\|K^d-\Pi\|_{s\to s} \le 2^{1-2/s}(1-\gamma)^{2d/s},\] which has the required form. For Thorin’s complex extension, see (Thorin 1948, Introduction, §0.1; Russian translation, pp. 43–45). ◻

Generating paths without retaining their endpoints

The coordinates in the preceding proof used the indicated identifier. The next lemma is the implementation form of the same averaging, and also supplies changes of variables for later error estimates.

Lemma 21 (Validated walk variants). A length-\(d\) word in the walk of Lemma 20 can be expanded into a bounded-alphabet family of variants described by \(O(d)\) bits such that:

  1. each fixed variant acts bijectively and measure-preservingly on the full environment before any endpoint tests;

  2. its transformation is independent of the indicated identifiers;

  3. for each walk word and each environment in the required conditional slice, exactly one variant passes the slice and wrap tests at the indicated identifiers, and this variant gives the conditional walk transformation;

  4. the sum of valid variants, each with the probability of its original walk word, is exactly the conditional average.

The tests can be performed when their indicated vertices are visited. In the admissible implementation range, field elements and identifiers have \(O(B)\) bits, the total auxiliary bit string has \(O(B)\) bits, and \(d=O(B)\). In that range the transformations and tests use \(O(B)\) shared scratch space, and the retained description of a variant uses \(O(d)\) bits.

Proof. Consider a fixed affine generator \((M,c)\) on a channel with retention side \(L\). Enumerate the transformations \[ A'=MA+(c+Lz)e_1^{\mathsf T}\pmod P, \tag{56}\] where \(e_1\) selects the constant column and \(z\) ranges over a fixed finite box of integer vectors. That box contains all possible wraps: for \(w\in\{0,\ldots,L-1\}^H\), every coordinate of \(Mw+c\) is bounded in absolute value by a fixed multiple of \(L\) plus a fixed constant. Since \(L\ge1\) and the generators are fixed, one bound for \(z\) works for every side length.

Each transformation in Equation (56) is invertible. Left multiplication preserves the frame condition, and the added term changes only the constant column. It therefore preserves uniform measure on the full channel. It does not refer to the indicated identifier.

At that identifier, let \(w=Av(y)\) in its integer representatives. Check first that \(w\) is in the retention box, and then that \(Mw+c+Lz\) lies in the same box coordinatewise, using integer arithmetic before reduction modulo \(P\). The unique valid wrap has coordinates \(z_i=-\lfloor(Mw+c)_i/L\rfloor\). For this vector the evaluation of \(A'\) is the representative of the intended affine action modulo \(L\). Applying these checks at every step gives the unique valid wrap sequence for a whole word.

The construction applies separately to the bounded number of channels. On a channel of side \(P\) the ordinary field reduction suffices; alternatively, the same wrap validation may be used at its public point. The free-bit torus has ordinary invertible affine operations and requires no unknown endpoint. Frozen bits remain fixed. Thus a word has only constantly many generator and wrap choices per step, giving the \(O(d)\)-bit description.

For each word and initial environment in the slice, exactly one variant survives. Its weight must be the weight of that word. In particular, the sum is not divided by the number of guessed wraps. This proves the asserted exact identity. All matrices have fixed dimensions and \(O(B)\)-bit entries; evaluating a bounded-degree polynomial or a word’s individual steps needs \(O(B)\) arithmetic scratch. A test depending on a vertex can be delayed until that vertex is current, with the required matrix state regenerated from the recorded transformations. Theorem 58 accounts for reversible restoration and the simultaneous space used when these operations are nested. ◻

Remark 22 (Effective environment coordinates). Conditional averaging is expressed in the channel and bit coordinates of the environment supplied to that invocation. A fixed variant from Lemma 21 is a measure-preserving change of variables on the full environment; its retention and wrap tests remain explicit masks. Thus free and frozen bits always refer to the supplied coordinates, which may differ from those before the change of variables.

Remark 23. Two distinct uses of the variants should be kept separate. For centered-noise estimates, one first recognizes the sum of valid variants as the conditional walk and applies Lemma 20. For crude bounds and changes of variables, each fixed variant is a measure-preserving bijection, and one may then pay for their \(2^{O(d)}\) number. Counting variants before a spectral estimate would needlessly weaken the useful contraction.

Additive estimation and conditional prefix bounds

The bounded stochastic tables of Section 5 represent the exact hierarchy through their conditional means. In this section, earlier estimates are input tables and vectors with those interfaces. We define their sampled raw prefixes and derive bounds under the corresponding child-error hypotheses. Two errors arise: the conditional mean can differ from the deterministic target, and a finite averaging walk leaves noise around that mean. We bound them separately.

These prefixes are intermediate arrays. Final row outputs result from applying the atom functions to completed, compressed rows; Section 9 completes the recursive definition, including its endpoint-grouping rule.

The scale and source-domain comparisons first determine the accuracy needed from each child. We then average differences between consecutive child budgets. Their conditional means telescope, while their small absolute norms allow finite conditional walks to suppress the noise. The resulting prefix estimates are the input to the compression analysis in Section 8. That section proves an error reserve under exact endpoint grouping; Section 9 constructs the comparison calculation to which the reserve will apply.

Tasks and dependency order

A row task specifies its input matrix, stage, iteration label, start rate \(u\), end scale \(a\), optional usage rate \(t\le a\), channel roles, and fixed padding suffixes. These parameters are common to its whole table. A vector task specifies the analogous vertex domain. Its argument is a vertex, rather than an additional parameter retained by another task. Each task has a nonnegative integer statistical budget \(M\); \(m\) will denote a child budget. These budgets are distinct from the numerical digit places introduced in Section 10. The ordinary and root discrepancy norms are defined in Equations (51) and (52). We always take \(s=8\).

The source weight \(X_u(x)/u\) has bounded expectation, and the factor \(1/N\) normalizes the sum over source vertices. Endpoint sums remain explicit: bounded row supports or incoming degrees, as appropriate, will control them. A bound at one specified start consequently costs a factor \(N\) in moment power; Lemma 71 makes and pays this conversion at the end of the proof.

Here is a dependency order, with only a fixed number of task types at each stage. First come the bin tasks for \(C_l\). At each pilot step come its input-score task \(\mathcal S_i\), the detector vector \(U_i\), and then the pilot correction task. Next come the actual correction tasks in iteration order, the routed bin tasks for \(D_l\), and the reward update. Place the unnormalized preparation immediately before each normalized task of the same input. The raw first channels of both these tasks refer only to earlier inputs in this order; the normalized task has in addition its preparation as the source of its second channel.

For a correction task or its preparation, first complete its raw arrays and pre-incoming scores, then form its incoming gate from those scores. This assembly uses the current task’s completed arrays; every recursive input to their raw calculation is an earlier task in the order just specified.

Raw sample terms

For a fixed row and endpoint let \[\mathcal C_{xy}^{u,a}=\{X_u(x)=1,\ Y_a(y)=1\}.\] We use conditional means and conditional walks on this slice. Write \(\mathcal B_b\) for a normalized bin table and \(\mathcal B_b^{(j)}\) for a bin with destination routing fraction already included; \(\mathcal F\) denotes a normalized correction table. Their values are in atom units and are zero off their specified membership domains. In all formulas below put \(\tau=\min(a,b)\).

The direct dense and correction–dense product terms, before division by the current first-channel scale \(a\), are respectively \[\begin{align*} T_b(x,y)&=b\frac{\pi_a}{\pi_\tau}\mathcal B_b(x,y), \tag{57}\\ T_b^{\rm prod}(x,y)&= b\frac{\pi_a}{\pi_\tau}\frac f{\pi_f} \sum_{z:Z_f(z)=1}\mathcal F(x,z)\mathcal B_b(z,y). \tag{58}\end{align*}\] In the product, the left task has roles \(X,Z\) and rates \((u;f,f)\), listed as start, end, and usage. The right task has roles \(Z,Y\) and rates \((f;b,\tau)\). In the direct term the start rate stays \(u\). The term \(-E\) in \(D\) is the negative of Equation (57) with \(b=f\) and \(\mathcal F\) in place of \(\mathcal B_b\). Summing the direct and product terms gives the raw inputs \(G_i,G_i'\); adding the negative term gives \(D\).

For a copy projection \(x_i,y_j\) from stage \(l+1\) to stage \(l\), the raw term is \[ r_i(x)b\frac{\pi_a}{\pi_\tau}\mathcal B_b^{(j)}(x,y). \tag{59}\] Partition by the fixed symbols \(i,j\). Append those symbols to the native hash arguments before the inherited suffixes on their respective channels. The parent-stage source multiplier is obtained from its detector vectors on the broad parent-stage vertex domain. No child tuple contains the old source or a particular right endpoint as an extra parameter.

For the second raw channel, use \(\zeta\) times the preparatory table, take its conditional average on \(\mathcal C_{xy}^{u,t}\), and extend the result by zero off \(Y_t(y)\). For a detector vector use the trial \[ \operatorname{rank}_{J} \{\mathcal S_i(x,y):X_f(x)=1\} \quad\hbox{on the slice }Y_f(y)=1. \tag{60}\] Ranks include arbitrary zero padding. For rewards, writing \(f=b_l\), the two terms for \(W_{l+1}(x_i)\) on its inherited source slice are \[ r_i(x)W_l(x),\qquad r_i(x)\frac f{\pi_f} \sum_{z:Z_f(z)=1}\mathcal F(x,z)W_l(z). \tag{61}\] The first \(W_l\) call keeps the inherited source domain; the second restarts on the midpoint domain. Vector answers are finally clipped into \([0,1]\).

Lemma 24 (True sample means). Using true stochastic tables in Equations (57)–(61) gives, after the indicated conditional averages, the prescribed raw inputs and vector values on their conditioning slices, with the stated zero extensions. In particular, the first row channel has target \(d_y=G(x,y)/a\) and the second has target \(\zeta q_y\mathbf1_{Y_t(y)}\).

Proof. If \(b\ge a\), the usage hit is already the conditioning hit. The conditional normalization of a bin therefore turns \(b\mathcal B_b\) into its deterministic bin contribution. If \(b<a\), nested origin boxes give \[\mathbb E\left[\frac{\pi_a}{\pi_b}\mathbf1_{Y_b(y)}V \,\middle|\,\mathcal C_{xy}^{u,a}\right] =\mathbb E[V\mid\mathcal C_{xy}^{u,b}]\] for every integrable \(V\) on the narrower slice. Thus the importance factor performs exactly the required change of conditioning.

For a fixed midpoint in Equation (58), first make this terminal change when necessary and then condition on \(Z_f(z)=1\). Apart from membership indicators, the true right bin gate and denominator use only channel \(Y\). The true left correction uses only channels \(X,Z\). These channels are independent, and the remaining true multipliers are deterministic. The two normalizations can therefore be applied separately. In particular, after the midpoint hit is fixed, the shared role \(Z\) does not occur in the right bin’s remaining random value. This is an identity for the true atoms; the estimates will instead be compared by their joint error norms. The factor \(f/\pi_f\) cancels midpoint sampling and converts its correction atom to the correction entry. Summing midpoints proves the product formula. The projection multiplies by deterministic \(r_i(x)\) and includes the routing fraction in the bin, so the same argument applies on each fixed-symbol part. Fixed-suffix invariance in Lemma 4 justifies all inherited padding conventions.

The second channel is its defining preparatory conditional mean, on its own narrower slice; it has no importance factor. The detector mean is the definition of \(U_i\). True reward factors are deterministic, so the same midpoint argument proves Equation (61). ◻

To form a budget-\(m\) sample term, replace all earlier inputs in the corresponding formula by budget-\(m\) estimates, with their domains and common task parameters unchanged. This includes the detector inputs to \(r_i\); the primary factor \(r_0=1\) remains constant. Products, source functions, and ranks are evaluated before averaging. We first compare this sample’s conditional mean with the same formula using true atoms. The scale and source-domain costs of that comparison will determine how much accuracy a budget must provide.

Mean comparisons and the cost of releasing a source

Lemma 24 identifies the exact conditional means. For estimated factors, the comparison has two distinct costs. The factor \(b/a\) converts a bin contribution to the current entry units. A right factor at a midpoint also replaces the old source domain of rate \(u\) by the stage domain of rate \(f\). The following estimates account for both changes without assuming independence of the errors.

For nonnegative \(H,H'\) define \[ \mathcal L(H,H')= \begin{cases}|H-H'|^s/(H+H')^{s-1},&H+H'>0,\\0,&H=H'=0. \end{cases} \tag{62}\]

Lemma 25 (Homogeneous loss). The loss in Equation (62) is comparable, by constants depending only on \(s\), to \(|H^{1/s}-(H')^{1/s}|^s\). It contracts under sums and nonnegative mixtures. If \(H\le C_0G\) for a nonnegative \(G\), then for every \(a>0\) and \(H'\ge0\), \[ a\min\left\{1,\frac{|H-H'|}{a+G}\right\}^{s} \le C\mathcal L(H,H'), \tag{63}\] where \(C\) depends only on \(s,C_0\).

Proof. Put \(v=\max(H^{1/s},(H')^{1/s})\) and \(w=\min(H^{1/s},(H')^{1/s})\). Factoring \(v^s-w^s\) bounds it between \(v^{s-1}(v-w)\) and \(sv^{s-1}(v-w)\); the denominator is between \(v^{s(s-1)}\) and \(2^{s-1}v^{s(s-1)}\). This proves comparability. For a mixture, apply Hölder to \(|H-H'|=\mathcal L(H,H')^{1/s}(H+H')^{(s-1)/s}\) and then integrate; the triangle inequality first bounds the absolute difference of the integrals. This proves contraction, including finite sums.

If \(H'\le2H+a\), then \(H+H'\le3C_0G+a\le C'(a+G)\), which proves Equation (63) directly. If \(H'>2H+a\), then \(|H-H'|>H'/2\), \(H+H'<3H'/2\), and \(H'>a\). Thus \(\mathcal L(H,H')\ge cH'\ge ca\), while the left side is at most \(a\). ◻

Every unsigned true term of the signed formula for \(D\) is bounded by a constant times \(D\), by Lemma 8; the corresponding assertion for the other raw inputs follows from their nonnegative sums. Consequently the preceding lemma can be applied separately to all true contributions. Triangle inequalities combine their clipped relative errors even when a contribution has a negative sign.

Lemma 26 (Scale and domain comparison). Compare a budget-\(m\) sample term with the same term using true stochastic inputs. For a first-channel mean contribution use its clipped relative error against \(1+d_y\). In the bound for its \(s\)th-power norm, the scale factor multiplying a child root-discrepancy power is, up to fixed constants, \[\begin{cases} b/a,&b\ge a,\\ (b/a)^{s-1},&b<a. \end{cases}\] A child that releases the old source and restarts on a stage-rate \(f\) domain costs in addition at most \(f/u\) in norm power. Direct children retaining that source have no such factor. The same-scale vector and second-channel comparisons have constant scale cost, with \(f/u\) only where a source is released. No independence of estimated factors is required.

Proof. First suppose \(b\ge a\). Apply Equation (63) after conditional expectation, and then the mixture contraction of Lemma 25 before it. Homogeneity supplies a factor \(b/a\) after division by \(a\). The factor \(f/\pi_f\) in a product is bounded above and below by fixed constants. For bounded nonnegative atoms a product-root difference splits into the factor-root differences, multiplied by fixed constants times support indicators of the unaffected factor in one of the two versions.

To determine the source-domain factor, write \(\Delta_s\mathcal F\) and \(\Delta_s\mathcal B_b\) for the root discrepancies. After integration on the source/end slice, the relevant sums, besides \(O(b/a)/N\), have the forms \[\begin{align*} &\mathbb E\sum_{x,z,y}\frac{X_u(x)}u Z_f(z)Y_a(y) |\Delta_s\mathcal F(x,z)|^s \mathbf1_{\mathcal B_b(z,y)\ne0},\\ &\mathbb E\sum_{x,z,y}\frac{X_u(x)}u Z_f(z)Y_a(y) \mathbf1_{\mathcal F(x,z)\ne0} |\Delta_s\mathcal B_b(z,y)|^s. \end{align*}\] Either support indicator can come from either compared version. In the first sum the right row cap eliminates \(y\) at constant cost. In the second sum the left incoming cap eliminates \(x\) at cost \(O(1/u)\). The right child’s row metric has weight \(1/f\), leaving exactly \(O(f/u)\). Its tuple is the same for every old source, so this elimination compares the same right error throughout. More explicitly, for either left support relation \(\mathcal P\) with incoming degree at most \(K_c\) and a common right discrepancy \(e(z,y)\), the relevant inequality is \[\begin{align*} &\frac1N\mathbb E\sum_x\frac{X_u(x)}u \sum_{z:(x,z)\in\mathcal P}Z_f(z) \sum_{y:Y_b(y)=1}|e(z,y)|^s\\ &\qquad\leq \frac{K_c}{Nu}\mathbb E\sum_z Z_f(z) \sum_{y:Y_b(y)=1}|e(z,y)|^s\\ &\qquad=K_c\frac fu\,\|e\|_{f,b;s}^s, \end{align*}\] where the source channel in the last norm is \(Z\). The incoming cap applies pointwise, so the discrepancy can depend on the same environment as \(\mathcal P\). Using \(Y_a\le Y_b\) enlarges the remaining endpoint domain to the child’s own domain. A direct term simply retains its source and endpoint sums. In a projected term, the source-multiplier error eliminates \(y\) by a row cap and changes the vertex weight from \(1/u\) to \(1/f_{\rm par}\); this costs \(O(f_{\rm par}/u)\). Partitioning the fixed copy symbols keeps every tuple common during these sums.

For \(b<a\), change the terminal conditioning to \(Y_b\) before taking a power. The normalized coefficient inside is then \(b/a\). There are only constantly many nonzero left paths in the union of the two rows. Ordinary bounded-factor difference inequalities and the finite-sum power inequality therefore apply with constant cost. The integrated change from \(Y_a\) to \(Y_b\) leaves \(\pi_a/\pi_b\) outside the power. Its total factor is \[(b/a)^s\frac{\pi_a}{\pi_b}=O((b/a)^{s-1}).\] The same row and incoming sums as above then apply, now with \(Y_b\). Ordinary differences of bounded nonnegative values are controlled by their root differences. Absolute mean error bounds clipped relative mean error, so this proves the fine-bin assertion.

For the second raw channel, Jensen is applied directly on \(\mathcal C_{xy}^{u,t}\). Its integrated endpoint mask is \(Y_t\), already contained in the preparatory norm’s \(Y_a\) domain. There is no importance factor and hence no \(\pi_a/\pi_t\) penalty. The fixed scaling \(\zeta\) does not increase the cost.

A rank discrepancy is at most the supremum of its input discrepancies. For the detector trial, summing its power over \(Y_f/f\) is bounded by the score error sum over \(X_f/f\) and \(Y_f\); the rates match exactly. For rewards, the finite left-path count gives an ordinary product comparison, and the same incoming cap eliminates old starts for a right-vector error. Source functions have bounded Lipschitz roots in their clipped detector inputs, so their costs are the broad vertex costs already computed. Every estimate here is pointwise or follows from Jensen and summation; none uses independence of its errors. ◻

The preceding comparison measures errors in a retained sample term. Omitting a fine bin instead discards its whole true contribution; the next bound measures that loss.

Lemma 27 (Fine omission). If one fine bin is omitted, the norm of its normalized contribution is at most \(C(b/a)^{(s-1)/s}\). This applies to direct terms, routed terms, and correction–bin products.

Proof. The deterministic contribution \(H\) has entries at most \(Cb\) and bounded row mass. For a product these bounds follow from Lemma 13. Therefore \[\sum_y\pi_a|H(x,y)/a|^s \le \frac{\pi_a}{a^s}(Cb)^{s-1}\sum_yH(x,y) \le C'(b/a)^{s-1}.\] Integrating source hits gives \(\pi_u/u=O(1)\), and summing at most \(N\) rows cancels the normalization \(1/N\). ◻

Choosing precision to pay for scale and domain changes

The preceding bounds explain the rate weights in the accuracy parameter. Write \(h_b=\log_2(b/a)\) for the offset of a bin scale. In norm, the coarse-bin cost is \((b/a)^{1/s}\) and the fine-bin cost is \((b/a)^{(s-1)/s}\). Since \(s=8\), a weight \(1/2\) between these two exponents leaves a geometric margin in both directions. The same weight pays for the source-release cost \((f/u)^{1/s}\). We now make this accuracy convention precise; the additive schedule will then trade lower child budgets for longer averaging walks.

Assign positive integer ranks \(i_*\) in the task order fixed above, so every dependency decreases rank by at least one. The finite number of pilot steps and of preparations is an absolute constant, so \(i_*=O(l+1)\). A task at integer budget \(M\ge0\) has implicit precision \[ p=\frac{M}{K_{\rm prec}}-\frac12\log_2\frac1a -\frac12\log_2\frac1u-A_*i_*. \tag{64}\] For a stage-\(l\) vector use \(a=f=b_l\); detector vectors also have \(u=f\), whereas rewards retain their inherited start rate. The constants \(K_{\rm prec}\) and \(A_*\) will be chosen below. If \(p\le0\), return the zero table or vector and empty access classes. Otherwise the task is called active. The intended induction bound is \(C_b2^{-p}\) for every required ordinary vector norm and every required root discrepancy norm, including the pre-incoming scores and witnesses. A fixed \(C_b\) makes this assertion trivial for inactive tasks: values and row supports are uniformly bounded, and \(\mathbb E X_u/u=\pi_u/u=O(1)\).

At a normalized correction task, the column multiplier is assembled from earlier detector-vector estimates at budget \(M-c_{\rm col}\), for a fixed positive integer \(c_{\rm col}\). This is a separate final assembly call. All other calls used in one additive sample term have the common budget denoted by \(m\).

Suppose a child at budget \(m\) has parameters \(a',u',i_*'\). Subtracting Equation (64) for the parent gives the exact identity \[ p'-p=-\frac{M-m}{K_{\rm prec}} +\frac12\log_2\frac{a'}a +\frac12\log_2\frac{u'}u +A_*(i_*-i_*'). \tag{65}\] In a bin comparison \(a'\ge c b\) for a fixed \(c>0\): a bin child has \(a'=b\), while a correction has \(a'=f\) and \(b\le K_f f\). The start is either retained or reset to the indicated broad stage rate. Fixed-stage changes in vector calls only change fixed constants. Thus, after taking \(s\)-th roots of the costs in Lemma 26, the scale contribution times child accuracy is bounded by \[ C C_b\,2^{-p-A_*+(M-m)/K_{\rm prec}} 2^{-3|h_b|/8}. \tag{66}\] Indeed for \(h_b\ge0\) the exponent is \(h_b/s-h_b/2=-3h_b/8\); for \(h_b<0\) it is \((s-1)h_b/s-h_b/2=3h_b/8\). For a child that restarts at rate \(f\), the source-domain comparison costs \((f/u)^{1/s}\) in norm, while its precision gains \(\frac12\log_2(f/u)\). Multiplying the cost by the resulting accuracy factor gives \((f/u)^{1/s-1/2}=(f/u)^{-3/8}\le C\), using the allowed start-rate condition \(u\le f\).

For the separate final column-vector call, \(a=f\), its child uses \(a'=u'=f\), and \(M-m=c_{\rm col}\). Thus its norm cost times child accuracy is at most \[C C_b2^{-p}\,2^{c_{\rm col}/K_{\rm prec}-A_*} (f/u)^{-3/8}.\] The fixed rank separation pays this fixed budget loss. This calculation also explains why the source-domain ratio is charged in the error analysis rather than inserted as a numerical coefficient in a sample.

The additive schedule

Every budget-zero task is inactive under the preceding precision rule. Hence a budget-zero sample term is zero: it contains at least one estimated table or vector factor. Exact base terms \(S(x,y)/a\) and \(W_0(x)\) are inserted separately.

For every first-channel term labelled by scale \(b\), set \(c_b=\lceil c_0(1+|h_b|)\rceil\), where \(c_0\) is a fixed positive constant. This rule includes the sparse direct term \(-E\) with \(b=f\), as well as direct bins, products, and projection terms. Every child factor in a term, including the detector-vector inputs to a projected source multiplier, uses that term’s common budget \(m\) and cutoff \(c_b\). Use zero scale offset only for second-channel preparatory means, detector-vector trials, and reward terms with fixed stage-rate changes. The separate final column assembly retains its fixed drop \(c_{\rm col}\). Denote the corresponding normalized sample term at budget \(m\) by \(T_{b,m}\), including division by \(a\) for the first row channel. For \(1\le m\le M-c_b\) form \[ T_{b,m}-T_{b,m-1}. \tag{67}\] Apply to this difference the conditional walk of length \(\ell(M-m)\), where \(\ell\) is a sufficiently large fixed integer. Use the same generator word and the same guessed wraps for its two parts.

Group increments by dyadic slack \(d\le M-m<2d\). Add groups in decreasing powers of two \(d\le M\) down to \(1\), including empty groups. Put exact base contributions in the first group. Let \(S_j\) be the uncompressed prefix through the group with slack \(d_j\). The conditional mean prefix is obtained by replacing each walk by uniform averaging on its slice; noise is the difference from that mean. For an available bin the last budget reached by this mean prefix is \[ m_j(b)=\max\{0,M-\max(d_j,c_b)\}. \tag{68}\] This exact cutoff will matter in the error bound. Bins with \(c_b\ge M\) have no layers and are omitted.

The walk is applied to an array indexed by actual endpoints. Thus, for each term and pair, its input is the single function \[F_{xy}(\sigma)=T_{b,m}(x,y;\sigma)-T_{b,m-1}(x,y;\sigma),\] with all contributions at that endpoint already summed. A path description is only a way to enumerate those contributions. In particular, the two budgets must be grouped at the same actual endpoint before their difference is measured. The following lemma establishes this identity for the path enumeration and also records the description length needed later by the controller.

Lemma 28 (Immediate paths). At a row group of slack \(d\), immediate candidates have descriptions of \(O(d)\) bits. With exact endpoint grouping their sums equal the specified conditional-walk increments. A product visits its midpoint by a paired correction move before using any factor at that midpoint. Its right-hand task tuple does not depend on the discarded old source. Every child estimator in such a candidate has budget at most \(M-d\).

Proof. A descriptor records the finitely many formula and copy-part choices, its scale and budget offsets, its sign in Equation (67), the generator word and guessed wraps, and a bounded number of constituent ports. An included scale obeys \(|h_b|=O(d)\), and the walk has length \(O(d)\). Lemma 21 therefore bounds all these records by \(O(d)\) bits. Port alphabets are fixed by Lemma 17. There are only \(O((1+d)^2)\) choices of budgets and scales before the exponential word/variant choices.

Check each constituent port, the outer source and endpoint memberships, and the guessed wraps at the vertices where their coordinates can be evaluated. For a fixed endpoint and original generator word, unique wrap validation leaves exactly its genuine conditional transformation. Give the surviving term its original walk-word probability, multiplied by the formula coefficients and sign. Guessed wraps do not change that probability. Constituent ports enumerate entries once. Hence summing over true equal endpoints is precisely the walk of the difference, even though its endpoint was not known when the variant was specified.

The first factor of every product is a correction, whose retained ports have bounded degree in both directions. Its paired move supplies the midpoint and a bounded return token. The right factor can then be queried using the common \(Z,Y\) tuple prescribed above; that tuple contains no old source. A direct unbalanced terminal motion has no further factor query at its endpoint. These are combinatorial access requirements; the controller section implements the pairing and terminal operations. Finally, the slack definition gives \(m\le M-d\), also for the \(m-1\) term. ◻

All reachable rate logarithms remain \(O(B)\) when \(M_{\rm top}=O(B)\): a scale extension at slack \(d\) changes a reciprocal-rate logarithm by \(O(d)\) and reduces budget by at least \(d\). Stage resets have only their fixed stage rates, usage rates are minima of available rates, and inherited padding has at most one symbol per stage above the task. Summing along a dependency chain bounds all extensions by \(O(M_{\rm top})\). The field in the shared environment can therefore be chosen large enough for every used scale, including budget-zero endpoints of available differences. Omitted finer deterministic contributions need no hash tables.

We can now estimate each mean prefix by telescoping the conditional means to budget \(m_j(b)\). Its deterministic omission error and its child-estimation error are treated separately. The remaining noise will be bounded by applying the conditional spectral estimate to the endpoint functions \(F_{xy}\) above.

Lemma 29 (Mean prefixes). Assume all earlier table and vector bounds with coefficient \(C_b\). For an active task there is a fixed \(B_1\), independent of the later choice of \(A_*\), such that every mean prefix has norm at most \(c_*2^{-p+B_1d_j}\) against its target. For the first channel the norm uses clipped relative discrepancy; for the second and for vectors it uses ordinary absolute discrepancy. The coefficient \(c_*\) can be made arbitrarily small by increasing \(A_*\).

Proof. Available scales. For an available scale, conditional expectations of Equation (67) telescope to its term at \(m=m_j(b)\) in Equation (68). If this budget is zero, the term is zero and the inactive-child bound still applies. More generally, the child bound remains valid whenever its precision is nonpositive, by the bounded-output estimate for inactive tasks. In every case \[M-m\le d_j+c_0(1+|h_b|)+1.\] Choose \(K_{\rm prec}\) so that \(c_0/K_{\rm prec}<1/8\). Equation (66) is then bounded by \[C' C_b\,2^{-p-A_*+d_j/K_{\rm prec}}2^{-|h_b|/4}.\] The rate grid is geometric, so summing this bound over available scales costs a fixed constant. There are only a fixed number of formula and factor types. Unbinned calls use the same calculation without the scale sum. Minkowski and subadditivity of clipped relative discrepancies combine all these terms. Exact base contributions are already present in every prefix and have no error.

Wholly omitted scales. A bin with \(c_b\ge M\) contributes no layer. We bound its deterministic contribution directly by Lemma 27. First we show that every such bin is fine. Put \(r=\log_2(1/a)\ge0\). A coarse bin has \(0\le h_b\le r\) because all grid rates are at most one. Activity gives \(M>K_{\rm prec}(r/2+A_*i_*)\), even after discarding the nonnegative start-rate cost. With \(K_{\rm prec}/2>c_0\) and \(A_*\) sufficiently large, this implies \(M>\lceil c_0(1+r)\rceil\), so a wholly omitted bin is fine. For the sparse direct term \(b=f\), the same argument gives availability when \(f\ge a\). When \(f<a\), the admissibility condition \(a\le K_f f\) gives \(|h_f|\le\log_2 K_f\), and increasing \(A_*\) ensures \(M>c_f\). Thus this term is never wholly omitted in an active task. The inequality \(c_b\ge M\) implies \[|h_b|\ge (M-1)/c_0-1.\] Lemma 27 and a geometric sum consequently bound all such omissions by \(C2^{-(s-1)M/(s c_0)}\). Since \(p\le M/K_{\rm prec}\) and \(K_{\rm prec}>s c_0/(s-1)\), this is at most \(C2^{-p-\eta M}\) for some fixed \(\eta>0\), after absorbing the fixed rounding offset. Finally \(M>K_{\rm prec}A_*i_*\ge K_{\rm prec}A_*\) makes its coefficient arbitrarily small as \(A_*\) increases. Taking, for example, any fixed \(B_1>1/K_{\rm prec}\) proves the claim, including every last-budget and no-budget cutoff loss. ◻

Absolute difference bounds and averaging noise

The conditional-mean errors are now bounded. To control the noise left by finite averaging, we need an absolute norm bound on each paired difference before the walk is applied. This coarser bound will also be used for the fingerprint comparison.

Lemma 30 (Absolute differences). At a slack-\(d\) group, the ordinary raw norm of an unaveraged paired difference is at most \[ C C_b\,2^{-p-A_*+C_0d}, \tag{69}\] where \(C_0\) is fixed independently of the walk length coefficient. For two coupled versions of a budget-\(r\) sample term, if the corresponding child discrepancies are bounded by \(2^{-p'_r}\lambda_r\), their normalized raw difference is bounded by \(2^{-p+C_0'd}\lambda_r\), summed over the finitely many child types. For a paired layer at budgets \(m\) and \(m-1\), apply this estimate to both parts. Its coupled discrepancy is consequently bounded by \(2^{-p+C_0'd}(\lambda_m+\lambda_{m-1})\), after enlarging the fixed constant, or by the corresponding maximum over all children in the group. Every child in either part has budget at most \(M-d\). For a fixed guessed measure-preserving variant and its valid endpoint restriction the same bounds hold. Counting all variants separately may increase \(C_0'\) by a constant depending on the chosen walk length.

Proof. Subtract the same stochastic true-atom term from the budget-\(m\) and budget-\((m-1)\) terms, and use the triangle inequality. A product has only constantly many left paths in the union of the compared rows. Ordinary bounded-factor differences and the row and incoming caps therefore give the sums used in Lemma 26, without any averaging. For a coarse bin the absolute coefficient costs \(b/a\) in the norm rather than its \(s\)-th root. For a fine bin the importance factor can stay inside the power. Both are bounded by \(2^{O(|h_b|)}\). Equation (65) pays the only separate possibly large domain ratio, \(f/u\), as before. All other scale and accuracy changes are \(2^{O(d)}\) because \(|h_b|=O(d)\) and \(M-m=O(d)\) in this group. Ranks and bounded source functions obey the same pointwise comparisons. This proves Equation (69). Replacing comparison to truth by a comparison of the two coupled child versions gives the second assertion by the identical inequalities; independence is still unnecessary.

For a fixed guessed variant, valid paths require the transformed memberships. Change variables by its full-environment measure-preserving bijection and then discard surplus validity guards. This leaves exactly the child-domain sums just bounded. When combining the paired terms, combine their contributions at the same actual endpoint before measuring the difference. The number of distinct variant descriptors is \(2^{O(d)}\), so treating them separately merely enlarges the crude exponent. The estimate before such separate counting is a statement about the single-seed sample function and has no dependence on the eventual walk length coefficient. ◻

Lemma 31 (Conditional prefix bounds). Assume exact local grouping, the earlier task error bounds, and that the sample functions depend only on variables varied by the conditional walk. Their uniform slice means are then deterministic at each indicated pair or vertex. Write \(S_j^{\rm mean}\) for a mean prefix and \(S_j^{\rm noise}=S_j-S_j^{\rm mean}\) for its noise. There are fixed \(B_1\) and arbitrarily large prescribed \(K_n\) such that \[\begin{align*} \left\|\min\left(1, \frac{|S_j^{1,\rm mean}-d|}{1+d}\right)\right\|_{u,a;s} &\le c_*2^{-p+B_1d_j}, \tag{70}\\ \|S_j^{\rm noise}\|_{u,a;s} &\le c_*2^{-p-K_nd_j}. \tag{71}\end{align*}\] The second-channel mean uses its ordinary absolute norm, extended by zero off its usage domain, and satisfies the first bound. Vector means and noises satisfy the corresponding vertex bounds. The coefficient \(c_*\) can be chosen as small as needed by task separation and walk constants.

Proof. The mean assertion is Lemma 29. For a difference layer, write \(F_{xy}\) for its sample function on \(\mathcal C_{xy}^{u,a}\), \(\mathcal A_{xy}\) for the conditional walk operator of length \(\ell(M-m)\), and \(\Pi_{xy}\) for uniform averaging on that slice. Put \(w_{xy}=\Pr(\mathcal C_{xy}^{u,a})/(Nu)\) and \(r_{M-m}=C2^{-c\ell(M-m)}\). Lemma 20 gives a uniform operator bound, so \[\sum_{x,y}w_{xy} \|(\mathcal A_{xy}-\Pi_{xy})F_{xy}\|_{L_s}^s \le r_{M-m}^s \sum_{x,y}w_{xy}\|F_{xy}\|_{L_s}^s.\] The slice weights make the sum on the right precisely the integrated raw norm to the \(s\)th power, bounded by Lemma 30. Thus the uniform operator estimate applies directly to the integrated input bound, even when different slices have different input norms. For the second channel replace \(Y_a\) by \(Y_t\) in the slice and its weight; the mask stays \(Y_t\) throughout. Vectors use the same calculation on their one-vertex slices. No extra conditioning factor appears.

There are \(O((1+d)^2)\) budget/scale terms in a group. Choose \(\ell\) so that the contraction dominates \(2^{C_0d}\) and this polynomial, leaving at least \(2^{-(K_n+1)d}\). Summing groups with slack at least \(d_j\) gives Equation (71). Exact base terms have no noise. This contraction applies to the paired difference as one function of the same seed. The sum of valid variants is the conditional walk, so the gap applies to that sum before any separate variant counting. Increasing \(A_*\) supplies an arbitrarily small coefficient in both this estimate and the mean bound. ◻

The determinism hypothesis in Lemma 31 applies at each fixed pair or vertex, in the task’s effective environment as in Remark 22. Uniform averaging must integrate every coordinate on which its sample function depends. In Section 9, the clean calculation removes dependence on the current block’s salt, while the parent walk varies all earlier salts used by lower-block calculations; Lemma 40 verifies this hypothesis. The resulting first-channel mean is a deterministic number, and the second-channel mean has only its prescribed \(Y_t\) membership mask. These are the means used in the conditional heavy-count argument of Lemma 35.

Compression of additive row estimates

The additive schedule produces a raw prefix on all retained endpoints after each slack group. We replace each array by a finite list of exceptions to a common default. Our objective is to control the bounded row functions, including their scores and support witnesses; a large raw entry need not have a small absolute error.

The update changes only small entries until its first overflow, when list capacity forces its threshold above the prescribed floor. Before then, the total earlier change is much smaller than the next floor. At the first overflow, either many retained endpoints are truly heavy or the prefix errors account for the excess. The heavy-count event is rare even after conditioning on one specified endpoint. We use this rarity for the deterministic conditional-mean errors, and use the absolute norm bound for the remaining noise. A stationary-reference comparison controls the subsequent updates.

We first fix the increments supplied by exact endpoint grouping and prove the comparison without range clamps. We then show that clamping preserves the relevant row functions and assemble the exact-grouping error reserve. Section 9 compares the actual and clean calculations with those clamps in place.

The row interface

Fix a row task, its start rate \(u\), and its first-channel endpoint rate \(a\). Write \(d_y\geq0\) for its true first raw input, in units of \(a\). Its second raw input, when present, is \(v_y=\zeta\mathbf1_{\{Y_t(y)=1\}}q_y\). Here \(q_y\) is the deterministic conditional mean, \(0\leq q_y\leq1\), and \(q_y=0\) whenever \(d_y\) is below the positive-start threshold of the preparatory function \(h\). Thus \(d_y\) and \(q_y\) are deterministic at each specified endpoint; \(v_y\) carries the prescribed narrower membership mask. There is a fixed constant \(C_{\mathrm{mass}}\) such that \[ a\sum_y d_y\leq C_{\mathrm{mass}}. \tag{72}\] The endpoint sets in this section are the retained sets, with the stipulated common padding on their hash arguments.

We apply a fixed finite family of pre-incoming row functions to a raw array. Denote this family by \(\mathcal H\). It includes the main pre-functions, auxiliary scores, and witnesses required by the preceding construction. Each function uses the first entry, possibly the second entry, and the \(K\)th largest first entry. For each member \(H\in\mathcal H\) there are fixed positive thresholds \(g_H,e_H\), with \(3g_H<e_H\), such that its row gate is zero when that rank is at least \(g_H\), and its entry factor is zero when the first entry is at most \(e_H\). The gate is constant and full in a neighborhood of zero. The \(s\)th roots of the output functions are bounded and Lipschitz in their relevant entry and rank arguments, where \(s=8\). All these arguments can be clipped to fixed sensitivity ranges without changing the functions. These are exactly the threshold buffers and saturations imposed when constructing the atoms; the thresholds for the required functions are listed in Table 1. In particular, for every real raw array, including arrays with negative entries, \[ \sum_{H\in\mathcal H}\#\{y:H(y)\ne0\}\leq C_{\mathrm{supp}}, \tag{73}\] where \(C_{\mathrm{supp}}\) is fixed. Indeed a nonzero gate implies that fewer than \(K\) entries exceed its close threshold, and every positive output entry exceeds the larger entry threshold. This observation applies also to the broader witnesses.

Choose \(\eta>0\) below all positive entry thresholds, all positive-start thresholds for rank sensitivity, and the threshold for \(h\), with enough fixed separation that perturbations of size several times \(\eta\) remain inside these flat regions. Choose \(\zeta\leq\eta/100\). These constants are fixed throughout the proof.

Let \(d_1>d_2>\cdots>d_{m_*}=1\) be the decreasing powers of two indexing slack groups. Empty groups are included. Let \(S_i=(S_i^1,S_i^2)\) be the uncompressed prefix through group \(i\), and let \(\Delta_i=S_i-S_{i-1}\), with \(S_0=0\). Define the clipped endpoint errors and their row supremum by \[ \begin{aligned} e_i(y)&=\max\left\{ \min\left(1,\frac{|S_i^1(y)-d_y|}{1+d_y}\right), \min(1,|S_i^2(y)-v_y|)\right\},\\ e_i&=\sup_y e_i(y). \end{aligned} \tag{74}\] omitting channel two when absent. Write \(\beta_i(y)\) for the maximum of the mean discrepancies, with the same clipping and first-channel normalization as in Equation (74), and \(\gamma_i(y)\) for the maximum absolute prefix noise. Thus \(e_i(y)\leq\beta_i(y)+\gamma_i(y)\). At a specified row and endpoint, the first-channel mean discrepancy is deterministic. The second-channel mean discrepancy is a deterministic magnitude times its narrower membership mask. Prefix noise may depend on the shared environment. The deterministic assertion is a hypothesis of the compression argument: the slice average defining a mean must vary every environment coordinate on which its sample function depends. Lemma 31 states this requirement, and Lemma 40 verifies it for the clean calculations to which we will apply the result. We will use conditional rarity to control these fixed mean-error magnitudes; the noise will be controlled by its unconditional norm bound.

For row integration use \[\int F\,d\mu=\frac1N\mathbb E\sum_x\frac{X_u(x)}u F(x).\] Endpoint integration means inserting the sum over retained endpoints in this expression. The prefix estimates used as hypotheses below are \[ \left(\int\sum_y\beta_i(y)^s\,d\mu\right)^{1/s} \leq c_*2^{-p+B_1d_i},\qquad \left(\int\sum_y\gamma_i(y)^s\,d\mu\right)^{1/s} \leq c_*2^{-p-K_nd_i}. \tag{75}\] Lemma 31 supplies these estimates for the additive schedule. Here and below endpoint sums use the relevant masks. In particular these bounds control the \(L_s(\mu)\) norms of the corresponding row suprema, which we denote by \(\beta_i\) and \(\gamma_i\).

The update map

Adjoin arbitrarily many dummy positions, all with target and increments zero. A compressed array has default \((z,0)\), shared also by the dummies, and a finite list of exceptions. Initially \(z=0\) and the list is empty. Order statistics of an array with this default mean the order statistics of its finite exceptional values with as many default values as needed. For a rank of order \(r\), the finite list together with at least \(r\) default copies gives exactly the prescribed rank. Additional real list positions whose values equal the default do not change this rank. Dummy positions serve only this rank computation; endpoint sums in the error bounds are over real retained IDs. The margin statement of Lemma 16 ensures that a default position cannot expose a row port. At group \(j\) add \(\Delta_j\) by actual endpoint equality, obtaining \(w_y\), and set \[ \begin{aligned} \rho_j&=2^{-c_1d_j},\qquad n_j=2^{c_2d_j},\qquad m_y=\max_\nu|w_y^\nu|,\\ u_0&=\max\{z,\rho_j,\operatorname{rank}_{n_j}(m)\},\\ q_0&=\max\{0,\operatorname{rank}_K(w^1)\},\\ z'&=\max\{u_0,\min(q_0,3u_0)\}. \end{aligned} \tag{76}\] Here \(c_1,c_2\) are fixed sufficiently large positive integers, and \(c_2\) is chosen also so that \(n_j\geq K\) in every group. With \(b=(z',0)\), the updated array is \[ T_j^\nu(y)=b^\nu+ \operatorname{clip}\bigl(w_y^\nu-b^\nu, \pm7(m_y-2u_0)_+\bigr). \tag{77}\] The notation \(\operatorname{clip}(r,\pm A)=\max(-A,\min(r,A))\) is used.

Lemma 32 (Local compression properties). The update has the following properties. It replaces every position of size at most \(2u_0\) by the new default, is identity at every position of size at least \(3u_0\), and lies coordinatewise between the input and the new default. The default is nondecreasing. The map from the input array and old default to the updated array and new default is uniformly Lipschitz in the sup norm. A list containing all exceptions can be chosen with at most \(n_j\) positions, using only the previous list and the new candidates.

Proof. The zero-radius assertion is immediate. Since \(z'\leq3u_0\), if \(m_y\geq3u_0\) then \[|w_y^1-z'|\leq m_y+3u_0\leq7(m_y-2u_0), \qquad |w_y^2|\leq7(m_y-2u_0),\] which proves identity there. Clipping a displacement toward zero proves the coordinatewise assertion. Also \(z'\geq u_0\geq z\).

If two input arrays and their defaults differ by at most \(\delta\), their size arrays, size ranks, \(u_0\)’s, and \(q_0\)’s differ by at most \(\delta\). Their new defaults differ by at most \(3\delta\), and their clipping radii by at most \(21\delta\). The elementary bound \[|\operatorname{clip}(r,\pm A)-\operatorname{clip}(r',\pm A')| \leq |r-r'|+|A-A'|\] therefore gives, for example, a Lipschitz constant 28 for the full update. This constant is independent of the number and magnitudes of entries.

Fewer than \(n_j\) positions have size strictly larger than \(u_0\). A coarse test supported above \(u_0\) and full above \(2u_0\) retains all possible exceptions within the stated cap. A position outside the old list and new candidate set has input \((z,0)\), has size at most \(u_0\), and hence becomes the new default. No other position need be inspected. The selected list may retain some real positions that become exactly default after compression. Such positions still represent the same full array. No padding dummy is selected, since its input size is at most \(u_0\). ◻

In particular an absolute perturbation introduced at group \(i\) propagates through the remaining updates and the final root functions with factor at most \[ C C_1^{m_*-i+1} \leq C(d_i+1)^{\log_2 C_1}, \tag{78}\] for a fixed \(C_1>1\), enlarged below as needed. This statement concerns full arrays and defaults; it does not require their exceptional lists to agree.

Deterministic comparison arguments

Call group \(j\) an overflow if its \(u_0\) exceeds \(\rho_j\). Before the first overflow every compression change has magnitude at most \(6\rho_i\) and occurs only at input size below \(3\rho_i\). Because successive slacks halve, choosing \(c_1\) large gives, uniformly in \(j\), \[ 6\sum_{i<j}\rho_i\leq\rho_j/100, \qquad 6\sum_i\rho_i\leq\eta/100. \tag{79}\] The same estimates cover dummy defaults. For example, before the first overflow the preceding default is at most \(3\rho_{j-1}\ll\rho_j\).

Lemma 33 (Preloss and stationary reference). Let \(L_{j-1}=T_{j-1}-S_{j-1}\) before the first overflow at \(j\). Then \(\|L_{j-1}\|_\infty\leq6\sum_{i<j}\rho_i\), and at every endpoint with \(d_y\geq\eta\), \[ \max_\nu|L_{j-1}^\nu(y)| \leq C\sum_{i<j}\rho_i e_i(y). \tag{80}\] Start a reference calculation at the input of group \(j\) with the true array \((d,v)+L_{j-1}\), the same preceding default, and no further increments. Every subsequent compression preserves its relevant row outputs. Their root discrepancy from the true outputs is bounded by \(C\sum_{i<j}\rho_i e_i\).

Proof. The uniform preloss bound follows by adding the changes described above. If a change occurs at group \(i\) at a position with \(d_y\geq\eta\), then \(|w_y^1|\leq3\rho_i\) and the preceding preloss is much smaller than \(\rho_i\). Thus \[\frac{|S_i^1(y)-d_y|}{1+d_y} \geq \frac{d_y-3\rho_i-6\sum_{k<i}\rho_k}{1+d_y} \geq c_\eta>0.\] The change of size at most \(6\rho_i\) is consequently bounded by \(C\rho_i e_i(y)\). Adding such changes proves Equation (80).

In the reference array negative first magnitudes and all second magnitudes are at most \(\eta\), by the choices of \(\zeta\) and the floors. Let \(q\) be its initial \(K\)th first-coordinate rank. The default is nonnegative, so \(q\geq0\). If \(q>\eta\), then \(n_i\geq K\) implies that the \(n_i\)th size rank is at most \(q\). The floor and preceding default are also at most \(q\), so \(u_0\leq q\) and \(z'\leq q\). Entries at most \(q\) stay at most \(q\). Entries at least \(q\) are unchanged if \(q\geq3u_0\); otherwise \(z'=q\), and they stay at least \(q\). The \(K\)th rank is therefore unchanged. Negative first magnitudes and second magnitudes do not grow, so this argument repeats at every subsequent group. If instead \(q\leq\eta\), the size rank, \(u_0\), and new default are at most \(\eta\), and the next \(K\)th rank is again at most \(\eta\).

Consider an individual row function. In the first case its rank stays fixed. If its gate is not already zero, \(q<g_H\), so every entry in its support has first coordinate greater than \(e_H>3q\geq3u_0\) and both channels there are untouched. A position below entry support cannot cross into support because the default is at most \(q<e_H\). In the second case the rank stays in the common full-gate region; the same identity argument applies to every position in entry support. Thus each reference output is preserved.

Finally, positions with \(d_y<\eta\) stay below every entry threshold and cannot change a rank within its sensitivity region. Perturbations at all other positions satisfy Equation (80). Clip first entries and ranks to the fixed sensitivity ranges and use root-Lipschitz continuity of the row functions. This proves the last assertion. ◻

If no overflow ever occurs, the same proof, together with the final prefix error, gives the bound \[ \sup_{H,y}|H(T_{m_*};y)^{1/s}-H(d,v;y)^{1/s}| \leq C\left(e_{m_*}+\sum_i\rho_i e_i\right). \tag{81}\] For completeness, if the expression on the right is not small, bounded output ranges give the assertion after enlarging \(C\). Otherwise low-target positions remain below sensitivity. At other positions clipped first entry discrepancies are bounded by a constant times their relative error, and second entry discrepancies by their absolute error. Adding the preloss estimate and applying the root-Lipschitz bounds proves the claim.

Lemma 34 (Stability after the first overflow). At a first overflow in group \(j\), the discrepancy of final row-root outputs from the stationary reference of Lemma 33 is at most \[ C\sum_{i\geq j} C_1^{m_*-i+1}e_i. \tag{82}\] The constant does not depend on the largest true first entry.

Proof. Write \(E\) for the sum in Equation (82) without the initial constant. If \(E\) exceeds a sufficiently small fixed constant, bounded output ranges suffice. Otherwise all clipped errors involved are strictly below 1. Hence \[|S_i^1(y)-d_y|\leq(1+d_y)e_i, \qquad |S_i^2(y)-v_y|\leq e_i,\] and later increments have first magnitude at most \((1+d_y)(e_i+e_{i-1})\) and second magnitude at most \(e_i+e_{i-1}\). The input difference at group \(j\) is precisely \(S_j-(d,v)\); both calculations include the same previous preloss. The factors \(1+d_y\) prevent a direct absolute-error bound with a uniform constant when a true entry is large. The following two cases use the row gate and saturation to remove that dependence.

Choose \(G\) larger than every gate close threshold. A baseline at least \(G\) closes all gates, since the dummy entries force every first-coordinate rank to be at least the baseline. Closure is permanent under the unclamped update. Choose \(L_0\) much larger than \(G\), all sensitivity thresholds, and 1. If at least \(K\) true first entries exceed \(L_0\), the reference gates are closed. In the perturbed calculation either some \(u_0\geq G\), which forces permanent closure, or every \(u_0<G\). In the latter case those \(K\) coordinates stay larger than half their true value: their initial error and the total later additive variation are at most \(C(1+d_y)\sum_{i\geq j}e_i\), and their first value is then larger than \(3G\). Inductively each such position lies in the compression identity region. These \(K\) positive entries close all gates as well.

It remains to consider the case of fewer than \(K\) true entries above \(L_0\). Lemma 33 bounds every reference rank and threshold by \(2L_0\). Fix \(L_1\) much larger than \(L_0\) and all upper saturation thresholds. At the at most \(K-1\) positions where \(d_y>L_1\), replace the first input at group \(j\), in both calculations, by one common positive sentinel larger than \(L_1\), and give that coordinate zero later first-channel increments. Do not modify second channels. The reference ranks, thresholds, and outputs are unchanged: these coordinates stay above the relevant ranks and in all upper saturation regions. The sentinels are comparison devices in this proof. They need not be identified or stored by the compression algorithm. Fewer than \(K\) such coordinates occur, so replacing them by common large values removes their possibly large absolute differences while preserving every rank that lies in a sensitivity range.

All remaining first-coordinate input errors and increments now have absolute bounds by \(C(1+L_1)e_i\) and \(C(1+L_1)(e_i+e_{i-1})\), respectively. Lemma 32, iterated from \(j\) to the last group, therefore bounds every array and threshold discrepancy by \(CE\), on taking \(C_1\) larger than the fixed update constant. If \(E\) is sufficiently small, all these thresholds remain below \(2L_0+1\). The final root-output discrepancy is at most \(CE\).

To remove the modification in the perturbed calculation, argue inductively. Every original coordinate with \(d_y>L_1\) remains above \(d_y/2\), because its total raw additive variation is at most \(C(1+d_y)\sum e_i\). It is thus above the bounded thresholds and in the identity and saturation regions. Replacing fewer than \(K\) such positive coordinates by sentinels cannot change either a \(K\)th first rank or an \(n_i\)th size rank whose sentinel-computed value is at most \(2L_0+1\): both the original and modified values are above that bound, and \(n_i\geq K\). The defaults and thresholds therefore agree, the compression is identity in both channels at these positions, and all other coordinates agree. The relevant output behavior is the same. This completes the induction and the proof. ◻

Charging overflow events

The deterministic comparisons leave two cases to integrate: no overflow, and the first overflow. To handle the latter, we separate a large retained set of true heavy endpoints from overflow caused by errors at light endpoints. For this subsection let \(\mathcal O_j\) denote the event that group \(j\) is the first overflow. These events are pairwise disjoint. The heavy-count event defined next depends only on the true row and its retained endpoints; it does not itself assert that any overflow occurs.

At first overflow in group \(j\), the previous default is smaller than \(\rho_j\). Thus at least \(n_j\) retained real positions have input size strictly greater than \(\rho_j\). Define the true heavy-count event \[\Theta_j=\left\{\#\{y\text{ retained}:d_y>\rho_j/4\}\geq n_j/2\right\}.\] The dummy positions do not contribute to this count, since their input is the preceding default. On \(\mathcal O_j\cap\Theta_j^c\), at least \(n_j/2\) of the overflowing positions have \(d_y\leq\rho_j/4\). Their second targets vanish, because the floors lie below the \(h\)-threshold. Equation (79) then implies \(e_j(y)\geq c\rho_j\) at all these positions. Consequently the indicator of this false-type first overflow is bounded by \[ \frac{C}{n_j\rho_j^s}\sum_y e_j(y)^s. \tag{83}\] A bounded row-output error may be charged to this indicator. On \(\mathcal O_j\cap\Theta_j\), we instead use the stability estimate and the conditional rarity of \(\Theta_j\) at each endpoint.

Lemma 35 (Conditional rarity of a true heavy count). At every specified endpoint, \[ \mathbb P(\Theta_j\mid X_u(x)=Y_a(y)=1) \leq\frac{C}{n_j\rho_j}. \tag{84}\] The same bound holds with the narrower endpoint condition \(Y_t(y)=1\). Moreover, for every prefix mean row supremum, \[ \|\mathbf1_{\Theta_j}\beta_i\|_{L_s(\mu)} \leq Cc_*2^{-p+B_1d_i}(n_j\rho_j)^{-1/s}. \tag{85}\]

Proof. By Equation (72), the deterministic heavy set \[A_j(x)=\{y:d_y>\rho_j/4\} \quad\text{satisfies}\quad |A_j(x)|\leq\frac{4C_{\mathrm{mass}}}{a\rho_j}.\] Conditional on one specified endpoint hit, each other endpoint in this set is retained with probability at most \(Ca\), by Lemma 4. This remains true if the specified hit uses the narrower rate. The forced endpoint contributes at most 1, so the expected retained heavy count is at most \(1+Ca|A_j(x)|\leq C'/\rho_j\). The independent start condition has no effect. Markov’s inequality at the threshold \(n_j/2\) proves Equation (84). This calculation uses the conditional pair bound and linearity of expectation; it permits dependence among the heavy-endpoint hits.

For the first-channel mean error at a fixed row and endpoint, write its deterministic magnitude as \(b_i(y)\). Then \[\mathbf1_{\Theta_j}\sup_{y:Y_a(y)=1}b_i(y)^s \leq\sum_y\mathbf1_{\Theta_j}Y_a(y)b_i(y)^s.\] Condition on the start and that endpoint before taking expectations. Equation (84) bounds the resulting expectation by \(C/(n_j\rho_j)\) times the endpoint-integrated mean-error power. Indeed the factor \(b_i(y)^s\) is fixed under this conditioning, so \[\mathbb E\bigl[\mathbf1_{\Theta_j}X_u(x)Y_a(y)b_i(y)^s\bigr] \leq\frac{C}{n_j\rho_j} \mathbb E\bigl[X_u(x)Y_a(y)b_i(y)^s\bigr].\] The second-channel mean error has a deterministic magnitude multiplied by \(Y_t(y)\); repeat the same argument using the narrower conditional bound. Taking roots and using Equation (75) proves Equation (85). The conditioning step uses the deterministic mean-error magnitude at the specified endpoint. Prefix noise can depend on the heavy-count event, so no analogous conditional estimate is asserted for it. ◻

Proposition 36 (Integrated compression bound). For fixed sufficiently large \(c_1\), then \(c_2\), and sufficiently rapid prefix-noise decay \(K_n\), compression preserves the prescribed row error budget, with an arbitrarily small fixed coefficient when \(c_*\) is chosen small. More precisely, \[ \left(\int\sum_{H\in\mathcal H}\sum_y |H(T_{m_*};y)^{1/s}-H(d,v;y)^{1/s}|^s\,d\mu\right)^{1/s} \leq C_{\mathrm{buf}}c_*2^{-p}, \tag{86}\] where \(C_{\mathrm{buf}}\) is fixed, independent of the number of groups, row width, and requested precision.

Proof. Equation (73) converts every row sup bound above into a bound for its summed output power with a fixed factor. In fact the support of a discrepancy lies in the union of the true and computed supports, whose total size over \(\mathcal H\) is at most \(2C_{\mathrm{supp}}\). Thus this step introduces no factor from the ambient number of endpoints. The preloss contributions have norm at most \[Cc_*2^{-p}\sum_i\rho_i \bigl(2^{B_1d_i}+2^{-K_nd_i}\bigr).\] On the event of no overflow, Equation (81) adds the final-prefix cost, at most \(Cc_*2^{-p}(2^{B_1}+2^{-K_n})\). By Equation (83) and the disjointness of the possible first-overflow events, the total false-type error power is at most \[ Cc_*^s2^{-sp}\sum_j (n_j\rho_j^s)^{-1} \bigl(2^{sB_1d_j}+2^{-sK_nd_j}\bigr). \tag{87}\]

For a true-type first overflow at \(j\), use Lemma 34 and split \(e_i\leq\beta_i+\gamma_i\). For the mean part, enlarge the actual first-overflow event to \(\Theta_j\), apply Equation (85), and use Minkowski’s inequality. Its total norm is bounded by \[ Cc_*2^{-p}\sum_j(n_j\rho_j)^{-1/s} \sum_{i\geq j}C_1^{m_*-i+1}2^{B_1d_i}. \tag{88}\] For the noise part, use the uniqueness of the first-overflow group on each row. Its tail noise sum is bounded pointwise by the single full sum \(\sum_iC_1^{m_*-i+1}\gamma_i\), regardless of dependence between noise and overflow. More explicitly, \[\sum_j\mathbf1_{\mathcal O_j} \sum_{i\geq j}C_1^{m_*-i+1}\gamma_i \leq\sum_i C_1^{m_*-i+1}\gamma_i,\] because at most one indicator on the left is nonzero. The absolute prefix-noise bound therefore gives norm at most \[ Cc_*2^{-p}\sum_iC_1^{m_*-i+1}2^{-K_nd_i}, \tag{89}\] without an extra sum over \(j\).

Choose \(c_1>B_1\) large enough for the floor conditions. Choose \(c_2>sc_1+sB_1\) with a fixed positive margin and also large enough that \(n_j\geq K\). Equation (87) then converges uniformly. For its mean term the exponential factor is exactly \(2^{-(c_2-sc_1-sB_1)d_j}\). For Equation (88), its inner sum is at most \(C(d_j+1)^A2^{B_1d_j}\) for a fixed \(A\), by Equation (78) and dyadic grouping; the stronger choice of \(c_2\) just made ensures convergence of the outer sum, whose exponential factor is \(2^{-((c_2-c_1)/s-B_1)d_j}\). Taking \(K_n\) sufficiently large handles Equation (89). The preloss sum converges by \(c_1>B_1\). All constants are fixed before choosing the arbitrarily small prefix coefficient \(c_*\), giving Equation (86) and the asserted spare coefficient. ◻

Proposition 36 bounds the compressed pre-incoming functions, including scores and witnesses. Before assembling the final atom estimates, we impose the numerical range bounds needed to evaluate these arrays.

Bounding numerical ranges

In exact endpoint grouping, the additive increments are independent of which exceptions are stored. The path description and bounded child ranges give a fixed \(D\) for which \(\|\Delta_i\|_\infty\leq2^{Dd_i}\), after increasing \(D\) to include immediate term counts. This also bounds a group when all its same-key increments are summed into one coordinate, with a possibly larger fixed \(D\). Fix a common gate closure level \(G\ge\max_H g_H\), and then choose \(L>10G\) above every sensitivity threshold. Let \(R_i=2^{B_2d_i}\) and choose \(B_2\) sufficiently large that \[ R_i\geq 10L+10\sum_{k>i}\|\Delta_k\|_\infty \quad\text{for every group }i. \tag{90}\] Dyadic grouping permits this uniform choice because \(d_{i+1}=d_i/2\). After each compression, clip every coordinate and the baseline componentwise to \([-R_i,R_i]\). An off-list coordinate is clipped exactly as the baseline is, so it remains a default coordinate. These clamps therefore preserve the finite-list representation.

Lemma 37 (Range clamping). For fixed common increments satisfying Equation (90), these clamps do not change the final relevant row functions. Each clamp is sup-norm nonexpansive, so in comparisons of two calculations it preserves the uniform Lipschitz propagation bound. The analogous assertion holds for vector sums clipped beyond their final sensitivity range and all remaining additive variation.

Proof. Compare the clamped and unclamped calculations with the same increments. For the fixed level \(G<L/10\), once both baselines are at least \(G\), all gates are zero because of the dummies, and remain so: compression cannot lower the baseline and every clamp bound exceeds \(G\). It remains to compare the calculations before this happens.

Maintain the following invariant. Their baselines agree while below \(G\). Their coordinates either agree, or have matching signs and magnitudes well beyond \(L\); any such differing component has enough remaining margin that the total future increments cannot bring it into \([-L,L]\). At the first clamp of a component this follows from Equation (90). A subsequent clamp restores the same margin using the current \(R_i\). Between clamps, common additions consume at most the available tail margin.

After an addition, entries whose components differ therefore have size larger than \(L\) in both calculations. Their size-rank statistics either are equal below \(G\) or are both at least \(G\): differing values lie above this comparison range and all other size values agree. In the latter case both new baselines close the gates. In the former case the two values of \(u_0\) agree and are below \(G\). Only the first-coordinate rank clipped at \(3u_0\) is needed to compute \(z'\). Differing positive first components are above this range; differing negative first components are below zero and do not affect its nonnegative clipped value. Consequently the two new baselines agree as well. If this baseline is at least \(G\), both calculations close. Otherwise every differing position lies in the compression identity region because its size exceeds \(L>3G\); equal coordinates at other positions undergo the same update. Thus the invariant continues.

Before closure, all rank arguments and all entry arguments within their sensitivity ranges agree. Differing large components have matching signs and identical saturated behavior. The final row functions therefore agree. A scalar clipping map is 1-Lipschitz, which gives the comparison assertion. For a vector sum, the same tail-margin argument ensures that a clipped value stays on the same saturated side until the final bounded clipping; no rank argument is needed. ◻

The equality assertion of Lemma 37 is used for the clean error bound: exact endpoint grouping supplies the same increment arrays regardless of the chosen list presentation. It preserves the values of every relevant \(H\in\mathcal H\), including scores and smooth witnesses. It makes no assertion that finite-accuracy access tests or selected descriptors agree with those of the unclamped calculation. In the adaptive comparison of Section 9, both the actual and clean calculations are clamped from their definition. Their possibly different bucket increments are compared directly, and the nonexpansiveness part of Lemma 37 preserves the local perturbation bound. This use allows different real representatives, including retained positions whose current value equals the default.

The error reserve with exact endpoint grouping

The prefix estimates and compression bound now give the current task’s error estimate under exact endpoint grouping. The hypotheses still require each conditional average to vary every random coordinate used by its sample function. We combine the compressed pre-incoming bounds with the atom assembly comparison; the next section will verify this averaging hypothesis for the block-clean calculation.

Proposition 38 (Error reserve with exact endpoint grouping). Suppose the earlier row and vector tasks have error bound \(C_b2^{-p'}\) and the support and common-parameter properties of the atom section. Use exact comparisons locally, and vary every variable on which the sample terms depend. After compression and the range clamps above, all active row outputs, pre-incoming scores, and witnesses have error at most \((C_b/2)2^{-p}\). The same holds for the final clipped vectors and the assembled correction outputs, after choosing the fixed constants in the order specified below.

Proof. To apply Proposition 36, use the deterministic-mean and absolute-noise bounds of Lemma 31. The first target is nonnegative with bounded deterministic row mass before division by \(a\), and its second target is bounded by a fixed multiple of \(\zeta\) and vanishes below the preparatory support threshold. The buffered row functions and their witnesses have fixed root Lipschitz constants and fixed support bounds. These verify its remaining hypotheses. The compression proposition shows that all their errors are a fixed multiple of \(c_*2^{-p}\) when its convergent series of floors, capacities, and noise weights are chosen. This includes the pre-incoming quantities required by the incoming comparison. Exact endpoint grouping gives common increment arrays for the clamped and unclamped calculations. Lemma 37 therefore preserves these row-function values, including scores and witnesses, so the same estimates hold after clamping.

Lemma 18 then charges the incoming gates by those same pre-output norms. Its separate column-vector calls have a fixed positive budget slack and decrease task rank. Equation (65) pays their broad-domain factor, and increasing \(A_*\) makes their remaining coefficient as small as required. A final vector needs no compression; its clipping is nonexpansive and the bounds at \(d_j=1\) apply directly.

Choose the fixed atom range/support bound \(C_b\) first. Choose \(c_0\) and \(K_{\rm prec}\) to retain the margin in Equation (66) and fix \(B_1\). The compression proof then chooses its floor exponent \(c_1\), capacity exponent \(c_2\), and required noise decay \(K_n\). Choose task separation and walk length sufficiently large to supply the needed small \(c_*\). The total coefficient can thus be made at most \(C_b/2\). Inactive tasks retain their separate trivial \(C_b\) bound. The remaining half of the active-task error bound is reserved for the coupled fingerprint error. ◻

The reserve applies to a local exact-grouping calculation with the stated averaging property. The actual algorithm will also use short endpoint keys. In Section 9 we define an exact-grouping comparison within each budget block, verify that it has this averaging property, and bound its difference from the algorithm.

Adaptive fingerprints within budget blocks

We now replace endpoint comparisons at small budget slacks by short keys. The endpoint universe has size \(2^{O(B)}\), whereas the key range has size \(B\). These comparisons change the computed arrays. To control the change, we compare the algorithm with a calculation that uses exact equality within one budget block and the actual algorithms below that block.

This comparison has two roles. Its conditional averages give the deterministic means needed by the compression reserve, and its candidate endpoints form a set independent of the current block’s key. We charge actual paths leaving that set to errors in lower calls, then bound collisions inside it. The resulting coupled induction proves the error bound for the actual estimator.

We use the root error metric \(\mathfrak d_{u,a}\) for bounded nonnegative row outputs, its ordinary version \(\|\cdot\|_{u,a;s}\) for raw arrays, and the ordinary vertex metric for vectors. For a row quantity \(Z(x,\sigma)\) define also \[ \|Z\|_{\mathrm{row},u;s}^s =\frac1N\mathbb E_\sigma\sum_x\frac{X_u(x)}u|Z(x,\sigma)|^s. \tag{91}\] The value of \(s\) remains \(8\). Supremums over endpoints below are over actual ID positions, with the appropriate retained-domain masks; virtual default positions are never counted with the multiplicity of the ambient universe.

Uniform salts and their access rule

The fixed-pair estimate below combines polynomial evaluation, as in classical algebraic fingerprinting (Schwartz 1980), with affine universal hashing (Carter and Wegman 1979). Its proof is included to specify the fair-bit distribution and space bound. The adaptive use of these fingerprints requires the separate block comparison that follows.

Lemma 39 (Short fingerprints). Suppose every endpoint has an injective, fixed-length binary encoding of length \(L\le c_{\mathrm{ID}}B\). There is a uniform construction from \(O(\log B)\) independent fair bits of a map \(k:\{\text{endpoint IDs}\}\longrightarrow\{1,\ldots,B\}\) such that, for distinct fixed IDs \(v,w\), \[ \Pr[k(v)=k(w)]\le C/B\le C B^{-1/2}. \tag{92}\] Evaluating the map at a presented endpoint needs \(O(\log B)\) arithmetic scratch in addition to access to its encoding.

Proof. Pad all encodings to the same length and put \(P_v(T)=\sum_{i=0}^{L-1}v_iT^i\). Choose a prime \(B^3\le q\le2B^3\), increasing the fixed lower bound on \(B\) if necessary. Thus \(L<q\). Let \(\ell=\lceil\log_2q\rceil\). Split a string of \(3\ell\) fair bits into independent integers \(R,A,C\) in \(\{0,\ldots,2^\ell-1\}\), and reduce them modulo \(q\) to obtain \(r,\alpha,\beta\). Each residue has probability at most \(2/q\). Let \(\tau:\mathbb F_q\longrightarrow\{1,\ldots,B\}\) partition the standard residues into consecutive sets of sizes differing by at most one, and set \[k(v)=\tau\bigl(\alpha P_v(r)+\beta\bigr).\] All computations are in \(\mathbb F_q\) before applying \(\tau\).

The nonzero polynomial \(P_v-P_w\) has degree at most \(L-1\), so it has at most \(L-1\) roots. This root bound follows by repeatedly factoring \(T-r_0\) at a root; it applies here because the coefficient difference \(1\) or \(-1\) is nonzero in \(\mathbb F_q\). Consequently \[\Pr[P_v(r)=P_w(r)]\le 2(L-1)/q.\] For distinct evaluations \(z,z'\), the map \((\alpha,\beta)\mapsto(\alpha z+\beta,\alpha z'+\beta)\) is a bijection of \(\mathbb F_q^2\). Under uniform parameters its two outputs are uniform and independent. The probability that their \(\tau\)-values agree is \[q^{-2}\sum_{j=1}^B|\tau^{-1}(j)|^2 \le \frac{\lceil q/B\rceil}{q}\le\frac1B+\frac1q.\] The actual joint law of \((\alpha,\beta)\) is pointwise at most four times the uniform law and is independent of \(r\). Adding the two cases gives \(2(L-1)/q+4/B+4/q=O(1/B)\). A polynomial evaluation by Horner’s rule and the residue arithmetic use \(O(\log q)=O(\log B)\) bits. The prime can be chosen uniformly by searching and trial division in that space. The existence of a prime in the indicated interval is the same elementary prime-interval fact used in choosing the environment field. ◻

Fix a small positive rational \(\epsilon\), to be chosen last. This coefficient controls the fingerprint block width. Write \[ h_*=\lfloor\epsilon\log_2 B\rfloor,\qquad M_b=bh_*,\qquad \mathcal I_b=\{M_b,\ldots,M_b+h_*-1\}. \tag{93}\] Enlarge the fixed minimum \(B\) so \(h_*\ge1\) and \(h_*\ge(\epsilon/2)\log_2B\). Give each block meeting \(\{0,\ldots,M_{\mathrm{top}}\}\) an independent salt from Lemma 39. Pad each salt to a multiple of the fixed integer \(H\) by unused independent bits. Since \(M_{\mathrm{top}}=O(B)\), the total number of bits is \[ O\bigl((1+B/h_*)(\log B+H)\bigr)=O(B). \tag{94}\] Here and throughout, constants may depend on the fixed \(\epsilon\) and \(H\). We enumerate the original fair-bit strings when enumerating environments; reducing chunks modulo \(q\) does not change that sampling convention.

The following access rule is part of the algorithm. A task whose budget lies in \(\mathcal I_b\) may inspect salts with indices at most \(b\). Its conditional walks vary all earlier salts as free bits and freeze the salt of block \(b\) and every later salt. The restriction applies to structural choices, port orders, validity tests, and predecessor tests as well as to numerical values. The controller construction below implements this rule for every access operation, including inverse queries. It is not enough to impose it only on the final output values.

This rule is stated in the coordinates of the task’s effective environment. Write \(\sigma\) for that environment when analyzing a task. A fixed ancestor variant may mix its free salt coordinates before supplying \(\sigma\) to a descendant; the descendant’s table depends on the resulting environment, not on the word that produced it. Thus the rule does not forbid reading a later-indexed bit of the original master environment while replaying that word. It requires the descendant’s values, port orders, and access operations to be functions only of its permitted effective coordinates. Lemma 68 implements this distinction. Each fixed variant is measure preserving by Lemma 21, so the integrated estimates may change variables to its effective environment. We do not assume that an adaptively selected variant leaves a uniform conditional law. At a block-\(b\) task, its own variants act only on effective salts earlier than \(b\) and freeze its effective block-\(b\) salt and all later ones.

The update rule and the clean calculation

We now complete the recursive definition of the actual task; inactive tasks still return zero outputs and empty access classes. For an active row task, the additive schedule of Section 7 supplies candidate increments from smaller-budget outputs. Process the groups in decreasing slack. In each group, distribute increments by the rule below, then use the compression, range clamps, and shortlist rule of Section 8. After the last group, form the pre-incoming row functions and precursor classes from the completed arrays, then assemble any incoming gates and column multipliers as in Section 5. The separate column-vector calls have smaller budget. Vector tasks use the same schedule to sum their scalar terms, with their prescribed range clamps and final clipping; they require no endpoint grouping or row-list compression.

The row-group distribution rule is as follows. For a dyadic group with slack \(d>2h_*\), compare endpoint IDs exactly. For a group with \(d\le2h_*\), use the key from the task’s block as follows. First discard invalid candidates, including invalid guessed wraps, individually. Sum all new candidate increments having the same key. Add a bucket’s sum to every old representative with that key. If the bucket has no old representative, introduce exactly its first valid new candidate, with its previous value equal to the old default, and add the bucket sum there. Do not combine or discard old representatives in this key step. Apply the same full-array compression, clamping, and shortlist rule as in the preceding section.

Distinct actual representatives remain distinct. Indeed two introduced representatives have different keys, and an introduced representative’s key differs from every old key. Increments at one and the same actual endpoint are always combined, even when other endpoints collide with it. Stored values are never summed by key. The two channels use the same keys and the same representative selection for the entire group.

For an analytical comparison at a budget \(M\in\mathcal I_b\), define the clean calculation for block \(b\) by replacing key comparisons by exact ID comparisons throughout that block, recursively in every child whose budget still belongs to \(\mathcal I_b\). Children below \(M_b\) are the actual algorithms, with exactly the same environments in the two calculations. This replacement applies to the entire numerical and structural calculation: candidate validity, representative selection, compression, exposed classes, and their port orders. Descriptor enumeration uses fixed index orders. Each finite-accuracy test for a shortlist or exposed class retains its prescribed accuracy level and tie convention, independently of any requested output accuracy. All these operations use the corresponding clean operands. Thus the clean calculation specifies its access choices as well as its represented arrays.

Denote the actual and clean versions by superscripts \(\mathrm{ac}\) and \(\mathrm{cl}\). Both still perform compression and clamping. The clean calculation is a finite analytical recursion; its exact comparisons are not asserted to satisfy the eventual space bound.

Clamping is part of both definitions. The clean error reserve uses Lemma 37 for exact-ID increments, which are independent of the stored exception list. That lemma preserves the relevant row functions and witnesses; it need not preserve the list or port presentation of an unclamped calculation. The coupling below compares the two clamped calculations directly. Its propagation bound uses the Lipschitz compression map and the nonexpansiveness of each common clamp, without introducing an unclamped actual calculation.

Lemma 40 (Independence of the clean block). Every clean value and structural choice for block \(b\) is independent of that block’s effective salt and of later effective salts. In a task of budget \(M\in\mathcal I_b\), the actual and clean row states immediately before the first group with \(d\le2h_*\) are identical and independent of that salt. Their stored lists at this boundary have size \(2^{O(h_*)}\). The deterministic slice-mean assertion required by Lemma 31 holds for the clean calculation.

Proof. Induct on the budget and, within a task, on its prescribed sequence of operations. Lower-block algorithms ignore the salt of block \(b\) and later salts by the access rule; same-block clean children ignore them by induction. Candidate tests use these children, exact endpoint grouping precedes compression, and the compressed row determines its exposed classes and pre-incoming scores. Incoming ranks use those completed scores. Separate column-vector queries likewise use a smaller-budget clean child in this block or an actual child below it. All statements here concern the current effective environment: each same-block transformation preserves its current and later salts, and each lower-block transformation acts on a subset of its earlier salts. Each numerical recipe uses the same fixed operand order, and each structural test uses its fixed accuracy level and tie rule. Its result therefore depends only on earlier clean operands, including their finite numerical approximations. The fixed index orders then give salt-independent representatives and ports.

In particular, exact full-array addition, compression, and the prescribed row functions are salt independent. Reordering paths or choosing a different representative of the same endpoint would not change an exact sum; retaining an extra default-valued position would not change the represented array. The fixed conventions above also remove these presentation choices from the structural calculation.

For a group with \(d>2h_*\), every child budget is at most \(M-d<M_b\). Both runs use the same lower-block children and exact comparisons there. All their list operations, including their deterministic tie conventions, ignore the present salt. Thus their boundary states agree. The last larger dyadic slack, if there is one, is at most \(4h_*\), so its shortlist has at most \(2^{O(h_*)}\) positions. If there is no larger group, the boundary list is empty.

Finally, the clean estimand depends only on rank channels and earlier salts. These are precisely the variables averaged by its conditional walk, subject to its stated vertex conditions. Same-block salts are irrelevant, and lower-block salts are free variables of this walk. Projection onto the uniform measure of the slice therefore gives a deterministic quantity at its indicated pair or vertex, with the prescribed retained-domain masks. In particular, no random dependence on a frozen salt remains in this mean. For the second channel, its deterministic mean magnitude is extended by its prescribed narrower endpoint mask, as required in Section 8. The salt rule is required of all access operations, including inverse queries and port orders; Lemma 68 supplies its implementation. Lemmas 20 and 21 therefore apply to this function. This verifies the deterministic-mean hypothesis of Proposition 38; it does not assert independence between errors of different children. ◻

Witness envelopes and escaped paths

To control adaptive candidates, we first bound a set of endpoints determined by the clean calculation. An actual path leaving this set will be charged to a child witness discrepancy.

We recall the relevant interfaces from Lemma 17. Every exposed actual path port has a broader smooth witness equal to one at that port. Each witness has bounded row support. The actual correction ports, including exposed ports of weight zero, have incoming multiplicity at most \(K_c\). Tables contain one entry per pair of actual IDs. After fixing a copy part, the child tuple is common over the whole table; in particular, a right child reached from a midpoint does not depend on the discarded old start. All these statements hold for arbitrary computed raw arrays, not just for accurate ones.

The incoming cap applies to the entire exposed correction relation, including its zero-weight entries. The broader auxiliary score class needs only a row bound: it is queried for scores and witnesses, whereas the left factor of a two-factor path is a capped correction table. This distinction permits the source sum below to be removed even when an offending candidate has zero numerical weight.

Let \(\mathcal G_b\) be the sigma-field generated by all coordinates of the task’s effective environment except the salt of block \(b\). Fix a row \(x\) and condition on \(\mathcal G_b\) with \(X_u(x)=1\). Take its boundary list from Lemma 40, including every stored exact-default slot. For every small group and every immediate descriptor variant, evaluate the clean factors and their witnesses in the environment obtained by that fixed variant. This transformation leaves the present and later salts frozen. Include the endpoints in each direct witness support and, for a two-factor path, the endpoints in the right witness supports at all midpoints in the left witness support. Include local base endpoints when applicable. Let \(\mathcal E_x\) be the union of these endpoints and the boundary list. Set \(\mathcal E_x=\varnothing\) when \(X_u(x)=0\). Unused or invalid clean descriptors may enlarge this set. Actual candidates must still satisfy their wrap tests and all transformed source, midpoint, and terminal memberships; every error estimate below retains the memberships needed for its child metrics.

Lemma 41 (Size and independence of the envelope). The set \(\mathcal E_x\) is independent of the present salt and has size at most \(2^{C_Eh_*}\) for a fixed \(C_E\). All clean nondefault representatives in the small groups belong to this set. Consequently, conditional on the other data, \[ \Pr[\text{two distinct IDs in $\mathcal E_x$ have the same key}] \le C\,2^{2C_Eh_*}B^{-1/2}. \tag{95}\]

Proof. Each fixed variant is independent of the present salt and leaves it unchanged. Lemma 40 therefore makes its clean witness supports \(\mathcal G_b\)-measurable. We use these supports as sets of actual IDs, not the order of a port enumeration. At a group of slack \(d\) there are \(2^{O(d)}\) immediate descriptors, including variants, difference parts, and copy choices. The number of endpoints contributed by a direct witness is bounded. A two-factor path contributes a bounded number of left witness midpoints and, for each, a bounded number of right witness endpoints. Summing over the dyadic groups with \(d\le2h_*\) and adding the boundary list proves the size bound. No incoming bound on a clean witness is used here.

Clean exposed ports lie in their clean witness supports. Apart from the boundary representatives, exact updates can therefore create nondefault values only at the included endpoints. Compression maps all untouched default positions to the new default. Any extra exact-default slots in a shortlist may be omitted when describing its full array. This proves the representative claim. Equation (95) is now the pair union bound from Lemma 39, conditional on the fixed envelope. ◻

Let \(\mathcal F_x\) be the event that some valid actual candidate in the small groups has endpoint outside \(\mathcal E_x\), whether or not its increment is zero. As representatives are only carried forward or chosen among new candidates, all actual representatives stay in the envelope outside this event. This includes default-valued representatives: they start in the full boundary list or inherit a candidate endpoint. Dummy copies used to compute ranks are not new real representatives. Thus omitting an exact-default clean slot when describing its full array does not omit an actual address from the collision analysis.

Lemma 42 (Charging an escaped candidate). Suppose actual–clean child discrepancies at budget \(m\) have their prescribed metric bounds \(2^{-p_m}\lambda_m\). The contribution of \(\mathcal F_x\) to any bounded final pre-incoming row output has norm at most \[ 2^{-p}\sum_{\substack{d\le2h_*\\m\in\mathcal I_b,\ m\le M-d}} 2^{C d}\lambda_m, \tag{96}\] where the sum includes the finitely many dependency types and their descriptor multiplicities in its constant \(C\). Equivalently, one may use the maximum of \(\lambda_m\) over the children of each slack group and increase \(C\).

Proof. An actual endpoint outside the envelope has an offending path factor. For a direct path its actual port is outside the clean witness support. For a two-factor path, either its actual left port is outside its clean witness support, or its midpoint lies in that support and its actual right port is outside the corresponding clean right witness support. At the offending port the actual witness is one and the clean witness is zero. Thus the indicator of this event is at most the \(s\)-th power of their root discrepancy there. Union bounding over candidate paths gives a sum of such nonnegative powers. The unit-witness property holds also for exposed zero-weight ports; no lower bound on the candidate’s increment is being assumed.

Each output under consideration is bounded and has bounded row support. The union of the two output supports is still bounded, by Equation (73). Hence its entire row error power on \(\mathcal F_x\) is at most a fixed multiple of \(\mathbf1_{\mathcal F_x}\). We charge this row-failure indicator, rather than the size of the offending increment. There is no endpoint sum left on the output side to pay for.

The right-factor charge is the case where the old start must be summed out. Fix a guessed variant and a two-factor path \(x\to z\to y\). Change variables by its measure-preserving transformation, as in Lemma 21; in the following display all memberships, ports, and witness values are evaluated in that transformed environment. Write \(e_R(z,y)\) for the root discrepancy of the two right witnesses and \(\mathcal P^{\mathrm{ac}}_L\) for the actual left exposed-port relation. The candidate’s validity retains \(X_u(x)\), \(Z_f(z)\), and the right child’s endpoint membership \(Y_{a'}(y)\). Dropping only surplus guards, the right-port contribution is at most a constant times \[\begin{align*} &\frac1N\mathbb E\sum_x\frac{X_u(x)}u \sum_{z:(x,z)\in\mathcal P^{\mathrm{ac}}_L}Z_f(z) \sum_{y:Y_{a'}(y)=1}|e_R(z,y)|^s \\ &\qquad\le\frac{K_c}{N u}\mathbb E\sum_z Z_f(z) \sum_{y:Y_{a'}(y)=1}|e_R(z,y)|^s =K_c\frac f u\,\mathfrak d_{f,a'}(W_R^{\mathrm{ac}}, W_R^{\mathrm{cl}})^s. \tag{97}\end{align*}\] The middle inequality uses the incoming cap of the actual left port relation. It does not use a clean incoming cap. The right task’s tuple is independent of \(x\), so the same discrepancy is being summed when the old start is removed. This condition is necessary for the displayed equality with the right metric. For the fixed variant and copy part, the right tuple has only the prescribed task parameters and transformed environment. The midpoint \(z\) is its vertex argument, and the old source \(x\) is not a parameter. Its final gate and column queries use that same effective tuple, rather than the ancestral descriptor that happened to discover the entry.

For a direct or left offending port, sum the child’s witness discrepancy over its own row and endpoint domain; the bounded number of possible right ports incurs only a fixed factor. Projected paths are partitioned by the fixed copy symbols, with constant multiplicity. Normalization and preparation dependencies make the same direct witness comparison. If a candidate arose from a narrower terminal slice, its validity mask may be dropped in this row-event estimate. Thus no factor \(\pi_a/\pi_t\) is incurred. In contrast, the transformed source hit cannot be omitted before the change of variables and the incoming-cap sum; it has been retained in Equation (97).

At slack \(d\), scale ratios, variants, and immediate path counts cost at most \(2^{O(d)}\), as in Lemma 30. The possible reset factor in Equation (97) is absorbed by the precision gain. Admissible row tasks have \(u\le f\), as specified in Section 5. Equation (65) gives, after the scale-dependent \(O(d)\) terms, \[\begin{align*} p_m&\ge p-O(d)+\tfrac12\log_2(f/u),\\ (f/u)^{1/s}2^{-p_m} &\le 2^{-p+O(d)}(f/u)^{1/8-1/2} \le 2^{-p+O(d)}. \end{align*}\] Take \(s\)-th roots and use Minkowski to sum the charges. A child below \(M_b\) is identical in the two versions and contributes zero. All remaining children satisfy \(m\le M-d\) and belong to \(\mathcal I_b\), which proves Equation (96). ◻

Bucket algebra and the collision charge

The escape estimate leaves rows whose candidates stay in \(\mathcal E_x\). On these rows we separate child discrepancies from collisions among the clean endpoints. We compare full arrays with their defaults, so different shortlist presentations require no extra error term. The following observation explains the key update rule.

Lemma 43 (Only increments are aliased). Consider one small group on \(\mathcal F_x^c\). Let \(\Delta_{\mathrm{ac}}(y),\Delta_{\mathrm{cl}}(y)\) be its unaliased increments, namely its true-ID sums using actual and clean children, respectively. Let \(k_d\le2^{C d}\) bound its number of immediate descriptors, and let \(\mathcal C_x\) denote a collision on \(\mathcal E_x\). The discrepancy between the actual bucket-distributed increment and the clean true-ID increment is at most \[ (k_d+1)\left( \sup_y|\Delta_{\mathrm{ac}}(y)-\Delta_{\mathrm{cl}}(y)| +\mathbf1_{\mathcal C_x}\sup_y|\Delta_{\mathrm{cl}}(y)| \right) \tag{98}\] in sup norm, simultaneously for both channels. The old full-array discrepancy is added to this bound with coefficient one.

Proof. When all envelope keys are distinct, the rule gives exactly the actual true-ID increment: each old occurrence is updated at its own ID, and each genuinely new ID has its own bucket. Thus the discrepancy is simply \(\Delta_{\mathrm{ac}}-\Delta_{\mathrm{cl}}\).

When there is a collision, each updated representative receives a sum of increments at at most \(k_d\) distinct actual addresses. Its magnitude is at most \(k_d\sup_y|\Delta_{\mathrm{ac}}(y)|\). At positions not updated it is zero. At either type of position subtracting the clean increment costs at most a further \(\sup_y|\Delta_{\mathrm{cl}}(y)|\). Apply \[\sup_y|\Delta_{\mathrm{ac}}(y)| \le\sup_y|\Delta_{\mathrm{ac}}(y)-\Delta_{\mathrm{cl}}(y)| +\sup_y|\Delta_{\mathrm{cl}}(y)|\] to obtain Equation (98). This also covers an increment duplicated at several old representatives and a missing increment at an unchosen new representative.

At no point are old values combined or moved to a different ID. An unlisted old position has exactly the old default, and a newly selected representative starts with that value. Therefore the full-array comparison before compression is the old discrepancy plus the increment discrepancy just bounded. There is no term involving the height of an old value multiplied by a collision indicator. ◻

In applying this lemma to a group, combine contributions at equal actual endpoints within each paired difference before measuring its clean height. At a fixed endpoint and original generator word, the outer slice and wrap tests are the same for budgets \(m\) and \(m-1\). Their constituent port sets may differ, but each set contains all nonzero entries and enumerates each actual ID pair once; an absent entry contributes zero. Their true-ID sums are therefore exactly the two terms of the paired difference. Subtracting the same stochastic true-atom term from both sides, as in Lemma 30, preserves its cancellation. One can then sum the resulting sup bounds over difference layers. This remains legitimate if a first-bucket representative is chosen across several layers: the proof of Lemma 43 applies to their total increment, and the triangle inequality bounds the total by the sum of the paired bounds. A descriptor with zero increment contributes only to \(k_d\) in this bound. It may affect representative selection, which is already covered by Lemma 43 and the escape event.

Lemma 44 (Direct cost of collisions). The direct collision contribution to any final pre-incoming output has norm at most \[ 2^{-p}\,2^{C h_*}B^{-1/(2s)}, \tag{99}\] for a fixed \(C\) independent of \(D_1,D_2,\epsilon\) introduced below. All remaining contributions on \(\mathcal F_x^c\) are bounded by actual–clean child discrepancy norms with an amplification at most \(2^{C d}\) at slack \(d\).

Proof. For each small dyadic group, let \(\Delta_d^{\mathrm{cl},\nu}(x,y)\) be its complete clean true-ID increment in channel \(\nu\in\{1,2\}\), including any exact base term assigned to that group. Take an absent second channel to be zero, and choose a deterministic descriptor bound \(k_d\le2^{Cd}\). By Lemma 30, the clean paired increments at a group of slack \(d\) obey their crude absolute estimates in the retained endpoint metric. These estimates hold for each fixed valid variant after its measure-preserving change of variables. Counting variants and summing difference layers by the triangle inequality costs \(2^{O(d)}\), without using walk contraction. Within each paired layer the true-ID cancellation is performed before this inequality is applied. If a local unpaired base term occurs in a small group, its first-group slack satisfies \(M=O(d)\), so the same crude bound holds in its normalized raw units. Since \(\sup_y|Z_y|^s\le\sum_y|Z_y|^s\) over actual ID positions, the row norm of \(\max_\nu\sup_y|\Delta_d^{\mathrm{cl},\nu}(x,y)|\) is at most \(2^{-p+O(d)}\).

Lemmas 32 and 37, together with Equation (78), propagate an input discrepancy at slack \(d\) to the final array with a factor polynomial in \(d+1\). At each step the discrepancy is measured between the clamped full arrays, including their defaults. The next bucket increment is compared by Lemma 43; compression then applies its uniform Lipschitz bound and the common clamp cannot enlarge the discrepancy. This recurrence does not require their later bucket distributions or stored lists to agree. Root-Lipschitz row functions and bounded final support then give the same type of bound for the output metric. Choose fixed \(C_P,A_P\) so that \(C_P(d+1)^{A_P}\) bounds this propagation. For a retained row define \[H_x=C_P\sum_{\substack{d\le2h_*\\d\text{ a group slack}}} (d+1)^{A_P}(k_d+1) \max_\nu\sup_y|\Delta_d^{\mathrm{cl},\nu}(x,y)|.\] Set \(H_x=0\) when \(X_u(x)=0\). This quantity bounds the propagated clean-height terms in Equation (98). The preceding estimates give \[\|H\|_{\mathrm{row},u;s}\le2^{-p+C_0h_*}.\] All increments defining \(H_x\) are clean, and its other factors are deterministic. Thus \(H_x\) and \(\mathcal E_x\) are \(\mathcal G_b\)-measurable. Conditioning on this sigma-field gives \[\mathbb E\bigl[\mathbf1_{\mathcal C_x}H_x^s\mid\mathcal G_b\bigr] =H_x^s\Pr(\mathcal C_x\mid\mathcal G_b) \le C\,2^{2C_Eh_*}B^{-1/2}H_x^s.\] The row membership \(X_u(x)\) is also \(\mathcal G_b\)-measurable. Integrating the last inequality in the row measure therefore yields \[\begin{align*} \|\mathbf1_{\mathcal C}H\|_{\mathrm{row},u;s}^s &\le C\,2^{2C_Eh_*}B^{-1/2} \|H\|_{\mathrm{row},u;s}^s. \end{align*}\] Taking roots proves Equation (99) after enlarging \(C\). We have not conditioned on \(\mathcal F_x^c\) in this independence calculation: the actual error there is bounded by \(\mathbf1_{\mathcal C_x}H_x\), which is then integrated on the whole space. Conditioning on the salt-dependent stay-inside event is unnecessary and could destroy the stated independence.

For the first term of Equation (98), split products using bounded-factor differences, actual or clean row caps, and the actual or clean incoming cap of the represented two-way factor, exactly as in Lemma 30. These are comparisons of unaliased sums using the respective children; they make no assumption about key success or independence of child errors. Fixed valid variants, scale costs, and counts are all at most \(2^{O(d)}\), with the possible source-release factor absorbed by the same precision gain as in Equation (97). Compression propagation is again absorbed in \(2^{C d}\). This proves the remaining assertion. ◻

Closing the coupled induction

Proposition 45 (Uniform block coupling). The fixed constants can be chosen so that, for every task of budget \(M\in\mathcal I_b\) and precision \(p\), each prescribed actual–clean output discrepancy, including every pre-incoming score and witness, is at most \[ 2^{-p}\lambda_M,\qquad \lambda_M=2^{-D_1h_*+D_2(M-M_b)}. \tag{100}\] The same assertion holds for the ordinary vector metrics. Adding these discrepancies to the clean estimates proves the target error bounds for the actual tables and vectors.

Proof. Perform strong induction on the statistical budget, simultaneously for the actual target estimates, the block-clean target estimates needed at that budget, and Equation (100). An inactive task returns the same zero outputs and empty paths in both versions, so its coupling error is zero. Lower-block children are identical in a block coupling and already have their actual target bounds. Same-block clean children have the clean target bounds at smaller budgets. Therefore Lemma 40, the prefix estimate, and the compression analysis give the clean active target bounds with the spare constants in Proposition 38.

For the coupling, Lemmas 42–44 bound the pre-incoming quantities. Incoming gates are then compared by their rank and root-Lipschitz bounds. In their multiplier effect, restrict to the union of the final supports of the two tables; these have bounded incoming multiplicity. Thus Lemma 18 costs only a fixed factor on pre-incoming errors. Its separately queried column vectors have a positive constant budget drop and their \(f/u\) metric factor is covered by their precision gain. Vector raw sums have no endpoint grouping, but their same-block children may differ. The same crude individual-variant comparisons and nonexpansive clamps bound those differences. Detector rank comparisons are already part of the assembly estimate. No simultaneous key-success event over an incoming search is required.

Collect the finite number of output types and absorb their fixed assembly factors into one fixed exponent \(C\). For a child budget \(m\in\mathcal I_b\) used at group slack \(d\), we have \[ \lambda_m=\lambda_M2^{-D_2(M-m)} \le\lambda_M2^{-D_2d}. \tag{101}\] For a separate post-assembly query with constant positive budget drop, use \(d=1\) in this inequality. Replacing the dyadic group sum by the larger sum over positive integers, all coupled child charges are at most \[ 2^{-p}\lambda_M\sum_{d=1}^{2h_*}2^{-(D_2-C)d}. \tag{102}\] Counts of budget layers and dependency types are included in \(2^{C d}\). Groups with larger slack have identical children and contribute zero. Together with Equation (99), this is the complete coupling recurrence.

Choose \(D_2>C\) so large that the infinite geometric sum in Equation (102) is at most \(1/4\). Then choose \(D_1>D_2+1\). Finally choose rational \(\epsilon>0\) so small that \[ \frac1{2s\epsilon}>C+D_1+2. \tag{103}\] Since \(h_*\le\epsilon\log_2B\), the direct relative charge satisfies \[2^{C h_*}B^{-1/(2s)} \le 2^{-(1/(2s\epsilon)-C)h_*} \le\tfrac14\,2^{-D_1h_*} \le\tfrac14\lambda_M,\] after increasing the fixed minimum \(B\) to absorb any fixed coefficient. The child and direct charges fit within Equation (100), even with room for fixed norm conventions already included in \(C\).

Moreover \(0\le M-M_b<h_*\) implies \[\lambda_M\le2^{-(D_1-D_2)h_*}.\] This is uniformly as small as required after the same minimum-size padding. The triangle inequality therefore adds the coupling bound to the spare clean active bound, for example \(C_b2^{-p}/2\), without exceeding \(C_b2^{-p}\). Inactive target bounds are the previously fixed bounded-range bounds. This completes all three inductions.

The order of choices is consequential. Fix the analytical hierarchy, table thresholds, and trivial-case bound first. Choose the precision and cutoff constants for the mean comparisons; then the compression floors, capacities, and walk lengths so the clean sums converge with their spare coefficients. Fix clamping ranges and the crude exponents in the fixed-variant comparisons next. These choices determine \(C\). Only after them choose \(D_2,D_1,\epsilon\) as above. Thus reducing \(\epsilon\) to obtain the direct collision gain does not alter an earlier clean-coefficient requirement.

Later choices of the common stage-size bound, top-budget coefficient, and environment-field exponent do not change this crude exponent \(C\). Immediate descriptor counts depend on relative slack, fixed support bounds and fixed walk length. The sampler gap is uniform in both field size and the length of its single free-bit component. Likewise, an enlargement of the fixed identifier-length coefficient only changes the polynomial-evaluation collision term in Lemma 39 to at most \(2c_{\mathrm{ID}}/B^2\); increasing the fixed minimum \(B\ge c_{\mathrm{ID}}\) bounds it by \(2/B\). These enlargements can therefore be absorbed after \(D_2,D_1,\epsilon\) have been fixed. ◻

Theorem 46 (Error bounds for the actual estimators). For every allowed task, with precision \(p\) defined in Equation (64), the actual row outputs, including the pre-incoming scores and support witnesses, satisfy \[\mathfrak d_{u,a}(\bar H,H)\le C_b2^{-p}.\] The clipped vectors have the corresponding ordinary error bound. In particular, every allowed reward query satisfies \[ \frac1N\mathbb E_\sigma\sum_{x\in V_l}\frac{X_u(x)}u |\bar W_l(x)-W_l(x)|^s\le C_b^s2^{-sp}. \tag{104}\] These estimates are uniform over the allowed task parameters and stage, and the constants are independent of the input and of \(B\).

Proof. Proposition 45 adds the actual–clean discrepancy to Proposition 38 with its spare constant. The inactive-task bounds follow from the fixed bounded ranges and supports. The induction in Proposition 45 includes both kinds of tasks and every prescribed output. Applying its ordinary vector statement to \(W_l\) and taking the \(s\)-th power gives Equation (104). ◻

Catalytic key equality

The key is short, but retaining a \(\log B\)-bit key or index at each tight recursive frame would still be too costly. The following exact identity gives a comparison with only a constant number of endpoint excursions. Its use in the space accounting requires the restoring endpoint interfaces constructed below. The cancellation of arbitrary initial contents also appears in transparent programs for catalytic computation (Buhrman et al. 2014, sec. 3). Here the shared bit vector is included in the ordinary work-space bound.

Lemma 47 (Restoring equality test). Let \(V\in\{0,1\}^B\) be a shared vector with arbitrary initial contents. Suppose an excursion to either of two deterministically specified endpoints returns the cursor, and its path and result are independent of the incoming contents of \(V\). Such excursions may use nested procedures which restore \(V\). There is a procedure which xors the bit \(\mathbf1_{k(v)=k(w)}\) into a local result bit, restores \(V\), and needs only constant local data across four endpoint excursions.

Proof. Starting with result bit \(r_0\), perform the following four excursions: flip \(V_{k(v)}\); xor \(V_{k(w)}\) into the result; flip \(V_{k(v)}\) again; xor \(V_{k(w)}\) into the result again. The key is evaluated only at the endpoint currently visited. If \(V^0\) is the incoming vector, the first read equals \(V^0_{k(w)}\mathbin\oplus\mathbf1_{k(v)=k(w)}\) and the second equals \(V^0_{k(w)}\). Thus the final result is \(r_0\mathbin\oplus\mathbf1_{k(v)=k(w)}\) and both flips restore \(V^0\).

For nesting, use induction on the well-founded procedure order. A nested call sees some arbitrary incoming vector, restores precisely that vector, and returns a result independent of it. Hence the four outer endpoint paths and their keys are the same on repeated visits, even between the two outer flips. The terminal evaluation, read, or flip is a local operation with no further structural call. Only a phase and a result bit must survive across the excursions; a key index is recomputed at each visit in shared scratch. The result bit being accumulated is not used to choose either endpoint path. These facts give the induction invariant and the claimed local storage bound. ◻

For \(d>2h_*\), exact bit-by-bit ID comparison is retained. A bit position then takes \(O(\log B)=O(d)\) suspended bits, with constants allowed to depend on the already fixed \(\epsilon\). For small slack, Lemma 47 avoids that per-frame index. The shared vector itself occupies \(B\) bits and, by the lemma, may be reused by all nested comparisons. This algebra establishes restoration; the controller and traversal section supplies the deterministic excursion and call-order interfaces needed to apply it.

The endpoint action evaluates the key with the invoking comparison’s effective salt, even when the excursion visits a lower-budget table. It makes no further access query and does not choose or alter the excursion. Thus the lower path remains independent of later salts. Lemmas 66 and 68 implement this separation for all restoring accesses, including inverse queries.

Finite numerical queries with bounded suspended storage

The preceding sections define finite estimator tables and bound their statistical errors. We now compute digits of those estimator values. For a fixed table tuple, including its statistical budget, the structural conventions are fixed and each entry has one exact mathematical value. Increasing the numerical clock requests a later digit of that same value. Each request is a finite deterministic computation.

A query returns one bounded digit. During its evaluation, the space needed by a suspended caller must be bounded by the size of its structural descriptors and a constant times the decrease from its output clock to the requested operand clock. An operand clock may exceed the output clock by an amount proportional to descriptor size; the structural term also pays for this overshoot. These differences will telescope when arithmetic queries are nested inside graph procedures. Thus the relevant question is which data remain live at each operand request.

Streams, records, and uniformity

We use integer clocks \(t,j\geq 0\). Known binary scales are separated from numerical values, so a normalized value belongs to \([-1,1]\). The fixed digit alphabet is \[\mathcal A=\{-6,-5,\ldots,5,6\}.\] Every digit is encoded by a fixed number of bits. An operand stream \((x_j)_{j\geq0}\) denotes the exact real number \[ x=\sum_{j\geq0}x_j2^{-j},\qquad x_j\in\mathcal A. \tag{105}\] In particular, for its truncation \(x^{[K]}=\sum_{j=0}^Kx_j2^{-j}\), \[ |x-x^{[K]}|\leq 6\,2^{-K}. \tag{106}\] The stream need not be a usual binary expansion, and no uniqueness of the expansion is required. Redundant signed-digit representations have a long history in arithmetic (Avizienis 1963). The redundancy lets consecutive finite approximations define an exact stream even when they use different truncations.

An arithmetic root has public parameters, its clock, and a description of its operand interfaces. We require these parameters to be regenerable from the outer controller records without operand or graph queries. They are fixed inputs to each numerical computation; their full binary encodings are not copied at every gate. An operand label specifies its interface and descriptor; it never specifies an unknown vertex address from a previous cursor position.

The following lemma reduces each digit query to a fixed number of residue bits.

Lemma 48 (Exact streams from approximation residues). Suppose \(|z|\leq1\) and, for each \(t\geq0\), a deterministic integer \(N_t(z)\) satisfies \[ |N_t(z)-2^tz|<2. \tag{107}\] Set \[ z_0=N_0(z),\qquad z_t=N_t(z)-2N_{t-1}(z)\quad(t\geq1). \tag{108}\] Then \(z_t\in\mathcal A\) and these digits represent \(z\) exactly. For any fixed integer \(q\geq4\), each output digit is determined by the required \(N_t\) residues modulo \(2^q\).

Proof. Writing \(e_t=N_t-2^tz\) gives \(z_t=e_t-2e_{t-1}\) for \(t>0\), hence \(|z_t|<6\); the weaker alphabet bound \(|z_t|\leq6\) is convenient throughout. At place zero, \(|N_0|<3\), so \(|N_0|\leq2\). The partial sum through place \(T\) is \(2^{-T}N_T\), which tends to \(z\) by Equation (107). Finally, reduction modulo \(2^q\) is injective on the possible signed digits because \(2^q>12\). Thus the signed difference can be recovered from the two residues without forming the full integers \(N_t\). Concretely, for \(t>0\) compute \((N_t\bmod 2^q)-2(N_{t-1}\bmod 2^q)\) modulo \(2^q\), then use the fixed finite table which sends each occurring residue to its unique representative in \(\mathcal A\). At \(t=0\), decode the residue of \(N_0\) in the same way. This final decoding has constant storage and introduces no comparison of unknown exact real numbers. ◻

The integer \(N_t\) is always defined before taking its residue. Different places may use different truncations or finite circuits. Only determinism and Equation (107) are required; consistency between the finite approximations is supplied by Equation (108). In particular, when a digit at \(t\) uses \(N_{t-1}\), the latter is evaluated by its own place-\((t-1)\) definition.

We compute the required residues by finite Boolean circuits. To bound what survives an operand query, we evaluate those circuits from paths of child choices and regenerate their absolute addresses as needed. Here \(d\geq1\) measures the size of the immediate structural descriptors; in the application it is the current group’s slack.

The local record of a suspended numerical query contains every local bit that must be retained across an operand query: its binary gate path, return modes, pending bounded results, stored descriptor choices, and any explicitly specified retained word. The input and master choice, the common vertex cursor, and one reusable scratch area are shared facilities. Temporary gate addresses, digit places and scan counters may occupy scratch while they are being computed. Any such data needed after a recursive graph query must either be in the charged local record or be reconstructible from the root parameters and that record. An operand-dependent word retained across the query is charged even if it was originally formed in scratch.

In particular, reconstructing an absolute digit place may use a binary counter in scratch, while the suspended record stores only the choices needed to recover that place. The next lemma makes this distinction precise for the finite circuits used below.

Lemma 49 (Evaluation from regenerated gate addresses). Consider a finite Boolean circuit of bounded fan-in, with a specified output gate. Conditional gates may evaluate their selector first and then only their selected branch. Suppose the type and immediate predecessors of a gate, or its operand label and digit place if it is an input gate, can be computed from the root parameters and a path of child choices from the output gate. Suppose these computations use \(O(B+t+d)\) scratch bits and make no operand or graph queries. Then depth-first evaluation retains \(O(1)\) bits per gate on its current path, and requires no saved absolute gate address per gate. Its suspended record above an operand input has length at most a constant times the length of that path, plus any explicitly retained root-local record.

Proof. At a gate, save its next-child or return mode and the already computed Boolean input, if any. At a conditional gate also save the selected branch bit. The bounded fan-in makes this a constant-size record. To locate the next child, scan the saved path from its root and regenerate successive gate descriptions in scratch. The final description determines the next child choice. On return, its Boolean result is combined with the saved bit and the local record is popped. All temporary addresses and scan counters are discarded before an operand query; a later scan recreates them from the same root and saved choices. Circuit acyclicity gives termination. This is an evaluation of the circuit unfolded into a formula; repeated gates are recomputed, so no array of previously computed values is stored. ◻

Truncations, comparisons, and binary arithmetic

Lemma 50 (Finite comparisons and bounded elementary operations). The normalized operations of taking a pair average, a maximum or minimum, and a fixed binary shift admit deterministic integers satisfying Equation (107). Their residue procedures, for any fixed \(q\), have the following properties. There are constants \(C_0,C_1\) such that an operand place \(j\) obeys \(j\leq t+C_0\), its suspended local record has length at most \(C_0+C_1(t-j)\), and the intrinsic local space is \(O(t+1)\). Fixed finite compositions of these normalized operations have the same bounds, with enlarged constants. Fixed Boolean operations on returned digit codes, including sign decoding and conditional selection of code bits, add only constant local overhead.

Proof. First consider comparing two truncations through \(K=t+h\), where \(h\) is a sufficiently large fixed integer. Their difference digits have magnitude at most 12. Scanning from place zero to \(K\), update a signed prefix by \(p\leftarrow2p+\delta\). While \(|p|\leq12\), retain \(p\) exactly. Once \(p>12\) or \(p<-12\), retain only the corresponding absorbing sign state. This sign is correct for every finite continuation: the remaining suffix, in the current scale, has magnitude strictly less than 12. The state space is finite, so each transition has a fixed bounded-fan-in Boolean circuit. The sign at the end determines the exact order of the two finite truncations. The path from that sign bit to an input at place \(j\) has length \(O(K-j+1)=O(t-j+1)\).

Let \(v_t\) be the average, maximum, or minimum of the two truncations, and put \(N_t=\lfloor2^tv_t\rfloor\). Each operation is 1-Lipschitz for the sup norm on its operands. Consequently \(|v_t-v|\leq6\,2^{-K}\), where \(v\) is the corresponding exact output. Choosing \(h\) so that \(6\,2^{-h}<1/2\) gives \(|N_t-2^tv|<3/2\).

To obtain \(N_t\bmod2^q\), only a fixed number of low places of the selected truncation are needed. More explicitly, at scale \(2^K\), terms of input place \(j\leq t-q\) are multiples of \(2^{q+h}\). Modulo this modulus the remaining \(q+h\) places suffice. For an average with integer input prefixes \(A_K,B_K\) at scale \(2^K\), use \(N_t=\lfloor(A_K+B_K)/2^{h+1}\rfloor\). Its residue needs modulus \(2^{q+h+1}\) and one additional retained place: only places \(j\leq t-q-1\) can be omitted. The finite number of remaining digits can be added and shifted by a constant-size Boolean circuit. A maximum or minimum first uses the comparison bit to select these residue bits. Its long comparison branch has the path bound proved above, and its residue branch reaches only places within a fixed distance of \(t\). A fixed shift changes these places by a fixed amount. Fixed normalization factors are handled in this way whenever an intermediate range exceeds \([-1,1]\).

The topology of the comparison chain and the low-residue circuits depends only on the root clock, the fixed margins, and the gate path. Regenerating a place by replaying that path needs no saved binary place counter at each transition. Lemma 49 therefore gives the suspended-space bounds. Their longest paths have length \(O(t+1)\).

For composition, use one common coefficient of \((t-j)\) for the finite primitive list. Increasing such a coefficient by \(\Delta\) requires increasing its constant term by at most \(\Delta C_0\), because the operand overshoot is at most \(C_0\). With these common coefficients, the clock differences telescope along the composed path. The finite number of constant terms remains a constant. ◻

We will also use the following explicit Boolean arithmetic facts. A carry bit in binary addition is described by a generate/propagate pair \((g,p)\). Composition of adjacent intervals is the associative operation \[(g_2,p_2)\circ(g_1,p_1) =(g_2\lor(p_2\land g_1),\ p_2\land p_1).\] Balanced trees of these pairs give every carry in \(O(\log(W+1))\) depth for \(W\)-bit words. A balanced tree adding \(R\) such words modulo \(2^W\) has depth \[ O\bigl(\log(R+1)\log(W+1)\bigr). \tag{109}\] The trees can be padded by zero words to powers of two, so their gate addresses are computable by ordinary index arithmetic. For counts of \(R\) Boolean bits one can do better: constant-depth full adders replace three same-width rows by two rows of their sum, with carries placed in the next column. In \(O(\log(R+1))\) rounds only two rows remain. The word width is \(O(\log(R+1))\); a final carry computation gives total depth \(O(\log(R+1))\). These constructions use bounded fan-in throughout.

Multiplication by minimum-index bands

Grouping products by their smaller input index leaves a gap between the output clock and the requested input clocks; that gap pays for the retained circuit path within each band.

Lemma 51 (Multiplication). For normalized exact input streams \(x\) and \(y\), multiplication has all the conclusions of Lemma 50. In particular, its suspended storage at an operand place \(j\) is \(O(1)+c(t-j)\), despite the potentially long product expansion.

Proof. Fix \(q\) and write \(A=6\). Partition ordered pairs \((i,j)\) of nonnegative places into a zero band, where \(\min(i,j)=0\), and positive bands \[\mathcal P_D=\{(i,j):D\leq\min(i,j)<2D\}, \qquad D=1,2,4,\ldots.\] Set \(h_0=h_\circ\) for the zero band and \(h_D=h_\circ+\lfloor D/4\rfloor\) for positive bands. The fixed integer \(h_\circ\) will be chosen below. Retain in band \(D\) only pairs with \(i+j\leq t+h_D\), and let \[ S_D=\sum_{\substack{(i,j)\in\mathcal P_D\\i+j\leq t+h_D}} x_i y_j2^{t+h_D-i-j}. \tag{110}\] For the zero band the same formula uses its defining pair set. These are mathematical integers; no bound on their full binary length is assumed in the residue computation.

At any fixed diagonal \(i+j=s\), a positive band contains at most \(2D\) ordered pairs. Indeed the smaller place has \(D\) possibilities, and an off-diagonal pair has two orientations. The zero band has at most two pairs. Thus, in units of \(2^{-t}\), the total omitted product terms have absolute sum at most \[ E_{\mathrm{omit}} \leq 2A^2\sum_{D\in\{0,1,2,4,\ldots\}}(D+1)2^{-h_D}. \tag{111}\] The geometric summation of the diagonal weights gives this bound also when the truncation threshold lies below the first possible diagonal. Both the series in Equation (111) and \(\sum_D2^{-h_D}\) converge. Choose \(h_\circ\) sufficiently large that their displayed combined error bound, including coefficient \(2A^2\), is less than \(1/2\).

There are only finitely many nonempty retained bands. For a positive one, \(2D\leq t+h_D\), whence \[ D\leq\frac47(t+h_\circ). \tag{112}\] The left side of the nonemptiness condition grows monotonically faster than its right side along the dyadic bands. We may therefore stop at the last band satisfying this inequality, retaining empty bands within the finite prefix if necessary.

Starting with the most distant retained band, define integers recursively by \[ F_D=S_D+ \left\lfloor 2^{h_D-h_{D^+}}F_{D^+}\right\rfloor, \qquad N_t=\left\lfloor2^{-h_0}F_0\right\rfloor, \tag{113}\] where \(D^+\) is the next band, and the last missing tail is zero. One internal floor, divided back to scale \(2^t\), changes the sum by less than \(2^{-h_D}\). Induction down the bands consequently shows that the difference between \(2^{-h_0}F_0\) and the sum of all retained terms in scale-\(t\) units is at most \(\sum_D2^{-h_D}\). Equation (111) and the choice of \(h_\circ\) make the error before the final floor less than \(1/2\). The final floor costs less than one more unit. Therefore Equation (107) holds for \(z=xy\).

We next implement Equation (113) using residues. At band \(D\) use modulus \(2^{q+h_D}\). An integer term in Equation (110) with \(i+j\leq t-q\) is a multiple of this modulus and can be omitted from the circuit. It remains present in the mathematical integer \(S_D\). Only the diagonals \[ t-q<i+j\leq t+h_D \tag{114}\] need to be processed, giving \(O((D+1)(h_D+q+1))=O((D+1)^2)\) candidate pairs.

Signed floor shifts are well defined on the required residues. For \(h'\geq h\) and \(a-b=k2^{q+h'}\), \[ \left\lfloor\frac{a}{2^{h'-h}}\right\rfloor -\left\lfloor\frac{b}{2^{h'-h}}\right\rfloor =k2^{q+h}. \tag{115}\] This identity holds for all signed integers \(a,b\). Thus a next-band residue modulo \(2^{q+h'}\) supplies the correctly floor-shifted tail modulo \(2^{q+h}\). Apply the same observation to the final shift by \(h_0\) to obtain \(N_t\bmod2^q\). In the finite binary circuit, this shift just discards the lowest \(h'-h\) bits of the canonical \((q+h')\)-bit residue. The displayed identity proves that this is valid also when the underlying integer is negative.

For completeness, the finite gate structure can be indexed without storing an absolute operand place at each gate. For a positive band enumerate triples \((r,s,o)\), where \(D\leq r<2D\), \(t-q<s\leq t+h_D\), and \(o\in\{0,1\}\). The triple supplies \((r,s-r)\) if \(s-r\geq r\) and \(o=0\); it supplies \((s-r,r)\) if \(s-r>r\) and \(o=1\); otherwise it supplies a zero term. This enumerates the relevant pairs exactly once. For the zero band use \(r=0\) with the same orientation rule. Negative candidate places are invalid and supply zero. The offsets \(r-D\) and \(s-(t-q+1)\) have \(O(\log(D+2))\) bits. They and \(o\) are obtained from paths in padded binary selection trees. The full places are then computed from the root clock in scratch.

Each nonzero term is a constant-bit signed product \(x_i y_j\), shifted into a word of width \(W_D=q+h_D=O(D+1)\) and interpreted modulo \(2^{W_D}\). A bit of this word is a fixed Boolean function of the two bounded digit codes, the bit index and the known shift; negative products use their ordinary residue, equivalently sign extension before reduction. Add these words and the shifted tail using Equation (109). A band contributes depth \(O(\log^2(D+2))\). Bit selection for the shifted tail is wiring and does not require evaluating a full tail word.

The sum of these depths through the bands up to \(D\) is \(O(\log^3(D+2))\), hence at most \(C(D+1)\) for a fixed \(C\). For an original input digit in a positive band, its partner place is at least \(D\). Equation (114) gives \[ j\leq t+h_D-D \leq t+h_\circ-\frac34D. \tag{116}\] Equivalently, \(t-j\geq\tfrac34D-h_\circ\), so the full path cost through these bands satisfies \[C(D+1)\leq C\left(1+\frac43h_\circ\right) +\frac{4C}{3}(t-j).\] This gives \(C_0+C_1(t-j)\) with fixed constants. The zero band has constant depth and only constant overshoot. All its gate descriptions, term coordinates, bit positions and tail shifts are computable by replaying paths and doing integer index arithmetic from \(t\) and \(q\). Lemma 49 converts these depth bounds to suspended storage. Finally, Equation (112) bounds every intrinsic path and local word width by \(O(t+1)\). ◻

Reciprocals and the early-word convention

A prefix through roughly half the output clock makes the remainder in a first-order reciprocal expansion quadratically small. Releasing that prefix before later input queries keeps its long word from surviving a call with little remaining clock gap.

Lemma 52 (Reciprocal on a positive compact interval). Fix constants \(0<a\leq b\). On inputs \(x\in[a,b]\), a reciprocal, after a fixed known normalization, has the same operand-clock, suspended-record, and intrinsic-space bounds as in Lemma 50. Its local record may contain an \(O(t+1)\)-bit word only while querying input places \(j\leq\lfloor t/2\rfloor+O(1)\).

Proof. The input can first be normalized by a fixed binary scale. It is enough to prove the result with a fixed digit alphabet and fixed positive interval; all constants below may depend on \(a,b\). Choose a fixed power of two \(C\) large enough that the desired output \(z=1/(Cx)\) lies in \([0,1]\) and that the coefficients below have absolute value at most one. Put \[k=\lfloor t/2\rfloor+c,\qquad x^0=x^{[k]},\qquad e=x-x^0,\] where \(c\) is a sufficiently large fixed integer. By Equation (106), \(|e|\leq 6\,2^{-k}\). Choosing \(c\) first ensures \(x^0\in[a/2,2b]\) and \(|e|\leq1\) at every \(t\). Increasing \(C\) if necessary gives bounded coefficients \[\alpha=\frac1{Cx^0},\qquad \beta=-\frac1{C(x^0)^2}.\] The exact algebraic identity \[ \frac1{Cx}=\alpha+\beta e+ \frac{e^2}{C(x^0)^2x} \tag{117}\] shows that the linear expression \(T_t=\alpha+\beta e\) differs from \(z\) by at most a fixed constant times \(2^{-2k}\). Choose \(c\) large enough that \[ 2^t|T_t-z|\leq\frac14 \quad\hbox{for every }t\geq0. \tag{118}\] Clipping \(T_t\) to \([-1,1]\) cannot increase this error; denote the clipped expression by \(\widehat T_t\).

For this fixed original \(t\), the coefficients \(\alpha,\beta\) are exact rational numbers. Their digits are produced by integer arithmetic, as described below. The suffix \(e\) has the exact stream with digit zero through \(k\) and digit \(x_j\) thereafter. Thus Lemmas 50 and 51 provide exact streams for \(\widehat T_t\), using only a fixed number of normalized operations. Explicitly, form the product \(\beta e\), then the pair average \(u=(\alpha+\beta e)/2\), clip \(u\) to \([-1/2,1/2]\) by a maximum and a minimum, and finally apply the fixed shift by 2. Every exported intermediate value is normalized, and the final value is \(\widehat T_t\). Forming this expression introduces no additional reciprocal operation.

Let \(P_s(\widehat T_t)\) be the integer approximation furnished by those operations at place \(s=t+3\), and define \[ N_t(z)=\left\lfloor P_{t+3}(\widehat T_t)/8\right\rfloor. \tag{119}\] The cutoff \(k\) and the coefficients remain those of the original \(t\), not those of \(s\). Equations (107) and (118) give \[|N_t(z)-2^tz|<1+\frac28+\frac14=\frac32<2.\] To obtain this integer modulo \(2^q\), ask for the approximation \(P_{t+3}\) modulo \(2^{q+3}\) and use Equation (115).

It remains to justify storage in the coefficient computations. The signed numerator of \(x^0\) in units \(2^{-k}\) has \(O(k+1)\) bits. Form it by querying places \(0,\ldots,k\) and updating the integer prefix. The numerator, a place counter, and all data that must survive the next query are part of the local record. To return a coefficient digit of place \(r\), numerator squaring and integer division require \(O(k+r+1)\) bits. Here \(r\leq t+O(1)\), because the enclosing expression has fixed depth and constant overshoot. After the input prefix is obtained, these integer computations make no further operand queries. Standard binary long multiplication and division, performed by bounded loops with their partial products or remainder, use only this linear number of bits. For \(r>0\), computing floors of the coefficient scaled at \(r\) and \(r-1\) gives its exact difference digit; at \(r=0\) use the place-zero floor alone, as in Lemma 48. Clipping is unnecessary because the coefficients were normalized in advance.

Every original-input query during this formation has \(j\leq k\leq t/2+O(1)\). Hence \[ t+1\leq C_2+C_3(t-j) \tag{120}\] for fixed constants, at all such queries. The enclosing gate path also has length \(O(t+1)\), so the complete suspended local record satisfies the required bound there. Equivalently, the coefficient branch is charged against the original reciprocal clock \(t\), not against its possibly small coefficient digit place \(r\).

On a suffix branch with \(j>k\), the coefficient word is absent. The Boolean evaluation asks for a coefficient digit at a separate leaf, returns its bounded code bits, and discards the coefficient word. If the other child of a gate subsequently asks for a late suffix digit, only the bounded pending bits and gate choices are retained. The path bound from the preceding arithmetic lemmas therefore gives \(O(1)+c(t-j)\) also on this branch. Repeated coefficient requests rebuild the word. The parameter \(k\) and the choice of coefficient are regenerated from the original \(t\) and the saved gate path; no earlier cursor address is required. All intrinsic records have \(O(t+1)\) bits. ◻

The coefficients in this proof are operand-dependent, even though they are rational once \(x^0\) is known. They therefore do not qualify as public coefficient leaves. A coefficient request first forms its charged early prefix, then uses scratch for integer arithmetic, returns only the requested bounded digit, and releases that prefix before a suffix branch is evaluated. The original reciprocal clock \(t\) remains regenerable throughout. This order of operations is the lifetime rule used in the suspended-record bound. In particular, \(j\le\lfloor t/2\rfloor+c\) gives the explicit payment \(t+1\le2(t-j)+2c+1\) for the early word.

Indexed ranks, sums, scales, and structural leaves

Lemma 53 (Indexed finite ranks). Fix \(\kappa>0\) and \(d\geq1\). Consider at most \(2^{\kappa d}\) normalized values, indexed by descriptors of length at most \(\kappa d\), and an order statistic specified by an \(O(d)\)-bit index. The semantic list length \(n\) is public; the requested index is in \(\{1,\ldots,n\}\), or an out-of-range request returns a prescribed normalized constant. Absent descriptors may be assigned prescribed constant dummy values by structural validity predicates. This order statistic has deterministic integer approximations \(N_t\) satisfying Equation (107). Their residue procedures have constants \(a,b,c\), depending only on the descriptor bounds, such that an operand digit query at \(j\) obeys \[ j\leq t+cd,\qquad L\leq ad+b(t-j), \tag{121}\] where \(L\) is the total suspended local record length. A structural validity query is assigned clock zero and has suspended length \(O(d+t)\). The intrinsic local length is \(O(d+t)\).

Proof. For \(1\le k\le n\), the \(k\)th largest of the semantic \(n\)-entry list is its \((n-k+1)\)st ascending order statistic, before any extra padding. Pad to a fixed descriptor domain whose cardinality is a power of two, adding entries of value \(+1\). These entries do not change any of the first \(n\) ascending order statistics. An empty list or an out-of-range request returns its specified default immediately. Replace invalid original descriptors by the prescribed dummies. An operand access evaluates validity first and enters the actual value branch only when valid; an invalid access returns the dummy’s digit. Thus a dummy does not require a query to an undefined value. For each valid numerical value use its truncation through \(K=t+h\) for fixed large \(h\). Order these finite rational values, breaking ties by descriptor index. Compare every pair using the signed comparison of Lemma 50. For each descriptor count how many descriptors precede it, then compare that count to the requested rank minus one. Exactly one descriptor is selected. Select its low-residue bits using Boolean AND and OR trees and finally floor to output scale \(2^t\). The selected descriptor is internal to this residue computation. It may change with \(t\); the exact order statistic being approximated is fixed by the semantic list and the requested index.

If all coordinates of two finite lists differ by at most \(\eta\), their corresponding order statistics differ by at most \(\eta\). Indeed, at least the specified number of the perturbed entries are no larger than the original order statistic plus \(\eta\); interchanging the lists gives the reverse inequality. Here \(\eta=6\,2^{-K}\). Thus the selected truncation, followed by the final floor, satisfies Equation (107) for sufficiently large fixed \(h\). Descriptor ties affect neither this estimate nor the uniqueness of the finite selection. This argument perturbs the entire padded exact list. Some truncated original entries may exceed \(+1\) slightly, but their error, and hence the error of the selected padded order statistic, is still at most \(\eta\). The padding argument requires no separate claim that every finite truncation remains in \([-1,1]\).

Each count has \(O(d)\) bits. Carry-save counting as described above, the comparison to the requested rank, and selection among \(2^{O(d)}\) descriptors add \(O(d)\) Boolean depth above the pairwise comparison bits and selected residue bits. A place-\(j\) input to a pairwise comparison consequently has path length \(O(d)+O(t-j)\); a residue input is within a fixed distance of \(t\) and has path length \(O(d)\). The paths to validity predicates have length at most the whole circuit depth \(O(d+t)\), and the predicates themselves are separate structural leaves of clock zero.

Each gate is specified by its pair of descriptor indices, its counting-tree or selection-tree position, and a comparator or bit position. The paths determine these indices in \(O(d+t)\) scratch, using only the fixed domain and root parameters. No numerical comparison is needed to determine the circuit topology. Apply Lemma 49. The constants in Equation (121) can absorb the fixed digit overshoot because \(d\geq1\). ◻

For the compression ranks, copies of a previously computed baseline are ordinary indexed operands: their digit interfaces query the baseline’s exact stream one bounded digit at a time. A slot that chooses between an entry and the baseline first evaluates its structural flag and then queries the selected, already-defined stream. Treating this as a defined operand interface leaves no absent value for the rank procedure to evaluate. The required baseline copies belong to the semantic list; the extra power-of-two padding in the proof is then supplied by the constant \(+1\). The structural choice adds only a fixed conditional branch, so the same digit-clock and suspended-record bounds apply.

A rank over the full incoming vertex universe need not have a \(2^{O(d)}\) descriptor domain. That operation uses the terminal configuration procedure in the next section.

Lemma 54 (Bounded sums and known scales). Suppose an immediate descriptor domain has at most \(2^{\kappa d}\) terms, all required binary scale exponents have magnitude at most \(\kappa d\), and every intermediate output has a known range bound \(2^{\kappa d}\). A sum over an empty domain is zero, and an average is taken over a public nonempty domain. Balanced sums or averages of these terms, followed by the required known rescalings and bounded clipping, satisfy Equation (121) and have intrinsic local space \(O(d+t)\).

Proof. Choose one common known power-of-two range bound for the immediate terms and divide each term by it, so every term lies in \([-1,1]\). Pad their number to a power of two \(2^r\), \(r=O(d)\), and use a complete binary tree of pair averages. Every internal value stays normalized. Along one path there are \(r\) elementary operations. Choose common coefficients in their bounds, as in the last paragraph of the proof of Lemma 50. The constant terms sum to \(O(r)\) and the differences of successive digit clocks telescope. Thus the local length is \(O(d)+b(t-j)\) and the total overshoot is \(O(d)\). Restore the common normalization scale only after the tree of pair averages. For a sum, combine this factor with the known padding count \(2^r\) in one binary shift of magnitude \(O(d)\). For an average over an original domain of size \(m\), combine the scale restoration with multiplication of the padded average by the public factor \(2^r/m\). With minimal padding the latter factor is in \([1,2)\), so it requires only one normalized public multiplication and a fixed shift in addition to the common scale restoration.

More explicitly, for a shift by \(2^v\), truncate the operand at \(K=\max(0,t+v+h)\). If \(t+v+h\geq0\), its scaled truncation error in scale-\(t\) units is at most \(6\,2^{-h}\). If the scaled output is known to be normalized, floor in the desired output scale. Otherwise, when a bounded clip is required, apply that clip to the scaled truncation before flooring. For the normalized clip to \([-1,1]\), for example, define the approximation using \(\min(1,\max(-1,2^v x^{[K]}))\); its error cannot increase. This shift-and-clip procedure does not first export an unbounded intermediate as a normalized stream. The low residue requires the constant number of places near \(t+v\), and a comparison needed for clipping has a constant-state chain through \(K\). Comparisons to the scaled clipping endpoints are equivalently comparisons of \(x^{[K]}\) to the known endpoints divided by \(2^v\); their finite constants and shifted place data are public. When \(v\geq0\) these endpoints belong to \([-1,1]\); when \(v<0\) the normalized operand is already in the clip interval after scaling, so no clipping comparison is needed. Its path cost is \(O(1)+O(\max(0,t+v-j))\), which is at most \(O(d)+O(t-j)\) on an actual operand query. If \(t+v+h<0\), zero is already a sufficiently accurate approximation after increasing \(h\), and no operand query is needed. A fixed number of such rescalings and clips preserves the bounds by telescoping. Known non-power-of-two coefficients can instead be handled as public normalized streams and bounded multiplication. ◻

Lemma 55 (Finite buffered tests). Let \(d\geq1\), let \(s\) and \(u\) have known magnitude bound \(2^{\kappa d}\), and suppose \(u\geq\rho\geq2^{-\kappa d}\). There is a deterministic finite test that accepts every \(s>2u\) and accepts only \(s>u\). It uses \(O(d)\)-place digit queries after normalization, with intrinsic and suspended local lengths \(O(d)\). Its result is independent of any subsequent numerical output clock.

Proof. Compute finite rational approximations \(\widetilde s,\widetilde u\) with individual errors at most \(\rho/16\). The known range and lower bound on \(\rho\) make \(O(d)\) normalized digits sufficient. Accept exactly when \(\widetilde s>\tfrac32\widetilde u\), using a fixed convention at equality. If \(s>2u\), the left side minus the right side is greater than \(u/2-5\rho/32>0\). If \(s\leq u\), it is at most \(-u/2+5\rho/32<0\). The finite comparison and all previous arithmetic have the bounds already proved. The test chooses its precision from \(d\) and \(\rho\), not from an enclosing numerical clock. With its structural starting clock defined as zero, its operand requests have \(j=O(d)\); increasing \(a\) in Equation (121) pays also for the term \(-bj\). ◻

The numerical interface

We state the composition result with its syntactic hypotheses. They concern finite arithmetic recipes and do not assert any property of the underlying graph operations.

Application to graph procedures requires a separate certificate. The semantic descriptor domains and structural validity conventions are fixed independently of the requested numerical clock; selected operands already have defined digit interfaces; reciprocal arguments lie in the stipulated positive intervals; and root parameters remain recoverable after lower calls. Clock-dependent truncations and comparison circuits may change \(N_t\), but not its target recipe value. Lemmas 67 and 68 verify these conditions for the estimator procedures. An incoming rank over the whole vertex universe has no prescribed \(2^{O(d)}\) descriptor bound and uses the separate construction of Lemma 64.

Theorem 56 (Uniform numerical interface). Fix the digit alphabet, positive reciprocal intervals, and constants bounding the following recipe class. A recipe at slack \(d\geq1\) has a fixed number of stages built from bounded arithmetic operations, fixed piecewise formulas using maxima, minima and clipping, the reciprocals of Lemma 52, indexed ranks as in Lemma 53, and sums and known scales as in Lemma 54. The stage and descriptor descriptions are generated from the root parameters by index computations using \(O(B+d+t)\) scratch and no operand queries, and the total non-reduction dependency depth is bounded independently of \(d\). Each indexed domain has size \(2^{O(d)}\) and descriptors of length \(O(d)\). Discrete domain and validity decisions are structural leaves or buffered tests with their own fixed precision conventions. Public coefficient leaves have exact normalized digit procedures using \(O(B+d+t)\) shared scratch and making no operand or recursive graph queries.

Then each normalized recipe value has a deterministic exact \(\mathcal A\)-stream. For its digit at clock \(t\) there are uniform constants \(a,b,c\) such that every requested operand digit \(j\) satisfies \[ j\leq t+cd,\qquad \text{total suspended local binary record length} \leq ad+b(t-j). \tag{122}\] An output code bit satisfies the same bounds. A structural leaf is requested at clock zero and has suspended local length at most \(ad+bt\). The maximum intrinsic local record length is \(O(d+t)\).

All gate, digit and descriptor addresses can be regenerated from root parameters and the saved local records using \(O(B+d+t)\) shared scratch and no recursive graph calls. When \(d,t=O(B)\) this is \(O(B)\) scratch. If a public rational coefficient has \(O(B)\)-bit numerator and denominator that can be generated in \(O(B)\) scratch, its queried digits in this clock range also require only that shared \(O(B)\) scratch.

Proof. The primitive constructions define deterministic integers \(N_t\). Lemma 48 turns them into exact streams. When composing primitives, their operands are these exact streams; numerical approximation errors are not accumulated along the recipe as if they were statistical errors. Instead, each primitive’s deterministic \(N_t\) is an approximation to its own exact mathematical value. The construction uses only residues of \(N_t\) and \(N_{t-1}\) and a constant number of output code bits, which adds fixed overshoot and fixed local overhead.

For the record bound, choose one common positive clock coefficient \(b\) for the finite primitive and stage types. In any bound \(L\leq a_0d+b_0(t-j)\) with \(j-t\leq c_0d\), replacing \(b_0\) by \(b\geq b_0\) is justified by replacing \(a_0\) by \(a_0+(b-b_0)c_0\). Thus the same \(b\) can be used everywhere. There are a bounded number of stages, each with overhead \(O(d)\); within a balanced reduction the elementary overheads have total \(O(d)\) by Lemma 54. Along any path the intermediate clock differences telescope, giving Equation (122) and total overshoot \(O(d)\). The same sum, terminated at an internal arithmetic operation rather than an operand call, gives the intrinsic bound explicitly as follows. If that operation has current clock \(u\), its own intrinsic record has length at most \(C(d+u)\), while the outer saved portion has length at most \(A d+b(t-u)\). Choose the common \(b\) also to satisfy \(b\geq C\). Their sum is then at most \((A+C)d+bt\), because \(u\geq0\). Increasing the stage constants as above makes this same choice of \(b\) valid in all local bounds. The early-word case is already included in Lemma 52: those retained bits are part of the local record and are paid for at the actual original-input clock by Equation (120).

Structural validity branches are finite branches in this circuit organization, with length at most its full intrinsic bound. They ask the structural operand at clock zero. Buffered tests are separate recipes with structural starting clock zero and precision fixed by Lemma 55; they are not altered when a surrounding value is requested more accurately. Thus no discontinuous comparison of exact reals at unbounded precision is part of the interface.

All circuit constructions were given with padded domains and index rules. A saved gate choice selects an operand, a subtree, a bit of a known-width word, or one of the finite primitive modes. Scanning these choices from the root recovers the current descriptor and digit places by integer arithmetic. The reciprocal’s extra records additionally specify its early-prefix formation state. None of these computations refers to an ancestor cursor address or requires an operand value to discover the next gate. Their integers have \(O(B+d+t)\) bits; repeated scans reuse one scratch area. Lemma 49 therefore applies at every Boolean portion of the recipe.

Finally, for a public coefficient \(P/Q\), \(Q>0\), producing a place-\(r\) approximation uses integer division of \(2^rP\) by \(Q\), with signed floors when needed. The inputs and intermediate remainders have \(O(B+r)\) bits. Standard finite binary division and multiplication use that much space, and extracting the low residue or difference digit requires no larger word. All operand bits in this computation are public, so it makes no recursive graph call and can use shared scratch throughout. For \(r=O(B)\) its scratch bound is \(O(B)\). In contrast, any word whose bits depend on recursive operands and must survive a new operand request has already been charged to the local record. ◻

Remark 57 (Use of clocks). Theorem 56 is a finite-request statement for every \(t\). Its exact-stream conclusion follows by considering these finite requests separately, and does not require a machine to choose structural domains or port orders at infinite precision. For the actual top-level computation all reached clocks will be \(O(B)\) by the allowance calculation in the next section. Known rate and field coefficients have \(O(B)\)-bit descriptions, and a fixed rational generator probability raised to a walk length \(O(d)\) has numerator and denominator of \(O(d)\) bits; these are instances of the public-coefficient provision above.

Uniform controller implementation

We now implement the table and vector queries defined in the preceding sections. Theorem 56 supplies their numerical operations. We must arrange the recursive graph accesses so that the records retained by all callers fit in logarithmic space at the same time.

We keep one vertex identifier, called the cursor. Consider a product term \(\mathcal F(x,z)\mathcal B_b(z,y)\) from Equation (58). For a selected valid product descriptor, a query begins with the cursor at \(x\). The correction port moves it to \(z\) and supplies a bounded inverse port. The right factor can then be queried at \(z\), using its own common table tuple; that tuple contains no copy of \(x\). Returning along the inverse port restores \(x\). Arithmetic requests the two factor digits separately, so it need not retain a complete left-factor value while computing a right-factor digit. If the operation instead asks to reach \(y\), the right factor’s terminal motion is the last recursive delegation; only local label completion follows it. These are the paths that Lemma 28 requires us to realize.

Incoming access reverses a different object: a finite program whose canonical initial states describe the forward row ports. We make the program deterministic with halting success states. The component of a specified success state is then a tree directed toward that state; its canonical initial states are precisely the valid row ports ending there. A contour of that tree enumerates them and returns to the success state without saving a second copy of its vertex address. This gives both capped incoming ports and incoming score ranks. To use the contour we must define program states and their predecessor queries carefully: working states of a completed lower query are not themselves part of the calling program graph.

The storage bound follows the same distinction. A caller retains its own descriptor choices and numerical path while a lower procedure runs. A slack-\(d\) layer pays \(O(d)\) structural bits; a query from digit place \(t\) to place \(j\) pays for its numerical records through the bound \(O(d)+c_{\rm num}(t-j)\) of Theorem 56, for a fixed coefficient \(c_{\rm num}\). We assign a decreasing allowance to the finite operation order within a prefix, the move to its preceding prefix of slack \(2d\), and the move to a child of budget at most \(M-d\). The resulting differences telescope over the physical stack. The proof below establishes these record bounds on the actual recipes before imposing guards on artificial program states.

Parameters and interfaces

Fix the input and a master environment. The latter includes the rank matrices and all the fingerprint salts. A table tuple consists of the estimator type and stage, its statistical budget, scale and usage rates, channel roles and suffixes, and its effective environment. The tuple specifies an entire table, rather than an individual row or entry. Usage rates belong to this tuple; the separate integer clock \(t\) specifies a requested numerical digit. A generator and wrap word is a call-record descriptor used to derive a child’s effective environment. It is not an additional semantic parameter of the child table: after this derivation, the child’s values, forward port orders and guards depend on the resulting effective tuple, not on the ancestral word which produced it. Distinct walk descriptors are still separate summands with their prescribed weights, even if they induce the same environment transformation. The variants of Lemma 21 can be generated without their conditioning vertices. A vector tuple is defined in the same way, omitting the unneeded terminal parameters. All admissible absolute tuples have \(O(B)\) bits.

The mutable cursor records a vertex and its stage interpretation. Other shared facilities are the input, the top root parameters, the master environment, a catalytic bit vector of length \(B\), and one reusable scratch area. Its size is \(O(B)\) in the actual computation, where the top numerical clock is \(O(B)\); for a general top clock \(t\) it may use \(O(B+M_{\rm top}+t+1)\) bits. Scratch may be used to perform arbitrary local integer computations in this space. It has no content which must survive a recursive graph operation. Every such operation starts its scratch calculations canonically. Before every input access, its requested position is regenerated in shared scratch. The physical input head stays between the real endmarkers; off-data simulated positions use the base symbol routine from Lemma 3. These rules also apply during inaccurate calculations and guarded artificial predecessor probes. The requested coordinate and scan counters coexist only within this pure routine and are discarded before a recursive graph operation. It is harmless to copy an address temporarily inside a scratch calculation, but such a copy is discarded before a recursive graph operation begins.

The finite interfaces to be constructed are as follows. A port is an index in a specified row class; the class may have holes.

  1. A restoring test returns one bit. A quantity query returns one bounded signed digit of a specified value. Both restore the cursor and the catalytic vector.

  2. A terminal motion starts at a row and a port and finishes at that entry’s endpoint. No return address is provided. An excursion makes the same visit, performs a prescribed local action there, and returns. The action may use scratch and the catalyst but makes no recursive graph query.

  3. A paired motion maps a valid row port to its endpoint and a bounded inverse port. Its inverse has the same table tuple. On invalid ports the restoring existence test returns false; an exported motion may instead return failure without moving the cursor.

At the final row interface all port alphabets are bounded by constants. At a prefix of slack \(d\), indices and immediate descriptors have \(O(d)\) bits. These two statements must be distinguished: an entire terminal path represented by one prefix port need not have a description of length \(O(d)\).

A port code may be in range and still name a hole in its class. Its existence test returns false, its quantity query returns the zero digit, and its motion or excursion fails with cursor and catalyst unchanged. Rank operands retain their separately prescribed baseline or padding values: an absent real port is not a virtual default coordinate.

Theorem 58 (Realization of the estimator interfaces). Fix the construction constants and an admissible top budget \(M_{\mathrm{top}}=O(B)\). For every actual task tuple reachable from it, the row, vector and prefix interfaces defined in the preceding sections admit deterministic uniform implementations by restoring queries and cursor motions. The following assertions hold for every master environment, without an accuracy promise on its estimated values.

  1. Every test or numerical digit query terminates, returns its defined value, and restores the cursor and catalyst. Its result is independent of incoming catalyst and scratch contents. Structural port and order choices are independent of a higher numerical request.

  2. Terminal motions give the specified endpoints. Their excursions restore the cursor. Capped correction motions are partial paired maps with bounded inverse ports and safe behavior on invalid ports. Incoming score ranks compute the stipulated zero-padded ranks.

  3. For a root numerical request at clock \(t\), all records retained by the call and its descendants have total length \(O(M_{\mathrm{top}}+t+1)\), and every descendant clock is \(O(M_{\mathrm{top}}+t+1)\); the shared facilities use \(O(B+M_{\mathrm{top}}+t+1)\) space. In particular an actual top request with \(t=O(B)\) uses total deterministic \(O(B)\) space.

These implementations use no estimator oracle and retain no second vertex address across recursive graph operations.

The proof also constructs every lower operation instance admitted by the finite parameter ranges and permitted parent-to-child recipes. This includes calls made while testing possible predecessor states, whether or not those states occur in a designated execution. The larger construction is needed to compute incoming accesses by inverse traversal.

Program states and record bounds

A terminal program has a fixed root tuple and root clock. Its initial configuration is the cursor \(x\), a distinguished initiation tag, and the original row port \(p\) in a separate recoverable field. Distinct ports give distinct initiation states. In a scoring program the success local state contains a quantized score; otherwise it has only a success tag. The starting row and port are not root parameters. The program universe thus includes all rows and all ports of this same table. Similarly, a scoring universe at a fixed precision includes all possible quantized scores.

There are two different forms of invocation inside this program. A restoring bit query, or a lower paired motion, is an opaque transition query: its temporary working configurations are not vertices of the terminal program graph. They are computed by the lower procedure when the transition is probed. A terminal delegation with no bounded inverse port is instead explicit: it pushes the lower terminal controller into the local state. Thus the vertices of the program graph comprise the cursor and a finite chain of explicit controllers, but not the working states of opaque transition queries. This distinction will also be used when graphs are traversed inversely.

Definition 59 (Regenerable record). A local record is a delimited binary word with a tag from a fixed finite alphabet. It may contain immediate descriptor choices, bounded port or copy symbols, offsets, Boolean evaluation paths and pending bits, and the early numerical words permitted by Theorem 56. Head positions in such words are marked locally. A root tuple, a root operation and its numerical clock, together with the active records, determine every descendant tuple and clock by a scratch computation. This computation does not call a graph procedure and does not use a cursor that has been left behind. We call such records regenerable.

The records use delimited encodings, with marked positions rather than an independent absolute stack pointer at every level. Self-delimiting encodings enlarge a nonempty word of length \(r\) by \(O(r+1)\) bits. Constant tags also identify a root boundary, the active candidate copy during a neighbor test, and the owner of a pending endpoint action. These tags permit a scratch scan to select the appropriate active view of the records. An inactive copy is not part of a descendant’s program state or call recipe.

A root boundary is a constant tag, not a stored copy of an absolute table tuple. The logical root parameters of a nested query are obtained by replay up to that boundary; the temporary absolute tuple is then discarded before another graph query. Thus the phrase “fixed root tuple” specifies which mathematical program graph is being used. It does not allocate a fresh \(O(B)\)-bit parameter word at every nesting level. Lemma 68 supplies the replay procedure.

Definition 60 (Raw allowance certificate). A raw controller is the caller’s own selection, arithmetic, scoring or terminal-path record, excluding the working states of a query it calls. In a terminal program graph an explicit chain contains only its terminal-path controllers and, in a scoring program, its score suffix or completed score word. The interiors of cap traversals, inverse score ranks, numerical digit queries and paired motions are opaque queries; they are never explicitly pushed working graphs in this chain.

A raw controller has an integer allowance \(w\ge1\) and a numerical clock \(t\ge0\). Fixed raw constants \(a_0,b_0,a'>0\) give the following certificate at a call to a lower query or an explicitly pushed raw terminal controller at \((w',j)\): \[ w'<w,\qquad j-t\le a'(w-w'),\qquad \ell\le a_0(w-w')+b_0(t-j). \tag{123}\] Here \(\ell\) counts only the suspended caller record, including its delimiters, return modes and pending data. It does not count the callee’s physical working space. Without a pending call its own raw records have length at most \(a_0w+b_0t\). A structural operation has a canonical starting clock zero; numerical queries used by its finite-margin tests are included in this certificate.

We will prove this bound for the actual recipes below, before using guards to restrict artificial program states. In particular, guard insertion will not be used as a substitute for proving that a designated execution fits its bound.

We use the following finite instruction types.

  1. Local edit. Read a bounded number of marked local symbols, change a bounded number of symbols and a finite mode, and advance the marked local heads by one position. A push opens a delimiter and writes a tag in a bounded number of steps. A pop is performed only after the child word has been erased down to that tag, also by local steps.

  2. Pure or restoring branch. Compute one bit, either in scratch without a graph call or by a lower restoring test or digit query, and change a bounded local symbol or mode. A bounded signed digit is handled as a fixed number of bits. A longer selection result is written one bit at a time, with its index stored in the local record.

  3. Paired cursor step. With its motion tuple and short token already selected, apply one lower partial paired map, or one elementary base or label map. Its output token and finite mode are the only simultaneous local changes. The linear-size motion tuple is regenerated from records; it is not written as an opaque payload.

  4. Explicit terminal call. Create a lower controller with a regenerable call recipe. Ordinary instructions then execute that controller. On its return, local erasure and the prescribed bounded label completion are performed by the preceding instruction types.

An endpoint action in an excursion may be a pure computation reading or toggling the catalyst. Such actions are not individual edges of a terminal program graph: that graph uses the complete restoring, catalyst-independent predicate which contains the excursion. Its other pure branches depend only on initialized scratch, the input, cursor and permitted tuple data. A value test is never combined with a cursor-changing step. Conditional validity can be a separate restoring branch; a partial paired step may simply have no edge on an invalid domain.

Every record position erased by a local edit is marked before the edit, or is a known end of a delimited word. In particular, an instruction cannot erase an arbitrarily chosen position and then forget its index. Counters, carries and loop heads are manipulated by these local steps. Pure computations and restoring tests can be recomputed to select one output bit; their scratch states and physical input heads are not included in the states whose predecessors are enumerated.

Predecessor probes can reach artificial local states whose stored flags were not produced by a designated execution. The lower query interfaces must therefore be safe independently of those flags. We use the following entry conventions before constructing the graph.

Total calls on invalid ports.

An argument is range-valid when its tuple, operation type, cursor type, port code and clock lie in the permitted finite ranges. A range-valid port can still name a hole. Every exported existence predicate is a total restoring test of its earlier defining masks. Every port-dependent quantity query freshly tests presence at structural clock zero before entering its value calculation. If presence fails, the query returns the zero digit of the zero stream. A forward motion or excursion instead returns failure, with cursor and catalyst unchanged and without performing an endpoint action. A paired inverse uses its target-side incoming-count and inverse-rank validation. Excursion validity is checked before building its chain of terminal-path records, not anew at the remote endpoint or during rollback. An endpoint comparison returns false if either freshly tested descriptor is invalid. Wrong types or out-of-range codes are rejected by a pure check before any lower call.

Baseline queries and rank-padding inputs retain their separately specified values: an absent port is not a virtual default coordinate. In particular, incoming score ranks are padded by zero, whereas a compressed-array rank uses copies of its baseline. Every opaque query tag invokes the canonical entry of one of these total procedures. No permitted call recipe targets an internal continuation after its presence test. A historical validity flag in a caller is never accepted as a certificate of presence. Lemma 67 proves that the fresh tests use earlier operations and fit the allowance bounds.

Lemma 61 (Guarded local graph). Suppose a fixed terminal recipe uses the preceding instruction types, has regenerable records, and obeys (123) on its designated executions. Suppose all its opaque transition queries are already implemented, are total on admissible arguments, and have the stated safe invalid-port behavior. Their returned bits, success flags, cursor outputs and port tokens are independent of incoming scratch and catalytic contents, and they restore the catalyst. Then there is a finite graph with outdegree at most one, uniformly associated with its root tuple, with these properties.

  1. Its designated executions are precisely the prescribed terminal executions. Success states have no successor, and initiation and success tags are distinct.

  2. Predecessor cases, ordered incidences, and paired edge movement are computable with a bounded number of case labels and one cursor. An opaque predecessor query is made only after its local bound has been checked.

  3. If a transition query at \((w',j)\) is probed while the terminal root has parameters \((w,t)\), the complete live local portion of this probe, including a fixed number of edited or saved controller copies, has length \[ A(w-w')+D(t-j) \tag{124}\] for fixed enlarged coefficients \(A,D\). Shared facilities and the outside stack are not copied.

The graph and all its guards are independent of an originating row, an outer stack’s data, and catalytic contents.

Independence from the outer stack is understood after fixing the logical root tuple: two invocations that regenerate the same tuple define the same graph. Their physical replay histories may differ. The hypothesis on designated executions is proved for the concrete recipes in Lemma 67; the guards below extend query safety to artificial states reached while enumerating predecessors.

Proof. Use as potential states the cursor and delimited words having the specified finite tags. Check absolute parameter ranges, descriptor ranges, stage and suffix types, instruction modes, and the allowed parent-to-child recipes. Check the intrinsic word bounds of Definition 60. For each explicit chain also check the local bounds along that chain. Before any opaque query, check the bound from every ancestor of its local explicit chain to that query. All absolute parameters, clocks and lengths needed for these checks are regenerated or counted in scratch. No graph query is needed for a guard. At a state with an unavailable instruction omit that edge.

The designated paths pass the guards by hypothesis: their boundary bounds add along an explicit chain and telescope. Enlarging the coefficients once covers fixed instruction tags and bounded intermediate edits. The resulting graph is finite since its local words have \(O(w+t)\) bits and its cursor has \(O(B)\) bits. It is deterministic since every branch bit and partial paired map is deterministic. Guards may create additional dead ends, but they cannot create a new path from a canonical initiation to a success state.

For a local edit, predecessor candidates are obtained by guessing the bounded overwritten symbols, preceding finite mode, and bounded head moves, then testing the instruction forward. Marked positions and incremental erasure make the number of cases bounded. For a branch, perform the same fresh restoring query at the unchanged cursor and test the guessed preceding symbols. Incorrect historical branch bits do not resume the interior of a lower procedure: the query is always started again on its regenerated arguments. The guards are checked before the query, including for a candidate predecessor.

For a cursor step, first construct the possible preceding local record by a bounded inverse edit, guessing the bounded preceding token if needed. Before any motion, replay that candidate’s active records from the fixed root tuple. Check the instruction and stratum, permitted child kind, budget, scale and suffix ranges, salt permission, numerical clock and every ancestor-to-query raw length bound. These allocation and recipe checks do not use the predecessor cursor. The lower motion tuple is then fixed by the candidate record and the present output token. Apply the partial inverse at the present cursor and output token. If it succeeds, the returned token and preceding local symbols specify the candidate predecessor. Static checks and token comparisons can be made there. A guessed preceding port is compared with the actual port returned by the inverse. Any remaining cursor-specific type or domain check is pure at that cursor; a domain condition requiring another restoring query belongs to a separate branch edge. Thus it does not conceal a second high-cost query inside this paired step. To restore the probe, apply the paired forward map with that same tuple and the actual returned token. No saved copy of the cursor is used. Conversely, an outgoing paired probe checks the proposed target guards after its motion; if a guard fails, or if the probe is only a restoring incidence test, its inverse with the actual output token restores the starting cursor. Both directions have been allowed the same lower-operation cost. No unrelated value query is part of this cursor step.

Label an edge by its finite instruction case and bounded erased symbols. Its forward instruction determines a canonical reverse case code. An inverse probe accepts only that code. This removes any duplicate descriptions of the same predecessor edge without comparing two retained addresses. The finite case codes give a fixed local order and a bounded incidence token; a forward edge knows its reverse token.

For the space assertion, first ignore saved probe copies. Adding the boundary inequalities along the current explicit chain gives the bound \(a_0(w-w')+b_0(t-j)\). A current state and a candidate predecessor differ in length by a constant: this follows from the bounded edit used to construct every candidate, before it is known to be an actual predecessor. Consequently retaining or editing both costs only a fixed multiple of that bound, with a constant adjustment absorbed by the strict allowance decrease. More explicitly, write \(\Delta w=w-w'>0\) and \(\Delta t=t-j\ge-a'\Delta w\). For a fixed copy factor \(k\) and additive bounded metadata cost \(h\), choose \[ D\ge kb_0,\qquad A\ge ka_0+(D-kb_0)a'+h. \tag{125}\] Then \(k(a_0\Delta w+b_0\Delta t)+h\le A\Delta w+D\Delta t\). This also handles a negative clock decrement. These coefficients are uniform over all finite instruction cases.

The constants \(a_0,b_0\) describe the raw skeleton uniformly over the whole recipe family. They have not been replaced by the physical probe constants \(A,D\) in a lower skeleton’s guards. Fix the maximum finite copy factor \(k\) and metadata cost \(h\) and choose \(A,D\) once. A lower compiled query contributes a new physical stack segment when it runs; its contour state and saved copies are not words in the calling graph. Thus this estimate copies raw records, not another \(A,D\)-sized physical segment.

Only the local explicit chain is saved or edited. A nested query has its own root boundary and appends its local portion below the existing stack; its program graph excludes the caller’s saved copies. A tag selects the candidate chain for parameter replay, ignoring an inactive saved copy. Therefore a later neighbor probe does not copy the older outside portions again. Finally, all guards and forward case codes use just the fixed tuple and its permitted records. Imposing the same salt-visibility convention as the underlying recipe preserves this independence even at synthetic states. ◻

Contours and incoming operations

The rotation–edge-flip traversal below is the forest permutation of Cook and McKenzie (Cook and McKenzie 1987, Proposition 1, p. 387), with a sentinel to include an isolated root. Traversal of configuration trees also underlies space-preserving reversible simulation (Lange et al. 2000). We prove the contour and incoming operations needed here, including their local record bounds. In particular, the use of a contour will not replace the separate allowance and guard argument.

Lemma 62 (Inverse contour). Let a finite deterministic graph of the form in Lemma 61 have a specified canonical success halt. Its undirected component is a tree directed toward that halt. There are mutually inverse next-mark and previous-mark operations on its contour, where the marks are the root sentinel and one distinguished occurrence of each canonical initiation in the component. These operations require the current program state and bounded incidence control, but no saved root or contour-length counter.

Proof. Following the unique successor from any node in the undirected component reaches its sink. Indeed a directed cycle in a graph of outdegree at most one cannot be joined by a path to a sink: at the first connection the unique outgoing edge is already used toward the cycle. Equivalently, on an undirected cycle all its vertices must use their outgoing edge on that cycle, so no sink lies in its component. The same argument excludes two sinks in one component. Thus the component is a finite tree with the given unique sink.

Take the incidences of its edges, with an additional sentinel incidence at the root. Edge flip exchanges the two incidences of an edge and fixes the sentinel. At each vertex cyclically rotate its incidences in the fixed local order. Rotation followed by flip, or its inverse, gives one contour cycle on these incidences. This follows by induction on the number of edges: inserting a leaf inserts its two incidences into the old cycle. The isolated-root case consists only of the sentinel.

At a canonical initiation mark its unique successor incidence; mark the sentinel at the root. Each initiation is therefore counted once. Iterating the contour successor to the next mark always terminates, since the finite contour contains the root mark. Iterating its inverse to the preceding mark is the inverse operation. A success tag and the sentinel token identify the root in this component; its address never has to be compared with an earlier address. Constant-degree adjacency and edge flip are provided by Lemma 61. ◻

Lemma 63 (Capped incoming ports). Let a fixed row class have a constant port alphabet, distinct entries within each row, and a deterministic terminal program which maps its valid row ports to their endpoints. Assume the program is of the form in Lemma 61. For every fixed \(K_c\) there are restoring tests and mutually inverse paired motions for precisely those class entries whose endpoint has at most \(K_c\) incoming class entries. The inverse ports belong to \(\{1,\ldots,K_c\}\). No additional endpoint address or unbalanced path is held across these operations.

Proof. Give invalid row ports a distinct failure execution. The canonical initiations in a success-root component then correspond exactly to valid class entries at that endpoint. Class validity is tested before using a forward initiation as a contour starting point.

Start at a valid initiation mark. Go forward toward the sentinel, counting other initiation marks, stopping either at the sentinel or after \(K_c\) other marks. Store the number of mark jumps, at most \(K_c+1\), and restore the start by that many inverse jumps. Do the same in the backward direction and restore again. Reaching \(K_c\) other marks certifies excess. Otherwise, let \(a\) and \(b\) be the numbers of other initiations before the sentinel in the forward and backward directions. The total is \(1+a+b\); accept exactly when this is at most \(K_c\). The initiation’s rank after the sentinel in forward order is \(r=b+1\). A forward paired motion goes to the sentinel in \(a+1\) mark jumps and returns the inverse port \(r\).

From a canonical endpoint halt, count forward marks until returning to the sentinel or encountering \(K_c+1\) initiations. In the latter case restore the sentinel by the bounded number of inverse jumps and reject. In the former case the total \(n\le K_c\) is known. For an inverse port \(r\in\{1,\ldots,n\}\), go forward by \(r\) mark jumps to its initiation and return the row port stored in that initiation state. Invalid ranks are rejected at the sentinel. A halt with no preimages has \(n=0\) and is harmless. Wrong vertex types can be rejected before creating a halt state.

Every unsuccessful inverse request therefore leaves its endpoint unchanged: an excessive incoming count is undone before failure, and an invalid rank is detected at the recovered sentinel. For a successful request the returned row port is the one in the reached initiation, not a guessed port. This is the token used by a later forward restoration in Lemma 61.

These are inverse maps by the common contour order. Existence tests make the corresponding visits and reverse them, so they restore the cursor. All counts, ranks and incidence tokens have a fixed alphabet. The terminal program graph has the common table tuple as its root; neither direction fixes the starting row or port as an external parameter. Consequently the two directions use exactly the same graph. Restoring lower queries restore the catalyst, and the contour itself does not otherwise change it. ◻

The incoming scores need not have bounded support in a column, so a bounded incoming-port list does not suffice to compute their rank. We instead give the terminal program a success state containing both the endpoint and a rounded score. A contour at that halt counts the entries with that rounded score; successive contours visit the scores in decreasing order. The endpoint and score belong to the current program state. During a tour they may be replaced by an earlier state of the score computation, and the contour recovers them when it returns to its unique halt. Thus neither is retained separately while an earlier state requests a score digit. The next proof combines this observation with an assembly order that stores only a short suffix before a late digit query.

Lemma 64 (Incoming score rank). Let a constant-size forward row class cover every positive entry of a bounded nonnegative score table, with distinct entries in each row. Suppose pre-incoming score digits are restoring queries, and the class has a terminal program with the preceding interfaces. A fixed incoming \(k\)th-largest value, with zero padding, has exact redundant digits computed by a restoring procedure. Its local record above a place-\(j\) score query at output approximation clock \(t\) is \(O(1+(t-j)_+)\), with a fixed overshoot \(j\le t+c\), in addition to the structural allowance costs. A complete \(O(t+1)\)-bit score word is retained only over structural calls.

Proof. Normalize the score range to \([0,1]\), using a fixed public factor if necessary. At approximation place \(t\), choose a fixed extra margin \(c\). For the digit truncation \(v^{[t+c]}\) define the integer \[ q_t(v)=\operatorname{clip}_{[0,2^t]} \left(\left\lfloor 2^t v^{[t+c]}\right\rfloor\right). \tag{126}\] The digit tail has absolute value at most a fixed constant times \(2^{-t-c}\). Choose \(c\) so that this is less than \(2^{-t}\). Clipping is nonexpansive and fixes \(2^t v\), so \(|q_t(v)-2^t v|<2\). The same margin gives \(q_t(0)=0\): a truncation of the zero value has scaled magnitude less than one, so flooring and nonnegative clipping give zero. Since the forward class contains every positive score, enumerating its entries and padding by zero therefore gives the rank of the whole quantized score column.

Use a forward terminal program which first tests the entire class and membership domain at its canonical initiation, sending invalid inputs to failure. Its success configuration contains the endpoint in the cursor and \(q_t(v)\) in its local state. Before making its final structural motion, assemble the score digits from least to most significant. In units \(2^{-t-c}\), before the place-\(j\) digit is queried the accumulated suffix and its carry have \(O(1+t+c-j)\) bits. A self-delimiting offset from \(t+c\) fits the same bound. After assembly, division by \(2^c\), flooring, clipping and loop setup are pure local computations. The complete word is then held only through structural motion and local clearing to success. Thus all its graph-query boundaries satisfy the claimed record bound. The success word is part of the graph state, not a parameter of the graph. At this fixed \(t\) the graph includes every allowed quantized value.

At the requested endpoint, initialize the success state at \(q=2^t\). Tour that root component and count canonical initiations, saturating at the desired fixed rank. Continue to the sentinel even if the saturated count has already been reached. At the sentinel the same endpoint and quantized value have been recovered; no other success halt occurs in that component. If the cumulative count has not yet reached the requested rank, decrement the value field by pure local steps and tour the next success component. The value field itself is the loop variable. No separate copy of it survives the tour. The saturated count and control bits have constant size. At value zero apply the stipulated zero padding.

If the desired rank has been reached, use the value field only after the tour has returned to the sentinel. The clock \(t\) is a root parameter throughout; the loop value \(q\) is not. Consequently an inverse step which reaches an earlier partial-suffix state replaces the current score state by that suffix state, and need not retain the final \(q\) or the endpoint address alongside it. Unique-sink recovery supplies both again when the component tour ends.

This computes the integer order statistic \(N_t\) of the values \(q_t(v)\). A uniform perturbation of a finite list changes any order statistic by at most the same perturbation, including after zero padding. Thus \(N_t\) has error less than \(2\) as an approximation to \(2^t\) times the desired rank. Keep only the fixed number of residue bits needed to form its redundant digit, clear the other local data without a graph query, and return at the unchanged endpoint. Computing \(N_{t-1}\) uses its own quantization convention; the difference-digit construction needs only the bounded errors, not nested quantizations.

During a tour the live program state may be a partial suffix state preceding a high-place query. It then contains that short suffix, rather than a saved copy of the final score. Lemma 61 therefore also gives the required bound for inverse neighbor queries. All root components are finite, including isolated success states, and the decreasing finite value loop terminates. Fresh lower procedures and the contour preserve the catalyst. ◻

Certificates for the estimator recipes

We next verify the hypotheses of the graph construction for each of the actual recipes. These certificates concern syntax, storage and dependencies; they do not use the statistical accuracy of an estimated entry. Thus they apply also when its numerical value is inaccurate.

Lemma 65 (Ordered terminal paths and excursions). Every unbalanced path required by the estimator has the following ordered form: restoring selection and validity tests; local label maps and, if necessary, a paired left-factor motion; further restoring tests at that midpoint; at most one final terminal delegation; and then only local label completion. An excursion to its endpoint and back, for a pure local action, has regenerable records. It preserves the same allowance and clock bounds as this ordered path, up to fixed factors. No unbalanced spine is retained over a new high-cost query at its endpoint.

Proof. For a direct summand the path is a local base path, a projected child path, or a child row port. A projection stores its bounded symbol before removing it and appends the prescribed destination symbol after the child terminal motion. For a product the left factor is a correction with paired access. Its inverse token is stored at the midpoint; right-factor validity and weight queries are restoring there. The right terminal path is then the last delegation. An old prefix port delegates to that preceding prefix; after its return it requires no new graph query. These are exactly the path cases of Lemma 28.

There are two compiled continuations of this same terminal syntax. A one-way continuation clears its path records after performing its prescribed local completions. A visiting continuation retains the immediate choices and inverse tokens on its explicit terminal spine. After the local endpoint action it undoes the bounded completions, undoes the last delegated spine, and then uses the paired inverse token of an earlier motion. Each deeper record is cleared before invoking an earlier inverse map. For a projection, the saved symbol and prescribed append symbol give its local inverse. The identical lower paired tuple and actual returned token give the inverse of a correction move.

This is an induction on terminal delegations, not an attempt to invert an arbitrary unbalanced macro in one step. A frame stores only its own immediate descriptor, bounded choices and phase. Its returned child spine is used for local completion and the endpoint action, which make no graph calls. When a further graph query is made during rollback, the deeper spine has already been discarded. Thus the suspended records are those at the corresponding forward boundary, enlarged by bounded phase and inverse-token data. The choices can be re-created or kept on their own frame; no concatenated description of all old prefixes is passed to a new tight query.

In a restoring use, intrinsic constituent validity is checked before the visit. A product can test its right port by crossing the paired left edge and returning; this test does not use the right terminal path. Outer slice and wrap tests then use a valid intrinsic excursion. Hence none of these validity checks recursively invokes the very bucket or shortlist membership it helps to define. The endpoint action’s owner is marked in the outer frame, so its tuple can be regenerated by a scratch scan even while a descendant path is retained. The action’s result does not alter the selected paths or their undo. ◻

Lemma 66 (Implemented endpoint comparison). At slack \(d\), comparisons of two valid candidate or old-list endpoints are restoring, catalyst-independent tests with \(O(d)\) suspended structural control, in addition to the costs of their lower endpoint visits. They use no saved endpoint address. Their choice between an exact comparison and the designated fingerprint comparison is independent of a surrounding numerical clock.

Proof. The two descriptors occupy \(O(d)\) bits. By Lemma 65, each endpoint can be visited and restored for a local bit or hash action. If \(d>2h_*\), compare the \(O(B)\) identifier bits in order, retaining only the bit-position counter and bounded flags. Its counter uses \(O(\log B)=O(d)\) bits, since \(h_*=\lfloor\epsilon\log_2 B\rfloor\) with a fixed positive \(\epsilon\) and sufficient minimum padding.

Otherwise use the four visits in Lemma 47: flip the first key’s bit, xor the second key’s bit into a result, and repeat. The final result is key equality for every incoming catalyst and that vector is restored. Key evaluation and the indicated vector-bit action are pure endpoint operations in scratch. No key index is saved across another visit. Nested structural procedures already restore the catalyst and give results independent of its contents, by their lower operation order. The outer descriptors and all path choices ignore the accumulated result bit. The two repeated visits therefore use identical paths. This proves the same independence and restoration for the completed comparison, supplying the induction hypothesis at the next layer. The comparison mode is determined by \(M,d\) and the public block boundaries, not by the requested numerical place. ◻

For the block comparison in Section 9, the clean calculation replaces designated key comparisons by exact endpoint equality. Endpoint comparison is an opaque restoring bit query in the terminal program graph: its internal states are not graph vertices. The exact comparison can therefore use a longer bit-position counter without enlarging the guarded raw graph states. It still returns its bit and restores the cursor and catalyst. Only the actual algorithm must satisfy the tight physical-space bound; the clean calculation supplies comparison values and paths for the error analysis.

Lemma 67 (Concrete recipe certificates). All actual estimator operations admit regenerable records and a fixed finite ordering of operation strata. At a prefix of slack \(d\), each ordinary layer has \(O(d)\) structural control, and a numerical path to an operand at clock \(j\) has suspended record \[ c_{\mathrm{str}}d+c_{\mathrm{num}}(t-j),\qquad j-t\le c_{\mathrm{dig}}d, \tag{127}\] with fixed coefficients. Beyond the fixed strata of this prefix, a call is to the preceding prefix of slack \(2d\), or to an estimator of budget at most \(M-d\). The final row and incoming wrappers add a fixed number of strata at \(d=1\). Every designated terminal execution obeys the corresponding boundary bounds before guards are imposed.

Proof. Table 2 summarizes the records to be charged. We verify the recipes and their dependencies in the order in which they are constructed. All constants depend only on the fixed construction constants and the base machine.

Suspended records at a lower call. The bounds exclude shared facilities and the callee’s physical state. Terminal-path frames are charged separately along their explicit chain; graph probes copy only that local raw chain.
Operation Data retained across a lower call Local space bound
Structural tests and list selection Immediate descriptors, loop indices, counts and flags \(O(d)\)
Numerical recipes Gate paths, pending bits and permitted early operand words \(c_{\rm str}d+c_{\rm num}(t-j)\)
Ordered terminal paths One frame’s choices, phase and bounded inverse tokens \(O(d)\) per immediate frame
Incoming score ranks A partial score suffix; the full word only over structural calls \(O(1+(t-j)_+)\)
plus structural costs
Graph transition probes A fixed number of local raw-state copies and bounded incidence data \(A(w-w')+D(t-j)\)

Immediate paths and their values.

By Lemma 28, a group has at most \(2^{O(d)}\) immediate descriptors: term and channel part, scale and budget offsets, a variant word and wrap choices, bounded copy symbols, and a bounded number of factor ports. All factor budgets are at most \(M-d\). A descriptor is stored literally in \(O(d)\) bits; it never stores a discovered endpoint. Index and range tests are pure computations. Constituent existence tests call only these lower factors. Products cross a paired left edge for right-factor tests or digits, retaining a bounded inverse port; they return after each such query. Factor digits can therefore be requested separately by an arithmetic evaluation path. No assembled left weight remains over a late right-factor digit. Source vector factors are queried at the current source, with the prescribed local projection if needed. Their numerical coefficients and the factor products use a fixed arithmetic recipe.

For terminal motion the right unbalanced path, or the direct child path, is the last delegation. This is the ordered path of Lemma 65. To avoid a validity cycle distinguish intrinsic constituent existence, intrinsic motion or excursion, and outer validity. Outer validity first checks the intrinsic constituents, then uses an intrinsic excursion for its own hit and wrap conditions. A subsequent outer-valid motion may repeat that validity check and then use the intrinsic motion. Each of these is a separate stratum; none tests the updated bucket or shortlist. Invalid descriptors return failure or a zero value without relying on a motion’s promised inverse.

All rate factors multiplying these terms have logarithmic range \(O(d)\): a variable bin has \(|\log_2(b/a)|=O(d)\), while \(f/\pi_f\) is bounded. Resetting a source rate to the public stage rate does not multiply the numerical term by \(f/u\); that ratio was a norm cost. It therefore causes neither a large numerical shift nor a new long choice field. The fixed number of bounded-factor products, coefficients and \(O(d)\) scale shifts obey Theorem 56.

Candidate domain and bucket sums.

The descriptor pool is the concatenation of the preceding list positions and this group’s unmerged paths, with invalid slots masked. A new descriptor is a representative exactly when it is the first valid new descriptor with its key and no old valid position has that key. Old valid positions remain separate. To decide these flags, enumerate the \(2^{O(d)}\) descriptors, keeping only the current indices, bounded flags and \(O(d)\)-bit counts. Each test uses lower validity and endpoint comparison operations. The comparisons use only those unmerged paths or the preceding prefix’s already-defined positions. In particular they never inspect the new value of the bucket being defined.

An intermediate entry is the old entry at an old representative, or the previous baseline at a new representative, plus the sum of all valid matching increments in each channel. That sum has a fixed \(O(d)\)-bit descriptor domain and boundedly many \(2^{O(d)}\) scale factors. It is evaluated by the balanced sum construction of Theorem 56, with the preceding domain tests as structural leaves. No high-precision sequential accumulator is held while requesting term digits. The bucket rule preserves distinct actual representatives, including when two old representatives have the same key. Thus it gives exactly the intermediate array used in Lemma 43.

Ranks, compression, and list selection.

The intermediate domain and its values are now earlier inputs. Add enough copies of the default for the size and first-signal ranks. A finite domain of \(2^{O(d)}\) indices suffices: the preceding list has size \(2^{O(2d)}\) and the new group has \(2^{O(d)}\) descriptors. The default is included as often as needed to represent the unlimited padding in the prescribed ranks. The rank, max, min, absolute-value and clipping operations of Lemma 32 are the finite arithmetic primitives. The ranges are \(2^{O(d)}\) by Lemma 37; their known normalizing shifts have magnitude \(O(d)\).

Here \(u_0\ge\rho_d=2^{-c_1d}\). A canonical finite approximation can test a threshold strictly between \(u_0\) and \(2u_0\): with error a sufficiently small fixed fraction of \(\rho_d\), accept every size greater than \(2u_0\) and accept only sizes greater than \(u_0\). The required numerical clock, after normalizing the range, is \(O(d)\). Fix this clock and tie convention once for that structural test. Consequently every nondefault compressed position is included, and fewer than \(n_d\) positions are selected, as in Lemma 32. Numerical equalities at the boundary need not be decided. No default can expose a positive row score: its rank dummies close the gate first, by the margins of Lemma 16.

Remap a requested list ordinal to its selected intermediate descriptor by enumerating the pool and counting the preceding accepted flags. The loop stores \(O(d)\)-bit indices and a count. Its only calls are the already-defined intermediate flags and finite-margin selection tests. Compressed entry digits then query that descriptor’s intermediate value and the common compression thresholds; baseline digits use those thresholds directly. A terminal list motion selects its descriptor before delegating to an old prefix path or a new valid path. A descriptor from many prefixes ago is therefore represented by a chain of delegations, not retained as one long word above a tight child.

Presence before value or motion.

Ordinal lookup returns a presence bit and a selected descriptor only when its finite enumeration reaches the requested ordinal. It starts with a false presence bit and returns false if enumeration ends first; the false branch supplies no descriptor. Every compressed-entry digit query and list motion branches on this fresh result before using a descriptor. The preceding-prefix lookup has the same convention. A hole denotes no real representative. Rank operands retain their prescribed structural selection of an entry, baseline, or padding value; they do not infer that value from the zero stream exported at an absent real port.

The final precursor \(A\) and auxiliary class \(A_I\) use the bounded final list-position alphabet, allowing holes. Their presence tests use list presence, role-membership masks, and the fixed-precision raw-entry and raw-rank tests of Lemma 16. Source membership is checked locally; endpoint membership uses the earlier list or intrinsic excursion with a pure membership action. It does not use the exposed-class motion whose validity is being defined. These tests precede the pre-incoming score formula and use neither its incoming gate nor its cap. Every pre-score digit invocation repeats the corresponding class test and returns zero on failure. Other uncapped row-value queries use their support-covering class in the same way. Each such class contains the entire nonzero support, so the extension agrees with the mathematical table. Existence predicates evaluate their earlier masks directly; they do not ask for their own existence. Analytical witness functions without an exported port interface retain their stated full-array values and are not restricted by this convention.

Immediate-path values and visits first use their earlier constituent and, where appropriate, outer-validity tests. Endpoint comparisons likewise test their two earlier descriptors before making visits. A capped correction-value query stops with zero if its forward paired move fails; only a successful move permits a gate or column-value query at the endpoint. The paired inverse keeps its target-side validator from Lemma 63. Incoming ranks and vector queries return zero outside their prescribed vertex type or membership domain. These checks are repeated at canonical entry, including when the caller is a synthetic post-initiation state with a fabricated validity flag. An excursion checks before constructing its spine; its endpoint action and rollback retain the order of Lemma 65.

Place these entry wrappers a fixed number of strata above their validity predicates and value or motion cores. A wrapper at clock \(t\) calls its structural test at clock zero and retains \(O(d)\) descriptor and control bits; its boundary is therefore bounded by \(O(d)+bt\). When it subsequently calls a numerical core at clock \(j\), the test has finished and only bounded status and descriptor data remain, so the bound \(ad+b(t-j)\) persists after enlarging \(a\). The predicates use the fixed structural precisions already proved. They do not call the score, gate, cap, new bucket or new shortlist whose value is being checked. The finite stratum list is enlarged before choosing the allowance constants. Recipe guards permit opaque calls only to total entry wrappers, so a synthetic state cannot select a continuation that skips validation. No semantic predicate is evaluated by a guard itself.

These wrappers preserve the represented relations, values and outgoing endpoints. Their insertion may change the contour ordering of incoming ports; both directions use the same updated graph, so pairing remains consistent. The tests use the same effective tuples and fixed structural clocks, and restore the catalyst on both outcomes. They require no accuracy promise and add no retained endpoint address.

Finite strata within a group.

For clarity, a permissible ordering, from an exported operation down to its dependencies, is:

  1. compressed entries, baseline and list-port operations;

  2. ordinal remapping, compression and exception selection;

  3. size and first-signal ranks of intermediate entries;

  4. intermediate domain flags and entry or bucket-sum values;

  5. first-representative and key-match predicates;

  6. endpoint comparisons of outer-valid paths;

  7. outer-valid visits, then outer path validity;

  8. intrinsic visits, constituent tests, factor digits and ordered motions, subdivided according to the preceding distinction.

The wording of an interface name does not force a call at the same stratum. For example, the “visit” used to verify outer validity is the intrinsic visit, while the visit used by key comparison is an outer-valid visit. Arrange these finite subtypes in their indicated dependency order: outer-valid visits use outer validity; outer validity uses intrinsic visits; key comparisons use outer-valid visits; and bucket predicates use those comparisons. Arithmetic on intermediate values uses only their domain flags and lower matching predicates. This refines the displayed list to a finite acyclic list. Repeated enumeration over descriptors stays in a loop of its layer, rather than creating one recursive frame per descriptor. Between these layers Theorem 56 gives (127); structural loops add \(O(d)\) bits. All remaining path calls go to the preceding prefix or to the lower-budget factors already described.

Final row classes and pre-incoming values.

At the final slack \(d=1\), the list size, port alphabet and all threshold gaps are constants. The classes and witnesses in Lemma 17 are defined by canonical finite-margin entry and rank tests on the final raw arrays. They can be enumerated without testing positivity of a computed score. Their tests and pre-incoming score digits query only the final buffer and a fixed arithmetic recipe. The same raw-array tuple is used for every row. A normalized dense entry uses its second raw argument, divided by the fixed scaling constant and clipped before the reciprocal. It does not start a new preparation experiment at its unbalanced endpoint.

The broader auxiliary class used to score an incoming gate depends only on the corresponding pre-incoming row calculation. Its validity and score queries do not depend on that incoming gate or cap. The precursor \(A\) for a correction has the same property. Its canonical terminal program begins with a restoring check of the full precursor class and membership domain, and sends every invalid initiation to failure. Thus Lemma 63 counts exactly its valid entries, not merely entries promised valid by an external caller. The auxiliary scoring program has the same initial check. Its score digits are pre-incoming, so Lemma 64 invokes only lower forward operations. By Lemma 17, discarding a precursor with more than \(K_c\) incoming entries preserves its value.

Final incoming and vector wrappers.

Place the capped-pair and inverse-score wrappers above their respective terminal programs, and those programs above their pre-class, selection and pre-score queries. They are separate terminal programs: the gate score program does not call this task’s cap pairing, and the cap program does not call its gate rank. A final correction-value query first completes the paired move. At the endpoint its motion-return data consist only of the bounded inverse port while it queries a gate-rank digit or a lower-budget column-vector digit, and then returns. This can be done separately for every arithmetic operand bit. Only numerical path bits and immediate descriptors coexist with a late query. The factor \(r_i\) at a projected source is handled in the same way at that source.

A detector trial is an inverse-score wrapper of its lower-budget input row task. Reward terms query source factors locally and cross paired corrections before midpoint vector queries. Vector prefix sums use the simpler balanced arithmetic form of the same dyadic schedule. These are all the vector and incoming uses in the additive formulas. There are only a fixed number of final wrapper types. Their short counts and return ports cost a fixed allowance gap. In an inverse score program the saved numerical suffix satisfies the clock bound in Lemma 64; its full word is held only over structural motion. The cap wrapper’s counts have fixed size.

Uniform generation and record bounds.

Each enumerated object above has a literal \(O(d)\)-bit index or is a Boolean path in one of the stated arithmetic recipes. Uniform index-generation uses finite integer loops on these indices, including rank comparison pairs and balanced-tree nodes. Constants, field/rate coefficients and walk probabilities have \(O(B)\)-bit rational descriptions. Their requested digits use integer arithmetic in scratch and make no recursive graph call. In particular, a walk word has length \(O(d)\) and a fixed rational generator probability, so its probability denominator has \(O(d)\) bits. Rate and field expressions involve fixed powers of \(O(B)\)-bit integers. The arithmetic theorem’s uniformity and public-coefficient hypotheses are thus satisfied.

The finite arithmetic list covers the formulas used here explicitly. A fixed polynomial switch is evaluated on an affine-rescaled clipped argument; positive parts and absolute values use maxima, and a variable clip is a minimum followed by a maximum. The normalizing reciprocals have their arguments clipped to \([1/2,1]\). The eighth roots used in the error estimates express regularity of the row functions and are not additional numerical operations requested by the controller. Ranks inside a prefix use the finite descriptor construction above; the unbounded incoming classes use Lemma 64. This accounts for the arithmetic primitives without adding a real-number comparison or an incoming-rank oracle.

Structural selection words and loop counters have \(O(d)\) bits. The only longer local numerical words are the early operand records of Theorem 56 and the just-described score suffix; they obey their explicit clock-decrement bounds. Late operand calls hold only numerical paths, pending bounded results, immediate descriptors and short inverse tokens. A fixed number of arithmetic and structural layers enlarges the constants in (127); its clock differences add. Ordered terminal delegations add their own immediate records, with the same estimate at each boundary, including on rollback as proved in Lemma 65. This verifies the record certificate on the designated executions, independently of any graph guards. ◻

Lemma 68 (Parameter replay and salt visibility). The records in Lemma 67 are regenerable in the sense of Definition 59. With respect to their own effective environments, actual procedures at a budget in block \(b\) ignore all salt blocks later than \(b\), including in structural orders, graph universes, guards and predecessor probes. Re-querying a table entry or using a paired port in either direction uses the same effective table tuple.

Proof. Start with the fixed top parameters and scan the active records. An estimator edge specifies a task type, a stage change prescribed by the formula, a budget offset, a bin-scale offset if present, a bounded copy choice and channel-role recipe. These determine the next absolute stage, rates and suffixes. A projection appends the stored symbol before the inherited suffix. A fresh midpoint or vector role uses the specified public stage rate and its prescribed suffix; it does not carry the old source. These are the common-tuple recipes in Lemma 17 and Lemma 28.

An averaging edge also carries its generator and wrap word. Apply these fixed transformations in order to the current environment. Their descriptors are independent of the indicated endpoint; membership and wrap consistency are tested separately on visits to that vertex. Thus no discarded conditioning address is needed to regenerate the environment. For an old-prefix edge the tuple stays that of the same row calculation. At final assembly a gate or column multiplier uses this common tuple, rather than the internal variant which happened to lead to a selected representative. An action-owner tag allows an endpoint-local calculation to replay just to that ancestor tuple.

Numerical paths specify arithmetic operands, gates, digit places and the early-prefix modes of a reciprocal. Their finite index rules are those of Theorem 56; an absolute place is recovered from the root clock and the path or offset, not saved afresh per frame. Descriptor selections which used the cursor were stored explicitly as their bounded choices before movement. No later tuple computation needs to repeat such a selection at an unknown old cursor.

At every step the tuple and the temporary integers have \(O(B)\) bits in the actual parameter and clock ranges. Keep a current tuple in scratch and overwrite it as each record is processed. Repeating the scan reconstructs an ancestor or a selected candidate chain. Delimiters and marked active-copy tags specify which chain is used. This costs one scratch facility, including its scanning counters, rather than one absolute pointer, input-head position or tuple per frame.

A pure endpoint action belongs to its outer invoking predicate and may read that owner’s permitted salt. It is excluded from the lower program graph and cannot affect the lower path or rollback. Nested lower queries still restore their arbitrary incoming catalyst contents.

Finally, the only salt read by a new group comparison is its task’s own block salt. Descendants do not increase their budget. Walks alter only the already designated earlier free salt blocks and leave the present and later ones frozen. All descriptor and instruction orders are fixed index orders; all range and length checks depend only on public bounds and accessible tuple fields. Give each recipe the syntactic permission to address only those salt blocks at most its budget block in its effective environment. This is a direct range check on a requested salt index, not a condition on the salt’s value. A synthetic record cannot change that permission, choose a higher-budget recipe, or bypass the fixed parent-to-child rules. The same replay and checks are used before predecessor queries. Consequently the entire lower-block structure is independent of later salts. The clean calculation in Lemma 40 uses the same fixed descriptor orders and structural clocks, with its stipulated replacements of comparison predicates. The error proof uses the salt independence of the actual lower algorithms and of the clean recipes. ◻

Remark 69 (Effective coordinates and physical replay). The salt assertion in Lemma 68 is a statement about a task as a function of its effective environment. For fixed other task parameters, write this function schematically as \(F_b(\eta)\). Its values, access orders and guarded program graph use only the rank channels and the permitted salt coordinates of \(\eta\). An ancestor can invoke it at \(\eta=\varphi(\sigma)\), where \(\sigma\) is the master environment and \(\varphi\) is a recorded environment transformation. Computing the permitted coordinates of \(\varphi(\sigma)\) can read physical master bits whose original block indices exceed \(b\). Those reads derive the effective argument; they do not give \(F_b\) a new semantic parameter or permission to inspect later coordinates of that argument.

For a clean calculation in the current block, the current task’s own walks vary the earlier effective salt coordinates and freeze its current and later ones. Same-block clean comparisons omit the current key, and lower-block procedures obey the local permission rule just proved. This is the coordinate convention used for the slice means in Lemma 40. A callback owned by a higher task remains the separate pure endpoint action described above; it cannot enter the lower graph, change its selected path or affect rollback through catalytic contents.

The uniform space theorem

Two linear quantities organize the final accounting. For positive physical-space coefficients \(A,D\), chosen below, define \[ \Phi(w,t)=Aw+Dt,\qquad \Psi(w,t)=t+a'w, \tag{128}\] where \(a'\) is the clock-overshoot coefficient in Definition 60. The decrease of \(\Phi\) will pay for a suspended physical segment, including its local probe copies. The nonincrease of \(\Psi\) will bound every descendant numerical clock and hence the shared arithmetic scratch.

Proof of Theorem 58. Let \(S\) exceed the finite number of operation strata in Lemma 67, including the subdivision of path validation and all final wrappers. Choose fixed integers \(K_0>S+1\) and \(K_1>2(K_0+S+1)\), increasing them as needed for the local constants. Assign to an operation at budget \(M\), prefix slack \(d\) and stratum \(s_0\in\{0,\ldots,S\}\) the allowance \[ w=K_1(1+M)-(K_0+s_0)d. \tag{129}\] Final wrappers use \(d=1\), with increasing \(s_0\) down their dependencies. For a nontrivial prefix, \(1\le d\le M\); trivial tasks use fixed top strata and terminate locally.

A same-slack descent lowers \(w\) by at least \(d\). A preceding-prefix descent to \(2d\), even with a reset to an earlier stratum, lowers it by at least \((K_0-S)d>d\). A statistical descent to \(m\le M-d\) lowers it by at least \((K_1-K_0-S)d\), since the child allowance is at most \(K_1(1+m)\) minus a nonnegative stratum term. Hence every required boundary has \(w-w'\ge c d\) for a fixed \(c>0\), and all allowances remain positive and \(O(1+M)\). The estimates (127) and the ordered-path record certificates give Definition 60, with common coefficients. When a common clock coefficient is increased, compensate the allowance coefficient using the bounded clock overshoot, just as in (125). Add a fixed allowance gap for a traversal wrapper’s bounded incidence, count and status data.

There are two distinct estimates in this induction. First, Lemma 67 certifies the raw caller words with the same \(a_0,b_0\) at every budget; a lower opaque operation contributes its call recipe, but none of its internal working states, to that certificate. Second, Lemma 61 converts a copied raw portion to its physical bound with one fixed pair \(A,D\). The induction below appends a lower operation’s physical portion; it never substitutes that physical portion into an upper raw graph state. The raw certificate therefore does not depend recursively on the already enlarged physical constants.

Construct the procedures in increasing allowance. The induction covers all recipe instances admitted by the finite range and parent-to-child checks: both designated calls and calls made when probing candidate predecessor records. It does not require such a candidate record to have occurred on a designated execution. Each opaque call still enters the lower procedure at its canonical entry, where its semantic presence tests are performed. At the bottom are local base ports, rank and salt evaluation, rational arithmetic without graph operands, and trivial zero tasks. Suppose lower operations are total, have their claimed cursor and catalyst behavior, and obey the local bounds. The finite restoring and numerical recipes of Lemma 67, using Theorem 56, then compute exactly their specified bit or digit. Their finite loops and lower calls terminate. Endpoint visits use Lemma 65, and completed endpoint comparisons use Lemma 66. The latter restores arbitrary catalyst contents and has an independent answer, preserving the induction hypothesis. The other recipes make no direct catalyst-dependent choice.

The canonical terminal programs have finite loops and only lower terminal delegations or opaque calls, so their prescribed executions terminate. Their class checks send invalid canonical initiations to failure. The certificates already proved show that all their designated boundaries fit the local bounds. By Lemma 68, their parameters and static guards are uniform and independent of a discarded cursor. We may therefore apply Lemma 61. Guards remove no designated execution. At synthetic states they ensure that every probed query has an admissible lower tuple and available space; fresh safe lower queries do not rely on the truth of an earlier stored flag.

The finite graph contours of Lemma 62 now implement the higher cap and inverse-rank wrappers in Lemmas 63 and 64. They terminate on success-root components even if other graph components contain cycles or guard-created dead ends. They count only the canonical initial executions of these programs, so their paired maps and ranks are the mathematical ones specified by the tables. By the common root tuple and salt convention the same maps are obtained in either direction and on every re-query. This completes the well-founded construction. An opaque query in this induction denotes its already constructed finite procedure; recursively substituting these procedures gives an ordinary deterministic machine.

The allowance is a runtime parameter of this fixed finite collection of instruction types, rather than an index for input-dependent program text. Operation tags select the finite recipe family; loop counters, gate paths and regenerable descriptors select its current instance. Thus the induction constructs one uniform interpreter for these instructions. It requires neither a separately compiled machine for each budget nor stored transition graphs for the contours.

It remains to add all simultaneously live data. By Lemma 61, a traversal root at \((w_i,t_i)\) with a pending query at \((w_{i+1},t_{i+1})\) contributes at most \[ A(w_i-w_{i+1})+D(t_i-t_{i+1}) =\Phi(w_i,t_i)-\Phi(w_{i+1},t_{i+1}) \tag{130}\] local bits, including its fixed number of neighbor-state copies. Ordinary restoring and numerical calls have the same bound after one enlargement of \(A,D\). Inverse traversal copies only its local explicit chain, not older stack segments or the internal configurations of a transition query. Choose \(A,D\) also to dominate the bottom operation’s intrinsic linear bound. If the active chain has boundaries \(0,\ldots,r\), with boundary zero at the top, put \(\Phi_i=\Phi(w_i,t_i)\). The total local space is at most \[\sum_{i=0}^{r-1}(\Phi_i-\Phi_{i+1})+\Phi_r+O(1) =\Phi_0+O(1) =Aw_{\mathrm{top}}+Dt_{\mathrm{top}}+O(1).\] Figure 1 shows which records enter each summand.

Physical stack space, with \(\ell_i\) the length of segment \(i\) and \(\Phi_i=\Phi(w_i,t_i)\). Opaque child work is appended outside raw copies. The proved record certificates and local graph bound, with enlarged \(A,D\), pay for copies and metadata. Shared facilities are counted separately. The displayed copies occur at traversal calls; older physical segments are not copied into the child. Ordinary restoring and numerical segments obey the same difference bound.

The clock inequality in (123) gives \[ \Psi(w_{i+1},t_{i+1})\le\Psi(w_i,t_i). \tag{131}\] Since \(t_i\le\Psi(w_i,t_i)\), all clocks are \(O(t_{\mathrm{top}}+w_{\mathrm{top}})\). This also bounds the public scratch calculations in the arithmetic theorem. An uncharged full coefficient or cursor copy cannot persist there: any operand-dependent early word crossing a query was included in its record certificate, while pure public computations have no recursive query. The master environment, cursor, catalyst, scratch and their global scan counters are single facilities. Their total space is \(O(B+t_{\mathrm{top}}+w_{\mathrm{top}})\).

For exact-stream semantics the induction is valid for each finite requested clock, allowing scratch to scale with that clock. Every structural domain uses its fixed finite-margin clock, so its definition does not change in this argument. In particular the integer \(N_t\) of a score rank or arithmetic operation is always evaluated by its own place convention, even when used in a request for the next digit. The actual application has \(M_{\mathrm{top}},t_{\mathrm{top}}=O(B)\), and hence all these facilities and active records use \(O(B)\) bits. ◻

Derandomization and quantitative consequences

We now complete the proof of \(\mathsf L=\mathsf{RL}=\mathsf{BPL}\). The hierarchy approximates the acceptance probability by a deterministic reward, and the actual estimator approximates that reward in an integrated error norm. We pass from this norm to the specified start, enumerate the finite conditional environment law, and take a median. The controller computes each estimator value in logarithmic space and is total on every environment. This gives the following approximation result without any decision-gap promise.

All transducers below have read-only input and a one-way write-only output tape. Output contents are never available as work memory. For a fixed randomized machine \(\mathcal M\), write \(p_{\mathcal M}(x)\) for its acceptance probability on input \(x\).

Theorem 70 (Fixed inverse-polynomial accuracy). Fix a randomized machine \(\mathcal M\) with polynomial worst-case running time and \(O(\log(n+2))\) work space, and fix an integer \(d\ge1\). There is a uniform deterministic transducer that, on every input \(x\) of length \(n\), outputs a dyadic rational \(z_x\in[0,1]\) with \[|z_x-p_{\mathcal M}(x)|\le(n+2)^{-d}.\] Its work space and output length are \(O(\log(n+2))\), and its running time is polynomial in \(n+2\). The constants may depend on \(\mathcal M\) and \(d\), but not on the input.

The fixed exponent in this theorem is enough to decide bounded-error computations. After proving it, we derive the class equality and its promise version. We then allow the requested accuracy to be part of the input and construct accepting computations under an inverse-polynomial probability promise.

Choosing constants and finite ranges

The construction has the following order of choices.

  1. Fix \(H,h_0,s=8\) and the fixed number of independent channels. The finite-action gap is uniform in every later modulus and bit length.

  2. Choose the correction thresholds and cutoff, then the iteration count \(k_*\) from the factorial bound and its resulting uniform magnitude constants. These estimates hold for every column multiplier in \([0,1]\).

  3. Choose \(D_0\) for the density reduction, the routing thresholds, and the finite ladders of entry and gate thresholds. The row-gate rank \(K\) can then be chosen from the row-mass bounds alone.

  4. Choose the detector rank \(J\) large enough for the incidence bound on extra copies. Choose \(R_*\) large enough for overload detection, and finally the incoming rank \(K_c\) large enough for correction normalization. The earlier row-gate choice does not depend on \(K_c\).

  5. Fix the bounded-output constant, the scale cutoff and precision coefficient, and the prefix mean exponent. Choose the buffer floor and capacity exponents and its noise coefficient so that all buffer sums converge. Choose task separation and walk length coefficients sufficiently large for the small constants in the clean estimation induction.

  6. Fix the crude range and propagation constants, then the block gain coefficients \(D_2,D_1\) and a sufficiently small rational block width coefficient \(\epsilon\). These choices close the coupled fingerprint induction of Theorem 46.

  7. Choose the fixed operation-stratum and allowance coefficients and the guard constants in Theorem 58. Their finite enlargement does not alter the analytical tables or the distribution of the environment.

  8. Choose the linear top-budget coefficient, the field-size exponent, and the minimum padding of \(B\).

Within each step only finitely many strict inequalities between fixed constants are needed. The proofs in the corresponding sections give these inequalities. In particular, the row rank is fixed before the incoming rank, the block constants follow the crude error constants, and the field is chosen after the reachable resource ranges. There is no dependence of an earlier sampler gap on that last choice.

The coefficient \(\epsilon\) in this order is the fixed fingerprint block-width coefficient. Once it has been chosen, increasing the fixed linear top budget increases the number of salt blocks but leaves the sampler’s gap uniform: all earlier salt bits form one free component in Lemma 20. The storage estimate (94) therefore remains \(O(B)\).

For clarity, one may bound the stage index by \(L\), the inherited suffix length by the stages remaining above a task, and logarithms of reciprocal requested rates by \[ HL+c(M_{\rm top}-M) \tag{132}\] with a sufficiently large fixed \(c\). A child at slack \(d\) can extend the variable scale range by only \(O(d)\) and loses at least \(d\) of budget. Stage-rate resets are already covered by \(HL\). The finitely many local task operations use the stated widths and fixed subranks. Thus Equation (132) is closed under all estimator recipes. The controller guards allow these designated executions.

There is consequently a fixed \(C_P\) such that a prime \[2^{C_PB}\le P\le 2^{C_PB+1}\] is larger than all identifiers, makes the radix invertible, and gives \(P2^{-k}\ge1\) for every rate ever requested. Bertrand’s postulate provides such a prime; testing consecutive integers by trial division finds one using \(O(B)\) space. The search range and running time are \(2^{O(B)}\). Field arithmetic and fixed-dimensional rank tests also use \(O(B)\) space.

A dyadic approximation at the specified start

Lemma 71 (Conditioning on one start). Let \(W\in[0,1]^V\), and suppose an estimator satisfies \[\frac1N\mathbb E\sum_{v\in V}\frac{X_f(v)}f |\bar W(v)-W(v)|^8\le C_b^8 2^{-8p}.\] For any specified \(z\in V\) and every \(\eta>0\), \[\begin{align*} \mathbb E\bigl[|\bar W(z)-W(z)|^8\mid X_f(z)=1\bigr] &\le N\frac f{\pi_f}C_b^8 2^{-8p}, \tag{133}\\ \Pr\bigl[|\bar W(z)-W(z)|>\eta\mid X_f(z)=1\bigr] &\le N\frac f{\pi_f}C_b^8 2^{-8p}\eta^{-8}. \tag{134}\end{align*}\]

Proof. Keep only the nonnegative summand at \(z\). Since \(\Pr[X_f(z)=1]=\pi_f>0\), \[\frac{\pi_f}{f} \mathbb E\bigl[|\bar W(z)-W(z)|^8\mid X_f(z)=1\bigr] =\mathbb E\!\left[\frac{X_f(z)}f|\bar W(z)-W(z)|^8\right] \le N C_b^8 2^{-8p}.\] Dividing by \(\pi_f/f\) proves Equation (133). In particular, choosing one specified coordinate costs the factor \(N\) from the integrated norm. Markov’s inequality gives Equation (134). ◻

Proof of Theorem 70. Apply Lemma 3 to \(\mathcal M\), with no promise on its acceptance probability. Set \[r=d\lceil\log_2(n+2)\rceil,\qquad \Delta=2^{-r},\qquad L=\left\lceil\frac{C_0B+r+3}{H-1}\right\rceil,\] where \(|V_0|\le2^{C_0B}\) is an explicit configuration bound. Since \(d\) is fixed, \(r,L=O(B)\). Let \(z\) be the all-primary copy of the start at stage \(L\). Proposition 11 gives \[ 0\le p_{\mathcal M}(x)-W_L(z)\le\Delta/8. \tag{135}\] This bound holds for every acceptance probability. Choose a fixed \(C_N\) large enough to bound all the resulting stage domains by \(N=2^{C_NB}\).

For the final reward task take \(a=u=f=b_L\). Its rank satisfies \(i_*=O(L+1)\), and Equation (64) reads \[p=M_{\rm top}/K_{\rm prec}-HL-A_*i_*.\] By Lemma 4, fix an integer \(C_{\rm tail}\ge1\) bounding \((f/\pi_f)C_b^8\) throughout the admissible range. Choose the fixed linear coefficient of \(M_{\rm top}=O(B)\) large enough that \[ p\ge r+3+\frac{C_NB+\lceil\log_2 C_{\rm tail}\rceil+3}{8}. \tag{136}\] The top-budget and field-size choices in the preceding subsection allow this enlargement; Equation (132) still bounds every descendant rate. Theorem 46 supplies the reward moment estimate. Lemma 71, with \(\eta=\Delta/8\), now shows that \[ \Pr\bigl[|\bar W_L(z)-W_L(z)|>\Delta/8 \mid X_f(z)=1\bigr]\le1/8. \tag{137}\] Thus the factor \(N\) lost on selecting one coordinate has been paid by the statistical precision. Explicitly, the bound in Equation (134) is at most \[C_{\rm tail}\,2^{C_NB-8p+8r+24}\le 2^{-3},\] by Equation (136).

We next turn the estimator values into a finite list of dyadic numbers. For each master environment, Theorem 56 gives the digits \(w_j\in\{-6,\ldots,6\}\) of its one fixed value \(\bar W_L(z)\). Structural choices do not depend on the requested digit. Put \(K=r+6\), compute \[A_\sigma=\sum_{j=0}^{K}w_j2^{K-j},\qquad a_\sigma=\min\{2^K,\max\{0,A_\sigma\}\},\qquad v_\sigma=a_\sigma2^{-K}.\] The integer \(A_\sigma\) is obtained by starting at zero and applying \(A\leftarrow2A+w_j\) in order, so it uses \(O(K)\) bits. By Equation (106), the truncation error is at most \(6\,2^{-K}<\Delta/8\). Clipping cannot increase error relative to \(W_L(z)\in[0,1]\). Equations (135) and (137) imply that, on more than three quarters of the environments conditioned on \(X_f(z)=1\), \[ |v_\sigma-p_{\mathcal M}(x)|<3\Delta/8. \tag{138}\]

Enumerate the conditional law exactly. Enumerate each field-matrix entry in \(\{0,\ldots,P-1\}\) and each raw salt bit string, retaining precisely those encodings whose positive-degree columns form a frame and for which \(X_f(z)=1\). The retained encodings are equally weighted and give the law in Equation (137). In particular, enumerate the original salt strings, not their reduced fingerprint parameters. Different raw strings that induce the same fingerprint remain separate encodings; their multiplicities reproduce the original probability law. At least one encoding is retained: a frame exists, and its independent constant column can place the designated evaluation in the nonempty retention box. The matrices have fixed dimensions and \(\log P=O(B)\); together with Equation (94), this gives \(O(B)\) bits per master environment and at most \(2^{O(B)}\) encodings to enumerate.

Let \(E\ge1\) be the number of retained encodings, found by one pass, and let \(h=\lfloor E/2\rfloor+1\). Binary-search the integer interval \([0,2^K]\) for the least \(a\) such that at least \(h\) retained encodings satisfy \(a_\sigma\le a\). Each count is computed by a fresh pass, recomputing each numerator \(a_\sigma\). This finds the upper median of the finite multiset, including all repetitions. Since more than three quarters of its members satisfy Equation (138), so does its median: fewer than \(E/4\) values can lie below the indicated interval, and fewer than \(E/4\) can lie above it. The algorithm retains only the current count and search registers; it never stores the multiset. Output the integers \((a,K)\) representing \(z_x=a2^{-K}\). Then \[|z_x-p_{\mathcal M}(x)|<3\Delta/8 \le(n+2)^{-d}.\]

We account for the memory held during this computation. The master environment, current cursor, catalyst, shared scratch, and suspended controller records together use \(O(B)\) bits by Theorem 58, because \(M_{\rm top},K=O(B)\). The outer computation keeps only a fixed number of additional \(O(B)\)-bit registers: its top parameters and designated start, the environment enumeration index, \(E\), \(h\), the current count, the two binary-search endpoints and candidate, the digit index, and the signed numerator accumulator. Each count is bounded by the number of environment encodings, so it also has \(O(B)\) bits. These outer registers remain live while a digit query runs; in particular, both its environment-enumeration state and its numerator accumulator are included. They are present once for the whole computation, not once per recursive level. Every query terminates on every environment, including the inaccurate environments, by the same realization theorem. The output is written only after the median is computed.

Each digit query is a halting deterministic computation with \(O(B)\) work bits and physical input heads confined between the real endmarkers. Counting its finite control, work-tape contents, and all work and input head positions gives at most \(2^{O(B)}\) complete configurations on its fixed input. A repeated complete configuration would force a deterministic loop, so totality bounds the query by \(2^{O(B)}\) steps. There are \(2^{O(B)}\) environment encodings, \(O(K+1)\) binary-search passes, and \(K+1\) digit queries per numerator. Their product is still \(2^{O(B)}=(n+2)^{O(1)}\). All finite parameters are generated uniformly, and all constants are fixed for \((\mathcal M,d)\). ◻

The class equality and promise separation

Proof of Theorem 1. Let \(\mathcal M\) recognize a language in \(\mathsf{BPL}\), with acceptance probability at most \(1/3\) on no inputs and at least \(2/3\) on yes inputs. Apply Theorem 70 with \(d=3\). For every input length, including \(n=0\), its error is at most \((n+2)^{-3}\le1/8\). Hence the dyadic answer is at most \(1/3+1/8=11/24\) on a no input and at least \(2/3-1/8=13/24\) on a yes input. Comparing this answer with \(1/2\) decides the language uniformly in logarithmic space. This proves \(\mathsf{BPL}\subseteq\mathsf L\) directly from probability approximation.

A deterministic machine may ignore its coins, giving \(\mathsf L\subseteq\mathsf{RL}\). Two independent repetitions of an \(\mathsf{RL}\) machine, accepting if either repetition accepts, have acceptance probability zero on no inputs and at least \(3/4\) on yes inputs. They use logarithmic space and polynomial time, so \(\mathsf{RL}\subseteq\mathsf{BPL}\). The three inclusions prove \(\mathsf L=\mathsf{RL}=\mathsf{BPL}\). ◻

A promise problem is a pair \((Y,N)\) of disjoint sets of inputs. Here \(\mathsf{PromiseBPL}\), denoted \(\mathsf{prBPL}\) in the introduction, requires one polynomial-time randomized logarithmic-space machine that halts on every input, accepts each input in \(Y\) with probability at least \(2/3\), and accepts each input in \(N\) with probability at most \(1/3\). The class \(\mathsf{PromiseL}\) consists of pairs separated by a total deterministic logarithmic-space decider; there is no correctness requirement outside \(Y\cup N\).

Corollary 72. Under these conventions, \(\mathsf{PromiseBPL}=\mathsf{PromiseL}\).

Proof. The same \(d=3\) approximation and threshold \(1/2\) used in the preceding proof accept every input in \(Y\) and reject every input in \(N\). The resulting machine halts also outside the promise because its approximation queries are total. The reverse inclusion follows by ignoring the coins. ◻

Accuracy supplied as part of the input

Increasing the numerical clock alone approximates the same estimator value more accurately; it does not reduce the statistical error. To handle an arbitrary requested accuracy while retaining the proved condition \(M_{\rm top}=O(B)\), we apply the fixed-accuracy theorem to a longer input represented implicitly.

Proof of Theorem 2. Fix a padded machine \(\mathcal M_{\rm pad}\) as follows. On a string beginning with \(1^n0x\), where \(|x|=n\), it simulates \(\mathcal M(x)\) and ignores all remaining symbols. If the delimiter or the following \(n\) symbols are absent, it rejects. This is one fixed randomized machine with polynomial worst-case running time and logarithmic work space in its full input length. On every string with the specified prefix its acceptance probability is \(p_{\mathcal M}(x)\). Fix once and for all the transducer of Theorem 70 for \(\mathcal M_{\rm pad}\) and \(d=1\).

On input \((x,1^q)\), with \(n=|x|\) and \(q\ge1\), set \[m=2^{q+\lceil\log_2(2n+4)\rceil+4}.\] Simulate that fixed transducer on the virtual length-\(m\) string \[y=1^n0x0^{m-(2n+1)}.\] An input query is answered by comparing its address with the prefix boundaries and either returning the appropriate bit of \(x\) or a fixed bit. The simulated input-head address and all these comparisons use \(O(\log m)\) bits; the string \(y\) is never stored. The simulated transducer returns a dyadic rational \(z\in[0,1]\) satisfying \[|z-p_{\mathcal M}(x)|\le(m+2)^{-1}<2^{-q-4}.\] Its numerator and exponent have \(O(\log m)\) bits and can be retained in work space while its finite output is simulated.

Compute by exact dyadic arithmetic \[a=\left\lfloor 2^{q+2}z+\tfrac12\right\rfloor.\] Then \(0\le a\le2^{q+2}\), and rounding gives \[\left|a2^{-(q+2)}-p_{\mathcal M}(x)\right| \le2^{-q-3}+2^{-q-4}<2^{-q}.\] Write the binary numerator \(a\) to the output tape. It has at most \(q+3\) bits and represents the denominator specified in the theorem. The simulator state, virtual input address, actual input counters, captured dyadic answer, and rounding operands occupy a fixed number of records of total size \[O(\log m)=O(\log(n+2)+q).\] The simulation and the synthesis of each input bit take polynomial time in \(m\), so the total time is \(m^{O(1)}=(n+2)^{O(1)}2^{O(q)}\). The padded machine and its exponent-one approximator were fixed before reading \(q\). Thus their constants are independent of the requested accuracy. Throughout their simulation, the original size convention is \(B=\Theta(\log(m+2))\), exactly as required by the estimator and controller theorems. ◻

Certified accepting computations

We return to fixed inverse-polynomial accuracy to construct an accepting computation when the machine’s acceptance probability has an inverse-polynomial lower bound. For a transducer, the conclusion concerns outputs certified by an accepting run: its output tape must be write-only, so past emitted symbols do not influence later transitions.

Proof. Give halted configurations constant self-transitions through time \(T(n)\). A configuration and its elapsed time use \(O(\log(n+2))\) bits. There is a fixed randomized logarithmic-space continuation machine which, on a valid encoding of \((x,v,t)\), starts from configuration \(v\) and simulates the remaining \(T(n)-t\) steps with fresh coins. It checks the fixed work-space box and time bound, rejecting malformed encodings and any attempted violation. Its acceptance probability on the reachable configurations is the continuation value \(p_t(v)\), and it has a polynomial worst-case time bound in its input length.

Write \(\gamma=(n+2)^{-c}\). Choose a fixed exponent \(d\) large enough that Theorem 70, applied to the continuation machine, approximates every queried value within \(\delta\le\gamma/(8T(n))\). Such a fixed exponent exists because \(T\) is a fixed polynomial and the continuation input contains \(x\). The extra fields have logarithmic length. We supply them virtually from saved configuration and time registers, so the approximation and those registers together use logarithmic space.

Starting at the initial configuration, at each time estimate the continuation value of every positive-probability successor and choose one with maximum estimated value, breaking ties by a fixed port order. Since \(p_t(v)\) is the weighted mean of the successor values, the chosen successor \(v'\) satisfies \[p_{t+1}(v')\ge\max_w p_{t+1}(w)-2\delta \ge p_t(v)-2\delta.\] On an input satisfying the stated lower bound, after \(T(n)\) choices the terminal continuation value is at least \(\gamma-2T(n)\delta\ge3\gamma/4>0\). A terminal continuation value is zero or one, so this computation accepts.

First run this construction without writing output. If its terminal state rejects, output the failure symbol. Otherwise repeat the same deterministic choices and stream the configurations through the original first halting state, or the symbols emitted by the simulated transducer. The artificial self-transitions after halting are omitted from the output. This second pass follows the same accepting computation. Consequently no rejecting computation is output, even outside the probability promise.

At any time the current, candidate, and best successor configurations, the time counter, a constant number of dyadic scores, the virtual input address, and the approximation workspace occupy a constant number of \(O(\log(n+2))\)-bit records. The finite successor lists have constant size, and there are only \(T(n)\) steps in each pass, so the total running time is polynomial. Each approximation has a deterministic guarantee at its queried configuration; no simultaneous accuracy event for all adaptive queries is needed. ◻

The effective compiler

The probability-approximation theorem fixes the source machine. We now construct one effective map from source descriptions and supplied resource bounds to deterministic simulations with explicit bounds.

Theorem 74 (Effective simulation with explicit bounds). There is a terminating algorithm which, given a finite randomized machine description \(M\) and positive integers \(a,b\), emits a deterministic machine \(D_M\) and positive integers \(K,H,c\) with the following guarantee. If, on every input of length \(n\) and every random tape, \(M\) halts within \((n+2)^a\) steps, uses at most \(b\lceil\log(n+2)\rceil\) work bits, and has acceptance probability at most \(1/3\) on no inputs and at least \(2/3\) on yes inputs, then \(D_M\) decides the same language. On each input, that same execution uses at most \(K\log(n+2)\) work bits and \(H(n+2)^c\) elementary bit operations.

The compiler’s own running time need only be finite, and its output exponent may depend on the source.

We prove Theorem 74 by specializing one fixed deterministic program. First we construct a randomized interpreter \(U\) that enforces polynomial-time and logarithmic-space bounds on every input. Theorem 70 supplies a deterministic machine \(A\) that separates the inputs on which \(U\) accepts with probability at least \(2/3\) from those on which it accepts with probability at most \(1/3\). Both \(U\) and \(A\) are fixed before the compiler receives a source description \(M\) and bounds \(a,b\).

When the source satisfies its supplied resource bounds, padding places each encoded computation within \(U\)’s caps without changing its acceptance probability. The generated machine runs \(A\) on that padded input, generating each requested symbol from the real input. We then compute the output bounds \(K,H,c\) from the generated finite transition table and the fixed library data. The compiler never has to select a new probability approximator or infer resource bounds from the source’s behavior.

Universal deterministic estimation has a substantial antecedent in Pyne, Raz and Zhan (Pyne et al. 2023, full version, Theorem 5.1). They construct one deterministic algorithm that, on input \((1^n,\mathcal B)\) with \(n\ge1\) and \(\mathcal B\) an ordered branching program of length and width at most \(n\), approximates its acceptance probability to error less than \(1/n\). For every space-constructible \(S(n)\ge\log n\), this single algorithm uses \(O(S(n))\) space exactly when \(\mathsf{prBPL}\subseteq\mathsf{SPACE}[O(S(n))]\). Here the preceding sections establish the separator used by one fixed library, and we compute explicit resource bounds for its source-specific specialization.

All quantities in this section are local to the compiler construction. In particular, \(P\) is a padding exponent, \(d_{\mathrm{code}}\) is a program length, \(B_n=\lceil\log_2(n+2)\rceil\), and \(N_{\mathrm{pad}}\) is a virtual input length. They do not denote the statistical budgets, field modulus, or domain sizes used in the preceding sections. We use the usual counted-cell convention for work tapes: traversed writable cells count toward space, whether or not they contain a nonblank symbol. The finite control and the read-only input have their standard treatment.

One fixed interpreter with enforced resource bounds

Fix a binary coding of ordinary finite transition lists. One concrete choice uses a fixed record grammar and the field encoding \(w\mapsto 1^{|w|}0w\). Lists specify the finite state set, tape and alphabet declarations, initial and halting states, blanks and endmarkers, and transition records. A transition record lists its tested state, scanned symbols and coin pattern, followed by its next state, written symbols and local head movements. The grammar has fixed nesting depth; its fields are actual listed data, not recursively executable descriptions. State and symbol IDs are canonical binary integers used only as names. The first matching transition is used; a missing transition causes rejection. Wildcards for ignored coins can equivalently be expanded into a finite explicit list.

Write \(z\) for a program in this coding and \(d_{\mathrm{code}}=|z|\ge1\). The coding permits at most \(d_{\mathrm{code}}\) work tapes and a block of at most \(d_{\mathrm{code}}\) fresh fair bits per simulated transition. Every ID has at most \(d_{\mathrm{code}}\) bits. A harmless padding field can enlarge the program to meet these inequalities. All tape movements are local. Work coordinates are measured from their initial origins; the ordinary model has one input head and one head per work tape. For a source input \(x\) of length \(n\), let \(E_{d_{\mathrm{code}}}(x)\) encode each of its symbols by its ID padded to exactly \(d_{\mathrm{code}}\) bits.

Preserving the supplied source bounds.

An ordinary source machine \(M\) can be converted effectively to this format by renaming its states and symbols and listing its transitions. Its tapes are represented directly, so a source transition retains its work writes and head movements and counts as one interpreted transition. The conversion does not pass through a slower single-tape simulation.

Under the optional one-way read-only random-tape convention, the source may reuse its current random bit. Store a cached bit and a validity flag in finite control, initially invalid. At each transition obtain one fresh bit. Use it if the cache is invalid; otherwise use the cached bit and discard the fresh one. Execute the source transition, then retain the used bit if the random head stayed and invalidate the cache if it advanced. This samples a newly visited random cell lazily on its next use. It preserves the joint distribution of the source run and needs no initialization transition. An initially halting state is recognized before the first transition. The same argument uses a fixed finite block for a convention allowing finitely many fresh bits per elementary transition, discarding unused bits. The finite cache changes neither counted work space nor the number of source transitions. Thus the supplied parameters \(a,b\) remain valid; only the static description length can increase.

Lemma 75 (A total resource-capped interpreter). There is one fixed randomized Turing machine \(U\) with polynomial worst-case running time and \(O(\log(m+2))\) counted work bits on every binary input of length \(m\), with the following behavior. Set \[ \lambda_m=\lceil\log_2(m+2)\rceil, \qquad R(m,d_{\mathrm{code}}) =\left\lfloor\frac{\lambda_m} {100(d_{\mathrm{code}}+1)^3}\right\rfloor. \tag{139}\] On an input with prefix \[ 1^{d_{\mathrm{code}}}0\,z\,1^n0\, E_{d_{\mathrm{code}}}(x), \tag{140}\] it first requires \(1\le d_{\mathrm{code}}\le\lambda_m\) and \(R(m,d_{\mathrm{code}})\ge1\). It then simulates the encoded machine for at most \(m\) transitions, allowing each counted work tape only coordinates in \([-R,R]\). It rejects malformed data, a requested move outside these bounds, or failure to halt by the end of the transition budget. Otherwise it returns the encoded machine’s decision. The suffix after the complete prefix has no parsing significance.

Proof. Count \(m\) before interpreting the prefix. Scan the unary code-length header with a bounded counter, rejecting if its length exceeds \(\lambda_m\) or the required delimiter is absent. Check the two eligibility conditions before copying \(z\) or allocating simulation storage. Then verify that the following \(d_{\mathrm{code}}\) code bits, the unary data-length header, and the required data bits are present. The parsed \(n\) is at most \(m\), and \(d_{\mathrm{code}}n\le\lambda_m m\), so even size calculations preceding a failed length check use only \(O(\lambda_m)\) bits. List counts, ID lengths, tape declarations, and movements are checked against the actual code and the stated bounds before any loops or allocations using them. There is no enumeration of an implicit domain of symbol or state names.

Store the encoded tapes sequentially on a fixed number of real work tapes of \(U\). At most \(d_{\mathrm{code}}\) strips, each with \(2R+1\) cells of \(d_{\mathrm{code}}\) bits, require \[ d_{\mathrm{code}}^2(2R+1)=O(\lambda_m) \tag{141}\] bits. Indeed, eligibility gives \(d_{\mathrm{code}}^3\le\lambda_m\) and \(d_{\mathrm{code}}^2R\le\lambda_m\). The copied program, current state, scanned and written symbol tuples, coin block, strip coordinates, step counter, and parsing scratch fit simultaneously in \(O(\lambda_m)\) bits. A single simulated input-head coordinate has magnitude \(O(m)\) by the transition cap and therefore also uses \(O(\lambda_m)\) bits. Read the corresponding source symbol by rescanning the encoded data, supplying the appropriate markers or off-data blanks according to the encoded input-tape convention.

All routines are bounded scans and ordinary binary arithmetic. They iterate over actual code positions, input positions, a bounded tape/head list, strip positions, and bit positions, with a fixed nesting depth. Each range is polynomially bounded in \(m+2\); binary arithmetic is implemented by bit scans, not unit-cost operations on growing words. Matching a listed transition, fetching input symbols, updating the strips, and obtaining its finite coin block therefore take polynomial time in \(m+2\), with one fixed exponent determined by these routines. The same bounds hold when a check fails. There are at most \(m\) interpreted transitions. Checking the initial halting state and the state reached by the last permitted transition gives the asserted total behavior, including \(m=0\). The resulting fixed \(U\) has the required polynomial worst-case time on every random tape. ◻

No bounded-error language promise is asserted for \(U\) on arbitrary inputs. Its acceptance probability may lie between \(1/3\) and \(2/3\). This is why we use probability approximation rather than a bare language-class containment.

Fixing one deterministic library

Apply Theorem 70 to the fixed machine \(U\) and accuracy exponent \(3\), and compare its dyadic answer with \(1/2\). The error on an input of length \(m\) is at most \((m+2)^{-3}\le1/8\). Hence the resulting deterministic machine accepts whenever \(\Pr[U\text{ accepts}]\ge2/3\) and rejects whenever that probability is at most \(1/3\), since \[\frac13+\frac18=\frac{11}{24}<\frac12 <\frac{13}{24}=\frac23-\frac18.\] It halts on every input, including inputs between the two thresholds, as also stated in Corollary 72. Its logarithmic space bound counts all simultaneously live work in the approximation algorithm. Store the \(O(\log(m+2))\)-bit dyadic output in work space for the threshold test.

We fix this library once in the following convenient form. It is an ordinary finite-tape deterministic decision machine \(A\) with one input head confined to the endmarked input, initially at its left endmarker, such that every work head stays in \[ [-(\lambda_m+1),\lambda_m+1] \tag{142}\] relative to its origin. To obtain this form, first consolidate any input heads by keeping their coordinates in logarithmic work space and scanning for each read. Even an off-data coordinate under an unbounded-input convention has logarithmic length because the original library has polynomial running time. A fixed change of initial control state handles its input-head start convention.

Next choose an integer \(w\) at least a coefficient bounding each counted work head’s distance from its origin by \(w\lambda_m\). Such one fixed \(w\) exists from the logarithmic bound, enlarged if necessary for the finitely many small lengths. Group consecutive work cells into blocks of width \(w\), storing within-block offsets in finite control. The resulting finite alphabet and finite transition list realize Equation (142), also for two-sided tapes. All symbol widths will be charged in bits below; grouping only exposes a uniform address bound and gives no uncharged bit compression.

This choice of \(A\), including its one blocking width, precedes every compiler input. More formally, the remaining construction is a computable syntactic function of a finite library list and \((M,a,b)\). Fixing one valid finite list yields one ordinary computable compiler on all those source inputs. It does not require an algorithm that discovers a separator or a logarithmic-space coefficient from an arbitrary machine description. The existence of the fixed list was proved by Theorem 70; no constant is selected as a noncomputable function of the source machine or its ordinary input.

The fixed finite data are the program grammar, the transition list of \(A\) with its blocked work alphabets and head conventions, and the finite callback templates described below. The construction constants, cutoffs, generator choices and finite instruction routines used by the separator are included in that list. Its blocking width is fixed at the same time. These are data specifying one compiler; subsequent source-dependent integers are computed from these data and the supplied \((M,a,b)\).

Padding and exact preservation of acceptance probability

Given the source description and its supplied positive integers \(a,b\), construct \(z\) as above and compute \[ \begin{split} P&=1000(d_{\mathrm{code}}+2)^4(a+b+2),\\ B_n&=\lceil\log_2(n+2)\rceil, \qquad N_{\mathrm{pad}}=2^{PB_n}. \end{split} \tag{143}\] The generated machine will run \(A\) on a virtual binary input \(y_x\) of length \(N_{\mathrm{pad}}\) whose prefix is Equation (140) and whose remaining bits are zero. Its prefix length is exactly \[ L_{\mathrm{pre}} =(d_{\mathrm{code}}+1)+d_{\mathrm{code}}+(n+1) +d_{\mathrm{code}}n =(d_{\mathrm{code}}+1)(n+2). \tag{144}\] Since \(B_n\ge1\), \(2^{B_n}\ge n+2\), and \(2^{P-1}\ge d_{\mathrm{code}}+1\), this length is at most \(N_{\mathrm{pad}}\). The interpreter’s length parameter is \[ \ell_{\mathrm{pad}} :=\lambda_{N_{\mathrm{pad}}}=PB_n+1. \tag{145}\]

All eligibility checks pass. In particular, \[\frac{P}{100(d_{\mathrm{code}}+1)^3} =10\frac{(d_{\mathrm{code}}+2)^4} {(d_{\mathrm{code}}+1)^3}(a+b+2) \ge30(a+b+2),\] so \[ R(N_{\mathrm{pad}},d_{\mathrm{code}}) \ge30(a+b+2)B_n-1\ge(a+b)B_n+4. \tag{146}\] Also \(P>a\), whence \((n+2)^a\le2^{aB_n}<N_{\mathrm{pad}}\). Local motion from an initial work origin to a cell entails visiting the intervening counted cells. The source bound of \(bB_n\) work bits therefore places every source work cell in the allocated strip; the slack in Equation (146) also covers a final one-cell move. This uses the source’s original cell costs, each at least one bit, and does not charge it for the larger interpreter IDs.

Consequently no resource cutoff changes a promised source run, and \[ \Pr[U(y_x)\text{ accepts}]=\Pr[M(x)\text{ accepts}]. \tag{147}\] There is no conditioning or additional approximation in this equality. The decision of \(A(y_x)\) is therefore the decision required for \(L_M\). Although \(N_{\mathrm{pad}}\) is given exponentially, it is at most \((n+2)^{2P}\) for the fixed source. More importantly, it need never be written as an input string.

Remark 76 (Optional head conventions). The construction also applies when the fixed coding and \(U\) are chosen to allow at most \(d_{\mathrm{code}}\) heads per tape, with the usual common initial origin. These conventions are fixed before choosing \(A\). There are then at most \(d_{\mathrm{code}}^2\) work-head coordinates and scanned symbols. Their storage is bounded by \[O\bigl(d_{\mathrm{code}}^2\log(2R+1) +d_{\mathrm{code}}^3\bigr)=O(\lambda_m).\] For multiple input heads, or an optional entirely blank tape whose cells are not charged, use signed \(R\)-bit coordinates and reject on overflow. At most \(d_{\mathrm{code}}^2\) such coordinates still cost \(O(d_{\mathrm{code}}^2R)\) bits. These optional conventions do not change the interpreter’s bounds. The padded instances just constructed never overflow those coordinates.

Under the optional head conventions, local motion bounds an input or uncharged blank-tape coordinate by the source time, up to the fixed endmarker offset. Its signed encoding fits because \(R\ge aB_n+4\).

Figure 2 shows how the generated machine supplies the virtual input requested by \(A\). The library’s work tapes remain live while the callback scans the real input; the resource bound below counts the library and callback storage simultaneously.

The fixed library reads the padded input through a callback. The callback regenerates prefix symbols from the real input and returns zero for an address in the padding interval.

Compiling the virtual input operations

Inline the finite transition list of \(A\), giving it its own work tapes. The generated machine \(D\) suspends its library control state in finite control during an input operation. Four auxiliary binary words suffice: the length count \(n+1\), the virtual right-endmarker address \(N_{\mathrm{pad}}+1\), the virtual input-head position \(J\), and a prefix position \(I\). Each has a separate tape, with origin and termination marks and digits in low-end-first order.

Initialize the length count to one and increment it once per real input symbol. Its bit length is exactly \[ \operatorname{bitlength}(n+1) =\lfloor\log_2(n+1)\rfloor+1 =\lceil\log_2(n+2)\rceil=B_n \qquad(n\ge0). \tag{148}\] For every bit place in that count, write \(P\) zero bits on the bound tape, then append a one. This produces the low-end-first representation of \(2^{PB_n}=N_{\mathrm{pad}}\); increment it to obtain \(N_{\mathrm{pad}}+1\). The repetition of \(P\) is hardwired finite control. Initially set \(J=0\), the library’s left-endmarker position.

At a virtual read, first test \(J=0\) and \(J=N_{\mathrm{pad}}+1\) and return the corresponding endmarker. Otherwise initialize \(I=1\) and regenerate the prefix in its natural order:

  1. emit the fixed string \(1^{d_{\mathrm{code}}}0z\);

  2. scan \(x\), emitting one \(1\) per symbol, and then emit \(0\);

  3. rewind and scan \(x\) again, emitting each fixed-width symbol code in \(E_{d_{\mathrm{code}}}(x)\).

Here “emit” means compare the current \(I\) with \(J\), return the bit if they agree, and otherwise increment \(I\) and continue. If the scan finishes without a match, return zero. The strings \(z\), the integer \(d_{\mathrm{code}}\), symbol codes and emission phases are fixed in the generated control. The real input head is rewound by scanning to its endmarker, without storing its position. In particular, \(J=N_{\mathrm{pad}}+1\) is answered before regeneration, while a padding address such as \(J=N_{\mathrm{pad}}\) returns zero after the finite prefix scan. The largest prefix counter value is \(L_{\mathrm{pre}}+1\le N_{\mathrm{pad}}+1\).

For explicit counter routines, an increment uses a finite carry flag. Compare two low-end-first words by scanning both, retaining only a finite comparison flag that is overwritten by each higher differing bit. After one terminator, treat its later bits as zero using a finite “ended” flag. Rewind each tape to its origin mark afterward. Zeroing and copying are likewise bounded bit scans. No growing arithmetic register beyond the named words is required. All words have at most \(PB_n+2\) bits; including delimiters and rewinds, their heads need visit only positions through \(PB_n+10\).

After the callback, perform the indicated library transition and update \(J\) in place for its input-head motion. The library’s work heads and contents are untouched by the callback. Its state, the returned input symbol, and any step-local scanned tuple are fixed finite-control data. The callback does not invoke \(A\) recursively or maintain a stack of callbacks. The real input head stays between its real endmarkers.

Every listed routine terminates. For example, a complete regeneration uses at most \(L_{\mathrm{pre}}\) emissions, each with an \(O(PB_n+1)\)-bit comparison and increment, together with two real-input scans. Initialization makes \(n\) increments and writes \(PB_n+1\) bound bits. These observations give direct polynomial bounds for the auxiliary operations. The whole simulation halts because \(A\) is total. Its more uniform explicit time bound, including every such operation, is derived next from the complete configurations of this same machine.

All the above instructions can be generated as finite transition tables, unrolling fixed repetitions and strings if necessary. Formation of the source list, the integer \(P\), and this transition table is a finite effective process. It makes no test of the source’s semantic error or resource promises. Syntactically invalid compiler inputs may be assigned a fixed default output; correctness is required on the stipulated class. The generated simulation is total and obeys the bounds below even when the supplied source promises fail; those promises are used to establish that its decision agrees with the source.

Explicit simultaneous bounds

Use the finite-alphabet description just constructed to compute \(t\), its number of work tapes; \(g\ge2\), the maximum of two and its work alphabet sizes, including all blank, origin and termination symbols; and \(q_D\), its number of states, including callback phases. Let \(u\ge1\) dominate its explicit transition-table bit length, all alphabet sizes (including the input alphabet), the number of states, and the number of tapes. Padding this integer is effective. Return the positive integer bounds \[ \begin{split} \rho&=100(P+2),\qquad S=1+(t+1)(2\rho+1)(g+2),\\ K&=4S,\qquad c=2S+1,\qquad H=2^{(u+100)^2}(q_D+1). \end{split} \tag{149}\] These depend on the fixed source and the fixed library, and are computed from finite syntax. They have no dependence on \(n\) or \(x\).

All live work bits.

Every work head stays in \([-\rho B_n,\rho B_n]\). For a library tape this follows from Equations (142) and (145), since \(\ell_{\mathrm{pad}}+1=PB_n+2\le\rho B_n\). For an auxiliary tape it follows from \(PB_n+10\le\rho B_n\), using \(B_n\ge1\). Thus the library’s live data and all callback data satisfy the same bound at the same time. Encoding a work alphabet of size \(g_i\) in \(\max\{1,\lceil\log_2 g_i\rceil\}\le g\) bits per symbol bounds their total data by \[ t(2\rho B_n+1)g\le SB_n. \tag{150}\] Even twice this quantity satisfies the stronger logarithmic bound: with \(L_n=\log_2(n+2)\ge1\) we have \(B_n\le2L_n\) and \[2SB_n\le4SL_n=K L_n.\]

All finite-alphabet transitions.

For a fixed real input \(x\), a complete configuration is specified by the state, the real input-head position, all work-head positions, and all work symbols in the stated intervals. Everything outside these intervals remains at its initial value. There are at most \[ q_D(n+2)(2\rho B_n+1)^t g^{t(2\rho B_n+1)} \le q_D(n+2)2^{SB_n} \le q_D(n+2)^c \tag{151}\] such configurations. For the first inequality use \[2\rho B_n+1\le(2\rho+1)B_n, \quad \log_2g\le g, \quad \log_2(2\rho B_n+1)\le(2\rho+1)B_n;\] the second uses \(B_n\le2\log_2(n+2)\) and \(c=2S+1\). A deterministic halting computation cannot repeat a complete configuration. Therefore Equation (151) bounds its total transitions, including initialization, counter arithmetic, every regenerated input bit, and every library transition. The virtual head position is already in the counted word \(J\); it is not an additional uncounted configuration component.

Elementary bit operations.

The preceding count refers to ordinary finite-alphabet transitions. To charge them in bits, encode each work symbol by a fixed-width block, mapping blank to the all-zero block. Read and write within the current block, return to its first bit, and move to the first bit of the target block. Within-block phases and scanned finite tuples are stored in finite control. No per-block writable delimiter is needed, and these operations visit only blocks whose macro positions occur in the proved work intervals. Existing origin and termination symbols have already been included in \(g\). Thus the block implementation obeys the bit-space estimate above; the factor two also allows the usual blank/bit tape convention. Input-symbol access has fixed width and is covered by the same finite description parameter \(u\).

There are at most \(u+1\) scanned symbols of at most \(u\) bits each. An explicit sequential comparison against at most \(u\) transition rows, followed by the fixed-width reads, writes and moves, uses at most a fixed polynomial in \(u+1\) bit actions per transition. For example, the deliberately loose bound \(100(u+1)^8\) suffices for these scans and comparisons, with all temporary phase and symbol information in finite control. It is dominated, for \(u\ge1\), by \[100(u+1)^8\le2^{(u+100)^2}.\] The binary microcode may have more states than \(q_D\), but no new configuration count is needed: multiply the number of original transitions by this per-transition bit cost. Equations (149) and (151) give the simultaneous bound \[ \operatorname{time}_{\mathrm{bit}}(D,x) \le H(n+2)^c. \tag{152}\] No growing-word arithmetic is charged as a single elementary bit action.

Completion of the proof of Theorem 74. The compiler returns the described machine \(D\) and the integers \(K,H,c\) in Equation (149). Its finite syntactic construction terminates. On every promised source input, Equation (147) and the fixed separator’s threshold guarantee give the correct decision. The totality of \(A\) and the finite callback loops prove halting. Equations (150) and (152) give at most \(K\log_2(n+2)\) work bits and \(H(n+2)^c\) bit operations for this same machine.

All claims include \(n=0\): then \(B_n=1\), the real-input scans are empty, the count \(n+1=1\) has one bit, and every displayed padding, strip, space, and time inequality remains valid. The library’s grouping width already includes its finitely many small-length exceptions. All construction constants are fixed before the ordinary input is read.

The work and time internal to probability approximation—environment enumeration, numerical digits, recursive interfaces, recomputation, and precision management—belong to the one complete library machine proved in Theorem 70. Its tapes remain counted throughout virtualization. The compiler adds only the exact input operations analyzed here, without an evaluation oracle or an external stack. Configuration counting gives polynomial time because this entire deterministic machine uses exact logarithmic space. ◻

Ahmadinejad, AmirMahdi, Jonathan Kelner, Jack Murtagh, John Peebles, Aaron Sidford, and Salil Vadhan. 2020. “High-Precision Estimation of Random Walks in Small Space.” 61st IEEE Annual Symposium on Foundations of Computer Science (FOCS 2020), 1295–306. https://doi.org/10.1109/FOCS46700.2020.00123.
Ajtai, Miklós, János Komlós, and Endre Szemerédi. 1987. “Deterministic Simulation in LOGSPACE.” Proceedings of the Nineteenth Annual ACM Symposium on Theory of Computing, 132–40. https://doi.org/10.1145/28395.28410.
Aleliunas, Romas, Richard M. Karp, Richard J. Lipton, László Lovász, and Charles Rackoff. 1979. “Random Walks, Universal Traversal Sequences, and the Complexity of Maze Problems.” 20th Annual Symposium on Foundations of Computer Science, 218–23. https://doi.org/10.1109/SFCS.1979.34.
Armoni, Roy, Amnon Ta-Shma, Avi Wigderson, and Shiyu Zhou. 2000. “An \(O(\log(n)^{4/3})\) Space Algorithm for \((s,t)\) Connectivity in Undirected Graphs.” Journal of the ACM 47 (2): 294–311. https://doi.org/10.1145/333979.333984.
Avizienis, Algirdas. 1963. “On a Flexible Implementation of Digital Computer Arithmetic.” In Information Processing 1962, edited by C. M. Popplewell. North-Holland. https://www.ece.ucdavis.edu/~vojin/CLASSES/EPFL/Papers/1-Avizienis-Flex-Implmnt%20of%20Digital%20Com%20Arith.pdf.
Babai, László, Noam Nisan, and Márió Szegedy. 1992. “Multiparty Protocols, Pseudorandom Generators for Logspace, and Time-Space Trade-Offs.” Journal of Computer and System Sciences 45 (2): 204–32. https://doi.org/10.1016/0022-0000(92)90047-M.
Borodin, Allan, Stephen Cook, and Nicholas Pippenger. 1983. “Parallel Computation for Well-Endowed Rings and Space-Bounded Probabilistic Machines.” Information and Control 58 (1–3): 113–36. https://doi.org/10.1016/S0019-9958(83)80060-6.
Braverman, Mark, Gil Cohen, and Sumegha Garg. 2020. “Pseudorandom Pseudo-Distributions with Near-Optimal Error for Read-Once Branching Programs.” SIAM Journal on Computing 49 (5): STOC18-242-STOC18-299. https://doi.org/10.1137/18M1197734.
Buhrman, Harry, Richard Cleve, Michal Koucký, Bruno Loff, and Florian Speelman. 2014. “Computing with a Full Memory: Catalytic Space.” Proceedings of the 46th Annual ACM Symposium on Theory of Computing, 857–66. https://doi.org/10.1145/2591796.2591874.
Cai, Jin-Yi, Venkatesan T. Chakaravarthy, and Dieter van Melkebeek. 2006. “Time-Space Tradeoff in Derandomizing Probabilistic Logspace.” Theory of Computing Systems 39 (1): 189–208. https://doi.org/10.1007/s00224-005-1264-9.
Carter, J. Lawrence, and Mark N. Wegman. 1979. “Universal Classes of Hash Functions.” Journal of Computer and System Sciences 18 (2): 143–54. https://doi.org/10.1016/0022-0000(79)90044-8.
Chen, Ben, Gil Cohen, Dean Doron, Yuval Khaskelberg, and Amnon Ta-Shma. 2026. “Improved Error Reduction for Weighted PRGs.” Approximation, Randomization, and Combinatorial Optimization. Algorithms and Techniques (APPROX/RANDOM 2026), Leibniz international proceedings in informatics, vol. 392: 39:1–23. https://doi.org/10.4230/LIPIcs.APPROX/RANDOM.2026.39.
Cheng, Kuan, and William M. Hoza. 2022. “Hitting Sets Give Two-Sided Derandomization of Small Space.” Theory of Computing 18 (21): 1–32. https://doi.org/10.4086/toc.2022.v018a021.
Cheng, Kuan, and Ruiyang Wu. 2026. SC Derandomization for Regular ROBPs and Models Beyond BPL. https://arxiv.org/abs/2609.23603.
Cohen, Gil, Dean Doron, Oren Renard, Ori Sberlo, and Amnon Ta-Shma. 2021. “Error Reduction for Weighted PRGs Against Read Once Branching Programs.” 36th Computational Complexity Conference (CCC 2021), Leibniz international proceedings in informatics, vol. 200: 22:1–17. https://doi.org/10.4230/LIPIcs.CCC.2021.22.
Cohen, Gil, Dean Doron, Ori Sberlo, and Amnon Ta-Shma. 2023. “Approximating Iterated Multiplication of Stochastic Matrices in Small Space.” Proceedings of the 55th Annual ACM Symposium on Theory of Computing, 35–45. https://doi.org/10.1145/3564246.3585181.
Cook, Stephen A., and Pierre McKenzie. 1987. “Problems Complete for Deterministic Logarithmic Space.” Journal of Algorithms 8 (3): 385–94. https://doi.org/10.1016/0196-6774(87)90018-6.
Gill, John. 1977. “Computational Complexity of Probabilistic Turing Machines.” SIAM Journal on Computing 6 (4): 675–95. https://doi.org/10.1137/0206049.
Hoza, William M. 2021. “Better Pseudodistributions and Derandomization for Space-Bounded Computation.” Approximation, Randomization, and Combinatorial Optimization. Algorithms and Techniques (APPROX/RANDOM 2021), Leibniz international proceedings in informatics, vol. 207: 28:1–23. https://doi.org/10.4230/LIPIcs.APPROX/RANDOM.2021.28.
Lange, Klaus-Jörn, Pierre McKenzie, and Alain Tapp. 2000. “Reversible Space Equals Deterministic Space.” Journal of Computer and System Sciences 60 (2): 354–67. https://doi.org/10.1006/jcss.1999.1672.
Nisan, Noam. 1992. “Pseudorandom Generators for Space-Bounded Computation.” Combinatorica 12 (4): 449–61. https://doi.org/10.1007/BF01305237.
Nisan, Noam. 1994. “\(\mathrm{RL}\subseteq\mathrm{SC}\).” Computational Complexity 4 (1): 1–11. https://doi.org/10.1007/BF01205052.
Nisan, Noam, Endre Szemerédi, and Avi Wigderson. 1992. “Undirected Connectivity in \(O(\log^{1.5} n)\) Space.” Proceedings of the 33rd Annual Symposium on Foundations of Computer Science, FOCS ’92, 24–29. https://doi.org/10.1109/SFCS.1992.267822.
Nisan, Noam, and David Zuckerman. 1996. “Randomness Is Linear in Space.” Journal of Computer and System Sciences 52 (1): 43–52. https://doi.org/10.1006/jcss.1996.0004.
Pyne, Edward, Ran Raz, and Wei Zhan. 2023. “Certified Hardness Vs. Randomness for Log-Space.” 64th IEEE Annual Symposium on Foundations of Computer Science (FOCS 2023), 989–1007. https://doi.org/10.1109/FOCS57990.2023.00061.
Pyne, Edward, and Roei Tell. 2026. Using Hardness Vs Randomness to Design Low-Space Algorithms. Electronic Colloquium on Computational Complexity, Report TR26-045. https://eccc.weizmann.ac.il/report/2026/045/.
Reingold, Omer. 2005. “Undirected ST-Connectivity in Log-Space.” Proceedings of the Thirty-Seventh Annual ACM Symposium on Theory of Computing, STOC ’05, 376–85. https://doi.org/10.1145/1060590.1060647.
Reingold, Omer. 2008. “Undirected Connectivity in Log-Space.” Journal of the ACM 55 (4): 17:1–24. https://doi.org/10.1145/1391289.1391291.
Reingold, Omer, Luca Trevisan, and Salil Vadhan. 2006. “Pseudorandom Walks on Regular Digraphs and the RL Vs. L Problem.” Proceedings of the 38th Annual ACM Symposium on Theory of Computing, 457–66. https://doi.org/10.1145/1132516.1132583.
Riesz, Marcel. 1927. “Sur Les Maxima Des Formes Bilinéaires Et Sur Les Fonctionnelles Linéaires.” Acta Mathematica 49: 465–97. https://archive.ymsc.tsinghua.edu.cn/pacm_download/117/5399-11511_2007_Article_BF02564121.pdf.
Saks, Michael, and Shiyu Zhou. 1999. “\(\mathrm{BP}_{\mathrm H}\mathrm{SPACE}(S)\subseteq \mathrm{DSPACE}(S^{3/2})\).” Journal of Computer and System Sciences 58 (2): 376–403. https://doi.org/10.1006/jcss.1998.1616.
Schwartz, Jacob T. 1980. “Fast Probabilistic Algorithms for Verification of Polynomial Identities.” Journal of the ACM 27 (4): 701–17. https://doi.org/10.1145/322217.322225.
Shalom, Yehuda. 1999. “Bounded Generation and Kazhdan’s Property (T).” Publications Mathématiques de l’IHÉS 90: 145–68. https://doi.org/10.1007/BF02698832.
Thorin, G. O. 1948. “Convexity Theorems Generalizing Those of M. Riesz and Hadamard with Some Applications.” PhD thesis, Lund University. https://www.mathnet.ru/eng/mat17.
LEVEL 1 COMPLETE!
You read 55,488 words and 2,976 formulas. Your math teacher would be proud.
Converted from the LaTeX source. Something look off? The original PDF is the real thing.

Cool Links: openai/math   Lean   Mathlib   arXiv   the real Coolmath Games