A
D
V
E
R
T
I
S
E
M
E
N
T
ADVERTISEMENT
A uniform Hilbert transform estimate for Lipschitz directions
expertly designed by an internal OpenAI model  ·  released 2026-09-25  ·  original PDF
Theorems: 6 Lemmas: 31 Proofs: 47
Formulas: 3,537 Words: 39,013 Play time: ~4 hours

>>> How to Play <<<
We prove a uniform L2 bound for the Hilbert transform along a Lipschitz unit vector field in the plane, with integration restricted to a fixed absolute multiple of the reciprocal Lipschitz seminorm. The bound is uniform in the inner truncation and holds for fields depending on both coordinates. This gives an affirmative answer to Stein's weak-$(2,2)$ conjecture.

>>> Level Map <<<
  1. Short transforms and band reduction
  2. The problem and its predecessors
  3. The analytic and geometric reductions
  4. The phase gain and its iteration
  5. Organization, inputs and uniformity
  6. Fourier conventions and the local reduction
  7. A maximal theorem and elementary analytic tools
  8. The local slope operator and its endpoints
  9. The band theorem and summation of amplitude levels
  10. A Lipschitz commutator
  11. Square-function reassembly
  12. Removing long and very short times
  13. From a spatial truncation to a finite scale sum
  14. Narrow vertical bins
  15. The admissible packet forms
  16. The precise remaining model estimates
  17. Ordinary tiles and the popular forest
  18. Density and the model estimates
  19. Maximal functions and separated synthesis
  20. Removing large size
  21. Removing large density
  22. The tree estimate
  23. Proof of the model estimates
  24. The improvement problem and popular tops
  25. The assignment to popular tops
  26. The adjoint sum and exceptional subsets
  27. A weak \(L^2\) estimate for the top count
  28. Parameters and masks
  29. Reference-line concentration and significant directions
  30. Positive estimates and the main coefficients
  31. Fourier sums and a dyadic martingale
  32. Every-line moment bounds
  33. Uniform top bounds and neighborhood derivatives
  34. The masked sum and the count threshold
  35. A simultaneous count of separated significant directions
  36. A counting lemma on a line
  37. Transferring witnesses to a reference line
  38. The common exceptional set and the choice of \(D\)
  39. A strict finite Fourier inequality with constant one
  40. Angular hierarchy and predicted factors
  41. Eligible masses and shifted blocks
  42. The analytic family and its block comparisons
  43. Full-root variation and amplitude bounds
  44. Spatial meshes for prediction
  45. Positive-mass motion and the locality of other shifts
  46. Variation of moving endpoint weights
  47. Amplitude motion from one good snapshot
  48. One deterministic residual sequence for each path
  49. Mass stopping and the exact path expansion
  50. Path drops and marked spatial cells
  51. Prediction and averaging on marked cells
  52. Snapshot data and denominator shrinkage
  53. Exact averaging of the prediction
  54. Comparison to the actual factor
  55. Absorbing the baseline error by blending
  56. Changing to uniform measure without a loss per level
  57. Step supremums and normalization
  58. Residual terms and integration along paths
  59. Marked residuals
  60. Unary segments and their energy telescope
  61. Common residual weights and variation
  62. The path product and its integral
  63. Interpolation, boundary averaging, and completion
  64. From the path estimate to analytic interpolation
  65. The central weight and the exponent budget
  66. Boundary averaging with the eligibility condition
  67. Recovering the original adjoint sum
  68. Choice of parameters and completion of the reductions

Short transforms and band reduction

Let \(v:\mathbb R^2\to S^1\) be a Lipschitz unit vector field. For \(0<\epsilon<a\), define the short directional Hilbert transform by \[H_{v,a}^{\epsilon}f(x) =\int_{\epsilon<|t|<a}f(x-tv(x))\,\frac{\mathrm dt}{t}.\] The direction is held fixed at the output point \(x\): the integral samples the straight line through \(x\) in direction \(v(x)\). For \(f\in\mathcal S(\mathbb R^2)\), the symmetric principal value \(H_{v,a}f=\lim_{\epsilon\downarrow0}H_{v,a}^{\epsilon}f\) exists at every point, since \(f(x-tv(x))-f(x+tv(x))=O_x(|t|)\) near zero.

Theorem 1. There are absolute constants \(a_*\in(0,1/2)\) and \(C_*<\infty\) such that every Lipschitz field \(v:\mathbb R^2\to S^1\) with \(\mathop{\mathrm{Lip}}(v)\le1\) satisfies \[\sup_{0<\epsilon<a_*} \left\|H_{v,a_*}^{\epsilon}f\right\|_{L^2(\mathbb R^2)} \le C_*\left\|f\right\|_{L^2(\mathbb R^2)} \qquad(f\in\mathcal S(\mathbb R^2)).\] Its symmetric principal value therefore satisfies \[\left\|H_{v,a_*}f\right\|_2\le C_*\left\|f\right\|_2, \qquad \big|\{x:|H_{v,a_*}f(x)|>\lambda\}\big| \le C_*^2\lambda^{-2}\left\|f\right\|_2^2 \quad(\lambda>0).\]

Both coordinates may enter the field arbitrarily subject to the Lipschitz bound. The supremum in the theorem is outside the norm, so its first assertion is a uniform operator bound for each inner truncation. The principal-value assertion is stated on Schwartz functions and gives an \(L^2\)-bounded extension by density.

The outer length has a natural scaling. If \(L_v=\mathop{\mathrm{Lip}}(v)>0\), put \(\widetilde v(X)=v(X/L_v)\) and \(\widetilde f(X)=f(X/L_v)\). Then \(\mathop{\mathrm{Lip}}(\widetilde v)=1\), and changing variables \(s=L_vt\) gives \[H_{v,a_*/L_v}^{\epsilon}f(x) =H_{\widetilde v,a_*}^{L_v\epsilon}\widetilde f(L_vx), \qquad \left\|\widetilde f\right\|_2=L_v\left\|f\right\|_2.\] The theorem thus has the same constant at outer length \(a_*/L_v\). If \(L_v=0\), the field is constant. Rotation and Fubini reduce its transform, for any finite outer length, to the one-dimensional Hilbert transform. These conventions distinguish the short-scale assertion from an estimate at arbitrary outer lengths for varying fields.

The problem and its predecessors

Stein’s question places Lipschitz regularity at the center of the planar directional singular-integral problem. In the formulation (Lacey and Li 2010, Conjecture 1.5), it asks for a uniform weak-\((2,2)\) bound with the outer length controlled by the reciprocal Lipschitz seminorm. 1 gives an affirmative answer with a strong \(L^2\) bound. The scale restriction limits the turning of the field, but it does not make the operator a Fourier multiplier: the line used by the integral still varies with the output point.

Lacey and Li trace Stein’s question to Zygmund’s conjecture on almost-everywhere differentiation by averages over shrinking straight segments in direction \(v(x)\). They also recall a Besicovitch obstruction showing that Hölder regularity of any exponent below one does not by itself give the corresponding uniform weak-\((2,2)\) estimate for the short positive averaging maximal operator (Lacey and Li 2010, Preface, Conjecture 1.2 and Chapter 2). Analytic fields admit additional structure: Bourgain proved an \(L^2\) theorem for their averaging maximal operator (Bourgain 1989), and Stein and Street proved local \(L^p\), \(1<p<\infty\), estimates for single-parameter singular Radon transforms associated with real-analytic maps (Stein and Street 2012). Their singular-integral theorem includes analytic frozen-direction Hilbert transforms, with localization depending on the analytic map. The uniform Lipschitz estimate considered here asks for constants controlled by the Lipschitz seminorm alone.

Several distinct advances explain the remaining difficulty. Lacey and Li proved an annular weak-\((2,2)\) estimate, together with strong \(L^p\) estimates for \(p>2\), for directional Hilbert transforms with measurable directions (M. T. Lacey and Li 2006, Theorem 1.1). An annular restriction isolates a frequency scale; it does not give a uniform strong \(L^2\) bound for the complete transform. For Lipschitz unit fields, Guo obtained annular weak-\((2,2)\) and strong \(L^p\), \(p>2\), bounds, uniformly in the annulus scale, for the \(r\)-variation, \(r>2\), of smooth dyadic Hilbert-transform partial sums over the short scale range determined by the reciprocal Lipschitz length; the varying dyadic cutoff is distinct from the hard inner time cutoff in 1 (Guo 2017b, Theorem 1.2 and (1.6)–(1.8) of the author version). Lacey and Li’s Lipschitz Kakeya theorem controls averages over short rectangles on which the field has a prescribed popular direction (M. Lacey and Li 2006, Theorem 1.4). For popularity \(\delta\), its weak-\(L^2\) norm costs \(C\delta^{-1/2}\). We use precisely this estimate.

The Lacey–Li memoir develops the relation between those maximal operators and the Hilbert-transform problem (Lacey and Li 2010). Its Theorem 1.18 assumes a stronger maximal estimate below \(L^2\) for the strong annular conclusion; its full-transform conclusion also requires \(C^{1+\eta}\) regularity. These are conditional results with hypotheses different from those of 1. For bounded measurable slopes depending on one coordinate, Bateman proved single-annulus \(L^p\) estimates for \(1<p<\infty\) (Bateman 2013). Bateman and Thiele proved strong \(L^p\) bounds for the full-line principal-value transform when \(3/2<p<\infty\), for every nonvanishing measurable field depending on one coordinate (Bateman and Thiele 2013, Theorem 1). The one-coordinate \(L^2\) case is already connected to Carleson’s theorem by a partial Fourier transform (Carleson 1966; Bateman and Thiele 2013); Bateman and Thiele obtain the stated wider \(p\)-range. The one-coordinate hypothesis remains a substantive restriction on the direction field. Guo extended the one-coordinate setting to fields constant along suitable Lipschitz foliations, with a uniform transversality assumption (Guo 2015, 2017a).

A different family of results combines Lipschitz control with lacunary directions. Guo and Thiele proved strong \(L^p\) bounds without a frequency restriction for short principal-value transforms associated with dyadic roundings of a positive Lipschitz slope (Guo and Thiele 2017); Di Plinio and Parissis extended this framework to finite lacunary order for fields generated by separated dyadic roundings of Lipschitz angular functions (Di Plinio and Parissis 2018). The rounded direction field can itself be discontinuous. These results exploit directional structure beyond that assumed in the present theorem.

A further reduction separates frequency reassembly from the main single-band estimate. Di Plinio, Guo, Thiele and Zorin-Kranich established that uniform strong \(L^2\) bounds on vertical frequency bands imply the full short principal-value bound under a small uniform Lipschitz bound in the vertical variable (Di Plinio et al. 2018, Corollaries 1.3–1.4 of the author version). We give a commutator proof adapted to the two truncation endpoints and to the narrow frequency bins used later. The new estimate to be proved is the logarithmically improved restricted bound on a band; the principle of band-to-full reassembly is already present in that work.

The analytic and geometric reductions

We describe the proof before fixing its technical parameters. Localization and rotation write a short piece of the operator in slope form \[H_{u,a}^{\epsilon}h(x,y) =\int_{\epsilon<|t|<a} h(x-t,y-tu(x,y))\,\frac{\mathrm dt}{t}, \qquad |u|\le1,\quad \mathop{\mathrm{Lip}}(u)\le C,\] with both input and test supported in a short horizontal strip. Writing \(\eta\) for the frequency dual to \(y\), a vertical frequency band at scale \(w\) means that \(|\eta|\) is comparable to \(w^{-1}\). For its smooth projection \(P_w\), the central restricted estimate is \[|\langle H_{u,a}^{\epsilon}P_w f,g\rangle| \lesssim \bigl(1+|\log(|F|/|E|)|\bigr)^{-3}\sqrt{|F||E|}, \qquad |f|\le\mathbf 1_F,\quad |g|\le\mathbf 1_E.\] Here \(F,E\) are sets of positive finite measure in the horizontal strip. Its constants are uniform in \(w\) and \(\epsilon\). The logarithmic factor makes amplitude-level summation converge. A Lipschitz commutator then permits Littlewood–Paley reassembly across vertical bands. These implications, including all endpoint errors, are proved in this section before the restricted estimate itself is used.

The difficult regime is \(|E|\gg|F|\). Write \(L=\log\sqrt{|E|/|F|}\). Positive operator bounds discard all but very short integration times. The vertical band is then divided into narrow intervals centered at frequencies \(\eta_j\), with common width comparable to \(W^{-1}\), where \(W=mw\) and \(m\) grows subexponentially in \(L\). The components form one finite-dimensional Hilbert space. A wave packet has horizontal length \(\ell\), transverse width \(W\), and a direction label with precision \(d=w/\ell\). Smooth packet forms retain the original field only through scalar selectors of its direction.

Ordinary size and density estimates remove the small-density packets. Each remaining packet is assigned to a top box whose direction is popular on a measured subset of the test set. The popularity witnesses are disjoint. Lacey–Li’s theorem controls a decaying weighted count of the tops, after every thin parallelogram has been enclosed by an admissible orthogonal rectangle. This step is geometric: it supplies a polynomial cost in \(L\), and does not provide the final oscillatory saving.

The phase gain and its iteration

Two statements make the oscillation usable. First, on every reference line the packet sums have polynomial moment bounds, square-function bounds and \(3\)-variation bounds. The variation argument combines Lépingle’s martingale inequality with the comparison between smooth Fourier truncations and conditional expectations developed by Jones, Seeger and Wright (Lépingle 1976; Jones et al. 2008). All three packet estimates hold for every fixed deterministic subcollection. Popularity and Lipschitz control then produce one exceptional set outside which each separated list of significant top directions in the required angular cluster and precision range has at most \(D=L^C\) members, simultaneously over angular resolutions.

Second, a finite Fourier sum with at most \(D\) separated real frequencies has a strict improvement over the usual coefficient exponent. For suitable band-limited probability measures \(\nu\), we prove \[\left\|\sum_{i=1}^d a_i e^{i\lambda_i t}\right\|_{L^p(\nu)} \le \left(\sum_{i=1}^d |a_i|^{r_*}\right)^{1/r_*}, \qquad \frac1{r_*}=1-\frac{1+c/\log(2D)}p, \quad p\ge4.\] The operator constant is exactly one, for complex coefficients and for Hilbert-valued coefficients acted on by diagonal phases. The proof separates the case of one dominant coefficient from its complement. In the latter case, separated quantiles give a strict fourth-moment deficit, which survives a small improvement of the coefficient exponent. An unspecified constant in this inequality would accumulate at every angular level and destroy the argument.

The iteration organizes directions into nested angular intervals. Each interval carries the positive mass of the tops whose direction precision is sufficiently fine for that interval; these are its eligible tops. It also carries an oscillatory amplitude assembled from its packets. The significant-direction count enters phase averaging through a weighted selection of separated block children, whose analytic weights cancel the probability cost of retaining them. We compare mass-normalized \(p\)-energies after adding the mass itself as a baseline. On an appropriate spatial interval, a finer amplitude can be predicted from one good point by diagonal phase motion. The finite Fourier estimate therefore controls the average of the predicted energy factor. The comparison errors are proportional to the relative mass lost on the path, so their sum is controlled by a logarithmic mass telescope.

Packets that stop between two angular endpoints create residual terms. Square summation handles blocks with a definite mass drop. Across nearly unary portions of a path, the energies telescope and the remaining residual sums are estimated by \(3\)-variation of one fixed sequence for that path. The sequence is chosen before the spatial point or mass stop is selected. Normalized factors on nested spatial cells can then be integrated without any cost proportional to the angular depth.

Finally we interpolate among shifted groupings of angular levels. At the central parameter the packet coefficients contain boundary weights. Averaging one common angular-grid translation removes these weights, using eligibility as well as angular proximity in the boundary telescope. With \(p=L^{1/2}\), the strict phase gain leaves a decrease of order \((\sqrt L\log L)^{-1}\) in the exponent of \(\sqrt{|E|/|F|}\). This dominates all fixed polynomial losses and supplies the logarithmic restricted estimate above.

Organization, inputs and uniformity

Section 1 gives the analytic reduction and the precise packet estimates needed for the band theorem. Section 2 proves the ordinary tile estimates and constructs the popular forest. Section 3 establishes every-line concentration and the simultaneous count of significant directions. Section 4 proves the constant-one finite Fourier inequality. Section 5 constructs the angular hierarchy and the predicted factors. Section 6 controls residuals and integrates path products. Section 7 performs interpolation, angular-grid averaging and the final strong-type assembly.

There are two specialized external inputs: Lacey–Li’s short Lipschitz-popularity maximal theorem, stated in 2, and the unweighted scalar martingale \(3\)-variation inequality of Lépingle, used in the form (Zorin-Kranich 2020, Theorem 1.1, \(p=2\), \(r=3\), weight one). The latter enters the locally proved Fourier/martingale comparison in 25; we credit its origin to (Lépingle 1976). Standard one-dimensional Hilbert-transform, maximal-function and interpolation facts are recalled where used. The commutator, tile, concentration, finite-phase and hierarchy arguments are included in full.

All packet collections, scale ranges and angular depths are finite until estimates independent of those truncations have been proved. The bin Hilbert spaces are finite-dimensional and complex, and every constant is independent of their dimension. Pairings are linear in the first variable. A factor \(L^C\) always has a fixed exponent; it never conceals the number of packets, scales or hierarchy levels.

Fourier conventions and the local reduction

We use the Fourier convention \[\widehat f(\xi,\eta) =\int_{\mathbb R^2}f(x,y)e^{-i(x\xi+y\eta)}\,\mathrm dx\,\mathrm dy, \qquad f(x,y)=\frac1{(2\pi)^2}\int_{\mathbb R^2} \widehat f(\xi,\eta)e^{i(x\xi+y\eta)}\,\mathrm d\xi\,\mathrm d\eta.\] The same convention is used in one dimension. Pairings are linear in the first entry. A constant depending on a fixed finite list of smooth cutoffs or on their prescribed seminorms is an absolute constant. No such list will depend on a frequency scale, an inner truncation, the number of packets, or the dimension of the bin Hilbert space.

A maximal theorem and elementary analytic tools

We first record precisely the specialized maximal theorem used later. If an orthogonal rectangle \(R\) has long side \(L(R)\), short side \(W(R)\), and long-side direction \(e_R\in S^1\), let \(\operatorname{EX}(R)\) be the arc centered at \(e_R\) of total length \(W(R)/L(R)\).

Theorem 2 (Lacey–Li). Let \(v:\mathbb R^2\to S^1\) be Lipschitz and \(0<\delta<1\). Restrict to rectangles satisfying \[L(R)\le \frac1{100\mathop{\mathrm{Lip}}(v)},\qquad \left|R\cap v^{-1}(\operatorname{EX}(R))\right|\ge\delta\left|R\right|,\] where the first restriction is void when \(\mathop{\mathrm{Lip}}(v)=0\). Then \[\mathcal M_{v,\delta}h(z) =\sup_R\frac{\mathbf 1_R(z)}{\left|R\right|}\int_R\left|h\right| \quad\text{satisfies}\quad \left\|\mathcal M_{v,\delta}h\right\|_{L^{2,\infty}} \le C\delta^{-1/2}\left\|h\right\|_2.\] Here the supremum is tested at every point of \(R\), and the constant is independent of \(v\) and \(\delta\).

This is (M. Lacey and Li 2006, Theorem 1.4). Requiring the field to lie in a smaller fixed subarc only restricts the collection of admissible rectangles. In the bounded slope range, the map \(t\mapsto(1,t)/\sqrt{1+t^2}\) changes slope distances into arc distances by comparable absolute factors. These observations explain the slope version of 2 used below. When an intermediate box is a parallelogram, the application of the theorem will explicitly enclose it by a comparable orthogonal rectangle of admissible length.

We also use the following familiar one-dimensional estimates, with their indicated uniformities. The Hardy–Littlewood maximal operator is bounded on \(L^p\), \(1<p<\infty\), and \[ \sup_{0<r<R<\infty} \left\|\int_{r<|t|<R}h(\,\cdot-t\,)\frac{\,\mathrm dt}{t}\right\|_p \le C_p\left\|h\right\|_p. \tag{1}\] The supremum in (1) is outside the norm; the usual maximal truncation theorem gives this assertion as well. See (Hajłasz 2012, Theorems 1.1 and 7.6 and the proof of (7.7)). Fubini gives the corresponding horizontal estimates on \(\mathbb R^2\). The \(L^2\) square-function identities used here are consequences of Plancherel and finite overlap of smooth frequency annuli. The standard smooth Littlewood–Paley and random-sign multiplier formulations are also recorded in (Tao 2006, Proposition 5.3, Corollary 5.4, and Remark 4.5). We give the particular Carleson embedding and commutator arguments needed for our reduction below. The John–Nirenberg and martingale variation estimates are stated in the sections where their local forms are used.

The local slope operator and its endpoints

Fix an absolute bound \(C_{\rm Lip}\) large enough for the slope fields obtained in the following localization; for example, a bound of ten suffices after taking the squares small enough. Choose an absolute horizontal length \(b_0>0\) with \(b_0 C_{\rm Lip}<1/10\), and subsequently choose \(0<a_*<b_0/100\). These constants may be decreased to accommodate the fixed geometric enlargements in later applications of 2. All such enlargements are chosen absolutely. For the remainder of the local argument write \(a=a_*\), and put \[ H_{u,a}^{\epsilon}h(x,y) =\int_{\epsilon<|t|<a} h(x-t,y-tu(x,y))\frac{\,\mathrm dt}{t}, \qquad 0<\epsilon<a, \tag{2}\] where \(\left|u\right|\le1\) and \(\mathop{\mathrm{Lip}}(u)\le C_{\rm Lip}\). The local inputs and tests will be supported in \(I\times\mathbb R\), where \(I\) is any interval of length at most \(b_0\).

Lemma 3. For \(\left|t\right|\,C_{\rm Lip}<1\), define \[A_t h(x,y)=h(x-t,y-tu(x,y)).\] For \(1\le p<\infty\), \[ \left\|A_t h\right\|_p \le (1-\left|t\right|\,C_{\rm Lip})^{-1/p}\left\|h\right\|_p, \tag{3}\] and \(A_t\) is a contraction on \(L^\infty\). Consequently, replacing either symmetric endpoint at scale \(r\) by a measurable endpoint between \(c r\) and \(C r\), with \(c,C>0\) fixed and \(\max(1,C)r C_{\rm Lip}<1/2\), costs a bounded \(L^p\) operator, independently of \(r\).

Proof. For each fixed \(x\), the map \(\varphi_{x,t}(y)=y-tu(x,y)\) is increasing, onto, and satisfies \[(1-\left|t\right|\,C_{\rm Lip})\left|y-y'\right| \le\left|\varphi_{x,t}(y)-\varphi_{x,t}(y')\right| \le(1+\left|t\right|\,C_{\rm Lip})\left|y-y'\right|.\] The change-of-variables formula for a one-dimensional bi-Lipschitz map, followed by translation in \(x\), proves (3). A difference of endpoint truncations is pointwise dominated by \[\int_{c' r<|t|<C' r}\left|A_t h\right|\frac{\,\mathrm dt}{|t|}\] with \(c'=\min(1,c)\) and \(C'=\max(1,C)\). Minkowski and (3) bound this by \(C_p\log(C'/c')\left\|h\right\|_p\). The same proof applies to inner and outer endpoints, and permits endpoints depending on \((x,y)\). ◻

We explain the localization of the original unit-field operator. Partition the plane into squares \(Q\) of a sufficiently small fixed side length. For output in \(Q\), only a fixed enlargement \(Q^+\) of that square contributes to a transform of length \(a_*\). Rotate coordinates so that the vector at the center of \(Q\) points along the positive first coordinate axis. Since \(\mathop{\mathrm{Lip}}(v)\le1\), on the rotated output square we have \[v=(v_1,v_2),\qquad v_1\ge c>0, \qquad u=v_2/v_1,\qquad \left|u\right|\le1, \qquad\mathop{\mathrm{Lip}}(u)\le C_{\rm Lip}.\] The last inequality follows by subtracting the two ratios and using the lower bound for \(v_1\). Extend \(u\) to the plane with the same Lipschitz bound, using \(\inf_{z\in Q}(u(z)+C_{\rm Lip}|\,\cdot-z|)\), and then clip the extension to \([-1,1]\). The infimum extension agrees with \(u\) on \(Q\), and clipping preserves both agreement and the Lipschitz bound.

At an output point, the substitution \(s=t v_1(x,y)\) changes \(\,\mathrm dt/t\) to \(\,\mathrm ds/s\) and changes the two endpoints to \(\epsilon v_1(x,y)\) and \(a_*v_1(x,y)\). By 3, each endpoint can be replaced by \(\epsilon\) or \(a_*\), respectively, at a uniform \(L^2\) cost. The rotated sets \(Q\) and \(Q^+\) fit in a common horizontal interval of length at most \(b_0\), after choosing the square side and \(a_*\) in that order. Thus a uniform bound for \(\mathbf 1_{I\times\mathbb R}H_{u,a}^{\epsilon}\mathbf 1_{I\times\mathbb R}\) implies the required bound on this output square. The output squares are disjoint, the sets \(Q^+\) have bounded overlap, and rotations preserve area. Squaring the local estimates and summing over \(Q\) therefore proves the global estimate. Notice that both endpoint errors are uniform in the inner radius \(\epsilon\).

The band theorem and summation of amplitude levels

A vertical band cutoff at length \(w>0\) is a multiplier \(P_w=p(wD_y)\), where \(p\) is smooth and supported in a fixed compact annulus \(\{r:c_\eta\le|r|\le C_\eta\}\), with \(0<c_\eta<C_\eta<\infty\) absolute. We allow a fixed bounded family of such cutoffs, including slightly enlarged cutoffs equal to one on the supports of others. All constants below are uniform when the required finitely many cutoff seminorms are bounded. The dyadic inverse frequency scale \(w\) belongs to \(2^{\mathbb Z}\).

Theorem 4 (Restricted estimate on a vertical band). For the fixed absolute constants in (2), there is an absolute \(C\) such that, for every field \(u\) with \(\left|u\right|\le1\) and \(\mathop{\mathrm{Lip}}(u)\le C_{\rm Lip}\), every \(0<\epsilon<a\), every dyadic \(w\), and every vertical band cutoff \(P_w\), \[ \left|\langle H_{u,a}^{\epsilon}P_w f,g\rangle\right| \le C\bigl(1+\left|\log(|F|/|E|)\right|\bigr)^{-3} \sqrt{|F||E|} \tag{4}\] whenever \(F,E\subset I\times\mathbb R\) have positive finite measure, \(\left|f\right|\le\mathbf 1_F\), \(\left|g\right|\le\mathbf 1_E\), and \(\left|I\right|\le b_0\).

The proof of 4 will be completed after the packet estimates and their iteration have been established. The next results show, without using that theorem as a premise elsewhere in its proof, that it implies 1. The last part of this section then gives explicit sufficient packet estimates for 4.

Lemma 5. Suppose a bilinear form \(\mathcal B\) satisfies \[\left|\mathcal B(f,g)\right| \le C_0(1+|\log(|F|/|E|)|)^{-3}\sqrt{|F||E|} \quad (|f|\le\mathbf 1_F,\ |g|\le\mathbf 1_E)\] on finite-measure sets in a fixed measurable domain. Then \[\left|\mathcal B(f,g)\right|\le C C_0\left\|f\right\|_2\left\|g\right\|_2\] for finite amplitude decompositions, and hence for all \(L^2\) inputs whenever the form is obtained by an \(L^2\)-continuous truncation or by completion from those inputs.

Proof. Use disjoint amplitude sets \(F_i=\{2^{i-1}<|f|\le2^i\}\) and write \(f=\sum_i2^i f_i\), where \(|f_i|\le\mathbf 1_{F_i}\). For the moment all sums are finite. Put \(a_i=2^i\sqrt{|F_i|}\), so \(\sum_i a_i^2\le4\left\|f\right\|_2^2\). Group the indices into measure classes \[\mathcal I_k=\{i:2^k\le|F_i|<2^{k+1}\},\qquad A_k=\left(\sum_{i\in\mathcal I_k}a_i^2\right)^{1/2}.\] If \(i_0\) is the largest index in a nonempty class, then \[\sum_{i\in\mathcal I_k}a_i \le 2^{(k+1)/2}\sum_{i\le i_0}2^i \le 2\sqrt2\,a_{i_0}\le2\sqrt2\,A_k.\] Make the analogous decomposition \(g=\sum_j2^j g_j\), with measure classes \(\mathcal J_l\) and square sums \(B_l\). For \(i\in\mathcal I_k\), \(j\in\mathcal J_l\), \[(1+|\log(|F_i|/|E_j|)|)^{-3} \le C(1+|k-l|)^{-3}.\] The assumed bound, followed by the within-class estimates, gives \[\left|\mathcal B(f,g)\right| \le C C_0\sum_{k,l}(1+|k-l|)^{-3}A_kB_l \le C C_0\left\|A\right\|_{\ell^2}\left\|B\right\|_{\ell^2}.\] The final step is Young’s inequality for the summable sequence \((1+|k|)^{-3}\). This proves the assertion. Truncating the amplitude decompositions and using the asserted continuity or completion gives the last statement, with no dependence on the number of classes. ◻

A Lipschitz commutator

The estimate below belongs to the Lipschitz commutator theory initiated by Calderón and developed by Coifman and Meyer (Calderón 1965; Coifman and Meyer 1978). We give a dyadic proof with the precise multiplier seminorms needed for Littlewood–Paley reassembly and for the later narrow-bin reduction. The argument separates low output frequencies from the annular interactions, because the constant must remain uniform as the number of frequency scales increases. Let \(\chi\) be a real even smooth cutoff equal to one on \([-1,1]\) and supported in \([-2,2]\), and define \[S_k=\chi(2^{-k}D),\qquad \Delta_k=S_k-S_{k-1}.\] Thus \(\Delta_k\) is supported in \(\{2^{k-1}\le|\zeta|\le2^{k+1}\}\), and Plancherel gives \[ \left\|h\right\|_2^2\asymp\sum_k\left\|\Delta_kh\right\|_2^2. \tag{5}\]

Lemma 6 (The required Carleson embedding). For every fixed integer \(c\), every \(q\in L^\infty(\mathbb R)\), and every \(h\in L^2(\mathbb R)\), \[ \sum_k\left\|(\Delta_kq)S_{k+c}h\right\|_2^2 \le C_c\left\|q\right\|_\infty^2\left\|h\right\|_2^2. \tag{6}\]

Proof. For a dyadic interval \(J\) of length \(2^{-j}\), split \(q=q\mathbf 1_{3J}+q\mathbf 1_{(3J)^c}\). By (5), the first part satisfies \[\sum_{k\ge j}\int_J |\Delta_k(q\mathbf 1_{3J})|^2 \le C\left\|q\right\|_\infty^2|J|.\] The kernels of \(\Delta_k\) have Schwartz decay at scale \(2^{-k}\). For \(x\in J\) and \(k\ge j\), this gives \[|\Delta_k(q\mathbf 1_{(3J)^c})(x)| \le C\left\|q\right\|_\infty 2^{j-k}.\] After summing the squares, we obtain \[ \sum_{k\ge j}\int_J|\Delta_kq|^2 \le C\left\|q\right\|_\infty^2|J|. \tag{7}\] For each dyadic interval \(I\) of length \(2^{-k}\), set \(\alpha_I=\int_I|\Delta_kq|^2\). Then (7) says \(\sum_{I\subset J}\alpha_I\le C\left\|q\right\|_\infty^2|J|\).

Write \(M\) for the Hardy–Littlewood maximal operator. Kernel decay, summed on dyadic annuli about a point of \(I\), gives \[\sup_{x\in I}|S_{k+c}h(x)| \le C_c\operatorname*{ess\,inf}_{y\in I}Mh(y).\] The shift \(c\) is fixed, so moving the kernel center across \(I\) changes the bound by only a fixed factor. If \(F\ge0\) belongs to \(L^2\), the dyadic Carleson condition implies \[\sum_I\alpha_I \bigl(\operatorname*{ess\,inf}_I F\bigr)^2 \le C\left\|q\right\|_\infty^2\left\|F\right\|_2^2.\] Indeed, for each \(t>0\), sum over the maximal dyadic intervals on which \(\operatorname*{ess\,inf}_I F>t\), apply the Carleson condition on each such interval, and integrate against \(2t\,\mathrm dt\). Their union lies, up to null sets, in \(\{F>t\}\). Apply this assertion with \(F=Mh\), and use the \(L^2\) maximal theorem. This proves (6). This is also the usual continuous Littlewood–Paley Carleson embedding; compare the proof of (Tao 2007, Proposition 4.1). ◻

Lemma 7. Let \(b\in W^{1,\infty}(\mathbb R)\), and let \(T=\mathfrak m(D)\), where \(\mathfrak m\) is smooth away from zero and \[A(\mathfrak m) =\max_{0\le r\le4}\sup_{\zeta\ne0} |\zeta|^r|\mathfrak m^{(r)}(\zeta)|<\infty.\] Then \[ \left\|[b,T]\partial h\right\|_2 \le C A(\mathfrak m)\left\|b'\right\|_\infty\left\|h\right\|_2. \tag{8}\] The commutator is initially defined on the Sobolev space \(W^{1,2}(\mathbb R)\); the estimate gives its bounded extension to \(L^2\). Its constant is independent of \(\left\|b\right\|_\infty\) and of all frequency truncations.

Proof. Put \(L_b=\left\|b'\right\|_\infty\) and \(b_k=\Delta_kb\). Cancellation and the first absolute moment of the smooth kernels give \[ \left\|b_k\right\|_\infty\le C2^{-k}L_b, \qquad \left\|\partial S_kb\right\|_\infty\le C L_b, \qquad b_k'=\Delta_kb'. \tag{9}\] Also \(\left\|S_Nb-b\right\|_\infty\le C2^{-N}L_b\). Consequently, for each fixed \(j\), the exact identity \[ b=S_{j-4}b+\sum_{k\ge j-3}b_k \tag{10}\] has a uniformly convergent high-frequency sum. This identity retains the constant and low-frequency part of \(b\); no homogeneous reconstruction of a bounded function at frequency zero is needed.

For any fixed smooth annular cutoff \(\rho\), let \(K_j\) be the kernel of \(\mathfrak m(D)\rho(2^{-j}D)\). Rescaling and three integrations by parts on its fixed compact frequency support show \[ \int_\mathbb R|z|\,|K_j(z)|\,\mathrm dz \le C A(\mathfrak m)2^{-j}. \tag{11}\] In fact the rescaled kernel is bounded by \(C A(\mathfrak m)(1+|z|)^{-3}\), which proves the assertion.

First suppose that \(\widehat h\) is smooth and compactly supported away from zero. Put \(h_j=\Delta_jh\); only finitely many \(h_j\) are nonzero. Apply (10) to each of them. There are three interactions.

For the low frequencies of \(b\), put \(B_j=S_{j-4}b\). The supports of both \(h_j\) and \(B_jh_j\) lie in a fixed annulus of radius \(2^j\). A slightly enlarged annular multiplier \(T_j\) therefore satisfies \([B_j,T]\partial h_j=[B_j,T_j]\partial h_j\). The kernel formula and (11) give \[\left\|[B_j,T_j]\partial h_j\right\|_2 \le \mathop{\mathrm{Lip}}(B_j)\int |z||K_j(z)|\,\mathrm dz\, \left\|\partial h_j\right\|_2 \le C A(\mathfrak m)L_b\left\|h_j\right\|_2.\] The outputs remain in fixed annuli of radius \(2^j\). Finite frequency overlap and (5) bound their sum by \(C A(\mathfrak m)L_b\left\|h\right\|_2\).

For \(k\ge j+4\), the output of \([b_k,T]\partial h_j\) lies in a fixed annulus of radius \(2^k\). Plancherel, Bernstein’s inequality, and (9) yield \[\left\|[b_k,T]\partial h_j\right\|_2 \le C A(\mathfrak m)L_b 2^{j-k}\left\|h_j\right\|_2.\] Group by \(k\), use frequency overlap, and apply Young’s inequality on \(\ell^2(\mathbb Z)\) to the summable sequence \(2^{-n}\mathbf 1_{n\ge4}\). This bounds the high–low sum by the same quantity and proves its convergence in \(L^2\).

It remains to consider \(|k-j|\le3\). Fix one such difference and write \(H_k=h_{k+r}\), \(|r|\le3\). The exact identity \[ [b_k,T]\partial H_k =\partial([b_k,T]H_k)-b_k'TH_k+T(b_k'H_k) \tag{12}\] treats the possible low output frequencies. Indeed, \(\partial([b_k,T]H_k)\) has spectrum in a ball of radius \(C2^k\). Its \(\Delta_i\) projection vanishes for \(i>k+c_0\), with \(c_0\) fixed, and otherwise \[\left\|\Delta_i\partial([b_k,T]H_k)\right\|_2 \le C A(\mathfrak m)L_b2^{i-k}\left\|H_k\right\|_2.\] Another application of (5) and Young’s inequality to \(2^{-n}\mathbf 1_{n\ge-c_0}\) bounds the sum of these derivative terms. Thus arbitrarily low output frequencies cause no logarithmic loss.

Choose a fixed \(c_1\) so that \(S_{k+c_1}\) equals one on the spectrum of every product \(b_k'H_k\). For \(g\in L^2\), \[\begin{split} \left|\left\langle\sum_k b_k'H_k,g\right\rangle\right| &\le \left(\sum_k\left\|H_k\right\|_2^2\right)^{1/2} \left(\sum_k\left\|(\Delta_kb')S_{k+c_1}g\right\|_2^2\right)^{1/2}\\ &\le C L_b \left(\sum_k\left\|H_k\right\|_2^2\right)^{1/2}\left\|g\right\|_2, \end{split}\] by 6. Since \(\left\|T\right\|_{2\to2}\le A(\mathfrak m)\), this bounds \(T\sum_k b_k'H_k\). Replacing \(H_k\) by \(TH_k\) in the same calculation bounds \(\sum_k b_k'TH_k\): the frequency supports are unchanged and \(\sum_k\left\|TH_k\right\|_2^2\le A(\mathfrak m)^2 \sum_k\left\|H_k\right\|_2^2\). Equation (12), summed over its seven possible differences, completes the proof on the chosen dense class.

For \(h\in W^{1,2}\), approximate by \((S_N-S_{-N})h\), followed if necessary by smooth compact Fourier approximations away from zero. The approximation converges in \(W^{1,2}\). Since \(b\) is bounded and \(T\) is \(L^2\)-bounded, the defining commutators converge in \(L^2\). The estimate passes to the limit, and then defines a bounded operator on all of \(L^2\). ◻

For a vertical multiplier \(T\), set \[C_t h(x,y)=h(x,y-tu(x,y)).\]

Corollary 8. For \(\left|t\right|\le b_0\) and \(T\) as in 7, \[ \left\|TC_t h-C_tTh\right\|_2 \le C A(\mathfrak m)\left|t\right|\,\left\|h\right\|_2. \tag{13}\] The constant is uniform over \(\left|u\right|\le1\), \(\mathop{\mathrm{Lip}}(u)\le C_{\rm Lip}\).

Proof. Work first with a smooth field and on one fixed \(x\)-fiber. Let \(\varphi_s(y)=y-su(x,y)\). Direct differentiation gives \[C_s^{-1}C_s'=b_s\partial_y, \qquad b_s=-u(x,\cdot)\circ\varphi_s^{-1},\] and hence, since \(T\) commutes with \(\partial_y\), \[ \frac{\,\mathrm d}{\,\mathrm ds}(C_sTC_s^{-1}) =C_s[b_s,T]\partial_y C_s^{-1}. \tag{14}\] The bi-Lipschitz estimates in 3 imply \[\left\|C_s\right\|_{2\to2}+\left\|C_s^{-1}\right\|_{2\to2}\le C, \qquad \mathop{\mathrm{Lip}}(b_s)\le \frac{C_{\rm Lip}}{1-|s|C_{\rm Lip}}\le C.\] Apply 7 inside (14), integrate from \(0\) to \(t\), and multiply by \(C_t\). The estimate is uniform in the fiber, so integration in \(x\) proves (13).

A bounded Lipschitz field is approximated by convolutions with smooth approximate identities, preserving the bounds on its size and Lipschitz seminorm. For each fixed small \(s\), the corresponding composition operators converge strongly on \(L^2\): this follows first for continuous compactly supported inputs by uniform convergence of the maps, and then by density and their uniform operator bounds. Passing to the limit proves the assertion for the original field. ◻

Square-function reassembly

The preceding composition estimate pays a factor of \(|t|\), which cancels the singular measure \(\,\mathrm dt/|t|\) when commutators are integrated. This is the mechanism that permits the band estimates to be assembled. Di Plinio, Guo, Thiele and Zorin-Kranich established this band-to-full reduction under Lipschitz regularity (Di Plinio et al. 2018, Corollaries 1.3–1.4 of the author version). We retain the argument here to track the fixed outer endpoint, arbitrary inner endpoint and horizontal localization used throughout this paper.

Choose a real smooth annular function \(p\) so that \(\sum_{w\in2^{\mathbb Z}}p(w\eta)^2=1\) for \(\eta\ne0\). Such a function is obtained by dividing an ordinary smooth dyadic partition by the square root of its sum of squares. Set \(P_w=p(wD_y)\). For every finite set \(\mathcal W\) of scales and every choice of signs \(\varepsilon_w\in\{-1,1\}\), the multiplier \(T=\sum_{w\in\mathcal W}\varepsilon_wP_w\) satisfies \(A(\mathfrak m_T)\le C\), by finite overlap of the annuli. Writing \(S_t^x h(x,y)=h(x-t,y)\), we have \[H_{u,a}^{\epsilon} =\int_{\epsilon<|t|<a} C_t S_t^x\frac{\,\mathrm dt}{t}.\] The vertical multiplier \(T\) commutes with \(S_t^x\). Thus 8 implies \[ \left\|[T,H_{u,a}^{\epsilon}]\right\|_{2\to2} \le C\int_{\epsilon<|t|<a}|t|\frac{\,\mathrm dt}{|t|} \le C a. \tag{15}\] All these operators are already bounded on \(L^2\) for fixed \(\epsilon>0\), with the elementary bound \(C\log(a/\epsilon)\); no conclusion of 4 is needed to form their square functions.

Randomize the signs independently and use their orthogonality. For \(f=\mathbf 1_{I\times\mathbb R}f\), Equation (15) and the triangle inequality in the random-sign \(L^2\) space give \[ \left(\sum_{w\in\mathcal W} \left\|\mathbf 1_{I\times\mathbb R}P_w H_{u,a}^{\epsilon}f\right\|_2^2 \right)^{1/2} \le C\left\|f\right\|_2+ \left(\sum_{w\in\mathcal W} \left\|\mathbf 1_{I\times\mathbb R}H_{u,a}^{\epsilon}P_w f\right\|_2^2 \right)^{1/2}. \tag{16}\] The horizontal restriction commutes with every vertical multiplier.

Proposition 9. The uniform restricted estimate in 4 implies 1.

Proof. Apply 5 to the localized band form. It gives a uniform local \(L^2\) bound for \(H_{u,a}^{\epsilon}\widetilde P_w\), where \(\widetilde P_w\) is an enlarged cutoff equal to one on the spectrum of \(P_w\). Since \(P_w f\) retains the same horizontal support, \[\left\|\mathbf 1_{I\times\mathbb R}H_{u,a}^{\epsilon}P_w f\right\|_2 =\left\|\mathbf 1_{I\times\mathbb R}H_{u,a}^{\epsilon} \widetilde P_w(P_w f)\right\|_2 \le C\left\|P_w f\right\|_2.\] Insert this in (16) and let \(\mathcal W\) increase to all dyadic scales. Plancherel and the square partition give \(\left\|\mathbf 1_{I\times\mathbb R}H_{u,a}^{\epsilon}f\right\|_2 \le C\left\|f\right\|_2\), uniformly in \(\epsilon\). The localization already proved above gives the asserted global bound for the unit field. Finally, for a Schwartz input the symmetric principal value exists pointwise; Fatou gives its \(L^2\) bound, and Chebyshev gives the weak estimate in 1. ◻

Removing long and very short times

We now reduce 4 to packet estimates. In this part only, abbreviate \(H_{u,a}^{\epsilon}\) to \(H\), and put \[L=\log\sqrt{|E|/|F|}.\] Fix a sufficiently large absolute threshold \(L_0\), which can be increased after the fixed constants in the argument are chosen. When \(L\ge L_0\), choose dyadic numbers satisfying \[ \tfrac12 e^{-L^{0.97}}<\tau\le e^{-L^{0.97}},\qquad e^{L^{0.9}}\le m<2e^{L^{0.9}},\qquad W=mw. \tag{17}\] Here \(\tau\) is the short-time cutoff. The parameter \(m\) sets the number of vertical frequency bins: their width will be comparable to \(W^{-1}\), and \(W=mw\) is the corresponding spatial width. For every fixed \(C\), the product \(m^C\tau\) tends to zero as \(L\to\infty\); this pays for the binning errors. We take \(L_0\) large enough that \(\tau<a\). For all other values of \(L\), let \(m=1\), \(W=w\), and let \(\tau\in(0,a)\) be a fixed small dyadic constant.

Lemma 10. For any fixed \(1<p<\infty\), the contribution of \(\max(\epsilon,\tau)<|t|<a\) to \(HP_w\) has \(L^p\) norm at most \(C_p(1+L)\) if \(L\ge L_0\), and at most a fixed \(C_p\) otherwise. On the remaining times below \(w\), replacing \(u\) by zero costs a uniformly bounded \(L^p\) operator. The resulting horizontal symmetric truncation is uniformly bounded on \(L^p\).

Proof. The first assertion follows from 3, the uniform \(L^1\) norm of the convolution kernel of \(P_w\), and \[\int_{\tau<|t|<a}\frac{\,\mathrm dt}{|t|} =2\log(a/\tau)\le C(1+L^{0.97}).\] For the second, the fundamental theorem of calculus gives \[\begin{split} &(P_wh)(x-t,y-tu(x,y))-(P_wh)(x-t,y)\\ &\qquad=-t u(x,y)\int_0^1 (\partial_yP_wh)(x-t,y-\theta t u(x,y))\,\mathrm d\theta. \end{split}\] The intermediate maps satisfy 3, and the kernel of \(\partial_yP_w\) has \(L^1\) norm \(O(w^{-1})\). Consequently, \[ \left\|(A_t-S_t^x)P_wh\right\|_p \le C_p |t|w^{-1}\left\|h\right\|_p. \tag{18}\] Integrating over any subinterval of \(0<|t|<w\) against \(\,\mathrm dt/|t|\) gives a constant. Equation (1) handles the horizontal part, uniformly in its inner and outer endpoints. ◻

Every error with \(L^p\) norm \(B_p\) contributes at most \[ B_p|F|^{1/p}|E|^{1-1/p} =B_p\sqrt{|F||E|} \left(\frac{|F|}{|E|}\right)^{1/p-1/2} \tag{19}\] to the restricted form. In the regime \(L\ge L_0\), choose one fixed \(p<2\), for example \(p=3/2\); the normalized error is then at most \(C(1+L)e^{-L/3}\). In the opposite regime \(|F|\ge|E|\), the fixed choice \(p=4\) gives the power \((|E|/|F|)^{1/4}\) for each uniformly bounded error. On a bounded range of \(L\), fixed operator bounds suffice. These choices are made once and do not approach the endpoint \(p=2\).

From a spatial truncation to a finite scale sum

Choose an even smooth cutoff \(\chi_0\), equal to one near zero and compactly supported, and set \[\vartheta(r)=\chi_0(r)-\chi_0(2r),\qquad \psi(r)=-i\pi\operatorname{sgn}(r)\vartheta(r).\] The function \(\psi\) is smooth and supported in a fixed compact annulus. Its inverse Fourier transform \(k\) is Schwartz, and \(k_\ell(t)=\ell^{-1}k(t/\ell)\) has multiplier \(\psi(\ell\zeta)\). Since \(\sum_{\ell\in2^{\mathbb Z}}\vartheta(\ell\zeta)=1\) for \(\zeta\ne0\), \[\sum_{\ell\in2^{\mathbb Z}}k_\ell(t)=\frac1t\qquad(t\ne0).\] The identity follows by Fourier inversion in tempered distributions; away from zero the kernel sum is locally uniformly convergent by Schwartz decay for small \(\ell\) and the bound \(C/\ell\) for large \(\ell\).

Lemma 11. For dyadic \(0<r\le R\), \[ \int_\mathbb R\left| \frac{\mathbf 1_{\{r<|t|<R\}}}{t} -\sum_{r\le\ell\le R}k_\ell(t) \right|\,\mathrm dt\le C. \tag{20}\] Also, \[ \int_\mathbb R|t|\left|\sum_{r\le\ell\le R}k_\ell(t)\right| \,\mathrm dt\le C R. \tag{21}\] Both constants are independent of the number of scales. Changing either spatial endpoint by a fixed factor preserves the first bound.

Proof. For \(r<|t|<R\), the omitted scales below \(r\) have sum bounded by \(C r/|t|^2\), using \(\left|k_\ell(t)\right|\le C\ell/|t|^2\) and a geometric sum. The omitted scales above \(R\) have sum at most \(C/R\). The integrals of these bounds over the middle interval are bounded. For \(|t|<r\), the retained sum is at most \(C\sum_{\ell\ge r}\ell^{-1}\le C/r\); for \(|t|>R\), it is at most \(C R/|t|^2\). These are integrable on their respective regions and prove (20). For the moment bound, use \[\int |t||k_\ell(t)|\,\mathrm dt =\ell\int |s||k(s)|\,\mathrm ds\] and \(\sum_{\ell\le R}\ell\le2R\). Finally, changing an endpoint by a fixed factor changes the spatial kernel only on an annulus having bounded \(\,\mathrm dt/|t|\) measure. ◻

After 10, the only possible remaining times are between \(r=\max(w,\epsilon)\) and \(\tau\). Round the lower endpoint to a comparable dyadic number; the resulting annular error has a uniform \(L^p\) bound. If this lower endpoint exceeds \(\tau\), there is no remaining model. Otherwise 11 reduces the multiplier, with \(u\) frozen at the output point, to \[ \mathfrak a_{u,\mathcal L}(x,y;\xi,\eta) =\sum_{\ell\in\mathcal L} \psi\bigl(\ell(\xi+\eta u(x,y))\bigr), \qquad \mathcal L\subset\{\ell\in2^{\mathbb Z}:w\le\ell\le\tau\}. \tag{22}\] Here \(\mathcal L\) is the consecutive set of retained scales; when \(\epsilon\le w\) it is the whole displayed range. Denote the associated operator by \(\mathcal A_{u,\mathcal L}\), or simply \(\mathcal A\). Its spatial representation is \[\mathcal A h(x,y) =\int_\mathbb RK_{\mathcal L}(t) h(x-t,y-tu(x,y))\,\mathrm dt, \qquad K_{\mathcal L}=\sum_{\ell\in\mathcal L}k_\ell.\]

The kernels in this representation are not compactly supported in \(t\). This causes no large-time change-of-variables problem: if both the input and the test are supported in \(I\times\mathbb R\), a nonzero contribution has \(t\in I-I\), hence \(|t|\le b_0\). Every vertical cutoff preserves that horizontal support. Thus 3 applies to the localized kernel error in (20), and its \(L^p\) cost is uniform. Only the original input and test carry horizontal restrictions; no spatial support is subsequently imposed on their convolutions.

Narrow vertical bins

Split the vertical cutoff into its positive and negative frequency parts. Reflection \(y\mapsto-y\) interchanges the two cases and replaces \(u\) by \(-u\), preserving all assumptions. We work with positive frequencies, so \(c_\eta\le w\eta\le C_\eta\).

Suppose first that \(L\ge L_0\) and \(m\) is as in (17). Choose real smooth cutoffs \(Q_j\), \(1\le j\le J\le C m\), supported in intervals of length \(C/W\) centered at \(\eta_j\asymp w^{-1}\), with bounded overlap and \[\sum_{j=1}^J Q_j(\eta)^2=1 \quad\text{on }\mathop{\mathrm{supp}}p(w\eta).\] One obtains them from a fixed smooth partition at spacing \(W^{-1}\) by dividing by the square root of its sum of squares on a slightly enlarged band. Their normalized profiles and all fixed derivatives are uniform. In particular, their convolution kernels have bounded \(L^1\) norms, and \[ A(Q_j)\le C m^4. \tag{23}\] Let \(\widetilde Q_j\) be a slightly enlarged cutoff equal to one on the spectrum of \(Q_jP_w\). These cutoffs again have the stated bin width and normalized smoothness bounds.

Set \(F_j^0=Q_jP_wf\) and \(G_j^0=Q_jg\). The identity on the band and self-adjointness of the \(Q_j\) give \[ \langle\mathcal A P_wf,g\rangle =\sum_j\langle\mathcal A\widetilde Q_jF_j^0,G_j^0\rangle +\sum_j\langle[\mathcal A,Q_j]F_j^0,g\rangle. \tag{24}\] Indeed, expand \(P_w=\sum_jQ_j^2P_w\), commute one \(Q_j\) past \(\mathcal A\), and use \(\widetilde Q_jF_j^0=F_j^0\). This identity does not require \(\mathcal A\) to preserve individual bins.

On localized inputs and tests, Equations (13), (21), and (23) imply \[\left|\langle[\mathcal A,Q_j]F_j^0,g\rangle\right| \le C m^4\tau\left\|F_j^0\right\|_2\left\|g\right\|_2.\] In applying the composition estimate, only \(|t|\le b_0\) are tested, as explained above. Bounded frequency overlap gives \(\sum_j\left\|F_j^0\right\|_2^2\le C\left\|f\right\|_2^2\), so the total error in (24) is at most \[ C m^5\tau\left\|f\right\|_2\left\|g\right\|_2. \tag{25}\] The harmless integer exponent five includes the factor \(\sqrt J\). More generally, any fixed multiplier seminorm choice would give \(C m^C\tau\), and for every fixed \(C,N\), \[ m^C\tau \le C\exp\bigl(C' L^{0.9}-L^{0.97}\bigr) =O(L^{-N}). \tag{26}\] The large-\(L\) threshold in this statement is absolute after the fixed exponents are chosen.

Let \(\mathcal H=\mathbb C^J\) with its Euclidean norm, and regard \(F^0=(F_j^0)_j\), \(G^0=(G_j^0)_j\) as \(\mathcal H\)-valued functions. For some fixed power \(Q=Cm^C\), their norms obey \[ \begin{aligned} \left\|F^0\right\|_{L^\infty(\mathcal H)}&\le Q,& \left\|F^0\right\|_{L^1(\mathcal H)}&\le Q|F|,& \left\|F^0\right\|_{L^2(\mathcal H)}^2&\le C|F|,\\ \left\|G^0\right\|_{L^\infty(\mathcal H)}&\le Q,& \left\|G^0\right\|_{L^1(\mathcal H)}&\le Q|E|,& \left\|G^0\right\|_{L^2(\mathcal H)}^2&\le C|E|. \end{aligned} \tag{27}\] For the first two columns, apply the uniform scalar kernel bounds and sum over at most \(Cm\) components. For the last column, use Plancherel and bounded overlap. No compact support in \(y\) is asserted after the projections. Horizontal support in \(I\) is preserved.

When \(m=1\), we do not perform this commutation or project the test. Instead retain the original positive-frequency band cutoff on the input and use one scalar component. It fits in a vertical frequency interval centered at a fixed \(\eta_1\asymp w^{-1}\) of length \(C/w=C/W\). Its positive annular support is retained.

Lemma 12 (Amplitude levels after bin projection). Let \(h:\mathbb R^2\to\mathcal H\) satisfy \[\left\|h\right\|_\infty\le Q,\qquad \left\|h\right\|_1\le Q M_0,\qquad \left\|h\right\|_2^2\le C_0M_0,\] where \(M_0>0\) and \(Q\le C\exp(C L^{0.9})\). Write \[A_i=\{2^{i-1}<\left\|h(\cdot)\right\|_{\mathcal H}\le2^i\}, \qquad a_i=2^i\sqrt{|A_i|/M_0}.\] For sufficiently large absolute \(L\), \[ \sum_i a_i\le C\sqrt{1+L},\qquad \sum_{\{i:\,|\log(|A_i|/M_0)|>L/4\}}a_i \le C e^{-cL}. \tag{28}\] Empty level sets are omitted. The constants depend on the fixed constants in the hypotheses, not on the dimension of \(\mathcal H\).

Proof. The amplitude upper bound implies \(2^i\le2Q\) on every nonempty level. If \(|A_i|<e^{-L/4}M_0\), then \(a_i\le2^i e^{-L/8}\), and summing the geometric series in \(i\) gives \(C Qe^{-L/8}\). If \(|A_i|>e^{L/4}M_0\), the \(L^1\) bound gives \[2^i\le2Qe^{-L/4},\qquad a_i^2=4^i|A_i|/M_0\le2Q\,2^i.\] Summing the square-root geometric series over the allowed amplitudes again gives \(C Qe^{-L/8}\). Since \(\log Q=O(L^{0.9})\), both tails are \(O(e^{-cL})\).

The remaining level sets occupy at most \(C(1+L)\) dyadic measure classes. Within a class, the geometric estimate in the proof of 5 bounds the sum of \(a_i\) by a constant times its square sum. Cauchy–Schwarz over the classes and \(\sum_i a_i^2\le4C_0\) give the first bound in (28). The already bounded tails can be included in that bound. ◻

Apply 12 to the two vectors in (27). Decomposing them into amplitude levels produces vector inputs and tests bounded by indicators of new sets. The extra cutoffs \(\widetilde Q_j\) in (24) remain part of the operator after this decomposition; thus no band support is being assumed for an individual amplitude slice. If both new sets have retained measures, then their logarithmic disparity is in \[ \left[\frac34L,\frac54L\right] \subset\left[\frac12L,2L\right]. \tag{29}\] Indeed, the two changes of logarithmic measure are each at most \(L/4\), and the logarithm of the square-root ratio uses one half of their difference. The product of the two sums of normalized amplitude weights is \(O(1+L)\).

The admissible packet forms

We next describe the packet model with all geometric and analytic uniformities that will be used. Fix \(w\in2^{\mathbb Z}\), \(m\ge1\), \(W=mw\), a finite dyadic scale set \(\mathcal L\subset[w,\tau]\), and at most \(Cm\) bin centers \(\eta_j\asymp w^{-1}\). The scalar case \(m=1\) just described is included. Constants in this definition are fixed throughout.

Definition 13. A tile \(s\) has a length \(\ell_s\in\mathcal L\), a bounded slope \(\theta_s\), a center \(z_s=(x_s,y_s)\), and precision \(d_s=w/\ell_s\). Its box is \[R_s=\left\{(x,y): |x-x_s|\le\frac{\ell_s}{2},\quad |(y-y_s)-\theta_s(x-x_s)|\le\frac W2\right\}, \qquad |R_s|=\ell_s W.\] A \(t\)-box is defined by the same formula with slope \(t\). The nominal horizontal length need not be the Euclidean long side when \(W>\ell_s\).

At a fixed scale, slope labels have bounded multiplicity on every interval of length \(d_s\). For each scale and label, the centers lie in one translated box lattice \[z_0+\{(n\ell_s,\theta_s n\ell_s+kW):n,k\in\mathbb Z\}.\] The lattice and center of a tile are common to every bin component. An admissible collection is any finite subcollection of such lattices.

Its analyzing and synthesizing waves, denoted \(\Phi_{s,j}\) and \(\Psi_{s,j}\), are bounded-normalized modulated Schwartz functions adapted to \(R_s\). More precisely, each can be written as \[e^{i\eta_j((y-y_s)-\theta_s(x-x_s))} \phi_{s,j}\left( \frac{x-x_s}{\ell_s}, \frac{(y-y_s)-\theta_s(x-x_s)}{W}\right),\] where, for all fixed nonnegative integers \(a,b,N\), \[ |\partial_X^a\partial_Y^b\phi_{s,j}(X,Y)| \le C_{a,b,N}(1+|X|+|Y|)^{-N}. \tag{30}\] The constants are uniform in \(s,j,w,m\). Both waves also have Fourier support in \[ |\eta-\eta_j|\le C W^{-1},\qquad c_\eta'\le w\eta\le C_\eta',\qquad |\xi+\theta_s\eta|\le c_*\ell_s^{-1}, \tag{31}\] with fixed positive \(c_\eta',C_\eta'\). The constant \(c_*\) is chosen sufficiently small compared with the selector separation constants below and the fixed annular support of \(\psi\).

The scalar selector \(b_{s,j}\) is smooth, with \[ \mathop{\mathrm{supp}}b_{s,j}\subset \{t:c_b d_s\le|t-\theta_s|\le C_b d_s\}, \qquad d_s^r\left\|b_{s,j}^{(r)}\right\|_\infty\le C_r \tag{32}\] for every fixed \(r\ge0\), where \(0<c_b<C_b\) are absolute. Finally, the collection uses one angular dyadic grid. Each tile has an interval \(\Omega_s\) in that grid, of length \(C_\Omega d_s\), where \(C_\Omega\) is a fixed large dyadic number, containing both the wave slope support \(\{-\xi/\eta:(\xi,\eta)\in\mathop{\mathrm{supp}}\widehat\Phi_{s,j} \cup\mathop{\mathrm{supp}}\widehat\Psi_{s,j}\}\) and every selector support, with a margin at least \(c_\Omega d_s\), where \(c_\Omega>0\) is fixed. The interval is common to all components. Its level depends only on \(\ell_s\).

The associated form is \[ \Lambda_{\mathcal S}(f,g) =\sum_{s\in\mathcal S}|R_s|^{-1}\sum_j \langle f_j,\Phi_{s,j}\rangle \langle b_{s,j}(u)\Psi_{s,j},g_j\rangle. \tag{33}\] It is linear in \(f\) and conjugate-linear in \(g\). The definition makes sense for an arbitrary measurable bounded field \(u\). In particular, the ordinary scalar and vector estimates in the next section do not require spatial Lipschitz regularity of \(u\). The stronger estimates later in the proof use the Lipschitz hypothesis from (2).

Proposition 14 (Packet representation). The diagonal form \[\sum_j\langle\mathcal A_{u,\mathcal L} \widetilde Q_j f_j,g_j\rangle\] is a sum, with absolutely summable scalar coefficients, of averages of forms (33). Each form satisfies 13, uniformly in \(w,m,\mathcal L\), and has one common angular grid and common spatial lattices across components. The same assertion holds for \(m=1\) with the original positive-frequency band cutoff in place of \(\widetilde Q_1\). The sums can be obtained as limits of finite packet collections.

Proof. We give the construction, including the normalization and the separation of the different series forms. At each length \(\ell\), partition the frequency slope \(-\xi/\eta\) smoothly into intervals of width \(c\,d\), \(d=w/\ell\), with centers \(\theta\) on a lattice of spacing a fixed small multiple of \(d\). Take \(c\) so small that on every such frequency piece \(\left|\ell(\xi+\theta\eta)\right|\le c_*/4\). Introduce normalized coordinates \[ X'=\ell(\xi+\theta\eta),\qquad Y'=W(\eta-\eta_j),\qquad U'=(u-\theta)/d. \tag{34}\] Then \[ \ell(\xi+u\eta)=X'+(w\eta)U',\qquad w\eta=w\eta_j+Y'/m. \tag{35}\] The vertical cutoff restricts \(Y'\) to a fixed compact interval and \(w\eta\) to a fixed positive annulus. On the support of \(\psi\), its argument is bounded away from zero. In view of \(\left|X'\right|\le c_*/4\), this forces \(c_b'\le|U'|\le C_b'\), with fixed positive bounds. Enlarge these bounds slightly, retaining separation from zero. Since \(d\le1\) and \(\left|u\right|\le1\), only bounded values of \(\theta\) can contribute.

In the coordinates (34), the product of the annular multiplier, the slope partition, and the vertical cutoff is smooth on a fixed compact coordinate region, with all fixed derivatives bounded uniformly. For example the slope partition is a smooth function of \(-X'/(w\eta)\), and (35) shows that its derivatives are uniform because \(w\eta\) stays away from zero and \(m\ge1\). Choose slightly enlarged smooth plateaus in \(X',Y',U'\), with the \(X'\)-plateau supported in \(|X'|<c_*\), the \(U'\)-plateau supported away from zero, and an additional vertical plateau keeping \(w\eta\) in a positive annulus. Fourier series on a larger fixed coordinate box now give \[\sum_{\nu\in\mathbb Z^3} a_{\nu,\ell,\theta,j} A_{\nu,\ell,\theta,j}(\xi,\eta) B_{\nu,\ell,\theta,j}((u-\theta)/d),\] where \(A_\nu\) is a localized exponential in \((X',Y')\), \(B_\nu\) is a localized exponential in \(U'\), and for every fixed \(N\), \[ \sup_{\ell,\theta,j,w,m}|a_{\nu,\ell,\theta,j}| \le C_N(1+|\nu|)^{-N}. \tag{36}\] This follows by integration by parts on the fixed coordinate box. Multiplication by the plateaus restores the original symbol exactly; their supports give the strict spectral and selector restrictions in 13.

Factor the frequency function \(A_\nu\) into two smooth frequency functions by inserting a plateau equal to one on its support. Both factors retain \(|X'|<c_*\), a vertical interval of length \(C/W\), and a positive annular vertical support. Their inverse Fourier transforms are modulated Schwartz kernels at the scales \(\ell,W\). Their fixed seminorms, and those of the selector factor, grow at most polynomially in \(|\nu|\). The coefficients in (36) absorb these losses uniformly. One can do this simultaneously for every fixed seminorm as follows. Let \(A_\nu^*=\sup_{\ell,\theta,j,w,m}|a_{\nu,\ell,\theta,j}|\) and omit zero values. Extract the common coefficient \(c_\nu=(A_\nu^*)^{1/4}\), multiply each of the two frequency factors by \((A_\nu^*)^{1/4}\), and multiply the selector factor by \(a_{\nu,\ell,\theta,j}/(A_\nu^*)^{3/4}\). The product is unchanged. Each of the three factors now has uniform fixed seminorms, by (36), and \(\sum_\nu c_\nu<\infty\). In particular, no derivative order in this construction grows with \(L\) or a truncation parameter.

For clarity, consider one such factored term \(b(u)A(D)B(D)\). Write \(K_A,K_B\) for its two convolution kernels and \(|R|=\ell W\). At an intermediate convolution center \(z\), put \[\Psi_z(x)=|R|K_A(x-z),\qquad \Phi_z(y)=|R|\overline{K_B(z-y)}.\] Then the contribution to the bilinear form is exactly \[ \frac1{|R|^2}\int_{\mathbb R^2} \langle f,\Phi_z\rangle \langle b(u)\Psi_z,g\rangle\,\mathrm dz. \tag{37}\] The conjugation in \(\Phi_z\) keeps its Fourier support in the same positive-frequency region as that of \(\Psi_z\). Both waves have the carrier \(\eta_j(-\theta,1)\) and satisfy (30)–(31).

Let \(\mathcal Q\) be a fundamental cell of the box lattice, of area \(|R|\). For any integrable function \(F\), \[\int_{\mathbb R^2}F(z)\,\mathrm dz =\int_{\mathcal Q}\sum_{n,k\in\mathbb Z} F\bigl(z+(n\ell,\theta n\ell+kW)\bigr)\,\mathrm dz.\] Applying this identity to (37) and normalizing the average over \(\mathcal Q\) leaves precisely the factor \(|R|^{-1}\) in (33). Use the same translation for every \(j\) at this scale and slope. There are only finitely many lengths and relevant slope labels, so a finite product of these normalized translation measures supplies a common average for their complete sum.

We also need one angular grid per form. The union of the wave slope supports and all selector supports for a label \(\theta\) is contained in \([\theta-Cd,\theta+Cd]\), with \(C\) fixed and independent of \(j\). Choose a sufficiently large fixed dyadic multiple \(H d\). Among three adjacent dyadic grids, this interval is contained in a cell of length \(H d\) with a margin comparable to \(d\). One may see this directly from the three grids \[\bigl\{2^{-k}([0,1)+n+(-1)^k\alpha):n,k\in\mathbb Z\bigr\}, \qquad \alpha\in\{0,1/3,2/3\}:\] they are nested, and their boundary positions at any fixed scale are spaced by one third of that scale. Taking \(H\) large gives the claimed containment and margin. This is the usual adjacent-grid principle; see (Lerner and Nazarov 2019, Theorem 3.1 and Remark 3.2). Assign each label to one suitable grid and split the sum into the three resulting forms. Since \(d=w/\ell\) is dyadic, the chosen cell level depends only on \(\ell\).

At a fixed length and label, the lattice sums converge absolutely in a bilinear pairing of \(L^2\) functions. Indeed, envelope decay and Cauchy–Schwarz give \[\sum_z |R|^{-1}|\langle f,\Phi_z\rangle|^2 \le C\left\|f\right\|_2^2, \qquad \sum_z |R|^{-1}|\langle b_z(u)\Psi_z,g\rangle|^2 \le C\left\|g\right\|_2^2.\] For example, bound each squared pairing by \(C|R|\) times the integral of \(|f|^2\) against a decaying box envelope, and then sum those envelopes over the lattice. Their sum is uniformly bounded. The same argument works with the bounded selector and with vector components, by a further Cauchy–Schwarz inequality in \(j\). Since the length and label sets are finite, the continuous representation is justified by these bounds and the summability in \(\nu\), first on regular inputs and then on \(L^2\). Exhausting each lattice by finite subsets gives the claimed finite packet approximations.

For each fixed series index and angular grid, all scales and labels are kept in one form before its estimate is applied. Different series indices remain separate forms and are summed only using \(\sum_\nu c_\nu<\infty\). In particular, their copies of the same spatial lattice are never combined into a collection with unbounded multiplicity. This completes the representation. ◻

The precise remaining model estimates

The following proposition states the forward dependency used to finish the band theorem. It also records exactly how much saving survives the bin and amplitude reductions.

Proposition 15. Suppose that the following bounds hold, uniformly for every admissible finite form \(\Lambda_{\mathcal S}\) in 13, with the fixed packet constants there. The field \(u:\mathbb R^2\to\mathbb R\) is bounded and measurable, the sets \(F,E\subset\mathbb R^2\) are measurable with positive finite measure, and the input and test satisfy \(\left\|f(z)\right\|_{\mathcal H}\le\mathbf 1_F(z)\), \(\left\|g(z)\right\|_{\mathcal H}\le\mathbf 1_E(z)\).

  1. For \(m\ge1\), there is a fixed exponent \(C_1\) such that \[ |\Lambda_{\mathcal S}(f,g)| \le C(1+\log(2m))^{C_1}\sqrt{|F||E|}. \tag{38}\]

  2. For \(L\ge L_0\), with \(\tau,m,W\) as in (17), assume in addition that \(\left|u\right|\le1\), \(\mathop{\mathrm{Lip}}(u)\le C_{\rm Lip}\), and \(F,E\subset I\times\mathbb R\) for an interval \(I\) with \(\left|I\right|\le b_0\). Whenever \(L/2\le\log\sqrt{|E|/|F|}\le2L\), \[ |\Lambda_{\mathcal S}(f,g)| \le C L^{-20}\sqrt{|F||E|}. \tag{39}\]

  3. For \(m=1\), one scalar component, and \(|E|\le|F|\), there is a fixed \(c>0\) such that \[ |\Lambda_{\mathcal S}(f,g)| \le C\sqrt{|F||E|}(|E|/|F|)^c. \tag{40}\]

Then 4 holds. The constants remain uniform in \(w\), \(\epsilon\), and the number of retained scales.

Proof. We apply the assumed bounds to the local field and sets of 4, for which \(\left|u\right|\le1\), \(\mathop{\mathrm{Lip}}(u)\le C_{\rm Lip}\), and \(F,E\subset I\times\mathbb R\) with \(\left|I\right|\le b_0\). Suppose first that \(L=\log\sqrt{|E|/|F|}\ge L_0\). The time and scale reductions give exponentially small normalized errors by (19). The bin error is smaller than any fixed power of \(L^{-1}\), by (26). It remains to estimate the diagonal form in (24).

Apply the amplitude decompositions from 12 to \(F^0,G^0\). The projections defining these vectors preserve horizontal support. After omitting levels of measure zero, their level sets have positive finite measure by (27) and lie in \(I\times\mathbb R\); the normalized slices satisfy the stated indicator bounds. First use finite decompositions. Each pair of slices has a packet representation by 14; its cutoff is \(\widetilde Q_j\), so the slice itself need not have narrow spectrum. Summing the common series coefficients and averaging the translations preserve each assumed model bound. Pairs for which both measures are retained satisfy (29); apply (39) to them. Their total contribution, divided by \(\sqrt{|F||E|}\), is at most \[C L^{-20}(1+L)\le C L^{-19}.\] If either measure is discarded, use (38). The sum of those normalized weights is at most \(C\sqrt{1+L}\,e^{-cL}+Ce^{-2cL}\); the factor \((1+\log(2m))^{C_1}\) is a fixed power of \(L\). These terms are therefore exponentially small. The estimates are uniform for the finite amplitude sums. Their tails converge absolutely by the same bounds, so they apply to the full diagonal form. All errors and the main term are bounded by \(C(1+L)^{-3}\sqrt{|F||E|}\), as required.

If \(L\) ranges over a fixed bounded interval, use \(m=1\). No bin projections or their amplitude losses are needed. The time and scale errors have bounded operator norms, and (38), with the scalar packet representation, bounds the remaining form. The factor \((1+|\log(|F|/|E|)|)^{-3}\) is bounded below by a positive absolute constant on this interval.

Finally suppose that \(L\) is large and negative. Again use \(m=1\), retain the original scalar input and test, and apply (40). It gives a positive power of \(|E|/|F|\). The errors give the power \((|E|/|F|)^{1/4}\) by the fixed choice \(p=4\) in (19). Every fixed positive power of this ratio is bounded by a constant times \((1+|\log(|F|/|E|)|)^{-3}\). This proves the remaining case.

All arguments include an arbitrary positive inner cutoff. If \(\epsilon>w\), the lower retained scale is raised to a dyadic scale comparable to \(\epsilon\), which simply selects a finite subcollection of the allowed scales. If it exceeds \(\tau\), the middle form is empty and the error estimates suffice. If \(w\ge\tau\), the short-time and horizontal bounds already cover the whole remaining integral. None of the bounds used here depends on how many scales lie between the endpoints. ◻

The basic and reverse-disparity assertions of 15 will follow from the ordinary tile estimates, in particular 16. The remaining large-positive-disparity assertion is the improvement proved by the popular-forest argument and assembled in 51. Thus the conditional use of 4 in 9 is separate from, and subsequent to, the proof of its model hypotheses.

Ordinary tiles and the popular forest

The purpose of this section is to separate the part of the packet form controlled by ordinary size and density estimates from the part requiring phase averaging. The size–density–tree decomposition follows the time-frequency framework of Lacey and Thiele for the Carleson operator (Lacey and Thiele 2000) and its planar directional adaptation by Lacey and Li (M. T. Lacey and Li 2006; Lacey and Li 2010). We prove the estimates in the present normalization, including arbitrary subcollections and the Hilbert-valued summation needed for narrow vertical bins.

For the ordinary estimates, write the packet form (33) as \(\Lambda_{\mathcal S}(f,g)\), where \(\mathcal S\) is a finite collection from one of the forms constructed in the preceding section. The constants in these estimates depend on fixed packet adaptation, separation, and multiplicity constants only. In particular, they do not depend on the number of tiles or lengths, on \(w\), or on \(m=W/w\geq1\). These ordinary estimates hold for every measurable real-valued slope function \(u\).

Density and the model estimates

Fix one of the angular dyadic grids. Denote by \(\Omega(t,l)\) the interval in this grid at the level corresponding to the dyadic length \(l\) that contains \(t\). The intervals are half open. Their length is a fixed constant times \(w/l\); in particular, \(\Omega(t,\ell_s)=\Omega_s\) whenever \(t\in\Omega_s\).

For a box \(R\) of length \(l\), width \(W\), slope \(t\), and center \(z_R=(x_R,y_R)\), put \[ B_R(x,y)= \left(1+\frac{|x-x_R|}{l} +\frac{|(y-y_R)-t(x-x_R)|}{W}\right)^{-D_0}. \tag{41}\] Here \(D_0\) is a sufficiently large fixed integer, chosen once below. For a measurable set \(E\) of finite measure define \[ D_E(s)= \sup_{\substack{\ell_s\leq l\leq\tau,\ l\ \text{dyadic}\\ t\in\Omega_s,\ z_s\in R}} \frac{1}{|R|} \int_{E\cap\{u\in\Omega(t,l)\}} B_R(z)\,\,\mathrm dz. \tag{42}\] The boxes in this supremum have the dimensions and slope used in (41). Since \(\int B_R\leq C|R|\), all densities are bounded by a fixed constant. Density is monotone in \(E\); passing to a subcollection of tiles does not change the density of any remaining tile.

Proposition 16 (Ordinary model estimates). Let the bin Hilbert space be \(\mathcal H=\mathbb C^J\), where \(J\leq C_{\rm bin}m\). If \(\left\|g(z)\right\|_{\mathcal H}\leq\mathbf 1_E(z)\) and \(D_E(s)\leq\gamma\) for every \(s\in\mathcal S\), with \(0<\gamma\leq C\), then \[ |\Lambda_{\mathcal S}(f,g)| \leq C(1+\log(2m))^C\gamma^{1/4} |E|^{1/2}\left\|f\right\|_{L^2(\mathcal H)}. \tag{43}\] In the scalar case, if \(|f|\leq\mathbf 1_F\), \(|g|\leq\mathbf 1_E\), and \(0<|E|<|F|<\infty\), then \[ |\Lambda_{\mathcal S}(f,g)| \leq C\sqrt{|F||E|} \left(\frac{|E|}{|F|}\right)^{9/20}. \tag{44}\] Both conclusions hold uniformly for every subcollection of \(\mathcal S\).

We first prove the scalar estimates, suppressing the bin index \(j\); size and density selection and the tree bound provide the scalar model bounds. The final component summation gives the stated logarithmic vector loss in \(m\).

Maximal functions and separated synthesis

For a slope \(t\), let \(\mathcal M_t\) be the uncentered strong maximal operator in the coordinates \((x,y-tx)\). We shall use \[ \left\|\mathcal M_t h\right\|_2\leq C\left\|h\right\|_2, \tag{45}\] with a constant independent of \(t\). Here and below maximal operators act on the absolute value of their argument. To recall a proof, the one-dimensional uncentered maximal operator \(M\) is of weak type \((1,1)\): a finite interval covering of a compact subset of \(\{Mh>\lambda\}\) has a disjoint subfamily whose threefold enlargements cover that subset, giving measure at most \(3\lambda^{-1}\left\|h\right\|_1\). Approximation gives the full inequality. Apply it to \(h\mathbf 1_{\{|h|>\lambda/2\}}\), since the complementary part has maximal function at most \(\lambda/2\), and integrate the distribution estimate against \(2\lambda\,\,\mathrm d\lambda\). Fubini’s theorem gives \(\left\|Mh\right\|_2^2\leq C\left\|h\right\|_2^2\). The strong maximal operator is bounded pointwise by the composition of the two one-dimensional maximal operators. Fubini and the determinant-one shear prove (45).

A consequence we use repeatedly is smooth reproduction on a box. Suppose a Schwartz function, after removal of a fixed carrier, has Fourier support \[|\xi+t\eta|\leq A/r,\qquad |\eta|\leq A/W.\] A product of smooth low-pass cutoffs that equals one on this set reproduces the function. Its kernel is bounded, for any fixed \(N\), by \[\frac{C_{A,N}}{rW} \left(1+\frac{|x|}{r}\right)^{-N} \left(1+\frac{|y-tx|}{W}\right)^{-N}.\] Decomposing this majorant into dilated rectangles proves the following statement. If \(Q\) is an \(r\times W\) box of slope \(t\), and \(F_\alpha\) is any finite family with these common spectral bounds, then \[ \sup_{z\in Q}\sup_\alpha |F_\alpha(z)| \leq C_A\inf_{z\in Q} \mathcal M_t\bigl(\sup_\alpha|F_\alpha|\bigr)(z). \tag{46}\] Indeed, replacing the center of any kernel rectangle by any other point of \(Q\) changes its required enlargement by only an absolute factor. A common carrier has modulus one and does not affect this argument.

A test tree has parameters \(T=(z_T,t,l)\), where \(l\leq\tau\) is dyadic, and a box \(R_T\) of length \(l\), width \(W\), and slope \(t\). Its tiles satisfy \[ \ell_s\leq l,\qquad z_s\in R_T,\qquad t\in\Omega_s. \tag{47}\] An arbitrary subset satisfying these conditions is also a tree. Choose a fixed \(c_1>0\) so that \(c_*\ll c_1\ll\min(c_b,c_\Omega)\), where \(c_b d_s\) is the lower distance of the selector support from \(\theta_s\). A tile is separated in this tree if \(|t-\theta_s|\geq c_1d_s\).

For a positive integer \(J_0\) define the top polynomial \[ \mathfrak p_{T,J_0}(z)= \left(1+\frac{(x-x_T)^2}{l^2} +\frac{((y-y_T)-t(x-x_T))^2}{W^2}\right)^{J_0}. \tag{48}\]

Lemma 17 (Separated synthesis). Let \(\mathcal T\) be separated tiles satisfying (47), and let \(\Upsilon_s\) be either of their packet families \(\Phi_s\) or \(\Psi_s\). For every fixed \(J_0\), \[\begin{align*} \left\|\mathfrak p_{T,J_0} \sum_{s\in\mathcal T}c_s\Upsilon_s\right\|_2 &\leq C_{J_0} \left(\sum_{s\in\mathcal T}|c_s|^2|R_s|\right)^{1/2}, \tag{49}\\ \left\|\mathfrak p_{T,J_0} \sup_{a<b}\left| \sum_{\substack{s\in\mathcal T\\a<\ell_s\leq b}} c_s\Upsilon_s\right|\right\|_2 &\leq C_{J_0} \left(\sum_{s\in\mathcal T}|c_s|^2|R_s|\right)^{1/2}. \tag{50}\end{align*}\] The same statements hold for Hilbert-valued sums \(\bigl(\sum_s c_{s,j}\Upsilon_{s,j}\bigr)_j\), with \(\sum_s\left\|c_s\right\|_{\mathcal H}^2|R_s|\) on the right and constants independent of \(J\). Bounded scalar factors depending on \(s,j\), such as \(b_{s,j}(t)\), can be incorporated into the coefficients.

Proof. From the packet support and \(t\in\Omega_s\), \[\xi+t\eta=(\xi+\theta_s\eta)+(t-\theta_s)\eta, \qquad \frac{a_0}{\ell_s}\leq|\xi+t\eta| \leq\frac{a_1}{\ell_s},\] where \(0<a_0<a_1<\infty\) are fixed. Positivity of \(a_0\) uses \(c_*\ll c_1\) and the two fixed bounds on \(w\eta\). For envelopes, replacing slope \(\theta_s\) by \(t\) changes the transverse coordinate over one tile length by at most \(C d_s\ell_s=Cw\leq CW\). Thus every fixed Schwartz envelope remains a uniform envelope in the \(t\)-coordinates.

At one length there are boundedly many labels, since all labels of the tree lie in an interval of length \(O(d_s)\). Within each label the centers lie on a translated box lattice. The inner product of two normalized packets at this length is bounded by a fixed rapidly decreasing function of their normalized center separation. Its row and column sums are bounded. The elementary Schur inequality, obtained by applying \(2|a_sa_q|\leq |a_s|^2+|a_q|^2\) to these row and column sums, proves the fixed-length synthesis bound.

This argument also applies after multiplying packets by \(\mathfrak p_{T,J_0}\). Their centers belong to \(R_T\), their lengths are at most \(l\), and their widths are \(W\). Consequently the polynomial can be absorbed into a fixed larger number of packet decay powers. Multiplication by a polynomial differentiates the Fourier transform and does not enlarge its support.

Choose a fixed integer \(q_0\) such that \(2^{q_0}a_0>4a_1\), and split dyadic lengths according to their exponents modulo \(q_0\). Within each resulting list, the displayed annuli are disjoint and have gaps. Orthogonality in the first \(t\)-coordinate therefore proves (49), also for the polynomially weighted packets.

On one such list, the sum over all lengths above a given endpoint is exactly a smooth low-pass convolution in that coordinate applied to the complete sum: choose the cutoff to equal one on the lower-frequency annuli retained at that endpoint and zero on all higher-frequency annuli. The gap permits dilates of one fixed smooth cutoff. Their absolute kernels are dominated by the one-dimensional maximal operator. An interval of lengths is a difference of two endpoint sums. Apply this observation to the full weighted sum, use its just-proved \(L^2\) bound, and sum over the fixed number of lists. This gives (50).

For Hilbert-valued sums, the fixed-length and orthogonality bounds follow by summing the scalar squared bounds over components. The same low-pass convolution acts on every component, since \(a_0,a_1\) are common. Its norm is pointwise bounded by the positive kernel acting on the Hilbert norm of the full sum. The maximal-function proof is therefore dimension free. Alternatively one can sum the squared scalar maximal estimates. ◻

For scalar coefficients \(v_s=\langle f,\Phi_s\rangle/|R_s|\), define their size on a collection by \[ \operatorname{size}(v)= \sup_T \left(\frac{1}{|R_T|} \sum_{\substack{s\in T\\|t-\theta_s|\geq c_1d_s}} |v_s|^2|R_s|\right)^{1/2}. \tag{51}\] A one-tile test gives \(|v_s|\leq C\operatorname{size}(v)\). If \(\Phi_s=0\), then \(v_s=0\). Otherwise choose \((\xi,\eta)\in\mathop{\mathrm{supp}}\widehat\Phi_s\) and put \(q=-\xi/\eta\). The angular margin and \(c_1\ll c_\Omega\) place both \(q-2c_1d_s\) and \(q+2c_1d_s\) in \(\Omega_s\). At least one is at distance at least \(2c_1d_s\) from \(\theta_s\); use it as the direction of a top centered at \(z_s\) with length \(\ell_s\). This is a separated one-tile test of area \(|R_s|\), which proves the bound. Moreover \[ \operatorname{size}(v)\leq C\left\|f\right\|_\infty. \tag{52}\] To prove this, for the separated portion of a fixed tree put \(B_T=\sum_s|v_s|^2|R_s|\) and \(A_T=\sum_s v_s\Phi_s\). The pairing is linear in its first argument, so \(B_T=\langle f,A_T\rangle\). By 17 and Cauchy–Schwarz with \(\mathfrak p_{T,1}^{-1}\), \[B_T\leq\left\|f\right\|_\infty\left\|A_T\right\|_1 \leq C\left\|f\right\|_\infty |R_T|^{1/2}B_T^{1/2}.\] Division, with the zero case understood, proves (52).

Removing large size

Lemma 18 (Size selection). Suppose \(\operatorname{size}(v)\leq C_*\sigma\), with \(C_*\) fixed. One can remove a collection of trees so that the remaining size is at most \(\sigma\) and \[ \sum_T|R_T|\leq C\sigma^{-2}\left\|f\right\|_2^2. \tag{53}\] The removed tiles partition into ordinary trees satisfying (47); their total top area has the same bound.

Proof. First consider positive separated tests, with \(t-\theta_s\geq c_1d_s\). While a test has energy larger than \(\sigma^2|R_T|/2\), choose one whose direction is within \(\varepsilon\) of the infimum of the directions of all such tests. Here \(\varepsilon>0\) is a sufficiently small fixed multiple of \(\min_{s\in\mathcal S}d_s\). Delete all remaining tiles with \(\ell_s\leq l_T\), \(t_T\in\Omega_s\), and \(z_s\in A_0R_T\), including nonseparated tiles. The fixed absolute dilation \(A_0\) is chosen below. Record the positive separated test tiles present at selection as \(\mathcal Q_T\). The procedure terminates because every selection deletes at least one tile. Thereafter perform the negative separated selection, using directions within \(\varepsilon\) of the supremum and reversing the order. At the end, each signed energy is at most \(\sigma^2|R_T|/2\); hence the size is at most \(\sigma\).

We prove the area bound for the positive selection. The \(\mathcal Q_T\) are disjoint and \[ B:=\sum_T\sum_{s\in\mathcal Q_T}|v_s|^2|R_s| \asymp \sigma^2\sum_T|R_T|. \tag{54}\] The lower comparison is the selection threshold and the upper one follows from the initial size bound, which persists on every remaining subcollection. We claim \[ \left\|\sum_T\sum_{s\in\mathcal Q_T}v_s\Phi_s\right\|_2^2 \leq CB. \tag{55}\]

Expand the squared norm. Pairs with lengths in a fixed ratio have total contribution \(O(B)\). Indeed, spectral overlap forces their labels to lie within \(O(d_s)\), so there are boundedly many neighboring labels. Their box shapes are comparable because the transverse discrepancy over either length is \(O(w)\leq O(W)\). Packet decay and the spatial lattices then give the bounded row and column sums for the normalized Gram matrix, exactly as in the fixed-length proof above.

For the other pairs fix the top \(T\) of the longer member \(q\in\mathcal Q_T\), and write \(t=t_T\). Let \(\mathcal A_T\) consist of the selected tiles \(s\) for which \(\ell_q>C_0\ell_s\) and \(\langle\Phi_s,\Phi_q\rangle\ne0\) for at least one such \(q\). Choose the fixed \(C_0\) sufficiently large. Spectral overlap and the common vertical frequency range give \[|\theta_s-\theta_q|\leq Cc_*(d_s+d_q),\qquad |\theta_q-t|\leq C_\Omega d_q.\] Thus, by the choices of \(c_*\) and \(C_0\), \[ |\theta_s-t|\leq \tfrac14 c_1d_s, \qquad t\in\Omega_s. \tag{56}\] The second assertion uses the fixed margin about the wave slope support inside \(\Omega_s\).

We next verify a bounded overlap fact for the boxes of \(\mathcal A_T\) in the coordinates of slope \(t\). These boxes can all be replaced by fixed comparable \(\ell_s\times W\) boxes in those coordinates. Suppose two of the latter overlap, with \(\ell_s\) much smaller than \(\ell_{s'}\). Let \(t_s,t_{s'}\) be the directions of their respective selecting tests. Because \(s\) was in a positive separated test, \[t_s\geq\theta_s+c_1d_s>t+\tfrac12c_1d_s, \qquad |t_{s'}-t|\leq C d_{s'}.\] After separating the lengths by another fixed factor, \(t_{s'}<t+\tfrac14c_1d_s\). If the test selecting \(s'\) had not yet been selected when \(s\) was selected, all its eventual test tiles would still have been present at that time. It would therefore already have been failing, since deleting tiles only decreases each fixed test energy. The near-infimum rule, with \(\varepsilon\ll c_1\min d_s\), therefore forces the test of \(s'\) to have been selected first.

Its length is at least \(\ell_{s'}\). Overlap of the comparable \(t\)-boxes bounds the horizontal center separation by \(C\ell_{s'}\). In the coordinates of \(t_{s'}\), the transverse center separation is \(O(W+w)=O(W)\): use \(|t_{s'}-t|\ell_{s'}=O(w)\), or equivalently compare each original tile over its own length. Also \(t_{s'}\in\Omega_s\), by the margin in (56). For an absolute \(A_0\) this places \(s\) in the enlarged deletion tree of \(s'\), a contradiction. At comparable lengths, bounded label and lattice multiplicity give bounded overlap directly. Splitting lengths into a fixed number of lists proves the asserted overlap bound for all of \(\mathcal A_T\). The same reasoning applies to any fixed enlargement of these boxes by increasing \(A_0\) once.

Let \(B_s^*\) be sufficiently rapidly decreasing uniform tile envelopes in the \(t\)-coordinates, and put \(D_T=\mathfrak p_{T,J_0}^{-1}\), with \(J_0\) a fixed large integer. Then \[ \left\|D_T\sum_{s\in\mathcal A_T}B_s^*\right\|_2 \leq C|R_T|^{1/2}. \tag{57}\] Here centers of \(\mathcal A_T\) need not belong to \(R_T\). To prove the estimate, test against \(\left\|h\right\|_2\leq1\). Since \(\ell_s\leq l_T\), the top weight can be passed across a tile envelope by spending fixed decay powers. Decomposition into dilated tile boxes, followed by the overlap just proved, gives \[\begin{align*} \sum_s\int |h|D_T B_s^* &\leq C\sum_s |R_s| \inf_{R_s^t}D_T\inf_{R_s^t}\mathcal M_t h\\ &\leq C\int D_T\mathcal M_t h \leq C\left\|D_T\right\|_2\left\|h\right\|_2 \leq C|R_T|^{1/2}. \end{align*}\] The notation \(R_s^t\) denotes the fixed comparable \(t\)-box; changing it by an absolute dilation is harmless. This proves (57).

For each \(s\in\mathcal A_T\), the sum of its cross terms with \(\mathcal Q_T\) can include all \(q\in\mathcal Q_T\) with \(\ell_q>C_0\ell_s\), since the added inner products are zero. Therefore its absolute value is bounded using \[U_T(z)=\sup_{r>0} \left|\sum_{\substack{q\in\mathcal Q_T\\\ell_q>r}} v_q\Phi_q(z)\right|.\] The coefficient bound \(|v_s|\leq C\sigma\), (57), and 17 show that all such cross terms for this top have sum at most \[C\sigma\int\left(\sum_{s\in\mathcal A_T}B_s^*\right)U_T \leq C\sigma |R_T|^{1/2}\left\|D_T^{-1}U_T\right\|_2 \leq C\sigma^2|R_T|.\] Summing and using (54) proves (55). Pairing its synthesis with \(f\) gives \(B\leq\left\|f\right\|_2(CB)^{1/2}\), hence \(B\leq C\left\|f\right\|_2^2\). The negative selection has the identical proof with order reversed.

Finally, partition each enlarged deletion box \(A_0R_T\) into a fixed number of translated ordinary top boxes of length \(l_T\) and width \(W\), assigning tiles by their centers. Their direction and length conditions are unchanged, and their total area is \(O(|R_T|)\). This proves the last assertion and (53). ◻

Removing large density

Lemma 19 (Density selection). For \(0<\eta\leq C\), all tiles with \(D_E(s)>\eta\) can be assigned to keeper boxes \(R_T\), of dyadic length \(l_T\leq\tau\), width \(W\), and slope \(t_T\), with the following properties. There are pairwise disjoint measurable sets \(E_T\) such that \[\begin{gather*} E_T\subset E\cap K_\eta R_T \cap\{u\in\Omega(t_T,l_T)\},\qquad |E_T|\geq c\eta|R_T|, \tag{58}\\ \sum_T|R_T|\leq C\eta^{-1}|E|,\qquad K_\eta\leq C(1+\eta^{-1/100}), \tag{59}\\ \ell_s\leq l_T,\qquad t_T\in\Omega_s,\qquad z_s\in C K_\eta R_T \quad\text{for each tile assigned to }T. \tag{60}\end{gather*}\] After partitioning by centers, the assigned tiles form ordinary test trees of total top area at most \[ C\eta^{-1}(1+\eta^{-1/20})|E|. \tag{61}\]

Proof. For each tile to be removed select one box in (42) witnessing an average greater than \(\eta\). Since \[\int_{\mathbb R^2\setminus KR}B_R\leq C K^{2-D_0}|R|,\] choosing the fixed \(D_0\) sufficiently large and then \(K=K_\eta\leq C(1+\eta^{-1/100})\) retains at least \(\eta|R|/2\) of the weighted test inside \(KR\). Its unweighted tested set has measure at least \(\eta|R|/2\), because \(B_R\leq1\).

Process these finitely many witnesses in decreasing length order. Keep a witness unless both its \(K\)-dilated spatial box and its dyadic direction interval intersect the corresponding objects of an earlier keeper. Assign it to any such keeper when it is rejected. The tested sets of the keepers are disjoint: a common tested point would force both intersections and preclude keeping the later one. This proves (58) and (59).

Let a rejected witness have length \(l'\), direction \(t'\), and original associated tile \(s\), and let its keeper have length \(l_T\geq l'\) and direction \(t_T\). Dyadic nesting gives \[\Omega(t_T,l_T)\subset\Omega(t',l')\subset\Omega_s, \qquad |t_T-t'|\leq Cw/l'.\] Choose a point common to the two dilated spatial boxes. Moving from it to the rejected box center, or to \(z_s\), costs horizontally \(O(Kl')\) and transversely in the \(t_T\)-coordinates \(O(KW+K l'|t_T-t'|)=O(KW)\). This proves (60). A keeper’s own associated tiles satisfy the same conclusion.

Partition \(CKR_T\) into \(O(K^2)\) translated boxes of length \(l_T\), width \(W\), and slope \(t_T\), assigning tiles by their centers. The length and angular requirements for a test tree persist. Their total area is at most \(C(1+\eta^{-1/50})\eta^{-1}|E|\), which is bounded by (61) after changing the absolute constant. ◻

The tree estimate

Lemma 20 (Tree bound). Suppose a scalar tree \(\mathcal T\) satisfies \(D_E(s)\leq\eta\) on its tiles and has size at most \(\sigma\). If \(|g|\leq\mathbf 1_E\), then \[ |\Lambda_{\mathcal T}(f,g)| \leq C\sigma\eta|R_T|. \tag{62}\] The size assumption can be replaced by the same bound on any encompassing collection.

Proof. The assertion is immediate for an empty tree, so assume the tree is nonempty. Work in centered coordinates \((x,y-tx)\) of the top, so its coordinate extents are \(l\times W\). For clarity write the transverse coordinate as \(y'\), with top center \((0,0)\). All tile centers satisfy \(|x_s|\leq l/2\) and \(|y_s'|\leq W/2\). The packet envelopes are uniform in these coordinates.

Partition the horizontal line into maximal dyadic intervals \(I\) for which there is no tile with \(x_s\in3I\) and \(\ell_s\leq |I|\). The tiles are finite and have positive minimum length, so these intervals cover the line up to endpoints and are pairwise disjoint. Each parent fails the condition, furnishing a tile \(q_I\) with \[ \ell_{q_I}\leq2|I|,\qquad \mathop{\mathrm{dist}}(x_{q_I},I)\leq C|I|. \tag{63}\] Fix a sufficiently large absolute dyadic constant \(C_{\mathrm W}\). We treat lengths \(\ell_s\leq C_{\mathrm W}|I|\) and longer lengths separately on the strip above \(I\).

For a short tile, density tested at its own box and wave decay give, for every fixed large \(N_0\), \[ \int_{E\cap\{x\in I\}} |v_s b_s(u)\Psi_s| \leq C\sigma\eta|R_s| \left(1+\frac{\mathop{\mathrm{dist}}(x_s,I)}{\ell_s}\right)^{-N_0}. \tag{64}\] Indeed the selector support lies in \(\Omega_s\), the density definition includes \(R_s\) with direction \(\theta_s\), and a fixed number of additional wave-decay powers extracts the last factor while retaining the density weight.

Here is an explicit summation of (64). At a fixed length there are \(O(1)\) labels, and each horizontal column contains \(O(1)\) centers within the transverse top strip. If \(\ell_s\leq r:=|I|\), the defining property of \(I\) gives \(\mathop{\mathrm{dist}}(x_s,I)\geq r\). Summing over the one-dimensional center lattice then bounds its contribution, divided by \(C\sigma\eta W\), by \(C r(\ell_s/r)^{N_0-1}\). For \(r<\ell_s\leq C_{\mathrm W}r\), the same lattice sum is \(O(r)\) after multiplication by \(\ell_s\), and there are only a fixed number of such lengths. Thus the total is \(O(\sigma\eta Wr)\). All intervals with \(r\leq l\) lie within a fixed enlargement of the top’s horizontal interval, by (63); their disjoint lengths sum to \(O(l)\). If \(r>l\), all tile lengths are at most \(l\), all centers are at distance at least \(r\) from \(I\), and their restriction \(|x_s|\leq l/2\) improves the preceding sum to \(C\sigma\eta Wl(l/r)^{N_0-1}\). For each such dyadic \(r\), only \(O(1)\) intervals can occur, again by the parent witness. Summing proves a total short-tile contribution \(C\sigma\eta lW\).

Now let \(I\) admit longer tiles. Choose a dyadic \(l'\asymp |I|\), greater than \(\ell_{q_I}\) and small enough that \(l'<\ell_s\) for every longer tile. Increasing \(C_{\mathrm W}\) once ensures this. It also ensures \(l'\leq\tau\). By (63), a box of slope \(t\), length \(l'\), and width \(W\) can contain \(z_{q_I}\) and the portion \(I\times[-W/2,W/2]\), with a fixed enlargement absorbed into the choice of \(l'\). Dyadic nesting puts every longer selector support inside \(\Omega(t,l')\). Consequently the original longer-tile sum vanishes outside \(A_I=\{u\in\Omega(t,l')\}\). We perform all following freezing and error decompositions after multiplying that sum by \(\mathbf 1_{A_I}\), and retain this indicator on both the main term and the error. Individually the frozen terms need not vanish outside \(A_I\). For the cells \[Q_{I,k}=I\times[(k-\tfrac12)W,(k+\tfrac12)W], \qquad k\in\mathbb Z,\] the density test at \(q_I\) therefore yields \[ \mu_{I,k}:= |Q_{I,k}\cap E\cap\{u\in\Omega(t,l')\}| \leq C\eta(1+|k|)^{D_1}|I|W, \tag{65}\] where the fixed \(D_1\) can be taken to be \(D_0\). The density weight on the cell is bounded below by a constant times \((1+|k|)^{-D_0}\).

Put \(a=|u-t|\) at the point of evaluation. For nonseparated tiles, selector support implies \(a\asymp d_s\), so only a fixed number of dyadic lengths occur. Their pointwise sum is bounded by \(C\sigma(1+|k|)^{-N_1}\) on \(Q_{I,k}\), for any fixed \(N_1\). For separated tiles replace \(b_s(u)\) by \(\mathbf 1_{\{d_s\geq a\}}b_s(t)\). If \(d_s\geq a\), the error is at most \(Ca/d_s\), using the derivative bound on \(b_s\). If \(d_s<a\), the actual selector vanishes unless \(d_s\geq a/C\), since \(|t-\theta_s|\leq C_\Omega d_s\). There are only a fixed number of these remaining scales. Summing the geometric series in \(a/d_s\), and using lattice overlap at each scale, bounds the entire error by \(C\sigma(1+|k|)^{-N_1}\). This argument also covers \(a=0\), when the separated error is zero.

The resulting main sum, at each point, is an interval sum of the separated packets with coefficients \(v_sb_s(t)\): its lengths obey \(C_{\mathrm W}|I|<\ell_s\leq w/a\). Let \(M_{I,k}\) be the supremum on \(Q_{I,k}\) of the absolute values of all such interval sums with arbitrary upper endpoint. We claim that, for every fixed sufficiently large \(N_2\), \[ \sum_{I,k} M_{I,k}^2(1+|k|)^{N_2}|I|W \leq C\sigma^2|R_T|, \tag{66}\] where only the \(I\) admitting longer tiles occur. To see this, let \(U\) be the maximal interval sum over all separated lengths of the tree. For each fixed \(I\), every interval sum under consideration has horizontal \(t\)-frequency at most \(C/|I|\) and vertical frequency, after the common carrier is removed, at most \(C/W\). The same is true after multiplication by any top polynomial. Apply (46) to the weighted sums and dominate their absolute values by \(\mathfrak p_{T,J_0}U\). Since the cells are disjoint, the squared cell suprema are bounded in sum by \(C\left\|\mathcal M_t (\mathfrak p_{T,J_0}U)\right\|_2^2\). For all the present cells \(|x|\leq Cl\), and \(\mathfrak p_{T,J_0}\asymp (1+|k|)^{2J_0}\) there. Taking \(4J_0\geq N_2\), (45) and 17 prove (66).

The intervals admitting longer tiles have total length \(O(l)\), by their parent witnesses and \(|I|<l/C_{\mathrm W}\). Choose \(N_2>2D_1+2\). Factoring the full density bound before Cauchy–Schwarz, \[\begin{align*} \sum_{I,k}M_{I,k}\mu_{I,k} &\leq C\eta \left(\sum_{I,k}M_{I,k}^2 (1+|k|)^{N_2}|I|W\right)^{1/2}\\ &\hspace{1.2cm}\cdot \left(\sum_{I,k}(1+|k|)^{2D_1-N_2}|I|W\right)^{1/2} \leq C\sigma\eta|R_T|. \end{align*}\] The previously bounded pointwise errors have the same total bound by (65), taking \(N_1>D_1+2\). Adding the short contribution proves (62). ◻

Proof of the model estimates

Proof of 16. We first prove the scalar part of (43). If \(\left\|f\right\|_2=0\) or \(|E|=0\), the form is zero; assume otherwise. Process thresholds \(\eta=2^{-n}\) in decreasing order, starting high enough that no density or size condition fails. At each step first remove densities exceeding \(\eta\) by 19, then lower size to \[\sigma_\eta=\sqrt\eta\, \left\|f\right\|_2/|E|^{1/2}\] by 18. At the start of the step, the remaining collection has density at most \(2\eta\), as well as the original bound \(\gamma\), and size at most \(\sqrt2\sigma_\eta\). These bounds persist under all removals within the step.

The density removals have total ordinary top area at most \(C(1+\eta^{-1/20})\eta^{-1}|E|\), and the size removals at most \(C\sigma_\eta^{-2}\left\|f\right\|_2^2 =C\eta^{-1}|E|\). Applying 20 to every resulting ordinary tree and summing bounds the form by \[ C\left\|f\right\|_2|E|^{1/2} \sum_{\eta\ \text{dyadic}} \min(\gamma,2\eta)\sqrt\eta\, (1+\eta^{-1/20})\eta^{-1}. \tag{67}\] There is no limiting remainder to estimate for a finite collection. Every positive-density tile is eventually removed. If \(D_E(s)=0\), testing at \(R_s\) shows that \(b_s(u)g=0\) almost everywhere, since the weight is strictly positive; such a term is zero. Zero-size terms are zero as well.

For \(0<\gamma\leq1\), the terms with \(\eta\leq\gamma\) sum to \(C\gamma^{9/20}\), and the terms with \(\eta\geq\gamma\) do also: the relevant powers are \(\eta^{1/2},\eta^{9/20}\) below the crossover and \(\gamma\eta^{-1/2}, \gamma\eta^{-11/20}\) above it. For bounded larger \(\gamma\) the sum is uniformly bounded. This proves the asserted scalar \(\gamma^{1/4}\) estimate, with room in its exponent.

For (44), put \(r=|E|/|F|<1\). In addition to the threshold size estimate, every tree has size at most a fixed constant by (52). Its actual size at step \(\eta\) is therefore at most \(C\min(1,\sqrt{\eta/r})\). No change in the removal algorithm is necessary. With the universal density bound, the sum of the tree contributions at that step is at most \[C|E|\min(1,\sqrt{\eta/r}) \min(1,\eta)\eta^{-1} (1+\eta^{-1/20}).\] For \(r\leq\eta\leq1\), this sums to \(C|E|r^{-1/20}\); the scales \(\eta>1\) cost \(C|E|\). For \(\eta\leq r\), the sum is at most \[C|E|r^{-1/2} \sum_{\eta\leq r} (\eta^{1/2}+\eta^{9/20}) \leq C|E|r^{-1/20} =C\sqrt{|F||E|}\,r^{9/20}.\] This proves the reverse-disparity claim, including the density-area inflation from the selection.

Finally consider \(\mathcal H=\mathbb C^J\). For a scalar measurable \(h\) supported in \(E\), write \[\left\|h\right\|_{2,1}^{\#} =\int_0^\infty |\{|h|>a\}|^{1/2}\,\,\mathrm da.\] Layer-cake decomposition, retaining the complex phase of \(h\), extends the scalar estimate to \[|\Lambda_j(f_j,g_j)| \leq C\gamma^{1/4}\left\|f_j\right\|_2 \left\|g_j\right\|_{2,1}^{\#}.\] Every sliced support lies in \(E\), so its density still has upper bound \(\gamma\). As \(|g_j|\leq1\), split the last integral at \(a_0=(2+m)^{-2}\). The portion below \(a_0\) is at most \(a_0|E|^{1/2}\); for the rest, Cauchy–Schwarz gives \[\int_{a_0}^1|\{|g_j|>a\}|^{1/2}\,\,\mathrm da \leq \left(\int_{a_0}^1 a|\{|g_j|>a\}|\,\,\mathrm da\right)^{1/2} \left(\int_{a_0}^1\frac{\,\mathrm da}{a}\right)^{1/2} \leq C\sqrt{1+\log(2m)}\left\|g_j\right\|_2.\] Since \(J\leq C_{\rm bin}m\) and \(\sum_j\left\|g_j\right\|_2^2\leq|E|\), \[\sum_j(\left\|g_j\right\|_{2,1}^{\#})^2 \leq C(1+\log(2m))|E|.\] Cauchy–Schwarz in \(j\) proves (43), even with a square root of this logarithmic factor. All arguments used only the stated packet hypotheses, which persist on every subcollection. Every choice of decay or derivative order above is fixed before the parameters and the finite tile collection are chosen. This completes the proof of 16. ◻

For later parameter choices we record a fixed numerical loss available from this proof. In the regime \(m\asymp\exp(L^{0.9})\), with \(L\) sufficiently large, the last vector estimate implies \[ |\Lambda_{\mathcal S}(f,g)| \leq C L^{b_4}\gamma^{c_4}|E|^{1/2} \left\|f\right\|_{L^2(\mathcal H)}, \qquad b_4=1,\quad c_4=\tfrac14. \tag{68}\] These exponents are fixed independently of all subsequent density, forest, and mask choices.

The improvement problem and popular tops

We now turn to the improvement estimate (39) in the local Lipschitz setting: \(|u|\le1\), \(\mathop{\mathrm{Lip}}(u)\le C_{\rm Lip}\), and the sets \(F,E\) have positive finite measure and lie in the horizontal testing strip \(\mathcal D=I\times\mathbb R\), where \(|I|\le b_0\). Write \(\mathcal H\) for the finite-dimensional Hilbert space of vertical bins; the input and test satisfy \(\left\|f(z)\right\|_{\mathcal H}\le\mathbf 1_F(z)\) and \(\left\|g(z)\right\|_{\mathcal H}\le\mathbf 1_E(z)\). We use the parameters in (17): \(L\) is sufficiently large, \(m\asymp\exp(L^{0.9})\), \(W=mw\), and every tile length is at most the dyadic cutoff \(\tau\le\exp(-L^{0.97})\). All tile collections in this part are finite. Every estimate is independent of the number of tiles, tops, scales, and bins.

The adjoint sum and exceptional subsets

Only the retained tiles, with the assignment \(T(s)\), are used henceforth. With pairings linear in the first entry, define \[ a_{s,j}=|R_s|^{-1} \langle g_j,b_{s,j}(u)\Psi_{s,j}\rangle, \qquad \phi_s=(a_{s,j}\Phi_{s,j})_j, \qquad P=\sum_s\phi_s. \tag{71}\] These conventions give precisely the form (33): \[\langle f,P\rangle =\sum_s|R_s|^{-1}\sum_j \langle f_j,\Phi_{s,j}\rangle \langle b_{s,j}(u)\Psi_{s,j},g_j\rangle.\] Indeed \(\langle f_j,a_{s,j}\Phi_{s,j}\rangle =\overline{a_{s,j}}\langle f_j,\Phi_{s,j}\rangle\). For fixed \(g\) satisfying \(\left\|g(z)\right\|_{\mathcal H}\le\mathbf 1_E(z)\), the supremum of the absolute value of this pairing over \(\left\|f(z)\right\|_{\mathcal H}\le\mathbf 1_F(z)\) equals \(\int_F\left\|P(z)\right\|_{\mathcal H}\,\mathrm dz\). Equality follows by taking \(f=P/\left\|P\right\|_{\mathcal H}\) on \(F\cap\{P\ne0\}\), and zero elsewhere.

Put \(N=\sqrt{|E|/|F|}\). The remaining objective, to be proved by the subsequent estimates, is \[ \int_F\left\|P(z)\right\|_{\mathcal H}\,\mathrm dz \le C L^{-20}N|F|, \qquad \frac L2\le\log N\le2L. \tag{72}\] Applying (68) with its absolute density upper bound and arbitrary \(L^2\) input, and then taking the Hilbert-space dual norm, gives the fixed exponent \(C_P=b_4=1\) such that \[ \left\|P\right\|_{L^2(\mathbb R^2;\mathcal H)} \le C L^{C_P}\sqrt{|E|}. \tag{73}\] Consequently, for every measurable \(Z\subset F\), \[ \int_Z\left\|P(z)\right\|_{\mathcal H}\,\mathrm dz \le C L^{C_P} \left(\frac{|Z|}{|F|}\right)^{1/2}N|F|. \tag{74}\] For any prescribed fixed \(Q_0>0\), an exceptional subset with \(|Z|\le L^{-2(C_P+Q_0+1)}|F|\) therefore costs at most \(C L^{-Q_0}N|F|\). This estimate is applied to the original sum \(P\) before any masks or other modifications. It does not require an \(L^2\) bound for those modified sums.

A weak \(L^2\) estimate for the top count

For a top \(T\), use normalized coordinates \[ X_T(z)=\frac{x-x_T}{l_T},\qquad Y_T(z)=\frac{(y-y_T)-s_T(x-x_T)}{W},\qquad \mathop{\mathrm{dist}}_T(z)=|X_T(z)|+|Y_T(z)|. \tag{75}\] Fix a sufficiently large absolute exponent \(D_2>4\). For now let \(K_0\ge1\) be any fixed power of \(L\); the choice used with the masks is specified below. Set \(C_3=2^{D_2}\) and \[ h_T(z)=C_3\left(1+\frac{\mathop{\mathrm{dist}}_T(z)}{K_0}\right)^{-D_2}, \qquad n(z)=\sum_{T\in\mathcal T}h_T(z). \tag{76}\] In particular, \(h_T\ge1\) on \(\{\mathop{\mathrm{dist}}_T\le K_0\}\). We use the convention \[\left\|a\right\|_{L^{2,\infty}(\mathcal D)} =\sup_{\lambda>0}\lambda |\{z\in\mathcal D:|a(z)|>\lambda\}|^{1/2}.\]

Proposition 21. For the tops in (70), \[ \left\|n\right\|_{L^{2,\infty}(\mathcal D)} \le C L^C\sqrt S. \tag{77}\] More precisely, if \[D_*\ge\max(2K_0,\,2L^{C_{\mathrm{as}}},\,2),\] then one may take the right-hand side to be \(C L^{C_{\mathrm{as}}}D_*^3\sqrt S\), provided \(D_*\) is chosen within an absolute factor of the displayed maximum. In particular, when \(K_0=L^k\), a permissible exponent in (77) is \[ c_{\mathrm{cnt}}(k) =C_{\mathrm{as}}+3\max(k,C_{\mathrm{as}}). \tag{78}\]

Once the fixed power \(K_0\) is chosen, this estimate together with \(S\le L^{C_{\mathrm{as}}}N^2|F|\) makes the part of \(F\) where \(n\) exceeds \(N\) times a sufficiently large fixed power of \(L\) small enough for (74).

Proof. If \(S=0\), the assertion is immediate, so assume \(S>0\). We first state the maximal-function input in the form used here. For a Lipschitz unit vector field \(v_0\), let \(M_{v_0,\delta}\) be the supremum of averages over Euclidean rectangles containing the evaluation point, with long side at most \(c_{\mathrm{LL}}/\mathop{\mathrm{Lip}}(v_0)\), for which the field belongs to the rectangle’s direction uncertainty arc on at least a \(\delta\) fraction of the rectangle. The arc is centered at the long-side direction and has total length equal to the short-to-long side ratio, as in 2. 2, namely (M. Lacey and Li 2006, Theorem 1.4), gives \[ \left\|M_{v_0,\delta}a\right\|_{2,\infty} \le C\delta^{-1/2}\left\|a\right\|_2, \qquad 0<\delta<1. \tag{79}\] The same estimate holds if the admissible rectangles are required to satisfy a smaller fixed uncertainty interval, since this restricts the family of rectangles. When \(\mathop{\mathrm{Lip}}(v_0)=0\), the length restriction is understood as absent. We apply this result to \(v_0=(1,u)/\sqrt{1+u^2}\), whose Lipschitz constant is bounded by an absolute constant. We also use the ordinary uncentered strong maximal operator \(M_{\mathrm{str}}\) over axis-parallel rectangles, which is bounded on \(L^2(\mathbb R^2)\).

Let \(Q\subset\mathcal D\) have finite measure, let \(D\ge D_*\), and put \[a_T(D)=\frac{|Q\cap DR_T|}{|DR_T|} =\frac{|Q\cap DR_T|}{D^2|R_T|}.\] Here and below dilation is about the box center in its slope coordinates. We prove, for \(0<\beta\le1\), that \[ \sum_{T:a_T(D)>\beta}|R_T| \le \min\left(S,\, C L^{2C_{\mathrm{as}}}D^2\beta^{-2}|Q|\right). \tag{80}\]

Fix a top occurring on the left-hand side. Set \[J_T=I\cap[x_T-Dl_T/2,x_T+Dl_T/2],\qquad a=|J_T|, \qquad b=C_{\mathrm{w}}DW,\] where \(C_{\mathrm{w}}\ge1\) is a fixed absolute constant. The witness set is contained in \(DR_T\cap\mathcal D\), so \(a>0\). The slope parallelogram \[\widetilde R_T=\{(x,y):x\in J_T,\quad |(y-y_T)-s_T(x-x_T)|\le b/2\}\] contains both \(E_T\) and \(Q\cap DR_T\). It satisfies \[ a\le\min(|I|,Dl_T),\qquad ab\le C_{\mathrm{w}}D^2|R_T|,\qquad \frac ba\ge C_{\mathrm{w}}\frac W{l_T} =C_{\mathrm{w}}m d_T. \tag{81}\] Thus the possibly very long horizontal dilation has been clipped before any Euclidean enclosing rectangle is constructed.

Suppose first that \(b/a\le\varepsilon_0\), for a fixed small absolute \(\varepsilon_0\). Projection onto the orthonormal directions \[e_T=\frac{(1,s_T)}{\sqrt{1+s_T^2}},\qquad e_T^\perp=\frac{(-s_T,1)}{\sqrt{1+s_T^2}}\] encloses \(\widetilde R_T\) in a Euclidean rectangle \(U_T\) whose side lengths are \[L_T=\sqrt{1+s_T^2}\,a +\frac{|s_T|}{\sqrt{1+s_T^2}}b, \qquad B_T=\frac{b}{\sqrt{1+s_T^2}}.\] The slopes are absolutely bounded. Consequently \(L_T\asymp a\), \(B_T\asymp b\), \(|U_T|\le C D^2|R_T|\), and \[L_T\le C|I|,\qquad \frac{B_T}{L_T}\ge c m d_T.\] On the witness set, normalization of the direction gives \(\left\|v_0-e_T\right\|\le C d_T\). Since \(m\to\infty\), this error is smaller than any required fixed small multiple of \(B_T/L_T\), for sufficiently large \(L\). The fixed widening constant \(C_{\mathrm{w}}\) can absorb the absolute aperture constants as well. Choose the horizontal localization constant so that \(C|I|\) meets the length restriction in (79). The rectangle \(U_T\) is then admissible with the common popularity parameter \[\delta_D=c L^{-C_{\mathrm{as}}}D^{-2},\] where \(c>0\) is sufficiently small and absolute. Indeed, \(|E_T|\ge L^{-C_{\mathrm{as}}}|R_T|\), and \(|U_T|\le CD^2|R_T|\). Moreover, \[\frac1{|U_T|}\int_{U_T}\mathbf 1_Q \ge\frac{|Q\cap DR_T|}{CD^2|R_T|}>c\beta.\] Every point of \(E_T\) therefore belongs to \(\{M_{v_0,\delta_D}\mathbf 1_Q>c\beta\}\).

If \(b/a>\varepsilon_0\), enclose \(\widetilde R_T\) in the axis-parallel rectangle with horizontal side \(J_T\) and vertical side of length \(b+|s_T|a\). Its area is at most \(C ab\le CD^2|R_T|\). It contains \(E_T\) and \(Q\cap DR_T\), so every point of \(E_T\) belongs to \(\{M_{\mathrm{str}}\mathbf 1_Q>c\beta\}\). This argument also covers boxes whose transverse side exceeds the horizontal side.

The sets \(E_T\) are pairwise disjoint. Summing their lower measure bounds and using the two maximal estimates yields \[\begin{align*} \sum_{T:a_T(D)>\beta}|R_T| &\le L^{C_{\mathrm{as}}} \left|\bigcup_{T:a_T(D)>\beta}E_T\right|\\ &\le C L^{C_{\mathrm{as}}} (\delta_D^{-1}+1)\beta^{-2}|Q|\\ &\le C L^{2C_{\mathrm{as}}}D^2\beta^{-2}|Q|. \end{align*}\] The bound by \(S\) is immediate, proving (80).

For positive numbers \(S_0,A_0\), direct integration, split at \(\min(1,\sqrt{A_0/S_0})\), gives \[\int_0^1\min(S_0,A_0\beta^{-2})\,\mathrm d\beta \le2\sqrt{S_0A_0}.\] Layer cake in (80) therefore gives \[\begin{align*} \sum_T|R_T|a_T(D) &\le C L^{C_{\mathrm{as}}}D\sqrt{S|Q|}, \tag{82}\\ \sum_T|Q\cap DR_T| &\le C L^{C_{\mathrm{as}}}D^3\sqrt{S|Q|}. \tag{83}\end{align*}\] In the second line the additional factor \(D^2\) is the dilation-area ratio.

Set \(D_k=2^kD_*\). The explicit formula for \(h_T\) implies the pointwise bound \[h_T(z)\le C_{D_2}\sum_{k\ge0}2^{-kD_2}\mathbf 1_{D_kR_T}(z).\] To verify it outside \(D_*R_T\), choose the least \(k\ge1\) such that \(z\in D_kR_T\). Then \(\mathop{\mathrm{dist}}_T(z)>2^{k-2}D_*\), while \(D_*\ge2K_0\), so \(h_T(z)\le C_{D_2}2^{-kD_2}\). Inside \(D_*R_T\), use \(h_T\le C_3\). Applying (83), \[ \int_Q n \le C L^{C_{\mathrm{as}}}D_*^3\sqrt{S|Q|} \sum_{k\ge0}2^{-k(D_2-3)} \le C L^{C_{\mathrm{as}}}D_*^3\sqrt{S|Q|}. \tag{84}\] The sum converges because \(D_2>4\). Finally take \(Q=\{z\in\mathcal D:n(z)>\lambda\}\). This set has finite measure, since the collection is finite and \(D_2>2\). The inequality \(\lambda|Q|\le\int_Q n\) proves the asserted weak \(L^2\) bound. All exponents in this argument are fixed: in particular the dilation cost before the final summation is exactly \(D^3\), independently of the number of tops and angular scales. ◻

Parameters and masks

All logarithms below are natural. Set \[ \begin{split} p&=L^{1/2},\qquad G=2^{\,2\left\lceil L^{0.7}/(2\log2)\right\rceil},\\ K&=\left\lceil(\log L)^2\right\rceil, \qquad B=G^K. \end{split} \tag{85}\] Thus \(G\) and \(\sqrt G\) are dyadic powers, \(\log G\asymp L^{0.7}\), and \(\log B\asymp L^{0.7}(\log L)^2=o(L^{0.9})\). In particular, \[ \frac{B^{100}}m\longrightarrow0\qquad(L\longrightarrow\infty). \tag{86}\]

The parameters have different tasks. The large bin number \(m\) permits fine angular phase separation while keeping changes in the packet envelopes small. The much smaller time cutoff \(\tau\) pays for the commutation errors. The scale ratio \(G\) will separate prediction and averaging lengths; grouping \(K\) angular transitions into a block gives the ratio \(B\). Their strict asymptotic ordering is what allows these tasks to coexist.

Let \(A>2\) be a fixed exponent to be chosen sufficiently large after the unmasked top estimates in (110). Their fixed exponents and dilation \(K_1=L^{A_1}\) are independent of \(A\). The comparison of the original and masked sums uses 21 at \(K_1\), so its loss \(c_{\mathrm{cnt}}(A_1)\) is fixed before \(A\) is chosen. The count at the \(A\)-dependent \(K_0\) below is used only after that choice. Constants in the mask estimates may depend on the fixed value of \(A\). Put \(r_L=2\lceil L\rceil\), and define, in the coordinates (75), \[ q_T(z)=\left( \operatorname{sinc}\frac{X_T(z)}{L^A}\, \operatorname{sinc}\frac{Y_T(z)}{L^A} \right)^{r_L},\qquad \operatorname{sinc}(a)= \begin{cases}\sin(a)/a,&a\ne0,\\1,&a=0.\end{cases} \tag{87}\] The exponent is even, so \(0\le q_T\le1\) and the mask is real. From now on take \[K_0=L^{A+2}\] in (76). A top is called significant at \(z\) if \(\mathop{\mathrm{dist}}_T(z)\le K_0\). Every significant top contributes at least one to \(n(z)\).

Lemma 22. The mask has Fourier support \[ \mathop{\mathrm{supp}}\widehat q_T\subset \left\{(\xi,\eta): |\xi+s_T\eta|\le\frac{r_L}{L^A l_T},\quad |\eta|\le\frac{r_L}{L^A W}\right\}. \tag{88}\] For sufficiently large \(L\), with all other exponents fixed, \[\begin{align*} |q_T(z)|&\le \exp(-L\log L) &&\text{if }\mathop{\mathrm{dist}}_T(z)>K_0, \tag{89}\\ |q_T(z)|h_T(z)^{-1}&\le C &&\text{for all }z, \tag{90}\\ |1-q_T(z)|&\le C L^{1-A} &&\text{if }\mathop{\mathrm{dist}}_T(z)\le L^{A/2}. \tag{91}\end{align*}\] In particular multiplication by \(q_T\) adds at most \(O(L^{1-A})\) times the bandwidths of the top box to a Fourier support.

Proof. For the Fourier convention \(e^{ix\xi}\), the transform of \(\operatorname{sinc}(x)\) is supported in \([-1,1]\). Taking its \(r_L\)-th power convolves that transform \(r_L\) times, so the transform of \(\operatorname{sinc}(X/L^A)^{r_L}\) is supported in \([-r_L/L^A,r_L/L^A]\). The product of the two one-dimensional functions has this support bound in each normalized frequency coordinate. The affine coordinate change (75) sends these coordinates to \(l_T(\xi+s_T\eta)\) and \(W\eta\), proving (88). The center translation contributes only a Fourier phase. The calculation can equivalently be performed first for compactly supported Fourier distributions and then their convolutions; the resulting mask is integrable for \(r_L>1\).

Write \(d=\mathop{\mathrm{dist}}_T(z)\). Since at least one of \(|X_T(z)|,|Y_T(z)|\) is at least \(d/2\), the elementary inequality \(|\operatorname{sinc}(a)|\le\min(1,|a|^{-1})\) gives, for \(d>0\), \[ |q_T(z)|\le\min\left(1,\left(\frac{2L^A}{d}\right)^{r_L}\right). \tag{92}\] If \(d>K_0=L^{A+2}\), this is at most \((2L^{-2})^{r_L}\le\exp(-L\log L)\), proving (89). To include arbitrary distances in the weighted bound, write \(d=tK_0\), with \(t>1\). Once \(r_L\ge D_2\), \[|q_T(z)|h_T(z)^{-1} \le 2^{-D_2}(1+t)^{D_2} (2L^{-2})^{r_L}t^{-r_L} \le (2L^{-2})^{r_L}.\] For \(d\le K_0\), use \(q_T\le1\) and \(h_T\ge1\). This proves (90), including the region arbitrarily far from the top.

Finally, if \(d\le L^{A/2}\), both arguments of the sinc functions have absolute value at most \(L^{-A/2}\). They lie in \([0,1]\) after evaluation by sinc, and \[0\le1- \operatorname{sinc}(X_T/L^A) \operatorname{sinc}(Y_T/L^A) \le C\frac{X_T^2+Y_T^2}{L^{2A}} \le C L^{-A}.\] The inequality \(1-a^{r_L}\le r_L(1-a)\), for \(0\le a\le1\), proves (91). Since \(r_L/L^A=O(L^{1-A})\), the bandwidth assertion follows from (88) and addition of Fourier supports under multiplication. ◻

For completeness, the spectral enlargement also respects the finer tile coordinates. If \(T=T(s)\), an added frequency \((\zeta,\nu)\in\mathop{\mathrm{supp}}\widehat q_T\) satisfies \[|\zeta+\theta_s\nu| \le\frac{r_L}{L^A} \left(\frac1{l_T}+\frac{|\theta_s-s_T|}{W}\right) \le\frac{C L^{1-A}}{\ell_s},\qquad |\nu|\le\frac{C L^{1-A}}{W}.\] Here we used \(\ell_s\le l_T\), \(|\theta_s-s_T|\le Cd_s\), and \(d_s/W=1/(m\ell_s)\). Thus all fixed spectral margins needed later remain available after taking \(L\) sufficiently large.

Reference-line concentration and significant directions

We use the finite packet forms, their assignment to popular tops, and the notation of the preceding sections. In particular, all estimates concern one of the separate forms in the packet expansion. Its label multiplicity and its spatial lattice multiplicity are bounded by absolute constants. The Hilbert space of bin components is denoted by \(\mathcal H\), and \(\left\|g(z)\right\|_{\mathcal H}\leq\mathbf 1_E(z)\leq1\). Every subcollection in this section is fixed before its functions are evaluated at a spatial point.

The first task is to control every interval of scales assigned to a top outside one small exceptional set. We prove moment bounds on each reference line, then obtain the uniform top and derivative bounds in 28. 29 uses these bounds to replace the original adjoint sum \(P\) by the masked sum \(P_q\) and impose the weighted top-count threshold.

For a slope \(t\), put \(Y_t(x,y)=y-tx\). A \(t\)-box of dimensions \(r\times W\) is a rectangle in the coordinates \((x,Y_t)\). The affine change to these coordinates preserves area. Fix a reference length \(l\) and consider tiles satisfying \[ \ell_s\leq l,\qquad \left|\theta_s-t\right|\leq C_* B^4d_s,\qquad d_s=w/\ell_s. \tag{93}\] Here and below the fixed constant \(C_*\) can be increased finitely many times. This only changes absolute constants in the conclusions. The inequality \(B^{100}\ll m\) is more than sufficient for all such choices.

The tiles assigned to one top satisfy the narrower bound \(\left|\theta_s-s_T\right|\leq C d_s\). The wider \(B^4\) window also covers the fixed packet families used for residual estimates, where the reference slope can be that of a different top; see (187) and (188).

For a fixed, sufficiently large integer \(n\), write \[ \mathcal B_{s,n}^{t}(z)= \left(1+\frac{(x-x_s)^2}{\ell_s^2}\right)^{-n} \left(1+\frac{(Y_t(z)-Y_t(z_s))^2}{W^2}\right)^{-n}. \tag{94}\] The change from the packet coordinates of slope \(\theta_s\) to these coordinates has normalized shear at most \(C_*B^4/m\). Thus packet bounds with any specified fixed number of derivatives and any specified fixed decay order imply the corresponding bounds by \(C_n\mathcal B_{s,n}^{t}\). We always choose the available packet decay orders larger than the envelope orders used below. These are fixed integers; none grows with \(L\), the number of scales, or \(\dim\mathcal H\).

Positive estimates and the main coefficients

For a reference slope \(t\), decompose the actual coefficient as \(a_s=a_s^{\rm main}+a_s^{\rm err}\), where componentwise \[ a_{s,j}^{\rm main} =\frac1{\left|R_s\right|}\int \mathbf 1_{\{\left|u(z)-t\right|\leq d_s\}} g_j(z)\overline{b_{s,j}(t)\Psi_{s,j}(z)}\,\,\mathrm dz. \tag{95}\] This convention agrees with the pairing, linear in its first argument, used to define \(a_s\). In particular the error analyzes \(\overline{b_{s,j}(u)}g_j- \mathbf 1_{\{\left|u-t\right|\leq d_s\}}\overline{b_{s,j}(t)}g_j\).

We first record a positive estimate which also applies when only one scale is present. For \(r>0\) let \(k_{r,n}(x)=r^{-1}(1+x^2/r^2)^{-n}\).

Lemma 23. For a fixed scale \(\ell=w/d\), and any subcollection of tiles satisfying (93) at this scale, \[ \sum_{\ell_s=\ell}\left\|a_s\right\|_{\mathcal H} \mathcal B_{s,n}^{t}(z)\leq C_n. \tag{96}\] There are fixed constants \(0<c<1<C'\) for which \[\begin{align*} f_d(z)&=C\left( \mathbf 1_{\{cd\leq\left|u(z)-t\right|\leq C'B^4d\}} +\frac{\left|u(z)-t\right|}{d} \mathbf 1_{\{\left|u(z)-t\right|\leq cd\}}\right), \tag{97}\\ \sum_{\ell_s=\ell}\left\|a_s^{\rm err}\right\|_{\mathcal H} \mathcal B_{s,n}^{t}(z) &\leq C_n \big(k_{\ell,n'}*_{x} k_{W,n'}*_{Y_t}f_d\big)(z) \tag{98}\end{align*}\] for a fixed \(n'>2\) with \(2n'<n\). The constants are independent of the number of labels permitted in (93).

Proof. Minkowski’s inequality bounds the norm of an analyzing coefficient by the integral of its scalar analyzing envelope times the largest modulus of its multiplier among the components, divided by \(\ell W\); the factor \(\left\|g\right\|_{\mathcal H}\) can be replaced by one. At a fixed source point the support condition for \(b_{s,j}(u)\) allows only a bounded number of labels at this scale. This remains true after taking the maximum over components, since the support condition is the common condition on the label slope at precision \(d\).

Choose the available analyzing decay order \(N\geq2n\), with \(n,n',N\) all fixed. For one label, the lattice gives the unequal-exponent bound \[\sum_{s\text{ in that lattice}} \frac{\mathcal B_{s,N}^{t}(z')\mathcal B_{s,n}^{t}(z)}{\ell W} \leq C_{n,N,n'} k_{\ell,n'}(x-x') k_{W,n'}(Y_t(z)-Y_t(z')).\] Indeed, in each normalized coordinate, \(1+\left|a-b\right|^2\leq2(1+\left|a-c\right|^2)(1+\left|b-c\right|^2)\) extracts the relative kernel with exponent \(n'\) from the two endpoint envelopes. The remaining exponents are \(N-n'\) at the source and \(n-n'\) at the output. Drop the source factor, which is at most one, and sum the output envelope. Its lattice sum is uniformly bounded since \(n-n'>1/2\): first sum the vertical centers in each horizontal row, then the horizontal rows. The transverse translations of the rows do not affect these one-dimensional bounds. The comparison with the reference coordinates already gives the required analyzing envelope. Summing the bounded number of labels active at the source and integrating proves (96).

For the error, if \(\left|u-t\right|\leq cd\) with \(c\) sufficiently small, the replacement indicator is one and smoothness gives \(\left|b_{s,j}(u)-b_{s,j}(t)\right|\leq C\left|u-t\right|/d\). The union of labels on which either multiplier is nonzero is bounded. Outside this region, any nonzero original multiplier requires \(\left|u-t\right|\leq C'B^4d\) by (93), whereas a nonzero replacement requires \(\left|u-t\right|\leq d\). The multipliers are bounded there. Again, only boundedly many labels are active at the source, counting both families. Applying the same lattice calculation with the common factor (97) proves (98). ◻

The smoothness and support assumptions on \(b_{s,j}\) also show that nonzero main coefficients use only a bounded number of labels per scale. In such a component, \(\left|t-\theta_s\right|\asymp d_s\). On a line of slope \(t\), the Fourier frequency of a packet is \(\xi+t\eta\), and hence its support is contained in \[ c_a\ell_s^{-1}\leq\left|\xi+t\eta\right| \leq C_a\ell_s^{-1} \tag{99}\] with absolute \(0<c_a<C_a\). This uses the separation of the selector from \(\theta_s\) and the small choice of the packet bandwidth \(c_*\).

Multiplying a synthesizing packet by its top mask preserves (99), with slightly relaxed absolute constants. Indeed the extra frequency on a \(t\)-line has magnitude at most \[ C L^{1-A}\left(l_{T(s)}^{-1} +\frac{\left|t-s_{T(s)}\right|}{W}\right) \leq C L^{1-A}(1+B^4/m)\ell_s^{-1}. \tag{100}\] Here \(\ell_s\leq l_{T(s)}\) and \(\left|t-s_{T(s)}\right|\leq C B^4d_s\). For \(A\geq2\) and sufficiently large \(L\), this is less than \(c_a/(10\ell_s)\). Moreover, for every fixed derivative order the masked packets have the same uniform envelope bounds. One can see this either from the explicit sinc product or from its compact Fourier support and bounded supremum. In particular \[ \left|(\partial_x+t\partial_y)q_T\right| \leq C L^{1-A}\big(l_T^{-1}+\left|t-s_T\right|/W\big),\qquad \left|\partial_yq_T\right|\leq C L^{1-A}/W. \tag{101}\] These absolute bounds are global and do not require a good-point assumption.

Lemma 24. For every \(t\)-box \(Q\) of width \(W\), the main coefficients satisfy \[ \sum_{\substack{z_s\in Q\\\ell_s\leq\mathrm{length}(Q)}} \left\|a_s^{\rm main}\right\|_{\mathcal H}^2\left|R_s\right| \leq C\left|Q\right|. \tag{102}\] The same estimate holds after restricting to any fixed subcollection.

Proof. Let \(\mathcal S_Q\) be the finite set in the sum and dualize its weighted coefficient norm against vectors \(c_s\) satisfying \(\sum_{s\in\mathcal S_Q}\left\|c_s\right\|^2\left|R_s\right|=1\). Using (95), the dual pairing is bounded by the integral of the norm of the synthesis \[\sum_{s\in\mathcal S_Q} \mathbf 1_{\{\left|u(z)-t\right|\leq d_s\}} \big(c_{s,j}b_{s,j}(t)\Psi_{s,j}(z)\big)_j.\] At each \(z\), the indicators choose a partial interval of the dyadic lengths: \(\left|u(z)-t\right|\leq w/\ell_s\) is monotone in \(\ell_s\). The norm of this synthesis is therefore at most its maximal partial length sum. All its nonzero component frequencies satisfy (99).

Write \(r=\mathrm{length}(Q)>0\). Apply the proof of the Hilbert-valued synthesis estimate in 17, including a fixed polynomial growth weight \(P_Q\) about \(Q\). This proof applies to the present reference box even if \(r\) is not dyadic or \(r>\tau\). For the weighted packet bounds it uses only center containment in \(Q\) and \(\ell_s\leq r\); multiplication by \(P_Q\) preserves frequency support and gives tile bounds with slower fixed decay. The annular separation and maximal partial sums depend on the dyadic tile lengths, not on the reference length \(r\). For example, a fixed sufficiently large integer \(q\) gives the admissible polynomial \[P_Q(z)=\left(1+\frac{(x-x_Q)^2}{r^2}\right)^q \left(1+\frac{(Y_t(z)-Y_t(z_Q))^2}{W^2}\right)^q, \qquad \int P_Q^{-2}\leq C_q rW=C_q\left|Q\right|.\] The main label count is bounded, so the same synthesis proof yields \[\int\left\|\text{selected synthesis}\right\| \leq \left(\int P_Q^{-2}\right)^{1/2} \left\|P_Q\,\text{maximal synthesis}\right\|_2 \leq C\left|Q\right|^{1/2}.\] Taking the supremum over \(c_s\) proves (102). Deleting packets preserves both the synthesis hypotheses and this proof. ◻

Fourier sums and a dyadic martingale

For a finite sequence \((F_k)\) of functions with values in \(\mathcal H\), define its interval-sum variation by \[ \mathcal V_3(F_k)(x) =\sup_{\mathcal I} \left(\sum_{I\in\mathcal I} \left\|\sum_{k\in I}F_k(x)\right\|_{\mathcal H}^{3}\right)^{1/3}, \tag{103}\] where \(\mathcal I\) ranges over finite families of disjoint integer intervals. Empty intervals contribute zero. This is bounded by the usual \(3\)-variation of partial sums with the initial zero partial sum included: insert the endpoints of the chosen intervals in order and retain the resulting nonnegative terms. This formulation allows the interval endpoints to depend on the evaluation point.

Lemma 25. Suppose \(F_k\in L^2(\mathbb R;\mathcal H)\) has Fourier support in \(\{\nu:c_a2^k\leq\left|\nu\right|\leq C_a2^k\}\), with fixed positive \(c_a,C_a\). Then \[ \left\|\mathcal V_3(F_k)\right\|_{L^2(\mathbb R)} \leq C_{c_a,C_a} \left(\sum_k\left\|F_k\right\|_2^2\right)^{1/2}. \tag{104}\] The constant is independent of the finite index range and of \(\dim\mathcal H\).

Proof. We give the smooth-convolution versus dyadic-expectation comparison explicitly; compare (Jones et al. 2008, Lemma 3.2 and Section 4). Split the integers into a fixed number \(q\) of residue classes, large enough that the annuli in one class have gaps of ratio greater than four. The variation of the sum is at most the sum of the variations of these classes, so it suffices to treat one class. Its supports are disjoint, and \(F=\sum_kF_k\) satisfies \(\left\|F\right\|_2^2=\sum_k\left\|F_k\right\|_2^2\).

Choose a smooth function \(\chi\) equal to one on \([-1,1]\) and supported in \([-2,2]\), and let \(P_i\) have multiplier \(\chi(2^{-i}\nu)\). At a sequence of integer indices lying in the annular gaps, \(P_iF\) is exactly the partial sum of all the preceding bands. Include an index below the first band, where this partial sum is zero.

For each integer \(i\), let \(\mathcal F_i\) be the sigma algebra generated by the intervals \([a2^{-i},(a+1)2^{-i})\), \(a\in\mathbb Z\), and let \(E_i\) be averaging on its atoms. The sigma algebras increase with \(i\) and \((E_iF)_i\) is a martingale. We prove the comparison bound \[ \sum_i\left\|(P_i-E_i)F\right\|_2^2\leq C\left\|F\right\|_2^2. \tag{105}\] Take a smooth dyadic partition \(F=\sum_jf_j\) with \(f_j\) supported in a fixed annulus of size \(2^j\) and \(\sum_j\left\|f_j\right\|_2^2\leq C\left\|F\right\|_2^2\). If \(j\leq i-C\), then \(P_if_j=f_j\), and Poincaré’s inequality on each interval of length \(2^{-i}\), followed by Bernstein’s inequality, gives \[\left\|(P_i-E_i)f_j\right\|_2\leq C2^{j-i}\left\|f_j\right\|_2.\] If \(j\geq i+C\), then \(P_if_j=0\). Define its band primitive \(A_j\) by \(\widehat A_j(\nu)=\widehat f_j(\nu)/(i\nu)\). Thus \(A_j'=f_j\) and \(\left\|A_j\right\|_2\leq C2^{-j}\left\|f_j\right\|_2\). With \(h=2^{-i}\), averaging on a cell gives \(E_if_j=(A_j((a+1)h)-A_j(ah))/h\) there. The sampling estimate \[ \sum_{a\in\mathbb Z}\left\|A_j(ah)\right\|_{\mathcal H}^{2} \leq C2^j\left\|A_j\right\|_2^2,\qquad h\geq c2^{-j}, \tag{106}\] follows directly from smooth reproduction. Namely, take a Schwartz kernel \(K_j\) whose Fourier transform equals one on the spectrum of \(A_j\) and which satisfies \(\left|K_j(x)\right|\leq C_N2^j(1+2^j\left|x\right|)^{-N}\). Cauchy–Schwarz in \(A_j=K_j*A_j\) gives \(\left\|A_j(ah)\right\|^2\leq C\int\left|K_j(ah-x)\right|\left\|A_j(x)\right\|^2\,\mathrm dx\); summing uses \(\sup_x\sum_a\left|K_j(ah-x)\right|\leq C2^j\). Consequently \[\left\|E_if_j\right\|_2^2 \leq C h^{-1}\sum_a\left\|A_j(ah)\right\|^2 \leq C2^{i-j}\left\|f_j\right\|_2^2.\] For \(\left|i-j\right|\leq C\) use the bounded \(L^2\) norms of \(P_i,E_i\). The resulting bounds have the form \(\left\|(P_i-E_i)f_j\right\|_2\leq a_{i-j}\left\|f_j\right\|_2\) with a fixed \(\ell^1(\mathbb Z)\) sequence \(a\). The triangle inequality followed by the \(\ell^1*\ell^2\) convolution inequality proves (105).

For any sequence \((D_i)\), its usual \(3\)-variation is bounded pointwise by \(2(\sum_i\left\|D_i\right\|^2)^{1/2}\): each endpoint of an increasing sequence is used at most twice, and \(\ell^3\) is bounded by \(\ell^2\). Thus (105) controls the variation of \((P_i-E_i)F\).

Lépingle’s martingale variation inequality (Lépingle 1976), in the unweighted \(L^2\) case and at variation exponent \(3\), gives \(\left\|V^3(E_iF)\right\|_2\leq C\left\|F\right\|_2\); see (Zorin-Kranich 2020, Theorem 1.1, with \(p=2\), \(r=3\), and weight one). For a finite index range on \(\mathbb R\), apply its probability-space version on every atom of the coarsest sigma algebra, normalizing Lebesgue measure there, and sum the squared inequalities. This proves the same statement on the line without a finite-measure assumption. For Hilbert-valued \(F\), apply the scalar inequality to its real and imaginary components. Minkowski’s inequality, in the form \(\ell^3(\ell^2)\leq\ell^2(\ell^3)\), and the outer \(L^2\) norm give the same constant independently of the number of components. Combining with the comparison and then taking the selected low-pass indices proves the estimate for one spaced class. Summing the fixed number of classes proves (104). ◻

Every-line moment bounds

We use the following precise local form of the John–Nirenberg inequality (John and Nirenberg 1961); a proof is included to specify the normalization.

Lemma 26. Let \(I\) be a finite interval and let \(f\) be a locally integrable scalar function on \(I\). Suppose that for each dyadic subinterval \(J\) of \(I\), \(\left|J\right|^{-1}\int_J\left|f-f_J\right|\leq A\), where \(f_J\) denotes its average. Then, for every \(b\geq1\), \[\left(\frac1{\left|I\right|}\int_I\left|f\right|^b\right)^{1/b} \leq\left|f_I\right|+CbA.\]

Proof. If \(A=0\), dyadic differentiation makes \(f\) constant almost everywhere, and the assertion is immediate. Otherwise take the maximal proper dyadic subintervals \(J\) of \(I\) on which the average of \(\left|f-f_I\right|\) exceeds \(2A\). Their total length is at most \(\left|I\right|/2\), the function has \(\left|f-f_I\right|\leq2A\) almost everywhere off their union, and dyadic parent maximality gives \(\left|f_J-f_I\right|\leq4A\) on each selected interval. Repeat the same selection inside each \(J\), now subtracting its own mean. At generation \(k\) the selected union has measure at most \(2^{-k}\left|I\right|\). Outside it, the successive differences of means and the final differentiation bound show \(\left|f-f_I\right|\leq(4k+2)A\). This proves an exponential distributional bound for \(\left|f-f_I\right|/A\). Integrating that bound by layer cake gives \(\left\|f-f_I\right\|_{L^b(I,\,\mathrm dx/\left|I\right|)}\leq CbA\); adding the mean proves the lemma. Finite stopping generations and then monotone convergence justify the argument for any locally integrable \(f\). ◻

Fix a sufficiently large envelope exponent \(n\). At each scale put \[U_\ell(z)=\sum_{\ell_s=\ell}\left\|a_s^{\rm main}\right\|\mathcal B_{s,n}^{t}(z), \qquad \mathcal E_t(z)=\sum_s\left\|a_s^{\rm err}\right\|\mathcal B_{s,n}^{t}(z), \qquad \mathcal S_t(z)=\left(\sum_\ell U_\ell(z)^2\right)^{1/2}.\] These quantities can be formed on any fixed subcollection; it is also permissible to dominate a subcollection by the positive quantities on the full collection.

Proposition 27. For every slope \(t\), every collection satisfying (93), every line \(Y_t=y_0\), every horizontal interval \(I\) of length \(l\) on that line, and every \(b\geq2\), the normalized \(L^b(I)\) norms of \(\mathcal E_t\), \(\mathcal S_t\), and \[\mathcal V_3\left(\sum_{\ell_s=2^k}\phi_s\right), \qquad \mathcal V_3\left(\sum_{\ell_s=2^k}q_{T(s)}\phi_s\right)\] are at most \(C b^C(1+\log B)^C\). The constants are uniform in the line, interval, finite collection, dimension, and number of scales. In particular the assertion holds for every fixed deterministic subcollection with the same constants.

Proof. We give the local mean and oscillation bounds, including their behavior on intervals shorter than \(l\). All convolutions in what follows are in the \((x,Y_t)\) coordinates.

First, (97) satisfies \[ 0\leq f_d\leq C,\qquad \sum_{d\text{ dyadic}}f_d\leq C(1+\log B). \tag{107}\] The first indicators are nonzero for only \(O(1+\log B)\) dyadic \(d\) at a fixed point, and the second terms sum geometrically. Average first with the common vertical kernel \(k_{W,n'}\). Its \(L^1\) norm is fixed, so (107) persists on each reference line. The right side of (98), summed over scales, therefore defines a nonnegative majorant \(\mathcal A_t\) for \(\mathcal E_t\).

Let \(J\) be any horizontal interval of length \(r\leq l\). If \(\ell\leq r\), \[\frac1r\int_J k_{\ell,n'}(x-x')\,\mathrm dx \leq \frac Cr\left(1+\frac{\mathop{\mathrm{dist}}(x',J)}r\right)^{-2}.\] Hence the local mean of the short-scale part of \(\mathcal A_t\) is at most \(C(1+\log B)\), by summing the source functions before integration. For \(\ell>r\), the derivative of its convolution with a bounded source has supremum at most \(C/\ell\). The total oscillation of these terms on \(J\) is at most \(Cr\sum_{\ell>r}\ell^{-1}\leq C\). It follows that \(\mathcal A_t\) has local BMO norm \(O(1+\log B)\). Its mean on an interval of length \(l\) has the same bound, since all scales then belong to the short-scale part. John–Nirenberg on that interval gives \(\left\|\mathcal E_t\right\|_{L^b(I,\,\mathrm dx/l)}\leq Cb(1+\log B)\). We only use domination by the BMO majorant here; no assertion that an arbitrary function below a BMO function is BMO is required.

For the square function, fixed-scale lattice overlap gives \[ U_\ell(z)^2 \leq C\sum_{\ell_s=\ell} \left\|a_s^{\rm main}\right\|^2\mathcal B_{s,n}^{t}(z). \tag{108}\] For the same interval \(J\), partition center space into \(t\)-boxes \(Q_{ab}\) of dimensions \(r\times W\), with horizontal offset \(a\) from \(J\) and transverse offset \(b\) from the reference line. For \(z_s\in Q_{ab}\) and \(\ell_s\leq r\), the envelope bound implies \[\int_J\mathcal B_{s,n}^{t}(x,y_0+tx)\,\mathrm dx \leq C\ell_s(1+\left|a\right|+\left|b\right|)^{-n_1}\] for a fixed \(n_1>4\), after choosing \(n\) sufficiently large. Indeed the transverse decay gives the \(b\) factor; for a distant horizontal box the integral is bounded by \(Cr(\ell_s/(r\left|a\right|))^{n_1}\), which is at most \(C\ell_s(1+\left|a\right|)^{-n_1}\) because \(\ell_s\leq r\). Equation (102) gives exactly \[\sum_{\substack{z_s\in Q_{ab}\\\ell_s\leq r}} \left\|a_s^{\rm main}\right\|^2\ell_s =W^{-1}\sum_{\substack{z_s\in Q_{ab}\\\ell_s\leq r}} \left\|a_s^{\rm main}\right\|^2\left|R_s\right|\leq Cr.\] Summing the offsets in (108) proves \[ \frac1r\int_J\sum_{\ell\leq r}U_\ell^2\,\mathrm dx\leq C. \tag{109}\] This argument applies on each line directly; it is not a restriction of an almost-everywhere plane estimate to a line. Notice also that the factor \(W\) has canceled, so the estimate has no dependence on \(m\). Since main coefficients are bounded, fixed-scale lattice bounds give \(U_\ell\leq C\) and \(\left|\partial_x U_\ell\right|\leq C/\ell\) on the line. The long-scale part of \(\mathcal S_t^2\) has oscillation at most \(C\) on \(J\). By (109), its local BMO norm is bounded; its mean on every length-\(l\) interval is bounded as well. John–Nirenberg applied to \(\mathcal S_t^2\) gives \(\left\|\mathcal S_t\right\|_{L^b(I,\,\mathrm dx/l)}\leq C\sqrt b\).

It remains to prove the main-term variation bound, with or without the mask. Fix \(J\) and its center boxes \(Q_{ab}\) as above, and first restrict to scales \(\ell_s\leq r\). Group the synthesis by those boxes. For a fixed box multiply its synthesis by \[P_{ab}(x)=\left(1+\frac{(x-x_{Q_{ab}})^2}{r^2}\right)^{q_0},\] where \(q_0\) is a sufficiently large fixed integer. Because centers lie in \(Q_{ab}\) and \(\ell_s\leq r\), these multiplied waves still have fixed tile envelope bounds. Their Fourier supports are preserved by polynomial multiplication. Using (108) and integrating on the entire line therefore gives \[\sum_\ell\left\|P_{ab} \sum_{\substack{\ell_s=\ell\\z_s\in Q_{ab}}} a_s^{\rm main}\Phi_s\right\|_{L^2(\mathbb R;\mathcal H)}^2 \leq Cr(1+\left|b\right|)^{-n_1},\] and the same bound holds with \(q_{T(s)}\Phi_s\) in place of \(\Phi_s\). Here and elsewhere \(a_s^{\rm main}\Phi_s\) denotes componentwise multiplication. The bound follows from (102) as above, retaining transverse decay while integrating the horizontal envelope. Apply 25 with the length index reversed; this preserves the interval-sum variation. The support hypotheses hold by (99) and (100). On \(J\), \(P_{ab}\geq c(1+\left|a\right|)^{2q_0}\), and multiplication by this one scalar commutes with every interval sum and with its norm. Thus the \(L^2(J)\) variation of the synthesis from \(Q_{ab}\) is at most \[Cr^{1/2}(1+\left|a\right|)^{-2q_0}(1+\left|b\right|)^{-n_1/2}.\] Minkowski and summation over \(a,b\) give a bound \(Cr^{1/2}\) for the short-scale main variation.

For each longer scale the derivative of main synthesis on the line is bounded by \(C/\ell\), also after masking by (101). The variation of the long-scale part therefore changes by at most \(Cr\sum_{\ell>r}\ell^{-1}\leq C\) on \(J\). The full main variation differs from the long-scale variation by at most the short-scale variation pointwise, by the triangle inequality for (103). This proves its bounded local BMO norm. Taking \(r=l\) proves its bounded mean on \(I\). John–Nirenberg gives normalized \(L^b\) norm at most \(Cb\). The variation of the error synthesis is bounded by the positive sum \(C\mathcal E_t\), since disjoint intervals use each scale at most once. For the masked error the same is true because \(\left|q_T\right|\leq1\). The already proved error bound completes the proposition. ◻

All arguments in 27 use only fixed support, decay, and derivative bounds on the synthesizing waves. They still hold if those waves are replaced by waves with the same supports and such bounds multiplied by a fixed constant \(C_0\); the conclusions are then multiplied by \(C C_0\). This observation includes fixed polynomial growth weights about a region containing the centers, provided the polynomial is absorbed into a slower fixed tile envelope. It also includes the componentwise operator \(W(\partial_y-i\eta_j)\) on the unmasked synthesizing waves. After demodulation its multiplier has bounded magnitude on the vertical packet support. These observations will be used only at fixed orders below.

Uniform top bounds and neighborhood derivatives

Choose a fixed exponent \(A_1\) so that \(K_1=L^{A_1}\) contains the polynomial dilation of \(R_T\) in which all assigned centers lie, with a fixed additional margin. This choice depends on the top assignment and the fixed weight exponent \(D_2\), and is made before choosing \(A\). Set \[h_T^{(1)}(z)=C_3\big(1+\mathop{\mathrm{dist}}_T(z)/K_1\big)^{-D_2}.\] For an interval \(\mathcal J\) of dyadic lengths define the top sum \(F_{T,\mathcal J}=\sum_{s:T(s)=T,\ \ell_s\in\mathcal J}\phi_s\). The family of these sums includes empty intervals and is finite for our finite packet collection. Let \(\mathcal D F=(W(\partial_y-i\eta_j)F_j)_j\).

Corollary 28. For every fixed \(J_*>0\) there are a fixed exponent \(C_{\mathrm{top}}\) and a measurable set \(\mathcal X_{\rm top}\) with \(\left|\mathcal X_{\rm top}\right|\leq S e^{-J_*L}\) such that, for all tops \(T\) and all \(z\notin\mathcal X_{\rm top}\), \[ \sup_{\mathcal J}\ \sup_{\left|v\right|\leq W} \left(\left\|F_{T,\mathcal J}(z+(0,v))\right\|_{\mathcal H} +\left\|\mathcal D F_{T,\mathcal J}(z+(0,v))\right\|_{\mathcal H}\right) \leq L^{C_{\mathrm{top}}}\big(h_T^{(1)}(z)\big)^8. \tag{110}\] Independently of this exceptional set, for every \(z\), every \(l_*>0\), and every interval \(\mathcal J\subset[l_*,\infty)\), \[ \left\|(\partial_x+s_T\partial_y)F_{T,\mathcal J}(z)\right\|_{\mathcal H} \leq \frac{L^{C_{\mathrm{top}}}}{l_*}\big(h_T^{(1)}(z)\big)^8. \tag{111}\] All packet orders, \(K_1\), and \(C_{\mathrm{top}}\) can be fixed before \(A\); they are independent of the later mass cutoff and hierarchy depth.

Proof. Use 27 on assigned top tiles, with \(t=s_T\) and \(l=l_T\). These satisfy (93) with a fixed angular constant. Write \(X=(x-x_T)/l_T\) and \(Y=(Y_{s_T}(z)-Y_{s_T}(z_T))/W\). For a sufficiently large fixed integer \(q\), let \[P_T(z)=\left(1+\frac{X^2+Y^2}{K_1^2}\right)^q.\] Its product with any assigned wave is bounded by a slower fixed tile envelope with uniform constants. Indeed at the centers \(X,Y\) are bounded by a fixed fraction of \(K_1\), and in moving away from a center the top scales are no smaller than the corresponding tile scales. The same statement holds for the fixed derivatives in use. Multiplication by \(P_T\) preserves spectrum. Applying the last observation after 27 gives a uniform polynomial moment bound on every length-\(l_T\) interval of every parallel line for \(\sup_{\mathcal J}\left\|P_TF_{T,\mathcal J}\right\|\) and \(\sup_{\mathcal J}\left\|P_T\mathcal D F_{T,\mathcal J}\right\|\).

We justify the required vertical neighborhood supremum. Demodulate component \(j\) by \(e^{-i\eta_jY_{s_T}}\). Each of the just described weighted sums has vertical Fourier support in \([-C/W,C/W]\). Choose one Schwartz reproduction kernel \(K_W\) with Fourier transform one on that interval. Its absolute values at all arguments shifted by at most \(W\) are bounded by \(C_nW^{-1}(1+\left|y\right|/W)^{-n}\). Reproduction, the Hilbert norm inequality, and the supremum over \(\mathcal J\) show that its vertical neighborhood supremum is bounded by convolution of the corresponding unshifted supremum with this positive integrable kernel. Minkowski in normalized \(L^b\) on horizontal intervals preserves the uniform moment estimate. Finally \(P_T(z+(0,v))\asymp P_T(z)\) for \(\left|v\right|\leq W\), uniformly in \(z\), because \(K_1\geq1\). We have therefore proved, on every top-scale box, polynomial moments for \(P_T\) times the entire left side of (110).

For clarity, this yields a global exceptional-area bound, not merely one inside the top box. Partition the plane into boxes of dimensions \(l_T\times W\), indexed by \((a,b)\in\mathbb Z^2\). On each such box, \(1+\mathop{\mathrm{dist}}_T/K_1\asymp1+(\left|a\right|+\left|b\right|)/K_1\). Choose \(q\) with \(2q>8D_2+4\). If the left side of (110) exceeds its right side in that box, its product with \(P_T\) exceeds \[c L^{C_{\mathrm{top}}} \big(1+(\left|a\right|+\left|b\right|)/K_1\big)^{2q-8D_2}.\] Take moment order \(b_L=\lceil L\rceil\). The available moment bound is at most a fixed power of \(L\), since \(1+\log B\) is such a power. Choose \(C_{\mathrm{top}}\) larger by a fixed amount. Markov’s inequality and summation over \((a,b)\) then give \[\left|\mathcal X_T\right| \leq C\left|R_T\right|K_1^2 e^{-(J_*+1)L} \leq \left|R_T\right|e^{-J_*L}\] for sufficiently large \(L\), after increasing the threshold exponent or the fixed lower bound on \(L\). Equivalently one can first arrange \(e^{-(J_*+2)L}\) to absorb the polynomial \(K_1^2\). Taking the union over tops proves the asserted bound by \(S=\sum_T\left|R_T\right|\). This argument uses fixed \(q\) and fixed packet orders, even though the moment order grows with \(L\).

For (111), coefficient norms are bounded, there are boundedly many assigned labels per scale, and each differentiated wave has a bound \(C\ell_s^{-1}\) times a fixed tile envelope in slope \(s_T\). At a fixed scale, summing this envelope over assigned centers is at most \(C\ell_s^{-1}(1+\mathop{\mathrm{dist}}_T/K_1)^{-8D_2}\). To see the decay, multiply by the inverse top factor, absorb it into a slower tile envelope as for \(P_T\), and use bounded lattice overlap. Summing over \(\ell_s\geq l_*\) costs \(\sum_{\ell\geq l_*}\ell^{-1}\leq C/l_*\). The same bound applies to any interval of these lengths by the triangle inequality. This proves the unconditional derivative estimate and completes the corollary. ◻

The masked sum and the count threshold

The choice of \(A\) in the preceding section can now be completed. Its dependence on the fixed constants just obtained is allowed; their construction did not depend on \(A\).

Proposition 29. Given any fixed \(H>0\), choose \(A\) sufficiently large after \(K_1\) and \(C_{\mathrm{top}}\), and then choose a sufficiently large fixed count threshold exponent \(C_n\). There is a measurable set \(F_{\rm conc}\subset F\) such that \[\left|F\setminus F_{\rm conc}\right|\leq L^{-H}\left|F\right|, \qquad \int_{F_{\rm conc}}\left\|P-P_q\right\|_{\mathcal H} \leq L^{-H}N\left|F\right|, \qquad P_q=\sum_sq_{T(s)}\phi_s,\] and (110) holds there for every top. On this set, \[ n(z)=\sum_T h_T(z)\leq NL^{C_n}. \tag{112}\] One has \(K_0=L^{A+2}>K_1\), so both (110) and (111) also hold with \(h_T\) in place of \(h_T^{(1)}\).

Proof. First fix, for example, \(J_*=10\) in 28, and fix the resulting \(C_{\mathrm{top}}\) before making any choice of \(A\). Its exceptional area is smaller than any prescribed fixed inverse power of \(L\) times \(\left|F\right|\), since \(S/\left|F\right|\leq L^C N^2\leq L^C e^{4L}\). On \(\mathop{\mathrm{dist}}_T\leq L^{A/2}\), the elementary estimate for the mask is \(\left|1-q_T\right|\leq C L^{1-A}\): use \(\left|1-\operatorname{sinc}v\right|\leq Cv^2\) near zero and the integer power \(2\lceil L\rceil\) in (87). Outside this region, \(\left|1-q_T\right|\leq2\), whereas \[\big(h_T^{(1)}\big)^7 \leq C L^{-7D_2(A/2-A_1)}.\] Equation (110), with the full length interval, therefore implies outside \(\mathcal X_{\rm top}\) that, for any prescribed fixed \(H_1\), \[ \left\|P-P_q\right\|\leq L^{-H_1}\sum_T h_T^{(1)} \tag{113}\] once \(A\) is chosen large enough. In the near region the factor \(L^{C_{\mathrm{top}}+1-A}\) suffices; in the far region use the displayed seventh power to reduce \((h_T^{(1)})^8\) to \(h_T^{(1)}\).

Apply (77) with dilation \(K_1\) to \(n_1=\sum_Th_T^{(1)}\). For any finite-measure set \(A_0\) in the testing domain, weak \(L^2\) and layer cake imply \[\int_{A_0}n_1\leq2\left\|n_1\right\|_{2,\infty}\left|A_0\right|^{1/2} \leq L^C S^{1/2}\left|A_0\right|^{1/2}.\] Since \(S\leq L^C\left|E\right|\) and \(\left|E\right|=N^2\left|F\right|\), choosing \(H_1\) larger than \(H\) by the fixed powers in these inequalities proves the claimed mask integral even when \(A_0=F\setminus\mathcal X_{\rm top}\).

After \(A\) has been fixed, (77) with dilation \(K_0\) gives \[\left|\{z\in F:n(z)>NL^{C_n}\}\right| \leq L^{C-2C_n} S/N^2 \leq L^{C-2C_n}\left|F\right|.\] Take \(C_n\) sufficiently large and define \(F_{\rm conc}\) by deleting this set and \(\mathcal X_{\rm top}\) from \(F\). Increasing the fixed thresholds slightly makes the sum of their areas at most \(L^{-H}\left|F\right|\). This proves all assertions. ◻

The original adjoint estimate (43) gives \(\left\|P\right\|_2\leq L^{C_P}\left|E\right|^{1/2}\) for a fixed exponent \(C_P\). Thus, by taking \(H\) as large as required, any discarded set above has arbitrarily small prescribed polynomial pairing with the original \(P\). The concentration estimates above hold for finite tile sets and scale ranges, with constants independent of their cardinalities. After the later assembly proves the uniform model bound for each finite form, the estimate passes to the convergent bilinear packet representation as in [red:packet-representation,red:model-suffices].

A simultaneous count of separated significant directions

On the set \(F_{\mathrm{conc}}\) from 29, the weighted threshold (112) bounds the total top mass by \(NL^{C_n}\). Because each significant top contributes at least one, this bounds their number by the same quantity, which may grow with \(N\). For the later finite Fourier gain we need a bound polynomial in \(L\) for significant directions that are \(\rho\)-separated inside an interval of length \(B\rho\), with top precisions at most \(\rho/B^2\). We derive that bound from the popularity witnesses, simultaneously for every \(\rho>0\) and every such list.

We retain the parameters and the finite family of tops from the preceding sections. In particular, every top has width \(W=mw\), length \(l_T\leq\tau\), and precision \(d_T=w/l_T\). Its witness satisfies \[ E_T\subset L^{a_0}R_T,\qquad |E_T|\geq L^{-a_0}l_TW,\qquad |u-s_T|\leq c_0d_T\quad\hbox{on }E_T, \tag{114}\] where \(a_0\) and \(c_0\) are fixed constants; increasing \(a_0\) absorbs the fixed constants in the first two bounds. The dilated box in (114) is understood in the coordinates of \(R_T\). The sum of the top areas is \(S=\sum_T l_TW\). A top is significant at \(z\) when \(\mathop{\mathrm{dist}}_T(z)\leq K_0\), with \(K_0=L^{A+2}\) as before.

Proposition 30. There is a fixed exponent \(C_D\) with the following property. For every fixed \(J>0\) and all sufficiently large \(L\), there is a measurable set \(\mathcal E_{\mathrm{count}}\subset\mathbb R^2\) such that \[ |\mathcal E_{\mathrm{count}}|\leq S e^{-JL}. \tag{115}\] Define \[ Z_0=F_{\mathrm{conc}}\setminus\mathcal E_{\mathrm{count}}, \qquad D=L^{C_D}. \tag{116}\] For every \(z\in Z_0\), every \(\rho>0\), and every list \(\mathcal L\) of tops significant at \(z\) satisfying \[ \begin{split} &\max_{T\in\mathcal L}s_T-\min_{T\in\mathcal L}s_T\leq B\rho, \qquad d_T\leq\rho/B^2\quad(T\in\mathcal L),\\ &|s_T-s_{T'}|\geq\rho\quad(T,T'\in\mathcal L,\ T\ne T'), \end{split} \tag{117}\] one has \(\#\mathcal L\leq D/100\). The exponent \(C_D\) depends only on the fixed geometric constants, including the already chosen mask exponent \(A\). It is independent of the number of tops, the number of scales, and the bin Hilbert space.

The margin in the conclusion permits a fixed number of additional directions and fixed multiplicity factors in later applications, after increasing the absolute lower bound on \(L\). The exceptional set will be constructed from one angular bin family for each reference top. It therefore applies to all \(\rho\) and all lists simultaneously.

We transfer the other tops’ popularity witnesses to a line in the direction of a longest top. Separated top directions then occupy distinct angular bins, and the following lemma bounds the number of bins each having a large average on some interval through the evaluation point. The reference top contributes one additional direction.

A counting lemma on a line

The popular-bin counting argument is inspired by Bateman’s one-variable covering/BMO method (Bateman 2008, sec. 3.3.2, Lemmas 5–6). The moment estimate below is proved directly; the Lipschitz transfer to reference lines and the simultaneous multiresolution formulation are established here, rather than imported from Bateman’s maximal theorem.

Lemma 31. Let \((f_b)_{b\in\mathcal B}\) be a countable family of measurable functions on \(\mathbb R\) such that \(0\leq f_b\leq1\) and \(\sum_b f_b\leq C_0\), where \(C_0\geq1\). Fix \(H>0\) and \(0<\vartheta\leq1\). There is a measurable function \(\mathcal N:\mathbb R\to[0,\infty]\) with the following properties. If \(\mathcal B_x\) is any collection of distinct indices such that, for every \(b\in\mathcal B_x\), some interval \(I_b\) containing \(x\) satisfies \[ 0<|I_b|\leq H,\qquad |I_b|^{-1}\int_{I_b}f_b\geq\vartheta, \tag{118}\] then \(\#\mathcal B_x\leq\mathcal N(x)\). Moreover, for every bounded interval \(U\) and every \(\lambda\geq0\), \[ |\{x\in U:\mathcal N(x)>\lambda\}| \leq C(|U|+H)\exp(-c\vartheta\lambda/C_0), \tag{119}\] where \(C,c>0\) are absolute. If the \(f_b\) depend jointly measurably on an additional parameter, \(\mathcal N\) can be chosen jointly measurable in that parameter and \(x\).

Proof. For \(t\in\{0,1/3,2/3\}\) use the dyadic grid whose intervals of length \(2^k\) are \[\mathcal D_k^t =\big\{2^k[j+(-1)^kt,j+1+(-1)^kt):j\in\mathbb Z\big\}.\] These grids are nested: a boundary at scale \(k+1\) is a boundary at scale \(k\) because \(3t\) is an integer. At a fixed scale their combined boundary points have spacing \(2^k/3\). Given an interval \(I\) of positive finite length, choose the least integer \(k\) with \(2^k>3|I|\). Its closure meets boundary points from at most one of the three grids. Hence \(I\) is contained in an interval \(Q\) from one of the other grids, and \(|Q|\leq6|I|\).

Choose an integer \(k_H\) with \(6H\leq2^{k_H}<12H\), and in each grid retain only scales \(k\leq k_H\). Put \(\beta=\vartheta/6\). For each \(b\) and each grid, call \(Q\) bad if \[|Q|^{-1}\int_Q f_b\geq\beta.\] Let \(\mathcal J_b^t\) be the intervals maximal under inclusion among these bad intervals. Every bad interval is contained in a maximal one, since it has only finitely many ancestors up to level \(k_H\). For fixed \(b,t\), the members of \(\mathcal J_b^t\) are pairwise disjoint. Define \[\mathcal N_t(x)=\sum_{b\in\mathcal B} \sum_{Q\in\mathcal J_b^t}\mathbf 1_Q(x), \qquad \mathcal N(x)=\sum_{t\in\{0,1/3,2/3\}}\mathcal N_t(x).\] Each interval in (118) produces a bad dyadic interval containing \(x\), at an allowed scale. Thus each distinct index in \(\mathcal B_x\) contributes at least once to \(\mathcal N(x)\).

For every dyadic interval \(P\) at an allowed scale, disjointness within each bin gives the Carleson estimate \[ \begin{split} \sum_b\sum_{\substack{Q\in\mathcal J_b^t\\Q\subseteq P}}|Q| &\leq\beta^{-1}\sum_b \sum_{\substack{Q\in\mathcal J_b^t\\Q\subseteq P}} \int_Q f_b\\ &\leq\beta^{-1}\int_P\sum_b f_b \leq\mathcal A|P|,\qquad \mathcal A=C_0/\beta. \end{split} \tag{120}\] All sums here are of nonnegative terms, so this argument also applies to countably many bins and intervals.

For completeness, (120) has an elementary exponential-tail consequence. Fix a root interval \(P_0\) of length \(2^{k_H}\). First take any finite subfamily of the selected interval occurrences inside \(P_0\), retaining multiplicities when different bins select the same interval, and let \(V\) count that subfamily. For an integer \(r\geq1\), expansion of \(V^r\) and the nesting of intersecting dyadic intervals give \[ \int_{P_0} V^r \leq r!\sum_{Q_1\supseteq\cdots\supseteq Q_r}|Q_r| \leq r!\mathcal A^r|P_0|. \tag{121}\] The chain sum counts occurrences with their multiplicities. To obtain the second inequality, sum first over \(Q_r\subseteq Q_{r-1}\) and apply (120); repeat until the remaining sum is bounded by \(\mathcal A|P_0|\). Repeated intervals are allowed throughout. Ordering the \(r\) occurrences accounts for the factor \(r!\) and can only overcount.

Enumerate all selected occurrences inside \(P_0\). Their finite initial sums increase to \(\mathcal N_t\) on \(P_0\), so monotone convergence and (121) show \[ \int_{P_0}\exp(\mathcal N_t/(2\mathcal A)) \leq \sum_{r=0}^{\infty}2^{-r}|P_0|=2|P_0|. \tag{122}\] In particular the count is finite almost everywhere. Markov’s inequality bounds its superlevel measure in \(P_0\) by \(2|P_0|\exp(-\lambda/(2\mathcal A))\). The root intervals of one grid meeting \(U\) have total length at most \(|U|+2^{k_H+1}\). Sum the root estimates and observe that \(\mathcal N>\lambda\) implies \(\mathcal N_t>\lambda/3\) for at least one grid. Since \(\mathcal A=6C_0/\vartheta\) and \(2^{k_H}<12H\), this proves (119).

There are countably many bins and dyadic intervals. For each fixed interval its badness is determined by a measurable integral. Its maximality adds only finitely many conditions on its ancestors before level \(k_H\). Consequently the displayed count is measurable. The same argument, using measurability of integrals with a parameter, proves the final assertion. No maximality selection or exceptional set depending on an uncountable family of intervals is required. ◻

Transferring witnesses to a reference line

Fix a top \(R\). We use the coordinates \[z=(x,y_R+s_R(x-x_R)+v).\] The significant region of \(R\) is contained in \[ \mathcal Q_R=\{|x-x_R|\leq K_0l_R,\ |v|\leq K_0W\}. \tag{123}\] The coordinate change \((x,v)\mapsto z\) has Jacobian one.

For each integer \(h\), partition each of the two angular annuli \[[s_R+2^h,s_R+2^{h+1}),\qquad (s_R-2^{h+1},s_R-2^h]\] into \(B^2\) intervals of length \(\delta_h=2^h/B^2\). Here \(B\) is a dyadic integer, as stipulated in its definition. Denote this countable family of base intervals by \(\mathcal B_R\). Enlarge every base interval concentrically by a fixed factor \(c_1\), chosen below, and write \(b^+\) for the resulting closed interval. For sufficiently large \(L\), and hence \(B\), these enlarged intervals satisfy \[ \sum_{b\in\mathcal B_R}\mathbf 1_{b^+}(a)\leq C_{\mathrm{bin}} \qquad(a\in\mathbb R), \tag{124}\] with a fixed \(C_{\mathrm{bin}}\). Indeed, enlargement moves an endpoint by at most \(c_1 2^h/B^2\). Once \(B^2\geq4c_1\), an enlarged interval from annulus \(h\) stays at distance comparable to \(2^h\) from \(s_R\). A point therefore meets only a fixed number of annuli, and at most a fixed number, depending on \(c_1\), of the enlarged equal-length intervals in each such annulus.

Choose a fixed power \(P_L=L^{a_1}\) sufficiently large to dominate the dilation constants in (114), \(K_0\), and their fixed multiples. It will not depend on \(R\) or on a list. Replacing a witness for this argument by a compact subset \(\widetilde E_T\subset E_T\) of at least half its measure is permitted by inner regularity. The slope condition can first be enforced after deleting a null subset if necessary. The horizontal projection \(A_T\) of \(\widetilde E_T\) is compact, and Fubini’s theorem gives \[ |A_T|\geq \frac{|\widetilde E_T|}{2L^{a_0}W} \geq \frac14L^{-2a_0}l_T. \tag{125}\] The denominator bounds the length of each vertical section of the dilated witness box. The original witnesses themselves remain unchanged elsewhere in the proof.

Suppose now that a list satisfies (117) at \(z\), and choose \(R\) from the list with maximal length. If the list is empty there is nothing to prove. For \(T\ne R\) put \(\Delta_T=|s_T-s_R|\). Then \[ \rho\leq\Delta_T\leq B\rho,\qquad d_T\leq\Delta_T/B^2,\qquad l_T\leq l_R\leq\tau. \tag{126}\] There is an interval \(I_T\) containing both \(A_T\) and the horizontal coordinate \(x\) of \(z\) with \[ |I_T|\leq2P_Ll_T\leq2P_Ll_R. \tag{127}\] For example use the interval centered at \(x_T\) of radius \(P_Ll_T\); significance gives \(|x-x_T|\leq K_0l_T\).

For \(x'\in A_T\), take any \((x',y')\in\widetilde E_T\), and let \(z'=(x',y+s_R(x'-x))\) be the point with horizontal coordinate \(x'\) on the reference line through \(z\). The top-coordinate representations of the witness point and \(z\) give \[ |y'-(y+s_R(x'-x))| \leq P_L\bigl(W+l_T\Delta_T\bigr). \tag{128}\] This bound holds for every possible witness point above \(x'\), so no measurable choice of such points is needed. By the Lipschitz bound on \(u\), \[ |u(z')-s_T| \leq c_0d_T+\mathop{\mathrm{Lip}}(u)P_L\bigl(W+l_T\Delta_T\bigr). \tag{129}\] The parameter choices ensure \[ \mathop{\mathrm{Lip}}(u)P_Lm\tau\leq B^{-100},\qquad B^2\leq m \tag{130}\] for sufficiently large \(L\). To check the first assertion, \(\log(P_Lm\tau)= -L^{0.97}+O(L^{0.9}+\log L)\), whereas \(\log B=O(L^{0.7}(\log L)^2)\); the fixed Lipschitz constant does not affect this comparison. Since \(W=ml_Td_T\leq m\tau\Delta_T/B^2\), the last two terms in (129) are at most \[\bigl(\mathop{\mathrm{Lip}}(u)P_Lm\tau+ \mathop{\mathrm{Lip}}(u)P_L\tau B^2\bigr)\frac{\Delta_T}{B^2} \leq 2B^{-100}\frac{\Delta_T}{B^2}.\] Consequently the transferred error is bounded by \((c_0+1)\Delta_T/B^2\).

Let \(b(T)\) be the unique base interval containing \(s_T\). If its annulus index is \(h\), then \(2^h\leq\Delta_T<2^{h+1}\). Choose \(c_1\) large enough in terms of \(c_0\) so that every point within \(2(c_0+1)\delta_h\) of any point of a base interval lies in its enlargement. Equations (129) and (130) now show \[u(z')\in b(T)^+\qquad(x'\in A_T).\] Define, on every reference line, the bin functions \[ f_{R,b}(x',v)= \mathbf 1_{b^+}\bigl(u(x',y_R+s_R(x'-x_R)+v)\bigr). \tag{131}\] They are jointly measurable, and their sum is bounded by \(C_{\mathrm{bin}}\) by (124). Combining (125) and (127), there is a fixed exponent \(a_2\) such that, with \(\vartheta=L^{-a_2}\), \[ \frac1{|I_T|}\int_{I_T} f_{R,b(T)}(x',v)\,\,\mathrm dx' \geq\vartheta. \tag{132}\] For example any fixed \(a_2>2a_0+a_1\) suffices once \(L\) is large. All choices so far precede the eventual choice of \(D\).

The base intervals \(b(T)\) for \(T\in\mathcal L\setminus\{R\}\) are distinct. Indeed, any one such interval has length \(2^h/B^2\leq\Delta_T/B^2\leq\rho/B<\rho\), so it cannot contain two directions from a \(\rho\)-separated list.

The common exceptional set and the choice of \(D\)

Apply 31 to the full family (131), with \[H_R=2P_Ll_R,\qquad \vartheta=L^{-a_2},\qquad C_0=C_{\mathrm{bin}}.\] Write \(\mathcal N_R(x,v)\) for its jointly measurable count. This construction depends on \(R\) but on neither \(\rho\) nor a list. Whenever \(R\) is a longest member of an admissible list at \(z=(x,y_R+s_R(x-x_R)+v)\), the preceding argument gives \[ \#\mathcal L\leq1+\mathcal N_R(x,v). \tag{133}\] Use the threshold \(\Lambda_{\mathrm{count}}=L^{a_2+2}\) and let \[\mathcal E_R= \{z\in\mathcal Q_R: \mathcal N_R(x,v)>\Lambda_{\mathrm{count}}\}.\] In (119), take \(U=[x_R-K_0l_R,x_R+K_0l_R]\). Since \(P_L\) dominates \(K_0\), that estimate gives, for each \(v\), \[|\{x\in U:\mathcal N_R(x,v)> \Lambda_{\mathrm{count}}\}| \leq CP_Ll_R e^{-cL^2}.\] Integrating over \(|v|\leq K_0W\) and using the Jacobian-one coordinates yields \[ |\mathcal E_R|\leq CP_LK_0l_RW e^{-cL^2}. \tag{134}\] In particular, for every fixed \(J>0\), the right-hand side is at most \(e^{-JL}l_RW\) for sufficiently large \(L\), because \(P_LK_0\) is a fixed power of \(L\).

Set \(\mathcal E_{\mathrm{count}}=\bigcup_R\mathcal E_R\). The top family is finite, and the sets are measurable. Summing (134) proves (115). The same proof applies to any countable family with finite area sum by countable subadditivity. The angular families were already countable, and their finite approximations were handled in the proof of 31; thus no cardinality or depth factor is introduced by a limiting operation.

Choose \[ C_D=a_2+4,\qquad D=L^{C_D}. \tag{135}\] At a point outside \(\mathcal E_{\mathrm{count}}\), a longest member \(R\) of any nonempty admissible list is significant, so the point lies in \(\mathcal Q_R\) and \(\mathcal N_R\leq\Lambda_{\mathrm{count}}\) there. Equation (133) gives \(\#\mathcal L\leq1+L^{a_2+2}\leq D/100\) for sufficiently large \(L\). This proves 30.

Finally, this additional discard is permissible in the original adjoint pairing. The ordinary model estimate gives \(\left\|P\right\|_2\leq L^C\sqrt{|E|}\), and \(S\leq L^C|E|\). Consequently \[ \int_{F_{\mathrm{conc}}\setminus Z_0}\left\|P\right\| \leq L^C|E|e^{-JL/2} \leq L^C e^{(2-J/2)L}\,N|F|, \tag{136}\] where the last inequality uses \(|E|=N^2|F|\) and \(N\leq e^{2L}\). Taking, for example, any fixed \(J>4\) makes this smaller than any prescribed inverse power of \(L\) times \(N|F|\), after increasing the absolute lower bound on \(L\). The estimates (110) and (112), together with the already established mask replacement, continue to hold on \(Z_0\subset F_{\mathrm{conc}}\).

A strict finite Fourier inequality with constant one

The next estimate improves the sequence exponent in the usual finite Fourier bound while keeping the operator norm at most one. Keeping this constant prevents a fixed loss from accumulating over the levels of the angular hierarchy. The distinction between a dominant coefficient and a strict norm deficit is related to the study of near-extremizers for discrete Young and Hausdorff–Young inequalities. Relevant precedents include Fournier’s work on sharpness, the fourth-moment case of Eisner and Tao, and the quantitative discrete results of Charalambides and Christ (Fournier 1977; Eisner and Tao 2012; Charalambides and Christ 2011). Here the frequencies are arbitrary separated real numbers and the averaging probability is specified by its Fourier support. We prove the needed quantitative gap in that setting, including the exact constant one after the coefficient exponent is improved.

Throughout this section the Fourier transform of a finite measure is \(\widehat\nu(\omega)=\int e^{-i\omega t}\,\,\mathrm d\nu(t)\).

Fix the absolute constants \[ \epsilon_0=\frac14, \qquad \kappa_0=2^{-10}, \qquad c_{\rm ph}=2^{-60}. \tag{137}\] For the bound \(D\ge1\) from the preceding section and for \(p\ge4\), put \[ \alpha=2+\frac{c_{\rm ph}}{\log(2D)}, \qquad \mu=\frac{\alpha-1}{p}, \qquad r_* = \frac{1}{1-\mu}. \tag{138}\] In particular, \(2<\alpha<7/3\), \(0<\mu<1/3\), and \(1<r_*\le3/2\). The numerical values in (137) are convenient fixed choices satisfying the estimates below.

Proposition 32 (Finite Fourier gain). Let \(\lambda_1,\ldots,\lambda_d\in\mathbb R\), where \(d\le D\), satisfy \(\left|\lambda_i-\lambda_j\right|\ge\Delta>0\) for \(i\ne j\). Let \(\nu\) be any positive probability measure on \(\mathbb R\) such that \(\mathop{\mathrm{supp}}\widehat\nu\subset(-\epsilon_0\Delta,\epsilon_0\Delta)\). Then, for every finite \(p\ge4\) and every complex coefficient sequence, \[ \left(\int_\mathbb R \left|\sum_{i=1}^d a_i e^{i\lambda_i t}\right|^p \,\,\mathrm d\nu(t)\right)^{1/p} \le \left(\sum_{i=1}^d |a_i|^{r_*}\right)^{1/r_*}. \tag{139}\] The same estimate holds for coefficients in a finite-dimensional Hilbert space \(\mathcal H\) when the exponentials act diagonally: in an orthonormal basis, the \(\xi\)-th component may use frequencies \(\lambda_{i,\xi}\), provided they are \(\Delta\)-separated as \(i\) varies, for each fixed \(\xi\). The right side is then \((\sum_i\left\|a_i\right\|_{\mathcal H}^{r_*})^{1/r_*}\). All these estimates have constant one.

We first prove a quantitative gap in the fourth moment when no coefficient dominates. The proof compares the law of three indices subject to an approximate additive relation with the independent triple law; separated lower and upper quantiles force a definite difference between these laws. We then handle a dominant coefficient directly, interpolate, and pass to Hilbert-valued coefficients.

Lemma 33 (A quantitative gap away from a single coefficient). For a separated frequency list as above, define \[\mathcal Q=\{(i,j,k,l): |\lambda_i+\lambda_j-\lambda_k-\lambda_l| <\epsilon_0\Delta\}.\] Let \(x_i\ge0\), \(\sum_i x_i=1\), and put \[Q(x)=\sum_{(i,j,k,l)\in\mathcal Q} (x_i x_j x_k x_l)^{3/4}.\] Then \(Q(x)\le1\). If \(0<\kappa<1\) and \(\max_i x_i\le1-\kappa\), then \[ Q(x)\le1-\delta_\kappa, \qquad \delta_\kappa= \frac{\kappa^4(1-\kappa/2)^2}{4096} \ge 2^{-14}\kappa^4. \tag{140}\]

Proof. Any three specified indices have at most one completion to an element of \(\mathcal Q\): two possible missing frequencies would be within \(2\epsilon_0\Delta<\Delta\) of each other. On the finite set of quadruples, consider the four nonnegative measures \[A=x_i x_j x_k\mathbf 1_{\mathcal Q},\quad B=x_i x_k x_l\mathbf 1_{\mathcal Q},\quad C=x_i x_j x_l\mathbf 1_{\mathcal Q},\quad E=x_j x_k x_l\mathbf 1_{\mathcal Q}.\] Each has total mass at most one, by unique completion. Their pointwise geometric mean is the summand defining \(Q(x)\), so Hölder’s inequality already gives \(Q(x)\le1\).

For nonnegative measures on a finite set write \(\mathsf H(f,g)=\sum\sqrt{fg}\) for their Hellinger affinity. Cauchy–Schwarz gives \[ Q(x)\le \bigl(\mathsf H(A,B)\mathsf H(C,E)\bigr)^{1/2} \le \mathsf H(A,B)^{1/2}. \tag{141}\] Project \(A\) to the three coordinates \((i,k,l)\), obtaining a measure \(P\) of mass at most one, and let \(P_0(i,k,l)=x_i x_k x_l\) be the independent triple law. For each \((i,k,l)\) there is at most one completing \(j\). Consequently \[ \mathsf H(P,P_0)=\mathsf H(A,B)\ge Q(x)^2. \tag{142}\] This identity also holds when a factor \(x_i\), \(x_k\), or \(x_l\) vanishes, since the corresponding terms on both sides are zero.

Adjoin one point to the triple space and place the missing mass of \(P\) there, giving a probability \(P^+\). Extend \(P_0\) by zero at that point. For any two probabilities \(U,V\), \[\frac12\sum |U-V| \le \frac12 \left(\sum(\sqrt U-\sqrt V)^2\right)^{1/2} \left(\sum(\sqrt U+\sqrt V)^2\right)^{1/2} =\sqrt{1-\mathsf H(U,V)^2}.\] Every event \(\mathcal A\) in the triple space therefore satisfies \[ P_0(\mathcal A)-P(\mathcal A) \le \sqrt{1-\mathsf H(P,P_0)^2} \le \sqrt{1-Q(x)^4}. \tag{143}\] Thus a fourth moment close to the ordinary bound would force the completed triple law close to the independent law, quantitatively and without a factor depending on the number of frequencies.

Let \(X\) be a random frequency with probabilities \(x_i\). Choose lower and upper \(\kappa/4\)-quantile frequencies \(a,b\), so that \[\Pr(X\le a),\Pr(X\ge b)\ge\kappa/4, \qquad \Pr(X<a),\Pr(X>b)\le\kappa/4.\] They can be chosen with \(a\le b\). If \(a=b\), that atom would have mass at least \(1-\kappa/2>1-\kappa\), a contradiction. Hence \(b-a\ge\Delta\), and \([a,b]\) has mass at least \(1-\kappa/2\). Of its two closed halves, choose one, denoted \([e,h]\), with mass at least \((1-\kappa/2)/2\).

Consider the triple event \[\mathcal A=\{\lambda_i\ge b,\ \lambda_k\le a,\ \lambda_l\le h\}.\] For any completion in \(\mathcal Q\) with these pair restrictions, \[\lambda_l-\lambda_j >\lambda_i-\lambda_k-\epsilon_0\Delta \ge b-a-\epsilon_0\Delta > (b-a)/2=h-e.\] Thus \(\lambda_l\le h\) forces \(\lambda_j<e\). Under the measure \(A\), the free coordinates \((i,j,k)\) carry their independent probabilities, with a possible deletion when no completion exists. It follows that \[P(\mathcal A) \le \Pr(X\ge b)\Pr(X\le a)\Pr(X<e),\] whereas \[P_0(\mathcal A) =\Pr(X\ge b)\Pr(X\le a)\Pr(X\le h).\] Their difference is at least \[g_\kappa= (\kappa/4)^2\frac{1-\kappa/2}{2} =\frac{\kappa^2(1-\kappa/2)}{32}.\] Combining this with (143) gives \(Q(x)^4\le1-g_\kappa^2\). Since \((1-t)^{1/4}\le1-t/4\) for \(0\le t\le1\), \(Q(x)\le1-g_\kappa^2/4\). This is (140). ◻

Proof of 32. The ordinary fourth-moment bound. Write \(Ta(t)=\sum_i a_i e^{i\lambda_i t}\) and \(s=4/3\). The Fourier support hypothesis and positivity of \(\nu\) give \[\int |Ta|^4\,\,\mathrm d\nu \le \sum_{\mathcal Q}|a_i a_j a_k a_l|.\] Indeed every other term in the fourth-power expansion averages to zero, and the absolute value of every surviving Fourier coefficient is at most one. If \(\left\|a\right\|_{\ell^s}=1\), the last sum is \(Q(x)\) for \(x_i=|a_i|^s\). Thus 33 proves the ordinary norm-one bound \[ \left\|Ta\right\|_{L^4(\nu)}\le\left\|a\right\|_{\ell^{4/3}}. \tag{144}\]

For the improved bound at four, put \[\theta=\frac{c_{\rm ph}}{\log(2D)}, \qquad r_4=\frac4{3-\theta}.\] Then \(4/3<r_4\le3/2\) and \[ \left\|a\right\|_{\ell^{4/3}} \le D^{\theta/4}\left\|a\right\|_{\ell^{r_4}}, \qquad D^\theta\le e^{c_{\rm ph}}. \tag{145}\] No dominant coefficient. For a nonzero sequence, normalize \(x_i=|a_i|^{4/3}/ \sum_k|a_k|^{4/3}\). If \(\max_i x_i\le1-\kappa_0\), (140) and (145) yield \[\left\|Ta\right\|_{L^4(\nu)}^4 \le (1-\delta_{\kappa_0})e^{c_{\rm ph}} \left\|a\right\|_{\ell^{r_4}}^4 \le \left\|a\right\|_{\ell^{r_4}}^4.\] Here \(\delta_{\kappa_0}\ge2^{-54}>c_{\rm ph}\), and \(1-\delta\le e^{-\delta}\).

A dominant coefficient. Suppose instead that one coefficient has normalized mass exceeding \(1-\kappa_0\). Divide all coefficients by that coefficient and remove its carrier from \(Ta\). The normalized sum is \(1+U(t)\), and, if \(b\) denotes the \(\ell^{r_4}\) norm of the other coefficients, \[ b\le\left(\frac{\kappa_0}{1-\kappa_0}\right)^{3/4}<\frac1{64}. \tag{146}\] Frequency separation gives the exact identities \[\int U\,\,\mathrm d\nu=0, \qquad \int |U|^2\,\,\mathrm d\nu =\sum_{i\ne i_0}|a_i/a_{i_0}|^2\le b^2.\] By (144) and (145), \(\int|U|^4\,\,\mathrm d\nu\le e^{c_{\rm ph}}b^4\), and hence \(\int|U|^3\,\,\mathrm d\nu\le e^{c_{\rm ph}/2}b^3\). Expanding \(|1+U|^4\), the linear term has mean zero. Bounding the remaining terms gives \[\begin{align*} \int|1+U|^4\,\,\mathrm d\nu &\le1+6b^2+4e^{c_{\rm ph}/2}b^3+e^{c_{\rm ph}}b^4\\ &\le1+14b^2 \le1+\frac83 b^{r_4} \le(1+b^{r_4})^{4/r_4}. \end{align*}\] For the penultimate inequality, use \(b^{2-r_4}\le\sqrt b\le1/8\); for the last one, use \(4/r_4\ge8/3\) and convexity. Restoring the dominant coefficient proves the norm-one \(\ell^{r_4}\)-to-\(L^4(\nu)\) estimate in this case too. The zero sequence is immediate.

Interpolation with constant one. The same operator has norm at most one from \(\ell^1\) to \(L^\infty(\nu)\). For \(0<t\le1\), interpolation between these two bounds gives \[\frac1{p_t}=\frac t4, \qquad \frac1{r_t}=1-t+\frac t{r_4}.\] One can see the normalization directly from the three-lines lemma (Hajłasz 2012, Lemma 3.3, p. 57). For \(\left\|a\right\|_{\ell^{r_t}}=\left\|g\right\|_{L^{p_t'}(\nu)}=1\), initially with \(g\) simple, interpolate their nonzero magnitudes by \[a_i(z)=\operatorname{sgn}(a_i) |a_i|^{r_t(1-z+z/r_4)}, \qquad g(z)=\operatorname{sgn}(g) |g|^{p_t'(1-z+3z/4)}.\] Here \(\operatorname{sgn}\) denotes complex phase, and the phase of \(g\) is chosen for a bilinear dual pairing. The zero values stay zero. On the two vertical boundaries, these functions have respective unit norms in \((\ell^1,L^1)\) and \((\ell^{r_4},L^{4/3})\). The analytic function \(\int (Ta(z))g(z)\,\,\mathrm d\nu\) is thus bounded by one there and at \(z=t\). Finite sums and simple functions justify analyticity and boundedness on the strip; density and duality remove the simplicity restriction. Taking \(t=4/p\) gives \[\frac1{r_t} =1-\frac{1+\theta}{p} =1-\frac{\alpha-1}{p},\] which is precisely \(r_t=r_*\). This proves the scalar form of (139) for every finite \(p\ge4\).

Hilbert-valued coefficients. Finally write the Hilbert space as \(\ell^2\) in an orthonormal basis, and let \(F_\xi(t)=\sum_i a_{i,\xi}e^{i\lambda_{i,\xi}t}\). Minkowski’s inequality, the scalar result, and Minkowski’s inequality again give \[\begin{align*} \left\|F\right\|_{L^p(\nu;\ell^2)} &\le\left(\sum_\xi\left\|F_\xi\right\|_{L^p(\nu)}^2\right)^{1/2}\\ &\le\left(\sum_\xi \left(\sum_i|a_{i,\xi}|^{r_*}\right)^{2/r_*} \right)^{1/2}\\ &\le\left(\sum_i \left(\sum_\xi|a_{i,\xi}|^2\right)^{r_*/2} \right)^{1/r_*}. \end{align*}\] The first inequality uses \(p\ge2\), and the last uses \(r_*\le2\). Each has constant one and is independent of the Hilbert dimension. Only the frequency list within an individual component enters the scalar estimate, completing the proof. ◻

Angular hierarchy and predicted factors

We continue with the finite popular forest and the set \(Z_0\) from 30, on which the top estimates, the bound (112), and the separated-top count hold. The parameters \(p,G,K,B\) are defined in (85), and \(\alpha,\mu,r_*\) in (138). We use \(i\) for individual angular levels and \(j\) for blocks of levels. Each shift chooses different block endpoints. The analytic family constructed below is shared by all \(K\) groupings. Its block identities explain how the phase estimate will compare a parent amplitude with its children, while its equal-coordinate value leaves the boundary weights removed in the final angular-grid average. We then prove the estimates that predict amplitudes at a finer angular level, up to a small error, from one good point on a much longer horizontal interval. No mass cutoff is used in those estimates.

All sums below are finite. We retain the strictly positive spatial weights \(h_T\) from (76); in particular, \[ 0<h_T\le C_3, \qquad \int_{\mathbb R^2}h_T\,\,\mathrm dz\le L^C |R_T|. \tag{147}\] The integral follows by the change of variables to the normalized coordinates of \(R_T\), since the decay exponent in that definition is greater than two. Its constant includes the fixed polynomial dilation \(K_0^2\). Constants denoted by \(L^C\) in this section are independent of the number of tiles, tops, levels, and Hilbert-space components.

Eligible masses and shifted blocks

Choose a dyadic number \(d(0)\) such that \(B^2\le d(0)<2B^2\), enlarging it by a fixed absolute factor if necessary to cover the initial precision range. For a translation parameter \(\vartheta\in[0,d(0))\), let \[ d(i)=d(0)G^{-i}, \qquad \mathcal G_i^\vartheta =\bigl\{[\vartheta+q d(i),\vartheta+(q+1)d(i)):q\in\mathbb Z\bigr\}. \tag{148}\] Since \(G\) is an integer dyadic power, these grids are nested. We fix \(\vartheta\) for now. The top slopes lie in a bounded interval, so only an absolute number of intervals in \(\mathcal G_0^\vartheta\) contain a top slope. These intervals are the roots. The half-open convention assigns each top \(T\) a unique interval \(v_i(T)\) at every level \(i\).

For \(v\in\mathcal G_i^\vartheta\) put \[ \mathcal T(v) =\{T:s_T\in v,\ d_T\le d(i)/B^2\}, \qquad n_v(z)=\sum_{T\in\mathcal T(v)}h_T(z). \tag{149}\] We call these tops eligible at \(v\). A node is nonempty if \(\mathcal T(v)\ne\varnothing\). For such a node \(n_v(z)>0\) at every spatial point. An empty node has mass and amplitude zero and is omitted from every expression involving a negative power or logarithm of its mass. Thus none of the logarithms below requires a spatial lower cutoff.

If \(v_1,\ldots,v_G\) are the immediate children of \(b\), then their eligible top sets are disjoint subsets of \(\mathcal T(b)\). More precisely, \[ n_b-\sum_{v\text{ child of }b}n_v =\sum_{\substack{T\in\mathcal T(b)\\d_T>d(i+1)/B^2}}h_T\ge0, \qquad b\in\mathcal G_i^\vartheta. \tag{150}\] The same assertion holds between any two levels, with the corresponding eligibility cutoff. In particular, every positive mass difference used below is itself a sum of the same positive functions \(h_T\).

For a packet \(s\) and a top \(T\), define their last eligible levels by \[ k_s=\max\{i\ge0:d_s\le d(i)/B^2\}, \qquad k_T=\max\{i\ge0:d_T\le d(i)/B^2\}. \tag{151}\] The initial precision choice makes these sets nonempty. The assigned-top relation \(\ell_s\le l_{T(s)}\) gives \(d_{T(s)}\le d_s\), and therefore \(k_s\le k_{T(s)}\). Thus a packet is stopped only while its assigned top is still eligible. Since the collection is finite and all precisions are positive, the maximum of the \(k_T\) is finite. The grids may be defined at every level, but no amplitude requires a level beyond this maximum. We also keep the first block endpoint past this maximum, so that terminal blocks have the same notation as other blocks.

For each \(\sigma\in\{0,\ldots,K-1\}\) use the endpoints \[ i_0^\sigma=0, \qquad i_j^\sigma=1+\sigma+(j-1)K\quad(j\ge1). \tag{152}\] The first block has between one and \(K\) transitions; all subsequent blocks have \(K\). The transition \(i\to i+1\) is the final transition of a block for precisely the shift \(\sigma\equiv i\pmod K\). For this shift, let \(a(i)\) be the starting level of that block. Given an immediate child \(v\in\mathcal G_{i+1}^\vartheta\) with immediate parent \(b\in\mathcal G_i^\vartheta\), denote its ancestor at level \(a(i)\) by \(J(i,v)\). Define \[ m_v=n_b+ \sum_{\substack{v'\in\mathcal G_{i+1}^\vartheta, \ v'\subset J(i,v),\ v'\not\subset b\\ \mathop{\mathrm{dist}}(v,v')<G^{-1/4}d(i)}}n_{v'}. \tag{153}\] Here the distance is the distance between the angular intervals. The summands outside \(b\) have disjoint top sets, and these sets are disjoint from \(\mathcal T(b)\). They all consist of tops eligible at \(J(i,v)\). Consequently, \[ n_v\le n_b\le m_v\le n_{J(i,v)}. \tag{154}\] In particular, \(m_v-n_v\) is a sum of positive top weights. The use of the larger block ancestor in (153) never changes the radius of the neighbor test: that radius is always tied to the immediate transition scale \(d(i)\). 1 shows the shifted block endpoints and the top sets entering \(m_v\).

Shifted blocks and boundary masses, schematically. The upper diagram uses \(K=3\). Each shift starts with \(1+\sigma\) transitions and then uses blocks of \(K\) transitions; colored arrows are the final transitions for that shift. At a final transition \(i\to i+1\), the mass \(m_v\) includes all of \(n_b\) and the masses of neighboring children outside \(b\) but inside \(J\). The neighbor radius is determined by the immediate scale \(d(i)\), even when \(J\) begins much earlier. Neither diagram is to scale.

For \(\zeta=(\zeta_0,\ldots,\zeta_{K-1})\in\mathbb C^K\) and nonempty \(v\) at level \(i+1\), put \[ w_v(\zeta,z) =\exp\!\left(\zeta_{i\bmod K}\log\frac{n_v(z)}{m_v(z)}\right). \tag{155}\] All logarithms in this formula are real logarithms of positive numbers. Thus the weights are entire functions of \(\zeta\) and have modulus at most one when all \(\Re\zeta_\sigma\ge0\).

The analytic family and its block comparisons

For a nonempty node \(v\) at level \(h\), define the \(\mathcal H\)-valued analytic amplitude by \[ A_v(\zeta,z) =\sum_{\substack{s:s_{T(s)}\in v\\ k_s\ge h}} q_{T(s)}(z)\phi_s(z)\, n_{v_{k_s}(T(s))}(z)^{-\mu/K} \prod_{i=h}^{k_s-1}w_{v_{i+1}(T(s))}(\zeta,z). \tag{156}\] The condition \(k_s\ge h\) already ensures eligibility of \(T(s)\) at \(v\). The empty product is one. This single definition is used for every shift. At fixed \(z\), it is a finite sum of entire \(\mathcal H\)-valued functions of \(\zeta\).

We first exhibit two algebraic uses of this common family. For the block comparison, temporarily fix a shift \(\sigma\) and parameters satisfying \[ \Re\zeta_\sigma=\mu, \qquad \Re\zeta_{\sigma'}=0\quad(\sigma'\ne\sigma), \qquad \left\|\Im\zeta\right\|_2\le L^2. \tag{157}\] We refer to this choice of real parts as the vertex corresponding to \(\sigma\). Write \(i_j=i_j^\sigma\) and \(\rho_j=d(i_j)\). Thus \[ G\le\frac{\rho_j}{\rho_{j+1}}\le B, \tag{158}\] including the first, possibly shorter, block.

For a nonempty \(J\in\mathcal G_{i_j}^\vartheta\), call its nonempty descendants at level \(i_{j+1}\) its block children. For such a block child \(c\), all packets entering \(A_c\) have the same transition factor from \(J\) to \(c\). Denote it by \(\lambda_c=\prod_{i=i_j}^{i_{j+1}-1}w_{v_{i+1}}\), along that descendant path. Only the last transition of this block has a nonzero real exponent. Separating packets according to whether their stopping level precedes \(i_{j+1}\) therefore gives the exact identity \[ A_J=\sum_{\substack{c\subset J\\c\in\mathcal G_{i_{j+1}}^\vartheta}} \lambda_c A_c+B_J, \qquad |\lambda_c|=\left(\frac{n_c}{m_c}\right)^\mu. \tag{159}\] The remainder \(B_J\) is precisely the sum in (156) over \(i_j\le k_s<i_{j+1}\). Its assigned tops need not lose eligibility at the end of the block. Thus residual packets and lost top mass are different parts of the block decomposition.

For each nonempty \(J\), we add the mass to the normalized \(p\)-th power of the amplitude. The resulting positive amount \(V_J\) and its multiplicative comparison factor \(\beta_J\) are \[ V_J=n_J+n_J^{\alpha-p}\left\|A_J\right\|_{\mathcal H}^p, \qquad \beta_J= \frac{V_J}{n_J+\sum_c n_c^{\alpha-p}\left\|A_c\right\|_{\mathcal H}^p}. \tag{160}\] We use \(\beta_J^0\) for the same ratio with \(B_J\) omitted only in the numerator. The mass baseline makes \(V_J\) and every denominator positive, including when all its block-child amplitudes vanish. The sum is over the nonempty descendants at the end of the block. Equivalently, the denominator is the lost eligibility mass \(n_J-\sum_c n_c\) plus the block-child amounts \(\sum_c V_c\). This is the comparison later used in the exact path expansion.

At any one spatial point, abbreviate \(n=n_J\) and normalize the child amplitudes by \[ l_c=\frac{n_c}{n},\qquad X_c=n_c^{\mu-1}A_c,\qquad \beta_J^0= \frac{1+\left\|\sum_c l_c^{1-\mu}\lambda_c X_c\right\|_{\mathcal H}^p} {1+\sum_c l_c\left\|X_c\right\|_{\mathcal H}^p}. \tag{161}\] The ratio follows by dividing (160) by \(n\) and using \(p\mu=\alpha-1\). Also \(\sum_c l_c\le1\). This is a pointwise identity: the masses, scalar factors, and vectors still vary with the spatial point. The vector sum in its numerator is the weighted child sum whose phase motion we will compare with the weighted child energies in the denominator.

The boundary mass \(m_c\) is designed to pay for selecting separated children in that comparison. At the final transition of this block put \(i=i_{j+1}-1\) and \[ Q=G^{-1/4}d(i)=G^{3/4}\rho_{j+1}. \tag{162}\] Fix one spatial point and choose any finite subfamily \(\mathcal C\) of nonempty block children. Join distinct \(c,c'\in\mathcal C\) when \(\mathop{\mathrm{dist}}(c,c')<Q\). Give the children independent exponential clocks of rates \(n_c\), and retain a child when its clock rings before those of all its neighbors. The rates are positive. The clocks have no ties with probability one, and the retained intervals have pairwise distance at least \(Q\). The retention probability of \(c\) is \[ \pi_c=\int_0^\infty n_c e^{-n_ct} \prod_{c'\sim c}e^{-n_{c'}t}\,\,\mathrm dt =\frac{n_c}{n_c+\sum_{c'\sim c}n_{c'}} \ge\frac{n_c}{m_c}. \tag{163}\] To see the inequality, let \(b\) be the immediate parent of \(c\). The eligible top sets of \(c\) and its competitors inside \(b\) are disjoint subsets of \(\mathcal T(b)\), so their masses sum to at most \(n_b\). Every competitor outside \(b\) lies inside \(J\) and within \(Q=G^{-1/4}d(i)\) of \(c\), so its mass is included among the additional terms of \(m_c\) in (153). Thus \(m_c\) bounds the entire competing mass even when only a subfamily of children is used.

Multiplying the coefficient of a retained child by \(\pi_c^{-1}\) reconstructs its original coefficient in expectation. The corresponding expected \(r_*\)-power cost is \[\mathbb E\!\left[\mathbf 1_{\{c\text{ retained}\}}\pi_c^{-1}\right]=1, \qquad \mathbb E\!\left[\mathbf 1_{\{c\text{ retained}\}}\pi_c^{-r_*}\right] =\pi_c^{1-r_*}.\] The coefficient exponent in (159) matches this cost exactly. The identities \(r_*(1-\mu)=1\) and \(\mu r_*=r_*-1\) give \[ \pi_c^{1-r_*}|\lambda_c|^{r_*} =\left(\frac{n_c/m_c}{\pi_c}\right)^{r_*-1}\le1. \tag{164}\] The first identity also turns \(l_c^{r_*(1-\mu)}\) into \(l_c\), the weight in the denominator of (161). The later snapshot construction will use this selection for the child vectors it retains and check the remaining hypotheses of the finite Fourier estimate.

The same analytic family has a second algebraic use at the equal-coordinate parameter \[\zeta^{\mathrm c}=(\mu/K,\ldots,\mu/K).\] Let \(v\) be a root and let \(s\) be a packet contributing to \(A_v\), with \(T=T(s)\). Apart from the factor \(q_T\phi_s\), the coefficient of \(s\) in \(n_v^{\mu/K}A_v(\zeta^{\mathrm c},z)\) is \[n_v^{\mu/K}n_{v_{k_s}(T)}^{-\mu/K} \prod_{i=0}^{k_s-1} \left(\frac{n_{v_{i+1}(T)}}{m_{v_{i+1}(T)}}\right)^{\mu/K}.\] Cancelling consecutive eligible masses identifies it as \[ \kappa_s =\prod_{i=0}^{k_s-1} \left(\frac{n_{v_i(T)}}{m_{v_{i+1}(T)}}\right)^{\mu/K}, \qquad T=T(s). \tag{165}\] The empty product is one, and each factor lies in \((0,1]\) by (153). Thus for each root the equal-coordinate value produces its portion of the masked packet sum with only these boundary ratios in its coefficients. The remaining tasks are to estimate the vertex values, interpolate those estimates to \(\zeta^{\mathrm c}\), and remove the boundary ratios by averaging the angular-grid translation.

Full-root variation and amplitude bounds

We now leave the temporary vertex specialization and return to the general range \(0\le\Re\zeta_\sigma\le\mu\) for every coordinate, with arbitrary imaginary parts. The following bounds control the endpoint weights without a loss proportional to the number of levels. Along one eligible top path, write \(n_i=n_{v_i(T)}\) and \[ P_i=\log\frac{n_i}{n_{i+1}}, \qquad R_i=\log\frac{m_{v_{i+1}(T)}}{n_{i+1}}. \tag{166}\] Both are nonnegative as long as the top remains eligible at level \(i+1\).

Lemma 34 (Full-root variation bound). Fix an eligible path from a root \(v_0(T)\) through level \(k\). At every spatial point, \[ \sum_{i=0}^{k-1}R_i \le K\log\frac{n_{v_0(T)}}{n_{v_k(T)}}, \qquad \sum_{i=0}^{k-1}P_i =\log\frac{n_{v_0(T)}}{n_{v_k(T)}}. \tag{167}\] For a fixed starting level \(h\le k_T\), define, for \(h\le k\le k_T\), \[\omega_{T,h,k}(\zeta,z) =n_{v_k(T)}(z)^{-\mu/K} \prod_{i=h}^{k-1}w_{v_{i+1}(T)}(\zeta,z).\] If \(0\le\Re\zeta_\sigma\le\mu\) for every \(\sigma\), then the supremum plus total variation of this sequence, with the endpoint \(k\) varying, is bounded by \[ \sup_{h\le k\le k_T}|\omega_{T,h,k}| +\sum_{k=h}^{k_T-1} |\omega_{T,h,k+1}-\omega_{T,h,k}| \le C h_T^{-\mu/K} \left[1+K(1+\left\|\zeta\right\|_2) \log\frac{n_{v_0(T)}}{h_T}\right]. \tag{168}\] The same bound holds upon restricting the endpoints to a subinterval.

Proof. Fix a shift \(\sigma\). For the transitions \(i\equiv\sigma\pmod K\) with \(i<k\), (154) gives \[R_i\le\log\frac{n_{a(i)}}{n_{i+1}}.\] For this fixed shift, the intervals of levels \([a(i),i+1]\) are successive blocks starting at zero. Their logarithms telescope. The endpoint of the last completed block is at most \(k\), so the sum is at most \(\log(n_0/n_k)\), by monotonicity of the masses. Summing over the \(K\) shifts proves the first inequality in (167). The second is an ordinary telescope. This is a full-root bound. In particular, it remains applicable to weights occurring on a later segment even when another shift’s block starts before that segment.

Set \(a=\mu/K\), solely for this proof, and use the specified logarithm \[L_k=-a\log n_k-\sum_{i=h}^{k-1}\zeta_{i\bmod K}R_i, \qquad \omega_{T,h,k}=e^{L_k}.\] Its increment is \[ L_{k+1}-L_k=aP_k-\zeta_{k\bmod K}R_k. \tag{169}\] Since \(n_k\ge h_T\) and all real transition exponents are nonnegative, \(|e^{L_k}|\le h_T^{-a}\). For any complex \(A,B\), \[|e^B-e^A| \le |B-A|\int_0^1|e^{(1-t)A+tB}|\,\,\mathrm dt \le |B-A|\max(|e^A|,|e^B|).\] Using this inequality for adjacent logarithms and then (167) proves (168). Notice that the exponential of the sum of the absolute logarithmic increments has not been used. Such an exponential would be unsuitable when \(h_T\) is very small. ◻

Lemma 35 (Amplitude bounds). On \(Z_0\), for \(0\le\Re\zeta_\sigma\le\mu\), \[ \left\|A_v(\zeta,z)\right\|_{\mathcal H} \le L^C(1+\left\|\zeta\right\|_2)n_v(z). \tag{170}\] More precisely, the contribution of any one eligible top \(T\) to \(A_v\), also when restricted to any interval of packet lengths, is bounded by \[ L^C(1+\left\|\zeta\right\|_2)|q_T(z)|h_T(z)^4. \tag{171}\] If all eligible tops of \(v\) are insignificant at \(z\in Z_0\), then \[ \left\|A_v(\zeta,z)\right\|_{\mathcal H} \le L^C(1+\left\|\zeta\right\|_2)e^{-L\log L}n_v(z). \tag{172}\]

Proof. For vectors \(f_h,\ldots,f_k\) and scalars \(\omega_h,\ldots,\omega_k\), summation by parts gives \[ \left\|\sum_{i=h}^k\omega_i f_i\right\| \le \left(|\omega_k|+\sum_{i=h}^{k-1}|\omega_{i+1}-\omega_i|\right) \max_{h\le j\le k}\left\|\sum_{i=h}^j f_i\right\|. \tag{173}\] Group the packets assigned to \(T\) according to their stopping level. The resulting initial sums, or the sums obtained after restricting to an interval of packet lengths, are interval partial sums of the original top’s length sequence. By (110) and the fact that the \(K_0\)-dilated weight is at least the corresponding \(K_1\)-dilated weight, their norm at \(z\in Z_0\) is at most \(L^C h_T(z)^8\). The mask \(q_T(z)\) is common to this whole top. Apply (173) and (168).

To see that the logarithmic factor can be absorbed uniformly, use (112) and \(\log N\le2L\). Since \(n_{v_0(T)}\le n\) and \(h_T\le C_3\), \[0\le\log\frac{n_{v_0(T)}}{h_T} \le C L+\log^+\frac1{h_T}.\] Here and below \(\log^+x=\max(0,\log x)\). Also \(\mu/K<1\). The elementary boundedness of \(x^\delta\log^q(1/x)\) on \((0,1]\), for fixed \(\delta>0\) and fixed \(q\), shows that \[h_T^{8-\mu/K} \left[1+K(1+\left\|\zeta\right\|_2) \log\frac{n_{v_0(T)}}{h_T}\right] \le L^C(1+\left\|\zeta\right\|_2)h_T^4.\] This proves (171). Summing over the eligible tops and using \(h_T^4\le C_3^3h_T\) proves (170). For insignificant tops, the mask estimate after (87) gives \(|q_T|\le e^{-L\log L}\). Applying the same summation proves (172). ◻

Spatial meshes for prediction

Return to a fixed shift \(\sigma\) and a vertex parameter satisfying (157). Write \(i_j=i_j^\sigma\) and \(\rho_j=d(i_j)\) as in the block comparison above, and take \(J\in\mathcal G_{i_j}^\vartheta\). Its block children and their factors are those in (159). Let \(R\) be any top eligible in \(J\), and parametrize a line of slope \(s_R\) by \(z(x)=(x,y_*+s_Rx)\). On this line use one nested dyadic grid in the parameter \(x\), at lengths \[ H_j=\frac{w}{\rho_{j+1}\sqrt G}, \qquad H_{-1}=\frac{w}{\rho_0\sqrt G}. \tag{174}\] These are dyadic lengths because \(w,d(0),G\), and \(\sqrt G\) are dyadic. They increase with \(j\). The grids can, for example, have common origin zero. Each \(H_j\) cell is a union of \(H_{j-1}\) cells.

Positive-mass motion and the locality of other shifts

Lemma 36 (Mass motion). Fix \(R,J,j\) as above and an interval \(I\) of horizontal length \(H_j\) on a line of slope \(s_R\). Consider the masses \(n_J,n_c,m_c\) in (159), the masses used in the factors \(\lambda_c\), and those used in any endpoint multiplier of a packet contributing to \(A_c\). Each such mass, and each of its positive subset-complement differences used below, is a sum of weights of tops \(T\) satisfying \[ |s_T-s_R|\le C\rho_j, \qquad d_T\le\rho_j/B^2. \tag{175}\] There is an absolute constant \(C\) such that, with \(\varepsilon_{\mathrm{mov}}=C G^{-1/2}\), every nonzero such mass \(a\) satisfies \[ |\partial_x\log a|\le\frac{\varepsilon_{\mathrm{mov}}}{H_j} \quad\text{almost everywhere on }I. \tag{176}\] If \(r=a/(a+b)\) is one of the ratios of a submass to its containing mass, then \[ |\partial_x\log r| \le\frac{2\varepsilon_{\mathrm{mov}}}{H_j}(1-r) \le\frac{2\varepsilon_{\mathrm{mov}}}{H_j}(-\log r). \tag{177}\] When \(b=0\), the ratio is identically one and its derivative is zero. In particular, nonzero \(n_J-n_c\) and \(m_c-n_c\) have the relative motion bound (176). All these masses vary by a factor \(1+O(G^{-1/2})\) across \(I\).

For the parameters (157), the block factors also satisfy, for \(x,x_0\in I\), \[ \left|\frac{\lambda_c(x)}{\lambda_c(x_0)}-1\right| \le G^{-c_0} \tag{178}\] for some absolute \(c_0>0\) and all sufficiently large \(L\).

Proof. All descendants of \(J\) have slopes inside \(J\), so their top slopes are within \(\rho_j\) of \(s_R\). Their eligibility cutoffs are at most \(\rho_j/B^2\). A mass in an additional boundary term (153) at a transition \(i\ge i_j\) can come from an interval outside \(J\), because the block associated to another shift can start before \(i_j\). However, that interval is within \(G^{-1/4}d(i)\) of the relevant child and has width \(d(i+1)\). The relevant child is a descendant of \(J\), and \(d(i)\le\rho_j\). Hence every added slope is within \(C\rho_j\) of \(s_R\). The tops in an added interval have precision at most \(d(i+1)/B^2\); those in the immediate-parent mass have precision at most \(d(i)/B^2\). This proves (175), including for weights below the block. Taking a difference of the relevant nested top sets preserves the same property. This argument uses the neighbor radius at the immediate transition, rather than the possibly larger width of the other shift’s block ancestor.

In normalized coordinates of a top \(T\), horizontal motion along \(s_R\) has velocities \(1/l_T\) and \((s_R-s_T)/W\). Differentiating (76), using the Lipschitz property of the absolute value at almost every point, gives \[|\partial_x\log h_T| \le C\left(\frac1{l_T}+\frac{|s_R-s_T|}{W}\right) \le C\left(\frac{\rho_j}{B^2w}+\frac{\rho_j}{mw}\right).\] The factor \(K_0^{-1}\) available from the dilation is not needed. Multiplying by \(H_j\) and using (158) yields \[H_j|\partial_x\log h_T| \le\frac C{\sqrt G} \left(\frac{B}{B^2}+\frac Bm\right) \le C G^{-1/2}.\] For a positive sum \(a=\sum_T h_T\), its logarithmic derivative is the weighted average of the logarithmic derivatives of the summands. This proves (176), also for positive differences, which are such sums by their set interpretation.

If \(a,b>0\), direct differentiation gives \[\partial_x\log\frac a{a+b} =\frac b{a+b}\bigl(\partial_x\log a-\partial_x\log b\bigr).\] This proves (177); the inequality \(1-r\le-\log r\) holds for \(0<r\le1\). If the complementary top set is empty, \(r=1\) everywhere. In particular, writing \(q_r=-\log r\ge0\), \[ |\partial_x q_r|\le\frac{2\varepsilon_{\mathrm{mov}}}{H_j}q_r, \qquad e^{-2\varepsilon_{\mathrm{mov}}}q_r(x_0) \le q_r(x)\le e^{2\varepsilon_{\mathrm{mov}}}q_r(x_0) \quad(x,x_0\in I). \tag{179}\] The inequalities remain valid for an identically zero loss.

Finally, a block factor is a product of at most \(K\) transition factors. The simpler bound \(|\partial_x\log r|\le2\varepsilon_{\mathrm{mov}}/H_j\) shows that its specified complex logarithm changes by at most \(C K(1+\left\|\zeta\right\|_2)G^{-1/2}\). For (157) this is \(L^C G^{-1/2}\) and tends to zero faster than any negative fixed power of \(L\). Exponentiating proves (178) after decreasing \(c_0\). This last argument does not require any bound on the size of a logarithmic loss itself. ◻

Variation of moving endpoint weights

The next estimate makes explicit why the motion of the analytic weights can be controlled from a single point, even for very small \(h_T\). For a finite scalar sequence \(a=(a_k)\) let \[\left\|a\right\|_{\mathrm{BV}} =\sup_k|a_k|+\sum_k|a_{k+1}-a_k|.\] Only this elementary finite-sequence norm is used here.

Lemma 37 (Moving weights). In the setting of 36, fix a top \(T\) contributing to \(A_c\), and take endpoint weights \(\omega_k=\omega_{T,i_{j+1},k}\). If \(z_0\in I\cap Z_0\) and the parameters satisfy (157), then, uniformly for \(z'\in I\), \[\begin{align*} \left\|(\omega_k(z'))_k\right\|_{\mathrm{BV}} &\le L^C h_T(z_0)^{-4}, \tag{180}\\ \left\|(\omega_k(z')-\omega_k(z_0))_k\right\|_{\mathrm{BV}} &\le G^{-1/2}L^C h_T(z_0)^{-4}. \tag{181}\end{align*}\] Both conclusions hold for any interval of endpoints. Neither assumes that \(z'\) belongs to \(Z_0\).

Proof. Parametrize the interval by its horizontal coordinate \(x\), and write \(h=i_{j+1}\) and \(a=\mu/K\) in this proof. Use \(P_i,R_i\) from (166) for \(h\le i<k_T\). By (179), each of these nonnegative losses is comparable throughout \(I\) to its value at \(x_0\), with factor \(e^{O(\varepsilon_{\mathrm{mov}})}\), and its absolute derivative is bounded by \(C\varepsilon_{\mathrm{mov}}/H_j\) times itself. Define \[\mathcal B(x)=\sum_{i=h}^{k_T-1} \bigl(aP_i(x)+|\zeta_{i\bmod K}|R_i(x)\bigr).\] It follows that \(\mathcal B(x)\le e^{C\varepsilon_{\mathrm{mov}}}\mathcal B(x_0)\). By the full-root bound, \[ \mathcal B(x_0) \le C K(1+\left\|\zeta\right\|_2) \log\frac{n_{v_0(T)}(z_0)}{h_T(z_0)}. \tag{182}\] The full-root mass is used only at the snapshot in this inequality; its spatial derivative is not needed.

Let \(L_k=\log\omega_k\) be the specified logarithm in the proof of 34, and put \(q_k=\partial_x L_k\). Equation (169) and the mass derivative bound imply \[\begin{align*} \sup_k|q_k| &\le\frac{C\varepsilon_{\mathrm{mov}}}{H_j} \bigl(a+\mathcal B(x)\bigr),\\ \sum_k|q_{k+1}-q_k| &\le\frac{C\varepsilon_{\mathrm{mov}}}{H_j}\mathcal B(x). \end{align*}\] For example, the initial term is \(q_h=-a\partial_x\log n_h\), and subsequent differences are the derivatives of \(aP_i-\zeta_{i\bmod K}R_i\). At every point of \(I\), \[\sup_k|\omega_k(x)|\le h_T(x)^{-a} \le e^{a\varepsilon_{\mathrm{mov}}}h_T(z_0)^{-a}.\] The complex-exponential estimate used for (168) therefore gives \[\left\|\omega(x)\right\|_{\mathrm{BV}} \le C h_T(z_0)^{-a}\bigl(1+\mathcal B(x_0)\bigr).\]

For finite sequences the product rule for differences gives \[\left\|ab\right\|_{\mathrm{BV}} \le \sup_k|a_k|\left\|b\right\|_{\mathrm{BV}} +\sup_k|b_k|\sum_k|a_{k+1}-a_k|.\] Using \(\partial_x\omega_k=\omega_k q_k\) and the preceding estimates yields \[ \left\|\partial_x\omega(x)\right\|_{\mathrm{BV}} \le\frac{C\varepsilon_{\mathrm{mov}}}{H_j} h_T(z_0)^{-a}\bigl(1+\mathcal B(x_0)\bigr)^2. \tag{183}\] The functions involved are absolutely continuous on the line. Integrate (183) between \(x_0\) and \(x'\), over distance at most \(H_j\), and use the triangle inequality for the finite-sequence norm. This proves the difference bound with right side \(C\varepsilon_{\mathrm{mov}} h_T(z_0)^{-a}(1+\mathcal B(x_0))^2\).

At the snapshot, (112), (157), and (182) bound the last expression by \(G^{-1/2}L^C h_T(z_0)^{-4}\). Indeed the only possible nonpolynomial quantity is \(\log^+(1/h_T(z_0))\), whose square is absorbed by the difference between exponents \(a<1\) and \(4\). The same argument proves (180). There is no assumption that \(G^{-1/2}\mathcal B(x_0)\) is small: very small \(h_T\) is allowed, and only a polynomial in its logarithm has appeared. ◻

Amplitude motion from one good snapshot

Lemma 38 (Predictable amplitude motion). Fix a block child \(c\) of \(J\), a path top \(R\) eligible in \(J\), and an \(H_j\) cell on a line of slope \(s_R\). Let \(t_c\) be the angular center of \(c\). If \(z_0\in Z_0\) lies in that cell and \(z'=z_0+(b,s_Rb)\) is any other point in the cell, then, for the parameters (157), \[ \left\|A_c(\zeta,z')-\mathcal D_c(b)A_c(\zeta,z_0)\right\|_{\mathcal H} \le G^{-c_0}n_c(z_0), \qquad \mathcal D_c(b)_{\nu\nu} =\exp\bigl(i(s_R-t_c)\eta_\nu b\bigr). \tag{184}\] Here \(c_0>0\) is absolute and may also be used, after decreasing it, in (178). The estimates hold at all displaced points in the cell, not just its points in \(Z_0\).

Proof. Fix one top \(T\) contributing to \(A_c\). Its contributing packets have \(k_s\ge i_{j+1}\), hence \[ \ell_s\ge l_*:=\frac{B^2w}{\rho_{j+1}}, \qquad l_T\ge l_*. \tag{185}\] Also \(s_T\in c\), so \(|s_T-t_c|\le\rho_{j+1}/2\), while \(|s_T-s_R|\le C\rho_j\).

Consider any interval partial sum \(S\) of these unweighted packets on \(T\). The good-snapshot estimates (110), including their demodulated vertical-derivative and vertical-neighborhood versions, bound \(S\) and \(W(\partial_y-i\eta_\nu)S_\nu\) by \(L^Ch_T(z_0)^8\) throughout the vertical neighborhood of radius \(W\) at \(z_0\). Put \(v=(s_R-s_T)b\) and first move to \(z_1=z_0+(0,v)\). Since \[|v|\le C\rho_jH_j\le\frac{CBw}{\sqrt G}\ll W,\] the fundamental theorem of calculus after componentwise demodulation gives \[\left\|S(z_1)-\operatorname{diag}(e^{i\eta_\nu v})S(z_0)\right\| \le L^Ch_T(z_0)^8\frac{|v|}{W} \le G^{-1/2}L^Ch_T(z_0)^8.\] Next move from \(z_1\) to \(z'\) along slope \(s_T\) by horizontal distance \(b\). The nonexceptional horizontal derivative estimate in 28, applied to lengths at least \(l_*\), bounds this change by \[L^Ch_T(z_0)^8\frac{|b|}{l_*} \le G^{-1/2}L^Ch_T(z_0)^8.\] To justify its spatial weight here, the two displacements change the normalized coordinates of \(T\) by at most \(C/(B^2\sqrt G)+CB/(m\sqrt G)\). Its weight at every intermediate point is therefore comparable to \(h_T(z_0)\). If the derivative estimate is expressed using the smaller dilation \(K_1\), that weight is bounded above by the \(K_0\)-dilated weight used here. No good-point estimate at \(z_1\) or \(z'\) has been invoked for this horizontal move.

The replacement of \(s_T\) by \(t_c\) in the diagonal phase costs at most \[C|s_T-t_c|\frac{|b|}{w}\left\|S(z_0)\right\| \le G^{-1/2}L^Ch_T(z_0)^8,\] because \(\eta_\nu\asymp w^{-1}\) uniformly in \(\nu\). This proves, uniformly over the interval partial sum, \[ \left\|S(z')-\mathcal D_c(b)S(z_0)\right\| \le G^{-1/2}L^Ch_T(z_0)^8. \tag{186}\]

The same bound holds with the mask \(q_T\) multiplying \(S\) at its own evaluation point. Indeed \(|q_T|\le1\), and its formula (87), together with boundedness of the derivative of sinc, gives \[|q_T(z')-q_T(z_0)| \le L^C\left(\frac{|b|}{l_T} +\frac{|s_R-s_T|\,|b|}{W}\right) \le G^{-1/2}L^C.\] Combining this with the snapshot bound for \(S\) proves the masked version of (186). This estimate uses absolute motion of the mask and does not divide by \(q_T\), which may vanish.

Group the masked packets by stopping level and denote their sums by \(f_k(z)\). Their weighted contribution to the difference in (184) is \[\begin{align*} &\sum_k\omega_k(z') \bigl(f_k(z')-\mathcal D_c(b)f_k(z_0)\bigr)\\ &\hspace{15mm} +\sum_k\bigl(\omega_k(z')-\omega_k(z_0)\bigr) \mathcal D_c(b)f_k(z_0). \end{align*}\] All interval partial sums of the first parenthesis have norm at most \(G^{-1/2}L^Ch_T(z_0)^8\), by the masked version of (186). The snapshot partial sums in the second line have norm at most \(L^Ch_T(z_0)^8\). Apply (173) to both lines, then (180) and (181). The result for this top is at most \(G^{-1/2}L^Ch_T(z_0)^4\). Sum over the tops eligible in \(c\). Since \(h_T^4\le C_3^3h_T\), this is at most \(G^{-1/2}L^Cn_c(z_0)\). Finally \(\log G\asymp L^{0.7}\) absorbs every fixed power of \(L\) into a smaller positive power of \(G\), proving (184). ◻

One deterministic residual sequence for each path

We now refine \(Z_0\) by a set which depends on the angular translation but not on \(\zeta\). Its purpose is to make the later residual estimates available on most of the path weight. The subsets of packets used in its definition must be fixed before evaluating a spatial point or choosing a mass cutoff.

For each possible path top \(R\), let \[ \mathcal P_R =\{s:\ell_s\le l_R,\quad |\theta_s-s_R|\le C B^4d_s\}, \tag{187}\] where the absolute constant is large enough to cover the assigned-top comparisons below. Apply the main/error decomposition of 27 with reference slope \(s_R\) to this fixed collection. Denote the resulting positive error sum by \(\mathcal E_R\) and positive main square function by \(\mathcal S_R\). The absolute envelopes can use fixed decay exponents slow enough to dominate the norms of the unweighted synthesis packets. There is no loss in doing so because the packet bounds are available with arbitrarily large fixed decay exponents.

For each shift \(\sigma\), define a second fixed collection \[ \mathcal Q_{R,\sigma} =\bigcup_{\substack{j\ge0\\ i_{j+1}^\sigma\le k_R}} \left\{s: \begin{array}{l} i_j^\sigma\le k_s<i_{j+1}^\sigma,\\ T(s)\in\mathcal T\bigl(v_{i_{j+1}^\sigma}(R)\bigr) \end{array}\right\}. \tag{188}\] For blocks at whose endpoint \(R\) is still eligible, these are the packets that stop within the block while their assigned tops remain eligible in the child \(v_{i_{j+1}^\sigma}(R)\). The union takes all such blocks at once, fixing one packet collection before any spatial point or mass cutoff is chosen. Let \(\mathcal V_{R,\sigma}\) be the \(3\)-variation of the sequence formed by summing \(q_{T(s)}\phi_s\) from this collection at each packet length. It permits any finite list of disjoint intervals of that one length sequence. Subsequent choices of those intervals may depend on the evaluation point; the collection and its length sequence remain fixed.

Every packet in \(\mathcal Q_{R,\sigma}\) satisfies the hypotheses of 27 with \(t=s_R\) and \(l=l_R\). Indeed, \(k_s<i_{j+1}^\sigma\le k_R\) gives \[ d_s>\frac{\rho_{j+1}}{B^2}\ge d_R, \qquad \ell_s<l_R. \tag{189}\] Since \(s_{T(s)}\) and \(s_R\) are in the same block-child interval of width \(\rho_{j+1}\), \[|\theta_s-s_R| \le |\theta_s-s_{T(s)}|+|s_{T(s)}-s_R| \le C d_s+\rho_{j+1}\le C B^2d_s\le C B^4d_s.\] For later use, if a residual packet only lies in the larger path node \(J=v_{i_j}(R)\) and has \(\ell_s\le l_R\), then \(|s_{T(s)}-s_R|\le\rho_j\le B\rho_{j+1}<B^3d_s\); thus it belongs to \(\mathcal P_R\) as well. These facts explain both choices of fixed collections.

Lemma 39 (Good paths). There is a fixed exponent \(C_{\mathrm{path}}\), chosen independently of the mass cutoff below, such that the following construction is valid. Call \(R\) bad at \(z\) if any of \(\mathcal E_R(z),\mathcal S_R(z),\mathcal V_{R,\sigma}(z)\), \(0\le\sigma<K\), exceeds \(L^{C_{\mathrm{path}}}\). Otherwise call it good at \(z\). For every fixed angular translation, \[ \sum_R\int_{\mathbb R^2}h_R(z)\mathbf 1_{\{R\text{ bad at }z\}}\,\,\mathrm dz \le S e^{-10L}. \tag{190}\] The constant \(10\) may be replaced by any prescribed absolute constant by increasing the fixed exponent. The measurable set \[ Z=Z_\vartheta =\left\{z\in Z_0: \sum_{R\text{ bad at }z}h_R(z)\le e^{-L}\right\} \tag{191}\] is independent of \(\zeta\) and satisfies \[ |Z_0\setminus Z|\le S e^{-9L} \le L^Ce^{-5L}|F|. \tag{192}\] Its removal is therefore allowable in the target estimate (72), using the original adjoint sum’s \(L^2\) bound.

Proof. Fix \(R\), a line of slope \(s_R\), and any horizontal interval \(I\) of length \(l_R\) on that line. By 27, the normalized \(L^b(I)\) norm of each indicated quantity is at most \(C b^C(1+\log B)^C\). The same bound holds for \(\mathcal V_{R,\sigma}\) because (188) is a fixed subset satisfying (93). As \(\log B=K\log G\) is bounded by a fixed power of \(L\), this bound is polynomial in \(L\) when \(b\) is a fixed positive multiple of \(L\).

For clarity, fix an arbitrarily large absolute number \(J_0\). Choose \(b=\lceil L\rceil\) and then a fixed exponent \(C_{\mathrm{path}}\) sufficiently larger than those in the moment bound. Markov’s inequality gives, for all large \(L\), \[\begin{aligned} &|\{x\in I:\mathcal E_R(x)>L^{C_{\mathrm{path}}}\}| +|\{x\in I:\mathcal S_R(x)>L^{C_{\mathrm{path}}}\}|\\ &\qquad +\sum_{\sigma=0}^{K-1} |\{x\in I:\mathcal V_{R,\sigma}(x)>L^{C_{\mathrm{path}}}\}| \le e^{-J_0L}|I|. \end{aligned}\] The factor \(K+2\) is polynomial and is absorbed by the same choice. Alternatively one may increase the fixed multiple in the choice of \(b\); neither choice depends on the number of scales or tops.

Partition the whole reference line into intervals of length \(l_R\). Along this line \(h_R\) changes by at most an absolute factor on each such interval, directly from (76). Hence the preceding estimate implies \[\int_I h_R\mathbf 1_{\{R\text{ bad}\}}\,\,\mathrm dx \le C e^{-J_0L}\int_I h_R\,\,\mathrm dx.\] Sum over the intervals and integrate in the transverse coordinate \(y-s_Rx\); this change of coordinates has determinant one. Using (147) gives \[\int_{\mathbb R^2}h_R\mathbf 1_{\{R\text{ bad}\}}\,\,\mathrm dz \le L^C e^{-J_0L}|R_R|.\] Sum over \(R\) and take \(J_0\) large enough to obtain (190). This establishes the estimate with its full spatial weight, not only on the central part of \(R\).

Markov’s inequality applied to the nonnegative sum defining (191) gives its first measure bound. Since \(S\le L^C|E|=L^CN^2|F|\) and \(N\le e^{2L}\), the second follows. For example, Cauchy–Schwarz and the previously proved adjoint estimate give \[\int_{Z_0\setminus Z}\left\|P\right\| \le |Z_0\setminus Z|^{1/2}\left\|P\right\|_2 \le L^Ce^{-5L/2}N|F|,\] which is smaller than any required fixed negative power of \(L\) times \(N|F|\). The estimate is applied to the original sum \(P\), as required for exceptional discards.

Finally, all packet collections used here depend only on the finite geometric data, the angular translation, and the shift. They do not depend on \(z\), \(\zeta\), significance, or a mass cutoff. Their finite length variations and the other thresholded quantities are measurable in \(z\). They are measurable jointly with \(\vartheta\) as well: at the finitely many relevant angular levels, membership of each of the finitely many top slopes in its interval changes at finitely many translation endpoints. On the complementary parameter intervals all the relevant membership data are fixed. This also proves the asserted independence from \(\zeta\). ◻

Mass stopping and the exact path expansion

Choose a cutoff \[ M=L^{A_M}, \tag{193}\] where the fixed exponent \(A_M\) will be taken sufficiently large in the residual estimates. Every threshold and exponent in the preceding lemmas has already been fixed without reference to \(A_M\). Statements below that use \(M\) hold for each such fixed exponent after increasing the lower threshold for \(L\).

For convenience set \[ C_M=L^{Cp}(1+M)^{\alpha-1}, \tag{194}\] with a fixed sufficiently large \(C\). Equation (170) and \(p\mu=\alpha-1\) imply that a node of mass \(n_J\le M\) satisfies \[ V_J\le C_M n_J. \tag{195}\] Indeed \(V_J/n_J\le1+L^{Cp}n_J^{\alpha-1}\) for the parameters.

Lemma 40 (Expansion into top paths). Fix \(z\in Z\), a root, a shift, and parameters as above. Starting from a node with \(n_J>M\), expand \[ V_J=\beta_J \left(n_J-\sum_c n_c+\sum_c V_c\right), \tag{196}\] and stop a branch on first reaching a node of mass at most \(M\). At each expanding node, split \(n_J-\sum_c n_c\) into the individual weights of the tops which do not remain eligible at the block endpoint. At a stopped node split \(V_J\) proportionally among its eligible top weights. This gives exactly one terminal path for each eligible root top \(R\), with a terminal amount \(e_R\) satisfying \[ 0<e_R\le C_Mh_R, \qquad V_{\mathrm{root}} =\sum_R e_R\prod_{J\text{ active on }R\text{'s path}}\beta_J. \tag{197}\] If \(n_{\mathrm{root}}>M\), at least one half of this sum comes from tops which are good at \(z\). Consequently, \[ V_{\mathrm{root}}(z) \le 2C_M\sum_{\substack{R\text{ eligible at root}\\R\text{ good at }z}} h_R(z)\prod_{J\text{ active on }R\text{'s path}}\beta_J(z). \tag{198}\] If \(n_{\mathrm{root}}\le M\), (195) applies instead.

Proof. Identity (196) follows directly from (160), because \(V_c-n_c=n_c^{\alpha-p}\left\|A_c\right\|^p\). The term \(n_J-\sum_c n_c\) is a sum of top weights by (150), now between the endpoints of the whole block. Every top either remains eligible in exactly one block child or loses eligibility and terminates at \(J\). The construction therefore branches by disjoint top sets, assigning each top exactly one path. It terminates after finitely many blocks even if no mass stop occurs, since no top is eligible beyond the final level. At such a final expanding node the block-child sum is empty and all tops terminate directly.

If \(R\) terminates directly at an expanding node, its terminal amount is \(h_R\). If it reaches a mass-stopped node \(c\), the terminal amount is \((V_c/n_c)h_R\), which is at most \(C_Mh_R\) by (195). Repeated use of (196) proves the equality in (197).

To verify the claim about bad paths, it is useful to normalize the expansion by actual probabilities. At an active node define \[D_J=n_J-\sum_c n_c+\sum_cV_c =n_J+\sum_c(V_c-n_c)>M.\] Choose a block child \(c\) with probability \(V_c/D_J\), or a directly terminating top \(R\) with probability \(h_R/D_J\). These probabilities sum to one. At a mass-stopped node choose an eligible top with probability \(h_R/n_c\). The probability of a given complete terminal path is exactly its summand in (197) divided by \(V_{\mathrm{root}}\). In fact, for an active edge \(J\to c\), \(V_c/D_J=\beta_JV_c/V_J\), and at a direct terminal edge \(h_R/D_J=\beta_Jh_R/V_J\). The successive \(V_c\) factors cancel; at a mass stop the last factor \(h_R/n_c\) leaves the amount \((V_c/n_c)h_R\). This proves the asserted probability identity without an independence assumption on the amplitudes.

For a directly terminating top, its last conditional probability is \(h_R/D_J\le h_R/M\le h_R\). If a top terminates in the next node by a mass stop, the combined last choice of that child and then that top has conditional probability \[\frac{V_c}{D_J}\frac{h_R}{n_c} \le C_M\frac{h_R}{M}\le C_Mh_R.\] Earlier conditional probabilities are at most one. Thus the unconditional probability of any specified terminal top is at most \(C_Mh_R\). This estimate also covers tops which terminate after many active ancestors; it uses only the last active step. By (191), the total probability of bad terminal tops is at most \[C_M\sum_{R\text{ bad at }z}h_R(z)\le C_Me^{-L}<\frac12\] for all sufficiently large \(L\). Here \(\log C_M=O(p\log L+\log(1+M))=o(L)\) for each fixed \(A_M\). The probability identity now implies that good tops carry at least half of the total amount, and the bound on \(e_R\) proves (198). ◻

The remaining path estimate, proved below in 46, is \[ \int_{Z\cap\{R\text{ good}\}} h_R\prod_{J\text{ active on }R\text{'s path}}\beta_J \le L^{Cp}(1+M)^C\int_{\mathbb R^2}h_R. \tag{199}\] [pred:section,res:section] prove (199) from the estimates established above. Once that estimate is available, (198) and (147) control the integral of \(V_{\mathrm{root}}\) by summing over the eligible root tops.

Path drops and marked spatial cells

Fix a top \(R\) and one shift. Let \(J_R\) be the last block index with \(i_{J_R}\le k_R\). For \(0\le j\le J_R\) put \[ J_j=v_{i_j}(R),\qquad n_j=n_{J_j},\qquad n_{J_R+1}=0,\qquad h_j^*=n_j-n_{j+1},\qquad u_j=\frac{h_j^*}{n_j}. \tag{200}\] For \(j<J_R\), the path block child is \(v_{i_{j+1}}(R)\). For \(j=J_R\) there is no path child: \(R\) ceases to be eligible at the next block endpoint. Other tops in \(J_j\) may still continue, and all their children are retained in (159) and (160). Only the path notation sets the next mass to zero. In particular, the last possible path drop has \(h_j^*=n_j\) and \(u_j=1\).

The actual active indices form the prefix before the first mass at most \(M\), equivalently the indices with \(n_j>M\). A path which stops by mass need not reach its final possible block \(J_R\). At the last active index of such a path we retain the actual eligible next mass in (200); it is not set to zero merely because the expansion stops there.

Lemma 41 (Drops and marking). At a point of \(Z_0\), the active path drops satisfy \[ \sum_{j:\,n_j>M}u_j \le 1+\log^+\frac{n_0}{M}\le C L. \tag{201}\] On each line of slope \(s_R\), mark an \(H_j\) cell \(I\) when \[ \inf_{z\in I}n_j(z)>M, \qquad \sup_{z\in I}h_j^*(z)\ge\frac12. \tag{202}\] Every point of a marked cell is active and has \(h_j^*\ge c\) for an absolute \(c>0\). At an active point in an unmarked cell with \(h_j^*\ge1\), one has \(M<n_j<M+1\) for all sufficiently large \(L\). There is at most one such index on each active path, and its factor is bounded at points of \(Z_0\) by \[ \beta_{J_j}\le L^{Cp}(1+M)^C. \tag{203}\] If an index is inactive at a point, the cell containing that point is unmarked. Finally, even in the last possible block, \[ H_j\le\frac{l_R}{B\sqrt G}. \tag{204}\]

Proof. If there are no active indices there is nothing to prove in (201). Otherwise let \(q\) be the last active index. For \(j<q\), the next mass is also greater than \(M\), and \[u_j=1-\frac{n_{j+1}}{n_j} \le\log\frac{n_j}{n_{j+1}}.\] The last term \(u_q\) is at most one, whether its next mass is positive or zero. Therefore the sum is at most \(1+\log(n_0/n_q)\le1+\log(n_0/M)\). This argument deliberately does not telescope through a possibly tiny mass below the stopping cutoff. Equation (112) gives the last inequality.

Both \(n_j\) and \(h_j^*\) are positive sums of weights of tops eligible at \(J_j\), except that the difference can be identically zero. In the last possible block it is simply \(n_j\). Thus 36 gives relative variation \(1+O(G^{-1/2})\) across the whole \(H_j\) cell. A marked cell has positive supremum of \(h_j^*\), so this difference is not identically zero there, and its infimum is at least \(\tfrac12 e^{-CG^{-1/2}}\ge\tfrac14\) for large \(L\). The first marking condition supplies activity everywhere on the cell.

Suppose instead that \(z\) is active, its cell is unmarked, and \(h_j^*(z)\ge1\). The second marking condition then holds, so the first one fails: the cell has infimum mass at most \(M\). The relative-motion bound implies \[M<n_j(z)\le e^{CG^{-1/2}}M<M+1,\] where the last inequality holds because \(M\) is a fixed polynomial in \(L\). Its drop of at least one gives \(n_{j+1}(z)<M\) (or zero at the last possible block), and all subsequent masses are at most that mass. This is consequently the last active index and can occur at most once. At a point of \(Z_0\), since the denominator of (160) is at least \(n_j\), (170) gives \[\beta_{J_j} \le\frac{V_{J_j}}{n_j} \le1+L^{Cp}n_j^{\alpha-1} \le L^{Cp}(1+M)^C,\] proving (203). If \(n_j(z)\le M\), its cell cannot have infimum greater than \(M\), proving the assertion about inactive indices. This observation will justify assigning the value one to the auxiliary phase factor on every unmarked cell.

Finally, eligibility at the starting level gives \(d_R\le\rho_j/B^2\), equivalently \(l_R\ge B^2w/\rho_j\). Together with (158), this yields \[\frac{H_j}{l_R} \le\frac{\rho_j}{B^2\rho_{j+1}\sqrt G} \le\frac1{B\sqrt G}.\] This only uses eligibility at the start of the block, and therefore applies to the final possible block as well. ◻

All constructions above are finite and measurable. Their constants are uniform in the finite depth, so no lower bound on a nonzero top weight, or on the number of angular levels retained, is part of the argument. The limits from finite packet collections and the angular-translation average are taken only after the subsequent sections have obtained the uniform estimates.

Prediction and averaging on marked cells

Fix a shift at a vertex of the analytic parameter simplex, with imaginary parameters of Euclidean norm at most \(L^2\). Retain the block notation, path top \(R\), and meshes from the preceding construction. Thus \(J\) is the path interval at level \(i_j\), its block children \(c\) have width \(\rho_{j+1}\), and \[G\rho_{j+1}\le\rho_j\le B\rho_{j+1},\qquad H_j=\frac{w}{\rho_{j+1}\sqrt G},\qquad H_{j-1}=\frac{w}{\rho_j\sqrt G}.\] The first inequalities include the initial, possibly shorter, block. The path child is denoted by \(a\) when it exists. At the last path block there is no path child, even when some other tops continue; in that case the relative path drop is \(u_j=1\).

Proposition 42 (Averaging factors on marked cells). There is an absolute \(c_{\rm av}>0\) with the following property, for all sufficiently large \(L\). On a fixed line of slope \(s_R\), each marked \(H_j\)-cell \(I\) meeting \(Z\) admits a nonnegative function \(\Gamma_j\) which is constant on the \(H_{j-1}\)-subcells of \(I\), satisfies \(|I|^{-1}\int_I\Gamma_j\le1\), and obeys \[ \beta_J^0(z)\le \exp\bigl(G^{-c_{\rm av}}u_j(z)\bigr)\Gamma_j(z) \qquad (z\in I). \tag{205}\] Set \(\Gamma_j=1\) on unmarked cells and on cells not meeting \(Z\). The functions can be chosen measurably on each fixed line. No choice jointly measurable in the family of parallel lines is needed.

We first construct a normalized prediction \(\widetilde\beta\) whose average is at most one for suitable probability measures. We compare it with \(\beta_J^0\), then blend it with one to obtain \(\mathcal B\), whose comparison error contains the relative path drop. Finally we pass to uniform cell measure, take the finer-cell supremum \(\mathcal B^\sharp\), and normalize it to obtain \(\Gamma_j\). The losses in these last operations also contain the relative path drop.

Snapshot data and denominator shrinkage

Fix a marked cell \(I\) meeting \(Z\), and choose one snapshot \(z_0\in I\cap Z\). Parameterize the line by horizontal displacement \(b\) from \(z_0\), and also write \(I\) for the translated parameter interval. All unprimed masses and amplitudes below are evaluated at \(z_0\); a prime denotes evaluation at \(z_0+(b,s_Rb)\). Let \(n=n_J(z_0)\) and \(\upsilon=u_j(z_0)\). If the path child exists, then \(\upsilon=1-n_a/n\); otherwise \(\upsilon=1\). The positive-mass motion bounds give \[ \tfrac12\upsilon\le u_j(z)\le2\upsilon\qquad(z\in I) \tag{206}\] for large \(L\). Marking ensures \(\upsilon>0\). Only nonempty children are used. Their masses are positive at every point, since the sets of eligible tops are fixed and their weights are strictly positive.

Use the pointwise normalization (161), with \(l_c\) and \(X_c\) evaluated at the snapshot and the primed quantities evaluated at the displaced point. That identity expresses \(\beta_J^0\) in either set of data. The next construction shrinks snapshot vectors to obtain one-sided denominator comparisons for those children. When \(\upsilon<1/2\), the path vector will instead remain unchanged so that its numerator and denominator variations can be compared together. All vector norms in this section are in \(\mathcal H\). The weights satisfy \(\sum_c l_c\le1\), and, when there is a path child, \(\sum_{c\ne a}l_c\le\upsilon\). When there is none, an off-path sum means the sum over all children.

Choose a fixed exponent \(C_*\), large enough for the estimates below, and put \[ K_*=L^{C_*}(1+N^\mu). \tag{207}\] Equations (112) and (170) imply, at the snapshot, \[ \left\|X_c\right\|\le K_*,\qquad n^\mu\le K_*,\qquad \sum_{c\ne a}l_c^{1-\mu}\left\|X_c\right\|\le K_*\upsilon. \tag{208}\] Enlarging the fixed exponent absorbs absolute constants here and below. For the last estimate use \[l_c^{1-\mu}X_c=n^{\mu-1}A_c =n^\mu l_c(A_c/n_c).\] The sum over all children in this estimate is at most \(K_*\).

Let \(\gamma>0\) be a sufficiently small fixed exponent below the motion exponent in 38, and set \(\varepsilon=G^{-\gamma}\). This choice is independent of the depth, dimension, and parameters \(L,p,D\). Decreasing \(\gamma\), if necessary, absorbs the fixed powers of \(L\) in the following consequences of the motion estimates: \[\begin{align*} \left\|l_c'^{1/p}X_c' -\mathcal D_c(b)l_c^{1/p}X_c\right\| &\le\varepsilon l_c^{1/p}n_c^\mu, \tag{209}\\ \left\|n'^{\mu-1}\lambda_c'A_c' -n^{\mu-1}\lambda_c\mathcal D_c(b)A_c\right\| &\le\varepsilon K_*l_c. \tag{210}\end{align*}\] Both estimates hold throughout \(I\), without a good-point hypothesis at the displaced point. For the scalar rescaling in the first one, write \(l_c^{1/p}X_c=n^{-1/p}n_c^{\alpha/p-1}A_c\). The exponents \(-1/p\) and \(\alpha/p-1\) are absolutely bounded. Their factors have the stated relative mass motion; the error in \(A_c\) is bounded by (184), and its snapshot size is at most \(L^C n_c\). This proves (209). For the second estimate use \(n^{\mu-1}\), the relative complex motion of \(\lambda_c\), \(\left|\lambda_c\right|\le1\), and \(n^{\mu-1}n_c=n^\mu l_c\le K_*l_c\).

Shrink \(A_c\) radially toward zero by a distance at most \(2\varepsilon n_c\), choosing \[\widetilde A_c= \begin{cases} (1-2\varepsilon n_c/\left\|A_c\right\|)_+ A_c,& A_c\ne0,\\ 0,& A_c=0. \end{cases}\] If no eligible top in \(c\) is significant at \(z_0\), set \(\widetilde A_c=0\). This fits within the same error: 35 and the mask suppression give \[\left\|A_c\right\|\le L^C e^{-L\log L}n_c\le\varepsilon n_c\] for such a child. There is one exception: when \(\upsilon<1/2\), leave the path vector unchanged, setting \(\widetilde A_a=A_a\). Define \(\widetilde X_c=n_c^{\mu-1}\widetilde A_c\). For every child to which shrinkage was applied, \[ l_c'\left\|X_c'\right\|^p\ge l_c\left\|\widetilde X_c\right\|^p \qquad(b\in I). \tag{211}\] Indeed radial shrinkage reduces the norm of \(l_c^{1/p}X_c\) by the smaller of its whole norm and \(2\varepsilon l_c^{1/p}n_c^\mu\), whereas (209) permits a reduction of at most \(\varepsilon l_c^{1/p}n_c^\mu\) at a displaced point.

The bounds in (208) continue to hold with \(\widetilde X_c\). Including the shrinkage error in (210) gives \[ \left\|l_c'^{1-\mu}\lambda_c'X_c' -l_c^{1-\mu}\lambda_c\mathcal D_c(b)\widetilde X_c\right\| \le C\varepsilon K_* l_c. \tag{212}\]

Exact averaging of the prediction

On the whole parameter line, define \[ F(b)=\sum_c l_c^{1-\mu}\lambda_c \mathcal D_c(b)\widetilde X_c,\qquad \widetilde\beta(b)= \frac{1+\left\|F(b)\right\|^p} {1+\sum_c l_c\left\|\widetilde X_c\right\|^p}. \tag{213}\] The masses, scalar weights, and vectors are fixed snapshot data. Only the diagonal unitaries \(\mathcal D_c(b)_{\xi\xi}=e^{i(s_R-t_c)\eta_\xi b}\) vary.

Lemma 43 (Fourier-supported averaging). There is a fixed sufficiently small \(\epsilon'>0\) such that, with \[ Q=G^{3/4}\rho_{j+1},\qquad \delta=\epsilon' Q/w, \tag{214}\] every positive probability \(\nu\) satisfying \(\mathop{\mathrm{supp}}\widehat\nu\subset(-\delta,\delta)\) obeys \(\int\widetilde\beta\,\,\mathrm d\nu\le1\).

Proof. Use only children for which \(\widetilde X_c\ne0\); an empty collection gives the assertion immediately. Apply the clock selection from (163) to this finite subfamily at the snapshot, with \(Q\) as in (214). The clock rates are fixed as \(b\) varies. Each retained set has pairwise interval distance at least \(Q\), and the retention probabilities satisfy \(\pi_c\ge n_c/m_c\).

Every used child other than the possible unshrunk path child contains a significant eligible top at \(z_0\). From each retained child other than that possible exception, select one such top. Their slopes are \(Q\)-separated, their precisions satisfy \[d_T\le\rho_{j+1}/B^2\le Q/B^2,\] and their diameter is at most \(\rho_j\le B\rho_{j+1}\le BQ\). The simultaneous count in 30, with \(D\) as chosen in [count:choice-D], bounds this list by \(D/100\). Adding the possible retained path child leaves at most \(D\) members. Clock exclusion separates that extra child too; it does not require a significant representative for that one child.

In each Hilbert component, the retained carrier frequencies are separated by at least \(c Q/w\), with \(c>0\) absolute, since \(\eta_\xi\asymp w^{-1}\). Choose \(\epsilon'\le\epsilon_0c\). For a retained set \(\mathcal S\), let \[F_{\mathcal S}(b)= \sum_{c\in\mathcal S}\pi_c^{-1} l_c^{1-\mu}\lambda_c\mathcal D_c(b)\widetilde X_c.\] Expectation over the clocks gives \(\mathbb E_{\mathcal S}F_{\mathcal S}=F\); no independence of retention indicators is asserted or needed. Put \(r=r_*\). Minkowski in that expectation, (139), and Jensen for the concave function \(t^{1/r}\) give \[\begin{align*} \left\|F\right\|_{L^p(\nu;\mathcal H)} &\le \mathbb E_{\mathcal S} \left(\sum_{c\in\mathcal S} \pi_c^{-r}l_c\left|\lambda_c\right|^r \left\|\widetilde X_c\right\|^r\right)^{1/r} \\ &\le\left(\sum_c\pi_c^{1-r}l_c\left|\lambda_c\right|^r \left\|\widetilde X_c\right\|^r\right)^{1/r} \\ &\le\left(\sum_c l_c\left\|\widetilde X_c\right\|^r\right)^{1/r} \le\left(\sum_c l_c\left\|\widetilde X_c\right\|^p\right)^{1/p}. \tag{215}\end{align*}\] The power of \(l_c\) uses \(r(1-\mu)=1\). The third line uses the retention cancellation (164) for the present child subfamily. The last inequality in (215) is Hölder with weights of total mass at most one and \(r\le p\); the missing mass can be placed at a zero vector. Raise to power \(p\), add one, and divide by the constant denominator in (213). ◻

2 summarizes the horizontal scales used in the comparison and discretization below.

The horizontal-length axis (schematic) for a marked block. Here \(Q=G^{3/4}\rho_{j+1}\) is the retained angular separation and \(G\rho_{j+1}\le\rho_j\le B\rho_{j+1}\), including the initial block. The quantity \(l_*=B^2w/\rho_{j+1}\) is the lower bound for predictable packet lengths in (185). On an interval of length \(H_j\), the separated relative phases oscillate many times while the packet envelopes move little. The preceding mesh also satisfies \(H_{j-1}\rho_j/w=G^{-1/2}\), which controls the step supremums. Absolute frequency constants are suppressed.

Comparison to the actual factor

On the underlying real Hilbert space put \(\mathcal L(Y)=\log(1+\left\|Y\right\|^p)\). For \(p\ge4\), \[ \left\|\nabla\mathcal L(Y)\right\|\le p,\qquad \left\|D^2\mathcal L(Y)\right\|\le C p^2. \tag{216}\] Indeed its radial first derivative is \(p r^{p-1}/(1+r^p)\), its radial second derivative is \[\frac{p(p-1)r^{p-2}}{1+r^p} -\frac{p^2r^{2p-2}}{(1+r^p)^2},\] and each tangential Hessian eigenvalue is \(p r^{p-2}/(1+r^p)\). Splitting into \(r\le1\) and \(r\ge1\) proves the bounds; the formulas extend continuously at zero. No dimension enters these estimates.

If \(\upsilon\ge1/2\), every child was shrunk. The denominator comparison (211), the sum of (212), and (216) yield \[ \log\beta_J^0(z)-\log\widetilde\beta(b) \le CpK_*\varepsilon\upsilon,\qquad \upsilon\ge1/2. \tag{217}\] The factor \(\upsilon\) here costs only an absolute factor of two.

For \(0<\upsilon<1/2\), retaining that factor requires the unshrunk path vector. Write \(\sigma_a=\lambda_a/\left|\lambda_a\right|\) and remove \(\mathcal D_a(b)\) and \(\sigma_a\) from the predicted numerator. With \[\begin{align*} Y&=l_a^{1/p}X_a,\\ C_a&=l_a^{1-\mu-1/p}\left|\lambda_a\right| =l_a^{1-\alpha/p}(n_a/m_a)^\mu,\\ Z_1(b)&=\sigma_a^{-1}\mathcal D_a(b)^{-1} \sum_{c\ne a}l_c^{1-\mu}\lambda_c \mathcal D_c(b)\widetilde X_c,\\ R_1&=\sum_{c\ne a}l_c\left\|\widetilde X_c\right\|^p, \end{align*}\] the prediction becomes \[ \widetilde\beta(b)= \frac{1+\left\|C_aY+Z_1(b)\right\|^p} {1+\left\|Y\right\|^p+R_1}. \tag{218}\] Here \(Y,C_a\), and the denominator are constant on the whole line. Since \(n_a\le m_a\le n\), both \(l_a\) and \(n_a/m_a\) lie in \([1-\upsilon,1]\). As \(1-\alpha/p\ge0\), we have \[ 0<C_a\le1,\qquad 1-C_a\le C\upsilon,\qquad \left\|Z_1(b)\right\|\le K_*\upsilon. \tag{219}\]

In the actual numerator remove \(\mathcal D_a(b)\) and \(\sigma_a'=\lambda_a'/\left|\lambda_a'\right|\). Define \[Y'=l_a'^{1/p}\mathcal D_a(b)^{-1}X_a',\qquad C_a'=l_a'^{1-\mu-1/p}\left|\lambda_a'\right|,\qquad R_1'=\sum_{c\ne a}l_c'\left\|X_c'\right\|^p,\] and let \(Z_1'\) be the remaining, similarly rotated off-path sum. The actual ratio has the form (218) with primes. We have \[\begin{align*} \left\|Y'-Y\right\|&\le C\varepsilon K_*, &\left|C_a'-C_a\right|&\le C\varepsilon\upsilon,\\ \left\|Z_1'-Z_1\right\|&\le C\varepsilon K_*\upsilon, &R_1'&\ge R_1. \tag{220}\end{align*}\] The first and last statements follow from (209) and (211). For the off-path sum, sum (212) over \(c\ne a\), whose weights sum to at most \(\upsilon\). Changing the removed scalar phase costs at most \(C\varepsilon\) times the off-path size, because \(\left|\sigma_a'-\sigma_a\right|\le C\varepsilon\).

The extra \(\upsilon\) in the coefficient estimate follows from the complementary-mass assertion of 36. If \(r=A/(A+B)\), with \(A>0\), \(B\ge0\), and the positive summands change relatively by at most \(C\varepsilon\), then \[|r'-r|\le C\varepsilon r(1-r),\qquad |\log r'-\log r|\le C\varepsilon(1-r).\] If \(B=0\), the ratio is identically one. Apply this to \(n_a/n\), with complement \(n-n_a\), and to \(n_a/m_a\), with complement \(m_a-n_a\). Both values of \(1-r\) are at most \(\upsilon\). The bounded nonnegative powers defining \(C_a\) give its claimed variation. This concerns the modulus of the path weight; its scalar phase has already been removed.

Set \(\mathfrak r=R_1/(1+\left\|Y\right\|^p+R_1)\). We claim \[\begin{align*} \log\beta_J^0(z)-\log\widetilde\beta(b) &\le C p^2K_*^2\varepsilon\upsilon +C pK_*\varepsilon\mathfrak r,\\ -\log\widetilde\beta(b) &\ge \mathfrak r-CpK_*\upsilon. \tag{221}\end{align*}\] Hold \(Z_1=Z_1(b)\) fixed when comparing \(Y\) with \(Y'\). The derivative in \(V\) of \(\mathcal L(C_aV+Z_1)-\mathcal L(V)\), along that segment, is bounded by \[p|1-C_a|+Cp^2\bigl(|1-C_a|\left\|V\right\|+\left\|Z_1\right\|\bigr) \le Cp^2K_*\upsilon.\] Equation (220) bounds its change by \(Cp^2K_*^2\varepsilon\upsilon\). The changes in \(C_a,Z_1\) fit the same bound by (216).

For the residual denominator, set \[J(V)=-\log(1+R_1e^{-\mathcal L(V)}).\] Its derivative has norm at most \(pR_1/(e^{\mathcal L(V)}+R_1)\). On the segment from \(Y\) to \(Y'\), \(\left|\mathcal L(V)-\mathcal L(Y)\right|\le CpK_*\varepsilon\). This tends to zero as \(L\) increases, so that residual ratio is at most \(2\mathfrak r\), and \(J(Y')-J(Y)\le CpK_*\varepsilon\mathfrak r\). The actual residual \(R_1'\ge R_1\) can only decrease the actual ratio. This proves the first inequality of (221). For the second, use \(C_a\le1\), (216), and \(\log(1+t)\ge t/(1+t)\): \[\begin{align*} -\log\widetilde\beta &=\mathcal L(Y)-\mathcal L(C_aY+Z_1) +\log(1+R_1e^{-\mathcal L(Y)})\\ &\ge-p\left\|Z_1\right\|+\mathfrak r \ge\mathfrak r-CpK_*\upsilon. \end{align*}\]

Absorbing the baseline error by blending

The parameter sizes give \[ pK_* = G^{o(1)}\quad\text{as }L\longrightarrow\infty, \tag{222}\] with every preceding constant fixed. Indeed \(\log N\le2L\), \(\mu=O(L^{-1/2})\), and \(\log G\asymp L^{7/10}\), so \(\log(pK_*)=O(\sqrt L+\log L)=o(\log G)\). All the preceding smallness requirements consequently follow from a single sufficiently large lower bound on \(L\).

Set \(\theta_{\mathrm{bl}}=\sqrt\varepsilon\) and blend the prediction with one: \[ \mathcal B(b)=(1-\theta_{\mathrm{bl}})\widetilde\beta(b)+\theta_{\mathrm{bl}}. \tag{223}\] It retains average at most one for every probability in 43. Concavity of the logarithm gives \[\log \mathcal B-\log\widetilde\beta \ge-\theta_{\mathrm{bl}}\log\widetilde\beta.\] For \(\upsilon<1/2\), Equation (221) yields \[\begin{align*} \log\beta_J^0-\log \mathcal B &\le\bigl(Cp^2K_*^2\varepsilon+CpK_*\theta_{\mathrm{bl}}\bigr)\upsilon +\bigl(CpK_*\varepsilon-\theta_{\mathrm{bl}}\bigr)\mathfrak r\\ &\le\bigl(Cp^2K_*^2\varepsilon +CpK_*\sqrt\varepsilon\bigr)\upsilon. \end{align*}\] The last coefficient of \(\mathfrak r\) is nonpositive for large \(L\). For \(\upsilon\ge1/2\), use (217) and \(\log\widetilde\beta\le\mathcal L(F)\le p\left\|F\right\|\le CpK_*\) to obtain the same last bound. Therefore \[ \beta_J^0(z)\le e^{a_L\upsilon}\mathcal B(b),\qquad a_L=Cp^2K_*^2\varepsilon+CpK_*\sqrt\varepsilon \le G^{-\gamma/3} \tag{224}\] for sufficiently large \(L\). The baseline has absorbed the term proportional to \(\mathfrak r\); the remaining error is proportional to \(\upsilon\).

Changing to uniform measure without a loss per level

Choose a fixed nonnegative Schwartz function \(k\) of integral one with \(\mathop{\mathrm{supp}}\widehat k\subset[-1/4,1/4]\). For example, square the absolute value of the inverse Fourier transform of a nonzero smooth function supported in \((-1/8,1/8)\) and normalize its integral. Let \(H=H_j\), \(h=G^{-1/16}\), \(s=HG^{-1/8}\), and \(k_s(t)=s^{-1}k(t/s)\). Let \(I^+\) be the concentric enlargement of \(I\) of length \((1+h)H\), and define \[ \,\mathrm d\nu_I(t)= \left(\frac{\mathbf 1_{I^+}}{|I^+|}*k_s\right)(t)\,\,\mathrm dt. \tag{225}\] This is a positive probability whose Fourier transform is supported in \([-1/(4s),1/(4s)]\). Since \(H\delta=\epsilon'G^{1/4}\), this support lies strictly inside \((-\delta,\delta)\) for large \(L\).

Every point of \(I\) is at distance at least \(hH/2\) from the complement of \(I^+\). The Schwartz tail beyond \((hH)/(2s)=G^{1/16}/2\) is \(O(G^{-1/8})\), by a fixed tail estimate. Thus the density in (225) is at least \(H^{-1}(1-CG^{-1/16})\) on \(I\). Writing \(\omega_I\) for uniform probability on \(I\), choose \(0<\eta\le CG^{-1/16}<1/2\) so that \[ \nu_I\ge(1-\eta)\omega_I,\qquad \nu_I=(1-\eta)\omega_I+\eta\nu_{\rm rem} \tag{226}\] for a positive probability \(\nu_{\rm rem}\).

The bound \(\int \mathcal B\,\,\mathrm d\nu_I\le1\) alone would lose an amount of order \(\eta\) on changing measure. We need a bound proportional to \(\upsilon\). If \(\widetilde\beta\le1\) on the whole line, then \(\mathcal B\le1\) and its uniform average is already at most one. Otherwise choose \(b_0\) with \(\widetilde\beta(b_0)>1\). When \(\upsilon<1/2\), the denominator and path vector in (218) are fixed. For all \(b\in\mathbb R\), \[\begin{align*} \log\widetilde\beta(b) &\ge\log\widetilde\beta(b_0) -p\left\|Z_1(b)-Z_1(b_0)\right\|\\ &\ge-2pK_*\upsilon. \end{align*}\] Consequently \[ (1-\widetilde\beta(b))_+\le CpK_*\upsilon \qquad(b\in\mathbb R). \tag{227}\] This does not assume an upper bound on \(R_1\): its contribution cancels in the logarithmic difference. Equivalently, a value exceeding one already forces \(R_1/(1+\left\|Y\right\|^p)\le e^{CpK_*\upsilon}-1\). For \(\upsilon\ge1/2\), Equation (227) follows from nonnegativity since \(pK_*\upsilon\ge1\). The blend inherits the same lower bound, as \((1-\mathcal B)_+=(1-\theta_{\mathrm{bl}})(1-\widetilde\beta)_+\).

In the case where the prediction exceeds one somewhere, (226) gives \[1\ge\int \mathcal B\,\,\mathrm d\nu_I \ge(1-\eta)\int \mathcal B\,\,\mathrm d\omega_I +\eta(1-CpK_*\upsilon).\] Including the easier case \(\mathcal B\le1\), this proves \[ \int \mathcal B\,\,\mathrm d\omega_I\le1+C G^{-1/16}pK_*\upsilon. \tag{228}\] The relative mass drop has survived the change of probability.

Step supremums and normalization

The points \(s_R,t_c\) lie in \(J\), so all component frequencies of \(\mathcal D_c\) have magnitude at most \(C\rho_j/w\). For \(\upsilon\ge1/2\), differentiating (213) and using (208) gives \(\left\|\partial_b F(b)\right\|\le C(\rho_j/w)K_*\). For \(\upsilon<1/2\), remove the path carrier as in (218). The off-path frequencies are \((t_a-t_c)\eta_\xi\), again of magnitude at most \(C\rho_j/w\), so \(\left\|\partial_b Z_1(b)\right\|\le C(\rho_j/w)K_*\upsilon\). In both cases, \[\left|\frac{\,\mathrm d}{\,\mathrm db}\log\widetilde\beta(b)\right| \le CpK_*\upsilon\frac{\rho_j}{w}.\] Blending cannot increase the absolute logarithmic derivative, because \[\frac{\,\mathrm d}{\,\mathrm db}\log \mathcal B =\frac{(1-\theta_{\mathrm{bl}})\widetilde\beta}{\mathcal B} \frac{\,\mathrm d}{\,\mathrm db}\log\widetilde\beta,\qquad 0\le\frac{(1-\theta_{\mathrm{bl}})\widetilde\beta}{\mathcal B}\le1.\]

On each \(H_{j-1}\)-subcell \(I'\subset I\), let \(\mathcal B^\sharp\) equal \(\sup_{b\in\overline{I'}}\mathcal B(b)\). The prediction is positive and continuous, and these compact-cell supremums are finite. The meshes are nested. Since \(H_{j-1}\rho_j/w=G^{-1/2}\), \[ \mathcal B\le \mathcal B^\sharp\le e^{b_L\upsilon}\mathcal B,\qquad b_L=CpK_*G^{-1/2}. \tag{229}\] Use the fixed half-open mesh convention at endpoints; taking supremums over closed subcells preserves the inequalities. Define \[ m_I=\max\left(1,\int \mathcal B^\sharp\,\,\mathrm d\omega_I\right),\qquad \Gamma_j=\mathcal B^\sharp/m_I\quad\text{on }I. \tag{230}\] This function is nonnegative, constant on the finer cells, and has average at most one. Equations (228) and (229) yield \[ \log m_I\le\bigl(b_L+CG^{-1/16}pK_*\bigr)\upsilon. \tag{231}\]

Proof of 42. Perform the construction on each marked cell meeting \(Z\). Equations (224), (229), and (230) give \[\beta_J^0(z)\le e^{a_L\upsilon}\mathcal B(b) \le e^{a_L\upsilon}\mathcal B^\sharp(b) =e^{a_L\upsilon}m_I\Gamma_j(b).\] By (231), the exponent is at most \[\bigl(G^{-\gamma/3}+CpK_*G^{-1/2} +CpK_*G^{-1/16}\bigr)\upsilon.\] Equations (222) and (206) bound this by \(G^{-c_{\rm av}}u_j(z)\), for a fixed positive exponent: choose \(c_{\rm av}<\min(\gamma/4,1/64)\) and then take \(L\) sufficiently large. This proves (205) with no error independent of the mass drop.

There are countably many cells on each fixed line and finitely many levels in the current truncation. Choose one snapshot in every relevant cell, use (230) there, and set \(\Gamma_j=1\) elsewhere. This is a measurable step function of the line parameter. The proof of 46 integrates the original integrand on lines where its slice is measurable, as holds for almost every transverse parameter. It therefore needs no jointly measurable snapshot choices across parallel lines. ◻

Residual terms and integration along paths

We work with a finite tile collection and a fixed angular translation. Fix a shift \(\sigma\) and a parameter \(\zeta\) whose real part is the corresponding vertex: \(\Re\zeta_\sigma=\mu\), all other real parts vanish, and \(\left\|\Im\zeta\right\|_2\le L^2\). All estimates in this section are uniform in these choices, the number of tiles, the angular depth, and \(\dim\mathcal H\). We take \(L\) sufficiently large that \[ p>\alpha,\qquad 0<\mu\le\frac1{12}. \tag{232}\] We retain the stopping threshold \(M=L^{A_M}\) from (193); its fixed exponent will be chosen at the end of the residual estimates.

Fix a path top \(R\). We retain the block notation of the hierarchy: \(J_j=v_{i_j}(R)\), \(n_j=n_{J_j}\), and \(A_j=A_{J_j}\). The last possible block is the last one whose starting node contains \(R\) as an eligible top. At this last block there is no path child, even if other tops continue; the next path mass is defined to be zero. At every other block write \(J_{j+1}\) for the path child and set \[h_j^*=n_j-n_{j+1},\qquad u_j=\frac{h_j^*}{n_j}.\] An index is active if \(n_j>M\). The active indices form a prefix. On \(Z\), the mass bound (112) and monotonicity give \[ \sum_{j\text{ active}}u_j\le C L. \tag{233}\] Indeed, except for the last active index, use \(1-n_{j+1}/n_j\le\log(n_j/n_{j+1})\) and telescope down to a mass larger than \(M\). The final summand is at most one. This proof also applies when the path stops by ineligibility.

Marked residuals

For \(n>0\) and \(A\in\mathcal H\), put \[V(n,A)=n+n^{\alpha-p}\left\|A\right\|^{p} =n\bigl(1+\left\|n^{\mu-1}A\right\|^{p}\bigr).\] The function \(\mathcal L_p(Y)=\log(1+\left\|Y\right\|^{p})\) is globally \(p\)-Lipschitz on the underlying real Hilbert space. Away from zero its derivative has norm \(p\left\|Y\right\|^{p-1}/(1+\left\|Y\right\|^{p})\le p\), and the assertion at zero follows by continuity. Since \(\beta_J\) and \(\beta_J^0\) have the same denominator, the decomposition (159) therefore implies \[ \log\beta_{J_j}-\log\beta_{J_j}^0 \le Cp\,n_j^{-1+\mu}\left\|B_{J_j}\right\|. \tag{234}\]

Lemma 44 (Square summation of residual blocks). At a point \(z\in Z\) where \(R\) is a good path, \[ \left(\sum_j\left\|B_{J_j}(z)\right\|^2\right)^{1/2}\le L^{C_b}. \tag{235}\] Here the sum can be restricted to any subset of the possible blocks. The fixed exponent \(C_b\) is independent of \(M\).

Proof. A residual packet in block \(j\) has \[ \frac{\rho_{j+1}}{B^2}<d_s\le\frac{\rho_j}{B^2}, \qquad i_j\le k_s<i_{j+1}. \tag{236}\] Because \(\rho_j/\rho_{j+1}\le B\), this block contains at most \(N_{\mathrm{len}}:=C(1+\log B)\) dyadic packet lengths. Also \(s_{T(s)},s_R\in J_j\), so the top assignment and (236) imply \[\left|\theta_s-s_R\right| \le C d_s+\rho_j \le C B^3 d_s\le C B^4d_s.\] Thus the reference-line estimates apply whenever \(\ell_s\le l_R\).

The coefficient multiplying \(q_{T(s)}\phi_s\) in \(B_{J_j}\) is a product of transition weights of modulus at most one, followed by \(n_{v_{k_s}(T(s))}^{-\mu/K}\). The ending mass is at least \(h_{T(s)}\). Consequently the modulus of the coefficient multiplying \(\phi_s\) is at most \(|q_{T(s)}|h_{T(s)}^{-\mu/K}\le C\). This last bound follows from the mask estimate (87): on the significant region \(h_T\ge1\), while off that region the mask has faster decay than this fixed small negative power of \(h_T\).

At \(z\) let \(e_\ell\) and \(a_\ell\) be the sums of the absolute error and main envelopes, respectively, at length \(\ell\) in the fixed family \(\mathcal P_R\) of (187). The preceding bounds place every residual packet with \(\ell_s\le l_R\) in this family, so its absolute envelopes are included in these positive sums. Goodness of \(R\), as specified in 39, gives \[\sum_{\ell\le l_R}e_\ell\le L^C, \qquad \left(\sum_{\ell\le l_R}a_\ell^2\right)^{1/2}\le L^C.\] The contribution of such lengths to \(\left\|B_{J_j}\right\|\) is bounded by \(C\sum_{\ell\in\mathcal I_j}(e_\ell+a_\ell)\), where \(\mathcal I_j\) is the set of residual lengths in that block. The sets \(\mathcal I_j\) are disjoint. Positivity and Cauchy–Schwarz within each block therefore give \[\begin{align*} \left\{\sum_j \left(\sum_{\ell\in\mathcal I_j,\,\ell\le l_R} e_\ell\right)^2\right\}^{1/2} &\le\sum_{\ell\le l_R}e_\ell,\\ \sum_j\left(\sum_{\ell\in\mathcal I_j,\,\ell\le l_R} a_\ell\right)^2 &\le N_{\mathrm{len}}\sum_{\ell\le l_R}a_\ell^2. \end{align*}\] The factor \(N_{\mathrm{len}}\) is polynomial in \(L\).

Before the last possible block, \(R\) remains eligible at the child level. Hence \(d_R\le\rho_{j+1}/B^2<d_s\), and consequently \(\ell_s<l_R\). Lengths larger than \(l_R\) can occur only in the last possible block. At each of its at most \(N_{\mathrm{len}}\) lengths, the total absolute coefficient envelope is bounded by a constant, by 23, applied with reference length equal to that scale. To recall the reason this bound has no factor counting the available labels, the condition \(|\theta_s-s_R|\le CB^4d_s\) makes the analyzing and synthesizing box envelopes comparable to common \(s_R\)-coordinate envelopes, since \(W\gg B^4w\). Summing centers first produces a common positive convolution kernel of integral \(O(1)\). At each source point only boundedly many selector labels are nonzero. Summing these labels then uses \(\left\|g\right\|_{\mathcal H}\le1\), and costs \(O(1)\). The uniformly bounded multipliers just considered preserve this absolute estimate. The last block therefore contributes at most \(C N_{\mathrm{len}}\) in norm. Combining the estimates proves (235). ◻

By 41, at any marked point one has \(h_j^*\ge c_m\) for a fixed \(c_m>0\). Monotonicity of the masses therefore implies \[ \#\{j:\ j\text{ marked at }z,\ t\le n_j<2t\} \le C(1+t),\qquad t\ge M. \tag{237}\] For completeness, all but the final marked index in this shell account for disjoint mass drops of size at least \(c_m\) while the mass lies in an interval of length at most \(t\); the final index costs one even if its drop leaves the shell. Dyadic summation gives \[ \sum_{j\text{ marked at }z}n_j^{-2+2\mu} \le C\sum_{r\ge0}(2^rM)^{-1+2\mu} \le C M^{-1+2\mu}. \tag{238}\] Combining (234), 44, and Cauchy–Schwarz yields \[ \sum_{j\text{ marked at }z} \bigl(\log\beta_{J_j}-\log\beta_{J_j}^0\bigr) \le CpL^{C_b}M^{-1/2+\mu}. \tag{239}\]

Unary segments and their energy telescope

Call an active index unary at \(z\) if its cell is unmarked and \(h_j^*(z)<1\). At such an index the path child exists: at the last possible block the drop would instead be \(n_j>M\). A significant top has mass at least one. Thus every significant top eligible in \(J_j\) must remain eligible in the path child, since the total mass removed from that child is \(h_j^*<1\).

At a unary index write \[ A_j=\lambda_j A_{j+1}+U_j+W_j. \tag{240}\] Here \(\lambda_j\) is the path-child coefficient in (159); \(U_j\) consists of the residual packets assigned to tops which remain eligible in the path child. These are exactly the packets of the fixed subset \(\mathcal Q_{R,\sigma}\) in (188) which stop in block \(j\). The term \(W_j\) collects all contributions of tops eligible in \(J_j\) which do not continue in the path child, including their residual packets and, when present, their contributions through other children.

Those latter tops are all insignificant at \(z\). Their contributions within \(A_{J_j}\) are whole interval sums on individual tops, so 35 bounds the contribution of each by \(L^C(1+\left\|\zeta\right\|)|q_T|h_T^4\). Since \(h_T\) is bounded above by an absolute constant and \(|q_T|\le e^{-L\log L}\) there, this is at most \(Ce^{-L}h_T\) for sufficiently large \(L\). The masses of these tops sum to \(h_j^*\). Therefore \[ \left\|W_j\right\|\le Ce^{-L}h_j^*. \tag{241}\] This estimate uses a sum of top masses, not the number of such tops.

Partition each consecutive run of unary indices into consecutive segments \([e,f)\). Starting at \(e\), end a completed segment as soon as its accumulated drop is at least one. Each individual drop is less than one, so every completed segment has total drop in \([1,2)\). At the end of the run, retain the final remainder as one segment if it is nonempty. Hence every segment satisfies \[ \Delta_{e,f}:=\sum_{e\le j<f}h_j^*=n_e-n_f<2, \qquad n_f>M-2, \qquad n_j\asymp n_e\quad(e\le j\le f). \tag{242}\] The endpoint remains an eligible path node, including when it is the node where the mass stopping rule first applies.

The number of segment starts in a mass shell obeys \[ \#\{[e,f):\ t\le n_e<2t\}\le C(1+t),\qquad t\ge M. \tag{243}\] Completed segments account for disjoint drops of size at least one. For each remainder other than the final remainder on the path, charge it to the next interruption of a unary run. Such an interruption is either marked, with drop at least \(c_m\), or is an active unmarked index with drop at least one. The mass at this interruption differs from the start mass of the remainder by less than two. Thus the interruptions charged by starts in \([t,2t)\) lie in the enlarged shell \([t-2,2t)\) and their fixed positive drops give \(O(1+t)\) choices. The one possible final remainder adds one.

For one segment define \[A_i^*=\left(\prod_{i\le j<f}\lambda_j\right)A_f, \qquad V_i^*=V(n_i,A_i^*),\qquad e\le i\le f,\] with the empty product equal to one. Since \(n_i\ge n_{i+1}\), \(\alpha-p<0\), and \(|\lambda_i|\le1\), we have \[ V_i^*\le h_i^*+V_{i+1}^*,\qquad V_f^*=V_f. \tag{244}\] Indeed, the baseline terms differ by \(h_i^*\), while the energy term on the left is at most \(n_{i+1}^{\alpha-p}\left\|A_{i+1}^*\right\|^p\). The denominator of \(\beta_{J_i}\) contains the energy of the path child and the nonnegative energies of all other children. It is therefore at least \(h_i^*+V_{i+1}\). It follows that \[\sum_{e\le i<f}\log\beta_{J_i} \le \log V_e-\log V_f -\sum_{e\le i<f}\log(1+h_i^*/V_{i+1}).\] On the other hand, (244) gives \[\log V_e^*-\log V_f \le\sum_{e\le i<f}\log(1+h_i^*/V_{i+1}^*).\] Subtracting these inequalities produces the useful estimate \[ \begin{aligned} \sum_{e\le i<f}\log\beta_{J_i} &\le \log(V_e/V_e^*) +\sum_{e\le i<f} \left[\log(1+h_i^*/V_{i+1}^*) -\log(1+h_i^*/V_{i+1})\right] \\ &\le Cp\, n_e^{-1+\mu} \max_{e\le i\le f}\left\|A_i-A_i^*\right\|. \end{aligned} \tag{245}\] To justify the last line, first use the Lipschitz estimate for \(\mathcal L_p\) to obtain \[|\log V(n,A)-\log V(n,A')| \le p n^{\mu-1}\left\|A-A'\right\|.\] For \(h\ge0\), the derivative of \(x\mapsto\log(1+he^{-x})\) has absolute value \(h/(e^x+h)\). Along the interval between \(\log V_{i+1}\) and \(\log V_{i+1}^*\), both endpoint amounts, and therefore every intermediate amount, are at least \(n_{i+1}\). The corresponding correction term is bounded by \[\frac{h_i^*}{n_{i+1}} p n_{i+1}^{\mu-1}\left\|A_{i+1}-A_{i+1}^*\right\|.\] Now use (242) and \(\sum_{e\le i<f}h_i^*/n_{i+1}\le2/(M-2)\). This proves (245), including its endpoint case.

Common residual weights and variation

We give the coefficient calculation needed to apply the good-path variation estimate. It is important here that the same angular grid underlies all shifts.

For integer levels \(k\) along the path put \(v_k=v_k(R)\) and \(a_k=n_{v_k}\), so \(a_{i_j}=n_j\). For a fixed segment \([e,f)\) and a block index \(j\in[e,f)\), define \[ \omega_j(k)=a_k^{-\mu/K} \prod_{r=i_j}^{k-1} \left(\frac{a_{r+1}}{m_{v_{r+1}}}\right)^{\zeta_{r\bmod K}}, \qquad i_j\le k<i_f. \tag{246}\] Here \(i_j\) denotes the angular level at the start of block \(j\).

Every packet in \(U_b\), for \(j\le b<f\), is assigned to a top in \(v_{i_{b+1}}(R)\). It follows by nesting that its top lies in \(v_k(R)\) at every intermediate level \(k\le i_{b+1}\), in particular at its stopping level \(k_s\). Multiplying its coefficient in \(U_b\) by the path propagation factor \(\prod_{j\le r<b}\lambda_r\) therefore gives exactly \(\omega_j(k_s)q_{T(s)}\phi_s\). This identity holds for every contributing top, whether or not that top is significant. Thus the scalar coefficient depends only on the stopping level, rather than on the packet or top inside that level.

Lemma 45 (Variation of the common coefficients). Uniformly over the segments and their starting indices, \[ \sup_{i_j\le k<i_f}|\omega_j(k)| +\sum_{k=i_j}^{i_f-2}|\omega_j(k+1)-\omega_j(k)| \le L^{C_w}. \tag{247}\] The exponent \(C_w\) is fixed before \(M\) is chosen.

Proof. Every ending mass in this sequence is at least \(a_{i_f}=n_f>M-2\). Since all real parts of \(\zeta\) are nonnegative, \(|\omega_j(k)|\le a_k^{-\mu/K}\le1\) when \(M\ge3\). Using the real logarithms of positive masses, a logarithmic increment of the coefficient is \[\log\omega_j(k+1)-\log\omega_j(k) =\frac{\mu}{K}\log\frac{a_k}{a_{k+1}} -\zeta_{k\bmod K}\log\frac{m_{v_{k+1}}}{a_{k+1}}.\] The logarithm here is the specified linear expression defining the coefficient, so no branch of a complex logarithm is being chosen. The first sum of absolute increments telescopes. For the second use 34 on the entire root-to-endpoint path. All its summands are nonnegative, whence \[\begin{align*} \sum_{k=i_j}^{i_f-2} |\log\omega_j(k+1)-\log\omega_j(k)| &\le \left(\frac{\mu}{K} +K\max_\sigma|\zeta_\sigma|\right) \log\frac{n_{\mathrm{root}}}{n_f} \\ &\le L^{C_w}. \end{align*}\] The last step uses \(n_{\mathrm{root}}\le NL^C\) on \(Z\), \(\log N\le2L\), \(n_f> M-2\ge1\), and \(\max_\sigma|\zeta_\sigma|\le\mu+L^2\). The full-root bound is needed: a denominator \(m_v\) belonging to a different shift may involve a block whose starting level precedes the segment. No estimate using only \(\log(n_e/n_f)\) is asserted.

For two successive logarithmic values \(z_0,z_1\) in this sequence, both \(|e^{z_0}|\) and \(|e^{z_1}|\) are at most one. Along the straight segment between them, \(|e^{(1-t)z_0+tz_1}|\le1\). Integration of the derivative of the exponential consequently yields \(|e^{z_1}-e^{z_0}|\le|z_1-z_0|\). The logarithmic-increment bound proves (247). ◻

Let \(\mathcal Q_{R,\sigma}\) denote the fixed residual subset in (188). For each unary segment \([e,f)\) define \[\Lambda_{e,f}(z)= \max_{i_e\le a\le b<i_f} \left\|\sum_{\substack{s\in\mathcal Q_{R,\sigma}\\ a\le k_s\le b}} q_{T(s)}(z)\phi_s(z)\right\|.\] An empty sum has value zero. Finite depth makes the maximum finite. The interval of stopping levels \([a,b]\) corresponds to an interval of dyadic packet lengths. The latter follows directly from the definition \(k_s=\max\{k:d_s\le d(k)/B^2\}\), which is monotone in \(\ell_s\). Since distinct unary segments have disjoint level intervals \([i_e,i_f)\), one may choose a maximizing interval in each segment and apply the single good-path \(3\)-variation bound to all those disjoint intervals. Thus \[ \left(\sum_{[e,f)}\Lambda_{e,f}(z)^3\right)^{1/3} \le L^{C_v}. \tag{248}\] The segments may depend on \(z\); the variation is a supremum over all choices of disjoint intervals, so this dependence does not alter the inequality. The underlying packet subset itself is the single deterministic set \(\mathcal Q_{R,\sigma}\).

Abel summation for a finite scalar sequence \((\omega_k)\) and vector sequence \((F_k)\) gives \[\left\|\sum_{k=a}^b\omega_k F_k\right\| \le \left(|\omega_b|+ \sum_{k=a}^{b-1}|\omega_k-\omega_{k+1}|\right) \max_{a\le t\le b} \left\|\sum_{k=a}^tF_k\right\|.\] Applying this identity with (246), and then iterating (240), yields \[ \max_{e\le i\le f}\left\|A_i-A_i^*\right\| \le L^{C_w}\Lambda_{e,f} +Ce^{-L}\Delta_{e,f}. \tag{249}\] Indeed the propagated \(W_j\) terms have modulus at most their unpropagated bounds because \(|\lambda_j|\le1\). There are no residual terms beyond \(i_f\) in this calculation: \(U_j\) uses only \(i_j\le k_s<i_{j+1}\) with \(j<f\). When \(f\) is a mass-stopping node it is still an eligible node and (242) applies. A final block ending by ineligibility cannot be unary and is excluded from this calculation.

Using (243) and dyadic mass shells, we get \[ \begin{split} \sum_{[e,f)}n_e^{(-1+\mu)3/2} &\le C\sum_{r\ge0}(2^rM)^{-1/2+3\mu/2} \le C M^{-1/2+3\mu/2},\\ \left(\sum_{[e,f)}n_e^{(-1+\mu)3/2}\right)^{2/3} &\le C M^{-1/3+\mu}. \end{split} \tag{250}\] Hölder with exponents \(3/2\) and \(3\), followed by (245), (248), and (249), therefore bounds the sum of the unary log factors by \[ \sum_{j\text{ unary at }z}\log\beta_{J_j} \le CpL^{C_u}M^{-1/3+\mu} +Cp e^{-L}\left(\max_{j\text{ active}}n_j^\mu\right) \sum_{j\text{ active}}u_j. \tag{251}\] For the error term, use \(n_e\ge n_j\) and \(\mu-1<0\) within a segment to obtain \(n_e^{\mu-1}h_j^*\le n_j^{\mu-1}h_j^*=n_j^\mu u_j\). By (112), \(\max n_j^\mu\le\exp(C\sqrt L)\). Together with (233), this makes the second term in (251) tend to zero.

Choose a fixed exponent \(C_{\mathrm{res}}\) larger than all polynomial exponents in (239) and (251), and then choose \[ A_M>4(C_{\mathrm{res}}+2),\qquad M=L^{A_M}. \tag{252}\] By (232), \(-1/3+\mu\le-1/4\) and \(-1/2+\mu\le-5/12\). Since \(p=L^{1/2}\), this choice makes both principal residual costs bounded by an absolute constant for sufficiently large \(L\). None of the good-path thresholds or the exponents used to select \(A_M\) depended on \(M\).

The path product and its integral

Theorem 46 (Integrated path product). With \(M\) chosen by (252), the estimate (199) holds for every path top \(R\), every shift, and every vertex parameter with \(\left\|\Im\zeta\right\|_2\le L^2\). Its constants are independent of the tile collection, angular depth, and Hilbert-space dimension.

Proof. Fix a line of slope \(s_R\), and on that line use the functions \(\Gamma_j\) constructed in 42. At a point of \(Z\) where \(R\) is good, the active indices fall into the marked indices, the unary indices, and at most one active unmarked index with \(h_j^*\ge1\). For the marked indices, combine (205), (239), and (233). Their phase-domination losses have sum at most \(CLG^{-c}\), which is bounded. The unary factors have bounded log product by (251) and (252).

At the possible remaining index, 41 gives \(M<n_j<M+1\). From (170) and the baseline \(n_j\) in the denominator of \(\beta_{J_j}\), \[\beta_{J_j}\le 1+L^{Cp}n_j^{\alpha-1} \le L^{Cp}(1+M)^C.\] Here \(C\) is fixed, since \(\alpha\) stays in a fixed compact interval below three. We have therefore proved the pointwise domination \[ \prod_{j\text{ active at }z}\beta_{J_j}(z) \le L^{Cp}(1+M)^C\prod_{j=0}^{J_R}\Gamma_j(z), \qquad z\in Z\cap\{R\text{ good}\}, \tag{253}\] where \(J_R\) is the last possible block index for \(R\). It is legitimate to include every possible index on the right: if \(n_j(z)\le M\), its mesh cell cannot satisfy \(\inf n_j>M\), so it is unmarked and \(\Gamma_j(z)=1\). Unmarked active indices likewise have \(\Gamma_j=1\).

Recall that \(\Gamma_j\) is constant on the mesh of length \(H_{j-1}\) and has average at most one on every cell of length \(H_j\). The meshes are nested. On any cell \(I\) of the largest mesh \(H_{J_R}\), successive averaging from the smallest mesh upward therefore gives \[ \int_I\prod_{j=0}^{J_R}\Gamma_j(x)\,\,\mathrm dx\le |I|. \tag{254}\] More explicitly, all later factors \(\Gamma_k\), \(k>j\), are constant on each \(H_j\) cell. Integrating \(\Gamma_j\) on those cells can thus only decrease the integral of the nonnegative product. Iterate this observation for \(j=0,1,\ldots,J_R\).

The weight \(h_R\) is uniformly comparable on \(I\). This remains true for a terminal block and for the short first block: eligibility of \(R\) at \(i_j\) gives \(\rho_j\ge B^2d_R\), and every block has \(\rho_j/\rho_{j+1}\le B\). Thus \[ H_j=\frac{w}{\rho_{j+1}\sqrt G} \le\frac{l_R}{B\sqrt G}\le l_R. \tag{255}\] Along the line of slope \(s_R\), the transverse normalized coordinate of \(R\) is fixed and the horizontal normalized coordinate changes by at most \(H_j/l_R\). The definition (76) consequently gives \(\sup_I h_R\le C\inf_Ih_R\), with an absolute constant depending only on the fixed decay exponent. By (254), \[\int_I h_R\prod_{j=0}^{J_R}\Gamma_j\,\,\mathrm dx \le (\sup_Ih_R)|I|\le C\int_Ih_R\,\,\mathrm dx.\] Sum over the largest-mesh cells and use (253). Whenever the original integrand in (199) has a measurable slice on this line, we obtain \[\int_{Z\cap\{R\text{ good}\}\text{ on the line}} h_R\prod_{j\text{ active}}\beta_{J_j}\,\,\mathrm dx \le L^{Cp}(1+M)^C \int_{\text{line}}h_R\,\,\mathrm dx.\]

The snapshots used to construct the \(\Gamma_j\) need only be chosen on this one line. There are countably many mesh cells, so these choices give measurable step functions on that line. No jointly measurable choice of snapshots across different lines is required. The original integrand in (199), expressed entirely through \(h_R\), the good sets, and the finite hierarchy factors \(\beta_{J_j}\), is nonnegative and Lebesgue measurable. In the coordinates \((x,y-s_Rx)\), it has a measurable slice for almost every transverse parameter. The preceding line estimate therefore holds for almost every such parameter. Integrating it and applying Tonelli, using the Jacobian-one coordinate change, proves (199). ◻

All sums and products in this section were finite. The shell counts, Abel estimates, and mesh averaging bounds contain no factor equal to the number of scales, tiles, or hierarchy levels. They therefore apply with identical constants to every finite truncation required in the packet reduction.

Interpolation, boundary averaging, and completion

We continue with a finite popular collection and its adjoint sum \(P=\sum_s\phi_s\). Thus \(N=(|E|/|F|)^{1/2}\), \(L/2\le\log N\le2L\), and \(S\lesssim L^C|E|\). All statements in this section are uniform in the finite collection, its depth, and the dimension of \(\mathcal H\). We write \(\vartheta\) for the translation of the angular grid. The set \(Z_0\) and the functions \(P,P_q,h_T,n\) do not depend on \(\vartheta\), whereas the smaller set \(Z=Z_\vartheta\), the angular masses, and the analytic sums do. Until the boundary-averaging step, we fix \(\vartheta\) and suppress it in the notation.

From the path estimate to analytic interpolation

Proposition 47 (Vertex estimate). For every nonempty root \(v\), every shift \(\sigma\), and every parameter with \[\Re\zeta=\mu e_\sigma, \qquad \|\Im\zeta\|_2\le L^2,\] one has \[ \big\|n_v^{\alpha/p-1}A_v(\zeta)\big\|_{L^p(Z;\mathcal H)} \le L^C S^{1/p}. \tag{256}\] The constant and exponent are independent of the shift and the grid translation.

Proof. Recall the nonnegative amount \[V_v=n_v+n_v^{\alpha-p}\|A_v\|^p.\] On the set where \(n_v\le M\), Equation (170) gives \[V_v\le L^{Cp}(1+M)^{\alpha-1}n_v.\] Here and below the bound on the imaginary parts is absorbed into \(L^{Cp}\). On the remaining set, apply Lemma 40. At least half of the root amount is carried by good paths, and the terminal amount of a path ending at \(R\) is at most \(L^{Cp}(1+M)^{\alpha-1}h_R\). Consequently, \[\begin{align*} \int_Z V_v &\le L^{Cp}(1+M)^{\alpha-1}\int n_v\\ &\quad+2L^{Cp}(1+M)^{\alpha-1} \sum_{R:\,s_R\in v} \int_{Z\cap\{\text{good }R\}} h_R\prod_{\text{active path of }R}\beta_J. \end{align*}\] The root and its descendants contain only their eligible tops; sums over tops here and subsequently are understood with that convention. Theorem 46, namely Equation (199), bounds each integral in the second line by \(L^{Cp}(1+M)^C\int h_R\). Moreover, \[\sum_R\int h_R\lesssim L^C S,\] by the definition of the dilated decaying weights. We conclude that \[\int_Z V_v\le L^{Cp}(1+M)^C S.\] Since \(M\) is a fixed power of \(L\), taking a \(p\)-th root and using \(n_v^{\alpha-p}\|A_v\|^p\le V_v\) proves the result. ◻

Lemma 48 (Interpolation of the shift parameters). Equation (256) also holds at \[\zeta^{\mathrm c}=(\mu/K,\ldots,\mu/K).\]

Proof. Let \[\mathfrak S_\mu =\{a\in[0,\infty)^K:\ \textstyle\sum_\sigma a_\sigma=\mu\}, \qquad \mathcal F_v(\zeta,z)=n_v(z)^{\alpha/p-1}A_v(\zeta,z).\] Every nonempty mass is positive. For each fixed \(z\), the defining sum is a finite sum of exponentials of linear functions of \(\zeta\), with real logarithms of positive ratios as coefficients. In particular it is entire in \(\zeta\).

We first give a uniform bound on the tube over \(\mathfrak S_\mu\), without using the vertex estimate. From Equations (170) and (112), and from \(\int n_v\le\int n\lesssim L^C S\), we have, for \(a\in\mathfrak S_\mu\) and \(\eta\in\mathbb R^K\), \[\begin{align*} \|\mathcal F_v(a+i\eta)\|_{L^p(Z)} &\le L^C(1+\mu+\|\eta\|_2) \left(\int_Z n_v^\alpha\right)^{1/p}\\ &\le L^C(1+\|\eta\|_2) (NL^C)^{(\alpha-1)/p}S^{1/p}. \tag{257}\end{align*}\] The factor involving \(N\) is at most \(\exp(C\sqrt L)\), since \(\log N\le2L\), \(p=L^{1/2}\), and \(2<\alpha<3\) for sufficiently large \(L\).

Multiply the analytic family by \[\mathcal D(\zeta)=\exp\left(\sum_\sigma\zeta_\sigma^2\right).\] On the tube under consideration, \[|\mathcal D(a+i\eta)| =\exp(\|a\|_2^2-\|\eta\|_2^2) \le\exp(\mu^2-\|\eta\|_2^2).\] At a vertex, Proposition 47 controls \(\mathcal D\mathcal F_v\) whenever \(\|\eta\|_2\le L^2\). For \(\|\eta\|_2>L^2\), Equation (257) gives the same bound: indeed \[(1+\|\eta\|_2) \exp(\mu^2-\|\eta\|_2^2+C\sqrt L)\] is bounded by an absolute constant in this range, after increasing the fixed lower threshold for \(L\). Thus the damped family has norm at most \(L^C S^{1/p}\) at every vertex, uniformly in all imaginary parts.

Here are details that justify applying scalar three-lines to this family. Test against any element \(b\) of the dual unit ball \(L^{p'}(Z;\mathcal H)\). One may first restrict the pairing to a bounded set on which every nonempty mass used in the finite family is bounded below by a positive constant. Such sets exhaust \(Z\), and the restriction does not increase the norm of \(b\). On each such set, integration defines an entire scalar function; logarithms of the finitely many masses are bounded there. The crude bound above and the damping make this scalar function bounded on every interpolating strip. If two real vectors \(a_0,a_1\) have established uniform bounds for all imaginary parts, apply the bounded scalar three-lines lemma (Hajłasz 2012, Lemma 3.3, p. 57) to \[z\longmapsto \int_Z \mathcal D\big((1-z)a_0+za_1+i\eta\big) \big\langle \mathcal F_v\big((1-z)a_0+za_1+i\eta,\cdot\big),b \big\rangle\,\,\mathrm dx,\] with the indicated preliminary restriction understood. The boundary lines have the required bounds because their imaginary vectors are still arbitrary. The real parts throughout the strip remain in the simplex. Starting with its vertices and successively taking convex combinations proves the same bound at their barycenter, with no factor depending on \(K\). Exhaustion of \(Z\), followed by duality, removes the preliminary restriction. Finally \(\mathcal D(\zeta^{\mathrm c})=\exp(\mu^2/K)\) is an absolute factor. This proves the lemma. ◻

The central weight and the exponent budget

The central coefficient identity (165) gives weights \(\kappa_s\in(0,1]\). Define \[\mathcal W_\vartheta =\sum_s\kappa_s q_{T(s)}\phi_s.\] By that identity, its portion in root \(v\) is \(n_v^{\mu/K}A_v(\zeta^{\mathrm c})\).

Put \(\gamma=1-\alpha/p+\mu/K\), which is positive for the values of \(L\) under consideration. Interpolation and Hölder’s inequality give, for each root, \[\begin{align*} \int_Z\big\|n_v^{\mu/K}A_v(\zeta^{\mathrm c})\big\| &\le \|n_v^{\alpha/p-1}A_v(\zeta^{\mathrm c})\|_{L^p(Z)} \left(\int_Z n_v^{\gamma p'}\right)^{1/p'}\\ &\le L^C S^{1/p}|F|^{1-1/p}(NL^C)^\gamma. \end{align*}\] Only an absolute number of roots meet the bounded set of top slopes. Since \(S\lesssim L^C|E|=L^C N^2|F|\), summing over them proves \[ \int_{Z_\vartheta}\|\mathcal W_\vartheta\| \le L^C|F|N^{\,1+(2-\alpha)/p+\mu/K}. \tag{258}\]

We record explicitly why this improves by every fixed negative power of \(L\). The net decrease of the exponent of \(N\) is \[\delta_L =\frac{\alpha-2}{p}-\frac\mu K =\frac1p\left(\alpha-2-\frac{\alpha-1}{K}\right).\] The separated-count bound is a fixed power of \(L\), so the definition of \(\alpha\) gives \(\alpha-2\ge c_1/\log L\) with a fixed \(c_1>0\). Also \(\alpha-1\le2\) and \(K\ge(\log L)^2\). For sufficiently large \(L\), therefore, \[\delta_L\ge\frac{c_1}{2\sqrt L\log L}, \qquad \delta_L\log N\ge\frac{c_1\sqrt L}{4\log L}.\] In particular, for every fixed \(Q>0\), after increasing the absolute lower threshold for \(L\), Equation (258) implies, uniformly in \(\vartheta\), \[ \int_{Z_\vartheta}\|\mathcal W_\vartheta\| \le L^{-Q}N|F|. \tag{259}\] This conclusion absorbs any fixed polynomial losses in the preceding estimates; none of their exponents is changed here.

Boundary averaging with the eligibility condition

We now restore the dependence on the uniformly translated angular grid. The next estimate is pointwise on the fixed set \(Z_0\).

Lemma 49 (Expected boundary loss). Let \(\epsilon=4G^{-1/4}\). For \(z\in Z_0\), \[ \mathbb E_\vartheta\|P_q(z)-\mathcal W_\vartheta(z)\| \le L^C\epsilon n(z). \tag{260}\] Here the expectation is normalized Lebesgue measure for the common translation modulo \(d(0)\); independent translations at different levels are neither assumed nor needed.

Proof. Fix \(z\in Z_0\) and a top \(T\), and suppress \(z\) from the masses. For \(i<k_T\) let \[r_i(\vartheta) =\log\frac{m_{v_{i+1}(T)}}{n_{v_i(T)}}\ge0, \qquad M_i(b)=\sum_{\substack{U:\,|s_U-s_T|\le b\,d(i)\\ d_U\le d(i)/B^2}}h_U.\] Unlike \(r_i\), the latter mass does not depend on the grid translation. Each mass in this display is at least \(h_T\) when \(i<k_T\), and is at most \(n\).

Let \(b_i(\vartheta)\in[0,1/2]\) be the distance from \(s_T\) to the nearer boundary of its level-\(i\) interval, divided by \(d(i)\). An extra top counted in \(m_{v_{i+1}(T)}-n_{v_i(T)}\) has slope within \(\epsilon d(i)\) of \(s_T\). Indeed the neighbor distance in Equation (153) is \(G^{-1/4}d(i)\), and the two level-\((i+1)\) interval lengths add at most \(2d(i)/G\). Such a top is eligible at level \(i+1\), hence also at level \(i\). It lies outside the immediate parent. It follows that \(r_i=0\) if \(b_i>\epsilon\). If \(0<b_i\le\epsilon\), the immediate parent contains every top counted by \(M_i(b_i)\), whereas its extra mass is disjoint from these inner tops and is at most \(M_i(\epsilon)-M_i(b_i)\). Consequently \[ r_i(\vartheta) \le\log\frac{M_i(\epsilon)}{M_i(b_i(\vartheta))}. \tag{261}\] Endpoint coincidences can be omitted: for the finite set of slopes and levels they form a null set of translations.

For a fixed number \(0<b<\epsilon\), set \(r=\lceil\log_G(\epsilon/b)\rceil\). The inclusion \[ M_{i+r}(\epsilon)\le M_i(b) \tag{262}\] uses both conditions in the definition of \(M_i\): the radius satisfies \(\epsilon d(i+r)\le b d(i)\), and the precision threshold \(d(i+r)/B^2\) is no larger than \(d(i)/B^2\). To handle the end of the path without losing a term, define \[a_i=\begin{cases} \log(M_i(\epsilon)/h_T),&0\le i<k_T,\\ 0,&i\ge k_T. \end{cases}\] If \(i+r<k_T\), Equation (262) applies. If \(i+r\ge k_T\), use \(M_i(b)\ge h_T\) instead. In both cases \[\log\frac{M_i(\epsilon)}{M_i(b)}\le a_i-a_{i+r}.\] Summing and cancelling gives \[ \sum_{i<k_T}\log\frac{M_i(\epsilon)}{M_i(b)} \le\sum_{0\le i<\min(r,k_T)}a_i \le r\log\frac n{h_T}. \tag{263}\] This argument also covers \(k_T=0\), with empty sums.

Because \(d(0)/d(i)=G^i\) is an integer, the common uniform translation is uniform modulo every \(d(i)\). Thus each \(b_i\) has density \(2\) on \([0,1/2]\). Linearity of expectation, Equation (261), and Equation (263) imply \[\begin{align*} \mathbb E_\vartheta\sum_{i<k_T}r_i(\vartheta) &\le2\int_0^\epsilon \sum_{i<k_T}\log\frac{M_i(\epsilon)}{M_i(b)}\,\,\mathrm db\\ &\le2\log(n/h_T) \int_0^\epsilon\lceil\log_G(\epsilon/b)\rceil\,\,\mathrm db \le C\epsilon\log(n/h_T). \tag{264}\end{align*}\] For completeness the integral in the last line equals \(\epsilon/(1-G^{-1})\): express the ceiling as a sum of indicators of \(b<\epsilon G^{-j}\), \(j\ge0\), and integrate. The dependence of the different \(b_i\)’s is absent from this calculation.

On the fixed top \(T\), the value of \(\kappa_s\) depends only on the stopping level \(k_s\). As packet length increases, \(k_s\) increases; Equation (165) shows that these values are nonincreasing. Their deficits from one are nondecreasing and satisfy \[0\le1-\kappa_s \le\frac\mu K\sum_{i<k_T}r_i(\vartheta),\] using \(k_s\le k_T\) and \(1-e^{-x}\le x\). Group the packets by stopping level. Every interval of these levels is an interval of packet lengths, so Equation (110) bounds its unweighted sum by \(L^C h_T^8\). Summation by parts against the monotone deficits, and \(|q_T|\le1\), therefore give \[\left\|\sum_{s:T(s)=T}(1-\kappa_s)q_T\phi_s\right\| \le L^C h_T^8\frac\mu K\sum_{i<k_T}r_i(\vartheta).\] For example, write the deficits as sums of their nonnegative increments and use the bound on the corresponding suffix sums; the sum of these increments is their largest value. Taking expectations, using \(\mu/K\le1\), and summing over tops proves \[\mathbb E_\vartheta\|P_q-\mathcal W_\vartheta\| \le L^C\epsilon\sum_T h_T^8\log(n/h_T).\] The weights have a fixed absolute upper bound. For every such weight, \[h_T^8\log(n/h_T) \le C h_T(1+\log_+ n),\] because \(h^7\log_+(1/h)\) is bounded on a fixed bounded interval. Equation (112) gives \(1+\log_+n\lesssim L\) on \(Z_0\). This proves Equation (260). ◻

Recovering the original adjoint sum

Theorem 50 (Improvement for the popular model). The popular adjoint sum satisfies Equation (72); namely, \[\int_F\|P\|\lesssim L^{-20}N|F|.\]

Proof. We make the exceptional-set comparison explicit because the good set depends on the grid. Write \(\|P\|_2\le L^{C_P}\sqrt{|E|}\), as supplied by Equation (43). Proposition 29 and Proposition 30 allow the fixed polynomial thresholds to be chosen so that \[ \int_{F\setminus Z_0}\|P\|\le L^{-30}N|F|, \qquad \int_{Z_0}\|P-P_q\|\le L^{-30}N|F|. \tag{265}\] Indeed a discarded area at most \(L^{-2C_P-60}|F|\) has the stated first bound by Cauchy–Schwarz. The exponentially small concentration and count exceptions meet this area bound for sufficiently large \(L\).

For each fixed translation the good-path construction gives \[|Z_0\setminus Z_\vartheta|\le S e^{-9L}.\] This is Markov’s inequality applied to the weighted bad-path mass: its integral is at most \(S e^{-10L}\), and the deletion threshold is \(e^{-L}\). Since \(S\lesssim L^C N^2|F|\) and \(N^2\le e^{4L}\), \[ \int_{Z_0\setminus Z_\vartheta}\|P\| \le L^{C_P}\sqrt{|E|}\,(S e^{-9L})^{1/2} \le L^C e^{-5L/2}N|F| \le L^{-30}N|F|. \tag{266}\] These are bounds for each translation; no union over translations is taken.

For every \(\vartheta\), the triangle inequality gives the following decomposition, whose left side is independent of \(\vartheta\): \[\begin{align*} \int_F\|P\| &\le \int_{F\setminus Z_\vartheta}\|P\| +\int_{Z_\vartheta}\|P-P_q\| +\int_{Z_\vartheta}\|P_q-\mathcal W_\vartheta\| +\int_{Z_\vartheta}\|\mathcal W_\vartheta\|. \tag{267}\end{align*}\] The first two terms are controlled by Equations (265) and (266). The final term is controlled by Equation (259) with \(Q=30\). Average the remaining term over translations. It is nonnegative and \(Z_\vartheta\subset Z_0\), so Tonelli’s theorem and Lemma 49 yield \[\begin{align*} \mathbb E_\vartheta\int_{Z_\vartheta} \|P_q-\mathcal W_\vartheta\| &\le\int_{Z_0}\mathbb E_\vartheta \|P_q-\mathcal W_\vartheta\|\\ &\le L^C\epsilon\int_{Z_0}n \le L^C\epsilon N|F| \le L^{-30}N|F|. \end{align*}\] The penultimate inequality uses Equation (112); the last uses \(\epsilon=4G^{-1/4}\) and \(\log G\asymp L^{.7}\). All functions being averaged are measurable: for finitely many tops, packets, and levels, the angular assignments change only at finitely many translation values, and the good sets were defined by measurable quantities. Thus this averaging is legitimate even though \(Z_\vartheta\) and \(\mathcal W_\vartheta\) use the same translation. Averaging Equation (267) now proves the asserted estimate. In particular, on every discarded set we have estimated the original sum \(P\). ◻

Choice of parameters and completion of the reductions

We collect the order of the absolute choices. First fix the packet adaptation, separation, and overlap constants and the fixed geometric enlargements used with 2. Fix the local slope bound \(C_{\rm Lip}\), strip length \(b_0\), and outer length \(a=a_*\) as in (2), with the required room for those enlargements. Choose the low-density exponent \(C_2\) by (68), before constructing the popular forest.

Next fix the decay exponent in the top-count weight and the concentration orders that determine \(K_1\). Choose the fixed mask-removal and exceptional-area powers, then choose \(A\) using the count at dilation \(K_1\), whose constants do not depend on \(A\). The count at the larger dilation \(K_0=L^{A+2}\) is used only after this comparison. Fix the threshold in (112), the separated-count exponent defining \(D\), and the good-path thresholds. These choices determine the polynomial powers in the motion, prediction, and residual estimates.

Now choose \(M=L^{A_M}\) as in (252). For large \(L\), \(\mu<1/12\), so this fixed exponent controls both \[pL^C M^{-1/2+\mu} \quad\text{and}\quad pL^C M^{-1/3+\mu}\] in the marked and unary residual estimates. The good-path thresholds were fixed before \(A_M\), and their required proportion remains valid: \[e^{-L}L^{Cp}(1+M)^{\alpha-1} =\exp\big(-L+O(\sqrt L\log L)+O(\log L)\big)<\tfrac12\] for sufficiently large \(L\).

Finally choose \(L_0\) after the outer length and all fixed exponents. It ensures \(\tau<a\), the geometric smallness requirements, and all the inequalities used above, because \[\log B=O(L^{.7}(\log L)^2)=o(L^{.9}),\qquad \log(N^\mu)=O(\sqrt L)=o(\log G),\] and, for every fixed exponent \(C\), \[\log(m^C\tau)=O(L^{.9})-\Theta(L^{.97})\longrightarrow-\infty.\] It also enforces the saving in (259). All choices are independent of the depth, tile count, and dimension.

Theorem 51 (The three model estimates). For the packet forms in (33), let \(u\) be a bounded measurable real slope, let \(F,E\subset\mathbb R^2\) be measurable sets of positive finite measure, and suppose \(\|f\|_{\mathcal H}\le\mathbf 1_F\) and \(\|g\|_{\mathcal H}\le\mathbf 1_E\) pointwise. The following estimates hold uniformly over every finite subcollection.

  1. For \(m\ge1\), \[|\Lambda(f,g)|\lesssim(1+\log(2m))^C\sqrt{|F||E|}.\]

  2. Suppose \(L\ge L_0\), with \(\tau,m,W\) chosen by (17), and \(L/2\le\log\sqrt{|E|/|F|}\le2L\). If \(\left|u\right|\le1\), \(\mathop{\mathrm{Lip}}(u)\le C_{\rm Lip}\), and \(F,E\subset I\times\mathbb R\) for an interval \(I\) of length at most \(b_0\), then \[|\Lambda(f,g)|\lesssim L^{-20}\sqrt{|F||E|}.\]

  3. For the scalar model with \(m=1\) and \(|E|\le|F|\), there is an absolute \(c>0\) such that \[|\Lambda(f,g)|\lesssim \sqrt{|F||E|}\left(\frac{|E|}{|F|}\right)^c.\]

Proof. The first and third assertions are precisely the ordinary estimates proved with (43). For the second, the fixed deletion of tiles with \(D_E(s)\le L^{-C_2}\) in 2.7 contributes at most \(CL^{-40}\sqrt{|F||E|}\), by (68). Assign the remaining tiles to the popular tops as already constructed. Their adjoint sum is \(P\), and the coefficient convention in Equation (33) gives \[|\Lambda_{\mathrm{popular}}(f,g)| \le\int_F\|P\|.\] Theorem 50 bounds this by \(L^{-20}N|F|=L^{-20}\sqrt{|F||E|}\). Adding the deleted contribution proves the second assertion. All preceding arguments apply to arbitrary finite tile sets and thus give the claimed uniformity. ◻

Completion of the proofs of [thm:band-restricted,thm:main]. Theorem 51 supplies exactly the three hypotheses of Proposition 15. That proposition therefore proves the band restricted estimate (4). The dependence here is forward: the reductions established the implication from the model estimates, which have now been proved without using the band restricted theorem.

All estimates were obtained first for finite packet collections and positive inner truncations, with constants independent of both. The summable smooth expansions and translated-lattice averages in the packet reduction consequently pass to their limits as specified in Proposition 15. The amplitude summation, vertical Littlewood–Paley reassembly, and spatial localization proved in the reductions give a uniform strong \(L^2\) estimate for the original fixed-field transform with a positive inner cutoff. For a Schwartz input its symmetric principal value exists pointwise; Fatou’s lemma gives the same strong bound for the principal-value operator, which extends by density. For the already chosen outer length \(a=a_*\), Chebyshev’s inequality gives the weak \((2,2)\) assertion of Theorem 1. The constants are independent of the vector field, the input, and the level in that inequality. The dilation calculation following 1 gives the outer length \(a_*/\mathop{\mathrm{Lip}}(v)\) for positive Lipschitz seminorm; constant directions have already been handled by the one-dimensional theorem. ◻

Bateman, Michael. 2008. \(L^p\) Estimates for Maximal Averages Along One-Variable Vector Fields in \(\mathbb R^2\). arXiv:0802.0183v1. https://arxiv.org/abs/0802.0183v1.
Bateman, Michael. 2013. “Single Annulus \(L^p\) Estimates for Hilbert Transforms Along Vector Fields.” Revista Matemática Iberoamericana 29 (3): 1021–69. https://doi.org/10.4171/RMI/748.
Bateman, Michael, and Christoph Thiele. 2013. “\(L^p\) Estimates for the Hilbert Transforms Along a One-Variable Vector Field.” Analysis & PDE 6 (7): 1577–600. https://doi.org/10.2140/apde.2013.6.1577.
Bourgain, Jean. 1989. “A Remark on the Maximal Function Associated to an Analytic Vector Field.” In Analysis at Urbana, Vol. I (Urbana, IL, 1986–1987), vol. 137. London Mathematical Society Lecture Note Series. Cambridge University Press. https://doi.org/10.1017/CBO9780511662294.006.
Calderón, Alberto P. 1965. “Commutators of Singular Integral Operators.” Proceedings of the National Academy of Sciences of the United States of America 53 (5): 1092–99. https://doi.org/10.1073/pnas.53.5.1092.
Carleson, Lennart. 1966. “On Convergence and Growth of Partial Sums of Fourier Series.” Acta Mathematica 116: 135–57. https://doi.org/10.1007/BF02392815.
Charalambides, Marcos, and Michael Christ. 2011. Near-Extremizers of Young’s Inequality for Discrete Groups. arXiv:1112.3716. https://arxiv.org/abs/1112.3716.
Coifman, Ronald R., and Yves Meyer. 1978. “Commutateurs d’intégrales Singulières Et Opérateurs Multilinéaires.” Annales de l’Institut Fourier 28 (3): 177–202. https://doi.org/10.5802/aif.708.
Di Plinio, Francesco, Shaoming Guo, Christoph Thiele, and Pavel Zorin-Kranich. 2018. “Square Functions for Bi-Lipschitz Maps and Directional Operators.” Journal of Functional Analysis 275 (8): 2015–58. https://doi.org/10.1016/j.jfa.2018.07.005.
Di Plinio, Francesco, and Ioannis Parissis. 2018. “A Sharp Estimate for the Hilbert Transform Along Finite Order Lacunary Sets of Directions.” Israel Journal of Mathematics 227 (1): 189–214. https://doi.org/10.1007/s11856-018-1724-y.
Eisner, Tanja, and Terence Tao. 2012. “Large Values of the Gowers–Host–Kra Seminorms.” Journal d’Analyse Mathématique 117: 133–86. https://doi.org/10.1007/s11854-012-0018-2.
Fournier, John J. F. 1977. “Sharpness in Young’s Inequality for Convolution.” Pacific Journal of Mathematics 72 (2): 383–97. https://doi.org/10.2140/pjm.1977.72.383.
Guo, Shaoming. 2015. “Hilbert Transform Along Measurable Vector Fields Constant on Lipschitz Curves: \(L^2\) Boundedness.” Analysis & PDE 8 (5): 1263–88. https://doi.org/10.2140/apde.2015.8.1263.
Guo, Shaoming. 2017a. “Hilbert Transform Along Measurable Vector Fields Constant on Lipschitz Curves: \(L^p\) Boundedness.” Transactions of the American Mathematical Society 369 (4): 2493–519. https://doi.org/10.1090/tran/6750.
Guo, Shaoming. 2017b. “Single Annulus Estimates for the Variation-Norm Hilbert Transforms Along Lipschitz Vector Fields.” Proceedings of the American Mathematical Society 145 (2): 601–15. https://arxiv.org/abs/1610.05233.
Guo, Shaoming, and Christoph Thiele. 2017. “Hilbert Transforms Along Lipschitz Direction Fields: A Lacunary Model.” Mathematika 63 (2): 351–63. https://doi.org/10.1112/S0025579316000280.
Hajłasz, Piotr. 2012. Harmonic Analysis. Author lecture notes, 28 March 2012. https://sites.pitt.edu/~hajlasz/Notatki/Harmonic%20Analysis2.pdf.
John, Fritz, and Louis Nirenberg. 1961. “On Functions of Bounded Mean Oscillation.” Communications on Pure and Applied Mathematics 14 (3): 415–26. https://doi.org/10.1002/cpa.3160140317.
Jones, Roger L., Andreas Seeger, and James Wright. 2008. “Strong Variational and Jump Inequalities in Harmonic Analysis.” Transactions of the American Mathematical Society 360 (12): 6711–42. https://doi.org/10.1090/S0002-9947-08-04538-8.
Lacey, Michael T., and Xiaochun Li. 2006. “Maximal Theorems for the Directional Hilbert Transform on the Plane.” Transactions of the American Mathematical Society 358 (9): 4099–117. https://doi.org/10.1090/S0002-9947-06-03869-4.
Lacey, Michael, and Xiaochun Li. 2006. On a Lipschitz Variant of the Kakeya Maximal Function. arXiv:math/0601213v2. https://arxiv.org/abs/math/0601213v2.
Lacey, Michael, and Xiaochun Li. 2010. On a Conjecture of E. M. Stein on the Hilbert Transform on Vector Fields. Vol. 205. Memoirs of the American Mathematical Society. American Mathematical Society. https://doi.org/10.1090/S0065-9266-10-00572-7.
Lacey, Michael, and Christoph Thiele. 2000. “A Proof of Boundedness of the Carleson Operator.” Mathematical Research Letters 7 (4): 361–70. https://doi.org/10.4310/MRL.2000.v7.n4.a1.
Lépingle, Dominique. 1976. “La Variation d’ordre \(p\) Des Semi-Martingales.” Zeitschrift für Wahrscheinlichkeitstheorie Und Verwandte Gebiete 36 (4): 295–316. https://doi.org/10.1007/BF00532696.
Lerner, Andrei K., and Fedor Nazarov. 2019. “Intuitive Dyadic Calculus: The Basics.” Expositiones Mathematicae 37 (3): 225–65. https://doi.org/10.1016/j.exmath.2018.01.001.
Stein, Elias M., and Brian Street. 2012. “Multi-Parameter Singular Radon Transforms III: Real Analytic Surfaces.” Advances in Mathematics 229 (4): 2210–38. https://doi.org/10.1016/j.aim.2011.11.016.
Tao, Terence. 2006. Lecture Notes 4 for 247A. UCLA course notes, Fall 2006. https://www.math.ucla.edu/~tao/247a.1.06f/notes4.pdf.
Tao, Terence. 2007. Lecture Notes 6 for 247B. UCLA course notes, Winter 2007. https://www.math.ucla.edu/~tao/247b.1.07w/notes6.pdf.
Zorin-Kranich, Pavel. 2020. “Weighted Lépingle Inequality.” Bernoulli 26 (3): 2311–18. https://doi.org/10.3150/20-BEJ1194.
LEVEL 1 COMPLETE!
You read 39,013 words and 3,537 formulas. Your math teacher would be proud.
Converted from the LaTeX source. Something look off? The original PDF is the real thing.

Cool Links: openai/math   Lean   Mathlib   arXiv   the real Coolmath Games