A
D
V
E
R
T
I
S
E
M
E
N
T
ADVERTISEMENT
The ℓ¹-Bass Conjecture for Discrete Groups
expertly designed by an internal OpenAI model  ·  released 2026-10-05  ·  original PDF
Theorems: 2 Lemmas: 13 Proofs: 18
Formulas: 1,238 Words: 13,476 Play time: ~1 hour

>>> How to Play <<<
We prove the ℓ1-Bass conjecture for every discrete group. The Hattori–Stallings trace of every idempotent matrix over the complex ℓ1 group algebra is supported on finitely many conjugacy classes of finite-order elements.

>>> Level Map <<<
  1. Introduction
  2. History and relation to earlier methods
  3. Proof strategy
  4. Reduction to exponential decay
  5. Cyclic simplex formulas for the trace
  6. Tuples and their weights
  7. Compatible forms
  8. Smooth summaries at a spatial scale
  9. Smooth selection of centers
  10. Predictions with an enlarged radius
  11. Accurate finite sample strings
  12. Smooth weights on strings
  13. Predictions at the original radius
  14. A connection from successive time averages
  15. Time averages and spatial summaries
  16. Lag pairings and their compatibility
  17. Contraction and the positive normalizer
  18. Differential bounds and the resulting connection
  19. Curvature estimates and the Pfaffian bound
  20. Smooth stages as weighted operators
  21. The terminal sign kernel
  22. The raw stage and nearby coordinates
  23. Expansion of the normalized curvature
  24. Proof of the trace-support theorem

Introduction

Let \(G\) be a discrete group. Its complex Banach group algebra \(\ell^1(G)\) consists of functions \(a\colon G\to\mathbb C\) with \(\|a\|_1=\sum_{g\in G}|a(g)|<\infty\), equipped with convolution \[(a*b)(g)=\sum_{h\in G}a(h)b(h^{-1}g).\] Write \(\mathcal C(G)\) for the set of conjugacy classes, and \(\mathcal C_{\rm fin}(G)\) for those consisting of finite-order elements. The subscript refers to element order, not to the size of a conjugacy class.

For an idempotent matrix \(e=(e_{ij})\in M_q(\ell^1(G))\), its Hattori–Stallings trace (Hattori 1965; Stallings 1965) has coefficients \[ \tau_C(e)=\sum_{g\in C}\sum_{i=1}^q e_{ii}(g), \qquad C\in\mathcal C(G). \tag{1}\] These sums converge absolutely, and \(\sum_C|\tau_C(e)|\le\sum_i\|e_{ii}\|_1\). The trace is additive under direct sums and depends only on the isomorphism class of the projective module \(e\ell^1(G)^q\). It therefore defines \[\operatorname{HS}^1\colon K_0(\ell^1(G))\longrightarrow \ell^1(\mathcal C(G)),\] where \(K_0\) is the Grothendieck group of finitely generated projective right modules. The \(\ell^1\)-Bass conjecture (Berrick et al. 2004, Conjecture 2.2) asserts that this map takes values in the algebraic direct sum over \(\mathcal C_{\rm fin}(G)\). We prove this assertion for all discrete groups.

Theorem 1. Let \(G\) be a discrete group, let \(q\ge1\), and let \(e\in M_q(\ell^1(G))\) be idempotent. There is a finite set \(F\subset\mathcal C_{\rm fin}(G)\) such that \(\tau_C(e)=0\) for every \(C\notin F\). Equivalently, \[\operatorname{HS}^1\bigl(K_0(\ell^1(G))\bigr) \subseteq \bigoplus_{C\in\mathcal C_{\rm fin}(G)}\mathbb C[C].\]

The finite-support assertion is part of the conclusion, in addition to vanishing on infinite-order classes. It is not an automatic consequence of absolute summability. In particular, the proof must make its geometric cutoff independent of the conjugacy class under consideration.

History and relation to earlier methods

The Hattori–Stallings trace refines the rank of a projective module by taking values in the quotient of its ring by additive commutators (Hattori 1965; Stallings 1965). For the algebraic group rings \(\mathbb ZG\) and \(\mathbb CG\), this quotient is indexed by conjugacy classes. Bass’s conjecture for the integral group ring asks that this trace be supported at the identity (Bass 1976); the complex-group-ring formulation permits finite-order classes. These support questions connect the algebraic structure of projective modules with the conjugacy structure of a group.

Cyclic homology gives one route to these questions. Burghelea’s computation separates group-ring cyclic homology into contributions from finite- and infinite-order conjugacy classes (Burghelea 1985). Eckmann used this description and the Chern character to prove Bass-type vanishing for several classes of groups (Eckmann 1986). Emmanouil extended this method through nilpotence properties of the cyclic periodicity operator (Emmanouil 1998). Our argument retains cyclic symmetry but works directly with weighted configurations and differential forms.

Berrick, Chatterji, and Mislin formulated the \(\ell^1\)-Bass conjecture in the finite-support form used here. They proved it for countable groups whose Bost assembly map is rationally surjective in degree zero (Berrick et al. 2004, Theorem 1.4). Together with Lafforgue’s Banach-algebra \(K\)-theory (Lafforgue 2002), this includes amenable groups; the passage from countable subgroups to arbitrary groups is explained in (Berrick et al. 2004, Lemma 1.5). Theorem 1 removes the assembly hypothesis while retaining the algebraic finite-support conclusion.

Ji, Ogle, and Ramsey developed related Bass conjectures for rapid-decay completions and higher Chern characters (Ji et al. 2010, 2014). Their work highlights the growth estimates needed when passing from group rings to analytic completions. Here the summability needed in the proof comes from a similarity reduction to an exponentially decaying idempotent. The subsequent estimates use a spatial scale uniform over all classes to be excluded.

The weighted cyclic configurations and Pfaffian detection strategy build on the complex-group-ring argument in (OpenAI 2026, secs. 7–8). In the group-ring setting, nonzero products of coefficients involve group elements from a fixed finite set. For an \(\ell^1\) idempotent these elements may have arbitrarily large word length, even after reduction to exponential decay. The estimates must retain summable length factors without losing control as the number of factors increases. We obtain them from smooth spatial summaries whose derivative bounds are independent of the size of the input support, multiscale time averaging with a positive translation component, and a final ordered-kernel bound uniform in the closing group element. A single spatial radius then excludes all but finitely many finite-order classes. All ingredients used here are proved below; the complex-group-ring theorem is not an input. Standard concentration and transgression arguments are credited at their points of use.

Proof strategy

The argument uses Banach algebra power series, differential forms on finite-dimensional simplices, and elementary estimates for integral operators. We give the needed estimates in full.

First, a finite-support approximation followed by an analytic idempotent correction replaces \(e\) by a similar idempotent \(p\) with an exponential moment in a finitely generated subgroup (Section 2). This step preserves every trace coefficient. It does not assert that \(p\) has finite support.

For a fixed \(g\in G\), consider cyclic lists of group elements whose closing edge is twisted by multiplication by \(g\). Products of the coefficients of \(p\) assign complex weights to these lists. Their total weight is \(\tau_{[g]}(p)\), and idempotency makes the weights compatible with deleting vertices. The ordered times of a list form a simplex. On these simplices, a one-form invariant under simultaneous time translation and equal to one on the translation vector will be called a connection form. A cyclic Stokes argument shows that the weighted average of its curvature Pfaffian equals the trace coefficient, up to sign (Section 3).

The task is then to construct compatible connection forms with small averaged curvature. Interpret a timed list as a step function with values in \(G\), and average its Dirac masses on a sequence of time scales. The resulting probability distributions are summarized by smooth subprobability distributions on auxiliary labels. These summaries record spatial concentration while controlling derivatives caused by moving one boundary of the step function (Section 4). Pairing their differentials with bounded odd time kernels gives one-forms. Their contractions with time translation telescope across the averaging scales. A final equality kernel detects the twist and leaves a strictly positive contraction, which allows normalization to a connection (Section 5).

The curvature estimates use three different features of this construction. At the initial scale, a curvature entry vanishes unless its two times are close; an occupancy estimate controls the resulting sparse Pfaffians on average. At intermediate scales, the short support of the time kernel gives a Hilbert–Schmidt bound. At the last scale, the equality kernel reduces to weighted order kernels, whose singular values satisfy \(s_\nu\le C/\nu\). Expanding the Pfaffian across these stages gives exponential decay in the simplex dimension (Section 6). A single spatial radius works for every infinite-order class and every finite-order class avoiding a fixed finite word ball. This uniformity proves the finite-support conclusion in Section 7.

Reduction to exponential decay

The first step replaces a general \(\ell^1\) idempotent by an idempotent whose coefficients have exponential decay in a finitely generated subgroup. The replacement preserves every trace coefficient exactly. This permits the polynomially weighted sums that arise from differential forms in arbitrarily large, but fixed, dimensions.

For a complex matrix \(A=(A_{ij})\), write \(|A|_1=\sum_{i,j}|A_{ij}|\). On \(M_q(\ell^1(G))\) we use \[\|a\|_1=\sum_{h\in G}|a(h)|_1.\] This norm is submultiplicative. The identity matrix has norm \(q\); none of the power-series arguments below requires its norm to be one.

Lemma 2 (Traciality). For every conjugacy class \(C\) of a discrete group \(G\), the functional \[T_C(a)=\sum_{h\in C}\mathop{\mathrm{tr}}a(h),\qquad a\in M_q(\ell^1(G)),\] is continuous and satisfies \(T_C(ab)=T_C(ba)\). It is therefore invariant under similarity by invertible matrices over \(\ell^1(G)\).

Proof. The bound \(|T_C(a)|\le\|a\|_1\) gives continuity. Expand \(T_C(ab)\) as the absolutely convergent sum \[\sum_{\substack{x,y\in G\\xy\in C}}\mathop{\mathrm{tr}}\bigl(a(x)b(y)\bigr).\] The conditions \(xy\in C\) and \(yx\in C\) are equivalent, and the scalar matrix trace is cyclic. Interchanging \(x\) and \(y\) proves the identity. If \(u\) is invertible, apply it to \(ua\) and \(u^{-1}\). ◻

The same rearrangement applies to rectangular matrices: if \(a\in M_{q,r}(\ell^1(G))\) and \(b\in M_{r,q}(\ell^1(G))\), then \(T_C(ab)=T_C(ba)\), with the trace in the appropriate size on each side. An isomorphism between two finitely generated projective modules, together with its inverse, can be extended by their defining idempotents to such matrices \(a,b\) with products equal to the two idempotents. This proves the module invariance used in the definition of \(\operatorname{HS}^1\).

Proposition 3 (Exponentially decaying representative). Let \(e\in M_q(\ell^1(G))\) satisfy \(e^2=e\). There are a finitely generated subgroup \(\Gamma\le G\), a word length \(|\cdot|\) on \(\Gamma\), a number \(c>0\), and an idempotent \(p\in M_q(\ell^1(\Gamma))\) such that \[ \sum_{h\in\Gamma}|p(h)|_1e^{c|h|}<\infty, \tag{2}\] and \(p\) is similar to \(e\) in \(M_q(\ell^1(G))\). If the trace of every such \(p\) has finite support on finite-order conjugacy classes of \(\Gamma\), the same conclusion holds for every \(e\) over every discrete group.

Proof. Choose a finite-support matrix \(a\) close to \(e\) and put \(b=4(a-a^2)\). We may require \(\|b\|_1<1\). Let \[Q=(1-b)^{-1/2}=\sum_{j=0}^\infty \frac{\binom{2j}{j}}{4^j}b^j, \qquad p=\frac12 I+\left(a-\frac12 I\right)Q.\] The series converges absolutely in the matrix Banach algebra. All its factors commute with \(a\), and the scalar power-series identity gives \(Q^2(1-b)=I\). Since \((a-I/2)^2=(I-b)/4\), we obtain \((p-I/2)^2=I/4\), hence \(p^2=p\). Moreover \(p\to e\) as \(a\to e\).

Let \(\Gamma\) be generated by the finite union of the supports of the entries of \(a\). Fix a finite symmetric generating set and its word length. The norm \[\|v\|_c=\sum_{h\in\Gamma}|v(h)|_1e^{c|h|}\] is submultiplicative by \(|xy|\le |x|+|y|\). The element \(b\) has finite support, so \(\|b\|_c\to\|b\|_1\) as \(c\downarrow0\). Choose \(c>0\) with \(\|b\|_c<1\). The same series defining \(Q\) then converges in this weighted algebra, proving (2).

Finally, set \[u=pe+(1-p)(1-e).\] Direct multiplication gives \(ue=pu\) and \(u-I=(p-e)(2e-I)\). For \(a\) sufficiently close to \(e\), the latter has norm less than one, so \(u\) is invertible and \(p=ueu^{-1}\). Lemma 2 preserves all the trace coefficients.

For a conjugacy class \(C\) of \(G\), its intersection with \(\Gamma\) is a union of conjugacy classes of \(\Gamma\). Absolute convergence gives \[T_C(p)=\sum_{\substack{D\in\mathcal C(\Gamma)\\D\subset C}}T_D(p).\] Thus passing from \(\Gamma\) to \(G\) merges classes and cannot enlarge a finite set of surviving classes. The order of an element does not change under subgroup inclusion. This proves the final assertion. ◻

Henceforth \(G\) is finitely generated, \(d(x,y)=|x^{-1}y|\) is a fixed left-invariant word metric, and \(p\) satisfies (2). We fix a constant \(\kappa_G\) such that \[ |B(x,s)|\le e^{\kappa_G s}\qquad(x\in G,\ s\ge1). \tag{3}\] Such a bound follows by counting words in the generators. It imposes no additional growth hypothesis on \(G\). Constants without a subscript may depend on this metric and on fixed construction choices. A constant denoted \(C_R\) may also depend on the radius parameter \(R\) introduced later, but will not depend on a tuple, its twist \(g\), or the dimension \(n\).

Cyclic simplex formulas for the trace

We now express a conjugacy-class trace coefficient as a weighted average of curvature Pfaffians. Idempotency provides compatibility of the weights under deletion of a vertex; cyclicity supplies the cancellation that makes the average independent of the connection form. The tuple weights have the same algebraic origin as the weighted paths in (OpenAI 2026, sec. 7); we give the identities needed here directly.

Tuples and their weights

Fix \(g\in G\) and write \(H=C_G(g)\) for its centralizer. For \(k\ge0\) let \[\mathcal T_k(g)=H\backslash G^{k+1},\] where \(H\) acts by simultaneous left multiplication. For a tuple \(h=(h_0,\ldots,h_k)\) extend its indices to all integers by \(h_{i+k+1}=gh_i\) and define \[ v_{k+1}(h)=\mathop{\mathrm{tr}}\left(\prod_{i=0}^k p(h_i^{-1}h_{i+1})\right), \qquad l_i=1+d(h_{i-1},h_i). \tag{4}\] The product is in the displayed order. Both expressions descend to \(\mathcal T_k(g)\), and \(l_i\) is periodic with period \(k+1\).

Lemma 4 (Weight identities). The following properties hold.

  1. The weights have total sum \(\tau_{[g]}(p)\), in every dimension.

  2. Twisted rotation \((h_0,\ldots,h_k)\mapsto(h_1,\ldots,h_k,gh_0)\) preserves the weights.

  3. For \(k\ge1\), summing the weights over all insertions of one entry into a fixed shortened tuple gives its weight in \(\mathcal T_{k-1}(g)\).

All these sums converge absolutely, even after multiplication by any fixed polynomial in the edge lengths. In particular, with \(M_p=\sum_{w\in G}|p(w)|_1(1+|w|)\), \[ \sum_{h\in\mathcal T_k(g)}|v_{k+1}(h)|\prod_{i=0}^k l_i \le M_p^{k+1}. \tag{5}\]

Proof. Put \(w_i=h_i^{-1}h_{i+1}\). Their product is \(h_0^{-1}gh_0\). Conversely, any list \((w_0,\ldots,w_k)\) whose product lies in \([g]\) determines a tuple: choose \(h_0\) realizing this conjugate and successively put \(h_{i+1}=h_iw_i\). Two choices of \(h_0\) differ by left multiplication by an element of \(H\), and this is the only ambiguity. We therefore have a bijection between \(\mathcal T_k(g)\) and such increment lists.

The inequality \(|\mathop{\mathrm{tr}}(A_0\cdots A_k)|\le\prod_i|A_i|_1\) gives (5) upon dropping the restriction on the product of the increments. The exponential moment gives the same conclusion for any fixed polynomial in their lengths. The unrestricted convolution sum with product in \([g]\) is the trace coefficient of \(p^{k+1}=p\), proving (i). Rotation is cyclic permutation of the matrix factors, so (ii) follows from their trace.

To delete an interior entry, sum the two adjacent factors over that entry. Their convolution is \(p^2=p\). The other deletions follow by rotation. One may fix a representative of the shortened tuple because its simultaneous left \(H\)-action has trivial stabilizer. Thus no multiplicity enters the insertion sum, and (iii) follows. ◻

Compatible forms

For each tuple use the ordered time domain \[\widetilde\Delta_k =\{(t_0,\ldots,t_k):t_0\le t_1\le\cdots\le t_k\le t_0+1\}.\] Extend times by \(t_{i+k+1}=t_i+1\), and set \(x_i=t_{i+1}-t_i\) for \(0\le i\le k\). Thus \(x_i\ge0\) and \(\sum_i x_i=1\). Simultaneous time translation is generated by \(V=\sum_{i=0}^k\partial_{t_i}\). Its quotient is the simplex \[\Delta_k=\{0\le t_1\le\cdots\le t_k\le1\},\] using the section \(t_0=0\), oriented by \(dt_1\wedge\cdots\wedge dt_k\). We regard \([t_i,t_{i+1})\) as carrying the label \(h_i\). The face \(x_i=0\) corresponds to deleting \((h_i,t_i)\): the interval carrying the label \(h_i\) has length zero. At the closing face use the periodic lifts to make this same convention.

Definition 5. A compatible connection family through dimension \(n\) assigns a real one-form \(\alpha_h\) to each \(h\in\mathcal T_k(g)\), \(0\le k\le n\), on \(\widetilde\Delta_k\), with the following properties:

  1. It is invariant under time translation, and \(\alpha_h(V)=1\).

  2. It is invariant under simultaneous left multiplication by \(H\) and under twisted rotation of both labels and times.

  3. Its tangential restriction to \(x_i=0\) agrees with the form of the shortened tuple obtained by deleting \((h_i,t_i)\).

  4. The form is smooth up to every face, by restriction of a smooth extension, and the coefficients of \(\alpha_h\) and \(d\alpha_h\) are bounded by polynomials in the edge lengths, for each fixed dimension and fixed construction parameters.

Translation invariance and \(\alpha(V)=1\) imply \(\iota_Vd\alpha=0\). Thus \(d\alpha\) is a form on \(\Delta_k\) independent of the choice of translation section. A translation-invariant form whose contraction with \(V\) is zero will be called basic.

For a real skew-symmetric \(2m\times2m\) matrix \(A\), our convention is \[\omega=\sum_{i<j}A_{ij}\,dt_i\wedge dt_j \quad\Longrightarrow\quad \omega^m=m!\mathop{\mathrm{Pf}}(A)\,dt_1\wedge\cdots\wedge dt_{2m}.\] In particular \(\mathop{\mathrm{Pf}}(A)^2=\det A\). This identity follows by reducing a skew form to two-dimensional blocks, with both sides transforming by the square of the determinant of a change of basis.

Proposition 6 (Trace as a curvature average). Let \(n=2m>0\) and let \(\alpha\) be a compatible connection family through dimension \(n\). Write \(A_h\) for the matrix of \(d\alpha_h\) on the section \(t_0=0\). If \(\mathbb E_{\Delta_n}\) denotes uniform probability on \(\Delta_n\), then \[ \sum_{h\in\mathcal T_n(g)}v_{n+1}(h)\, \mathbb E_{\Delta_n}\mathop{\mathrm{Pf}}(A_h) =(-1)^m\tau_{[g]}(p). \tag{6}\] The series converges absolutely.

Proof. First compare two compatible families \(\alpha\) and \(\alpha'\). Their difference is basic. The following transgression form is the elementary abelian instance of the connection-invariance mechanism in Chern–Weil theory (Chern and Simons 1974, secs. 2–3): \[\eta=(\alpha-\alpha')\wedge \sum_{j=0}^{m-1}(d\alpha)^j\wedge(d\alpha')^{m-1-j}\] is basic, has degree \(n-1\), and satisfies \(d\eta=(d\alpha)^m-(d\alpha')^m\). It respects all the face and rotation identifications. Stokes’ theorem and Lemma 4(iii) express the weighted sum of \(\int_{\Delta_n}d\eta\) as a signed sum of copies of \[I=\sum_{h\in\mathcal T_{n-1}(g)}v_n(h) \int_{\Delta_{n-1}}\eta_h.\] Each face identification has a fixed orientation sign, independent of the labels. The polynomial bounds and the exponential moment justify Stokes term by term and all the rearrangements.

On the simplex of dimension \(n-1\), rotation cyclically permutes its \(n\) gap coordinates. It has orientation sign \((-1)^{n-1}=-1\). It preserves the weights and the forms. Basicness ensures that returning to the section \(t_0=0\) has no additional effect on the integral. Consequently \(I=-I\), so \(I=0\). The weighted integral of \((d\alpha)^m\) is therefore independent of the compatible family.

Use now the particular connection \[\alpha^{\rm can}=\sum_{i=0}^k x_i\,dt_i.\] Its contraction with \(V\) is one, and all other conditions are immediate from the gap coordinates: on a zero gap the corresponding summand vanishes and adjacent time differentials agree tangentially. On \(\widetilde\Delta_k\) its curvature is \[d\alpha^{\rm can}=\sum_{i=0}^k dt_{i+1}\wedge dt_i.\] For \(k=n\) in the section \(t_0=0\), the terms at either end vanish, leaving the path of adjacent pairs. Its Pfaffian is \((-1)^m\): the only complete pairing uses \((1,2),(3,4),\ldots,(n-1,n)\), each with coefficient \(-1\).

Finally, \(\Delta_n\) has Euclidean coordinate volume \(1/n!\), so \[ \frac{n!}{m!}\int_{\Delta_n}(d\alpha_h)^m =\mathbb E_{\Delta_n}\mathop{\mathrm{Pf}}(A_h). \tag{7}\] Evaluation at \(\alpha^{\rm can}\) and the total-weight identity give (6). ◻

The proposition reduces trace vanishing to a quantitative construction. We shall choose connections, in arbitrarily large even dimensions, for which the average of the absolute Pfaffian is at most a small exponential factor times \(\prod_{i=1}^n l_i\). Equation (5) then permits summation against the tuple weights. The next section constructs the spatial summaries used to obtain these connections.

Smooth summaries at a spatial scale

The connection forms will be built from smooth summaries of probability distributions on the group. A summary must retain information about spatial concentration while changing in a controlled way when mass moves between two group elements. We construct two versions. The first allows the spatial radius to increase by a factor of six and has a derivative bound independent of that radius. The second retains the radius and has a polynomial dependence on it. Their different bounds will be used at different stages of the averaging construction.

Throughout this section, \(G\) is a finitely generated group with a fixed left invariant word metric \(d\). Write \(B(y,s)=\{h:d(h,y)\le s\}\). Counting words in a finite symmetric generating set gives a constant \(c_d\) such that \[ |B(y,s)|\le \exp(c_d s)\qquad(s\ge1). \tag{8}\] Indeed, if the generating set has \(m\) elements, the ball is covered by words of length at most \(\lfloor s\rfloor\), and their number is at most \((m+1)^{\lfloor s\rfloor+1}\). Increasing \(c_d\) also covers the trivial group. Constants denoted by \(C\) may depend on this word metric but not on a radius, an accuracy parameter, or the size of an input support.

Let \(\mathcal P_f(G)\) be the set of finitely supported probability distributions on \(G\). For \(r\ge1\) set \[ \begin{split} \chi_r(s)&=\min\{1,\max\{0,2-s/r\}\},\qquad s\ge0,\\ E_r(P,y)&=\sum_{h\in G}P_h\chi_r(d(h,y)),\\ q_r(P)&=\max_{y\in G}E_r(P,y),\qquad \phi_r(P)=(q_r(P)-a_*)_+,\qquad a_*=0.80. \end{split} \tag{9}\] The maximum exists: the score vanishes outside the finite \(2r\) neighborhood of \(\mathop{\mathrm{supp}}P\). It lies in \([0,1]\), so \(0\le\phi_r\le0.20\). The functions \(q_r\) and \(\phi_r\) need not be smooth; they will only be used in inequalities.

Definition 7. A channel with label set \(\Lambda\), on which \(G\) acts, assigns to each \(P\in\mathcal P_f(G)\) a finitely supported subprobability \(Y(P)=(Y(P)_\omega)_{\omega\in\Lambda}\). Thus \(Y(P)_\omega\ge0\) and \(\sum_\omega Y(P)_\omega\le1\). Translation of weights means \((sP)_h=P_{s^{-1}h}\) and \((sY)_\omega=Y_{s^{-1}\omega}\). We require left equivariance, \(Y(sP)=sY(P)\) for \(s\in G\), and the following smoothness property. For each finite \(S\subset G\), the restriction to probabilities supported in \(S\) has its labels in a fixed finite subset of \(\Lambda\), and its coordinate functions extend smoothly to a neighborhood of that probability simplex in \(\mathbb R^S\).

In a transfer direction \(\dot P=\delta_h-\delta_{h'}\), derivatives are computed after enlarging \(S\) to contain \(h,h'\). The channels we construct will have compatible formulas on these enlarged spaces, and their derivatives will vanish at zero output coordinates. For such a channel define its Fisher seminorm in this direction by \[ \mathcal F_Y(P;h,h') =\left(\sum_{\omega:Y(P)_\omega>0} \frac{|\dot Y(P)_\omega|^2}{Y(P)_\omega}\right)^{1/2}. \tag{10}\] The zero-coordinate assertion is part of the construction, and will be proved; it is not a convention for discarding a nonzero derivative.

If \(A\) has label set \(G\) and \(L\) has label set \(\Lambda\), a cross kernel is a function \(\kappa:G\times\Lambda\to[0,1]\) satisfying \(\kappa(sy,s\omega)=\kappa(y,\omega)\). It is fixed independently of \(P\). It gives the predictions \[ \operatorname{pred}(P,y) =\sum_{\omega\in\Lambda}\kappa(y,\omega)L(P)_\omega. \tag{11}\] Exact predictions alone would be easy: \(L(P)=P\) and \(\kappa(y,h)=\chi_r(d(h,y))\) give \(\operatorname{pred}(P,y)=E_r(P,y)\). When \(G\) has at least two elements, this choice has no uniform Fisher bound: a transfer between distinct elements, adding mass to an entry of size \(t>0\), contributes \(1/t\) to the squared seminorm, which diverges as \(t\downarrow0\). The approximation allowed in the next proposition permits uniform derivative control, including at probability faces.

Proposition 8 (Spatial summary channels). Fix \(r\ge1\) and \(0<\epsilon\le10^{-3}\). There are channels \(A,L\) and a cross kernel \(\kappa\) with either of the following two choices of radius: the early version has \(r'=6r\), and the later version has \(r'=r\). In both versions \(A\) is labelled by \(G\), and for every \(P\in\mathcal P_f(G)\) and \(y\in G\), \[\begin{align*} \sum_{z\in G}A(P)_z\bigl(E_r(P,z)-a_*\bigr) &\ge\phi_r(P)-10\epsilon, \tag{12}\\ \operatorname{pred}(P,y) &\le E_{r'}(P,y)+10\epsilon, \tag{13}\\ \operatorname{pred}(P,y) &\ge E_r(P,y)-10\epsilon \qquad\text{if }A(P)_y>0. \tag{14}\end{align*}\] For \(Y=A\) or \(Y=L\), derivatives in every transfer direction vanish at zero coordinates and satisfy \[ \mathcal F_Y(P;h,h')\le C(1+d(h,h')) \begin{cases} \epsilon^{-10},&\text{early version},\\ r^{10}\epsilon^{-10},&\text{later version}. \end{cases} \tag{15}\] Here \(C\) depends only on the fixed word metric. The numerical error constant \(10\) in (12)–(14) is independent of the group and of the parameters.

To see why these guarantees are suited to averaging, assume them for the moment. Take a finite family \(P^{(\nu)}\) of input probabilities and weights \(\theta_\nu\ge0\) with \(\sum_\nu\theta_\nu=1\), and write \[\overline P=\sum_\nu\theta_\nu P^{(\nu)},\qquad \overline A_y=\sum_\nu\theta_\nu A(P^{(\nu)})_y,\qquad \overline{\operatorname{pred}}_y =\sum_\nu\theta_\nu\operatorname{pred}(P^{(\nu)},y).\] The selection and lower prediction bounds give \(\sum_yA(P)_y(\operatorname{pred}(P,y)-a_*)\ge\phi_r(P)-20\epsilon\). On the other hand, the upper prediction bound and linearity of \(E_{r'}\) give \(\overline{\operatorname{pred}}_y \le E_{r'}(\overline P,y)+10\epsilon\). Since \(\overline A\) is a subprobability, its pairing with \(E_{r'}(\overline P,\cdot)-a_*\) is at most \(\phi_{r'}(\overline P)\); if the maximal centered score is negative, that pairing is nonpositive. Finally, \(\sum_\nu\theta_\nu\sum_yA(P^{(\nu)})_y=\sum_y\overline A_y\), so the terms containing \(a_*\) cancel and \[ \begin{split} &\sum_\nu\theta_\nu\sum_y A(P^{(\nu)})_y\operatorname{pred}(P^{(\nu)},y) -\sum_y\overline A_y\overline{\operatorname{pred}}_y\\ &\hspace{2em}\ge \sum_\nu\theta_\nu\phi_r(P^{(\nu)}) -\phi_{r'}(\overline P)-30\epsilon. \end{split} \tag{16}\] Thus loss of concentration under averaging gives a lower bound for the difference of the two pairings. The same calculation applies to time averages of fields with a common finite support, which is how these estimates will enter the connection construction.

We first construct \(A\) and a second distribution \(D\) on possible centers. The channel \(A\) concentrates near a maximizing center when the maximal score exceeds \(a_*\). The thresholds for \(D\) are lower, so that \(D\) already has almost full mass whenever \(A\) is nonzero. The two versions of \(L\) will then use this same \(D\) to predict scores at the selected centers.

Smooth selection of centers

The number of centers that can carry substantial score has a bound independent of \(|\mathop{\mathrm{supp}}P|\).

Lemma 9. For \(P\in\mathcal P_f(G)\), the set \(\{y:E_r(P,y)>1/2\}\) has diameter at most \(4r\) and cardinality at most \(\exp(\gamma r)\), where \(\gamma\) depends only on the word metric.

Proof. The inequality \(E_r(P,y)>1/2\) implies \(P(\{h:d(h,y)<2r\})>1/2\). Any two such sets intersect, since their masses add to more than one. Their centers consequently have distance less than \(4r\). If the set of centers is nonempty, choose one of them. All the others lie in its \(4r\) ball, and (8) gives the claimed bound, for example with \(\gamma=4c_d\). The empty case is immediate. ◻

Choose smooth functions \(\psi_A,\psi_D:\mathbb R\to[0,1]\) such that \[ \begin{array}{c|cc} &\text{identically zero on}&\text{identically one on}\\ \hline \psi_A&(-\infty,0.70]&[0.75,\infty)\\ \psi_D&(-\infty,0.58]&[0.62,\infty). \end{array} \tag{17}\] For completeness, a smooth increasing cutoff from zero to one on \([0,1]\) is \(\eta(t)/(\eta(t)+\eta(1-t))\), where \(\eta(t)=e^{-1/t}\) for \(t>0\) and \(\eta(t)=0\) for \(t\le0\). Every right derivative of \(\eta\) at zero vanishes: away from zero it is \(e^{-1/t}\) times a polynomial in \(1/t\), which tends to zero as \(t\downarrow0\). The denominator in the displayed quotient is everywhere positive, and affine changes of variable give (17). Their first derivatives are bounded by a fixed numerical constant, since they vanish outside fixed compact transition intervals.

Take \(N=M_0r\epsilon^{-2}\) and define \[ \begin{split} w^A_y&=\psi_A(E_r(P,y))^2 e^{N(E_r(P,y)-0.80)},\\ w^D_y&=\psi_D(E_r(P,y))^2 e^{N(E_r(P,y)-0.66)},\\ A(P)_y&=\frac{w^A_y}{1+\sum_z w^A_z},\qquad D(P)_y=\frac{w^D_y}{1+\sum_z w^D_z}. \end{split} \tag{18}\] Here \(M_0\) is fixed in terms of the metric, large enough that for all allowed \(r,\epsilon\), \[ (1+e^{\gamma r})e^{-N\epsilon}\le\epsilon, \qquad e^{\gamma r-0.04N}\le1. \tag{19}\] For example, any \(M_0\ge25\gamma+4\) suffices. Indeed, \(\log(1+e^{\gamma r})+\log(1/\epsilon) \le\gamma r+\log2+1/\epsilon\le(\gamma+2)r/\epsilon\), and \(N\ge25\gamma r\).

The additional \(1\) in each denominator may be regarded as the weight of an unused null label. It allows the channel to have small total mass when no center reaches the exponential threshold. Each displayed sum is finite by Lemma 9.

Lemma 10 (Selector estimates). The formulas (18) define smooth equivariant channels. They satisfy \[ \sum_y A(P)_y(E_r(P,y)-a_*)\ge\phi_r(P)-3\epsilon. \tag{20}\] Whenever \(q_r(P)>0.70\), \[ \sum_zD(P)_z\ge1-\epsilon, \qquad \sum_zD(P)_zE_r(P,z)\ge q_r(P)-2\epsilon. \tag{21}\] Both channels have zero derivatives at zero coordinates, and \[ \mathcal F_A(P;h,h')+\mathcal F_D(P;h,h') \le C\epsilon^{-2}d(h,h'). \tag{22}\]

Proof. Fix \(P\) and abbreviate \(E_r(P,y)\) to \(E_y\) and \(q_r(P)\) to \(q\). Suppose first that \(q>a_*+\epsilon\). At a maximizing center the cutoff \(\psi_A\) is one, and its weight is \(e^{N(q-a_*)}\). The normalized mass of centers with \(E_y<q-\epsilon\), together with the null mass, is at most \((1+e^{\gamma r})e^{-N\epsilon}\le\epsilon\). On the remaining centers \(E_y-a_*\ge q-a_*-\epsilon\); on all centers \(E_y-a_*\ge-1\). Their weighted sum is therefore at least \(q-a_*-3\epsilon\). If \(q\le a_*+\epsilon\), the total mass at centers with \(E_y<a_*-\epsilon\) is at most \(e^{\gamma r}e^{-N\epsilon}\le\epsilon\), using the \(1\) in the denominator. All other negative contributions are bounded below by \(-\epsilon\) times their mass. The sum is at least \(-2\epsilon\), whereas \(\phi_r(P)\le\epsilon\). This proves (20) in both cases.

Now assume \(q>0.70\). The cutoff \(\psi_D\) is one at a maximizing center. The same comparison, with threshold \(0.66\), shows that the total mass of the null label and the centers with \(E_z<q-\epsilon\) is at most \(\epsilon\). Indeed the null-weight ratio is \(e^{-N(q-0.66)}\le e^{-N\epsilon}\). Thus the nonnull mass is at least \(1-\epsilon\), and the score average is at least \((1-\epsilon)(q-\epsilon)\ge q-2\epsilon\). This proves (21).

We next prove the derivative estimate. Put \(d_0=d(h,h')\). The function \(\chi_r\) is \(1/r\)-Lipschitz, so \[ |\dot E_y| =|\chi_r(d(h,y))-\chi_r(d(h',y))| \le d_0/r. \tag{23}\] The following calculation applies to either cutoff, with its corresponding threshold \(\theta\in\{0.80,0.66\}\). Set \[f_y=\psi(E_y)e^{N(E_y-\theta)/2},\qquad F=(1,(f_y)_y),\qquad Z=\|F\|^2=1+\sum_y f_y^2.\] The nonnull coordinates of \(F/\|F\|\) are the smooth nonnegative square roots of the normalized weights. Differentiating a normalized Euclidean vector gives \[ \left\|\frac{d}{dt}\frac{F}{\|F\|}\right\|^2 \le\frac{\|\dot F\|^2}{Z}. \tag{24}\] To see this, the derivative is \(\|F\|^{-1}\) times the orthogonal projection of \(\dot F\) onto \(F^\perp\).

Differentiating \(f_y\) and using \((u+v)^2\le2u^2+2v^2\) gives \[ \frac{\|\dot F\|^2}{Z} \le C\frac{d_0^2}{r^2} \left( \frac{\sum_y|\psi'(E_y)|^2e^{N(E_y-\theta)}}{Z} +N^2\frac{\sum_y f_y^2}{Z} \right). \tag{25}\] The second fraction is at most one. A nonzero cutoff derivative requires \(E_y\le0.75\) for \(A\) or \(E_y\le0.62\) for \(D\). In either case \(E_y-\theta\le-0.04\). Lemma 9 bounds the number of these terms by \(e^{\gamma r}\). The first fraction is consequently at most \(Ce^{\gamma r-0.04N}\le C\). Equations (24) and (25) now bound the squared norm of the square-root derivative by \(C(1+N^2)d_0^2/r^2\). At a positive normalized weight \(Y_y\), \(|\dot Y_y|^2/Y_y=4|\frac{d}{dt}\sqrt{Y_y}|^2\). Substituting \(N=M_0r\epsilon^{-2}\) proves (22).

These computations also justify regularity at zero coordinates. A smooth nonnegative function has zero derivative wherever it vanishes. Thus \(\psi(E_y)=0\) implies \(\dot f_y=0\), even in a transfer direction that leaves a probability face. Hence the normalized weight and its square root both have zero derivative there.

Finally, on distributions supported in a fixed finite \(S\), only centers within its \(2r\) neighborhood can have a nonzero score. The formulas in (18) are therefore finite smooth formulas on \(\mathbb R^S\), with positive denominators. Enlarging \(S\) adds the same formulas for the additional coordinates. This proves the stated smoothness and compatibility at faces. Left equivariance follows from the left invariance of \(d\). ◻

The final averaging stage will only need \(D\), together with its concentration on a bounded set of labels. We record that consequence separately.

Lemma 11 (Terminal channel). For each \(r\ge1\) and \(0<\epsilon\le10^{-3}\), the channel \(D\) in (18) is smooth, left equivariant, and has total mass at most one. Its nonempty support has diameter at most \(4r\) and cardinality at most \(e^{\gamma r}\), with \(\gamma\) depending only on the word metric. If \(q_r(P)>0.70\), its mass is at least \(1-\epsilon\ge1/2\). Its transfer derivatives vanish at zero coordinates and satisfy \[\mathcal F_D(P;h,h')\le C\epsilon^{-2}d(h,h') \le C\epsilon^{-10}(1+d(h,h')).\] In particular, if \(\phi_r(P)>0\), then \[ \sum_zD(P)_z^2\ge\frac14e^{-\gamma r}. \tag{26}\]

Proof. Every center in the support of \(D\) has score greater than \(0.58>1/2\). Apply Lemma 9 and Lemma 10. If \(\phi_r(P)>0\), then \(q_r(P)>0.80\), so \(\sum_zD_z\ge1/2\). The Cauchy–Schwarz inequality on the finite support gives \(\sum_zD_z^2\ge(\sum_zD_z)^2/|\mathop{\mathrm{supp}}D|\), proving (26). ◻

Predictions with an enlarged radius

The early channel can use the average score of the centers chosen by \(D\). All these centers lie near any center selected by \(A\). Enlarging the radius allows us to bound each score contributing at \(y\) by \(E_{6r}(P,y)\). Define \[ \Lambda=G,\qquad L(P)_z=D(P)_zE_r(P,z),\qquad \kappa(y,z)=\mathbf1_{\{d(y,z)\le4r\}}. \tag{27}\] This is a subprobability channel, since \(0\le E_r(P,z)\le1\). When \(d(y,z)\le4r\), every \(h\) with \(\chi_r(d(h,z))>0\) has \(d(h,y)<6r\). On this set \(\chi_{6r}(d(h,y))=1\), and therefore \[ E_r(P,z)\le E_{6r}(P,y). \tag{28}\] After summing against \(D\), this proves \(\operatorname{pred}(P,y)\le E_{6r}(P,y)\).

If \(A(P)_y>0\), then \(E_r(P,y)>0.70\) and \(q_r(P)>0.70\). For every \(z\) with \(D(P)_z>0\), Lemma 9 gives \(d(y,z)\le4r\). Hence every such \(z\) contributes to the prediction, and (21) gives \[\operatorname{pred}(P,y) =\sum_zD(P)_zE_r(P,z) \ge q_r(P)-2\epsilon \ge E_r(P,y)-2\epsilon.\]

To estimate derivatives, write \(E_z=E_r(P,z)\) again. Where \(D_z>0\), \(E_z>0.58\). Differentiating \(L_z=D_zE_z\) yields \[ \sum_{z:L_z>0}\frac{|\dot L_z|^2}{L_z} \le 2\sum_{z:D_z>0}E_z\frac{|\dot D_z|^2}{D_z} +2\sum_{z:D_z>0}D_z\frac{|\dot E_z|^2}{E_z} \le C\epsilon^{-4}d(h,h')^2. \tag{29}\] In the second sum only \(D_z>0\) is relevant; its denominators are thus bounded below by \(0.58\). We used (23), (22), and \(r\ge1\). At a zero coordinate \(D_z=0\) and \(\dot D_z=0\), so \(\dot L_z=0\). The same finite formulas prove smoothness at faces. This establishes all the early-version claims of Proposition 8.

Accurate finite sample strings

To keep the radius equal to \(r\), an average of center scores is insufficient: we must predict the score separately at each \(y\). A finite string of sample points approximates all scores in a bounded region simultaneously. Random sampling will prove the existence of an accurate string. We will then put smooth positive weights on all strings in a fixed finite alphabet. This distinction is useful at probability faces: the defining weights will not be products of possibly vanishing input probabilities.

Fix a center \(z\in G\). Introduce a null symbol \(\partial\), fixed by the group action, and the alphabet \[\mathcal A_z=B(z,20r)\cup\{\partial\}.\] Define a probability \(Q^z\) on this alphabet by \[ P^z_h=P_h\chi_{10r}(d(h,z)),\qquad Q^z_h=P^z_h\ (h\ne\partial),\qquad Q^z_\partial=1-\sum_hP^z_h. \tag{30}\] The notation \(E_r(P^z,y)\) has the same linear meaning as in (9), although \(P^z\) can have mass less than one. This truncation provides a fixed finite alphabet while preserving every score needed for the lower prediction bound. Indeed, if \(A(P)_yD(P)_z>0\), Lemma 9 gives \(d(y,z)\le4r\). Every \(h\) contributing to \(E_r(P,y)\) then has \(d(h,z)<6r\), where the factor \(\chi_{10r}(d(h,z))\) is one. Consequently \[ E_r(P^z,y)=E_r(P,y) \qquad\text{if }A(P)_yD(P)_z>0. \tag{31}\] For arbitrary \(y\), we still have \(E_r(P^z,y)\le E_r(P,y)\), which is the direction needed for the upper bound. Thus accurate predictions for \(P^z\) will suffice. The Lipschitz taper in (30) will also keep their transfer derivatives proportional to \(d(h,h')/r\).

For a positive integer \(b\) and a string \(w=(w_1,\ldots,w_b)\in\mathcal A_z^b\) let \[ H_y(w)=\frac1b\sum_{i:w_i\ne\partial}\chi_r(d(w_i,y)), \qquad \Delta(P;z,w)=\max_{y\in B(z,22r)} |H_y(w)-E_r(P^z,y)|. \tag{32}\] Both scores vanish when \(d(y,z)>22r\), by the triangle inequality. Thus \(\Delta\) also controls the discrepancy at every \(y\in G\). Always \(0\le\Delta\le1\).

Lemma 12. There is a constant \(C\), depending only on the word metric, and an integer \(b\) depending only on \(r,\epsilon\), with \(1\le b\le Cr\epsilon^{-4}\), such that for every \(P\) and \(z\) some \(w\in\mathcal A_z^b\) satisfies \(\Delta(P;z,w)\le\epsilon\).

Proof. Take \(b\) independent samples with distribution \(Q^z\). For a fixed test center \(y\), each summand in \(H_y\) lies in \([0,1]\) and has mean \(E_r(P^z,y)\). The bounded-variable concentration estimate of Hoeffding (Hoeffding 1963) gives the needed probability bound. We include a direct exponential-moment proof, with sufficient constants. If a real random variable \(X\) has mean zero and \(|X|\le1\), then, for \(|t|\le1\), \[\mathbf E e^{tX} \le1+\sum_{k\ge2}\frac{|t|^k}{k!} \le1+t^2\le e^{t^2}.\] For independent copies \(X_1,\ldots,X_b\), Markov’s inequality and independence, with \(t=\epsilon/2\), imply \[\mathbf P\!\left(\frac1b\sum_iX_i>\epsilon\right) \le e^{-tb\epsilon}\prod_i\mathbf E e^{tX_i} \le e^{-b\epsilon^2/4}.\] Applying the same argument to \(-X_i\) gives the bound \(2e^{-b\epsilon^2/4}\) for absolute error exceeding \(\epsilon\).

By (8), there is a fixed \(c_1\) such that the number of test centers is at most \(e^{c_1r}\). A union bound therefore gives \[\mathbf P\bigl(\Delta(P;z,w)>\epsilon\bigr) \le2e^{c_1r-b\epsilon^2/4}.\] Choose \(b=\lceil M_1r\epsilon^{-4}\rceil\) with a sufficiently large fixed \(M_1\). The right side is strictly less than one for all permitted \(r,\epsilon\). At least one string is accurate. Since \(r\epsilon^{-4}\ge1\), rounding preserves the asserted upper bound on \(b\). ◻

Smooth weights on strings

Fix the integer \(b\) supplied by Lemma 12. To assign smooth weights, replace the maximum of the signed errors by the logarithm of a sum of exponentials. Put \[ S(P;z,w)=\frac1\lambda\log \left(\sum_{y\in B(z,22r)}\sum_{\sigma\in\{-1,1\}} \exp\bigl(\lambda\sigma(H_y(w)-E_r(P^z,y))\bigr)\right). \tag{33}\] Choose \(\lambda>0\), depending only on \(r,\epsilon\), so large that the logarithm of the number of signed tests is at most \(\lambda\epsilon\). The number of tests is independent of \(z\), since left translation identifies the balls. The elementary inequalities \(e^{\lambda\max x_i}\le\sum_i e^{\lambda x_i} \le M e^{\lambda\max x_i}\) give \[ \Delta(P;z,w)\le S(P;z,w)\le\Delta(P;z,w)+\epsilon. \tag{34}\]

The first derivative of this smoothed maximum has a particularly simple bound. The test function \(h\mapsto\chi_{10r}(d(h,z))\chi_r(d(h,y))\) is \(11/(10r)\)-Lipschitz: both factors have absolute value at most one, and their Lipschitz constants add. Thus every signed error in (33) has transfer derivative of absolute value at most \(11d(h,h')/(10r)\). Differentiation of the logarithm gives a convex combination of these derivatives; the factor \(\lambda\) cancels. Consequently \[ |\dot S(P;z,w)|\le\frac{11}{10r}d(h,h'). \tag{35}\] In particular this estimate does not grow with the number of tests or the smoothing parameter.

Now define a probability distribution on the entire fixed set of strings by \[ \pi_z(P)_w =\frac{e^{-N_1S(P;z,w)}}{\sum_{v\in\mathcal A_z^b}e^{-N_1S(P;z,v)}}, \qquad N_1=M_2rb\epsilon^{-2}. \tag{36}\] Every coordinate is strictly positive. The constant \(M_2\) will be fixed in terms of the metric.

Lemma 13. For \(M_2\) sufficiently large, uniformly in \(P,z,y\), \[\begin{align*} \sum_w\pi_z(P)_w\Delta(P;z,w)&\le5\epsilon, \tag{37}\\ \left|\sum_w\pi_z(P)_wH_y(w)-E_r(P^z,y)\right| &\le5\epsilon. \tag{38}\end{align*}\] The transfer derivatives satisfy \[ \left(\sum_w\frac{|\dot\pi_z(P)_w|^2}{\pi_z(P)_w}\right)^{1/2} \le C N_1\frac{d(h,h')}{r}. \tag{39}\]

Proof. There are at most \(e^{c_2rb}\) strings, for a constant \(c_2\) fixed by the metric. By Lemma 12 and (34), at least one string has \(S\le2\epsilon\). Every string with \(\Delta>4\epsilon\) has \(S>4\epsilon\), and hence a weight at most \(e^{-2N_1\epsilon}\) times the weight of that particular string. Therefore \[\sum_{w:\Delta(P;z,w)>4\epsilon}\pi_z(P)_w \le e^{c_2rb-2N_1\epsilon}\le\epsilon\] once \(M_2\) is sufficiently large. One fixed choice works for all \(r\ge1\), \(b\ge1\), and \(\epsilon\le10^{-3}\), as is seen by taking logarithms and using \(\log(1/\epsilon)\le1/\epsilon\). Since \(\Delta\le1\), splitting the expectation at \(4\epsilon\) proves (37). The discrepancy controls each test center, and both scores vanish outside the test ball, proving (38) for every \(y\).

Differentiating the normalized exponential in (36) gives \[\frac{\dot\pi_z(P)_w}{\pi_z(P)_w} =-N_1\left(\dot S(P;z,w) -\sum_v\pi_z(P)_v\dot S(P;z,v)\right).\] Equation (35) bounds its absolute value by \(C N_1d(h,h')/r\). Squaring and summing against the probability \(\pi_z(P)\) proves (39). ◻

Predictions at the original radius

We can now finish the later channel. Its label set consists of pairs \((z,w)\), with \(z\in G\) and \(w\in\mathcal A_z^b\). Translation acts on the center and on every nonnull entry of the string. Define \[ L(P)_{z,w}=D(P)_z\pi_z(P)_w, \qquad \kappa(y,(z,w))=H_y(w). \tag{40}\] The total mass of \(L\) is \(\sum_zD(P)_z\le1\), and the kernel takes values in \([0,1]\). All formulas commute with left translation.

By (38), and since \(P^z_h\le P_h\), \[\operatorname{pred}(P,y) \le\sum_zD(P)_z\bigl(E_r(P^z,y)+5\epsilon\bigr) \le E_r(P,y)+5\epsilon.\] Suppose next that \(A(P)_y>0\). For every \(z\) with \(D(P)_z>0\), (31) identifies the predicted truncated score with \(E_r(P,y)\). Moreover \(q_r(P)>0.70\), and (21) gives \(\sum_zD(P)_z\ge1-\epsilon\). It follows that \[\operatorname{pred}(P,y) \ge E_r(P,y)\sum_zD(P)_z-5\epsilon\sum_zD(P)_z \ge E_r(P,y)-6\epsilon.\] These are the required prediction bounds with \(r'=r\).

For the derivative bound it is useful to retain the product structure \(L_{z,w}=D_z\pi_z(w)\). Since \(\sum_w\pi_z(w)=1\) and \(\sum_w\dot\pi_z(w)=0\), expansion of the square gives the exact identity \[ \sum_{z,w:L_{z,w}>0}\frac{|\dot L_{z,w}|^2}{L_{z,w}} =\sum_{z:D_z>0}\frac{|\dot D_z|^2}{D_z} +\sum_zD_z\sum_w\frac{|\dot\pi_z(w)|^2}{\pi_z(w)}. \tag{41}\] All weights and derivatives here are real; the mixed term vanishes by the displayed zero-sum identity. Using (22), (39), and \(b\le Cr\epsilon^{-4}\) yields \[ \mathcal F_L(P;h,h') \le C\bigl(\epsilon^{-2}+N_1/r\bigr)d(h,h') \le Cr\epsilon^{-6}d(h,h'). \tag{42}\]

It remains to check the regularity promised in the channel definition. For a fixed finite input support \(S\), only finitely many centers \(z\) can have \(D_z\ne0\), all in the \(2r\) neighborhood of \(S\). For each such center, the alphabet and string set are finite and independent of \(P\). The functions \(S(P;z,w)\) and \(\pi_z(P)_w\) are smooth functions of the coordinates of \(P\): the scores are linear, and each exponential sum has a strictly positive denominator. Thus (40) is a fixed finite smooth formula on each enlarged support space. If \(L_{z,w}=0\), strict positivity of \(\pi_z(w)\) implies \(D_z=0\); also \(\dot D_z=0\) by Lemma 10. Consequently \(\dot L_{z,w}=0\). This proves smoothness at faces and the required behavior at zero entries.

Completion of the proof of Proposition 8. Use the same selector \(A\) in both versions. Inequality (20) gives (12). The early construction (27) gives the two prediction inequalities with errors \(0\) and \(2\epsilon\); the later construction (40) gives errors \(5\epsilon\) and \(6\epsilon\). All are bounded by \(10\epsilon\). The Fisher estimates (22), (29), and (42) imply (15), since \(r\ge1\) and \(\epsilon\le1\). The preceding constructions also verify the channel and cross-kernel requirements, including the assertions at zero coordinates. ◻

For use with differential forms, regard the labels of \(A\) and \(L\) as disjoint even when their underlying sets happen to coincide. On their disjoint union define the symmetric kernel \(K\) by \[ \begin{split} K(y,\omega)=K(\omega,y)&=\tfrac12\kappa(y,\omega) \quad(y\in G,\ \omega\in\Lambda),\\ K&=0\quad\text{within either channel}. \end{split} \tag{43}\] For vectors on this union, write \(K(U,V)=\sum_{\xi,\eta}K(\xi,\eta)U_\xi V_\eta\). The combined field \(Y=A\oplus L\) has mass at most two, and \[ K(Y(P),Y(P)) =\sum_yA(P)_y\operatorname{pred}(P,y). \tag{44}\] The kernel is bounded and invariant under simultaneous left translation. The expression in (10) also makes sense for this field of mass at most two, and its square is \(\mathcal F_A(P;h,h')^2+\mathcal F_L(P;h,h')^2\). Thus Proposition 8 controls its derivatives as well. These facts permit both channels to enter a single bilinear pairing in the averaging construction.

A connection from successive time averages

Proposition 6 allows us to compute the trace with any connection satisfying Definition 5. We now construct a connection whose curvature can be estimated at several time scales. The construction first gives a one-form \(B\) with uniformly positive contraction \(B(V)\); division by this contraction will produce the connection. The summary channels of Proposition 8 make the contributions of successive scales telescope. The final contribution uses the twist by \(g\).

Throughout this section \(G\) is finitely generated, \(d\) is its fixed left invariant word metric, and \(g\in G\). Constants denoted by \(C\) may depend on the word metric and on the fixed smooth functions used below. After the parameter choice at the end of the construction, a subscript \(R\) will permit dependence on the radius \(R\), but never on the tuple, its dimension, or the element \(g\).

Time averages and spatial summaries

Begin with a finite number \(b_*\geq2\) of averaging stages, an even ambient dimension \(n\), and positive time scales \(\delta_1<\cdots<\delta_{b_*}\) satisfying \[ \sum_{i=1}^j\delta_i\leq2\delta_j, \qquad 10\delta_{b_*}<1. \tag{45}\] The construction through positivity uses only these conditions on the time scales. Their precise dependence on \(R\) and \(n\) will be selected afterwards, when preparing the quantitative estimates. The same scales are kept fixed in every dimension \(0\leq k\leq n\). Thus passing to a face never changes the averaging operation.

Choose a nonnegative, even \(C^\infty\) function \(\mu\), supported in \([-1,1]\), with integral one, and put \(\mu_j(s)=\delta_j^{-1}\mu(s/\delta_j)\). On a tuple of dimension \(k\), write its step function and its averages as \[ \begin{split} f(u)&=h_i\quad(t_i\leq u<t_{i+1}),\qquad f(u+1)=gf(u),\\ P_0(u)&=\delta_{f(u)},\qquad P_j(u)=\int_{\mathbb R}\mu_j(s)P_{j-1}(u-s)\,ds. \end{split} \tag{46}\] Zero-length intervals are omitted. These are finitely supported probability distributions at each time, and \(P_j(u+1)=gP_j(u)\). If \(\eta_j=\mu_j*\cdots*\mu_1\), then \[ P_j=\eta_j*P_0,\qquad \mathop{\mathrm{supp}}\eta_j\subset[-2\delta_j,2\delta_j],\qquad \|\eta_j\|_\infty\leq C/\delta_j. \tag{47}\] The support bound follows from (45), and the supremum bound follows by convolving \(\mu_j\) with probability densities.

To summarize these averages spatially, choose \(R\geq1\) and an integer \(2\leq J_*\leq b_*\). Put \[ r_j=6^{\min(j,J_*)-1}R,\qquad r=6^{J_*-1}R. \tag{48}\] The radius increases by a factor of six during the first \(J_*-1\) steps and is then fixed at \(r\). Use accuracies \[ \epsilon_j=\epsilon_*(1+j)^{-2}\quad(j\geq1), \qquad 30\sum_{j\geq1}\epsilon_j\leq\frac1{10}, \tag{49}\] where \(0<\epsilon_*\leq10^{-3}\) is fixed sufficiently small. The summable errors will leave a positive quantity after telescoping.

For \(1\leq j<b_*\), apply Proposition 8 to \(P_j(u)\) at parameters \((r_j,\epsilon_j)\). Use its early version when \(j<J_*\) and its later version when \(j\geq J_*\). Denote the resulting channels by \(A_j,L_j\) and their prediction kernel by \(\kappa_j\). On the disjoint union of the two label sets put \[ Y_j(u)=A_j(P_j(u))\oplus L_j(P_j(u)). \tag{50}\] Define a symmetric bilinear kernel \(K_j\) to be zero between labels in the same channel and to have value \(\kappa_j(y,\omega)/2\) between the two channels. Thus \[ K_j(Y_j(u),Y_j(u)) =\sum_y A_j(P_j(u))_y\operatorname{pred}_j(P_j(u),y). \tag{51}\] The total mass of \(Y_j(u)\) is at most two, and \(|K_j|\leq1\).

At the first and last stages use \[ \begin{array}{lll} Y_0=P_0,& K_0(h,h')=\chi_R(d(h,h')),\\[2pt] Y_{b_*}(u)=D(P_{b_*}(u)),& K_{b_*}(h,h')=\mathbf1_{h=h'}, \end{array} \tag{52}\] where \(D\) is the channel of Lemma 11 at parameters \((r,\epsilon_{b_*})\). These two fields have mass at most one. All the fields obey \(Y_j(u+1)=gY_j(u)\), and every kernel is invariant under simultaneous left translation of its labels.

The last channel needs a multiplier to compensate for its possible spread over many labels. Lemma 11 gives \(\sum_zD(P)_z^2\geq\frac14e^{-C_Dr}\) when \(\phi_r(P)>0\), for a fixed constant \(C_D\). Choose a fixed \(C_*\geq C_D\) and set \[ W=\exp(C_*r). \tag{53}\] Since \(0\leq\phi_r\leq1-a_*=1/5\), this gives, for every accuracy parameter in (49), \[ W\sum_zD(P)_z^2\geq\phi_r(P). \tag{54}\] We shall construct the connection under either of the assumptions \[ \begin{array}{ll} \textnormal{(i)}&g\text{ has infinite order},\\[2pt] \textnormal{(ii)}&g\text{ has finite order and } \displaystyle\min_{h\in G}d(h,gh)>4r. \end{array} \tag{55}\] Only the final time-lag kernel and its contraction distinguish these two cases.

Lag pairings and their compatibility

We next turn the fields into one-forms. For \(0\leq j<b_*\) let \(\rho_j=\mu_{j+1}*\mu_{j+1}\) and define the odd function \[ U_j(s)=\operatorname{sign}(s)-2\int_0^s\rho_j(a)\,da. \tag{56}\] It is bounded in absolute value by one and vanishes outside \([-2\delta_{j+1},2\delta_{j+1}]\). In the sense of distributions, \[ U_j'=2(\delta_0-\rho_j). \tag{57}\] Here \(\delta_0\) denotes the Dirac measure at zero, rather than a time scale. At the final stage set \[ \begin{array}{lll} U_{b_*}(s)=\operatorname{sign}(s),&\rho_{b_*}=0, &\text{in case (i)},\\[2pt] U_{b_*}(s)=\operatorname{sign}(s)\mathbf1_{|s|<1},& \rho_{b_*}=\tfrac12(\delta_{-1}+\delta_1), &\text{in case (ii)}. \end{array} \tag{58}\] Formula (57) holds at this stage as well. Values at the jump points have no effect on any integral below.

The differential \(dY_j(u)\) in the following definition acts on the tuple times \(t_0,\ldots,t_k\), with \(u\) fixed. Write \(\int_{\rm per}\) for integration over one period, using a half-open interval when integrating the raw impulse measure, and put \[ \begin{split} B_j&=-\frac12\int_{\rm per}du\int_{\mathbb R}dv\, U_j(v-u)K_j(dY_j(u),Y_j(v)),\\ B&=\sum_{j=0}^{b_*-1}B_j+WB_{b_*}. \end{split} \tag{59}\] For \(j=0\) the differential of the step field is a measure in \(u\). Specifically, set \[ J_i=\delta_{h_{i-1}}-\delta_{h_i},\qquad dP_0(u)=\sum_{i=0}^k\sum_{a\in\mathbb Z} g^aJ_i\,\delta_{t_i+a}(u)\,dt_i. \tag{60}\] The periodic convention gives \(h_{-1}=g^{-1}h_k\). Thus the coefficient of \(dt_i\) in the raw form is \[ (B_0)_i=-\frac12\int_{\mathbb R}U_0(v-t_i)K_0(J_i,P_0(v))\,dv. \tag{61}\] The expression in (60) is a distributional identity in the interior of the ordered time domain. Formula (61) will also specify its smooth extension to the faces.

Lemma 14. The forms in (59) are well defined, independent of the period cut, and smooth up to every simplex face by one-sided extension. They are invariant under simultaneous translation of all times, under the centralizer action, and under twisted rotation. Their tangential restrictions to faces equal the forms for the shortened tuples, using the same parameters and the same fixed integer \(n\).

Proof. We first treat \(j>0\). On a fixed tuple the smoothed probabilities have the explicit expression \[ P_j(u)=\sum_{a\in\mathbb Z}\sum_{i=0}^k \left(\int_{t_i+a}^{t_{i+1}+a}\eta_j(u-v)\,dv\right)\delta_{g^ah_i}. \tag{62}\] For \(u\) in a bounded interval and times in a neighborhood of any point of the closed ordered domain, only a fixed finite set of summands occurs. Each coefficient is smooth in \(u\) and the times, including when the endpoints coincide. The channel maps are smooth on the finite probability simplices, with smooth extensions at their faces, so \(Y_j\) has the same regularity. Its possible labels on one period lie in a fixed finite set locally in the time variables.

When \(U_j\) has compact support, the two integrals in (59) can therefore be differentiated under fixed integration domains. The final full-sign kernel in case (i) requires an additional observation. Write \(v=w+a\) with \(w\) in one period. If \(y,z\) are possible labels in that period, then \[K_{b_*}(y,g^az)\ne0\quad\Longleftrightarrow\quad y=g^az.\] Since \(g\) has infinite order, this equation has at most one solution \(a\in\mathbb Z\) for each pair \((y,z)\). There are only finitely many such pairs locally. The full-sign integral is therefore locally a finite sum of smooth integrals over bounded domains, even though the possible matching integers need not be uniformly bounded as the tuple varies.

To check the period cut, regard \[F_j(u)=\int_{\mathbb R}U_j(v-u)K_j(dY_j(u),Y_j(v))\,dv\] as a one-form in the time parameters. The identities \(Y_j(u+1)=gY_j(u)\) and \(dY_j(u+1)=g\,dY_j(u)\), followed by the change of variable \(v\mapsto v+1\), give \(F_j(u+1)=F_j(u)\). Hence its integral over one period is independent of the cut. This also shows why a moving cut does not introduce an extra exterior-derivative term: if the cut is \(c=c(t)\), the endpoint contribution is \[ dc\wedge\bigl(F_j(c+1)-F_j(c)\bigr)=0. \tag{63}\]

For the raw stage, work first away from the faces and choose a fixed period cut away from its finitely many jumps. Evaluating the impulse measure gives (61); moving an impulse to another period leaves this coefficient unchanged, by equivariance and a change of variable. To prove one-sided smoothness explicitly, let \[Q(s)=\int_0^sU_0(w)\,dw =|s|-2\int_0^s\int_0^w\rho_0(a)\,da\,dw.\] Then \((B_0)_i\) is a locally finite sum of terms \[ -\frac12K_0(J_i,\delta_{g^ah_q}) \bigl(Q(t_{q+1}+a-t_i)-Q(t_q+a-t_i)\bigr), \quad 0\leq q\leq k. \tag{64}\] Although \(Q\) is constant rather than zero outside the lag support, each displayed endpoint difference vanishes when its interval misses that support. Since all boundaries lie in one period and \(2\delta_1<1\), only the shifts \(a=-1,0,1\) can contribute, with the same finite choice valid in a sufficiently small neighborhood of any point of the closed domain. On the entire closed ordered domain, each affine expression \(t_q+a-t_i\) has a fixed weak sign: the order of the boundaries and all their period translates is prescribed. The same is true for \(t_{q+1}+a-t_i\). Thus every absolute value in (64) equals a fixed signed affine function on that domain. The remaining part of \(Q\) is smooth on the whole line. Consequently the coefficient has a smooth extension from the prescribed side at each face, including intersections of faces. Compact support of \(U_0\) allows the finite set of summands to be fixed in a neighborhood.

It remains to identify the restrictions. On the face \(t_i=t_{i+1}\), where \(0\leq i<k\), the interval labelled \(h_i\) disappears and the adjoining impulses combine tangentially as \[J_i\,dt_i+J_{i+1}\,dt_{i+1} =\bigl(\delta_{h_{i-1}}-\delta_{h_{i+1}}\bigr)dt_i.\] This is precisely the impulse for the shortened tuple. On the closing face \(t_k=t_0+1\), combine the impulses at that common lift: \[J_k+gJ_0 =\delta_{h_{k-1}}-\delta_{gh_0}.\] Their time differentials agree tangentially, and the displayed sum is the closing impulse after deleting \(h_k\). These identities also handle corners by successive merging. In the smoothed stages, the underlying step function itself is unchanged by deleting a zero-length interval; (62) then gives equality of the fields and their tangential differentials. Hence every \(B_j\) respects the faces.

Finally, translating all times by \(s\) replaces \(Y_j(u)\) by \(Y_j(u-s)\). Changing both integration variables by \(s\) and using cut independence proves translation invariance. The centralizer action preserves the twist and all pairings. Twisted rotation describes the very same step function with a different first boundary, so it too preserves the forms. These arguments also establish cut independence of the raw form on the faces, by continuity. ◻

Contraction and the positive normalizer

The forms are now compatible on all tuple simplices. To normalize them to a connection, we must show that their sum has positive contraction with the translation vector \(V\). The choice of the lag derivative in (57) converts this contraction into a difference between a pairing at one time and its average at the next scale.

Lemma 15. For every stage \(0\leq j\leq b_*\), \[ C_j:=B_j(V) =\int_{\rm per}\left( K_j(Y_j(u),Y_j(u)) -\int_{\mathbb R}K_j(Y_j(u),Y_j(u+s))\,d\rho_j(s)\right)du. \tag{65}\] For \(j<b_*\), the second integrated pairing is also \[ \int_{\rm per} K_j\bigl((\mu_{j+1}*Y_j)(u),(\mu_{j+1}*Y_j)(u)\bigr)\,du. \tag{66}\]

Proof. Translation of the times translates the field in the opposite direction, so \(dY_j(V)=-\partial_uY_j\). Define the scalar function \[H_j(u)=\int_{\mathbb R}U_j(v-u)K_j(Y_j(u),Y_j(v))\,dv.\] As in the proof of Lemma 14, it is periodic. For a smooth stage its distributional derivative is \[H_j'(u) =-\int U_j'(v-u)K_j(Y_j(u),Y_j(v))\,dv +\int U_j(v-u)K_j(\partial_uY_j(u),Y_j(v))\,dv.\] Its derivative integrates to zero over a period. Substitution into (59), followed by (57), gives (65), with the factor \(1/2\) cancelling the factor two in \(U_j'\).

The same computation is valid for \(Y_0\) with its jumps. Indeed, for each label \(y\), the function \[u\longmapsto\int U_0(v-u)K_0(\delta_y,P_0(v))\,dv\] is continuous and has bounded weak derivative obtained by replacing \(U_0\) by \(-U_0'\). Integration by parts against the jump measure is therefore valid. A cut away from the jumps makes the boundary cancellation ordinary periodicity; the result extends to the faces by Lemma 14.

For (66), expand both convolutions. If their lags are \(a,b\), simultaneous translation in the period integral changes \(K_j(Y_j(u-a),Y_j(u-b))\) into \(K_j(Y_j(u),Y_j(u+a-b))\). This is legitimate because the scalar pairing is invariant under simultaneous integer shifts of its two time arguments. The distribution of \(a-b\) is \(\mu_{j+1}*\mu_{j+1}=\rho_j\), since \(\mu_{j+1}\) is even. This gives the asserted identity. ◻

Write \[\Phi_j=\int_{\rm per}\phi_{r_j}(P_j(u))\,du \qquad(1\leq j\leq b_*).\] These are the concentration quantities that will telescope.

Lemma 16. Under (55), the scalar function \[ C_{\rm nor}:=B(V)=\sum_{j=0}^{b_*-1}C_j+WC_{b_*} \tag{67}\] satisfies \(C_{\rm nor}\geq c_0\), where \(c_0=1/10\) is independent of all tuples, \(n\), \(R\), and \(g\).

Proof. Consider first \(1\leq j<b_*\). The concentration-loss calculation (16) takes the following form for our time averages. In this paragraph abbreviate \(A_j(P_j(u))\) to \(A(u)\) and the prediction \(\operatorname{pred}_j(P_j(u),y)\) to \(p(u,y)\). Put \(\overline A=\mu_{j+1}*A\) and \(\overline p(\cdot,y)=\mu_{j+1}*p(\cdot,y)\). Equations (51) and (66) give \[C_j=\int_{\rm per}\sum_y A(u)_y p(u,y)\,du -\int_{\rm per}\sum_y\overline A(u)_y\overline p(u,y)\,du.\] The total mass of \(A(u)\) is a periodic scalar function. Averaging preserves its integral, so we may subtract \(a_*\) from both predictions: \[ C_j=\int_{\rm per}\sum_y A(u)_y(p(u,y)-a_*)\,du -\int_{\rm per}\sum_y\overline A(u)_y (\overline p(u,y)-a_*)\,du. \tag{68}\] This cancellation uses the integrated mass; it does not require either channel to have mass one at each time.

The first and third inequalities of Proposition 8 imply \[\sum_y A(u)_y(p(u,y)-a_*) \geq\phi_{r_j}(P_j(u))-20\epsilon_j.\] For the other term, the prediction upper bound, linearity of \(E_{r_{j+1}}\) in its probability argument, and \(P_{j+1}=\mu_{j+1}*P_j\) imply \[\overline p(u,y)-a_* \leq E_{r_{j+1}}(P_{j+1}(u),y)-a_*+10\epsilon_j.\] Here the radius is \(r_{j+1}=6r_j\) for the early version and \(r_{j+1}=r_j\) for the later version, exactly as required by the two channel bounds. Since \(\overline A(u)\) is a subprobability distribution, pairing this inequality with it yields an upper bound \(\phi_{r_{j+1}}(P_{j+1}(u))+10\epsilon_j\). Consequently \[ C_j\geq\Phi_j-\Phi_{j+1}-30\epsilon_j \qquad(1\leq j<b_*). \tag{69}\]

The raw stage supplies the initial positive quantity. Its same-time pairing is one, while (66) gives \(K_0(P_1(u),P_1(u))\) for the averaged pairing. For every probability distribution \(P\), \[K_0(P,P)=\sum_yP_yE_R(P,y)\leq q_R(P) \leq a_*+\phi_R(P).\] Therefore \[ C_0\geq1-a_*-\Phi_1=\frac15-\Phi_1. \tag{70}\]

At the final stage the negative term in (65) vanishes. This is immediate from \(\rho_{b_*}=0\) in case (i). In case (ii), let \(S=\mathop{\mathrm{supp}}D(P_{b_*}(u))\). Lemma 11 gives \(\operatorname{diam}S\leq4r\). If \(z=gw\) with \(z,w\in S\), then \(d(w,gw)\leq4r\), contrary to (55). Thus \(S\cap gS=S\cap g^{-1}S=\varnothing\), and twist equivariance makes both lag pairings at \(s=\pm1\) zero. In either case, \[C_{b_*}=\int_{\rm per}\sum_zD(P_{b_*}(u))_z^2\,du, \qquad WC_{b_*}\geq\Phi_{b_*}\] by (54). Adding this inequality to (69) and (70) leaves \[C_{\rm nor}\geq\frac15-30\sum_{j=1}^{b_*-1}\epsilon_j \geq\frac1{10},\] as claimed. ◻

Differential bounds and the resulting connection

Positivity now allows the normalization. Before applying the trace identity, we also need growth bounds that justify summation over all tuples. More precise bounds on the individual channel differentials will provide the input for the curvature estimates of the next section.

We now specialize the parameters of the preceding construction. At the shortest time scales we use the early channels, whose derivative bound is independent of the spatial radius. The later channels keep that radius fixed; this also fixes the terminal multiplier \(W\) while the time scales continue to increase. Their longer time scales will absorb the polynomial factor \(r^{10}\) in the later derivative bound.

Choose once and for all an integer \(J_*\) with \((3/2)^{J_*-1}>200\). For sufficiently large \(R\) retain the spatial radii, accuracies, and multiplier from (48), (49), and (53), and set \[ \beta_1=R,\qquad \beta_{j+1}=\beta_j^{3/2}. \tag{71}\] Choose \(b_*\) to be the smallest integer such that \[ b_*\geq J_*,\qquad \beta_{b_*}\geq W^2R^{100}. \tag{72}\] For each even integer \(n>10\beta_{b_*}\), use \[ \delta_j=\beta_j/n\qquad(1\leq j\leq b_*). \tag{73}\] The ratio \(\beta_{j+1}/\beta_j\) is at least two when \(R\) is sufficiently large, so these scales satisfy (45). Thus the construction and positivity proved above apply. For each such \(n\), the same actual scales \(\beta_j/n\) are used in every dimension \(k\leq n\); face restriction never changes the denominator to \(k\).

The stage count, spatial radii, accuracies, multiplier, and numbers \(\beta_j\) now depend only on \(R\) and the fixed choices. Since \(\log\beta_j=(3/2)^{j-1}\log R\) and \(\log(W^2R^{100})=2C_*r+100\log R\), with \(r\) a fixed multiple of \(R\), \[ b_*=O(1+\log R). \tag{74}\] The particular exponents in (71) and (72) will enter the curvature bounds. They leave room for the polynomial losses in the channel derivatives.

Choose a sufficiently large fixed constant \(C\) and set \[ a_j=C\epsilon_j^{-10} \begin{cases} 1,&1\leq j<J_*,\\ r^{10},&J_*\leq j\leq b_*. \end{cases} \tag{75}\] As before, \(l_i=1+d(h_{i-1},h_i)\), with periodic indexing, and \(\partial_i=\partial/\partial t_i\).

Proposition 17. Let \(R\) be sufficiently large, choose the parameters in (71) and (72), and let \(n>10\beta_{b_*}\) be even, with time scales (73). Suppose that \(g\) satisfies (55). The forms \(B_j\) and \(B\) in (59), constructed in all dimensions \(0\leq k\leq n\) with the same parameters, have the following properties. For every smooth stage \(1\leq j\leq b_*\) and every coordinate \(t_i\), \[ \left(\int_{\rm per}\sum_\omega \frac{|\partial_iY_j(u)_\omega|^2}{Y_j(u)_\omega}\,du \right)^{1/2} \leq\frac{a_jl_i}{\sqrt{\delta_j}}. \tag{76}\] Terms with zero denominator are defined to be zero. Moreover, \[ |B(\partial_i)|+|\partial_iC_{\rm nor}|\leq C_Rl_i. \tag{77}\] The scalar \(C_{\rm nor}\) is at least \(c_0=1/10\), and \[ \alpha=\frac{B}{C_{\rm nor}} \tag{78}\] is a compatible connection family in the sense of Definition 5. Its coefficients and those of its exterior derivative are bounded by polynomials in the lengths, uniformly over the tuples. In particular, the family may be used in Proposition 6.

Proof. Differentiating (62) at a boundary gives the smeared impulse formula \[ \partial_iP_j(u) =\sum_{a\in\mathbb Z}\eta_j(u-t_i-a)g^aJ_i. \tag{79}\] For each \(u\), at most one summand can be nonzero, because \(4\delta_j<1\). The derivative is supported within \(2\delta_j\) of \(t_i\) modulo periods. The distance between the two labels in \(g^aJ_i\) is \(d(h_{i-1},h_i)\), independently of \(a\). The derivative estimate of Proposition 8, or Lemma 11 at the last stage, therefore bounds the square root of the pointwise sum in (76) by \(C\epsilon_j^{-10}l_i\eta_j(u-t_i-a)\) for \(j<J_*\), with an additional factor \(r^{10}\) for \(j\geq J_*\), where \(a\) is the contributing integer. Since \[\int_{\rm per}\sum_a\eta_j(u-t_i-a)^2\,du =\int_{\mathbb R}\eta_j(s)^2\,ds \leq\|\eta_j\|_\infty\leq C/\delta_j,\] we obtain (76) after fixing the constant in (75). The two disjoint channels only alter this constant. At a zero channel weight the first derivative is zero, as stated in Proposition 8, so the convention in (76) is consistent.

The support information is useful as well as the square-integral bound. The support of \(\partial_iY_j\) has measure at most \(4\delta_j\) in a period, and the field has mass at most two. Cauchy–Schwarz in time and labels consequently gives \[ \int_{\rm per}\sum_\omega|\partial_iY_j(u)_\omega|\,du \leq C a_jl_i. \tag{80}\] For example, the square of the factor paired with (76) is at most \(\int_{\mathop{\mathrm{supp}}\partial_iY_j}\sum_\omega Y_j(u)_\omega\,du \leq8\delta_j\). This explains why the integrated derivative loses no power of the time scale.

We now bound the coefficients of \(B\). For a compact lag kernel, the inner integral in (59) has length at most two, its kernel is bounded, and the field has bounded mass at every time. Thus (80) gives \(|B_j(\partial_i)|\leq Ca_jl_i\). At the full-sign stage there is no bound on the length of the integration interval, but equality of labels gives a stronger estimate. For each fixed group label \(y\), \[ \int_{\mathbb R}Y_{b_*}(v)_y\,dv =\int_0^1\sum_{a\in\mathbb Z}Y_{b_*}(v)_{g^{-a}y}\,dv \leq1. \tag{81}\] The labels \(g^{-a}y\) are distinct because \(g\) has infinite order, and the total mass at every time is at most one. Combining (81) with (80) proves the same coefficient bound. In particular, the estimate does not depend on how far apart the finitely many matching periods are.

Differentiating (65) gives \(|\partial_iC_j|\leq Ca_jl_i\) for \(j>0\). Indeed, each differentiation puts a derivative on one of the two factors. The other factor has bounded total mass, the measure \(\rho_j\) has mass at most one, and (80) is unchanged by a time shift. The last observation follows because the sum of absolute derivatives is a periodic scalar function.

For the raw stage, (61) bounds \(|(B_0)_i|\leq C\), since \(\|J_i\|_1\leq2\) and the lag support has length at most \(4\delta_1<1\). Its contraction satisfies the exact identity \[C_0=1-\int_{\rm per}K_0(P_1(u),P_1(u))\,du.\] Formula (79) for \(P_1\), before passing through any channel map, gives \(\int_{\rm per}\|\partial_iP_1(u)\|_1\,du\leq2\). Hence \(|\partial_iC_0|\leq4\). Summing these estimates with the coefficients \(1,\ldots,1,W\) proves (77), because \[1+\sum_{j=1}^{b_*-1}a_j+Wa_{b_*}\] depends only on \(R\) and the fixed choices.

For completeness, polynomial control of exterior derivatives can be obtained here without using the finer estimates of the next section. For a smooth stage, differentiation at a fixed period cut gives \[ dB_j=\frac12\int_{\rm per}du\int_{\mathbb R}dv\, U_j(v-u)K_j(dY_j(u)\wedge dY_j(v)). \tag{82}\] There is no term from the cut by (63). For a compact lag kernel, reduction of both time variables to one period introduces at most three relevant period shifts. Boundedness of the kernel and (80) therefore imply \[|dB_j(\partial_i,\partial_{i'})| \leq Ca_j^2l_il_{i'}.\] For the full-sign equality kernel, at most one shift contributes for each pair of labels, and the same estimate follows. In the raw stage, differentiate (61) with respect to a distinct boundary. The resulting terms have factors \(U_0(t_{i'}+a-t_i)K_0(J_i,g^aJ_{i'})\); boundedness of \(U_0,K_0\) and the jump masses gives a uniform bound. Differentiation of a coefficient in its own time variable contributes no exterior two-form. These calculations hold in the interior and extend to the closed simplex by Lemma 14. Summation over stages gives \[ |dB(\partial_i,\partial_{i'})|\leq C_Rl_il_{i'}. \tag{83}\]

Finally Lemma 16 permits division by \(C_{\rm nor}\). The scalar normalizer inherits all symmetries and face restrictions of \(B\), so (78) does also, and \(\alpha(V)=1\). Equations (77), (83), and \[d\alpha=C_{\rm nor}^{-1}dB -C_{\rm nor}^{-2}dC_{\rm nor}\wedge B\] give the required polynomial coefficient bounds. The resulting family is therefore a compatible connection family. Its curvature is basic, as also follows directly from translation invariance and \(\alpha(V)=1\). ◻

The construction separates two estimates that the trace argument will use differently. The bounds in Proposition 17 justify the weighted simplex identities, while the scale-dependent bound (76) retains the information needed to make high-dimensional curvature Pfaffians small. We turn to those sharper estimates next.

Curvature estimates and the Pfaffian bound

We estimate the curvature of the connection constructed in Proposition 17. The smooth stages are controlled by operators on weighted Hilbert spaces. The raw stage is instead controlled by the number of nearby simplex coordinates. These two estimates will be combined only after averaging over the coordinate subsets in the Pfaffian expansion.

Throughout this section, the word metric and the fixed construction constants are understood. The constant \(C\) may depend on these data but not on \(R,n,g\), or the tuple. A constant \(C_R\) may additionally depend on \(R\). All parameters \(r,W,b_*,\beta_j,\delta_j\), and \(a_j\) retain their definitions from Section 5. In particular, \(\delta_j=\beta_j/n\) and \(10\delta_{b_*}<1\).

We work in the section \(t_0=0\), on the simplex \[\Delta_n=\{0\le t_1\le\cdots\le t_n\le1\}, \qquad n=2m.\] Expectation with respect to uniform simplex probability will be denoted by \(\mathbb E_{\Delta_n}\); its density in these coordinates is \(n!\). If \(Q\) is a two-form, its matrix is defined by \(Q=\sum_{i<i'}Q_{ii'}\,dt_i\wedge dt_{i'}\). Our convention is \[ \frac{Q^m}{m!} =\mathop{\mathrm{Pf}}(Q)\,dt_1\wedge\cdots\wedge dt_n. \tag{84}\] Equation (7) identifies the expectation below with the normalized integral in Proposition 6. For a subset \(I\subset\{1,\ldots,n\}\) of even cardinality, \(\mathop{\mathrm{Pf}}_I(Q)\) means the Pfaffian of the principal matrix on \(I\), listed in increasing order. We set \(\mathop{\mathrm{Pf}}_{\varnothing}(Q)=1\).

Theorem 18 (Averaged curvature bound). Fix a sufficiently large \(R\), and set \(S=b_*+2\). Let \(n\) be an even integer with \(n>10\beta_{b_*}\). Suppose either that \(g\) has infinite order, or that \(g\) has finite order and \[\min_{h\in G}d(h,gh)>4r.\] For every tuple \((h_0,\ldots,h_n)\), the connection \(\alpha=B/C_{\rm nor}\) of Proposition 17 satisfies \[ \mathbb E_{\Delta_n}|\mathop{\mathrm{Pf}}(d\alpha)| \le C_R\bigl(CS^{3/2}R^{-1/10}\bigr)^n \prod_{i=1}^n l_i, \qquad l_i=1+d(h_{i-1},h_i). \tag{85}\] The constants are independent of \(n,g\), and the tuple. Moreover, \(S=O(1+\log R)\) as \(R\to\infty\).

We prove the theorem in three steps: estimates for principal Pfaffians of the smooth stages, an averaged estimate for the raw stage, and the expansion of the normalized curvature.

Smooth stages as weighted operators

Fix a stage \(j>0\) and a point of the simplex. Let \(\Lambda_j\) be its label set and equip \((0,1)\times\Lambda_j\) with the measure \[d\vartheta_j(u,y)=Y_j(u)_y\,du.\] We use the real Hilbert space \(\mathcal H_j=L^2(\vartheta_j)\). Recall that the total mass of \(Y_j(u)\) is at most two, and is at most one at the last stage. Define the operator \(T_j\) by the kernel \[ T_j(u,y;v,z) =\sum_{a\in\mathbb Z}U_j(v+a-u)K_j(y,g^a z), \tag{86}\] where the action of \(g\) on labels is the one used in the channel construction. Oddness of \(U_j\), symmetry of \(K_j\), and simultaneous equivariance of \(K_j\) show that this kernel is skew. Indeed, exchange \((u,y)\) with \((v,z)\) and replace \(a\) by \(-a\) in the sum.

For \(1\le i\le n\), put \[X_i(u,y)=\frac{\partial_iY_j(u)_y}{Y_j(u)_y} \quad\text{when }Y_j(u)_y>0,\] and give \(X_i\) the value zero at the remaining points. The derivatives of the zero weights vanish, and (76) gives \[ \|X_i\|_{\mathcal H_j} \le \frac{a_jl_i}{\sqrt{\delta_j}}. \tag{87}\] The exterior-derivative identity (82) is \[dB_j=\frac12\int_{\rm per}du\int_{\mathbb R}dv\, U_j(v-u)K_j\bigl(dY_j(u)\wedge dY_j(v)\bigr).\] The integral over a period can be taken with fixed endpoints because its integrand is periodic. Reducing \(v\) to \((0,1)\) gives (86); the two antisymmetric terms then agree. Consequently, \[ (dB_j)_{ii'}=\langle X_i,T_jX_{i'}\rangle_{\mathcal H_j}. \tag{88}\]

Here is the finite-dimensional estimate we shall apply to this pairing. Let \(I\) have even size \(k>0\), let \(E\) be the span of the vectors \(X_i\), \(i\in I\), and let \(s_1\ge s_2\ge\cdots\) be the singular values of the compression \(P_ET_j|_E\), completed by zeros. Then \[ |\mathop{\mathrm{Pf}}_I(dB_j)| \le \prod_{i\in I}\|X_i\| \left(\prod_{\nu=1}^k s_\nu\right)^{1/2}. \tag{89}\] To see this, the matrix in (88) has rank less than \(k\) if \(\dim E<k\), so its Pfaffian is zero. Otherwise write the \(X_i\) as columns of a square matrix \(V\) in an orthonormal basis of \(E\). The curvature matrix is \(V^{\mathsf T}T_EV\). Taking determinants and using \(\mathop{\mathrm{Pf}}(A)^2=\det A\) for a skew matrix gives the right side of (89), after Hadamard’s inequality \(|\det V|\le\prod_i\|X_i\|\).

For \(1\le j<b_*\), the support of \(U_j\) is contained in \([-2\delta_{j+1},2\delta_{j+1}]\). At most one integer \(a\) in (86) contributes for a fixed pair \((u,v)\), because \(4\delta_{j+1}<1\). Thus the kernel is bounded by a fixed constant and vanishes unless the circular distance between \(u\) and \(v\) is at most \(2\delta_{j+1}\). The mass bound on the channels implies \[\|T_j\|_{\rm HS}^2 =\int\!\int |T_j(u,y;v,z)|^2\, d\vartheta_j(u,y)d\vartheta_j(v,z) \le C\delta_{j+1}.\] Bessel’s inequality gives the same upper bound for the squared Hilbert–Schmidt norm of every compression. Applying the arithmetic mean inequality to \(s_1^2,\ldots,s_k^2\) gives \[\left(\prod_{\nu=1}^k s_\nu\right)^{1/2} \le C^k\delta_{j+1}^{k/4}k^{-k/4}.\] Combining this with (87) and (89), and using \(\beta_{j+1}=\beta_j^{3/2}\), proves \[ |\mathop{\mathrm{Pf}}_I(dB_j)| \le\bigl(Ca_j\beta_j^{-1/8}\bigr)^k (n/k)^{k/4}\prod_{i\in I}l_i, \qquad 1\le j<b_*. \tag{90}\] In fact, the powers of the time scales simplify as \[\delta_j^{-k/2}\delta_{j+1}^{k/4}k^{-k/4} =\beta_j^{-k/8}(n/k)^{k/4}.\]

The terminal sign kernel

The last lag kernel is not supported at a small time scale. Its useful property is instead that an order comparison is constant between two disjoint ordered intervals. The following elementary approximation works even when the source and target measures differ.

Lemma 19 (Approximation of order kernels). Let \(\mu,\nu\) be finite nonatomic positive measures on \(\mathbb R\), and set \(M=\mu(\mathbb R)+\nu(\mathbb R)\). Let \(A:L^2(\nu)\to L^2(\mu)\) have any one of the kernels \[\operatorname{sign}(s-t),\qquad \mathbf1_{s<t},\qquad \mathbf1_{s>t},\] where \(t\) is the output variable and \(s\) the input variable. For each integer \(N\ge1\), there is an operator \(A_N\) of rank at most \(N\) such that \[\|A-A_N\|\le \frac{M}{2N}.\] Also \(\|A\|\le\sqrt{\mu(\mathbb R)\nu(\mathbb R)}\le M/2\).

Proof. If \(M=0\), the assertion is immediate. Otherwise partition \(\mathbb R\) into \(N\) successive intervals \(I_1,\ldots,I_N\), each having \((\mu+\nu)\)-measure \(M/N\). Such a partition exists because the measure is nonatomic. Endpoint conventions do not affect the operators. Let \(A_N\) retain the kernel of \(A\) when \(t\) and \(s\) lie in different intervals, and set the kernel to zero when they lie in the same interval. On \(I_a\times I_b\) with \(a\ne b\), the order of \(t\) and \(s\) is fixed. The range of \(A_N\) is therefore contained in the span of the \(N\) functions \(\mathbf1_{I_a}\) in \(L^2(\mu)\).

The remainder \(A-A_N\) maps \(L^2(\nu|_{I_a})\) to \(L^2(\mu|_{I_a})\) for each \(a\), without mixing different intervals. On that block its kernel is bounded by one, so Cauchy–Schwarz gives the operator norm bound \[\sqrt{\mu(I_a)\nu(I_a)} \le \frac{\mu(I_a)+\nu(I_a)}2=\frac{M}{2N}.\] Both the source blocks and the target blocks are orthogonal. Hence the norm of their direct sum is the maximum of their norms, proving the approximation claim. Applying the same Cauchy–Schwarz estimate on the whole line gives the asserted bound on \(\|A\|\). ◻

Lemma 20 (Singular values at the last stage). For the operator \(T_{b_*}\), every finite-dimensional orthogonal compression has singular values satisfying \[ s_\nu\le \frac C\nu,\qquad \nu\ge1. \tag{91}\] The constant is independent of \(n,R,g\), and the tuple, in either of the two cases of Theorem 18.

Proof. We will approximate \(T_{b_*}\), for every \(x>0\), by an operator of rank at most \(C/x\) and norm error at most \(x\).

First suppose that \(g\) has infinite order. The final label set is \(G\) and the label kernel is equality. It follows that \(\mathcal H_{b_*}\) splits orthogonally according to the left \(g\)-orbits in \(G\). In one orbit, choose a representative \(y_0\) and write a label as \(g^a y_0\), \(a\in\mathbb Z\). Associate to the pair \((u,g^a y_0)\) the real coordinate \(u-a\). If the second label is \(g^b y_0\), the equality kernel in (86) selects precisely the shift \(a-b\). The resulting kernel is \[\operatorname{sign}(v+a-b-u) =\operatorname{sign}\bigl((v-b)-(u-a)\bigr).\] Thus this orbit gives a sign operator on a nonatomic measure on the line. Let \(b\) be the mass of that measure. The orbit masses sum to at most one, because the final channel is a subprobability channel. Lemma 19, with the same measure in source and target, gives norm at most \(b\) and a rank-\(N\) approximation with error at most \(b/N\).

Discard each orbit block with \(b\le x\). On every other block choose \(N=\lceil b/x\rceil\). The error of the resulting direct-sum approximation is at most \(x\). Moreover, \[\sum_{b>x}\lceil b/x\rceil \le \frac1x\sum_{b>x}b+\#\{b>x\} \le \frac2x.\] This proves the desired rank bound in the infinite-order case.

Now suppose that \(g\) has finite order. Put \(d\mu_y(u)=Y_{b_*}(u)_y\,du\) on \((0,1)\), extend this measure by zero to \(\mathbb R\), and write \(\mathcal H_y=L^2(\mu_y)\). Since the final lag kernel is \(\operatorname{sign}(s)\mathbf1_{|s|<1}\), the only possible shifts for \(0<u,v<1\) are \(0,1,-1\). The operator is the sum of the three kernels \[ \begin{split} &\operatorname{sign}(v-u)\mathbf1_{y=z},\\ &\mathbf1_{v<u}\mathbf1_{y=gz},\\ &-\mathbf1_{v>u}\mathbf1_{y=g^{-1}z}. \end{split} \tag{92}\] The shifted contributions are direct sums of maps \(\mathcal H_z\to\mathcal H_{gz}\) and \(\mathcal H_z\to\mathcal H_{g^{-1}z}\), respectively. After reindexing the target summands, each is an orthogonal direct sum: its source blocks are mutually orthogonal and its target blocks are mutually orthogonal. The identity contribution has the same property. This remains true even though the measures in a paired block need not coincide; the shifted summands need not be invariant subspaces.

For a paired block, let \(b=\mu_y(\mathbb R)+\mu_z(\mathbb R)\). By Lemma 19, its norm is at most \(b\), and it has a rank-\(N\) approximation with error at most \(b/N\). The sums of these \(b\)’s over a single contribution are at most two: every label mass appears once as a source mass and once as a target mass. Discarding blocks of mass at most \(\varepsilon\) and taking \(N=\lceil b/\varepsilon\rceil\) on the others therefore gives error at most \(\varepsilon\) and total rank at most \(4/\varepsilon\). Use \(\varepsilon=x/3\) for each contribution in (92) and add the three approximations. Their total error is at most \(x\) and their total rank is at most \(C/x\). If \(g\) has order two, the two shift permutations coincide; the separate estimates still apply and are added in exactly the same way.

Finally, orthogonal compression cannot increase either the rank or the error of an approximation. Singular values beyond the dimension of the compressed space are zero, so consider an index \(\nu\) no larger than that dimension. If the compressed matrix has an approximation of rank less than \(\nu\) with error \(x\), its \(\nu\)th singular value is at most \(x\). Indeed, the span of its first \(\nu\) right singular vectors contains a unit vector annihilated by the approximating matrix, and the original matrix has norm at least \(s_\nu\) on that vector. If the rank bound above is \(C_0/x\), take \(x=2C_0/\nu\). The approximating rank is then at most \(\nu/2<\nu\), which proves (91). ◻

For an even principal set \(I\) of size \(k>0\), insert (91) in (89). Multiplication of the two-form by \(W\) multiplies its size-\(k\) Pfaffian by \(W^{k/2}\). We obtain \[|\mathop{\mathrm{Pf}}_I(WdB_{b_*})| \le W^{k/2}\left(\frac{a_{b_*}}{\sqrt{\delta_{b_*}}}\right)^k \frac{C^{k/2}}{\sqrt{k!}}\prod_{i\in I}l_i.\] For \(k\ge1\), integration of \(\log x\) gives \(\log(k!)\ge\int_1^k\log x\,dx\ge k\log k-k\), and hence \(k!\ge(k/e)^k\). We conclude that \[ |\mathop{\mathrm{Pf}}_I(WdB_{b_*})| \le\bigl(Ca_{b_*}\sqrt W\,\beta_{b_*}^{-1/2}\bigr)^k (n/k)^{k/2}\prod_{i\in I}l_i. \tag{93}\] The estimates (90) and (93) control every smooth stage at every point of the simplex. It remains to control the raw stage in the average required for the full Pfaffian.

The raw stage and nearby coordinates

Recall \(J_i=\delta_{h_{i-1}}-\delta_{h_i}\). The coefficient of \(dt_i\) in \(B_0\) is \[-\frac12\int_{\mathbb R}U_0(v-t_i)K_0(J_i,Y_0(v))\,dv.\] For distinct \(i,i'\) in the interior of the simplex, differentiating at the boundary \(t_{i'}\) and antisymmetrizing gives \[(dB_0)_{ii'} =\sum_{a\in\mathbb Z}U_0(t_{i'}+a-t_i)K_0(J_i,g^aJ_{i'}).\] There is at most one nonzero summand. Since \(K_0(h,h')=\chi_R(d(h,h'))\) and \(\chi_R\) is \(1/R\)-Lipschitz, \[|K_0(J_i,g^aJ_{i'})| \le\frac{2d(h_{i-1},h_i)}R \le\frac{2l_i l_{i'}}R.\] The support of \(U_0\) also shows that this matrix entry is zero if the circular distance between \(t_i\) and \(t_{i'}\) exceeds \(2R/n\).

For \(I\subset\{1,\ldots,n\}\) and \(i\in I\), let \(N_i(I)\) be the number of indices in \(I\), including \(i\) itself, whose times have circular distance at most \(2R/n\) from \(t_i\). If \(k=|I|\) is even, divide row and column \(i\) of the principal curvature matrix by \(l_i\). Each resulting row has Euclidean norm at most \((C/R)\sqrt{N_i(I)}\). Hadamard’s inequality, followed by the Pfaffian–determinant identity, yields \[ |\mathop{\mathrm{Pf}}_I(dB_0)| \le (C/R)^{k/2}\prod_{i\in I}l_i \prod_{i\in I}N_i(I)^{1/4}. \tag{94}\]

We next average the last product. The needed assertion is an average over subsets of a fixed size; it is not a bound asserted separately for each fixed subset of the ordered coordinates.

Lemma 21 (Average of the neighbor counts). Let \(R\ge1\), \(n\ge8R\), and \(0\le k\le n\). Choose \((t_1,\ldots,t_n)\) with uniform probability on \(\Delta_n\), and independently choose \(I\) uniformly among the \(k\)-element subsets of \(\{1,\ldots,n\}\). Then \[ \mathbb E_{\Delta_n,I}\prod_{i\in I}N_i(I)^{1/4} \le C^k R^{k/4}, \tag{95}\] with the empty product interpreted as one.

Proof. The statement is immediate for \(k=0\). Uniform simplex coordinates are the order statistics of \(n\) independent uniform points in \((0,1)\). Selecting a uniform set of \(k\) ranks is equivalent, as an unordered set of points, to selecting a uniform set of \(k\) of those independent points. The selected unordered set therefore has the law of \(k\) independent uniform points on the circle. Notice that the fixed endpoint \(t_0=0\) is not one of the selectable coordinates.

Partition the circle into \(m'=\lfloor n/(2R)\rfloor\) equal cells. Their length is at least \(2R/n\), and \(m'\ge n/(4R)\). Let \(p_a\) be the number of selected points in cell \(a\), with cell indices read cyclically. A point in cell \(a\) can have counted neighbors only in that cell or its two adjacent cells. Consequently \[\prod_{i\in I}N_i(I)^{1/4} \le\prod_{a:p_a>0}(p_a+p_{a-1}+p_{a+1})^{p_a/4}.\] The comparison with the product of the occupancies themselves follows from \[\begin{split} \sum_{a:p_a>0}p_a \log\frac{p_a+p_{a-1}+p_{a+1}}{p_a} &\le\sum_{a:p_a>0}(p_{a-1}+p_{a+1})\\ &\le2k. \end{split}\] Thus the preceding product is at most \(e^{k/2}\prod_a p_a^{p_a/4}\), with \(0^0=1\).

The occupancy vector has multinomial probabilities \(k!(m')^{-k}\prod_a(p_a!)^{-1}\). The factorial lower bound implies \(p^{p/4}\le e^{p/4}(p!)^{1/4}\), so \[\mathbb E\prod_a p_a^{p_a/4} \le e^{k/4}\frac{k!}{(m')^k} \sum_{\sum p_a=k}\left(\prod_a\frac1{p_a!}\right)^{3/4}.\] There are \(\binom{k+m'-1}{k}\) nonnegative occupancy vectors. Applying Hölder’s inequality to the last sum, and then the multinomial theorem, gives \[\sum_{\sum p_a=k}\left(\prod_a\frac1{p_a!}\right)^{3/4} \le \left(\frac{(m')^k}{k!}\right)^{3/4} \binom{k+m'-1}{k}^{1/4}.\] After substitution, the remaining factor apart from \(C^k\) is \[\left[\frac{k!}{(m')^k}\binom{k+m'-1}{k}\right]^{1/4} =\left[\prod_{q=0}^{k-1}(1+q/m')\right]^{1/4} \le(1+k/m')^{k/4}.\] Finally, \(k\le n\) and \(m'\ge n/(4R)\) give \(k/m'\le4R\). Since \(R\ge1\), all factors are bounded as in (95). ◻

Combining (94) with Lemma 21 gives a factor of size \((CR^{-1/4})^k\), once the length factors have been extracted and the coordinate subset has been averaged. The remaining task is to combine this subset average with the pointwise estimates for the smooth stages. We first make their scale dependence uniform.

Expansion of the normalized curvature

First record the dependence of the preceding estimates on \(R\). Since \(r\) is a fixed multiple of \(R\), the stopping condition compares \[\log\beta_j=(3/2)^{j-1}\log R, \qquad \log(W^2R^{100})=O(R+\log R).\] Taking \(j\) of order \(1+\log R\) makes the first expression at least the second. As in (74), minimality therefore gives \[ b_*=O(1+\log R). \tag{96}\] For \(j<J_*\), the constants \(a_j\) are bounded independently of \(R\), because this is a fixed finite set of stages. Their bases in (90) are at most \(CR^{-1/8}\). At the later stages, \[a_j\le CR^{10}(1+b_*)^{20}, \qquad \beta_j\ge R^{200},\] so the intermediate bases are at most \(CR^{-15}(1+\log R)^{20}\). At the last stage, the stopping condition \(\beta_{b_*}\ge W^2R^{100}\) gives \[Ca_{b_*}\sqrt W\,\beta_{b_*}^{-1/2} \le CR^{-40}(1+\log R)^{20}W^{-1/2}.\] Together with the raw factor just obtained, all these bases are at most \[ \varepsilon_R=C_1R^{-1/10}\le1 \tag{97}\] for sufficiently large \(R\), after choosing a fixed \(C_1\) large enough.

The normalization of the connection gives \[ d\alpha=C_{\rm nor}^{-1}\sum_{j<b_*}dB_j +C_{\rm nor}^{-1}W\,dB_{b_*}+Q, \qquad Q=-C_{\rm nor}^{-2}dC_{\rm nor}\wedge B. \tag{98}\] The lower bound \(C_{\rm nor}\ge c_0>0\) from Proposition 17 changes the bases in (97) only by a fixed factor; enlarge \(C_1\) accordingly. The last term in (98) is decomposable, so \(Q\wedge Q=0\). Moreover, (77) implies \[ |Q_{ii'}|\le C_Rl_i l_{i'}. \tag{99}\] In particular, this stage can contribute to a Pfaffian partition only on zero or two indices. The constant \(C_R\) in (99) is independent of \(n\).

There are \(S=b_*+2\) stages in (98): the raw stage, the \(b_*\) smooth stages, and \(Q\). Denote their matrices by \(M_0,\ldots,M_{S-1}\), with \(M_{S-1}=Q\). The Pfaffian expansion is \[ \mathop{\mathrm{Pf}}\left(\sum_{j=0}^{S-1}M_j\right) =\sum_{I_0\sqcup\cdots\sqcup I_{S-1}=\{1,\ldots,n\}} \sigma(I_0,\ldots,I_{S-1}) \prod_{j=0}^{S-1}\mathop{\mathrm{Pf}}_{I_j}(M_j), \tag{100}\] where all \(|I_j|\) are even and each sign \(\sigma\) is \(1\) or \(-1\). There is no multiplicity factor for a fixed ordered partition. To check this normalization, expand the divided exterior power \((\sum_jM_j)^m/m!\) as \(\sum_{\sum m_j=m}\prod_j M_j^{m_j}/m_j!\), and then extract the coordinate volume coefficient using (84) on each coordinate subset.

Take absolute values in (100). In each term the length factors multiply to \(\prod_{i=1}^nl_i\), independently of the partition. Fix now the even sizes \(k_j=|I_j|\). For any given raw subset \(I_0\) of size \(k_0\), the number of completions to a partition of these sizes is \[\frac{(n-k_0)!}{\prod_{j=1}^{S-1}k_j!},\] which is independent of \(I_0\). Every factor other than the raw neighbor counts has already been bounded independently of the simplex point and of the actual subset. Thus, upon summing over these partitions and integrating over the simplex, the count product in (94) is averaged uniformly over all raw subsets of size \(k_0\). Lemma 21 applies without any assertion about an individual fixed subset.

The smooth estimates also contain factors \((n/k_j)^{k_j/4}\) or \((n/k_j)^{k_j/2}\). Their product is at most \[ \prod_{j:k_j>0}(n/k_j)^{k_j/2}\le S^{n/2}. \tag{101}\] For completeness, put \(p_j=k_j/n\) on the positive sizes. Concavity of the logarithm gives \[\sum_jp_j\log(1/p_j) \le\log\left(\sum_jp_j/p_j\right)\le\log S;\] exponentiating \(n/2\) times this inequality proves (101). Factors for the raw and decomposable stages may be inserted here because they are at least one.

Let \(q=k_{S-1}\). A nonzero term has \(q=0\) or \(q=2\). The small bases in the other stages give \(\varepsilon_R^{n-q}\), while the decomposable stage costs at most \(C_R\). Since \(0<\varepsilon_R\le1\), both possibilities are bounded by \(C_R\varepsilon_R^n\) after absorbing the single factor \(\varepsilon_R^{-2}\) into \(C_R\). This absorption introduces no dependence on \(n\).

For fixed sizes, the number of ordered partitions is \(n!/\prod_j k_j!\). Summed over all admissible sizes this is at most \(S^n\), the number of all assignments of the \(n\) indices to \(S\) stages. Combining this count, (101), and the preceding uniform subset average proves \[\mathbb E_{\Delta_n}|\mathop{\mathrm{Pf}}(d\alpha)| \le C_R\varepsilon_R^n S^{3n/2}\prod_{i=1}^n l_i.\] Substitution of (97) proves (85), and (96) supplies the assertion about \(S\). This completes the proof of Theorem 18.

Proof of the trace-support theorem

We now sum the curvature estimate against the cyclic tuple weights. The order of the two limiting choices is essential: the spatial radius is fixed first, uniformly in the conjugacy class, and only then does the simplex dimension tend to infinity.

Proof of Theorem 1. By Proposition 3, it suffices to treat a finitely generated group with a fixed word metric and an idempotent \(p\) having an exponential moment. The zero idempotent is immediate, so assume \(M_p>0\), with \(M_p\) as in Lemma 4.

For the construction of Sections 5 and 6, write \[Q_R=C S^{3/2}R^{-1/10},\qquad S=b_*+2.\] The fixed constants here may depend on the word metric, but not on \(g\), the tuple, or \(n\). Since \(S=O(1+\log R)\), we have \(Q_R\to0\). Choose \(R\) once and for all so large that the construction applies and \[ Q_RM_p<1. \tag{102}\] This fixes the final spatial radius \(r=6^{J_*-1}R\), the scale ladder, and all constants denoted \(C_R\).

Let \(g\) have infinite order, or have finite order with \(\min_h d(h,gh)>4r\). Proposition 17 provides a compatible connection family through every sufficiently large even dimension \(n\). Combining Proposition 6, Theorem 18, and (5), we obtain \[\begin{align*} |\tau_{[g]}(p)| &\le \sum_{h\in\mathcal T_n(g)}|v_{n+1}(h)| \mathbb E_{\Delta_n}|\mathop{\mathrm{Pf}}(d\alpha_h)|\\ &\le C_RQ_R^n\sum_{h\in\mathcal T_n(g)}|v_{n+1}(h)| \prod_{i=1}^n l_i\\ &\le C_RM_p\,(Q_RM_p)^n. \tag{103}\end{align*}\] In the last step we used \(l_0\ge1\) to insert the missing closing-edge factor. The constant \(C_R\) is independent of \(n\) and \(g\). Letting \(n\) tend to infinity through even integers in (103) gives \(\tau_{[g]}(p)=0\).

It remains to identify the possible surviving classes. Every infinite-order class has just been eliminated. If a finite-order class does not satisfy the displacement condition, some conjugate of its representative has word length at most \(4r\), because \[d(h,gh)=|h^{-1}gh|.\] Thus each surviving class meets the finite ball \(B(1,4r)\). There are only finitely many such classes. This proves the result for \(p\). Proposition 3 transfers it to the original \(e\) and group. Finally, an element of \(K_0(\ell^1(G))\) is a difference of idempotent classes; taking the union of their two finite supports proves the asserted inclusion for \(\operatorname{HS}^1\). ◻

Bass, Hyman. 1976. “Euler Characteristics and Characters of Discrete Groups.” Inventiones Mathematicae 35: 155–96. https://doi.org/10.1007/BF01390137.
Berrick, A. J., Indira Chatterji, and Guido Mislin. 2004. “From Acyclic Groups to the Bass Conjecture for Amenable Groups.” Mathematische Annalen 329 (4): 597–621. https://doi.org/10.1007/s00208-004-0521-6.
Burghelea, Dan. 1985. “The Cyclic Homology of the Group Rings.” Commentarii Mathematici Helvetici 60 (3): 354–65. https://doi.org/10.1007/BF02567420.
Chern, Shiing-Shen, and James Simons. 1974. “Characteristic Forms and Geometric Invariants.” Annals of Mathematics, 2nd series, vol. 99 (1): 48–69. https://doi.org/10.2307/1971013.
Eckmann, Beno. 1986. “Cyclic Homology of Groups and the Bass Conjecture.” Commentarii Mathematici Helvetici 61: 193–202. https://doi.org/10.1007/BF02621911.
Emmanouil, Ioannis. 1998. “On a Class of Groups Satisfying Bass’ Conjecture.” Inventiones Mathematicae 132 (2): 307–30. https://doi.org/10.1007/s002220050225.
Hattori, Akira. 1965. “Rank Element of a Projective Module.” Nagoya Mathematical Journal 25: 113–20. https://doi.org/10.1017/S002776300001148X.
Hoeffding, Wassily. 1963. “Probability Inequalities for Sums of Bounded Random Variables.” Journal of the American Statistical Association 58 (301): 13–30. https://doi.org/10.1080/01621459.1963.10500830.
Ji, Ronghui, Crichton Ogle, and Bobby Ramsey. 2010. “Relatively Hyperbolic Groups, Rapid Decay Algebras and a Generalization of the Bass Conjecture.” Journal of Noncommutative Geometry 4 (1): 83–124. https://doi.org/10.4171/JNCG/50.
Ji, Ronghui, Crichton Ogle, and Bobby W. Ramsey. 2014. “On the Hochschild and Cyclic (Co)homology of Rapid Decay Group Algebras.” Journal of Noncommutative Geometry 8 (1): 45–59. https://doi.org/10.4171/JNCG/148.
Lafforgue, Vincent. 2002. “\(K\)-Théorie Bivariante Pour Les Algèbres de Banach Et Conjecture de Baum–Connes.” Inventiones Mathematicae 149 (1): 1–95. https://doi.org/10.1007/s002220200213.
OpenAI. 2026. The Bass trace conjecture and the characteristic-zero Kaplansky idempotent conjecture. OpenAI Math Release preprint OAI:The-Bass-trace-conjecture-for-complex-group-rings-September-24-2026.
Stallings, John. 1965. “Centerless Groups—an Algebraic Formulation of Gottlieb’s Theorem.” Topology 4 (2): 129–34. https://doi.org/10.1016/0040-9383(65)90060-1.
LEVEL 1 COMPLETE!
You read 13,476 words and 1,238 formulas. Your math teacher would be proud.
Converted from the LaTeX source. Something look off? The original PDF is the real thing.

Cool Links: openai/math   Lean   Mathlib   arXiv   the real Coolmath Games