We prove sharp contraction of the information carried by a binary channel under independent symmetric noise on a uniform discrete cube. At fixed initial information, a noisy coordinate retains the most information. The Boolean specialization resolves the Courtade–Kumar conjecture and gives an output-entropy refinement. We also establish a stronger mean-dependent entropy-production bound. The proof combines an explicit three-point optimizer for the local joining problem, two entropy capacities, and a common-output thinning inequality, followed by dimension induction and integration along the noise semigroup.
A one-bit summary of independent fair bits can be chosen before those bits pass through independent binary symmetric channels. Which summary retains the most information about the noisy observation? Kumar and Courtade formulated the conjecture that a single input coordinate is optimal at every noise level and in every dimension (Kumar and Courtade 2013; Courtade and Kumar 2014). We prove a sharp comparison for arbitrary randomized binary summaries, together with a mean-dependent bound on their instantaneous information loss.
Let \(\Omega_n=\{-1,1\}^n\) carry uniform measure. All information is measured with natural logarithms. For \(r\in[-1,1]\), define \[L=\log2,\qquad
H(r)=-\frac{1+r}{2}\log\frac{1+r}{2}
-\frac{1-r}{2}\log\frac{1-r}{2},\qquad \psi(r)=L-H(r),\] with \(0\log0=0\). Thus \(H(r)\) is the entropy of a sign of mean \(r\). For a soft binary channel \(u:\Omega_n\to[-1,1]\), define \[I(u)=H(\mathbb Eu)-\mathbb EH(u)=\mathbb E\psi(u)-\psi(\mathbb Eu).\] Let \(X\) be uniform on \(\Omega_n\). If an auxiliary sign \(U\) has conditional mean \(\mathbb E[U\mid X]=u(X)\), then \(I(u)=I(U;X)\). For \(-1\le\rho\le1\), let \(Y_i=X_iZ_i\), where the signs \(Z_i\) are independent of \((X,U)\) and each other, with \(\mathbb EZ_i=\rho\). The noise operator is \(T_\rho u(y)=\mathbb E[u(X)\mid Y=y]\). Thus \(\mathbb E[U\mid Y]=T_\rho u(Y)\) and \(I(T_\rho u)=I(U;Y)\).
Theorem 1 (Sharp soft-channel contraction). For every finite \(n\ge0\), every \(u:\Omega_n\to[-1,1]\), and \(-1\le\rho\le1\), \[
I(T_\rho u)
\le\psi\!\left(|\rho|\,\psi^{-1}(I(u))\right).
\tag{1}\] The inverse is the nonnegative increasing branch on \([0,L]\). For \(n\ge1\), every soft coordinate \(u(x)=a x_i\), \(|a|\le1\), attains equality.
The parameter \(\psi^{-1}(I(u))\) is the amplitude of a fair soft coordinate with the same information as \(u\). The theorem says that this effective amplitude contracts at least as fast as a coordinate. Sharpness concerns all soft channels at a given initial information; it does not assert that every biased Boolean source attains the bound.
Corollary 2 (The Boolean bound with output bias). For every finite \(n\ge0\), \(\rho\in[-1,1]\), and \(f:\Omega_n\to\{-1,1\}\), put \(m=\mathbb Ef\) and \(r_m=\psi^{-1}(H(m))\). Then \[
I(f(X);Y)\le\psi(|\rho|r_m)\le\psi(|\rho|).
\tag{2}\] For \(n\ge1\), signed coordinate functions attain the bound at \(m=0\). If \(0<|m|<1\) and \(0<|\rho|\le1\), the refined upper curve is strictly below the universal curve \(\psi(|\rho|)\).
For noise crossover probability \(\varepsilon\in[0,1/2]\), one has \(\rho=1-2\varepsilon\). Writing \(h_2(p)=-p\log_2p-(1-p)\log_2(1-p)\), the universal bound in bits is \[I_2(f(X);Y)\le\frac{\psi(\rho)}{L}=1-h_2(\varepsilon).\] This is the Courtade–Kumar conjecture. The refinement uses the actual entropy of the Boolean output, while the input distribution remains uniform.
Prior work and related results
Ordentlich, Shayevitz, and Weinstein (Ordentlich et al. 2016) obtained asymptotically sharp high-noise bounds for balanced Boolean functions. Samorodnitsky (Samorodnitsky 2016) proved the conjecture for all Boolean functions in a dimension-independent high-noise range. Yu’s work on \(\Phi\)-stability and local optimality (Yu 2023, 2026) enlarged the proved range for balanced functions; the latter gives \(0\le\rho\le0.914\) by a computer-assisted argument. Javanmard and Woodruff (Javanmard and Woodruff 2026) extended Courtade and Kumar’s coordinatewise sum bound (Courtade and Kumar 2014, Theorem 1) from balanced functions to arbitrary output bias and refined high-noise entropy estimates.
Several nearby formulations clarify the role of the full noisy vector. Pichler, Piantanida, and Matz (Pichler et al. 2018) proved coordinate optimality when the observation is also reduced to one Boolean output. At fixed output mean, Li and Médard (Li and Médard 2021) identified lexicographic optimizers at sufficiently low noise and related high-noise optimizers to first-level Fourier weight, in dimension-dependent ranges. Barnes and Özgür (Barnes and Özgür 2020) proved the equivalence of the balanced conjecture and a symmetrized Li–Médard moment inequality. Kindler, O’Donnell, and Witmer (Kindler et al. 2016) proved fixed-mean Gaussian and spherical analogues. These fixed-mean questions are distinct from attainability of the refined upper curve in Corollary 2.
Chen, Gohari, and Nair (Chen et al. 2025) developed the differential approach that we use: prove a local entropy-production inequality, lift it by induction on the cube dimension, and integrate along the noise semigroup. Their approach builds on the auxiliary receiver method of Gohari and Nair (Gohari and Nair 2022). The August 2026 revision of Chen–Gohari–Nair disproves one proposed auxiliary inequality (Chen et al. 2025, Appendix B); the local comparisons below are proved directly. Their symmetric optimizer is a predecessor of our arbitrary-mean joining problem, as explained in Section 3.
Recent preprints announce proofs of the Courtade–Kumar conjecture by several routes. Chen, Gohari, Javanmard, Lin, Mirrokni, Nair, and Woodruff (Chen et al. 2026) use a computer-assisted Bellman inequality; Ky and Tran (Ky and Tran 2026) develop entropy-production and spectral estimates with computer-assisted verification; Mahdavifar and Beirami (Mahdavifar and Beirami 2026) use entropy flow, Fourier analysis, and bounds for pivotal sets; and Kramer and Saglam (Kramer and Saglam 2026) prove a finite-noise bound that retains the Boolean output bias. The overlap with the first work extends beyond the Boolean conjecture: its Lemma 9.2 applies to arbitrary binary channels, and retaining the initial information when integrating its production bound gives Theorem 1 and Corollary 2. Our proof solves the joining problem with a common entropy budget and uses two capacities and common-output thinning. It also proves the additional mean-dependent production estimate in Theorem 3; we compare the two production profiles in Section 8.1.
From information loss to a local joining inequality
For an interior soft field \(g:\Omega_n\to(-1,1)\), put \[
m=\mathbb Eg,\qquad e=\mathbb EH(g),\qquad I(g)=H(m)-e,
\qquad D(g)=\mathbb E[V(g)Ng],
\tag{3}\] where \(V(r)=\operatorname{atanh}r\) and \[N=\frac12\sum_{i=1}^n(\mathrm{Id}-\mathrm{flip}_i),
\qquad P_t=T_{e^{-t}}=e^{-tN}.\] Here \(\mathrm{flip}_i\) flips coordinate \(i\). Mean conservation and finite-dimensional differentiation give \[
\frac{d}{dt}I(P_tg)=-D(P_tg).
\tag{4}\] The production of a soft coordinate \(g(x)=r x_i\) is \(rV(r)\). We therefore define its rate by \(F(\psi(r))=rV(r)\) for \(0\le r<1\) and seek the inequality \(D(g)\ge F(I(g))\). Once this is proved, \(r_t=\psi^{-1}(I(P_tg))\) satisfies \(r_t'\le-r_t\) whenever \(I(P_tg)>0\). Integrating gives Theorem 1; Section 8 treats zero information and boundary-valued inputs.
The local-to-global program is due to Chen–Gohari–Nair (Chen et al. 2025, Definition 1 and Theorem 2). To see the required local comparison, split the cube along one coordinate into fields \(g_+,g_-\) on \(\Omega_{n-1}\). Their productions satisfy the exact identity \[
D(g)=\frac{D(g_+)+D(g_-)}2+\mathbb Ec(g_+,g_-),\qquad
c(a,b)=\frac{a-b}{4}\{V(a)-V(b)\},
\tag{5}\] where the last expectation is over the remaining coordinates. Write \(m_\pm=\mathbb Eg_\pm\) and \(I_\pm=I(g_\pm)\). For a profile with \(\Phi_m(0)=0\), the required induction step is that the joining edges pay its increase: \[\mathbb Ec(g_+,g_-)
\ge\Phi_m(I(g))-
\frac{\Phi_{m_+}(I_+)+\Phi_{m_-}(I_-)}2.\] The information-only choice \(\Phi_m(I)=F(I)\) does not satisfy the corresponding local comparison for every coupled pair law (Chen et al. 2025, Remark 5). The proof must retain the mean as well as the information.
Two profiles and the geometry of the joining cost
The linear estimate \(D(g)\ge2I(g)\) follows from the classical cube logarithmic Sobolev inequality and its entropy-dissipation consequence, within the hypercontractive framework of Bonami, Gross, and Beckner (Bonami 1970; Gross 1975; Beckner 1975); see also Bobkov–Tetali (Bobkov and Tetali 2006, Example 3.7) for the modified logarithmic Sobolev formulation. Mrs. Gerber’s lemma (Wyner and Ziv 1973) contains the same scalar one-bit entropy curve, but its direct vector form has dimension-normalized arguments. The nonlinear surplus above \(2I\) is the part that needs a new local comparison.
Since \(F'(0)=2\), define \[S(I)=F(I)-2I\quad(0\le I<L),\qquad S(I)=0\quad(I<0).\] For \(|m|<1\) and \(0\le I<H(m)\), set \[
\mathcal B_m(I)=2I+(1-m^2)S\!\left(
L-\frac{H(m)-I}{1-m^2}\right).
\tag{6}\] The factor \(1-m^2\) is useful because the average of the two child variance factors at means \(m\pm\delta\) is \(1-m^2-\delta^2\). Convexity therefore pools their surplus terms in the same form. Section 5 explains the two-capacity allocation behind this profile and proves its local joining inequality. The same two capacities occur in the Boolean finite-noise bound of Kramer and Saglam (Kramer and Saglam 2026); that section gives the comparison.
To prove the local joining inequality for \(\mathcal B\), regard \((g_+(X'),g_-(X'))\) as a coupled pair of random variables, where \(X'\) is uniform on the remaining cube. We minimize the joining cost over all coupled pair laws with the same child means and average conditional entropy, extending \(c\) by zero at the equal corners and by \(+\infty\) at every other boundary pair. The exact optimizer has one interior pair \((a_c,b_c)\) and possibly the two equal-output corners \((1,1)\) and \((-1,-1)\). Section 3 proves this description by assigning a nonnegative price to entropy. Section 4 proves the one-bit inequality that pays the interior pair. Adding either common output reduces the payment required by \(\mathcal B\) at least in proportion to the retained mass. Composing these two operations pays the full optimizer. Throughout this minimization, the original face entropies remain in the child profiles; only their average enters the cost problem.
At mean zero, \(\mathcal B_0=F\), but at other means neither profile can be discarded. We prove local joining for their maximum, \[
\widehat{\mathcal B}_m(I)
=\max\{\mathcal B_m(I),F(I)\}.
\tag{7}\] Two comparisons make this possible. First, \[
I\ge L/2\quad\Longrightarrow\quad
\mathcal B_m(I)\ge F(I).
\tag{8}\] Thus only parents below half information can require the \(F\) branch. At such a parent, convexity at the original child states reduces the unpaid amount to \(F(I)-F(j)\), where \(j=(I_++I_-)/2\).
For the second comparison, reflect the output signs if necessary so that the parent mean is nonnegative. An optimizer with one corner then combines a core of nonnegative center with \((1,1)\). Section 6 shows that this addition reduces the increment at least in proportion to the retained mass, and so the one-bit inequality pays the mixture. When both corners occur, the core is reflected, \((r,-r)\). Holding the core and its mass fixed preserves the separation of the child means, their average conditional entropy, and the joining cost. We increase the common mean until one corner disappears, as in Figure 1. Below half information the required increment increases during this motion. The one-corner endpoint therefore pays the original state as well.
Section 7 combines these comparisons and performs dimension induction to prove the following strengthening of the scalar rate inequality.
Theorem 3 (Sharp soft production). For every finite \(n\ge0\) and every \(g:\{-1,1\}^n\to(-1,1)\), with \(m=\mathbb Eg\), \[
D(g)\ge\widehat{\mathcal B}_m(I(g))
=\max\{\mathcal B_m(I(g)),F(I(g))\}
\ge F(I(g)).
\tag{9}\]
The scalar facts needed by the local comparisons are proved in Section 2. Section 8 integrates the production bound and compares its mean-dependent part with the profile of Chen et al. (Chen et al. 2026).
The scalar rate and its curvature
A soft coordinate fixes the rate used throughout the proof. We need its convexity, two bounds on its entropy-scaled curvature, and one comparison of logarithmic slopes. These facts follow from a positive hyperbolic integral and two elementary moment comparisons.
All logarithms are natural. For a sign with mean \(r\in[-1,1]\), put \[
L=\log 2,\qquad
\psi(r)=\frac{(1+r)\log(1+r)+(1-r)\log(1-r)}2,\qquad
H(r)=L-\psi(r),
\tag{10}\] with \(0\log0=0\). Thus \(H\) is binary entropy and \(\psi\) is its gap from the entropy of a fair sign. Write \(V(r)=\operatorname{atanh}r\) for \(-1<r<1\). Then \[
\psi'(r)=V(r),\qquad \psi''(r)=\frac1{1-r^2},\qquad H'(r)=-V(r).
\tag{11}\] The restriction of \(\psi\) to \([0,1)\) increases bijectively onto \([0,L)\). Define the production rate and its surplus over the linear rate by \[
F(\psi(r))=rV(r)\qquad(0\le r<1),
\tag{12}\]\[
S(I)=
\begin{cases}
F(I)-2I,&0\le I<L,\\
0,&I<0.
\end{cases}
\tag{13}\]
Lemma 4 (Elementary entropy bounds). For \(|r|\le1\), \[
\frac{r^2}{2}\le\psi(r)\le Lr^2,\qquad H(r)\ge L(1-r^2).
\tag{14}\] The quotient \(\psi(r)/r^2\) is strictly increasing for \(0<r<1\). Moreover, \[
\frac23<L<\frac{25}{36}<\frac7{10},\qquad
\frac12<M:=\frac{4L^2}{3}<\frac23.
\tag{15}\]
Proof. Integrating the geometric series for \(\psi''\) twice gives \[
\psi(r)=\sum_{n\ge1}\frac{r^{2n}}{2n(2n-1)}.
\tag{16}\] The first term gives the lower bound. The coefficients are positive and sum to \(\psi(1)=L\), so \(r^{2n}\le r^2\) gives the upper bound. The same series proves the quotient’s strict increase. Subtracting from \(L\) gives the bound for \(H\). Finally, \(L=2\operatorname{atanh}(1/3)\) implies \[\frac23<L
=\frac23+2\sum_{n\ge1}\frac{3^{-(2n+1)}}{2n+1}
<\frac23+\frac23\sum_{n\ge1}3^{-(2n+1)}
=\frac{25}{36}.\] Hence \(16/27<M<625/972<2/3\), as required. ◻
Lemma 5 (Convexity of the rate and surplus). The function \(F\) is smooth on \([0,L)\), with right derivatives at zero, and \[
F(0)=0,\qquad F'(0)=2,\qquad F''(0)=\frac43,\qquad
F''(I)\ge\frac43,\qquad F'''(I)\ge0.
\tag{17}\] The zero-extended surplus \(S\) is nonnegative, nondecreasing, convex, and \(C^1\) on \((-\infty,L)\), with \(S(0)=S'(0)=0\). It is not \(C^2\) at zero.
Proof. Put \(r=\tanh v\) and \(I=\psi(r)\). Direct differentiation gives \[I=v\tanh v-\log\cosh v,\qquad
\frac{dI}{dv}=v\operatorname{sech}^2v,\qquad
F'(I)=1+\frac{\sinh(2v)}{2v}.\] For \(v>0\), another differentiation yields \[\begin{align*}
F''(\psi(r))
&=\frac{(1+r^2)V(r)-r}{(1-r^2)^2V(r)^3}\tag{18}\\
&=\cosh^2v\,A(v),\qquad
A(v):=2\int_0^1(1-s^2)\cosh(2vs)\,ds.
\tag{19}\end{align*}\] Indeed, integration by parts evaluates the integral as \([2v\cosh(2v)-\sinh(2v)]/(2v^3)\). The integral gives \(A(v)\ge A(0)=4/3\) and \(A'(v)>0\) for \(v>0\). Thus \(F''\ge4/3\) and \(F'''>0\) for \(I>0\). Both \(A(v)\) and \(I\) are analytic in \(v^2\), and \(I=v^2/2+O(v^4)\), so the inverse function theorem gives smooth extension to \(I=0\). Continuity gives \(F'''(0)\ge0\). The identities \(S(0)=S'(0)=0\) and \(S''=F''>0\) for \(I>0\) prove the assertions about its zero extension. Its second derivative jumps from \(0\) to \(4/3\) at zero. ◻
Lemma 6 (Growth of the curvature). The function \(I\mapsto e^{-2I}F''(I)\) is strictly increasing on \([0,L)\). For \(I>0\) it satisfies \[\frac{F'''(I)}{F''(I)}>2.\]
Proof. Using (19), for \(v>0\), \[\frac{d}{dI}\log F''(I)
=\frac{\sinh(2v)}v
+\frac{\cosh^2v}{v}\frac{A'(v)}{A(v)}>2.\] Here \(A'(v)\ge0\) and \(\sinh(2v)>2v\). Continuity at zero extends the asserted strict increase to \([0,L)\). ◻
Lemma 7 (Scaled entropy curvature). For \(0<h\le L\), put \[
P(h)=h^2F''(L-h).
\tag{20}\] Then \[
\frac12\le P(h)\le M=\frac{4L^2}{3},\qquad P(L)=M.
\tag{21}\]
Proof. Write \(h=H(r)\), \(t=r^2\), and \(x=V(r)/r\), with \(x=1\) at \(r=0\). Elementary integration gives \[
x=\int_0^1\frac{du}{1-tu^2},\qquad
\frac{H(r)}{1-t}=\int_0^1\frac{du}{(1+u)(1-tu^2)}.
\tag{22}\] For the second identity, use the partial fraction identity \[\frac1{(1+u)(1-tu^2)}
=\frac1{1-t}\left(\frac1{1+u}+\frac{t(u-1)}{1-tu^2}\right)\] and \(H(r)=L-\frac12\log(1-r^2)-rV(r)\). Let \(\mu_t\) have density \([x(1-tu^2)]^{-1}\) on \([0,1]\), and set \[b_t=\mathbb E_{\mu_t}\frac1{1+u},\qquad n_t=\mathbb E_{\mu_t}u^2.\] Then \(b_t=H(r)/[(1-t)x]\) and \(n_t=(V(r)-r)/(r^2V(r))\), with continuous values at zero. Substitution into (18) gives \[P(H(r))=b_t^2(1+n_t).\] Jensen’s inequality and Cauchy–Schwarz imply \[b_t\ge\frac1{1+\mathbb E_{\mu_t}u}\ge\frac1{1+\sqrt{n_t}},
\qquad
P(H(r))\ge\frac{1+n_t}{(1+\sqrt{n_t})^2}\ge\frac12.\]
For the upper bound, a density proportional to \(w(u)/(1-tu^2)\), with \(w\) independent of \(t\), satisfies \[
\frac{d}{dt}\mathbb E_t f(u)
=\operatorname{Cov}_t\left(f(u),\frac{u^2}{1-tu^2}\right).
\tag{23}\] Oppositely monotone functions have nonpositive covariance, since for independent copies \(U,U'\), \[2\operatorname{Cov}(f(U),g(U))
=\mathbb E[(f(U)-f(U'))(g(U)-g(U'))].\] Applying this with \(w=1\) shows \(1/2\le b_t\le b_0=L\). The same identity shows that \(n_t\) increases, so we must control its growth together with the decrease of \(b_t\). Compare \(b_t-1/2\) with \(1-n_t\) by reweighting \(\mu_t\) by \(1-u^2\) to obtain \(\nu_t\). The identity \[\beta_t:=\frac{b_t-1/2}{1-n_t}
=\frac12\mathbb E_{\nu_t}(1+u)^{-2},\qquad
d\nu_t(u)\ \propto\ \frac{1-u^2}{1-tu^2}\,du\] and (23) now give \(0<\beta_t\le\beta_0=\frac32(L-\frac12)\). Consequently \[1+n_t\le2-\frac{b_t-1/2}{\beta_0},\] and hence \[P(H(r))\le f(b_t),\qquad
f(b)=b^2\left[2-\frac{b-1/2}{\beta_0}\right].\] For \(1/2\le b\le L\), \[f'(b)=\frac b{\beta_0}(4\beta_0+1-3b)
\ge\frac b{\beta_0}(3L-2)>0.\] Thus \(P(H(r))\le f(L)=4L^2/3\). At \(r=0\), the uniform law has \(b_0=L\) and \(n_0=1/3\), which also gives \(P(L)=4L^2/3\). ◻
A joining cost with one entropy budget
The energy decomposition (5) leaves the cost of joining the two faces of a coordinate split. We minimize this cost at fixed child means and average conditional entropy. Allowing arbitrary coupled pair laws is a valid relaxation because only a lower bound on the actual joining cost is needed; no independence of the faces is assumed. The symmetric fixed-entropy problem of Chen–Gohari–Nair (Chen et al. 2025, Theorem 4 and Section IV-C) has an optimizer supported on a reflected interior pair and the two equal-output corners. Chen et al. (Chen et al. 2026, sec. 3.1, Theorem 3.10) characterize related three-point optimizers when both means and both entropy moments are prescribed, under explicit contact and mass conditions. Here we prescribe the two means but constrain only the average entropy. Pricing this common entropy budget gives the optimizer throughout the domain, including the transition from two corners to one.
For interior \(x,y\), use the joining cost \[
c(x,y)=\frac{x-y}{4}\,[V(x)-V(y)].
\tag{24}\] It is symmetric and nonnegative. Extend it by zero at the equal corners \((1,1),(-1,-1)\) and by \(+\infty\) at every other boundary pair. This is a lower semicontinuous, jointly convex function. Indeed, \[c(x,y)=\frac18\{J(1+x,1+y)+J(1-x,1-y)\},
\qquad J(s,t)=(s-t)\log(s/t),\] and \(J\) is the sum of two relative-entropy perspectives.
Fix \(-1<b<a<1\), and write \[m=\frac{a+b}{2},\qquad \delta=\frac{a-b}{2},\qquad
e_{\max}=\frac{H(a)+H(b)}2.\] Jensen’s inequality makes \(e_{\max}\) the largest possible average entropy; the deterministic pair attains it. We call this boundary the capacity state of the joining problem. For \(0<e\le e_{\max}\), define \[
\mathcal K(m,\delta,e)=
\inf\left\{\mathbb Ec(A,B):
\mathbb EA=a,\ \mathbb EB=b,\
\frac{\mathbb EH(A)+\mathbb EH(B)}2\le e\right\}.
\tag{25}\] The infimum is over joint laws on \([-1,1]^2\). Only the average entropy is constrained.
Constructing and certifying the priced core
We first construct candidate laws of the form \[
\lambda\delta_{(x,y)}
+\pi_+\delta_{(1,1)}+\pi_-\delta_{(-1,-1)},\qquad -1<y<x<1.
\tag{26}\] The prescribed mean difference forces \(\lambda(x-y)=2\delta\); the mean sum and total mass then determine the two corner masses. We will minimize a priced objective within this family, and then prove optimality against all coupled laws by an affine lower bound.
For an ordered interior pair \(x>y\), set \[z=\log\frac{1+x}{1+y},\quad
w=\log\frac{1-y}{1-x},\quad
u=(e^z-1)^{-1},\quad v=(e^w-1)^{-1},\quad n=1+u+v.\] Then \[
x=\frac{1+u-v}{n},\qquad
y=\frac{u-v-1}{n},\qquad x-y=\frac2n,\qquad
\frac{c(x,y)}{x-y}=\frac{z+w}{8}.
\tag{27}\] In these coordinates, the masses required by the two means are \[
\lambda=\delta n,\qquad
\pi_+=\frac{1+b-2\delta u}{2},\qquad
\pi_-=\frac{1-a-2\delta v}{2}.
\tag{28}\] They sum to one, and the corner masses are nonnegative exactly when \[
z\ge z_0:=\log\frac{1+a}{1+b},\qquad
w\ge w_0:=\log\frac{1-b}{1-a}.
\tag{29}\] For a price \(p\ge0\), define \[\mathcal H(z,w)=\frac{H(x)+H(y)}{x-y},\qquad
F_p(z,w)=\frac{z+w}{8}+p\mathcal H(z,w).\] The corners contribute neither cost nor entropy. Hence the priced objective of the candidate law (26) is \[
\mathbb E\{c(A,B)+p[H(A)+H(B)]\}=2\delta F_p(z,w).
\tag{30}\] This explains the normalization by \(x-y\): the prescribed mean difference absorbs the core mass. The remaining task is to minimize \(F_p\) subject to (29) and certify that its value also bounds every other pair law.
The normalized entropy has the positive series \[
\begin{split}
\mathcal H(z,w)
&=\frac{q(z)+q(w)}2+
\sum_{i,j\ge1}\frac{e^{-iz-jw}}{\max(i,j)},\\
q(t)&=\frac{t}{e^t-1}-\log(1-e^{-t}).
\end{split}
\tag{31}\] Here is a direct verification. The normalized entropy of \(x\) is \([n\log n-(u+1)\log(u+1)-v\log v]/2\), and that of \(y\) is its transpose. Put \(\xi=e^{-z}\) and \(\zeta=e^{-w}\). After subtracting \([q(z)+q(w)]/2\), the sum becomes \[n\log(1-\xi\zeta)-v\log(1-\xi)-u\log(1-\zeta).\] The coefficient of \(\xi^i\zeta^j\) is \(1/i+1/j-1/\min(i,j)=1/\max(i,j)\), which proves (31). The series and its derivatives converge uniformly on compact subsets of the positive quadrant.
The function \(\mathcal H\) is symmetric and strictly convex. In fact, with \(v_t=(e^t-1)^{-1}\), \[q''(t)=v_t(v_t+1)\{t(2v_t+1)-1\}>0,\] because \(t(2v_t+1)=t\coth(t/2)>2\). Each term in the double series is convex; its exponential direction vectors span the plane.
Lemma 8 (The priced core). For every \(p\ge0\), the minimum of \[\mathbb E\{c(A,B)+p[H(A)+H(B)]\},
\qquad \mathbb EA=a,\quad \mathbb EB=b,\] has a unique optimizer. It is one ordered interior pair \((x,y)\) and possible equal corners, as in (26). Its core is given by the unique minimizer of \(F_p\) subject to (29), using (27); its masses are (28). At \(p=0\), the optimizer is the deterministic pair \((a,b)\).
Proof. For \(p>0\), strict convexity and the linear term give a unique minimizer on the floored quadrant. At \(p=0\), the unique minimum is \((z_0,w_0)\). We now turn this constrained minimum into an affine lower bound in the pair values, valid for every coupled law. At the selected point put \[\gamma_+=\frac{(F_p)_z}{2u(u+1)},\qquad
\gamma_-=\frac{(F_p)_w}{2v(v+1)}.\] Both are nonnegative; a derivative vanishes unless its floor is active. The function \[\mathscr Q(z,w)=F_p(z,w)+\gamma_+(1+2u)+\gamma_-(1+2v)\] has zero gradient at the selected point and is strictly convex on the whole positive quadrant. At \(p=0\), both added coefficients are positive, and \(u''=u(u+1)(2u+1)>0\), with the same identity for \(v\), so strict convexity still holds. Let \(\kappa>0\) be its minimum; positivity follows from the positive linear term in \(F_p\) and the nonnegative remaining terms. Multiplying \(\mathscr Q\ge\kappa\) by \(x-y\) gives the global support \[
\begin{split}
c(x,y)+p[H(x)+H(y)]-\kappa(x-y)
+\gamma_+(2+x+y)+\gamma_-(2-x-y)\ge0.
\end{split}
\tag{32}\] For \(x>y\) this is the convex minimum just proved. For \(x<y\), every term is nonnegative and \(-\kappa(x-y)>0\). On the interior diagonal, the entropy term is positive if \(p>0\); if \(p=0\), both corner coefficients are positive. At the equal corners, equality is allowed exactly when the corresponding \(\gamma\) vanishes. All other boundary pairs have infinite cost.
The candidate masses in (28) have means \(a,b\), and complementarity gives \(\gamma_\pm\pi_\pm=0\). Thus (26) is supported at equality points of (32). Averaging the support proves optimality. Its only possible interior contact is the unique convex minimizer; the mean difference fixes its mass, and the mean sum fixes the two corners. This proves uniqueness. ◻
Proposition 9 (Exact entropy coverage). For every \(0<e\le e_{\max}\), a law of the form (26) attains (25) and has average entropy exactly \(e\).
Proof. Let \(e(p)\) be the average entropy of the unique priced optimizer. It is continuous for \(p\ge0\): on each bounded price interval the minimizers lie in a common compact subset of the floored quadrant, since \(F_p\ge(z+w)/8\); uniqueness then gives continuity. Also \(e(0)=e_{\max}\).
There are finite-cost laws with the prescribed means and arbitrarily small entropy. Choose a reflected core \((r,-r)\) with \(r\) sufficiently close to one, mass \(\delta/r\), and the two equal-corner masses needed to match \(m\). They are nonnegative because \(\delta<1-|m|\). Its entropy \((\delta/r)H(r)\) tends to zero. If such a comparison law has cost \(C_\varepsilon\) and entropy at most \(\varepsilon\), priced optimality and nonnegativity of cost imply \[2p\,e(p)\le C_\varepsilon+2p\varepsilon.\] Hence \(e(p)\to0\) as \(p\to\infty\). Continuity supplies a price with \(e(p)=e\).
For any competing law whose average entropy \(\widetilde e\) is at most \(e\), the global priced support gives \[\mathbb Ec(A,B)+2p\widetilde e
\ge \mathbb Ec(A_p,B_p)+2pe.\] Since \(p\ge0\), its cost is at least that of the priced optimizer. This proves the proposition without a signed entropy multiplier or an abstract optimizer replacement. ◻
The two canonical geometries
By reflection assume \(m\ge0\), so \(z_0\le w_0\). Symmetry and uniqueness imply \(z\le w\) at the priced minimizer: otherwise swapping the coordinates is feasible and has the same value. The case \(p=0\) is already deterministic, so assume \(p>0\) in the following strict-convexity argument. If \(w>w_0\) and \(z<w\), then \((F_p)_w=0\) while strict convexity and symmetry give \((F_p)_z<(F_p)_w\), contradicting constrained optimality. Thus either \(z=w\) or \(w=w_0\).
In the first case, the core is \((r,-r)\), with \[
k=\frac{\delta}{r},\qquad
e=kH(r),\qquad
\mathcal K=\delta V(r),\qquad 0\le m\le1-k.
\tag{33}\] At fixed \(\delta,e\), the radius \(r\) and the cost do not depend on \(m\). Uniqueness of \(r\) follows because \(H(r)/r\) strictly decreases.
In the second case, \(\pi_-=0\). With core mass \(k\), \[
\begin{split}
1-m\le k\le1,\qquad
x=1-\frac{1-a}{k},\quad y=1-\frac{1-b}{k},\\
e=\frac{k}{2}[H(x)+H(y)],\qquad
\mathcal K=k\,c(x,y).
\end{split}
\tag{34}\] Indeed \(z\le w\) means the core center is nonnegative, and \[\frac{x+y}{2}=1-\frac{1-m}{k}\ge0.\] The two geometries meet at \(k=1-m\), where the reflected core has only a positive common corner. Thus their entropy threshold is \[
e_0(m,\delta)=(1-m)H\!\left(\frac{\delta}{1-m}\right).
\tag{35}\] To check the direction of the threshold, differentiate the one-corner entropy in (34) at fixed \(a,b\): \[\frac{d}{dk}\frac{k}{2}[H(x)+H(y)]
=\frac12\left[\log\frac2{1+x}+\log\frac2{1+y}\right]>0.\] The reflected entropy \(kH(\delta/k)\) also strictly increases in \(k\), since its derivative is \(H(r)+rV(r)>0\). Consequently the reflected geometry has \(e\le e_0\), and the one-corner geometry has \(e\ge e_0\). This is a support threshold for the exact cost optimizer, separate from the analytic threshold (8) for the production profile. At capacity \(k=1\) and the law is deterministic. If \(\delta=0\), taking \(A=B\) with mean \(m\) and arbitrarily small entropy gives \(\mathcal K=0\); that case requires no canonical classification.
Mixture weights for a fixed reflected core \((r,-r)\). The horizontal segment keeps its mass \(k\) fixed, so the separation \(\delta=kr\), entropy \(e=kH(r)\), and joining cost \(\delta V(r)\) stay fixed. Moving right increases the mean \(m=\pi_+-\pi_-\) until the negative common output disappears. Theorem 20 compares the low-information demand with this endpoint. The triangle is schematic; its points represent probability weights, not channel values.
Figure 1 displays the motion within the reflected geometry. The path preserves the optimizer’s cost and its average entropy; the original law’s individual entropies are retained separately when applying the local production inequalities.
The deterministic capacity comparison
The scalar profile is calibrated at balanced one-bit channels. At a biased one-bit channel, the two output labels have unequal posterior correlations. Their separation provides the extra reserve needed by the variance-weighted profile. The inequality below also follows by differentiating at correlation one the pointwise one-bit comparison in the proof of Kramer–Saglam (Kramer and Saglam 2026, Lemma 2.4). We give a direct proof using size-biased posterior information.
Lemma 10 (A variance-refined one-bit inequality). Let \(g(X)=m+\delta X\), where \(X\) is a uniform sign and \(|m|+|\delta|<1\). Write \(\alpha=1-m^2\) and \[J=\frac{\psi(m+\delta)+\psi(m-\delta)}2-\psi(m),
\qquad D=\frac\delta2\{V(m+\delta)-V(m-\delta)\}.\] Then \(J<\alpha L\) and \[
D\ge\alpha F(J/\alpha).
\tag{36}\]
Proof. Reflection and relabeling allow \(m\ge0\), \(\delta\ge0\). The case \(\delta=0\) is immediate, so assume \(\delta>0\). Set \[w_\pm=\frac{1\pm m}{2},\qquad
r_\pm=\frac\delta{1\pm m},\qquad i_\pm=\psi(r_\pm).\] The two posterior representations of the one-bit channel give \[
J=w_+i_++w_-i_-,\qquad
D=w_+F(i_+)+w_-F(i_-).
\tag{37}\] For the first identity, introduce a sign \(U\) with conditional mean \(g(X)\). Its probabilities are \(w_\pm\), and the conditional means of \(X\) given \(U=\pm1\) are \(r_+\) and \(-r_-\). Evaluating \(I(U;X)\) in either order gives the identity. For the second, use \(w_\pm r_\pm=\delta/2\) and \[V(m+\delta)-V(m-\delta)=V(r_+)+V(r_-).\]
To extract the variance refinement from these identities, use the secant slope \(\varphi(s)=F(s)/s\), with \(\varphi(0)=2\). It is increasing by convexity of \(F\) and \(F(0)=0\). It is also convex: the scalar fact \(F'''\ge0\) gives, for \(s>0\), \[\varphi''(s)
=\frac{s^2F''(s)-2sF'(s)+2F(s)}{s^3}
=\frac1{s^3}\int_0^s t^2F'''(t)\,dt\ge0.\] Since \(\delta>0\), both \(i_\pm\) and \(J\) are positive. The size-biased weights \(w_\pm i_\pm/J\) therefore sum to one. Their mean information is \[s_*:=\frac{w_+i_+^2+w_-i_-^2}{J},\qquad
0<s_*\le\max\{i_+,i_-\}<L.\] Jensen’s inequality now gives \[
\frac DJ
=\frac{w_+i_+}{J}\varphi(i_+)
+\frac{w_-i_-}{J}\varphi(i_-)
\ge\varphi(s_*).
\tag{38}\] It remains to show \(s_*\ge J/\alpha\). This will also establish \(J/\alpha<L\) before we evaluate the rate there.
The positive series for \(\psi\) shows that \(\psi(r)/r^2\) increases. Consequently \[q:=\frac{i_+}{i_-}
\le q_0:=\left(\frac{1-m}{1+m}\right)^2.\] For \(0\le q\le1\), the ratio \[R(q)=\frac{w_+q^2+w_-}{(w_+q+w_-)^2}\] decreases, since \[R'(q)=\frac{2w_+w_-(q-1)}{(w_+q+w_-)^3}\le0.\] Direct substitution at \(q_0\) therefore gives \[
\frac{w_+i_+^2+w_-i_-^2}{J^2}
\ge\frac{1+3m^2}{1-m^2}\ge\frac1\alpha.
\tag{39}\]
Thus \(J/\alpha\le s_*<L\). Combining (38) with monotonicity of \(\varphi\) yields \[\frac DJ
\ge\varphi(s_*)
\ge\varphi(J/\alpha).\] This proves both assertions. ◻
The ordinary one-bit bound follows at once: \[c(m+\delta,m-\delta)\ge\alpha F(J/\alpha)\ge F(J).\] The last inequality is convexity with \(F(0)=0\) and \(0<\alpha\le1\). We retain the variance-refined form because it pays the deterministic base of the reserve profile, while its ordinary consequence pays the information increment at a core.
A reserve that survives adding either common output
The linear rate \(2I\) leaves a remainder that has a simple behavior under mixing. Write \(S=F-2I\), as in (13), and put \[
R(v,e)=vS\!\left(L-\frac ev\right),\qquad v,e>0,
\tag{40}\] using the zero extension of \(S\) at negative arguments. The profile (6) is \[\mathcal B_m(I)=2I+R(1-m^2,H(m)-I).\] Its construction can be read as an allocation of information between two capacities. Set \(\alpha=1-m^2\) and \(c_m=H(m)-\alpha L\ge0\). Then \[
\mathcal B_m(I)=
\inf_{\substack{u+v=I\\0\le u\le c_m\\0\le v<\alpha L}}
\{2u+\alpha F(v/\alpha)\},\qquad 0\le I<H(m).
\tag{41}\] Indeed, \(F'\ge2\) makes the optimum \(u=\min\{I,c_m\}\). The capacity \(c_m\) is paid at the linear rate, while the remaining capacity \(\alpha L\) uses the exact coordinate rate. This is an identity for the profile, not a decomposition of an actual channel. The finite-noise bound of Kramer and Saglam (Kramer and Saglam 2026, Theorem 1.1) reads, in the present normalization, \[I(T_\rho f)\le\rho^2c_m+\alpha\psi(\rho)
\qquad(0\le\rho\le1)\] for a Boolean field \(f\) on the uniform cube with \(\mathbb Ef=m\). It uses the same coefficients \(\alpha\) and \(c_m\); here they define the local production profile \(\mathcal B_m\).
To determine the payment required by a join, consider a physical state \[|m|+\delta<1,\quad\delta\ge0,\quad
0<e\le\bar h:=\frac{H(m+\delta)+H(m-\delta)}2,\] and write \[\alpha=1-m^2,\quad\beta=\alpha-\delta^2,
\qquad J=H(m)-\bar h,\quad j=\bar h-e.\] Set \(a=m+\delta\) and \(b=m-\delta\). Let \(e_a,e_b>0\) satisfy \(e_a\le H(a)\), \(e_b\le H(b)\), and \((e_a+e_b)/2=e\), and put \(I_a=H(a)-e_a\), \(I_b=H(b)-e_b\). Their average information is \(j\), and their average variance capacity is \([(1-a^2)+(1-b^2)]/2=\beta\). Since \(R\) is a convex perspective, the child profiles satisfy \[
\frac{\mathcal B_a(I_a)+\mathcal B_b(I_b)}2
\ge2j+R(\beta,e).
\tag{42}\] The parent information is \(I=H(m)-e=J+j\). Subtracting the pooled bound from \(\mathcal B_m(I)=2(J+j)+R(\alpha,e)\) gives the required payment \[
W(m,\delta,e)=2J+R(\alpha,e)-R(\beta,e).
\tag{43}\] Thus \[\mathcal B_m(I)-\frac{\mathcal B_a(I_a)+\mathcal B_b(I_b)}2
\le W(m,\delta,e).\] It is enough to prove \(W\le\mathcal K\). The deterministic capacity comparison pays the interior core; the next two lemmas show that this payment survives adding either common output.
A logarithmic bound for the corner geometry
Lemma 11. For \(-1<m<1\) and \(0\le\delta<1-|m|\), put \[A=1-m,\qquad
\kappa=\frac12\log\frac1{1-\delta^2/(1+m)^2}.\] Then \[
A^2\log\frac\alpha\beta\le3\kappa.
\tag{44}\]
Proof. The case \(\delta=0\) is immediate. Otherwise set \[h=\frac{1-m}{1+m}>0,\qquad
t=\frac{\delta^2}{(1+m)^2}<\min\{1,h^2\},\qquad
f(t)=-\log(1-t).\] The ratio \(f(t/h)/f(t)\) decreases in \(t\) when \(h\ge1\) and increases when \(h\le1\). Indeed, the ratio of the corresponding positive integrands is \((1-t)/(h-t)\), whose derivative has sign \(1-h\); the same monotonicity passes to their integral ratio.
If \(h\ge1\), the ratio is at most its limit \(1/h\) at zero. Thus \(A^2f(t/h)/f(t)\le A^2/h=1-m^2\le1\). If \(0<h<1\), it suffices to bound the ratio at \(t=h^2\). The function \[G(h)=(2-h)f(h)-(2+h)\log(1+h)\] has \(G(0)=0\) and \[G'(h)=2\left[\frac h{1-h^2}-V(h)\right]>0.\] The last sign follows by bounding the integral defining \(V(h)\) by its endpoint integrand. Consequently \[\frac{f(h^2)}{f(h)}
=1-\frac{\log(1+h)}{f(h)}\ge\frac{2h}{2+h},\] and \[A^2\frac{f(h)}{f(h^2)}
\le\frac{2h(2+h)}{(1+h)^2}
=2-\frac2{(1+h)^2}\le\frac32.\] Since \(f(t)=2\kappa\) and \(f(t/h)=\log(\alpha/\beta)\), both cases give (44). ◻
The common-output contraction
Lemma 12 (Adding either common output). For every physical state \((m,\delta,e)\) and \(0<\tau\le1\), \[
W(1-\tau+\tau m,\tau\delta,\tau e)
\le\tau W(m,\delta,e).
\tag{45}\] The same inequality holds with parent mean \(-1+\tau+\tau m\), corresponding to the negative common output. There is no sign restriction on the original mean.
Proof. The positive-corner ray scales \(A=1-m\), \(\delta\), and \(e\) by the same factor. Hold \(\delta/A\) and \(e/A\) fixed, and put \(\eta=1+\delta^2/A^2\). Entropy differentiation gives \[
A\frac{dJ}{dA}-J=\kappa.
\tag{46}\] To check this identity, use \(H(m)-(1-m)V(m)=\log[2/(1+m)]\); the average of the two child logarithms is the parent logarithm plus \(\kappa\).
For \(h>0\), define the entropy intercept \[\gamma(h)=S(L-h)+hS'(L-h)
=\begin{cases}
\displaystyle\int_h^L\frac{P(u)}u\,du,&0<h<L,\\
0,&h\ge L,
\end{cases}
\qquad P(u)=u^2F''(L-u).\] By the scalar curvature bound, \(P\le M=4L^2/3\) and \(2>3M\). Thus \(\gamma\ge0\), and \[
0\le\gamma(e/\alpha)-\gamma(e/\beta)
\le M\log\frac\alpha\beta.
\tag{47}\] If an argument crosses the clipping point, the integral is simply truncated at \(L\).
Since \(\alpha=A(2-A)\) and \(\beta=A(2-\eta A)\), differentiation of the perspectives in (43) gives \[
A\frac{dW}{dA}-W
=2\kappa+A^2\{\eta\gamma(e/\beta)-\gamma(e/\alpha)\}.
\tag{48}\] For example, with \(e=Az\), one has \(R(\alpha,e)/A=(2-A)S(L-z/(2-A))\), whose derivative is \(-\gamma(e/\alpha)\). The second perspective gives the factor \(\eta\). The clipped remainder is \(C^1\), because \(S(0)=S'(0)=0\), so these identities hold across the seam; no second derivative of the zero extension is needed.
Now \(\eta\ge1\), (47), and Lemma 11 give the whole comparison: \[
A\frac{dW}{dA}-W
\ge2\kappa-MA^2\log\frac\alpha\beta
\ge(2-3M)\kappa\ge0.
\tag{49}\] Therefore \(W/A\) is nondecreasing. Comparing \(A\) and \(\tau A\) proves (45). The intermediate states remain physical by entropy concavity, and all child means stay interior for \(\tau>0\). Finally \(W\) is even in \(m\), so reflection proves the negative-corner statement. For \(\delta=0\), every displayed payment and margin is zero. ◻
From one core to the local inequality
Theorem 13 (Local joining for the linear reserve). Let \((A,B)\) have any joint law on \([-1,1]^2\) with \(e_a=\mathbb EH(A)>0\) and \(e_b=\mathbb EH(B)>0\). Put \[a=\mathbb EA,\quad b=\mathbb EB,\quad m=\frac{a+b}{2},
\quad e=\frac{e_a+e_b}{2},\] and \(I=H(m)-e\), \(I_a=H(a)-e_a\), \(I_b=H(b)-e_b\). Then \[
\mathbb Ec(A,B)\ge
\mathcal B_m(I)-\frac{\mathcal B_a(I_a)+\mathcal B_b(I_b)}2.
\tag{50}\]
Proof. An infinite cost is immediate. Otherwise positive child entropies give interior means. Swap the children if necessary so that \(\delta=(a-b)/2\ge0\); entropy concavity makes \((m,\delta,e)\) physical.
First pay the deterministic capacity state. For an interior pair \(x,y\), let \(m_0=(x+y)/2\), \(\delta_0=|x-y|/2\), \(e_0=[H(x)+H(y)]/2\), and \(J_0=H(m_0)-e_0\). Write \(\alpha_0=1-m_0^2\), \(\beta_0=\alpha_0-\delta_0^2\). The entropy lower bound gives \(e_0\ge\beta_0L\) and \[[L-e_0/\alpha_0]_+\le J_0/\alpha_0<L.\] The last strict inequality and the variance-refined one-bit bound are Lemma 10. Hence \[
\begin{aligned}
W(m_0,\delta_0,e_0)
&=2J_0+\alpha_0S(L-e_0/\alpha_0)\\
&\le2J_0+\alpha_0S(J_0/\alpha_0)
=\alpha_0F(J_0/\alpha_0)\le c(x,y).
\end{aligned}
\tag{51}\]
For \(\delta>0\), Proposition 9 supplies an exact optimizer with interior core \((x,y)\) of mass \(k>0\) and equal-corner masses \(\pi_+,\pi_-\), where \(k+\pi_++\pi_-=1\). Its cost is \(k c(x,y)\). Starting from the core, add the positive corner with retention factor \(k/(k+\pi_+)\), then the negative corner with retention factor \(k+\pi_+\). These two operations produce exactly the canonical law. Lemma 12 and (51) give \[
W(m,\delta,e)
\le(k+\pi_+)\frac{k}{k+\pi_+}W(m_0,\delta_0,e_0)
\le k c(x,y)=\mathcal K(m,\delta,e).
\tag{52}\] Both retention factors are positive; a missing corner gives a factor one. No orientation or separate case for the core is required. For \(\delta=0\), \(W=\mathcal K=0\) directly.
Finally the actual pair law has cost at least \(\mathcal K(m,\delta,e)\). Combining (52) with the original-child pooling bound (42) gives \[\mathbb Ec(A,B)+\frac{\mathcal B_a(I_a)+\mathcal B_b(I_b)}2
\ge2(J+j)+R(\alpha,e)=\mathcal B_m(I).\] All child profiles use their original individual entropies; only the joining cost was minimized with their common average budget. ◻
The pooling step has a useful exact interpretation. If \(e<\beta L\), its minimum over the entropy split is attained at \(e_a/(1-a^2)=e_b/(1-b^2)=e/\beta\); these entropies are physical because \(H(a)\ge(1-a^2)L\), and similarly for \(b\). If \(e\ge\beta L\), allocate both child entropies in the intervals \([(1-a^2)L,H(a)]\) and \([(1-b^2)L,H(b)]\) with the prescribed average. Both remainders vanish. Thus the perspective in (42) is the exact common-budget value of the child reserve, rather than an additional approximation.
Paying the information increment by thinning
The canonical optimizer with one positive common corner has a useful feature: its cost is linear in the mass of its interior core. We will show that the information-rate increment is at most the core’s increment multiplied by that mass. At full mass the increment is paid by the one-bit posterior comparison. This will settle the whole high-entropy branch of the canonical cost problem.
The scalar inputs already proved here are \[
F(0)=0,\qquad F'(I)\ge2,\qquad
0<F''(I)\le\frac{M}{(L-I)^2},\qquad
M=\frac{4L^2}{3},\qquad 2>3M.
\tag{53}\] Indeed, the slope follows from Lemma 5, the curvature bound from Lemma 7, and the strict margin from Lemma 4. The proof isolates the roles of these two bounds: the slope supplies a positive term, while the curvature controls the possible loss over an entropy interval.
The common-output path
Fix an interior ordered pair \(a_c>b_c\) with nonnegative center, and write \[x_c=\frac{a_c+b_c}{2}\ge0,\qquad
d_c=\frac{a_c-b_c}{2}>0,\qquad
h_c=\frac{H(a_c)+H(b_c)}2,\qquad
J_c=H(x_c)-h_c.\] For \(0<\lambda\le1\), put mass \(\lambda\) on this core and mass \(1-\lambda\) on \((1,1)\). The second atom is a common output: its coordinates agree, and it contributes neither entropy nor cost. The mixture has means \(x+d\) and \(x-d\), where \[
x=1-\lambda(1-x_c),\qquad d=\lambda d_c,
\qquad e=\lambda h_c.
\tag{54}\] For this path define \[
\bar h=\frac{H(x+d)+H(x-d)}2,\qquad
\nu=\bar h-e,\qquad J=H(x)-\bar h,
\qquad I=\nu+J=H(x)-e.
\tag{55}\] The average child information is \(\nu\), and \(J\) is the entropy gained by averaging the two child means. Since \(0<d_c<1-x_c\le1\), the mixture has \(x\ge0\) and \(0<d<1-x\). Entropy concavity gives \(\nu\ge0\); strict concavity gives \(J>0\); and \(e>0\) gives \(I<L\). Thus every rate argument lies in \([0,L)\).
These quantities have a channel interpretation. Let \(X\) be a fair sign and let \(B\in\{0,1\}\) be an independent flag with \(\Pr(B=1)=\lambda\). Given \(X,B\), draw a sign \(U\) with conditional mean \[\mathbb E[U\mid X,B]=(1-B)+B(x_c+d_cX).\] Thus \(B=0\) selects the common output and \(B=1\) selects the core. The three relevant entropies are \(H(U)=H(x)\), \(H(U\mid X)=\bar h\), and \(H(U\mid X,B)=e\). The chain rule therefore reads \[
I=I(U;X,B),\qquad \nu=I(U;B\mid X),\qquad
J=I(U;X),\qquad I=\nu+J.
\tag{56}\] In particular, \(\nu\) is the information about the flag after the input has been revealed.
Theorem 14 (Common-output payment). For the core and mass in (54), \[
F(\nu+J)-F(\nu)\le\lambda F(J_c).
\tag{57}\]
We prove this by showing that \([F(\nu+J)-F(\nu)]/\lambda\) increases with \(\lambda\). This normalization matches the mixture’s cost \(\lambda c(a_c,b_c)\). We first compute its derivative and identify the entropy-interval estimate that makes it nonnegative.
The normalized derivative
For the state (54), set \[
\begin{gathered}
p=1+x,\qquad \sigma=1-x,\qquad b=\log(1+x),\\
q=e+\psi(x)=L-I,\qquad s=e+L-\bar h=L-\nu,\\
\kappa=-\frac12\log\left(1-\frac{d^2}{p^2}\right)>0.
\end{gathered}
\tag{58}\] In particular \(s-q=J\) and \(0<q<s\le L\). Write \[\Delta(\lambda)=F(I)-F(\nu),\qquad
A=\log\frac2{1+x}=L-b.\] For \(h_{\mathrm{bin}}(t)=H(1-2t)\), the entropy identity is \[t h_{\mathrm{bin}}'(t)-h_{\mathrm{bin}}(t)=\log(1-t).\] Apply it first to \(I=h_{\mathrm{bin}}(\lambda(1-x_c)/2)-\lambda h_c\), then to the two child expressions with initial probabilities \((1-a_c)/2\) and \((1-b_c)/2\), and average. The result is \[
\lambda I'=I-A,\qquad
\lambda\nu'=\nu-A-\kappa.
\tag{59}\] For the second identity, the two logarithms have average \[\frac12\log\frac{(1+x)^2-d^2}{4}=-A-\kappa.\]
Substitution gives the exact Euler derivative \[
\lambda\Delta'-\Delta
=\int_\nu^I(u-A)F''(u)\,du+F'(\nu)\kappa.
\tag{60}\] Indeed, at the chosen value of \(\lambda\), the derivative in \(u\) of \((u-A)F'(u)-F(u)\) is \((u-A)F''(u)\). The value of \(A\) is held fixed only during this auxiliary integration; both path derivatives have already been included.
The last term in (60) is at least \(2\kappa\). Convexity makes the integrand nonnegative where \(u\ge A\). On the remaining part, the curvature bound \(F''(u)\le M/(L-u)^2\) and the substitution \(h=L-u\) give \[
\lambda\Delta'-\Delta
\ge2\kappa-M\int_q^s\frac{(h-b)_+}{h^2}\,dh.
\tag{61}\] Thus it suffices to prove that this integral is at most \(3\kappa\): the strict margin \(2>3M\) will then make \((\Delta/\lambda)'\) nonnegative. The cutoff \(b\) marks exactly the part of the entropy interval where the curvature term can be negative. We will first prove an interval estimate for this kernel, then check its hypotheses using the entropy of the core.
The positive term also has a channel interpretation. For the channel in (56), the posterior of the fair input at the positive output is \[\Pr(X=z\mid U=1)=\frac12\left(1+\frac{zd}{p}\right),
\qquad z\in\{-1,1\}.\] Consequently \[\kappa=\frac12\sum_{z=\pm1}
\log\frac{1/2}{\Pr(X=z\mid U=1)}
=D_{\mathrm{KL}}\!\left(
\operatorname{Unif}\{-1,1\}\,\middle\|\,
\mathcal L(X\mid U=1)\right).\] This is the relative entropy from the prior to the positive-output posterior, with the flag marginalized. It is an auxiliary path quantity, distinct from the canonical cost \(\mathcal K\).
An estimate for entropy intervals
Lemma 15 (An interval estimate with a logarithmic cutoff). Let \(0\le x<1\), and put \[p=1+x,\qquad \sigma=1-x,\qquad
\gamma=\frac{\sigma}{p},\qquad
\beta=\frac{\log(1+x)}L.\] Suppose \[0\le z<\gamma,\qquad \sigma-pz\le u\le1-z.\] For \(\beta>0\), define \(k_\beta(v)=(v-\beta)_+/v^2\) for \(v>0\) and \(k_\beta(0)=0\). For \(\beta=0\), set \(k_0(v)=1/v\) for \(v>0\). Then \[
\int_u^{u+z}k_\beta(v)\,dv
\le-\frac32\log(1-\gamma z).
\tag{62}\] If \(x>0\), the endpoint \(z=\gamma\) is also allowed. At \(z=0\), both sides are zero.
Proof. We first bound the cutoff by a rational function: \[
\beta\ge b_*:=\frac{4x}{3+x},\qquad
1-\beta\le\frac{3\sigma}{3+x}\le\frac32\gamma.
\tag{63}\] For \(0<x\le1\), let \(R(x)=(3+x)\log(1+x)/x\). Direct differentiation gives \[x^2R'(x)=N(x):=\frac{x(3+x)}{1+x}-3\log(1+x),
\qquad N'(x)=-\frac{x(1-x)}{(1+x)^2}\le0.\] Since \(N(0)=0\), \(R\) decreases and \(R(x)\ge R(1)=4L\). This proves the first inequality, including \(x=0\) by continuity. The other two follow from \(1-b_*=3\sigma/(3+x)\) and \(x\le1\).
Consider an admissible interval \([u,u+z]\) with \(z<\gamma\). The constraints imply \[a:=\frac{2-u}{1+z}\in[1,p].\] Keep this \(a\) and the cutoff \(\beta\) fixed, and consider the affine family of intervals \[
u_a(s)=2-a(1+s),\qquad v_a(s)=u_a(s)+s,
\qquad 0\le s\le z.
\tag{64}\] Because \(1\le a\le p\), these intervals satisfy \[0<\sigma-ps\le u_a(s)\le v_a(s)\le1.\] The initial interval has length zero; the final one is \([u,u+z]\). The Leibniz rule gives \[\frac{d}{ds}\int_{u_a(s)}^{v_a(s)}k_\beta(v)\,dv
=D_a(s):=a k_\beta(u_a(s))-(a-1)k_\beta(v_a(s)).\] We claim that \[
D_a(s)\le\frac{3\gamma}{2(1-\gamma s)}.
\tag{65}\]
For this estimate, abbreviate \(u=u_a(s)\) and \(v=v_a(s)\). Above the cutoff, \(k_\beta'(v)=(2\beta-v)/v^3\); thus \(k_\beta\) first increases and then decreases, with the decreasing case \(\beta=0\) included. Its minimum on \([u,1]\) is at an endpoint. Consequently \[k_\beta(v)\ge\min\{k_\beta(u),1-\beta\}.\] If \(k_\beta(u)\le1-\beta\), this gives \(D_a\le k_\beta(u)\le1-\beta\), and (65) follows from (63).
Otherwise set \(A_*=k_\beta(u)-(1-\beta)>0\). Then \(u>\beta\), and \[D_a\le aA_*+(1-\beta),\qquad
Y:=\frac{1-\gamma s}{\gamma}
=\frac2\sigma-\frac{2-u}{a}>0.\] To bound their product, hold \(u,\beta,x\) fixed and vary \(a\) in the algebraic expression \[\Theta(a)=\{aA_*+(1-\beta)\}
\left\{\frac2\sigma-\frac{2-u}{a}\right\}.\] It increases with \(a\), since \[\Theta'(a)=\frac{2A_*}{\sigma}
+\frac{(2-u)(1-\beta)}{a^2}>0.\] The actual value has \(a\le p\), so \[
YD_a\le\Theta(p)
=\{p k_\beta(u)-x(1-\beta)\}
\frac{4x+\sigma u}{\sigma p}.
\tag{66}\] This last expression decreases with \(\beta\): its positive factor is independent of \(\beta\), and the first factor has derivative \(-p/u^2+x<0\). Since \(u>\beta\ge b_*\), replacing \(\beta\) by \(b_*\) keeps \(u\) above the cutoff and gives an upper bound. The hard-case condition also gives \[u^2\{k_\beta(u)-(1-\beta)\}
=(1-u)\{(1-\beta)u-\beta\}>0.\] Thus \(u<1\) and \(u>\beta/(1-\beta)\ge b_*/(1-b_*)\). Putting \(t=b_*/u\) now yields \(0\le t<1/(1+u)\). Using \(x=3b_*/(4-b_*)\), substitution and a common denominator give \[\frac32-\left.\Theta(p)\right|_{\beta=b_*}
=\frac{\mathcal P(t,u)}{2(1-tu)(2+tu)},\] where the numerator satisfies \[
\begin{split}
\mathcal P(t,u):={}&3(2-tu-t^2u^2)\\
&-\{2(2+tu)(1-t)-3tu^2(1-tu)\}\{1+t(3-u)\}\ge0,
\qquad 0\le u\le1,\quad 0\le t\le\frac1{1+u}.
\end{split}
\tag{67}\] To prove this polynomial inequality, put \[C_2=12-8u+8u^2-6u^3,\qquad C_1=3u^2-u-8,\qquad
C_3=u(3-u)(2-3u^2).\] Expansion gives \[
\mathcal P=2+C_1t+C_2t^2+C_3t^3.
\tag{68}\] Since \[C_3+u(1+u)=u(1-u)(7+6u-3u^2)\ge0,
\qquad t(1+u)\le1,\] we have \(C_3t\ge-u\): this is immediate for \(C_3\ge0\), while for \(C_3<0\) it follows from \(C_3\ge-u(1+u)\) and \(t\le1/(1+u)\). Hence \(\mathcal P\ge2+C_1t+(C_2-u)t^2\). The exact identity \[8(C_2-u)-C_1^2
=4+(1-u)\{2(5u-3)^2+u^2(1+9u)+10\}\ge4\] therefore gives \[8\mathcal P\ge(4+C_1t)^2+4t^2>0.\] This proves (67), including \(t=0\). The denominator in the preceding gap is positive because \(tu=b_*<1\). Thus \(YD_a\le3/2\), which is (65). Integrating from \(0\) to \(z\) proves (62).
For \(x=0\), admissibility forces \(a=1\) and the family is \([1-s,1]\), whose lower endpoint is positive for \(s<1\). The integrand \(k_0(v)=1/v\) therefore causes no endpoint problem. For \(x>0\), the cutoff is positive and \(k_\beta\) is continuous at zero with the stated extension. Also \(\gamma<1\), so continuity along the same family extends the bound to \(z=\gamma\). The zero-length case is immediate. ◻
The entropy interval of the channel
We now return to the quantities in (58). To apply the interval estimate to \([q,s]\), we need a lower bound on its position and an upper bound on its length \(J\). Both follow from the entropy of the oriented core.
Lemma 16 (The entropy interval on a thinning path). Every state on (54) satisfies \[
\int_q^s\frac{(h-b)_+}{h^2}\,dh\le3\kappa.
\tag{69}\]
Proof. The nonnegative center of the core supplies the essential lower entropy bound. Write \(h_{\mathrm{bin}}(t)=H(1-2t)\), let \(\sigma_c=1-x_c\le1\), and put \[t=\frac{1-d_c/\sigma_c}{2}\in(0,1/2).\] The two negative-output probabilities at the core are \(\sigma_c t\) and \(\sigma_c(1-t)\). Concavity and \(h_{\mathrm{bin}}(0)=0\) imply \[h_{\mathrm{bin}}(\sigma_c t)\ge\sigma_c h_{\mathrm{bin}}(t),\qquad
h_{\mathrm{bin}}(\sigma_c(1-t))
\ge\sigma_c h_{\mathrm{bin}}(1-t)=\sigma_c h_{\mathrm{bin}}(t).\] Averaging and multiplying by \(\lambda\), with \(\sigma=\lambda\sigma_c\) and \(d/\sigma=d_c/\sigma_c\), gives \[
e\ge\sigma H(d/\sigma).
\tag{70}\] This is a consequence of the oriented common-output path; it is not an assertion about every pair law with the same means. The bound \(H(r)\ge L(1-r^2)\) in Lemma 4 now yields \[
q\ge e\ge L\left(\sigma-\frac{d^2}{\sigma}\right).
\tag{71}\]
Expanding the binary entropies in \(J\) gives the posterior identity \[J=\frac p2\psi(d/p)+\frac\sigma2\psi(d/\sigma).\] Using \(\psi(r)\le Lr^2\) on both terms, we obtain \[J\le\frac{Ld^2}{\sigma p}=:Lz,\qquad
0<z<\frac\sigma p.\] The upper bound on \(z\) follows from \(d<\sigma\). Together with (71), these estimates say \[q\ge L(\sigma-pz),\qquad q+J=s\le L,
\qquad L(\sigma-pz)+Lz=L(\sigma-xz)\le L.\] We may therefore enlarge \([q,s]\) to an interval of length \(Lz\) inside \([L(\sigma-pz),L]\). More explicitly, choose its lower endpoint as \(q_*:=\min\{q,L-Lz\}\). Both candidates are at least \(L(\sigma-pz)\), and \(q_*\le q\); its upper endpoint is either \(q+Lz\ge s\) or \(L\ge s\). The integrand is nonnegative, so this enlargement can only increase the integral.
Rescale \(h=Lv\). With \(\beta=b/L\) the integrand becomes \(k_\beta(v)\,dv\), with no additional factor of \(L\). The lower endpoint \(q_*/L\) and length \(z\) satisfy Lemma 15 for this \(x\) and \(\gamma=\sigma/p\). Hence \[\int_q^s\frac{(h-b)_+}{h^2}\,dh
\le-\frac32\log(1-\gamma z)
=-\frac32\log\left(1-\frac{d^2}{p^2}\right)
=3\kappa.\] This includes \(x=0\), which on the path can occur only at \(\lambda=1\), \(x_c=0\). ◻
Proof of Theorem 14. Combine the derivative reduction (61) with Lemma 16: \[\begin{align*}
\lambda\Delta'-\Delta
&\ge2\kappa-M\int_q^s\frac{(h-b)_+}{h^2}\,dh\\
&\ge(2-3M)\kappa\ge0.
\end{align*}\] Thus \((\Delta/\lambda)'=(\lambda\Delta'-\Delta)/\lambda^2\ge0\). At \(\lambda=1\), the flag is deterministic, so \(\nu=0\), \(I=J_c\), and \(\Delta=F(J_c)\). Comparing with this endpoint proves (57). No limit as \(\lambda\) tends to zero is needed. ◻
The canonical high-entropy value
Theorem 17 (The canonical high-entropy value comparison). Let \(x\ge0\), \(d>0\), and \(x+d<1\). Define \[\bar h(x,d)=\frac{H(x+d)+H(x-d)}2,\qquad
e_0(x,d)=(1-x)H\!\left(\frac d{1-x}\right).\] For every \(e_0(x,d)\le e\le\bar h(x,d)\), \[
\mathcal K(x,d,e)\ge
J_0(x,d,e):=F(H(x)-e)-F(\bar h(x,d)-e).
\tag{72}\]
Proof. The canonical classification at (35) represents every state in this closed high-entropy branch by (54), with a nonnegative core center and \(\mathcal K(x,d,e)=\lambda c(a_c,b_c)\). Put \(\alpha_c=1-x_c^2\). Lemma 10 and convexity of \(F\), with \(F(0)=0\), give \[c(a_c,b_c)\ge\alpha_c F(J_c/\alpha_c)\ge F(J_c).\] Theorem 14 therefore implies \[J_0(x,d,e)=F(\nu+J)-F(\nu)
\le\lambda F(J_c)
\le\lambda c(a_c,b_c)=\mathcal K(x,d,e).\] ◻
The canonical optimizer is used here only to evaluate a lower bound on joining cost. For an original pair law with means \(x\pm d\) and individual entropies \(e_a,e_b\), the value of \(e\) in (72) is their actual average. The original child information values remain \(H(x+d)-e_a\) and \(H(x-d)-e_b\); their average is \(\bar h(x,d)-e\). Convexity of \(F\) therefore gives the common floor at those original child states: \[\frac{F(H(x+d)-e_a)+F(H(x-d)-e_b)}2
\ge F(\bar h(x,d)-e).\] No individual entropy moment of the canonical law is substituted for either original child entropy.
From the linear reserve to the sharp rate
The local reserve inequality does not by itself give the sharp rate \(F(I)\) at every state. We now join their maximum \[\widehat{\mathcal B}_m(I)
=\max\{\mathcal B_m(I),F(I)\}.\] Closure under a maximum is a useful feature of the local-production framework (Chen et al. 2025, Lemma 3), but \(F\) alone does not satisfy a joining inequality for every law. We therefore prove the needed comparison in the part of the domain where the reserve is insufficient. Two comparisons organize the proof. Above half information, the reserve already dominates \(F\). Below half information, the demand from \(F\) increases with the common mean of the child states. Thus every reflected core can be paid at its one-corner wall by the same thinning theorem used for a one-corner optimizer.
The reserve above half information
Lemma 18 (The high-information reserve). For \(|m|<1\) and \(L/2\le I<H(m)\), \[
\mathcal B_m(I)\ge F(I).
\tag{73}\] The inequality is strict if \(m\ne0\). Consequently a state with \(F(I)>\mathcal B_m(I)\) must have \(I<L/2\).
Proof. Recall \(S(I)=F(I)-2I\) on \([0,L)\) and put \[k_0=L-\frac12,\qquad
\Xi(I)=(I-k_0)S'(I)-S(I).\] Since \(S(0)=S'(0)=0\) and \(S''=F''\) on this interval, \[
\Xi(I)=\int_0^I(s-k_0)F''(s)\,ds.
\tag{74}\] Lemma 6 says that \(w(s)=e^{-2s}F''(s)\) strictly increases. Also \(0<k_0<L/2\), and the binary identity \(e^L=2\) gives the exact balance \[
\int_0^{L/2}(s-k_0)e^{2s}\,ds
=\left[\frac12e^{2s}(s-L)\right]_0^{L/2}=0.
\tag{75}\] On the part of the integral below \(k_0\), the factor \(s-k_0\) is negative and \(w(s)\le w(k_0)\). Above \(k_0\), one has \(s-k_0>0\) and \(w(s)\ge w(k_0)\). Replacing \(w(s)\) by \(w(k_0)\) therefore gives a lower bound on both parts, strictly so on intervals of positive length. Equations (74)– (75) imply \[\Xi(L/2)
=\int_0^{L/2}(s-k_0)e^{2s}w(s)\,ds>0.\] Since \(\Xi'(I)=(I-k_0)F''(I)>0\) for \(I\ge L/2\), we have \(\Xi(I)>0\) throughout that range.
For the mean-dependent comparison, set \[\alpha=1-m^2,\qquad
c_m=H(m)-\alpha L=Lm^2-\psi(m).\] The elementary entropy bounds give \(0\le c_m\le m^2k_0\). Hence, for \(I\ge L/2\), the argument \[s=L-\frac{H(m)-I}{\alpha}
=\frac{I-c_m}{\alpha}\] is positive; it is less than \(L\) because \(I<H(m)\). No clipping occurs in this comparison. The tangent inequality for the convex function \(S\), and \(S'\ge0\), now yield \[\begin{align*}
\mathcal B_m(I)-F(I)
&=\alpha S(s)-S(I)\\
&\ge\alpha[S(I)+(s-I)S'(I)]-S(I)\\
&=m^2[IS'(I)-S(I)]-c_mS'(I)\\
&\ge m^2[(I-k_0)S'(I)-S(I)]
=m^2\Xi(I)\ge0.
\end{align*}\] The inequality is strict at a nonzero mean. At mean zero, \(\alpha=1\), \(c_m=0\), and the two profiles agree identically. ◻
A direct comparison with the one-corner wall
Lemma 19 (The demand below half information). Let \(m,d>0\), \(m+d<1\), and \[0<e\le\bar h(m,d):=\frac{H(m+d)+H(m-d)}2.\] Put \(I=H(m)-e\), \(j=\bar h(m,d)-e\), and \(J=I-j\). If \(I\le L/2\), then at fixed \(d,e\), \[
\partial_m\{F(I)-F(j)\}>0.
\tag{76}\] At \(j=0\), the derivative is interpreted from the physical side.
Proof. We compare the rate slopes with the entropy slopes. First, \[
\frac{F''(u)}{F'(u)}<2\qquad(0\le u\le L/2).
\tag{77}\] Indeed, set \(h=L-u\ge L/2\). The lower bound \(F''(u)\ge1/[2(L-u)^2]\) from Lemma 7 gives \[F'(u)\ge2+\frac12\left(\frac1h-\frac1L\right).\] Using the upper bound from the same lemma, we get \[2h^2F'(u)\ge(4-L^{-1})h^2+h
\ge L^2+\frac L4>\frac43L^2\ge h^2F''(u).\] The quadratic in \(h\) increases for \(h>0\), and the strict comparison uses \(L<3/4\). This proves (77). Since \(0\le j<I\le L/2\), integration gives \[
\log\frac{F'(I)}{F'(j)}<2(I-j)=2J.
\tag{78}\]
For the entropy slopes, write \[A=-\partial_m\bar h(m,d)
=\frac{V(m+d)+V(m-d)}2>0.\] We claim the reverse comparison \[
\log\frac{A}{V(m)}\ge2J.
\tag{79}\] Hold \(m>0\) fixed and set \(v_\pm=V(m\pm d)\). The function \(G(v)=\cosh^2v-v^2\) is even and strictly increasing for \(v>0\), because \(G'(v)=\sinh(2v)-2v>0\). Direct differentiation gives \[\partial_d\left\{\log\frac{A}{V(m)}-2J\right\}
=\frac{G(v_+)-G(v_-)}{v_++v_-}>0.\] Indeed, \(A_d=(\cosh^2v_+-\cosh^2v_-)/2\) and \(2J_d=v_+-v_-\). The numerator and denominator are positive since \(m+d>|m-d|\), so \(v_+>|v_-|\). At \(d=0\) the expression in braces is zero. Integration proves (79).
Finally \(I_m=-V(m)\) and \(j_m=-A\). Therefore \[\partial_m\{F(I)-F(j)\}
=-F'(I)V(m)+F'(j)A>0\] by (78)–(79). The regular right derivative \(F'(0)=2\) includes a capacity endpoint. ◻
The local inequality and dimension induction
Theorem 20 (Joining the strengthened profile). Let \((A,B)\) have a joint law on \([-1,1]^2\) with \(e_a=\mathbb EH(A)>0\) and \(e_b=\mathbb EH(B)>0\). Put \[a=\mathbb EA,\quad b=\mathbb EB,\quad
m=\frac{a+b}{2},\quad e=\frac{e_a+e_b}{2},\quad
I=H(m)-e,\] and \(I_a=H(a)-e_a\), \(I_b=H(b)-e_b\). Then \[
\mathbb Ec(A,B)\ge
\widehat{\mathcal B}_m(I)
-\frac{\widehat{\mathcal B}_a(I_a)
+\widehat{\mathcal B}_b(I_b)}2.
\tag{80}\]
Proof. An infinite joining cost makes the assertion immediate. Otherwise positive child entropies give \(|a|,|b|<1\), and entropy concavity makes all the profile arguments physical.
If the parent maximum is attained by \(\mathcal B_m(I)\), apply the local reserve inequality of Theorem 13 and enlarge the two child terms to their \(\widehat{\mathcal B}\) values. This includes every tie at the parent.
It remains to suppose \(F(I)>\mathcal B_m(I)\). Put \(d=|a-b|/2\) and \[\bar h(m,d)=\frac{H(m+d)+H(m-d)}2,
\qquad j=\bar h(m,d)-e=\frac{I_a+I_b}{2}.\] Since \(\widehat{\mathcal B}\ge F\), convexity at the original child states gives \[
\frac{\widehat{\mathcal B}_a(I_a)
+\widehat{\mathcal B}_b(I_b)}2
\ge\frac{F(I_a)+F(I_b)}2\ge F(j).
\tag{81}\] For \(d=0\), one has \(I=j\) and nonnegativity of cost proves the claim. Otherwise reflect to \(m>0\): mean zero is impossible in this strict-parent case, because \(\mathcal B_0=F\). Let \[J_0(x,d,e)=F(H(x)-e)-F(\bar h(x,d)-e).\] The actual law is feasible in the common-budget cost problem, so it is enough to prove \[
\mathcal K(m,d,e)\ge J_0(m,d,e).
\tag{82}\]
Use the exact optimizer from Proposition 9. In its one-positive-corner geometry, the entropy is at least the transition value (35); Theorem 17 therefore proves (82).
In the reflected geometry (33), write \[d=kr,\qquad e=kH(r),\qquad \ell=1-k,
\qquad 0<r<1,\quad0<k\le1.\] Then \(0<m\le\ell\), and the cost is \(dV(r)\). Hold \(d,e,k,r\) fixed and increase the common mean from \(m\) to \(\ell\). The law with mass \(k\) at \((r,-r)\) and masses \((1-k+x)/2\), \((1-k-x)/2\) at the equal corners makes every \(x\in[m,\ell]\) physical at the same entropy and cost. Its means remain interior, since \(\ell+d=1-k(1-r)<1\).
Lemma 18 gives \(H(m)-e<L/2\). The function \(H(x)-e\) decreases along this interval, so Lemma 19 applies throughout. Consequently \[J_0(m,d,e)\le J_0(\ell,d,e).\] At \(x=\ell\), the negative-corner mass is zero, and the state is at the one-corner transition. Theorem 17 gives \[J_0(\ell,d,e)\le dV(r)=\mathcal K(m,d,e).\] This proves (82). A zero-length interval requires only the last comparison. The case \(k=1\) would force \(m=0\) and is already covered by the reserve branch.
Only the information bound is retained while the mean increases; no claim is made about which profile stays active there. The child terms in (81) always use the original entropies \(e_a,e_b\), not the entropies of a replacement optimizer. Together with that pooling inequality, the cost bound completes (80). ◻
Proof of Theorem 3. Induct on dimension. At dimension zero, a constant field has \(D=I=0\) and \(\widehat{\mathcal B}_m(0)=0\). For a coordinate split \(g_+,g_-\), the exact energy decomposition is \[D(g)=\frac{D(g_+)+D(g_-)}2+\mathbb Ec(g_+,g_-).\] Each within-face term has its face probability \(1/2\); symmetry of the joining cost combines the two orientations of every joining edge into the last expectation. The induction hypothesis bounds the within-face productions, and Theorem 20 supplies the remaining difference on the joining edges. Their sum is the parent \(\widehat{\mathcal B}\) value. Every face is interior-valued, so all child entropies are positive and all local uses are physical. ◻
Sharp entropy interpolation under noise
The sharp production bound has a direct dynamical meaning: a soft field loses information at least as fast as a noisy coordinate with the same information. The appropriate coordinate for this comparison is the inverse entropy deficit, rather than information itself. In that coordinate the comparison flow is simply exponential decay.
Proof of Theorem 1. Write \(g=u\) for the field in the theorem. First suppose \(0<\rho<1\), and write \(\rho=e^{-t}\) with \(t>0\). If \(g\) is constant, both sides of (1) are zero. Otherwise every value of \(g_s=T_{e^{-s}}g\) lies in \((-1,1)\) for \(s>0\): the noise kernel gives strictly positive weight to every input, and a nonconstant \([-1,1]\)-valued input cannot average to either endpoint. The mean stays fixed. By (4) and Theorem 3, \[
\frac{d}{ds}I(g_s)=-D(g_s)\le-F(I(g_s))
\qquad(s>0).
\tag{83}\]
For a nonconstant input, \(g_s\) is nonconstant at every finite time, because the Walsh characters \(\chi_A(x)=\prod_{i\in A}x_i\), \(A\subseteq\{1,\ldots,n\}\), form a basis and satisfy \(T_{e^{-s}}\chi_A=e^{-s|A|}\chi_A\). All these eigenvalues are positive. Strict convexity of \(\psi\) therefore gives \(0<I(g_s)<L\). Set \(r(s)=\psi^{-1}(I(g_s))\in(0,1)\). The identity \(F(\psi(r))=rV(r)\) and \(\psi'(r)=V(r)\) turn (83) into \[V(r(s))r'(s)\le-r(s)V(r(s)),\qquad r'(s)\le-r(s).\] Thus, for \(0<\varepsilon<t\), \[r(t)\le e^{-(t-\varepsilon)}r(\varepsilon).\] On a finite cube the noise operator and the continuous functional \(I\) converge to their initial values as \(\varepsilon\downarrow0\). Continuity of \(\psi^{-1}\), including at \(L\), yields \[\psi^{-1}(I(T_{e^{-t}}g))
\le e^{-t}\psi^{-1}(I(g)).\] Applying the increasing function \(\psi\) proves the claim for \(0<\rho<1\). This positive-time argument does not require the initial entropy production to be finite.
At \(\rho=0\), the field is constant; at \(\rho=1\), the assertion is an identity. For negative correlation, \(T_{-\rho}g(x)=T_\rho g(-x)\), and the change of variables \(x\mapsto-x\) preserves the uniform measure and information. This proves the full stated range.
For \(g(x)=a x_i\) with \(|a|\le1\), the initial information is \(\psi(|a|)\) and the information after noise is \(\psi(|\rho a|)\). Thus every soft coordinate attains equality. ◻
Proof of Corollary 2. Since \(f\) is Boolean, \(H(f(x))=0\) at every input, so \(I(f)=H(m)\). Also \(I(f(X);Y)=I(T_\rho f)\). The first inequality is therefore Theorem 1. The second follows from \(r_m\le1\) and monotonicity of \(\psi\). If \(0<|m|<1\), then \(0<H(m)<L\) and \(0<r_m<1\), giving the stated strict comparison with the unbiased bound. For a signed coordinate \(f\), \(T_\rho f=\rho f\), so its information is exactly \(\psi(|\rho|)\). ◻
Comparison of mean-dependent production bounds
Theorem 3 gives more than the scalar rate needed for Theorem 1. For \(0<m<1\) and \(0\le I<H(m)\), put \(h=H(m)-I\). The hybrid production profile of Chen et al. (Chen et al. 2026, sec. 9 and Lemma 9.2), converted to natural logarithms and the generator \(N\) used here, is \[C_m(I)=\max\{F(L-h)-mV(s),F(I)\},\qquad mH(s)=hs,
\qquad m\le s<1.\] The equation for \(s\) has a unique solution: \(H(s)/s\) decreases strictly from \(H(m)/m\) to zero on \([m,1)\). In particular, the \(F(I)\) branch yields Theorem 1 by the inverse-entropy integration above, with the initial information left unchanged.
The mean-dependent estimate in Theorem 3 has a different strength. At every fixed \(m\in(0,1)\), \[
\mathcal B_m(H(m)-h)\sim\frac{1-m^2}{2}\log\frac1h,
\qquad C_m(H(m)-h)\sim\frac{1-m}{2}\log\frac1h
\quad(h\downarrow0).
\tag{84}\] Here is a direct verification. For \(r=1-\varepsilon\), \[H(r)=\frac\varepsilon2\log\frac{2\exp(1)}{\varepsilon}
+O(\varepsilon^2),\qquad
V(r)=\frac12\log\frac2\varepsilon+O(\varepsilon).\] Consequently \(F(L-h)\sim\tfrac12\log(1/h)\). The equation \(H(s)=hs/m\) gives \(s\to1\) and \(V(s)\sim\tfrac12\log(1/h)\). As \(F(H(m)-h)\) remains bounded, these facts give the second equivalent in (84). The first follows from \[\mathcal B_m(H(m)-h)
=2(H(m)-h)+(1-m^2)S\left(L-\frac h{1-m^2}\right)\] and \(S(u)=F(u)-2u\) for positive \(u\). Thus the displayed profile \(C_m\) does not dominate \(\mathcal B_m\): the ratio \(\mathcal B_m(H(m)-h)/C_m(H(m)-h)\) tends to \(1+m>1\) as \(h\downarrow0\). This comparison concerns the two explicit production profiles.
Barnes, Leighton Pate, and Ayfer Özgür. 2020. “The Courtade–Kumar Most Informative Boolean Function Conjecture and a Symmetrized Li–Médard Conjecture Are Equivalent.”2020 IEEE International Symposium on Information Theory (ISIT), 2205–9. https://doi.org/10.1109/ISIT44484.2020.9174063.
Beckner, William. 1975. “Inequalities in Fourier Analysis.”Annals of Mathematics 102 (1): 159–82. https://doi.org/10.2307/1970980.
Bobkov, Sergey G., and Prasad Tetali. 2006. “Modified Logarithmic Sobolev Inequalities in Discrete Settings.”Journal of Theoretical Probability 19 (2): 289–336. https://doi.org/10.1007/s10959-006-0016-3.
Bonami, Aline. 1970. “Étude Des Coefficients de Fourier Des Fonctions de \(L^p(G)\).”Annales de l’Institut Fourier 20 (2): 335–402. https://doi.org/10.5802/aif.357.
Chen, Zijie, Amin Gohari, and Chandra Nair. 2025. A Differential Equation Approach to the Most-Informative Boolean Function Conjecture. Https://arxiv.org/abs/2502.10019v2.
Courtade, Thomas A., and Gowtham R. Kumar. 2014. “Which Boolean Functions Maximize Mutual Information on Noisy Inputs?”IEEE Transactions on Information Theory 60 (8): 4515–25. https://doi.org/10.1109/TIT.2014.2326877.
Gohari, Amin, and Chandra Nair. 2022. “Outer Bounds for Multiuser Settings: The Auxiliary Receiver Approach.”IEEE Transactions on Information Theory 68 (2): 701–36. https://doi.org/10.1109/TIT.2021.3128136.
Javanmard, Adel, and David P. Woodruff. 2026. Progress on the Courtade–Kumar Conjecture: Optimal High-Noise Entropy Bounds and Generalized Coordinate-Wise Mutual Information. Https://arxiv.org/abs/2601.09679v1.
Kindler, Guy, Ryan O’Donnell, and David Witmer. 2016. Continuous Analogues of the Most Informative Function Problem. Https://arxiv.org/abs/1506.03167v3.
Kramer, Dimiter, and Mert Saglam. 2026. On the Most Informative Boolean Function Under Noise. Preprint.
Kumar, Gowtham R., and Thomas A. Courtade. 2013. “Which Boolean Functions Are Most Informative?”2013 IEEE International Symposium on Information Theory (ISIT), 226–30. https://doi.org/10.1109/ISIT.2013.6620221.
Li, Jiange, and Muriel Médard. 2021. “Boolean Functions: Noise Stability, Non-Interactive Correlation Distillation, and Mutual Information.”IEEE Transactions on Information Theory 67 (2): 778–89. https://doi.org/10.1109/TIT.2020.3041028.
Ordentlich, Or, Ofer Shayevitz, and Omri Weinstein. 2016. “An Improved Upper Bound for the Most Informative Boolean Function Conjecture.”2016 IEEE International Symposium on Information Theory (ISIT), 500–504. https://doi.org/10.1109/ISIT.2016.7541349.
Pichler, Georg, Pablo Piantanida, and Gerald Matz. 2018. “Dictator Functions Maximize Mutual Information.”The Annals of Applied Probability 28 (5): 3094–101. https://doi.org/10.1214/18-AAP1384.
Samorodnitsky, Alex. 2016. “On the Entropy of a Noisy Function.”IEEE Transactions on Information Theory 62 (10): 5446–64. https://doi.org/10.1109/TIT.2016.2584625.
Wyner, Aaron D., and Jacob Ziv. 1973. “A Theorem on the Entropy of Certain Binary Sequences and Applications—I.”IEEE Transactions on Information Theory 19 (6): 769–72. https://doi.org/10.1109/TIT.1973.1055107.
Yu, Lei. 2023. “On the \(\Phi\)-Stability and Related Conjectures.”Probability Theory and Related Fields 186 (3–4): 1045–80. https://doi.org/10.1007/s00440-023-01209-5.
Yu, Lei. 2026. “Local Optimality of Dictator Functions with Applications to Courtade–Kumar and Li–Médard Conjectures.”The Annals of Applied Probability.
LEVEL 1 COMPLETE!
You read 9,346 words and 975 formulas. Your math teacher would be proud. Converted from the LaTeX source. Something look off? The original PDF is the real thing.