Leveraged Learning: entropy cleared per bit received
An expository essay containing some new results. A learner holds a prior belief over Boolean maps that answer a finite set of $Q$ questions, and receives answers one by one. Each answer carries surprisal and can clear predictive uncertainty about both the question asked and questions still unasked. We quantify this effect by the leverage: table entropy cleared per bit of surprisal received. Taking expectations over the truth prior and the question order, define the aggregate leverage as the ratio of expected uncertainty cleared to expected surprisal received. It equals one for independent answers and can exceed one for correlated answers. At finite size, this ratio is determined exactly by a single sequence: the mean entropy $G_\ell$ of the answers to $\ell$ questions. As the number of input bits grows at fixed asked fraction $t=\ell/Q$, a limiting increment profile $γ(t)$ determines the macroscopic learning curve. With $η_0$ the limiting initial entropy per question, the leverage becomes $L(t)= \frac{η_0-(1-t)γ(t)} {\int_0^tγ(x),dx}$. For exchangeable priors, de Finetti's representation gives a constant bulk profile $γ(t)$: deduction is confined to a boundary layer at $t=0$, and the leverage is forced to a hyperbolic form, as surprisal grows linearly. By contrast, we construct a simplicity prior with nontrivial bulk learning by grading Boolean maps by their polynomial degree over $\mathbb{F}_2$ and allocating weight across degree classes through a CDF $F$. Reed-Muller capacity then yields $γ(t)=1-F(t)$. This realizes any nonincreasing profile taking values in $[0,1]$, together with the corresponding macroscopic leverage curve.