- Boxes
- definitions
- Ellipses
- theorems and lemmas
- Blue border
- the statement of this result is ready to be formalized; all prerequisites are done
- Orange border
- the statement of this result is not ready to be formalized; the blueprint needs more work
- Blue background
- the proof of this result is ready to be formalized; all prerequisites are done
- Green border
- the statement of this result is formalized
- Green background
- the proof of this result is formalized
- Dark green background
- the proof of this result and all its ancestors are formalized
- Dark green border
- this is in Mathlib
Uniform output-layer regularity. The output-layer class \(H \subseteq \mathbb R^{\mathcal X}\) has finite uniform covering numbers \(N(H, \| \cdot \| _\infty , u) {\lt} \infty \) for all \(u {\gt} 0\), and there are constants \(B_H, L_H\) with \(\| h\| _\infty \le B_H\) and \(|h(x) - h(x')| \le L_H d(x,x')\) for all \(h \in H\).
Sample output-layer regularity. For a sample \(S\) and a hidden class \(F\), the output-layer class \(H \subseteq \mathbb R^{\mathcal X}\) has finite uniform pushed-forward covering numbers \(N_{S,F}(H, u) {\lt} \infty \) for all \(u {\gt} 0\), and there are constants \(B_H\), \(L_H\) such that every \(h \in H\) satisfies \(|h(x)| \le B_H\) and \(|h(x) - h(x')| \le L_H d(x, x')\) for all \(x, x' \in \mathcal R_S(F)\) (boundedness and the Lipschitz estimate are only required on the points reached by the sample; compare ass:ent-readout, where both hold on all of \(\mathcal X\) and \(H\) is covered in the sup norm).
Output-layer realization of hidden geometry. For a sample \(S\), an output-layer class \(H\), a hidden class \(B_k\) and constants \(\kappa , R_{\mathrm{out}}\): there is, for each \(g \in B_k\), an output layer \(h_g \in H\) such that the map \(\Psi _k(g) = h_g \circ g\) satisfies \(\| \Psi _k(g) - \Psi _k(g')\| _S \ge \kappa \, d_S(g,g')\) for \(g, g' \in B_k\) and \(\| \Psi _k(g)\| _{S,\infty } \le R_{\mathrm{out}}\) for \(g \in B_k\).
Sub-Gaussian output-layer increments. Let \(\mathfrak F\) be a hidden-layer class and \(A_H, L\) constants. For all \(f, g \in \mathfrak F\) and \(t {\gt} 0\),
where \(\mathbb P_\sigma \) is the uniform distribution on \(\{ \pm 1\} ^n\); and when \(d_S(f,g) = 0\) the requirement is \(Z_f = Z_g\) (for every \(\sigma \)), which is the limiting interpretation of the display (in Lean the display with \(d_S(f,g)=0\) reads \(\ldots \le 2\exp (0)\), hence the separate conjunct).
(E1: free semigroup with one-point uniform separation.) Let \(F = \{ f_1, \dots , f_r\} \), \(r \ge 2\), and suppose there are a base point \(x_* \in \mathcal X\) and \(\delta {\gt} 0\) such that for every \(k\) and all distinct words \(u \ne v \in [r]^k\),
Then for every \(k\) and every \(\varepsilon {\lt} \delta /2\), \(N^{\mathrm{ext}}(B(k,F), d_\infty , \varepsilon ) \ge r^k\). (The paper’s hypothesis (1), that \(u \mapsto f_u\) is injective on \([r]^k\), follows from the separation hypothesis since \(\delta {\gt} 0\), so it is omitted; the hypothesis \(r \ge 2\) is not used in the proof.)
(E1’: equal-length coding.) Let \(F = \{ f_1, \dots , f_r\} \), \(r \ge 2\), and suppose there are \(x_* \in \mathcal X\) and \(\delta {\gt} 0\) with \(d(f_u(x_*), f_v(x_*)) \ge \delta \) for all distinct \(u, v \in [r]^k\) and all \(k\). Then \(N^{\mathrm{ext}}(B(k,F), d_\infty , \varepsilon ) \ge r^k\) for all \(k\) and all \(\varepsilon {\lt} \delta /2\). (This generalizes the paper, which assumes in addition that the generators are isometries, that \(d_\infty \) is finite on \(\langle F\rangle \) and that \(u \mapsto f_u\) is injective on \([r]^k\): none of these is needed for the lower bound, which follows from ‘cond:e1-free-iso‘ in one line; the isometry hypothesis only serves, in the paper, to make the covering numbers meaningful.)
(E2: ping–pong coding.) Let \(F = \{ f_1, \dots , f_r\} \), \(r \ge 2\), and suppose there are sets \(V_i \subseteq U_i \subseteq \mathcal X\), anchors \(a_1, \dots , a_r \in \mathcal X\), a marker \(q \in \mathcal X\) and a constant \(\alpha {\gt} 0\) such that, with \(A = \{ a_1, \dots , a_r\} \) and \(Q = \{ q\} \cup \bigcup _j V_j\):
(disjoint chambers) \(U_i \cap U_j = \emptyset \) for \(i \ne j\);
(coding cores) \(Q \subseteq f_i(V_i)\) for every \(i\);
(reset) \(f_i(x) = a_i\) for \(x \notin U_i\), and \(f_i(A) \subseteq A\);
(marker separation) \(d(q, a_i) \ge \alpha \) for every \(i\).
Then (a) \(N^{\mathrm{ext}}(B(k,F), d_\infty , \varepsilon ) \ge r^k\) for every \(k\) and every \(\varepsilon {\lt} \alpha /2\), and (b) \(u \mapsto f_u\) is injective on \([r]^k\) for every \(k\). (This generalizes the paper, which assumes in (1) a uniform separation \(d(U_i, U_j) \ge \Delta {\gt} 0\) and additionally \(V_i \ne \emptyset \): the proof uses of (1) only the disjointness of the chambers, and \(V_i \ne \emptyset \) follows from (2) since \(q \in f_i(V_i)\); the hypothesis \(r \ge 2\) is not used.)
(E3: memory-preserving expansion grows super- or double-exponentially.) Let \(E\) be a normed space, \(G \subseteq E\), \(\lambda {\gt} 1\), \(X = E^{\mathbb N}\) with the bounded sup metric, \(F = \{ r, A\} \cup \{ g_u : u \in G\} \) and \(W_k = \{ w_u : u \in G^k\} \subseteq B(2k+1, F)\). Then (i) each \(w_u\) is the constant map with value \((\lambda u_{k-1}, \dots , \lambda ^k u_0, 0, \dots )\) (‘cond:e3-i‘), (ii) \(d_\infty (w_u, w_v) = \max _{j{\lt}k} \min \{ 1, \lambda ^{j+1}\| u_{k-1-j} - v_{k-1-j}\| \} \) (‘cond:e3-ii‘), and (iii) for every \(0 {\lt} \varepsilon {\lt} 1/2\),
Consequently \(N^{\mathrm{ext}}(B(2k+1,F), d_\infty , \varepsilon ) \ge N^{\mathrm{ext}}(W_k, d_\infty , \varepsilon )\).
(P1: equicontinuous semigroup on a compact domain saturates, hypothesis 1.) Assume \(\mathcal X\) is compact and the semigroup \(\langle F \rangle \) is precompact (totally bounded) in \(d_\infty \). Then, with \(G := \overline{\langle F\rangle }^{d_\infty }\), for all \(\varepsilon {\gt} 0\) and all \(k\),
hence no dependence on \(k\). (Compactness of \(\mathcal X\) is not needed for this hypothesis.)
(P1, hypothesis 2a.) Assume \(\mathcal X\) is compact and the semigroup \(\langle F \rangle \) is equicontinuous. Then for all \(\varepsilon {\gt} 0\) and all \(k\), \(N^{\mathrm{ext}}(B(k,F), d_\infty , \varepsilon ) \le N^{\mathrm{ext}}(\overline{\langle F\rangle }, d_\infty , \varepsilon ) {\lt} \infty \).
(P1, hypothesis 2b.) Assume \(\mathcal X\) is compact and the semigroup \(\langle F \rangle \) is uniformly Lipschitz: every \(g \in \langle F \rangle \) is \(K\)-Lipschitz for a common \(K\). Then for all \(\varepsilon {\gt} 0\) and all \(k\), \(N^{\mathrm{ext}}(B(k,F), d_\infty , \varepsilon ) \le N^{\mathrm{ext}}(\overline{\langle F\rangle }, d_\infty , \varepsilon ) {\lt} \infty \).
(P1, hypothesis 2c.) Assume \(\mathcal X\) is compact and the generators are non-expanding: \(\mathrm{lip}\, F \le 1\). Then for all \(\varepsilon {\gt} 0\) and all \(k\), \(N^{\mathrm{ext}}(B(k,F), d_\infty , \varepsilon ) \le N^{\mathrm{ext}}(\overline{\langle F\rangle }, d_\infty , \varepsilon ) {\lt} \infty \).
(P1’: contraction to a compact invariant set.) Assume
(uniform contraction) every \(f \in F\) is \(c\)-Lipschitz with \(0 {\lt} c {\lt} 1\);
(invariant set) there is a nonempty \(F\)-invariant set \(A \subseteq \mathcal X\) (\(f(A) \subseteq A\) for all \(f \in F\));
(bounded absorbing set) there are \(L \in \mathbb N\) and a bounded set \(K \subseteq \mathcal X\) with \(f(\mathcal X) \subseteq K\) for every word \(f \in F^L\) of length \(L\).
Then for every \(\varepsilon {\gt} 0\), with \(m(\varepsilon ) := L + \lceil \log _{1/c}(2\, \mathrm{diam}(K)/\varepsilon ) \rceil \), for every \(k \ge 0\),
In particular the right-hand side is independent of \(k\). (This generalizes the paper, which assumes \(A\) compact, \(K\) compact and \(k \ge m(\varepsilon )\): the proof uses of \(A\) only that it is nonempty and invariant, of \(K\) only its boundedness, and the bound holds for every \(k\). The right-hand side is finite whenever each \(N^{\mathrm{ext}}(F^l, d_\infty , \varepsilon )\), \(l {\lt} m(\varepsilon )\), is finite, which the paper does not assume.)
(P2: nilpotent control grows polynomially.) Let \((H, d_H)\) be a group with a pseudo-emetric, identity \(e\), and assume the length \(g \mapsto d_H(e,g)\) is subadditive: \(d_H(e,gh) \le d_H(e,g) + d_H(e,h)\). Assume its balls have polynomial entropy of degree \(D \ge 0\): there is \(1 \le C_H {\lt} \infty \) such that for all \(R \ge 0\) and \(\delta {\gt} 0\),
Suppose \(H\) acts on \(\mathcal X\) through a homomorphism \(\alpha : H \to \mathcal X^{\mathcal X}\), that there is a bounded set \(S \subseteq H\) (\(d_H(e,s) \le R_S\) for \(s \in S\)) with \(F \subseteq \alpha (S)\), and that the orbit map is Lipschitz in the uniform metric: \(d_\infty (\alpha (g), \alpha (h)) \le L_\alpha d_H(g,h)\) with \(0 {\lt} L_\alpha {\lt} \infty \). Then with
one has, for every \(\varepsilon {\gt} 0\) and every \(k \ge 0\),
(In Lean the ball-entropy hypothesis and the conclusion are stated in \([0,\infty ]\) as \(N \le \operatorname {ofReal}(\cdots )\), which in particular asserts finiteness; the bound holds for \(k = 0\) as well.)
(Real-valued form of P2.) Under the hypotheses of ‘cond:p2-nilp‘, for every \(\varepsilon {\gt} 0\) and \(k \ge 0\), \(N^{\mathrm{ext}}(B(k,F), d_\infty , \varepsilon ) \le C_H \max (1, R_S L_\alpha )^D (1 + k/\varepsilon )^D\) as real numbers (the covering number being finite).
(Saturation for compact equicontinuous semigroups.) Let \((\mathcal X, d)\) be compact and \(F \subseteq \mathcal X^{\mathcal X}\). If the generated semigroup \(\langle F \rangle \) is equicontinuous, then for every \(\varepsilon {\gt} 0\) and every \(k\)
In particular the covering number of the word ball does not grow with the depth \(k\).
(Internal version of ‘cor:aa-semigroup-saturation‘.) Under the same hypotheses, for every \(\varepsilon {\gt} 0\) and every \(k\), \(N\bigl(B(k,F), d_\infty , \varepsilon \bigr) \le N\bigl(\overline{\langle F\rangle }^{d_\infty }, d_\infty , \varepsilon /2\bigr) {\lt} \infty \) (internal covering numbers are not monotone in the set, whence the loss of a factor \(2\) in the radius).
(Balanced value \(O(n^{-1/2})\).) Under the hypotheses of ‘prop:cot-append‘, at the depth \(k^\ast = \lceil \log n / (2\log (1/\theta ))\rceil \) the bound reads
with probability at least \(1 - \delta \): every term is of order \(n^{-1/2}\) (up to the \(\sqrt{\log (1/\delta )}\) factor of the deviation term), so the balanced value is \(O(n^{-1/2})\).
Consequently (via ‘prop:hilbert-sg‘) the sub-Gaussian increment condition holds for the linear readout class \(H_L = \{ x \mapsto \langle w, \Phi _L(x)\rangle : \| w\| \le 1\} \) and every hidden-layer class on \(\mathcal A^{\mathbb N}\), with \(A_H = 1\) and \(L = \sqrt2\, \theta ^{1-L}\).
(Double-exponential regime.) Assume there are \(p, c_-, c_+, \delta _0 {\gt} 0\) with \(c_- \delta ^{-p} \le \log M(G, \delta )\) and \(\log N^{\mathrm{ext}}(G, \delta ) \le c_+ \delta ^{-p}\) for all \(0 {\lt} \delta {\lt} \delta _0\), and that \(N^{\mathrm{ext}}(G, \delta ) {\lt} \infty \) for all \(\delta {\gt} 0\). Then for every fixed \(\varepsilon \) with \(0 {\lt} \varepsilon {\lt} 1/2\) and \(\varepsilon {\lt} \lambda \delta _0/2\) there are \(C_1, C_2 {\gt} 0\) and \(k_0\) such that \(C_1 \lambda ^{pk} \le \log N^{\mathrm{ext}}(W_k, d_\infty , \varepsilon ) \le C_2 \lambda ^{pk}\) for all \(k \ge k_0\).
(Profiles from the envelope, ‘cor:envelope-profiles‘.) Let every \(f \in F\) be \(\Lambda \)-Lipschitz, \(\overline D {\gt} 0\), \(k \ge 1\) and \(D_k(S) \le \overline D\).
(Parametric layers.) If \(N^{\mathrm{ext}}(F, d_\infty , \varepsilon ) {\lt} \infty \) and \(\log N^{\mathrm{ext}}(F, d_\infty , \varepsilon ) \le p\log (C/\varepsilon )\) for all \(0 {\lt} \varepsilon \le C\), with \(p \ge 0\) and \(C \ge \overline D\), then for \(0 {\lt} \varepsilon \le \overline D\), \(\log N^{\mathrm{ext}}(B(k,F), d_\infty , \varepsilon ) \le \log (k+1) + kp[\log (C/\varepsilon ) + \log S_k(\Lambda )]\), and \(\mathsf V_k(S) \le \overline D\bigl(\sqrt{\log (k+1)} + \sqrt{kp\log k} + k\sqrt{p\log \Lambda _+} + \sqrt{kp}\, (\sqrt{\log (2C/\overline D)} + \sqrt\pi /2)\bigr)\).
(Nonparametric layers.) If \(N^{\mathrm{ext}}(F, d_\infty , \varepsilon ) {\lt} \infty \) and \(\log N^{\mathrm{ext}}(F, d_\infty , \varepsilon ) \le c\, \varepsilon ^{-q}\) for all \(\varepsilon {\gt} 0\), with \(c \ge 0\) and \(0 {\lt} q {\lt} 2\), then for \(\varepsilon {\gt} 0\), \(\log N^{\mathrm{ext}}(B(k,F), d_\infty , \varepsilon ) \le \log (k+1) + kc\, S_k(\Lambda )^q\varepsilon ^{-q}\), and \(\mathsf V_k(S) \le \overline D\sqrt{\log (k+1)} + \sqrt{kc}\, (2S_k(\Lambda ))^{q/2}\, \overline D^{\, 1-q/2}/(1-q/2)\).
See ‘cor:envelope-profiles-a‘ and ‘cor:envelope-profiles-b‘ for the factors \(2\).
(Profiles from the envelope: parametric layers, ‘cor:envelope-profiles‘(a).) Let every \(f \in F\) be \(\Lambda \)-Lipschitz, let \(\overline D {\gt} 0\), \(C \ge \overline D\), \(p \ge 0\), and suppose \(N^{\mathrm{ext}}(F, d_\infty , \varepsilon ) {\lt} \infty \) and \(\log N^{\mathrm{ext}}(F, d_\infty , \varepsilon ) \le p\log (C/\varepsilon )\) for all \(0 {\lt} \varepsilon \le C\). If \(D_k(S) \le \overline D\) then, for \(k \ge 1\), with \(\Lambda _+ := \max \{ 1,\Lambda \} \),
(The paper has \(C\) in place of \(2C\): the factor \(2\) is the price of comparing the internal covering number in \(d_S\) defining \(\mathsf V_k(S)\) with the external one in \(d_\infty \), ‘lem:covering-empSpace-le-external-unifMaps‘.)
(Parametric layers, entropy form of ‘cor:envelope-profiles‘(a).) Let every \(f \in F\) be \(\Lambda \)-Lipschitz and suppose \(N^{\mathrm{ext}}(F, d_\infty , \varepsilon ) {\lt} \infty \) and \(\log N^{\mathrm{ext}}(F, d_\infty , \varepsilon ) \le p\log (C/\varepsilon )\) for all \(0 {\lt} \varepsilon \le C\), with \(p \ge 0\). Then for \(k \ge 1\) and \(0 {\lt} \varepsilon \le C\),
(Parametric layers, ‘cor:envelope-profiles‘(a) for the hypothesis \(\log N(F, \varepsilon ) \le p\log (1 + C/\varepsilon )\).) Let every \(f \in F\) be \(\Lambda \)-Lipschitz, let \(\overline D {\gt} 0\), \(C \ge 0\), \(p \ge 0\), and suppose \(N^{\mathrm{ext}}(F, d_\infty , \varepsilon ) {\lt} \infty \) and \(\log N^{\mathrm{ext}}(F, d_\infty , \varepsilon ) \le p\log (1 + C/\varepsilon )\) for all \(\varepsilon {\gt} 0\). If \(D_k(S) \le \overline D\) then, for \(k \ge 1\), with \(\Lambda _+ := \max \{ 1,\Lambda \} \),
(Same proof as ‘cor:envelope-profiles-a‘, with \(1 + 2CS_k/\varepsilon \le (1 + 2C/\overline D)\, S_k\, \overline D/\varepsilon \) for \(\varepsilon \le \overline D\).)
(Profiles from the envelope: nonparametric layers, ‘cor:envelope-profiles‘(b).) Let every \(f \in F\) be \(\Lambda \)-Lipschitz, let \(\overline D {\gt} 0\), \(c \ge 0\), \(0 {\lt} q {\lt} 2\), and suppose \(N^{\mathrm{ext}}(F, d_\infty , \varepsilon ) {\lt} \infty \) and \(\log N^{\mathrm{ext}}(F, d_\infty , \varepsilon ) \le c\, \varepsilon ^{-q}\) for all \(\varepsilon {\gt} 0\). If \(D_k(S) \le \overline D\) then, for \(k \ge 1\),
(The paper has \(S_k(\Lambda )^{q/2}\) in place of \((2S_k(\Lambda ))^{q/2}\): the factor \(2\) is the price of comparing the internal covering number in \(d_S\) with the external one in \(d_\infty \).)
(Nonparametric layers, entropy form of ‘cor:envelope-profiles‘(b).) Let every \(f \in F\) be \(\Lambda \)-Lipschitz and suppose \(N^{\mathrm{ext}}(F, d_\infty , \varepsilon ) {\lt} \infty \) and \(\log N^{\mathrm{ext}}(F, d_\infty , \varepsilon ) \le c\, \varepsilon ^{-q}\) for all \(\varepsilon {\gt} 0\), with \(c \ge 0\). Then for \(k \ge 1\) and \(\varepsilon {\gt} 0\),
(Finite-dimensional feature map.) Let \(\mathcal H = \mathbb R^m\) and \(\Phi : \mathcal X \to \mathbb R^m\). If for each \(j\) the feature matrix \(\Phi _j = (\Phi (f_j(x_i)))_{i \in [n]} \in \mathbb R^{m \times n}\) has full column rank, quantified as \(\sigma _{\min }(\Phi _j)^2 = \lambda _{\min }(\Phi _j^\top \Phi _j) \ge \lambda _{\min ,j} {\gt} 0\), i.e. \(c^\top \Phi _j^\top \Phi _j c \ge \lambda _{\min ,j} \| c\| _2^2\), then ‘prop:linear-interpolation‘ applies with \(\Lambda _j = \lambda _{\min ,j}^{-1/2} = \sigma _{\min }(\Phi _j)^{-1}\). (This is ‘cor:rkhs-readout‘ for \(\mathcal H = \mathbb R^m\), since \(\Phi _j^\top \Phi _j\) is the Gram matrix.)
(Readout realization from interpolation.) Under the hypotheses of ‘prop:linear-interpolation‘, if moreover \(\kappa \, d_S(f_j, f_\ell ) \le \rho \) for all \(j \ne \ell \), then the readout-realization assumption holds for \(H_{R_H}(\Phi )\) on \(B_k = \{ f_1, \dots , f_M\} \) with constants \(\kappa \) and \(R_{\mathrm{out}}\).
Matching depth dependence (i). There is a universal constant \(c {\gt} 0\) such that under Assumption ass:readout-realization-main for \(B_k\) (constants \(\kappa , R_{\mathrm{out}}\)), the boundedness hypothesis of thm:sudakov-type and \(\varepsilon _0 {\gt} 0\): if \(M(B_k, d_S, 2\varepsilon _0) \ge e^{\alpha k}\) for some \(\alpha {\gt} 0\), then \(\hat{\mathfrak R}_S(\mathcal H_k) \ge c\, \kappa \varepsilon _0 \sqrt{\alpha k / n}\) whenever \(n \ge R_{\mathrm{out}}^2 \alpha k / (\kappa ^2 \varepsilon _0^2)\).
Matching depth dependence (ii). There is a universal constant \(c {\gt} 0\) such that under Assumption ass:readout-realization-main for \(B_k\) (constants \(\kappa , R_{\mathrm{out}}\)), the boundedness hypothesis of thm:sudakov-type, \(k \ge 1\) and \(\varepsilon _0 {\gt} 0\): if \(M(B_k, d_S, 2\varepsilon _0) \ge k^{\beta }\) for some \(\beta {\gt} 0\), then \(\hat{\mathfrak R}_S(\mathcal H_k) \ge c\, \kappa \varepsilon _0 \sqrt{\beta \log k / n}\) whenever \(n \ge R_{\mathrm{out}}^2 \beta \log k / (\kappa ^2 \varepsilon _0^2)\).
(Balanced value \(O(n^{-1/2})\).) Under the hypotheses of ‘prop:ode-fixedpoint‘ with \(h\mu {\lt} 1\), at the depth \(k^\ast = \lceil \log n / (2\log (1/(1-h\mu )))\rceil \),
with probability at least \(1 - \delta \): the balanced value is of order \(n^{-1/2}\).
(Balanced value \(O(n^{-1/2})\).) Under the hypotheses of ‘prop:ode-horizon‘ with \(k^\ast = \lceil \sqrt n\rceil \) (and \(T/k^\ast \le h_0\), projection inactive at step size \(T/k^\ast \)),
with probability at least \(1 - \delta \): the balanced value is of order \(n^{-1/2}\).
(Finite hidden-layer classes; table row “E1’, E2 with \(|F| = r\)”.) If \(F\) is finite with \(|F| = r \ge 2\) and the state metric is bounded, \(d(x,y) \le D_{\mathcal X}\) (\(D_{\mathcal X} \ge 0\)), then \(|B(k,F)| \le r^{k+1}\), hence \(N(B(k,F), d_S, \varepsilon ) \le r^{k+1}\) for every \(\varepsilon {\gt} 0\), and
for every \(k\): profile (iii), \(\mathrm{var}(k,n) = O(\sqrt{k/n})\).
(P1 \(\Rightarrow \) saturation; table row “P1 (compact, equicontinuous)”.) Let \(\mathcal X\) be compact and the semigroup \(\langle F\rangle \) equicontinuous. With \(N_\infty (\varepsilon ) := N^{\mathrm{ext}}\bigl(\overline{\langle F\rangle }, d_\infty , \varepsilon /2\bigr)\) and \(\mathsf V_\infty := \int _0^{\mathrm{diam}(\mathcal X)} \sqrt{\log N_\infty (\varepsilon )}\, d\varepsilon {\lt} \infty \) (assumed interval-integrable), \(D_k(S) \le \mathrm{diam}(\mathcal X)\) and \(\mathsf V_k(S) \le \mathsf V_\infty \) for all \(k\): profile (i), \(\mathsf V_k = O(1)\), \(\mathrm{var}(k,n) = O(n^{-1/2})\).
(P1 \(\Rightarrow \) saturation; general form.) Let the state metric be bounded, \(d(x,y) \le D_{\mathcal X}\) (\(D_{\mathcal X} \ge 0\)), and let the semigroup \(\langle F\rangle \) be totally bounded in \(d_\infty \) (‘cond:p1‘). Put \(N_\infty (\varepsilon ) := N^{\mathrm{ext}}\bigl(\overline{\langle F\rangle }, d_\infty , \varepsilon /2\bigr)\) (finite for \(\varepsilon {\gt} 0\); the factor \(\tfrac 12\) comes from comparing internal with external covers) and assume \(\mathsf V_\infty := \int _0^{D_{\mathcal X}} \sqrt{\log N_\infty (\varepsilon )}\, d\varepsilon {\lt} \infty \) (interval-integrability of the majorant). Then for every \(k\) and every sample \(S\), \(\mathsf V_k(S) \le \mathsf V_\infty \): profile (i), \(\mathrm{var}(k,n) = O(n^{-1/2})\).
(P2 on a bounded state space; table row “P2, compact \(\mathcal X\)”.) Under the hypotheses of ‘cond:p2-nilp‘, if moreover \(d(x,y) \le D_{\mathcal X}\) for all \(x, y\) (\(D_{\mathcal X} {\gt} 0\)), then \(D_k(S) \le D_{\mathcal X}\) and for every \(k \ge 1\)
with \(C_0 = 2^D C_H\max (1, R_SL_\alpha )^D\): profile (ii), \(\mathrm{var}(k,n) = O(\sqrt{D\log k/n})\).
(P2 on a general state space; table row “P2, non-compact \(\mathcal X\)”.) Under the hypotheses of ‘cond:p2-nilp‘ with \(R_S {\gt} 0\), \(D_k(S) \le 2L_\alpha R_S k\) and for every \(k \ge 1\)
with \(C_0 = 2^D C_H\max (1, R_SL_\alpha )^D\): profile (iv), \(\mathrm{var}(k,n) = O(k\sqrt{D/n})\).
The sup-norm theorem from the sample theorem. Under the sup-norm Assumptions ass:ent-readout and ass:ent-transition with \(L_H {\gt} 0\), \(B_H {\gt} 0\), if \(x \mapsto \mathcal E_H(x/4) + \mathcal E_F(x/(8L_H))\) is integrable on \([0, B_H/2]\), then for every sample \(S\) of size \(n \ge 1\),
This is thm:rad.decomp.ent.ent with the scale \(x/(8L_H)\) in place of \(x/(4L_H)\) for \(F\) (the price of the internal cover of \(F\) in the sample version), obtained from thm:rad.decomp.ent.ent-sample-foml: the sup-norm assumptions imply the sample assumptions, and \(\mathcal E_{H,S} \le \mathcal E_H\), \(\mathcal E_{F,S} \le \mathcal E_F\) pointwise at positive scales.
(RKHS / kernel output layer.) Let \(\Phi : \mathcal X \to \mathcal H\) be a feature map (e.g. the canonical feature map of a positive definite kernel \(K\), so that \(\langle \Phi (a), \Phi (a')\rangle = K(a, a')\)) and fix \(f_1, \dots , f_M\) and a sample \(S\). Suppose that for each \(j\) the Gram matrix \(G_j = (\langle \Phi (f_j(x_i)), \Phi (f_j(x_{i'})) \rangle )_{i,i'}\) satisfies \(c^\top G_j c \ge \lambda _{\min ,j} \| c\| _2^2\) with \(\lambda _{\min ,j} {\gt} 0\). Then ‘prop:linear-interpolation‘ applies with \(\Lambda _j = \lambda _{\min ,j}^{-1/2}\): for all codes \(u^{(j)}\) with \(\lambda _{\min ,j}^{-1/2} \| u^{(j)}\| _2 \le R_H\), \(\rho \)-separated and \(R_{\mathrm{out}}\)-bounded, there are \(h_j \in H_{R_H}(\Phi )\) with \(h_j(f_j(x_i)) = u^{(j)}_i\), \(\| h_j \circ f_j - h_\ell \circ f_\ell \| _S \ge \rho \) (\(j \ne \ell \)) and \(\| h_j \circ f_j\| _{S,\infty } \le R_{\mathrm{out}}\).
Rates under fixed-scale packing lower bounds (exponential). There is a universal constant \(c {\gt} 0\) such that under Assumption ass:readout-realization-main for \(B_k\): if \(M(B_k, d_S, 2\varepsilon _0) \ge e^{\alpha k}\) for some \(\varepsilon _0, \alpha {\gt} 0\), then
Rates under fixed-scale packing lower bounds (polynomial). There is a universal constant \(c {\gt} 0\) such that under Assumption ass:readout-realization-main for \(B_k\): if \(k \ge 1\) and \(M(B_k, d_S, 2\varepsilon _0) \ge k^{\beta }\) for some \(\varepsilon _0, \beta {\gt} 0\), then
Conditional Sudakov-type lower bound for a bounded hypothesis class. Under the hypotheses of thm:sudakov-type, the boundedness hypothesis holds as soon as \(\mathcal H_k\) is uniformly bounded on the sample, \(|g(x_i)| \le C\) for all \(g \in \mathcal H_k\) and \(i \le n\); hence the same lower bound.
(Super-exponential regime.) Assume there are \(c_-, c_+, \delta _0 {\gt} 0\) with \(c_- \log (1/\delta ) \le \log M(G, \delta )\) and \(\log N^{\mathrm{ext}}(G, \delta ) \le c_+ \log (1/\delta )\) for all \(0 {\lt} \delta {\lt} \delta _0\), and that \(N^{\mathrm{ext}}(G, \delta ) {\lt} \infty \) for all \(\delta {\gt} 0\) (automatic for compact \(G\)). Then for every fixed \(\varepsilon \) with \(0 {\lt} \varepsilon {\lt} 1/2\) and \(\varepsilon {\lt} \lambda \delta _0/2\) there are \(C_1, C_2 {\gt} 0\) and \(k_0\) such that \(C_1 k^2 \le \log N^{\mathrm{ext}}(W_k, d_\infty , \varepsilon ) \le C_2 k^2\) for all \(k \ge k_0\). (The smallness condition \(\varepsilon {\lt} \lambda \delta _0/2\) makes every radius \(2\varepsilon /\lambda ^{j+1} \le 2\varepsilon /\lambda \) fall below \(\delta _0\); it cannot be replaced by "large \(k\)" since the \(j = 0\) radius does not shrink with \(k\).)
(Plugging a profile into the decomposition.) Under the hypotheses of ‘thm:hidden-decomp-depth‘ (including the integrability of the entropy integrand), if \(\mathsf V_k(S) \le B\) then \(\hat{\mathfrak R}_S(\mathcal H_k) \le \hat{\mathfrak R}_S(H) + \frac{12A_HL}{\sqrt n}\, B\). The statements about \(\mathrm{var}(k,n)\) in ‘prop:profiles‘ follow by multiplying the profile bounds with \(12A_HL/\sqrt n\).
(Estimation term for finite hidden-layer classes.) Under the hypotheses of ‘cor:profile-finite‘ and ‘thm:hidden-decomp-depth‘, \(\hat{\mathfrak R}_S(\mathcal H_k) \le \hat{\mathfrak R}_S(H) + \frac{12A_HL}{\sqrt n}\, D_{\mathcal X}\sqrt{(k+1)\log r}\) for every \(k\): \(\mathrm{var}(k,n) = O(\sqrt{k/n})\).
(Estimation term under P1.) Let \(\mathcal X\) be compact, the semigroup \(\langle F\rangle \) equicontinuous, and \(\mathsf V_\infty \) as in ‘cor:profile-p1‘. Under the hypotheses of ‘thm:hidden-decomp-depth‘, \(\hat{\mathfrak R}_S(\mathcal H_k) \le \hat{\mathfrak R}_S(H) + \frac{12A_HL}{\sqrt n} \mathsf V_\infty \) for every \(k\): \(\mathrm{var}(k,n) = O(n^{-1/2})\) uniformly in the depth.
(Estimation term under P2, bounded state space.) Under the hypotheses of ‘cor:profile-p2-compact‘ and ‘thm:hidden-decomp-depth‘, for \(k \ge 1\), \(\hat{\mathfrak R}_S(\mathcal H_k) \le \hat{\mathfrak R}_S(H) + \frac{12A_HL}{\sqrt n}\Bigl( D_{\mathcal X}\sqrt D\bigl(\sqrt{\log (1 + k/D_{\mathcal X})} + \tfrac {\sqrt\pi }{2}\bigr) + D_{\mathcal X}\sqrt{\log C_0}\Bigr)\): \(\mathrm{var}(k,n) = O(\sqrt{D\log k/n})\).
(Estimation term under P2, general state space.) Under the hypotheses of ‘cor:profile-p2-noncompact‘ and ‘thm:hidden-decomp-depth‘, for \(k \ge 1\), \(\hat{\mathfrak R}_S(\mathcal H_k) \le \hat{\mathfrak R}_S(H) + \frac{12A_HL}{\sqrt n}\, 2L_\alpha R_Sk\Bigl(\sqrt D\bigl(\sqrt{\log (1 + 1/(2L_\alpha R_S))} + \tfrac {\sqrt\pi }{2}\bigr) + \sqrt{\log C_0}\Bigr)\): \(\mathrm{var}(k,n) = O(k\sqrt{D/n})\).
The estimation-plus-deviation term of ‘thm:bv‘ (with the constants of ‘thm:bv-general‘), as a function of an upper bound \(B\) for \(\hat{\mathfrak R}_S(\mathcal H)\):
(All regime propositions below are stated in terms of this quantity, so that the constants of ‘thm:bv‘ enter in one place only.)
For a list of layers \(u = [f_m, \dots , f_1]\) (an explicit representation of a word) the composed map is \(f_u = f_m \circ \cdots \circ f_1\), with \(f_{[]} = \mathrm{id}\). In Lean: ‘compList [] = id‘ and ‘compList (f :: u) = f ∘ compList u‘ (the head of the list acts last, matching the recursion \(B(k+1,F) = B(k,F) \cup F \circ B(k,F)\)).
The success cylinder of a program \(u = (u_1, \dots , u_j)\): the set \([u]\) of inputs whose first \(j\) symbols are \(u_1, \dots , u_j\). In Lean, membership is defined by recursion on \(u\): \(x \in [\, ]\) always, and \(x \in [a :: u]\) iff \(x_0 = a\) and \(\sigma (x) \in [u]\).
The scratchpad space is \(\mathcal X = \mathcal A^{\mathbb N} = \{ x = (x_0, x_1, \dots ) : x_j \in \mathcal A\} \), the space of infinite symbol sequences over the alphabet \(\mathcal A\) (a type synonym of \(\mathbb N \to \mathcal A\) carrying the parameter \(\theta \) of the metric below).
For \(u = (u_0, \dots , u_{k-1}) \in E^k\) the word \(w_u = A \circ g_{u_{k-1}} \circ A \circ g_{u_{k-2}} \circ \cdots \circ A \circ g_{u_0} \circ r\) (so \(u_0\) is written first and \(u_{k-1}\) last). In Lean it is defined by recursion on \(k\): \(w_{()} = r\) and \(w_{(u, a)} = A \circ g_a \circ w_u\).
The pseudometric space \((\mathbb R^{\mathcal X}, \| \cdot \| _S)\) of real-valued functions equipped with the empirical \(L^2\) distance \(\| u - v\| _S = \bigl(\frac1n \sum _i |u(x_i) - v(x_i)|^2\bigr)^{1/2}\) of the sample \(S\) (a type synonym of \(\mathbb R^{\mathcal X}\)).
The empirical Rademacher complexity of \(G \subseteq \mathbb R^{\mathcal X}\) on the sample \(S = (x_1,\dots ,x_n)\) is \(\hat{\mathfrak R}_S(G) = \mathbb E_\sigma \sup _{g \in G} \frac1n \sum _{i=1}^n \sigma _i g(x_i) = 2^{-n} \sum _{\sigma \in \{ \pm 1\} ^n} \sup _{g \in G} \frac1n \sum _{i=1}^n \sigma _i g(x_i)\) (no absolute value). In Lean this is FoML’s ‘empiricalRademacherComplexity_without_abs‘ for the class indexed by \(G\) itself.
The entropy integral of \(A\) up to scale \(D\) is \(\mathsf V(D, A) = \int _0^{D} \sqrt{\log N(A, \varepsilon )}\, d\varepsilon \). In the paper, \(\mathsf V_k(S) = \int _0^{D_k(S)} \sqrt{\log N(B(k,F), d_S, \varepsilon )}\, d\varepsilon \) with \(D_k(S) = \operatorname {diam}_S B(k,F)\) is \(\mathsf V(\operatorname {diam}_S B(k,F), B(k,F))\) for the empirical pseudometric \(d_S\).
A vector field \(s : K \to \mathbb R^d\) is \(\mu \)-strongly monotone and \(\Lambda \)-co-coercive (the properties of \(s = \nabla \phi \) for \(\phi \) \(\mu \)-strongly concave and \(\Lambda \)-smooth) if for all \(x, y \in K\)
(The first is the gradient characterization of strong concavity; the second is the standard consequence of strong concavity and smoothness. We take both as the definition.)
The depth-dependent part of the excess-risk bound of theorem 98: \(\mathsf{gen}(k,n) = \mathsf{bias}(k) + \mathsf{var}(k,n)\).
For a real Hilbert space \(\mathcal H\), a feature map \(\Phi : \mathcal X \to \mathcal H\) and a radius \(R_H \ge 0\), the norm-bounded linear output-layer class is
The class of equal-step schemes of at most \(k\) steps, \(E_T(k) := \{ \Phi _{s,m} : s \in \mathcal S_T,\ 0 \le m \le k\} \), where \(\Phi _{s,m} = T_{s,T/m,\tau _m} \circ \cdots \circ T_{s,T/m,\tau _1}\) with \(\tau _i = (i-1)T/m\) is the \(m\)-step equal-step Euler scheme and \(\Phi _{s,0} = \mathrm{id}\).
A map \(\Pi _K : \mathbb R^d \to \mathbb R^d\) is a (Euclidean) projection onto \(K\) if \(\Pi _K(x) \in K\) for all \(x\), \(\Pi _K(x) = x\) for \(x \in K\), and \(\Pi _K\) is \(1\)-Lipschitz. (For a nonempty closed convex \(K\) the nearest-point map has these properties; Mathlib provides only the existence of nearest points, so we take the projection and its properties as hypotheses.)
A scheme of \(n\) steps with a single drift \(s\), step sizes \(h = (h_0, \dots , h_{n-1})\) and time stamps \(\tau = (\tau _0, \dots , \tau _{n-1})\) is the composition \(T_{s,h_{n-1},\tau _{n-1}} \circ \cdots \circ T_{s,h_0,\tau _0}\) (the empty scheme is \(\mathrm{id}\)).
The class \(B_T(k)\) of schemes of at most \(k\) steps with a single drift \(s \in \mathcal S_T\), step sizes \(h_i \in [0,h_0]\), time stamps \(\tau _i \in [0,T]\) and total time \(\sum _i h_i \le T\) (it contains \(\mathrm{id}\), the empty scheme). The consistency condition \(\tau _i = \sum _{j{\lt}i}h_j\) of the paper is dropped; the bounds below hold for this larger class.
The uniform pushed-forward covering number of the output-layer class: \(N_{S,F}(H, u) = \sup _{f \in F} N^{\mathrm{ext}}(H, \| \cdot \| _{f \circ S}, u)\), the largest external covering number of \(H\) in the empirical metric of a pushed-forward sample \(f \circ S = (f(x_1), \dots , f(x_n))\), \(f \in F\). (Any uniform bound \(N^{\mathrm{ext}}(H, \| \cdot \| _{f \circ S}, u) \le \bar N(u)\) for all \(f \in F\) dominates it.)
The (population) Rademacher complexity is \(\mathfrak R_n(G) = \mathbb E_{S \sim P^{\otimes n}} \hat{\mathfrak R}_S(G)\). In Lean we take FoML’s ‘rademacherComplexity‘, i.e. the expectation over \(S \sim P^{\otimes n}\) of the absolute version \(\mathbb E_\sigma \sup _{g \in G} \bigl|\frac1n \sum _i \sigma _i g(x_i)\bigr| \ge \hat{\mathfrak R}_S(G)\); this is the quantity for which FoML’s deviation bounds are stated, and it dominates the paper’s \(\mathfrak R_n(G)\).
The covering constant of one ReLU block: \(C_F := 2 L_F \max \{ \beta _W, \beta \} = 2(2\beta _W R_K + \beta _W + \beta + 1) \max \{ \beta _W,\beta \} \). (The paper’s \(C_F = 8\sqrt{\max \{ m,w\} }\max \{ \beta _W,1\} \max \{ \beta _W R_K + \beta , \beta _W, 1\} \) has the same shape; the factor \(\sqrt{\max \{ m,w\} }\) comes from bounding the operator norm by the Frobenius norm, which the volumetric argument in the operator norm avoids.)
The first expand-and-reset map of the computed illustration (‘sec:relu-computed‘), \(\eta = 1/32\): the piecewise-linear map with breakpoints \((0, 3/8), (\eta , 0), (1/4 - \eta , 1), (1/4, 3/8), (1, 3/8)\), i.e. \(g_0(x) = 3/8 - 12x\) on \([0, \eta ]\), \(= \tfrac {16}{3}(x - \eta )\) on \([\eta , 1/4 - \eta ]\), \(= 1 - 20(x - 1/4 + \eta )\) on \([1/4 - \eta , 1/4]\) and \(= 3/8\) on \([1/4, 1]\).
The second expand-and-reset map, with breakpoints \((0, 5/8), (3/4, 5/8), (3/4 + \eta , 0), (1 - \eta , 1), (1, 5/8)\), i.e. \(g_1(x) = 5/8\) on \([0, 3/4]\), \(= 5/8 - 20(x - 3/4)\) on \([3/4, 3/4 + \eta ]\), \(= \tfrac {16}{3}(x - 3/4 - \eta )\) on \([3/4 + \eta , 1 - \eta ]\) and \(= 1 - 12(x - 1 + \eta )\) on \([1 - \eta , 1]\).
A one-hidden-layer ReLU layer on \(\mathbb R^d\) with hidden index set \(m\) (width \(|m|\)) is \(x \mapsto W_2\, \mathrm{relu}(W_1 x) + b\), with \(W_1 \in \mathbb R^{m \times d}\), \(W_2 \in \mathbb R^{d \times m}\), \(b \in \mathbb R^d\) and \(\mathrm{relu}\) applied coordinatewise.
The parameter space of a ReLU block of width \(w\) on \(\mathbb R^m\) (\(=E\)): \(\vartheta = (W, b, V, c) \in (\mathbb R^m \to \mathbb R^w) \times \mathbb R^w \times (\mathbb R^w \to \mathbb R^m) \times \mathbb R^m\), with the operator norm on the linear maps and the sup (product) norm on the tuple.
The \(k\)-independent majorant of case (i) at the scale of the empirical metric: \(N_\infty (\varepsilon ) := N^{\mathrm{ext}}(K, \varepsilon /4) + m(\varepsilon /2)\, N^{\mathrm{ext}}(F_\Lambda , d_\infty , (1-\Lambda )\varepsilon /2)^{m(\varepsilon /2)}\) (the factor \(2\) from \(N(A, d_S, \varepsilon ) \le N^{\mathrm{ext}}(A, d_\infty , \varepsilon /2)\)).
The parameters of a one-dimensional ReLU block with \(w\) units of input weights \(u_i\), thresholds \(t_i\), output weights \(v_i\) and bias \(c\): \(W = (u_i)_i\), \(b = (-t_i)_i\), \(V = (v_i)_i\) (as a row), so that \(V\, \mathrm{relu}(Wx + b) + c = \sum _i v_i\, \mathrm{relu}(u_i x - t_i) + c\).
A class \(\mathcal H \subseteq \mathbb R^{\mathcal X}\) is (sup-norm) separable if it has a countable subset \(\mathcal D \subseteq \mathcal H\) which is dense for the uniform norm: for every \(f \in \mathcal H\) and \(\varepsilon {\gt} 0\) there is \(g \in \mathcal D\) with \(\sup _x |f(x) - g(x)| \le \varepsilon \).
Let \(F \subseteq \mathcal X^{\mathcal X}\) be a hidden-layer class. The depth-\(k\) hidden class is the word ball \(B(k,F) = \{ f_m \circ \cdots \circ f_1 : 0 \le m \le k,\ f_i \in F\} \), with \(B(0,F) = \{ \mathrm{id}\} \). It is defined recursively by \(B(k+1,F) = B(k,F) \cup F \circ B(k,F)\).
For a finite family \(f = (f_1, \dots , f_r)\) of self-maps and a word \(u = (i_1, \dots , i_k) \in [r]^k\), the associated map is \(f_u = f_{i_k} \circ \cdots \circ f_{i_1}\) (the first letter acts first); \(f_{\emptyset } = \mathrm{id}\). In Lean: ‘wordOf f [] = id‘ and ‘wordOf f (i :: u) = wordOf f u ∘ f i‘.
(Exact realization of an affine map.) For \(A \in \mathbb R^{d \times d}\) and \(b \in \mathbb R^d\), with \(W_1 = [I_d; -I_d] \in \mathbb R^{2d \times d}\) and \(W_2 = [A, -A] \in \mathbb R^{d \times 2d}\),
i.e. \(x \mapsto Ax + b\) is one ReLU layer of width \(2d\).
(Approximation transfer, ‘eq:approx-transfer‘.) Let \(\ell \) be \(\beta _\ell \)-Lipschitz in its first argument, \(\mathcal H\) and \(\mathcal C \ne \emptyset \) classes of measurable functions, and \(B \in \mathbb R\). If every \(c \in \mathcal C\) is uniformly approximable from \(\mathcal H\) within \(B + \varepsilon \) for every \(\varepsilon {\gt} 0\) (in particular if \(\sup _{c \in \mathcal C}\inf _{f \in \mathcal H} \| f - c\| _\infty \le B\)), then
Balancing principle. Let \(S \subseteq \mathbb R\) be a set of depths on which \(\mathsf{bias}\ge 0\) is nonincreasing and \(\mathsf{var}(\cdot ,n) \ge 0\) is nondecreasing. If \(k_0 \in S\) balances the two terms, \(\mathsf{bias}(k_0) = \mathsf{var}(k_0,n)\), then \(\mathsf{gen}(k,n) \ge \mathsf{bias}(k_0)\) for every \(k \in S\), while \(\mathsf{gen}(k_0,n) = 2\mathsf{bias}(k_0)\); hence \(k_0\) minimizes the bound over \(S\) up to the factor \(2\).
Deterministic excess-risk bound. Fix a sample \(\mathcal D\) and suppose \(|L[f] - \hat L[f]| \le \Delta _0\) for every \(f \in \mathcal H\). If \(d_T(\iota f, f) \le \varepsilon _{\mathrm{imp}}\) and \(|L[\iota f] - L[f]| \le \beta _L\, d_T(\iota f, f)\) for all \(f \in \mathcal H\) (\(\beta _L \ge 0\)), then every \(\eta \)-empirical minimizer \(\hat f \in \mathcal H\) satisfies, with \(\hat h = \iota (\hat f)\), \(L[\hat h] - \inf _{\mathcal C} L \le \beta _L \varepsilon _{\mathrm{imp}} + \varepsilon _{\mathrm{model}} + \eta + 2\Delta _0\).
Deterministic gap bound. Fix a sample \(\mathcal D\) and suppose \(|L[f] - \hat L[f]| \le \Delta _0\) for every \(f \in \mathcal H\). Under the implementation hypotheses (\(d_T(\iota f, f) \le \varepsilon _{\mathrm{imp}}\), \(|L[\iota f] - L[f]| \le \beta _L d_T(\iota f,f)\), \(|\hat L[\iota f] - \hat L[f]| \le \beta _{\hat L} d_T(\iota f, f)\)), every \(f \in \mathcal H\) satisfies \(L[\iota f] - \hat L[\iota f] \le (\beta _L + \beta _{\hat L}) \varepsilon _{\mathrm{imp}} + \Delta _0\).
One-sided Rademacher complexity of the loss class. Let \(\mathcal H \ne \emptyset \) be pointwise bounded and \(\ell : \mathbb R \times \mathcal Y \to \mathbb R\) be \(\beta _\ell \)-Lipschitz in its first argument (\(\beta _\ell \ge 0\)). Then on every sample \(\mathcal D = ((x_i,y_i))_i\) with \(S = (x_i)_i\),
where \(\hat{\mathfrak R}\) is the one-sided (no absolute value) empirical Rademacher complexity. Proof: the one-sided contraction lem:contraction-without-abs applied to \(\psi (z,u) = \pm \ell (u,y)\), which is \(\beta _\ell \)-Lipschitz in \(u\) (no vanishing condition at \(u = 0\) is needed).
Uniform deviation for a sup-separable Lipschitz-loss class. Let \(n \ge 1\), \(\mathcal H\) be a sup-norm separable, pointwise bounded class of measurable functions, \(\ell : \mathbb R \times \mathcal Y \to [0,b]\) measurable and \(\beta _\ell \)-Lipschitz in its first argument (\(b {\gt} 0\), \(\beta _\ell \ge 0\)), \(\delta \in (0,1)\). Then with probability at least \(1 - \delta \) over \(\mathcal D \sim P^{\otimes n}\),
Proof: apply the two-sided bound lem:two-sided-tail-empirical (one-sided symmetrization for \(\ell \circ \mathcal H\) and \(-\ell \circ \mathcal H\)) with \(C(\mathcal D) = \beta _\ell \hat{\mathfrak R}_S(\mathcal H)\) from lem:bv-loss-rademacher and \(\varepsilon = b\sqrt{2\log (4/\delta )/n}\), so that \(4\exp (-n\varepsilon ^2/(2b^2)) = \delta \); the topology on \(\mathcal H\) is the uniform-convergence topology, which is separable and first countable by lem:sup-dense-separableSpace, and evaluations are continuous.
Composition of covers. Let every \(h \in H\) be \(L_H\)-Lipschitz, let \(C_H\) be an \(r_1\)-cover of \(H\) for \(\| \cdot \| _\infty \) and \(C_F\) an \(r_2\)-cover of \(F\) for \(d_\infty \) (centres anywhere). Then \(\{ h_a \circ f_b : h_a \in C_H, f_b \in C_F\} \) is an \((r_1 + L_H r_2)\)-cover of \(H \circ F\) for \(\| \cdot \| _\infty \): \(|h(f(x)) - h_a(f_b(x))| \le |h(f(x)) - h(f_b(x))| + |h(f_b(x)) - h_a(f_b(x))| \le L_H d_\infty (f, f_b) + \| h - h_a\| _\infty \).
Composition of covers on the sample. Let every \(h \in H\) be \(L_H\)-Lipschitz on \(\mathcal R_S(F)\), let \(C_F \subseteq F\) be an \(r_2\)-cover of \(F\) for \(d_S\) with centres in \(F\), and for every \(f_b \in C_F\) let \(C_H(f_b)\) be an \(r_1\)-cover of \(H\) for \(\| \cdot \| _{f_b \circ S}\) (centres anywhere). Then \(\{ h_a \circ f_b : f_b \in C_F,\ h_a \in C_H(f_b)\} \) is an \((r_1 + L_H r_2)\)-cover of \(H \circ F\) for \(\| \cdot \| _S\): \(\| h \circ f - h_a \circ f_b\| _S \le \| h \circ f - h \circ f_b\| _S + \| h \circ f_b - h_a \circ f_b\| _S \le L_H d_S(f, f_b) + \| h - h_a\| _{f_b \circ S}\), where the first estimate uses the Lipschitz property of \(h\) at the reachable points \(f(x_i), f_b(x_i)\) (this is where \(f_b \in F\) is needed).
(Sup-norm separability of \(H_R(\Phi ) \circ \mathfrak F\).) Let \(\mathcal H\) be a separable real inner product space, \(\Phi : \mathcal X \to \mathcal H\) be \(L_\Phi \)-Lipschitz with \(\| \Phi \| \le M\), \(R \ge 0\), and let \(\mathfrak F\) have a countable uniformly dense subset. Then \(H_R(\Phi ) \circ \mathfrak F\) is sup-norm separable: the countable set \(\{ \langle w, \Phi (g(\cdot ))\rangle : w \in W,\ g \in D\} \), with \(W\) a countable dense subset of the ball of radius \(R\), is uniformly dense, since \(|\langle w, \Phi (f x)\rangle - \langle w', \Phi (g x)\rangle | \le \| w - w'\| M + R L_\Phi \, d(f x, g x)\).
(Bias for append-only steps.) Let \(\Phi \) be \(L_\Phi \)-Lipschitz and \(R \ge 0\). Against the teacher–student class \(\mathcal C = \overline{H_R(\Phi ) \circ \langle F_{\rm w}\rangle }^{\, d_\infty }\),
for a teacher \(c = h_w \circ g_u \circ g_v\) with \(|u| = k\), \(|c(x) - h_w(g_u(x))| \le R L_\Phi \, d_\theta (g_u(g_v x), g_u x) \le R L_\Phi \theta ^k\) since \(\mathrm{lip}(g_u) = \theta ^k\) and \(\mathrm{diam} = 1\), and the same bound holds on the uniform closure; conclude with ‘lem:approx-transfer‘.
(The short words cover the word ball.) For every \(\varepsilon {\gt} 0\) and \(k\), the word ball \(B(\min \{ k, \ell (\varepsilon )\} , F_{\rm w}) \subseteq B(k, F_{\rm w})\) is an \(\varepsilon \)-cover of \(B(k,F_{\rm w})\) in \(d_\infty \): a word \(g_u\) with \(|u| {\gt} \ell \) is within \(\theta ^{\ell } \le \varepsilon \) of the word \(g_{u'}\) formed by its last \(\ell \) letters.
For \(0 {\lt} \varepsilon \le 1\) and \(m \ge 2\), \(\sqrt{\log N(B(k,F_{\rm w}), d_S, \varepsilon )} \le \sqrt{\log m}\Bigl(\frac{\sqrt{\log (1/\varepsilon )}}{\sqrt{\log (1/\theta )}} + \sqrt2\Bigr)\), using \(\ell (\varepsilon ) + 1 \le \log (1/\varepsilon )/\log (1/\theta ) + 2\) and \(\sqrt{a+b} \le \sqrt a + \sqrt b\).
(Saturation for append-only steps.) Let \(|\mathcal A| = m \ge 2\). For \(\varepsilon {\gt} 0\) let \(\ell (\varepsilon ) := \lceil \log (1/\varepsilon )/\log (1/\theta )\rceil \). Then for all \(k \ge 0\),
(Saturated variance profile for append-only steps.) Let \(|\mathcal A| = m \ge 2\). For every sample \(S\) and every \(k \ge 0\),
(Case (i) of ‘prop:profiles‘: the entropy integral does not depend on the depth \(k\).)
(Estimation term for append-only steps.) Let \(|\mathcal A| = m \ge 2\), \(\Phi \) be \(L_\Phi \)-Lipschitz with \(\| \Phi \| \le M_\Phi \) and \(R {\gt} 0\). Then for every sample \(S\) of size \(n \ge 1\) and every depth \(k\),
(‘thm:hidden-decomp-depth‘ with \(A\_ H = 1\), ‘lem:cot-append-profile‘ and ‘lem:cot-readout-rademacher‘).
(Exponential growth for branching steps.) For \(r \ge 2\), every \(k \ge 0\) and every \(\varepsilon {\lt} 1/2\),
The lower bound is the ping–pong bound ‘cond:e2-pingpong‘; the upper bound holds because \(F_{\rm b}\) is finite, so \(|B(k,F_{\rm b})| \le \sum _{j \le k} r^j \le r^{k+1}\).
(The branching family is a ping–pong family, ‘ex:e2-pingpong-subshift‘.) With chambers \(U_a = V_a = \{ x : x_0 = a\} \), anchors \(\bar a\), marker \(\bar\bullet \) and \(\alpha = 1\), \(F_{\rm b}\) satisfies the hypotheses of ‘cond:e2-pingpong‘ (the chambers are disjoint; they are in fact \(1\)-separated, which the paper’s form of E2 requires). Consequently, for \(r \ge 2\), \(N^{\mathrm{ext}}(B(k,F_{\rm b}), d_\infty , \varepsilon ) \ge r^k\) for every \(k\) and every \(\varepsilon {\lt} 1/2\), and \(u \mapsto f_u\) is injective on \([r]^k\).
(Empirical saturation for branching steps.) For every sample \(S = (X_1, \dots , X_n)\), every \(k \ge 0\) and every \(\varepsilon {\gt} 0\) (the paper states \(\varepsilon \in (0,1]\); the bound holds for all \(\varepsilon {\gt} 0\)),
Centres: the identity, the \(r\) constant maps \(\bar a\), and for each \(1 \le j \le k\) the programs \(u\) of length \(j\) with \(\hat P_n([u]) {\gt} \varepsilon ^2\) (fewer than \(\varepsilon ^{-2}\) of them, at most \(n\), at most \(r^j\)); every other program of length \(j\) is within \(d_S\)-distance \(\varepsilon \) of the constant map \(\bar u_j\).
(Empirical saturation for branching steps, ‘lem:cot-branch-sample‘.) For every sample \(S = (X_1,\dots ,X_n)\), every \(k \ge 0\) and every \(\varepsilon {\gt} 0\) (the paper states \(\varepsilon \in (0,1]\)),
and consequently \(\mathsf V_k(S) \le \sqrt{\log (k+1)} + \sqrt{\log (1+r)} + \sqrt{2\log 2} + \sqrt{\pi /2}\), the root-logarithmic profile, for every sample and without any assumption on the input distribution.
(Root-logarithmic empirical entropy integral for branching steps.) For every sample \(S\) and every \(k\),
(the paper’s bound without the term \(\sqrt{2\log 2}\): this additive constant is the price of comparing the internal covering number in \(d_S\), which defines \(\mathsf V_k(S)\), with the external one at half the scale, \(N(B, d_S, \varepsilon ) \le N^{\mathrm{ext}}(B, d_S, \varepsilon /2) \le 1 + r + 4k\varepsilon ^{-2}\)).
(Sup-norm separability.) For a separable feature space, a Lipschitz feature map \(\Phi \) with \(\| \Phi \| \le M_\Phi \) and \(R \ge 0\), the class \(\mathcal H_k = H_R(\Phi ) \circ B(k, F_{\rm w})\) is sup-norm separable (\(B(k,F_{\rm w})\) is finite and the ball of radius \(R\) has a countable dense subset).
(Window output features.) \(\Phi _L\) is \(\sqrt2\, \theta ^{1-L}\)-Lipschitz with respect to \(d_\theta \), and \(\| \Phi _L\| \le 1\). (If \(d_\theta (x,y) \le \theta ^L\) the first \(L\) symbols agree and \(\Phi _L(x) = \Phi _L(y)\); otherwise \(d_\theta (x,y) \ge \theta ^{L-1}\) and \(\| \Phi _L(x) - \Phi _L(y)\| = \sqrt2 \le \sqrt2\, \theta ^{1-L} d_\theta (x,y)\).)
(Programs are close to constants on the sample.) For every program \(u\) (with last letter \(u_j\), or any \(b\) if \(u\) is empty),
since \(f_u = \bar u_j\) off \([u]\) and \(\mathrm{diam}(\mathcal A^{\mathbb N}) = 1\).
The bridge used by all profiles: for every \(A \subseteq \mathcal X^{\mathcal X}\) and \(\varepsilon \ge 0\), \(N(A, d_S, \varepsilon ) \le N^{\mathrm{ext}}(A, d_\infty , \varepsilon /2)\) (internal covering number on the left, external on the right; the factor \(2\) is the price of comparing internal with external covers).
For every \(A \subseteq \mathcal X^{\mathcal X}\) and \(\varepsilon \ge 0\), \(N^{\mathrm{ext}}(A, d_S, \varepsilon ) \le N^{\mathrm{ext}}(A, d_\infty , \varepsilon )\): covering numbers in the empirical metric are dominated by those in the uniform metric, so all hypotheses may be verified in \(d_\infty \).
(Upper bound in sum form.) Suppose \(N^{\mathrm{ext}}(G,\delta ) {\lt} \infty \) for all \(\delta {\gt} 0\), \(G \ne \emptyset \), and \(\log N^{\mathrm{ext}}(G, \delta ) \le \varphi (\delta )\) for \(0 {\lt} \delta {\lt} \delta _0\). If \(0 {\lt} \varepsilon {\lt} \lambda \delta _0\) then \(\log N^{\mathrm{ext}}(W_k, \varepsilon ) \le \sum _{j{\lt}k} \varphi (\varepsilon /\lambda ^{j+1})\) for every \(k\).
(Lower bound in sum form.) Suppose \(N^{\mathrm{ext}}(G,\delta ) {\lt} \infty \) for all \(\delta {\gt} 0\), \(G \ne \emptyset \), and \(\varphi (\delta ) \le \log M(G, \delta )\) for \(0 {\lt} \delta {\lt} \delta _0\). If \(0 {\lt} \varepsilon {\lt} 1/2\) and \(2\varepsilon {\lt} \lambda \delta _0\) then \(\sum _{j{\lt}k} \varphi (2\varepsilon /\lambda ^{j+1}) \le \log N^{\mathrm{ext}}(W_k, \varepsilon )\) for every \(k\).
(EL balance with saturated variance.) For \(0 {\lt} \theta {\lt} 1\) and \(n \ge 1\), the depth \(k := \lceil \log n / (2\log (1/\theta ))\rceil \) satisfies \(\theta ^k \le n^{-1/2}\): the bias \(\theta ^k\) is balanced against the saturated variance \(n^{-1/2}\), and \(k = \frac{\log n}{2\log (1/\theta )} + O(1)\).
Composition covering lemma. If every \(h \in H\) is \(L_H\)-Lipschitz then for every \(\varepsilon \ge 0\),
(For \(L_H = 0\) the second radius is \(0\) by the convention \(x/0 = 0\), and the inequality still holds.)
Composition covering lemma on the sample. If every \(h \in H\) is \(L_H\)-Lipschitz on \(\mathcal R_S(F)\) then for every \(\varepsilon \ge 0\),
(For \(L_H = 0\) the second radius is \(0\) by the convention \(x/0 = 0\).) Compared with the sup-norm lemma lem:ent-composition-covering, the radius of \(F\) is \(\varepsilon /(4L_H)\) instead of \(\varepsilon /(2L_H)\): the cover of \(F\) must have its centres in \(F\), and \(N(F, d_S, 2r) \le N^{\mathrm{ext}}(F, d_S, r)\).
For \(H\) \(L_H\)-Lipschitz on \(\mathcal R_S(F)\) and radii \(r_1, r_2 \ge 0\), \(N^{\mathrm{ext}}(H \circ F, \| \cdot \| _S, r_1 + L_H r_2) \le N_{S,F}(H, r_1) \cdot N(F, d_S, r_2)\), where \(N(F, d_S, \cdot )\) is the internal covering number (centres in \(F\)). Proof: choose a minimal internal cover \(C_F\) of \(F\) and, for each \(f_b \in C_F\), a minimal external cover of \(H\) for \(\| \cdot \| _{f_b \circ S}\); apply lem:comp-cover-sample and count.
Entropy of the composition class. Let every \(h \in H\) be \(L_H\)-Lipschitz with \(L_H {\gt} 0\), and let \(\varepsilon \in \mathbb R\) be such that \(N^{\mathrm{ext}}(H, \| \cdot \| _\infty , \varepsilon /2)\) and \(N^{\mathrm{ext}}(F, d_\infty , \varepsilon /(2L_H))\) are finite. Then
Entropy of the composition class on the sample. Let every \(h \in H\) be \(L_H\)-Lipschitz on \(\mathcal R_S(F)\) with \(L_H {\gt} 0\), and let \(\varepsilon \in \mathbb R\) be such that \(N_{S,F}(H, \varepsilon /2)\) and \(N^{\mathrm{ext}}(F, d_S, \varepsilon /(4L_H))\) are finite. Then
(Polynomial growth, bounded diameter; single-set form.) If \(0 \le D \le \overline D\), \(\overline D {\gt} 0\), \(C_0 \ge 1\), \(D \ge 0\) (the exponent) and \(N(A, \varepsilon ) \le C_0 (1 + k/\varepsilon )^{D}\) for all \(\varepsilon {\gt} 0\), then \(\mathsf V(D, A) \le \overline D\sqrt{D}\bigl(\sqrt{\log (1 + k/\overline D)} + \tfrac {\sqrt\pi }{2}\bigr) + \overline D\sqrt{\log C_0}\).
Almost-everywhere comparison of the entropy integrands. Let \(H \ne \emptyset \), \(F \ne \emptyset \) satisfy Assumptions ass:ent-readout, ass:ent-transition with \(L_H {\gt} 0\), \(n \ge 1\), and let \(N_S(x)\) be FoML’s open-ball covering number of \(H \circ F\) for \(\| \cdot \| _S\). Then for every \(\varepsilon {\gt} 0\) and Lebesgue-a.e. \(x {\gt} \varepsilon \),
Proof: for every \(y {\lt} x\), \(N_S(x) \le N^{\mathrm{ext}}(H \circ F, \| \cdot \| _\infty , y/2) \le N^{\mathrm{ext}}(H, y/4) N^{\mathrm{ext}}(F, y/(4L_H))\); the right-hand side \(\psi (y) = \mathcal E_H(y/4) + \mathcal E_F(y/(4L_H))\) is antitone in \(y\), hence continuous outside a countable set, and letting \(y \uparrow x\) at a continuity point gives the claim.
Almost-everywhere comparison of the entropy integrands on the sample. Let \(H, F \ne \emptyset \) satisfy ass:ent-readout-sample and ass:ent-transition-sample with \(L_H {\gt} 0\), and let \(N_S(x)\) be FoML’s open-ball covering number of \(H \circ F\) for \(\| \cdot \| _S\). Then for every \(\varepsilon {\gt} 0\) and Lebesgue-a.e. \(x {\gt} \varepsilon \),
Proof: for every \(y {\lt} x\), \(N_S(x) \le N^{\mathrm{ext}}(H \circ F, \| \cdot \| _S, y/2) \le N_{S,F}(H, y/4)\, N^{\mathrm{ext}}(F, d_S, y/(8L_H))\); the right-hand side is antitone in \(y\), hence continuous outside a countable set, and we let \(y \uparrow x\) at a continuity point.
(Logarithmic form of the envelope.) If every \(f \in F\) is \(\Lambda \)-Lipschitz, \(k \ge 1\) and \(N := N^{\mathrm{ext}}(F, \varepsilon /S_k(\Lambda )) {\lt} \infty \), then \(\log N^{\mathrm{ext}}(B(k,F), \varepsilon ) \le \log (k+1) + k\log N\) (using \(1 + kN^k \le (k+1)N^k\) for \(N \ge 1\); for \(N = 0\) both sides are \(\ge 0 = \log 1\)).
(Transfer to the empirical metric.) If every \(f \in F\) is \(\Lambda \)-Lipschitz and \(N^{\mathrm{ext}}(F, d_\infty , \varepsilon /(2S_k(\Lambda ))) {\lt} \infty \), then \(\log N(B(k,F), d_S, \varepsilon ) \le \log N^{\mathrm{ext}}(B(k,F), d_\infty , \varepsilon /2)\) for every sample \(S\) (‘lem:covering-empSpace-le-external-unifMaps‘).
(First inequality of ‘prop:envelope‘.) If every \(f \in F\) is \(\Lambda \)-Lipschitz then for every \(k \ge 0\) and \(\varepsilon \ge 0\),
(Second inequality of ‘prop:envelope‘.) For every \(k \ge 0\) and \(\varepsilon \ge 0\),
because \(S_m(\Lambda ) \le S_k(\Lambda )\) for \(m \le k\), covering numbers are nonincreasing in the radius, and \(N^m \le N^k\) for \(N \ge 1\) (or \(N = 0\), \(m, k \ge 1\)).
(Uniform form of the telescoping estimate ‘lem:impl-word-error‘.) If every \(f \in F\) is \(\Lambda \)-Lipschitz and \(d_\infty (f, \tilde f) \le \delta \) for \(f \in F\), then \(d_\infty (f_u, \tilde f_u) \le \delta \sum _{i{\lt}|u|} \Lambda ^i\) for every list \(u\) of layers in \(F\).
(Covering the words of length \(m\).) If every \(f \in F\) is \(\Lambda \)-Lipschitz then for every \(\varepsilon \ge 0\) and \(m \ge 0\),
The \(M^m\) compositions \(c_{i_m} \circ \cdots \circ c_{i_1}\) of the centres of an \(\varepsilon /S_m(\Lambda )\)-cover \(\{ c_1, \dots , c_M\} \) of \(F\) form an \(\varepsilon \)-cover of \(F^m\) by the telescoping estimate.
For \(\kappa , \varepsilon _0, R_{\mathrm{out}} {\gt} 0\), \(n \ge 1\) and \(q \ge 0\), the first branch of the minimum is the smaller one as soon as \(n \ge R_{\mathrm{out}}^2 q / (\kappa ^2 \varepsilon _0^2)\): \(\kappa \varepsilon _0 \sqrt{q/n} \le \kappa ^2 \varepsilon _0^2 / R_{\mathrm{out}}\).
FoML covering numbers versus empirical external covering numbers. Let \(G \subseteq \mathbb R^{\mathcal X}\) be nonempty and totally bounded for \(\| \cdot \| _S\), and let \(0 \le 2r {\lt} x\). Then FoML’s open-ball covering number of \(G\) for \(\| \cdot \| _S\) satisfies \(N^{\mathrm{open}}_S(G, x) \le N(G, \| \cdot \| _S, 2r) \le N^{\mathrm{ext}}(G, \| \cdot \| _S, r)\) (the two metrics coincide; the factor \(2\) is the price of the internal cover).
FoML covering numbers versus uniform external covering numbers. Let \(G \subseteq \mathbb R^{\mathcal X}\) be nonempty and totally bounded for \(\| \cdot \| _S\), and let \(0 \le 2r {\lt} x\). Then FoML’s open-ball covering number of \(G\) for \(\| \cdot \| _S\) satisfies \(N^{\mathrm{open}}_S(G, x) \le N(G, \| \cdot \| _\infty , 2r) \le N^{\mathrm{ext}}(G, \| \cdot \| _\infty , r)\) (internal closed \(2r\)-balls for the sup norm with centres in \(G\) are contained in open \(x\)-balls for \(\| \cdot \| _S\)).
(Bias for fixed-point refinement.) Let \(K\) be compact with \(D_K = \mathrm{diam}(K)\) and \(R \ge 0\). Against the teacher–student class \(\mathcal C = \overline{H_R \circ \langle F_{\rm fp}\rangle }^{\, d_\infty }\),
by the truncation argument ‘lem:truncation-bias‘ (\(\mathrm{lip}(h) \le R\), \(\mathrm{lip}(u) \le (1-h\mu )^k\) for \(u \in F_{\rm fp}^{\, k}\)) and ‘lem:approx-transfer‘.
(Contraction of projected gradient steps.) Let \(0 {\lt} \mu \le \Lambda \), let \(s\) be \(\mu \)-strongly monotone and \(\Lambda \)-co-coercive, and let \(0 {\lt} h \le 2/(\mu +\Lambda )\). Then \(T_s = \Pi _K(\cdot + h s(\cdot ))\) is \(\lambda \)-Lipschitz on \(K\) with \(\lambda := 1 - h\mu \in [0,1)\).
(Explicit entropy bound via ‘cond:p1-ucont‘.) Assume moreover \(h\mu {\lt} 1\) and \(K \ne \emptyset \), and let \(\lambda = 1 - h\mu \), \(m(\varepsilon ) = \lceil \log _{1/\lambda }(2 D_K/\varepsilon )\rceil \) with \(D_K = \mathrm{diam}(K)\). Then for every \(\varepsilon {\gt} 0\) and every \(k\) (generalizing the paper, which assumes \(k \ge m(\varepsilon )\); see ‘cond:p1-ucont‘),
(In Lean the absorbing set is \(K\) itself, with \(L = 0\).)
(Saturated variance profile for fixed-point refinement.) Under the hypotheses of ‘lem:fp-saturation‘, with \(N_\infty (\varepsilon ) := N(\overline{\langle F_{\rm fp}\rangle }, d_\infty , \varepsilon /2) {\lt} \infty \) and \(D_K = \mathrm{diam}(K)\): if \(\int _0^{D_K}\sqrt{\log N_\infty (\varepsilon )}\, d\varepsilon {\lt} \infty \) then \(\mathsf V_k(S) \le \mathsf V_\infty := \int _0^{D_K}\sqrt{\log N_\infty (\varepsilon )}\, d\varepsilon \) for all \(k\) and all samples \(S\) (case (i) of ‘prop:profiles‘).
(Saturation for fixed-point refinement, ‘cond:p1‘.) If \(K\) is compact then for every \(\varepsilon {\gt} 0\) and every \(k\), \(N^{\mathrm{ext}}(B(k,F_{\rm fp}), d_\infty , \varepsilon ) \le N^{\mathrm{ext}}(\overline{\langle F_{\rm fp}\rangle }, d_\infty , \varepsilon ) {\lt} \infty \); in particular \(\sup _k N^{\mathrm{ext}}(B(k,F_{\rm fp}), d_\infty , \varepsilon ) {\lt} \infty \).
(Estimation term for fixed-point refinement.) Let \(K\) be compact with \(\| x\| \le M_K\) on \(K\), \(R {\gt} 0\), and assume \(\mathsf V_\infty (F_{\rm fp}) {\lt} \infty \) (‘def:sat-integrable‘). Then for every sample \(S\) of size \(n \ge 1\) and every depth \(k\),
(‘cor:var-profiles-p1‘ with \(A\_ H = 1\), \(L = R\), and ‘lem:linear-readouts-rademacher‘).
(Right inverse from a well-conditioned Gram matrix.) Let \(\varphi _1, \dots , \varphi _n \in \mathcal H\) and suppose the Gram matrix \(G = (\langle \varphi _i, \varphi _{i'} \rangle )_{i,i'}\) satisfies \(c^\top G c \ge \lambda _{\min } \| c\| _2^2\) for all \(c \in \mathbb R^n\), with \(\lambda _{\min } {\gt} 0\). Then for every \(c \in \mathbb R^n\) there is \(w \in \mathcal H\) with \(\langle w, \varphi _i\rangle = c_i\) for all \(i\) and \(\| w\| \le \lambda _{\min }^{-1/2} \| c\| _2\).
(Existence of the layerwise implementation map.) Given implementations \(f \mapsto \tilde f\) of the transitions and \(h \mapsto \tilde h\) of the output layers, there is a map \(\iota : \mathbb R^{\mathcal X} \to \mathbb R^{\mathcal X}\) such that every \(g \in \mathcal H_k\) has a representation \(g = h \circ f_m \circ \cdots \circ f_1\) with \(h \in H\), \(f_i \in F\), \(m \le k\) and \(\iota (g) = \tilde h \circ \tilde f_m \circ \cdots \circ \tilde f_1\).
(Telescoping error propagation.) Let every \(f \in F\) be \(\Lambda \)-Lipschitz and let \(\tilde f\) satisfy \(d_\infty (f, \tilde f) \le \delta \) for \(f \in F\), where \(\delta \ge 0\). Then for every representation \(u = [f_m, \dots , f_1]\) of layers in \(F\) and every \(x\),
(Rademacher complexity of the linear readout class.) If \(\| \Phi (x)\| \le M\) for all \(x\) (\(R, M \ge 0\)), then for every sample \(S\) of size \(n\),
(Proof: \(\hat{\mathfrak R}_S(H_R(\Phi )) \le R\, \mathbb E_\sigma \| n^{-1}\sum _i\sigma _i \Phi (x_i)\| \le R\sqrt{\sum _i\| \Phi (x_i)\| ^2}/n \le RM/\sqrt n\), via FoML’s hilbertPredictor_empiricalRademacherComplexity_le.)
(‘prop:hilbert-sg‘ for the radius-\(R\) class.) If \(\Phi \) is \(L_\Phi \)-Lipschitz and \(R {\gt} 0\), then \(H_R(\Phi )\) satisfies the sub-Gaussian increment condition ‘ass:sg-increment-main‘ with \(A_H = 1\) and \(L = R L_\Phi \), for every hidden-layer class \(\mathfrak F\).
(Modulus extraction on compact domains.) Let \((\mathcal X,d)\) be compact and \(H \subseteq \mathcal X^{\mathcal X}\) equicontinuous. Then \(H\) admits a common monotone modulus of continuity \(\omega \) with \(\omega (r) \to 0\) as \(r \downarrow 0\) and \(d(f(x), f(y)) \le \omega (d(x,y))\) for all \(f \in H\), \(x, y \in \mathcal X\). In particular ‘thm:maa‘ applies.
(Covering bound for equal-step schemes.) If every drift \(s \in \mathcal S_T\) is \(\Lambda _s\)-Lipschitz in \(x\) and \(T \ge 0\), then for every \(k \ge 0\) and \(\varepsilon \ge 0\),
(the radius is \(\varepsilon /(Te^{\Lambda _sT})\), equal to \(0\) when \(T = 0\)).
(Pointwise stability of equal-step schemes.) If \(s(\cdot ,\tau )\) is \(\Lambda _s\)-Lipschitz for every \(\tau \), \(T \ge 0\), and \(\| s(x,\tau ) - s'(x,\tau )\| \le \delta \) for all \(x \in K\), \(\tau \in [0,T]\), then \(\| \Phi _{s,m}(x) - \Phi _{s',m}(x)\| \le Te^{\Lambda _sT}\delta \) for all \(m \ge 0\) and \(x\).
(Covering the \(m\)-step schemes.) Under the hypotheses of ‘lem:ode-equal-scheme-uniformdist‘ for every \(s \in \mathcal S_T\), the image of an \(\varepsilon e^{-\Lambda _sT}/T\)-cover of \(\mathcal S_T\) (in \(\| \cdot \| _\infty \) on \(K \times [0,T]\), centres extended by \(0\) outside \([0,T]\)) under \(s \mapsto \Phi _{s,m}\) is an \(\varepsilon \)-cover of \(\{ \Phi _{s,m} : s \in \mathcal S_T\} \); hence \(N^{\mathrm{ext}}(\{ \Phi _{s,m} : s \in \mathcal S_T\} , d_\infty , \varepsilon ) \le N^{\mathrm{ext}}(\mathcal S_T, \| \cdot \| _\infty , \varepsilon e^{-\Lambda _sT}/T)\).
(Stability of equal-step schemes in \(d_\infty \).) If \(s(\cdot ,\tau )\) is \(\Lambda _s\)-Lipschitz for every \(\tau \) and \(T \ge 0\), then for every drift \(s'\) and every \(m \ge 0\),
(Only the Lipschitz constant of \(s\) is used; \(s'\) may be any drift.)
(Global error of the explicit Euler scheme.) Let \(s\) be \(\Lambda _s\)-Lipschitz in \(x\), \(\Lambda _\tau \)-Lipschitz in \(\tau \) and bounded by \(M_s\), let \(x\) solve \(\dot x(t) = s(x(t),t)\) on \([0,T]\) (\(T \ge 0\)), and let \(y_0, \dots , y_k\) be the Euler iterates with \(k \ge 1\) equal steps \(h = T/k\) started at \(y_0 = x(0)\). Then
(Discrete Grönwall: \(e_{i+1} \le (1 + h\Lambda _s) e_i + c h^2/2\) with \(c = \Lambda _s M_s + \Lambda _\tau \), hence \(e_k \le \frac{c h^2}{2}\sum _{j{\lt}k}(1+h\Lambda _s)^j \le \frac{c h^2}{2}\, k\, e^{\Lambda _s T}\).)
(One-step consistency of the Euler scheme.) Let \(s\) be \(\Lambda _s\)-Lipschitz in \(x\), \(\Lambda _\tau \)-Lipschitz in \(\tau \) and bounded by \(M_s\), and let \(x\) solve \(\dot x(t) = s(x(t), t)\) on \([t_0, t_0 + h]\) (\(h \ge 0\)). Then
(The function \(g(t) = x(t) - x(t_0) - (t - t_0) s(x(t_0),t_0)\) has \(\| g'(t)\| \le \Lambda _s\| x(t) - x(t_0)\| + \Lambda _\tau (t - t_0) \le (\Lambda _s M_s + \Lambda _\tau )(t - t_0)\), and the mean value inequality with the quadratic boundary \(B(t) = (\Lambda _s M_s + \Lambda _\tau )(t-t_0)^2/2\) gives the claim.)
(Bias for fixed-horizon integration.) Let the target class be the readouts of the endpoint map of the exact flow of \(s^\ast \in \mathcal S_T\): \(\mathcal C = \{ x_0 \mapsto \langle w, \Phi _T(x_0)\rangle : \| w\| \le R\} \), where for every \(x_0 \in K\) the flow \(\Phi _T(x_0) = x(T)\) is given by a solution of \(\dot x = \tilde s(x, t)\), \(x(0) = x_0\), for an extension \(\tilde s\) of \(s^\ast \) which is \(\Lambda _s\)-Lipschitz in \(x\), \(\Lambda _\tau \)-Lipschitz in \(\tau \) and bounded by \(M_s\). Assume the projection is inactive along the Euler steps of size \(T/k\) (\(K\) is invariant), \(k \ge 1\) and \(T/k \le h_0\). Then
by ‘lem:ode-euler-error‘, ‘lem:ode-scheme-eq-euler‘ and ‘lem:approx-transfer‘.
(Estimation term for fixed-horizon schemes.) Let \(K\) be compact with \(\| x\| \le M_K\) on \(K\), every drift in \(\mathcal S_T \ne \emptyset \) be \(\Lambda _s\)-Lipschitz in \(x\), \(T \ge 0\), \(R {\gt} 0\), and assume the entropy condition of ‘lem:ode-saturation-profile‘. Then for every sample \(S\) of size \(n \ge 1\) and every \(k\),
(‘thm:hidden-decomp‘ on the scheme class, ‘lem:ode-saturation-profile‘ and ‘lem:linear-readouts-rademacher‘).
(Integrability of the entropy integrand from total boundedness.) Under the hypotheses of ‘lem:ode-profile-saturation-generic‘, the entropy integrand \(\varepsilon \mapsto \sqrt{\log N(A_k, d_S, \varepsilon )}\) is interval-integrable on \([0, \mathrm{diam}_S(A_k)]\) (it is dominated by the integrable majorant).
(Saturated profile from total boundedness.) Let \(\mathcal X\) be compact, let \(G \subseteq \mathcal X^{\mathcal X}\) be totally bounded in \(d_\infty \) and \(A_k \subseteq G\) for all \(k\). With \(N_\infty (\varepsilon ) := N(\overline G, d_\infty , \varepsilon /2) {\lt} \infty \) and \(\overline D := \mathrm{diam}(\mathcal X)\), if \(\int _0^{\overline D}\sqrt{\log N_\infty (\varepsilon )}\, d\varepsilon {\lt} \infty \) then \(\mathsf V(\mathrm{diam}_S(A_k), A_k) \le \int _0^{\overline D} \sqrt{\log N_\infty (\varepsilon )}\, d\varepsilon \) for every sample \(S\) and every \(k\) (case (i) of ‘prop:profiles‘).
(Stability and saturation.) Let \(K\) be compact and let every drift \(s \in \mathcal S_T\) be \(\Lambda _s\)-Lipschitz in \(x\). Every scheme in \(\bigcup _k B_T(k)\) is \(e^{\Lambda _sT}\)-Lipschitz on \(K\). Consequently \(\bigcup _k B_T(k)\) is equicontinuous on the compact set \(K\), hence totally bounded in \(d_\infty \) (‘thm:caa‘), and
(Saturated variance profile for Euler schemes.) Under the hypotheses of ‘lem:ode-saturation‘, with \(N_\infty (\varepsilon ) := N(\overline{\bigcup _k B_T(k)}, d_\infty , \varepsilon /2) {\lt} \infty \) and \(D_K = \mathrm{diam}(K)\): if \(\int _0^{D_K}\sqrt{\log N_\infty (\varepsilon )}\, d\varepsilon {\lt} \infty \) then \(\mathsf V_k(S) \le \mathsf V_\infty := \int _0^{D_K}\sqrt{\log N_\infty (\varepsilon )}\, d\varepsilon \) for all \(k\) and all samples \(S\) (case (i) of ‘prop:profiles‘).
(Stability with respect to the drift; discrete Grönwall.) Let \(s(\cdot ,\tau )\) be \(\Lambda _s\)-Lipschitz for every \(\tau \), let \(h_i \ge 0\), and suppose \(\| s(x,\tau _i) - s'(x,\tau _i)\| \le \delta \) for all \(x \in K\) and all stamps \(\tau _i\) used by the scheme. Then for every \(x\),
where \(y_n, y_n'\) are the schemes of \(s\) and \(s'\) with the same steps and stamps started at \(x\). Indeed \(\| y_{i} - y_{i}'\| \le (1 + h_i\Lambda _s)\| y_{i-1} - y_{i-1}'\| + h_i\delta \) and \(1 + h\Lambda _s \le e^{h\Lambda _s}\).
(Explicit entropy of equal-step schemes.) Let every drift \(s \in \mathcal S_T\) be \(\Lambda _s\)-Lipschitz in \(x\) and \(T \ge 0\). Then for all \(s, s' \in \mathcal S_T\) and \(m \ge 0\),
and consequently, for every \(k \ge 0\) and \(\varepsilon \ge 0\),
The entropy-integral consequence is ‘lem:ode-scheme-entropy-profile‘.
(Entropy integral of equal-step schemes.) Let \(K\) be compact, every drift \(s \in \mathcal S_T\) be \(\Lambda _s\)-Lipschitz in \(x\), \(T {\gt} 0\), and suppose \(N^{\mathrm{ext}}(\mathcal S_T, \| \cdot \| _\infty , \rho ) {\lt} \infty \) for all \(\rho {\gt} 0\) and that \(\varepsilon \mapsto \sqrt{\log N^{\mathrm{ext}}(\mathcal S_T, \| \cdot \| _\infty , \varepsilon e^{-\Lambda _sT}/(2T))}\) is integrable on \([0, D_K]\), \(D_K = \mathrm{diam}(K)\). Then for every sample \(S\) and every \(k\), for the class \(E_T(k)\),
(The paper has \(\varepsilon e^{-\Lambda _sT}/T\): the factor \(2\) is the price of comparing the internal covering number in \(d_S\) defining \(\mathsf V_k(S)\) with the external one in \(d_\infty \), ‘lem:covering-empSpace-le-external-unifMaps‘.)
(Inactive projection.) Let \(\tilde s : \mathbb R^d \times \mathbb R \to \mathbb R^d\) extend the drift \(s\) from \(K\), and suppose \(K\) is invariant under the Euler steps, \(x + h\, s(x,\tau ) \in K\) for \(x \in K\). Then the projection is inactive and the scheme with \(k\) equal steps of size \(h\) coincides with the Euler iterates: \(T_{s,h,(k-1)h} \circ \cdots \circ T_{s,h,0}(x_0) = y_k(x_0)\).
(Integrability under P1.) Under the hypotheses of ‘cor:profile-p1-totallyBounded‘ (in particular the interval-integrability of the majorant \(\sqrt{\log N_\infty }\) on \([0, D_{\mathcal X}]\)), the entropy integrand \(\varepsilon \mapsto \sqrt{\log N(B(k,F), d_S, \varepsilon )}\) is interval-integrable on \([0, D_k(S)]\) for every \(k\): the hypothesis of ‘thm:hidden-decomp-depth‘ is automatic.
(Long words are almost constant.) Under the hypotheses of ‘cond:p1-ucont‘, for every word \(w \in F^l\) with \(l \ge m(\varepsilon )\) and all \(x, y \in \mathcal X\),
In particular \(d_\infty (w, \mathrm{const}_{w(a_0)}) \le \varepsilon /2\) for every \(a_0\).
(Long words.) Under the hypotheses of ‘cond:p1-ucont‘, the set of all words of length \(\ge m(\varepsilon )\) satisfies
Indeed each such word \(w\) is \(\varepsilon /2\)-close to the constant map \(\mathrm{const}_{w(a_0)}\) with \(w(a_0) \in A\), so the constant maps at the points of an \(\varepsilon /2\)-net of \(A\) form an \(\varepsilon \)-cover.
(Short words; subadditivity over finitely many sets.) For a finite family \((A_i)_{i \in s}\), \(N^{\mathrm{ext}}\bigl(\bigcup _{i \in s} A_i, \varepsilon \bigr) \le \sum _{i \in s} N^{\mathrm{ext}}(A_i, \varepsilon )\). Applied to \(A_l = F^l\), \(l {\lt} m(\varepsilon )\), this covers the short words.
(Diameter envelope.) Under the hypotheses of ‘lem:p2-wordball-subset-ball‘ and with the orbit map \(L_\alpha \)-Lipschitz, for all \(f, g \in B(k,F)\),
i.e. \(\mathrm{diam}_\infty B(k,F) \le 2L_\alpha R_S k\) (and a fortiori \(D_k(S) \le 2 L_\alpha R_S k\) for every sample \(S\)).
(P2 in the empirical metric.) Under the hypotheses of ‘cond:p2-nilp‘, for every \(k\) and \(\varepsilon {\gt} 0\),
as real numbers. (The factor \(2^D\) comes from \(N(\cdot , d_S, \varepsilon ) \le N^{\mathrm{ext}}(\cdot , d_\infty , \varepsilon /2)\) and \(1 + 2k/\varepsilon \le 2(1 + k/\varepsilon )\).)
(Integrability under P2.) Under the hypotheses of ‘cond:p2-nilp‘, if \(D_k(S) \le \overline D\) with \(\overline D {\gt} 0\), then the entropy integrand of \(B(k,F)\) is interval-integrable on \([0, D_k(S)]\) (it is dominated by \(\sqrt{\log C_0} + \sqrt D\sqrt{\log (1 + k/\overline D)} + \sqrt D\sqrt{\log (\overline D/\varepsilon )}\), ‘lem:sqrt-metric-entropy-le-of-poly‘).
Let \(\alpha : H \to \mathcal X^{\mathcal X}\) be a homomorphism, assume the length \(g \mapsto d_H(e,g)\) is subadditive, \(d_H(e, gh) \le d_H(e,g) + d_H(e,h)\), and let \(S \subseteq H\) satisfy \(d_H(e,s) \le R_S\) for all \(s \in S\) and \(F \subseteq \alpha (S)\). Then for every \(k\),
Let \(f = (f_1,\dots ,f_r)\) and let \(P\) be a finite family of probes such that the evaluations \(\mathrm{ev}_P(f_u)\), \(u \in [r]^k\), are pairwise at sup-distance \(\ge \delta {\gt} 0\). Then \(N^{\mathrm{ext}}(B(k,F), d_\infty , \varepsilon ) \ge r^k\) for every \(\varepsilon \) with \(2\varepsilon {\lt} \delta \).
(Probes and packing.) Let \(A \subseteq (\mathcal X^{\mathcal X}, d_\infty )\), let \(P\) be a finite family of probes and let \(T \subseteq \mathrm{ev}_P(A)\) be \(\delta \)-separated: \(d_{\max }(y,z) \ge \delta \) for all distinct \(y, z \in T\). Then for every \(\varepsilon \) with \(2\varepsilon {\lt} \delta \), \(|T| \le N^{\mathrm{ext}}(A, d_\infty , \varepsilon )\).
(Indexed form of ‘lem:probes-packing‘.) Let \((g_i)_{i \in I}\) be a family of maps in \(A \subseteq (\mathcal X^{\mathcal X}, d_\infty )\) and \(P\) a finite family of probes such that \(d_{\max }(\mathrm{ev}_P(g_i), \mathrm{ev}_P(g_j)) \ge \delta {\gt} 0\) for all \(i \ne j\). Then \(|I| \le N^{\mathrm{ext}}(A, d_\infty , \varepsilon )\) whenever \(2\varepsilon {\lt} \delta \).
(Reachable radius.) Let every \(f \in F\) be \(\Lambda \)-Lipschitz, \(x_0 \in \mathcal X\), \(R \ge 0\), and suppose \(d(f(x_0), x_0) \le c\) for all \(f \in F\) (with \(c \ge 0\)). Then every \(g \in B(k,F)\) maps the ball \(B(x_0,R)\) into \(B\bigl(x_0, \Lambda ^kR + c\, S_k(\Lambda )\bigr)\) when \(\Lambda \ge 1\), and into \(B\bigl(x_0, R + c\, S_k(\Lambda )\bigr)\) when \(\Lambda \le 1\). Consequently, if \(S \subset B(x_0,R)\), then \(D_k(S) \le 2\bigl(\Lambda _+^kR + c\, S_k(\Lambda )\bigr)\) with \(\Lambda _+ = \max \{ 1,\Lambda \} \): bounded in \(k\) if \(\Lambda {\lt} 1\), at most linear if \(\Lambda = 1\), and at most exponential if \(\Lambda {\gt} 1\).
(Reachable radius, unified form.) Let every \(f \in F\) be \(\Lambda \)-Lipschitz, \(x_0 \in \mathcal X\), \(c \ge 0\) with \(d(f(x_0), x_0) \le c\) for all \(f \in F\), and \(R \ge 0\). Then every \(g \in B(k,F)\) maps \(B(x_0, R)\) into \(B\bigl(x_0, \Lambda _+^kR + c\, S_k(\Lambda )\bigr)\), \(\Lambda _+ := \max \{ 1,\Lambda \} \): for \(d(x,x_0) \le \rho \), \(d(f(x), x_0) \le d(f(x), f(x_0)) + d(f(x_0), x_0) \le \Lambda \rho + c\); iterate.
(Diameter of the reachable states.) Under the hypotheses of ‘lem:reachable-radius-core‘, if the sample lies in \(B(x_0, R)\) then \(D_k(S) \le 2\bigl(\Lambda _+^kR + c\, S_k(\Lambda )\bigr)\): bounded in \(k\) if \(\Lambda {\lt} 1\), at most linear if \(\Lambda = 1\), at most exponential if \(\Lambda {\gt} 1\).
A bounded set \(K\) of a finite-dimensional space has finite covering numbers: \(N^{\mathrm{ext}}(K, \varepsilon ) {\lt} \infty \) for \(\varepsilon {\gt} 0\) (external covering number of the state space \(K\) by points of \(K\); \(K\) is totally bounded since its closure is compact).
(Envelope profile of the ReLU class, all \(\Lambda \).) Let \(K\) be bounded with \(\| x\| \le R_K\) on \(K\), \(D_K \le \overline D\), \(\overline D {\gt} 0\), \(p = 2mw + w + m\) and \(C_F\) as in ‘lem:relu-layer-covering‘. Then for \(k \ge 1\) and every sample \(S\), with \(\Lambda _+ = \max \{ 1, \Lambda \} \),
(‘cor:envelope-profiles-a-one-add‘ with ‘lem:relu-layer-covering‘).
(The expand-and-reset maps satisfy ‘cond:e2-pingpong‘.) With the chambers, cores, anchors and marker above and separation \(\alpha = 1/8\): \(N^{\mathrm{ext}}(B(k, \{ g_0, g_1\} ), d_\infty , \varepsilon ) \ge 2^k\) for all \(k\) and \(2\varepsilon {\lt} 1/8\), and the words of each length are pairwise distinct.
(Covering numbers of one ReLU block.) Let \(K \subseteq \mathbb R^m\) with \(\| x\| \le R_K\) on \(K\) (\(R_K \ge 0\)), let \(\Pi _K\) be a \(1\)-Lipschitz retraction onto \(K\), and let \(F_\Lambda \) be the class of ReLU blocks of width \(w\) with \(p = 2mw + w + m\) parameters. (a) For admissible \(\vartheta , \vartheta '\),
(b) for every \(\varepsilon {\gt} 0\), \(N^{\mathrm{ext}}(F_\Lambda , d_\infty , \varepsilon ) {\lt} \infty \) and \(\log N^{\mathrm{ext}}(F_\Lambda , d_\infty , \varepsilon ) \le p\log (1 + C_F/\varepsilon )\) with \(C_F = 2(2\beta _W R_K + \beta _W + \beta + 1)\max \{ \beta _W, \beta \} \) (‘def:relu-cover-const‘; the paper’s constant has the same shape, see there).
(Parameter-Lipschitz estimate.) Let \(\| x\| \le R_K\) on \(K\), \(\| V\| _{\rm op} \le \beta _W\), \(\| W'\| _{\rm op} \le \beta _W\) and \(\| b'\| \le \beta \). Then
(Only these three parameter bounds are used, as in the paper’s proof.)
(Covering bound, cardinality form.) Let \(\| x\| \le R_K\) on \(K\) with \(R_K \ge 0\) and \(p = 2mw + w + m\). For every \(\varepsilon {\gt} 0\), \(N^{\mathrm{ext}}(F_\Lambda , d_\infty , \varepsilon ) \le (1 + C_F/\varepsilon )^p\) with \(C_F = 2(2\beta _W R_K + \beta _W + \beta + 1)\max \{ \beta _W, \beta \} \): \(N^{\mathrm{ext}}(F_\Lambda , \varepsilon ) \le N(P, \varepsilon /L_F) \le M(P, \varepsilon /L_F) \le M(\overline B(0, \max \{ \beta _W,\beta \} ), \varepsilon /L_F) \le (1 + 2L_F\max \{ \beta _W,\beta \} /\varepsilon )^p\) by ‘lem:relu-lipschitz-on-image‘, ‘lem:packing-covering‘ and ‘lem:relu-volumetric‘.
(Covering bound, entropy form.) Under the hypotheses of ‘lem:relu-layer-covering-b‘, for every \(\varepsilon {\gt} 0\), \(N^{\mathrm{ext}}(F_\Lambda , d_\infty , \varepsilon ) {\lt} \infty \) and \(\log N^{\mathrm{ext}}(F_\Lambda , d_\infty , \varepsilon ) \le p\log (1 + C_F/\varepsilon )\).
\(\vartheta \mapsto f_\vartheta \) is \(L_F\)-Lipschitz on the admissible parameter set, from \((\mathrm{parameters}, \| \cdot \| _{\sup })\) to \((\mathcal X^{\mathcal X}, d_\infty )\): each of the four differences in ‘lem:relu-layer-covering-a‘ is at most \(\| \vartheta - \vartheta '\| _{\sup }\).
(Volumetric bound.) In a real normed space \(E\) of dimension \(q\), for \(\rho \ge 0\) and \(\delta {\gt} 0\), the packing number of the closed ball satisfies \(M(\overline B(x, \rho ), \delta ) \le \lfloor (1 + 2\rho /\delta )^q \rfloor \); in particular a \(\delta \)-cover of the ball of size at most \((1 + 2\rho /\delta )^q\) exists.
(Volumetric bound, finite sets.) Let \(E\) be a real normed space of dimension \(q\), \(\rho \ge 0\), \(\delta {\gt} 0\). If \(s \subseteq \overline B(x, \rho )\) is a finite \(\delta \)-separated set (distinct points at distance \({\gt} \delta \)) then \(|s| \le (1 + 2\rho /\delta )^q\): the open balls \(B(y, \delta /2)\), \(y \in s\), are pairwise disjoint and contained in \(B(x, \rho + \delta /2)\), and Haar measure scales like \(r^q\).
(Telescoping estimate for contractive layers.) If every \(f \in F\) is \(\Lambda \)-Lipschitz with \(\Lambda {\lt} 1\) then \(N^{\mathrm{ext}}(F^l, d_\infty , \varepsilon ) \le N^{\mathrm{ext}}(F, d_\infty , (1 - \Lambda )\varepsilon )^l\), since \(S_l(\Lambda ) = \sum _{i{\lt}l}\Lambda ^i \le 1/(1-\Lambda )\) in ‘lem:envelope-words‘.
If \(N(A, \varepsilon ) \le C_0 (1 + k/\varepsilon )^{D}\) for \(\varepsilon {\gt} 0\), with \(C_0 \ge 1\), \(D \ge 0\), then for \(0 {\lt} \varepsilon \le \overline D\), \(\sqrt{\log N(A,\varepsilon )} \le \sqrt{\log C_0} + \sqrt{D}\sqrt{\log (1 + k/\overline D)} + \sqrt{D}\sqrt{\log (\overline D/\varepsilon )}\).
(Sub-multiplicativity.) Let \(F, G \subseteq \mathcal X^{\mathcal X}\) and suppose every \(f \in F\) is \(\lambda \)-Lipschitz. Then for all \(\varepsilon , \delta \ge 0\),
(This is the paper’s \(N(FG, \varepsilon + \delta ) \le N(F, \varepsilon /\rho ) N(G, \delta /\lambda )\) with \(\rho = 1\), since right composition is \(1\)-Lipschitz for \(d_\infty \); we write \(\lambda \delta \) in place of \(\delta /\lambda \) to avoid dividing by \(\lambda = 0\).)
(Substitution step.) There is a universal constant \(c {\gt} 0\) such that under Assumption ass:readout-realization-main for \(B_k\) and the boundedness hypothesis of thm:sudakov-type, for every \(\varepsilon _0 {\gt} 0\) and every \(q \ge 0\) with \(q \le \log M(B_k, d_S, 2\varepsilon _0)\), \(\hat{\mathfrak R}_S(\mathcal H_k) \ge c \min \{ \kappa \varepsilon _0 \sqrt{q/n}, \kappa ^2 \varepsilon _0^2 / R_{\mathrm{out}}\} \).
Let \(F \ne \emptyset \) and fix a sign pattern \(\sigma \). If for every \(f \in F\) the Rademacher averages \(\{ \frac1n\sum _i \sigma _i h(f(x_i)) : h \in H\} \) are bounded above, and \(f \mapsto Z_f(\sigma )\) is bounded above on \(F\), then \(\sup _{u \in H \circ F} \frac1n \sum _i \sigma _i u(x_i) = \sup _{f \in F} Z_f(\sigma )\).
If \(G \subseteq \mathbb R^{\mathcal X}\) is nonempty with finite uniform covering numbers \(N^{\mathrm{ext}}(G, \| \cdot \| _\infty , r) {\lt} \infty \) for all \(r {\gt} 0\), then \(G\) is totally bounded for the empirical pseudometric \(\| \cdot \| _S\) (FoML’s ‘EmpiricalFunctionSpace‘), \(n \ge 1\).
(Truncation for the uniform closure.) Under the hypotheses of ‘lem:truncation-bias-exact‘, every \(g\) in the uniform closure \(\overline{H \circ \langle F\rangle }^{\, d_\infty }\) satisfies, for every \(\varepsilon {\gt} 0\), \(\inf _{f \in \mathcal H_k}\| f - g\| _\infty \le L_H c^k D_{\mathcal X} + \varepsilon \).
(Truncation, Sec. ‘sec:examples-regime‘.) Let every \(h \in H\) be \(L_H\)-Lipschitz, every \(f \in F\) be \(c\)-Lipschitz, and \(d(x,y) \le D_{\mathcal X}\) on \(\mathcal X\). Then every \(g = h \circ w \in H \circ \langle F\rangle \) is within uniform distance \(L_H c^k D_{\mathcal X}\) of \(\mathcal H_k = H \circ B(k,F)\): if \(w = w_2 \circ w_1\) with \(|w_2| = k\) then \(|h(w_2(w_1 x)) - h(w_2 x)| \le L_H c^k d(w_1 x, x)\).
(Append-only scratchpad, rigorous form.) Let \(|\mathcal A| = m \ge 2\), let \(\Phi : \mathcal A^{\mathbb N} \to \mathcal H\) be an \(L_\Phi \)-Lipschitz feature map into a separable Hilbert space with \(\| \Phi \| \le M_\Phi \) (e.g. the window features \(\Phi _L\)), \(R {\gt} 0\), \(H = H_R(\Phi )\), \(F = F_{\rm w}\), and let the target class be \(\mathcal C = \overline{H \circ \langle F_{\rm w}\rangle }^{\, d_\infty }\). For a measurable loss \(\ell : \mathbb R \times \mathcal Y \to [0,b]\), \(\beta _\ell \)-Lipschitz in its first argument, \(n \ge 1\), \(\eta \ge 0\) and \(\delta \in (0,1)\): with probability at least \(1 - \delta \) over \(\mathcal D \sim P^{\otimes n}\), every \(\eta \)-empirical minimizer \(\hat h \in \mathcal H_k = H \circ B(k, F_{\rm w})\) satisfies
where \(\mathrm{dev}_{\ell ,n,\delta }(B) = 4\beta _\ell B + 6b\sqrt{2\log (4/\delta )/n}\) is the estimation-plus-deviation term of ‘thm:bv‘ (‘def:bv-dev‘). This is the EL regime with \(\alpha = \log (1/\theta )\) and saturated variance. (The implementation map is the identity, \(\varepsilon _{\mathrm{imp}} = 0\).)
(Covering envelope, ‘prop:envelope‘ restated.) Let \(F \subseteq \mathcal X^{\mathcal X}\) consist of \(\Lambda \)-Lipschitz maps, \(\Lambda \ge 0\), and \(S_m(\Lambda ) = \sum _{i=0}^{m-1}\Lambda ^i\). For every \(k \ge 0\) and \(\varepsilon \ge 0\),
The centres of the covers of \(F\) need not belong to \(F\) nor be Lipschitz; only the Lipschitz constant of the maps in \(F\) is used. (In Lean \(\varepsilon /S_0 = 0\) and the statement holds for \(k = 0\) as well.)
Finite Lipschitz scalar output layers. Let \(H = \{ h_1, \dots , h_m\} \) be a finite class of real-valued functions on \(\mathcal X\), each \(L\)-Lipschitz. Then Assumption ass:sg-increment-main holds with \(A_H = (1 + \log m / \log 2)^{1/2}\), for every hidden-layer class \(\mathfrak F\).
(Global scalar observable.) Fix a sample \(S = (x_i)_{i=1}^n\), a hidden class \(B_k\) and \(U_{k,S} = \{ f(x_i) : f \in B_k, i \in [n]\} \). Let \(\Phi : \mathcal X \to \mathcal H\) and assume there are \(u \in \mathcal H\) with \(\| u\| \le R_H\) and constants \(\kappa , R_\Phi {\gt} 0\) such that
Then the readout-realization assumption holds for \(H_{R_H}(\Phi )\) on \(B_k\) with constants \(\kappa \) and \(R_{\mathrm{out}} = R_H R_\Phi \), with the single choice \(h_f = h_u = \langle u, \Phi (\cdot )\rangle \) for all \(f \in B_k\): \(\| h_u \circ f - h_u \circ g\| _S \ge \kappa \, d_S(f,g)\) and \(\| h_u \circ f\| _{S,\infty } \le R_H R_\Phi \) for \(f, g \in B_k\).
Hilbert output layers satisfy the increment condition. Let \(\mathcal H\) be a real Hilbert space and \(\Phi : \mathcal X \to \mathcal H\) be \(L\)-Lipschitz. Then Assumption ass:sg-increment-main holds for \(H = H_\Phi \) with \(A_H = 1\) (and the same \(L\)), for every hidden-layer class \(\mathfrak F\).
Hilbert output layers satisfy the increment condition, given the vector Hoeffding inequality. Let \(\mathcal H\) be a real inner product space and \(\Phi : \mathcal X \to \mathcal H\) be \(L\)-Lipschitz. Assume the Rademacher tail bound lem:rademacher-hilbert-tail holds in \(\mathcal H\) for samples of size \(n\): \(\mathbb P_\sigma (\| \sum _i \sigma _i v_i\| {\gt} t) \le 2\exp (-t^2/(2\sum _i\| v_i\| ^2))\). Then Assumption ass:sg-increment-main holds for \(H = H_\Phi \) with \(A_H = 1\) and the same \(L\), for every hidden-layer class \(\mathfrak F\). Proof: \(|Z_f - Z_g| \le \sup _{\| w\| \le 1} |\langle w, \frac1n\sum _i \sigma _i v_i\rangle | = \| \frac1n \sum _i \sigma _i v_i\| \) with \(v_i = \Phi (f(x_i)) - \Phi (g(x_i))\), and \(\sum _i \| v_i\| ^2 \le L^2 n d_S(f,g)^2\).
One-dimensional Hilbert output layers. Let \(\Phi : \mathcal X \to \mathbb R\) be \(L\)-Lipschitz and \(H_\Phi = \{ x \mapsto w\, \Phi (x) : |w| \le 1\} \). Then Assumption ass:sg-increment-main holds for \(H_\Phi \) with \(A_H = 1\) and the same \(L\) (this is prop:hilbert-sg for \(\mathcal H = \mathbb R\), fully proved from the real Hoeffding inequality).
(Net-based implementation.) Let \(\mathcal H \subseteq \mathbb R^{\mathcal X}\) be totally bounded in the uniform norm and let \(\mathcal A \subseteq \mathbb R^{\mathcal X}\) be uniformly dense on \(\mathcal H\) (for every \(g \in \mathcal H\) and \(\eta {\gt} 0\) there is \(a \in \mathcal A\) with \(\| g - a\| _\infty \le \eta \)). Then for every \(\varepsilon {\gt} 0\) there exist a finite class \(\mathcal H_{\mathrm{imp}}^\varepsilon \subseteq \mathcal A\) with \(|\mathcal H_{\mathrm{imp}}^\varepsilon | \le N(\mathcal H, \| \cdot \| _\infty , \varepsilon /2)\) and an implementation map \(\iota : \mathcal H \to \mathcal H_{\mathrm{imp}}^\varepsilon \) with \(\sup _{g \in \mathcal H} \| g - \iota (g)\| _\infty \le \varepsilon \). (The paper’s compact domain \(K\) and the continuity of the functions are only used to guarantee total boundedness and density, which are the hypotheses here.)
(Layerwise implementation.) Let \(F \subseteq \mathcal X^{\mathcal X}\) with \(\mathrm{lip}(f) \le \Lambda \) for all \(f \in F\), and \(H \subseteq \mathbb R^{\mathcal X}\) with \(\mathrm{lip}(h) \le L_H\) for all \(h \in H\). Suppose every \(f \in F\) is assigned an implemented map \(\tilde f\) with \(d_\infty (f, \tilde f) \le \delta \) (\(\delta \ge 0\)) and every \(h \in H\) an implemented \(\tilde h\) with \(\| h - \tilde h\| _\infty \le \delta _H\). Let \(\iota \) be an implementation map that sends each \(g \in \mathcal H_k\), for one representation \(g = h \circ f_m \circ \cdots \circ f_1\) with \(m \le k\), to \(\tilde h \circ \tilde f_m \circ \cdots \circ \tilde f_1\). Then
(Stated for an arbitrary \(\iota \) with this property; such an \(\iota \) exists by ‘lem:impl-map-exists‘.)
(Exact implementation of affine transitions.) Let \(\mathcal X = \mathbb R^d\) and let every \(f \in F\) be affine, \(f(x) = Ax + b\). Then every element of \(B(k,F)\) is a ReLU network of depth \(\le k\) and width \(2d\), and with the identity implementation map \(\iota = \mathrm{id}\) (the abstract class \(\mathcal H_k\) is itself the implemented class) one has \(\varepsilon _{\mathrm{imp}}(k) = 0\).
(Map-dependent finite-set interpolation criterion for linear heads.) Fix \(f_1, \dots , f_M \in B_k\) and a sample \(S = (x_i)_{i=1}^n\), and let \(\Phi : \mathcal X \to \mathcal H\). Assume that for each \(j\) the evaluation operator \(T_j : w \mapsto (\langle w, \Phi (f_j(x_i))\rangle )_{i \in [n]}\) admits a right inverse of norm \(\le \Lambda _j\): for every \(c \in \mathbb R^n\) there is \(w \in \mathcal H\) with \(\| w\| \le \Lambda _j \| c\| _2\) and \(\langle w, \Phi (f_j(x_i))\rangle = c_i\) for all \(i\). Then for every family of code vectors \(u^{(1)}, \dots , u^{(M)} \in \mathbb R^n\) with \(\Lambda _j \| u^{(j)}\| _2 \le R_H\) there are output layers \(h_j \in H_{R_H}(\Phi )\) with \(h_j(f_j(x_i)) = u^{(j)}_i\). Consequently, if \(\bigl(\frac1n \sum _i |u^{(j)}_i - u^{(\ell )}_i|^2\bigr)^{1/2} \ge \rho \) for \(j \ne \ell \) then \(\| h_j \circ f_j - h_\ell \circ f_\ell \| _S \ge \rho \) for \(j \ne \ell \), and if \(\max _{j,i} |u^{(j)}_i| \le R_{\mathrm{out}}\) (with \(R_{\mathrm{out}} \ge 0\)) then \(\| h_j \circ f_j\| _{S,\infty } \le R_{\mathrm{out}}\).
(Fixed-point refinement, rigorous form.) Let \(K \subseteq \mathbb R^d\) (a separable Hilbert space) be compact with \(\| x\| \le M_K\) on \(K\) and \(D_K = \mathrm{diam}(K)\), let \(F_{\rm fp}\) be the projected gradient steps of ‘lem:fp-contraction‘ with \(0 {\lt} h \le 2/(\mu +\Lambda )\), \(H = H_R\) (\(R {\gt} 0\)), and let the target class be \(\mathcal C = \overline{H_R \circ \langle F_{\rm fp}\rangle }^{\, d_\infty }\). Assume \(\mathsf V_\infty (F_{\rm fp}) {\lt} \infty \). For a measurable loss \(\ell : \mathbb R \times \mathcal Y \to [0,b]\), \(\beta _\ell \)-Lipschitz in its first argument, \(n \ge 1\), \(\eta \ge 0\), \(\delta \in (0,1)\): with probability at least \(1 - \delta \) over \(\mathcal D \sim P^{\otimes n}\), every \(\eta \)-empirical minimizer \(\hat h \in H_R \circ B(k, F_{\rm fp})\) satisfies
with \(\mathrm{dev}_{\ell ,n,\delta }(B) = 4\beta _\ell B + 6b\sqrt{2\log (4/\delta )/n}\) (‘def:bv-dev‘): the EL regime with \(\alpha = \log (1/(1-h\mu )) \ge h\mu \) and saturated variance.
(Fixed-horizon integration, rigorous form, \(p = 1\).) Let \(K \subseteq \mathbb R^d\) (a separable Hilbert space) be compact with \(\| x\| \le M_K\) on \(K\), let every drift in \(\mathcal S_T\) be \(\Lambda _s\)-Lipschitz in \(x\), \(H = H_R\) (\(R {\gt} 0\)), and let the target class be the flow readouts \(\mathcal C = \{ \langle w, \Phi _T(\cdot )\rangle : \| w\| \le R\} \) of a teacher \(s^\ast \in \mathcal S_T\) as in ‘lem:ode-horizon-bias‘ (projection inactive, \(k \ge 1\), \(T/k \le h_0\)). Assume the entropy condition \(\mathsf V_\infty {\lt} \infty \) of ‘lem:ode-saturation-profile‘. For a measurable loss \(\ell : \mathbb R \times \mathcal Y \to [0,b]\), \(\beta _\ell \)-Lipschitz in its first argument, \(n \ge 1\), \(\eta \ge 0\), \(\delta \in (0,1)\): with probability at least \(1 - \delta \) over \(\mathcal D \sim P^{\otimes n}\), every \(\eta \)-empirical minimizer \(\hat h \in H_R \circ B_T(k)\) satisfies
with \(C_E = (\Lambda _s M_s + \Lambda _\tau )e^{\Lambda _s T}/2\) and \(\mathrm{dev}_{\ell ,n,\delta }(B) = 4\beta _\ell B + 6b\sqrt{2\log (4/\delta )/n}\) (‘def:bv-dev‘): the PL regime with \(\beta = p = 1\) and saturated variance.
(Saturation.) Let \(A_k\) be sets with diameters \(0 \le D_k \le \overline D\) and \(N(A_k, \varepsilon ) \le N_\infty (\varepsilon )\) for all \(k\) and \(\varepsilon \), with \(N_\infty (\varepsilon ) {\lt} \infty \) for \(\varepsilon {\gt} 0\) and \(\mathsf V_\infty := \int _0^{\overline D} \sqrt{\log N_\infty (\varepsilon )}\, d\varepsilon {\lt} \infty \) (the integrand is interval-integrable). Then \(\mathsf V(D_k, A_k) \le \mathsf V_\infty \) for all \(k\).
(Polynomial growth, bounded diameter.) If \(0 \le D_k \le \overline D\), \(\overline D {\gt} 0\), and \(N(A_k, \varepsilon ) \le C_0 (1 + k/\varepsilon )^{D}\) for all \(k \ge 1\) and \(\varepsilon {\gt} 0\) (with \(C_0 \ge 1\), \(D \ge 0\)), then for \(k \ge 1\)
(Exponential growth, bounded diameter.) If \(0 \le D_k \le \overline D\) and \(\log N(A_k, \varepsilon ) \le \alpha k + \psi (\varepsilon )\) for all \(k\) and \(\varepsilon {\gt} 0\), with \(\Psi := \int _0^{\overline D} \sqrt{\psi (\varepsilon )}\, d\varepsilon {\lt} \infty \) (i.e. \(\sqrt\psi \) is interval-integrable on \([0,\overline D]\)), then \(\mathsf V(D_k, A_k) \le \overline D\sqrt{\alpha k} + \Psi = O(\sqrt k)\).
(Polynomial growth, linearly growing diameter.) If \(0 \le D_k \le D_1 k\) with \(D_1 {\gt} 0\), and \(N(A_k, \varepsilon ) \le C_0 (1 + k/\varepsilon )^{D}\) for all \(k \ge 1\) and \(\varepsilon {\gt} 0\) (with \(C_0 \ge 1\), \(D \ge 0\)), then for \(k \ge 1\)
(Depth profiles of ReLU networks by Lipschitz constant, upper bounds.) Let \(K \ne \emptyset \) be bounded with \(\| x\| \le R_K\) on \(K\) (\(R_K \ge 0\)), \(F = F_\Lambda \), \(p = 2mw + w + m\) and \(C_F\) as in ‘lem:relu-layer-covering‘.
(Contractive, \(0 {\lt} \Lambda {\lt} 1\).) For every \(\varepsilon {\gt} 0\) and \(k\), \(N^{\mathrm{ext}}(B(k,F), d_\infty , \varepsilon ) \le N^{\mathrm{ext}}(K, \varepsilon /2) + m(\varepsilon ) N^{\mathrm{ext}}(F, d_\infty , (1-\Lambda )\varepsilon )^{m(\varepsilon )} {\lt} \infty \) with \(m(\varepsilon ) = \lceil \log (2D_K/\varepsilon )/\log (1/\Lambda )\rceil \), so \(\sup _k N^{\mathrm{ext}}(B(k,F), d_\infty , \varepsilon ) {\lt} \infty \).
(Non-expanding, \(\Lambda \le 1\).) If \(K\) is compact then for every \(\varepsilon {\gt} 0\) and \(k\), \(N^{\mathrm{ext}}(B(k,F), d_\infty , \varepsilon ) \le N^{\mathrm{ext}}(\overline{\langle F\rangle }, d_\infty , \varepsilon ) {\lt} \infty \); and for \(D_K \le \overline D\), \(\overline D {\gt} 0\), \(k \ge 1\), the envelope gives \(\mathsf V_k(S) \le \overline D(\sqrt{\log (k+1)} + \sqrt{kp\log k} + \sqrt{kp}(\sqrt{\log (1 + 2C_F/\overline D)} + \sqrt\pi /2)) = O(\sqrt{kp\log k})\).
(Expanding, \(\Lambda {\gt} 1\).) For \(D_K \le \overline D\), \(\overline D {\gt} 0\), \(k \ge 1\), \(\mathsf V_k(S) \le \overline D(\sqrt{\log (k+1)} + \sqrt{kp\log k} + k\sqrt{p\log \Lambda } + \sqrt{kp}(\sqrt{\log (1 + 2C_F/\overline D)} + \sqrt\pi /2)) = O(k\sqrt{p\log \Lambda })\).
The lower bound of (iii) is ‘prop:relu-regimes-iii-lower‘; the saturated profiles \(\mathsf V_k(S) = O(1)\) of (i) and (ii) are ‘prop:relu-regimes-i-profile‘ and ‘prop:relu-regimes-ii-profile‘.
(‘prop:relu-regimes‘(i): contractive layers, \(\Lambda {\lt} 1\).) Let \(K \ne \emptyset \) be bounded with \(\| x\| \le R_K\) on \(K\), and \(0 {\lt} \Lambda {\lt} 1\). Then ‘cond:p1-ucont‘ applies with the invariant set \(K\) itself (\(L = 0\), absorbing set \(K\)), and with \(m(\varepsilon ) = \lceil \log (2D_K/\varepsilon )/\log (1/\Lambda )\rceil \) (\(D_K = \mathrm{diam}\, K\)), for every \(\varepsilon {\gt} 0\) and every \(k\),
which does not depend on \(k\); in particular \(\sup _k N^{\mathrm{ext}}(B(k,F_\Lambda ), d_\infty , \varepsilon ) {\lt} \infty \). (The paper’s \(k \ge m(\varepsilon )\) is not needed, cf. ‘cond:p1-ucont‘; the right-hand side is finite by ‘lem:relu-layer-covering‘ and the total boundedness of \(K\).)
(Case (i), saturated profile: \(\mathsf V_k(S) = O(1)\).) Under the hypotheses of ‘prop:relu-regimes-i‘, if \(\varepsilon \mapsto \sqrt{\log N_\infty (\varepsilon )}\) is interval-integrable on \([0, D_K]\) then \(\mathsf V_k(S) \le \mathsf V_\infty := \int _0^{D_K}\sqrt{\log N_\infty (\varepsilon )}\, d\varepsilon \) for all \(k\) and all samples \(S\) (case (i) of ‘prop:profiles‘).
(‘prop:relu-regimes‘(ii): non-expanding layers, \(\Lambda \le 1\).) If \(K\) is compact and \(\Lambda \le 1\) then ‘cond:p1‘ (2c, non-expanding generators on a compact state space) applies: for every \(\varepsilon {\gt} 0\) and every \(k\), \(N^{\mathrm{ext}}(B(k,F_\Lambda ), d_\infty , \varepsilon ) \le N^{\mathrm{ext}}(\overline{\langle F_\Lambda \rangle }, d_\infty , \varepsilon ) {\lt} \infty \); in particular \(\sup _k N^{\mathrm{ext}}(B(k,F_\Lambda ), d_\infty , \varepsilon ) {\lt} \infty \).
(Case (ii), explicit envelope: \(\mathsf V_k(S) = O(\sqrt{kp\log k})\).) For \(\Lambda \le 1\), under the hypotheses of ‘lem:relu-envelope-profile‘,
(Case (ii), saturated profile: \(\mathsf V_k(S) = O(1)\).) Under the hypotheses of ‘prop:relu-regimes-ii‘, with \(N_\infty (\varepsilon ) := N(\overline{\langle F_\Lambda \rangle }, d_\infty , \varepsilon /2) {\lt} \infty \): if \(\int _0^{D_K}\sqrt{\log N_\infty (\varepsilon )}\, d\varepsilon {\lt} \infty \) then \(\mathsf V_k(S) \le \int _0^{D_K}\sqrt{\log N_\infty (\varepsilon )}\, d\varepsilon \) for all \(k\) and all samples \(S\) (case (i) of ‘prop:profiles‘, as in ‘lem:fp-profile‘).
(‘prop:relu-regimes‘(iii), upper bound: expanding layers, \(\Lambda {\gt} 1\): \(\mathsf V_k(S) = O(k\sqrt{p\log \Lambda })\).) Under the hypotheses of ‘lem:relu-envelope-profile‘, for \(\Lambda {\gt} 1\),
(‘prop:relu-regimes‘(iii), lower bound.) For \(m = 1\), \(K = [0,1]\) with the clipping retraction, width \(w \ge 4\), \(\Lambda \ge 20\), and \(\beta _W \ge 41\), \(\beta \ge 2\) (at least the weight norms of the representations of the two expand-and-reset maps), the class \(F_\Lambda \) contains \(g_0, g_1\), which satisfy ‘cond:e2-pingpong‘ with two generators and separation \(1/8\); hence
(The paper states \(w \ge 5\); the two maps have three interior breakpoints each, so width \(4\) suffices.)
Implementation-free bias–variance decomposition. Fix \(k \ge 0\), \(\eta \ge 0\), \(n \ge 1\). Assume \(\mathcal H_k = H \circ B(k,F)\) consists of measurable functions, is sup-norm separable and pointwise bounded, \(\mathcal C\) is a nonempty class of measurable benchmarks, the loss \(\ell : \mathbb R \times \mathcal Y \to [0,b]\) is measurable and \(\beta _\ell \)-Lipschitz in its first argument (\(b {\gt} 0\), \(\beta _\ell \ge 0\)), the implementation map \(\iota \) produces measurable functions, and \(\varepsilon _{\mathrm{imp}}(k) = \sup _{f \in \mathcal H_k} \| f - \iota f\| _\infty \le \varepsilon _{\mathrm{imp}}\) for a real \(\varepsilon _{\mathrm{imp}} \ge 0\). Then for every \(\delta \in (0,1)\), with probability at least \(1-\delta \) over \(\mathcal D \sim P^{\otimes n}\), every \(\eta \)-empirical minimizer \(\hat f \in \mathcal H_k\) satisfies, with \(\hat h := \iota (\hat f)\),
This is thm:bv-general with \(d_T(f,g) = \| f - g\| _\infty \) and \(\beta _L = \beta _{\hat L} = \beta _\ell \) (the loss is \(\beta _\ell \)-Lipschitz); the Rademacher constant \(4\beta _\ell \) is the paper’s and the deviation constant is explicit (as in thm:bv-general); \(\mathcal H_k\) is assumed pointwise bounded.
Implementation-free generalization gap. Under the hypotheses of thm:bv, with probability at least \(1-\delta \) over \(\mathcal D \sim P^{\otimes n}\), every \(f \in \mathcal H_k\) (in particular every \(\eta \)-empirical minimizer) satisfies, with \(h := \iota (f)\),
(The Rademacher constant \(2\beta _\ell \) is the paper’s; the deviation constant is explicit, as in thm:bv-general-gap.)
General bias–variance decomposition (excess risk). Let \(\mathcal H\) be a sup-norm separable, pointwise bounded class of measurable functions \(\mathcal X \to \mathbb R\), \(\mathcal C\) a nonempty class of measurable benchmark functions, and \(\ell : \mathbb R \times \mathcal Y \to [0,b]\) measurable and \(\beta _\ell \)-Lipschitz in its first argument (\(b {\gt} 0\), \(\beta _\ell \ge 0\)). Let \(\iota \) be an implementation map, \(d_T\) a nonnegative function (pseudo-metric) and assume, for constants \(\beta _L, \beta _{\hat L} \ge 0\) and \(\varepsilon _{\mathrm{imp}} \ge 0\) (any upper bound for \(\sup _{f\in \mathcal H} d_T(\iota f, f)\)), that for every \(f \in \mathcal H\): \(d_T(\iota f, f) \le \varepsilon _{\mathrm{imp}}\), \(|L[\iota f] - L[f]| \le \beta _L d_T(\iota f, f)\) and \(|\hat L[\iota f] - \hat L[f]| \le \beta _{\hat L} d_T(\iota f, f)\) for every sample. Let \(n \ge 1\), \(\eta \ge 0\) and \(\delta \in (0,1)\). Then with probability at least \(1 - \delta \) over \(\mathcal D \sim P^{\otimes n}\), every \(\eta \)-empirical minimizer \(\hat f \in \mathcal H\) satisfies, with \(\hat h := \iota (\hat f)\),
The Rademacher constant \(4\beta _\ell \) is the paper’s; the deviation constant is explicit (\(C b\sqrt{\log (1/\delta )/n}\) in the paper, with unspecified universal \(C\), becomes \(6b\sqrt{2\log (4/\delta )/n}\), from lem:bv-uniform-deviation-onesided). The pointwise boundedness of \(\mathcal H\) makes \(\hat{\mathfrak R}_S(\mathcal H)\) a genuine supremum.
General bias–variance decomposition (generalization gap). Under the hypotheses of thm:bv-general, with probability at least \(1-\delta \) over \(\mathcal D \sim P^{\otimes n}\), every \(f \in \mathcal H\) (in particular every \(\eta \)-empirical minimizer \(\hat f\)) satisfies, with \(h := \iota (f)\),
(The paper states this for the \(\eta \)-empirical minimizer only, with \(2\beta _\ell \hat{\mathfrak R}_S(\mathcal H) + Cb\sqrt{\log (1/\delta )/n}\); its proof gives it for every \(f \in \mathcal H\), which is what we state. The deviation constant is explicit, see thm:bv-general.)
(‘thm:bv‘ with \(\iota = \mathrm{id}\) and explicit bounds.) Let \(\mathcal H\) be a sup-norm separable, pointwise bounded class of measurable functions, \(\mathcal C \ne \emptyset \) a class of measurable benchmarks, \(\ell \) measurable, bounded by \(b {\gt} 0\) and \(\beta _\ell \)-Lipschitz, \(n \ge 1\), \(\eta \ge 0\), \(\delta \in (0,1)\). Suppose \(\hat{\mathfrak R}_S(\mathcal H) \le B\) for every sample \(S\) of size \(n\) and \(\varepsilon _{\mathrm{model}} \le \mathrm{bias}\). Then with probability at least \(1 - \delta \) over \(\mathcal D \sim P^{\otimes n}\), every \(\eta \)-empirical minimizer \(\hat f \in \mathcal H\) satisfies
Hidden–output decomposition under a sub-Gaussian increment condition. Let \(n \ge 1\), let \(F \subseteq \mathcal X^{\mathcal X}\) be totally bounded in \(d_S\), let \(f_0 \in F\) be an arbitrary anchor, and suppose Assumption ass:sg-increment-main holds on \(\mathfrak F \supseteq F\) with constants \(A_H, L {\gt} 0\). (We also assume that the suprema defining \(Z_f(\sigma )\), \(f \in F\), are finite, i.e. the Rademacher averages over \(H\) are bounded above; this is implicit in the paper.) Then
where \(\hat{\mathfrak R}_S(H \circ \{ f_0\} ) = \mathbb E_\sigma Z_{f_0}(\sigma )\) (lem:emp-rademacher-comp-singleton). This generalizes the paper, which assumes \(\mathrm{id} \in F\) and anchors at \(f_0 = \mathrm{id}\), giving \(\hat{\mathfrak R}_S(H)\) on the right-hand side (thm:hidden-decomp-id); the anchor is arbitrary because the Dudley bound thm:dudley-subgaussian-finite-space is anchored at any point. We assume in addition that the entropy integrand \(\varepsilon \mapsto \sqrt{\log N(F,d_S,\varepsilon )}\) is integrable on \([0,\mathrm{diam}_S(F)]\); the paper’s bound is trivially true with right-hand side \(+\infty \) otherwise, whereas Lean’s Bochner integral of a non-integrable function is \(0\).
If Assumption ass:sg-increment-main holds on the semigroup \(\langle F_0 \rangle \) with constants \((A_H, L)\), then for every depth \(k \ge 0\) with \(B(k,F_0)\) totally bounded in \(d_S\),
with the same constants \(A_H, L\) (assuming, as in thm:hidden-decomp, that the entropy integrand of \(B(k,F_0)\) is integrable on \([0, \mathrm{diam}_S(B(k,F_0))]\)).
Hidden–output decomposition, anchored at the identity (the paper’s form). Under the hypotheses of thm:hidden-decomp with \(\mathrm{id} \in F\),
Proof: thm:hidden-decomp with \(f_0 = \mathrm{id}\) and \(H \circ \{ \mathrm{id}\} = H\).
(Metric modulus criterion.) Let \((\mathcal X, d)\) be a totally bounded (pseudo)metric space. If \(H \subseteq \mathcal X^{\mathcal X}\) has a common modulus of continuity, namely there is \(\omega : \mathbb R \to \mathbb R\) with \(\omega (r) \to 0\) as \(r \downarrow 0\) (for every \(\varepsilon {\gt} 0\) there is \(r {\gt} 0\) with \(\omega (t) \le \varepsilon \) for \(0 \le t \le r\)) and
then \(H\) is totally bounded in \(d_\infty \).
Deterministic entropy decomposition. The paper states: under Assumptions ass:ent-readout and ass:ent-transition, for every sample \(S = (x_1,\dots ,x_n)\) with \(n \ge 1\),
The formalized statement is: under Assumptions ass:ent-readout and ass:ent-transition with \(L_H {\gt} 0\) and \(B_H {\gt} 0\), if the entropy integrand \(x \mapsto \mathcal E_H(x/4) + \mathcal E_F(x/(4L_H))\) is integrable on \([0, B_H/2]\), then for every sample \(S\) of size \(n \ge 1\),
where \(\mathcal E_H(u) = \sqrt{\log N^{\mathrm{ext}}(H, \| \cdot \| _\infty , u)}\) and \(\mathcal E_F(v) = \sqrt{\log N^{\mathrm{ext}}(F, d_\infty , v)}\) are the root entropies for the external covering numbers by closed balls of the assumptions.
Deviations from the paper. (i) We assume in addition that the entropy integrand is integrable on \([0, B_H/2]\); the paper’s bound is trivially true with right-hand side \(+\infty \) otherwise (e.g. Lipschitz classes on a two-dimensional domain), whereas Lean’s Bochner integral of a non-integrable function is \(0\), so the statement without this hypothesis is false as transcribed. (ii) FoML’s Dudley integral (dudley_entropy_integral’) uses internal open-ball covers, and converting the external closed-ball covering numbers of the assumptions costs a factor \(2\) in the radius (lem:foml-covering-le-external-unifFun), so the scales are \(x/4\) and \(x/(4L_H)\) on \([0, B_H/2]\) instead of \(\varepsilon /2\), \(\varepsilon /(2L_H)\) on \([0, 2B_H]\). Up to these constant rescalings the statement is the paper’s.
Proof: let \(\varepsilon \to 0\) in thm:rad.decomp.ent.ent-foml, using the integrability hypothesis to compare the integrals over \([\varepsilon , B_H/2]\) and \([0, B_H/2]\).
Deterministic entropy decomposition (FoML form). Under Assumptions ass:ent-readout and ass:ent-transition with \(L_H {\gt} 0\), for every sample \(S\) of size \(n \ge 1\) and every \(0 {\lt} \varepsilon {\lt} B_H/2\),
where \(\mathcal E_H(u) = \sqrt{\log N^{\mathrm{ext}}(H, \| \cdot \| _\infty , u)}\) and \(\mathcal E_F(v) = \sqrt{\log N^{\mathrm{ext}}(F, d_\infty , v)}\) (external covering numbers by closed balls). Proof: FoML’s Dudley integral for the class \(H \circ F\) with the empirical pseudometric \(\| \cdot \| _S\) (dudley_entropy_integral’), whose internal open-ball covering number at scale \(x\) is at most \(N^{\mathrm{ext}}(H \circ F, \| \cdot \| _\infty , y/2)\) for every \(y {\lt} x\) (lem:foml-covering-le-external-unifFun, losing a factor \(2\) in the radius when passing from external to internal covers), followed by the composition covering lemma lem:ent-composition-entropy at scale \(y/2\) and the limit \(y \uparrow x\) almost everywhere (lem:entropy-integrand-ae). Compared with the paper’s statement (quoted in thm:rad.decomp.ent.ent), the scales \(x/2\), \(x/(2L_H)\) are replaced by \(x/4\), \(x/(4L_H)\) (the paper implicitly uses covers with centres in the class, while the assumptions here are stated with external covering numbers), the integration range is \([\varepsilon , B_H/2]\) instead of \([0, 2B_H]\), and there is the additive term \(4\varepsilon \) from FoML’s Dudley bound.
Deterministic entropy decomposition on the sample. Let \(S\) be a sample of size \(n \ge 1\), and suppose ass:ent-readout-sample and ass:ent-transition-sample hold with \(L_H {\gt} 0\), \(B_H {\gt} 0\). If the entropy integrand \(x \mapsto \mathcal E_{H,S}(x/4) + \mathcal E_{F,S}(x/(8L_H))\) is integrable on \([0, B_H/2]\), then
where \(\mathcal E_{H,S}(u) = \sqrt{\log N_{S,F}(H, u)}\) with \(N_{S,F}(H, u) = \sup _{f \in F} N^{\mathrm{ext}}(H, \| \cdot \| _{f \circ S}, u)\), and \(\mathcal E_{F,S}(v) = \sqrt{\log N^{\mathrm{ext}}(F, d_S, v)}\).
Comparison with the paper (thm:rad.decomp.ent.ent, Appendix D). The paper assumes \(H\) uniformly \(L_H\)-Lipschitz and bounded on \(\mathcal X\) and covered in the sup norm, and \(F\) covered in \(d_\infty \); its proof passes from \(\| \cdot \| _S\) to \(\| \cdot \| _\infty \) at the very first step. Here (W8 of the plan) every hypothesis is on the sample only: \(H\) is bounded and \(L_H\)-Lipschitz on the reachable points \(\mathcal R_S(F) = \{ f(x_i)\} \), \(F\) is covered in \(d_S\), and \(H\) is covered in the empirical metrics \(\| \cdot \| _{f \circ S}\) of the pushed-forward samples, uniformly in \(f \in F\) (this is the honest sample analogue of the sup norm cover: in the composition estimate the output layer is evaluated at \(f_b(x_i)\), not at \(x_i\), so a cover of \(H\) for \(\| \cdot \| _S\) alone would not suffice). The sup-norm assumptions imply the sample assumptions for every \(S\) (cor:rad.decomp.ent.ent-of-sample). The scales are \(x/4\), \(x/(8L_H)\) on \([0, B_H/2]\) in place of the paper’s \(\varepsilon /2\), \(\varepsilon /(2L_H)\) on \([0, 2B_H]\): one factor \(2\) comes from FoML’s internal open-ball covers (as in thm:rad.decomp.ent.ent) and, for \(F\) only, a second factor \(2\) from the fact that the cover of \(F\) must have its centres in \(F\) (so that the pushed-forward samples \(f_b \circ S\) are admissible), see lem:ent-composition-covering-sample. As in the sup-norm theorem, integrability of the integrand is an additional hypothesis.
Proof: let \(\varepsilon \to 0\) in thm:rad.decomp.ent.ent-sample-foml.
Deterministic entropy decomposition on the sample (FoML form). Under Assumptions ass:ent-readout-sample and ass:ent-transition-sample with \(L_H {\gt} 0\), for every sample \(S\) of size \(n \ge 1\) and every \(0 {\lt} \varepsilon {\lt} B_H/2\),
where \(\mathcal E_{H,S}(u) = \sqrt{\log N_{S,F}(H, u)}\) and \(\mathcal E_{F,S}(v) = \sqrt{\log N^{\mathrm{ext}}(F, d_S, v)}\). Proof: FoML’s Dudley integral for \(H \circ F\) with the pseudometric \(\| \cdot \| _S\) (dudley_entropy_integral’); its internal open-ball covering number at scale \(x\) is at most \(N^{\mathrm{ext}}(H \circ F, \| \cdot \| _S, y/2)\) for every \(y {\lt} x\) (lem:foml-covering-le-external-empFun), then the composition covering lemma on the sample lem:ent-composition-entropy-sample at scale \(y/2\) and the limit \(y \uparrow x\) almost everywhere (lem:entropy-integrand-ae-sample).
Conditional Sudakov-type lower bound. There is a universal constant \(c {\gt} 0\) such that the following holds. Let \((\mathcal X, d)\) be a pseudometric space, \(S = (x_1,\dots ,x_n)\) a sample with \(n \ge 1\), \(H\) an output-layer class, \(B \subseteq \mathcal X^{\mathcal X}\) an arbitrary hidden class, and suppose Assumption ass:readout-realization-main holds for \(B\) with constants \(\kappa , R_{\mathrm{out}} {\gt} 0\). Then for every \(\varepsilon {\gt} 0\),
equivalently the bound with \(\sup _{\varepsilon {\gt} 0}\) on the right. (In Lean \(\log M\) is read as \(0\) when \(M = \infty \), which only weakens the inequality.) This generalizes the paper, which states the bound for the word ball \(B = B(k,F)\) (thm:sudakov-type-wordball); the word-ball structure is not used in the proof, and \(B = \emptyset \) is allowed (both sides are then \(0\)). In Lean we add the hypothesis that, for every sign pattern, the Rademacher averages over \(H \circ B\) are bounded above (equivalently, that the supremum defining \(\hat{\mathfrak R}_S(H \circ B)\) is finite); without it the Lean statement is false because an unbounded supremum evaluates to \(0\). The proof is complete modulo the Bernoulli–Sudakov minoration thm:bernoulli-sudakov. Proof: take a maximal \(2\varepsilon \)-packing \(g_1, \dots , g_M\) of \(B\) for \(d_S\) (\(M = M(B, d_S, 2\varepsilon )\); if \(M = \infty \) the right-hand side is \(0\) in Lean and the claim is lem:emp-rademacher-nonneg); the transported vectors \(u_j = (h_{g_j}(g_j(x_i)))_i\) are \(2\kappa \varepsilon \)-separated in \(\| \cdot \| _S\) and bounded by \(R_{\mathrm{out}}\), so thm:bernoulli-sudakov gives \(\mathbb E_\sigma \max _j \frac1n\sum _i \sigma _i u_{j,i} \ge c\min \{ 2\kappa \varepsilon \sqrt{\log M/n}, 4\kappa ^2\varepsilon ^2/R_{\mathrm{out}}\} \), and the left-hand side is at most \(\hat{\mathfrak R}_S(H \circ B)\) since \(\{ h_{g_j} \circ g_j\} \subseteq H \circ B\).
Conditional Sudakov-type lower bound for word balls (the paper’s form). There is a universal constant \(c {\gt} 0\) such that, for \(F\) a hidden-layer class, \(k \ge 0\) and Assumption ass:readout-realization-main for \(B_k = B(k,F)\) with constants \(\kappa , R_{\mathrm{out}} {\gt} 0\) (and the boundedness hypothesis of thm:sudakov-type), for every \(\varepsilon {\gt} 0\),
Proof: thm:sudakov-type with \(B = B(k,F)\), since \(\mathcal H_k = H \circ B(k,F)\).
EP regime (??). Let \(\alpha ,\gamma {\gt} 0\) and \({k^{\ast }}(n) = \frac{1}{2\alpha }(\log n - \gamma \log \log n)\). Then \(\mathsf{gen}({k^{\ast }},n) = e^{-\alpha {k^{\ast }}} + \sqrt{{k^{\ast }}^{\gamma }/n} \asymp n^{-1/2}(\log n)^{\gamma /2}\) as \(n \to \infty \) (more precisely, between \(1\) and \(1 + (2\alpha )^{-\gamma /2}\) times the rate).
PP regime (??). Let \(\beta ,\gamma {\gt} 0\) and \(n {\gt} 0\). At the balancing depth \({k^{\ast }}= n^{1/(2\beta +\gamma )}\), \(\mathsf{gen}({k^{\ast }},n) = {k^{\ast }}^{-\beta } + \sqrt{{k^{\ast }}^{\gamma }/n} = 2\, n^{-\beta /(2\beta +\gamma )}\), and for every depth \(k {\gt} 0\), \(\mathsf{gen}(k,n) = k^{-\beta } + \sqrt{k^{\gamma }/n} \ge n^{-\beta /(2\beta +\gamma )}\). Thus \({k^{\ast }}\) is optimal up to the factor \(2\) and \(\mathsf{gen}({k^{\ast }},n) \asymp n^{-\beta /(2\beta +\gamma )}\).