- Boxes
- definitions
- Ellipses
- theorems and lemmas
- Blue border
- the statement of this result is ready to be formalized; all prerequisites are done
- Orange border
- the statement of this result is not ready to be formalized; the blueprint needs more work
- Blue background
- the proof of this result is ready to be formalized; all prerequisites are done
- Green border
- the statement of this result is formalized
- Green background
- the proof of this result is formalized
- Dark green background
- the proof of this result and all its ancestors are formalized
- Dark green border
- this is in Mathlib
Assume the setting of Theorem 328, (SC\(_a\)) with \(0{\lt}a\le 1\), the complexity bounds of Lemmas 314 and 315, and (S\(_\infty \)) for the family \(\lambda \le 1\), \(\beta =\beta _0\) with constant \(B_\infty \); let \(C_1\) be Definition 145 with \(\kappa _0=1/\beta _0\), and \(D'=D'(\delta )\), \(D''=D''(\delta )\) as in ??. Put \(\lambda _1:=\min \{ 1,\sqrt{\beta _0}/B_\infty \} \) and \(\lambda _N:=N^{-1/(2(a+2))}\). Then for \(N\ge N_{(i)}:=\max \{ \lambda _1^{-2(a+2)},(64D''{}^2)^{(a+2)/a}\} \), with probability at least \(1-\delta \),
the rate is \(N^{-1/6}\) for \(a=1\) and \(N^{-1/10}\) for \(a=1/2\). (For \(\lambda \le \lambda _1\), \(c_1(B_\infty )=1/\lambda \) and \(E_*=B_\infty +{Y_{\max }}\) in Theorem 380, so \(\sup |m_N-m^*|\le 12(B_\infty +{Y_{\max }})D'\lambda ^{-2}N^{-1/2}\); the bias is Corollary 473; balance \(\lambda ^{-2}N^{-1/2}=\lambda ^a\).)
Assume the setting of Theorem 328, (SC\(_a\)) with \(0{\lt}a\le 1\) and the complexity bounds of Lemmas 314 and 315; let \(D'=D'(\delta )\), \(D''=D''(\delta )\) be as in ??. Put \(c_u:=\frac12e^{\| u^\dagger \| ^2/4}\| u^\dagger \| \), \(\lambda _2:=\min \{ 1,c_u^{-2}\} \), \(\lambda _N:=N^{-1/(2a+5)}\) and \(\beta _N:=\lambda _N^{-a}\). Then for \(N\ge N_{(ii)}:=\max \{ \lambda _2^{-(2a+5)},(64D''{}^2)^{(2a+5)/(2a+1)}\} \), with probability at least \(1-\delta \),
where \(C_2\) is Definition 147 with \(\beta _0=\kappa _0=1\); the rate is \(N^{-1/7}\) for \(a=1\). ((SC\(_a\)) contains (R), so Lemma 467 with \(\zeta _0=\frac\kappa 2\| u^\dagger \| ^2\le \frac12\| u^\dagger \| ^2\) gives \(M_*\le c_u\lambda ^{-1/2}\); for \(\lambda \le \lambda _2\) and \(\beta _N\ge 1\), \(c_1(M_*)=1/\lambda \) and \(E_*\le (c_u+{Y_{\max }})\lambda ^{-1/2}\), so \(\sup |m_N-m^*|\le 12(c_u+{Y_{\max }})D'\lambda ^{-5/2}N^{-1/2}\); the bias is Corollary 476 with \(\beta _0=\kappa _0=1\); balance \(\lambda ^{-5/2}N^{-1/2}=\lambda ^a\).)
Assume the setting of Theorem 328, (SC\(_a\)) with \(0{\lt}a\le 1\), the complexity bounds of Lemmas 314 and 315, and let \(D'=D'(\delta )\), \(D''=D''(\delta )\) be as in ??. Put \(\lambda _N:=N^{-1/(2(a+3))}\) and \(\beta _N:=\lambda _N^{-a}\). Then for \(N^{(a+1)/(a+3)}\ge 64\max \{ 1,{Y_{\max }}\} ^4D''{}^2\) (which is \(N\ge N_B(\delta )\) along the schedule, since \(c_1\le \max \{ 1,{Y_{\max }}\} /\lambda _N\)), with probability at least \(1-\delta \),
the rate is \(N^{-1/8}\) for \(a=1\). (?? with ??: along the schedule \(c_1\le \max \{ 1,{Y_{\max }}\} /\lambda \) and \(E_*\le 2{Y_{\max }}/\lambda \), so \(\sup |m_N-m^*|\le 24\max \{ 1,{Y_{\max }}\} {Y_{\max }}D'\lambda ^{-3}N^{-1/2}\); balance \(\lambda ^{-3}N^{-1/2}=\lambda ^a\).)
Assume the setting of Theorem 328, (SC\(_a\)) with \(0{\lt}a\le 1\), the complexity bounds of Lemmas 314 and 315, and (S\(_\infty \)) for the family \(\lambda \le 1\), \(\beta =\beta _0\); let \(C_1\) be Definition 145 with \(\kappa _0=1/\beta _0\) and \(M:=\max \{ 1,{Y_{\max }}/\sqrt{\beta _0}\} \). With \(\lambda _N:=N^{-1/(2(a+3))}\), for \(N^{(a+1)/(a+3)}\ge 64M^4D''{}^2\), with probability at least \(1-\delta \),
(As Corollary 491, with \(c_1\le M/\lambda \) and the bias of Corollary 473. The formalized constants \(c_1\), \(E_*\) of Theorem 381 use \(M_*={Y_{\max }}/\lambda \), whence the rate \(N^{-a/(2(a+3))}\) in place of the paper’s \(N^{-a/(2(a+2))}\).)
Let \(0{\lt}\alpha {\lt}1\), \(\nu _0=\nu _0^{(L)}\) (\(L\ge 2\)), \((\lambda ,\beta )\) fixed and \(f\in H^1(-1,1)\), and let \(\rho ^*_L\) be the corresponding minimizer of the free energy. Then
??. (Corollary 760 with \(\| K^{(L)}-K^{(\infty )}\| \le {\varepsilon }_L\to 0\), Theorem 712, and Lemma 716. In Lean the reference law is \(\nu _0^{(L)}\) itself, so no smoothing \(\tilde\nu _0^{(L)}\) and no correction \(1/L\) of \({\varepsilon }_L\) is needed.)
Under the hypotheses of Corollary 720, with \(m^*_L\) the conditional mean amplitude of \(\rho ^*_L\), \(\limsup _L\| m^*_L\| ^2_{L^2(\nu ^*_L)}\le \| f\| ^2_{{\mathcal{H}}_\infty }\): the bounds of Theorem 397 hold asymptotically with \(f\in H^1\) in place of (R) and \(\| f\| _{{\mathcal{H}}_\infty }\) in place of \(\| u_0^\dagger \| \) (Corollary 762).
Let \((\nu _L)_L\) be hidden laws on \(Z\) whose kernel operators converge in operator norm, \(\| K_{\nu _L}-K_\infty \| \to 0\), to the kernel operator \(K_\infty =K_{\varphi ',\nu '}\) of another feature/hidden-law pair on the same \(L^2(P_X)\); let \((\lambda ,\beta )\) be fixed, \(f\in L^2(P_X)\), and \(\rho ^*_L\) the minimizer of the free energy with reference hidden law \(\nu _L\). Then
In the setting of Corollary 720, \(\nu _L=\tilde\nu _0^{(L)}\), \(K_\infty =K^{(\infty )}\) and \(Q^{(\infty )}_\lambda (f)\le \| f\| ^2_{{\mathcal{H}}_\infty }\) for \(f\in H^1(-1,1)\). (Proposition 757 for each \(L\) and Theorem 742.)
Under the hypotheses of Corollary 720, with \(\nu ^*_L\) the hidden marginal of \(\rho ^*_L\) and \(\kappa =\lambda /\beta \), \(\limsup _L\operatorname {KL}(\nu ^*_L\| \nu _0^{(L)})\le \frac\kappa 2\| f\| ^2_{{\mathcal{H}}_\infty }\); the right-hand side does not depend on \(\lambda ,\beta \) (Corollary 761).
In the setting of Theorem 328 let \(\delta \in (0,1)\), \(\Delta _N:=\Delta _N(\delta /2)\) and \(\Delta '_N\) as in 86. With probability at least \(1-\delta \) the following hold simultaneously: (a) \(\| \rho _N^*-\rho ^*\| _{\mathrm{TV}}\le \sqrt{\Delta _N/(2\beta )}\) and \(\| \nu _N^*-\nu ^*\| _{\mathrm{TV}}\le \sqrt{\Delta _N/(2\beta )}\); (b) \(\sup _z|g_N(z)-g^*(z)|\le \sqrt{2\Delta _N}+\Delta '_N\), i.e. \(\sup _z|(R^N_{\lambda ,\nu _N^*}y)(z)-(R_{\lambda ,\nu ^*}f)(z)| \le \frac1\lambda [\sqrt{2\Delta _N}+\Delta '_N]\); (c) \(\sup _x|F_{\rho _N^*}(x)-F_{\rho ^*}(x)|\le \frac1\lambda [\sqrt{2\Delta _N}+\Delta '_N] +\frac{2{Y_{\max }}}\lambda \sqrt{\Delta _N/(2\beta )}\); (d) \(\| \Pi \rho _N^*-\Pi \rho ^*\| _{\mathrm{TV}}\le \frac1{2\lambda }[\sqrt{2\Delta _N}+\Delta '_N] +\frac{{Y_{\max }}}\lambda \sqrt{\Delta _N/(2\beta )}\).
Assume the setting of Theorem 370, \(\delta \in (0,1)\) and \(N\ge N_B(\delta )\). On the event \(\Omega _\delta \) the following hold simultaneously.
\(\| \rho _N^*-\rho ^*\| _{\mathrm{TV}}\le \frac{9c_1}{2\sqrt\beta }\eta _N'\) and \(\| \nu _N^*-\nu ^*\| _{\mathrm{TV}}\le \frac{9c_1}{2\sqrt\beta }\eta _N'\).
\(\sup _z|(R^N_{\lambda ,\nu _N^*}y)(z)-(R_{\lambda ,\nu ^*}f)(z)|=\sup _z|m_N(z)-m^*(z)| \le \frac1\lambda [\| h\| _N+\eta _N']\le \frac{10c_1}\lambda \eta _N'\).
\(\sup _x|F_{\rho _N^*}(x)-F_{\rho ^*}(x)|\le D_\varsigma \le 10c_1^2\eta _N'\), since \(|h(x)|=|\int \varphi _z(x)\, \mathrm d\Pi \varsigma |\le D_\varsigma \).
\(\| \Pi \rho _N^*-\Pi \rho ^*\| _{\mathrm{TV}}=\frac12D_\varsigma \le 5c_1^2\eta _N'\).
Assume the setting of Theorem 369. For any upper bounds \(\eta '\ge G_N\), \(\eta ''\ge H_N\) with \(c_1(M_1)^2\eta ''\le \frac18\): \(\| \rho _N^*-\rho ^*\| _{\mathrm{TV}},\| \nu _N^*-\nu ^*\| _{\mathrm{TV}}\le \frac{9c_1(M_1)}{2\sqrt\beta }\eta '\), \(\sup _z|m_N-m^*|\le \frac1\lambda [\| h\| _N+\eta ']\le \frac{10c_1(M_1)}\lambda \eta '\), \(\sup _x|h|\le D_\varsigma \le 10c_1(M_1)^2\eta '\) and \(\| \Pi \rho _N^*-\Pi \rho ^*\| _{\mathrm{TV}}=\frac12D_\varsigma \le 5c_1(M_1)^2\eta '\).
By Lemma 314, \(\Delta _N\le D(\delta /2)/\sqrt N\) and \(\Delta _N'\le D'(\delta )/\sqrt N\), so that on \(E'_\delta \), \(\| \rho _N^*-\rho ^*\| _{\mathrm{TV}}\le \sqrt{D(\delta /2)/(2\beta \sqrt N)}\) and \(\sup _z|(R^N_{\lambda ,\nu _N^*}y)(z)-(R_{\lambda ,\nu ^*}f)(z)| \le \frac1\lambda \bigl[\sqrt{2D(\delta /2)/\sqrt N}+D'(\delta )/\sqrt N\bigr]\): for fixed \((\lambda ,\beta ,\delta ,m,{Y_{\max }})\), (b) is \(O(N^{-1/4})\) and (a), (c), (d) are \(O(\beta ^{-1/2}N^{-1/4})\).
On \(E'_\delta \), with \(\Delta _N:=\Delta _N(\delta /2)\) and \(\Delta '_N\) as in 86,
where \(R^N_{\lambda ,\nu _N^*}y=-g_N/\lambda =m_{\rho _N^*}\) (\(\nu _N^*\)-a.e.) and \(R_{\lambda ,\nu ^*}f=-g^*/\lambda =m_{\rho ^*}\) (\(\nu ^*\)-a.e.).
On \(E_{\delta /2}\) (in particular on \(E'_\delta \)), with \(\Delta _N:=\Delta _N(\delta /2)\), \(\| \rho _N^*-\rho ^*\| _{\mathrm{TV}}\le \sqrt{\Delta _N/(2\beta )}\) and \(\| \nu _N^*-\nu ^*\| _{\mathrm{TV}}\le \sqrt{\Delta _N/(2\beta )}\) (Pinsker’s inequality and ??; the total variation of the marginals is at most that of the joint laws).
In the setting of Theorem 328, let \({\varepsilon }{\gt}0\), \(\delta \in (0,1)\), \({\mathfrak R}_N(\Phi )\le C_{\mathrm P}\sqrt{(m+1)/N}\) and \(N_0({\varepsilon },\delta )\) as in 85. If \(N\ge N_0({\varepsilon },\delta )\) then on \(E_\delta \) (probability at least \(1-\delta \)), \(\operatorname {KL}(\rho _N^*\| \rho ^*)\le {\varepsilon }\) and \(\operatorname {KL}(\rho ^*\| \rho _N^*)\le {\varepsilon }\).
\(L(\rho _n^*)=o(\lambda _n)\), \(\operatorname {KL}(\nu _n^*\| \nu _0)+\operatorname {KL}(\nu _0\| \nu _n^*)=o(\kappa _n)\) and \(\| u_0^\dagger -m_n^*\| _{L^2(\nu _0)}\to 0\). (The Pythagorean identity ?? with \(u_0=u_0^\dagger \) and \((\lambda _n,\beta _n)\) gives \(\frac4{\lambda _n}L(\rho _n^*)+\| u_0^\dagger -m_n^*\| ^2_{L^2(\nu _0)} +\frac2{\kappa _n}[\operatorname {KL}(\nu _n^*\| \nu _0)+\operatorname {KL}(\nu _0\| \nu _n^*)] =\| u_0^\dagger \| ^2-\| m_n^*\| ^2_{L^2(\nu _n^*)}\to 0\) by Theorem 420; the three terms on the left are nonnegative.)
Assume (A1), (A3), (A5), (A8), \(f\in L^2(P_X)\) and (A9). Let \((\rho _t)_{t\ge 0}\) be the flow of laws of the solution of Theorem 525 and \(\alpha _*:=\alpha _*({\mathcal F}(\rho _0))\). Then for all \(t\ge 0\)
(Necessity of amplitude regularization.) For \(\lambda =0\), \(\beta =0\) and frozen hidden parameters, the endpoint of the fixed-feature flow is \(\gamma _\infty =S_N^\dagger y+P_{\ker S_N}\gamma _0\): the canonical empirical ridgelet transform plus the \(\ker S_N\) component of the initialization, which is data independent, invisible to the training loss and conserved by the flow. With amplitude regularization \(\lambda {\gt}0\) the \(\ker S_N\) component decays like \(e^{-\lambda t}\): \(P_{\ker S_N}\gamma _t=e^{-\lambda t}P_{\ker S_N}\gamma _0\).
Assume the hypotheses of Theorem 427 and let, for every sample and every \(M\ge 1\), \((\mu ^{M,N}_t)_{t\ge 0}\) be a family of laws on \(\Theta ^M\) ergodic to the empirical \(M\)-particle Gibbs measure (Assumption (E); for the \(M\)-particle empirical noisy gradient descent this is the ergodic theorem). Let \(g\) be bounded measurable, \({\varepsilon }{\gt}0\) and \(\delta \in (0,1)\). Then for the \((\lambda ,\beta )\) and \(N_1\) of Theorem 427, on the event of probability at least \(1-\delta \) over the sample, for every \(M\ge 1\)
in the form: for every \({\varepsilon }'{\gt}0\), eventually in \(t\), \({\mathbb E}_{\mu _t}|\int g\, \mathrm d\Pi \rho _\theta -\int g\, u_0^\dagger \, \mathrm d\nu _0| \le \frac{2{\varepsilon }}3+\frac{C_g({Y_{\max }},\lambda ,\beta )}{\sqrt M}+{\varepsilon }'\), with the constant of Definition 135. (The triangle inequality ?? at fixed \((M,t)\), integrated over \(\mu _t\): the second and third terms are deterministic on the event of stage (2) and bounded by \({\varepsilon }/3\) each, and the first is Theorem 554.)
Under the hypotheses of Corollary 555, since \(C_g({Y_{\max }},\lambda ,\beta )\) does not depend on the sample, there is \(M_0=M_0({\varepsilon };{Y_{\max }},\lambda ,\beta ,g)\) such that on the same event of probability at least \(1-\delta \), for every \(M\ge M_0\), eventually in \(t\), \({\mathbb E}_{\mu _t}\bigl|\int g\, \mathrm d\Pi \rho _\theta -\int g\, u_0^\dagger \, \mathrm d\nu _0\bigr|\le {\varepsilon }\); that is,
(\(M_0:=\max \{ \lceil (6C_g/{\varepsilon })^2\rceil ,1\} \) gives \(C_g/\sqrt M\le {\varepsilon }/6\); take \({\varepsilon }'={\varepsilon }/6\).)
For \(\beta \ge \beta _0\) and \(\kappa \le \kappa _0\), \(\| m^*-u^\dagger \| _{L^2(\nu _0)}\le 2C_2(\beta ^{-1}+\kappa )+\lambda ^{\min (a,1)}\| g_0\| \), and likewise for the coefficient measure and the realization as in Corollary 473, with \(C_1\kappa \) replaced by \(C_2(\beta ^{-1}+\kappa )\).
Assume only the hypotheses of Theorem 468 ((S\(_\infty \)) is not used), and let \(\zeta _\beta =e^{\zeta _0}\| u^\dagger \| ^2/(8\beta )\). Then
and for \(\beta \ge \beta _0{\gt}0\), \(\kappa \le \kappa _0\), \(\Xi \le C_2(\beta ^{-1}+\kappa )\) with \(C_2\) as in Definition 147. (Pointwise \(|w-1|\le \zeta e^\zeta +\zeta _0\le \zeta _\beta e^{\zeta _\beta }+\zeta _0\) by Lemma 467, and \(\| m^*\| _{L^2(\nu _0)}\le e^{\zeta _0/2}\| u^\dagger \| \).)
(Hidden marginal.) Pointwise \(|w-1|\le \zeta _\infty e^{\zeta _\infty }+\zeta _0\), so for \(\kappa \le \kappa _0\), \(\| w-1\| _{L^\infty (\nu _0)}\le C_1'\kappa \) with \(C_1'\) as in Definition 146,
(\(\| \nu ^*-\nu _0\| _{\mathcal M}=\int |w-1|\, \mathrm d\nu _0\), and \(\operatorname {KL}\le \chi ^2\) by Lemma 905.)
For \(\kappa \le \kappa _0\),
Assume the hypotheses of Theorem 468 and (S\(_\infty \)) with constant \(B_\infty \), and put \(\zeta _\infty :=\frac\kappa 2B_\infty ^2\). Then
and for \(\kappa \le \kappa _0\), \(\Xi \le C_1\kappa \) with \(C_1\) as in Definition 145. (In the bound of Lemma 462, \(|w-1|\le \zeta e^\zeta +\zeta _0 \le \zeta _\infty e^{\zeta _\infty }+\zeta _0\), and \(\| m^*\| _{L^2(\nu _0)}\le e^{\zeta _0/2}\| u^\dagger \| \) by Lemma 465.)
Under the hypotheses of Proposition 544, \({\mathbb E}_{\pi _M}|\bar X-m|\le C_g/\sqrt M\) with the explicit constant \(C_g=\frac A{2\beta }+\log 2+4\exp \bigl(2C(\| f\| /\lambda +\sqrt{2\beta /(\pi \lambda )})\bigr) \exp \bigl(\tfrac \beta {2\lambda }(2C+\| f\| /\beta )^2\bigr)\), which depends only on \(C,\| f\| ,\lambda ,\beta \).
For \(k\ge 1\), \(\frac{4c_\alpha }{\pi ^2(k+1)^2}\le \mu _k(K^{(\infty )})\le \frac{4c_\alpha }{\pi ^2k^2}\), \(c_\alpha =\frac\alpha {1+\alpha }\) ??, and \(\frac1{1+\alpha }\le \mu _0{\lt}1\). In particular \(\mu _k\asymp k^{-2}\), and the exponent does not depend on \(\alpha \).
For \(E\in {\mathbb R}\), on the sublevel set \({\mathcal D}_E:=\{ \rho \in {\mathcal D}:{\mathcal F}(\rho )\le E\} \) and with \(c_E:=\sqrt{2(E+\beta \log Z_U)}\),
i.e. \(\hat\mu _\rho \) satisfies \(\mathrm{LSI}(\alpha _*(E))\) for every \(\rho \in {\mathcal D}_E\).
For \(E\in {\mathbb R}\) let \(c_E:=\sqrt{2(E+\beta \log Z_U)}\) (here \(c_E=\sqrt{2E}\), the constant of the free energy being dropped) and
For \(r\in L^2(P_X)\) the analysis map is \((S^*r)(z):=\langle \varphi _z,r\rangle _{L^2(P_X)}=\int _{{\mathcal X}}\varphi _z(x)r(x)\, P_X(\, \mathrm dx)\), a function of \(z\in Z\) that does not depend on the base measure \(\nu \). By Lemma ?? it is the adjoint of \(S_\nu \) for every \(\nu \).
Under (R), the canonical (minimum-norm) ridgelet transform of \(f\) with respect to \(\nu _0\) is \(u_0^\dagger :=S_{\nu _0}^\dagger f:=P_{(\ker S_{\nu _0})^\perp }u_0\), the orthogonal projection onto \((\ker S_{\nu _0})^\perp \) of any representer \(u_0\) of \(f\) (Lemma 382: it does not depend on \(u_0\)). Even if \(\operatorname {ran}S_{\nu _0}\) is not closed, \(S_{\nu _0}^\dagger f\) is defined in this sense as long as \(f\in \operatorname {ran}S_{\nu _0}\). (In Lean, \(u_0^\dagger :=0\) when \(f\notin \operatorname {ran}S_{\nu _0}\).)
Let \(\sigma \colon {\mathbb R}\to {\mathbb R}\) be measurable with \(|\sigma |\le 1\) (\(\tanh \) or \(\operatorname {erf}\)). The feature on \(S^1\) with hidden parameter \(z=(w,b)\in {\mathbb R}^2\times {\mathbb R}\) is \(\varphi _z(x(\theta )):=\sigma (w\cdot x(\theta )-b)=\sigma (w_1\cos \theta +w_2\sin \theta -b)\), a feature map in the sense of Definition 1.
The coefficient measure of \(\rho \in {\mathcal P}(\Theta )\) with \(\int |a|\, \mathrm d\rho {\lt}\infty \) is the signed measure \(\Pi \rho (\, \mathrm dz):=\int _{{\mathbb R}} a\, \rho (\, \mathrm da\, \mathrm dz)\) on \(Z\), i.e. \(\Pi \rho (A)=\int _{{\mathbb R}\times A}a\, \, \mathrm d\rho \); by disintegration \(\Pi \rho =m_\rho \, \nu _\rho \).
\(r_N:=\frac1{2\lambda }\Bigl[\sqrt{2D(\delta /2)/\sqrt N} +D'(\delta )/\sqrt N\Bigr]+\frac{{Y_{\max }}}\lambda \sqrt{\frac{D(\delta /2)}{2\beta \sqrt N}}\), the bound of Corollary 340 with \(\Delta _N\le D(\delta /2)/\sqrt N\) and \(\Delta _N'\le D'(\delta )/\sqrt N\) (Corollary 341); \(r_N\to 0\) as \(N\to \infty \).
The conditional mean amplitude of \(\rho \in {\mathcal P}(\Theta )\) is \(m_\rho (z):={\mathbb E}_\rho [a\mid z]\), defined \(\nu _\rho \)-a.e. through the disintegration \(\rho (\, \mathrm da\, \mathrm dz)=\nu _\rho (\, \mathrm dz)\, \kappa (z,\, \mathrm da)\) as \(m_\rho (z)=\int _{{\mathbb R}} a\, \kappa (z,\, \mathrm da)\).
Let \(e=(e_i)_{i\in I}\) be a Hilbert basis of the real Hilbert space \(H\) and \(c\colon I\to {\mathbb R}\) a bounded sequence. The multiplication operator by \(c\) in the coordinates of \(e\) is the bounded operator \(D_c x:=\sum _ic_i\langle e_i,x\rangle e_i\), of norm \(\| D_c\| \le \sup _i|c_i|\); it is the unique bounded operator with \(D_ce_i=c_ie_i\). (In Lean, \(D_c:=0\) when \(c\) is unbounded.)
For a finite measure \(\nu \) on \(Z\) (in the paper \(\nu \in {\mathcal P}(Z)\)) the empirical kernel operator \(K_{N,\nu }\) on \(({\mathbb R}^N,\langle \cdot ,\cdot \rangle _N)\) is \((K_{N,\nu }h)_i:=\int _Z\varphi _z(x_i)\langle \varphi _z,h\rangle _N\, \nu (\, \mathrm dz)\); for \(\nu _M=\frac1M\sum _\ell \delta _{z_\ell }\), \((K_{N,\nu _M}h)_i=\frac1M\sum _\ell \varphi _{z_\ell }(x_i)\langle \varphi _{z_\ell },h\rangle _N\).
Let \({\mathcal X}\) and \(Z\) be measurable spaces. A feature map is a jointly measurable function \(\varphi \colon Z\times {\mathcal X}\to {\mathbb R}\), \((z,x)\mapsto \varphi _z(x)\), with \(|\varphi _z(x)|\le 1\) for all \(z\in Z\) and \(x\in {\mathcal X}\). Assumption (A3) of the paper, \(\varphi _{(w,b)}(x)=\tanh (w^\top x-b)\) on \(Z={\mathbb R}^m\times {\mathbb R}\), is the instance ‘tanhFeature‘.
For a feature map \(\varphi \) on \(Z\times {\mathcal X}\) and a measurable map \(e\colon Z'\to Z\), the reparametrized feature map on \(Z'\times {\mathcal X}\) is \((z',x)\mapsto \varphi _{e(z')}(x)\). It is used to identify the hidden-parameter space \({\mathbb R}^m\times {\mathbb R}\) of (A3) with the Euclidean space \({\mathbb R}^{m+1}\).
Let \(Z\) be a metric space. The feature map \(\varphi \) is Lipschitz in the hidden parameter if there is \(L\ge 0\) such that \(|\varphi _z(x)-\varphi _{z'}(x)|\le L\, d(z,z')\) for all \(x\in {\mathcal X}\) and \(z,z'\in Z\). Under (A1) and (A3) this holds with \(L=\sqrt{R_X^2+1}\).
For a feature map \(\varphi \) on \(Z\times {\mathcal X}\) and a subset \(A\subseteq {\mathcal X}\), the restriction of \(\varphi \) to \(A\) is the feature map \(Z\times A\to {\mathbb R}\), \((z,x)\mapsto \varphi _z(x)\); it is used with \(A=\{ |x|\le R_X\} \), which is assumption (A1).
The feature map is smooth with bounded derivatives if \(z\mapsto \varphi _z(x)\) is \(C^\infty \) for every \(x\) and \(\sup _{x,z}\| \nabla _z^n\varphi _z(x)\| {\lt}\infty \) for every \(n\ge 1\). Under (A1) and (A3) this holds with \(\sup _{x,z}\| \nabla _z\varphi _z(x)\| \le \sqrt{R_X^2+1}\) and \(\sup _{x,z}\| \nabla _z^2\varphi _z(x)\| \le c_2(R_X^2+1)\).
For a sample \((x_i,y_i)_{i=1}^N\), \(\lambda \ge 0\) and a potential \(V\colon Z\to {\mathbb R}\), the finite-width objective is \(J_N(\theta ):=L_N(\theta )+\frac1M\sum _{\ell =1}^MU(\theta _\ell )\), where \(L_N(\theta )=\frac1{2N}\sum _{i=1}^N(F_\theta (x_i)-y_i)^2\) and \(U(a,z)=\frac\lambda 2a^2+V(z)\).
For fixed hidden parameters \(z_1,\dots ,z_M\in Z\) and the sample \((x_i)_{i=1}^N\), the fixed-feature synthesis operator is \(S_N\colon ({\mathbb R}^M,\langle \cdot ,\cdot \rangle _M)\to ({\mathbb R}^N,\langle \cdot ,\cdot \rangle _N)\), \((S_N\gamma )_i:=\frac1M\sum _{\ell =1}^M\gamma _\ell \varphi _{z_\ell }(x_i)\), with adjoint \((S_N^*r)_\ell =\frac1N\sum _i\varphi _{z_\ell }(x_i)r_i=(S_N^*r)(z_\ell )\).
The real Fourier modes \(\cos (\ell \theta )\), \(\sin (\ell \theta )\) (\(\ell \in \mathbb Z\)) as elements of \(L^2(\tau )\); \({\mathcal{H}}_0=\operatorname {span}\{ 1\} \) and \({\mathcal{H}}_\ell =\operatorname {span}\{ \cos \ell \theta ,\sin \ell \theta \} \) for \(\ell \ge 1\).
The free energy on \({\mathcal D}\) is \({\mathcal F}(\rho ):=L(\rho )+\beta \, \operatorname {KL}(\rho \| \mu _U)\) (the paper’s definition contains the additional constant \(-\beta \log Z_U\), which is dropped here). By Lemma 183, whenever \(\int U\, \mathrm d\rho {\lt}\infty \) one has \({\mathcal F}(\rho )=L(\rho )+\int U\, \mathrm d\rho +\beta \operatorname {Ent}(\rho )-\beta \log Z_U\).
For \(\nu \in {\mathcal P}(Z)\), \(u\colon Z\to {\mathbb R}\) measurable and \(\sigma ^2:=\beta /\lambda \) put \(\rho _{\nu ,u}:=\nu (\, \mathrm dz)\otimes {\mathcal N}(u(z),\sigma ^2)(\, \mathrm da)\), a probability measure on \(\Theta \) with hidden marginal \(\nu \) and conditional mean amplitude \(u\). The competitor of the paper is \(\tilde\rho _u:=\rho _{\nu _0,u}\) for \(u\in L^2(\nu _0)\).
For a reference measure \(\mu \) on \(\Theta \), a potential \(W\colon \Theta \to {\mathbb R}\) and \(\beta {\gt}0\) with \(Z_W:=\int e^{-W/\beta }\, \mathrm d\mu {\lt}\infty \), the Gibbs measure is \(\hat\mu _W(\, \mathrm d\theta ):=Z_W^{-1}e^{-W(\theta )/\beta }\, \mu (\, \mathrm d\theta )\). (The paper takes \(\mu \) to be Lebesgue measure on \(\Theta ={\mathbb R}\times Z\).)
Let \(\lambda ,\beta {\gt}0\) and \(\nu _0\in {\mathcal P}(Z)\) (in the paper \(\nu _0(\, \mathrm dz)=Z_0^{-1}e^{-V(z)/\beta }\, \mathrm dz\)). With \(U(\theta )=\frac\lambda 2a^2+V(z)\) the Gibbs reference measure is \(\mu _U(\, \mathrm d\theta ):=Z_U^{-1}e^{-U(\theta )/\beta }\, \mathrm d\theta ={\mathcal N}(0,\beta /\lambda )(\, \mathrm da)\otimes \nu _0(\, \mathrm dz)\).
A curve \(\theta \colon [0,\infty )\to \Theta ^M\) is a solution of the gradient flow ??, \(\dot\theta _t=-M\nabla J_N(\theta _t)\), if it is differentiable on \([0,\infty )\) with \(\dot a_\ell =-[(S_N^*r_t)(z_\ell )+\lambda a_\ell ]\) and \(\dot z_\ell =-[a_\ell \nabla _z(S_N^*r_t)(z_\ell )+\nabla V(z_\ell )]\).
For \(A{\gt}0\) and \(c\in \mathbb N\), \(\mathrm{HC}(m,A,c)\) is the statement: for every \(N\), every \(x_1,\dots ,x_N\in {\mathbb R}^m\) and every \(0{\lt}{\varepsilon }\le 1/2\), \(\mathcal N({\varepsilon },\Phi ,L_2(P_N))\le (A/{\varepsilon })^{c(m+1)}\). Haussler’s bound gives \(\mathrm{HC}(m,A,2)\) with an absolute \(A\).
\(f\in H^1(-1,1)\) with weak derivative \(g\) if \(g\in L^2(P_X)\) and \(f(x)=f(-1)+\int _{-1}^xg(t)\, \, \mathrm dt\) for all \(x\in [-1,1]\) (absolutely continuous with an \(L^2\) derivative; \(g=f'\) is unique a.e.). \(H^1(-1,1)\) is the set of \(f\) admitting such a \(g\).
For \(\nu \in {\mathcal P}(Z)\), \(\lambda {\gt}0\) and \(f\in L^2(P_X)\) the resolvent quadratic form is \(Q^\nu _\lambda (f):=\langle f,(K_\nu +\lambda )^{-1}f\rangle _{L^2(P_X)}\), where \(K_\nu =S_\nu S_\nu ^*\) is the kernel operator and \((K_\nu +\lambda )^{-1}\) its resolvent (Definition 19). Equivalently, with \(g=(K_\nu +\lambda )^{-1}f\), \(Q^\nu _\lambda (f)=\| S^*g\| ^2_{L^2(\nu )}+\lambda \| g\| ^2_{L^2(P_X)}\); and \(\frac\lambda 2Q^\nu _\lambda (f)=\min _{u\in L^2(\nu )}\bigl[\frac12\| S_\nu u-f\| ^2 +\frac\lambda 2\| u\| ^2\bigr]\) (Proposition ??).
For \(\lambda {\gt}0\) and \(f\in L^2(P_X)\), \(R^{(\infty )}_\lambda f:=S^*(K^{(\infty )}+\lambda )^{-1}f\), i.e. \((R^{(\infty )}_\lambda f)(w,b)=\int \tanh (wx-b)\, \bigl[(K^{(\infty )}+\lambda )^{-1}f\bigr](x)\, P_X(\, \mathrm dx)\), where \(S^*\) is the analysis map of the tanh feature and \(K^{(\infty )}\) the kernel operator of the sign network ??.
For \(\lambda {\gt}0\) the resolvent \((K_\nu +\lambda )^{-1}\) is the inverse of the bounded operator \(K_\nu +\lambda \) on \(L^2(P_X)\); by Lemma ?? it exists and \(\| (K_\nu +\lambda )^{-1}\| \le 1/\lambda \). (In Lean it is defined as the inverse when \(K_\nu +\lambda \) is invertible and as \(0\) otherwise.)
With \(\mu :=\nu _N^*+\nu ^*\), \(w_N:=\, \mathrm d\nu _N^*/\, \mathrm d\mu \), \(w^*:=\, \mathrm d\nu ^*/\, \mathrm d\mu \), the density of \(\Pi \varsigma =m_N\nu _N^*-m^*\nu ^*\) with respect to \(\mu \) is \(\psi :=w_Nm_N-w^*m^*\), so that \(D_\varsigma =\int |\psi |\, \mathrm d\mu \).
A probability measure \(\mu \) on \({\mathbb R}^d\) satisfies \(\mathrm{LSI}(\alpha )\) with \(\alpha {\gt}0\) if for all \(g\in C_c^\infty ({\mathbb R}^d)\)
(Here the test functions are \(C^1_c\); by standard approximation the two conventions agree.)
For \(R{\gt}0\) let \(\chi _R(x):=\phi (2-|x|^2/R^2)\), where \(\phi \in C^\infty ({\mathbb R})\) is nondecreasing with \(\phi =0\) on \((-\infty ,0]\) and \(\phi =1\) on \([1,\infty )\). Then \(0\le \chi _R\le 1\), \(\chi _R=1\) on \(\{ |x|\le R\} \), \(\chi _R=0\) on \(\{ |x|\ge \sqrt2R\} \) and \(|\nabla \chi _R|\le C/R\) with \(C=4\sup |\phi '|\).
A probability measure \(\mu \) satisfies the KL form of \(\mathrm{LSI}(\alpha )\) if for every probability measure \(\rho \ll \mu \) with a smooth density and \(\operatorname {KL}(\rho \| \mu ){\lt}\infty \),
This is (L1) of the paper, obtained from \(\mathrm{LSI}(\alpha )\) by substituting \(g=\sqrt{\, \mathrm d\rho /\, \mathrm d\mu }\).
A probability measure \(\mu \) on \(\Theta \) satisfies the KL form of \(\mathrm{LSI}(\alpha )\) if \(\operatorname {KL}(\rho \| \mu )\le \frac1{2\alpha }I(\rho |\mu )\) for every probability measure \(\rho \) on \(\Theta \) with a smooth density with respect to \(\mu \) and \(\operatorname {KL}(\rho \| \mu ){\lt}\infty \).
For \(c\ge 0\) let \(\alpha (c):=\dfrac {\min \{ \lambda ,\ \lambda _z e^{-c^2/(2\lambda \beta )}\} } {\beta \, (1+\sqrt{R_X^2+1}\, c/\lambda )^2}\), the right-hand side of the LSI constant of \(\hat\mu _\rho \) in Lemma ??(iv) with \(c=c_\rho \); it is nonincreasing in \(c\).
The setting of Section ??: \(Q\in {\mathcal P}({\mathcal X}\times {\mathbb R})\) with \(|Y|\le {Y_{\max }}\) a.s., \(f={\mathbb E}[Y\mid X]\) a regression function with \(|f|\le {Y_{\max }}\), \(\lambda ,\beta {\gt}0\), and \(\rho ^*\in {\mathcal D}\) a minimizer of \({\mathcal F}\) on \({\mathcal D}\) (with target \(f\in L^2(P_X)\)).
Assume (A1), (A3), (A5), \(f\in L^2(P_X)\), (R) and the convention ?? (\(\nu _0\) fixed). Let \((\lambda _n,\beta _n)_{n\ge 1}\) satisfy \(\lambda _n\to 0\) and \(\kappa _n:=\lambda _n/\beta _n\to 0\), and let \(\rho _n^*\) be the minimizer of \({\mathcal F}\) for \((\lambda _n,\beta _n)\); write \(\nu _n^*:=\nu _{\rho _n^*}\), \(m_n^*:=m_{\rho _n^*}\), \(\gamma _n:=\Pi \rho _n^*=m_n^*\nu _n^*\) and \(\gamma _\infty :=u_0^\dagger \nu _0\).
For \(\rho \) with \(\int |a|\, \mathrm d\rho {\lt}\infty \) and \(\lambda {\gt}0\), \(m_\rho :=-s_\rho /\lambda =-\lambda ^{-1}S^*r_\rho \colon Z\to {\mathbb R}\), a bounded measurable function; for the minimizer \(\rho ^*\) it is the conditional mean amplitude (Theorem 228).
For \(\rho \) with \(\int |a|\, \mathrm d\rho {\lt}\infty \), \(\lambda {\gt}0\) and a finite measure \(\nu \) on \(Z\), the bounded measurable function \(-s_\rho /\lambda =-\lambda ^{-1}S^*r_\rho \) is an element of \(L^2(\nu )\); for \(\nu =\nu ^*\) it is \(m_{\rho ^*}\) (Theorem 228).
(Hypothesis W as the definition of a solution in law.) A curve \((\rho _t)_{t\ge 0}\) of probability measures on \(\Theta \) is an MFLD flow if
\(\rho _t\in {\mathcal D}\) for all \(t\ge 0\), \(t\mapsto \rho _t\) is weakly continuous, \(\sup _{t\le T}\int |\theta |^2\, \mathrm d\rho _t{\lt}\infty \) for every \(T\), and for a.e. \(t\ge 0\) the measure \(\rho _t\) has a smooth density with respect to \(\hat\mu _{\rho _t}\) with \(\nabla \log \frac{\, \mathrm d\rho _t}{\, \mathrm d\hat\mu _{\rho _t}}\in L^2(\rho _t)\);
\(u\mapsto I(\rho _u|\hat\mu _{\rho _u})\) is locally integrable on \([0,\infty )\) and for all \(0\le s\le t\)
\[ {\mathcal F}(\rho _t)={\mathcal F}(\rho _s)-\beta ^2\int _s^tI(\rho _u|\hat\mu _{\rho _u})\, \mathrm du . \]
(W2) is the integrated form of \(\frac{\, \mathrm d}{\, \mathrm dt}{\mathcal F}(\rho _t)=-\beta ^2I(\rho _t|\hat\mu _{\rho _t})\) for a locally absolutely continuous \(t\mapsto {\mathcal F}(\rho _t)\).
For \(\mu \in {\mathcal P}(\Theta ^M)\) the \(M\)-particle free energy is \({\mathcal F}^M(\mu ):=M\int _{\Theta ^M}L(\rho _\theta )\, \mu (\, \mathrm d\theta ) +\beta \, \operatorname {KL}(\mu \| \mu _U^{\otimes M})\) (the paper’s \({\mathcal F}^M\) up to the dropped constant \(-\beta \log Z_{\pi _M}\)).
The \(M\)-particle Gibbs measure is \(\pi _M(\, \mathrm d\theta ):=Z_M^{-1}e^{-ML(\rho _\theta )/\beta }\, \mu _U^{\otimes M}(\, \mathrm d\theta )\), \(Z_M:=\int e^{-ML(\rho _\theta )/\beta }\, \mathrm d\mu _U^{\otimes M}\). Since \(\mu _U^{\otimes M}\propto e^{-\sum _\ell U(\theta _\ell )/\beta }\, \mathrm d\theta \), this is \(\pi _M\propto e^{-MJ(\theta )/\beta }\, \mathrm d\theta \) with \(J(\theta )=L(\rho _\theta )+\frac1M\sum _\ell U(\theta _\ell )\), the invariant measure of the \(M\)-particle noisy gradient descent.
For \(\theta =(\theta _\ell )_{\ell =1}^M\in \Theta ^M\) the particle risk is \(L(\rho _\theta )=\tfrac 12\| F_{\rho _\theta }-f\| _{L^2(P_X)}^2\), where \(\rho _\theta =\frac1M\sum _\ell \delta _{\theta _\ell }\) and \(F_{\rho _\theta }(x)=\frac1M\sum _\ell a_\ell \varphi _{z_\ell }(x)\).
For \(\gamma \ne 0\) with \(c:=\| \gamma \| _{\mathcal M}\) and Jordan decomposition \(\gamma =\gamma ^+-\gamma ^-\), the lift \(\bar\rho :=c^{-1}\bigl(\gamma ^+(\, \mathrm dz)\otimes \delta _{c}(\, \mathrm da)+ \gamma ^-(\, \mathrm dz)\otimes \delta _{-c}(\, \mathrm da)\bigr)\) is the measure \(\frac{|\gamma |(\, \mathrm dz)}{c}\otimes \delta _{c\vartheta (z)}(\, \mathrm da)\) of 159, with \(\vartheta =\, \mathrm d\gamma /\, \mathrm d|\gamma |\in \{ \pm 1\} \).
\({\widehat{\mathfrak R}}_N(\Psi ):={\mathbb E}_\sigma \sup _{z,z'\in Z} \bigl|\tfrac 1N\sum _{k=1}^N\sigma _k\varphi _z(x_k)\varphi _{z'}(x_k)\bigr|\), and \({\mathfrak R}_N(\Psi ):={\mathbb E}\, {\widehat{\mathfrak R}}_N(\Psi )\) over the inputs of an i.i.d. sample from \(Q\).
For \(\rho \in {\mathcal D}\) let \(r_\rho =F_\rho -f\), \(s_\rho :=S^*r_\rho \) and \(W_\rho (\theta ):=a\, s_\rho (z)+U(\theta )\). The proximal Gibbs measure is \(\hat\mu _\rho (\, \mathrm d\theta ):=Z_\rho ^{-1}e^{-W_\rho (\theta )/\beta }\, \mathrm d\theta =Z_\rho ^{-1}e^{-a s_\rho (z)/\beta }\mu _U(\, \mathrm d\theta )\), where \(Z_\rho =\int e^{-W_\rho /\beta }\, \mathrm d\theta {\lt}\infty \) since \(|s_\rho |\le c_\rho \).
For a class \(\mathcal G=\{ g_i:i\in \iota \} \) of functions on \(\Xi \) and points \(\xi _1,\dots ,\xi _N\in \Xi \), with i.i.d. Rademacher signs \(\sigma _k\in \{ \pm 1\} \), the empirical Rademacher complexity is \({\widehat{\mathfrak R}}_N(\mathcal G):={\mathbb E}_\sigma \sup _{i}\bigl|\tfrac 1N\sum _{k=1}^N\sigma _kg_i(\xi _k)\bigr|\) (absolute convention, FoML’s ‘empiricalRademacherComplexity‘); the one-sided version \({\mathbb E}_\sigma \sup _i\frac1N\sum _k\sigma _kg_i(\xi _k)\) of the paper is ‘empiricalRademacherOneSided‘. For \(\mathcal G=-\mathcal G\) the two agree.
For \(\rho \in {\mathcal P}(\Theta )\), \(\Theta ={\mathbb R}\times Z\), with \(\int |a|\, \mathrm d\rho {\lt}\infty \) the realization is \(F_\rho (x):=\int _\Theta a\, \varphi _z(x)\, \rho (\, \mathrm da\, \mathrm dz)\), so that \(|F_\rho (x)|\le \int |a|\, \mathrm d\rho \).
For \(\lambda {\gt}0\), \(\nu \in {\mathcal P}(Z)\) and \(f\in L^2(P_X)\) the regularized ridgelet transform is \(R_{\lambda ,\nu }f:=S^*(K_\nu +\lambda )^{-1}f\), i.e. \((R_{\lambda ,\nu }f)(z)=\langle \varphi _z,(K_\nu +\lambda )^{-1}f\rangle _{L^2(P_X)}\), a bounded continuous function of \(z\) that lies in \(L^2(\nu )\).
For \(\rho \ll \mu \) with a smooth density \(p=\frac{\, \mathrm d\rho }{\, \mathrm d\mu }\) the relative Fisher information is \(I(\rho |\mu ):=\int \bigl|\nabla \log \tfrac {\, \mathrm d\rho }{\, \mathrm d\mu }\bigr|^2\, \mathrm d\rho =\int |\nabla \log p|^2\, \mathrm d\rho \).
For a sample \(\omega =(x_i,y_i)_{i=1}^N\) the empirical \(M\)-particle Gibbs measure is \(\pi _M(\, \mathrm d\theta )\propto e^{-ML_N(\rho _\theta )/\beta }\mu _U^{\otimes M} (\, \mathrm d\theta )\), the invariant measure of the \(M\)-particle empirical noisy gradient descent (\(J=J_N\)); it is the particle Gibbs measure of the empirical feature for the uniform measure on \(\{ 1,\dots ,N\} \) and the target \(y\) (Lemma 271).
For a finite signed Borel measure \(\gamma \) on \(Z\) the synthesis is \((S\gamma )(x):=\int _Z\varphi _z(x)\, \gamma (\, \mathrm dz)\); if \(\gamma =h\nu \) with \(h\in L^2(\nu )\) then \(S\gamma =S_\nu h\), and \(F_\rho =S\, \Pi \rho \) when \(\int |a|\, \, \mathrm d\rho {\lt}\infty \).
(S\(_\infty \)) Uniform sup-norm bound. For the family of \((\lambda ,\beta )\) under consideration (for instance \(\lambda \le \lambda _0\), \(\kappa \le \kappa _0\)), \(\sup _{(\lambda ,\beta )}\sup _{z\in Z}|m^*_{\lambda ,\beta }(z)|\le B_\infty {\lt}\infty \), where \(m^*_{\lambda ,\beta }\) is the conditional mean amplitude of the minimizer of \({\mathcal F}_{\lambda ,\beta }\).
(SC\(_a\)) Source condition of order \(a{\gt}0\). There is \(g_0\in L^2(\nu _0)\) with \(u^\dagger =T_0^ag_0\). Since \(\operatorname {ran}T_0^a\subset \overline{\operatorname {ran}T_0}=(\ker S_{\nu _0})^\perp \), such a \(u^\dagger \) automatically satisfies the minimum-norm condition, and (SC\(_a\)) contains (R) (\(f=S_{\nu _0}u^\dagger \)). For \(a=1/2\), \(\operatorname {ran}T_0^{1/2}=\operatorname {ran}S_{\nu _0}^*\), so (SC\(_{1/2}\)) is equivalent to \(u^\dagger =S^*g\) with \(g\in L^2(P_X)\), i.e. to \(f=K_{\nu _0}g\). (In Lean the condition is formalized for \(a\in \{ 1/2,1\} \) with a bound \(G\) on the norm of the source element: \(f=K_{\nu _0}g\), \(\| g\| \le G\), resp. \(f=S_{\nu _0}T_0g_0\), \(\| g_0\| \le G\); the spectral power \(T_0^a\) of an operator on a real Hilbert space is not available in Mathlib.)
(SC\(_a\)) Source condition of order \(a{\gt}0\). There is \(g_0\in L^2(\nu _0)\) with \(u^\dagger =T_0^ag_0\), where \(T_0^a\) is the spectral power of Definition 164. Since \(\operatorname {ran}T_0^a\subset (\ker S_{\nu _0})^\perp \), such a \(u^\dagger \) automatically satisfies the minimum-norm condition, and (SC\(_a\)) contains (R) (\(f=S_{\nu _0}u^\dagger \)). (In Lean, with a bound \(G\) on the source norm: \(f=S_{\nu _0}T_0^ag_0\) with \(\| g_0\| \le G\).)
Let \(T\) be a bounded operator on the real Hilbert space \(H\) that is diagonal in a Hilbert basis \(e=(e_i)_{i\in I}\), \(Te_i=\mu _ie_i\) with \(0\le \mu _i\le M\). For \(a\ge 0\) the spectral power \(T^a\) is the multiplication operator by \((\mu _i^a)_i\) in the coordinates of \(e\): \(T^ax:=\sum _i\mu _i^a\langle e_i,x\rangle e_i\) (Definition 788), with the convention \(0^a=0\) for \(a{\gt}0\). It is the operator \(f(T)\) of the continuous functional calculus for \(f(t)=t^a\) on \([0,M]\).
\(C_g({Y_{\max }},\lambda ,\beta ):=\frac1{2\beta }\Bigl(\frac{{Y_{\max }}^2}{\lambda ^2} +\frac\beta \lambda \Bigr)+\log 2+4\exp \Bigl(2C\Bigl(\frac{{Y_{\max }}}\lambda +\sqrt{\frac{2\beta }{\pi \lambda }}\Bigr)\Bigr)\exp \Bigl(\frac\beta {2\lambda } \bigl(2C+\frac{{Y_{\max }}}\beta \bigr)^2\Bigr)\), a bound on the constant \(C^N_g\) of Theorem 550 that depends on the sample only through \({Y_{\max }}\) (\(A_N\le {Y_{\max }}^2/\lambda ^2+\beta /\lambda \), and ?? with \(\| f\| \to {Y_{\max }}\)).
For \(\nu \in {\mathcal P}(Z)\) the synthesis operator is \(S_\nu \colon L^2(\nu )\to L^2(P_X)\), \((S_\nu u)(x):=\int _Z u(z)\varphi _z(x)\, \nu (\, \mathrm dz)\). It is a well-defined bounded linear operator with \(\| S_\nu \| \le 1\), since \(|(S_\nu u)(x)|\le \int |u|\, \mathrm d\nu \le \| u\| _{L^2(\nu )}\) pointwise.
(A3) Let \(m\ge 1\), \({\mathcal X}={\mathbb R}^m\) and \(Z={\mathbb R}^m\times {\mathbb R}\). The hidden parameter is \(z=(w,b)\) and \(\varphi _z(x):=\tanh (w^\top x-b)\). Then \(|\varphi _z|\le 1\) and \((z,x)\mapsto \varphi _z(x)\) is continuous, so \(\varphi \) is a feature map in the sense of Definition 1.
(A3) with the Euclidean hidden-parameter space: the tanh feature \(\varphi _{(w,b)}(x)=\tanh (w^\top x-b)\) of Definition 5 with \(z=(w,b)\) in \({\mathbb R}^{m+1}={\mathbb R}^m\times {\mathbb R}\) carrying the Euclidean norm \(|(w,b)|=\sqrt{|w|^2+b^2}\) (the reparametrization, Definition 4, along the identity \({\mathbb R}^{m+1}\to {\mathbb R}^m\times {\mathbb R}\)).
A class \(F=\{ f_i:i\in \iota \} \) of real functions on \({\mathcal X}\) is separable if there is a countable \(D\subseteq \iota \) such that every \(f_i\) is the pointwise limit of a sequence \((f_{u_n})_n\) with \(u_n\in D\). Countable classes are separable, and so are classes with \(\iota \) a separable first countable topological space and \(i\mapsto f_i(x)\) continuous for every \(x\).
The real trigonometric system \((e_\ell )_{\ell \in \mathbb Z}\) of \(L^2(\tau )\): \(e_0:=1\), \(e_\ell :=\sqrt2\cos (\ell \theta )\) for \(\ell {\gt}0\) and \(e_\ell :=\sqrt2\sin (\ell \theta )\) for \(\ell {\lt}0\). It is an orthonormal basis of \(L^2(\tau )\), \(\{ e_\ell ,e_{-\ell }\} \) is an orthonormal basis of \({\mathcal{H}}_\ell \) (\(\ell \ge 1\)), and \(\| P_\ell f\| ^2=\langle f,e_\ell \rangle ^2+\langle f,e_{-\ell }\rangle ^2\).
Let \(0{\lt}\alpha {\lt}1\) and \(L\ge 2\). The truncated homogeneous law is \(\nu _0^{(L)}(\, \mathrm dw\, \mathrm db):=Z_L^{-1}|w|^{\alpha -1}\mathbf1_{B_L}(w,b)\, \mathrm dw\, \mathrm db\) with \(B_L=\{ (w,b):\ 1/L\le |w|\le L,\ |b|\le L\} \) and \(Z_L=\int _{B_L}|w|^{\alpha -1}\, \mathrm dw\, \mathrm db\) ??.
The zonal function of the kernel is \(\kappa (\phi ):=k_{\nu _0}(x(\phi ),x(0))\), \(\phi \in {\mathbb R}/2\pi \mathbb Z\); by rotation invariance of \(\nu _0\), \(k_{\nu _0}(x(\theta ),x(\theta '))=\kappa (\theta -\theta ')\), and \(\kappa (\phi )=\tilde\kappa (\cos \phi )\) in the notation of Theorem ??.
For \(\gamma \ne 0\), \(\bar\rho (\, \mathrm da\, \mathrm dz):=\frac{|\gamma |(\, \mathrm dz)}{c}\otimes \delta _{c\, \vartheta (z)}(\, \mathrm da)\) with \(c=\| \gamma \| _{\mathcal M}\) and \(\vartheta =\, \mathrm d\gamma /\, \mathrm d|\gamma |\in \{ \pm 1\} \); for \(\gamma =0\), \(\bar\rho :=\delta _0\otimes \nu \) with an arbitrary \(\nu \in {\mathcal P}(Z)\).
(E) A family \((\mu _t)_{t\ge 0}\subset {\mathcal P}(\Theta ^M)\) is ergodic to \(\pi \in {\mathcal P}(\Theta ^M)\) if \(\| \mu _t-\pi \| _{{\mathrm{TV}}}\to 0\) as \(t\to \infty \) and \(\sup _{t\ge 0}\int \frac1M\sum _\ell a_\ell ^2\, \mu _t(\, \mathrm d\theta ){\lt}\infty \). For the laws of the \(M\)-particle noisy gradient descent and \(\pi =\pi _M\) the first property is the ergodic theorem for non-degenerate Langevin diffusions with an invariant probability measure and the second is the moment bound of the particle system.
Let \(\gamma \) be a finite signed measure on \(Z\) (and \(\nu \in {\mathcal P}(Z)\) arbitrary, used only when \(\gamma =0\)). Then \(\inf \bigl\{ \int a^2\, \mathrm d\rho :\rho \in {\mathcal P}_2,\ \Pi \rho =\gamma \bigr\} =\| \gamma \| ^2_{\mathcal M}\) and the infimum is attained, at \(\bar\rho \) of 159.
For \(u_1,u_2\in {\mathbb R}\) and \(s:=u_1-u_2\), \(1-\tanh (u_1-b)\tanh (u_2-b)=\frac{\cosh s}{\cosh (b-u_1)\cosh (b-u_2)}\), \(\int _{\mathbb R}\frac{\, \mathrm db}{\cosh (b-u_1)\cosh (b-u_2)}=\frac{2s}{\sinh s}\), hence \(\int _{\mathbb R}\bigl[1-\tanh (u_1-b)\tanh (u_2-b)\bigr]\, \mathrm db=2s\coth s =2|s|+\frac{4|s|}{e^{2|s|}-1}\) (\(s\ne 0\); \(=2\) for \(s=0\)).
Let \(\rho \in {\mathcal P}(\Omega )\) and \(Y\) be measurable with \({\mathbb E}_\rho Y=0\) and \(K:={\mathbb E}_\rho [Y^2e^{|Y|}]{\lt}\infty \). Then for \(|u|\le 1\), \({\mathbb E}_\rho e^{uY}\le 1+Ku^2\), hence \(\log {\mathbb E}_\rho e^{uY}\le Ku^2\). (Integrate \(e^{uY}\le 1+uY+u^2Y^2e^{|Y|}\).)
Under Theorem ??(c), (d) and Lemma ??(a), deterministically,
(\(\Pi \rho _N^*-\Pi \rho ^*=(m_N-m^*)\nu _N^*+m^*(\nu _N^*-\nu ^*)\), subadditivity of the total variation, \(|\nu _N^*-\nu ^*|(Z)=2\| \nu _N^*-\nu ^*\| _{\mathrm{TV}}\le 2\| \rho _N^*-\rho ^*\| _{\mathrm{TV}}\le \sqrt{{\mathcal{K}}_N}\) by Lemma 852.)
Let \(u_0\in L^2(\nu _0)\) be a representer of \(f\) and \(\tilde\rho :=\tilde\rho _{u_0}=\rho _{\nu _0,u_0}\). Then \(\tilde\rho \in {\mathcal D}\), \(F_{\tilde\rho }=S_{\nu _0}u_0=f\), \(L(\tilde\rho )=0\), \(\operatorname {KL}(\tilde\rho \| \mu _U)=\frac\lambda {2\beta }\| u_0\| ^2_{L^2(\nu _0)}\) and \({\mathcal F}(\tilde\rho )=\frac\lambda 2\| u_0\| ^2_{L^2(\nu _0)}\) (Lemma 391 with \(\nu =\nu _0\)).
\(\int _Z|m_n^*||w_n-1|\, \mathrm d\nu _0\to 0\); in particular \(\sup _x|F_{\rho _n^*}(x)-(S_{\nu _0}m_n^*)(x)|\le \int |m_n^*||w_n-1|\, \mathrm d\nu _0\to 0\). (Truncation, Lemma 404: for every \(T{\gt}0\), \(\int |m_n^*||w_n-1|\, \mathrm d\nu _0\le \epsilon _n(T)\sqrt{\tilde Z_n}\| u_0^\dagger \| +\frac1T(1+\tilde Z_n)\| u_0^\dagger \| ^2\) with \(\epsilon _n(T)\to 0\) and \(\tilde Z_n\to 1\), so \(\limsup _n\int |m_n^*||w_n-1|\, \mathrm d\nu _0\le 2\| u_0^\dagger \| ^2/T\); let \(T\to \infty \).)
For \(T{\gt}0\) put \(\epsilon (T):=(1-\tilde Z^{-1})+(e^{\kappa T^2/2}-1)\). Where \(|m^*|\le T\) one has \(\tilde Z^{-1}\le w\le e^{\kappa T^2/2}\), hence \(|w-1|\le \epsilon (T)\); where \(|m^*|{\gt}T\) one has \(|w-1|\le w+1\) and \(|m^*|\le m^{*2}/T\). Hence pointwise \(|m^*||w-1|\le \epsilon (T)|m^*|+\frac{m^{*2}}T(w+1)\) and
Let \(K\ge 0\) and \(\Phi \colon [0,\infty )\to {\mathbb R}\) with \(\Phi (t)\, (1+K(t-s))\le \Phi (s)\) for all \(0\le s\le t\). Then \(\Phi (t)\le e^{-Kt}\Phi (0)\) for all \(t\ge 0\). (Iterating on the partition \(kt/n\) gives \(\Phi (t)(1+Kt/n)^n\le \Phi (0)\); let \(n\to \infty \).)
Let \(\rho , \mu \) be probability measures with \(\mathrm{KL}(\rho \| \mu ) {\lt} \infty \), and let \(g \ge 0\) be measurable with \(\int e^{g} \, d\mu {\lt} \infty \) (Lebesgue integral). Then \(\int g \, d\rho \le \mathrm{KL}(\rho \| \mu ) + \log \int e^{g} \, d\mu \), where the left-hand side is the Bochner integral (equal to \(0\) if \(g \notin L^1(\rho )\)).
Let \(\rho ,\mu \) be probability measures with \(\operatorname {KL}(\rho \| \mu ){\lt}\infty \) and let \(g\ge 0\) be measurable with \(\int e^g\, \mathrm d\mu {\lt}\infty \). Then \(\int g\, \mathrm d\rho \le \operatorname {KL}(\rho \| \mu )+\log \int e^g\, \mathrm d\mu \) in \([0,\infty ]\), where the left-hand side is the Lebesgue integral of \(g\); in particular \(g\in L^1(\rho )\). The proof applies the bounded Donsker–Varadhan inequality to \(g\wedge n\) and lets \(n\to \infty \) by monotone convergence.
Let \(\rho ,\mu \) be probability measures on a metrizable space with its Borel \(\sigma \)-algebra. Then \(\operatorname {KL}(\rho \| \mu )=\sup _{g\in C_b}\Bigl[\int g\, \mathrm d\rho -\log \int e^g\, \mathrm d\mu \Bigr]\) in \([0,\infty ]\) (the supremum of the nonnegative parts; the inequality \(\ge \) is the Donsker–Varadhan inequality, and \(\le \) is Lemma 867).
Let \(F\) be a class with \(|F_i(S_k)|\le 1\) on the sample and suppose \(\mathcal N(x,F,L_2(P_N))\le (A/x)^d\) for \(0{\lt}x\le 1/2\), with \(d\ge 1\), \(A{\gt}0\). Then for every \(0{\lt}\alpha {\lt}1/2\), \({\widehat{\mathfrak R}}_N(F)\le 4\alpha +\frac6{\sqrt N}\sqrt{d\log (2A/\alpha )}\) (absolute convention).
Let \(F\) be a class with \(|F_i(S_k)|\le 1\) on the sample and suppose \(\mathcal N(x,F,L_2(P_N))\le (A/x)^d\) for \(0{\lt}x\le 1/2\), with \(d\ge 1\) and \(2A\ge 1\). Then for every \(0{\lt}\alpha {\lt}1/2\), \({\widehat{\mathfrak R}}_N(F)\le 4\alpha +\frac{12}{\sqrt N}\sqrt d\Bigl(\tfrac 12\sqrt{\log (2A)}+\sqrt2\Bigr)\) (absolute convention).
Let \(\rho ,\mu \) be probability measures on a metrizable space with its Borel \(\sigma \)-algebra, and \(K\in [0,\infty ]\) with \(\mathrm{DV}_{\rho ,\mu }(g)\le K\) for every \(g\in C_b\). Then \(\mathrm{DV}_{\rho ,\mu }(g)\le K\) for every bounded measurable \(g\). Proof: given \(|g|\le C\) and \({\varepsilon }{\gt}0\), choose \(g_1\in C_b\) with \(\int |g-g_1|\, \mathrm d(\rho +\mu )\le {\varepsilon }\) (density of \(C_b\) in \(L^1\) of a finite measure on a metrizable space) and clamp it to \(g_2:=\max (-C,\min (C,g_1))\in C_b\), which still satisfies \(\int |g-g_2|\, \mathrm d(\rho +\mu )\le {\varepsilon }\). Then \(|\int g\, \mathrm d\rho -\int g_2\, \mathrm d\rho |\le {\varepsilon }\) and \(|\int e^g\, \mathrm d\mu -\int e^{g_2}\, \mathrm d\mu |\le e^C{\varepsilon }\), so \(\mathrm{DV}(g)\le \mathrm{DV}(g_2)+{\varepsilon }+\log (\int e^g\, \mathrm d\mu +e^C{\varepsilon })-\log \int e^g\, \mathrm d\mu \), and the error tends to \(0\) as \({\varepsilon }\to 0\).
\(\nu _N^*(\, \mathrm dz) \propto \exp \bigl(\frac{g_N(z)^2}{2\lambda \beta }\bigr)\nu _0(\, \mathrm dz)\) (with \(\nu _0\propto e^{-V/\beta }\, \mathrm dz\) this is \(\nu _N^*(\, \mathrm dz) \propto \exp \bigl(\frac{g_N(z)^2}{2\lambda \beta }-\frac{V(z)}\beta \bigr)\, \mathrm dz\)).
With \(K:=K_{N,\nu _N^*}\),
and \(K+\lambda \) is injective on \({\mathbb R}^N\), so that \(r_N^*=-\lambda (K+\lambda )^{-1}y\) and \(F_{\rho _N^*}|_x=K(K+\lambda )^{-1}y\).
The minimizer \(\rho _N^*\) of \({\mathcal F}_N\) on \({\mathcal D}\) is the minimizer of the free energy of the empirical feature \((z,i)\mapsto \varphi _z(x_i)\) for the uniform measure on \(\{ 1,\dots ,N\} \) and the target \(y\in {\mathbb R}^N=L^2(\text{uniform})\); all of Lemma ?? and Theorem ?? apply to it.
If \(|y_i|\le {Y_{\max }}\) for all \(i\) and \(T\ge {Y_{\max }}/\lambda \), then deterministically \(\int _{\{ |a|{\gt}T\} }|a|\, \mathrm d\rho _N^* \le B\exp \bigl(-\frac\lambda {2\beta }(T-\frac{{Y_{\max }}}\lambda )^2\bigr)\), \(B=\frac{{Y_{\max }}}\lambda +\sqrt{\frac{2\beta }{\pi \lambda }}\).
Let \(\Psi =\{ \Psi _i:i\in I\} \) be a class of measurable functions on \({\mathcal X}\times {\mathbb R}\) bounded by one, \(\kappa \) a finite measure on \(I\), \(c\in L^1(\kappa )\) and \(e\in L^1(P)\). Then, with \(H(p):=\int _Ic(i)\Psi _i(p)\, \kappa (\, \mathrm di)\), \((P_N-P)[He]=\int _Ic(i)\, (P_N-P)(\Psi _ie)\, \kappa (\, \mathrm di)\) (Fubini, since \(\iint |c(i)\Psi _i(p)e(p)|\, \kappa (\, \mathrm di)P(\, \mathrm dp) \le \| c\| _{L^1(\kappa )}\| e\| _{L^1(P)}{\lt}\infty \)).
Assume (A3), (A4), (A5), \(f\in L^2(P_X)\), and let \(\rho ^*\) be the minimizer of \({\mathcal F}\) on \({\mathcal D}\). For \(\rho \in {\mathcal D}\), \(\operatorname {KL}(\rho \| \hat\mu _\rho ){\lt}\infty \) and
For \(\sigma =\operatorname {erf}\): \(a_{2j}(s)=0\) and \(a_{2j+1}(s)=\frac2{\sqrt\pi }\frac{(-1)^j(2j)!}{j!\sqrt{(2j+1)!}} \frac{s^{2j+1}}{(1+2s^2)^{j+1/2}}\) (\(j\ge 0\)); in particular \(a_1(s)=\frac2{\sqrt\pi }\frac s{\sqrt{1+2s^2}}\). (Gaussian integration by parts, \(\sqrt{k!}\, a_k(s)=s^k{\mathbb E}[\sigma ^{(k)}(sG)]\), and the generating function of the physicists’ Hermite polynomials.)
For \(\sigma =\operatorname {erf}\), \(\kappa _s(\rho )=\frac2\pi \arcsin \frac{2s^2\rho }{1+2s^2}\), hence on the circle \(\kappa (\phi )=\tilde\kappa (\cos \phi )=\frac2\pi \arcsin \frac{2(\sigma _w^2\cos \phi +\sigma _b^2)}{1+2s^2}\), \(s^2=\sigma _w^2+\sigma _b^2\) (the arcsine kernel of Williams 1998; the series of Lemma 595 is the Taylor series of the arcsine).
If \(\nu \ll \nu '\) and \(u,u'\) are measurable then \(\rho _{\nu ,u}\ll \rho _{\nu ',u'}\) with \(\frac{\, \mathrm d\rho _{\nu ,u}}{\, \mathrm d\rho _{\nu ',u'}}(a,z) =\frac{\, \mathrm d\nu }{\, \mathrm d\nu '}(z)\exp \Bigl(\frac{-(a-u(z))^2+(a-u'(z))^2}{2\sigma ^2}\Bigr) =\frac{\, \mathrm d\nu }{\, \mathrm d\nu '}(z) \exp \Bigl(\kappa \bigl(a(u-u')(z)-\tfrac 12(u^2-u'{}^2)(z)\bigr)\Bigr)\).
For \(f\in C^1({\mathbb R})\) with \(f\) and \(f'\) bounded and every \(n\ge 0\), \(\int f\, He_{n+1}\, \, \mathrm d\gamma =\int f'\, He_n\, \, \mathrm d\gamma \): since \(He_{n+1}=XHe_n-He_n'\) and \((He_n\gamma )'=(He_n'-XHe_n)\gamma \), this is the integration by parts \(\int f\, (XHe_n-He_n')\gamma =\int f'He_n\gamma \) (the boundary terms vanish by the Gaussian decay).
For \(v{\gt}0\) and \(t\ge 0\), \({\mathcal N}(0,v)(\{ |x|{\gt}t\} )=2\bar\Phi (t/\sqrt v) \le e^{-t^2/(2v)}\). (For \(x\ge t\ge 0\), \(x^2\ge t^2+(x-t)^2\), so \(\phi _v(x)\le e^{-t^2/(2v)}\phi _v(x-t)\); integrate over \(x{\gt}t\) and use \({\mathcal N}(0,v)(0,\infty )\le \tfrac 12\).)
Let \(v{\gt}0\), \(|m|\le m_0\le T\) and \(X\sim {\mathcal N}(m,v)\). Then \({\mathbb E}\bigl[|X|\, 1_{\{ |X|{\gt}T\} }\bigr]\le \bigl(m_0+\sqrt{2v/\pi }\bigr)e^{-(T-m_0)^2/(2v)}\). (Write \(X=m+W\) with \(W\sim {\mathcal N}(0,v)\): \(|X|{\gt}T\) forces \(|W|{\gt}t:=T-m_0\) and \(|X|\le m_0+|W|\), so the left side is at most \(m_0{\mathbb P}(|W|{\gt}t)+{\mathbb E}[|W|1_{\{ |W|{\gt}t\} }]\), and Lemmas 906 and 907 bound the two terms.)
If \(\int e^{-W/\beta }\, \mathrm d\mu {\lt}\infty \) then \(\hat\mu _W=Z_W^{-1}e^{-W/\beta }\mu \) is the exponential tilt of \(\mu \) by \(-W/\beta \) (Mathlib’s ‘Measure.tilted‘), i.e. the measure with density \(e^{-W/\beta }/\int e^{-W/\beta }\, \mathrm d\mu \) with respect to \(\mu \).
Under the hypotheses of Lemma 218, \(\int |a|\, \mathrm d\rho _s\le \int |s/\lambda |\, \mathrm d\nu _{\rho _s}+\sqrt{2\beta /(\pi \lambda )} \le C/\lambda +\sqrt{2\beta /(\pi \lambda )}\) (the Gaussian bound \({\mathbb E}|X|\le |\mu |+\sigma \sqrt{2/\pi }\) for \(X\sim {\mathcal N}(\mu ,\sigma ^2)\)).
Let \(s\) be measurable with \(|s|\le C\) and \(c:=C^2/(2\lambda \beta )\). Then \(\nu [s]=w\, \nu _0\) with \(w=Z_*^{-1}e^{s^2/(2\lambda \beta )}\), \(1\le Z_*\le e^c\), hence \(e^{-c}\le w\le e^{c}\) everywhere; in particular \(\nu [s]\) and \(\nu _0\) are mutually absolutely continuous.
Let \(s\) be measurable with \(|s|\le C\), \(W_s(\theta )=a\, s(z)\) and \(\rho _s:=\hat\mu _{W_s}\) relative to \(\mu _U\). Then \(\rho _s(\, \mathrm da\, \mathrm dz)=\nu [s](\, \mathrm dz)\, {\mathcal N}(-s(z)/\lambda ,\beta /\lambda )(\, \mathrm da)\) with \(\nu [s](\, \mathrm dz)=Z_*^{-1}e^{s(z)^2/(2\lambda \beta )}\nu _0(\, \mathrm dz)\), \(Z_*=\int e^{s^2/(2\lambda \beta )}\, \mathrm d\nu _0\). (Completing the square, \(a s+\frac\lambda 2a^2=\frac\lambda 2(a+s/\lambda )^2-\frac{s^2}{2\lambda }\).)
Let \(s\) be measurable with \(|s|\le C\) and \(\rho _s=\hat\mu _{W_s}\), \(W_s(\theta )=a\, s(z)\). Then \(\int e^{c|a|}\, \mathrm d\rho _s{\lt}\infty \) for every \(c\in {\mathbb R}\), since \(e^{-as(z)/\beta }e^{c|a|}\le e^{(C/\beta +c)|a|}\) and the \(a\)-marginal of \(\mu _U\) is Gaussian.
Let \(s\) be measurable with \(|s|\le C\), \(\rho _s=\hat\mu _{W_s}\) as in Lemma 218, \(B_C:=\frac C\lambda +\sqrt{\frac{2\beta }{\pi \lambda }}\) and \(T\ge C/\lambda \). Then \(\int _{\{ |a|{\gt}T\} }|a|\, \mathrm d\rho _s\le B_C\exp \bigl(-\frac\lambda {2\beta }(T-\frac C\lambda )^2\bigr)\). (Conditionally on \(z\), \(a\sim {\mathcal N}(-s(z)/\lambda ,\beta /\lambda )\) with \(|s(z)/\lambda |\le C/\lambda \); apply Lemma 908 with \(m_0=C/\lambda \), \(v=\beta /\lambda \) and integrate over \(\nu _{\rho _s}\).)
Let \(\mu \) be a \(\sigma \)-finite measure on \(\Theta \), \(\beta {\gt}0\), \(W\colon \Theta \to {\mathbb R}\) measurable with \(Z_W:=\int _\Theta e^{-W/\beta }\, \mathrm d\mu {\lt}\infty \), and \(\hat\mu _W(\, \mathrm d\theta ):=Z_W^{-1}e^{-W(\theta )/\beta }\mu (\, \mathrm d\theta )\). If \(\rho \in {\mathcal P}(\Theta )\), \(\rho \ll \mu \), \(\int |W|\, \mathrm d\rho {\lt}\infty \) and \(\operatorname {Ent}_\mu (\rho )=\int \log \frac{\, \mathrm d\rho }{\, \mathrm d\mu }\, \mathrm d\rho \) is finite, then
(The paper takes \(\mu \) to be Lebesgue measure; the case \(\operatorname {Ent}_\mu (\rho )=+\infty \) is Lemma 181.)
Let \(\mu \) be \(\sigma \)-finite, \(\beta {\gt}0\), \(W\) measurable with \(Z_W{\lt}\infty \), and let \(\rho \in {\mathcal P}(\Theta )\) with \(\rho \ll \mu \) and \(\int |W|\, \mathrm d\rho {\lt}\infty \). Then \(\operatorname {KL}(\rho \| \hat\mu _W){\lt}\infty \) if and only if \(\operatorname {Ent}_\mu (\rho )=\int \log \frac{\, \mathrm d\rho }{\, \mathrm d\mu }\, \mathrm d\rho \) is finite, i.e. \(\log \frac{\, \mathrm d\rho }{\, \mathrm d\mu }\in L^1(\rho )\).
Let \(\mu \in {\mathcal P}(\Theta )\), \(\beta {\gt}0\), \(W\) measurable with \(Z_W{\lt}\infty \), and \(\rho \in {\mathcal P}(\Theta )\) with \(\rho \ll \mu \) and \(\int |W|\, \mathrm d\rho {\lt}\infty \). Then \(\operatorname {KL}(\rho \| \hat\mu _W){\lt}\infty \) if and only if \(\operatorname {KL}(\rho \| \mu ){\lt}\infty \).
Let \(\mu \in {\mathcal P}(\Theta )\), \(\beta {\gt}0\), \(W\) measurable with \(Z_W{\lt}\infty \), and \(\rho \in {\mathcal P}(\Theta )\) with \(\rho \ll \mu \), \(\int |W|\, \mathrm d\rho {\lt}\infty \) and \(\operatorname {KL}(\rho \| \mu ){\lt}\infty \). Then \(\operatorname {KL}(\rho \| \hat\mu _W)=\operatorname {KL}(\rho \| \mu )+\frac1\beta \int W\, \mathrm d\rho +\log Z_W\), i.e. \(\int W\, \mathrm d\rho +\beta \operatorname {KL}(\rho \| \mu )=\beta \operatorname {KL}(\rho \| \hat\mu _W)-\beta \log Z_W\).
Assume (A3), an i.i.d. sample of size \(N\ge 1\) and a measurable \(e\colon {\mathcal X}\times {\mathbb R}\to {\mathbb R}\) with \(|e|\le E_*\), \(E_*{\gt}0\) (in the paper \(e(x,y)=y-f(x)\)). For every \(\delta _1\in (0,1)\), with probability at least \(1-\delta _1\), \(G_N\le 2E_*\, {\mathfrak R}_N(\Phi )+E_*\sqrt{2\log (2/\delta _1)/N}\).
Fix \(\delta \in (0,1)\) and apply Lemma 305 with \(\delta _1=\delta /2\) and Lemma 303 with \(\delta _2=\delta /2\): with \(\eta _N'=2E_*{\mathfrak R}_N(\Phi )+E_*\sqrt{2\log (4/\delta )/N}\) and \(\eta _N''=2{\mathfrak R}_N(\Psi )+\sqrt{2\log (4/\delta )/N}\), \({\mathbb P}(\Omega _\delta )={\mathbb P}(\{ G_N\le \eta _N'\} \cap \{ H_N\le \eta _N''\} )\ge 1-\delta \).
If \(|e|\le E\) \(Q\)-a.s. for some \(E{\gt}0\), then with \(\eta _N'=2E{\mathfrak R}_N(\Phi )+E\sqrt{2\log (4/\delta )/N}\) and \(\eta _N''=2{\mathfrak R}_N(\Psi )+\sqrt{2\log (4/\delta )/N}\), \({\mathbb P}(\{ G_N\le \eta _N'\} \cap \{ H_N\le \eta _N''\} )\ge 1-\delta \) (Lemmas 305 and 303 at level \(\delta /2\) each, and a union bound).
(Gradient of \(J_N\).) Under (A3) and (A6), \(J_N\in C^1(\Theta ^M)\) and, with \(r_i=F_\theta (x_i)-y_i\), \(\partial _{a_\ell }J_N(\theta )=\frac1M\bigl[(S_N^*r)(z_\ell )+\lambda a_\ell \bigr]\), \(\nabla _{z_\ell }J_N(\theta )=\frac1M\bigl[a_\ell \nabla _z(S_N^*r)(z_\ell )+\nabla V(z_\ell )\bigr]\). In Lean the two partial derivatives are packaged as the Fréchet derivative applied to a direction \(v=(v^a_\ell ,v^z_\ell )_\ell \): \(\partial J_N(\theta )v=\frac1M\sum _\ell \bigl([(S_N^*r)(z_\ell )+\lambda a_\ell ]v^a_\ell +[a_\ell \, \partial (S_N^*r)(z_\ell )+\partial V(z_\ell )]v^z_\ell \bigr)\).
Let \(F\) be a class with \(|F_i(S_k)|\le 1\) on the sample \(S_1,\dots ,S_N\) and \(\operatorname {Pdim}(F)\le d\), \(d\ge 1\). Then for every \(0{\lt}{\varepsilon }\le 1\), \(\mathcal N({\varepsilon },F,L_2(P_N))\le (8/{\varepsilon })^{6d}\), with no dependence on \(N\) or on the sample (Haussler’s bound; the sharp form is \(e(d+1)(2e/{\varepsilon })^d\) in \(L_1(P_N)\)).
If \(g\in L^2(\gamma )\) and \(\int gHe_n\, \, \mathrm d\gamma =0\) for all \(n\ge 0\), then \(g=0\) \(\gamma \)-a.e.: the Hermite polynomials are complete in \(L^2(\gamma )\). (Every monomial is a combination of Hermite polynomials, so all moments of \(g\gamma \) vanish; by dominated convergence its Fourier transform vanishes, and the positive and negative parts of \(g\gamma \) have the same characteristic function, hence coincide.)
For \(|\rho |\le 1\) and standard Gaussians \((X,Y)\) with correlation \(\rho \), \({\mathbb E}[He_m(X)He_n(Y)]=\rho ^n\, n!\, \delta _{mn}\). (Proof: \({\mathbb E}[He_n(\rho x+\sqrt{1-\rho ^2}Z)] =\rho ^nHe_n(x)\) by the three-term recursion and Gaussian integration by parts, then orthogonality.)
For \(b\ge 0\) and \(v=(1+2b)^{-1}\), \(\int He_n(y)e^{-by^2}\, \gamma (\, \mathrm dy)=\sqrt v\, (1-v)^{n/2}He_n(0)\): the density \(\gamma (y)e^{-by^2}\) is \(\sqrt v\) times the density of \({\mathcal N}(0,v)\), and \({\mathbb E}[He_n(\sqrt vG)]=(1-v)^{n/2}He_n(0)\) by the Gaussian smoothing identity.
For all \(x,x'\in {\mathbb R}^m\), \(k_{\nu _0}(x,x')=\int \varphi _z(x)\varphi _z(x')\nu _0(\, \mathrm dz) =\sum _{k\ge 0}a_k(s_x)a_k(s_{x'})\rho _{xx'}^{\, k}\), the series converging absolutely; on the circle \(s_x\equiv s\) and \(\kappa (\phi )=\tilde\kappa (\cos \phi ) =\kappa _s\bigl((\sigma _w^2\cos \phi +\sigma _b^2)/s^2\bigr)\). (In Lean: the case \(\| x\| =\| x'\| =1\).)
For \(s^2=\sigma _w^2+\sigma _b^2{\gt}0\) and \(\theta ,\theta '\in S^1\), \(k_{\nu _0}(x(\theta ),x(\theta '))=\sum _{k\ge 0}a_k(s)^2\rho ^k=\kappa _s(\rho )\) with \(\rho =(\sigma _w^2\, x(\theta )\cdot x(\theta ')+\sigma _b^2)/s^2\): the pre-activations \(U=(w\cdot x(\theta )-b)/s\), \(U'=(w\cdot x(\theta ')-b)/s\) form a standard Gaussian pair with correlation \(\rho \) (Lemma 570) and \({\mathbb E}[\phi (U)\phi (U')]=\sum _ka_k(\phi )^2\rho ^k\) for \(\phi =\sigma (s\cdot )\in L^2(\gamma )\) (Lemma 919).
For \(F,G\in L^2(\gamma )\) and \(|\rho |\le 1\), \({\mathbb E}[F(X)G(Y)]=\sum _{n\ge 0}\langle He_n/\sqrt{n!},F\rangle \langle He_n/\sqrt{n!},G\rangle \rho ^n\). (Expand \(F\) and \(G\) in the Hermite basis, use \({\mathbb E}[He_m(X)He_n(Y)]=\rho ^n\, n!\, \delta _{mn}\) and the continuity of \((F,G)\mapsto {\mathbb E}[F(X)G(Y)]\) on \(L^2(\gamma )\times L^2(\gamma )\).)
For \(\rho ^2+c^2=1\) and every \(x\in {\mathbb R}\), \(\int He_n(\rho x+cz)\, \gamma (\, \mathrm dz)=\rho ^nHe_n(x)\): the Hermite polynomials are the eigenfunctions of the Ornstein–Uhlenbeck semigroup. (Two-step induction with the recursion \(He_{n+2}=XHe_{n+1}-(n+1)He_n\) and Gaussian integration by parts.)
Assume (A3) and an i.i.d. sample of size \(N\ge 1\) (only the \(x_i\) are used). For every \(\delta _2\in (0,1)\), with probability at least \(1-\delta _2\), \(H_N\le 2\, {\mathfrak R}_N(\Psi )+\sqrt{2\log (2/\delta _2)/N}\). (FoML’s bound has the sharper radius \(\sqrt{2\log (1/\delta _2)/N}\), obtained without a union bound.)
If \(\sup _{\lambda {\gt}0}Q^\nu _\lambda (f){\lt}\infty \), then \(f\in \operatorname {ran}S_\nu \). (Let \(\ell :=\sup _{\lambda {\gt}0}Q^\nu _\lambda (f)=\lim _{\lambda \downarrow 0}Q^\nu _\lambda (f)\). For \(0{\lt}\lambda '\le \lambda \le \lambda _0\), Lemma 747 gives \(\| u_{\lambda '}-u_\lambda \| ^2\le Q_{\lambda '}-Q_\lambda \le \ell -Q_{\lambda _0}\), so \(u_{1/n}\) is Cauchy in \(L^2(\nu )\) with limit \(u\); and \(\| S_\nu u_\lambda -f\| ^2\le \lambda Q_\lambda (f) \le \lambda \sup Q\to 0\) gives \(S_\nu u=f\). No weak compactness is needed.)
For \(h\in L^2(P_X)\) and \(u=S^*h\) (so \(|u|\le \| h\| \)), \(\| \Pi \rho ^*-u\, \nu _0\| _{\mathcal M}\le \| m^*-u\| _{L^1(\nu ^*)}+\| h\| \, \| \nu ^*-\nu _0\| _{\mathcal M}\le \| m^*-u\| _{L^2(\nu ^*)}+\| h\| \, \| \nu ^*-\nu _0\| _{\mathcal M}\), since \(\Pi \rho ^*=m^*\nu ^*\) (Theorem 229) and \(m^*\nu ^*-u\nu _0=(m^*-u)\nu ^*+u(w-1)\nu _0\).
Let \(u\in L^2(\nu _0)\) be arbitrary and \(\tilde\rho _u:=\rho _{\nu _0,u}=\nu _0\otimes {\mathcal N}(u,\beta /\lambda )\) the competitor of ??. Then \(\tilde\rho _u\in {\mathcal D}\), \(F_{\tilde\rho _u}=S_{\nu _0}u\), \(L(\tilde\rho _u)=\frac12\| S_{\nu _0}u-f\| ^2\), \(\operatorname {KL}(\tilde\rho _u\| \mu _U)=\frac\lambda {2\beta }\| u\| ^2_{L^2(\nu _0)}\) and \({\mathcal F}(\tilde\rho _u)=\frac12\| S_{\nu _0}u-f\| ^2+\frac\lambda 2\| u\| ^2_{L^2(\nu _0)}\) (Lemma ?? with \(\nu =\nu _0\)).
For \(\alpha {\gt}0\) and \(j\ge 0\) the equation \(\varkappa \tan \varkappa =\alpha \) has exactly one root \(\varkappa ^e_j\) in \((j\pi ,j\pi +\frac\pi 2)\): \(\varkappa \tan \varkappa \) is continuous and strictly increasing from \(0\) to \(+\infty \) there (proof of Theorem ??(iii)).
For \(g\in L^2(P_X)\) and \(P_X\)-a.e. \(x\), \((K_\nu g)(x)=\int _{{\mathcal X}}k_\nu (x,x')g(x')\, P_X(\, \mathrm dx')\) with \(k_\nu (x,x')=\int _Z\varphi _z(x)\varphi _z(x')\, \nu (\, \mathrm dz)\), \(|k_\nu |\le 1\). (Fubini in \((K_\nu g)(x)=\int _Z\varphi _z(x)\int _{{\mathcal X}}\varphi _z(x')g(x')\, P_X(\, \mathrm dx')\, \nu (\, \mathrm dz)\), the integrand being dominated by \(|g(x')|\).)
Let \((\varphi ,\nu )\) and \((\varphi ',\nu ')\) be two feature/hidden-law pairs on the same \(L^2(P_X)\) with kernels \(k_\nu (x,x')=\int \varphi _z(x)\varphi _z(x')\, \nu (\, \mathrm dz)\) and \(k_{\nu '}\). If \(\sup _{x,x'}|k_\nu (x,x')-k_{\nu '}(x,x')|\le {\varepsilon }\), then \(\| K_\nu -K_{\nu '}\| _{L^2(P_X)\to L^2(P_X)}\le {\varepsilon }\). (By Lemma 753, \(|(K_\nu g-K_{\nu '}g)(x)|\le {\varepsilon }\| g\| _{L^1(P_X)}\le {\varepsilon }\| g\| _{L^2(P_X)}\) for a.e. \(x\), and \(P_X\) is a probability measure.) In the setting of Theorem 676, \(\| K^{(L)}-K^{(\infty )}\| \le \sup _{x,x'\in [-1,1]}|K^{(L)}(x,x')-K^{(\infty )}(x,x')| \le {\varepsilon }_L\).
Let \((\varphi ,\nu )\) and \((\varphi ',\nu ')\) be two feature/hidden-law pairs on the same \(L^2(P_X)\) with kernels \(k_\nu \), \(k_{\nu '}\). If \(|k_\nu (x,x')-k_{\nu '}(x,x')|\le {\varepsilon }\) for \(P_X\)-a.e. \(x\) and \(P_X\)-a.e. \(x'\), then \(\| K_\nu -K_{\nu '}\| _{L^2(P_X)\to L^2(P_X)}\le {\varepsilon }\) (the proof of Lemma 754 only uses the bound almost everywhere).
Under the hypotheses of Lemma 766, for every \(h\in L^2(P_X)\),
where \(\| \nu ^*-\nu _0\| _{\mathcal M}=|\nu ^*-\nu _0|(Z)=\| w-1\| _{L^1(\nu _0)}\). (Expand \(\| m^*-u\| ^2_{L^2(\nu ^*)}\) for \(u=S^*h\); Lemma 769 for the cross term, Lemma 770 for \(\| u\| ^2_{L^2(\nu ^*)}\) and Lemma 767 with this \(u\) for \(\| m^*\| ^2_{L^2(\nu ^*)}\); then \(-\frac1\lambda \| r^*\| ^2-2\langle r^*,h\rangle =-\frac1\lambda \| r^*+\lambda h\| ^2 +\lambda \| h\| ^2\le \lambda \| h\| ^2\) and \(\frac1\lambda \| K_{\nu _0}h-f\| ^2+2\langle K_{\nu _0}h-f,h\rangle +\lambda \| h\| ^2 =\frac1\lambda \| (K_{\nu _0}+\lambda )h-f\| ^2\).)
For \(h=(K_{\nu _0}+\lambda )^{-1}f\) (so that \(S^*h=R_{\lambda ,\nu _0}f\)), the first term of ?? vanishes and
since \(\| \nu ^*-\nu _0\| _{\mathcal M}\le \sqrt{\kappa Q^{\nu _0}_\lambda (f)}\) (Lemma 768) and \(\lambda \| (K_{\nu _0}+\lambda )^{-1}f\| ^2\le Q^{\nu _0}_\lambda (f)\) (Lemma 743).
Assume (A1), (A3), (A4), (A5) and \(f\in L^2(P_X)\) (no (R)). With \(q:=Z_*/Z_0=\int e^{\kappa m^{*2}/2}\, \mathrm d\nu _0\),
??. (Theorem 226, the logarithm of the density integrated against \(\nu ^*\), and \(e^{\kappa m^{*2}/2}\ge 1\).)
Let \(\varkappa {\gt}0\), \(g\in C[-1,1]\) with \(g'\in C[-1,1]\), \(g''=-\varkappa ^2g\) on \((-1,1)\) and \(g'(\pm 1)=\mp \frac\alpha 2(g(1)+g(-1))\). Then \(g=A\sin \varkappa x+B\cos \varkappa x\) with \(A\cos \varkappa =0\) and \(B(\varkappa \sin \varkappa -\alpha \cos \varkappa )=0\) (proof of Theorem ??(iii)).
Without (R), \(1\le q\le \exp \bigl(\frac\kappa 2\| m^*\| ^2_{L^2(\nu ^*)}\bigr) \le \exp \bigl(\frac\kappa 2Q^{\nu _0}_\lambda (f)\bigr)\): \(\log q=\frac\kappa 2\| m^*\| ^2_{L^2(\nu ^*)} -\operatorname {KL}(\nu ^*\| \nu _0)\le \frac\kappa 2\| m^*\| ^2_{L^2(\nu ^*)}\) (Lemma 763) and \(\| m^*\| ^2_{L^2(\nu ^*)}\le Q^{\nu _0}_\lambda (f)\) (Proposition 759).
Assume (A1), (A3), (A4), (A5) and \(f\in L^2(P_X)\); let \(\rho ^*\) be the minimizer of \({\mathcal F}\) for \((\lambda ,\beta ,\nu _0)\), \(r^*=F_{\rho ^*}-f\), \(\nu ^*\) its hidden marginal, \(m^*\) its conditional mean amplitude and \(\kappa =\lambda /\beta \). For \(u\in L^2(\nu _0)\) put \(E(u):=\frac1\lambda \| S_{\nu _0}u-f\| ^2_{L^2(P_X)}+\| u\| ^2_{L^2(\nu _0)}\) ??. Then for every \(u\in L^2(\nu _0)\),
all terms on the right-hand side being nonnegative. If \(S_{\nu _0}u=f\) this is ??. (The strong-convexity identity ?? of Lemma 212 at the competitor \(\tilde\rho _u=\nu _0\otimes {\mathcal N}(u,\beta /\lambda )\), Lemma 755, with the chain rules \(\operatorname {KL}(\rho ^*\| \mu _U)=\operatorname {KL}(\nu ^*\| \nu _0)+\frac\kappa 2\| m^*\| ^2_{L^2(\nu ^*)}\) and \(\operatorname {KL}(\tilde\rho _u\| \rho ^*)=\operatorname {KL}(\nu _0\| \nu ^*)+\frac\kappa 2\| u-m^*\| ^2_{L^2(\nu _0)}\), Lemma 390; multiply by \(2/\lambda \).)
For every \(u\in L^2(\nu _0)\),
(drop nonnegative terms in ??).
\(\| \nu ^*-\nu _0\| _{\mathcal M}\le \sqrt{\kappa \, Q^{\nu _0}_\lambda (f)}\) and \(1\le q\le e^{\kappa Q^{\nu _0}_\lambda (f)/2}\): Pinsker’s inequality \(\| \nu ^*-\nu _0\| _{\mathcal M}\le \sqrt{2\operatorname {KL}(\nu ^*\| \nu _0)}\) (Lemma 765) with \(\operatorname {KL}(\nu ^*\| \nu _0)\le \frac\kappa 2Q^{\nu _0}_\lambda (f)\) (Proposition 759), and Lemma 764.
For \(\lambda {\gt}0\) and \(f\in L^2(P_X)\), with \(g=(K_\nu +\lambda )^{-1}f\) and \(u_\lambda =S_\nu ^*g=R_{\lambda ,\nu }f\), \(Q^\nu _\lambda (f)=\langle K_\nu g,g\rangle +\lambda \| g\| ^2=\| u_\lambda \| ^2_{L^2(\nu )} +\lambda \| g\| ^2_{L^2(P_X)}\ge 0\). (\(f=(K_\nu +\lambda )g\) and \(\langle K_\nu g,g\rangle =\| S_\nu ^*g\| ^2\).)
For \(r\in L^1(P_X)\) and \(t\in {\mathbb R}\), \(\lim _{w\to +\infty }(S^*r)(w,wt)=\lim _{w\to +\infty }\int \tanh (w(x-t))r(x)\, P_X(\, \mathrm dx) =\int \operatorname {sgn}(x-t)r(x)\, P_X(\, \mathrm dx)=(S_\varpi ^*r)(t)\), and \(\lim _{w\to -\infty }(S^*r)(w,wt)=-(S_\varpi ^*r)(t)\) (dominated convergence of \(\tanh (w(x-t))\to \operatorname {sgn}(w)\operatorname {sgn}(x-t)\) for \(x\ne t\), with the dominating function \(|r|\)).
Let \(K\ge 0\) be a bounded positive operator on a Hilbert space and \(x\in \overline{\operatorname {ran}K}\). Then \(\lambda (K+\lambda )^{-1}x\to 0\) as \(\lambda \downarrow 0\). (For \({\varepsilon }{\gt}0\) write \(x=(x-Kw)+Kw\) with \(\| x-Kw\| {\lt}{\varepsilon }\); then \(\| \lambda (K+\lambda )^{-1}(x-Kw)\| \le {\varepsilon }\) and \(\| \lambda (K+\lambda )^{-1}Kw\| \le \lambda \| w\| \), Lemma 197.)
Let \(\nu \in {\mathcal P}(Z)\) and \(f\in L^2(P_X)\) with \(Q^\nu _s(f)\le H_f\) for all \(s{\gt}0\) (for \(f=K_\nu ^{1/2}\tilde h\) one may take \(H_f=\| \tilde h\| ^2\)). Then for every \(\lambda {\gt}0\), \(\| (K_\nu +\lambda )^{-1}f\| ^2\le \frac{H_f}{4\lambda }\). (With \(z:=(K_\nu +\lambda )^{-1}(K_\nu +s)^{-1}f\) one has \((K_\nu +\lambda )^{-1}f=(K_\nu +s)z\) and \(Q^\nu _s(f)=\langle (K_\nu +s)(K_\nu +\lambda )^2z,z\rangle \), so Lemma 781 gives \(4\lambda \| (K_\nu +\lambda )^{-1}f\| ^2 \le (1+4s/\lambda )Q^\nu _s(f)\le (1+4s/\lambda )H_f\) for every \(s{\gt}0\); let \(s\downarrow 0\).)
For \(u\in L^1(\varpi )\) and \(x\in [-\ell _\alpha ,\ell _\alpha ]\), \((S_\varpi u)(x)=\int \operatorname {sgn}(x-t)u(t)\, \varpi (\, \mathrm dt) =\frac1{2\ell _\alpha }\Bigl[\int _{-\ell _\alpha }^xu-\int _x^{\ell _\alpha }u\Bigr]\) (the integrand is \(u\) for \(t{\lt}x\) and \(-u\) for \(t{\gt}x\)).
Let \(A\ge 0\) be a positive operator on a real inner product space, \(\lambda {\gt}0\), \(s\ge 0\) and \(x:=4s/\lambda \). Then for every \(z\), \(4\lambda \| (A+s)z\| ^2\le (1+x)\langle (A+s)(A+\lambda )^2z,z\rangle \). (The scalar identity \((1+x)(t+\lambda )^2-4\lambda (t+s)=(t-\lambda )^2+x\, t(t+2\lambda )\) gives \((1+x)\langle (A+s)(A+\lambda )^2z,z\rangle -4\lambda \| (A+s)z\| ^2 =\langle (A+s)v,v\rangle +x\langle (A+s)A(A+2\lambda )z,z\rangle \ge 0\) with \(v=(A-\lambda )z\), all operators being polynomials in \(A\) with nonnegative coefficients or squares.)
For \(0{\lt}\lambda '\le \lambda \), \(\| u_{\lambda '}-u_\lambda \| ^2_{L^2(\nu )}\le Q^\nu _{\lambda '}(f)-Q^\nu _\lambda (f)\). (The identity \(J_\lambda (u)=J_\lambda (u_\lambda )+\frac12\| S_\nu (u-u_\lambda )\| ^2 +\frac\lambda 2\| u-u_\lambda \| ^2\) of Lemma ??(4) with \(u=u_{\lambda '}\), \(J_\lambda (u_\lambda )=\frac\lambda 2Q^\nu _\lambda (f)\) and \(J_\lambda (u_{\lambda '})\le \frac\lambda {\lambda '}J_{\lambda '}(u_{\lambda '}) =\frac\lambda 2Q^\nu _{\lambda '}(f)\).)
Assume (R). Then \(u_\lambda =R_{\lambda ,\nu }f\to S_\nu ^\dagger f\) in \(L^2(\nu )\) as \(\lambda \downarrow 0\). (By Lemma 453, \(u_\lambda -u^\dagger =-\lambda (T_0+\lambda )^{-1}u^\dagger \); \(u^\dagger \in (\ker S_\nu )^\perp =\overline{\operatorname {ran}T_0}\), so for \({\varepsilon }{\gt}0\) write \(u^\dagger =(u^\dagger -T_0w)+T_0w\) with \(\| u^\dagger -T_0w\| {\lt}{\varepsilon }\); then \(\| \lambda (T_0+\lambda )^{-1}(u^\dagger -T_0w)\| \le {\varepsilon }\) and \(\| \lambda (T_0+\lambda )^{-1}T_0w\| \le \lambda \| w\| \).)
For a \(\sigma \)-finite \(\rho \), measurable \(Y\) and \(t\in {\mathbb R}\), \(\int e^{t\sum _iY(x_i)}\, \rho ^{\otimes \iota }(\, \mathrm dx)=\bigl(\int e^{tY}\, \mathrm d\rho \bigr)^{|\iota |}\) (Fubini; no integrability is needed, both sides being \(0\) when the right-hand factor is not integrable).
For \(\nu \in {\mathcal P}(Z)\) and \(u\in L^2(\nu )\) measurable, \(\int |a|\, \mathrm d\rho _{\nu ,u}\le \| u\| _{L^1(\nu )}+\sigma \sqrt{2/\pi }{\lt}\infty \) (\({\mathbb E}|X|\le |\mu |+\sigma \sqrt{2/\pi }\) for \(X\sim {\mathcal N}(\mu ,\sigma ^2)\), and \(u\in L^2(\nu )\subset L^1(\nu )\)).
For \(\nu \in {\mathcal P}(Z)\) and \(u\in L^2(\nu )\) (measurable), \(\int |a|\, \mathrm d\rho _{\nu ,u}\le \| u\| _{L^1(\nu )}+\sigma \sqrt{2/\pi }{\lt}\infty \), \(\int a^2\, \mathrm d\rho _{\nu ,u}=\| u\| ^2_{L^2(\nu )}+\sigma ^2\), \(F_{\rho _{\nu ,u}}=S_\nu u\) and \(\Pi \rho _{\nu ,u}=u\, \nu \). (For \(X\sim {\mathcal N}(\mu ,\sigma ^2)\), \({\mathbb E}|X|\le |\mu |+\sigma \sqrt{2/\pi }\) and \({\mathbb E}X^2=\mu ^2+\sigma ^2\); apply this with \(\mu =u(z)\) and integrate against \(\nu \); Fubini gives \(F_{\rho _{\nu ,u}}(x)=\int _Z[\int _{\mathbb R}a\, {\mathcal N}(u(z),\sigma ^2)(\, \mathrm da)]\varphi _z(x)\nu (\, \mathrm dz) =(S_\nu u)(x)\) and likewise \(\Pi \rho _{\nu ,u}=u\nu \).)
If \(\nu ,\nu '\in {\mathcal P}(Z)\), \(u\in L^2(\nu )\), \(u'\) measurable, \(\nu \ll \nu '\), \(\operatorname {KL}(\nu \| \nu '){\lt}\infty \) and \(u'\in L^2(\nu )\), then \(\operatorname {KL}(\rho _{\nu ,u}\| \rho _{\nu ',u'}){\lt}\infty \) and
(By Lemma 389, \(\log \frac{\, \mathrm d\rho _{\nu ,u}}{\, \mathrm d\rho _{\nu ',u'}}=\log \frac{\, \mathrm d\nu }{\, \mathrm d\nu '}(z) +\kappa (a(u-u')-\tfrac 12(u^2-u'{}^2))\), which is \(\rho _{\nu ,u}\)-integrable since \(|a(u-u')|\le \frac12(a^2+(u-u')^2)\) with \(\int a^2\, \mathrm d\rho _{\nu ,u}{\lt}\infty \); and \(\int a\, {\mathcal N}(u,\sigma ^2)(\, \mathrm da)=u\) gives \(\int \kappa (u(u-u')-\tfrac 12(u^2-u'{}^2))\, \mathrm d\nu =\frac\kappa 2\| u-u'\| ^2_{L^2(\nu )}\).)
For \(\nu \in {\mathcal P}(Z)\) and \(u\in L^2(\nu )\) measurable, \(F_{\rho _{\nu ,u}}(x)=\int _Z\bigl[\int _{\mathbb R}a\, {\mathcal N}(u(z),\sigma ^2)(\, \mathrm da)\bigr] \varphi _z(x)\, \nu (\, \mathrm dz)=(S_\nu u)(x)\) for every \(x\) (Fubini, since \(\int |a|\, \mathrm d\rho _{\nu ,u}{\lt}\infty \)).
If \(\nu \in {\mathcal P}(Z)\), \(u\in L^2(\nu )\) measurable, \(\nu \ll \nu _0\) and \(\operatorname {KL}(\nu \| \nu _0){\lt}\infty \), then \(\rho _{\nu ,u}\in {\mathcal D}\) and
(Lemma 390 with \(\nu '=\nu _0\) and \(u'=0\), since \(\mu _U=\rho _{\nu _0,0}\).)
Let \(\mu \) be a probability measure on a metrizable space with its Borel \(\sigma \)-algebra, and let \(\rho _i\to \rho \) weakly (along a filter) in \({\mathcal P}\). Then \(\operatorname {KL}(\rho \| \mu )\le \liminf _i\operatorname {KL}(\rho _i\| \mu )\). Proof: for \(g\in C_b\), \(\mathrm{DV}_{\rho ,\mu }(g)=\lim _i\mathrm{DV}_{\rho _i,\mu }(g)\le \liminf _i\operatorname {KL}(\rho _i\| \mu )\) by the Donsker–Varadhan inequality and the weak convergence, and the supremum over \(g\) is \(\operatorname {KL}(\rho \| \mu )\) by Lemma 868.
Let \(\rho ,\mu \) be probability measures with \(\operatorname {KL}(\rho \| \mu ){\lt}\infty \), let \(A\) be measurable and \(t\ge 0\). Then \(t\, \rho (A)\le \operatorname {KL}(\rho \| \mu )+\log \bigl(1+(e^t-1)\mu (A)\bigr)\). (Donsker–Varadhan with \(g=t\, \mathbf1_A\), for which \(\int e^g\, \mathrm d\mu =1+(e^t-1)\mu (A)\).) With \(t=\log (1+1/\mu (A))\) this is \(\rho (A)\le [\operatorname {KL}(\rho \| \mu )+\log 2]/\log (1+1/\mu (A))\).
Let \(\mu \) be a probability measure on a Hausdorff space such that \(\{ \mu \} \) is tight (e.g. any probability measure on a Polish space), and let \(C\in {\mathbb R}\). Then the sublevel set \(\{ \rho \in {\mathcal P}:\operatorname {KL}(\rho \| \mu )\le C\} \) is tight. Indeed, given \({\varepsilon }{\gt}0\) let \(t:=(\max (C,0)+\log 2)/{\varepsilon }\) and let \(K\) be compact with \(\mu (K^c)\le e^{-t}\); by Lemma 858, \(t\rho (K^c)\le C+\log (1+(e^t-1)e^{-t})\le C+\log 2\), so \(\rho (K^c)\le {\varepsilon }\).
Let \(\rho \) be a finite measure on \(\alpha \times \Omega \) with \(\Omega \) standard Borel, \(\nu _1\) finite and \(\nu _2\) a probability measure. Then \(\operatorname {KL}(\rho ^{(1)}\| \nu _1)\le \operatorname {KL}(\rho \| \nu _1\otimes \nu _2)\), where \(\rho ^{(1)}\) is the first marginal. (Chain rule: \(\operatorname {KL}(\rho \| \nu _1\otimes \nu _2)=\operatorname {KL}(\rho ^{(1)}\| \nu _1) +\operatorname {KL}(\rho \| \rho ^{(1)}\otimes \nu _2)\) with \(\rho =\rho ^{(1)}\otimes _{\mathrm m}\kappa \) its disintegration.)
Let \(\rho ,\mu \) be probability measures and \(K\in [0,\infty ]\) with \(\mathrm{DV}_{\rho ,\mu }(g)\le K\) for every bounded measurable \(g\). Then \(\operatorname {KL}(\rho \| \mu )\le K\). Proof: if \(\rho \not\ll \mu \), the functions \(t\mathbf1_A\) with \(\mu (A)=0{\lt}\rho (A)\) give \(\mathrm{DV}=t\rho (A)\to \infty \). Otherwise let \(r=d\rho /d\mu \) and \(g_n:=\min (n,\log \max (r,e^{-n}))\), so that \(|g_n|\le n\) and \(\int e^{g_n}\, \mathrm d\mu \le \int (r+e^{-n})\, \mathrm d\mu =1+e^{-n}\); hence \(\int g_n\, \mathrm d\rho \le K+\log (1+e^{-n})\). Since \(\min (n,(\log r)^+)\le g_n+(\log r)^-\) and \((\log r)^-\in L^1(\rho )\) (Lemma 864), monotone convergence gives \((\log r)^+\in L^1(\rho )\), so \(\operatorname {KL}(\rho \| \mu )=\int \log r\, \mathrm d\rho =\lim \int g_n\, \mathrm d\rho \le K\) by dominated convergence.
Let \(\rho ,\rho ',\mu \) be probability measures with \(\operatorname {KL}(\rho \| \mu ){\lt}\infty \) and \(\operatorname {KL}(\rho '\| \mu ){\lt}\infty \), and let \({\varepsilon }\in [0,1]\). Then \(\operatorname {KL}((1-{\varepsilon })\rho +{\varepsilon }\rho '\| \mu )\le (1-{\varepsilon })\operatorname {KL}(\rho \| \mu )+{\varepsilon }\operatorname {KL}(\rho '\| \mu ){\lt}\infty \). (Proof: \(\frac{\, \mathrm d\rho _{\varepsilon }}{\, \mathrm d\mu }=(1-{\varepsilon })\frac{\, \mathrm d\rho }{\, \mathrm d\mu }+{\varepsilon }\frac{\, \mathrm d\rho '}{\, \mathrm d\mu }\) and \(\operatorname {KL}(\rho \| \mu )=\int \phi (\frac{\, \mathrm d\rho }{\, \mathrm d\mu })\, \mathrm d\mu \) with \(\phi (x)=x\log x+1-x\) convex.)
Let \(\nu _0,\nu \in {\mathcal P}(Z)\) with \(\nu =w\, \nu _0\) for a measurable \(w\) with \(e^{-c}\le w\le e^{c}\). Then \(\operatorname {KL}(\nu \| \nu _0){\lt}\infty \) and \(\operatorname {KL}(\nu _0\| \nu ){\lt}\infty \), since \(\log \frac{\, \mathrm d\nu }{\, \mathrm d\nu _0}=\log w\) and \(\log \frac{\, \mathrm d\nu _0}{\, \mathrm d\nu }=-\log w\) are bounded by \(|c|\).
Let \(\Omega \) be standard Borel, \(\nu \in {\mathcal P}(\Omega )\) and \(\mu \in {\mathcal P}(\Omega ^n)\) with marginals \(\mu ^{(i)}\). Then \(\sum _{i=1}^n\operatorname {KL}(\mu ^{(i)}\| \nu )\le \operatorname {KL}(\mu \| \nu ^{\otimes n})\). (Induction on \(n\) with the chain rule and Lemmas 877, 878.)
Let \(\rho \ll \mu \) be finite measures. Then \((\log \frac{d\rho }{d\mu })^-\) is \(\rho \)-integrable: \(\int (\log \frac{d\rho }{d\mu })^-\, \mathrm d\rho =\int r\, (\log r)^-\, \mathrm d\mu \le \mu (\alpha )\) with \(r=\frac{d\rho }{d\mu }\), since \(r(\log r)^-\le 1-r\le 1\) on \((0,1)\).
For \(\rho \in {\mathcal D}\) (indeed for any \(\rho \) with \(\int |a|\, \mathrm d\rho {\lt}\infty \)), \(\hat\mu _\rho (\, \mathrm da\, \mathrm dz)=\hat\nu _\rho (\, \mathrm dz)\, {\mathcal N}(-s_\rho (z)/\lambda ,\beta /\lambda )(\, \mathrm da)\), where \(\hat\nu _\rho \) is the hidden marginal of \(\hat\mu _\rho \).
For \(f\ge 0\) and every \(t{\gt}0\), \(\operatorname {Ent}_\mu (f)\le \int f\log f\, \mathrm d\mu -\bigl(\int f\, \mathrm d\mu \bigr)\log t-\int f\, \mathrm d\mu +t\); for a probability measure this is \(\operatorname {Ent}_\mu (f)\le \int \bigl(f\log \tfrac ft-f+t\bigr)\, \mathrm d\mu \), with equality for \(t=\int f\, \mathrm d\mu \).
Under (A8), \(\hat\nu _\rho \) satisfies \(\mathrm{LSI}(\hat\alpha _\rho )\) with \(\hat\alpha _\rho \ge \frac{\lambda _z}\beta \exp \bigl(-\frac{\operatorname {osc}(s_\rho ^2)}{2\lambda \beta }\bigr) \ge \frac{\lambda _z}\beta \exp \bigl(-\frac{c_\rho ^2}{2\lambda \beta }\bigr)\).
The hidden marginal of \(\hat\mu _\rho \) is \(\hat\nu _\rho (\, \mathrm dz) =\hat Z_\rho ^{-1}\exp \bigl(\frac{s_\rho (z)^2}{2\lambda \beta }\bigr)\nu _0(\, \mathrm dz)\), with \(\hat Z_\rho =\int e^{s_\rho ^2/(2\lambda \beta )}\, \mathrm d\nu _0 \in [1,e^{c_\rho ^2/(2\lambda \beta )}]\).
(L1.) If \(\mu \) satisfies \(\mathrm{LSI}(\alpha )\) on a finite-dimensional space, \(\rho \ll \mu \) is a probability measure with \(\operatorname {KL}(\rho \| \mu ){\lt}\infty \), \(\log \frac{\, \mathrm d\rho }{\, \mathrm d\mu }\) is \(C^1\) and \(\nabla \log \frac{\, \mathrm d\rho }{\, \mathrm d\mu }\in L^2(\rho )\), then \(\operatorname {KL}(\rho \| \mu )\le \frac1{2\alpha }I(\rho |\mu )\) (substitute \(g=\sqrt{\, \mathrm d\rho /\, \mathrm d\mu }\), cut off by compactly supported functions, in the LSI).
Let \(\mu \) satisfy \(\mathrm{LSI}(\alpha )\) on a finite-dimensional space, and let \(\rho =p\, \mu \) be a probability measure with \(p{\gt}0\) of class \(C^1\), \(\nabla \log p\in L^2(\rho )\) and \(\operatorname {KL}(\rho \| \mu ){\lt}\infty \). Then \(\operatorname {KL}(\rho \| \mu )\le \frac1{2\alpha }\int |\nabla \log p|^2\, \mathrm d\rho \).
(L1.) If \(\mu \) satisfies \(\mathrm{LSI}(\alpha )\) on a finite-dimensional space and \(\rho \ll \mu \) is a probability measure with a smooth density and \(\operatorname {KL}(\rho \| \mu ){\lt}\infty \), then \(\operatorname {KL}(\rho \| \mu )\le \frac1{2\alpha }I(\rho |\mu )\).
Under (A1), (A3), (A5), (A8), and assuming that \(S^*r\) is \(C^1\) for every \(r\in L^2(P_X)\) (Lemma 510), \(\hat\mu _\rho \) satisfies \(\mathrm{LSI}(\alpha _\rho )\) with
where \(\ell _\rho =\sup |\nabla s_\rho |\le \sqrt{R_X^2+1}\, c_\rho \). The right-hand side does not depend on the dimension \(m\).
(F2, tensorization.) If \(\mu _1\) satisfies \(\mathrm{LSI}(\alpha _1)\) on \(E_1\) and \(\mu _2\) satisfies \(\mathrm{LSI}(\alpha _2)\) on \(E_2\), then \(\mu _1\otimes \mu _2\) satisfies \(\mathrm{LSI}(\min \{ \alpha _1,\alpha _2\} )\) on the Euclidean product \(E_1\times E_2\).
On a sample with \(|y_i|\le {Y_{\max }}\), for every \(z\in Z\),
hence \(|g_N(z)-g^*(z)|\le \| F_{\rho _N^*}|_x-F_{\rho ^*}|_x\| _N+G_N(\omega )\).
On a sample with \(|y_i|\le {Y_{\max }}\), \({\mathcal F}_N(\rho ^*)-{\mathcal F}_N(\rho _N^*) =[L_N(\rho ^*)-\tilde L(\rho ^*)]+[\tilde{{\mathcal F}}(\rho ^*)-\tilde{{\mathcal F}}(\rho _N^*)] +[\tilde L(\rho _N^*)-L_N(\rho _N^*)]\le G_++G_-\), since the middle term is \(\le 0\) by minimality of \(\rho ^*\).
On a sample with \(|y_i|\le {Y_{\max }}\), \(\tilde{{\mathcal F}}(\rho _N^*)-\tilde{{\mathcal F}}(\rho ^*) =[\tilde L(\rho _N^*)-L_N(\rho _N^*)]+[{\mathcal F}_N(\rho _N^*)-{\mathcal F}_N(\rho ^*)] +[L_N(\rho ^*)-\tilde L(\rho ^*)]\le G_++G_-\), since the middle term is \(\le 0\) by minimality of \(\rho _N^*\) and \(\rho _N^*,\rho ^*\in {\mathcal P}_B\).
\(\langle m_n^*,u_0^\dagger \rangle _{L^2(\nu _0)}\to \| u_0^\dagger \| ^2\). (For \({\varepsilon }{\gt}0\) pick \(g\in L^2(P_X)\) with \(\| u_0^\dagger -S_{\nu _0}^*g\| \le {\varepsilon }\), possible since \(u_0^\dagger \in (\ker S_{\nu _0})^\perp =\overline{\operatorname {ran}S_{\nu _0}^*}\); then \(\langle m_n^*,S_{\nu _0}^*g\rangle =\langle S_{\nu _0}m_n^*,g\rangle \to \langle f,g\rangle =\langle u_0^\dagger ,S_{\nu _0}^*g\rangle \) by Theorem 417, and the remaining terms are bounded by \((\sup _n\| m_n^*\| _{L^2(\nu _0)}+\| u_0^\dagger \| ){\varepsilon }\) by Lemma 415.)
For \(s^2=\sigma _w^2+\sigma _b^2{\gt}0\), \(\mu _\ell =\sum _{k\ge 0}a_k(s)^2c_{k,\ell }\) with \(c_{k,\ell }:=\int \rho (\phi )^k\cos (\ell \phi )\, \tau (\, \mathrm d\phi )\), \(\rho (\phi )=(\sigma _w^2\cos \phi +\sigma _b^2)/s^2\) (dominated convergence in \(\mu _\ell =\int \kappa (\phi )\cos (\ell \phi )\, \tau (\, \mathrm d\phi )\) with \(\kappa (\phi )=\sum _ka_k(s)^2\rho (\phi )^k\), the terms being bounded by the summable \(a_k(s)^2\)).
Let \(s\) be measurable with \(|s|\le C\), \(W(\theta )=a\, s(z)\) and \(\hat\mu :=\hat\mu _W\) (relative to \(\mu _U\)). For every \(\rho \in {\mathcal D}\), \(\operatorname {KL}(\rho \| \hat\mu ){\lt}\infty \) and
(Chain rule \(\operatorname {KL}(\rho \| \mu _U)=\operatorname {KL}(\rho \| \hat\mu )+\int \log \frac{\, \mathrm d\hat\mu }{\, \mathrm d\mu _U}\, \mathrm d\rho \), legitimate since \(\int |a||s|\, \mathrm d\rho \le C\int |a|\, \mathrm d\rho {\lt}\infty \) by Lemma 168.)
Let \(Z\) be a Polish space with its Borel \(\sigma \)-algebra, let \(z\mapsto \varphi _z(x)\) be continuous for every \(x\) (both hold under (A3)–(A5)), and let \(f\in L^2(P_X)\). Then the free energy \({\mathcal F}\) has a minimizer on \({\mathcal D}\): there is \(\rho ^*\in {\mathcal D}\) with \({\mathcal F}(\rho ^*)\le {\mathcal F}(\rho )\) for all \(\rho \in {\mathcal D}\). (Direct method: tightness of the sublevel sets of \(\operatorname {KL}(\cdot \| \mu _U)\), lower semicontinuity of \(\operatorname {KL}\), Prokhorov, and continuity of \(L\) along weakly convergent sequences with uniformly bounded second amplitude moments.)
For \(\rho ,\rho '\in {\mathcal D}\) (indeed for any \(\rho ,\rho '\) with \(\int |a|\, \mathrm d\rho ,\int |a|\, \mathrm d\rho '{\lt}\infty \)), with \(s_\rho :=S^*r_\rho \),
(Proof: \(r_{\rho '}=r_\rho +(F_{\rho '}-F_\rho )\), expand the square, and use Lemma 203.)
For \(\rho \) with \(\int |a|\, \mathrm d\rho {\lt}\infty \) and \(g\in L^2(P_X)\), \(\int _\Theta a\, (S^*g)(z)\, \rho (\, \mathrm d\theta )=\langle F_\rho ,g\rangle _{L^2(P_X)}\) (Fubini, since \(\iint |a||\varphi _z(x)||g(x)|\, \rho (\, \mathrm d\theta )P_X(\, \mathrm dx) \le \int |a|\, \mathrm d\rho \, \| g\| _{L^1(P_X)}{\lt}\infty \)).
Let \(\rho ^*\in {\mathcal D}\) be a minimizer of \({\mathcal F}\) on \({\mathcal D}\), \(s^*:=S^*r_{\rho ^*}\) and \(W^*(\theta ):=a\, s^*(z)\). Then \(Z_{W^*}{\lt}\infty \) and \(\rho ^*=\hat\mu _{W^*}\), i.e. \(\rho ^*(\, \mathrm d\theta )=Z_{W^*}^{-1}e^{-a s^*(z)/\beta }\, \mu _U(\, \mathrm d\theta )\); with \(\mu _U\propto e^{-U/\beta }\, \mathrm d\theta \) this is \(\rho ^*(\, \mathrm da\, \mathrm dz)\propto \exp \bigl(-\frac{a s^*(z)+\frac\lambda 2a^2+V(z)}\beta \bigr)\, \mathrm da\, \mathrm dz\). (Proof: with \(\hat\mu :=\hat\mu _{W^*}\in {\mathcal D}\) and \(\rho _{\varepsilon }:=(1-{\varepsilon })\rho ^*+{\varepsilon }\hat\mu \), Lemmas 208, 204 and 904 give \(0\le {\mathcal F}(\rho _{\varepsilon })-{\mathcal F}(\rho ^*)\le \frac{{\varepsilon }^2}2\| F_{\hat\mu }-F_{\rho ^*}\| ^2 -{\varepsilon }\beta \operatorname {KL}(\rho ^*\| \hat\mu )\); let \({\varepsilon }\downarrow 0\).)
Let \(s\) be measurable with \(|s|\le C\) and \(W(\theta )=a\, s(z)\). Then \(\hat\mu _W=Z_W^{-1}e^{-W/\beta }\mu _U\) is a probability measure with \(\operatorname {KL}(\hat\mu _W\| \mu _U){\lt}\infty \), since \(\log \frac{\, \mathrm d\hat\mu _W}{\, \mathrm d\mu _U}=-\frac{as(z)}\beta -\log Z_W\) is \(\hat\mu _W\)-integrable.
If \(\rho _0\in {\mathcal D}\) solves the self-consistency equation \(\rho _0=\hat\mu _{W_{\rho _0}}\) with \(W_{\rho _0}(\theta )=a\, (S^*r_{\rho _0})(z)\), then \({\mathcal F}(\rho )\ge {\mathcal F}(\rho _0)+\beta \operatorname {KL}(\rho \| \rho _0)\ge {\mathcal F}(\rho _0)\) for every \(\rho \in {\mathcal D}\), so \(\rho _0\) is a minimizer of \({\mathcal F}\) on \({\mathcal D}\).
Let \(\rho _0\in {\mathcal D}\) satisfy the self-consistency equation \(\rho _0=\hat\mu _{W_{\rho _0}}\), \(W_{\rho _0}(\theta )=a\, (S^*r_{\rho _0})(z)\). Then for every \(\rho \in {\mathcal D}\), \(\operatorname {KL}(\rho \| \rho _0){\lt}\infty \) and \({\mathcal F}(\rho )-{\mathcal F}(\rho _0)=\tfrac 12\| F_\rho -F_{\rho _0}\| _{L^2(P_X)}^2+\beta \operatorname {KL}(\rho \| \rho _0)\). (Lemmas 208 and 204 with \(\hat\mu =\rho _0\).)
Let \(\rho ^*\in {\mathcal D}\) be a minimizer of \({\mathcal F}\) on \({\mathcal D}\). For every \(\rho \in {\mathcal D}\), \(\operatorname {KL}(\rho \| \rho ^*){\lt}\infty \) and
(Lemma 209 with \(\rho _0=\rho ^*\), which is self-consistent by Lemma 211.)
The minimizer of \({\mathcal F}\) on \({\mathcal D}\) is unique: if \(\rho _1,\rho _2\in {\mathcal D}\) both minimize \({\mathcal F}\) then \(\rho _1=\rho _2\). (By Lemma 212, \(0\ge {\mathcal F}(\rho _2)-{\mathcal F}(\rho _1)\ge \beta \operatorname {KL}(\rho _2\| \rho _1)\), so \(\operatorname {KL}(\rho _2\| \rho _1)=0\).)
Assume (R). \(u_0^\dagger :=P_{(\ker S_{\nu _0})^\perp }u_0\) is a representer and does not depend on the choice of the representer \(u_0\): for every representer \(u_0\) of \(f\), \(S_{\nu _0}^\dagger f=P_{(\ker S_{\nu _0})^\perp }u_0\) and \(S_{\nu _0}u_0^\dagger =f\). (Two representers differ by an element of \(\ker S_{\nu _0}\), so their projections onto \((\ker S_{\nu _0})^\perp \) coincide; and \(u_0-u_0^\dagger \in \ker S_{\nu _0}\).)
Assume (R). For every representer \(u\) of \(f\), \(\| u\| ^2_{L^2(\nu _0)}=\| u_0^\dagger \| ^2_{L^2(\nu _0)}+\| u-u_0^\dagger \| ^2_{L^2(\nu _0)}\). In particular \(\| u_0^\dagger \| \le \| u\| \), with equality if and only if \(u=u_0^\dagger \). (\(u-u_0^\dagger \in \ker S_{\nu _0}\perp u_0^\dagger \), and Pythagoras.)
Assume (R). \(u_0^\dagger \) is the unique representer in \((\ker S_{\nu _0})^\perp =\overline{\operatorname {ran}S_{\nu _0}^*}\), and if \(h\in L^2(\nu _0)\) satisfies \(S_{\nu _0}h=f\) and \(\| h\| _{L^2(\nu _0)}\le \| u_0^\dagger \| _{L^2(\nu _0)}\) then \(h=u_0^\dagger \). (A representer \(h\in (\ker S_{\nu _0})^\perp \) satisfies \(h-u_0^\dagger \in \ker S_{\nu _0}\cap (\ker S_{\nu _0})^\perp =\{ 0\} \); for the last claim, Pythagoras with \(\| h\| ^2\le \| u_0^\dagger \| ^2\) forces \(\| h-u_0^\dagger \| =0\).)
Let \(\lambda ,\beta {\gt}0\), \(\nu _0\in {\mathcal P}(Z)\), and let \(\rho \in {\mathcal D}\), i.e. \(\rho \in {\mathcal P}(\Theta )\) with \(\operatorname {KL}(\rho \| \mu _U){\lt}\infty \). For every \(\eta \in (0,1)\),
Only \(|\varphi _z|\le 1\) of (A3) and \(\nu _0\in {\mathcal P}(Z)\) of (A4), (A5) enter.
For every \(\nu \in {\mathcal P}(Z)\) and \(r\in L^2(P_X)\) the adjoint of the synthesis operator is the analysis map: \((S_\nu ^*r)(z)=\langle \varphi _z,r\rangle _{L^2(P_X)}=(S^*r)(z)\) for \(\nu \)-a.e. \(z\). In particular \(S_\nu ^*r\) does not depend on \(\nu \) (as a function).
For \(\lambda {\gt}0\), \(\| \lambda (K_\nu +\lambda )^{-1}\| \le 1\) and \(\| K_\nu (K_\nu +\lambda )^{-1}\| \le 1\). (Proof without spectral calculus: for \(g=(K_\nu +\lambda )^{-1}f\), \(\| f\| ^2=\| (K_\nu +\lambda )g\| ^2 =\| K_\nu g\| ^2+2\lambda \langle K_\nu g,g\rangle +\lambda ^2\| g\| ^2\) dominates both \(\lambda ^2\| g\| ^2\) and \(\| K_\nu g\| ^2\) since \(\langle K_\nu g,g\rangle \ge 0\).)
For \(r\in L^2(P_X)\), \(P_X\)-a.e. in \(x\), \((K_\nu r)(x)=\int _Z\varphi _z(x)\, (S^*r)(z)\, \nu (\, \mathrm dz) =\int _Z\varphi _z(x)\langle \varphi _z,r\rangle _{L^2(P_X)}\, \nu (\, \mathrm dz)\), i.e. \(K_\nu \) is the integral operator with kernel \(k_\nu (x,x')=\int \varphi _z(x)\varphi _z(x')\, \nu (\, \mathrm dz)\).
For \(\lambda {\gt}0\) the operator \(K_\nu +\lambda \) is a bijection of \(L^2(P_X)\) with bounded inverse: \((K_\nu +\lambda )(K_\nu +\lambda )^{-1}=\mathrm{id}\) and \((K_\nu +\lambda )^{-1}(K_\nu +\lambda )=\mathrm{id}\), where \((K_\nu +\lambda )^{-1}\) is ‘kernelResolvent‘. (Proof: \(K_\nu \ge 0\) gives \(\langle (K_\nu +\lambda )r,r\rangle \ge \lambda \| r\| ^2\), so \(K_\nu +\lambda \) is injective with closed range whose orthogonal complement is trivial.)
For \(\lambda {\gt}0\) and \(f\in L^2(P_X)\), \(\| R_{\lambda ,\nu }f\| _{L^2(\nu )}\le \| f\| _{L^2(P_X)}/(2\sqrt\lambda )\). (Proof without the spectral measure: with \(g=(K_\nu +\lambda )^{-1}f\), \(\| S_\nu ^*g\| ^2=\langle K_\nu g,g\rangle \) and \(\| f\| ^2-4\lambda \langle K_\nu g,g\rangle =\| (K_\nu -\lambda )g\| ^2\ge 0\).)
For \(\lambda {\gt}0\) and \(f\in L^2(P_X)\), \(u_0=S_\nu ^*(K_\nu +\lambda )^{-1}f\) is the unique minimizer over \(L^2(\nu )\) of the Tikhonov functional \(J(u)=\frac12\| S_\nu u-f\| _{L^2(P_X)}^2+\frac\lambda 2\| u\| _{L^2(\nu )}^2\): \(J(u_0)\le J(u)\) for all \(u\), with equality only for \(u=u_0\). (Proof: by the normal equation, \(J(u)-J(u_0)=\frac12\| S_\nu (u-u_0)\| ^2+\frac\lambda 2\| u-u_0\| ^2\).)
Since \(0\le L(\rho _\theta )\), \(Z_M=\int e^{-ML(\rho _\theta )/\beta } \, \mathrm d\mu _U^{\otimes M}\in (0,1]\), so \(\pi _M\) is a probability measure, and \(\operatorname {KL}(\pi _M\| \mu _U^{\otimes M})=-\frac M\beta \int L(\rho _\theta )\, \mathrm d\pi _M-\log Z_M{\lt}\infty \) (the potential \(ML(\rho _\theta )\) is \(\pi _M\)-integrable because \(ue^{-u/\beta }\le \beta \)).
For every permutation \(\sigma \) of \(\{ 1,\dots ,M\} \), the image of \(\pi _M\) under \(\theta \mapsto \theta \circ \sigma \) is \(\pi _M\): \(\mu _U^{\otimes M}\) is permutation invariant and \(L(\rho _{\theta \circ \sigma })=L(\rho _\theta )\). In particular all one-particle marginals of \(\pi _M\) coincide.
For \(\mu \in {\mathcal P}(\Theta ^M)\) with \(\operatorname {KL}(\mu \| \mu _U^{\otimes M}){\lt}\infty \), \(\operatorname {KL}(\mu \| \pi _M){\lt}\infty \) and \({\mathcal F}^M(\mu )=\beta \, \operatorname {KL}(\mu \| \pi _M)-\beta \log Z_M\), \(Z_M:=\int e^{-ML(\rho _\theta )/\beta }\, \mathrm d\mu _U^{\otimes M}\) (Lemma 187 with \(W=ML(\rho _\theta )\) and reference measure \(\mu _U^{\otimes M}\); \(W\in L^1(\mu )\) by Lemma 529).
Let \(\mu \in {\mathcal P}(\Theta ^M)\) with \(\operatorname {KL}(\mu \| \mu _U^{\otimes M}){\lt}\infty \). Then \(\int a_\ell ^2\, \mu (\, \mathrm d\theta ){\lt}\infty \) for every \(\ell \), hence \(\int |a_\ell |\, \mathrm d\mu {\lt}\infty \) and \(L(\rho _\theta )\), \(\sum _\ell a_\ell s(z_\ell )\) (\(s\) bounded measurable) are \(\mu \)-integrable. (Donsker–Varadhan with \(g(\theta )=\frac{\lambda }{4\beta }a_\ell ^2\), \(\int e^g\, \mathrm d\mu _U^{\otimes M}=\int e^{\lambda a^2/(4\beta )}\, \mathrm d\mu _U=\sqrt2\).)
Let \(Z\) be first countable, \(Z_0\subseteq Z\) dense, \(\varphi \) continuous in \(z\), \(Q\in {\mathcal P}({\mathcal X}\times {\mathbb R})\), \(B\ge 0\) and \(x_1,\dots ,x_N\in {\mathcal X}\). For every \(\rho \in {\mathcal P}_B\) and \(\varepsilon {\gt}0\) there is \(\rho '\in {\mathcal P}_B^0\) with \(\int |F_\rho -F_{\rho '}|\, \mathrm dQ\le \varepsilon \) and \(|F_\rho (x_k)-F_{\rho '}(x_k)|\le \varepsilon \) for all \(k\). The proof samples \(M\) particles from \(\rho \) reweighted by \(|a|\) (Monte Carlo, expected squared error \(\le B^2/M\)), then moves the atoms into \(Z_0\) and the amplitude into \({\mathbb Q}\).
Let \(E\) be a real inner product space of dimension \(d\) and let \(n\ge d+2\). For all points \((x_i,s_i)\in E\times {\mathbb R}\), \(i=1,\dots ,n\), there is a pattern \(T\subseteq [n]\) such that no \((w,b)\in E\times {\mathbb R}\) satisfies \(s_i\le \langle w,x_i\rangle -b\iff i\in T\) for all \(i\). In other words, the affine class \(\{ x\mapsto \langle w,x\rangle -b\} \) has pseudo-dimension at most \(d+1\).
Let \(g\colon {\mathbb R}\to {\mathbb R}\) be such that for every \(t\in {\mathbb R}\) the level set \(\{ u:t\le g(u)\} \) is of the form \(\{ u:s\le u\} \), or \({\mathbb R}\), or \(\emptyset \) (e.g. \(g=\tanh \), with \(s=\operatorname {artanh}t\) for \(|t|{\lt}1\)). If \(\operatorname {Pdim}(F)\le d\) then \(\operatorname {Pdim}(g\circ F)\le d\), where \(g\circ F=\{ g\circ F_i:i\in \iota \} \).
Let \(\mu \) be a nonzero \(\sigma \)-finite measure on \(\Theta \), \(\beta {\gt}0\) and \(W\) measurable with \(Z_W=\int e^{-W/\beta }\, \mathrm d\mu {\lt}\infty \). Then \(\hat\mu _W^{\otimes M}=\widehat{(\mu ^{\otimes M})}_{W_M}\) with \(W_M(\theta )=\sum _{\ell }W(\theta _\ell )\), i.e. \(\frac{\, \mathrm d\hat\mu _W^{\otimes M}}{\, \mathrm d\mu ^{\otimes M}}(\theta ) =\prod _\ell \frac{e^{-W(\theta _\ell )/\beta }}{Z_W}=\frac{e^{-W_M(\theta )/\beta }}{Z_W^M}\), and \(Z_{W_M}=Z_W^M\).
For probability measures \(\rho ,\rho '\) with \(\operatorname {KL}(\rho \| \rho ')\) and \(\operatorname {KL}(\rho '\| \rho )\) finite, \(\| \rho -\rho '\| _{\mathrm{TV}}^2 \le \frac12\min \{ \operatorname {KL}(\rho \| \rho '),\operatorname {KL}(\rho '\| \rho )\} \), hence \(2\| \rho -\rho '\| _{\mathrm{TV}}\le \sqrt{\operatorname {KL}(\rho \| \rho ')+\operatorname {KL}(\rho '\| \rho )}\).
For \(\| x\| =\| x'\| =1\) and \(s^2=\sigma _w^2+\sigma _b^2{\gt}0\), the pair \(U=(w\cdot x-b)/s\), \(U'=(w\cdot x'-b)/s\) is, under \(\nu _0={\mathcal N}(0,\sigma _w^2I_2)\otimes {\mathcal N}(0,\sigma _b^2)\), a standard Gaussian pair with correlation \(\rho =(\sigma _w^2\, x\cdot x'+\sigma _b^2)/s^2\). (Both are centered Gaussian vectors with unit variances and covariance \(\rho \); compare the characteristic functions.)
With \(s^*=S^*r_{\rho ^*}\) and \(W^*(\theta )=a\, s^*(z)\), \(\rho ^{*\otimes M}=\widehat{(\mu _U^{\otimes M})}_{W^*_M}\), \(W^*_M(\theta )=\sum _\ell a_\ell s^*(z_\ell )=M\int as^*\, \mathrm d\rho _\theta \), and \(\log \frac{\, \mathrm d\rho ^{*\otimes M}}{\, \mathrm d\mu _U^{\otimes M}}(\theta ) =-\frac1\beta \sum _\ell a_\ell s^*(z_\ell )-M\log Z_*\).
For \(\rho \) with \(\int |a|\, \mathrm d\rho {\lt}\infty \) and every \(\rho '\in {\mathcal D}\), \(\operatorname {KL}(\rho '\| \hat\mu _\rho ){\lt}\infty \) and
(The chain rule \(\operatorname {KL}(\rho '\| \mu _U)=\operatorname {KL}(\rho '\| \hat\mu _\rho )-\frac1\beta \int a\, s_\rho \, \mathrm d\rho ' +\log \frac{Z_U}{Z_\rho }\), legitimate since \(\int |a s_\rho |\, \mathrm d\rho '\le c_\rho \int |a|\, \mathrm d\rho '{\lt}\infty \).)
Let \(\rho _N^*\) be the minimizer of \({\mathcal F}_N\) on \({\mathcal D}\) for the sample \((x_i,y_i)_{i=1}^N\). Then for every \(\rho \in {\mathcal D}\), \(\operatorname {KL}(\rho \| \rho _N^*){\lt}\infty \) and
(Lemma 212 for the empirical feature, through the dictionary of Lemma 61.)
Let \(f\) be a regression function of \(Q\) with \(|f|\le C\), \(|Y|\le {Y_{\max }}\) a.s., \(P_X\) the first marginal of \(Q\), and \(\rho ^*\) the minimizer of \({\mathcal F}\) on \({\mathcal D}\) (for the target \(f\in L^2(P_X)\)). Then for every \(\rho \in {\mathcal D}\), \(\operatorname {KL}(\rho \| \rho ^*){\lt}\infty \) and
(Lemma 212 and Lemma 288: \(\tilde{{\mathcal F}}-{\mathcal F}\) is the constant \(\tfrac 12{\mathbb E}(Y-f(X))^2\).)
For a sample \(\omega =(x_i,y_i)_{i=1}^N\) and the minimizer \(\rho _N^*\) of \({\mathcal F}_N\) on \({\mathcal D}\), every \(\rho \in {\mathcal D}\) satisfies \({\mathcal F}_N(\rho )-{\mathcal F}_N(\rho _N^*)=\tfrac 12\| F_\rho |_x-F_{\rho _N^*}|_x\| ^2_N +\beta \, \operatorname {KL}(\rho \| \rho _N^*)\).
For \(\lambda {\gt}0\) and every \(u\in L^2(\nu )\) with \(S_\nu u=f\), \(Q^\nu _\lambda (f)\le \| u\| ^2_{L^2(\nu )}\); under (R), \(Q^\nu _\lambda (f)\le \| S_\nu ^\dagger f\| ^2\). (Take \(u\) in the Tikhonov minimum, Proposition 744: \(\frac\lambda 2Q^\nu _\lambda (f)\le \frac12\| S_\nu u-f\| ^2+\frac\lambda 2\| u\| ^2 =\frac\lambda 2\| u\| ^2\).)
For \(0{\lt}\lambda \le \lambda '\), \(Q^\nu _{\lambda '}(f)\le Q^\nu _\lambda (f)\): the map \(\lambda \mapsto Q^\nu _\lambda (f)\) is nonincreasing. (Without spectral theory: \(Q^\nu _\lambda (f)=\min _u[\lambda ^{-1}\| S_\nu u-f\| ^2+\| u\| ^2]\) by Proposition 744, and the functional is pointwise nonincreasing in \(\lambda \).)
If \(f\notin \operatorname {ran}S_\nu \), then \(Q^\nu _\lambda (f)\uparrow +\infty \) as \(\lambda \downarrow 0\): for every \(M\) there is \(\lambda _0{\gt}0\) with \(Q^\nu _\lambda (f)\ge M\) for all \(0{\lt}\lambda \le \lambda _0\). Together with Lemma 749, \(\lim _{\lambda \downarrow 0}Q^\nu _\lambda (f){\lt}\infty \) if and only if (R) holds for \(\nu \). (Otherwise \(Q^\nu _\lambda (f)\) is bounded on \((0,\infty )\) by monotonicity, and Lemma 750 gives \(f\in \operatorname {ran}S_\nu \).)
Assume (R). Then \(Q^\nu _\lambda (f)\uparrow \| S_\nu ^\dagger f\| ^2_{L^2(\nu )}\) as \(\lambda \downarrow 0\): \(\lambda \mapsto Q^\nu _\lambda (f)\) is nonincreasing, bounded by \(\| S_\nu ^\dagger f\| ^2\), and converges to it. (\(Q^\nu _\lambda (f)=\langle f,g_\lambda \rangle =\langle S_\nu u^\dagger ,g_\lambda \rangle =\langle u^\dagger ,u_\lambda \rangle \) and \(u_\lambda \to u^\dagger \), Lemma 748.)
Let \(Z\) be a topological space with its Borel \(\sigma \)-algebra, let \(z\mapsto \varphi _z(x)\) be continuous for every \(x\), and let \(\rho _n\to \rho \) weakly in \({\mathcal P}(\Theta )\) with \(\sup _n\int a^2\, \mathrm d\rho _n\le C\) and \(\int a^2\, \mathrm d\rho \le C\). Then \(F_{\rho _n}(x)\to F_\rho (x)\) for every \(x\in {\mathcal X}\). Proof: for \(T{\gt}0\) the function \(\theta \mapsto \max (-T,\min (T,a))\, \varphi _z(x)\) is bounded and continuous, and replacing \(a\) by its clamped value changes \(F_\rho (x)\) by at most \(\int a^2\, \mathrm d\rho /T\le C/T\) (Lemma 214).
Let \(T\) be diagonal in \(e\) with eigenvalues \(0\le \mu _i\le 1\), \(\lambda {\gt}0\), \(a{\gt}0\) and \((T+\lambda )y=\lambda T^ag\), i.e. \(y=\lambda (T+\lambda )^{-1}T^ag\). Then \(\| y\| \le \lambda ^{\min (a,1)}\| g\| \) and \(\langle Ty,y\rangle \le \bigl(\lambda ^{\min (a+1/2,1)}\| g\| \bigr)^2\). (Coefficientwise \(\langle e_i,y\rangle =\frac{\lambda \mu _i^a}{\mu _i+\lambda }\langle e_i,g\rangle \) and \(\mu _i\langle e_i,y\rangle ^2=\bigl(\frac{\lambda \mu _i^{a+1/2}}{\mu _i+\lambda }\bigr)^2 \langle e_i,g\rangle ^2\); apply Lemma 795 and Parseval.)
For \(\lambda {\gt}0\), \(a{\gt}0\) and \(t\in [0,1]\), \(\frac{\lambda t^a}{t+\lambda }\le \lambda ^{\min (a,1)}\). For \(0{\lt}a\le 1\) the weighted arithmetic–geometric mean inequality \(\lambda ^{1-a}t^a\le (1-a)\lambda +at\le \lambda +t\) gives \(\frac{\lambda t^a}{t+\lambda }=\lambda ^a\frac{\lambda ^{1-a}t^a}{t+\lambda }\le \lambda ^a\); for \(a{\gt}1\), \(t^a\le t\) on \([0,1]\) gives \(\frac{\lambda t^a}{t+\lambda }\le \frac{\lambda t}{t+\lambda }\le \lambda \).
Under the hypotheses of Lemma 215, for \(P_X\in {\mathcal P}({\mathcal X})\) and \(f\in L^2(P_X)\), \(L(\rho _n)\to L(\rho )\). Proof: \(F_{\rho _n}(x)\to F_\rho (x)\) pointwise and \(|F_{\rho _n}|\le \sqrt C\), so \((F_{\rho _n}-f)^2\le 2C+2f^2\in L^1(P_X)\) and dominated convergence applies.
Let \(f\) be a regression function of \(Q\) with \(|f|\le C\), \(|Y|\le {Y_{\max }}\) a.s., and \(\rho \) with \(\int |a|\, \mathrm d\rho {\lt}\infty \). Then \(\tilde L(\rho )=L(\rho )+\tfrac 12\, {\mathbb E}(Y-f(X))^2=L(\rho )+\tfrac 12\, {\mathbb E}\, \mathrm{Var}(Y\mid X)\), since \({\mathbb E}[(F_\rho (X)-f(X))(f(X)-Y)]=0\) by the tower property.
Let \(F\) be a class with \(|F_i(S_k)|\le 1\) on the sample \(S_1,\dots ,S_N\) and \(\operatorname {Pdim}(F)\le d\). For every \({\varepsilon }{\gt}0\) the \(L_2(P_N)\) covering number of \(F\) on the sample satisfies \(\mathcal N({\varepsilon },F,L_2(P_N))\le \bigl(N(\lfloor 2/{\varepsilon }\rfloor +1)+1\bigr)^d\); in particular \(\mathcal N({\varepsilon },F,L_2(P_N))\le (4N/{\varepsilon })^d\) for \(0{\lt}{\varepsilon }\le 1\), \(N\ge 1\).
Let \(u,u_P,v,D,G,H\ge 0\) and \(c\ge 1\) with \(c^2H\le \frac18\), \(\frac12u^2+\frac12u_P^2+v^2\le DG+\frac12D^2H\) and \(D\le c(u+v+G)\). Then \(u+v\le 9cG\), \(u_P\le 6cG\) and \(D\le 10c^2G\). (With \(s=u+v\): \(\frac14s^2\le \frac12u^2+v^2\), \(D^2H\le \frac18(s+G)^2\), hence \(s^2\le 8cGs+9cG^2\) and \(s\le 9cG\).)
Let \(F\) be a class with \(|F_i(S_k)|\le 1\) on the sample \(S_1,\dots ,S_N\) and \(\operatorname {Pdim}(F)\le d\), \(d\ge 1\), and let \(0{\lt}{\varepsilon }\le 1\). Every \({\varepsilon }\)-separated finite subset of \(F\) in \(L_2(P_N)\) has at most \((8/{\varepsilon })^{6d}\) elements.
Let \((e_i)_{i\in I}\) be an orthonormal basis of \(L^2(P_X)\) with \(K_\nu e_i=\mu _ie_i\) (so \(\mu _i=\| S_\nu ^*e_i\| ^2\in [0,1]\)). Then \(f\in \operatorname {ran}S_\nu \iff \sum _i\mu _i^{-1}\langle f,e_i\rangle ^2{\lt}\infty \) and \(\langle f,e_i\rangle =0\) whenever \(\mu _i=0\). In that case \(u_0^\dagger =S_\nu ^\dagger f=\sum _i\mu _i^{-1}\langle f,e_i\rangle S_\nu ^*e_i\), the series converging in \(L^2(\nu )\) with mutually orthogonal terms, and \(\| u_0^\dagger \| ^2_{L^2(\nu )}=\sum _i\mu _i^{-1}\langle f,e_i\rangle ^2\). (Singular system: \(v_i:=\mu _i^{-1/2}S_\nu ^*e_i\) is orthonormal with \(S_\nu v_i=\mu _i^{1/2}e_i\); “\(\Rightarrow \)” is Bessel’s inequality for \(\mu _i^{-1/2}\langle S_\nu u,e_i\rangle =\langle u,v_i\rangle \), “\(\Leftarrow \)” is the termwise synthesis of \(u=\sum _i\mu _i^{-1/2}\langle f,e_i\rangle v_i\), which lies in \(\overline{\operatorname {ran}S_\nu ^*}=(\ker S_\nu )^\perp \).)
The source condition of order \(1/2\) of Definition 139, \(f=K_{\nu _0}g\) with \(\| g\| \le G\), implies (SC\(_{1/2}\)) of Definition 165 with the same bound: \(\operatorname {ran}S_{\nu _0}^*\subseteq \operatorname {ran}T_0^{1/2}\) with \(S_{\nu _0}^*g=T_0^{1/2}g_0\), \(\| g_0\| \le \| g\| \) (Lemma 797).
Let \(T\) be diagonal in two Hilbert bases \(e=(e_i)_{i\in I}\) and \(e'=(e'_j)_{j\in J}\) of \(H\), \(Te_i=\mu _ie_i\) with \(0\le \mu _i\le M\) and \(Te'_j=\mu '_je'_j\). Then for every \(a\ge 0\) the spectral powers of Definition 789 defined through \(e\) and through \(e'\) coincide: by Lemma 793 both map \(e'_j\) to \(\mu _j'{}^ae'_j\), and a bounded operator is determined by its values on a Hilbert basis.
Let \(T\) be diagonal in the Hilbert basis \(e\) with eigenvalues \(0\le \mu _i\le M\), and \(a\ge 0\). If \(Tv=\lambda v\), then \(T^av=\lambda ^av\). (Coordinatewise: \(\langle e_i,T^av\rangle =\mu _i^a\langle e_i,v\rangle \) and \(\langle e_i,v\rangle =0\) unless \(\mu _i=\lambda \), since eigenvectors of the self-adjoint \(T\) for distinct eigenvalues are orthogonal.)
Let \(T=B^*B\) be diagonal in \(e\) with eigenvalues \(\mu _i\in [0,M]\). Then \(\operatorname {ran}B^*\subseteq \operatorname {ran}T^{1/2}\): for \(g\) in the target space of \(B\), \(g_0:=\sum _{\mu _i\ne 0}\mu _i^{-1/2}\langle Be_i,g\rangle e_i\) satisfies \(T^{1/2}g_0=B^*g\) and \(\| g_0\| \le \| g\| \) (Bessel’s inequality for the orthonormal family \((\mu _i^{-1/2}Be_i)_{\mu _i\ne 0}\); \(Be_i=0\) when \(\mu _i=\| Be_i\| ^2=0\)).
For \(a{\gt}0\), \(\operatorname {ran}T^a\subseteq (\ker T)^\perp =\overline{\operatorname {ran}T}\): if \(Tv=0\) then \(\mu _i\langle e_i,v\rangle =0\) for every \(i\), so each term of \(\langle v,T^ag\rangle =\sum _i\mu _i^a\langle v,e_i\rangle \langle e_i,g\rangle \) vanishes (\(0^a=0\)).
For every \(\nu \in {\mathcal P}(Z)\) and \(f\in \operatorname {ran}S_\nu \), \(\sup _{z\in Z}|R_{\lambda ,\nu }f(z)|\le \| S_\nu ^\dagger f\| _{L^2(\nu )}/(2\sqrt\lambda )\); more generally the bound holds with \(\| h\| _{L^2(\nu )}\) for every representer \(h\) of \(f\). (With \(g:=(K_\nu +\lambda )^{-1}f\), \(|R_{\lambda ,\nu }f(z)|\le \| g\| \) and \(\| S_\nu ^*g\| ^2+\lambda \| g\| ^2=\langle (K_\nu +\lambda )g,g\rangle =\langle h,S_\nu ^*g\rangle \le \| h\| \| S_\nu ^*g\| \), hence \(\lambda \| g\| ^2\le \| h\| ^2/4\).)
For \(m^*=R_{\lambda ,\nu ^*}f\),
(\(h:=u^\dagger /w\in L^2(\nu ^*)\) with \(\| h\| ^2_{L^2(\nu ^*)}=\int u^{\dagger 2}w^{-1}\, \mathrm d\nu _0 \le q\| u^\dagger \| ^2\) and \(S_{\nu ^*}h=S_{\nu _0}u^\dagger =f\); apply Lemma 466.)
Let \(Z\) be standard Borel. For every \(\ell \), the one-particle marginal \(\pi _M^{(\ell )}\) of \(\pi _M\) satisfies \(M\, \operatorname {KL}(\pi _M^{(\ell )}\| \rho ^*)=\sum _{\ell '}\operatorname {KL}(\pi _M^{(\ell ')}\| \rho ^*) \le \operatorname {KL}(\pi _M\| \rho ^{*\otimes M})\), by exchangeability and the superadditivity \(\operatorname {KL}(\mu \| \nu ^{\otimes M})\ge \sum _\ell \operatorname {KL}(\mu ^{(\ell )}\| \nu )\).
Let \(\rho \in {\mathcal P}(\Theta )\) with \(\int a^2\, \mathrm d\rho {\lt}\infty \) and \(M\ge 1\). Then
(For each \(x\), \(F_{\rho _\theta }(x)=\frac1M\sum _\ell a_\ell \varphi _{z_\ell }(x)\) is a mean of i.i.d. variables of mean \(F_\rho (x)\); Fubini is justified by \(\int a^2\, \mathrm d\rho {\lt}\infty \).)
Under the assumptions of Theorem 328, for every realization of the sample, deterministically,
(Add the two quadratic expansions of Lemma ??; the entropy terms \(\beta \operatorname {KL}(\cdot \| \mu _U)\) cancel, and \(\ell _{F_{\rho _N^*}}-\ell _{F_{\rho ^*}}=he+\frac12h^2\).)
For each direction separately, \(\tfrac 12\| h\| _N^2+\beta \operatorname {KL}(\rho ^*\| \rho _N^*)\le -(P_N-P)[he+\tfrac 12h^2]\) and \(\tfrac 12\| h\| ^2+\beta \operatorname {KL}(\rho _N^*\| \rho ^*)\le -(P_N-P)[he+\tfrac 12h^2]\) (drop one of the two nonnegative expansions).
\(S_{\nu _0}\) is Hilbert–Schmidt with \(\| S_{\nu _0}\| _{HS}^2 =\sum _i\| S_{\nu _0}e_i\| ^2\le 1\) along any Hilbert basis \((e_i)\) of \(L^2(\nu _0)\): with \(\varphi _x:=\varphi _\cdot (x)\in L^2(\nu _0)\), \(\| S_{\nu _0}u\| ^2=\int \langle u,\varphi _x\rangle ^2 \, P_X(\, \mathrm dx)\), and Bessel’s inequality gives \(\sum _{i\in s}\langle e_i,\varphi _x\rangle ^2 \le \| \varphi _x\| ^2\le 1\) for every finite \(s\). Consequently \(\sum _i\langle T_0e_i,e_i\rangle =\sum _i\| S_{\nu _0}e_i\| ^2\le 1\) and \(T_0\) is compact.
(a) The pseudo-dimension of the class \(\Phi =\{ \varphi _z:z\in Z\} \) of \(\tanh \) ridge functions \(x\mapsto \tanh (w^\top x-b)\), \((w,b)\in {\mathbb R}^m\times {\mathbb R}\), is at most \(m+1\). (b) There is an absolute constant \(C_{\mathrm P}\) such that for every \(N\) and every \(x_1,\dots ,x_N\), \({\widehat{\mathfrak R}}_N(\Phi )\le C_{\mathrm P}\sqrt{(m+1)/N}\), hence \({\mathfrak R}_N(\Phi )\le C_{\mathrm P}\sqrt{(m+1)/N}\). A crude evaluation through Haussler’s covering bound and Dudley’s entropy integral gives \(C_{\mathrm P}\le 43\).
Assume \(\mathrm{HC}(m,A,c)\) for all \(m\), with \(2A\ge 1\), \(c\ge 1\). Then for every \(m\), \(N\) and \(x_1,\dots ,x_N\in {\mathbb R}^m\), \({\widehat{\mathfrak R}}_N(\Phi )\le \bigl(1+6\sqrt{\log (2A)}+12\sqrt2\bigr)\sqrt c\, \sqrt{(m+1)/N}\), i.e. Lemma 314(b) with \(C_{\mathrm P}=(1+6\sqrt{\log (2A)}+12\sqrt2)\sqrt c\).
For the tanh feature (A3) with \(z=(w,b)\in {\mathbb R}^{m+1}\) (Euclidean norm), \(\| \nabla _z\varphi _z(x)\| \le \sqrt{|x|^2+1}\) for all \(x\in {\mathbb R}^m\) and \(z\in {\mathbb R}^{m+1}\). Indeed \(\nabla _z\varphi _z(x)=\tanh '(w^\top x-b)\, (x,-1)\) with \(0{\lt}\tanh '\le 1\) and \(|(x,-1)|=\sqrt{|x|^2+1}\). In particular, under (A1) the feature is Lipschitz in the hidden parameter with constant \(\sqrt{R_X^2+1}\).
For the tanh feature (A3), with \(z=(w,b)\in {\mathbb R}^{m+1}\) (Euclidean norm, Definition 6), and inputs \(|x|\le R_X\) (A1), the map \(z\mapsto \varphi _z(x)\) is \(C^\infty \) and, for every \(n\ge 0\),
where \(\| \tanh ^{(n)}\| _\infty \le \sup _{|s|\le 1}|p_n(s)|\) for the polynomials \(p_0=X\), \(p_{n+1}=p_n'\, (1-X^2)\) with \(\tanh ^{(n)}=p_n(\tanh )\). Hence the restriction of the tanh feature to \(\{ |x|\le R_X\} \) (Definition 3) is smooth with bounded derivatives in the sense of Definition 127.
Assume (A1), (A3) and (SC\(_a\)) (hence (R)). Then
(For \(a\in \{ 1/2,1\} \): Lemmas 454 and 455, with \(\lambda ^{1/2}=\sqrt\lambda \) and \(\min (a+1/2,1)=1\).)
Under (SC\(_{1/2}\)), \(f=K_{\nu _0}g\) with \(g\in L^2(P_X)\): \(\| u_\lambda -u^\dagger \| _{L^2(\nu _0)}\le \frac{\sqrt\lambda }2\| g\| \le \sqrt\lambda \| g\| \) and \(\| S_{\nu _0}(u_\lambda -u^\dagger )\| _{L^2(P_X)}\le \lambda \| g\| \). (With \(y:=(T_0+\lambda )^{-1}S_{\nu _0}^*g\) one has \(u_\lambda -u^\dagger =-\lambda y\), and the normal equation \((T_0+\lambda )y=S_{\nu _0}^*g\) gives \(\| S_{\nu _0}y\| ^2+\lambda \| y\| ^2\le \| g\| \, \| S_{\nu _0}y\| \), hence \(\lambda \| y\| ^2\le \| g\| ^2/4\) and \(\| S_{\nu _0}y\| \le \| g\| \).)
Assume (A1), (A3) and (R). Then \(u_\lambda -u^\dagger =-\lambda (T_0+\lambda )^{-1}u^\dagger \), i.e. \((T_0+\lambda )(u_\lambda -u^\dagger )=-\lambda u^\dagger \). (\(f=S_{\nu _0}u^\dagger \) gives \(S_{\nu _0}^*f=T_0u^\dagger \), so \((T_0+\lambda )u_\lambda -(T_0+\lambda )u^\dagger =T_0u^\dagger -(T_0+\lambda )u^\dagger =-\lambda u^\dagger \).)
Under (SC\(_1\)), \(u^\dagger =T_0g_0\) with \(g_0\in L^2(\nu _0)\): \(\| u_\lambda -u^\dagger \| _{L^2(\nu _0)}\le \lambda \| g_0\| \) and \(\| S_{\nu _0}(u_\lambda -u^\dagger )\| _{L^2(P_X)}\le \lambda \| g_0\| \). (\(u_\lambda -u^\dagger =-\lambda (T_0+\lambda )^{-1}T_0g_0\) and \(\| T_0(T_0+\lambda )^{-1}\| \le 1\); the residual bound follows from \(\| S_{\nu _0}\| \le 1\).)
Assume (A1), (A3) and (SC\(_a\)) with \(a{\gt}0\) (hence (R)). Then
(\(u_\lambda -u^\dagger =-\lambda (T_0+\lambda )^{-1}T_0^ag_0\) (Lemma 453), the eigenvalues of \(T_0\) lie in \([0,1]\), and Lemma 796 applies; \(\| S_{\nu _0}h\| ^2=\langle T_0h,h\rangle \).)
Let \(F=\{ F_i:i\in \iota \} \) be a nonempty class of functions on \({\mathcal X}\) with \(|F_i(S_k)|\le M\) on the sample \(S_1,\dots ,S_N\), and let \(\psi _x\colon {\mathbb R}\to {\mathbb R}\) be \(L\)-Lipschitz for every \(x\) (\(L\ge 0\)). Then, in the one-sided convention, \({\widehat{\mathfrak R}}_N(\{ \psi \circ F_i\} )\le L\, {\widehat{\mathfrak R}}_N(F)\), where \((\psi \circ F_i)(x)=\psi _x(F_i(x))\). This extends FoML’s finite-class contraction principle (‘empiricalRademacherComplexity_without_abs_contraction_finite‘) to an arbitrary index set: for every sign vector \(\sigma \) pick a near-maximizer \(i_\sigma \) of \(\sum _k\sigma _k\psi _{S_k}(F_i(S_k))\); the finite subclass \(\{ F_{i_\sigma }\} _\sigma \) realizes the left-hand side up to \(\varepsilon \) and its right-hand side is dominated by that of \(F\).
Let \(G\) be a measurable function of \(N\) i.i.d. samples \(\omega _1,\dots , \omega _N\sim \mu \) such that replacing any single \(\omega _k\) changes \(G\) by at most \(c{\gt}0\). Then for every \(\varepsilon \ge 0\), \({\mathbb P}\bigl(G-{\mathbb E}G\ge \varepsilon \bigr)\le \exp \bigl(-2\varepsilon ^2/(Nc^2)\bigr)\) and \({\mathbb P}\bigl(G-{\mathbb E}G\le -\varepsilon \bigr)\le \exp \bigl(-2\varepsilon ^2/(Nc^2)\bigr)\) (FoML’s ‘mcdiarmid_inequality_pos_iid_of_const‘ with \(t=1/(Nc^2)\)).
Let \(F=\{ f_i\} \) be a class with a pointwise dense sequence \((f_{e_n})_n\) and let \(\Phi \colon \iota \to {\mathbb R}\) be such that \(\Phi (u_n)\to \Phi (i)\) whenever \(f_{u_n}\to f_i\) pointwise. Then \(\sup _{i\in \iota }\Phi (i)=\sup _n\Phi (e_n)\) (with the convention that both sides are \(0\) if \(\Phi \) is unbounded). In particular the empirical Rademacher complexities and the uniform deviations of a bounded separable class are those of the countable subclass \(\{ f_{e_n}\} \).
Let \(F=\{ f_i:i\in \iota \} \) be a nonempty separable class of measurable functions on \(\mathcal Z\) with \(|f_i|\le b\), let \(\mu \in {\mathcal P}(\mathcal Z)\), \(N\ge 1\), and let \(\omega _1,\dots ,\omega _N\sim \mu \) be i.i.d. Then, in the one-sided convention,
This is the one-sided form of FoML’s symmetrization (‘expectation_le_rademacher‘, from ‘symmetrization_equation‘) with the countability of the class relaxed to separability: all suprema are computed on a pointwise dense sequence.
Let \(F=\{ f_i:i\in \iota \} \) be a nonempty separable class of measurable functions on \({\mathcal X}\) with \(|f_i|\le b\), \(b{\gt}0\), let \(X\colon \Omega \to {\mathcal X}\) be measurable, \(\mu \in {\mathcal P}(\Omega )\) and \(\omega _1,\dots ,\omega _n\sim \mu \) i.i.d. Then for every \(\varepsilon \ge 0\),
This is FoML’s ‘uniform_deviation_tail_bound_separable_of_pos‘ with its topological hypotheses replaced by the separability of the class.
Under the hypotheses of Lemma 802 with \(b{\gt}0\), for every \(\varepsilon \ge 0\),
and the same for \(\sup _i(\frac1N\sum _kf_i(\omega _k)-\mu (f_i))\). The one-sided deviation is measurable (as a supremum over the dense sequence) and has bounded differences \(2b/N\), so McDiarmid’s inequality (Lemma 799) applies.
Let \(Z\) be separable and first countable, \(\varphi \) continuous in \(z\), \(Q\in {\mathcal P}({\mathcal X}\times {\mathbb R})\) with \(|Y|\le {Y_{\max }}\) a.s., \(B{\gt}0\), \({Y_{\max }}\ge 0\) and \(N\ge 1\). Let \(G:=\sup _{\rho \in {\mathcal P}_B}|L_N(\rho )-\tilde L(\rho )|\). Then for every \(\delta \in (0,1)\), with probability at least \(1-\delta \) over the i.i.d. sample of size \(N\) from \(Q\),
The proof is symmetrization (one-sided), the contraction principle, the identity \({\widehat{\mathfrak R}}_N({\mathcal{H}}_B)=B{\widehat{\mathfrak R}}_N(\Phi )\), McDiarmid’s inequality for \(G_\pm \) with bounded differences \((B+{Y_{\max }})^2/(2N)\) and a union bound; the suprema over \({\mathcal P}_B\) are computed on the countable dense subclass \({\mathcal P}_B^0\).
Let \(Z\) be separable and first countable, \(\varphi \) continuous in \(z\), \(Q\in {\mathcal P}({\mathcal X}\times {\mathbb R})\) with \(|Y|\le {Y_{\max }}\) a.s., \(B\ge 0\), \({Y_{\max }}\ge 0\) and \(N\ge 1\). Then \({\mathbb E}\, G_\pm \le 2B(B+{Y_{\max }})\, {\mathfrak R}_N(\Phi )\).
Let \(B\ge 0\) and \(x_1,\dots ,x_N\in {\mathcal X}\). Then, in the one-sided convention of the paper, \({\widehat{\mathfrak R}}_N({\mathcal{H}}_B)=B\, {\widehat{\mathfrak R}}_N(\Phi )\) for \({\mathcal{H}}_B=\{ F_\rho :\rho \in {\mathcal P}_B\} \), where \({\widehat{\mathfrak R}}_N(\Phi )={\mathbb E}_\sigma \sup _z|\psi _\sigma (z)|\) is ‘featureRademacher‘.
Let \(B\ge 0\), \(x_1,\dots ,x_N\in {\mathcal X}\) and \(\sigma \in \{ \pm 1\} ^N\), and let \(\psi _\sigma (z):=\frac1N\sum _i\sigma _i\varphi _z(x_i)\). Then \(\sup _{\rho \in {\mathcal P}_B}\frac1N\sum _i\sigma _iF_\rho (x_i)=B\sup _{z\in Z}|\psi _\sigma (z)|\) (and the same with \(\bigl|\frac1N\sum _i\sigma _iF_\rho (x_i)\bigr|\) on the left).
Let \(E\) be a real inner product space of dimension \(d\) and let \(S\subset E\) have \(d+2\) points. There is \(T\subseteq S\) such that for no \((w,c)\in E\times {\mathbb R}\) one has \(\{ v\in S:\langle w,v\rangle +c{\gt}0\} =T\). Hence the VC dimension of the class of affine halfspaces of \(E\) is at most \(d+1\).
Assume the hypotheses of Lemma 459 and (R). Then \(w=q^{-1}e^{\zeta }\) with \(q=Z_*/Z_0=\tilde Z\), and pointwise
(\(1\le q\le e^{\zeta _0}\); \(e^{\zeta }-1\le \zeta e^{\zeta }\); \(1-q^{-1}\le \log q\le \zeta _0\); \(1-e^{-\zeta }\le \zeta \).)
(Difference from \(u_\lambda \).) \((T_0+\lambda )(v-u_\lambda )=\lambda (1-w^{-1})v=\lambda (w-1)m^*\) and \(\| v-u_\lambda \| _{L^2(\nu _0)}\le \| (w-1)m^*\| _{L^2(\nu _0)}=\Xi \). (Subtract \((T_0+\lambda )u_\lambda =S_{\nu _0}^*f\) from the weighted normal equation, and use \(\| \lambda (T_0+\lambda )^{-1}\| \le 1\).)
(Weighted normal equation.) With \(M_{w^{-1}}\) the multiplication operator by \(w^{-1}\), in \(L^2(\nu _0)\) and pointwise, \((T_0+\lambda M_{w^{-1}})\, v=S_{\nu _0}^*f\), where \(w^{-1}v=m^*\). (By Theorem 228, \(s^*=S^*r^*=-\lambda m^*\) for every \(z\); since \(r^*=F_{\rho ^*}-f=S_{\nu _0}v-f\), \(S^*r^*=T_0v-S_{\nu _0}^*f\), hence \(T_0v-S_{\nu _0}^*f=-\lambda m^*=-\lambda w^{-1}v\).)
With \(v:=w\, m^*\): \(v\) is bounded, \(v\in L^2(\nu _0)\) and \(S_{\nu _0}v=S_{\nu ^*}m^*=F_{\rho ^*}\) (by \(|m^*\varphi _z|\le \| m^*\| _\infty \) and Fubini, \(S_{\nu ^*}m^*(x)=\int m^*(z)\varphi _z(x)w(z)\nu _0(\, \mathrm dz)=(S_{\nu _0}v)(x)\), which is \(F_{\rho ^*}\) by Theorem 231).
(Variational interpretation.) \(v\) is the unique minimizer of the weighted Tikhonov functional \({\mathcal{J}}_w(u):=\tfrac 12\| S_{\nu _0}u-f\| ^2_{L^2(P_X)}+\frac\lambda 2\int _Zu^2w^{-1}\, \mathrm d\nu _0\) (\(u\in L^2(\nu _0)\)): for every \(u\), \({\mathcal{J}}_w(u)={\mathcal{J}}_w(v)+\tfrac 12\| S_{\nu _0}(u-v)\| ^2+\frac\lambda 2\int (u-v)^2w^{-1}\, \mathrm d\nu _0\), since the cross terms vanish by the weighted normal equation.
For \(\sigma =\operatorname {erf}\) and \(\sigma _b{\gt}0\) there are \(\eta {\gt}0\) and \(C{\lt}\infty \) with \(\mu _\ell \le Ce^{-\eta |\ell |}\): \(\phi \mapsto \tilde\kappa (\cos \phi )\) is \(2\pi \)-periodic and holomorphic on a strip \(|\Im \phi |{\lt}\eta \), since \(|2(\sigma _w^2t+\sigma _b^2)/(1+2s^2)|{\lt}1\) on \([-1,1]\), and its Fourier coefficients decay geometrically.
For \(\sigma =\operatorname {erf}\), \(\sigma _b{\gt}0\), (R) implies \(\sum _\ell e^{\eta |\ell |}\| P_\ell f\| ^2{\lt}\infty \) for some \(\eta {\gt}0\), so \(f\) is real analytic in \(\theta \) (holomorphic extension to a strip \(|\Im \theta |{\lt}\eta /2\)): \(\| P_\ell f\| ^2\le \| u_0^\dagger \| ^2\mu _\ell \le C\| u_0^\dagger \| ^2e^{-\eta |\ell |}\).
For \(\sigma =\tanh \) and \(\sigma _b{\gt}0\), \(\mu _\ell =O(|\ell |^{-p})\) for every \(p{\gt}0\): \(\sum _kk^{2p}a_k(s)^2=\| N^p\tanh (s\cdot )\| ^2_{L^2(\gamma )}{\lt}\infty \) for the Ornstein–Uhlenbeck number operator \(N\), so \(\phi \mapsto \tilde\kappa (\cos \phi )\) is \(C^\infty \) and its Fourier coefficients decay faster than any power.
Assume (A1), (A3), (A4), (A5) and \(f\in L^2(P_X)\). Let \(\rho ^*\) be the minimizer of \({\mathcal F}\) (Lemma ??). For every \(u\in L^2(\nu _0)\),
(In the proof of Theorem 397, take the competitor \(\tilde\rho _u=\nu _0\otimes {\mathcal N}(u,\beta /\lambda )\) for an arbitrary \(u\): \(L(\tilde\rho _u)=\frac12\| S_{\nu _0}u-f\| ^2\) and \(\operatorname {KL}(\tilde\rho _u\| \mu _U)=\frac\lambda {2\beta }\| u\| ^2\), Lemma 755; minimality \({\mathcal F}(\rho ^*)\le {\mathcal F}(\tilde\rho _u)\).)
For \(\lambda {\gt}0\), \(f\in L^2(P_X)\) and \(u_\lambda :=S_\nu ^*(K_\nu +\lambda )^{-1}f =R_{\lambda ,\nu }f\), one has \(S_\nu u_\lambda -f=-\lambda (K_\nu +\lambda )^{-1}f\) and \(\frac12\| S_\nu u_\lambda -f\| ^2+\frac\lambda 2\| u_\lambda \| ^2 =\frac\lambda 2\langle f,(K_\nu +\lambda )^{-1}(\lambda +K_\nu )(K_\nu +\lambda )^{-1}f\rangle =\frac\lambda 2Q^\nu _\lambda (f)\); by Lemma ??(4) this is the minimum of the Tikhonov functional \(u\mapsto \frac12\| S_\nu u-f\| ^2+\frac\lambda 2\| u\| ^2\) over \(L^2(\nu )\).
Under the hypotheses of Proposition 756, with \(\nu ^*\) the hidden marginal of \(\rho ^*\), \(m^*\) its conditional mean amplitude and \(\kappa =\lambda /\beta \),
and \(\| m^*\| ^2_{L^2(\nu ^*)}+\frac{2\beta }\lambda \operatorname {KL}(\nu ^*\| \nu _0)+\frac2\lambda L(\rho ^*) \le Q^{\nu _0}_\lambda (f)\). (The chain rule \(\operatorname {KL}(\rho ^*\| \mu _U)=\operatorname {KL}(\nu ^*\| \nu _0) +\frac\kappa 2\| m^*\| ^2_{L^2(\nu ^*)}\) of Theorem 396, whose proof uses only the Gaussian conditional structure of Theorem 225 and not (R), substituted into Proposition 757.)
Under the hypotheses of Proposition 756, the infimum of the right-hand side of ?? over \(u\in L^2(\nu _0)\) is attained at \(u=R_{\lambda ,\nu _0}f\) with value \(\frac\lambda 2Q^{\nu _0}_\lambda (f)\) (Proposition 744), and therefore \(L(\rho ^*)+\beta \, \operatorname {KL}(\rho ^*\| \mu _U)\le \frac\lambda 2Q^{\nu _0}_\lambda (f)\).
Assume (A1), (A3), (A4), (A5), and that \(P_X\) has finitely many atoms \(x_1,\dots ,x_n\), so that \(L^2(P_X)\cong {\mathbb R}^n\). Suppose \(K_{\nu _0}\) is invertible on \(L^2(P_X)\) with smallest eigenvalue \(\sigma _0{\gt}0\). Then (R) holds automatically, and for \(\kappa \le \kappa _0\),
(\(K_{\nu ^*}=\int \Phi _z\Phi _z^\top w\, \mathrm d\nu _0\) and \(w\ge e^{-\zeta _0}\) give \(K_{\nu ^*}\succeq e^{-\zeta _0}K_{\nu _0}\succeq e^{-\zeta _0}\sigma _0I\), hence \(\| (K_{\nu ^*}+\lambda )^{-1}\| \le e^{\zeta _0}/\sigma _0\) and \(|m^*(z)|\le \| (K_{\nu ^*}+\lambda )^{-1}f\| \). In Lean the hypothesis is the coercivity \(\sigma _0\| h\| ^2\le \langle K_{\nu _0}h,h\rangle \) for a general \(P_X\).)
Suppose \(K_{\nu _0}\) is coercive on \(L^2(P_X)\), \(\langle K_{\nu _0}h,h\rangle \ge \sigma _0\| h\| ^2\) with \(\sigma _0{\gt}0\) (in the paper: \(P_X\) has finitely many atoms and \(K_{\nu _0}\) is invertible with smallest eigenvalue \(\sigma _0\)). Then (R) holds for every \(f\): \(f=S_{\nu _0}(S_{\nu _0}^*K_{\nu _0}^{-1}f)\).
Let \(\nu _0\), \(\delta =\| K_{\nu _0}-K_\infty \| \) be as in Theorem 776 and let \(f\in L^2(P_X)\) satisfy \(Q^{(\infty )}_s(f)\le H_f\) for all \(s{\gt}0\) (for the sign network and \(f\in H^1(-1,1)\), \(H_f=\| f\| ^2_{{\mathcal{H}}_\infty }\) by the Sobolev characterization of the limit RKHS); put \(h_\lambda :=(K_\infty +\lambda )^{-1}f\). Then \(\| h_\lambda \| \le \frac{\sqrt{H_f}}{2\sqrt\lambda }\) and
(Lemma 782; the resolvent identity gives \(Q^{\nu _0}_\lambda (f)-Q^{(\infty )}_\lambda (f)=\langle (K_{\nu _0}+\lambda )^{-1}f, (K_\infty -K_{\nu _0})h_\lambda \rangle \le \sqrt{Q^{\nu _0}_\lambda (f)/\lambda }\; \delta \, \frac{\sqrt{H_f}}{2\sqrt\lambda }\), and \(x\le A+c\sqrt{Ax}\) implies \(\sqrt x\le (1+c)\sqrt A\).)
(Realization and hidden marginal, unconditionally.) Under the hypotheses of Proposition 783, for the minimizer \(\rho ^*\) for \((\lambda ,\beta ,\nu _0)\),
Hence \(F_{\rho ^*_L}\to f\) in \(L^2(P_X)\) along \(\lambda _L\to 0\), \(\delta _L^2/\lambda _L\to 0\) for every \(\beta \), and the symmetric relative entropy tends to zero if moreover \(\kappa _L\to 0\). (Proposition 759, Lemma 767 with \(u=R_{\lambda ,\nu _0}f\), and (a).)
Under the hypotheses of Proposition 784, along a sequence \((\nu _L,\lambda _L,\beta _L)\) with \(\lambda _L\to 0\) and \(\delta _L^2/\lambda _L\to 0\), \(F_{\rho ^*_L}\to f\) in \(L^2(P_X)\) for every choice of \(\beta _L\); if moreover \(\kappa _L(1+\delta _L/(2\lambda _L))^2\to 0\) (e.g. \(\kappa _L\to 0\) with \(\delta _L/\lambda _L\) bounded) then \(\operatorname {KL}(\nu ^*_L\| \nu _L)+\operatorname {KL}(\nu _L\| \nu ^*_L)\to 0\).
(Amplitude, high-temperature regime.) Under the hypotheses of Proposition 783, with \(R^{(\infty )}_\lambda f=S^*h_\lambda \),
so that \(\| m^*_L-R^{(\infty )}_{\lambda _L}f\| _{L^2(\nu ^*_L)}\to 0\) whenever \(\delta _L/\lambda _L\to 0\) and \(\sqrt{\kappa _L}/\lambda _L\to 0\). (Lemma 771 with \(h=h_\lambda \): \((K_{\nu _0}+\lambda )h_\lambda -f=(K_{\nu _0}-K_\infty )h_\lambda \) has norm at most \(\delta \| h_\lambda \| \); then (a).)
Under the hypotheses of Proposition 785, along a sequence with \(\delta _L/\lambda _L\to 0\) and \(\sqrt{\kappa _L}/\lambda _L=(\lambda _L\beta _L)^{-1/2}\to 0\), \(\| m^*_L-R^{(\infty )}_{\lambda _L}f\| _{L^2(\nu ^*_L)}\to 0\) (?? with \(Q^{\nu _L}_{\lambda _L}(f)\le H_f(1+\delta _L/(2\lambda _L))^2\) bounded).
For \(0{\lt}\alpha {\lt}1\), \(L\ge 2\) and \(x,x'\in [-1,1]\), \(0\le {\mathcal{E}}_L(x,x')\le 2^{3-\alpha }L^{\alpha -1}+8e^{-L}M_L\) ??: by ??, \({\mathcal{E}}_L\le 8\int _{1/L}^Le^{-2(L-w)}w^{\alpha -1}\, \mathrm dw\); on \([1/L,L/2]\), \(e^{-2(L-w)}\le e^{-L}\), and on \([L/2,L]\), \(w^{\alpha -1}\le (L/2)^{\alpha -1}\) and \(\int _{L/2}^Le^{-2(L-w)}\, \mathrm dw\le \frac12\).
For \(\sigma {\gt}0\) with \(\sigma \| u\| ^2\le \langle u,Bu\rangle \) on \((\ker S_N)^\perp \) and \(\lambda {\gt}0\), \(\| \gamma _\lambda -S_N^\dagger y\| _M\le \lambda \| S_N^\dagger y\| _M/\sigma \); in particular \(\gamma _\lambda \to S_N^\dagger y\) as \(\lambda \downarrow 0\).
If the components of \(\gamma _0\) are centered, pairwise uncorrelated with variance \(\tau ^2\) (e.g. i.i.d.) and independent of \(z_1,\dots ,z_M\), then \({\mathbb E}\bigl[\| P_{\ker S_N}\gamma _0\| _M^2\, \big|\, z\bigr]=\frac{\tau ^2\dim \ker S_N}{M} \ge \tau ^2\bigl(1-\frac NM\bigr)\) ??.
For \(\lambda =0\), a solution \(\gamma _t\) of ??, \(\gamma _\infty :=S_N^\dagger y+P_{\ker S_N}\gamma _0\) and any \(\sigma {\gt}0\) with \(\sigma \| u\| ^2\le \langle u,Bu\rangle \) on \((\ker S_N)^\perp \) (in particular \(\sigma =\sigma _{\min }\)), \(\| \gamma _t-\gamma _\infty \| _M\le e^{-\sigma t}\| P_{(\ker S_N)^\perp }\gamma _0-S_N^\dagger y\| _M\) for \(t\ge 0\).
Let \(g\colon Z\to {\mathbb R}\) be measurable with \(|g|\le C\), \(m:=\int ag(z)\, \mathrm d\rho ^* =\int g\, \mathrm d\Pi \rho ^*\), \(\bar X(\theta ):=\frac1M\sum _\ell a_\ell g(z_\ell )=\int g\, \mathrm d\Pi \rho _\theta \) and \(K_g:=\int (ag(z)-m)^2e^{|ag(z)-m|}\, \mathrm d\rho ^*{\lt}\infty \). Then
(Donsker–Varadhan on \(\Theta ^M\) with \(h=\sqrt M|\bar X-m|\): with \(u=1/\sqrt M\) and \(Y_\ell =a_\ell g(z_\ell )-m\) i.i.d. centered, \({\mathbb E}_{\rho ^{*\otimes M}}e^{\pm u\sum _\ell Y_\ell }=({\mathbb E}e^{\pm uY})^M\le (1+K_gu^2)^M\le e^{K_g}\), so \(\log {\mathbb E}_{\rho ^{*\otimes M}}e^h\le \log 2+K_g\).)
Let \(\lambda \ge 0\), \(P\) a measure on \({\mathcal X}\), \(f\colon {\mathcal X}\to {\mathbb R}\) and \(\nu \in {\mathcal P}(Z)\) (so that \({\mathcal P}_2\ne \emptyset \)). Then \(\inf _{\rho \in {\mathcal P}_2}\bigl\{ L(\rho )+\frac\lambda 2\int a^2\, \mathrm d\rho \bigr\} =\inf _{\gamma \in {\mathcal M}(Z)}\bigl\{ L(\gamma )+\frac\lambda 2\| \gamma \| ^2_{\mathcal M}\bigr\} \), both sides being infima of nonnegative sets.
(Theorem ??(b).) Let \(\mu _k(K^{(L)})\) satisfy \(|\mu _k(K^{(L)})-\mu _k(K^{(\infty )})|\le {\varepsilon }_L\) for all \(k\) (part (a)). Then for \(1\le k\le k_{\max }(L)=\lfloor \sqrt{2c_\alpha /(\pi ^2{\varepsilon }_L)}\rfloor -1\), equivalently \({\varepsilon }_L\le \frac{2c_\alpha }{\pi ^2(k+1)^2}\), \(\frac{2c_\alpha }{\pi ^2(k+1)^2}\le \mu _k(K^{(L)})\le \frac{6c_\alpha }{\pi ^2k^2}\) ??.
Let \(\nu _0\in {\mathcal P}(Z)\) satisfy (A4), \(K_\infty =K_{\varphi ',\nu '}\), \(\delta :=\| K_{\nu _0}-K_\infty \| _{L^2(P_X)\to L^2(P_X)}\), \(f=K_\infty g\), \(G:=\| g\| \), \(u^\dagger _\infty :=S^*g\); let \(\rho ^*\), \(\nu ^*\), \(m^*\), \(q\) be the quantities of the minimizer for \((\lambda ,\beta ,\nu _0)\) and \(Q_\lambda :=Q^{\nu _0}_\lambda (f)\). Then (Amplitude.)
(Lemma 771 with \(h=g\): \((K_{\nu _0}+\lambda )g-f=(K_{\nu _0}-K_\infty )g+\lambda g\) has norm at most \((\delta +\lambda )G\), and \(\frac1\lambda (\delta +\lambda )^2 =(\delta /\sqrt\lambda +\sqrt\lambda )^2\); Lemma 768 for \(\| \nu ^*-\nu _0\| _{\mathcal M}\) and \(q\); Lemma 772.)
(Realization, hidden marginal.) Under the hypotheses of Theorem 776, with \(H:=\langle f,g\rangle \) and \(\eta :=\| (K_{\nu _0}-K_\infty )g\| \le \delta G\),
where \(E(u^\dagger _\infty )=\frac{\eta ^2}\lambda +\langle K_{\nu _0}g,g\rangle \) (Lemma 767 with \(u=u^\dagger _\infty \) and Theorem 775).
(Joint schedule.) Let \((\nu _L)_L\) be reference measures with \(\delta _L:=\| K_{\nu _L}-K_\infty \| \to 0\) (for \(\nu _L=\tilde\nu _0^{(L)}\), \(\delta _L\le (C_\alpha +1)L^{-\min (1,2\alpha )}\)), \(f=K_\infty g\), \(u^\dagger _\infty =S^*g\), and let \((\lambda _L,\beta _L)\) satisfy \(\lambda _L\to 0\), \(\kappa _L=\lambda _L/\beta _L\to 0\) and \(\delta _L^2/\lambda _L\to 0\). Then, for the minimizers \(\rho ^*_L\) of the free energies for \((\lambda _L,\beta _L,\nu _L)\), \(m^*_L\to u^\dagger _\infty \) in \(L^2(\nu ^*_L)\) and in \(L^2(\nu _L)\), \(\| \Pi \rho ^*_L-u^\dagger _\infty \nu _L\| _{\mathcal M}\to 0\), \(F_{\rho ^*_L}\to f\) in \(L^2(P_X)\) and \(\operatorname {KL}(\nu ^*_L\| \nu _L)+\operatorname {KL}(\nu _L\| \nu ^*_L)\to 0\). (Insert the schedule into Theorems 776, 777 and 778: \(Q_{\lambda _L}\) and \(E(u^\dagger _\infty )\) stay bounded by Theorem 775, and every right-hand side tends to zero.)
(Coefficient measure.) Under the hypotheses of Theorem 776, \(\| \Pi \rho ^*-u^\dagger _\infty \nu _0\| _{\mathcal M}\le \| m^*-u^\dagger _\infty \| _{L^2(\nu ^*)} +G\sqrt{\kappa E(u^\dagger _\infty )}\), with \(E(u^\dagger _\infty )=\frac{\eta ^2}\lambda +\langle K_{\nu _0}g,g\rangle \) (Lemma 774 with \(h=g\), \(|u^\dagger _\infty |\le G\), and \(\| \nu ^*-\nu _0\| _{\mathcal M}\le \sqrt{2\operatorname {KL}(\nu ^*\| \nu _0)}\le \sqrt{\kappa E(u^\dagger _\infty )}\)).
Let \(\nu _0\in {\mathcal P}(Z)\), \(K_\infty =K_{\varphi ',\nu '}\) a second kernel operator on \(L^2(P_X)\), \(\delta :=\| K_{\nu _0}-K_\infty \| \), \(f=K_\infty g\), \(G:=\| g\| \), \(H:=\langle f,g\rangle \) and \(u^\dagger _\infty :=S^*g\). Then \(S_{\nu _0}u^\dagger _\infty -f=(K_{\nu _0}-K_\infty )g\), \(\| u^\dagger _\infty \| ^2_{L^2(\nu _0)}=\langle K_{\nu _0}g,g\rangle \) and
?? (the minimum ?? and \(\langle K_{\nu _0}g,g\rangle -H=\langle (K_{\nu _0}-K_\infty )g,g\rangle \le \eta G\)).
Under the hypotheses of Theorem 776, for the choice \(\lambda =\delta \) one has \(E(u^\dagger _\infty )\le H+2\delta G^2\) and
and if \(\kappa (H+2\delta G^2)\le 1\) the same bound times \(e^{1/4}\) holds in \(L^2(\nu _0)\) ??. (For \(\lambda =\delta \), \((\delta /\sqrt\lambda +\sqrt\lambda )^2=4\delta \), \(E(u^\dagger _\infty )\le H+2\delta G^2\), \(\sqrt{a+b}\le \sqrt a+\sqrt b\) and \(e^{\kappa Q_\lambda /2}\le e^{1/2}\).)
Assume the setting of Theorem 328, (SC\(_a\)) with \(0{\lt}a\le 1\) and the complexity bound. With \(\lambda _N:=N^{-1/(4(a+2))}\) and \(\beta _N:=\lambda _N^{-a}\), for every \(N\ge 1\), with probability at least \(1-\delta \),
where \(C_2\) is Definition 147 with \(\beta _0=\kappa _0=1\). (\(\beta _N\ge 1\) and \(\kappa _N=\lambda _N^{1+a}\le 1\), so Corollary 476 applies with \(\beta _0=\kappa _0=1\) and \(2C_2(\beta _N^{-1}+\kappa _N)\le 4C_2\lambda _N^a\); moreover \(B+{Y_{\max }}\le (2{Y_{\max }}+1)/\lambda \) for \(\lambda \le 1\), \(a\le 1\).)
Assume the setting of Theorem 328, (SC\(_a\)) with \(0{\lt}a\le 1\), the complexity bound \({\mathfrak R}_N(\Phi )\le C_{\mathrm P}\sqrt{(m+1)/N}\) and (S\(_\infty \)) for the family \(\lambda \le 1\), \(\beta =\beta _0\); let \(C_1\) be Definition 145 with \(\kappa _0=1/\beta _0\). With \(\lambda _N:=N^{-1/(4(a+2))}\), for every \(N\ge 1\), with probability at least \(1-\delta \),
(??: the first term is Corollary 341 and Lemma 488 with \(B+{Y_{\max }}\le (2{Y_{\max }}+\sqrt{2\beta _0/\pi })/\lambda \), the second is Corollary 473; at \(\lambda =\lambda _N\), \(\lambda _N^{-2}N^{-1/4}=\lambda _N^a\) and \(\lambda _N\le \lambda _N^a\).)
Let \(K^{(\infty )}(x,x'):=1-c_\alpha |x-x'|\). For \(0{\lt}\alpha {\lt}1\), \(L\ge 2\) and \(x,x'\in [-1,1]\), \(|K^{(L)}(x,x')-K^{(\infty )}(x,x')|\le C_\alpha L^{-\min (1,2\alpha )}\) with \(C_\alpha =\frac1{1-2^{-2\alpha }}\bigl(2+\frac{2\alpha }{1+\alpha } +\frac{4\alpha }{3(2+\alpha )}+2^{1-\alpha }\alpha \bigr)+1\) ??: the sum of the bounds on \(T_1,\dots ,T_5\) with \(1-L^{-2\alpha }\ge 1-2^{-2\alpha }\), \(|d|\le 2\) and \(2e^{-L}\le 1\); on the diagonal \(K^{(L)}(x,x)-1=-\frac1L+T_5\).
For \(0{\lt}\alpha {\lt}1\) and \(L\ge 2\), \({\varepsilon }_L:=\| K^{(L)}-K^{(\infty )}\| _{L^2(P_X)\to L^2(P_X)} \le \sup _{x,x'\in [-1,1]}|K^{(L)}(x,x')-K^{(\infty )}(x,x')|\le C_\alpha L^{-\min (1,2\alpha )}\) ??, where \(K^{(L)}=S_{\nu _0^{(L)}}S^*_{\nu _0^{(L)}}\) is the kernel operator of the tanh feature and \(K^{(\infty )}=S_\varpi S_\varpi ^*\) that of the sign network (Remark 655). (Theorem 676, Remark 655 and Lemma 710, since \(P_X\)-a.e. \(x\) lies in \([-1,1]\).)
For \(0{\lt}\alpha \), \(L{\gt}1\) and \(|d|\le 2\), \(0\le T_4\le \frac{4\alpha L^{-3-2\alpha }}{3(2+\alpha )(1-L^{-2\alpha })}\) (from \(I_\alpha (S)\le \frac{S^{2+\alpha }}{3(2+\alpha )}\)); this is sharper than the bound \(\frac{4\alpha L^{-2-2\alpha }}{3(2+\alpha )(1-L^{-2\alpha })}\) stated in Theorem 676.
For \(\lambda {\gt}0\), \(f\in L^2(P_X)\) and two hidden laws \(\nu ,\nu '\in {\mathcal P}(Z)\), \(R_{\lambda ,\nu }f=S^*(K_\nu +\lambda )^{-1}f\) and \(R_{\lambda ,\nu '}f=S^*(K_{\nu '}+\lambda )^{-1}f\) are bounded functions on \(Z\) with \(\sup _{z\in Z}|R_{\lambda ,\nu }f(z)-R_{\lambda ,\nu '}f(z)|\le \lambda ^{-2}\| K_\nu -K_{\nu '}\| \, \| f\| \). With \(\nu =\nu _0^{(L)}\), \(\nu '\) the limit law and \(\| K^{(L)}-K^{(\infty )}\| \le {\varepsilon }_L\) this is ??. (\(\| S^*r\| _\infty \le \| r\| _{L^2(P_X)}\), Lemma ??(2), and Theorem 740.)
If \(f=K^{(\infty )}g\) with \(g\in L^2(P_X)\), then \(R^{(\infty )}_\lambda f\to u^\dagger _\infty =S^*g\) uniformly on \(Z\) as \(\lambda \downarrow 0\). (\(K^{(\infty )}\) is self-adjoint and injective, so \(\overline{\operatorname {ran}K^{(\infty )}}=L^2(P_X)\) and \((K^{(\infty )}+\lambda )^{-1}f=g-\lambda (K^{(\infty )}+\lambda )^{-1}g\to g\) in \(L^2\) by Lemma 723; the bound \(\| S^*r\| _\infty \le \| r\| \) of Lemma ??(2) gives uniform convergence.)
If \(g=-\frac{1+\alpha }\alpha f''\) \(P_X\)-a.e. (i.e. \(f=K^{(\infty )}g\), ??), then \(u^\dagger _\infty (w,b)=-\frac{1+\alpha }{2\alpha }\int _{-1}^1\tanh (wx-b)f''(x)\, \, \mathrm dx\) ??: the limiting canonical ridgelet transform is the ridgelet transform, with the activation itself as filter, of the target sharpened by \(-\Delta \).
If \(f\in H^2(-1,1)\) satisfies the boundary conditions ?? and \(g:=-\frac{1+\alpha }\alpha f''\), then \(f=K^{(\infty )}g\), \(R^{(\infty )}_\lambda f\to u^\dagger _\infty \) uniformly on \(Z\) as \(\lambda \downarrow 0\), and \(u^\dagger _\infty (w,b)=-\frac{1+\alpha }{2\alpha }\int _{-1}^1\tanh (wx-b)f''(x)\, \mathrm dx\) (Theorems 709, 726, 727).
For \(f=K^{(\infty )}g\) and every \(t\in {\mathbb R}\), \(\lim _{w\to \pm \infty }u^\dagger _\infty (w,wt)=\pm (S_\varpi ^*g)(t)\); that is, \(u^\dagger _\infty (w,wt)\to \operatorname {sgn}(w)\, u^\dagger _\varpi (t)\), where \(S_\varpi ^*g=u^\dagger _\varpi \) is the representer ?? of the sign network (extended by the same formula to \(|t|{\gt}\ell _\alpha \); the explicit values are Theorem 735). (Lemma 733.)
For \(f\in H^2(-1,1)\) with ?? and \(g=-\frac{1+\alpha }\alpha f''\): for \(|t|{\lt}1\), \(\lim _{w\to \pm \infty }u^\dagger _\infty (w,wt)=\pm \frac{1+\alpha }\alpha f'(t)\), and for \(|t|{\gt}1\), \(\lim _{w\to +\infty }u^\dagger _\infty (w,wt)=-\operatorname {sgn}(t)(1+\alpha )\frac{f(1)+f(-1)}2\) (Theorems 734, 735).
For \(f\in H^2(-1,1)\) with the boundary conditions ?? and \(g=-\frac{1+\alpha }\alpha f''\): \((S_\varpi ^*g)(t)=\frac{1+\alpha }\alpha f'(t)\) for \(|t|{\lt}1\), and \((S_\varpi ^*g)(t)=\mp (1+\alpha )\frac{f(1)+f(-1)}2\) for \(t\gtrless \pm 1\); hence \(\lim _{w\to \pm \infty }u^\dagger _\infty (w,wt)=\pm \frac{1+\alpha }\alpha f'(t)\) for \(|t|{\lt}1\) and \(\mp \operatorname {sgn}(t)(1+\alpha )\frac{f(1)+f(-1)}2\) for \(|t|{\gt}1\). (\(\int _{-1}^1\operatorname {sgn}(x-t)f''\, \mathrm dx=f'(1)+f'(-1)-2f'(t)=-2f'(t)\) for \(|t|{\lt}1\), using \(f'(1)+f'(-1)=0\), and \(\int \operatorname {sgn}(x-t)f''=-\operatorname {sgn}(t)(f'(1)-f'(-1)) =\operatorname {sgn}(t)\alpha (f(1)+f(-1))\) for \(|t|{\gt}1\).)
If \(f=K^{(\infty )}g\) with \(g\in L^2(P_X)\), then for \(L\ge 2\), \(\| S_{\nu _0^{(L)}}u^\dagger _\infty -f\| _{L^2(P_X)}\le {\varepsilon }_L\| g\| \) \(\le C_\alpha L^{-\min (1,2\alpha )}\| g\| \) ??, since \(S_{\nu _0^{(L)}}S^*g=K^{(L)}g\) and \(\| K^{(L)}g-K^{(\infty )}g\| \le {\varepsilon }_L\| g\| \) (Theorem 711).
If \(f=K^{(\infty )}g\) with \(f\in H^2(-1,1)\) satisfying ?? and \(g=-\frac{1+\alpha }\alpha f''\), then \(\langle f,g\rangle =\langle K^{(\infty )}g,g\rangle =\| S_\varpi ^*g\| ^2 =\| S_\varpi ^\dagger f\| ^2=\| f\| ^2_{{\mathcal{H}}_\infty }\) (Theorem 705; \(S_\varpi ^*g\) is the minimum-norm representer of \(f=S_\varpi S_\varpi ^*g\)).
(Theorem ??(ii).) If \(\mu {\gt}0\) and \(g\in C[-1,1]\) satisfy \(K^{(\infty )}g=\mu g\) on \([-1,1]\), then \(g\) is twice differentiable on \((-1,1)\) and \(\mu \, g''=-c_\alpha \, g\) there ??. (The paper states \(g\in C^\infty [-1,1]\); the Lean statement records \(C^2\) on the open interval, which is what the exhaustion uses.)
(Theorem ??(iii), exhaustion.) If \(\mu {\gt}0\) and \(g\in C[-1,1]\), \(g\not\equiv 0\), satisfy \(K^{(\infty )}g=\mu g\) on \([-1,1]\), then \(\mu =c_\alpha /\varkappa ^2\) for some \(\varkappa {\gt}0\) and either \(\cos \varkappa =0\) (\(\varkappa \in {\mathcal{K}}^o\)) and \(g=A\sin \varkappa x\), or \(\varkappa \tan \varkappa =\alpha \) (\(\varkappa \in {\mathcal{K}}^e\)) and \(g=B\cos \varkappa x\), with \(A,B\ne 0\). Hence the eigenvalues and eigenfunctions are exhausted by Theorem ??(iii) and all eigenvalues are simple.
(Theorem ??(v).) \(\sum _k\mu _k=\operatorname {Tr}K^{(\infty )}=1\): the odd part is \(\sum _j\frac1{(j+1/2)^2\pi ^2}=\frac12\), the even part \(\sum _j(\varkappa ^e_j)^{-2}=\frac1\alpha +\frac12\) by the Hadamard factorization of \(\varkappa \sin \varkappa -\alpha \cos \varkappa \).
With the notation of Theorem 225, \(\nu ^*(\, \mathrm dz)=Z_*^{-1}\exp \bigl(\frac{s^*(z)^2}{2\lambda \beta }\bigr)\nu _0(\, \mathrm dz)\) with \(Z_*=\int _Z\exp \bigl(\frac{s^*(z)^2}{2\lambda \beta }\bigr)\nu _0(\, \mathrm dz){\lt}\infty \); with \(\nu _0=Z_0^{-1}e^{-V/\beta }\, \mathrm dz\) this is \(\nu ^*(\, \mathrm dz)\propto \exp \bigl(\frac{s^*(z)^2}{2\lambda \beta }-\frac{V(z)}\beta \bigr)\, \mathrm dz\).
With \(K^*:=K_{\nu ^*}\),
(Proof: \(F_{\rho ^*}=S_{\nu ^*}m_{\rho ^*}=-\lambda ^{-1}K^*r^*\) by Lemma 221 and Theorem 230, so \((K^*+\lambda )r^*=-\lambda f\), and \(K^*+\lambda \) is invertible.)
\(\int |a|\, \mathrm d\rho ^*\le \frac{\| f\| }\lambda +\sqrt{\frac{2\beta }{\pi \lambda }}\) and \(\int a^2\, \mathrm d\rho ^*=\| m_{\rho ^*}\| _{L^2(\nu ^*)}^2+\frac\beta \lambda \le \min \bigl\{ \frac{\| f\| ^2}{\lambda ^2},\frac{\| f\| ^2}{4\lambda }\bigr\} +\frac\beta \lambda \).
The self-consistency equation \(s=-\lambda \, S^*(K_{\nu [s]}+\lambda )^{-1}f\), \(\nu [s](\, \mathrm dz)\propto \exp \bigl(\frac{s(z)^2}{2\lambda \beta }\bigr)\nu _0(\, \mathrm dz)\), has exactly one bounded measurable solution, namely \(s^*\): \(s^*\) solves it by Theorems 226 and 231, and any bounded measurable solution \(s\) satisfies \(s=s^*\). (Proof: \(\rho _s:=\hat\mu _{W_s}\) with \(W_s=a\, s(z)\) has \(\nu _{\rho _s}=\nu [s]\), \(F_{\rho _s}=K_{\nu [s]}(K_{\nu [s]}+\lambda )^{-1}f\) and \(S^*r_{\rho _s}=s\), so \(\rho _s\) solves the self-consistency equation of Lemma 210 and is the minimizer.)
In the setting of 240, with \(\rho _M^\circ =\frac1M\sum _\ell \delta _{\theta _\ell ^\circ }\) and \(\nu _M^\circ =\frac1M\sum _\ell \delta _{z_\ell ^\circ }\), \(\Pi \rho _M^\circ =\frac1M\sum _\ell a_\ell ^\circ \delta _{z_\ell ^\circ } =-\lambda ^{-1}s^\circ \, \nu _M^\circ \); in Lean: for every measurable \(A\subset Z\), \(\Pi \rho _M^\circ (A)=-\lambda ^{-1}\frac1M\sum _{\ell :z_\ell ^\circ \in A} s^\circ (z_\ell ^\circ )\).
Assume (A3), (A6) and (A7): \(z\mapsto \varphi _z(x)\) and \(V\) are real analytic, \(V\ge 0\) is coercive and \(Z\) is finite dimensional. For every \(\theta _0\in \Theta ^M\) the gradient flow ?? has a unique solution on \([0,\infty )\), which is bounded, has finite length and converges to a stationary point \(\theta ^\circ \): \(\int _0^\infty \| \dot\theta _t\| \, \mathrm dt{\lt}\infty \), \(\lim _{t\to \infty }\theta _t=\theta ^\circ \), \(\nabla J_N(\theta ^\circ )=0\).
In the setting of 244, for every \(x\in {\mathcal X}\), \(F_{\theta ^\circ }(x)=\frac1N\sum _{i=1}^Nk_{\nu _M^\circ }(x,x_i)\, \alpha _i^\circ \) with \(\alpha ^\circ :=(K^\circ +\lambda )^{-1}y\) and \(k_{\nu _M^\circ }(x,x'):=\frac1M\sum _\ell \varphi _{z_\ell ^\circ }(x) \varphi _{z_\ell ^\circ }(x')\).
Assume the setting of Section ?? (\(|Y|\le {Y_{\max }}\), i.i.d. sample, \(\lambda ,\beta {\gt}0\), \(Z\) separable and first countable, \(\varphi \) continuous in \(z\)) and Lemma ?? for \(L\) and \(L_N\). Let \(B\) be as in 82 and \(\Delta _N(\delta )\) as in 83. For every \(\delta \in (0,1)\), with probability at least \(1-\delta \) over the sample, the following hold simultaneously:
Assume the setting of Theorem 328. Put \(c_1:=\max \{ 1,1/\lambda ,M_*/\sqrt\beta \} \) and \(N_B(\delta ):=\lceil 64c_1^4D''(\delta )^2\rceil \) (\(\Leftrightarrow c_1^2\eta _N''\le \frac18\)). Let \(\delta \in (0,1)\) and \(N\ge N_B(\delta )\). On the event \(\Omega _\delta \) the following hold simultaneously.
\(\| h\| _N+\sqrt{\beta {\mathcal{K}}_N}\le 9c_1\eta _N'\), \(\| h\| _{L^2(P_X)}\le 6c_1\eta _N'\), \(D_\varsigma \le 10c_1^2\eta _N'\).
\({\mathcal{K}}_N\le \frac{81c_1^2}\beta \eta _N'{}^2\) and \(\| \rho _N^*-\rho ^*\| _{\mathrm{TV}}\le \frac{9c_1}{2\sqrt\beta }\eta _N'\).
(In Lean the statement is deterministic: for any upper bounds \(\eta '\ge G_N\), \(\eta ''\ge H_N\) with \(c_1^2\eta ''\le \frac18\).)
Assume the setting of Theorem 370 and \(\sup _z|m^*(z)|\le M_1\); put \(c_1(M_1):=\max \{ 1,1/\lambda ,M_1/\sqrt\beta \} \). For any upper bounds \(\eta '\ge G_N\), \(\eta ''\ge H_N\) with \(c_1(M_1)^2\eta ''\le \frac18\): \(\| h\| _N+\sqrt{\beta {\mathcal{K}}_N}\le 9c_1(M_1)\eta '\), \(\| h\| \le 6c_1(M_1)\eta '\), \(D_\varsigma \le 10c_1(M_1)^2\eta '\), \({\mathcal{K}}_N\le \frac{81c_1(M_1)^2}\beta \eta '{}^2\) and \(\| \rho _N^*-\rho ^*\| _{\mathrm{TV}}\le \frac{9c_1(M_1)}{2\sqrt\beta }\eta '\) (the proof of Theorem 370 with Lemma 366).
On \(\Omega _\delta \), for \(N\ge N_B(\delta )\) and under Lemmas 314 and 315: with \(D'(\delta )=2C_{\mathrm P}\sqrt{m+1}+\sqrt{2\log (4/\delta )}\), \({\mathcal{K}}_N\le \frac{121c_1^2E_*^2D'(\delta )^2}{\beta N}=O(1/N)\), \(\| h\| _N,\| h\| \le \frac{11c_1E_*D'(\delta )}{\sqrt N}\), \(\| \rho _N^*-\rho ^*\| _{\mathrm{TV}}\le \frac{11c_1E_*D'(\delta )}{2\sqrt{\beta N}}\), \(\sup _z|m_N-m^*|\le \frac{12c_1E_*D'(\delta )}{\lambda \sqrt N}\) and \(\sup _x|F_{\rho _N^*}-F_{\rho ^*}|\le \frac{12c_1^2E_*D'(\delta )}{\sqrt N}\), all of order \(N^{-1/2}\) with no logarithm.
Assume the setting of Theorem 369, the complexity bounds \({\mathfrak R}_N(\Phi )\le C_{\mathrm P}\sqrt{(m+1)/N}\), \({\mathfrak R}_N(\Psi )\le C'_{\mathrm P}\sqrt{(m+1)/N}\) and \(N\ge 64c_1(M_1)^4D''(\delta )^2\). On the event \(\Omega _\delta \) (with \(E:=M_1+{Y_{\max }}\) in \(\eta '_N\)), with \(D'(\delta )=2C_{\mathrm P}\sqrt{m+1}+\sqrt{2\log (4/\delta )}\): \({\mathcal{K}}_N\le \frac{121c_1(M_1)^2E^2D'(\delta )^2}{\beta N}\), \(\| h\| _N,\| h\| \le \frac{11c_1(M_1)ED'(\delta )}{\sqrt N}\), \(\| \rho _N^*-\rho ^*\| _{\mathrm{TV}}\le \frac{11c_1(M_1)ED'(\delta )}{2\sqrt{\beta N}}\), \(\sup _z|m_N-m^*|\le \frac{12c_1(M_1)ED'(\delta )}{\lambda \sqrt N}\) and \(\sup _x|F_{\rho _N^*}-F_{\rho ^*}|\le \frac{12c_1(M_1)^2ED'(\delta )}{\sqrt N}\).
Chain rule.
and consequently, under (R),
(By Theorem 225, \(\rho ^*=\rho _{\nu ^*,m^*}\); \(\nu ^*\ll \nu _0\) with \(\operatorname {KL}(\nu ^*\| \nu _0){\lt}\infty \) and \(m^*\in L^2(\nu ^*)\), so Lemma 391 applies; then substitute into Theorem 397.)
Energy comparison. Assume (A1), (A3), (A4), (A5), \(f\in L^2(P_X)\) and (R). Let \(u_0\) be any representer of \(f\) and \(u_0^\dagger \) the minimum-norm representer. Then
(Minimality of \(\rho ^*\) against the competitor \(\tilde\rho _{u_0}\), Lemma 393.)
Under the hypotheses of Theorem 397,
(The first two from \(L(\rho ^*)=\frac12\| r^*\| ^2\), \(\operatorname {KL}\ge 0\) and \(L\ge 0\); the third from the chain rule of Theorem 396 and \(\| m^*\| ^2\ge 0\).)
Normalizing constant. With \(Z_*\), \(Z_0\) as in Theorem 226 and \(\tilde Z:=Z_*/Z_0=\int e^{\kappa m^{*2}/2}\, \mathrm d\nu _0\),
(By Theorem 226 and \(s^*=-\lambda m^*\), \(\frac{\, \mathrm d\nu ^*}{\, \mathrm d\nu _0}=\frac{Z_0}{Z_*}\exp (\frac{s^{*2}}{2\lambda \beta }) =\frac{Z_0}{Z_*}\exp (\frac{\lambda m^{*2}}{2\beta })\); integrating the logarithm against \(\nu ^*\) gives the formula for \(\operatorname {KL}(\nu ^*\| \nu _0)\); \(\tilde Z\ge 1\) since the integrand is \(\ge 1\); and \(\operatorname {KL}\ge 0\) with Theorem 396 give \(\log \tilde Z\le \frac\kappa 2\| m^*\| ^2_{L^2(\nu ^*)}\le \frac\kappa 2\| u_0^\dagger \| ^2\).)
Pythagorean identity. For every representer \(u_0\),
(Apply the strong-convexity identity of Lemma 212 to the competitor \(\tilde\rho =\tilde\rho _{u_0}\): \({\mathcal F}(\tilde\rho )-{\mathcal F}(\rho ^*) =L(\rho ^*)+\beta \operatorname {KL}(\tilde\rho \| \rho ^*)\), so \(\frac\lambda 2\| u_0\| ^2=2L(\rho ^*)+\beta \operatorname {KL}(\rho ^*\| \mu _U)+\beta \operatorname {KL}(\tilde\rho \| \rho ^*)\); then Theorem 396 and Lemma 390 with \(\nu =\nu _0\), \(u=u_0\), \(\nu '=\nu ^*\), \(u'=m^*\) give \(\operatorname {KL}(\tilde\rho \| \rho ^*)=\operatorname {KL}(\nu _0\| \nu ^*)+\frac\kappa 2\| u_0-m^*\| ^2_{L^2(\nu _0)}\); multiply by \(2/\lambda \).)
Since every term on the right-hand side of the Pythagorean identity is nonnegative, with \(u_0=u_0^\dagger \),
Assume (A1), (A3), (A4), (A5) and (SC\(_a\)) (hence (R)). Fix \((\lambda ,\beta )\) and let \(\Xi =\| (w-1)m^*\| _{L^2(\nu _0)}\). Then
(Write \(m^*-u^\dagger =(m^*-v)+(v-u_\lambda )+(u_\lambda -u^\dagger )\): the first term is \((1-w)m^*\) of norm \(\Xi \), the second is at most \(\Xi \) by Lemma 461, the third is Lemma 456.)
With \(\Pi \rho ^*=m^*\nu ^*=v\nu _0\),
(\(\| v-u^\dagger \| _{L^1}\le \| v-m^*\| _{L^1}+\| m^*-u^\dagger \| _{L^1} \le \| v-m^*\| _{L^2}+\| m^*-u^\dagger \| _{L^2}\), \(\nu _0\) being a probability measure.)
Assume (A1), (A3), (A4), (A5) and (SC\(_a\)) with \(a{\gt}0\) (hence (R)). Fix \((\lambda ,\beta )\) and let \(\Xi =\| (w-1)m^*\| _{L^2(\nu _0)}\). Then \(\| m^*-u^\dagger \| _{L^2(\nu _0)}\le 2\, \Xi +\lambda ^{\min (a,1)}\| g_0\| \) (Theorem 468 with Lemma 484 in place of Lemma 456).
\(\| F_{\rho ^*}-f\| _{L^2(P_X)}\le \frac{\sqrt\lambda }2\, \Xi +\lambda ^{\min (a+1/2,1)}\| g_0\| \). (\(F_{\rho ^*}-f=S_{\nu _0}(v-u^\dagger )=S_{\nu _0}(v-u_\lambda )+S_{\nu _0}(u_\lambda -u^\dagger )\); \(\| S_{\nu _0}(v-u_\lambda )\| ^2=\langle T_0e,e\rangle \) with \((T_0+\lambda )e=\lambda (w-1)m^*\), and \(4\lambda \langle T_0e,e\rangle \le \| (T_0+\lambda )e\| ^2=\lambda ^2\Xi ^2\).)
\(m_n^*\to u_0^\dagger \) in \(L^2(\nu _0)\), i.e. \(\| m_n^*-u_0^\dagger \| _{L^2(\nu _0)}\to 0\). (In Lean, without compactness: \(\| m_n^*-u_0^\dagger \| ^2=\| m_n^*\| ^2-2\langle m_n^*,u_0^\dagger \rangle +\| u_0^\dagger \| ^2 \le \tilde Z_n\| u_0^\dagger \| ^2-2\langle m_n^*,u_0^\dagger \rangle +\| u_0^\dagger \| ^2\to 0\) by Lemmas 415 and 418.)
\(\| m_n^*\| _{L^2(\nu _0)}\to \| u_0^\dagger \| \), \(\| m_n^*\| _{L^2(\nu _n^*)}\to \| u_0^\dagger \| \) and \(\| S_{\nu _0}m_n^*-f\| _{L^2(P_X)}\to 0\). (The first from (a); for the second, \(\tilde Z_n^{-1}\| m_n^*\| ^2_{L^2(\nu _0)} \le \| m_n^*\| ^2_{L^2(\nu _n^*)}\le \| u_0^\dagger \| ^2\) with \(\tilde Z_n\to 1\); the third is ??.)
With \(h_n:=m_n^*w_n=\frac{\, \mathrm d\gamma _n}{\, \mathrm d\nu _0}\in L^1(\nu _0)\), \(\| h_n-u_0^\dagger \| _{L^1(\nu _0)}=\| \gamma _n-\gamma _\infty \| _{\mathcal M}\to 0\): \(\| h_n-u_0^\dagger \| _{L^1(\nu _0)}\le \int |m_n^*||w_n-1|\, \mathrm d\nu _0 +\| m_n^*-u_0^\dagger \| _{L^2(\nu _0)}\to 0\) by Lemma 416 and (a).
\(\gamma _n=\Pi \rho _n^*\rightharpoonup \gamma _\infty =u_0^\dagger \nu _0\), and \(\| m_n^*\| _{L^2(\nu _n^*)}\to \| u_0^\dagger \| _{L^2(\nu _0)}\): the learned coefficient measure converges to the canonical (minimum-norm) ridgelet transform with respect to the reference measure \(\nu _0\). (In Lean this is deduced from Theorem 421: convergence in total mass implies weak convergence.)
Assume (A3), (A5), \(f\in L^2(P_X)\), let \((\rho _t)_{t\ge 0}\) be an MFLD flow (Hypothesis W) and assume that \(\hat\mu _{\rho _t}\) satisfies the KL form of \(\mathrm{LSI}(\alpha _*)\) for every \(t\ge 0\). Then for all \(t\ge 0\)
Assume (A1), (A3), (A5), (A8), \(f\in L^2(P_X)\) and (A9): \(\rho _0\in {\mathcal D}\), \(\int |z|^2\, \mathrm d\rho _0{\lt}\infty \) and \(\int e^{c_0|a|}\, \mathrm d\rho _0{\lt}\infty \) for some \(c_0{\gt}0\). Then the MFLD has a solution on \([0,\infty )\) (unique in law), and its flow of laws \((\rho _t)_{t\ge 0}\) satisfies (W1) and (W2), i.e. it is an MFLD flow with \(\rho _0\) as initial law.
Let \(K_\nu =S_\nu S_\nu ^*\) and \(K_{\nu '}=S_{\nu '}S_{\nu '}^*\) be the kernel operators of two (feature map, hidden law) pairs on the same \(L^2(P_X)\), \(\lambda {\gt}0\) and \(f\in L^2(P_X)\). Then \(|Q^\nu _\lambda (f)-Q^{\nu '}_\lambda (f)|\le \lambda ^{-2}\| K_\nu -K_{\nu '}\| \, \| f\| ^2\). In the setting of Theorem ??, \(K_\nu =K^{(L)}\), \(K_{\nu '}=K^{(\infty )}\) and \(\| K^{(L)}-K^{(\infty )}\| \le {\varepsilon }_L\), this is \(|Q^{(L)}_\lambda (f)-Q^{(\infty )}_\lambda (f)|\le \lambda ^{-2}{\varepsilon }_L\| f\| ^2\). (Resolvent identity \((A+\lambda )^{-1}-(B+\lambda )^{-1}=(A+\lambda )^{-1}(B-A)(B+\lambda )^{-1}\) and \(\| (K+\lambda )^{-1}\| \le \lambda ^{-1}\).)
For every \(\nu \in {\mathcal P}(Z)\) and \(f\in L^2(P_X)\), \(\lambda \mapsto Q^\nu _\lambda (f)\) is nonincreasing on \((0,\infty )\), and as \(\lambda \downarrow 0\), \(Q^\nu _\lambda (f)\to \| S_\nu ^\dagger f\| ^2_{L^2(\nu )}\) if \(f\in \operatorname {ran}S_\nu \), while \(Q^\nu _\lambda (f)\to +\infty \) otherwise. For the limit kernel \(K^{(\infty )}\) of Theorem ??, \(\| S_\nu ^\dagger f\| ^2=\| f\| ^2_{{\mathcal{H}}_\infty }\) and \(\operatorname {ran}S_\nu =H^1(-1,1)\) (Theorem ??).
For \(0{\lt}\alpha {\lt}1\) and \(f\in L^2(P_X)\),
in the sense that \(\lambda \mapsto \lim _{L\to \infty }Q^{(L)}_\lambda (f)=Q^{(\infty )}_\lambda (f)\) (nonincreasing in \(\lambda \)) is bounded on \((0,\infty )\) if and only if \(f\) has an \(H^1(-1,1)\) representative. This is the precise meaning of “in the limit \(L\to \infty \), (R) is equivalent to \(f\in H^1\)”. (Theorems 713, 714, 715.)
For \(0{\lt}\alpha {\lt}1\), \(\lambda {\gt}0\) and \(f\in L^2(P_X)\), \(Q^{(L)}_\lambda (f)\to Q^{(\infty )}_\lambda (f)\) as \(L\to \infty \), where \(Q^{(L)}_\lambda (f)=\langle f,(K^{(L)}+\lambda )^{-1}f\rangle \) and \(Q^{(\infty )}_\lambda (f)=\langle f,(K^{(\infty )}+\lambda )^{-1}f\rangle \) (Theorem 740 with \({\varepsilon }_L\to 0\), Theorem 712).
Assume (R). Let \(g\) be bounded measurable (e.g. \(g\in C_b(Z)\)) and \({\varepsilon }{\gt}0\). There are \(\lambda _0=\lambda _0({\varepsilon },g){\gt}0\) and \(\kappa _0=\kappa _0({\varepsilon },g){\gt}0\) such that every \((\lambda ,\beta )\) with \(\lambda \le \lambda _0\) and \(\kappa =\lambda /\beta \le \kappa _0\) satisfies \(\bigl|\int g\, \mathrm d\Pi \rho ^*_{\lambda ,\beta }-\int g\, u_0^\dagger \, \mathrm d\nu _0\bigr|\le {\varepsilon }\) (the paper writes \({\varepsilon }/3\)). (If the claim were false, there would be minimizers \(\rho _k^*\) for \((\lambda _k,\beta _k)\) with \(\lambda _k,\kappa _k\le 1/(k+1)\) violating the bound; this contradicts \(\Pi \rho _k^*\rightharpoonup u_0^\dagger \nu _0\) (Theorem 422) along that schedule.)
In the setting of Theorem 328 with the complexity bound \({\mathfrak R}_N(\Phi )\le C\sqrt{(m+1)/N}\), for fixed \((\lambda ,\beta )\), a bounded measurable \(g\), \({\varepsilon }{\gt}0\) and \(\delta \in (0,1)\) there is \(N_1=N_1({\varepsilon },\delta ,g;\lambda ,\beta )\) such that for \(N\ge N_1\), with probability at least \(1-\delta \), \(\bigl|\int g\, \mathrm d\Pi \rho _N^*-\int g\, \mathrm d\Pi \rho ^*_{\lambda ,\beta }\bigr| \le 2\| g\| _\infty \| \Pi \rho _N^*-\Pi \rho ^*_{\lambda ,\beta }\| _{\mathrm{TV}}\le 2\| g\| _\infty r_N\le {\varepsilon }\) (the paper writes \({\varepsilon }/3\)). (Corollary 340 on the event \(E'_\delta \) of Lemma 339, the rate Corollary 341, and \(r_N\to 0\).)
Fix the sample \((x_i,y_i)_{i=1}^N\) with \(|y_i|\le {Y_{\max }}\), \((\lambda ,\beta )\) and \(M\ge 1\); let \(\rho _N^*\) be the minimizer of \({\mathcal F}_N\) and let \((\mu _t)_{t\ge 0}\) be ergodic to the empirical \(M\)-particle Gibbs measure \(\pi _M\) (Assumption (E)). For every bounded measurable \(g\) with \(\| g\| _\infty \le C\) and every \({\varepsilon }{\gt}0\) there is \(t_{\varepsilon }\) such that for \(t\ge t_{\varepsilon }\)
with the constant of Definition 135, which depends on the sample only through \({Y_{\max }}\). (Theorem 550 for the empirical feature, \(\| y\| _N\le {Y_{\max }}\), and \(\int g\, \mathrm d\Pi \rho _N^*=\int a\, g(z)\, \mathrm d\rho _N^*\).)
Assume the setting of Theorem 328 ((A1)–(A5), i.i.d. sample, \(|Y|\le {Y_{\max }}\), \(f={\mathbb E}[Y\mid X]\)), (R), the convention ?? and \({\mathfrak R}_N(\Phi )\le C\sqrt{(m+1)/N}\). Let \(g\) be bounded measurable, \({\varepsilon }{\gt}0\) and \(\delta \in (0,1)\). Then there are \(\lambda _0,\kappa _0{\gt}0\) such that for every \((\lambda ,\beta )\) with \(\lambda \le \lambda _0\), \(\lambda /\beta \le \kappa _0\) there is \(N_1=N_1({\varepsilon },\delta ,g;\lambda ,\beta )\) such that for \(N\ge N_1\), with probability at least \(1-\delta \),
(Stages (1) and (2) of Theorem ?? and the triangle inequality; the third stage, the reachability \(e^{(g)}_{M,t}\), is the subject of the dynamics.)
If \(f=K^{(\infty )}g_0\) with \(g_0\in L^2(P_X)\), then \(f\in H^2(-1,1)\) with \(f''=-c_\alpha g_0\), i.e. \(g_0=(K^{(\infty )})^{-1}f=-\frac{1+\alpha }\alpha f''\), and \(f'(1)=-\frac\alpha 2(f(1)+f(-1))\), \(f'(-1)=+\frac\alpha 2(f(1)+f(-1))\) ??. (\(f=S_\varpi v\) with \(v=S_\varpi ^*g_0\), so \(f'=v/\ell _\alpha \) by Theorem ??(ii) and \(v'=-g_0\); the boundary values come from \(v=\pm \frac12\int g_0\) outside \([-1,1]\).)
If \(f\in H^2(-1,1)\) with \(f'(1)=-\frac\alpha 2(f(1)+f(-1))\) and \(f'(-1)=+\frac\alpha 2(f(1)+f(-1))\), then \(f=K^{(\infty )}g_0\) for \(g_0:=-\frac{1+\alpha }\alpha f''\in L^2(P_X)\) ??. (\(S_\varpi ^*g_0=\ell _\alpha f'-\frac{\ell _\alpha }2(f'(1)+f'(-1))=\ell _\alpha f'\) on \((-1,1)\) and \(=\pm \frac{\ell _\alpha }2(f'(-1)-f'(1))=\pm (1+\alpha )c_0\) outside, i.e. \(S_\varpi ^*g_0=u_\varpi ^\dagger \), and \(S_\varpi u_\varpi ^\dagger =f\) by Theorem ??(ii).)
\(K^{(\infty )}=S_\varpi S_\varpi ^*\): for \(r\in L^2(P_X)\) and \(P_X\)-a.e. \(x\), \((S_\varpi S_\varpi ^*r)(x)=\int _{-1}^1K^{(\infty )}(x,x')r(x')\, P_X(\, \mathrm dx')\) with \(K^{(\infty )}(x,x')=1-c_\alpha |x-x'|\) (Remark 655 and the integral representation of \(S_\varpi S_\varpi ^*\)).
\(\operatorname {ran}S_\varpi =H^1(-1,1)\) as classes in \(L^2(P_X)\): \(F\in L^2(P_X)\) is of the form \(S_\varpi u\) with \(u\in L^2(\varpi )\) if and only if \(F\) has a representative \(f\in H^1(-1,1)\). (\(S_\varpi u\) is the primitive of \(u/\ell _\alpha \); conversely \(f=S_\varpi u_\varpi ^\dagger \) with the representer ??.)
For \(f\in H^1(-1,1)\) with \(f'=g\) and the representer \(u_\varpi ^\dagger \) of ??, \((S_\varpi u_\varpi ^\dagger )(x)=f(x)\) for every \(x\in [-1,1]\): with \(a=(1+\alpha )c_0\), \(S_\varpi u_\varpi ^\dagger (x)=\frac1{2\ell _\alpha }\bigl[2a(\ell _\alpha -1) +\ell _\alpha (2f(x)-f(1)-f(-1))\bigr]=f(x)\) (Theorem ??(ii), \(H^1\subseteq \operatorname {ran}S_\varpi \)).
For \(f\in H^1(-1,1)\) with \(c_0=\frac{f(1)+f(-1)}2\), \(\| S_\varpi ^\dagger f\| ^2_{L^2(\varpi )}=\| f\| ^2_{{\mathcal{H}}_\infty } =\frac{1+\alpha }\alpha \| f'\| ^2_{L^2(P_X)}+(1+\alpha )c_0^2\) ??. (\(\| u_\varpi ^\dagger \| ^2=\frac1{2\ell _\alpha }\bigl[\ell _\alpha ^2\int _{-1}^1f'{}^2\, \mathrm dx +\frac{2\ell _\alpha ^2c_0^2}{\ell _\alpha -1}\bigr]\) with \(\ell _\alpha =\frac{1+\alpha }\alpha \), \(\frac{\ell _\alpha }{\ell _\alpha -1}=1+\alpha \).)
\(K^{(\infty )}\) is injective on \(L^2(P_X)\): if \(K^{(\infty )}r=0\) then \(\| S_\varpi ^*r\| ^2=\langle K^{(\infty )}r,r\rangle =0\), so \(v:=S^*r\) vanishes \(\varpi \)-a.e.; \(v\) is continuous on \([-1,1]\) with \(v'=-r\) (Lemma 698), so \(v=0\) on \([-1,1]\), \(\int _{-1}^xr=0\) for all \(x\), and \(r=0\) a.e. by Lebesgue differentiation.
\(\ker S_\varpi =\{ u\in L^2(\varpi ):u=0\text{ a.e.\ on }(-1,1), \int _{-\ell _\alpha }^{-1}u=\int _1^{\ell _\alpha }u\} \). (\(S_\varpi u=0\) if and only if the \(H^1\) function \(h=S_\varpi u\) vanishes on \([-1,1]\), i.e. \(h'=u/\ell _\alpha =0\) a.e. on \((-1,1)\) and \(h(1)+h(-1)=\frac1{\ell _\alpha }[\int _{-\ell _\alpha }^{-1}u-\int _1^{\ell _\alpha }u]=0\); the first condition uses the Lebesgue differentiation theorem.)
\(K_{\nu _0}\) is the convolution operator \(g\mapsto \int \kappa (\theta -\theta ')g(\theta ')\, \tau (\, \mathrm d\theta ')\) on \(L^2(\tau )\), so \(\cos (\ell \theta )\) and \(\sin (\ell \theta )\) are eigenfunctions with eigenvalue \(\mu _\ell =\frac1{2\pi }\int _0^{2\pi }\tilde\kappa (\cos \phi )\cos (\ell \phi )\, \mathrm d\phi \), the \(\ell \)th Fourier (cosine, by evenness) coefficient of \(\kappa \): \(K_{\nu _0}=\sum _\ell \mu _\ell P_\ell \).
\(\mu _\ell =\sum _{n\ge \ell ,\ n\equiv \ell \ (2)}c_n2^{-n}\binom n{(n-\ell )/2}\) with \(c_n:=\sum _{k\ge n}a_k(s)^2\binom kn(\sigma _w^2/s^2)^n(\sigma _b^2/s^2)^{k-n}\): expand \(\tilde\kappa (t)=\sum _nc_nt^n\) (nonnegative coefficients, Tonelli) and insert \(\frac1{2\pi }\int _0^{2\pi }\cos ^n\phi \cos (\ell \phi )\, \mathrm d\phi =2^{-n}\binom n{(n-\ell )/2}\).
If \(\sigma _b{\gt}0\) and \(\sigma \) is not (a.e. equal to) a polynomial, then \(\mu _\ell {\gt}0\) for all \(\ell \), \(K_{\nu _0}\) is injective and \(\overline{\operatorname {ran}S_{\nu _0}}=L^2(\tau )\). (A non-polynomial \(\sigma \) has \(a_k(s)\ne 0\) for infinitely many \(k\), so every \(c_n{\gt}0\) and \(\mu _\ell \ge c_\ell 2^{-\ell }{\gt}0\).)
If \(\sigma _b=0\) and \(\sigma \) is odd, then \(\mu _\ell =0\) for even \(\ell \), and \(\operatorname {ran}S_{\nu _0}\) consists of odd functions. (\(\kappa (\phi +\pi )=-\kappa (\phi )\), so \(\mu _\ell =\int \kappa (\phi +\pi )\cos (\ell (\phi +\pi ))\, \tau (\, \mathrm d\phi )=-\mu _\ell \) for even \(\ell \).)
Assume \(\mu _\ell {\gt}0\) for all \(\ell \) (by (ii) this holds when \(\sigma _b{\gt}0\) and \(\sigma \) is not a polynomial). Then \(f\in \operatorname {ran}S_{\nu _0}\iff \sum _{\ell \ge 0}\mu _\ell ^{-1}\| P_\ell f\| ^2_{L^2(\tau )}{\lt}\infty \), where \(\| P_\ell f\| ^2=\langle f,e_\ell \rangle ^2+\langle f,e_{-\ell }\rangle ^2\) in the real trigonometric basis \((e_\ell )_{\ell \in \mathbb Z}\) of Definition 565, i.e. \(\sum _{\ell \in \mathbb Z}\mu _\ell ^{-1}\langle f,e_\ell \rangle ^2{\lt}\infty \). (Lemma 584 for the eigenbasis of Theorem 580.)
Under (R), \(u_0^\dagger =S_{\nu _0}^\dagger f=\sum _{\ell \in \mathbb Z}\mu _\ell ^{-1} \langle f,e_\ell \rangle S^*e_\ell =\sum _{\ell \ge 0}\mu _\ell ^{-1}S^*P_\ell f\), with \((S^*e_\ell )(w,b)=\frac1{2\pi }\int _0^{2\pi }\sigma (w_1\cos \theta +w_2\sin \theta -b) e_\ell (\theta )\, \mathrm d\theta \), the series converging in \(L^2(\nu _0)\) with mutually orthogonal terms, and \(\| u_0^\dagger \| ^2_{L^2(\nu _0)} =\sum _{\ell \in \mathbb Z}\mu _\ell ^{-1}\langle f,e_\ell \rangle ^2 =\sum _{\ell \ge 0}\mu _\ell ^{-1}\| P_\ell f\| ^2\).
Assume \(\mu _\ell {\gt}0\) for all \(\ell \). For \(a\in \{ 1/2,1\} \), \(u_0^\dagger \in \operatorname {ran}\bigl((S_{\nu _0}^*S_{\nu _0})^a\bigr)\iff \sum _{\ell \in \mathbb Z}\mu _\ell ^{-2a-1}\langle f,e_\ell \rangle ^2{\lt}\infty \) (\(=\sum _{\ell \ge 0}\mu _\ell ^{-2a-1}\| P_\ell f\| ^2\)), the source condition (SC\(_a\)) of Definition 139 for some source norm \(G\). (In the singular system of Lemma 584: \(a=1/2\) is \(f\in \operatorname {ran}K_{\nu _0}\), \(a=1\) is \(f\in K_{\nu _0}(\operatorname {ran}S_{\nu _0})\), both read off the coordinates \(\langle e_\ell ,K_{\nu _0}g\rangle =\mu _\ell \langle e_\ell ,g\rangle \).)
Assume \(\mu _\ell {\gt}0\) for all \(\ell \). For every \(a{\gt}0\), \(u_0^\dagger \in \operatorname {ran}\bigl((S_{\nu _0}^*S_{\nu _0})^a\bigr)\iff \sum _{\ell \in \mathbb Z}\mu _\ell ^{-2a-1}\langle f,e_\ell \rangle ^2{\lt}\infty \), where \(T_0^a\) is the spectral power of Definition 164. (In the singular system, \(T_0v_\ell =\mu _\ell v_\ell \) and \((T_0)^a\) maps \(v_\ell \mapsto \mu _\ell ^av_\ell \) on \((\ker S_{\nu _0})^\perp \) and vanishes on \(\ker S_{\nu _0}\).)
Under (R), Theorem ?? (\(\lambda _n\to 0\), \(\kappa _n\to 0\)) gives \(m_n^*\to u_0^\dagger \) in \(L^2(\nu _0)\) and \(\Pi \rho _n^*\to u_0^\dagger \nu _0\) in total variation: the learned conditional mean amplitude converges to the explicit weighted ridgelet transform \(u_0^\dagger =\sum _\ell \mu _\ell ^{-1}S^*P_\ell f\) of Theorem 587, the \(\tau \)-weighted ridgelet transform \(S^*\) with the activation as filter applied to the sharpened target \(K_{\nu _0}^{-1}f=\sum _\ell \mu _\ell ^{-1}P_\ell f\).
Let \(\rho ^*\) be the minimizer of \({\mathcal F}\) on \({\mathcal D}\) and \(M\ge 1\). Then
(Theorem 534 at \(\mu =\pi _M\) and at \(\mu =\rho ^{*\otimes M}\), the minimality of \(\pi _M\) for \({\mathcal F}^M\), and Lemma 535.)
Let \(\rho ^*\) be the minimizer of \({\mathcal F}\) on \({\mathcal D}\). For every \(\mu \in {\mathcal P}(\Theta ^M)\) with \(\operatorname {KL}(\mu \| \mu _U^{\otimes M}){\lt}\infty \), \(\operatorname {KL}(\mu \| \rho ^{*\otimes M}){\lt}\infty \) and
(Chain rule for \(\rho ^{*\otimes M}=\widehat{(\mu _U^{\otimes M})}_{W^*_M}\), the identity \({\mathcal F}(\rho ^*)=L(\rho ^*)-\int as^*\, \mathrm d\rho ^*-\beta \log Z_*\) and the expansion \(L(\rho _\theta )-L(\rho ^*)=\int as^*\, \mathrm d(\rho _\theta -\rho ^*) +\frac12\| F_{\rho _\theta }-F_{\rho ^*}\| ^2\).)
Assume (E): \(\| \mu _t-\pi _M\| _{{\mathrm{TV}}}\to 0\) and \(\sup _t\int \frac1M\sum _\ell a_\ell ^2\, \mathrm d\mu _t\le Q{\lt}\infty \). Then for every bounded measurable \(g\), with \(\bar X(\theta )=\int g\, \mathrm d\Pi \rho _\theta \), \(m=\int g\, \mathrm d\Pi \rho ^*\) and \(C_g\) as in Proposition 544,
hence \(\limsup _{M\to \infty }\limsup _{t\to \infty }{\mathbb E}\bigl|\int g\, \mathrm d\Pi \rho ^{M}_t -\int g\, \mathrm d\Pi \rho ^*\bigr|=0\). (Truncation: \(h\wedge T\) converges by total variation, and \({\mathbb E}_{\mu _t}[h-h\wedge T]\le {\mathbb E}_{\mu _t}h^2/T\le (2C^2Q+2m^2)/T\) uniformly in \(t\).)