Shallow Learning Tends to Ridgelet Transform

8 Dynamics: mean-field Langevin, log-Sobolev inequalities

Section 4 of the paper, restricted to what does not depend on stochastic calculus: the entropy sandwich (lem:entropy-sandwich), the exponential convergence of the mean-field Langevin flow from the energy identity and a uniform log-Sobolev inequality of the proximal Gibbs measures (thm:mfld-convergence), the decomposition, shear transport and log-Sobolev inequality of the proximal Gibbs measure (lem:lsi, cor:uniform-lsi). Of the standard facts about log-Sobolev inequalities, tensorization, Holley–Stroock, the transport by proper \(C^1\) Lipschitz maps and the KL form (L1) are proved in the chapter “Material for Mathlib”; the Gaussian log-Sobolev inequality and the well-posedness of the dynamics are stated and left as sorry. The transport by a \(C^1\) map is what the shear of lem:lsi needs, so lem:lsi(iv) assumes that the analysis maps \(S^*r\) are \(C^1\); this is proved for a smooth feature with bounded derivatives by differentiation under the integral sign.

8.1 Entropy sandwich

For every \(\rho \) with \(\int |a|\, \mathrm d\rho {\lt}\infty \) (in particular for \(\rho \in {\mathcal D}\)), \(\hat\mu _\rho \in {\mathcal D}\).

Proof ▶

For \(\rho \) with \(\int |a|\, \mathrm d\rho {\lt}\infty \) and every \(\rho '\in {\mathcal D}\), \(\operatorname {KL}(\rho '\| \hat\mu _\rho ){\lt}\infty \) and

\[ {\mathcal F}(\rho ')=G_\rho (\rho ')+\beta \, \operatorname {KL}(\rho '\| \hat\mu _\rho )-\beta \log Z_\rho ,\qquad G_\rho (\rho '):=L(\rho ')-\int a\, s_\rho (z)\, \mathrm d\rho ' . \]

(The chain rule \(\operatorname {KL}(\rho '\| \mu _U)=\operatorname {KL}(\rho '\| \hat\mu _\rho )-\frac1\beta \int a\, s_\rho \, \mathrm d\rho ' +\log \frac{Z_U}{Z_\rho }\), legitimate since \(\int |a s_\rho |\, \mathrm d\rho '\le c_\rho \int |a|\, \mathrm d\rho '{\lt}\infty \).)

Proof ▶

Assume (A3), (A4), (A5), \(f\in L^2(P_X)\), and let \(\rho ^*\) be the minimizer of \({\mathcal F}\) on \({\mathcal D}\). For \(\rho \in {\mathcal D}\), \(\operatorname {KL}(\rho \| \hat\mu _\rho ){\lt}\infty \) and

\[ \beta \, \operatorname {KL}(\rho \| \rho ^*)\ \le \ {\mathcal F}(\rho )-{\mathcal F}(\rho ^*)\ \le \ \beta \, \operatorname {KL}(\rho \| \hat\mu _\rho ). \]
Proof ▶

The left inequality is the strong convexity identity (Lemma 212). For the right one, write the decomposition of Lemma 496 for \(\rho '=\rho \) and \(\rho '=\rho ^*\); by the expansion of Lemma 204 around \(\rho \), \(G_\rho (\rho ^*)-G_\rho (\rho )=\frac12\| F_{\rho ^*}-F_\rho \| ^2\ge 0\), so \({\mathcal F}(\rho )-{\mathcal F}(\rho ^*)=\beta \operatorname {KL}(\rho \| \hat\mu _\rho )-\beta \operatorname {KL}(\rho ^*\| \hat\mu _\rho ) -\frac12\| F_{\rho ^*}-F_\rho \| ^2\le \beta \operatorname {KL}(\rho \| \hat\mu _\rho )\).

8.2 Log-Sobolev inequality of the proximal Gibbs measures

Theorem 498
✓
#

If \(\mu \) satisfies \(\mathrm{LSI}(\alpha )\) and \(0{\lt}\alpha '\le \alpha \), then \(\mu \) satisfies \(\mathrm{LSI}(\alpha ')\).

Proof ▶
Theorem 499

(F1, Gross.) For \(m\in {\mathbb R}\) and \(v{\gt}0\), \({\mathcal N}(m,v)\) satisfies \(\mathrm{LSI}(1/v)\).

Proof ▶

(F1, Gross.) Under (A8), \(\nu _0={\mathcal N}(0,(\beta /\lambda _z)I)\) satisfies \(\mathrm{LSI}(\lambda _z/\beta )\).

Proof ▶
Theorem 501
✓

(F2, tensorization.) If \(\mu _1\) satisfies \(\mathrm{LSI}(\alpha _1)\) and \(\mu _2\) satisfies \(\mathrm{LSI}(\alpha _2)\), then \(\mu _1\otimes \mu _2\) satisfies \(\mathrm{LSI}(\min \{ \alpha _1,\alpha _2\} )\).

Proof ▶

(F3, Holley–Stroock.) If \(\mu \) satisfies \(\mathrm{LSI}(\alpha )\), \(\psi \) is bounded measurable and \(\nu :=Z^{-1}e^{-\psi }\mu \), then \(\nu \) satisfies \(\mathrm{LSI}(\alpha e^{-\operatorname {osc}\psi })\), \(\operatorname {osc}\psi :=\sup \psi -\inf \psi \).

Proof ▶
Theorem 503
✓

(F4, transport.) If \(\mu \) satisfies \(\mathrm{LSI}(\alpha )\) and \(T\) is a proper \(C^1\) map which is Lipschitz with constant \(L_T\), then \(T_\# \mu \) satisfies \(\mathrm{LSI}(\alpha /L_T^2)\).

Proof ▶
Theorem 504
✓

If \(|s(z)-s(z')|\le \ell \, d(z,z')\) for all \(z,z'\), then \(T_s\) and \(T_s^{-1}\) are Lipschitz with constant \(1+\ell /\lambda \) for the metric \(d((a,z),(a',z'))=\max \{ |a-a'|,d(z,z')\} \) on \(\Theta \).

Proof ▶
Theorem 505
✓

If \(|s(z)-s(z')|\le \ell \, |z-z'|\) for all \(z,z'\), then \(T_s^{-1}\) is Lipschitz with constant \(1+\ell /\lambda \) for the Euclidean metric of \(\Theta \) (Lemma ??(ii): \(\| DT_s^{-1}\| _{\mathrm{op}}\le \kappa _0(\ell /\lambda )\le 1+\ell /\lambda \)).

Proof ▶

Step 1: it suffices to compare the squared norms.

Step 2: the components of ‘T⁻¹x − T⁻¹y‘ are ‘(a − a’) − (s z − s z’)/λ‘ and ‘z − z’‘.

Step 3: with ‘u = |a − a’|‘, ‘w = ‖z − z’‖‘, the first component is bounded by ‘u + k w‘, and ‘(u + k w)² + w² ≤ (1 + k)² (u² + w²)‘.

For \(\rho \in {\mathcal D}\) (indeed for any \(\rho \) with \(\int |a|\, \mathrm d\rho {\lt}\infty \)), \(\hat\mu _\rho (\, \mathrm da\, \mathrm dz)=\hat\nu _\rho (\, \mathrm dz)\, {\mathcal N}(-s_\rho (z)/\lambda ,\beta /\lambda )(\, \mathrm da)\), where \(\hat\nu _\rho \) is the hidden marginal of \(\hat\mu _\rho \).

Proof ▶

The hidden marginal of \(\hat\mu _\rho \) is \(\hat\nu _\rho (\, \mathrm dz) =\hat Z_\rho ^{-1}\exp \bigl(\frac{s_\rho (z)^2}{2\lambda \beta }\bigr)\nu _0(\, \mathrm dz)\), with \(\hat Z_\rho =\int e^{s_\rho ^2/(2\lambda \beta )}\, \mathrm d\nu _0 \in [1,e^{c_\rho ^2/(2\lambda \beta )}]\).

Proof ▶

For \(T_\rho (a,z):=(a+s_\rho (z)/\lambda ,\ z)\), \((T_\rho )_\# \hat\mu _\rho ={\mathcal N}(0,\beta /\lambda )\otimes \hat\nu _\rho \).

Proof ▶

By Lemma 506, under \(\hat\mu _\rho \) one has \(z\sim \hat\nu _\rho \) and \(a\mid z\sim {\mathcal N}(-s_\rho (z)/\lambda ,\beta /\lambda )\), so \(a+s_\rho (z)/\lambda \mid z\sim {\mathcal N}(0,\beta /\lambda )\) does not depend on \(z\).

Theorem 509
✓

If \(z\mapsto \varphi _z(x)\) is \(C^1\) for every \(x\) and \(Z\) is finite dimensional, then \(x\mapsto D_z\varphi _z(x)\) is measurable for every \(z\) (a limit of difference quotients along a basis).

Proof ▶

Step 1: the difference quotients along the basis are measurable in ‘x‘.

Step 2: a linear form is the sum of its values on the basis times the coordinates.

Step 3: the difference quotients converge to the derivative.

If the feature map is smooth with bounded derivatives (Definition 127), \(Z\) is finite dimensional and \(r\in L^1(P_X)\), then \(S^*r\colon z\mapsto \int \varphi _z(x)r(x)P_X(\, \mathrm dx)\) is \(C^1\) with \(D(S^*r)(z)=\int r(x)D_z\varphi _z(x)P_X(\, \mathrm dx)\).

Proof ▶

Differentiation under the integral sign: \(|r(x)D_z\varphi _z(x)|\le L|r(x)|\) with \(L=\sup \| \nabla _z\varphi \| \) is an integrable bound, \(x\mapsto D_z\varphi _z(x)\) is measurable (Lemma 509), and the derivative is continuous by dominated convergence.

Under (A8), \(\hat\nu _\rho \) satisfies \(\mathrm{LSI}(\hat\alpha _\rho )\) with \(\hat\alpha _\rho \ge \frac{\lambda _z}\beta \exp \bigl(-\frac{\operatorname {osc}(s_\rho ^2)}{2\lambda \beta }\bigr) \ge \frac{\lambda _z}\beta \exp \bigl(-\frac{c_\rho ^2}{2\lambda \beta }\bigr)\).

Proof ▶

\(\nu _0={\mathcal N}(0,(\beta /\lambda _z)I)\) satisfies \(\mathrm{LSI}(\lambda _z/\beta )\) by (F1), and \(\hat\nu _\rho =Z^{-1}e^{-\psi }\nu _0\) with \(\psi =-s_\rho ^2/(2\lambda \beta )\), \(-c_\rho ^2/(2\lambda \beta )\le \psi \le 0\); apply (F3).

Under (A1), (A3), (A5), (A8), and assuming that \(S^*r\) is \(C^1\) for every \(r\in L^2(P_X)\) (Lemma 510), \(\hat\mu _\rho \) satisfies \(\mathrm{LSI}(\alpha _\rho )\) with

\[ \alpha _\rho \ \ge \ \frac{\min \{ \lambda ,\ \lambda _z e^{-c_\rho ^2/(2\lambda \beta )}\} }{\beta \, (1+\ell _\rho /\lambda )^2} \ \ge \ \frac{\min \{ \lambda ,\ \lambda _z e^{-c_\rho ^2/(2\lambda \beta )}\} }{\beta \, (1+\sqrt{R_X^2+1}\, c_\rho /\lambda )^2}, \]

where \(\ell _\rho =\sup |\nabla s_\rho |\le \sqrt{R_X^2+1}\, c_\rho \). The right-hand side does not depend on the dimension \(m\).

Proof ▶

\({\mathcal N}(0,\beta /\lambda )\) satisfies \(\mathrm{LSI}(\lambda /\beta )\) (F1), so by (F2) and Lemma 511 the product \(\pi _\rho ={\mathcal N}(0,\beta /\lambda )\otimes \hat\nu _\rho \) satisfies \(\mathrm{LSI}(\min \{ \lambda /\beta ,\hat\alpha _\rho \} )\). By Lemma 508, \(\hat\mu _\rho =(T_\rho ^{-1})_\# \pi _\rho \), and \(T_\rho ^{-1}\) is Lipschitz with constant \(1+\ell _\rho /\lambda \) (Lemma 505, with \(\ell _\rho \le \sqrt{R_X^2+1}\, c_\rho \) by Lemma 191); it is \(C^1\) since \(s_\rho \) is, and proper since it is a homeomorphism (its inverse is \(T_\rho \)); conclude by (F4).

For \(\rho \in {\mathcal D}\), \(c_\rho ^2=2L(\rho )\le 2\bigl({\mathcal F}(\rho )+\beta \log Z_U\bigr)\) (here \(c_\rho ^2\le 2{\mathcal F}(\rho )\), the constant of the free energy being dropped).

Proof ▶
Theorem 514
✓
#

The map \(c\mapsto \alpha (c)\) of Definition 125 is nonincreasing on \([0,\infty )\).

Proof ▶

For \(E\in {\mathbb R}\), on the sublevel set \({\mathcal D}_E:=\{ \rho \in {\mathcal D}:{\mathcal F}(\rho )\le E\} \) and with \(c_E:=\sqrt{2(E+\beta \log Z_U)}\),

\[ \inf _{\rho \in {\mathcal D}_E}\alpha _\rho \ \ge \ \alpha _*(E) =\frac{\min \{ \lambda ,\ \lambda _z\, e^{-c_E^2/(2\lambda \beta )}\} }{\beta \, (1+\sqrt{R_X^2+1}\, c_E/\lambda )^2}{\gt}0, \]

i.e. \(\hat\mu _\rho \) satisfies \(\mathrm{LSI}(\alpha _*(E))\) for every \(\rho \in {\mathcal D}_E\).

Proof ▶

\(c_\rho \le c_E\) on \({\mathcal D}_E\) by Corollary 513, and the constant of Lemma 512 is nonincreasing in \(c_\rho \) (Lemma 514); conclude by Lemma 498.

For \(\rho _0=\mu _U={\mathcal N}(0,\beta /\lambda )\otimes \nu _0\), \(F_{\mu _U}=0\), \(L(\mu _U)=\| f\| ^2/2\) and \({\mathcal F}(\mu _U)=\| f\| ^2/2-\beta \log Z_U\) (here \(\| f\| ^2/2\)).

Proof ▶

Step 1: ‘F_μ_U(x) = (∫ a d𝒩(0, β/λ)) (∫ φ_z(x) dν₀) = 0‘.

Step 2: ‘L(μ_U) = ½ ∫ f² dP = ½ ‖f‖²‘.

Step 3: ‘KL(μ_U‖μ_U) = 0‘.

For \(\rho _0=\mu _U\), \(c_{{\mathcal F}(\mu _U)}=\| f\| \) and

\[ \alpha _*:=\alpha _*({\mathcal F}(\mu _U)) =\frac{\min \{ \lambda ,\ \lambda _z\, e^{-\| f\| ^2/(2\lambda \beta )}\} }{\beta \, (1+\sqrt{R_X^2+1}\, \| f\| /\lambda )^2}. \]
Proof ▶

Since \(\rho ^*\in {\mathcal D}_{{\mathcal F}(\mu _U)}\), \(\rho ^*=\hat\mu _{\rho ^*}\) satisfies \(\mathrm{LSI}(\alpha _*)\).

Proof ▶

8.3 Exponential convergence of the mean-field Langevin dynamics

(L1.) If \(\mu \) satisfies \(\mathrm{LSI}(\alpha )\) on a finite-dimensional space, \(\rho \ll \mu \) is a probability measure with \(\operatorname {KL}(\rho \| \mu ){\lt}\infty \), \(\log \frac{\, \mathrm d\rho }{\, \mathrm d\mu }\) is \(C^1\) and \(\nabla \log \frac{\, \mathrm d\rho }{\, \mathrm d\mu }\in L^2(\rho )\), then \(\operatorname {KL}(\rho \| \mu )\le \frac1{2\alpha }I(\rho |\mu )\) (substitute \(g=\sqrt{\, \mathrm d\rho /\, \mathrm d\mu }\), cut off by compactly supported functions, in the LSI).

Proof ▶

If a probability measure \(\mu \) on \(\Theta \) (\(Z\) finite dimensional) satisfies \(\mathrm{LSI}(\alpha )\), then it satisfies the KL form of \(\mathrm{LSI}(\alpha )\).

Proof ▶

Along an MFLD flow, \(t\mapsto {\mathcal F}(\rho _t)\) is nonincreasing on \([0,\infty )\).

Proof ▶

Assume (A3), (A5), \(f\in L^2(P_X)\), let \((\rho _t)_{t\ge 0}\) be an MFLD flow (Hypothesis W) and assume that \(\hat\mu _{\rho _t}\) satisfies the KL form of \(\mathrm{LSI}(\alpha _*)\) for every \(t\ge 0\). Then for all \(t\ge 0\)

\[ {\mathcal F}(\rho _t)-{\mathcal F}(\rho ^*)\le e^{-2\alpha _*\beta t}\bigl({\mathcal F}(\rho _0)-{\mathcal F}(\rho ^*)\bigr). \]
Proof ▶

Put \(\Phi (t):={\mathcal F}(\rho _t)-{\mathcal F}(\rho ^*)\ge 0\). By (W2), \(\Phi \) is nonincreasing. For a.e. \(u\ge 0\), the right inequality of Lemma 497 and (L1) give \(\Phi (u)\le \beta \operatorname {KL}(\rho _u\| \hat\mu _{\rho _u}) \le \frac{\beta }{2\alpha _*}I(\rho _u|\hat\mu _{\rho _u})\). Hence, for \(0\le s\le t\), \(\Phi (t)=\Phi (s)-\beta ^2\int _s^tI(\rho _u|\hat\mu _{\rho _u})\, \mathrm du \le \Phi (s)-2\alpha _*\beta \int _s^t\Phi (u)\, \mathrm du\le \Phi (s)-2\alpha _*\beta (t-s)\Phi (t)\), and Lemma 909 (Grönwall) concludes.

Under the hypotheses of Theorem 522, for all \(t\ge 0\), \(\operatorname {KL}(\rho _t\| \rho ^*) \le \frac1\beta e^{-2\alpha _*\beta t}\bigl({\mathcal F}(\rho _0)-{\mathcal F}(\rho ^*)\bigr)\). (The left inequality of Lemma 497.)

Proof ▶

Under the hypotheses of Theorem 522, for all \(t\ge 0\), \(\| F_{\rho _t}-F_{\rho ^*}\| _{L^2(P_X)}^2 \le 2e^{-2\alpha _*\beta t}\bigl({\mathcal F}(\rho _0)-{\mathcal F}(\rho ^*)\bigr)\). (The strong convexity identity of Lemma 212.)

Proof ▶

Assume (A1), (A3), (A5), (A8), \(f\in L^2(P_X)\) and (A9): \(\rho _0\in {\mathcal D}\), \(\int |z|^2\, \mathrm d\rho _0{\lt}\infty \) and \(\int e^{c_0|a|}\, \mathrm d\rho _0{\lt}\infty \) for some \(c_0{\gt}0\). Then the MFLD has a solution on \([0,\infty )\) (unique in law), and its flow of laws \((\rho _t)_{t\ge 0}\) satisfies (W1) and (W2), i.e. it is an MFLD flow with \(\rho _0\) as initial law.

Proof ▶

Assume (A1), (A3), (A5), (A8), \(f\in L^2(P_X)\) and (A9). Let \((\rho _t)_{t\ge 0}\) be the flow of laws of the solution of Theorem 525 and \(\alpha _*:=\alpha _*({\mathcal F}(\rho _0))\). Then for all \(t\ge 0\)

\[ {\mathcal F}(\rho _t)-{\mathcal F}(\rho ^*) \le e^{-2\alpha _*\beta t}\bigl({\mathcal F}(\rho _0)-{\mathcal F}(\rho ^*)\bigr),\quad \operatorname {KL}(\rho _t\| \rho ^*) \le \frac1\beta e^{-2\alpha _*\beta t}\bigl({\mathcal F}(\rho _0)-{\mathcal F}(\rho ^*)\bigr), \]
\[ \| F_{\rho _t}-F_{\rho ^*}\| _{L^2(P_X)}^2 \le 2e^{-2\alpha _*\beta t}\bigl({\mathcal F}(\rho _0)-{\mathcal F}(\rho ^*)\bigr). \]
Proof ▶

Theorem 525 gives the MFLD flow; since \({\mathcal F}(\rho _t)\le {\mathcal F}(\rho _0)\) (Lemma 521), Corollary 515 gives \(\mathrm{LSI}(\alpha _*)\) for every \(\hat\mu _{\rho _t}\) (the analysis maps are \(C^1\) by Lemma 510, the feature being smooth with bounded derivatives), hence its KL form (Lemma 520), and Theorems 522, 523, 524 apply.

8.4 Static chaos of the finite-particle Gibbs measure

The \(M\)-particle free energy decomposes exactly relative to the product of the mean-field minimizer, so that the relative entropy of the \(M\)-particle Gibbs measure with respect to \(\rho ^{*\otimes M}\) is bounded uniformly in \(M\); this closes the order “\(t\to \infty \) at fixed \(M\) first” of the order-of-limits theorem under the ergodicity of the finite-particle dynamics (manuscript note notes-14).

Theorem 527
✓

Let \(\mu \) be a nonzero \(\sigma \)-finite measure on \(\Theta \), \(\beta {\gt}0\) and \(W\) measurable with \(Z_W=\int e^{-W/\beta }\, \mathrm d\mu {\lt}\infty \). Then \(\hat\mu _W^{\otimes M}=\widehat{(\mu ^{\otimes M})}_{W_M}\) with \(W_M(\theta )=\sum _{\ell }W(\theta _\ell )\), i.e. \(\frac{\, \mathrm d\hat\mu _W^{\otimes M}}{\, \mathrm d\mu ^{\otimes M}}(\theta ) =\prod _\ell \frac{e^{-W(\theta _\ell )/\beta }}{Z_W}=\frac{e^{-W_M(\theta )/\beta }}{Z_W^M}\), and \(Z_{W_M}=Z_W^M\).

Proof ▶

Step 1: ‘e^−W_M/β = ∏ᵢ e^−W(θᵢ)/β‘ and ‘Z_W_M = Z_W^|ι|‘.

Step 2: both sides are ‘μ^⊗ι‘ with the product density (‘lem:pi-withDensity‘).

For every \(\theta \in \Theta ^M\), \(0\le L(\rho _\theta )\le \tfrac 12\bigl(\tfrac 1M\sum _\ell |a_\ell |+\| f\| \bigr)^2 \le \tfrac 1M\sum _\ell a_\ell ^2+\| f\| ^2\).

Proof ▶

Let \(\mu \in {\mathcal P}(\Theta ^M)\) with \(\operatorname {KL}(\mu \| \mu _U^{\otimes M}){\lt}\infty \). Then \(\int a_\ell ^2\, \mu (\, \mathrm d\theta ){\lt}\infty \) for every \(\ell \), hence \(\int |a_\ell |\, \mathrm d\mu {\lt}\infty \) and \(L(\rho _\theta )\), \(\sum _\ell a_\ell s(z_\ell )\) (\(s\) bounded measurable) are \(\mu \)-integrable. (Donsker–Varadhan with \(g(\theta )=\frac{\lambda }{4\beta }a_\ell ^2\), \(\int e^g\, \mathrm d\mu _U^{\otimes M}=\int e^{\lambda a^2/(4\beta )}\, \mathrm d\mu _U=\sqrt2\).)

Proof ▶

Since \(0\le L(\rho _\theta )\), \(Z_M=\int e^{-ML(\rho _\theta )/\beta } \, \mathrm d\mu _U^{\otimes M}\in (0,1]\), so \(\pi _M\) is a probability measure, and \(\operatorname {KL}(\pi _M\| \mu _U^{\otimes M})=-\frac M\beta \int L(\rho _\theta )\, \mathrm d\pi _M-\log Z_M{\lt}\infty \) (the potential \(ML(\rho _\theta )\) is \(\pi _M\)-integrable because \(ue^{-u/\beta }\le \beta \)).

Proof ▶

For \(\mu \in {\mathcal P}(\Theta ^M)\) with \(\operatorname {KL}(\mu \| \mu _U^{\otimes M}){\lt}\infty \), \(\operatorname {KL}(\mu \| \pi _M){\lt}\infty \) and \({\mathcal F}^M(\mu )=\beta \, \operatorname {KL}(\mu \| \pi _M)-\beta \log Z_M\), \(Z_M:=\int e^{-ML(\rho _\theta )/\beta }\, \mathrm d\mu _U^{\otimes M}\) (Lemma 187 with \(W=ML(\rho _\theta )\) and reference measure \(\mu _U^{\otimes M}\); \(W\in L^1(\mu )\) by Lemma 529).

Proof ▶

\(\pi _M\) is the minimizer of \({\mathcal F}^M\) on \(\{ \mu \in {\mathcal P}(\Theta ^M):\operatorname {KL}(\mu \| \mu _U^{\otimes M}){\lt}\infty \} \): \({\mathcal F}^M(\pi _M)=-\beta \log Z_M\le {\mathcal F}^M(\mu )\).

Proof ▶

With \(s^*=S^*r_{\rho ^*}\) and \(W^*(\theta )=a\, s^*(z)\), \(\rho ^{*\otimes M}=\widehat{(\mu _U^{\otimes M})}_{W^*_M}\), \(W^*_M(\theta )=\sum _\ell a_\ell s^*(z_\ell )=M\int as^*\, \mathrm d\rho _\theta \), and \(\log \frac{\, \mathrm d\rho ^{*\otimes M}}{\, \mathrm d\mu _U^{\otimes M}}(\theta ) =-\frac1\beta \sum _\ell a_\ell s^*(z_\ell )-M\log Z_*\).

Proof ▶

Let \(\rho ^*\) be the minimizer of \({\mathcal F}\) on \({\mathcal D}\). For every \(\mu \in {\mathcal P}(\Theta ^M)\) with \(\operatorname {KL}(\mu \| \mu _U^{\otimes M}){\lt}\infty \), \(\operatorname {KL}(\mu \| \rho ^{*\otimes M}){\lt}\infty \) and

\[ {\mathcal F}^M(\mu )-M{\mathcal F}(\rho ^*)=\beta \, \operatorname {KL}(\mu \| \rho ^{*\otimes M}) +\frac M2\int _{\Theta ^M}\| F_{\rho _\theta }-F_{\rho ^*}\| _{L^2(P_X)}^2\, \mu (\, \mathrm d\theta ). \]

(Chain rule for \(\rho ^{*\otimes M}=\widehat{(\mu _U^{\otimes M})}_{W^*_M}\), the identity \({\mathcal F}(\rho ^*)=L(\rho ^*)-\int as^*\, \mathrm d\rho ^*-\beta \log Z_*\) and the expansion \(L(\rho _\theta )-L(\rho ^*)=\int as^*\, \mathrm d(\rho _\theta -\rho ^*) +\frac12\| F_{\rho _\theta }-F_{\rho ^*}\| ^2\).)

Proof ▶

Step 0: notation ‘s* = S* r_ρ*‘, ‘W(θ) = a s*(z)‘, ‘W_M(θ) = ∑_ℓ W(θ_ℓ)‘.

Step 1: ‘Z_W_M = Z_*^M‘.

Step 2: the chain rule ‘KL(μ‖ρ*^⊗M) = KL(μ‖μ_U^⊗M) + (1/β) ∫ W_M dμ + M log Z_*‘.

Step 3: ‘ℱ(ρ*) = L(ρ*) − ∫ W dρ* − β log Z_*‘.

Step 4: the pointwise expansion ‘M L(ρ_θ) − W_M(θ) − M (L(ρ*) − ∫ W dρ*) = (M/2) ‖F_ρ_θ − F_ρ*‖²‘.

Step 5: integrate the pointwise identity against ‘μ‘.

Let \(\rho \in {\mathcal P}(\Theta )\) with \(\int a^2\, \mathrm d\rho {\lt}\infty \) and \(M\ge 1\). Then

\[ \int _{\Theta ^M}\| F_{\rho _\theta }-F_\rho \| _{L^2(P_X)}^2\, \rho ^{\otimes M}(\, \mathrm d\theta ) =\frac1M\int _{{\mathcal X}}\mathrm{Var}_\rho \bigl(a\varphi _z(x)\bigr)\, P_X(\, \mathrm dx) \le \frac1M\int a^2\, \mathrm d\rho . \]

(For each \(x\), \(F_{\rho _\theta }(x)=\frac1M\sum _\ell a_\ell \varphi _{z_\ell }(x)\) is a mean of i.i.d. variables of mean \(F_\rho (x)\); Fubini is justified by \(\int a^2\, \mathrm d\rho {\lt}\infty \).)

Proof ▶

Step A: ‘‖F_ρ_θ − F_ρ‖² = ∫ (F_ρ_θ(x) − F_ρ(x))² P(dx)‘ for every ‘θ‘.

Step B: the integrand is dominated by ‘2 (1/M) ∑_ℓ a_ℓ² + 2 (∫ |a| dρ)²‘, so Fubini applies.

Step C: for each ‘x‘, the variance of the mean of ‘M‘ i.i.d. copies of ‘a φ_z(x)‘ is ‘(1/M) Var_ρ(a φ_z(x)) ≤ (1/M) ∫ a² dρ‘.

Step D: Fubini and integration of the pointwise bound over ‘x‘.

\(\operatorname {KL}(\rho ^{*\otimes M}\| \mu _U^{\otimes M})=M\, \operatorname {KL}(\rho ^*\| \mu _U){\lt}\infty \).

Proof ▶

Let \(\rho ^*\) be the minimizer of \({\mathcal F}\) on \({\mathcal D}\) and \(M\ge 1\). Then

\[ \beta \, \operatorname {KL}(\pi _M\| \rho ^{*\otimes M}) +\frac M2\, {\mathbb E}_{\pi _M}\| F_{\rho _\theta }-F_{\rho ^*}\| _{L^2(P_X)}^2 \le \frac12\int a^2\, \mathrm d\rho ^* . \]

(Theorem 534 at \(\mu =\pi _M\) and at \(\mu =\rho ^{*\otimes M}\), the minimality of \(\pi _M\) for \({\mathcal F}^M\), and Lemma 535.)

Proof ▶

Step 1: the identity at ‘μ = π_M‘ and at ‘μ = ρ*^⊗M‘.

Step 2: minimality of ‘π_M‘ and the i.i.d. variance bound.

Under the hypotheses of Theorem 537, \(\operatorname {KL}(\pi _M\| \rho ^{*\otimes M})\le \frac1{2\beta }\int a^2\, \mathrm d\rho ^*\) (independent of \(M\)) and \({\mathbb E}_{\pi _M}\| F_{\rho _\theta }-F_{\rho ^*}\| ^2\le \frac1M\int a^2\, \mathrm d\rho ^*\).

Proof ▶

With \(A:=\min \{ \| f\| ^2/\lambda ^2,\| f\| ^2/(4\lambda )\} +\beta /\lambda \ge \int a^2\, \mathrm d\rho ^*\) (Theorem 233), \(\operatorname {KL}(\pi _M\| \rho ^{*\otimes M})\le A/(2\beta )\) and \({\mathbb E}_{\pi _M}\| F_{\rho _\theta }-F_{\rho ^*}\| ^2\le A/M\).

Proof ▶

Let \(s\) be measurable with \(|s|\le C\) and \(\rho _s=\hat\mu _{W_s}\), \(W_s(\theta )=a\, s(z)\). Then \(\int e^{c|a|}\, \mathrm d\rho _s{\lt}\infty \) for every \(c\in {\mathbb R}\), since \(e^{-as(z)/\beta }e^{c|a|}\le e^{(C/\beta +c)|a|}\) and the \(a\)-marginal of \(\mu _U\) is Gaussian.

Proof ▶

Under the hypotheses of Lemma 540, for \(c\ge 0\), \(\int e^{c|a|}\, \mathrm d\rho _s\le Z_s^{-1}\int e^{(c+C/\beta )|a|}\, \mathrm d\mu _U \le 2\exp \bigl(\tfrac {\beta }{2\lambda }(c+C/\beta )^2\bigr)\), using \(Z_s\ge 1\) and the Gaussian moment generating function.

Proof ▶

Step 1: ‘∫ e^c|a| dρ_s = Z⁻¹ ∫ e^−a s/β e^c|a| dμ_U ≤ Z⁻¹ ∫ e^t|a| dμ_U‘.

Step 2: ‘∫ e^t|a| dμ_U ≤ ∫ (e^ta + e^−ta) dμ_U = 2 e^t² (β/λ)/2‘ (Gaussian moment generating function).

Let \(g\) be measurable with \(|g|\le C\), \(m:=\int ag(z)\, \mathrm d\rho ^*\) and \(Y(\theta ):=ag(z)-m\). Then \(K_g:=\int Y^2e^{|Y|}\, \mathrm d\rho ^*{\lt}\infty \), since \(|Y|\le C|a|+|m|\), \(u^2\le 2e^u\) and \(\int e^{2C|a|}\, \mathrm d\rho ^*{\lt}\infty \).

Proof ▶

Under the hypotheses of Lemma 542, \(K_g\le 4\exp \bigl(2C(\| f\| /\lambda +\sqrt{2\beta /(\pi \lambda )})\bigr) \exp \bigl(\tfrac \beta {2\lambda }(2C+\| f\| /\beta )^2\bigr)\), a constant depending only on \(C,\| f\| ,\lambda ,\beta \).

Proof ▶

Step 1: ‘|m| ≤ C ∫ |a| dρ* ≤ C (‖f‖/λ + √(2β/(πλ)))‘.

Step 2: the pointwise bound ‘Y² e^|Y| ≤ 2 e^2|m| e^2C|a|‘.

Step 3: integrate and use the Gaussian exponential-moment bound.

Let \(g\colon Z\to {\mathbb R}\) be measurable with \(|g|\le C\), \(m:=\int ag(z)\, \mathrm d\rho ^* =\int g\, \mathrm d\Pi \rho ^*\), \(\bar X(\theta ):=\frac1M\sum _\ell a_\ell g(z_\ell )=\int g\, \mathrm d\Pi \rho _\theta \) and \(K_g:=\int (ag(z)-m)^2e^{|ag(z)-m|}\, \mathrm d\rho ^*{\lt}\infty \). Then

\[ {\mathbb E}_{\pi _M}\Bigl|\int g\, \mathrm d\Pi \rho _\theta -\int g\, \mathrm d\Pi \rho ^*\Bigr| \le \frac{\operatorname {KL}(\pi _M\| \rho ^{*\otimes M})+\log 2+K_g}{\sqrt M} \le \frac{C_g}{\sqrt M},\qquad C_g:=\frac{A}{2\beta }+\log 2+K_g . \]

(Donsker–Varadhan on \(\Theta ^M\) with \(h=\sqrt M|\bar X-m|\): with \(u=1/\sqrt M\) and \(Y_\ell =a_\ell g(z_\ell )-m\) i.i.d. centered, \({\mathbb E}_{\rho ^{*\otimes M}}e^{\pm u\sum _\ell Y_\ell }=({\mathbb E}e^{\pm uY})^M\le (1+K_gu^2)^M\le e^{K_g}\), so \(\log {\mathbb E}_{\rho ^{*\otimes M}}e^h\le \log 2+K_g\).)

Proof ▶

Step 0: notation. ‘Y(p) = a g(z) − m‘ is centered with ‘K = ∫ Y² e^|Y| dρ* < ∞‘; ‘S(θ) = ∑_ℓ Y(θ_ℓ)‘, ‘u = 1/√M‘, ‘h(θ) = u |S(θ)| = √M |X̄(θ) − m|‘.

Step 1: ‘X̄(θ) − m = (1/M) S(θ)‘.

Step 2: the exponential moments of ‘± u S‘ under ‘ρ*^⊗M‘ are at most ‘e^K‘.

Step 3: ‘e^u|S| ≤ e^uS + e^−uS‘, so ‘∫ e^u|S| dρ*^⊗M ≤ 2 e^K‘.

Step 4: Donsker–Varadhan on ‘Θ^M‘: ‘∫ u|S| dπ_M ≤ KL(π_M‖ρ*^⊗M) + log 2 + K‘.

Step 5: ‘|X̄ − m| = u · (u |S|)‘ since ‘u² = 1/M‘; conclude.

Under the hypotheses of Proposition 544, with \(A=\min \{ \| f\| ^2/\lambda ^2,\| f\| ^2/(4\lambda )\} +\beta /\lambda \), \({\mathbb E}_{\pi _M}|\bar X-m|\le C_g/\sqrt M\) with \(C_g=A/(2\beta )+\log 2+K_g\).

Proof ▶

Under the hypotheses of Proposition 544, \({\mathbb E}_{\pi _M}|\bar X-m|\le C_g/\sqrt M\) with the explicit constant \(C_g=\frac A{2\beta }+\log 2+4\exp \bigl(2C(\| f\| /\lambda +\sqrt{2\beta /(\pi \lambda )})\bigr) \exp \bigl(\tfrac \beta {2\lambda }(2C+\| f\| /\beta )^2\bigr)\), which depends only on \(C,\| f\| ,\lambda ,\beta \).

Proof ▶

For every permutation \(\sigma \) of \(\{ 1,\dots ,M\} \), the image of \(\pi _M\) under \(\theta \mapsto \theta \circ \sigma \) is \(\pi _M\): \(\mu _U^{\otimes M}\) is permutation invariant and \(L(\rho _{\theta \circ \sigma })=L(\rho _\theta )\). In particular all one-particle marginals of \(\pi _M\) coincide.

Proof ▶

Let \(Z\) be standard Borel. For every \(\ell \), the one-particle marginal \(\pi _M^{(\ell )}\) of \(\pi _M\) satisfies \(M\, \operatorname {KL}(\pi _M^{(\ell )}\| \rho ^*)=\sum _{\ell '}\operatorname {KL}(\pi _M^{(\ell ')}\| \rho ^*) \le \operatorname {KL}(\pi _M\| \rho ^{*\otimes M})\), by exchangeability and the superadditivity \(\operatorname {KL}(\mu \| \nu ^{\otimes M})\ge \sum _\ell \operatorname {KL}(\mu ^{(\ell )}\| \nu )\).

Proof ▶

Under the hypotheses of Theorem 537 with \(Z\) standard Borel, \(\operatorname {KL}(\pi _M^{(\ell )}\| \rho ^*)\le \frac{A}{2\beta M}\) for every \(\ell \).

Proof ▶

Assume (E): \(\| \mu _t-\pi _M\| _{{\mathrm{TV}}}\to 0\) and \(\sup _t\int \frac1M\sum _\ell a_\ell ^2\, \mathrm d\mu _t\le Q{\lt}\infty \). Then for every bounded measurable \(g\), with \(\bar X(\theta )=\int g\, \mathrm d\Pi \rho _\theta \), \(m=\int g\, \mathrm d\Pi \rho ^*\) and \(C_g\) as in Proposition 544,

\[ \limsup _{t\to \infty }{\mathbb E}_{\mu _t}\bigl|\bar X-m\bigr|\le \frac{C_g}{\sqrt M}, \]

hence \(\limsup _{M\to \infty }\limsup _{t\to \infty }{\mathbb E}\bigl|\int g\, \mathrm d\Pi \rho ^{M}_t -\int g\, \mathrm d\Pi \rho ^*\bigr|=0\). (Truncation: \(h\wedge T\) converges by total variation, and \({\mathbb E}_{\mu _t}[h-h\wedge T]\le {\mathbb E}_{\mu _t}h^2/T\le (2C^2Q+2m^2)/T\) uniformly in \(t\).)

Proof ▶

Step 0: notation: ‘m‘, the statistic ‘h(θ) = |X̄(θ) − m|‘, its bound ‘C S(θ) + |m|‘ with ‘S(θ) = (1/M) ∑_ℓ |a_ℓ|‘, the bound ‘B‘ of Proposition 14.4 and the second-moment bound ‘Q₂ = 2 C² Q + 2 m²‘ of ‘h‘.

Step 1: the truncation level ‘T‘ with ‘Q₂/T ≤ ε/2‘.

Step 2: integrability of ‘h‘, ‘h²‘ and ‘h ∧ T‘ under each ‘μ_t‘ and under ‘π_M‘.

Step 3: the tail bound ‘∫ h dμ_t ≤ ∫ h ∧ T dμ_t + Q₂/T‘ (Chebyshev: ‘h − h ∧ T ≤ h²/T‘).

Step 4: total-variation convergence for the truncated statistic, and ‘∫ h ∧ T dπ_M ≤ B‘.

Step 5: eventually ‘2 T ‖μ_t − π_M‖_TV ≤ ε/2‘.

For the tanh feature (A3) with \(z=(w,b)\in {\mathbb R}^{m+1}\) (Euclidean norm), \(\| \nabla _z\varphi _z(x)\| \le \sqrt{|x|^2+1}\) for all \(x\in {\mathbb R}^m\) and \(z\in {\mathbb R}^{m+1}\). Indeed \(\nabla _z\varphi _z(x)=\tanh '(w^\top x-b)\, (x,-1)\) with \(0{\lt}\tanh '\le 1\) and \(|(x,-1)|=\sqrt{|x|^2+1}\). In particular, under (A1) the feature is Lipschitz in the hidden parameter with constant \(\sqrt{R_X^2+1}\).

Proof ▶

Step 1: the chain rule bound ‘‖∇_z φ_z(x)‖ ≤ |tanh’(⟪w, x⟫ − b)| · ‖(x, −1)‖‘.

Step 2: ‘|tanh’| = 1 − tanh² ≤ 1‘ and ‘‖(x, −1)‖ = √(‖x‖² + 1)‘.

For the tanh feature (A3), with \(z=(w,b)\in {\mathbb R}^{m+1}\) (Euclidean norm, Definition 6), and inputs \(|x|\le R_X\) (A1), the map \(z\mapsto \varphi _z(x)\) is \(C^\infty \) and, for every \(n\ge 0\),

\[ \sup _{|x|\le R_X,\ z}\| \nabla _z^n\varphi _z(x)\| \le \| \tanh ^{(n)}\| _\infty \, (R_X^2+1)^{n/2}{\lt}\infty , \]

where \(\| \tanh ^{(n)}\| _\infty \le \sup _{|s|\le 1}|p_n(s)|\) for the polynomials \(p_0=X\), \(p_{n+1}=p_n'\, (1-X^2)\) with \(\tanh ^{(n)}=p_n(\tanh )\). Hence the restriction of the tanh feature to \(\{ |x|\le R_X\} \) (Definition 3) is smooth with bounded derivatives in the sense of Definition 127.

Proof ▶

Step 1: smoothness. ‘z ↦ tanh(⟪w, x⟫ − b)‘ is the composition of the ‘C^∞‘ function ‘tanh‘ with the continuous linear map ‘(w, b) ↦ ⟪(x, −1), (w, b)⟫‘.

Step 2: the ‘n‘-th derivative of ‘tanh‘ is bounded by a constant ‘c_n‘ (‘tanh⁽ⁿ⁾ = p_n(tanh)‘ with ‘|tanh| ≤ 1‘).

Step 3: the chain rule for the linear map ‘(w, b) ↦ ⟪(x, −1), (w, b)⟫‘ of norm ‘√(‖x‖² + 1) ≤ √(R_X² + 1)‘ gives ‘‖∇_z^n φ_z(x)‖ ≤ c_n (R_X² + 1)^n/2‘.

8.5 The third stage of the order of the limits: \(t\to \infty \) at fixed \(M\) first

If \(\int |a|\, \mathrm d\rho {\lt}\infty \) and \(g\) is bounded measurable, then \(\int _Zg\, \mathrm d\Pi \rho =\int _\Theta a\, g(z)\, \rho (\, \mathrm da\, \mathrm dz)\).

Proof ▶

Fix the sample \((x_i,y_i)_{i=1}^N\) with \(|y_i|\le {Y_{\max }}\), \((\lambda ,\beta )\) and \(M\ge 1\); let \(\rho _N^*\) be the minimizer of \({\mathcal F}_N\) and let \((\mu _t)_{t\ge 0}\) be ergodic to the empirical \(M\)-particle Gibbs measure \(\pi _M\) (Assumption (E)). For every bounded measurable \(g\) with \(\| g\| _\infty \le C\) and every \({\varepsilon }{\gt}0\) there is \(t_{\varepsilon }\) such that for \(t\ge t_{\varepsilon }\)

\[ {\mathbb E}_{\mu _t}\, e^{(g)}_{M,t}={\mathbb E}_{\mu _t}\Bigl|\int g\, \mathrm d\Pi \rho _\theta -\int g\, \mathrm d\Pi \rho _N^*\Bigr| \le \frac{C_g({Y_{\max }},\lambda ,\beta )}{\sqrt M}+{\varepsilon }, \]

with the constant of Definition 135, which depends on the sample only through \({Y_{\max }}\). (Theorem 550 for the empirical feature, \(\| y\| _N\le {Y_{\max }}\), and \(\int g\, \mathrm d\Pi \rho _N^*=\int a\, g(z)\, \mathrm d\rho _N^*\).)

Proof ▶

Step 1: the dictionary: ‘ρ_N^*‘ minimizes the free energy of the empirical feature with target ‘y ∈ L²(uniform)‘, and ‘‖y‖_N ≤ Ymax‘.

Step 2: ‘thm:static-chaos-order‘ for the empirical feature.

Step 3: the constant is bounded uniformly in the sample.

Assume the hypotheses of Theorem 427 and let, for every sample and every \(M\ge 1\), \((\mu ^{M,N}_t)_{t\ge 0}\) be a family of laws on \(\Theta ^M\) ergodic to the empirical \(M\)-particle Gibbs measure (Assumption (E); for the \(M\)-particle empirical noisy gradient descent this is the ergodic theorem). Let \(g\) be bounded measurable, \({\varepsilon }{\gt}0\) and \(\delta \in (0,1)\). Then for the \((\lambda ,\beta )\) and \(N_1\) of Theorem 427, on the event of probability at least \(1-\delta \) over the sample, for every \(M\ge 1\)

\[ \limsup _{t\to \infty }{\mathbb E}\Bigl|\int g\, \mathrm d\Pi \rho _t^{M,N}-\int g\, u_0^\dagger \, \mathrm d\nu _0\Bigr| \le \frac{C^N_g}{\sqrt M}+\frac{2{\varepsilon }}3 , \]

in the form: for every \({\varepsilon }'{\gt}0\), eventually in \(t\), \({\mathbb E}_{\mu _t}|\int g\, \mathrm d\Pi \rho _\theta -\int g\, u_0^\dagger \, \mathrm d\nu _0| \le \frac{2{\varepsilon }}3+\frac{C_g({Y_{\max }},\lambda ,\beta )}{\sqrt M}+{\varepsilon }'\), with the constant of Definition 135. (The triangle inequality ?? at fixed \((M,t)\), integrated over \(\mu _t\): the second and third terms are deterministic on the event of stage (2) and bounded by \({\varepsilon }/3\) each, and the first is Theorem 554.)

Proof ▶

Stage (1) at ‘(λ, β)‘.

Stage (2) at ‘(λ, β)‘, on the event ‘E’_δ‘.

Stage (3) for the sample ‘ω‘ (labels bounded by ‘Ymax‘ on ‘E’_δ‘).

The triangle inequality ‘eq:three-stages‘, integrated over ‘μ_t‘.

Under the hypotheses of Corollary 555, since \(C_g({Y_{\max }},\lambda ,\beta )\) does not depend on the sample, there is \(M_0=M_0({\varepsilon };{Y_{\max }},\lambda ,\beta ,g)\) such that on the same event of probability at least \(1-\delta \), for every \(M\ge M_0\), eventually in \(t\), \({\mathbb E}_{\mu _t}\bigl|\int g\, \mathrm d\Pi \rho _\theta -\int g\, u_0^\dagger \, \mathrm d\nu _0\bigr|\le {\varepsilon }\); that is,

\[ \limsup _{M\to \infty }\limsup _{t\to \infty }{\mathbb E}\Bigl|\int g\, \mathrm d\Pi \rho _t^{M,N} -\int g\, u_0^\dagger \, \mathrm d\nu _0\Bigr|\le \frac{2{\varepsilon }}3{\lt}{\varepsilon }. \]

(\(M_0:=\max \{ \lceil (6C_g/{\varepsilon })^2\rceil ,1\} \) gives \(C_g/\sqrt M\le {\varepsilon }/6\); take \({\varepsilon }'={\varepsilon }/6\).)

Proof ▶