8 Dynamics: mean-field Langevin, log-Sobolev inequalities
Section 4 of the paper, restricted to what does not depend on stochastic calculus: the entropy sandwich (lem:entropy-sandwich), the exponential convergence of the mean-field Langevin flow from the energy identity and a uniform log-Sobolev inequality of the proximal Gibbs measures (thm:mfld-convergence), the decomposition, shear transport and log-Sobolev inequality of the proximal Gibbs measure (lem:lsi, cor:uniform-lsi). Of the standard facts about log-Sobolev inequalities, tensorization, Holley–Stroock, the transport by proper \(C^1\) Lipschitz maps and the KL form (L1) are proved in the chapter “Material for Mathlib”; the Gaussian log-Sobolev inequality and the well-posedness of the dynamics are stated and left as sorry. The transport by a \(C^1\) map is what the shear of lem:lsi needs, so lem:lsi(iv) assumes that the analysis maps \(S^*r\) are \(C^1\); this is proved for a smooth feature with bounded derivatives by differentiation under the integral sign.
8.1 Entropy sandwich
For every \(\rho \) with \(\int |a|\, \mathrm d\rho {\lt}\infty \) (in particular for \(\rho \in {\mathcal D}\)), \(\hat\mu _\rho \in {\mathcal D}\).
For \(\rho \) with \(\int |a|\, \mathrm d\rho {\lt}\infty \) and every \(\rho '\in {\mathcal D}\), \(\operatorname {KL}(\rho '\| \hat\mu _\rho ){\lt}\infty \) and
(The chain rule \(\operatorname {KL}(\rho '\| \mu _U)=\operatorname {KL}(\rho '\| \hat\mu _\rho )-\frac1\beta \int a\, s_\rho \, \mathrm d\rho ' +\log \frac{Z_U}{Z_\rho }\), legitimate since \(\int |a s_\rho |\, \mathrm d\rho '\le c_\rho \int |a|\, \mathrm d\rho '{\lt}\infty \).)
Assume (A3), (A4), (A5), \(f\in L^2(P_X)\), and let \(\rho ^*\) be the minimizer of \({\mathcal F}\) on \({\mathcal D}\). For \(\rho \in {\mathcal D}\), \(\operatorname {KL}(\rho \| \hat\mu _\rho ){\lt}\infty \) and
The left inequality is the strong convexity identity (Lemma 212). For the right one, write the decomposition of Lemma 496 for \(\rho '=\rho \) and \(\rho '=\rho ^*\); by the expansion of Lemma 204 around \(\rho \), \(G_\rho (\rho ^*)-G_\rho (\rho )=\frac12\| F_{\rho ^*}-F_\rho \| ^2\ge 0\), so \({\mathcal F}(\rho )-{\mathcal F}(\rho ^*)=\beta \operatorname {KL}(\rho \| \hat\mu _\rho )-\beta \operatorname {KL}(\rho ^*\| \hat\mu _\rho ) -\frac12\| F_{\rho ^*}-F_\rho \| ^2\le \beta \operatorname {KL}(\rho \| \hat\mu _\rho )\).
8.2 Log-Sobolev inequality of the proximal Gibbs measures
If \(\mu \) satisfies \(\mathrm{LSI}(\alpha )\) and \(0{\lt}\alpha '\le \alpha \), then \(\mu \) satisfies \(\mathrm{LSI}(\alpha ')\).
(F1, Gross.) For \(m\in {\mathbb R}\) and \(v{\gt}0\), \({\mathcal N}(m,v)\) satisfies \(\mathrm{LSI}(1/v)\).
(F2, tensorization.) If \(\mu _1\) satisfies \(\mathrm{LSI}(\alpha _1)\) and \(\mu _2\) satisfies \(\mathrm{LSI}(\alpha _2)\), then \(\mu _1\otimes \mu _2\) satisfies \(\mathrm{LSI}(\min \{ \alpha _1,\alpha _2\} )\).
(F3, Holley–Stroock.) If \(\mu \) satisfies \(\mathrm{LSI}(\alpha )\), \(\psi \) is bounded measurable and \(\nu :=Z^{-1}e^{-\psi }\mu \), then \(\nu \) satisfies \(\mathrm{LSI}(\alpha e^{-\operatorname {osc}\psi })\), \(\operatorname {osc}\psi :=\sup \psi -\inf \psi \).
(F4, transport.) If \(\mu \) satisfies \(\mathrm{LSI}(\alpha )\) and \(T\) is a proper \(C^1\) map which is Lipschitz with constant \(L_T\), then \(T_\# \mu \) satisfies \(\mathrm{LSI}(\alpha /L_T^2)\).
If \(|s(z)-s(z')|\le \ell \, d(z,z')\) for all \(z,z'\), then \(T_s\) and \(T_s^{-1}\) are Lipschitz with constant \(1+\ell /\lambda \) for the metric \(d((a,z),(a',z'))=\max \{ |a-a'|,d(z,z')\} \) on \(\Theta \).
If \(|s(z)-s(z')|\le \ell \, |z-z'|\) for all \(z,z'\), then \(T_s^{-1}\) is Lipschitz with constant \(1+\ell /\lambda \) for the Euclidean metric of \(\Theta \) (Lemma ??(ii): \(\| DT_s^{-1}\| _{\mathrm{op}}\le \kappa _0(\ell /\lambda )\le 1+\ell /\lambda \)).
Step 1: it suffices to compare the squared norms.
Step 2: the components of ‘T⁻¹x − T⁻¹y‘ are ‘(a − a’) − (s z − s z’)/λ‘ and ‘z − z’‘.
Step 3: with ‘u = |a − a’|‘, ‘w = ‖z − z’‖‘, the first component is bounded by ‘u + k w‘, and ‘(u + k w)² + w² ≤ (1 + k)² (u² + w²)‘.
For \(\rho \in {\mathcal D}\) (indeed for any \(\rho \) with \(\int |a|\, \mathrm d\rho {\lt}\infty \)), \(\hat\mu _\rho (\, \mathrm da\, \mathrm dz)=\hat\nu _\rho (\, \mathrm dz)\, {\mathcal N}(-s_\rho (z)/\lambda ,\beta /\lambda )(\, \mathrm da)\), where \(\hat\nu _\rho \) is the hidden marginal of \(\hat\mu _\rho \).
For \(T_\rho (a,z):=(a+s_\rho (z)/\lambda ,\ z)\), \((T_\rho )_\# \hat\mu _\rho ={\mathcal N}(0,\beta /\lambda )\otimes \hat\nu _\rho \).
By Lemma 506, under \(\hat\mu _\rho \) one has \(z\sim \hat\nu _\rho \) and \(a\mid z\sim {\mathcal N}(-s_\rho (z)/\lambda ,\beta /\lambda )\), so \(a+s_\rho (z)/\lambda \mid z\sim {\mathcal N}(0,\beta /\lambda )\) does not depend on \(z\).
If \(z\mapsto \varphi _z(x)\) is \(C^1\) for every \(x\) and \(Z\) is finite dimensional, then \(x\mapsto D_z\varphi _z(x)\) is measurable for every \(z\) (a limit of difference quotients along a basis).
Step 1: the difference quotients along the basis are measurable in ‘x‘.
Step 2: a linear form is the sum of its values on the basis times the coordinates.
Step 3: the difference quotients converge to the derivative.
If the feature map is smooth with bounded derivatives (Definition 127), \(Z\) is finite dimensional and \(r\in L^1(P_X)\), then \(S^*r\colon z\mapsto \int \varphi _z(x)r(x)P_X(\, \mathrm dx)\) is \(C^1\) with \(D(S^*r)(z)=\int r(x)D_z\varphi _z(x)P_X(\, \mathrm dx)\).
Differentiation under the integral sign: \(|r(x)D_z\varphi _z(x)|\le L|r(x)|\) with \(L=\sup \| \nabla _z\varphi \| \) is an integrable bound, \(x\mapsto D_z\varphi _z(x)\) is measurable (Lemma 509), and the derivative is continuous by dominated convergence.
\(\nu _0={\mathcal N}(0,(\beta /\lambda _z)I)\) satisfies \(\mathrm{LSI}(\lambda _z/\beta )\) by (F1), and \(\hat\nu _\rho =Z^{-1}e^{-\psi }\nu _0\) with \(\psi =-s_\rho ^2/(2\lambda \beta )\), \(-c_\rho ^2/(2\lambda \beta )\le \psi \le 0\); apply (F3).
Under (A1), (A3), (A5), (A8), and assuming that \(S^*r\) is \(C^1\) for every \(r\in L^2(P_X)\) (Lemma 510), \(\hat\mu _\rho \) satisfies \(\mathrm{LSI}(\alpha _\rho )\) with
where \(\ell _\rho =\sup |\nabla s_\rho |\le \sqrt{R_X^2+1}\, c_\rho \). The right-hand side does not depend on the dimension \(m\).
\({\mathcal N}(0,\beta /\lambda )\) satisfies \(\mathrm{LSI}(\lambda /\beta )\) (F1), so by (F2) and Lemma 511 the product \(\pi _\rho ={\mathcal N}(0,\beta /\lambda )\otimes \hat\nu _\rho \) satisfies \(\mathrm{LSI}(\min \{ \lambda /\beta ,\hat\alpha _\rho \} )\). By Lemma 508, \(\hat\mu _\rho =(T_\rho ^{-1})_\# \pi _\rho \), and \(T_\rho ^{-1}\) is Lipschitz with constant \(1+\ell _\rho /\lambda \) (Lemma 505, with \(\ell _\rho \le \sqrt{R_X^2+1}\, c_\rho \) by Lemma 191); it is \(C^1\) since \(s_\rho \) is, and proper since it is a homeomorphism (its inverse is \(T_\rho \)); conclude by (F4).
For \(\rho \in {\mathcal D}\), \(c_\rho ^2=2L(\rho )\le 2\bigl({\mathcal F}(\rho )+\beta \log Z_U\bigr)\) (here \(c_\rho ^2\le 2{\mathcal F}(\rho )\), the constant of the free energy being dropped).
The map \(c\mapsto \alpha (c)\) of Definition 125 is nonincreasing on \([0,\infty )\).
For \(E\in {\mathbb R}\), on the sublevel set \({\mathcal D}_E:=\{ \rho \in {\mathcal D}:{\mathcal F}(\rho )\le E\} \) and with \(c_E:=\sqrt{2(E+\beta \log Z_U)}\),
i.e. \(\hat\mu _\rho \) satisfies \(\mathrm{LSI}(\alpha _*(E))\) for every \(\rho \in {\mathcal D}_E\).
For \(\rho _0=\mu _U={\mathcal N}(0,\beta /\lambda )\otimes \nu _0\), \(F_{\mu _U}=0\), \(L(\mu _U)=\| f\| ^2/2\) and \({\mathcal F}(\mu _U)=\| f\| ^2/2-\beta \log Z_U\) (here \(\| f\| ^2/2\)).
Step 1: ‘F_μ_U(x) = (∫ a d𝒩(0, β/λ)) (∫ φ_z(x) dν₀) = 0‘.
Step 2: ‘L(μ_U) = ½ ∫ f² dP = ½ ‖f‖²‘.
Step 3: ‘KL(μ_U‖μ_U) = 0‘.
For \(\rho _0=\mu _U\), \(c_{{\mathcal F}(\mu _U)}=\| f\| \) and
Since \(\rho ^*\in {\mathcal D}_{{\mathcal F}(\mu _U)}\), \(\rho ^*=\hat\mu _{\rho ^*}\) satisfies \(\mathrm{LSI}(\alpha _*)\).
8.3 Exponential convergence of the mean-field Langevin dynamics
(L1.) If \(\mu \) satisfies \(\mathrm{LSI}(\alpha )\) on a finite-dimensional space, \(\rho \ll \mu \) is a probability measure with \(\operatorname {KL}(\rho \| \mu ){\lt}\infty \), \(\log \frac{\, \mathrm d\rho }{\, \mathrm d\mu }\) is \(C^1\) and \(\nabla \log \frac{\, \mathrm d\rho }{\, \mathrm d\mu }\in L^2(\rho )\), then \(\operatorname {KL}(\rho \| \mu )\le \frac1{2\alpha }I(\rho |\mu )\) (substitute \(g=\sqrt{\, \mathrm d\rho /\, \mathrm d\mu }\), cut off by compactly supported functions, in the LSI).
If a probability measure \(\mu \) on \(\Theta \) (\(Z\) finite dimensional) satisfies \(\mathrm{LSI}(\alpha )\), then it satisfies the KL form of \(\mathrm{LSI}(\alpha )\).
Along an MFLD flow, \(t\mapsto {\mathcal F}(\rho _t)\) is nonincreasing on \([0,\infty )\).
Assume (A3), (A5), \(f\in L^2(P_X)\), let \((\rho _t)_{t\ge 0}\) be an MFLD flow (Hypothesis W) and assume that \(\hat\mu _{\rho _t}\) satisfies the KL form of \(\mathrm{LSI}(\alpha _*)\) for every \(t\ge 0\). Then for all \(t\ge 0\)
Put \(\Phi (t):={\mathcal F}(\rho _t)-{\mathcal F}(\rho ^*)\ge 0\). By (W2), \(\Phi \) is nonincreasing. For a.e. \(u\ge 0\), the right inequality of Lemma 497 and (L1) give \(\Phi (u)\le \beta \operatorname {KL}(\rho _u\| \hat\mu _{\rho _u}) \le \frac{\beta }{2\alpha _*}I(\rho _u|\hat\mu _{\rho _u})\). Hence, for \(0\le s\le t\), \(\Phi (t)=\Phi (s)-\beta ^2\int _s^tI(\rho _u|\hat\mu _{\rho _u})\, \mathrm du \le \Phi (s)-2\alpha _*\beta \int _s^t\Phi (u)\, \mathrm du\le \Phi (s)-2\alpha _*\beta (t-s)\Phi (t)\), and Lemma 909 (Grönwall) concludes.
Assume (A1), (A3), (A5), (A8), \(f\in L^2(P_X)\) and (A9): \(\rho _0\in {\mathcal D}\), \(\int |z|^2\, \mathrm d\rho _0{\lt}\infty \) and \(\int e^{c_0|a|}\, \mathrm d\rho _0{\lt}\infty \) for some \(c_0{\gt}0\). Then the MFLD has a solution on \([0,\infty )\) (unique in law), and its flow of laws \((\rho _t)_{t\ge 0}\) satisfies (W1) and (W2), i.e. it is an MFLD flow with \(\rho _0\) as initial law.
Assume (A1), (A3), (A5), (A8), \(f\in L^2(P_X)\) and (A9). Let \((\rho _t)_{t\ge 0}\) be the flow of laws of the solution of Theorem 525 and \(\alpha _*:=\alpha _*({\mathcal F}(\rho _0))\). Then for all \(t\ge 0\)
Theorem 525 gives the MFLD flow; since \({\mathcal F}(\rho _t)\le {\mathcal F}(\rho _0)\) (Lemma 521), Corollary 515 gives \(\mathrm{LSI}(\alpha _*)\) for every \(\hat\mu _{\rho _t}\) (the analysis maps are \(C^1\) by Lemma 510, the feature being smooth with bounded derivatives), hence its KL form (Lemma 520), and Theorems 522, 523, 524 apply.
8.4 Static chaos of the finite-particle Gibbs measure
The \(M\)-particle free energy decomposes exactly relative to the product of the mean-field minimizer, so that the relative entropy of the \(M\)-particle Gibbs measure with respect to \(\rho ^{*\otimes M}\) is bounded uniformly in \(M\); this closes the order “\(t\to \infty \) at fixed \(M\) first” of the order-of-limits theorem under the ergodicity of the finite-particle dynamics (manuscript note notes-14).
Let \(\mu \) be a nonzero \(\sigma \)-finite measure on \(\Theta \), \(\beta {\gt}0\) and \(W\) measurable with \(Z_W=\int e^{-W/\beta }\, \mathrm d\mu {\lt}\infty \). Then \(\hat\mu _W^{\otimes M}=\widehat{(\mu ^{\otimes M})}_{W_M}\) with \(W_M(\theta )=\sum _{\ell }W(\theta _\ell )\), i.e. \(\frac{\, \mathrm d\hat\mu _W^{\otimes M}}{\, \mathrm d\mu ^{\otimes M}}(\theta ) =\prod _\ell \frac{e^{-W(\theta _\ell )/\beta }}{Z_W}=\frac{e^{-W_M(\theta )/\beta }}{Z_W^M}\), and \(Z_{W_M}=Z_W^M\).
Step 1: ‘e^−W_M/β = ∏ᵢ e^−W(θᵢ)/β‘ and ‘Z_W_M = Z_W^|ι|‘.
Step 2: both sides are ‘μ^⊗ι‘ with the product density (‘lem:pi-withDensity‘).
For every \(\theta \in \Theta ^M\), \(0\le L(\rho _\theta )\le \tfrac 12\bigl(\tfrac 1M\sum _\ell |a_\ell |+\| f\| \bigr)^2 \le \tfrac 1M\sum _\ell a_\ell ^2+\| f\| ^2\).
Let \(\mu \in {\mathcal P}(\Theta ^M)\) with \(\operatorname {KL}(\mu \| \mu _U^{\otimes M}){\lt}\infty \). Then \(\int a_\ell ^2\, \mu (\, \mathrm d\theta ){\lt}\infty \) for every \(\ell \), hence \(\int |a_\ell |\, \mathrm d\mu {\lt}\infty \) and \(L(\rho _\theta )\), \(\sum _\ell a_\ell s(z_\ell )\) (\(s\) bounded measurable) are \(\mu \)-integrable. (Donsker–Varadhan with \(g(\theta )=\frac{\lambda }{4\beta }a_\ell ^2\), \(\int e^g\, \mathrm d\mu _U^{\otimes M}=\int e^{\lambda a^2/(4\beta )}\, \mathrm d\mu _U=\sqrt2\).)
Since \(0\le L(\rho _\theta )\), \(Z_M=\int e^{-ML(\rho _\theta )/\beta } \, \mathrm d\mu _U^{\otimes M}\in (0,1]\), so \(\pi _M\) is a probability measure, and \(\operatorname {KL}(\pi _M\| \mu _U^{\otimes M})=-\frac M\beta \int L(\rho _\theta )\, \mathrm d\pi _M-\log Z_M{\lt}\infty \) (the potential \(ML(\rho _\theta )\) is \(\pi _M\)-integrable because \(ue^{-u/\beta }\le \beta \)).
For \(\mu \in {\mathcal P}(\Theta ^M)\) with \(\operatorname {KL}(\mu \| \mu _U^{\otimes M}){\lt}\infty \), \(\operatorname {KL}(\mu \| \pi _M){\lt}\infty \) and \({\mathcal F}^M(\mu )=\beta \, \operatorname {KL}(\mu \| \pi _M)-\beta \log Z_M\), \(Z_M:=\int e^{-ML(\rho _\theta )/\beta }\, \mathrm d\mu _U^{\otimes M}\) (Lemma 187 with \(W=ML(\rho _\theta )\) and reference measure \(\mu _U^{\otimes M}\); \(W\in L^1(\mu )\) by Lemma 529).
\(\pi _M\) is the minimizer of \({\mathcal F}^M\) on \(\{ \mu \in {\mathcal P}(\Theta ^M):\operatorname {KL}(\mu \| \mu _U^{\otimes M}){\lt}\infty \} \): \({\mathcal F}^M(\pi _M)=-\beta \log Z_M\le {\mathcal F}^M(\mu )\).
With \(s^*=S^*r_{\rho ^*}\) and \(W^*(\theta )=a\, s^*(z)\), \(\rho ^{*\otimes M}=\widehat{(\mu _U^{\otimes M})}_{W^*_M}\), \(W^*_M(\theta )=\sum _\ell a_\ell s^*(z_\ell )=M\int as^*\, \mathrm d\rho _\theta \), and \(\log \frac{\, \mathrm d\rho ^{*\otimes M}}{\, \mathrm d\mu _U^{\otimes M}}(\theta ) =-\frac1\beta \sum _\ell a_\ell s^*(z_\ell )-M\log Z_*\).
Let \(\rho ^*\) be the minimizer of \({\mathcal F}\) on \({\mathcal D}\). For every \(\mu \in {\mathcal P}(\Theta ^M)\) with \(\operatorname {KL}(\mu \| \mu _U^{\otimes M}){\lt}\infty \), \(\operatorname {KL}(\mu \| \rho ^{*\otimes M}){\lt}\infty \) and
(Chain rule for \(\rho ^{*\otimes M}=\widehat{(\mu _U^{\otimes M})}_{W^*_M}\), the identity \({\mathcal F}(\rho ^*)=L(\rho ^*)-\int as^*\, \mathrm d\rho ^*-\beta \log Z_*\) and the expansion \(L(\rho _\theta )-L(\rho ^*)=\int as^*\, \mathrm d(\rho _\theta -\rho ^*) +\frac12\| F_{\rho _\theta }-F_{\rho ^*}\| ^2\).)
Step 0: notation ‘s* = S* r_ρ*‘, ‘W(θ) = a s*(z)‘, ‘W_M(θ) = ∑_ℓ W(θ_ℓ)‘.
Step 1: ‘Z_W_M = Z_*^M‘.
Step 2: the chain rule ‘KL(μ‖ρ*^⊗M) = KL(μ‖μ_U^⊗M) + (1/β) ∫ W_M dμ + M log Z_*‘.
Step 3: ‘ℱ(ρ*) = L(ρ*) − ∫ W dρ* − β log Z_*‘.
Step 4: the pointwise expansion ‘M L(ρ_θ) − W_M(θ) − M (L(ρ*) − ∫ W dρ*) = (M/2) ‖F_ρ_θ − F_ρ*‖²‘.
Step 5: integrate the pointwise identity against ‘μ‘.
Let \(\rho \in {\mathcal P}(\Theta )\) with \(\int a^2\, \mathrm d\rho {\lt}\infty \) and \(M\ge 1\). Then
(For each \(x\), \(F_{\rho _\theta }(x)=\frac1M\sum _\ell a_\ell \varphi _{z_\ell }(x)\) is a mean of i.i.d. variables of mean \(F_\rho (x)\); Fubini is justified by \(\int a^2\, \mathrm d\rho {\lt}\infty \).)
Step A: ‘‖F_ρ_θ − F_ρ‖² = ∫ (F_ρ_θ(x) − F_ρ(x))² P(dx)‘ for every ‘θ‘.
Step B: the integrand is dominated by ‘2 (1/M) ∑_ℓ a_ℓ² + 2 (∫ |a| dρ)²‘, so Fubini applies.
Step C: for each ‘x‘, the variance of the mean of ‘M‘ i.i.d. copies of ‘a φ_z(x)‘ is ‘(1/M) Var_ρ(a φ_z(x)) ≤ (1/M) ∫ a² dρ‘.
Step D: Fubini and integration of the pointwise bound over ‘x‘.
\(\operatorname {KL}(\rho ^{*\otimes M}\| \mu _U^{\otimes M})=M\, \operatorname {KL}(\rho ^*\| \mu _U){\lt}\infty \).
Let \(\rho ^*\) be the minimizer of \({\mathcal F}\) on \({\mathcal D}\) and \(M\ge 1\). Then
(Theorem 534 at \(\mu =\pi _M\) and at \(\mu =\rho ^{*\otimes M}\), the minimality of \(\pi _M\) for \({\mathcal F}^M\), and Lemma 535.)
Step 1: the identity at ‘μ = π_M‘ and at ‘μ = ρ*^⊗M‘.
Step 2: minimality of ‘π_M‘ and the i.i.d. variance bound.
Under the hypotheses of Theorem 537, \(\operatorname {KL}(\pi _M\| \rho ^{*\otimes M})\le \frac1{2\beta }\int a^2\, \mathrm d\rho ^*\) (independent of \(M\)) and \({\mathbb E}_{\pi _M}\| F_{\rho _\theta }-F_{\rho ^*}\| ^2\le \frac1M\int a^2\, \mathrm d\rho ^*\).
With \(A:=\min \{ \| f\| ^2/\lambda ^2,\| f\| ^2/(4\lambda )\} +\beta /\lambda \ge \int a^2\, \mathrm d\rho ^*\) (Theorem 233), \(\operatorname {KL}(\pi _M\| \rho ^{*\otimes M})\le A/(2\beta )\) and \({\mathbb E}_{\pi _M}\| F_{\rho _\theta }-F_{\rho ^*}\| ^2\le A/M\).
Let \(s\) be measurable with \(|s|\le C\) and \(\rho _s=\hat\mu _{W_s}\), \(W_s(\theta )=a\, s(z)\). Then \(\int e^{c|a|}\, \mathrm d\rho _s{\lt}\infty \) for every \(c\in {\mathbb R}\), since \(e^{-as(z)/\beta }e^{c|a|}\le e^{(C/\beta +c)|a|}\) and the \(a\)-marginal of \(\mu _U\) is Gaussian.
Under the hypotheses of Lemma 540, for \(c\ge 0\), \(\int e^{c|a|}\, \mathrm d\rho _s\le Z_s^{-1}\int e^{(c+C/\beta )|a|}\, \mathrm d\mu _U \le 2\exp \bigl(\tfrac {\beta }{2\lambda }(c+C/\beta )^2\bigr)\), using \(Z_s\ge 1\) and the Gaussian moment generating function.
Step 1: ‘∫ e^c|a| dρ_s = Z⁻¹ ∫ e^−a s/β e^c|a| dμ_U ≤ Z⁻¹ ∫ e^t|a| dμ_U‘.
Step 2: ‘∫ e^t|a| dμ_U ≤ ∫ (e^ta + e^−ta) dμ_U = 2 e^t² (β/λ)/2‘ (Gaussian moment generating function).
Let \(g\) be measurable with \(|g|\le C\), \(m:=\int ag(z)\, \mathrm d\rho ^*\) and \(Y(\theta ):=ag(z)-m\). Then \(K_g:=\int Y^2e^{|Y|}\, \mathrm d\rho ^*{\lt}\infty \), since \(|Y|\le C|a|+|m|\), \(u^2\le 2e^u\) and \(\int e^{2C|a|}\, \mathrm d\rho ^*{\lt}\infty \).
Under the hypotheses of Lemma 542, \(K_g\le 4\exp \bigl(2C(\| f\| /\lambda +\sqrt{2\beta /(\pi \lambda )})\bigr) \exp \bigl(\tfrac \beta {2\lambda }(2C+\| f\| /\beta )^2\bigr)\), a constant depending only on \(C,\| f\| ,\lambda ,\beta \).
Step 1: ‘|m| ≤ C ∫ |a| dρ* ≤ C (‖f‖/λ + √(2β/(πλ)))‘.
Step 2: the pointwise bound ‘Y² e^|Y| ≤ 2 e^2|m| e^2C|a|‘.
Step 3: integrate and use the Gaussian exponential-moment bound.
Let \(g\colon Z\to {\mathbb R}\) be measurable with \(|g|\le C\), \(m:=\int ag(z)\, \mathrm d\rho ^* =\int g\, \mathrm d\Pi \rho ^*\), \(\bar X(\theta ):=\frac1M\sum _\ell a_\ell g(z_\ell )=\int g\, \mathrm d\Pi \rho _\theta \) and \(K_g:=\int (ag(z)-m)^2e^{|ag(z)-m|}\, \mathrm d\rho ^*{\lt}\infty \). Then
(Donsker–Varadhan on \(\Theta ^M\) with \(h=\sqrt M|\bar X-m|\): with \(u=1/\sqrt M\) and \(Y_\ell =a_\ell g(z_\ell )-m\) i.i.d. centered, \({\mathbb E}_{\rho ^{*\otimes M}}e^{\pm u\sum _\ell Y_\ell }=({\mathbb E}e^{\pm uY})^M\le (1+K_gu^2)^M\le e^{K_g}\), so \(\log {\mathbb E}_{\rho ^{*\otimes M}}e^h\le \log 2+K_g\).)
Step 0: notation. ‘Y(p) = a g(z) − m‘ is centered with ‘K = ∫ Y² e^|Y| dρ* < ∞‘; ‘S(θ) = ∑_ℓ Y(θ_ℓ)‘, ‘u = 1/√M‘, ‘h(θ) = u |S(θ)| = √M |X̄(θ) − m|‘.
Step 1: ‘X̄(θ) − m = (1/M) S(θ)‘.
Step 2: the exponential moments of ‘± u S‘ under ‘ρ*^⊗M‘ are at most ‘e^K‘.
Step 3: ‘e^u|S| ≤ e^uS + e^−uS‘, so ‘∫ e^u|S| dρ*^⊗M ≤ 2 e^K‘.
Step 4: Donsker–Varadhan on ‘Θ^M‘: ‘∫ u|S| dπ_M ≤ KL(π_M‖ρ*^⊗M) + log 2 + K‘.
Step 5: ‘|X̄ − m| = u · (u |S|)‘ since ‘u² = 1/M‘; conclude.
Under the hypotheses of Proposition 544, with \(A=\min \{ \| f\| ^2/\lambda ^2,\| f\| ^2/(4\lambda )\} +\beta /\lambda \), \({\mathbb E}_{\pi _M}|\bar X-m|\le C_g/\sqrt M\) with \(C_g=A/(2\beta )+\log 2+K_g\).
Under the hypotheses of Proposition 544, \({\mathbb E}_{\pi _M}|\bar X-m|\le C_g/\sqrt M\) with the explicit constant \(C_g=\frac A{2\beta }+\log 2+4\exp \bigl(2C(\| f\| /\lambda +\sqrt{2\beta /(\pi \lambda )})\bigr) \exp \bigl(\tfrac \beta {2\lambda }(2C+\| f\| /\beta )^2\bigr)\), which depends only on \(C,\| f\| ,\lambda ,\beta \).
For every permutation \(\sigma \) of \(\{ 1,\dots ,M\} \), the image of \(\pi _M\) under \(\theta \mapsto \theta \circ \sigma \) is \(\pi _M\): \(\mu _U^{\otimes M}\) is permutation invariant and \(L(\rho _{\theta \circ \sigma })=L(\rho _\theta )\). In particular all one-particle marginals of \(\pi _M\) coincide.
Let \(Z\) be standard Borel. For every \(\ell \), the one-particle marginal \(\pi _M^{(\ell )}\) of \(\pi _M\) satisfies \(M\, \operatorname {KL}(\pi _M^{(\ell )}\| \rho ^*)=\sum _{\ell '}\operatorname {KL}(\pi _M^{(\ell ')}\| \rho ^*) \le \operatorname {KL}(\pi _M\| \rho ^{*\otimes M})\), by exchangeability and the superadditivity \(\operatorname {KL}(\mu \| \nu ^{\otimes M})\ge \sum _\ell \operatorname {KL}(\mu ^{(\ell )}\| \nu )\).
Under the hypotheses of Theorem 537 with \(Z\) standard Borel, \(\operatorname {KL}(\pi _M^{(\ell )}\| \rho ^*)\le \frac{A}{2\beta M}\) for every \(\ell \).
Assume (E): \(\| \mu _t-\pi _M\| _{{\mathrm{TV}}}\to 0\) and \(\sup _t\int \frac1M\sum _\ell a_\ell ^2\, \mathrm d\mu _t\le Q{\lt}\infty \). Then for every bounded measurable \(g\), with \(\bar X(\theta )=\int g\, \mathrm d\Pi \rho _\theta \), \(m=\int g\, \mathrm d\Pi \rho ^*\) and \(C_g\) as in Proposition 544,
hence \(\limsup _{M\to \infty }\limsup _{t\to \infty }{\mathbb E}\bigl|\int g\, \mathrm d\Pi \rho ^{M}_t -\int g\, \mathrm d\Pi \rho ^*\bigr|=0\). (Truncation: \(h\wedge T\) converges by total variation, and \({\mathbb E}_{\mu _t}[h-h\wedge T]\le {\mathbb E}_{\mu _t}h^2/T\le (2C^2Q+2m^2)/T\) uniformly in \(t\).)
Step 0: notation: ‘m‘, the statistic ‘h(θ) = |X̄(θ) − m|‘, its bound ‘C S(θ) + |m|‘ with ‘S(θ) = (1/M) ∑_ℓ |a_ℓ|‘, the bound ‘B‘ of Proposition 14.4 and the second-moment bound ‘Q₂ = 2 C² Q + 2 m²‘ of ‘h‘.
Step 1: the truncation level ‘T‘ with ‘Q₂/T ≤ ε/2‘.
Step 2: integrability of ‘h‘, ‘h²‘ and ‘h ∧ T‘ under each ‘μ_t‘ and under ‘π_M‘.
Step 3: the tail bound ‘∫ h dμ_t ≤ ∫ h ∧ T dμ_t + Q₂/T‘ (Chebyshev: ‘h − h ∧ T ≤ h²/T‘).
Step 4: total-variation convergence for the truncated statistic, and ‘∫ h ∧ T dπ_M ≤ B‘.
Step 5: eventually ‘2 T ‖μ_t − π_M‖_TV ≤ ε/2‘.
For the tanh feature (A3) with \(z=(w,b)\in {\mathbb R}^{m+1}\) (Euclidean norm), \(\| \nabla _z\varphi _z(x)\| \le \sqrt{|x|^2+1}\) for all \(x\in {\mathbb R}^m\) and \(z\in {\mathbb R}^{m+1}\). Indeed \(\nabla _z\varphi _z(x)=\tanh '(w^\top x-b)\, (x,-1)\) with \(0{\lt}\tanh '\le 1\) and \(|(x,-1)|=\sqrt{|x|^2+1}\). In particular, under (A1) the feature is Lipschitz in the hidden parameter with constant \(\sqrt{R_X^2+1}\).
Step 1: the chain rule bound ‘‖∇_z φ_z(x)‖ ≤ |tanh’(⟪w, x⟫ − b)| · ‖(x, −1)‖‘.
Step 2: ‘|tanh’| = 1 − tanh² ≤ 1‘ and ‘‖(x, −1)‖ = √(‖x‖² + 1)‘.
For the tanh feature (A3), with \(z=(w,b)\in {\mathbb R}^{m+1}\) (Euclidean norm, Definition 6), and inputs \(|x|\le R_X\) (A1), the map \(z\mapsto \varphi _z(x)\) is \(C^\infty \) and, for every \(n\ge 0\),
where \(\| \tanh ^{(n)}\| _\infty \le \sup _{|s|\le 1}|p_n(s)|\) for the polynomials \(p_0=X\), \(p_{n+1}=p_n'\, (1-X^2)\) with \(\tanh ^{(n)}=p_n(\tanh )\). Hence the restriction of the tanh feature to \(\{ |x|\le R_X\} \) (Definition 3) is smooth with bounded derivatives in the sense of Definition 127.
Step 1: smoothness. ‘z ↦ tanh(⟪w, x⟫ − b)‘ is the composition of the ‘C^∞‘ function ‘tanh‘ with the continuous linear map ‘(w, b) ↦ ⟪(x, −1), (w, b)⟫‘.
Step 2: the ‘n‘-th derivative of ‘tanh‘ is bounded by a constant ‘c_n‘ (‘tanh⁽ⁿ⁾ = p_n(tanh)‘ with ‘|tanh| ≤ 1‘).
Step 3: the chain rule for the linear map ‘(w, b) ↦ ⟪(x, −1), (w, b)⟫‘ of norm ‘√(‖x‖² + 1) ≤ √(R_X² + 1)‘ gives ‘‖∇_z^n φ_z(x)‖ ≤ c_n (R_X² + 1)^n/2‘.
8.5 The third stage of the order of the limits: \(t\to \infty \) at fixed \(M\) first
If \(\int |a|\, \mathrm d\rho {\lt}\infty \) and \(g\) is bounded measurable, then \(\int _Zg\, \mathrm d\Pi \rho =\int _\Theta a\, g(z)\, \rho (\, \mathrm da\, \mathrm dz)\).
Fix the sample \((x_i,y_i)_{i=1}^N\) with \(|y_i|\le {Y_{\max }}\), \((\lambda ,\beta )\) and \(M\ge 1\); let \(\rho _N^*\) be the minimizer of \({\mathcal F}_N\) and let \((\mu _t)_{t\ge 0}\) be ergodic to the empirical \(M\)-particle Gibbs measure \(\pi _M\) (Assumption (E)). For every bounded measurable \(g\) with \(\| g\| _\infty \le C\) and every \({\varepsilon }{\gt}0\) there is \(t_{\varepsilon }\) such that for \(t\ge t_{\varepsilon }\)
with the constant of Definition 135, which depends on the sample only through \({Y_{\max }}\). (Theorem 550 for the empirical feature, \(\| y\| _N\le {Y_{\max }}\), and \(\int g\, \mathrm d\Pi \rho _N^*=\int a\, g(z)\, \mathrm d\rho _N^*\).)
Step 1: the dictionary: ‘ρ_N^*‘ minimizes the free energy of the empirical feature with target ‘y ∈ L²(uniform)‘, and ‘‖y‖_N ≤ Ymax‘.
Step 2: ‘thm:static-chaos-order‘ for the empirical feature.
Step 3: the constant is bounded uniformly in the sample.
Assume the hypotheses of Theorem 427 and let, for every sample and every \(M\ge 1\), \((\mu ^{M,N}_t)_{t\ge 0}\) be a family of laws on \(\Theta ^M\) ergodic to the empirical \(M\)-particle Gibbs measure (Assumption (E); for the \(M\)-particle empirical noisy gradient descent this is the ergodic theorem). Let \(g\) be bounded measurable, \({\varepsilon }{\gt}0\) and \(\delta \in (0,1)\). Then for the \((\lambda ,\beta )\) and \(N_1\) of Theorem 427, on the event of probability at least \(1-\delta \) over the sample, for every \(M\ge 1\)
in the form: for every \({\varepsilon }'{\gt}0\), eventually in \(t\), \({\mathbb E}_{\mu _t}|\int g\, \mathrm d\Pi \rho _\theta -\int g\, u_0^\dagger \, \mathrm d\nu _0| \le \frac{2{\varepsilon }}3+\frac{C_g({Y_{\max }},\lambda ,\beta )}{\sqrt M}+{\varepsilon }'\), with the constant of Definition 135. (The triangle inequality ?? at fixed \((M,t)\), integrated over \(\mu _t\): the second and third terms are deterministic on the event of stage (2) and bounded by \({\varepsilon }/3\) each, and the first is Theorem 554.)
Stage (1) at ‘(λ, β)‘.
Stage (2) at ‘(λ, β)‘, on the event ‘E’_δ‘.
Stage (3) for the sample ‘ω‘ (labels bounded by ‘Ymax‘ on ‘E’_δ‘).
The triangle inequality ‘eq:three-stages‘, integrated over ‘μ_t‘.
Under the hypotheses of Corollary 555, since \(C_g({Y_{\max }},\lambda ,\beta )\) does not depend on the sample, there is \(M_0=M_0({\varepsilon };{Y_{\max }},\lambda ,\beta ,g)\) such that on the same event of probability at least \(1-\delta \), for every \(M\ge M_0\), eventually in \(t\), \({\mathbb E}_{\mu _t}\bigl|\int g\, \mathrm d\Pi \rho _\theta -\int g\, u_0^\dagger \, \mathrm d\nu _0\bigr|\le {\varepsilon }\); that is,
(\(M_0:=\max \{ \lceil (6C_g/{\varepsilon })^2\rceil ,1\} \) gives \(C_g/\sqrt M\le {\varepsilon }/6\); take \({\varepsilon }'={\varepsilon }/6\).)