Shallow Learning Tends to Ridgelet Transform

3 Basic lemmas of the setting

The lemmas of Section 2 of the paper: finiteness of the amplitude moments on the domain of the free energy (lem:moment), the Gibbs variational principle relative to a reference measure (lem:gibbs-variational), and the properties of the synthesis operator, its adjoint, the kernel operator, the resolvent and the regularized ridgelet transform (lem:operators).

3.1 Second moments on the domain

Theorem 166
✓

For \(\nu _0\in {\mathcal P}(Z)\) and \(f\colon {\mathbb R}\to {\mathbb R}\) measurable, \(\int _\Theta f(a)\, \mu _U(\, \mathrm da\, \mathrm dz)=\int _{{\mathbb R}} f(a)\, {\mathcal N}(0,\beta /\lambda )(\, \mathrm da)\).

Proof ▶

Let \(\lambda ,\beta {\gt}0\), \(\nu _0\in {\mathcal P}(Z)\) and \(2s\beta /\lambda {\lt}1\). Then \(a\mapsto e^{sa^2}\) is \(\mu _U\)-integrable and \(\int _\Theta e^{sa^2}\, \mu _U(\, \mathrm d\theta )=(1-2s\beta /\lambda )^{-1/2}\).

Proof ▶

Let \(\lambda ,\beta {\gt}0\), \(\nu _0\in {\mathcal P}(Z)\), and let \(\rho \in {\mathcal D}\), i.e. \(\rho \in {\mathcal P}(\Theta )\) with \(\operatorname {KL}(\rho \| \mu _U){\lt}\infty \). For every \(\eta \in (0,1)\),

\[ \int _\Theta a^2\, \rho (\, \mathrm da\, \mathrm dz)\le \frac{2\beta }{\eta \lambda } \Bigl[\operatorname {KL}(\rho \| \mu _U)+\tfrac 12\log \tfrac 1{1-\eta }\Bigr]. \]

Only \(|\varphi _z|\le 1\) of (A3) and \(\nu _0\in {\mathcal P}(Z)\) of (A4), (A5) enter.

Proof ▶

Step 1: the test function ‘g(θ) = c a²‘ with ‘c = ηλ/(2β)‘, so that ‘2 c β/λ = η < 1‘.

Step 2: the Gaussian exponential moment ‘∫ e^g dμ_U = (1-η)^-1/2‘.

Step 3: ‘g ∈ L¹(ρ)‘ (Lebesgue-integral Donsker–Varadhan), hence ‘a² ∈ L¹(ρ)‘.

Step 4: the Donsker–Varadhan inequality ‘c ∫ a² dρ ≤ KL + log (1-η)^-1/2‘.

Step 5: ‘log (1-η)^-1/2 = ½ log (1/(1-η))‘, and divide by ‘c‘.

Under the hypotheses of Lemma 168, taking \(\eta =1/2\), \(\int _\Theta a^2\, \mathrm d\rho \le \frac{4\beta }{\lambda } \bigl[\operatorname {KL}(\rho \| \mu _U)+\tfrac 12\log 2\bigr]\).

Proof ▶

Under the hypotheses of Lemma 168, \(\int _\Theta a^2\, \mathrm d\rho {\lt}\infty \), i.e. \(a\mapsto a^2\) is \(\rho \)-integrable.

Proof ▶

The test function ‘g(θ) = (λ/(4β)) a²‘ of the case ‘η = 1/2‘.

Theorem 171
✓

If \(\rho \) is a finite measure on \(\Theta \) with \(\int a^2\, \mathrm d\rho {\lt}\infty \), then \(\int |a|\, \mathrm d\rho {\lt}\infty \) (since \(|a|\le a^2+1\)).

Proof ▶

Under the hypotheses of Lemma 168, \(\int |a|\, \mathrm d\rho {\lt}\infty \) and \(\int a^2\, \mathrm d\rho {\lt}\infty \).

Proof ▶
Theorem 173
✓

For \(\rho \in {\mathcal P}(\Theta )\) with \(\int a^2\, \mathrm d\rho {\lt}\infty \), \(\int |a|\, \mathrm d\rho \le \bigl(\int a^2\, \mathrm d\rho \bigr)^{1/2}\) (Cauchy–Schwarz).

Proof ▶

‘Var[|a|] = ∫ a² − (∫ |a|)² ≥ 0‘.

Under the hypotheses of Lemma 168, \(\int |a|\, \mathrm d\rho \le \bigl(\int a^2\, \mathrm d\rho \bigr)^{1/2}{\lt}\infty \).

Proof ▶

Under the hypotheses of Lemma 168, for every \(x\in {\mathcal X}\), \(|F_\rho (x)|\le \int |a|\, \mathrm d\rho \le \bigl(\int a^2\, \mathrm d\rho \bigr)^{1/2}\).

Proof ▶

3.2 Gibbs variational principle

Theorem 176
✓
#

If \(\int e^{-W/\beta }\, \mathrm d\mu {\lt}\infty \) then \(\hat\mu _W=Z_W^{-1}e^{-W/\beta }\mu \) is the exponential tilt of \(\mu \) by \(-W/\beta \) (Mathlib’s ‘Measure.tilted‘), i.e. the measure with density \(e^{-W/\beta }/\int e^{-W/\beta }\, \mathrm d\mu \) with respect to \(\mu \).

Proof ▶
Theorem 177
✓

If \(\int e^{-W/\beta }\, \mathrm d\mu {\lt}\infty \) then \(\hat\mu _W(\, \mathrm d\theta )=\dfrac {e^{-W(\theta )/\beta }}{\int e^{-W/\beta }\, \mathrm d\mu }\, \mu (\, \mathrm d\theta )\).

Proof ▶
Theorem 178
✓

If \(\mu \ne 0\) and \(Z_W=\int e^{-W/\beta }\, \mathrm d\mu {\lt}\infty \) then \(\hat\mu _W\) is a probability measure.

Proof ▶
Theorem 179
✓

If \(Z_W{\lt}\infty \) then \(\hat\mu _W\ll \mu \).

Proof ▶
Theorem 180
✓

If \(Z_W{\lt}\infty \) then \(\mu \ll \hat\mu _W\), so \(\hat\mu _W\) and \(\mu \) are equivalent.

Proof ▶
Theorem 181
✓

Let \(\mu \) be \(\sigma \)-finite, \(\beta {\gt}0\), \(W\) measurable with \(Z_W{\lt}\infty \), and let \(\rho \in {\mathcal P}(\Theta )\) with \(\rho \ll \mu \) and \(\int |W|\, \mathrm d\rho {\lt}\infty \). Then \(\operatorname {KL}(\rho \| \hat\mu _W){\lt}\infty \) if and only if \(\operatorname {Ent}_\mu (\rho )=\int \log \frac{\, \mathrm d\rho }{\, \mathrm d\mu }\, \mathrm d\rho \) is finite, i.e. \(\log \frac{\, \mathrm d\rho }{\, \mathrm d\mu }\in L^1(\rho )\).

Proof ▶

Under the hypotheses of Lemma 181, if \(\operatorname {Ent}_\mu (\rho )\) is finite then \(\operatorname {KL}(\rho \| \hat\mu _W)=\operatorname {Ent}_\mu (\rho )+\frac1\beta \int W\, \mathrm d\rho +\log Z_W\).

Proof ▶

‘KL(ρ‖μ̂_W) = ∫ llr ρ μ̂_W dρ‘ for probability measures, and Mathlib’s computation of the log-likelihood ratio with respect to a tilted measure.

Theorem 183
✓

Let \(\mu \) be a \(\sigma \)-finite measure on \(\Theta \), \(\beta {\gt}0\), \(W\colon \Theta \to {\mathbb R}\) measurable with \(Z_W:=\int _\Theta e^{-W/\beta }\, \mathrm d\mu {\lt}\infty \), and \(\hat\mu _W(\, \mathrm d\theta ):=Z_W^{-1}e^{-W(\theta )/\beta }\mu (\, \mathrm d\theta )\). If \(\rho \in {\mathcal P}(\Theta )\), \(\rho \ll \mu \), \(\int |W|\, \mathrm d\rho {\lt}\infty \) and \(\operatorname {Ent}_\mu (\rho )=\int \log \frac{\, \mathrm d\rho }{\, \mathrm d\mu }\, \mathrm d\rho \) is finite, then

\[ \int W\, \mathrm d\rho +\beta \, \operatorname {Ent}_\mu (\rho )=\beta \, \operatorname {KL}(\rho \| \hat\mu _W)-\beta \log Z_W . \]

(The paper takes \(\mu \) to be Lebesgue measure; the case \(\operatorname {Ent}_\mu (\rho )=+\infty \) is Lemma 181.)

Proof ▶
Theorem 184
✓

Under the hypotheses of Lemma 183, \(\int W\, \mathrm d\rho +\beta \, \operatorname {Ent}_\mu (\rho )\ge -\beta \log Z_W\), with equality if and only if \(\rho =\hat\mu _W\) (Lemma 185).

Proof ▶

Under the hypotheses of Lemma 183, \(\int W\, \mathrm d\rho +\beta \, \operatorname {Ent}_\mu (\rho )=-\beta \log Z_W\) if and only if \(\rho =\hat\mu _W\).

Proof ▶

Let \(\mu \in {\mathcal P}(\Theta )\), \(\beta {\gt}0\), \(W\) measurable with \(Z_W{\lt}\infty \), and \(\rho \in {\mathcal P}(\Theta )\) with \(\rho \ll \mu \) and \(\int |W|\, \mathrm d\rho {\lt}\infty \). Then \(\operatorname {KL}(\rho \| \hat\mu _W){\lt}\infty \) if and only if \(\operatorname {KL}(\rho \| \mu ){\lt}\infty \).

Proof ▶

Let \(\mu \in {\mathcal P}(\Theta )\), \(\beta {\gt}0\), \(W\) measurable with \(Z_W{\lt}\infty \), and \(\rho \in {\mathcal P}(\Theta )\) with \(\rho \ll \mu \), \(\int |W|\, \mathrm d\rho {\lt}\infty \) and \(\operatorname {KL}(\rho \| \mu ){\lt}\infty \). Then \(\operatorname {KL}(\rho \| \hat\mu _W)=\operatorname {KL}(\rho \| \mu )+\frac1\beta \int W\, \mathrm d\rho +\log Z_W\), i.e. \(\int W\, \mathrm d\rho +\beta \operatorname {KL}(\rho \| \mu )=\beta \operatorname {KL}(\rho \| \hat\mu _W)-\beta \log Z_W\).

Proof ▶

Under the hypotheses of Lemma 187, \(\int W\, \mathrm d\rho +\beta \operatorname {KL}(\rho \| \mu )\ge -\beta \log Z_W\), with equality if and only if \(\rho =\hat\mu _W\).

Proof ▶

3.3 Synthesis operator, adjoint, kernel and regularized ridgelet transform

Theorem 189
✓

For \(r\in L^2(P_X)\) and every \(z\in Z\), \(|S^*r(z)|\le \| \varphi _z\| _{L^2(P_X)}\| r\| _{L^2(P_X)}\le \| r\| _{L^2(P_X)}\).

Proof ▶
Theorem 190
✓

For \(r\in L^2(P_X)\) and every \(\nu \in {\mathcal P}(Z)\), \(S^*r\in L^2(\nu )\): it is measurable (Fubini) and bounded by \(\| r\| _{L^2(P_X)}\).

Proof ▶
Theorem 191
✓

Let \(Z\) be a metric space and assume \(|\varphi _z(x)-\varphi _{z'}(x)|\le L\, d(z,z')\) for all \(x,z,z'\). Then for \(r\in L^2(P_X)\), \(|S^*r(z)-S^*r(z')|\le L\, \| r\| _{L^2(P_X)}\, d(z,z')\); under (A1) and (A3) this is ?? with \(L=\sqrt{R_X^2+1}\).

Proof ▶

Step 1: the difference is the integral of ‘(φ_z − φ_z’) r‘.

Step 2: ‘L · dist z z’ ≥ 0‘ (the input space is nonempty since ‘P‘ is a probability measure).

Step 3: bound the integrand by ‘L · dist z z’ · |r|‘ and use ‘‖r‖_L¹ ≤ ‖r‖_L²‘.

For every \(\nu \in {\mathcal P}(Z)\) and \(r\in L^2(P_X)\) the adjoint of the synthesis operator is the analysis map: \((S_\nu ^*r)(z)=\langle \varphi _z,r\rangle _{L^2(P_X)}=(S^*r)(z)\) for \(\nu \)-a.e. \(z\). In particular \(S_\nu ^*r\) does not depend on \(\nu \) (as a function).

Proof ▶
Theorem 193
✓

\(K_\nu =S_\nu S_\nu ^*\) is a self-adjoint positive semidefinite operator on \(L^2(P_X)\): \(\langle K_\nu r,r\rangle _{L^2(P_X)}=\| S_\nu ^*r\| _{L^2(\nu )}^2\ge 0\) for every \(r\), and \(\| K_\nu \| \le \| S_\nu \| \| S_\nu ^*\| \le 1\).

Proof ▶

For \(r\in L^2(P_X)\), \(P_X\)-a.e. in \(x\), \((K_\nu r)(x)=\int _Z\varphi _z(x)\, (S^*r)(z)\, \nu (\, \mathrm dz) =\int _Z\varphi _z(x)\langle \varphi _z,r\rangle _{L^2(P_X)}\, \nu (\, \mathrm dz)\), i.e. \(K_\nu \) is the integral operator with kernel \(k_\nu (x,x')=\int \varphi _z(x)\varphi _z(x')\, \nu (\, \mathrm dz)\).

Proof ▶

For \(\lambda {\gt}0\) the operator \(K_\nu +\lambda \) is a bijection of \(L^2(P_X)\) with bounded inverse: \((K_\nu +\lambda )(K_\nu +\lambda )^{-1}=\mathrm{id}\) and \((K_\nu +\lambda )^{-1}(K_\nu +\lambda )=\mathrm{id}\), where \((K_\nu +\lambda )^{-1}\) is ‘kernelResolvent‘. (Proof: \(K_\nu \ge 0\) gives \(\langle (K_\nu +\lambda )r,r\rangle \ge \lambda \| r\| ^2\), so \(K_\nu +\lambda \) is injective with closed range whose orthogonal complement is trivial.)

Proof ▶

For \(\lambda {\gt}0\), \(\| (K_\nu +\lambda )^{-1}\| \le 1/\lambda \): for \(g=(K_\nu +\lambda )^{-1}f\), \(\lambda \| g\| ^2\le \langle (K_\nu +\lambda )g,g\rangle =\langle f,g\rangle \le \| f\| \| g\| \).

Proof ▶

For \(\lambda {\gt}0\), \(\| \lambda (K_\nu +\lambda )^{-1}\| \le 1\) and \(\| K_\nu (K_\nu +\lambda )^{-1}\| \le 1\). (Proof without spectral calculus: for \(g=(K_\nu +\lambda )^{-1}f\), \(\| f\| ^2=\| (K_\nu +\lambda )g\| ^2 =\| K_\nu g\| ^2+2\lambda \langle K_\nu g,g\rangle +\lambda ^2\| g\| ^2\) dominates both \(\lambda ^2\| g\| ^2\) and \(\| K_\nu g\| ^2\) since \(\langle K_\nu g,g\rangle \ge 0\).)

Proof ▶

For \(\lambda {\gt}0\) and \(f\in L^2(P_X)\), \(\| R_{\lambda ,\nu }f\| _{L^2(\nu )}\le \| f\| _{L^2(P_X)}/(2\sqrt\lambda )\). (Proof without the spectral measure: with \(g=(K_\nu +\lambda )^{-1}f\), \(\| S_\nu ^*g\| ^2=\langle K_\nu g,g\rangle \) and \(\| f\| ^2-4\lambda \langle K_\nu g,g\rangle =\| (K_\nu -\lambda )g\| ^2\ge 0\).)

Proof ▶

For \(\lambda {\gt}0\), \(f\in L^2(P_X)\) and every \(z\in Z\), \(|R_{\lambda ,\nu }f(z)|\le \| (K_\nu +\lambda )^{-1}f\| _{L^2(P_X)}\le \| f\| _{L^2(P_X)}/\lambda \).

Proof ▶

For \(\lambda {\gt}0\) and \(f\in L^2(P_X)\), \(u=S_\nu ^*(K_\nu +\lambda )^{-1}f\) solves \((S_\nu ^*S_\nu +\lambda )u=S_\nu ^*f\); i.e. the push-through identity \(S_\nu ^*(K_\nu +\lambda )^{-1}f=(S_\nu ^*S_\nu +\lambda )^{-1}S_\nu ^*f\) holds.

Proof ▶

For \(\lambda {\gt}0\) and \(f\in L^2(P_X)\), \(u_0=S_\nu ^*(K_\nu +\lambda )^{-1}f\) is the unique minimizer over \(L^2(\nu )\) of the Tikhonov functional \(J(u)=\frac12\| S_\nu u-f\| _{L^2(P_X)}^2+\frac\lambda 2\| u\| _{L^2(\nu )}^2\): \(J(u_0)\le J(u)\) for all \(u\), with equality only for \(u=u_0\). (Proof: by the normal equation, \(J(u)-J(u_0)=\frac12\| S_\nu (u-u_0)\| ^2+\frac\lambda 2\| u-u_0\| ^2\).)

Proof ▶

3.4 Characterization of the minimizer of the free energy

For \(\rho \) with \(\int |a|\, \mathrm d\rho {\lt}\infty \), \(L(\rho )=\tfrac 12\| r_\rho \| _{L^2(P_X)}^2\) with \(r_\rho =F_\rho -f\in L^2(P_X)\).

Proof ▶

For \(\rho \) with \(\int |a|\, \mathrm d\rho {\lt}\infty \) and \(g\in L^2(P_X)\), \(\int _\Theta a\, (S^*g)(z)\, \rho (\, \mathrm d\theta )=\langle F_\rho ,g\rangle _{L^2(P_X)}\) (Fubini, since \(\iint |a||\varphi _z(x)||g(x)|\, \rho (\, \mathrm d\theta )P_X(\, \mathrm dx) \le \int |a|\, \mathrm d\rho \, \| g\| _{L^1(P_X)}{\lt}\infty \)).

Proof ▶

Step 1: the integrand ‘a φ_z(x) g(x)‘ is integrable on ‘P ⊗ ρ‘, being bounded by ‘|g(x)| |a|‘.

Step 2: write the inner product as an integral and swap the order of integration.

For \(\rho ,\rho '\in {\mathcal D}\) (indeed for any \(\rho ,\rho '\) with \(\int |a|\, \mathrm d\rho ,\int |a|\, \mathrm d\rho '{\lt}\infty \)), with \(s_\rho :=S^*r_\rho \),

\[ L(\rho ')-L(\rho )=\int _\Theta a\, s_\rho (z)\, \, \mathrm d(\rho '-\rho )(\theta ) +\tfrac 12\| F_{\rho '}-F_\rho \| _{L^2(P_X)}^2 . \]

(Proof: \(r_{\rho '}=r_\rho +(F_{\rho '}-F_\rho )\), expand the square, and use Lemma 203.)

Proof ▶

‘r_ρ’ = r_ρ + (F_ρ’ − F_ρ)‘ and ‘‖r + d‖² = ‖r‖² + 2⟪r, d⟫ + ‖d‖²‘.

Under the hypotheses of Lemma 204, \(L(\rho ')\ge L(\rho )+\int a\, s_\rho (z)\, \, \mathrm d(\rho '-\rho )\), with equality if and only if \(F_{\rho '}=F_\rho \) in \(L^2(P_X)\). In particular \(L\) is convex on \({\mathcal D}\).

Proof ▶
Theorem 206
✓

Let \(s\colon Z\to {\mathbb R}\) be measurable with \(|s|\le C\) and \(W(\theta ):=a\, s(z)\). Then \(\int _\Theta e^{-W/\beta }\, \mathrm d\mu _U{\lt}\infty \), since \(e^{-as(z)/\beta }\le e^{C|a|/\beta }\) and the \(a\)-marginal of \(\mu _U\) is Gaussian.

Proof ▶

Let \(s\) be measurable with \(|s|\le C\) and \(W(\theta )=a\, s(z)\). Then \(\hat\mu _W=Z_W^{-1}e^{-W/\beta }\mu _U\) is a probability measure with \(\operatorname {KL}(\hat\mu _W\| \mu _U){\lt}\infty \), since \(\log \frac{\, \mathrm d\hat\mu _W}{\, \mathrm d\mu _U}=-\frac{as(z)}\beta -\log Z_W\) is \(\hat\mu _W\)-integrable.

Proof ▶

Let \(s\) be measurable with \(|s|\le C\), \(W(\theta )=a\, s(z)\) and \(\hat\mu :=\hat\mu _W\) (relative to \(\mu _U\)). For every \(\rho \in {\mathcal D}\), \(\operatorname {KL}(\rho \| \hat\mu ){\lt}\infty \) and

\[ {\mathcal F}(\rho )=L(\rho )-\int _\Theta a\, s(z)\, \mathrm d\rho +\beta \, \operatorname {KL}(\rho \| \hat\mu )-\beta \log Z_W . \]

(Chain rule \(\operatorname {KL}(\rho \| \mu _U)=\operatorname {KL}(\rho \| \hat\mu )+\int \log \frac{\, \mathrm d\hat\mu }{\, \mathrm d\mu _U}\, \mathrm d\rho \), legitimate since \(\int |a||s|\, \mathrm d\rho \le C\int |a|\, \mathrm d\rho {\lt}\infty \) by Lemma 168.)

Proof ▶

Let \(\rho _0\in {\mathcal D}\) satisfy the self-consistency equation \(\rho _0=\hat\mu _{W_{\rho _0}}\), \(W_{\rho _0}(\theta )=a\, (S^*r_{\rho _0})(z)\). Then for every \(\rho \in {\mathcal D}\), \(\operatorname {KL}(\rho \| \rho _0){\lt}\infty \) and \({\mathcal F}(\rho )-{\mathcal F}(\rho _0)=\tfrac 12\| F_\rho -F_{\rho _0}\| _{L^2(P_X)}^2+\beta \operatorname {KL}(\rho \| \rho _0)\). (Lemmas 208 and 204 with \(\hat\mu =\rho _0\).)

Proof ▶

Step 1: the decomposition of ‘ℱ‘ relative to ‘μhat = ρ₀‘ at ‘ρ‘ and at ‘ρ₀‘.

Step 2: the expansion of ‘L‘ around ‘ρ₀‘.

If \(\rho _0\in {\mathcal D}\) solves the self-consistency equation \(\rho _0=\hat\mu _{W_{\rho _0}}\) with \(W_{\rho _0}(\theta )=a\, (S^*r_{\rho _0})(z)\), then \({\mathcal F}(\rho )\ge {\mathcal F}(\rho _0)+\beta \operatorname {KL}(\rho \| \rho _0)\ge {\mathcal F}(\rho _0)\) for every \(\rho \in {\mathcal D}\), so \(\rho _0\) is a minimizer of \({\mathcal F}\) on \({\mathcal D}\).

Proof ▶

Let \(\rho ^*\in {\mathcal D}\) be a minimizer of \({\mathcal F}\) on \({\mathcal D}\), \(s^*:=S^*r_{\rho ^*}\) and \(W^*(\theta ):=a\, s^*(z)\). Then \(Z_{W^*}{\lt}\infty \) and \(\rho ^*=\hat\mu _{W^*}\), i.e. \(\rho ^*(\, \mathrm d\theta )=Z_{W^*}^{-1}e^{-a s^*(z)/\beta }\, \mu _U(\, \mathrm d\theta )\); with \(\mu _U\propto e^{-U/\beta }\, \mathrm d\theta \) this is \(\rho ^*(\, \mathrm da\, \mathrm dz)\propto \exp \bigl(-\frac{a s^*(z)+\frac\lambda 2a^2+V(z)}\beta \bigr)\, \mathrm da\, \mathrm dz\). (Proof: with \(\hat\mu :=\hat\mu _{W^*}\in {\mathcal D}\) and \(\rho _{\varepsilon }:=(1-{\varepsilon })\rho ^*+{\varepsilon }\hat\mu \), Lemmas 208, 204 and 904 give \(0\le {\mathcal F}(\rho _{\varepsilon })-{\mathcal F}(\rho ^*)\le \frac{{\varepsilon }^2}2\| F_{\hat\mu }-F_{\rho ^*}\| ^2 -{\varepsilon }\beta \operatorname {KL}(\rho ^*\| \hat\mu )\); let \({\varepsilon }\downarrow 0\).)

Proof ▶

Step 1: ‘μhat = gibbs μ_U β (a s(z))‘ is a probability measure in the domain.

Step 2: the decomposition at ‘ρ₀‘, with ‘K = KL(ρ₀‖μhat)‘ and ‘D = ‖F_μhat − F_ρ₀‖²‘.

Step 3: for ‘ε ∈ (0,1)‘ and the mixture ‘ρ_ε = (1−ε) ρ₀ + ε μhat ∈ 𝒟‘, minimality, the decomposition, the expansion and the convexity of ‘KL‘ give ‘β K ≤ ε D / 2‘.

Step 4: letting ‘ε ↓ 0‘ gives ‘K = 0‘, hence ‘ρ₀ = μhat‘.

Let \(\rho ^*\in {\mathcal D}\) be a minimizer of \({\mathcal F}\) on \({\mathcal D}\). For every \(\rho \in {\mathcal D}\), \(\operatorname {KL}(\rho \| \rho ^*){\lt}\infty \) and

\[ {\mathcal F}(\rho )-{\mathcal F}(\rho ^*)=\tfrac 12\| F_\rho -F_{\rho ^*}\| _{L^2(P_X)}^2+\beta \, \operatorname {KL}(\rho \| \rho ^*) \ \ge \ \beta \, \operatorname {KL}(\rho \| \rho ^*). \]

(Lemma 209 with \(\rho _0=\rho ^*\), which is self-consistent by Lemma 211.)

Proof ▶

The minimizer of \({\mathcal F}\) on \({\mathcal D}\) is unique: if \(\rho _1,\rho _2\in {\mathcal D}\) both minimize \({\mathcal F}\) then \(\rho _1=\rho _2\). (By Lemma 212, \(0\ge {\mathcal F}(\rho _2)-{\mathcal F}(\rho _1)\ge \beta \operatorname {KL}(\rho _2\| \rho _1)\), so \(\operatorname {KL}(\rho _2\| \rho _1)=0\).)

Proof ▶

3.5 Existence of the minimizer

Theorem 214
✓

For \(T{\gt}0\) and \(a\in {\mathbb R}\), \(|a-\max (-T,\min (T,a))|\le a^2/T\).

Proof ▶

Let \(Z\) be a topological space with its Borel \(\sigma \)-algebra, let \(z\mapsto \varphi _z(x)\) be continuous for every \(x\), and let \(\rho _n\to \rho \) weakly in \({\mathcal P}(\Theta )\) with \(\sup _n\int a^2\, \mathrm d\rho _n\le C\) and \(\int a^2\, \mathrm d\rho \le C\). Then \(F_{\rho _n}(x)\to F_\rho (x)\) for every \(x\in {\mathcal X}\). Proof: for \(T{\gt}0\) the function \(\theta \mapsto \max (-T,\min (T,a))\, \varphi _z(x)\) is bounded and continuous, and replacing \(a\) by its clamped value changes \(F_\rho (x)\) by at most \(\int a^2\, \mathrm d\rho /T\le C/T\) (Lemma 214).

Proof ▶

Step 1: the truncation level ‘T = 4 max(C,1)/ε‘ and the bounded continuous test function ‘h_T(θ) = max (−T) (min T a) φ_z(x)‘.

Step 2: for every probability measure ‘σ‘ with ‘∫ a² dσ ≤ C‘, ‘|F_σ(x) − ∫ h_T dσ| ≤ ∫ a²/T dσ ≤ C₁/T = ε/4‘.

Step 3: weak convergence gives ‘|∫ h_T dρₙ − ∫ h_T dρ| < ε/2‘ for ‘n‘ large, and the triangle inequality concludes.

Under the hypotheses of Lemma 215, for \(P_X\in {\mathcal P}({\mathcal X})\) and \(f\in L^2(P_X)\), \(L(\rho _n)\to L(\rho )\). Proof: \(F_{\rho _n}(x)\to F_\rho (x)\) pointwise and \(|F_{\rho _n}|\le \sqrt C\), so \((F_{\rho _n}-f)^2\le 2C+2f^2\in L^1(P_X)\) and dominated convergence applies.

Proof ▶

Step 1: the uniform bound ‘F_σ(x)² ≤ ∫ a² dσ ≤ C‘.

Step 2: dominated convergence with the bound ‘2C + 2f²‘.

Let \(Z\) be a Polish space with its Borel \(\sigma \)-algebra, let \(z\mapsto \varphi _z(x)\) be continuous for every \(x\) (both hold under (A3)–(A5)), and let \(f\in L^2(P_X)\). Then the free energy \({\mathcal F}\) has a minimizer on \({\mathcal D}\): there is \(\rho ^*\in {\mathcal D}\) with \({\mathcal F}(\rho ^*)\le {\mathcal F}(\rho )\) for all \(\rho \in {\mathcal D}\). (Direct method: tightness of the sublevel sets of \(\operatorname {KL}(\cdot \| \mu _U)\), lower semicontinuity of \(\operatorname {KL}\), Prokhorov, and continuity of \(L\) along weakly convergent sequences with uniformly bounded second amplitude moments.)

Proof ▶

Step 1: the free energy is nonnegative, the domain contains ‘μ_U‘; let ‘m‘ be the infimum of ‘ℱ‘ on ‘𝒟‘.

Step 2: a minimizing sequence ‘ρₙ ∈ 𝒟‘ with ‘ℱ(ρₙ) < m + 1/(n+1)‘, so ‘ℱ(ρₙ) → m‘.

Step 3: the uniform bounds ‘KL(ρₙ‖μ_U) ≤ C := (m + 1)/β‘ (since ‘ℱ ≥ β KL‘) and ‘∫ a² dρₙ ≤ C’ := (4β/λ)(C + ½ log 2)‘ (‘lem:moment‘).

Step 4: tightness of ‘ρₙ‘ and Prokhorov’s theorem: a subsequence ‘ρ_φ n‘ converges weakly to some ‘ρlim‘.

Step 5: lower semicontinuity, ‘KL(ρlim‖μ_U) ≤ liminf KL(ρ_φ n‖μ_U) ≤ C‘: ‘ρlim ∈ 𝒟‘.

Step 6: continuity of the risk along the subsequence, ‘L(ρ_φ n) → L(ρlim)‘.

Step 7: ‘KL(ρ_φ n‖μ_U) = (ℱ(ρ_φ n) − L(ρ_φ n))/β → k := (m − L(ρlim))/β‘, and by lower semicontinuity ‘KL(ρlim‖μ_U) ≤ k‘.

Step 8: ‘ℱ(ρlim) = L(ρlim) + β KL(ρlim‖μ_U) ≤ L(ρlim) + β k = m ≤ ℱ(ρ’)‘ for all ‘ρ’ ∈ 𝒟‘.