6 Regularization limits and the geometry of the limit
Section 6 of the paper: representability, the canonical (minimum-norm) ridgelet transform, the energy comparison with a competitor built on the reference measure (thm:m4-energy), the convergence of the conditional mean amplitude to the canonical ridgelet transform as \(\lambda \to 0\) and \(\kappa =\lambda /\beta \to 0\), the order of the limits, and the reduction of the zero-temperature problem to total-variation regularization.
6.1 Reference measure and representability
Assume (R). \(u_0^\dagger :=P_{(\ker S_{\nu _0})^\perp }u_0\) is a representer and does not depend on the choice of the representer \(u_0\): for every representer \(u_0\) of \(f\), \(S_{\nu _0}^\dagger f=P_{(\ker S_{\nu _0})^\perp }u_0\) and \(S_{\nu _0}u_0^\dagger =f\). (Two representers differ by an element of \(\ker S_{\nu _0}\), so their projections onto \((\ker S_{\nu _0})^\perp \) coincide; and \(u_0-u_0^\dagger \in \ker S_{\nu _0}\).)
The chosen representer and ‘u₀‘ have the same synthesis ‘f‘.
Assume (R). For every representer \(u\) of \(f\), \(\| u\| ^2_{L^2(\nu _0)}=\| u_0^\dagger \| ^2_{L^2(\nu _0)}+\| u-u_0^\dagger \| ^2_{L^2(\nu _0)}\). In particular \(\| u_0^\dagger \| \le \| u\| \), with equality if and only if \(u=u_0^\dagger \). (\(u-u_0^\dagger \in \ker S_{\nu _0}\perp u_0^\dagger \), and Pythagoras.)
Pythagoras for the orthogonal projections onto ‘K = (ker S)ᗮ‘ and ‘Kᗮ‘, where the projection onto ‘Kᗮ‘ is ‘u − P_K u‘.
Assume (R). \(u_0^\dagger \) is the unique representer in \((\ker S_{\nu _0})^\perp =\overline{\operatorname {ran}S_{\nu _0}^*}\), and if \(h\in L^2(\nu _0)\) satisfies \(S_{\nu _0}h=f\) and \(\| h\| _{L^2(\nu _0)}\le \| u_0^\dagger \| _{L^2(\nu _0)}\) then \(h=u_0^\dagger \). (A representer \(h\in (\ker S_{\nu _0})^\perp \) satisfies \(h-u_0^\dagger \in \ker S_{\nu _0}\cap (\ker S_{\nu _0})^\perp =\{ 0\} \); for the last claim, Pythagoras with \(\| h\| ^2\le \| u_0^\dagger \| ^2\) forces \(\| h-u_0^\dagger \| =0\).)
6.2 Energy comparison
For \(\nu \in {\mathcal P}(Z)\) and \(u\in L^2(\nu )\) measurable, \(\int a^2\, \mathrm d\rho _{\nu ,u}=\| u\| ^2_{L^2(\nu )}+\sigma ^2\) (\({\mathbb E}X^2=\mu ^2+\sigma ^2\) for \(X\sim {\mathcal N}(\mu ,\sigma ^2)\), integrated against \(\nu \)).
For \(\nu \in {\mathcal P}(Z)\) and \(u\in L^2(\nu )\) measurable, \(\int |a|\, \mathrm d\rho _{\nu ,u}\le \| u\| _{L^1(\nu )}+\sigma \sqrt{2/\pi }{\lt}\infty \) (\({\mathbb E}|X|\le |\mu |+\sigma \sqrt{2/\pi }\) for \(X\sim {\mathcal N}(\mu ,\sigma ^2)\), and \(u\in L^2(\nu )\subset L^1(\nu )\)).
Step 1: the inner Gaussian moment bound ‘∫ |a| 𝒩(u(z), σ²)(da) ≤ |u(z)| + σ√(2/π)‘.
Step 2: integrate against ‘ν‘.
For \(\nu \in {\mathcal P}(Z)\) and \(u\in L^2(\nu )\) measurable, \(F_{\rho _{\nu ,u}}(x)=\int _Z\bigl[\int _{\mathbb R}a\, {\mathcal N}(u(z),\sigma ^2)(\, \mathrm da)\bigr] \varphi _z(x)\, \nu (\, \mathrm dz)=(S_\nu u)(x)\) for every \(x\) (Fubini, since \(\int |a|\, \mathrm d\rho _{\nu ,u}{\lt}\infty \)).
For \(\nu \in {\mathcal P}(Z)\) and \(u\in L^2(\nu )\) (measurable), \(\int |a|\, \mathrm d\rho _{\nu ,u}\le \| u\| _{L^1(\nu )}+\sigma \sqrt{2/\pi }{\lt}\infty \), \(\int a^2\, \mathrm d\rho _{\nu ,u}=\| u\| ^2_{L^2(\nu )}+\sigma ^2\), \(F_{\rho _{\nu ,u}}=S_\nu u\) and \(\Pi \rho _{\nu ,u}=u\, \nu \). (For \(X\sim {\mathcal N}(\mu ,\sigma ^2)\), \({\mathbb E}|X|\le |\mu |+\sigma \sqrt{2/\pi }\) and \({\mathbb E}X^2=\mu ^2+\sigma ^2\); apply this with \(\mu =u(z)\) and integrate against \(\nu \); Fubini gives \(F_{\rho _{\nu ,u}}(x)=\int _Z[\int _{\mathbb R}a\, {\mathcal N}(u(z),\sigma ^2)(\, \mathrm da)]\varphi _z(x)\nu (\, \mathrm dz) =(S_\nu u)(x)\) and likewise \(\Pi \rho _{\nu ,u}=u\nu \).)
Step 1: transport the set integral through the disintegration.
Step 2: the inner integral is the Gaussian mean ‘u(z)‘.
If \(\nu \ll \nu '\) and \(u,u'\) are measurable then \(\rho _{\nu ,u}\ll \rho _{\nu ',u'}\) with \(\frac{\, \mathrm d\rho _{\nu ,u}}{\, \mathrm d\rho _{\nu ',u'}}(a,z) =\frac{\, \mathrm d\nu }{\, \mathrm d\nu '}(z)\exp \Bigl(\frac{-(a-u(z))^2+(a-u'(z))^2}{2\sigma ^2}\Bigr) =\frac{\, \mathrm d\nu }{\, \mathrm d\nu '}(z) \exp \Bigl(\kappa \bigl(a(u-u')(z)-\tfrac 12(u^2-u'{}^2)(z)\bigr)\Bigr)\).
The exponent in the Gaussian shift identity, with ‘σ² = β/λ‘.
For fixed ‘z‘, the Gaussian shift identity ‘E(·, z) 𝒩(u’(z), σ²) = 𝒩(u(z), σ²)‘.
If \(\nu ,\nu '\in {\mathcal P}(Z)\), \(u\in L^2(\nu )\), \(u'\) measurable, \(\nu \ll \nu '\), \(\operatorname {KL}(\nu \| \nu '){\lt}\infty \) and \(u'\in L^2(\nu )\), then \(\operatorname {KL}(\rho _{\nu ,u}\| \rho _{\nu ',u'}){\lt}\infty \) and
(By Lemma 389, \(\log \frac{\, \mathrm d\rho _{\nu ,u}}{\, \mathrm d\rho _{\nu ',u'}}=\log \frac{\, \mathrm d\nu }{\, \mathrm d\nu '}(z) +\kappa (a(u-u')-\tfrac 12(u^2-u'{}^2))\), which is \(\rho _{\nu ,u}\)-integrable since \(|a(u-u')|\le \frac12(a^2+(u-u')^2)\) with \(\int a^2\, \mathrm d\rho _{\nu ,u}{\lt}\infty \); and \(\int a\, {\mathcal N}(u,\sigma ^2)(\, \mathrm da)=u\) gives \(\int \kappa (u(u-u')-\tfrac 12(u^2-u'{}^2))\, \mathrm d\nu =\frac\kappa 2\| u-u'\| ^2_{L^2(\nu )}\).)
Step 1: the log-likelihood ratio, ‘ρ‘-a.e. (where ‘dν/dν’ ∈ (0, ∞)‘).
Step 2: integrability of the log-likelihood ratio.
Step 3: integrate, using Fubini through the disintegration for the term ‘a (u − u’)‘.
If \(\nu \in {\mathcal P}(Z)\), \(u\in L^2(\nu )\) measurable, \(\nu \ll \nu _0\) and \(\operatorname {KL}(\nu \| \nu _0){\lt}\infty \), then \(\rho _{\nu ,u}\in {\mathcal D}\) and
(Lemma 390 with \(\nu '=\nu _0\) and \(u'=0\), since \(\mu _U=\rho _{\nu _0,0}\).)
Let \(\nu _0,\nu \in {\mathcal P}(Z)\) with \(\nu =w\, \nu _0\) for a measurable \(w\) with \(e^{-c}\le w\le e^{c}\). Then \(\operatorname {KL}(\nu \| \nu _0){\lt}\infty \) and \(\operatorname {KL}(\nu _0\| \nu ){\lt}\infty \), since \(\log \frac{\, \mathrm d\nu }{\, \mathrm d\nu _0}=\log w\) and \(\log \frac{\, \mathrm d\nu _0}{\, \mathrm d\nu }=-\log w\) are bounded by \(|c|\).
‘log w‘ is bounded by ‘|c|‘, hence integrable for every probability measure.
‘llr ν ν₀ = log w‘ and ‘llr ν₀ ν = −log w‘.
Let \(u_0\in L^2(\nu _0)\) be a representer of \(f\) and \(\tilde\rho :=\tilde\rho _{u_0}=\rho _{\nu _0,u_0}\). Then \(\tilde\rho \in {\mathcal D}\), \(F_{\tilde\rho }=S_{\nu _0}u_0=f\), \(L(\tilde\rho )=0\), \(\operatorname {KL}(\tilde\rho \| \mu _U)=\frac\lambda {2\beta }\| u_0\| ^2_{L^2(\nu _0)}\) and \({\mathcal F}(\tilde\rho )=\frac\lambda 2\| u_0\| ^2_{L^2(\nu _0)}\) (Lemma 391 with \(\nu =\nu _0\)).
Step 1: ‘KL(ρ̃‖μ_U) = (κ/2)‖u₀‖²‘.
Step 2: ‘F_ρ̃ = S_ν₀ u₀ = f‘, so ‘L(ρ̃) = 0‘.
\(\rho ^*=\rho _{\nu ^*,m^*}\) with \(m^*=-s^*/\lambda \) and variance \(\beta /\lambda \) (Theorem 225).
Chain rule.
and consequently, under (R),
(By Theorem 225, \(\rho ^*=\rho _{\nu ^*,m^*}\); \(\nu ^*\ll \nu _0\) with \(\operatorname {KL}(\nu ^*\| \nu _0){\lt}\infty \) and \(m^*\in L^2(\nu ^*)\), so Lemma 391 applies; then substitute into Theorem 397.)
Energy comparison. Assume (A1), (A3), (A4), (A5), \(f\in L^2(P_X)\) and (R). Let \(u_0\) be any representer of \(f\) and \(u_0^\dagger \) the minimum-norm representer. Then
(Minimality of \(\rho ^*\) against the competitor \(\tilde\rho _{u_0}\), Lemma 393.)
Under the hypotheses of Theorem 397,
(The first two from \(L(\rho ^*)=\frac12\| r^*\| ^2\), \(\operatorname {KL}\ge 0\) and \(L\ge 0\); the third from the chain rule of Theorem 396 and \(\| m^*\| ^2\ge 0\).)
Pythagorean identity. For every representer \(u_0\),
(Apply the strong-convexity identity of Lemma 212 to the competitor \(\tilde\rho =\tilde\rho _{u_0}\): \({\mathcal F}(\tilde\rho )-{\mathcal F}(\rho ^*) =L(\rho ^*)+\beta \operatorname {KL}(\tilde\rho \| \rho ^*)\), so \(\frac\lambda 2\| u_0\| ^2=2L(\rho ^*)+\beta \operatorname {KL}(\rho ^*\| \mu _U)+\beta \operatorname {KL}(\tilde\rho \| \rho ^*)\); then Theorem 396 and Lemma 390 with \(\nu =\nu _0\), \(u=u_0\), \(\nu '=\nu ^*\), \(u'=m^*\) give \(\operatorname {KL}(\tilde\rho \| \rho ^*)=\operatorname {KL}(\nu _0\| \nu ^*)+\frac\kappa 2\| u_0-m^*\| ^2_{L^2(\nu _0)}\); multiply by \(2/\lambda \).)
Step 1: the strong-convexity identity at the competitor.
Step 2: ‘KL(ρ*‖μ_U)‘ by the chain rule and ‘KL(ρ̃‖ρ*)‘ by the pair formula.
Step 3: rearrange and multiply by ‘2/λ‘.
Since every term on the right-hand side of the Pythagorean identity is nonnegative, with \(u_0=u_0^\dagger \),
Normalizing constant. With \(Z_*\), \(Z_0\) as in Theorem 226 and \(\tilde Z:=Z_*/Z_0=\int e^{\kappa m^{*2}/2}\, \mathrm d\nu _0\),
(By Theorem 226 and \(s^*=-\lambda m^*\), \(\frac{\, \mathrm d\nu ^*}{\, \mathrm d\nu _0}=\frac{Z_0}{Z_*}\exp (\frac{s^{*2}}{2\lambda \beta }) =\frac{Z_0}{Z_*}\exp (\frac{\lambda m^{*2}}{2\beta })\); integrating the logarithm against \(\nu ^*\) gives the formula for \(\operatorname {KL}(\nu ^*\| \nu _0)\); \(\tilde Z\ge 1\) since the integrand is \(\ge 1\); and \(\operatorname {KL}\ge 0\) with Theorem 396 give \(\log \tilde Z\le \frac\kappa 2\| m^*\| ^2_{L^2(\nu ^*)}\le \frac\kappa 2\| u_0^\dagger \| ^2\).)
The exponent identity ‘κ m*²/2 = s*²/(2λβ) = −W‘.
Step 1: the density.
Step 2: the relative entropy, by the Gibbs variational identity with ‘β = 1‘.
Step 3: the bounds on ‘Z̃‘.
6.3 Convergence to the canonical ridgelet transform
\(w^{-1}=\tilde Z\exp (-\frac\kappa 2m^{*2})\le \tilde Z\), so \(\| m^*\| ^2_{L^2(\nu _0)}=\int m^{*2}w^{-1}\, \mathrm d\nu ^*\le \tilde Z\| m^*\| ^2_{L^2(\nu ^*)} \le \exp (\frac\kappa 2\| u_0^\dagger \| ^2)\| u_0^\dagger \| ^2\).
Step 1: ‘m² ≤ Z̃ w m²‘ pointwise, since ‘w ≥ 1/Z̃‘.
Step 2: ‘‖m*‖²_L²(ν*) ≤ ‖u₀†‖²‘ (energy bound) and ‘Z̃ ≤ exp((κ/2)‖u₀†‖²)‘.
For \(T{\gt}0\) put \(\epsilon (T):=(1-\tilde Z^{-1})+(e^{\kappa T^2/2}-1)\). Where \(|m^*|\le T\) one has \(\tilde Z^{-1}\le w\le e^{\kappa T^2/2}\), hence \(|w-1|\le \epsilon (T)\); where \(|m^*|{\gt}T\) one has \(|w-1|\le w+1\) and \(|m^*|\le m^{*2}/T\). Hence pointwise \(|m^*||w-1|\le \epsilon (T)|m^*|+\frac{m^{*2}}T(w+1)\) and
Step 1: the pointwise inequality.
Step 2: integrate.
\(F_{\rho ^*}=S_{\nu ^*}m^*=\int m^*\varphi _z\, w\, \mathrm d\nu _0\) and \(|\varphi _z|\le 1\) give \(\sup _x|F_{\rho ^*}(x)-(S_{\nu _0}m^*)(x)|\le \int |m^*||w-1|\, \mathrm d\nu _0\).
\(\Pi \rho ^*=m^*\nu ^*=h\, \nu _0\) with \(h:=m^*w=\frac{\, \mathrm d\Pi \rho ^*}{\, \mathrm d\nu _0} \in L^1(\nu _0)\) (Theorem 229 and \(\nu ^*=w\nu _0\)).
For \(u\in L^2(\nu _0)\), \(\Pi \rho ^*-u\nu _0=(m^*w-u)\nu _0\), so \(\| \Pi \rho ^*-u\nu _0\| _{\mathcal M}=\| m^*w-u\| _{L^1(\nu _0)} \le \int |m^*||w-1|\, \mathrm d\nu _0+\| m^*-u\| _{L^1(\nu _0)} \le \int |m^*||w-1|\, \mathrm d\nu _0+\| m^*-u\| _{L^2(\nu _0)}\).
Step 1: the total mass is the ‘L¹(ν₀)‘ norm of the density ‘w m* − u‘.
Step 2: ‘|w m − u| ≤ |m| |w − 1| + |m − u|‘ and ‘∫ |m − u| dν₀ ≤ ‖m − u‖_L²(ν₀)‘.
For \(g\) measurable with \(|g|\le C\) and \(u\in L^2(\nu _0)\), \(\bigl|\int g\, \, \mathrm d\Pi \rho ^*-\int g\, u\, \, \mathrm d\nu _0\bigr| =\bigl|\int g\, (m^*w-u)\, \mathrm d\nu _0\bigr|\le C\, \| \Pi \rho ^*-u\nu _0\| _{\mathcal M}\).
\(\operatorname {KL}(\nu _n^*\| \nu _0)\le \frac{\kappa _n}2\| u_0^\dagger \| ^2\to 0\) (Theorem 398 with \((\lambda _n,\beta _n)\)).
\(\| \nu _n^*-\nu _0\| _{\mathrm{TV}}\le \sqrt{\operatorname {KL}(\nu _n^*\| \nu _0)/2}\to 0\) (Pinsker’s inequality, Lemma 850).
\(\| F_{\rho _n^*}-f\| ^2_{L^2(P_X)}\le \lambda _n\| u_0^\dagger \| ^2\to 0\) (Theorem 398).
\(\| m_n^*\| _{L^2(\nu _n^*)}\le \| u_0^\dagger \| _{L^2(\nu _0)}\) for all \(n\) (Theorem 396).
\(1\le \frac{Z_{*,n}}{Z_0}=\tilde Z_n\le \exp (\frac{\kappa _n}2\| u_0^\dagger \| ^2) \to 1\) (Theorem 401).
\(\| m_n^*\| ^2_{L^2(\nu _0)}\le \tilde Z_n\| m_n^*\| ^2_{L^2(\nu _n^*)} \le \exp (\frac{\kappa _n}2\| u_0^\dagger \| ^2)\| u_0^\dagger \| ^2\). In particular \((m_n^*)\) is bounded in \(L^2(\nu _0)\) and \(\limsup _n\| m_n^*\| _{L^2(\nu _0)}\le \| u_0^\dagger \| \).
\(\int _Z|m_n^*||w_n-1|\, \mathrm d\nu _0\to 0\); in particular \(\sup _x|F_{\rho _n^*}(x)-(S_{\nu _0}m_n^*)(x)|\le \int |m_n^*||w_n-1|\, \mathrm d\nu _0\to 0\). (Truncation, Lemma 404: for every \(T{\gt}0\), \(\int |m_n^*||w_n-1|\, \mathrm d\nu _0\le \epsilon _n(T)\sqrt{\tilde Z_n}\| u_0^\dagger \| +\frac1T(1+\tilde Z_n)\| u_0^\dagger \| ^2\) with \(\epsilon _n(T)\to 0\) and \(\tilde Z_n\to 1\), so \(\limsup _n\int |m_n^*||w_n-1|\, \mathrm d\nu _0\le 2\| u_0^\dagger \| ^2/T\); let \(T\to \infty \).)
Step 1: the truncation bound at level ‘T‘, uniformly in ‘n‘.
Step 2: the bound tends to ‘2‖u₀†‖²/T‘ as ‘n → ∞‘.
Step 3: conclude, choosing ‘T‘ with ‘2‖u₀†‖²/T < ε‘.
\(\langle m_n^*,u_0^\dagger \rangle _{L^2(\nu _0)}\to \| u_0^\dagger \| ^2\). (For \({\varepsilon }{\gt}0\) pick \(g\in L^2(P_X)\) with \(\| u_0^\dagger -S_{\nu _0}^*g\| \le {\varepsilon }\), possible since \(u_0^\dagger \in (\ker S_{\nu _0})^\perp =\overline{\operatorname {ran}S_{\nu _0}^*}\); then \(\langle m_n^*,S_{\nu _0}^*g\rangle =\langle S_{\nu _0}m_n^*,g\rangle \to \langle f,g\rangle =\langle u_0^\dagger ,S_{\nu _0}^*g\rangle \) by Theorem 417, and the remaining terms are bounded by \((\sup _n\| m_n^*\| _{L^2(\nu _0)}+\| u_0^\dagger \| ){\varepsilon }\) by Lemma 415.)
Step 1: ‘‖m_n^*‖_L²(ν₀) ≤ 2‖u₀†‖‘ eventually (‘Z̃_n ≤ 4‘ eventually).
Step 2: ‘⟪m_n^*, S* g⟫ → ⟪u₀†, S* g⟫‘ for every ‘g ∈ L²(P)‘.
Step 3: the ‘ε/3‘ argument.
\(m_n^*\to u_0^\dagger \) in \(L^2(\nu _0)\), i.e. \(\| m_n^*-u_0^\dagger \| _{L^2(\nu _0)}\to 0\). (In Lean, without compactness: \(\| m_n^*-u_0^\dagger \| ^2=\| m_n^*\| ^2-2\langle m_n^*,u_0^\dagger \rangle +\| u_0^\dagger \| ^2 \le \tilde Z_n\| u_0^\dagger \| ^2-2\langle m_n^*,u_0^\dagger \rangle +\| u_0^\dagger \| ^2\to 0\) by Lemmas 415 and 418.)
\(\| m_n^*\| _{L^2(\nu _0)}\to \| u_0^\dagger \| \), \(\| m_n^*\| _{L^2(\nu _n^*)}\to \| u_0^\dagger \| \) and \(\| S_{\nu _0}m_n^*-f\| _{L^2(P_X)}\to 0\). (The first from (a); for the second, \(\tilde Z_n^{-1}\| m_n^*\| ^2_{L^2(\nu _0)} \le \| m_n^*\| ^2_{L^2(\nu _n^*)}\le \| u_0^\dagger \| ^2\) with \(\tilde Z_n\to 1\); the third is ??.)
With \(h_n:=m_n^*w_n=\frac{\, \mathrm d\gamma _n}{\, \mathrm d\nu _0}\in L^1(\nu _0)\), \(\| h_n-u_0^\dagger \| _{L^1(\nu _0)}=\| \gamma _n-\gamma _\infty \| _{\mathcal M}\to 0\): \(\| h_n-u_0^\dagger \| _{L^1(\nu _0)}\le \int |m_n^*||w_n-1|\, \mathrm d\nu _0 +\| m_n^*-u_0^\dagger \| _{L^2(\nu _0)}\to 0\) by Lemma 416 and (a).
For every bounded measurable \(g\), \(|\int g\, \mathrm d\Pi \rho _n^*-\int g\, u_0^\dagger \, \mathrm d\nu _0| \le \| g\| _\infty \| \Pi \rho _n^*-u_0^\dagger \nu _0\| _{\mathcal M}\to 0\).
\(\gamma _n=\Pi \rho _n^*\rightharpoonup \gamma _\infty =u_0^\dagger \nu _0\), and \(\| m_n^*\| _{L^2(\nu _n^*)}\to \| u_0^\dagger \| _{L^2(\nu _0)}\): the learned coefficient measure converges to the canonical (minimum-norm) ridgelet transform with respect to the reference measure \(\nu _0\). (In Lean this is deduced from Theorem 421: convergence in total mass implies weak convergence.)
\(L(\rho _n^*)=o(\lambda _n)\), \(\operatorname {KL}(\nu _n^*\| \nu _0)+\operatorname {KL}(\nu _0\| \nu _n^*)=o(\kappa _n)\) and \(\| u_0^\dagger -m_n^*\| _{L^2(\nu _0)}\to 0\). (The Pythagorean identity ?? with \(u_0=u_0^\dagger \) and \((\lambda _n,\beta _n)\) gives \(\frac4{\lambda _n}L(\rho _n^*)+\| u_0^\dagger -m_n^*\| ^2_{L^2(\nu _0)} +\frac2{\kappa _n}[\operatorname {KL}(\nu _n^*\| \nu _0)+\operatorname {KL}(\nu _0\| \nu _n^*)] =\| u_0^\dagger \| ^2-\| m_n^*\| ^2_{L^2(\nu _n^*)}\to 0\) by Theorem 420; the three terms on the left are nonnegative.)
The defect ‘‖u₀†‖² − ‖m_n^*‖²_L²(ν_n^*) → 0‘.
6.4 Order of the limits
Assume (R). Let \(g\) be bounded measurable (e.g. \(g\in C_b(Z)\)) and \({\varepsilon }{\gt}0\). There are \(\lambda _0=\lambda _0({\varepsilon },g){\gt}0\) and \(\kappa _0=\kappa _0({\varepsilon },g){\gt}0\) such that every \((\lambda ,\beta )\) with \(\lambda \le \lambda _0\) and \(\kappa =\lambda /\beta \le \kappa _0\) satisfies \(\bigl|\int g\, \mathrm d\Pi \rho ^*_{\lambda ,\beta }-\int g\, u_0^\dagger \, \mathrm d\nu _0\bigr|\le {\varepsilon }\) (the paper writes \({\varepsilon }/3\)). (If the claim were false, there would be minimizers \(\rho _k^*\) for \((\lambda _k,\beta _k)\) with \(\lambda _k,\kappa _k\le 1/(k+1)\) violating the bound; this contradicts \(\Pi \rho _k^*\rightharpoonup u_0^\dagger \nu _0\) (Theorem 422) along that schedule.)
Step 1: a violating minimizer for ‘λ₀ = κ₀ = 1/(k+1)‘, for every ‘k‘.
Step 2: the violating family is a schedule in the sense of ‘def:m4-schedule‘.
Step 3: along the schedule the integrals converge, contradicting the violation.
In the setting of Theorem 328 with the complexity bound \({\mathfrak R}_N(\Phi )\le C\sqrt{(m+1)/N}\), for fixed \((\lambda ,\beta )\), a bounded measurable \(g\), \({\varepsilon }{\gt}0\) and \(\delta \in (0,1)\) there is \(N_1=N_1({\varepsilon },\delta ,g;\lambda ,\beta )\) such that for \(N\ge N_1\), with probability at least \(1-\delta \), \(\bigl|\int g\, \mathrm d\Pi \rho _N^*-\int g\, \mathrm d\Pi \rho ^*_{\lambda ,\beta }\bigr| \le 2\| g\| _\infty \| \Pi \rho _N^*-\Pi \rho ^*_{\lambda ,\beta }\| _{\mathrm{TV}}\le 2\| g\| _\infty r_N\le {\varepsilon }\) (the paper writes \({\varepsilon }/3\)). (Corollary 340 on the event \(E'_\delta \) of Lemma 339, the rate Corollary 341, and \(r_N\to 0\).)
Step 1: ‘2 ‖g‖_∞ r_N ≤ ε‘ for ‘N ≥ N₀‘.
Step 2: on the event ‘E’_δ‘ (probability ‘≥ 1 − δ‘), the three bounds.
Assume the setting of Theorem 328 ((A1)–(A5), i.i.d. sample, \(|Y|\le {Y_{\max }}\), \(f={\mathbb E}[Y\mid X]\)), (R), the convention ?? and \({\mathfrak R}_N(\Phi )\le C\sqrt{(m+1)/N}\). Let \(g\) be bounded measurable, \({\varepsilon }{\gt}0\) and \(\delta \in (0,1)\). Then there are \(\lambda _0,\kappa _0{\gt}0\) such that for every \((\lambda ,\beta )\) with \(\lambda \le \lambda _0\), \(\lambda /\beta \le \kappa _0\) there is \(N_1=N_1({\varepsilon },\delta ,g;\lambda ,\beta )\) such that for \(N\ge N_1\), with probability at least \(1-\delta \),
(Stages (1) and (2) of Theorem ?? and the triangle inequality; the third stage, the reachability \(e^{(g)}_{M,t}\), is the subject of the dynamics.)
Stage (1) at ‘(λ, β)‘.
Stage (2) at ‘(λ, β)‘.
6.5 The zero-temperature limit: total-variation regularization
If \(\gamma =h\nu \) with \(h\in L^1(\nu )\) then \((S\gamma )(x)=\int _Z h(z)\varphi _z(x)\, \nu (\, \mathrm dz)=(S_\nu h)(x)\) for every \(x\).
If \(\gamma =h\nu \) with \(h\in L^1(\nu )\) then \(S\gamma =S_\nu h\).
For every finite signed measure \(\gamma \) on \(Z\) and every \(x\in {\mathcal X}\), \(|(S\gamma )(x)|\le \| \gamma \| _{\mathcal M}\), since \(|\varphi _z(x)|\le 1\).
For every finite signed measure \(\gamma \) on \(Z\) the synthesis \(S\gamma \colon {\mathcal X}\to {\mathbb R}\) is measurable.
\(S(\gamma _1+\gamma _2)=S\gamma _1+S\gamma _2\) pointwise.
\(S(c\gamma )=c\, S\gamma \) pointwise, for \(c\in {\mathbb R}\).
If \(\int |a|\, \mathrm d\rho {\lt}\infty \) then \(F_\rho (x)=\int _Z\varphi _z(x)\, \Pi \rho (\, \mathrm dz)=(S\, \Pi \rho )(x)\) for every \(x\in {\mathcal X}\).
If \(\int |a|\, \mathrm d\rho {\lt}\infty \) then \(\| \Pi \rho \| _{\mathcal M}\le \int |a|\, \mathrm d\rho \), so that \(\sup _x|F_\rho (x)|\le \| \Pi \rho \| _{\mathcal M}\le \int |a|\, \mathrm d\rho \).
Step 1: ‘Πρ = (snd)_*(a⁺ ρ) − (snd)_*(a⁻ ρ)‘, so ‘‖Πρ‖_ℳ ≤ ∫ a⁺ dρ + ∫ a⁻ dρ‘.
Step 2: ‘∫ a⁺ dρ + ∫ a⁻ dρ = ∫ |a| dρ‘.
If \(\int |a|\, \mathrm d\rho {\lt}\infty \) then \(L(\rho )=L(\Pi \rho )\), where \(L(\gamma )=\tfrac 12\| S\gamma -f\| ^2_{L^2(P_X)}\).
Let \(\rho \in {\mathcal P}_2\) with \(\Pi \rho =\gamma \). Then \(\| \gamma \| _{\mathcal M}\le \int |a|\, \mathrm d\rho \le \bigl(\int a^2\, \mathrm d\rho \bigr)^{1/2}\), i.e. \(\| \gamma \| ^2_{\mathcal M}\le \int a^2\, \mathrm d\rho \).
The measure \(\bar\rho \) of 159 is a probability measure in \({\mathcal P}_2\) with \(\Pi \bar\rho =\gamma \) and \(\int a^2\, \mathrm d\bar\rho =\| \gamma \| ^2_{\mathcal M}\); for \(\gamma =0\) the same holds for \(\delta _0\otimes \nu \).
Let \(\gamma \) be a finite signed measure on \(Z\) (and \(\nu \in {\mathcal P}(Z)\) arbitrary, used only when \(\gamma =0\)). Then \(\inf \bigl\{ \int a^2\, \mathrm d\rho :\rho \in {\mathcal P}_2,\ \Pi \rho =\gamma \bigr\} =\| \gamma \| ^2_{\mathcal M}\) and the infimum is attained, at \(\bar\rho \) of 159.
\(\inf \bigl\{ \int a^2\, \mathrm d\rho :\rho \in {\mathcal P}_2,\ \Pi \rho =\gamma \bigr\} =\| \gamma \| ^2_{\mathcal M}\).
For \(\rho \in {\mathcal P}_2\) and \(\lambda \ge 0\), \(L(\rho )+\frac\lambda 2\int a^2\, \mathrm d\rho \ge L(\Pi \rho )+\frac\lambda 2\| \Pi \rho \| ^2_{\mathcal M}\).
For every \(\gamma \in {\mathcal M}(Z)\) the measure \(\bar\rho \) of 159 satisfies \(L(\bar\rho )+\frac\lambda 2\int a^2\, \mathrm d\bar\rho =L(\gamma )+\frac\lambda 2\| \gamma \| ^2_{\mathcal M}\).
Let \(\lambda \ge 0\), \(P\) a measure on \({\mathcal X}\), \(f\colon {\mathcal X}\to {\mathbb R}\) and \(\nu \in {\mathcal P}(Z)\) (so that \({\mathcal P}_2\ne \emptyset \)). Then \(\inf _{\rho \in {\mathcal P}_2}\bigl\{ L(\rho )+\frac\lambda 2\int a^2\, \mathrm d\rho \bigr\} =\inf _{\gamma \in {\mathcal M}(Z)}\bigl\{ L(\gamma )+\frac\lambda 2\| \gamma \| ^2_{\mathcal M}\bigr\} \), both sides being infima of nonnegative sets.
‘≤‘: for every ‘γ‘, ‘ρ̄ ∈ 𝒫₂‘ has ‘ℰ₀(ρ̄) = ℰ_TV(γ)‘.
‘≥‘: for every ‘ρ ∈ 𝒫₂‘, ‘ℰ_TV(Πρ) ≤ ℰ₀(ρ)‘.
If \(\rho \in {\mathcal P}_2\) minimizes \(L(\rho )+\frac\lambda 2\int a^2\, \mathrm d\rho \) over \({\mathcal P}_2\) then \(\Pi \rho \) minimizes \(L(\gamma )+\frac\lambda 2\| \gamma \| ^2_{\mathcal M}\) over \({\mathcal M}(Z)\).
If \(\gamma \) minimizes \(L(\gamma )+\frac\lambda 2\| \gamma \| ^2_{\mathcal M}\) over \({\mathcal M}(Z)\) then \(\bar\rho \) of 159 lies in \({\mathcal P}_2\) and minimizes \(L(\rho )+\frac\lambda 2\int a^2\, \mathrm d\rho \) over \({\mathcal P}_2\).
For \(f\in L^2(P_X)\) the functional \(\gamma \mapsto L(\gamma ) =\frac12\| S\gamma -f\| ^2_{L^2(P_X)}\) is convex on \({\mathcal M}(Z)\).
Step 1: ‘(Sγ − f)² ∈ L¹(P)‘ for every ‘γ‘, since ‘Sγ‘ is bounded and ‘f ∈ L²‘.
Step 2: the pointwise convexity inequality ‘(a s₁ + b s₂ − f)² ≤ a (s₁ − f)² + b (s₂ − f)²‘, integrated.
\(\gamma \mapsto \| \gamma \| ^2_{\mathcal M}\) is convex on \({\mathcal M}(Z)\), since \(\| \cdot \| _{\mathcal M}\) is a norm.
For \(f\in L^2(P_X)\) and \(\lambda \ge 0\) the functional \(\gamma \mapsto L(\gamma )+\frac\lambda 2\| \gamma \| ^2_{\mathcal M}\) is convex on \({\mathcal M}(Z)\).