10 Heavy-tailed hidden laws and Sobolev-type representability
Section 12 of the paper: for \(m=1\) and \(\sigma =\tanh \), the truncated homogeneous hidden law \(\nu _0^{(L)}\) with density \(|w|^{\alpha -1}\) on a box, whose kernel converges to the sign-network kernel \(K^{(\infty )}(x,x')=1-c_\alpha |x-x'|\) as \(L\to \infty \); representability becomes \(f\in H^1\) in the limit.
For a feature map \(\varphi \) and \(\nu \in {\mathcal P}(Z)\), the kernel function is \(K_\nu (x,x'):=\int \varphi _z(x)\varphi _z(x')\, \nu (\, \mathrm dz)\), the integral kernel of \(K_\nu =S_\nu S_\nu ^*\) (Definition 18).
For \(m=1\) the feature \(\varphi _{(w,b)}(x)=\tanh (w^\top x-b)\) of Definition 5 is \(\tanh (w_0x_0-b)\) under \({\mathbb R}^1\cong {\mathbb R}\).
\(c_\alpha :=\frac\alpha {1+\alpha }\) ??.
\(\ell _\alpha :=\frac1{c_\alpha }=\frac{1+\alpha }\alpha \) ??.
\(M_L:=\int _{1/L}^Lw^{\alpha -1}\, \mathrm dw=\frac{L^\alpha -L^{-\alpha }}\alpha \) (Lemma 649).
\(Z_L:=4LM_L=\frac{4L(L^\alpha -L^{-\alpha })}\alpha \), the value of \(\int _{B_L}|w|^{\alpha -1}\, \mathrm dw\, \mathrm db\) (Lemma 649).
The annulus \(A_L:=\{ w\in {\mathbb R}:\ 1/L\le |w|\le L\} \) in the scale variable, so that \(B_L=A_L\times [-L,L]\).
\(B_L:=\{ (w,b):\ 1/L\le |w|\le L,\ |b|\le L\} \) ??.
Let \(0{\lt}\alpha {\lt}1\) and \(L\ge 2\). The truncated homogeneous law is \(\nu _0^{(L)}(\, \mathrm dw\, \mathrm db):=Z_L^{-1}|w|^{\alpha -1}\mathbf1_{B_L}(w,b)\, \mathrm dw\, \mathrm db\) with \(B_L=\{ (w,b):\ 1/L\le |w|\le L,\ |b|\le L\} \) and \(Z_L=\int _{B_L}|w|^{\alpha -1}\, \mathrm dw\, \mathrm db\) ??.
\(K^{(L)}(x,x'):=K_{\nu _0^{(L)}}(x,x')=\int \tanh (wx-b)\tanh (wx'-b)\, \nu _0^{(L)}(\, \mathrm dw\, \mathrm db)\) (Definition 610).
The sign feature with threshold \(t\in {\mathbb R}\) is \(\varphi _t(x):=\operatorname {sgn}(x-t)\) (\(=0\) at \(x=t\)), a feature map on \(Z={\mathcal X}={\mathbb R}\) (Remark 655).
\(\varpi :=U[-\ell _\alpha ,\ell _\alpha ]\), the uniform law \(\frac{\, \mathrm dt}{2\ell _\alpha }\) on \([-\ell _\alpha ,\ell _\alpha ]\) (Remark 655).
(Conventions of Section 10.) \({\mathcal X}=[-1,1]\) and \(P_X(\, \mathrm dx)=\frac12\, \mathrm dx\) on \([-1,1]\), so that (A1) holds with \(R_X=1\).
\(K^{(\infty )}(x,x'):=1-c_\alpha |x-x'|=1-\frac\alpha {1+\alpha }|x-x'|\) (Theorem 676).
\(s\coth s=\frac{s\cosh s}{\sinh s}\) for \(s\ne 0\), extended by its limit \(1\) at \(s=0\) (the value of the bias integral ?? divided by \(2\)).
\(I_\alpha (S):=\int _0^S(s\coth s-1)s^{\alpha -1}\, \mathrm ds\) for \(S{\gt}0\) ??.
\(J_\alpha (S):=\int _0^S\frac{2s^\alpha }{e^{2s}-1}\, \mathrm ds\) ??.
\(D_\alpha :=\int _0^\infty \frac{2s^\alpha }{e^{2s}-1}\, \mathrm ds =2^{-\alpha }\Gamma (1+\alpha )\zeta (1+\alpha )\) ??.
\(\Psi _L(d):=\int _{1/L}^Lwd\coth (wd)\, w^{\alpha -1}\, \mathrm dw\) for \(d\ne 0\) and \(\Psi _L(0):=M_L\) ??.
\(E_L(w;x,x'):=\int _{|b|{\gt}L}\bigl[1-\tanh (wx-b)\tanh (wx'-b)\bigr]\, \mathrm db\) ??.
\({\mathcal{E}}_L(x,x'):=\int _{1/L\le |w|\le L}E_L(w;x,x')\, |w|^{\alpha -1}\, \mathrm dw\) ??.
\(T_2:=\frac{\alpha L^{-2\alpha }}{(1+\alpha )(1-L^{-2\alpha })}|d|\) (Theorem 676).
\(T_3:=\frac{|d|^{-\alpha }J_\alpha (L|d|)}{LM_L}\) (Theorem 676); in Lean \(0^{-\alpha }=0\), so \(T_3=0\) at \(d=0\) (the paper takes the limit \(\frac1L+T_1\) there).
\(C_\alpha :=\frac1{1-2^{-2\alpha }}\Bigl(2+\frac{2\alpha }{1+\alpha } +\frac{4\alpha }{3(2+\alpha )}+2^{1-\alpha }\alpha \Bigr)+1\) ??.
For \(L\in {\mathbb R}\) put \(\nu _0^{[L]}:=\nu _0^{(\max (L,2))}\) (Definition 610); this is \(\nu _0^{(L)}\) for \(L\ge 2\) and a probability measure for every \(L\), so that the limits \(L\to \infty \) can be taken along all real (or natural) \(L\).
For \(\lambda {\gt}0\) and \(f\in L^2(P_X)\), \(R^{(\infty )}_\lambda f:=S^*(K^{(\infty )}+\lambda )^{-1}f\), i.e. \((R^{(\infty )}_\lambda f)(w,b)=\int \tanh (wx-b)\, \bigl[(K^{(\infty )}+\lambda )^{-1}f\bigr](x)\, P_X(\, \mathrm dx)\), where \(S^*\) is the analysis map of the tanh feature and \(K^{(\infty )}\) the kernel operator of the sign network ??.
For \(f=K^{(\infty )}g\in \operatorname {ran}K^{(\infty )}\) (with \(g\in L^2(P_X)\) unique by injectivity), \(u^\dagger _\infty :=S^*(K^{(\infty )})^{-1}f=S^*g\), i.e. \(u^\dagger _\infty (w,b)=\int \tanh (wx-b)g(x)\, P_X(\, \mathrm dx)\) ??.
The integral operator of the limit kernel on functions on \([-1,1]\), \(K^{(\infty )}g(x):=\int _{-1}^1K^{(\infty )}(x,x')g(x')\, P_X(\, \mathrm dx') =\frac12\int _{-1}^1(1-c_\alpha |x-x'|)g(x')\, \mathrm dx'\) (Theorem ??).
\(Tg(x):=\int _{-1}^1|x-y|g(y)\, \mathrm dy\), so that \(K^{(\infty )}g=\frac12\langle g,1\rangle _{L^2(\, \mathrm dx)}-\frac{c_\alpha }2Tg\) (proof of Theorem ??).
\((K^{(\infty )}g)'(x)=-\frac{c_\alpha }2\Bigl(\int _{-1}^xg-\int _x^1g\Bigr)\), the derivative of \(K^{(\infty )}g=\frac12\int g-\frac{c_\alpha }2Tg\) with \((Tg)'(x)=\int _{-1}^xg-\int _x^1g\) (proof of Theorem ??).
\(\varkappa ^e_j\), the unique root of \(\varkappa \tan \varkappa =\alpha \) in \((j\pi ,j\pi +\frac\pi 2)\) (Theorem ??(iii)); \({\mathcal{K}}^e=\{ \varkappa ^e_j\} _{j\ge 0}\).
\(\varkappa _{(2j)}:=\varkappa ^e_j\) and \(\varkappa _{(2j+1)}:=(j+\frac12)\pi \), the increasing enumeration of \({\mathcal{K}}^e\cup {\mathcal{K}}^o\) (Theorem ??(iv)).
\(\mu _k:=\frac{c_\alpha }{\varkappa _{(k)}^2}\), the eigenvalues of \(K^{(\infty )}\) in decreasing order (Theorem ??(iv)).
\(k_{\max }(L):=\bigl\lfloor \sqrt{2c_\alpha /(\pi ^2{\varepsilon }_L)}\bigr\rfloor -1\) (Theorem ??(b)).
For a bounded operator \(A\) on a real Hilbert space and a subspace \(V\), \(m_V(A):=\inf _{g\in V,\ \| g\| =1}\langle Ag,g\rangle \) (proof of Theorem ??(a)).
\(\mu _k(A):=\max _{\dim V=k+1}\ \min _{g\in V,\ \| g\| =1}\langle Ag,g\rangle \), the min–max value; for a compact self-adjoint positive operator it is the \(k\)-th eigenvalue in decreasing order (proof of Theorem ??(a)).
For \(\nu \in {\mathcal P}(Z)\), \(\lambda {\gt}0\) and \(f\in L^2(P_X)\) the resolvent quadratic form is \(Q^\nu _\lambda (f):=\langle f,(K_\nu +\lambda )^{-1}f\rangle _{L^2(P_X)}\), where \(K_\nu =S_\nu S_\nu ^*\) is the kernel operator and \((K_\nu +\lambda )^{-1}\) its resolvent (Definition 19). Equivalently, with \(g=(K_\nu +\lambda )^{-1}f\), \(Q^\nu _\lambda (f)=\| S^*g\| ^2_{L^2(\nu )}+\lambda \| g\| ^2_{L^2(P_X)}\); and \(\frac\lambda 2Q^\nu _\lambda (f)=\min _{u\in L^2(\nu )}\bigl[\frac12\| S_\nu u-f\| ^2 +\frac\lambda 2\| u\| ^2\bigr]\) (Proposition ??).
\(f\in H^1(-1,1)\) with weak derivative \(g\) if \(g\in L^2(P_X)\) and \(f(x)=f(-1)+\int _{-1}^xg(t)\, \, \mathrm dt\) for all \(x\in [-1,1]\) (absolutely continuous with an \(L^2\) derivative; \(g=f'\) is unique a.e.). \(H^1(-1,1)\) is the set of \(f\) admitting such a \(g\).
\(f\in H^2(-1,1)\) with \(f'=g\) and \(f''=h\) if \(f\in H^1(-1,1)\) with weak derivative \(g\) and \(g\in H^1(-1,1)\) with weak derivative \(h\).
For \(f\) with \(f'=g\) continuous on \([-1,1]\), the boundary conditions of ?? are \(f'(1)=-\frac\alpha 2\bigl(f(1)+f(-1)\bigr)\) and \(f'(-1)=+\frac\alpha 2\bigl(f(1)+f(-1)\bigr)\).
For \(f\in H^1(-1,1)\) with \(f'=g\) and \(c_0:=\frac{f(1)+f(-1)}2\), \(\| f\| ^2_{{\mathcal{H}}_\infty }:=\frac{1+\alpha }\alpha \| f'\| ^2_{L^2(P_X)}+(1+\alpha )c_0^2\) ??, where \(\| f'\| ^2_{L^2(P_X)}=\frac12\int _{-1}^1f'(x)^2\, \mathrm dx\).
For \(f\in H^1(-1,1)\) with \(f'=g\) and \(c_0:=\frac{f(1)+f(-1)}2\), \(u_\varpi ^\dagger (t):=\frac{1+\alpha }\alpha g(t)\mathbf1_{(-1,1)}(t) +(1+\alpha )c_0\bigl[\mathbf1_{(-\ell _\alpha ,-1)}(t)-\mathbf1_{(1,\ell _\alpha )}(t)\bigr]\) ??.
For \(s\ne 0\), \(2s\coth s=2|s|+\frac{4|s|}{e^{2|s|}-1}\) ??.
For \(\alpha {\gt}0\) and \(L\ge 1\), \(\int _{1/L\le |w|\le L}|w|^{\alpha -1}\, \mathrm dw=2M_L\) (Lemma 649).
Step 1: split the annulus into the two intervals.
Step 2: the positive part is ‘∫_1/L^L w^α−1 dw = M_L‘.
Step 3: the negative part equals the positive one by ‘w ↦ −w‘.
With \(M_L:=\int _{1/L}^Lw^{\alpha -1}\, \mathrm dw=\frac{L^\alpha -L^{-\alpha }}\alpha \), \(Z_L=\int _{B_L}|w|^{\alpha -1}\, \mathrm dw\, \mathrm db=2L\cdot 2M_L=4LM_L =\frac{4L(L^\alpha -L^{-\alpha })}\alpha \) ??.
Step 1: the box is a product and the integrand depends on ‘w‘ alone.
Step 2: the bias interval has length ‘2L‘.
For \(\alpha {\gt}0\) and \(L{\gt}1\), \(\nu _0^{(L)}\) is a probability measure: \(\int _{B_L}|w|^{\alpha -1}\, \mathrm dw\, \mathrm db=Z_L\) (Lemma 649).
Step 1: the lower integral is the Bochner integral of the nonnegative density.
Step 2: ‘∫_B_L |w|^α−1 = Z_L‘.
For \(\alpha {\gt}0\) and \(L{\gt}1\), \(\nu _0^{(L)}\bigl(\{ |b|\le |w|\} \bigr)=\frac\alpha {1+\alpha }\cdot \frac{1-L^{-2-2\alpha }}{1-L^{-2\alpha }}\) ??; it tends to \(c_\alpha \) as \(L\to \infty \).
Step 1: the mass is the Bochner integral of the density over ‘|b| ≤ |w| ∩ B_L‘.
Step 2: Fubini: the inner integral in ‘b‘ over ‘[−|w|, |w|]‘ gives ‘2|w|^α‘ on the annulus.
Step 3: ‘∫_1/L ≤ |w| ≤ L 2|w|^α dw = 4 M_α+1,L‘.
For \(u_1,u_2,b\in {\mathbb R}\) and \(s:=u_1-u_2\), \(1-\tanh (u_1-b)\tanh (u_2-b)=\frac{\cosh s}{\cosh (b-u_1)\cosh (b-u_2)}\) ??.
For \(u_1,u_2\in {\mathbb R}\) with \(s:=u_1-u_2\ne 0\), \(\int _{\mathbb R}\frac{\, \mathrm db}{\cosh (b-u_1)\cosh (b-u_2)}=\frac{2s}{\sinh s}\) ??.
For \(u_1,u_2\in {\mathbb R}\) and \(s:=u_1-u_2\), \(1-\tanh (u_1-b)\tanh (u_2-b)=\frac{\cosh s}{\cosh (b-u_1)\cosh (b-u_2)}\), \(\int _{\mathbb R}\frac{\, \mathrm db}{\cosh (b-u_1)\cosh (b-u_2)}=\frac{2s}{\sinh s}\), hence \(\int _{\mathbb R}\bigl[1-\tanh (u_1-b)\tanh (u_2-b)\bigr]\, \mathrm db=2s\coth s =2|s|+\frac{4|s|}{e^{2|s|}-1}\) (\(s\ne 0\); \(=2\) for \(s=0\)).
Step 1 (‘s = 0‘): the integrand is ‘1 − tanh²(b − u₁) = (tanh(b − u₁))’‘ and ‘tanh(±∞) = ±1‘.
Step 2 (‘s ≠ 0‘): ‘cosh s · 2s / sinh s = 2 s coth s‘.
For \(\varpi :=U[-\ell _\alpha ,\ell _\alpha ]\) and \(x,x'\in [-1,1]\), \(\int _{-\ell _\alpha }^{\ell _\alpha }\operatorname {sgn}(x-t)\operatorname {sgn}(x'-t)\, \frac{\, \mathrm dt}{2\ell _\alpha }=1-\frac{|x-x'|}{\ell _\alpha }=K^{(\infty )}(x,x')\) ??.
For \(|x|,|x'|\le 1\), \(|w|\le L\) and \(g_w(b):=\tanh (wx-b)\tanh (wx'-b)\), \(E_L(w;x,x'):=\int _{|b|{\gt}L}\bigl[1-g_w(b)\bigr]\, \mathrm db\in \bigl[0,\ 4e^{-2(L-|w|)}\bigr]\) ??.
Step 1: split ‘|b| > L‘ into the two half lines.
Step 2: the half line ‘b > L‘: ‘∫_L^∞ 4 e^2|w| e^−2b db = 2 e^2|w| e^−2L‘.
Step 3: the half line ‘b < −L‘: ‘∫_−∞^−L 4 e^2|w| e^2b db = 2 e^2|w| e^−2L‘.
Step 4: sum the two bounds.
For \(S\ge 0\), \(J_\alpha (S)\ge 0\).
For \(0{\lt}\alpha \) and \(S\ge 0\), \(J_\alpha (S)\le S^\alpha /\alpha \) (from \(\frac{2s}{e^{2s}-1}\le 1\)).
For \(0{\lt}\alpha \) and \(S\ge 0\), \(J_\alpha (S)\le D_\alpha =\int _0^\infty \frac{2s^\alpha }{e^{2s}-1}\, \mathrm ds{\lt}\infty \) ??.
For \(S\ge 0\), \(I_\alpha (S)\ge 0\) (the integrand \((s\coth s-1)s^{\alpha -1}\) is nonnegative).
For \(0{\lt}\alpha \) and \(S\ge 0\), \(I_\alpha (S)\le \frac{S^{2+\alpha }}{3(2+\alpha )}\) (from the Langevin bound \(s\coth s-1\le s^2/3\)).
For \(0{\lt}\alpha \) and \(S\ge 0\), \(I_\alpha (S)=\frac{S^{1+\alpha }}{1+\alpha }-\frac{S^\alpha }\alpha +J_\alpha (S)\) ??: integrate \((s\coth s-1)s^{\alpha -1}=s^\alpha -s^{\alpha -1} +\frac{2s^\alpha }{e^{2s}-1}\), each term being integrable at \(0\).
Step 1: the pointwise identity of the integrands on ‘[0, S]‘ (at ‘s = 0‘ both sides vanish).
Step 2: integrate term by term.
For \(L\ge 0\) and \(g_w(b)=\tanh (wx-b)\tanh (wx'-b)\), \(\int _{-L}^Lg_w(b)\, \mathrm db=2L-\int _{\mathbb R}(1-g_w)\, \mathrm db+E_L(w;x,x')=2L-2wd\coth (wd)+E_L(w;x,x')\) with \(d=x-x'\) (??, ??).
Step 1: ‘∫_ℝ (1 − g_w) = ∫_|b| > L (1 − g_w) + ∫_[−L, L] (1 − g_w)‘.
Step 2: ‘∫_[−L, L] g_w = 2L − ∫_[−L, L] (1 − g_w)‘.
\(Z_LK^{(L)}(x,x')=\int _{B_L}\tanh (wx-b)\tanh (wx'-b)\, |w|^{\alpha -1}\, \mathrm dw\, \mathrm db\) (Definition 610).
For \(0{\lt}\alpha \), \(L{\gt}1\) and \(x,x'\in [-1,1]\), \(d=x-x'\): \(Z_LK^{(L)}(x,x')=2L\cdot 2M_L-4\Psi _L(d)+{\mathcal{E}}_L(x,x')\) (Fubini over \(B_L\), the inner integral of Lemma 663, and the evenness of \(wd\coth (wd)\) in \(w\)).
Step 1: Fubini over ‘B_L = A_L × [−L, L]‘.
Step 2: the inner integral (‘lem:ht-inner-bias-integral‘).
Step 3: split the outer integral into its three parts.
Step 4: ‘∫_A_L (wd) coth(wd) |w|^α−1 dw = 2 Ψ_L(d)‘ by evenness.
For \(0{\lt}\alpha \), \(L{\gt}1\), \(x,x'\in [-1,1]\) and \(d=x-x'\), \(K^{(L)}(x,x')=1-\frac{\Psi _L(d)}{LM_L}+\frac{{\mathcal{E}}_L(x,x')}{4LM_L}\) ??.
For \(0{\lt}\alpha {\lt}1\), \(L\ge 2\) and \(x,x'\in [-1,1]\), \(0\le {\mathcal{E}}_L(x,x')\le 2^{3-\alpha }L^{\alpha -1}+8e^{-L}M_L\) ??: by ??, \({\mathcal{E}}_L\le 8\int _{1/L}^Le^{-2(L-w)}w^{\alpha -1}\, \mathrm dw\); on \([1/L,L/2]\), \(e^{-2(L-w)}\le e^{-L}\), and on \([L/2,L]\), \(w^{\alpha -1}\le (L/2)^{\alpha -1}\) and \(\int _{L/2}^Le^{-2(L-w)}\, \mathrm dw\le \frac12\).
Step 1: ‘E_L ≤ 4 e^−2(L − |w|)‘ pointwise (‘lem:bias-truncation‘).
Step 2: evenness and ‘|w| = w‘ on ‘[1/L, L]‘.
Step 3: split at ‘L/2‘.
Step 4: on ‘[1/L, L/2]‘, ‘e^−2(L − w) ≤ e^−L‘ and ‘∫_1/L^L/2 w^α−1 ≤ M_L‘.
Step 5: on ‘[L/2, L]‘, ‘w^α−1 ≤ (L/2)^α−1‘ and ‘∫_L/2^L e^−2(L − w) ≤ 1/2‘.
\(\Psi _L(0)=M_L\) ??.
For \(0{\lt}\alpha \), \(L\ge 1\) and \(d\ne 0\), \(\Psi _L(d)=\int _{1/L}^Lwd\coth (wd)\, w^{\alpha -1}\, \mathrm dw =M_L+|d|^{-\alpha }\bigl[I_\alpha (L|d|)-I_\alpha (|d|/L)\bigr]\) ?? (substitute \(s=w|d|\)).
Step 1: on ‘[1/L, L]‘, ‘(wd) coth(wd) w^α−1 = w^α−1 + |d|^1−α (s coth s − 1) s^α−1‘ with ‘s = |d| w‘.
Step 2: the first part is ‘M_L‘.
Step 3: the substitution ‘s = |d| w‘.
For \(0{\lt}\alpha {\lt}1\), \(L{\gt}1\), \(x,x'\in [-1,1]\) with \(d=x-x'\ne 0\), \(K^{(L)}(x,x')-K^{(\infty )}(x,x')=T_1-T_2-T_3+T_4+T_5\) ??: insert ?? and ?? into ?? and use \(LM_L=L^{1+\alpha }(1-L^{-2\alpha })/\alpha \).
Step 1: the powers of ‘L‘ and ‘|d|‘ in terms of ‘A = L^α‘ and ‘B = |d|^α‘.
Step 2: clear denominators.
For \(0{\lt}\alpha \), \(L{\gt}1\) and \(x\in [-1,1]\), \(K^{(L)}(x,x)=1-\frac1L+T_5\) (Theorem 676; \(\Psi _L(0)=M_L\)).
For \(0{\lt}\alpha \) and \(L{\gt}1\), \(0\le T_3\le \frac1{L(1-L^{-2\alpha })}\) (from \(J_\alpha (S)\le S^\alpha /\alpha \)).
Step 1: ‘|d|^−α J_α(L|d|) ≤ |d|^−α (L|d|)^α / α = L^α / α‘.
Step 2: ‘1 / (L (1 − L^−2α)) = (L^α / α) / (L M_L)‘.
For \(0{\lt}\alpha \) and \(L{\gt}1\), \(T_3\le \frac{\alpha D_\alpha |d|^{-\alpha }}{L^{1+\alpha }(1-L^{-2\alpha })}\) (from \(J_\alpha \le D_\alpha \)).
For \(0{\lt}\alpha \), \(L{\gt}1\) and \(|d|\le 2\), \(0\le T_4\le \frac{4\alpha L^{-3-2\alpha }}{3(2+\alpha )(1-L^{-2\alpha })}\) (from \(I_\alpha (S)\le \frac{S^{2+\alpha }}{3(2+\alpha )}\)); this is sharper than the bound \(\frac{4\alpha L^{-2-2\alpha }}{3(2+\alpha )(1-L^{-2\alpha })}\) stated in Theorem 676.
Step 1: ‘|d|^−α I_α(|d|/L) ≤ |d|² L^−2−α / (3(2+α)) ≤ 4 L^−2−α / (3(2+α))‘.
Step 2: divide by ‘L M_L = L^1+α(1 − L^−2α)/α‘.
For \(0{\lt}\alpha {\lt}1\), \(L\ge 2\) and \(x,x'\in [-1,1]\), \(0\le T_5\le \frac{2^{1-\alpha }\alpha L^{-2}}{1-L^{-2\alpha }}+\frac{2e^{-L}}L\) (from ??).
Let \(K^{(\infty )}(x,x'):=1-c_\alpha |x-x'|\). For \(0{\lt}\alpha {\lt}1\), \(L\ge 2\) and \(x,x'\in [-1,1]\), \(|K^{(L)}(x,x')-K^{(\infty )}(x,x')|\le C_\alpha L^{-\min (1,2\alpha )}\) with \(C_\alpha =\frac1{1-2^{-2\alpha }}\bigl(2+\frac{2\alpha }{1+\alpha } +\frac{4\alpha }{3(2+\alpha )}+2^{1-\alpha }\alpha \bigr)+1\) ??: the sum of the bounds on \(T_1,\dots ,T_5\) with \(1-L^{-2\alpha }\ge 1-2^{-2\alpha }\), \(|d|\le 2\) and \(2e^{-L}\le 1\); on the diagonal \(K^{(L)}(x,x)-1=-\frac1L+T_5\).
Step 1: ‘1/(1 − L^−2α) ≤ Q := 1/(1 − 2^−2α)‘, ‘Q ≥ 1‘.
Step 2: ‘L^−k ≤ P := L^−min(1, 2α)‘ for ‘k ≥ min(1, 2α)‘.
Step 3: the five bounds in the form ‘Tᵢ ≤ cᵢ Q P‘.
Step 4: sum the bounds (off the diagonal) or use ‘K^(L)(x, x) = 1 − 1/L + T₅‘.
For \(0{\lt}\alpha {\lt}1\) and \(L\ge 2\), \(\sup _{x,x'\in [-1,1]}|K^{(L)}(x,x')-K^{(\infty )}(x,x')|\le C_\alpha L^{-\min (1,2\alpha )}\) ??.
For \(\varkappa \ne 0\) and \(x\in [-1,1]\), \(T[\sin \varkappa \, \cdot ](x)=-\frac2{\varkappa ^2}\sin \varkappa x +\frac{2\cos \varkappa }\varkappa x\) (proof of Theorem ??).
For \(\varkappa \ne 0\) and \(x\in [-1,1]\), \(T[\cos \varkappa \, \cdot ](x)=-\frac2{\varkappa ^2}\cos \varkappa x+\frac{2\sin \varkappa }\varkappa +\frac{2\cos \varkappa }{\varkappa ^2}\) (proof of Theorem ??).
(Theorem ??(iii), odd case.) If \(\varkappa {\gt}0\) and \(\cos \varkappa =0\), i.e. \(\varkappa \in {\mathcal{K}}^o=\{ (j+\frac12)\pi \} \), then \(K^{(\infty )}[\sin \varkappa \, \cdot ](x)=\frac{c_\alpha }{\varkappa ^2}\sin \varkappa x\) for \(x\in [-1,1]\).
(Theorem ??(iii), even case.) If \(\varkappa {\gt}0\) and \(\varkappa \sin \varkappa =\alpha \cos \varkappa \), i.e. \(\varkappa \tan \varkappa =\alpha \), then \(K^{(\infty )}[\cos \varkappa \, \cdot ](x)=\frac{c_\alpha }{\varkappa ^2}\cos \varkappa x\) for \(x\in [-1,1]\).
For \(\alpha {\gt}0\) and \(j\ge 0\) the equation \(\varkappa \tan \varkappa =\alpha \) has exactly one root \(\varkappa ^e_j\) in \((j\pi ,j\pi +\frac\pi 2)\): \(\varkappa \tan \varkappa \) is continuous and strictly increasing from \(0\) to \(+\infty \) there (proof of Theorem ??(iii)).
\(\varkappa ^e_j\in (j\pi ,j\pi +\frac\alpha {j\pi })\) for \(j\ge 1\) (Theorem ??(iii)).
\((\varkappa ^e_0)^2\le \alpha \) (Theorem ??(iii)).
(Theorem ??(iv), ??.) \(\frac{k\pi }2\le \varkappa _{(k)}\le \frac{(k+1)\pi }2\) for all \(k\ge 0\), and \(\varkappa _{(k)}\) is strictly increasing (interlacing \(\varkappa ^e_0{\lt}\frac\pi 2{\lt}\varkappa ^e_1{\lt}\frac{3\pi }2{\lt}\cdots \)).
For \(k\ge 1\), \(\frac{4c_\alpha }{\pi ^2(k+1)^2}\le \mu _k(K^{(\infty )})\le \frac{4c_\alpha }{\pi ^2k^2}\), \(c_\alpha =\frac\alpha {1+\alpha }\) ??, and \(\frac1{1+\alpha }\le \mu _0{\lt}1\). In particular \(\mu _k\asymp k^{-2}\), and the exponent does not depend on \(\alpha \).
(Theorem ??(i), injectivity.) If \(g\in C[-1,1]\) and \(K^{(\infty )}g=0\) on \([-1,1]\), then \(g=0\): \(K^{(\infty )}g=\langle g,1\rangle \mathbf1 -\frac{c_\alpha }2Tg\) with \((Tg)''=2g\), so differentiating twice gives \(-c_\alpha g=0\).
Step 1: ‘(K^(∞) g)’ = 0‘ on ‘(−1, 1)‘.
Step 2: ‘(K^(∞) g)” = −c_α g = 0‘ on ‘(−1, 1)‘.
Step 3: extend to ‘[−1, 1]‘ by continuity.
(Theorem ??(ii).) If \(\mu {\gt}0\) and \(g\in C[-1,1]\) satisfy \(K^{(\infty )}g=\mu g\) on \([-1,1]\), then \(g\) is twice differentiable on \((-1,1)\) and \(\mu \, g''=-c_\alpha \, g\) there ??. (The paper states \(g\in C^\infty [-1,1]\); the Lean statement records \(C^2\) on the open interval, which is what the exhaustion uses.)
Step 1: ‘g’ = (K^(∞) g)’/μ‘ on ‘(−1, 1)‘.
Step 2: ‘deriv g‘ coincides with ‘(K^(∞) g)’/μ‘ near every interior point, whose derivative is ‘−(c_α/μ) g‘.
Let \(\varkappa {\gt}0\), \(g\in C[-1,1]\) with \(g'\in C[-1,1]\), \(g''=-\varkappa ^2g\) on \((-1,1)\) and \(g'(\pm 1)=\mp \frac\alpha 2(g(1)+g(-1))\). Then \(g=A\sin \varkappa x+B\cos \varkappa x\) with \(A\cos \varkappa =0\) and \(B(\varkappa \sin \varkappa -\alpha \cos \varkappa )=0\) (proof of Theorem ??(iii)).
Step 1: the difference ‘h = g − A sin κx − B cos κx‘ solves the same equation.
Step 2: the energy ‘E = h’² + κ² h²‘ is constant on ‘(−1, 1)‘ and vanishes at ‘0‘.
Step 3: hence ‘h = 0‘ and ‘h’ = 0‘ on ‘(−1, 1)‘, and on ‘[−1, 1]‘ by continuity.
Step 4: the boundary conditions at ‘±1‘ give the constraints on ‘A‘, ‘B‘.
(Theorem ??(iii), exhaustion.) If \(\mu {\gt}0\) and \(g\in C[-1,1]\), \(g\not\equiv 0\), satisfy \(K^{(\infty )}g=\mu g\) on \([-1,1]\), then \(\mu =c_\alpha /\varkappa ^2\) for some \(\varkappa {\gt}0\) and either \(\cos \varkappa =0\) (\(\varkappa \in {\mathcal{K}}^o\)) and \(g=A\sin \varkappa x\), or \(\varkappa \tan \varkappa =\alpha \) (\(\varkappa \in {\mathcal{K}}^e\)) and \(g=B\cos \varkappa x\), with \(A,B\ne 0\). Hence the eigenvalues and eigenfunctions are exhausted by Theorem ??(iii) and all eigenvalues are simple.
Step 1: ‘g’ = (K^(∞) g)’/μ‘ on ‘(−1, 1)‘, continuous on ‘[−1, 1]‘, with ‘g” = −(c_α/μ) g = −κ² g‘.
Step 2: the boundary values ‘g’(±1) = ∓(α/2)(g(1) + g(−1))‘, from ‘μ(g(1) + g(−1)) = (1 − c_α) ∫ g‘ and ‘(K^(∞) g)’(±1) = ∓(c_α/2) ∫ g‘.
Step 3: ‘g = A sin κx + B cos κx‘ with ‘A cos κ = 0‘, ‘B(κ sin κ − α cos κ) = 0‘.
Step 4: ‘cos κ = 0‘ and ‘κ sin κ = α cos κ‘ cannot both hold, so exactly one of ‘A‘, ‘B‘ is nonzero.
(Theorem ??(i), positivity.) For \(g\in L^1(P_X)\), \(\iint K^{(\infty )}(x,x')g(x)g(x')\, P_X(\, \mathrm dx)P_X(\, \mathrm dx') =\int \Bigl(\int \operatorname {sgn}(x-t)g(x)\, P_X(\, \mathrm dx)\Bigr)^2\varpi (\, \mathrm dt)\ge 0\) (Remark 655 and Fubini).
Step 1: ‘sgn(x − t) sgn(x’ − t) g(x) g(x’)‘ is integrable on ‘P_X ⊗ P_X ⊗ ϖ‘.
Step 2: ‘K^(∞)(x, x’) g(x) g(x’) = ∫ sgn(x − t) sgn(x’ − t) g(x) g(x’) ϖ(dt)‘ for ‘P_X ⊗ P_X‘-a.e. ‘(x, x’)‘ (‘rem:sign-network‘).
Step 3: Fubini, and the inner integral factorizes into a square.
(Theorem ??(v).) \(\sum _k\mu _k=\operatorname {Tr}K^{(\infty )}=1\): the odd part is \(\sum _j\frac1{(j+1/2)^2\pi ^2}=\frac12\), the even part \(\sum _j(\varkappa ^e_j)^{-2}=\frac1\alpha +\frac12\) by the Hadamard factorization of \(\varkappa \sin \varkappa -\alpha \cos \varkappa \).
For bounded operators \(A,B\) on a real Hilbert space and all \(k\ge 0\), \(|\mu _k(A)-\mu _k(B)|\le \| A-B\| \) for the min–max values: the min–max principle and \(\langle Ag,g\rangle \le \langle Bg,g\rangle +\| A-B\| \) (proof of Theorem ??(a)).
The min–max values of \(K^{(\infty )}\) on \(L^2(P_X)\) are \(\mu _k(K^{(\infty )})=\frac{c_\alpha }{\varkappa _{(k)}^2}\), the eigenvalues of Theorem ?? in decreasing order (min–max principle for compact self-adjoint operators).
(Theorem ??(a).) If \(\| K^{(L)}-K^{(\infty )}\| \le {\varepsilon }_L\) on \(L^2(P_X)\), then \(|\mu _k(K^{(L)})-\mu _k(K^{(\infty )})|\le {\varepsilon }_L\) for all \(k\ge 0\).
(Theorem ??(b).) Let \(\mu _k(K^{(L)})\) satisfy \(|\mu _k(K^{(L)})-\mu _k(K^{(\infty )})|\le {\varepsilon }_L\) for all \(k\) (part (a)). Then for \(1\le k\le k_{\max }(L)=\lfloor \sqrt{2c_\alpha /(\pi ^2{\varepsilon }_L)}\rfloor -1\), equivalently \({\varepsilon }_L\le \frac{2c_\alpha }{\pi ^2(k+1)^2}\), \(\frac{2c_\alpha }{\pi ^2(k+1)^2}\le \mu _k(K^{(L)})\le \frac{6c_\alpha }{\pi ^2k^2}\) ??.
For \(u\in L^1(\varpi )\) and \(x\in [-\ell _\alpha ,\ell _\alpha ]\), \((S_\varpi u)(x)=\int \operatorname {sgn}(x-t)u(t)\, \varpi (\, \mathrm dt) =\frac1{2\ell _\alpha }\Bigl[\int _{-\ell _\alpha }^xu-\int _x^{\ell _\alpha }u\Bigr]\) (the integrand is \(u\) for \(t{\lt}x\) and \(-u\) for \(t{\gt}x\)).
For \(r\in L^1(P_X)\) and \(t\in [-1,1]\), \((S^*r)(t)=\int \operatorname {sgn}(x-t)r(x)\, P_X(\, \mathrm dx) =\frac12\Bigl[\int _t^1r-\int _{-1}^tr\Bigr]\); for \(t{\lt}-1\) it is \(\frac12\int _{-1}^1r\) and for \(t{\gt}1\) it is \(-\frac12\int _{-1}^1r\).
For \(u\in L^2(\varpi )\), \(h:=S_\varpi u\) is in \(H^1(-1,1)\) with weak derivative \(h'=u/\ell _\alpha \) on \((-1,1)\): \(h(x)-h(-1)=\frac1{\ell _\alpha }\int _{-1}^xu\) for \(x\in [-1,1]\) (Theorem ??(ii), \(\operatorname {ran}S_\varpi \subseteq H^1\)).
Both values are interval integrals over ‘[−ℓ_α, ℓ_α]‘ split at ‘x‘ and at ‘−1‘.
For \(f\in H^1(-1,1)\) with \(f'=g\) and the representer \(u_\varpi ^\dagger \) of ??, \((S_\varpi u_\varpi ^\dagger )(x)=f(x)\) for every \(x\in [-1,1]\): with \(a=(1+\alpha )c_0\), \(S_\varpi u_\varpi ^\dagger (x)=\frac1{2\ell _\alpha }\bigl[2a(\ell _\alpha -1) +\ell _\alpha (2f(x)-f(1)-f(-1))\bigr]=f(x)\) (Theorem ??(ii), \(H^1\subseteq \operatorname {ran}S_\varpi \)).
Split ‘∫_−ℓ^x‘ at ‘−1‘ and ‘∫_x^ℓ‘ at ‘1‘, and evaluate the four pieces.
\(K^{(\infty )}=S_\varpi S_\varpi ^*\): for \(r\in L^2(P_X)\) and \(P_X\)-a.e. \(x\), \((S_\varpi S_\varpi ^*r)(x)=\int _{-1}^1K^{(\infty )}(x,x')r(x')\, P_X(\, \mathrm dx')\) with \(K^{(\infty )}(x,x')=1-c_\alpha |x-x'|\) (Remark 655 and the integral representation of \(S_\varpi S_\varpi ^*\)).
\(\operatorname {ran}S_\varpi =H^1(-1,1)\) as classes in \(L^2(P_X)\): \(F\in L^2(P_X)\) is of the form \(S_\varpi u\) with \(u\in L^2(\varpi )\) if and only if \(F\) has a representative \(f\in H^1(-1,1)\). (\(S_\varpi u\) is the primitive of \(u/\ell _\alpha \); conversely \(f=S_\varpi u_\varpi ^\dagger \) with the representer ??.)
\(\ker S_\varpi =\{ u\in L^2(\varpi ):u=0\text{ a.e.\ on }(-1,1), \int _{-\ell _\alpha }^{-1}u=\int _1^{\ell _\alpha }u\} \). (\(S_\varpi u=0\) if and only if the \(H^1\) function \(h=S_\varpi u\) vanishes on \([-1,1]\), i.e. \(h'=u/\ell _\alpha =0\) a.e. on \((-1,1)\) and \(h(1)+h(-1)=\frac1{\ell _\alpha }[\int _{-\ell _\alpha }^{-1}u-\int _1^{\ell _\alpha }u]=0\); the first condition uses the Lebesgue differentiation theorem.)
Step 1: ‘S_ϖ u = 0‘ a.e. on ‘[−1, 1]‘, hence everywhere on ‘[−1, 1]‘ by continuity.
Step 2: the primitive ‘∫_−1^x u = ℓ_α (S_ϖ u (x) − S_ϖ u (−1))‘ vanishes on ‘[−1, 1]‘.
Step 3: ‘S_ϖ u (1) = (1/(2ℓ_α)) [∫_−ℓ^−1 u + ∫_−1^1 u − ∫_1^ℓ u] = 0‘.
The condition ‘u = 0‘ a.e. on ‘(−1, 1)‘, as a Lebesgue-a.e. statement on ‘[−1, 1]‘.
‘S_ϖ u (x) = (1/(2ℓ_α)) [∫_−ℓ^−1 u + ∫_−1^x u − ∫_x^1 u − ∫_1^ℓ u] = 0‘ on ‘[−1, 1]‘.
For \(f\in H^1(-1,1)\) the minimum-norm representer is \(S_\varpi ^\dagger f=u_\varpi ^\dagger \) ??: \(u_\varpi ^\dagger \) is a representer of \(f\) and is orthogonal to \(\ker S_\varpi \) (by (v)), hence is the unique representer in \((\ker S_\varpi )^\perp \) (Lemma 384).
For \(f\in H^1(-1,1)\) with \(c_0=\frac{f(1)+f(-1)}2\), \(\| S_\varpi ^\dagger f\| ^2_{L^2(\varpi )}=\| f\| ^2_{{\mathcal{H}}_\infty } =\frac{1+\alpha }\alpha \| f'\| ^2_{L^2(P_X)}+(1+\alpha )c_0^2\) ??. (\(\| u_\varpi ^\dagger \| ^2=\frac1{2\ell _\alpha }\bigl[\ell _\alpha ^2\int _{-1}^1f'{}^2\, \mathrm dx +\frac{2\ell _\alpha ^2c_0^2}{\ell _\alpha -1}\bigr]\) with \(\ell _\alpha =\frac{1+\alpha }\alpha \), \(\frac{\ell _\alpha }{\ell _\alpha -1}=1+\alpha \).)
\(\| f\| ^2_{{\mathcal{H}}_\infty }\le \frac{1+\alpha }\alpha \| f'\| ^2_{L^2(P_X)} +(1+\alpha )\| f\| _\infty ^2\), since \(|c_0|\le \| f\| _\infty \).
\(\| f\| ^2_{L^2(P_X)}\le 4\| f'\| ^2_{L^2(\, \mathrm dx)}+2c_0^2\), since \(|f(x)-c_0|\le \frac12\bigl(|f(x)-f(1)|+|f(x)-f(-1)|\bigr)\le \frac12\int _{-1}^1|f'| \le \frac1{\sqrt2}\| f'\| _{L^2(\, \mathrm dx)}\).
Pointwise bound on ‘[−1, 1]‘: ‘|f(x) − c₀| ≤ A/2‘, hence ‘f(x)² ≤ 2 (f x − c₀)² + 2c₀² ≤ A²/2 + 2c₀² ≤ G + 2c₀²‘.
Integrate the pointwise bound against the probability measure ‘P_X‘.
If \(f=K^{(\infty )}g_0\) with \(g_0\in L^2(P_X)\), then \(f\in H^2(-1,1)\) with \(f''=-c_\alpha g_0\), i.e. \(g_0=(K^{(\infty )})^{-1}f=-\frac{1+\alpha }\alpha f''\), and \(f'(1)=-\frac\alpha 2(f(1)+f(-1))\), \(f'(-1)=+\frac\alpha 2(f(1)+f(-1))\) ??. (\(f=S_\varpi v\) with \(v=S_\varpi ^*g_0\), so \(f'=v/\ell _\alpha \) by Theorem ??(ii) and \(v'=-g_0\); the boundary values come from \(v=\pm \frac12\int g_0\) outside \([-1,1]\).)
‘f’ = v/ℓ_α ∈ H¹‘ with derivative ‘−g₀/ℓ_α = −c_α g₀‘.
‘K^(∞) g₀ = S_ϖ (S_ϖ* g₀)‘ and ‘S_ϖ* g₀ = v‘ a.e.
Boundary values: ‘f’(±1) = v(±1)/ℓ_α = ∓ m/(2ℓ_α)‘ and ‘f(1) + f(−1) = (1/ℓ_α)[∫_−ℓ^−1 v − ∫_1^ℓ v] = ((ℓ_α − 1)/ℓ_α) m‘, ‘m = ∫_−1^1 g₀‘.
If \(f\in H^2(-1,1)\) with \(f'(1)=-\frac\alpha 2(f(1)+f(-1))\) and \(f'(-1)=+\frac\alpha 2(f(1)+f(-1))\), then \(f=K^{(\infty )}g_0\) for \(g_0:=-\frac{1+\alpha }\alpha f''\in L^2(P_X)\) ??. (\(S_\varpi ^*g_0=\ell _\alpha f'-\frac{\ell _\alpha }2(f'(1)+f'(-1))=\ell _\alpha f'\) on \((-1,1)\) and \(=\pm \frac{\ell _\alpha }2(f'(-1)-f'(1))=\pm (1+\alpha )c_0\) outside, i.e. \(S_\varpi ^*g_0=u_\varpi ^\dagger \), and \(S_\varpi u_\varpi ^\dagger =f\) by Theorem ??(ii).)
‘S* G₀ = S* (−ℓ_α h)‘ pointwise (the analysis map only sees the ‘P_X‘-class).
‘S* G₀ = u_ϖ†‘ on ‘(−ℓ, −1) ∪ (−1, 1) ∪ (1, ℓ)‘, hence ‘ϖ‘-a.e.
‘t ∈ (−ℓ_α, −1)‘: ‘S* G₀ (t) = −(ℓ_α/2)(f’(1) − f’(−1)) = (1+α) c₀‘.
‘t ∈ (−1, 1)‘: ‘S* G₀ (t) = ℓ_α f’(t) − (ℓ_α/2)(f’(1) + f’(−1)) = ℓ_α f’(t)‘.
‘t ∈ (1, ℓ_α)‘: ‘S* G₀ (t) = (ℓ_α/2)(f’(1) − f’(−1)) = −(1+α) c₀‘.
Let \((\varphi ,\nu )\) and \((\varphi ',\nu ')\) be two feature/hidden-law pairs on the same \(L^2(P_X)\) with kernels \(k_\nu \), \(k_{\nu '}\). If \(|k_\nu (x,x')-k_{\nu '}(x,x')|\le {\varepsilon }\) for \(P_X\)-a.e. \(x\) and \(P_X\)-a.e. \(x'\), then \(\| K_\nu -K_{\nu '}\| _{L^2(P_X)\to L^2(P_X)}\le {\varepsilon }\) (the proof of Lemma 754 only uses the bound almost everywhere).
For \(0{\lt}\alpha {\lt}1\) and \(L\ge 2\), \({\varepsilon }_L:=\| K^{(L)}-K^{(\infty )}\| _{L^2(P_X)\to L^2(P_X)} \le \sup _{x,x'\in [-1,1]}|K^{(L)}(x,x')-K^{(\infty )}(x,x')|\le C_\alpha L^{-\min (1,2\alpha )}\) ??, where \(K^{(L)}=S_{\nu _0^{(L)}}S^*_{\nu _0^{(L)}}\) is the kernel operator of the tanh feature and \(K^{(\infty )}=S_\varpi S_\varpi ^*\) that of the sign network (Remark 655). (Theorem 676, Remark 655 and Lemma 710, since \(P_X\)-a.e. \(x\) lies in \([-1,1]\).)
For \(0{\lt}\alpha {\lt}1\), \(\| K^{(L)}-K^{(\infty )}\| _{L^2(P_X)\to L^2(P_X)}\to 0\) as \(L\to \infty \) (\({\varepsilon }_L\le C_\alpha L^{-\min (1,2\alpha )}\to 0\), Theorem 711).
For \(0{\lt}\alpha {\lt}1\), \(\lambda {\gt}0\) and \(f\in L^2(P_X)\), \(Q^{(L)}_\lambda (f)\to Q^{(\infty )}_\lambda (f)\) as \(L\to \infty \), where \(Q^{(L)}_\lambda (f)=\langle f,(K^{(L)}+\lambda )^{-1}f\rangle \) and \(Q^{(\infty )}_\lambda (f)=\langle f,(K^{(\infty )}+\lambda )^{-1}f\rangle \) (Theorem 740 with \({\varepsilon }_L\to 0\), Theorem 712).
If \(f\in H^1(-1,1)\), then \(Q^{(\infty )}_\lambda (f)\to \| f\| ^2_{{\mathcal{H}}_\infty }\) as \(\lambda \downarrow 0\) (Lemma 749 and Theorem ??(ii),(iii): \(f\in \operatorname {ran}S_\varpi \) and \(\| S_\varpi ^\dagger f\| ^2=\| f\| ^2_{{\mathcal{H}}_\infty }\)).
If \(f\notin H^1(-1,1)\), then \(Q^{(\infty )}_\lambda (f)\to +\infty \) as \(\lambda \downarrow 0\) (Lemma 751 and Theorem ??(ii)).
For \(\lambda {\gt}0\) and \(f\in L^2(P_X)\), \(\lim _{L\to \infty }Q^{(L)}_\lambda (f)\) denotes the limit of \(L\mapsto Q^{(L)}_\lambda (f)\) along \(L\to \infty \) (Theorem 713: it exists and equals \(Q^{(\infty )}_\lambda (f)\)).
For \(0{\lt}\alpha {\lt}1\) and \(f\in L^2(P_X)\),
in the sense that \(\lambda \mapsto \lim _{L\to \infty }Q^{(L)}_\lambda (f)=Q^{(\infty )}_\lambda (f)\) (nonincreasing in \(\lambda \)) is bounded on \((0,\infty )\) if and only if \(f\) has an \(H^1(-1,1)\) representative. This is the precise meaning of “in the limit \(L\to \infty \), (R) is equivalent to \(f\in H^1\)”. (Theorems 713, 714, 715.)
For \(f\in H^1(-1,1)\), \(\lim _{\lambda \downarrow 0}\lim _{L\to \infty }Q^{(L)}_\lambda (f)=\| f\| ^2_{{\mathcal{H}}_\infty }\); for \(f\notin H^1(-1,1)\) the iterated limit is \(+\infty \).
Let \(0{\lt}\alpha {\lt}1\), \(\nu _0=\nu _0^{(L)}\) (\(L\ge 2\)), \((\lambda ,\beta )\) fixed and \(f\in H^1(-1,1)\), and let \(\rho ^*_L\) be the corresponding minimizer of the free energy. Then
??. (Corollary 760 with \(\| K^{(L)}-K^{(\infty )}\| \le {\varepsilon }_L\to 0\), Theorem 712, and Lemma 716. In Lean the reference law is \(\nu _0^{(L)}\) itself, so no smoothing \(\tilde\nu _0^{(L)}\) and no correction \(1/L\) of \({\varepsilon }_L\) is needed.)
Under the hypotheses of Corollary 720, with \(\nu ^*_L\) the hidden marginal of \(\rho ^*_L\) and \(\kappa =\lambda /\beta \), \(\limsup _L\operatorname {KL}(\nu ^*_L\| \nu _0^{(L)})\le \frac\kappa 2\| f\| ^2_{{\mathcal{H}}_\infty }\); the right-hand side does not depend on \(\lambda ,\beta \) (Corollary 761).
Under the hypotheses of Corollary 720, with \(m^*_L\) the conditional mean amplitude of \(\rho ^*_L\), \(\limsup _L\| m^*_L\| ^2_{L^2(\nu ^*_L)}\le \| f\| ^2_{{\mathcal{H}}_\infty }\): the bounds of Theorem 397 hold asymptotically with \(f\in H^1\) in place of (R) and \(\| f\| _{{\mathcal{H}}_\infty }\) in place of \(\| u_0^\dagger \| \) (Corollary 762).
Let \(K\ge 0\) be a bounded positive operator on a Hilbert space and \(x\in \overline{\operatorname {ran}K}\). Then \(\lambda (K+\lambda )^{-1}x\to 0\) as \(\lambda \downarrow 0\). (For \({\varepsilon }{\gt}0\) write \(x=(x-Kw)+Kw\) with \(\| x-Kw\| {\lt}{\varepsilon }\); then \(\| \lambda (K+\lambda )^{-1}(x-Kw)\| \le {\varepsilon }\) and \(\| \lambda (K+\lambda )^{-1}Kw\| \le \lambda \| w\| \), Lemma 197.)
Step 1: approximate ‘x‘ by ‘K w‘ within ‘ε/2‘.
Step 2: ‘λ R x = λ R (x − K w) + λ R (K w)‘ with the two bounds.
Step 3: ‘‖λ R x‖ ≤ ε/2 + λ‖w‖ < ε‘.
\(K^{(\infty )}\) is injective on \(L^2(P_X)\): if \(K^{(\infty )}r=0\) then \(\| S_\varpi ^*r\| ^2=\langle K^{(\infty )}r,r\rangle =0\), so \(v:=S^*r\) vanishes \(\varpi \)-a.e.; \(v\) is continuous on \([-1,1]\) with \(v'=-r\) (Lemma 698), so \(v=0\) on \([-1,1]\), \(\int _{-1}^xr=0\) for all \(x\), and \(r=0\) a.e. by Lebesgue differentiation.
Step 1: ‘S*_ϖ r = 0‘ in ‘L²(ϖ)‘, from ‘⟪K r, r⟫ = ‖S* r‖²‘.
Step 2: ‘v = S* r‘ vanishes Lebesgue-a.e. on ‘[−1, 1]‘.
Step 3: ‘v‘ is continuous on ‘[−1, 1]‘ (an ‘H¹‘ function), so ‘v = 0‘ on ‘[−1, 1]‘.
Step 4: ‘∫_−1^x r = v(−1) − v(x) = 0‘ for all ‘x ∈ [−1, 1]‘.
Step 5: ‘r = 0‘ a.e. on ‘(−1, 1)‘, hence in ‘L²(P_X)‘.
\(\overline{\operatorname {ran}K^{(\infty )}}=(\ker K^{(\infty )})^\perp =L^2(P_X)\), since \(K^{(\infty )}\) is self-adjoint and injective (Theorem 724).
If \(f=K^{(\infty )}g\) with \(g\in L^2(P_X)\), then \(R^{(\infty )}_\lambda f\to u^\dagger _\infty =S^*g\) uniformly on \(Z\) as \(\lambda \downarrow 0\). (\(K^{(\infty )}\) is self-adjoint and injective, so \(\overline{\operatorname {ran}K^{(\infty )}}=L^2(P_X)\) and \((K^{(\infty )}+\lambda )^{-1}f=g-\lambda (K^{(\infty )}+\lambda )^{-1}g\to g\) in \(L^2\) by Lemma 723; the bound \(\| S^*r\| _\infty \le \| r\| \) of Lemma ??(2) gives uniform convergence.)
If \(g=-\frac{1+\alpha }\alpha f''\) \(P_X\)-a.e. (i.e. \(f=K^{(\infty )}g\), ??), then \(u^\dagger _\infty (w,b)=-\frac{1+\alpha }{2\alpha }\int _{-1}^1\tanh (wx-b)f''(x)\, \, \mathrm dx\) ??: the limiting canonical ridgelet transform is the ridgelet transform, with the activation itself as filter, of the target sharpened by \(-\Delta \).
If \(f\in H^2(-1,1)\) satisfies the boundary conditions ?? and \(g:=-\frac{1+\alpha }\alpha f''\), then \(f=K^{(\infty )}g\), \(R^{(\infty )}_\lambda f\to u^\dagger _\infty \) uniformly on \(Z\) as \(\lambda \downarrow 0\), and \(u^\dagger _\infty (w,b)=-\frac{1+\alpha }{2\alpha }\int _{-1}^1\tanh (wx-b)f''(x)\, \mathrm dx\) (Theorems 709, 726, 727).
If \(f=K^{(\infty )}g\) with \(g\in L^2(P_X)\), then for \(L\ge 2\), \(\| S_{\nu _0^{(L)}}u^\dagger _\infty -f\| _{L^2(P_X)}\le {\varepsilon }_L\| g\| \) \(\le C_\alpha L^{-\min (1,2\alpha )}\| g\| \) ??, since \(S_{\nu _0^{(L)}}S^*g=K^{(L)}g\) and \(\| K^{(L)}g-K^{(\infty )}g\| \le {\varepsilon }_L\| g\| \) (Theorem 711).
For \(g\in L^2(P_X)\) and every \(L\), \(\| S^*g\| ^2_{L^2(\nu _0^{(L)})}=\langle K^{(L)}g,g\rangle \) ??.
If \(f=K^{(\infty )}g\), then \(\| u^\dagger _\infty \| ^2_{L^2(\nu _0^{(L)})}=\langle K^{(L)}g,g\rangle \to \langle f,g\rangle \) as \(L\to \infty \) ?? (\(|\langle (K^{(L)}-K^{(\infty )})g,g\rangle |\le {\varepsilon }_L\| g\| ^2\)).
If \(f=K^{(\infty )}g\) with \(f\in H^2(-1,1)\) satisfying ?? and \(g=-\frac{1+\alpha }\alpha f''\), then \(\langle f,g\rangle =\langle K^{(\infty )}g,g\rangle =\| S_\varpi ^*g\| ^2 =\| S_\varpi ^\dagger f\| ^2=\| f\| ^2_{{\mathcal{H}}_\infty }\) (Theorem 705; \(S_\varpi ^*g\) is the minimum-norm representer of \(f=S_\varpi S_\varpi ^*g\)).
For \(r\in L^1(P_X)\) and \(t\in {\mathbb R}\), \(\lim _{w\to +\infty }(S^*r)(w,wt)=\lim _{w\to +\infty }\int \tanh (w(x-t))r(x)\, P_X(\, \mathrm dx) =\int \operatorname {sgn}(x-t)r(x)\, P_X(\, \mathrm dx)=(S_\varpi ^*r)(t)\), and \(\lim _{w\to -\infty }(S^*r)(w,wt)=-(S_\varpi ^*r)(t)\) (dominated convergence of \(\tanh (w(x-t))\to \operatorname {sgn}(w)\operatorname {sgn}(x-t)\) for \(x\ne t\), with the dominating function \(|r|\)).
For \(f=K^{(\infty )}g\) and every \(t\in {\mathbb R}\), \(\lim _{w\to \pm \infty }u^\dagger _\infty (w,wt)=\pm (S_\varpi ^*g)(t)\); that is, \(u^\dagger _\infty (w,wt)\to \operatorname {sgn}(w)\, u^\dagger _\varpi (t)\), where \(S_\varpi ^*g=u^\dagger _\varpi \) is the representer ?? of the sign network (extended by the same formula to \(|t|{\gt}\ell _\alpha \); the explicit values are Theorem 735). (Lemma 733.)
For \(f\in H^2(-1,1)\) with the boundary conditions ?? and \(g=-\frac{1+\alpha }\alpha f''\): \((S_\varpi ^*g)(t)=\frac{1+\alpha }\alpha f'(t)\) for \(|t|{\lt}1\), and \((S_\varpi ^*g)(t)=\mp (1+\alpha )\frac{f(1)+f(-1)}2\) for \(t\gtrless \pm 1\); hence \(\lim _{w\to \pm \infty }u^\dagger _\infty (w,wt)=\pm \frac{1+\alpha }\alpha f'(t)\) for \(|t|{\lt}1\) and \(\mp \operatorname {sgn}(t)(1+\alpha )\frac{f(1)+f(-1)}2\) for \(|t|{\gt}1\). (\(\int _{-1}^1\operatorname {sgn}(x-t)f''\, \mathrm dx=f'(1)+f'(-1)-2f'(t)=-2f'(t)\) for \(|t|{\lt}1\), using \(f'(1)+f'(-1)=0\), and \(\int \operatorname {sgn}(x-t)f''=-\operatorname {sgn}(t)(f'(1)-f'(-1)) =\operatorname {sgn}(t)\alpha (f(1)+f(-1))\) for \(|t|{\gt}1\).)
‘|t| < 1‘: ‘½[−ℓ(g(1) − g(t)) + ℓ(g(t) − g(−1))] = ℓ g(t)‘ since ‘g(1) + g(−1) = 0‘.
‘t > 1‘: ‘−½ ∫_−1^1 (−ℓ h) = (ℓ/2)(g(1) − g(−1)) = −(1+α) c₀‘.
‘t < −1‘: ‘½ ∫_−1^1 (−ℓ h) = −(ℓ/2)(g(1) − g(−1)) = (1+α) c₀‘.
For \(f\in H^2(-1,1)\) with ?? and \(g=-\frac{1+\alpha }\alpha f''\): for \(|t|{\lt}1\), \(\lim _{w\to \pm \infty }u^\dagger _\infty (w,wt)=\pm \frac{1+\alpha }\alpha f'(t)\), and for \(|t|{\gt}1\), \(\lim _{w\to +\infty }u^\dagger _\infty (w,wt)=-\operatorname {sgn}(t)(1+\alpha )\frac{f(1)+f(-1)}2\) (Theorems 734, 735).
Let \(A_1,A_2,R_1,R_2\) be bounded operators on a Hilbert space and \(\lambda \in {\mathbb R}\) with \(R_1(A_1+\lambda )=\mathrm{id}\) and \((A_2+\lambda )R_2=\mathrm{id}\). Then \(R_1-R_2=R_1(A_2-A_1)R_2\).
Under the hypotheses of Lemma 737 with \(\lambda {\gt}0\) and \(\| R_1\| ,\| R_2\| \le 1/\lambda \), \(\| R_1-R_2\| \le \| R_1\| \, \| A_1-A_2\| \, \| R_2\| \le \lambda ^{-2}\| A_1-A_2\| \).
Under the hypotheses of Lemma 738, for every \(f\), \(|\langle f,R_1f\rangle -\langle f,R_2f\rangle |\le \lambda ^{-2}\| A_1-A_2\| \, \| f\| ^2\).
Let \(K_\nu =S_\nu S_\nu ^*\) and \(K_{\nu '}=S_{\nu '}S_{\nu '}^*\) be the kernel operators of two (feature map, hidden law) pairs on the same \(L^2(P_X)\), \(\lambda {\gt}0\) and \(f\in L^2(P_X)\). Then \(|Q^\nu _\lambda (f)-Q^{\nu '}_\lambda (f)|\le \lambda ^{-2}\| K_\nu -K_{\nu '}\| \, \| f\| ^2\). In the setting of Theorem ??, \(K_\nu =K^{(L)}\), \(K_{\nu '}=K^{(\infty )}\) and \(\| K^{(L)}-K^{(\infty )}\| \le {\varepsilon }_L\), this is \(|Q^{(L)}_\lambda (f)-Q^{(\infty )}_\lambda (f)|\le \lambda ^{-2}{\varepsilon }_L\| f\| ^2\). (Resolvent identity \((A+\lambda )^{-1}-(B+\lambda )^{-1}=(A+\lambda )^{-1}(B-A)(B+\lambda )^{-1}\) and \(\| (K+\lambda )^{-1}\| \le \lambda ^{-1}\).)
For \(\lambda {\gt}0\), \(f\in L^2(P_X)\) and two hidden laws \(\nu ,\nu '\in {\mathcal P}(Z)\), \(R_{\lambda ,\nu }f=S^*(K_\nu +\lambda )^{-1}f\) and \(R_{\lambda ,\nu '}f=S^*(K_{\nu '}+\lambda )^{-1}f\) are bounded functions on \(Z\) with \(\sup _{z\in Z}|R_{\lambda ,\nu }f(z)-R_{\lambda ,\nu '}f(z)|\le \lambda ^{-2}\| K_\nu -K_{\nu '}\| \, \| f\| \). With \(\nu =\nu _0^{(L)}\), \(\nu '\) the limit law and \(\| K^{(L)}-K^{(\infty )}\| \le {\varepsilon }_L\) this is ??. (\(\| S^*r\| _\infty \le \| r\| _{L^2(P_X)}\), Lemma ??(2), and Theorem 740.)
If \(\| K_{ u_L}-K_{ u'}\| o0\) along a filter in \(L\), then for every \(\lambda {\gt}0\) and \(f\in L^2(P_X)\), \(Q^{ u_L}_\lambda (f) o Q^{ u'}_\lambda (f)\) (Theorem efthm:order-of-limits-L-a).
For \(\lambda {\gt}0\) and \(f\in L^2(P_X)\), with \(g=(K_\nu +\lambda )^{-1}f\) and \(u_\lambda =S_\nu ^*g=R_{\lambda ,\nu }f\), \(Q^\nu _\lambda (f)=\langle K_\nu g,g\rangle +\lambda \| g\| ^2=\| u_\lambda \| ^2_{L^2(\nu )} +\lambda \| g\| ^2_{L^2(P_X)}\ge 0\). (\(f=(K_\nu +\lambda )g\) and \(\langle K_\nu g,g\rangle =\| S_\nu ^*g\| ^2\).)
For \(\lambda {\gt}0\), \(f\in L^2(P_X)\) and \(u_\lambda :=S_\nu ^*(K_\nu +\lambda )^{-1}f =R_{\lambda ,\nu }f\), one has \(S_\nu u_\lambda -f=-\lambda (K_\nu +\lambda )^{-1}f\) and \(\frac12\| S_\nu u_\lambda -f\| ^2+\frac\lambda 2\| u_\lambda \| ^2 =\frac\lambda 2\langle f,(K_\nu +\lambda )^{-1}(\lambda +K_\nu )(K_\nu +\lambda )^{-1}f\rangle =\frac\lambda 2Q^\nu _\lambda (f)\); by Lemma ??(4) this is the minimum of the Tikhonov functional \(u\mapsto \frac12\| S_\nu u-f\| ^2+\frac\lambda 2\| u\| ^2\) over \(L^2(\nu )\).
For \(\lambda {\gt}0\) and every \(u\in L^2(\nu )\) with \(S_\nu u=f\), \(Q^\nu _\lambda (f)\le \| u\| ^2_{L^2(\nu )}\); under (R), \(Q^\nu _\lambda (f)\le \| S_\nu ^\dagger f\| ^2\). (Take \(u\) in the Tikhonov minimum, Proposition 744: \(\frac\lambda 2Q^\nu _\lambda (f)\le \frac12\| S_\nu u-f\| ^2+\frac\lambda 2\| u\| ^2 =\frac\lambda 2\| u\| ^2\).)
For \(0{\lt}\lambda \le \lambda '\), \(Q^\nu _{\lambda '}(f)\le Q^\nu _\lambda (f)\): the map \(\lambda \mapsto Q^\nu _\lambda (f)\) is nonincreasing. (Without spectral theory: \(Q^\nu _\lambda (f)=\min _u[\lambda ^{-1}\| S_\nu u-f\| ^2+\| u\| ^2]\) by Proposition 744, and the functional is pointwise nonincreasing in \(\lambda \).)
Step 1: test the ‘λ’‘-Tikhonov minimum with ‘u_λ‘.
Step 2: ‘λ’⁻¹ ≤ λ⁻¹‘ and the identity for ‘Q_λ‘.
For \(0{\lt}\lambda '\le \lambda \), \(\| u_{\lambda '}-u_\lambda \| ^2_{L^2(\nu )}\le Q^\nu _{\lambda '}(f)-Q^\nu _\lambda (f)\). (The identity \(J_\lambda (u)=J_\lambda (u_\lambda )+\frac12\| S_\nu (u-u_\lambda )\| ^2 +\frac\lambda 2\| u-u_\lambda \| ^2\) of Lemma ??(4) with \(u=u_{\lambda '}\), \(J_\lambda (u_\lambda )=\frac\lambda 2Q^\nu _\lambda (f)\) and \(J_\lambda (u_{\lambda '})\le \frac\lambda {\lambda '}J_{\lambda '}(u_{\lambda '}) =\frac\lambda 2Q^\nu _{\lambda '}(f)\).)
‘½‖S u_λ’ − f‖² ≤ (λ/(2λ’))‖S u_λ’ − f‖²‘ since ‘λ’ ≤ λ‘.
Assume (R). Then \(u_\lambda =R_{\lambda ,\nu }f\to S_\nu ^\dagger f\) in \(L^2(\nu )\) as \(\lambda \downarrow 0\). (By Lemma 453, \(u_\lambda -u^\dagger =-\lambda (T_0+\lambda )^{-1}u^\dagger \); \(u^\dagger \in (\ker S_\nu )^\perp =\overline{\operatorname {ran}T_0}\), so for \({\varepsilon }{\gt}0\) write \(u^\dagger =(u^\dagger -T_0w)+T_0w\) with \(\| u^\dagger -T_0w\| {\lt}{\varepsilon }\); then \(\| \lambda (T_0+\lambda )^{-1}(u^\dagger -T_0w)\| \le {\varepsilon }\) and \(\| \lambda (T_0+\lambda )^{-1}T_0w\| \le \lambda \| w\| \).)
Step 1: approximate ‘u†‘ by ‘T₀ w‘ within ‘ε/2‘.
Step 2: the decomposition ‘u_λ − u† = d₁ + d₂‘ with ‘(T₀ + λ) d₁ = −λ(u† − T₀ w)‘ and ‘(T₀ + λ) d₂ = −λ T₀ w‘.
Step 3: ‘‖u_λ − u†‖ ≤ ε/2 + λ‖w‖ < ε‘.
Assume (R). Then \(Q^\nu _\lambda (f)\uparrow \| S_\nu ^\dagger f\| ^2_{L^2(\nu )}\) as \(\lambda \downarrow 0\): \(\lambda \mapsto Q^\nu _\lambda (f)\) is nonincreasing, bounded by \(\| S_\nu ^\dagger f\| ^2\), and converges to it. (\(Q^\nu _\lambda (f)=\langle f,g_\lambda \rangle =\langle S_\nu u^\dagger ,g_\lambda \rangle =\langle u^\dagger ,u_\lambda \rangle \) and \(u_\lambda \to u^\dagger \), Lemma 748.)
If \(\sup _{\lambda {\gt}0}Q^\nu _\lambda (f){\lt}\infty \), then \(f\in \operatorname {ran}S_\nu \). (Let \(\ell :=\sup _{\lambda {\gt}0}Q^\nu _\lambda (f)=\lim _{\lambda \downarrow 0}Q^\nu _\lambda (f)\). For \(0{\lt}\lambda '\le \lambda \le \lambda _0\), Lemma 747 gives \(\| u_{\lambda '}-u_\lambda \| ^2\le Q_{\lambda '}-Q_\lambda \le \ell -Q_{\lambda _0}\), so \(u_{1/n}\) is Cauchy in \(L^2(\nu )\) with limit \(u\); and \(\| S_\nu u_\lambda -f\| ^2\le \lambda Q_\lambda (f) \le \lambda \sup Q\to 0\) gives \(S_\nu u=f\). No weak compactness is needed.)
Step 1: the sequence ‘v n = u_1/(n+1)‘ is Cauchy.
For ‘0 < λ’ ≤ λ ≤ λ₀‘, ‘‖u_λ’ − u_λ‖² ≤ Q_λ’ − Q_λ ≤ ℓ − Q_λ₀ < 岑.
Step 2: ‘S_ν v n → f‘, since ‘‖S_ν v n − f‖² ≤ λ_n Q_λ_n(f) ≤ λ_n M → 0‘.
Step 3: ‘S_ν u = f‘ by uniqueness of limits.
If \(f\notin \operatorname {ran}S_\nu \), then \(Q^\nu _\lambda (f)\uparrow +\infty \) as \(\lambda \downarrow 0\): for every \(M\) there is \(\lambda _0{\gt}0\) with \(Q^\nu _\lambda (f)\ge M\) for all \(0{\lt}\lambda \le \lambda _0\). Together with Lemma 749, \(\lim _{\lambda \downarrow 0}Q^\nu _\lambda (f){\lt}\infty \) if and only if (R) holds for \(\nu \). (Otherwise \(Q^\nu _\lambda (f)\) is bounded on \((0,\infty )\) by monotonicity, and Lemma 750 gives \(f\in \operatorname {ran}S_\nu \).)
By monotonicity, ‘Q_λ(f) < M‘ for every ‘λ > 0‘.
For every \(\nu \in {\mathcal P}(Z)\) and \(f\in L^2(P_X)\), \(\lambda \mapsto Q^\nu _\lambda (f)\) is nonincreasing on \((0,\infty )\), and as \(\lambda \downarrow 0\), \(Q^\nu _\lambda (f)\to \| S_\nu ^\dagger f\| ^2_{L^2(\nu )}\) if \(f\in \operatorname {ran}S_\nu \), while \(Q^\nu _\lambda (f)\to +\infty \) otherwise. For the limit kernel \(K^{(\infty )}\) of Theorem ??, \(\| S_\nu ^\dagger f\| ^2=\| f\| ^2_{{\mathcal{H}}_\infty }\) and \(\operatorname {ran}S_\nu =H^1(-1,1)\) (Theorem ??).
For \(g\in L^2(P_X)\) and \(P_X\)-a.e. \(x\), \((K_\nu g)(x)=\int _{{\mathcal X}}k_\nu (x,x')g(x')\, P_X(\, \mathrm dx')\) with \(k_\nu (x,x')=\int _Z\varphi _z(x)\varphi _z(x')\, \nu (\, \mathrm dz)\), \(|k_\nu |\le 1\). (Fubini in \((K_\nu g)(x)=\int _Z\varphi _z(x)\int _{{\mathcal X}}\varphi _z(x')g(x')\, P_X(\, \mathrm dx')\, \nu (\, \mathrm dz)\), the integrand being dominated by \(|g(x')|\).)
Let \((\varphi ,\nu )\) and \((\varphi ',\nu ')\) be two feature/hidden-law pairs on the same \(L^2(P_X)\) with kernels \(k_\nu (x,x')=\int \varphi _z(x)\varphi _z(x')\, \nu (\, \mathrm dz)\) and \(k_{\nu '}\). If \(\sup _{x,x'}|k_\nu (x,x')-k_{\nu '}(x,x')|\le {\varepsilon }\), then \(\| K_\nu -K_{\nu '}\| _{L^2(P_X)\to L^2(P_X)}\le {\varepsilon }\). (By Lemma 753, \(|(K_\nu g-K_{\nu '}g)(x)|\le {\varepsilon }\| g\| _{L^1(P_X)}\le {\varepsilon }\| g\| _{L^2(P_X)}\) for a.e. \(x\), and \(P_X\) is a probability measure.) In the setting of Theorem 676, \(\| K^{(L)}-K^{(\infty )}\| \le \sup _{x,x'\in [-1,1]}|K^{(L)}(x,x')-K^{(\infty )}(x,x')| \le {\varepsilon }_L\).
Let \(u\in L^2(\nu _0)\) be arbitrary and \(\tilde\rho _u:=\rho _{\nu _0,u}=\nu _0\otimes {\mathcal N}(u,\beta /\lambda )\) the competitor of ??. Then \(\tilde\rho _u\in {\mathcal D}\), \(F_{\tilde\rho _u}=S_{\nu _0}u\), \(L(\tilde\rho _u)=\frac12\| S_{\nu _0}u-f\| ^2\), \(\operatorname {KL}(\tilde\rho _u\| \mu _U)=\frac\lambda {2\beta }\| u\| ^2_{L^2(\nu _0)}\) and \({\mathcal F}(\tilde\rho _u)=\frac12\| S_{\nu _0}u-f\| ^2+\frac\lambda 2\| u\| ^2_{L^2(\nu _0)}\) (Lemma ?? with \(\nu =\nu _0\)).
Step 1: ‘KL(ρ̃_u‖μ_U) = (κ/2)‖u‖²‘.
Step 2: ‘F_ρ̃_u = S_ν₀ u‘, so ‘L(ρ̃_u) = ½‖S_ν₀ u − f‖²‘.
Assume (A1), (A3), (A4), (A5) and \(f\in L^2(P_X)\). Let \(\rho ^*\) be the minimizer of \({\mathcal F}\) (Lemma ??). For every \(u\in L^2(\nu _0)\),
(In the proof of Theorem 397, take the competitor \(\tilde\rho _u=\nu _0\otimes {\mathcal N}(u,\beta /\lambda )\) for an arbitrary \(u\): \(L(\tilde\rho _u)=\frac12\| S_{\nu _0}u-f\| ^2\) and \(\operatorname {KL}(\tilde\rho _u\| \mu _U)=\frac\lambda {2\beta }\| u\| ^2\), Lemma 755; minimality \({\mathcal F}(\rho ^*)\le {\mathcal F}(\tilde\rho _u)\).)
Under the hypotheses of Proposition 756, the infimum of the right-hand side of ?? over \(u\in L^2(\nu _0)\) is attained at \(u=R_{\lambda ,\nu _0}f\) with value \(\frac\lambda 2Q^{\nu _0}_\lambda (f)\) (Proposition 744), and therefore \(L(\rho ^*)+\beta \, \operatorname {KL}(\rho ^*\| \mu _U)\le \frac\lambda 2Q^{\nu _0}_\lambda (f)\).
Under the hypotheses of Proposition 756, with \(\nu ^*\) the hidden marginal of \(\rho ^*\), \(m^*\) its conditional mean amplitude and \(\kappa =\lambda /\beta \),
and \(\| m^*\| ^2_{L^2(\nu ^*)}+\frac{2\beta }\lambda \operatorname {KL}(\nu ^*\| \nu _0)+\frac2\lambda L(\rho ^*) \le Q^{\nu _0}_\lambda (f)\). (The chain rule \(\operatorname {KL}(\rho ^*\| \mu _U)=\operatorname {KL}(\nu ^*\| \nu _0) +\frac\kappa 2\| m^*\| ^2_{L^2(\nu ^*)}\) of Theorem 396, whose proof uses only the Gaussian conditional structure of Theorem 225 and not (R), substituted into Proposition 757.)
The energy bound in the form ‘2 L(ρ*) + 2β KL(ν*‖ν₀) + λ‖m*‖² ≤ λ Q‘.
Let \((\nu _L)_L\) be hidden laws on \(Z\) whose kernel operators converge in operator norm, \(\| K_{\nu _L}-K_\infty \| \to 0\), to the kernel operator \(K_\infty =K_{\varphi ',\nu '}\) of another feature/hidden-law pair on the same \(L^2(P_X)\); let \((\lambda ,\beta )\) be fixed, \(f\in L^2(P_X)\), and \(\rho ^*_L\) the minimizer of the free energy with reference hidden law \(\nu _L\). Then
In the setting of Corollary 720, \(\nu _L=\tilde\nu _0^{(L)}\), \(K_\infty =K^{(\infty )}\) and \(Q^{(\infty )}_\lambda (f)\le \| f\| ^2_{{\mathcal{H}}_\infty }\) for \(f\in H^1(-1,1)\). (Proposition 757 for each \(L\) and Theorem 742.)
Assume (A1), (A3), (A4), (A5) and \(f\in L^2(P_X)\) (no (R)). With \(q:=Z_*/Z_0=\int e^{\kappa m^{*2}/2}\, \mathrm d\nu _0\),
??. (Theorem 226, the logarithm of the density integrated against \(\nu ^*\), and \(e^{\kappa m^{*2}/2}\ge 1\).)
The exponent identity ‘κ m*²/2 = s*²/(2λβ) = −W‘.
Step 1: the density.
Step 2: the relative entropy, by the Gibbs variational identity with ‘β = 1‘.
Step 3: ‘1 ≤ q‘.
Without (R), \(1\le q\le \exp \bigl(\frac\kappa 2\| m^*\| ^2_{L^2(\nu ^*)}\bigr) \le \exp \bigl(\frac\kappa 2Q^{\nu _0}_\lambda (f)\bigr)\): \(\log q=\frac\kappa 2\| m^*\| ^2_{L^2(\nu ^*)} -\operatorname {KL}(\nu ^*\| \nu _0)\le \frac\kappa 2\| m^*\| ^2_{L^2(\nu ^*)}\) (Lemma 763) and \(\| m^*\| ^2_{L^2(\nu ^*)}\le Q^{\nu _0}_\lambda (f)\) (Proposition 759).
Without (R), \(\| \nu ^*-\nu _0\| _{\mathcal M}=\| w-1\| _{L^1(\nu _0)} \le \sqrt{2\operatorname {KL}(\nu ^*\| \nu _0)}\) (Pinsker’s inequality, Lemma 850, in the convention \(\| \mu \| _{\mathcal M}=|\mu |(Z)=2\| \mu \| _{\mathrm{TV}}\)).
Assume (A1), (A3), (A4), (A5) and \(f\in L^2(P_X)\); let \(\rho ^*\) be the minimizer of \({\mathcal F}\) for \((\lambda ,\beta ,\nu _0)\), \(r^*=F_{\rho ^*}-f\), \(\nu ^*\) its hidden marginal, \(m^*\) its conditional mean amplitude and \(\kappa =\lambda /\beta \). For \(u\in L^2(\nu _0)\) put \(E(u):=\frac1\lambda \| S_{\nu _0}u-f\| ^2_{L^2(P_X)}+\| u\| ^2_{L^2(\nu _0)}\) ??. Then for every \(u\in L^2(\nu _0)\),
all terms on the right-hand side being nonnegative. If \(S_{\nu _0}u=f\) this is ??. (The strong-convexity identity ?? of Lemma 212 at the competitor \(\tilde\rho _u=\nu _0\otimes {\mathcal N}(u,\beta /\lambda )\), Lemma 755, with the chain rules \(\operatorname {KL}(\rho ^*\| \mu _U)=\operatorname {KL}(\nu ^*\| \nu _0)+\frac\kappa 2\| m^*\| ^2_{L^2(\nu ^*)}\) and \(\operatorname {KL}(\tilde\rho _u\| \rho ^*)=\operatorname {KL}(\nu _0\| \nu ^*)+\frac\kappa 2\| u-m^*\| ^2_{L^2(\nu _0)}\), Lemma 390; multiply by \(2/\lambda \).)
Step 1: the strong-convexity identity at the competitor ‘ρ̃_u‘.
Step 2: ‘KL(ρ*‖μ_U)‘ by the chain rule and ‘KL(ρ̃_u‖ρ*)‘ by the pair formula.
Step 3: rearrange and multiply by ‘2/λ‘.
For every \(u\in L^2(\nu _0)\),
(drop nonnegative terms in ??).
\(\| \nu ^*-\nu _0\| _{\mathcal M}\le \sqrt{\kappa \, Q^{\nu _0}_\lambda (f)}\) and \(1\le q\le e^{\kappa Q^{\nu _0}_\lambda (f)/2}\): Pinsker’s inequality \(\| \nu ^*-\nu _0\| _{\mathcal M}\le \sqrt{2\operatorname {KL}(\nu ^*\| \nu _0)}\) (Lemma 765) with \(\operatorname {KL}(\nu ^*\| \nu _0)\le \frac\kappa 2Q^{\nu _0}_\lambda (f)\) (Proposition 759), and Lemma 764.
For \(h\in L^2(P_X)\), \(S^*h=S^*_{\nu ^*}h\) in \(L^2(\nu ^*)\) and \(S_{\nu ^*}m^*=F_{\rho ^*}\) (Theorem 231), so \(\langle m^*,S^*h\rangle _{L^2(\nu ^*)}=\langle F_{\rho ^*},h\rangle =\langle f,h\rangle +\langle r^*,h\rangle \).
For \(h\in L^2(P_X)\), with \(u=S^*h\), \(|u|\le \| h\| \) (Lemma 192) and \(\| u\| ^2_{L^2(\nu ^*)}=\int u^2w\, \mathrm d\nu _0\le \| u\| ^2_{L^2(\nu _0)}+\| h\| ^2\| w-1\| _{L^1(\nu _0)}\), where \(\| u\| ^2_{L^2(\nu _0)}=\langle K_{\nu _0}h,h\rangle \).
Under the hypotheses of Lemma 766, for every \(h\in L^2(P_X)\),
where \(\| \nu ^*-\nu _0\| _{\mathcal M}=|\nu ^*-\nu _0|(Z)=\| w-1\| _{L^1(\nu _0)}\). (Expand \(\| m^*-u\| ^2_{L^2(\nu ^*)}\) for \(u=S^*h\); Lemma 769 for the cross term, Lemma 770 for \(\| u\| ^2_{L^2(\nu ^*)}\) and Lemma 767 with this \(u\) for \(\| m^*\| ^2_{L^2(\nu ^*)}\); then \(-\frac1\lambda \| r^*\| ^2-2\langle r^*,h\rangle =-\frac1\lambda \| r^*+\lambda h\| ^2 +\lambda \| h\| ^2\le \lambda \| h\| ^2\) and \(\frac1\lambda \| K_{\nu _0}h-f\| ^2+2\langle K_{\nu _0}h-f,h\rangle +\lambda \| h\| ^2 =\frac1\lambda \| (K_{\nu _0}+\lambda )h-f\| ^2\).)
Step 1: expand ‘‖m* − u‖²‘ and collect the three estimates.
Step 2: the two completions of the square.
\(\| m^*-S^*h\| ^2_{L^2(\nu _0)}\le q\, \| m^*-S^*h\| ^2_{L^2(\nu ^*)}\), since \(\, \mathrm d\nu _0=w^{-1}\, \mathrm d\nu ^*\) with \(w^{-1}=qe^{-\kappa m^{*2}/2}\le q\).
For \(h=(K_{\nu _0}+\lambda )^{-1}f\) (so that \(S^*h=R_{\lambda ,\nu _0}f\)), the first term of ?? vanishes and
since \(\| \nu ^*-\nu _0\| _{\mathcal M}\le \sqrt{\kappa Q^{\nu _0}_\lambda (f)}\) (Lemma 768) and \(\lambda \| (K_{\nu _0}+\lambda )^{-1}f\| ^2\le Q^{\nu _0}_\lambda (f)\) (Lemma 743).
For \(h\in L^2(P_X)\) and \(u=S^*h\) (so \(|u|\le \| h\| \)), \(\| \Pi \rho ^*-u\, \nu _0\| _{\mathcal M}\le \| m^*-u\| _{L^1(\nu ^*)}+\| h\| \, \| \nu ^*-\nu _0\| _{\mathcal M}\le \| m^*-u\| _{L^2(\nu ^*)}+\| h\| \, \| \nu ^*-\nu _0\| _{\mathcal M}\), since \(\Pi \rho ^*=m^*\nu ^*\) (Theorem 229) and \(m^*\nu ^*-u\nu _0=(m^*-u)\nu ^*+u(w-1)\nu _0\).
Step 1: ‘Πρ* = m* ν*‘ and ‘u ν* = (w u) ν₀‘.
Step 2: ‘Πρ* − u ν₀ = (m* − u) ν* + (w − 1) u ν₀‘.
Step 3: the two total variations.
Let \(\nu _0\in {\mathcal P}(Z)\), \(K_\infty =K_{\varphi ',\nu '}\) a second kernel operator on \(L^2(P_X)\), \(\delta :=\| K_{\nu _0}-K_\infty \| \), \(f=K_\infty g\), \(G:=\| g\| \), \(H:=\langle f,g\rangle \) and \(u^\dagger _\infty :=S^*g\). Then \(S_{\nu _0}u^\dagger _\infty -f=(K_{\nu _0}-K_\infty )g\), \(\| u^\dagger _\infty \| ^2_{L^2(\nu _0)}=\langle K_{\nu _0}g,g\rangle \) and
?? (the minimum ?? and \(\langle K_{\nu _0}g,g\rangle -H=\langle (K_{\nu _0}-K_\infty )g,g\rangle \le \eta G\)).
Let \(\nu _0\in {\mathcal P}(Z)\) satisfy (A4), \(K_\infty =K_{\varphi ',\nu '}\), \(\delta :=\| K_{\nu _0}-K_\infty \| _{L^2(P_X)\to L^2(P_X)}\), \(f=K_\infty g\), \(G:=\| g\| \), \(u^\dagger _\infty :=S^*g\); let \(\rho ^*\), \(\nu ^*\), \(m^*\), \(q\) be the quantities of the minimizer for \((\lambda ,\beta ,\nu _0)\) and \(Q_\lambda :=Q^{\nu _0}_\lambda (f)\). Then (Amplitude.)
(Lemma 771 with \(h=g\): \((K_{\nu _0}+\lambda )g-f=(K_{\nu _0}-K_\infty )g+\lambda g\) has norm at most \((\delta +\lambda )G\), and \(\frac1\lambda (\delta +\lambda )^2 =(\delta /\sqrt\lambda +\sqrt\lambda )^2\); Lemma 768 for \(\| \nu ^*-\nu _0\| _{\mathcal M}\) and \(q\); Lemma 772.)
‘(K_ν₀ + λ) g − f = (K_ν₀ − K_∞) g + λ g‘, of norm at most ‘(δ + λ) G‘.
(Realization, hidden marginal.) Under the hypotheses of Theorem 776, with \(H:=\langle f,g\rangle \) and \(\eta :=\| (K_{\nu _0}-K_\infty )g\| \le \delta G\),
where \(E(u^\dagger _\infty )=\frac{\eta ^2}\lambda +\langle K_{\nu _0}g,g\rangle \) (Lemma 767 with \(u=u^\dagger _\infty \) and Theorem 775).
(Coefficient measure.) Under the hypotheses of Theorem 776, \(\| \Pi \rho ^*-u^\dagger _\infty \nu _0\| _{\mathcal M}\le \| m^*-u^\dagger _\infty \| _{L^2(\nu ^*)} +G\sqrt{\kappa E(u^\dagger _\infty )}\), with \(E(u^\dagger _\infty )=\frac{\eta ^2}\lambda +\langle K_{\nu _0}g,g\rangle \) (Lemma 774 with \(h=g\), \(|u^\dagger _\infty |\le G\), and \(\| \nu ^*-\nu _0\| _{\mathcal M}\le \sqrt{2\operatorname {KL}(\nu ^*\| \nu _0)}\le \sqrt{\kappa E(u^\dagger _\infty )}\)).
Under the hypotheses of Theorem 776, for the choice \(\lambda =\delta \) one has \(E(u^\dagger _\infty )\le H+2\delta G^2\) and
and if \(\kappa (H+2\delta G^2)\le 1\) the same bound times \(e^{1/4}\) holds in \(L^2(\nu _0)\) ??. (For \(\lambda =\delta \), \((\delta /\sqrt\lambda +\sqrt\lambda )^2=4\delta \), \(E(u^\dagger _\infty )\le H+2\delta G^2\), \(\sqrt{a+b}\le \sqrt a+\sqrt b\) and \(e^{\kappa Q_\lambda /2}\le e^{1/2}\).)
Step 1: ‘Q_λ ≤ H + 2δG²‘ for ‘λ = δ‘.
Step 2: ‘(δ/√λ + √λ)² = 4λ‘ and ‘√(κ Q_λ) ≤ √(κ(H + 2δG²))‘.
Step 3: ‘√(G²(4λ + S)) ≤ G(2√λ + √S)‘.
Step 4: ‘e^κ Q_λ/2 ≤ e^1/2‘ when ‘κ(H + 2δG²) ≤ 1‘.
(Joint schedule.) Let \((\nu _L)_L\) be reference measures with \(\delta _L:=\| K_{\nu _L}-K_\infty \| \to 0\) (for \(\nu _L=\tilde\nu _0^{(L)}\), \(\delta _L\le (C_\alpha +1)L^{-\min (1,2\alpha )}\)), \(f=K_\infty g\), \(u^\dagger _\infty =S^*g\), and let \((\lambda _L,\beta _L)\) satisfy \(\lambda _L\to 0\), \(\kappa _L=\lambda _L/\beta _L\to 0\) and \(\delta _L^2/\lambda _L\to 0\). Then, for the minimizers \(\rho ^*_L\) of the free energies for \((\lambda _L,\beta _L,\nu _L)\), \(m^*_L\to u^\dagger _\infty \) in \(L^2(\nu ^*_L)\) and in \(L^2(\nu _L)\), \(\| \Pi \rho ^*_L-u^\dagger _\infty \nu _L\| _{\mathcal M}\to 0\), \(F_{\rho ^*_L}\to f\) in \(L^2(P_X)\) and \(\operatorname {KL}(\nu ^*_L\| \nu _L)+\operatorname {KL}(\nu _L\| \nu ^*_L)\to 0\). (Insert the schedule into Theorems 776, 777 and 778: \(Q_{\lambda _L}\) and \(E(u^\dagger _\infty )\) stay bounded by Theorem 775, and every right-hand side tends to zero.)
Step 1: ‘0 ≤ Q_L ≤ E_L ≤ Bnd_L → H‘, so ‘κ_L Q_L → 0‘, ‘κ_L E_L → 0‘.
Step 2: the right-hand side of Theorem (a) tends to zero.
Step 3: the coefficient measure, the realization and the relative entropies.
Let \(A\ge 0\) be a positive operator on a real inner product space, \(\lambda {\gt}0\), \(s\ge 0\) and \(x:=4s/\lambda \). Then for every \(z\), \(4\lambda \| (A+s)z\| ^2\le (1+x)\langle (A+s)(A+\lambda )^2z,z\rangle \). (The scalar identity \((1+x)(t+\lambda )^2-4\lambda (t+s)=(t-\lambda )^2+x\, t(t+2\lambda )\) gives \((1+x)\langle (A+s)(A+\lambda )^2z,z\rangle -4\lambda \| (A+s)z\| ^2 =\langle (A+s)v,v\rangle +x\langle (A+s)A(A+2\lambda )z,z\rangle \ge 0\) with \(v=(A-\lambda )z\), all operators being polynomials in \(A\) with nonnegative coefficients or squares.)
‘‖(A + s)z‖²‘ and the square ‘v = (A − λ)z‘.
The remaining polynomial in ‘A‘ has nonnegative coefficients.
Let \(\nu \in {\mathcal P}(Z)\) and \(f\in L^2(P_X)\) with \(Q^\nu _s(f)\le H_f\) for all \(s{\gt}0\) (for \(f=K_\nu ^{1/2}\tilde h\) one may take \(H_f=\| \tilde h\| ^2\)). Then for every \(\lambda {\gt}0\), \(\| (K_\nu +\lambda )^{-1}f\| ^2\le \frac{H_f}{4\lambda }\). (With \(z:=(K_\nu +\lambda )^{-1}(K_\nu +s)^{-1}f\) one has \((K_\nu +\lambda )^{-1}f=(K_\nu +s)z\) and \(Q^\nu _s(f)=\langle (K_\nu +s)(K_\nu +\lambda )^2z,z\rangle \), so Lemma 781 gives \(4\lambda \| (K_\nu +\lambda )^{-1}f\| ^2 \le (1+4s/\lambda )Q^\nu _s(f)\le (1+4s/\lambda )H_f\) for every \(s{\gt}0\); let \(s\downarrow 0\).)
Step 1: for every ‘s > 0‘, ‘4λ‖(K + λ)⁻¹ f‖² ≤ (1 + 4s/λ) Q_s(f)‘.
Step 2: let ‘s ↓ 0‘.
Let \(\nu _0\), \(\delta =\| K_{\nu _0}-K_\infty \| \) be as in Theorem 776 and let \(f\in L^2(P_X)\) satisfy \(Q^{(\infty )}_s(f)\le H_f\) for all \(s{\gt}0\) (for the sign network and \(f\in H^1(-1,1)\), \(H_f=\| f\| ^2_{{\mathcal{H}}_\infty }\) by the Sobolev characterization of the limit RKHS); put \(h_\lambda :=(K_\infty +\lambda )^{-1}f\). Then \(\| h_\lambda \| \le \frac{\sqrt{H_f}}{2\sqrt\lambda }\) and
(Lemma 782; the resolvent identity gives \(Q^{\nu _0}_\lambda (f)-Q^{(\infty )}_\lambda (f)=\langle (K_{\nu _0}+\lambda )^{-1}f, (K_\infty -K_{\nu _0})h_\lambda \rangle \le \sqrt{Q^{\nu _0}_\lambda (f)/\lambda }\; \delta \, \frac{\sqrt{H_f}}{2\sqrt\lambda }\), and \(x\le A+c\sqrt{Ax}\) implies \(\sqrt x\le (1+c)\sqrt A\).)
Step 1: the resolvent identity, ‘Q^ν₀_λ − Q^∞_λ = ⟪(K_ν₀ + λ)⁻¹ f, (K_∞ − K_ν₀) h_λ⟫‘.
Step 2: with ‘x = √Q^ν₀_λ‘, ‘a = √Hf‘, ‘c = δ/(2λ)‘: ‘x² ≤ a² + c a x‘, hence ‘x ≤ (1 + c) a‘.
(Realization and hidden marginal, unconditionally.) Under the hypotheses of Proposition 783, for the minimizer \(\rho ^*\) for \((\lambda ,\beta ,\nu _0)\),
Hence \(F_{\rho ^*_L}\to f\) in \(L^2(P_X)\) along \(\lambda _L\to 0\), \(\delta _L^2/\lambda _L\to 0\) for every \(\beta \), and the symmetric relative entropy tends to zero if moreover \(\kappa _L\to 0\). (Proposition 759, Lemma 767 with \(u=R_{\lambda ,\nu _0}f\), and (a).)
‘λ(1 + δ/(2λ))² = (√λ + δ/(2√λ))²‘.
(Amplitude, high-temperature regime.) Under the hypotheses of Proposition 783, with \(R^{(\infty )}_\lambda f=S^*h_\lambda \),
so that \(\| m^*_L-R^{(\infty )}_{\lambda _L}f\| _{L^2(\nu ^*_L)}\to 0\) whenever \(\delta _L/\lambda _L\to 0\) and \(\sqrt{\kappa _L}/\lambda _L\to 0\). (Lemma 771 with \(h=h_\lambda \): \((K_{\nu _0}+\lambda )h_\lambda -f=(K_{\nu _0}-K_\infty )h_\lambda \) has norm at most \(\delta \| h_\lambda \| \); then (a).)
Under the hypotheses of Proposition 784, along a sequence \((\nu _L,\lambda _L,\beta _L)\) with \(\lambda _L\to 0\) and \(\delta _L^2/\lambda _L\to 0\), \(F_{\rho ^*_L}\to f\) in \(L^2(P_X)\) for every choice of \(\beta _L\); if moreover \(\kappa _L(1+\delta _L/(2\lambda _L))^2\to 0\) (e.g. \(\kappa _L\to 0\) with \(\delta _L/\lambda _L\) bounded) then \(\operatorname {KL}(\nu ^*_L\| \nu _L)+\operatorname {KL}(\nu _L\| \nu ^*_L)\to 0\).
‘‖r*_L‖ ≤ √Hf (√λ_L + √(δ_L²/λ_L)/2) → 0‘.
Under the hypotheses of Proposition 785, along a sequence with \(\delta _L/\lambda _L\to 0\) and \(\sqrt{\kappa _L}/\lambda _L=(\lambda _L\beta _L)^{-1/2}\to 0\), \(\| m^*_L-R^{(\infty )}_{\lambda _L}f\| _{L^2(\nu ^*_L)}\to 0\) (?? with \(Q^{\nu _L}_{\lambda _L}(f)\le H_f(1+\delta _L/(2\lambda _L))^2\) bounded).
The bound ‘(Hf/4)(δ/λ)² + (Hf√Hf/4)(√κ/λ)(1 + δ/(2λ))‘, which tends to zero.