7 Convergence rates for the regularization limit
Section 10 of the paper: the learned-base-measure problem rewritten as a weighted Tikhonov problem over the reference measure, the pointwise deviation of the weight, the sup-norm bound of order \(\lambda ^{-1/2}\), the Tikhonov rate under a source condition (of order \(1/2\) or \(1\)), the rate theorem, the rates under the uniform sup-norm bound and in the high-temperature regime, and the finite-dimensional case.
7.1 Rewriting as a weighted Tikhonov problem
For \(\lambda {\gt}0\), \(u_\lambda =S_{\nu _0}^*(K_{\nu _0}+\lambda )^{-1}f\) satisfies \((T_0+\lambda )u_\lambda =S_{\nu _0}^*f\) (Lemma ??(4)).
If \(f=K_{\nu _0}g=S_{\nu _0}(S_{\nu _0}^*g)\) with \(g\in L^2(P_X)\), then (R) holds and \(u^\dagger =S_{\nu _0}^*g\), since \(S_{\nu _0}^*g\in \operatorname {ran}S_{\nu _0}^* \subseteq (\ker S_{\nu _0})^\perp \) (Lemma 384).
If \(f=S_{\nu _0}T_0g_0\) with \(g_0\in L^2(\nu _0)\), then (R) holds and \(u^\dagger =T_0g_0\), since \(T_0g_0=S_{\nu _0}^*(S_{\nu _0}g_0)\in (\ker S_{\nu _0})^\perp \).
(SC\(_a\)) implies (R): \(f=S_{\nu _0}u^\dagger \).
Assume (A1), (A3) and (R). Then \(u_\lambda -u^\dagger =-\lambda (T_0+\lambda )^{-1}u^\dagger \), i.e. \((T_0+\lambda )(u_\lambda -u^\dagger )=-\lambda u^\dagger \). (\(f=S_{\nu _0}u^\dagger \) gives \(S_{\nu _0}^*f=T_0u^\dagger \), so \((T_0+\lambda )u_\lambda -(T_0+\lambda )u^\dagger =T_0u^\dagger -(T_0+\lambda )u^\dagger =-\lambda u^\dagger \).)
Under (SC\(_{1/2}\)), \(f=K_{\nu _0}g\) with \(g\in L^2(P_X)\): \(\| u_\lambda -u^\dagger \| _{L^2(\nu _0)}\le \frac{\sqrt\lambda }2\| g\| \le \sqrt\lambda \| g\| \) and \(\| S_{\nu _0}(u_\lambda -u^\dagger )\| _{L^2(P_X)}\le \lambda \| g\| \). (With \(y:=(T_0+\lambda )^{-1}S_{\nu _0}^*g\) one has \(u_\lambda -u^\dagger =-\lambda y\), and the normal equation \((T_0+\lambda )y=S_{\nu _0}^*g\) gives \(\| S_{\nu _0}y\| ^2+\lambda \| y\| ^2\le \| g\| \, \| S_{\nu _0}y\| \), hence \(\lambda \| y\| ^2\le \| g\| ^2/4\) and \(\| S_{\nu _0}y\| \le \| g\| \).)
‘y := −λ⁻¹ d‘ solves the normal equation ‘(T₀ + λ) y = S_ν₀* g‘.
‘‖d‖² = λ²‖y‖² ≤ λ‖g‖²/4‘.
Under (SC\(_1\)), \(u^\dagger =T_0g_0\) with \(g_0\in L^2(\nu _0)\): \(\| u_\lambda -u^\dagger \| _{L^2(\nu _0)}\le \lambda \| g_0\| \) and \(\| S_{\nu _0}(u_\lambda -u^\dagger )\| _{L^2(P_X)}\le \lambda \| g_0\| \). (\(u_\lambda -u^\dagger =-\lambda (T_0+\lambda )^{-1}T_0g_0\) and \(\| T_0(T_0+\lambda )^{-1}\| \le 1\); the residual bound follows from \(\| S_{\nu _0}\| \le 1\).)
Assume (A1), (A3) and (SC\(_a\)) (hence (R)). Then
(For \(a\in \{ 1/2,1\} \): Lemmas 454 and 455, with \(\lambda ^{1/2}=\sqrt\lambda \) and \(\min (a+1/2,1)=1\).)
With \(v:=w\, m^*\): \(v\) is bounded, \(v\in L^2(\nu _0)\) and \(S_{\nu _0}v=S_{\nu ^*}m^*=F_{\rho ^*}\) (by \(|m^*\varphi _z|\le \| m^*\| _\infty \) and Fubini, \(S_{\nu ^*}m^*(x)=\int m^*(z)\varphi _z(x)w(z)\nu _0(\, \mathrm dz)=(S_{\nu _0}v)(x)\), which is \(F_{\rho ^*}\) by Theorem 231).
\(\Pi \rho ^*=m^*\nu ^*=v\, \nu _0\) and \(\| m^*\| ^2_{L^2(\nu ^*)}=\int m^{*2}w\, \mathrm d\nu _0=\int v^2w^{-1}\, \mathrm d\nu _0\).
(Weighted normal equation.) With \(M_{w^{-1}}\) the multiplication operator by \(w^{-1}\), in \(L^2(\nu _0)\) and pointwise, \((T_0+\lambda M_{w^{-1}})\, v=S_{\nu _0}^*f\), where \(w^{-1}v=m^*\). (By Theorem 228, \(s^*=S^*r^*=-\lambda m^*\) for every \(z\); since \(r^*=F_{\rho ^*}-f=S_{\nu _0}v-f\), \(S^*r^*=T_0v-S_{\nu _0}^*f\), hence \(T_0v-S_{\nu _0}^*f=-\lambda m^*=-\lambda w^{-1}v\).)
(Variational interpretation.) \(v\) is the unique minimizer of the weighted Tikhonov functional \({\mathcal{J}}_w(u):=\tfrac 12\| S_{\nu _0}u-f\| ^2_{L^2(P_X)}+\frac\lambda 2\int _Zu^2w^{-1}\, \mathrm d\nu _0\) (\(u\in L^2(\nu _0)\)): for every \(u\), \({\mathcal{J}}_w(u)={\mathcal{J}}_w(v)+\tfrac 12\| S_{\nu _0}(u-v)\| ^2+\frac\lambda 2\int (u-v)^2w^{-1}\, \mathrm d\nu _0\), since the cross terms vanish by the weighted normal equation.
Integrability of the weighted squares and of the cross term.
Step 1: the weighted integral splits, since ‘v w⁻¹ = m‘.
Step 2: the cross term is ‘⟪d, T₀ v − S* f + λ m⟫ = 0‘ by the normal equation.
Step 3: expand the square of the residual.
(Difference from \(u_\lambda \).) \((T_0+\lambda )(v-u_\lambda )=\lambda (1-w^{-1})v=\lambda (w-1)m^*\) and \(\| v-u_\lambda \| _{L^2(\nu _0)}\le \| (w-1)m^*\| _{L^2(\nu _0)}=\Xi \). (Subtract \((T_0+\lambda )u_\lambda =S_{\nu _0}^*f\) from the weighted normal equation, and use \(\| \lambda (T_0+\lambda )^{-1}\| \le 1\).)
Assume the hypotheses of Lemma 459 and (R). Then \(w=q^{-1}e^{\zeta }\) with \(q=Z_*/Z_0=\tilde Z\), and pointwise
(\(1\le q\le e^{\zeta _0}\); \(e^{\zeta }-1\le \zeta e^{\zeta }\); \(1-q^{-1}\le \log q\le \zeta _0\); \(1-e^{-\zeta }\le \zeta \).)
‘e^−ζ₀ ≤ 1/Z̃‘ since ‘Z̃ ≤ e^ζ₀‘.
‘1 − 1/Z̃ ≤ log Z̃ ≤ ζ₀‘.
‘w⁻¹ = Z̃ e^−ζ‘.
Pointwise, \(|(w-1)m^*|(z)\le \frac\kappa 2|m^*(z)|^3e^{\kappa m^*(z)^2/2} +\frac\kappa 2\| u^\dagger \| ^2|m^*(z)|\) (insert \(\zeta =\frac\kappa 2m^{*2}\), \(\zeta _0=\frac\kappa 2\| u^\dagger \| ^2\) into the bound on \(|w-1|\) and multiply by \(|m^*|\)).
\(\Xi \le \frac\kappa 2\bigl(\| m^{*3}e^{\kappa m^{*2}/2}\| _{L^2(\nu _0)} +\| u^\dagger \| ^2\, \| m^*\| _{L^2(\nu _0)}\bigr)\) (triangle inequality in \(L^2(\nu _0)\) applied to the pointwise bound of Lemma 463).
Step 1: ‘Ξ ≤ ‖c₁ φ₁ + c₂ φ₂‖‘ by pointwise domination.
Step 2: the triangle inequality and ‘‖ |m*| ‖ = ‖m*‖‘.
\(\| m^*\| _{L^2(\nu _0)}\le q^{1/2}\| m^*\| _{L^2(\nu ^*)} \le e^{\zeta _0/2}\| u^\dagger \| \) (\(\int m^{*2}\, \mathrm d\nu _0=\int m^{*2}w^{-1}\, \mathrm d\nu ^*\le q\| m^*\| ^2_{L^2(\nu ^*)}\), \(\| m^*\| _{L^2(\nu ^*)}\le \| u^\dagger \| \) and \(q\le e^{\zeta _0}\)).
For every \(\nu \in {\mathcal P}(Z)\) and \(f\in \operatorname {ran}S_\nu \), \(\sup _{z\in Z}|R_{\lambda ,\nu }f(z)|\le \| S_\nu ^\dagger f\| _{L^2(\nu )}/(2\sqrt\lambda )\); more generally the bound holds with \(\| h\| _{L^2(\nu )}\) for every representer \(h\) of \(f\). (With \(g:=(K_\nu +\lambda )^{-1}f\), \(|R_{\lambda ,\nu }f(z)|\le \| g\| \) and \(\| S_\nu ^*g\| ^2+\lambda \| g\| ^2=\langle (K_\nu +\lambda )g,g\rangle =\langle h,S_\nu ^*g\rangle \le \| h\| \| S_\nu ^*g\| \), hence \(\lambda \| g\| ^2\le \| h\| ^2/4\).)
For \(m^*=R_{\lambda ,\nu ^*}f\),
(\(h:=u^\dagger /w\in L^2(\nu ^*)\) with \(\| h\| ^2_{L^2(\nu ^*)}=\int u^{\dagger 2}w^{-1}\, \mathrm d\nu _0 \le q\| u^\dagger \| ^2\) and \(S_{\nu ^*}h=S_{\nu _0}u^\dagger =f\); apply Lemma 466.)
Step 1: ‘h := u†/w ∈ L²(ν*)‘, via ‘∫ w (u†/w)² dν₀ = ∫ u†² w⁻¹ dν₀ ≤ Z̃ ‖u†‖²‘.
Step 2: ‘S_ν* h = S_ν₀ u† = f‘.
Step 3: apply the general bound and insert ‘Z̃ ≤ e^ζ₀‘.
‘ζ = (κ/2) m*² ≤ (κ/2) Z̃‖u†‖²/(4λ) = Z̃‖u†‖²/(8β) ≤ e^ζ₀‖u†‖²/(8β)‘.
7.2 The rate theorem and its corollaries
Assume (A1), (A3), (A4), (A5) and (SC\(_a\)) (hence (R)). Fix \((\lambda ,\beta )\) and let \(\Xi =\| (w-1)m^*\| _{L^2(\nu _0)}\). Then
(Write \(m^*-u^\dagger =(m^*-v)+(v-u_\lambda )+(u_\lambda -u^\dagger )\): the first term is \((1-w)m^*\) of norm \(\Xi \), the second is at most \(\Xi \) by Lemma 461, the third is Lemma 456.)
With \(\Pi \rho ^*=m^*\nu ^*=v\nu _0\),
(\(\| v-u^\dagger \| _{L^1}\le \| v-m^*\| _{L^1}+\| m^*-u^\dagger \| _{L^1} \le \| v-m^*\| _{L^2}+\| m^*-u^\dagger \| _{L^2}\), \(\nu _0\) being a probability measure.)
‘∫ |m*| |w − 1| dν₀ = ∫ |(w − 1) m*| dν₀ ≤ Ξ‘.
\(\| F_{\rho ^*}-f\| _{L^2(P_X)}\le \frac{\sqrt\lambda }2\, \Xi +\lambda ^{\min (a+1/2,1)}\| g_0\| \). (\(F_{\rho ^*}-f=S_{\nu _0}(v-u^\dagger )=S_{\nu _0}(v-u_\lambda )+S_{\nu _0}(u_\lambda -u^\dagger )\); \(\| S_{\nu _0}(v-u_\lambda )\| ^2=\langle T_0e,e\rangle \) with \((T_0+\lambda )e=\lambda (w-1)m^*\), and \(4\lambda \langle T_0e,e\rangle \le \| (T_0+\lambda )e\| ^2=\lambda ^2\Xi ^2\).)
‘‖S_ν₀ e‖² = ⟪T₀ e, e⟫ ≤ ‖(T₀ + λ) e‖²/(4λ) = λ Ξ²/4‘.
\(\Xi \le \frac\kappa 2\bigl(\| m^{*3}e^{\kappa m^{*2}/2}\| _{L^2(\nu _0)} +\| u^\dagger \| ^2\| m^*\| _{L^2(\nu _0)}\bigr)\) (Lemma 464).
Assume the hypotheses of Theorem 468 and (S\(_\infty \)) with constant \(B_\infty \), and put \(\zeta _\infty :=\frac\kappa 2B_\infty ^2\). Then
and for \(\kappa \le \kappa _0\), \(\Xi \le C_1\kappa \) with \(C_1\) as in Definition 145. (In the bound of Lemma 462, \(|w-1|\le \zeta e^\zeta +\zeta _0 \le \zeta _\infty e^{\zeta _\infty }+\zeta _0\), and \(\| m^*\| _{L^2(\nu _0)}\le e^{\zeta _0/2}\| u^\dagger \| \) by Lemma 465.)
Monotonicity in ‘κ‘ of the exponential factors.
For \(\kappa \le \kappa _0\),
Step 1: the pointwise bound ‘|w − 1| ≤ C₁’ κ‘.
Step 2: the total mass and the ‘χ²‘ distance.
Assume only the hypotheses of Theorem 468 ((S\(_\infty \)) is not used), and let \(\zeta _\beta =e^{\zeta _0}\| u^\dagger \| ^2/(8\beta )\). Then
and for \(\beta \ge \beta _0{\gt}0\), \(\kappa \le \kappa _0\), \(\Xi \le C_2(\beta ^{-1}+\kappa )\) with \(C_2\) as in Definition 147. (Pointwise \(|w-1|\le \zeta e^\zeta +\zeta _0\le \zeta _\beta e^{\zeta _\beta }+\zeta _0\) by Lemma 467, and \(\| m^*\| _{L^2(\nu _0)}\le e^{\zeta _0/2}\| u^\dagger \| \).)
Step 2: insert ‘ζ₀ ≤ ζ₀₀ = κ₀ U²/2‘ and ‘1/β ≤ 1/β₀‘.
‘ζ_β ≤ (1/β) e^ζ₀₀ U²/8‘ and ‘ζ_β ≤ e^ζ₀₀ U²/(8β₀)‘.
The first term of ‘C e^ζ₀/2 U‘ is at most ‘(1/β) U³ e^ζ₀₀/2 M‘, the second at most ‘κ U³ e^ζ₀₀/2 M‘.
For \(\beta \ge \beta _0\) and \(\kappa \le \kappa _0\), \(\| m^*-u^\dagger \| _{L^2(\nu _0)}\le 2C_2(\beta ^{-1}+\kappa )+\lambda ^{\min (a,1)}\| g_0\| \), and likewise for the coefficient measure and the realization as in Corollary 473, with \(C_1\kappa \) replaced by \(C_2(\beta ^{-1}+\kappa )\).
Suppose \(K_{\nu _0}\) is coercive on \(L^2(P_X)\), \(\langle K_{\nu _0}h,h\rangle \ge \sigma _0\| h\| ^2\) with \(\sigma _0{\gt}0\) (in the paper: \(P_X\) has finitely many atoms and \(K_{\nu _0}\) is invertible with smallest eigenvalue \(\sigma _0\)). Then (R) holds for every \(f\): \(f=S_{\nu _0}(S_{\nu _0}^*K_{\nu _0}^{-1}f)\).
Assume (A1), (A3), (A4), (A5), and that \(P_X\) has finitely many atoms \(x_1,\dots ,x_n\), so that \(L^2(P_X)\cong {\mathbb R}^n\). Suppose \(K_{\nu _0}\) is invertible on \(L^2(P_X)\) with smallest eigenvalue \(\sigma _0{\gt}0\). Then (R) holds automatically, and for \(\kappa \le \kappa _0\),
(\(K_{\nu ^*}=\int \Phi _z\Phi _z^\top w\, \mathrm d\nu _0\) and \(w\ge e^{-\zeta _0}\) give \(K_{\nu ^*}\succeq e^{-\zeta _0}K_{\nu _0}\succeq e^{-\zeta _0}\sigma _0I\), hence \(\| (K_{\nu ^*}+\lambda )^{-1}\| \le e^{\zeta _0}/\sigma _0\) and \(|m^*(z)|\le \| (K_{\nu ^*}+\lambda )^{-1}f\| \). In Lean the hypothesis is the coercivity \(\sigma _0\| h\| ^2\le \langle K_{\nu _0}h,h\rangle \) for a general \(P_X\).)
Step 1: ‘⟪K_ν* g, g⟫ = ∫ w (S* g)² dν₀ ≥ e^−ζ₀ ⟪K_ν₀ g, g⟫‘.
Step 2: ‘e^−ζ₀ σ₀ ‖g‖² ≤ ⟪(K_ν* + λ) g, g⟫ = ⟪f, g⟫ ≤ ‖f‖ ‖g‖‘.
Step 3: ‘|m*(z)| = |(S* g)(z)| ≤ ‖g‖‘.
7.3 Spectral powers of \(T_0\) and the source condition of general order
For \(a{\gt}0\), \(\operatorname {ran}T_0^a\subseteq (\ker T_0)^\perp =(\ker S_{\nu _0})^\perp \) (Lemma 792; \(\ker T_0=\ker S_{\nu _0}\) since \(\langle T_0u,u\rangle =\| S_{\nu _0}u\| ^2\)).
If \(f=S_{\nu _0}T_0^ag_0\) with \(a{\gt}0\), then (R) holds and \(u^\dagger =T_0^ag_0\), since \(T_0^ag_0\in (\ker S_{\nu _0})^\perp \) (Lemma 384).
(SC\(_a\)) implies (R): \(f=S_{\nu _0}u^\dagger \).
The source condition of order \(1/2\) of Definition 139, \(f=K_{\nu _0}g\) with \(\| g\| \le G\), implies (SC\(_{1/2}\)) of Definition 165 with the same bound: \(\operatorname {ran}S_{\nu _0}^*\subseteq \operatorname {ran}T_0^{1/2}\) with \(S_{\nu _0}^*g=T_0^{1/2}g_0\), \(\| g_0\| \le \| g\| \) (Lemma 797).
Assume (A1), (A3) and (SC\(_a\)) with \(a{\gt}0\) (hence (R)). Then
(\(u_\lambda -u^\dagger =-\lambda (T_0+\lambda )^{-1}T_0^ag_0\) (Lemma 453), the eigenvalues of \(T_0\) lie in \([0,1]\), and Lemma 796 applies; \(\| S_{\nu _0}h\| ^2=\langle T_0h,h\rangle \).)
‘(T₀ + λ) d = λ T₀^a (−g₀)‘: the resolvent bounds apply with ‘g = −g₀‘.
With \(\Pi \rho ^*=m^*\nu ^*=v\nu _0\), \(\| \Pi \rho ^*-u^\dagger \nu _0\| _{\mathcal M}=\| v-u^\dagger \| _{L^1(\nu _0)} \le 3\, \Xi +\lambda ^{\min (a,1)}\| g_0\| \) (Theorem 469).
‘∫ |m*| |w − 1| dν₀ = ∫ |(w − 1) m*| dν₀ ≤ Ξ‘.
\(\| F_{\rho ^*}-f\| _{L^2(P_X)}\le \frac{\sqrt\lambda }2\, \Xi +\lambda ^{\min (a+1/2,1)}\| g_0\| \) (Theorem 470).
‘‖S_ν₀ e‖² = ⟪T₀ e, e⟫ ≤ ‖(T₀ + λ) e‖²/(4λ) = λ Ξ²/4‘.
7.4 Joint schedule of the sample size and the regularization
With \(D(\delta /2)\le (B+{Y_{\max }})^2[4C_{\mathrm P}\sqrt{m+1} +\sqrt{\tfrac 12\log (4/\delta )}]\) and \(N^{-1/2}\le N^{-1/4}\) for \(N\ge 1\), \(\sqrt{2D(\delta /2)/\sqrt N}+D'(\delta )/\sqrt N \le (B+{Y_{\max }})\, C_{\mathrm{stat}}(m,\delta )\, N^{-1/4}\).
Step 1: ‘D(δ/2) ≤ E² S₁‘.
Step 2: ‘√(2D/√N) ≤ E √(2S₁) N^−1/4‘.
Step 3: ‘D’/√N ≤ E S₂ N^−1/4‘.
Assume the setting of Theorem 328, (SC\(_a\)) with \(0{\lt}a\le 1\), the complexity bound \({\mathfrak R}_N(\Phi )\le C_{\mathrm P}\sqrt{(m+1)/N}\) and (S\(_\infty \)) for the family \(\lambda \le 1\), \(\beta =\beta _0\); let \(C_1\) be Definition 145 with \(\kappa _0=1/\beta _0\). With \(\lambda _N:=N^{-1/(4(a+2))}\), for every \(N\ge 1\), with probability at least \(1-\delta \),
(??: the first term is Corollary 341 and Lemma 488 with \(B+{Y_{\max }}\le (2{Y_{\max }}+\sqrt{2\beta _0/\pi })/\lambda \), the second is Corollary 473; at \(\lambda =\lambda _N\), \(\lambda _N^{-2}N^{-1/4}=\lambda _N^a\) and \(\lambda _N\le \lambda _N^a\).)
The balancing: ‘λ_N^−2 N^−1/4 = λ_N^a = N^−a/(4(a+2))‘ and ‘λ_N ≤ λ_N^a‘.
The bias term: ‘cor:rate-sinf‘ with ‘κ₀ = 1/β₀‘.
The sample term: ‘cor:m3-convergence‘ (b), (e) and ‘eq:C-stat‘.
The triangle inequality ‘eq:joint-triangle‘.
Assume the setting of Theorem 328, (SC\(_a\)) with \(0{\lt}a\le 1\) and the complexity bound. With \(\lambda _N:=N^{-1/(4(a+2))}\) and \(\beta _N:=\lambda _N^{-a}\), for every \(N\ge 1\), with probability at least \(1-\delta \),
where \(C_2\) is Definition 147 with \(\beta _0=\kappa _0=1\). (\(\beta _N\ge 1\) and \(\kappa _N=\lambda _N^{1+a}\le 1\), so Corollary 476 applies with \(\beta _0=\kappa _0=1\) and \(2C_2(\beta _N^{-1}+\kappa _N)\le 4C_2\lambda _N^a\); moreover \(B+{Y_{\max }}\le (2{Y_{\max }}+1)/\lambda \) for \(\lambda \le 1\), \(a\le 1\).)
The balancing: ‘λ_N^−2 N^−1/4 = λ_N^a = N^−a/(4(a+2))‘.
The high-temperature schedule: ‘β_N ≥ 1‘, ‘κ_N ≤ 1‘, ‘1/β_N = λ_N^a‘, ‘κ_N = λ_N^1+a ≤ λ_N^a‘.
The bias term: ‘cor:rate-high-temperature‘ with ‘β₀ = κ₀ = 1‘.
The sample term: ‘cor:m3-convergence‘ (b), (e) and ‘eq:C-stat‘.
The triangle inequality ‘eq:joint-triangle‘.
Assume the setting of Theorem 328, (SC\(_a\)) with \(0{\lt}a\le 1\), the complexity bounds of Lemmas 314 and 315, and let \(D'=D'(\delta )\), \(D''=D''(\delta )\) be as in ??. Put \(\lambda _N:=N^{-1/(2(a+3))}\) and \(\beta _N:=\lambda _N^{-a}\). Then for \(N^{(a+1)/(a+3)}\ge 64\max \{ 1,{Y_{\max }}\} ^4D''{}^2\) (which is \(N\ge N_B(\delta )\) along the schedule, since \(c_1\le \max \{ 1,{Y_{\max }}\} /\lambda _N\)), with probability at least \(1-\delta \),
the rate is \(N^{-1/8}\) for \(a=1\). (?? with ??: along the schedule \(c_1\le \max \{ 1,{Y_{\max }}\} /\lambda \) and \(E_*\le 2{Y_{\max }}/\lambda \), so \(\sup |m_N-m^*|\le 24\max \{ 1,{Y_{\max }}\} {Y_{\max }}D'\lambda ^{-3}N^{-1/2}\); balance \(\lambda ^{-3}N^{-1/2}=\lambda ^a\).)
The balancing: ‘λ_N^−3 N^−1/2 = λ_N^a = N^−a/(2(a+3))‘.
The high-temperature schedule: ‘β_N ≥ 1‘, ‘κ_N ≤ 1‘, ‘1/β_N = λ_N^a‘, ‘κ_N ≤ λ_N^a‘.
The bias term: ‘cor:rate-high-temperature‘ with ‘β₀ = κ₀ = 1‘.
The constants of ‘thm:m3-double-prime‘ along the schedule: ‘c₁ ≤ max1, Ymax/λ‘ and ‘E_* ≤ 2Ymax/λ‘, and the threshold ‘N ≥ N_B(δ)‘.
The sample term: ‘cor:m3-convergence-prime‘ (b) in the ‘√N‘ form.
The triangle inequality ‘eq:joint-triangle‘.
Assume the setting of Theorem 328, (SC\(_a\)) with \(0{\lt}a\le 1\), the complexity bounds of Lemmas 314 and 315, and (S\(_\infty \)) for the family \(\lambda \le 1\), \(\beta =\beta _0\); let \(C_1\) be Definition 145 with \(\kappa _0=1/\beta _0\) and \(M:=\max \{ 1,{Y_{\max }}/\sqrt{\beta _0}\} \). With \(\lambda _N:=N^{-1/(2(a+3))}\), for \(N^{(a+1)/(a+3)}\ge 64M^4D''{}^2\), with probability at least \(1-\delta \),
(As Corollary 491, with \(c_1\le M/\lambda \) and the bias of Corollary 473. The formalized constants \(c_1\), \(E_*\) of Theorem 381 use \(M_*={Y_{\max }}/\lambda \), whence the rate \(N^{-a/(2(a+3))}\) in place of the paper’s \(N^{-a/(2(a+2))}\).)
The balancing: ‘λ_N^−3 N^−1/2 = λ_N^a = N^−a/(2(a+3))‘ and ‘λ_N ≤ λ_N^a‘.
The bias term: ‘cor:rate-sinf‘ with ‘κ₀ = 1/β₀‘.
The constants of ‘thm:m3-double-prime‘ along the schedule: ‘c₁ ≤ max1, Ymax/√β₀/λ‘, ‘E_* ≤ 2Ymax/λ‘, and the threshold ‘N ≥ N_B(δ)‘.
The sample term: ‘cor:m3-convergence-prime‘ (b) in the ‘√N‘ form.
The triangle inequality ‘eq:joint-triangle‘.
Assume the setting of Theorem 328, (SC\(_a\)) with \(0{\lt}a\le 1\), the complexity bounds of Lemmas 314 and 315, and (S\(_\infty \)) for the family \(\lambda \le 1\), \(\beta =\beta _0\) with constant \(B_\infty \); let \(C_1\) be Definition 145 with \(\kappa _0=1/\beta _0\), and \(D'=D'(\delta )\), \(D''=D''(\delta )\) as in ??. Put \(\lambda _1:=\min \{ 1,\sqrt{\beta _0}/B_\infty \} \) and \(\lambda _N:=N^{-1/(2(a+2))}\). Then for \(N\ge N_{(i)}:=\max \{ \lambda _1^{-2(a+2)},(64D''{}^2)^{(a+2)/a}\} \), with probability at least \(1-\delta \),
the rate is \(N^{-1/6}\) for \(a=1\) and \(N^{-1/10}\) for \(a=1/2\). (For \(\lambda \le \lambda _1\), \(c_1(B_\infty )=1/\lambda \) and \(E_*=B_\infty +{Y_{\max }}\) in Theorem 380, so \(\sup |m_N-m^*|\le 12(B_\infty +{Y_{\max }})D'\lambda ^{-2}N^{-1/2}\); the bias is Corollary 473; balance \(\lambda ^{-2}N^{-1/2}=\lambda ^a\).)
The balancing: ‘λ_N^−2 N^−1/2 = λ_N^a = N^−a/(2(a+2))‘ and ‘λ_N ≤ λ_N^a‘.
The bias term: ‘cor:rate-sinf‘ with ‘κ₀ = 1/β₀‘, and the sup norm ‘|m*| ≤ B∞‘.
The threshold ‘N ≥ N_(i)‘: ‘λ_N ≤ λ₁‘ gives ‘c₁(B∞) = 1/λ_N‘, and ‘N^a/(a+2) ≥ 64 D”²‘ gives ‘N ≥ N_B(δ) = 64 c₁⁴ D”²‘.
The sample term: ‘thm:m3-double-prime-rate-general‘ with ‘M₁ = B∞‘.
The triangle inequality ‘eq:joint-triangle‘.
Assume the setting of Theorem 328, (SC\(_a\)) with \(0{\lt}a\le 1\) and the complexity bounds of Lemmas 314 and 315; let \(D'=D'(\delta )\), \(D''=D''(\delta )\) be as in ??. Put \(c_u:=\frac12e^{\| u^\dagger \| ^2/4}\| u^\dagger \| \), \(\lambda _2:=\min \{ 1,c_u^{-2}\} \), \(\lambda _N:=N^{-1/(2a+5)}\) and \(\beta _N:=\lambda _N^{-a}\). Then for \(N\ge N_{(ii)}:=\max \{ \lambda _2^{-(2a+5)},(64D''{}^2)^{(2a+5)/(2a+1)}\} \), with probability at least \(1-\delta \),
where \(C_2\) is Definition 147 with \(\beta _0=\kappa _0=1\); the rate is \(N^{-1/7}\) for \(a=1\). ((SC\(_a\)) contains (R), so Lemma 467 with \(\zeta _0=\frac\kappa 2\| u^\dagger \| ^2\le \frac12\| u^\dagger \| ^2\) gives \(M_*\le c_u\lambda ^{-1/2}\); for \(\lambda \le \lambda _2\) and \(\beta _N\ge 1\), \(c_1(M_*)=1/\lambda \) and \(E_*\le (c_u+{Y_{\max }})\lambda ^{-1/2}\), so \(\sup |m_N-m^*|\le 12(c_u+{Y_{\max }})D'\lambda ^{-5/2}N^{-1/2}\); the bias is Corollary 476 with \(\beta _0=\kappa _0=1\); balance \(\lambda ^{-5/2}N^{-1/2}=\lambda ^a\).)
The balancing: ‘λ_N^−5/2 N^−1/2 = λ_N^a = N^−a/(2a+5)‘.
The high-temperature schedule: ‘β_N ≥ 1‘, ‘κ_N ≤ 1‘, ‘1/β_N = λ_N^a‘, ‘κ_N ≤ λ_N^a‘.
The bias term: ‘cor:rate-high-temperature‘ with ‘β₀ = κ₀ = 1‘.
‘lem:sqrt-lambda-sup‘: ‘|m*| ≤ e^ζ₀/2‖u†‖/(2√λ) ≤ c_u/√λ‘ since ‘ζ₀ ≤ ‖u†‖²/2‘.
The threshold ‘N ≥ N_(ii)‘: ‘λ_N ≤ λ₂‘ gives ‘c₁(c_u/√λ_N) = 1/λ_N‘, and ‘N^(2a+1)/(2a+5) ≥ 64 D”²‘ gives ‘N ≥ N_B(δ) = 64 c₁⁴ D”²‘.
The sample term: ‘thm:m3-double-prime-rate-general‘ with ‘M₁ = c_u/√λ_N‘.
The triangle inequality ‘eq:joint-triangle‘.