Shallow Learning Tends to Ridgelet Transform

13 Material for Mathlib

General lemmas used by the formalization that belong upstream (ToMathlib/): the Donsker–Varadhan inequality for the Kullback–Leibler divergence, moments of the real Gaussian, the completion of the square behind the Gibbs decomposition, the total variation distance with Pinsker’s inequality, and the standard facts about log-Sobolev inequalities (ToMathlib/LogSobolev/): the entropy functional and its variational formulas, the Holley–Stroock perturbation, the transport by proper \(C^1\) maps, the tensorization on a Euclidean product, and the KL form (L1) by cut-off approximation.

Theorem 826
✓

Let \(\rho \ll \mu \) be probability measures with \(\log \frac{d\rho }{d\mu } \in L^1(\rho )\), and let \(g \in L^1(\rho )\) with \(e^g \in L^1(\mu )\). Then \(\int g \, d\rho \le \int \log \frac{d\rho }{d\mu } \, d\rho + \log \int e^g \, d\mu \).

Proof ▶

Step 1: the tilted measure ‘μ.tilted g‘ is a probability measure and ‘ρ ≪ μ.tilted g‘.

Step 2: Gibbs’ inequality ‘0 ≤ ∫ llr ρ (μ.tilted g) dρ‘ for the two probability measures.

Step 3: ‘∫ llr ρ (μ.tilted g) dρ = ∫ llr ρ μ dρ - ∫ g dρ + log ∫ e^g dμ‘.

Theorem 827
✓

Let \(\rho , \mu \) be probability measures with \(\mathrm{KL}(\rho \| \mu ) {\lt} \infty \), and let \(g \in L^1(\rho )\) with \(\int e^{g} \, d\mu {\lt} \infty \). Then \(\int g \, d\rho \le \mathrm{KL}(\rho \| \mu ) + \log \int e^{g} \, d\mu \).

Proof ▶

Step 1: finiteness of the divergence gives ‘ρ ≪ μ‘ and integrability of ‘llr ρ μ‘.

Step 2: ‘KL(ρ ‖ μ) = ∫ llr ρ μ dρ‘ for probability measures, and apply the llr version.

Let \(\rho , \mu \) be probability measures with \(\mathrm{KL}(\rho \| \mu ) {\lt} \infty \), and let \(g \ge 0\) be measurable with \(\int e^{g} \, d\mu {\lt} \infty \) (Lebesgue integral). Then \(\int g \, d\rho \le \mathrm{KL}(\rho \| \mu ) + \log \int e^{g} \, d\mu \), where the left-hand side is the Bochner integral (equal to \(0\) if \(g \notin L^1(\rho )\)).

Proof ▶

Step 1: ‘e^g ∈ L¹(μ)‘, since ‘e^g‘ is measurable with finite Lebesgue integral.

Step 2: the Bochner integral of ‘e^g‘ is the Lebesgue integral.

Step 3: if ‘g ∈ L¹(ρ)‘, apply the Donsker–Varadhan inequality.

Step 4: otherwise the left-hand side is ‘0‘, and ‘∫ e^g dμ ≥ ∫ 1 dμ = 1‘ gives ‘log ∫ e^g dμ ≥ 0‘.

Let \(\rho , \mu \) be probability measures with \(\mathrm{KL}(\rho \| \mu ) {\lt} \infty \), and let \(g\) be measurable with \(|g| \le C\). Then \(\int g \, d\rho \le \mathrm{KL}(\rho \| \mu ) + \log \int e^{g} \, d\mu \).

Proof ▶

Step 1: ‘g ∈ L¹(ρ)‘ and ‘e^g ∈ L¹(μ)‘ (bounded by ‘C‘ and ‘e^C‘).

Theorem 830
✓

Let \(\rho ,\mu \) be probability measures with \(\operatorname {KL}(\rho \| \mu ){\lt}\infty \) and let \(g\ge 0\) be measurable with \(\int e^g\, \mathrm d\mu {\lt}\infty \). Then \(\int g\, \mathrm d\rho \le \operatorname {KL}(\rho \| \mu )+\log \int e^g\, \mathrm d\mu \) in \([0,\infty ]\), where the left-hand side is the Lebesgue integral of \(g\); in particular \(g\in L^1(\rho )\). The proof applies the bounded Donsker–Varadhan inequality to \(g\wedge n\) and lets \(n\to \infty \) by monotone convergence.

Proof ▶

Step 1: the truncations ‘gₙ = min g n‘ are bounded by ‘n‘, so the bounded Donsker–Varadhan inequality gives ‘∫ gₙ dρ ≤ KL + log ∫ e^gₙ dμ ≤ KL + log ∫ e^g dμ‘.

Step 2: ‘⨆ n, min (g x) n = g x‘, so by monotone convergence ‘∫⁻ g dρ = ⨆ n, ∫⁻ gₙ dρ = ⨆ n, ofReal (∫ gₙ dρ)‘.

Theorem 831
✓

Let \(\rho ,\mu \) be probability measures with \(\operatorname {KL}(\rho \| \mu ){\lt}\infty \) and let \(g\ge 0\) be measurable with \(\int e^g\, \mathrm d\mu {\lt}\infty \). Then \(g\in L^1(\rho )\).

Proof ▶
Theorem 832
✓

For \(v \ge 0\) and \(2 s v {\lt} 1\), \(x \mapsto e^{s x^2}\) is integrable with respect to \(\mathcal{N}(0, v)\).

Proof ▶

Case ‘v = 0‘: ‘𝒩(0, 0)‘ is the Dirac mass at ‘0‘.

Case ‘v ≠ 0‘: unfold the density and complete the square in the exponent.

Theorem 833
✓

For \(v \ge 0\) and \(2 s v {\lt} 1\), \(\int e^{s x^2} \, \mathcal{N}(0, v)(dx) = (1 - 2 s v)^{-1/2}\).

Proof ▶

Case ‘v = 0‘: both sides are ‘1‘.

Case ‘v ≠ 0‘: unfold the density, complete the square, and use ‘integral_gaussian‘.

Theorem 834
✓

For \(v \ge 0\), \(\int |x| \, \mathcal{N}(0, v)(dx) = \sqrt{2 v / \pi }\).

Proof ▶

Step 1: ‘|x| = max x 0 + max (-x) 0‘, both parts integrable.

Step 2: by symmetry ‘x ↦ -x‘ of ‘𝒩(0, v)‘, both parts have the same integral ‘√(v / (2π))‘ (‘integral_max_zero_gaussianReal‘).

Step 3: ‘2 √(v / (2π)) = √(2 v / π)‘.

Theorem 835
✓

For \(\mu \in \mathbb {R}\) and \(v \ge 0\), \(\int |x| \, \mathcal{N}(\mu , v)(dx) \le |\mu | + \sqrt{2 v / \pi }\).

Proof ▶

Step 1: the centered absolute moment ‘∫ |x - μ| 𝒩(μ, v)(dx) = ∫ |y| 𝒩(0, v)(dy)‘.

Step 2: ‘|x| ≤ |μ| + |x - μ|‘ and integrate.

Theorem 836
✓

For \(\mu \in \mathbb {R}\) and \(v \ge 0\), \(\int x^2 \, \mathcal{N}(\mu , v)(dx) = \mu ^2 + v\).

Proof ▶

‘Var = ∫ x² - (∫ x)²‘ with ‘Var = v‘ and ‘∫ x = μ‘.

Theorem 837
✓

For \(\lambda {\gt} 0\), \(\beta {\gt} 0\) and \(s \in \mathbb {R}\), \(\int _{\mathbb {R}} e^{-(a s + \lambda a^2/2)/\beta } \, da = \sqrt{2\pi \beta /\lambda } \, e^{s^2/(2\lambda \beta )}\).

Proof ▶

Step 1: complete the square and pull out the constant.

Step 2: translation invariance of Lebesgue measure and ‘integral_gaussian‘.

Theorem 838
✓

For \(\lambda {\gt} 0\), \(\beta {\gt} 0\) and \(s \in \mathbb {R}\), as measures on \(\mathbb {R}\), \(e^{-(a s + \lambda a^2/2)/\beta } \, da = \sqrt{2\pi \beta /\lambda } \, e^{s^2/(2\lambda \beta )} \, \mathcal{N}(-s/\lambda , \beta /\lambda )(da)\).

Proof ▶

Step 1: the variance ‘β/λ‘ is positive, so ‘𝒩(-s/λ, β/λ)‘ has the Gaussian density.

Step 2: compare the densities pointwise.

Theorem 839
✓

For \(\lambda ,\beta {\gt}0\) and \(c\in {\mathbb R}\), as measures on \({\mathbb R}\), \(e^{-ac/\beta }\, {\mathcal N}(0,\beta /\lambda )(\, \mathrm da) =e^{c^2/(2\lambda \beta )}\, {\mathcal N}(-c/\lambda ,\beta /\lambda )(\, \mathrm da)\).

Proof ▶

Step 1: the product of the densities is ‘(√(2πβ/λ))⁻¹ e^−(a c + λ a²/2)/β‘.

Step 2: ‘lem:gaussian-complete-square‘ and the cancellation of the constants.

Theorem 840
✓

For \(v{\gt}0\) and \(m,m'\in {\mathbb R}\), as measures on \({\mathbb R}\), \({\mathcal N}(m,v)(\, \mathrm da)=\exp \Bigl(\frac{a(m-m')-(m^2-m'{}^2)/2}{v}\Bigr){\mathcal N}(m',v)(\, \mathrm da)\).

Proof ▶
Definition 841
✓
#

The total variation distance of two (finite) measures \(\mu , \nu \) is \(\| \mu - \nu \| _{\mathrm{TV}} := \sup _{A} |\mu (A) - \nu (A)|\), the supremum over measurable sets \(A\).

Theorem 842
✓

Let \(\mu , \nu \) be finite measures and \(g\) a measurable function with \(|g| \le C\). Then \(\left|\int g \, d\mu - \int g \, d\nu \right| \le 2 C \, \| \mu - \nu \| _{\mathrm{TV}}\).

Proof ▶

Step 1: the Hahn decomposition ‘P‘ of ‘μ - ν‘.

Step 2: ‘∫ g dμ - ∫ g dν = ∫ g d(μ - ν)|_P - ∫ g d(ν - μ)|_Pᶜ‘.

Step 3: each of the two integrals is bounded by ‘C ‖μ - ν‖_TV‘.

Theorem 843
✓
#

For finite measures \(\mu ,\nu \) and a measurable map \(g\), \(\| g_*\mu -g_*\nu \| _{\mathrm{TV}}\le \| \mu -\nu \| _{\mathrm{TV}}\), since \(g_*\mu (A)-g_*\nu (A) =\mu (g^{-1}A)-\nu (g^{-1}A)\).

Proof ▶
Definition 844
✓
#

For a finite signed measure \(\mu \) on \(Z\), \(\| \mu \| _{\mathrm{TV}}:=\tfrac 12|\mu |(Z)\), where \(|\mu |=\mu ^++\mu ^-\) is the total variation measure. For \(\mu =\rho -\rho '\) with \(\rho ,\rho '\) probability measures this is \(\sup _A|\rho (A)-\rho '(A)|\).

Theorem 845
✓
#

If \(\mu (A)-\mu (A^c)\le 2C\) for every measurable \(A\) then \(\| \mu \| _{\mathrm{TV}}\le C\): on a Hahn set \(S\) of \(\mu \), \(|\mu |(Z)=\mu (S^c)-\mu (S)\).

Proof ▶
Theorem 846
✓

For a finite measure \(\mu \) and \(\varphi \in L^1(\mu )\) measurable, the total variation of the signed measure \(\varphi \, \mu \) is \(\int |\varphi |\, \mathrm d\mu \) (its Jordan decomposition is \((\varphi ^+\mu ,\varphi ^-\mu )\)).

Proof ▶
Theorem 847
✓

With \(w_i=\, \mathrm d\nu _i/\, \mathrm d(\nu _1+\nu _2)\), \(\int |w_1-w_2|\, \mathrm d(\nu _1+\nu _2)\le 2\| \nu _1-\nu _2\| _{\mathrm{TV}}\) (split on \(\{ w_1\ge w_2\} \)).

Proof ▶

Step 1: ‘∫ |w₁ − w₂| = ∫_A (w₁ − w₂) + ∫_Aᶜ (w₂ − w₁)‘.

Step 2: each piece is a difference of masses, bounded by the total variation distance.

Theorem 848
✓

For \(p \in [0,1]\) and \(q \in (0,1)\), \(p \log \frac{p}{q} + (1-p) \log \frac{1-p}{1-q} \ge 2 (p - q)^2\).

Proof ▶

Step 1: the difference ‘h p‘ of the two sides, in a form continuous on ‘[0, 1]‘.

Step 2: ‘h‘ is continuous on ‘[0, 1]‘, differentiable on ‘(0, 1)‘ with derivative ‘h’(p) = log p - log q - log (1 - p) + log (1 - q) - 4 (p - q)‘, which is monotone on ‘(0, 1)‘ and vanishes at ‘q‘.

Step 3: ‘h‘ is antitone on ‘[0, q]‘ and monotone on ‘[q, 1]‘, hence ‘h p ≥ h q = 0‘.

Theorem 849
✓

Let \(\rho , \mu \) be probability measures with \(\mathrm{KL}(\rho \| \mu ) {\lt} \infty \). Then for every measurable set \(A\), \(|\rho (A) - \mu (A)| \le \sqrt{\mathrm{KL}(\rho \| \mu ) / 2}\).

Proof ▶

Step 1: ‘ρ ≪ μ‘, ‘llr ρ μ‘ is ‘ρ‘-integrable and ‘KL(ρ ‖ μ) = ∫ klFun (dρ/dμ) dμ‘.

Step 2: Jensen’s inequality on ‘A‘ and on ‘Aᶜ‘ (data processing along the indicator of ‘A‘): ‘∫ klFun (dρ/dμ) dμ ≥ q klFun (p/q) + (1 - q) klFun ((1 - p)/(1 - q))‘.

Step 3: the edge cases ‘q = 0‘ (then ‘p = 0‘) and ‘q = 1‘ (then ‘p = 1‘) are trivial; otherwise the binary Pinsker inequality applies.

Theorem 850
✓

Let \(\rho , \mu \) be probability measures with \(\mathrm{KL}(\rho \| \mu ) {\lt} \infty \). Then \(\| \rho - \mu \| _{\mathrm{TV}} \le \sqrt{\mathrm{KL}(\rho \| \mu ) / 2}\).

Proof ▶
Theorem 851
✓

Let \(\rho , \mu \) be probability measures with \(\mathrm{KL}(\rho \| \mu ) {\lt} \infty \) and \(g\) a measurable function with \(|g| \le C\). Then \(\left|\int g \, d\rho - \int g \, d\mu \right| \le 2 C \sqrt{\mathrm{KL}(\rho \| \mu ) / 2}\).

Proof ▶
Theorem 852
✓

For probability measures \(\rho ,\rho '\) with \(\operatorname {KL}(\rho \| \rho ')\) and \(\operatorname {KL}(\rho '\| \rho )\) finite, \(\| \rho -\rho '\| _{\mathrm{TV}}^2 \le \frac12\min \{ \operatorname {KL}(\rho \| \rho '),\operatorname {KL}(\rho '\| \rho )\} \), hence \(2\| \rho -\rho '\| _{\mathrm{TV}}\le \sqrt{\operatorname {KL}(\rho \| \rho ')+\operatorname {KL}(\rho '\| \rho )}\).

Proof ▶
Definition 853
✓
#

For a finite signed measure \(\gamma =\gamma ^+-\gamma ^-\) (Jordan decomposition) and \(f\) integrable for \(|\gamma |=\gamma ^++\gamma ^-\), \(\int f\, \, \mathrm d\gamma :=\int f\, \, \mathrm d\gamma ^+-\int f\, \, \mathrm d\gamma ^-\).

Definition 854
✓
#

The total mass of a finite signed measure is \(\| \gamma \| _{{\mathcal M}}:=|\gamma |(Z)=\gamma ^+(Z)+\gamma ^-(Z)\); it equals \(2\| \gamma \| _{TV}\) for the convention \(\| \gamma \| _{TV}=\sup _A|\gamma (A)|\) when \(\gamma (Z)=0\).

Theorem 855
✓

If \(\gamma =\mu _1-\mu _2\) with \(\mu _1,\mu _2\) finite measures and \(f\) is integrable for \(\mu _1\) and \(\mu _2\), then \(\int f\, \, \mathrm d\gamma =\int f\, \, \mathrm d\mu _1-\int f\, \, \mathrm d\mu _2\).

Proof ▶

Step 1: ‘γ⁺ + μ₂ = γ⁻ + μ₁‘ as (finite) measures.

Step 2: integrate ‘f‘ against both sides and rearrange.

Theorem 856
✓

For \(h\in L^1(\nu )\) and \(g\) measurable with \(gh\in L^1(\nu )\), \(\int g\, \, \mathrm d(h\nu )=\int gh\, \, \mathrm d\nu \).

Proof ▶
Theorem 857
✓

On a metrizable space, two finite signed measures with the same integrals \(\int g\, \, \mathrm d\gamma =\int g\, \, \mathrm d\gamma '\) for every \(g\in C_b\) are equal.

Proof ▶

Step 1: the finite measures ‘γ⁺ + γ’⁻‘ and ‘γ’⁺ + γ⁻‘ have the same integrals of bounded continuous functions, hence are equal.

Step 2: evaluate on a measurable set.

Theorem 858
✓

Let \(\rho ,\mu \) be probability measures with \(\operatorname {KL}(\rho \| \mu ){\lt}\infty \), let \(A\) be measurable and \(t\ge 0\). Then \(t\, \rho (A)\le \operatorname {KL}(\rho \| \mu )+\log \bigl(1+(e^t-1)\mu (A)\bigr)\). (Donsker–Varadhan with \(g=t\, \mathbf1_A\), for which \(\int e^g\, \mathrm d\mu =1+(e^t-1)\mu (A)\).) With \(t=\log (1+1/\mu (A))\) this is \(\rho (A)\le [\operatorname {KL}(\rho \| \mu )+\log 2]/\log (1+1/\mu (A))\).

Proof ▶

Step 1: the bounded Donsker–Varadhan inequality for ‘g = t 1_A‘, ‘|g| ≤ t‘.

Step 2: ‘∫ g dρ = t ρ(A)‘ and ‘e^g = 1 + (e^t − 1) 1_A‘, so ‘∫ e^g dμ = 1 + (e^t − 1) μ(A)‘.

Theorem 859
✓

Let \(\mu \) be a probability measure on a Hausdorff space such that \(\{ \mu \} \) is tight (e.g. any probability measure on a Polish space), and let \(C\in {\mathbb R}\). Then the sublevel set \(\{ \rho \in {\mathcal P}:\operatorname {KL}(\rho \| \mu )\le C\} \) is tight. Indeed, given \({\varepsilon }{\gt}0\) let \(t:=(\max (C,0)+\log 2)/{\varepsilon }\) and let \(K\) be compact with \(\mu (K^c)\le e^{-t}\); by Lemma 858, \(t\rho (K^c)\le C+\log (1+(e^t-1)e^{-t})\le C+\log 2\), so \(\rho (K^c)\le {\varepsilon }\).

Proof ▶

Step 1: the case ‘ε = ∞‘ is trivial; otherwise ‘ε’ = ε.toReal > 0‘.

Step 2: ‘t = (max C 0 + log 2)/ε’‘ and a compact ‘K‘ with ‘μ(Kᶜ) ≤ e^−t‘.

Step 3: for ‘ρ‘ in the sublevel set, ‘t ρ(Kᶜ) ≤ C + log 2‘, hence ‘ρ(Kᶜ) ≤ ε’‘.

Theorem 860
✓

Let \(\mu \) be a probability measure on a Polish space and \(C\in {\mathbb R}\). Then \(\{ \rho \in {\mathcal P}:\operatorname {KL}(\rho \| \mu )\le C\} \) is tight (Lemma 859, since \(\{ \mu \} \) is tight on a Polish space).

Proof ▶
Definition 861
✓
#

For measures \(\rho ,\mu \) and \(g\colon \alpha \to {\mathbb R}\) the Donsker–Varadhan functional is \(\mathrm{DV}_{\rho ,\mu }(g):=\int g\, \mathrm d\rho -\log \int e^g\, \mathrm d\mu \).

Theorem 862
✓

For probability measures \(\rho ,\mu \) and bounded measurable \(g\), \(\mathrm{DV}_{\rho ,\mu }(g)\le \operatorname {KL}(\rho \| \mu )\) in \([0,\infty ]\) (Lemma 829).

Proof ▶

For probability measures \(\rho ,\mu \) on a topological space with its Borel \(\sigma \)-algebra and \(g\in C_b\), \(\mathrm{DV}_{\rho ,\mu }(g)\le \operatorname {KL}(\rho \| \mu )\).

Proof ▶
Theorem 864
✓
#

Let \(\rho \ll \mu \) be finite measures. Then \((\log \frac{d\rho }{d\mu })^-\) is \(\rho \)-integrable: \(\int (\log \frac{d\rho }{d\mu })^-\, \mathrm d\rho =\int r\, (\log r)^-\, \mathrm d\mu \le \mu (\alpha )\) with \(r=\frac{d\rho }{d\mu }\), since \(r(\log r)^-\le 1-r\le 1\) on \((0,1)\).

Proof ▶

Step 1: ‘r (−log r)⁺ ≤ 1‘ for ‘r ≥ 0‘, from ‘log (1/r) ≤ 1/r − 1‘.

Step 2: transfer to ‘μ‘ through the density and bound the integrand by ‘1‘.

Let \(\rho ,\mu \) be probability measures and \(K\in [0,\infty ]\) with \(\mathrm{DV}_{\rho ,\mu }(g)\le K\) for every bounded measurable \(g\). Then \(\operatorname {KL}(\rho \| \mu )\le K\). Proof: if \(\rho \not\ll \mu \), the functions \(t\mathbf1_A\) with \(\mu (A)=0{\lt}\rho (A)\) give \(\mathrm{DV}=t\rho (A)\to \infty \). Otherwise let \(r=d\rho /d\mu \) and \(g_n:=\min (n,\log \max (r,e^{-n}))\), so that \(|g_n|\le n\) and \(\int e^{g_n}\, \mathrm d\mu \le \int (r+e^{-n})\, \mathrm d\mu =1+e^{-n}\); hence \(\int g_n\, \mathrm d\rho \le K+\log (1+e^{-n})\). Since \(\min (n,(\log r)^+)\le g_n+(\log r)^-\) and \((\log r)^-\in L^1(\rho )\) (Lemma 864), monotone convergence gives \((\log r)^+\in L^1(\rho )\), so \(\operatorname {KL}(\rho \| \mu )=\int \log r\, \mathrm d\rho =\lim \int g_n\, \mathrm d\rho \le K\) by dominated convergence.

Proof ▶

Step 0: the case ‘K = ∞‘ is trivial.

Step 1: ‘ρ ≪ μ‘; otherwise a ‘μ‘-null set ‘A‘ with ‘ρ(A) > 0‘ and the test functions ‘t 1_A‘, for which ‘DV = t ρ(A)‘, show that ‘K = ∞‘.

Step 2: the truncations ‘gₙ = min n (log (max r e^−n))‘ of the log-likelihood ratio, with ‘|gₙ| ≤ n‘ and ‘e^gₙ ≤ r + e^−n‘, so that ‘∫ e^gₙ dμ ≤ 1 + e^−n‘.

Step 3: the bound ‘∫ gₙ dρ ≤ K + log (1 + e^−n)‘.

Step 4: ‘ρ‘-a.e. ‘r > 0‘, and there ‘gₙ = min n (max (log r) (−n))‘ is the clamped log-likelihood ratio: ‘|gₙ| ≤ |log r|‘, ‘gₙ → log r‘, ‘min n (log r)⁺ ≤ gₙ + (log r)⁻‘.

Step 5: the negative part of ‘log r‘ is ‘ρ‘-integrable, and the uniform bound of Step 3 gives, by monotone convergence, that the positive part is ‘ρ‘-integrable too.

Step 6: ‘KL(ρ‖μ) = ∫ log r dρ = lim ∫ gₙ dρ ≤ K‘ by dominated convergence (with the bound ‘|log r|‘), since ‘∫ gₙ dρ ≤ K + log (1 + e^−n) → K‘.

Let \(\rho ,\mu \) be probability measures on a metrizable space with its Borel \(\sigma \)-algebra, and \(K\in [0,\infty ]\) with \(\mathrm{DV}_{\rho ,\mu }(g)\le K\) for every \(g\in C_b\). Then \(\mathrm{DV}_{\rho ,\mu }(g)\le K\) for every bounded measurable \(g\). Proof: given \(|g|\le C\) and \({\varepsilon }{\gt}0\), choose \(g_1\in C_b\) with \(\int |g-g_1|\, \mathrm d(\rho +\mu )\le {\varepsilon }\) (density of \(C_b\) in \(L^1\) of a finite measure on a metrizable space) and clamp it to \(g_2:=\max (-C,\min (C,g_1))\in C_b\), which still satisfies \(\int |g-g_2|\, \mathrm d(\rho +\mu )\le {\varepsilon }\). Then \(|\int g\, \mathrm d\rho -\int g_2\, \mathrm d\rho |\le {\varepsilon }\) and \(|\int e^g\, \mathrm d\mu -\int e^{g_2}\, \mathrm d\mu |\le e^C{\varepsilon }\), so \(\mathrm{DV}(g)\le \mathrm{DV}(g_2)+{\varepsilon }+\log (\int e^g\, \mathrm d\mu +e^C{\varepsilon })-\log \int e^g\, \mathrm d\mu \), and the error tends to \(0\) as \({\varepsilon }\to 0\).

Proof ▶

Step 0: the case ‘K = ∞‘ is trivial; otherwise it suffices to show ‘DV(g) ≤ K + δ‘ for every ‘δ > 0‘.

Step 1: choose ‘ε > 0‘ with ‘ε + log (I + e^C ε) < log I + δ‘, ‘I = ∫ e^g dμ‘.

Step 2: a bounded continuous ‘g₁‘ with ‘∫ |g − g₁| d(ρ + μ) ≤ ε‘, clamped to ‘[−C, C]‘ as ‘g₂‘; then ‘∫ |g − g₂| dρ ≤ ε‘ and ‘∫ |g − g₂| dμ ≤ ε‘.

Step 3: ‘∫ g₂ dρ ≥ ∫ g dρ − ε‘ and ‘∫ e^g₂ dμ ≤ I + e^C ε‘.

Step 4: combine with ‘DV(g₂) ≤ K‘.

Let \(\rho ,\mu \) be probability measures on a metrizable space with its Borel \(\sigma \)-algebra and \(K\in [0,\infty ]\) with \(\mathrm{DV}_{\rho ,\mu }(g)\le K\) for every \(g\in C_b\). Then \(\operatorname {KL}(\rho \| \mu )\le K\) (Lemmas 866 and 865).

Proof ▶

Let \(\rho ,\mu \) be probability measures on a metrizable space with its Borel \(\sigma \)-algebra. Then \(\operatorname {KL}(\rho \| \mu )=\sup _{g\in C_b}\Bigl[\int g\, \mathrm d\rho -\log \int e^g\, \mathrm d\mu \Bigr]\) in \([0,\infty ]\) (the supremum of the nonnegative parts; the inequality \(\ge \) is the Donsker–Varadhan inequality, and \(\le \) is Lemma 867).

Proof ▶
Theorem 869
✓

Let \(\mu \) be a probability measure on a metrizable space with its Borel \(\sigma \)-algebra, and let \(\rho _i\to \rho \) weakly (along a filter) in \({\mathcal P}\). Then \(\operatorname {KL}(\rho \| \mu )\le \liminf _i\operatorname {KL}(\rho _i\| \mu )\). Proof: for \(g\in C_b\), \(\mathrm{DV}_{\rho ,\mu }(g)=\lim _i\mathrm{DV}_{\rho _i,\mu }(g)\le \liminf _i\operatorname {KL}(\rho _i\| \mu )\) by the Donsker–Varadhan inequality and the weak convergence, and the supremum over \(g\) is \(\operatorname {KL}(\rho \| \mu )\) by Lemma 868.

Proof ▶

‘DV_ρᵢ(g) → DV_ρ(g)‘ by weak convergence, and ‘DV_ρᵢ(g) ≤ KL(ρᵢ‖μ)‘.

Theorem 870
✓

Let \(\mu _i\) be \(\sigma \)-finite measures and \(f_i\ge 0\) be \(\mu _i\)-integrable. Then \(\bigotimes _i(f_i\mu _i)=\bigl(\prod _if_i(x_i)\bigr)\bigotimes _i\mu _i\). (Both sides agree on measurable boxes by Fubini.)

Proof ▶
Theorem 871
✓
#

For a \(\sigma \)-finite \(\mu \) and a permutation \(\sigma \) of the finite index set, the image of \(\mu ^{\otimes \iota }\) under \(\theta \mapsto \theta \circ \sigma \) is \(\mu ^{\otimes \iota }\).

Proof ▶
Theorem 872
✓

If \(T\) is measurable with \(T_*\nu =\nu \) and \(D\circ T=D\) for a measurable \(D\ge 0\), then \(T_*(D\nu )=D\nu \).

Proof ▶
Theorem 873
✓
#

For a \(\sigma \)-finite \(\rho \), measurable \(Y\) and \(t\in {\mathbb R}\), \(\int e^{t\sum _iY(x_i)}\, \rho ^{\otimes \iota }(\, \mathrm dx)=\bigl(\int e^{tY}\, \mathrm d\rho \bigr)^{|\iota |}\) (Fubini; no integrability is needed, both sides being \(0\) when the right-hand factor is not integrable).

Proof ▶
Theorem 874
✓
#

For all \(x\in {\mathbb R}\), \(e^x\le 1+x+x^2e^{|x|}\). (For \(|x|\le 1\) this is the Taylor bound \(|e^x-1-x|\le x^2\); for \(x\ge 1\), \(e^x\le x^2e^x\); for \(x\le -1\), \(e^x\le 1\le 1+x+x^2\).)

Proof ▶
Theorem 875
✓

Let \(\rho \in {\mathcal P}(\Omega )\) and \(Y\) be measurable with \({\mathbb E}_\rho Y=0\) and \(K:={\mathbb E}_\rho [Y^2e^{|Y|}]{\lt}\infty \). Then for \(|u|\le 1\), \({\mathbb E}_\rho e^{uY}\le 1+Ku^2\), hence \(\log {\mathbb E}_\rho e^{uY}\le Ku^2\). (Integrate \(e^{uY}\le 1+uY+u^2Y^2e^{|Y|}\).)

Proof ▶

‘|Y| ≤ 1 + Y² ≤ 1 + Y² e^|Y|‘, so ‘Y ∈ L¹(ρ)‘.

The pointwise bound ‘e^uY ≤ 1 + uY + u² Y² e^|Y|‘ for ‘|u| ≤ 1‘.

Theorem 876
✓

For \(\sigma \)-finite \(\mu ,\nu \) and a measurable isomorphism \(e\), \(\operatorname {KL}(e_*\mu \| e_*\nu )=\operatorname {KL}(\mu \| \nu )\).

Proof ▶
Theorem 877
✓
#

Let \(\rho \) be a finite measure on \(\alpha \times \Omega \) with \(\Omega \) standard Borel, \(\nu _1\) finite and \(\nu _2\) a probability measure. Then \(\operatorname {KL}(\rho ^{(1)}\| \nu _1)\le \operatorname {KL}(\rho \| \nu _1\otimes \nu _2)\), where \(\rho ^{(1)}\) is the first marginal. (Chain rule: \(\operatorname {KL}(\rho \| \nu _1\otimes \nu _2)=\operatorname {KL}(\rho ^{(1)}\| \nu _1) +\operatorname {KL}(\rho \| \rho ^{(1)}\otimes \nu _2)\) with \(\rho =\rho ^{(1)}\otimes _{\mathrm m}\kappa \) its disintegration.)

Proof ▶
Theorem 878
✓
#

Under the hypotheses of Lemma 877 with the roles of the factors exchanged (\(\alpha \) standard Borel, \(\nu _1\) a probability measure), \(\operatorname {KL}(\rho ^{(2)}\| \nu _2)\le \operatorname {KL}(\rho \| \nu _1\otimes \nu _2)\).

Proof ▶
Theorem 879
✓

Let \(\Omega \) be standard Borel, \(\nu \in {\mathcal P}(\Omega )\) and \(\mu \in {\mathcal P}(\Omega ^n)\) with marginals \(\mu ^{(i)}\). Then \(\sum _{i=1}^n\operatorname {KL}(\mu ^{(i)}\| \nu )\le \operatorname {KL}(\mu \| \nu ^{\otimes n})\). (Induction on \(n\) with the chain rule and Lemmas 877, 878.)

Proof ▶
Theorem 880
✓
#

Let \(E\) be a real inner product space of dimension \(d\) and let \(n\ge d+2\). For all points \((x_i,s_i)\in E\times {\mathbb R}\), \(i=1,\dots ,n\), there is a pattern \(T\subseteq [n]\) such that no \((w,b)\in E\times {\mathbb R}\) satisfies \(s_i\le \langle w,x_i\rangle -b\iff i\in T\) for all \(i\). In other words, the affine class \(\{ x\mapsto \langle w,x\rangle -b\} \) has pseudo-dimension at most \(d+1\).

Proof ▶

The \(n{\gt}d+1\) vectors \((x_i,1)\in E\times {\mathbb R}\) are linearly dependent: there is \(c\ne 0\) with \(\sum _ic_ix_i=0\) and \(\sum _ic_i=0\), hence \(\sum _ic_i(\langle w,x_i\rangle -b-s_i)=-\kappa \) with \(\kappa :=\sum _ic_is_i\) for all \((w,b)\). Replacing \(c\) by \(-c\) we may assume \(\kappa \ge 0\). Since \(c\ne 0\) sums to zero, some \(c_{i_0}\) is negative. If \((w,b)\) realized the pattern \(T=\{ i:c_i\ge 0\} \), every term \(c_i(\langle w,x_i\rangle -b-s_i)\) would be nonnegative and the term \(i_0\) positive, so the sum would be positive, contradicting \(-\kappa \le 0\).

Theorem 881
✓
#

Let \(E\) be a real inner product space of dimension \(d\) and let \(S\subset E\) have \(d+2\) points. There is \(T\subseteq S\) such that for no \((w,c)\in E\times {\mathbb R}\) one has \(\{ v\in S:\langle w,v\rangle +c{\gt}0\} =T\). Hence the VC dimension of the class of affine halfspaces of \(E\) is at most \(d+1\).

Proof ▶

Enumerate \(S=\{ x_1,\dots ,x_{d+2}\} \). A halfspace pattern \(\langle w,x_i\rangle +c{\gt}0\iff i\in T\) is the complementary affine pattern \(0\le \langle -w,x_i\rangle -c\iff i\notin T\), so shattering \(S\) by halfspaces would pseudo-shatter the points \((x_i,0)\) by the affine class, contradicting Lemma 880.

Definition 882
✓
#

For a measure \(\mu \) and \(f\ge 0\) with \(f,\ f\log f\in L^1(\mu )\),

\[ \operatorname {Ent}_\mu (f):=\int f\log f\, \mathrm d\mu -\Bigl(\int f\, \mathrm d\mu \Bigr)\log \Bigl(\int f\, \mathrm d\mu \Bigr) \]

(with \(0\log 0=0\)).

Theorem 883
✓

\(\mu \) satisfies \(\mathrm{LSI}(\alpha )\) if and only if \(\operatorname {Ent}_\mu (g^2)\le \frac2\alpha \int |\nabla g|^2\, \mathrm d\mu \) for all \(g\in C^1_c\).

Proof ▶
Theorem 884
✓

For \(u\ge 0\) and \(v\in {\mathbb R}\), \(uv\le u\log u-u+e^v\).

Proof ▶
Theorem 885
✓
#

For \(u\ge 0\) and \(t{\gt}0\), \(u\log u-u\log t-u+t\ge 0\).

Proof ▶
Theorem 886
✓

For \(0\le u\le v\), \(|u\log u|\le 1+|v\log v|\).

Proof ▶
Theorem 887
✓

For \(f\ge 0\) and every \(t{\gt}0\), \(\operatorname {Ent}_\mu (f)\le \int f\log f\, \mathrm d\mu -\bigl(\int f\, \mathrm d\mu \bigr)\log t-\int f\, \mathrm d\mu +t\); for a probability measure this is \(\operatorname {Ent}_\mu (f)\le \int \bigl(f\log \tfrac ft-f+t\bigr)\, \mathrm d\mu \), with equality for \(t=\int f\, \mathrm d\mu \).

Proof ▶
Theorem 888
✓

For a probability measure \(\mu \), \(f\ge 0\) and \(h\) with \(f,\ f\log f,\ fh,\ e^h \in L^1(\mu )\), \(\int fh\, \mathrm d\mu \le \operatorname {Ent}_\mu (f)+\bigl(\int f\, \mathrm d\mu \bigr)\bigl(\int e^h\, \mathrm d\mu -1\bigr)\).

Proof ▶

Step 1: if ‘∫ f dμ = 0‘ then ‘f = 0‘ a.e. and both sides vanish.

Step 2: Young’s inequality with ‘v = h + log ∫ f dμ‘, integrated.

Theorem 889
✓

For a probability measure \(\mu \) and \(f\ge 0\) with \(f,\ f\log f\in L^1(\mu )\), \(\operatorname {Ent}_\mu (f)\ge 0\).

Proof ▶

Let \(\mu ,\nu \) be probability measures with \(\, \mathrm d\nu =h\, \mathrm d\mu \), \(0\le h\le K\). Then \(\operatorname {Ent}_\nu (f)\le K\, \operatorname {Ent}_\mu (f)\) for every \(f\ge 0\) with \(f,\ f\log f\in L^1(\mu )\cap L^1(\nu )\).

Proof ▶

With \(t=\int f\, \mathrm d\mu {\gt}0\) and \(\varphi _t(u)=u\log u-u\log t-u+t\ge 0\), \(\operatorname {Ent}_\nu (f)\le \int \varphi _t(f)\, \mathrm d\nu =\int h\, \varphi _t(f)\, \mathrm d\mu \le K\int \varphi _t(f)\, \mathrm d\mu =K\, \operatorname {Ent}_\mu (f)\). If \(\int f\, \mathrm d\mu =0\) both entropies vanish.

Theorem 891
✓

Let \(\, \mathrm d\nu =h\, \mathrm d\mu \) with \(0{\lt}c\le h\le K\). Then \(\int \psi \, \mathrm d\mu \le c^{-1}\int \psi \, \mathrm d\nu \) for every integrable \(\psi \ge 0\).

Proof ▶

(F3, Holley–Stroock.) If \(\mu \) satisfies \(\mathrm{LSI}(\alpha )\), \(\psi \) is measurable with \(m\le \psi \le M\) and \(\nu :=Z^{-1}e^{-\psi }\mu \), then \(\nu \) satisfies \(\mathrm{LSI}(\alpha e^{-(M-m)})\).

Proof ▶

The density \(h=e^{-\psi }/Z\) of \(\nu \) satisfies \(e^{-M}/Z\le h\le e^{-m}/Z\). By Lemma 890, \(\operatorname {Ent}_\nu (g^2)\le \frac{e^{-m}}Z\operatorname {Ent}_\mu (g^2) \le \frac{e^{-m}}Z\frac2\alpha \int |\nabla g|^2\, \mathrm d\mu \), and by Lemma 891, \(\int |\nabla g|^2\, \mathrm d\mu \le Ze^{M}\int |\nabla g|^2 \, \mathrm d\nu \); the normalization cancels and the constants multiply to \(e^{M-m}\).

Theorem 893
✓

If \(g\) and \(T\) are differentiable and \(\| DT(x)\| \le L\) for all \(x\), then \(|\nabla (g\circ T)(x)|\le L\, |\nabla g(T(x))|\).

Proof ▶

(F4, transport.) If \(\mu \) satisfies \(\mathrm{LSI}(\alpha )\) and \(T\) is a proper \(C^1\) map with \(\| DT\| \le L\), then \(T_\# \mu \) satisfies \(\mathrm{LSI}(\alpha /L^2)\).

Proof ▶

For \(g\in C^1_c\) on the target, \(g\circ T\in C^1_c\) and \(|\nabla (g\circ T)|\le L\, |\nabla g|\circ T\) (Lemma 893), so \(\operatorname {Ent}_{T_\# \mu }(g^2)=\operatorname {Ent}_\mu ((g\circ T)^2)\le \frac2\alpha \int |\nabla (g\circ T)|^2\, \mathrm d\mu \le \frac{2L^2}\alpha \int |\nabla g|^2\, \mathrm dT_\# \mu \).

(F4, transport.) If \(\mu \) satisfies \(\mathrm{LSI}(\alpha )\) and \(T\) is a proper \(C^1\) map which is Lipschitz with constant \(L\), then \(T_\# \mu \) satisfies \(\mathrm{LSI}(\alpha /L^2)\).

Proof ▶

For probability measures \(\mu ,\nu \) and a bounded measurable \(f\ge 0\) on \(X\times Y\),

\[ \operatorname {Ent}_{\mu \otimes \nu }(f)\le \int \operatorname {Ent}_\mu (f(\cdot ,y))\, \mathrm d\nu (y)+\int \operatorname {Ent}_\nu (f(x,\cdot ))\, \mathrm d\mu (x). \]
Proof ▶

With \(F(y)=\int f(x,y)\, \mathrm d\mu (x)\) and \(a=\int f\, \mathrm d\mu \otimes \nu \), Fubini gives \(\operatorname {Ent}_{\mu \otimes \nu }(f)=\int \operatorname {Ent}_\mu (f(\cdot ,y))\, \mathrm d\nu (y)+\operatorname {Ent}_\nu (F)\). For \(M{\gt}0\) let \(h=\log \max \{ F/a,e^{-M}\} \), so that \(\int e^h\, \mathrm d\nu \le 1+e^{-M}\) and \(\operatorname {Ent}_\nu (F)\le \int Fh\, \mathrm d\nu =\int \! \! \int f(x,y)h(y)\, \mathrm d\nu (y)\, \mathrm d\mu (x)\). By the variational inequality \(\int f(x,\cdot )h\, \mathrm d\nu \le \operatorname {Ent}_\nu (f(x,\cdot ))+e^{-M}\int f(x,\cdot )\, \mathrm d\nu \) (Lemma 888), hence \(\operatorname {Ent}_\nu (F)\le \int \operatorname {Ent}_\nu (f(x,\cdot ))\, \mathrm d\mu (x) +ae^{-M}\); let \(M\to \infty \).

Theorem 897
✓

For a differentiable \(g\) on the Euclidean product \(E_1\times E_2\), \(|\nabla g(x,y)|^2=|\nabla _xg(\cdot ,y)(x)|^2+|\nabla _yg(x,\cdot )(y)|^2\).

Proof ▶

(F2, tensorization.) If \(\mu _1\) satisfies \(\mathrm{LSI}(\alpha _1)\) on \(E_1\) and \(\mu _2\) satisfies \(\mathrm{LSI}(\alpha _2)\) on \(E_2\), then \(\mu _1\otimes \mu _2\) satisfies \(\mathrm{LSI}(\min \{ \alpha _1,\alpha _2\} )\) on the Euclidean product \(E_1\times E_2\).

Proof ▶

For \(g\in C^1_c(E_1\times E_2)\), Lemma 896 with \(f=g^2\), the LSI of \(\mu _1\) and \(\mu _2\) applied to the slices \(g(\cdot ,y)\), \(g(x,\cdot )\), Fubini and Lemma 897 give \(\operatorname {Ent}_{\mu _1\otimes \mu _2}(g^2)\le \frac2{\alpha _1}\int |\nabla _xg|^2+\frac2{\alpha _2}\int |\nabla _yg|^2\le \frac2{\min \{ \alpha _1,\alpha _2\} }\int |\nabla g|^2\, \mathrm d\mu _1\otimes \mu _2\).

(F2, tensorization.) If \(\mu _1\) satisfies \(\mathrm{LSI}(\alpha _1)\) and \(\mu _2\) satisfies \(\mathrm{LSI}(\alpha _2)\), then \(\mu _1\otimes \mu _2\) satisfies \(\mathrm{LSI}(\min \{ \alpha _1,\alpha _2\} )\) on \(\Theta ={\mathbb R}\times Z\).

Proof ▶
Definition 900
✓
#

For \(R{\gt}0\) let \(\chi _R(x):=\phi (2-|x|^2/R^2)\), where \(\phi \in C^\infty ({\mathbb R})\) is nondecreasing with \(\phi =0\) on \((-\infty ,0]\) and \(\phi =1\) on \([1,\infty )\). Then \(0\le \chi _R\le 1\), \(\chi _R=1\) on \(\{ |x|\le R\} \), \(\chi _R=0\) on \(\{ |x|\ge \sqrt2R\} \) and \(|\nabla \chi _R|\le C/R\) with \(C=4\sup |\phi '|\).

Theorem 901
✓

There is \(C\ge 0\) such that \(|\nabla \chi _R(x)|\le C/R\) for all \(R{\gt}0\) and \(x\).

Proof ▶

\(\nabla \chi _R(x)=-\phi '(2-|x|^2/R^2)\, 2x/R^2\) vanishes for \(|x|{\gt}2R\) (where the argument of \(\phi '\) is negative) and is bounded by \(\sup |\phi '|\cdot 4R/R^2\) otherwise.

Let \(\mu \) satisfy \(\mathrm{LSI}(\alpha )\) on a finite-dimensional space, and let \(\rho =p\, \mu \) be a probability measure with \(p{\gt}0\) of class \(C^1\), \(\nabla \log p\in L^2(\rho )\) and \(\operatorname {KL}(\rho \| \mu ){\lt}\infty \). Then \(\operatorname {KL}(\rho \| \mu )\le \frac1{2\alpha }\int |\nabla \log p|^2\, \mathrm d\rho \).

Proof ▶

Apply the LSI to \(g_n=\chi _n\sqrt p\) with the cutoffs \(\chi _n=\chi _{n+1}\) of Definition 900. By dominated convergence (\(|u\log u|\le 1+|p\log p|\) for \(0\le u\le p\), and \(p\log p\in L^1(\mu )\) since \(\operatorname {KL}(\rho \| \mu ){\lt}\infty \)), \(\operatorname {Ent}_\mu (g_n^2)\to \int p\log p\, \mathrm d\mu =\operatorname {KL}(\rho \| \mu )\). Since \(\nabla g_n=\chi _n\nabla \sqrt p+\sqrt p\, \nabla \chi _n\) and \(|\nabla \chi _n|\le C/(n+1)\) (Lemma 901), \(|\nabla g_n|^2\le |\nabla \sqrt p|^2+\frac C{n+1}(p+|\nabla \sqrt p|^2)+\frac{C^2}{(n+1)^2}p\), whose integral tends to \(\int |\nabla \sqrt p|^2\, \mathrm d\mu =\frac14\int |\nabla \log p|^2\, \mathrm d\rho \).

(L1.) If \(\mu \) satisfies \(\mathrm{LSI}(\alpha )\) on a finite-dimensional space and \(\rho \ll \mu \) is a probability measure with a smooth density and \(\operatorname {KL}(\rho \| \mu ){\lt}\infty \), then \(\operatorname {KL}(\rho \| \mu )\le \frac1{2\alpha }I(\rho |\mu )\).

Proof ▶

Lemma 902 applied to the density defining \(I(\rho |\mu )\).

Theorem 904
✓
#

Let \(\rho ,\rho ',\mu \) be probability measures with \(\operatorname {KL}(\rho \| \mu ){\lt}\infty \) and \(\operatorname {KL}(\rho '\| \mu ){\lt}\infty \), and let \({\varepsilon }\in [0,1]\). Then \(\operatorname {KL}((1-{\varepsilon })\rho +{\varepsilon }\rho '\| \mu )\le (1-{\varepsilon })\operatorname {KL}(\rho \| \mu )+{\varepsilon }\operatorname {KL}(\rho '\| \mu ){\lt}\infty \). (Proof: \(\frac{\, \mathrm d\rho _{\varepsilon }}{\, \mathrm d\mu }=(1-{\varepsilon })\frac{\, \mathrm d\rho }{\, \mathrm d\mu }+{\varepsilon }\frac{\, \mathrm d\rho '}{\, \mathrm d\mu }\) and \(\operatorname {KL}(\rho \| \mu )=\int \phi (\frac{\, \mathrm d\rho }{\, \mathrm d\mu })\, \mathrm d\mu \) with \(\phi (x)=x\log x+1-x\) convex.)

Proof ▶

Step 1: the density of the mixture is the mixture of the densities.

Step 2: pointwise convexity of ‘klFun‘.

Step 3: integrability of both sides.

Step 4: integrate.

Theorem 905
✓

For \(\nu =w\, \nu _0\) with \(w{\gt}0\), \(\operatorname {KL}(\nu \| \nu _0)=\int (w\log w+1-w)\, \mathrm d\nu _0\le \int (w-1)^2\, \mathrm d\nu _0=\chi ^2(\nu \| \nu _0)\), since \(\log w\le w-1\).

Proof ▶
Theorem 906
✓

For \(v{\gt}0\) and \(t\ge 0\), \({\mathcal N}(0,v)(\{ |x|{\gt}t\} )=2\bar\Phi (t/\sqrt v) \le e^{-t^2/(2v)}\). (For \(x\ge t\ge 0\), \(x^2\ge t^2+(x-t)^2\), so \(\phi _v(x)\le e^{-t^2/(2v)}\phi _v(x-t)\); integrate over \(x{\gt}t\) and use \({\mathcal N}(0,v)(0,\infty )\le \tfrac 12\).)

Proof ▶
Theorem 907
✓

For \(v{\gt}0\) and \(t\ge 0\), \(\int _t^\infty x\, {\mathcal N}(0,v)(\, \mathrm dx)=\sqrt{v/(2\pi )}\, e^{-t^2/(2v)}=v\, \phi _v(t)\).

Proof ▶
Theorem 908
✓

Let \(v{\gt}0\), \(|m|\le m_0\le T\) and \(X\sim {\mathcal N}(m,v)\). Then \({\mathbb E}\bigl[|X|\, 1_{\{ |X|{\gt}T\} }\bigr]\le \bigl(m_0+\sqrt{2v/\pi }\bigr)e^{-(T-m_0)^2/(2v)}\). (Write \(X=m+W\) with \(W\sim {\mathcal N}(0,v)\): \(|X|{\gt}T\) forces \(|W|{\gt}t:=T-m_0\) and \(|X|\le m_0+|W|\), so the left side is at most \(m_0{\mathbb P}(|W|{\gt}t)+{\mathbb E}[|W|1_{\{ |W|{\gt}t\} }]\), and Lemmas 906 and 907 bound the two terms.)

Proof ▶
Theorem 909
✓
#

Let \(K\ge 0\) and \(\Phi \colon [0,\infty )\to {\mathbb R}\) with \(\Phi (t)\, (1+K(t-s))\le \Phi (s)\) for all \(0\le s\le t\). Then \(\Phi (t)\le e^{-Kt}\Phi (0)\) for all \(t\ge 0\). (Iterating on the partition \(kt/n\) gives \(\Phi (t)(1+Kt/n)^n\le \Phi (0)\); let \(n\to \infty \).)

Proof ▶

Step 1: iterate on the partition ‘k t / n‘: ‘Φ(k t/n) (1 + K t/n)^k ≤ Φ(0)‘.

Step 2: ‘Φ t ≤ Φ 0 / (1 + K t / n)^n‘ for every ‘n ≥ 1‘.

Step 3: let ‘n → ∞‘, using ‘(1 + x/n)^n → e^x‘.

Theorem 910
✓

For \(\rho ^2+c^2=1\) and every \(x\in {\mathbb R}\), \(\int He_n(\rho x+cz)\, \gamma (\, \mathrm dz)=\rho ^nHe_n(x)\): the Hermite polynomials are the eigenfunctions of the Ornstein–Uhlenbeck semigroup. (Two-step induction with the recursion \(He_{n+2}=XHe_{n+1}-(n+1)He_n\) and Gaussian integration by parts.)

Proof ▶
Definition 911
✓
#

For \(\rho \in [-1,1]\), the law of \((X,\rho X+\sqrt{1-\rho ^2}Z)\) with \(X,Z\) independent standard Gaussians: the standard Gaussian pair with correlation \(\rho \).

Theorem 912
✓

Both marginals of the correlated standard Gaussian pair are \(\gamma \).

Proof ▶

For \(|\rho |\le 1\) and standard Gaussians \((X,Y)\) with correlation \(\rho \), \({\mathbb E}[He_m(X)He_n(Y)]=\rho ^n\, n!\, \delta _{mn}\). (Proof: \({\mathbb E}[He_n(\rho x+\sqrt{1-\rho ^2}Z)] =\rho ^nHe_n(x)\) by the three-term recursion and Gaussian integration by parts, then orthogonality.)

Proof ▶
Theorem 914
✓

If \(g\in L^2(\gamma )\) and \(\int gHe_n\, \, \mathrm d\gamma =0\) for all \(n\ge 0\), then \(g=0\) \(\gamma \)-a.e.: the Hermite polynomials are complete in \(L^2(\gamma )\). (Every monomial is a combination of Hermite polynomials, so all moments of \(g\gamma \) vanish; by dominated convergence its Fourier transform vanishes, and the positive and negative parts of \(g\gamma \) have the same characteristic function, hence coincide.)

Proof ▶
Definition 915
✓
#

\(a_n(f):=\frac1{\sqrt{n!}}\int f\, He_n\, \, \mathrm d\gamma \), the coefficient of \(f\in L^2(\gamma )\) on the normalized Hermite polynomial \(He_n/\sqrt{n!}\).

Definition 916
✓

The normalized Hermite polynomials \(He_n/\sqrt{n!}\), \(n\ge 0\), form an orthonormal basis of \(L^2(\gamma )\) (orthogonality and completeness).

For \(f\in L^2(\gamma )\), \(\| f\| _{L^2(\gamma )}^2=\sum _{n\ge 0}a_n(f)^2\), the series converging.

Proof ▶

For \(F,G\in L^2(\gamma )\) and \(|\rho |\le 1\), \({\mathbb E}[F(X)G(Y)]=\sum _{n\ge 0}\langle He_n/\sqrt{n!},F\rangle \langle He_n/\sqrt{n!},G\rangle \rho ^n\). (Expand \(F\) and \(G\) in the Hermite basis, use \({\mathbb E}[He_m(X)He_n(Y)]=\rho ^n\, n!\, \delta _{mn}\) and the continuity of \((F,G)\mapsto {\mathbb E}[F(X)G(Y)]\) on \(L^2(\gamma )\times L^2(\gamma )\).)

Proof ▶

For \(f,g\in L^2(\gamma )\), \(|\rho |\le 1\) and standard Gaussians \((X,Y)\) with correlation \(\rho \), \({\mathbb E}[f(X)g(Y)]=\sum _{n\ge 0}a_n(f)\, a_n(g)\, \rho ^n\), the series converging (absolutely, by Parseval and \(|\rho |\le 1\)).

Proof ▶