Two Different Ways for Functions to Converge
Convergence in \(L^p\) measures the total size of the error through its \(p\)th-power integral. Almost-everywhere convergence instead asks what happens at nearly every individual point. The previous tutorial established that finite-\(p\) convergence implies convergence in measure, but convergence in measure is not the same as almost-everywhere convergence. Here we compare these modes directly.
Throughout the discussion of finite exponents, take \(1\leq p<\infty\), and let \(f_n\) and \(f\) be measurable functions on a measure space \((X,\mathcal{F},\mu)\). Convergence almost everywhere means that there is a measurable null set \(N\) such that \(f_n(x)\to f(x)\) for every \(x\notin N\). As with \(L^p\) convergence, changes on null sets do not affect the equivalence classes in \(L^p\); pointwise statements are understood using representatives.
The essential difference is the order of the quantifiers. For almost-everywhere convergence, after excluding one null set, the index \(n_0\) may depend on the point \(x\). In \(L^p\) convergence, the integral error must tend to zero across the whole space. Neither condition, by itself, requires the pointwise errors to be uniformly small.
An Almost-Everywhere Convergent Example
Worked Example: Powers on the Unit Interval
On \([0,1]\) with Lebesgue measure, let \(f_n(x)=x^n\) and \(f(x)=0\). For every \(x\in[0,1)\), \(x^n\to0\), while at \(x=1\) the sequence is constantly \(1\). Since the exceptional singleton \(\{1\}\) has measure zero, \(f_n\to0\) almost everywhere.
For any finite \(p\geq1\), direct integration gives $$ \|f_n\|_p^p=\int_0^1 x^{np}\,dx=\frac{1}{np+1}. $$ The right-hand side tends to zero, so \(\|f_n\|_p=(np+1)^{-1/p}\to0\). Thus this sequence converges both almost everywhere and in \(L^p\). The calculation illustrates one situation in which the two modes agree; it does not show that almost-everywhere convergence always controls the integral error.
Every \(L^p\)-Convergent Sequence Has an Almost-Everywhere Convergent Subsequence
Although \(L^p\) convergence need not imply almost-everywhere convergence of the full sequence, it always allows a useful extraction. The idea is to choose a subsequence whose errors exceed successively smaller thresholds on sets with summable measures. Countable subadditivity then shows that almost every point belongs to only finitely many of those exceptional sets.
Proof. Since \(\|f_n-f\|_p\to0\), choose strictly increasing indices \(n_k\) such that $$ \|f_{n_k}-f\|_p^p\leq 2^{-2k} $$ for every positive integer \(k\). This is possible because the \(p\)th powers of the norms also tend to zero. Define $$ E_k=\{x\in X:|f_{n_k}(x)-f(x)|>2^{-k/p}\}. $$ The level-set estimate for \(L^p\), established in the previous tutorial, gives $$ \mu(E_k) \leq \frac{\|f_{n_k}-f\|_p^p}{(2^{-k/p})^p} \leq \frac{2^{-2k}}{2^{-k}} =2^{-k}. $$ For each \(m\), countable subadditivity yields $$ \mu\left(\bigcup_{k=m}^{\infty}E_k\right) \leq \sum_{k=m}^{\infty}2^{-k} =2^{1-m}. $$ The set of points belonging to infinitely many \(E_k\) is $$ \bigcap_{m=1}^{\infty}\bigcup_{k=m}^{\infty}E_k. $$ It is contained in \(\bigcup_{k=m}^{\infty}E_k\) for every \(m\), so its measure is at most \(2^{1-m}\) for every \(m\). Letting \(m\to\infty\), its measure is zero. Outside this null set, each point belongs to only finitely many \(E_k\). At each such point \(x\), there is an index \(K\) such that for every \(k\geq K\), $$ |f_{n_k}(x)-f(x)|\leq 2^{-k/p}. $$ Since \(2^{-k/p}\to0\), this proves \(f_{n_k}(x)\to f(x)\) outside a null set. Therefore the subsequence converges almost everywhere. \(\square\)
The conclusion is about a subsequence, not necessarily the original sequence. A different subsequence may be needed for each convergence problem; the theorem guarantees existence, not a prescribed choice. The proof uses the level-set estimate and summable bounds, and does not require the measure of \(X\) to be finite.
The Full Sequence Need Not Converge Almost Everywhere
The subsequence theorem cannot generally be strengthened to say that the entire \(L^p\)-convergent sequence converges almost everywhere. A sequence can have small \(L^p\) norm at every stage while its nonzero values visit almost every point infinitely often.
Worked Example: Small Norms with Repeated Pointwise Excursions
Work on \([0,1]\) with Lebesgue measure, ignoring the null set of points with two binary expansions. Divide the binary digits into consecutive disjoint blocks, where block \(j\) has length \(k_j=\lceil\log_2(j+1)\rceil\). Let \(A_j\) be the set of points whose digits in block \(j\) are all zero. The elementary lengths of binary intervals give $$ \mu(A_j)=2^{-k_j}. $$ Because \(j+1\leq 2^{k_j}<2(j+1)\), these measures satisfy $$ \frac{1}{2(j+1)}<\mu(A_j)\leq\frac{1}{j+1}. $$ In particular, \(\mu(A_j)\to0\) and \(\sum_{j=1}^{\infty}\mu(A_j)=\infty\).
Events determined by disjoint finite blocks of binary digits are independent: prescribing any finite selection of these blocks fixes the corresponding distinct digits, and the measure of the resulting binary intervals is the product of the prescribed block probabilities. Consequently, for fixed \(m\leq N\), $$ \mu\left(\bigcap_{j=m}^{N}A_j^c\right) =\prod_{j=m}^{N}(1-\mu(A_j)) \leq \exp\left(-\sum_{j=m}^{N}\mu(A_j)\right). $$ The right-hand side tends to zero as \(N\to\infty\), since the sum diverges. Thus, for each \(m\), the set of points that belong to no \(A_j\) for \(j\geq m\) has measure zero. Taking the countable union over \(m\) shows that almost every point belongs to infinitely many \(A_j\).
Set \(f_j=\mathbf{1}_{A_j}\). For each finite \(p\geq1\), the indicator integral formula gives $$ \|f_j\|_p^p=\int_0^1\mathbf{1}_{A_j}\,d\mu=\mu(A_j)\longrightarrow0. $$ Therefore \(f_j\to0\) in \(L^p\). But \(f_j(x)=1\) for infinitely many \(j\) at almost every \(x\), so the full sequence does not converge to \(0\) at almost every point. This example makes the distinction between a convergent subsequence and convergence of the full sequence explicit.
Almost-Everywhere Convergence Need Not Give \(L^p\) Convergence
The reverse implication fails as well. A function can become very large on a set of very small measure. Those sets may shrink enough for the values to approach zero pointwise, while the integral of the \(p\)th power remains fixed.
Worked Example: Shrinking Supports with Fixed \(L^p\) Norm
On \((0,1)\) with Lebesgue measure, fix \(1\leq p<\infty\) and define $$ f_n(x)=n^{1/p}\mathbf{1}_{(0,1/n)}(x). $$ For each fixed \(x\in(0,1)\), eventually \(1/n\leq x\), so \(x\notin(0,1/n)\) and \(f_n(x)=0\) for all sufficiently large \(n\). Hence \(f_n\to0\) everywhere on \((0,1)\). On the other hand, $$ \|f_n\|_p^p =\int_0^1 n\,\mathbf{1}_{(0,1/n)}(x)\,dx =n\cdot\frac{1}{n} =1. $$ The norms do not tend to zero. Thus almost-everywhere convergence does not imply convergence in \(L^p\), even on a space of finite measure.
A standard way to rule out this concentration of error is an integrable dominating function. The Dominated Convergence Theorem, established earlier in the course, then converts almost-everywhere convergence into \(L^p\) convergence.
Proof. Excluding the union of the null set in the domination assumptions and the null set where convergence fails, we have \(|f_n(x)|\leq g(x)\) for every \(n\), and \(f_n(x)\to f(x)\). Taking the limit gives \(|f(x)|\leq g(x)\) there. Monotonicity of the integral under almost-everywhere order implies \(\int_X|f|^p\,d\mu\leq\int_Xg^p\,d\mu<\infty\), so \(f\in L^p\). For almost every \(x\), the scalar triangle inequality and the bound \(|f(x)|\leq g(x)\) give $$ |f_n(x)-f(x)|^p \leq (|f_n(x)|+|f(x)|)^p \leq (2g(x))^p =2^p g(x)^p. $$ The function \(2^p g^p\) is integrable, and \(|f_n-f|^p\to0\) almost everywhere. The Dominated Convergence Theorem therefore gives $$ \int_X|f_n-f|^p\,d\mu\longrightarrow0. $$ Taking \(p\)th roots proves \(\|f_n-f\|_p\to0\). \(\square\)
Why \(L^\infty\) Is Different
For finite \(p\), the integral can be small because the error is concentrated on small sets. The essential supremum does not average the error in this way. In fact, convergence in \(L^\infty\) implies almost-everywhere convergence of the full sequence, for any chosen measurable representatives.
Proof. Write \(a_n=\|f_n-f\|_\infty\), so \(a_n\to0\). By the definition of essential supremum, for each \(n\) there is a null set \(N_n\) outside which \(|f_n-f|\leq a_n+1/n\). The union \(N=\bigcup_{n=1}^{\infty}N_n\) is null. For each \(x\notin N\), the inequality holds for every \(n\), and its right-hand side tends to zero. Therefore \(f_n(x)\to f(x)\). \(\square\)
This implication does not mean that almost-everywhere convergence implies \(L^\infty\) convergence; the shrinking-support example has essential supremum \(n^{1/p}\), not a quantity tending to zero. The modes of convergence must be compared with their defining measures of error, rather than treated as interchangeable notions.
Check Your Understanding
Use the definitions, examples, and results above to distinguish the two modes of convergence.
- Why does the subsequence theorem require selecting indices with especially small \(L^p\) errors?
- In the binary-block example, why does the measure of the set of points lying in no \(A_j\) after a fixed index equal zero?
- For the shrinking-support example, which calculation shows that almost-everywhere convergence does not imply \(L^p\) convergence?
- Why does an \(L^p\) dominating function allow the Dominated Convergence Theorem to be applied to \(|f_n-f|^p\)?
- What feature of the essential supremum makes \(L^\infty\) convergence imply almost-everywhere convergence of the full sequence?