Tutorials › Real Analysis › Almost-Everywhere Convergence Versus Lp Convergence

Lp Spaces · Tutorial 895 of 1000

Almost-Everywhere Convergence Versus Lp Convergence

See how almost-everywhere and \(L^p\) convergence differ, when one yields the other, and why \(L^p\) convergence guarantees an almost-everywhere convergent subsequence.

Advanced 9 min read

What You'll Learn

  • Distinguish almost-everywhere convergence from convergence in finite \(L^p\).
  • Prove that every \(L^p\)-convergent sequence has an almost-everywhere convergent subsequence.
  • Construct an \(L^p\)-convergent sequence whose full sequence fails to converge almost everywhere.
  • See why almost-everywhere convergence alone need not imply \(L^p\) convergence.
  • Apply an integrable dominating function to obtain \(L^p\) convergence.
  • Understand why \(L^\infty\) convergence has a stronger pointwise consequence.

Two Different Ways for Functions to Converge

Convergence in \(L^p\) measures the total size of the error through its \(p\)th-power integral. Almost-everywhere convergence instead asks what happens at nearly every individual point. The previous tutorial established that finite-\(p\) convergence implies convergence in measure, but convergence in measure is not the same as almost-everywhere convergence. Here we compare these modes directly.

Throughout the discussion of finite exponents, take \(1\leq p<\infty\), and let \(f_n\) and \(f\) be measurable functions on a measure space \((X,\mathcal{F},\mu)\). Convergence almost everywhere means that there is a measurable null set \(N\) such that \(f_n(x)\to f(x)\) for every \(x\notin N\). As with \(L^p\) convergence, changes on null sets do not affect the equivalence classes in \(L^p\); pointwise statements are understood using representatives.

Definition (Almost-Everywhere Convergence): The sequence \((f_n)\) converges to \(f\) almost everywhere if there exists \(N\in\mathcal{F}\) with \(\mu(N)=0\) such that, for every \(x\in X\setminus N\) and every \(\varepsilon>0\), there is an index \(n_0\) for which $$ n\geq n_0\quad\Longrightarrow\quad |f_n(x)-f(x)|<\varepsilon. $$

The essential difference is the order of the quantifiers. For almost-everywhere convergence, after excluding one null set, the index \(n_0\) may depend on the point \(x\). In \(L^p\) convergence, the integral error must tend to zero across the whole space. Neither condition, by itself, requires the pointwise errors to be uniformly small.

An Almost-Everywhere Convergent Example

Worked Example: Powers on the Unit Interval

On \([0,1]\) with Lebesgue measure, let \(f_n(x)=x^n\) and \(f(x)=0\). For every \(x\in[0,1)\), \(x^n\to0\), while at \(x=1\) the sequence is constantly \(1\). Since the exceptional singleton \(\{1\}\) has measure zero, \(f_n\to0\) almost everywhere.

For any finite \(p\geq1\), direct integration gives $$ \|f_n\|_p^p=\int_0^1 x^{np}\,dx=\frac{1}{np+1}. $$ The right-hand side tends to zero, so \(\|f_n\|_p=(np+1)^{-1/p}\to0\). Thus this sequence converges both almost everywhere and in \(L^p\). The calculation illustrates one situation in which the two modes agree; it does not show that almost-everywhere convergence always controls the integral error.

Every \(L^p\)-Convergent Sequence Has an Almost-Everywhere Convergent Subsequence

Although \(L^p\) convergence need not imply almost-everywhere convergence of the full sequence, it always allows a useful extraction. The idea is to choose a subsequence whose errors exceed successively smaller thresholds on sets with summable measures. Countable subadditivity then shows that almost every point belongs to only finitely many of those exceptional sets.

Theorem (Almost-Everywhere Convergent Subsequence): Let \(1\leq p<\infty\), and suppose \(f_n\to f\) in \(L^p(X,\mu)\). There is a subsequence \((f_{n_k})\) such that \(f_{n_k}\to f\) almost everywhere.

Proof. Since \(\|f_n-f\|_p\to0\), choose strictly increasing indices \(n_k\) such that $$ \|f_{n_k}-f\|_p^p\leq 2^{-2k} $$ for every positive integer \(k\). This is possible because the \(p\)th powers of the norms also tend to zero. Define $$ E_k=\{x\in X:|f_{n_k}(x)-f(x)|>2^{-k/p}\}. $$ The level-set estimate for \(L^p\), established in the previous tutorial, gives $$ \mu(E_k) \leq \frac{\|f_{n_k}-f\|_p^p}{(2^{-k/p})^p} \leq \frac{2^{-2k}}{2^{-k}} =2^{-k}. $$ For each \(m\), countable subadditivity yields $$ \mu\left(\bigcup_{k=m}^{\infty}E_k\right) \leq \sum_{k=m}^{\infty}2^{-k} =2^{1-m}. $$ The set of points belonging to infinitely many \(E_k\) is $$ \bigcap_{m=1}^{\infty}\bigcup_{k=m}^{\infty}E_k. $$ It is contained in \(\bigcup_{k=m}^{\infty}E_k\) for every \(m\), so its measure is at most \(2^{1-m}\) for every \(m\). Letting \(m\to\infty\), its measure is zero. Outside this null set, each point belongs to only finitely many \(E_k\). At each such point \(x\), there is an index \(K\) such that for every \(k\geq K\), $$ |f_{n_k}(x)-f(x)|\leq 2^{-k/p}. $$ Since \(2^{-k/p}\to0\), this proves \(f_{n_k}(x)\to f(x)\) outside a null set. Therefore the subsequence converges almost everywhere. \(\square\)

The conclusion is about a subsequence, not necessarily the original sequence. A different subsequence may be needed for each convergence problem; the theorem guarantees existence, not a prescribed choice. The proof uses the level-set estimate and summable bounds, and does not require the measure of \(X\) to be finite.

The Full Sequence Need Not Converge Almost Everywhere

The subsequence theorem cannot generally be strengthened to say that the entire \(L^p\)-convergent sequence converges almost everywhere. A sequence can have small \(L^p\) norm at every stage while its nonzero values visit almost every point infinitely often.

Worked Example: Small Norms with Repeated Pointwise Excursions

Work on \([0,1]\) with Lebesgue measure, ignoring the null set of points with two binary expansions. Divide the binary digits into consecutive disjoint blocks, where block \(j\) has length \(k_j=\lceil\log_2(j+1)\rceil\). Let \(A_j\) be the set of points whose digits in block \(j\) are all zero. The elementary lengths of binary intervals give $$ \mu(A_j)=2^{-k_j}. $$ Because \(j+1\leq 2^{k_j}<2(j+1)\), these measures satisfy $$ \frac{1}{2(j+1)}<\mu(A_j)\leq\frac{1}{j+1}. $$ In particular, \(\mu(A_j)\to0\) and \(\sum_{j=1}^{\infty}\mu(A_j)=\infty\).

Events determined by disjoint finite blocks of binary digits are independent: prescribing any finite selection of these blocks fixes the corresponding distinct digits, and the measure of the resulting binary intervals is the product of the prescribed block probabilities. Consequently, for fixed \(m\leq N\), $$ \mu\left(\bigcap_{j=m}^{N}A_j^c\right) =\prod_{j=m}^{N}(1-\mu(A_j)) \leq \exp\left(-\sum_{j=m}^{N}\mu(A_j)\right). $$ The right-hand side tends to zero as \(N\to\infty\), since the sum diverges. Thus, for each \(m\), the set of points that belong to no \(A_j\) for \(j\geq m\) has measure zero. Taking the countable union over \(m\) shows that almost every point belongs to infinitely many \(A_j\).

Set \(f_j=\mathbf{1}_{A_j}\). For each finite \(p\geq1\), the indicator integral formula gives $$ \|f_j\|_p^p=\int_0^1\mathbf{1}_{A_j}\,d\mu=\mu(A_j)\longrightarrow0. $$ Therefore \(f_j\to0\) in \(L^p\). But \(f_j(x)=1\) for infinitely many \(j\) at almost every \(x\), so the full sequence does not converge to \(0\) at almost every point. This example makes the distinction between a convergent subsequence and convergence of the full sequence explicit.

Almost-Everywhere Convergence Need Not Give \(L^p\) Convergence

The reverse implication fails as well. A function can become very large on a set of very small measure. Those sets may shrink enough for the values to approach zero pointwise, while the integral of the \(p\)th power remains fixed.

Worked Example: Shrinking Supports with Fixed \(L^p\) Norm

On \((0,1)\) with Lebesgue measure, fix \(1\leq p<\infty\) and define $$ f_n(x)=n^{1/p}\mathbf{1}_{(0,1/n)}(x). $$ For each fixed \(x\in(0,1)\), eventually \(1/n\leq x\), so \(x\notin(0,1/n)\) and \(f_n(x)=0\) for all sufficiently large \(n\). Hence \(f_n\to0\) everywhere on \((0,1)\). On the other hand, $$ \|f_n\|_p^p =\int_0^1 n\,\mathbf{1}_{(0,1/n)}(x)\,dx =n\cdot\frac{1}{n} =1. $$ The norms do not tend to zero. Thus almost-everywhere convergence does not imply convergence in \(L^p\), even on a space of finite measure.

A standard way to rule out this concentration of error is an integrable dominating function. The Dominated Convergence Theorem, established earlier in the course, then converts almost-everywhere convergence into \(L^p\) convergence.

Theorem (Dominated Almost-Everywhere Convergence Implies \(L^p\) Convergence): Let \(1\leq p<\infty\). Suppose \(f_n\to f\) almost everywhere and there is a function \(g\in L^p(X,\mu)\) such that \(|f_n|\leq g\) almost everywhere for every \(n\). Then \(f\in L^p(X,\mu)\) and \(\|f_n-f\|_p\to0\).

Proof. Excluding the union of the null set in the domination assumptions and the null set where convergence fails, we have \(|f_n(x)|\leq g(x)\) for every \(n\), and \(f_n(x)\to f(x)\). Taking the limit gives \(|f(x)|\leq g(x)\) there. Monotonicity of the integral under almost-everywhere order implies \(\int_X|f|^p\,d\mu\leq\int_Xg^p\,d\mu<\infty\), so \(f\in L^p\). For almost every \(x\), the scalar triangle inequality and the bound \(|f(x)|\leq g(x)\) give $$ |f_n(x)-f(x)|^p \leq (|f_n(x)|+|f(x)|)^p \leq (2g(x))^p =2^p g(x)^p. $$ The function \(2^p g^p\) is integrable, and \(|f_n-f|^p\to0\) almost everywhere. The Dominated Convergence Theorem therefore gives $$ \int_X|f_n-f|^p\,d\mu\longrightarrow0. $$ Taking \(p\)th roots proves \(\|f_n-f\|_p\to0\). \(\square\)

Why \(L^\infty\) Is Different

For finite \(p\), the integral can be small because the error is concentrated on small sets. The essential supremum does not average the error in this way. In fact, convergence in \(L^\infty\) implies almost-everywhere convergence of the full sequence, for any chosen measurable representatives.

Proposition (Essential-Supremum Convergence Implies Almost-Everywhere Convergence): If \(f_n\to f\) in \(L^\infty(X,\mu)\), then \(f_n(x)\to f(x)\) almost everywhere.

Proof. Write \(a_n=\|f_n-f\|_\infty\), so \(a_n\to0\). By the definition of essential supremum, for each \(n\) there is a null set \(N_n\) outside which \(|f_n-f|\leq a_n+1/n\). The union \(N=\bigcup_{n=1}^{\infty}N_n\) is null. For each \(x\notin N\), the inequality holds for every \(n\), and its right-hand side tends to zero. Therefore \(f_n(x)\to f(x)\). \(\square\)

This implication does not mean that almost-everywhere convergence implies \(L^\infty\) convergence; the shrinking-support example has essential supremum \(n^{1/p}\), not a quantity tending to zero. The modes of convergence must be compared with their defining measures of error, rather than treated as interchangeable notions.

Key takeaway: For finite \(p\), \(L^p\) convergence guarantees an almost-everywhere convergent subsequence, but not necessarily convergence of the full sequence. Almost-everywhere convergence alone does not guarantee \(L^p\) convergence; an integrable \(L^p\) dominating function is one useful sufficient condition. In contrast, \(L^\infty\) convergence implies almost-everywhere convergence of the full sequence.

Check Your Understanding

Use the definitions, examples, and results above to distinguish the two modes of convergence.

  1. Why does the subsequence theorem require selecting indices with especially small \(L^p\) errors?
  2. In the binary-block example, why does the measure of the set of points lying in no \(A_j\) after a fixed index equal zero?
  3. For the shrinking-support example, which calculation shows that almost-everywhere convergence does not imply \(L^p\) convergence?
  4. Why does an \(L^p\) dominating function allow the Dominated Convergence Theorem to be applied to \(|f_n-f|^p\)?
  5. What feature of the essential supremum makes \(L^\infty\) convergence imply almost-everywhere convergence of the full sequence?