Tutorials › Real Analysis › Measure Theory Mastery II

Measure Theory · Tutorial 850 of 1000

Measure Theory Mastery II

Learn how convergence in measure is detected by exceptional sets, how to extract almost-everywhere convergent subsequences, and why finite measure changes the relationship between convergence modes.

Advanced 10 min read

What You'll Learn

  • Define convergence in measure using the measures of deviation sets
  • Extract an almost-everywhere convergent subsequence from convergence in measure
  • Use Egorov’s theorem to relate almost-everywhere convergence to convergence in measure on finite spaces
  • Characterize convergence in measure through almost-everywhere convergent subsequences
  • Distinguish convergence in measure from pointwise and uniform convergence
  • Identify why finite total measure is essential for one direction

A Weaker Form of Convergence

The previous tutorial studied when almost-everywhere convergence can be made uniform after removing a set of small measure. We now consider a different question: can a sequence be close to its limit except on a set whose measure becomes small? This leads to convergence in measure. It is weaker than pointwise convergence in some respects, but it has a powerful relationship with almost-everywhere convergence through subsequences.

Let \((X,\mathcal{F},\mu)\) be a measure space, and let \(f_n,f:X\to\mathbb{R}\) be measurable. For a tolerance \(\varepsilon>0\), the set where the \(n\)th function deviates from \(f\) by more than \(\varepsilon\) is measurable: it is the preimage of an open set under the measurable function \(|f_n-f|\). Denote this set by

$$ D_n(\varepsilon)=\{x\in X:|f_n(x)-f(x)|>\varepsilon\}. $$
Definition: The sequence \((f_n)\) converges in measure to \(f\) if, for every \(\varepsilon>0\), \(\mu(D_n(\varepsilon))\to 0\) as \(n\to\infty\).

The tolerance \(\varepsilon\) is fixed while \(n\) increases. The definition says that the measure of the set where the error exceeds that tolerance tends to zero. It does not require the error to become small at every point, nor does it impose one bound that works uniformly across the entire space.

Convergence in Measure Yields an Almost-Everywhere Subsequence

Although convergence in measure need not imply almost-everywhere convergence of the entire sequence, it always permits extraction of a subsequence that converges almost everywhere. This implication does not require the whole space to have finite measure. Its proof uses a sequence of increasingly strict error tolerances and the First Borel–Cantelli Lemma from earlier in the course.

Theorem (Subsequence Extraction from Convergence in Measure): Let \((X,\mathcal{F},\mu)\) be a measure space. If \(f_n\) converges in measure to \(f\), then there is a subsequence \((f_{n_k})\) that converges to \(f\) almost everywhere.

Proof. For each positive integer \(k\), convergence in measure with tolerance \(1/k\) allows us to choose an index \(n_k\), larger than \(n_{k-1}\) when \(k>1\), such that

$$ \mu\bigl(D_{n_k}(1/k)\bigr)<2^{-k}. $$

The indices can be chosen strictly increasing because convergence in measure ensures the displayed bound for every sufficiently large index. By Countable Subadditivity, the measure of the union of the deviation sets from index \(K\) onward satisfies

$$ \mu\left(\bigcup_{k=K}^{\infty}D_{n_k}(1/k)\right) \leq \sum_{k=K}^{\infty}\mu\bigl(D_{n_k}(1/k)\bigr) <\sum_{k=K}^{\infty}2^{-k}. $$

The final sum tends to zero as \(K\to\infty\). Thus the set of points that belong to infinitely many of the sets \(D_{n_k}(1/k)\), their limsup, has measure zero. This is also the conclusion of the First Borel–Cantelli Lemma. Outside that null set, each point \(x\) belongs to only finitely many such deviation sets. Hence, for all sufficiently large \(k\),

$$ |f_{n_k}(x)-f(x)|\leq 1/k. $$

Since \(1/k\to0\), this inequality implies \(f_{n_k}(x)\to f(x)\) at every point outside the null set. Therefore the subsequence converges to \(f\) almost everywhere. \(\square\)

The subsequence is chosen to make the measures of its deviation sets summable. That summability is what lets Borel–Cantelli control how often any fixed point can be a large-error point. Merely knowing that each individual deviation set has small measure, without arranging a summable sequence of bounds, would not give this conclusion directly.

Worked Example: Uniform Convergence Gives Convergence in Measure

Let \(X=[0,2]\) with Lebesgue measure, and define \(f_n(x)=x/n\) and \(f(x)=0\). For every \(x\in[0,2]\),

$$ |f_n(x)-f(x)|=\frac{x}{n}\leq\frac{2}{n}. $$

Fix \(\varepsilon>0\). If \(n>2/\varepsilon\), then \(2/n<\varepsilon\), so \(D_n(\varepsilon)=\varnothing\). Consequently, \(\lambda(D_n(\varepsilon))=0\) for all sufficiently large \(n\), and \(f_n\) converges in measure to \(f\). In fact, the same bound holds at every point, so the convergence is uniform as well. This example shows how a uniform error estimate can settle convergence in measure immediately.

Almost-Everywhere Convergence on a Finite-Measure Space

The reverse implication requires care. On a finite-measure space, the Egorov’s Theorem from the previous tutorial gives the needed uniform control away from a small exceptional set. On a space of infinite measure, almost-everywhere convergence alone need not force convergence in measure.

Theorem (Almost-Everywhere Convergence Implies Convergence in Measure on a Finite-Measure Space): Suppose \(\mu(X)<\infty\), and \(f_n,f:X\to\mathbb{R}\) are measurable. If \(f_n\to f\) almost everywhere, then \(f_n\to f\) in measure.

Proof. Fix \(\varepsilon>0\), the tolerance in the definition of convergence in measure. Let \(\eta>0\). By Egorov’s Theorem, there is a measurable set \(E\) with \(\mu(E)<\eta\) such that \(f_n\to f\) uniformly on \(X\setminus E\). Uniform convergence gives an index \(N\) such that, for every \(n\geq N\) and every \(x\in X\setminus E\),

$$ |f_n(x)-f(x)|<\varepsilon. $$

Therefore, for \(n\geq N\), every point of \(D_n(\varepsilon)\) must lie in \(E\), so \(D_n(\varepsilon)\subseteq E\). Monotonicity of measure gives

$$ 0\leq\mu(D_n(\varepsilon))\leq\mu(E)<\eta. $$

Since \(\eta>0\) was arbitrary, \(\mu(D_n(\varepsilon))\to0\). This holds for every \(\varepsilon>0\), proving convergence in measure. \(\square\)

The proof uses finite total measure through Egorov’s Theorem: the almost-everywhere convergence can be made uniform outside a set of arbitrarily small measure. Once the convergence is uniform there, the deviation set is contained in the small exceptional set. No claim is made that the deviation sets themselves decrease with \(n\); the argument only bounds each sufficiently late one.

Worked Example: Moving Intervals on the Real Line

On \(\mathbb{R}\) with Lebesgue measure, set \(f_n=\mathbf{1}_{[n,n+1)}\) and \(f=0\). Any fixed \(x\in\mathbb{R}\) belongs to at most one of the intervals \([n,n+1)\). Thus \(f_n(x)=0\) for all sufficiently large \(n\), and \(f_n(x)\to0\) at every \(x\).

For the tolerance \(\varepsilon=1/2\), however, the deviation set is exactly the interval on which the indicator is one:

$$ D_n(1/2)=[n,n+1),\qquad \lambda(D_n(1/2))=1. $$

The measures of these deviation sets do not tend to zero, so \(f_n\) does not converge in measure to zero. This does not contradict the theorem: \(\mathbb{R}\) has infinite measure. It illustrates why almost-everywhere convergence need not imply convergence in measure without a finite-measure hypothesis.

A Subsequence Characterization on Finite Spaces

On a finite-measure space, the two results above give a useful way to recognize convergence in measure by examining subsequences. The characterization distinguishes convergence in measure from almost-everywhere convergence of the original sequence: every subsequence must contain a suitable almost-everywhere convergent further subsequence, but the original sequence itself need not converge almost everywhere.

Theorem (Subsequence Characterization of Convergence in Measure): Suppose \(\mu(X)<\infty\), and let \(f_n,f:X\to\mathbb{R}\) be measurable. Then \(f_n\to f\) in measure if and only if every subsequence of \((f_n)\) has a further subsequence that converges to \(f\) almost everywhere.

Proof. First suppose \(f_n\to f\) in measure. Every subsequence also converges in measure to \(f\): for a fixed \(\varepsilon>0\), the measures \(\mu(D_n(\varepsilon))\) tend to zero, and hence they still tend to zero along any subsequence. The Subsequence Extraction from Convergence in Measure Theorem then supplies a further subsequence converging to \(f\) almost everywhere.

Conversely, suppose every subsequence has a further subsequence converging to \(f\) almost everywhere. Assume, for contradiction, that \(f_n\) does not converge in measure to \(f\). Then there are an \(\varepsilon_0>0\), a \(\delta>0\), and a subsequence \((f_{n_j})\) such that

$$ \mu\bigl(D_{n_j}(\varepsilon_0)\bigr)\geq\delta $$

for every \(j\). By the assumed subsequence property, this subsequence has a further subsequence converging to \(f\) almost everywhere. Since \(\mu(X)<\infty\), the Almost-Everywhere Convergence Implies Convergence in Measure Theorem shows that this further subsequence converges in measure to \(f\). Its deviation-set measures for the fixed tolerance \(\varepsilon_0\) must therefore tend to zero. But each of those measures is at least \(\delta\), a contradiction. Thus \(f_n\to f\) in measure. \(\square\)

Worked Example: Convergence in Measure Without Almost-Everywhere Convergence

The following sequence on \([0,1]\) shows that convergence in measure need not give almost-everywhere convergence of the whole sequence. For each integer \(j\geq0\) and \(r\in\{0,\ldots,2^j-1\}\), define the interval

$$ I_{j,r}=[r/2^j,(r+1)/2^j). $$

List the indicators of these intervals level by level: at level \(j\), include the \(2^j\) functions \(\mathbf{1}_{I_{j,0}},\ldots,\mathbf{1}_{I_{j,2^j-1}}\), in that order. This defines one sequence \((f_n)\). The first index at level \(j\) is \(2^j\), since the preceding levels contain

$$ 1+2+\cdots+2^{j-1}=2^j-1 $$

functions. At every index in level \(j\), the function is the indicator of an interval of length \(2^{-j}\). For any fixed \(0<\varepsilon<1\), its deviation set from zero is that interval, so its measure is \(2^{-j}\). As \(n\to\infty\), the level \(j\) of the function at index \(n\) also tends to infinity. Hence these deviation measures tend to zero, and \(f_n\to0\) in measure.

Nevertheless, for every \(x\in[0,1)\), exactly one interval at each level contains \(x\), so \(f_n(x)=1\) at least once in every level. At every level \(j\geq1\), there are other intervals that do not contain \(x\), so \(f_n(x)=0\) at least once in that level as well. Both values therefore occur infinitely often, and \(f_n(x)\) does not converge. At \(x=1\), all the half-open intervals exclude the point, so the sequence is zero there. Thus convergence fails at every point of \([0,1)\), even though it holds in measure.

The subsequence theorem still applies. For instance, choosing the first interval at each level gives indicators of \([0,2^{-j})\). These converge to zero at every \(x>0\), and fail only at \(x=0\), a null set. The full sequence need not converge almost everywhere for convergence in measure to hold; an almost-everywhere convergent further subsequence is guaranteed instead.

How to Keep the Convergence Modes Distinct

Convergence in measure describes the size of error sets, while pointwise convergence describes the behavior at each individual point. Neither should be substituted for the other without a theorem supplying the implication. In particular, the moving-interval example shows that even everywhere convergence need not imply convergence in measure on an infinite-measure space. The interval-indicator example shows that convergence in measure on a finite space need not imply pointwise convergence almost everywhere for the entire sequence.

On a finite-measure space, almost-everywhere convergence does imply convergence in measure by Egorov’s Theorem, and convergence in measure allows extraction of an almost-everywhere convergent subsequence. On arbitrary measure spaces, the subsequence-extraction implication remains valid, but the reverse implication from almost-everywhere convergence can fail. These distinctions will be useful when convergence is studied alongside the Lebesgue integral.

Key takeaway: Convergence in measure means that every fixed positive error tolerance is exceeded only on sets whose measures tend to zero. It always yields an almost-everywhere convergent subsequence; on finite-measure spaces, almost-everywhere convergence also implies convergence in measure.

Check Your Understanding

Use the deviation-set definition and the proved results to check your understanding.

  1. For a fixed \(\varepsilon>0\), what must happen to \(\mu(D_n(\varepsilon))\) when \(f_n\to f\) in measure?
  2. In the subsequence-extraction proof, why are the bounds \(2^{-k}\) useful?
  3. Where does the finite-measure hypothesis enter the proof that almost-everywhere convergence implies convergence in measure?
  4. Why do the moving indicators \(\mathbf{1}_{[n,n+1)}\) converge pointwise to zero but not in measure on \(\mathbb{R}\)?
  5. What does the subsequence characterization guarantee, and what does it not guarantee, about the original sequence?