Measuring the Size of the Error
Almost-everywhere convergence asks whether the values \(f_n(x)\) approach \(f(x)\) at nearly every individual point. Convergence in \(L^p\), by contrast, controls an integral of the error. Convergence in measure takes a different approach: for each fixed error threshold, it measures the set of points where the error exceeds that threshold. This perspective is useful even when the functions are not in any \(L^p\) space and the whole space has infinite measure.
Let \((X,\mathcal{F},\mu)\) be a measure space, and suppose \(f_n\) and \(f\) are measurable, finite-valued real functions. The exceptional set at threshold \(\varepsilon\) is \(\{x:|f_n(x)-f(x)|>\varepsilon\}\). Convergence in measure says that, for every positive threshold, the measure of this set tends to zero. The measure of \(X\) itself need not be finite.
The threshold is fixed before \(n\) tends to infinity. The definition does not say that the error is small at every point, nor does it require the exceptional sets for different indices to be nested. Their measures simply have to become small at each fixed positive threshold. On a probability space, this mode of convergence is also called convergence in probability.
First Tests: Indicators and Shrinking Sets
Indicator functions make the definition especially concrete. For measurable sets \(A_n,A\), the values of \(\mathbf{1}_{A_n}-\mathbf{1}_A\) have absolute value either zero or one. Consequently, for \(0<\varepsilon<1\), the set where this difference exceeds \(\varepsilon\) is precisely the symmetric difference \(A_n\mathbin{\triangle}A\). For \(\varepsilon\geq1\), that set is empty. Thus convergence of indicators in measure is exactly convergence of the measures of the symmetric differences to zero.
Worked Example: Indicators of Shrinking Intervals
On \(\mathbb{R}\) with Lebesgue measure, let \(A_n=(2n,2n+1/n)\), and set \(f_n=\mathbf{1}_{A_n}\) and \(f=0\). If \(0<\varepsilon<1\), then $$ \{x:|f_n(x)-f(x)|>\varepsilon\}=A_n, \qquad \mu(A_n)=\frac{1}{n}. $$ If \(\varepsilon\geq1\), the exceptional set is empty. In either case its measure tends to zero, so \(f_n\to0\) in measure. In fact, for any fixed \(x\), \(x\) belongs to only finitely many of these intervals, so \(f_n(x)\to0\) pointwise as well.
The intervals in this example move across an infinite-measure space, but the definition still applies: it is the measure of each error set that matters, not whether \(\mu(X)\) is finite. On a finite-measure space, the same calculation often gives a simple route from a shrinking support to convergence in measure.
Convergence in Measure Does Not Mean Pointwise Convergence
Convergence in measure allows exceptional sets to change with \(n\). Even if each such set has small measure, the sets can revisit the same points infinitely many times. The following construction shows that convergence in measure need not imply almost-everywhere convergence of the full sequence.
Worked Example: Dyadic Intervals and Repeated Values
Work on \([0,1]\) with Lebesgue measure. At level \(k\), take the \(2^k\) half-open dyadic intervals $$ \left[\frac{j}{2^k},\frac{j+1}{2^k}\right), \qquad j=0,\ldots,2^k-1. $$ Enumerate all the intervals level by level, and let each function in the level-\(k\) block be the indicator of one of that level's intervals. The measure of the set where such an indicator exceeds \(\varepsilon\) is \(2^{-k}\) if \(0<\varepsilon<1\), and is zero if \(\varepsilon\geq1\). As the sequence proceeds, its level \(k\) tends to infinity, so these measures tend to zero. The sequence therefore converges in measure to zero.
However, every \(x\in[0,1)\) belongs to exactly one interval at every level. Its function values are therefore one once in each level block, and zero on the other intervals in that block. Since there are \(2^k\) intervals at level \(k\), these values do not converge as the sequence runs through the blocks. At \(x=1\), all the half-open intervals have value zero. Thus the sequence fails to converge pointwise at every \(x\in[0,1)\), a set of measure one.
There is a useful subsequence result in the other direction. If \(f_n\to f\) in measure, recursively choose \(n_k>n_{k-1}\) so that $$ \mu\bigl(\{x:|f_{n_k}(x)-f(x)|>2^{-k}\}\bigr)\leq2^{-k}. $$ The measures are summable, so countable subadditivity implies that almost every point belongs to only finitely many of these sets; hence \(f_{n_k}(x)\to f(x)\) almost everywhere. The dyadic example explains why the conclusion is about a subsequence and cannot generally be upgraded to the full sequence.
Almost-Everywhere Convergence Can Also Fail to Give Convergence in Measure
On a finite-measure space, almost-everywhere convergence does imply convergence in measure. Indeed, for fixed \(\varepsilon>0\), the indicators of the error sets tend to zero almost everywhere and are bounded by the integrable function \(1\); the Dominated Convergence Theorem then shows that the measures of those sets tend to zero. Finiteness of the space is essential to this argument. On an infinite-measure space, the constant bound \(1\) need not be integrable.
Worked Example: Pointwise Convergence with Infinite Error-Set Measures
On \(\mathbb{R}\) with Lebesgue measure, define \(f_n=\mathbf{1}_{[n,\infty)}\) and \(f=0\). For each fixed \(x\in\mathbb{R}\), every integer \(n>x\) satisfies \(x\notin[n,\infty)\), so \(f_n(x)=0\) from then on. Hence \(f_n\to0\) everywhere.
But for \(0<\varepsilon<1\), $$ \{x:|f_n(x)-f(x)|>\varepsilon\}=[n,\infty), \qquad \mu([n,\infty))=\infty. $$ These measures do not tend to zero. Thus pointwise, and therefore almost-everywhere, convergence does not imply convergence in measure on an arbitrary measure space.
This example also prevents a common mistake: small pointwise errors at each fixed point do not by themselves give small measures for the sets on which the errors are large. On infinite spaces those sets may have infinite measure, even when every individual point eventually leaves them.
Limits in Measure Are Unique Almost Everywhere
A sequence cannot have two genuinely different limits in measure. As with limits in other modes of convergence, uniqueness is the first structural property to establish. Since measurable functions are treated up to almost-everywhere equality in \(L^p\) spaces, uniqueness up to a null set is exactly the appropriate conclusion.
Proof. Fix \(\delta>0\). If \(|f(x)-g(x)|>\delta\), the triangle inequality implies that at least one of \(|f_n(x)-f(x)|>\delta/2\) or \(|f_n(x)-g(x)|>\delta/2\) must hold. Therefore $$ \{x:|f-g|>\delta\} \subseteq \{x:|f_n-f|>\delta/2\}\cup \{x:|f_n-g|>\delta/2\}. $$ By subadditivity, $$ \mu(\{|f-g|>\delta\}) \leq \mu(\{|f_n-f|>\delta/2\}) + \mu(\{|f_n-g|>\delta/2\}). $$ Both terms on the right tend to zero, so the set on the left has measure zero. In particular, for every positive integer \(m\), the set \(\{|f-g|>1/m\}\) is null. If \(f(x)\ne g(x)\), then \(|f(x)-g(x)|>1/m\) for some positive integer \(m\). Hence $$ \{x:f(x)\ne g(x)\} = \bigcup_{m=1}^{\infty}\{x:|f-g|>1/m\}. $$ This is a countable union of null sets, so it is null. Thus \(f=g\) almost everywhere. \(\square\)
Lipschitz Transformations Preserve Convergence in Measure
Some operations preserve convergence in measure because they cannot magnify differences without bound. A globally Lipschitz function has precisely this control. This result includes scalar multiplication and, by taking the transformation \(\phi(t)=t\), the identity operation.
Proof. If \(L=0\), then \(\phi\) is constant, so \(\phi(f_n)=\phi(f)\) everywhere and convergence is immediate. Suppose \(L>0\). For every \(\varepsilon>0\), the Lipschitz inequality gives the set inclusion $$ \{x:|\phi(f_n(x))-\phi(f(x))|>\varepsilon\} \subseteq \{x:|f_n(x)-f(x)|>\varepsilon/L\}. $$ Taking measures and using convergence in measure of \(f_n\) to \(f\) shows that the measure of the set on the left tends to zero. This is exactly convergence in measure of \(\phi(f_n)\) to \(\phi(f)\). \(\square\)
For example, \(\phi(t)=3t-2\) is Lipschitz with \(L=3\), so \(3f_n-2\) converges in measure to \(3f-2\) whenever \(f_n\to f\) in measure. The global hypothesis matters: a merely continuous function need not be uniformly continuous on all of \(\mathbb{R}\), and errors can occur where \(f\) takes arbitrarily large values. Additional control may be needed to handle such transformations on an infinite-measure space.
A Cauchy Criterion for Convergence in Measure
Convergence in measure also supports a completeness principle. It is useful to formulate the Cauchy condition directly through differences: for each positive error threshold, the measures of the sets where two sufficiently late terms differ by more than that threshold must be small. Unlike completeness of \(L^p\), this argument does not require norms or integrability.
Proof. By the Cauchy condition, choose a strictly increasing sequence of indices \((n_k)\) such that $$ \mu(E_k)\leq 2^{-k}, \qquad E_k=\{x:|f_{n_{k+1}}(x)-f_{n_k}(x)|>2^{-k}\}. $$ For each \(k\), let \(M_k\) be a Cauchy cutoff for threshold \(2^{-k}\) and measure bound \(2^{-k}\). Choose the indices recursively so that \(n_k>n_{k-1}\) and \(n_k\geq\max\{M_1,\ldots,M_k\}\); then both \(n_k\) and \(n_{k+1}\) are at least \(M_k\), giving \(\mu(E_k)\leq2^{-k}\). Define $$ E=\bigcap_{m=1}^{\infty}\bigcup_{k=m}^{\infty}E_k. $$ For every \(m\), subadditivity gives $$ \mu\left(\bigcup_{k=m}^{\infty}E_k\right) \leq\sum_{k=m}^{\infty}2^{-k}=2^{1-m}. $$ Since \(E\) is contained in each of these tail unions, \(\mu(E)=0\).
For \(x\notin E\), there is \(K\) such that \(x\notin E_k\) for all \(k\geq K\). Thus \(|f_{n_{k+1}}(x)-f_{n_k}(x)|\leq2^{-k}\) for \(k\geq K\). If \(r>s\geq K\), telescoping yields $$ |f_{n_r}(x)-f_{n_s}(x)| \leq\sum_{k=s}^{r-1}2^{-k} \leq 2^{1-s}. $$ The subsequence is therefore Cauchy in \(\mathbb{R}\) at every \(x\notin E\), and has a finite limit there. Define \(f\) to be this limit outside \(E\), and set \(f=0\) on \(E\). This function is measurable: the pointwise limit on the measurable set \(X\setminus E\), extended by zero on the measurable set \(E\), is measurable.
For any \(\varepsilon>0\), choose \(k\) large enough that \(2^{1-k}\leq\varepsilon\). If \(x\notin E\) and \(x\notin\bigcup_{j=k}^{\infty}E_j\), then the same telescoping bound, followed by the limit \(r\to\infty\), gives \(|f_{n_k}(x)-f(x)|\leq2^{1-k}\leq\varepsilon\). Consequently, $$ \mu(\{|f_{n_k}-f|>\varepsilon\}) \leq\mu\left(\bigcup_{j=k}^{\infty}E_j\right) \leq2^{1-k}. $$ In particular, the selected subsequence converges to \(f\) in measure.
It remains to show that the entire sequence converges to \(f\) in measure. Fix \(\varepsilon>0\) and \(\eta>0\). The Cauchy condition supplies \(N\) such that, whenever \(n,m\geq N\), $$ \mu(\{|f_n-f_m|>\varepsilon/2\})<\eta. $$ Choose \(k\) so large that \(n_k\geq N\) and \(\mu(\{|f_{n_k}-f|>\varepsilon/2\})<\eta\), which is possible by the subsequence estimate. For every \(n\geq N\), the triangle inequality gives $$ \{|f_n-f|>\varepsilon\} \subseteq \{|f_n-f_{n_k}|>\varepsilon/2\} \cup \{|f_{n_k}-f|>\varepsilon/2\}. $$ Its measure is therefore less than \(2\eta\). Since \(\eta\) was arbitrary, \(\mu(\{|f_n-f|>\varepsilon\})\to0\). This holds for every \(\varepsilon>0\), proving convergence in measure. \(\square\)
The definition of convergence in measure is simple, but its consequences depend on the underlying space and on the quantity being controlled. On finite-measure spaces, almost-everywhere convergence implies convergence in measure; on arbitrary spaces it need not. Convergence in measure has unique limits up to almost-everywhere equality, is preserved by Lipschitz transformations, and has its own Cauchy completeness principle. These facts make it a distinct and useful mode of convergence, rather than a substitute name for pointwise or \(L^p\) convergence.
Check Your Understanding
Use the level-set definition and the proved results to answer the following questions.
- Why does the definition of convergence in measure not require \(\mu(X)\) to be finite?
- For indicator functions, which set's measure determines whether \(\mathbf{1}_{A_n}\) converges in measure to \(\mathbf{1}_A\)?
- How does the dyadic-interval example show that convergence in measure need not imply almost-everywhere convergence of the full sequence?
- Where does the triangle inequality enter the proof that limits in measure are unique almost everywhere?
- Why does the Lipschitz stability theorem require a single finite constant that works for all real inputs?
- In the completeness proof, what role do the summable measures of the sets \(E_k\) play?