From Simple Functions to Continuous Functions
The previous tutorial established that simple functions supported on sets of finite measure are dense in \(L^p\) when \(1\leq p<\infty\). A simple function is still generally discontinuous, so the next question is whether those approximations can be replaced by continuous ones. On Euclidean space with Lebesgue measure, the answer is yes: measurable sets of finite measure can be approximated from inside by compact sets and from outside by open sets. A continuous function can then be made equal to one on the compact set and zero outside the open set.
We work on \(\mathbb{R}^d\) with Lebesgue measure \(m\), where \(d\) is a positive integer. Write \(C_c(\mathbb{R}^d)\) for the continuous real-valued functions with compact support. The compact-support condition is useful on infinite-measure spaces: it ensures that each approximant belongs to \(L^p\) for finite \(p\).
Regularity and Continuous Cutoffs
The measure-theoretic input is the regularity of Lebesgue measure. For a measurable set \(E\) of finite measure, outer regularity provides an open set containing \(E\) with little excess measure. Inner regularity provides a compact subset of \(E\) with little missing measure. We recall the finite-measure version and give the construction needed here.
Proof. By the construction of Lebesgue measure from coverings by boxes, for every \(\rho>0\) there is an open set \(U\supset E\) with \(m(U)<m(E)+\rho\). Since \(E\) is measurable and has finite measure, this gives \(m(U\setminus E)=m(U)-m(E)<\rho\).
For the inner approximation, choose a closed cube \(Q\) large enough that \(m(E\setminus Q)<\rho\). Such cubes increase to \(\mathbb{R}^d\), so continuity of measure from below, applied to their intersections with \(E\), gives this choice. The set \(Q\setminus E\) is measurable and has finite measure. By outer regularity, there is an open set \(V\supset Q\setminus E\) such that \(m(V\setminus(Q\setminus E))<\rho\). Put \(K=Q\setminus V\). This set is closed and bounded, hence compact, and \(K\subset E\). Moreover, $$ E\setminus K\subset (E\setminus Q)\cup\bigl(E\cap Q\cap V\bigr). $$ The first set on the right has measure less than \(\rho\); the second is contained in \(V\setminus(Q\setminus E)\), so it also has measure less than \(\rho\). Thus \(m(E\setminus K)<2\rho\). Choose \(\rho\) small enough that \(3\rho<\eta\), and combine this compact set with an outer open approximation whose excess is less than \(\rho\). Since \(K\subset E\subset U\), $$ m(U\setminus K)=m(U\setminus E)+m(E\setminus K)<3\rho<\eta. $$ This proves the lemma. \(\square\)
Given \(K\subset E\subset U\) as in the lemma, a continuous cutoff separates \(K\) from the outside of \(U\). For a nonempty compact \(K\), define \(\operatorname{dist}(x,K)=\inf_{y\in K}|x-y|\). This distance function is continuous: the triangle inequality gives $$ \bigl|\operatorname{dist}(x,K)-\operatorname{dist}(z,K)\bigr|\leq |x-z|. $$ If \(U^c\) is nonempty, the distance between \(K\) and \(U^c\) is positive. Choose \(r>0\) smaller than that distance and set $$ \phi(x)=\max\left(0,1-\frac{\operatorname{dist}(x,K)}{r}\right). $$ If \(U=\mathbb{R}^d\), take any \(r>0\). The function \(\phi\) is continuous, equals one on \(K\), and vanishes whenever \(\operatorname{dist}(x,K)\geq r\). Its support is contained in the closed \(r\)-neighborhood of \(K\), which is compact and, when \(U^c\ne\varnothing\), lies inside \(U\). If \(K\) is empty, the zero function can be used; in that case \(E\subset U\setminus K\), so the same error bound below applies.
Proof. Choose \(K\subset E\subset U\) with \(m(U\setminus K)<\varepsilon^p\), using the regularity lemma, and form the cutoff just described. On \(K\), both \(\mathbf{1}_E\) and \(\phi\) equal one; outside \(U\), both equal zero. Everywhere else their difference has absolute value at most one. Consequently, $$ \int_{\mathbb{R}^d}|\mathbf{1}_E-\phi|^p\,dm \leq m(U\setminus K) <\varepsilon^p. $$ Taking \(p\)th roots proves the claim. \(\square\)
Worked Examples of the Cutoff Method
Worked Example: Approximating the Indicator of an Interval
On \(\mathbb{R}\), let \(E=[0,1]\) and choose \(\delta>0\). Define \(\phi_\delta\) to be one on \([0,1]\), zero outside \([-\delta,1+\delta]\), and linear on the two remaining intervals: $$ \phi_\delta(x)= \begin{cases} 1+x/\delta,&-\delta\leq x<0,\\ 1-(x-1)/\delta,&1<x\leq 1+\delta. \end{cases} $$ Together with the values one and zero just specified, this defines a continuous compactly supported function. It agrees with \(\mathbf{1}_{[0,1]}\) on \([0,1]\) and outside \([-\delta,1+\delta]\). The error on each ramp is a linear function taking values between zero and one. Therefore, for \(1\leq p<\infty\), $$ \|\mathbf{1}_{[0,1]}-\phi_\delta\|_p^p =2\int_0^\delta (t/\delta)^p\,dt =\frac{2\delta}{p+1}. $$ The integral tends to zero with \(\delta\), so the indicator is approximated in every finite-\(p\) norm. The transition regions become narrow; the error is not required to be small at every point uniformly.
Worked Example: Approximating a Finite Simple Function
Let \(s=2\mathbf{1}_{[0,1]}-\mathbf{1}_{[2,4]}\) on \(\mathbb{R}\), and fix \(1\leq p<\infty\) and \(\varepsilon>0\). The indicator lemma gives \(\phi_1,\phi_2\in C_c(\mathbb{R})\) with $$ \|\mathbf{1}_{[0,1]}-\phi_1\|_p<\frac{\varepsilon}{6}, \qquad \|\mathbf{1}_{[2,4]}-\phi_2\|_p<\frac{\varepsilon}{6}. $$ Set \(g=2\phi_1-\phi_2\). A finite linear combination of continuous compactly supported functions is again continuous and compactly supported. By the triangle inequality in \(L^p\) and homogeneity of the norm, $$ \|s-g\|_p \leq 2\|\mathbf{1}_{[0,1]}-\phi_1\|_p +\|\mathbf{1}_{[2,4]}-\phi_2\|_p <2\frac{\varepsilon}{6}+\frac{\varepsilon}{6} =\frac{\varepsilon}{2} <\varepsilon. $$ Thus the indicator approximation transfers directly to this simple function. The same argument applies to any finite sum once each of its finite-measure level sets has been approximated.
Density of Compactly Supported Continuous Functions
Proof. By the Simple Functions Are Dense in \(L^p\) theorem from the previous tutorial, choose a simple function \(s\in L^p(\mathbb{R}^d)\), supported on a set of finite measure, such that $$ \|f-s\|_p<\frac{\varepsilon}{2}. $$ Write \(s=\sum_{j=1}^N c_j\mathbf{1}_{A_j}\), where the \(A_j\) are disjoint measurable sets of finite measure and the \(c_j\) are nonzero real numbers. Terms with zero coefficient can be omitted. For each \(j\), the indicator lemma supplies \(\phi_j\in C_c(\mathbb{R}^d)\) with $$ \|\mathbf{1}_{A_j}-\phi_j\|_p <\frac{\varepsilon}{2N(|c_j|+1)}. $$ Define \(g=\sum_{j=1}^N c_j\phi_j\). Since the sum is finite, \(g\in C_c(\mathbb{R}^d)\). Applying the triangle inequality and homogeneity, $$ \|s-g\|_p \leq\sum_{j=1}^N |c_j|\|\mathbf{1}_{A_j}-\phi_j\|_p <\sum_{j=1}^N\frac{|c_j|\varepsilon}{2N(|c_j|+1)} \leq\frac{\varepsilon}{2}. $$ The last inequality follows because \(|c_j|/(|c_j|+1)<1\) for every \(j\). A further application of the triangle inequality gives $$ \|f-g\|_p\leq\|f-s\|_p+\|s-g\|_p<\frac{\varepsilon}{2}+\frac{\varepsilon}{2}=\varepsilon. $$ This proves the density theorem. \(\square\)
Worked Example: Approximating a Singularity in \(L^p\)
Let \(0<\alpha<1/p\), and define \(f(x)=x^{-\alpha}\) on \((0,1)\), with \(f(x)=0\) elsewhere. The function belongs to \(L^p(\mathbb{R})\), since $$ \int_{\mathbb{R}}|f|^p\,dx =\int_0^1 x^{-\alpha p}\,dx =\frac{1}{1-\alpha p}<\infty. $$ For \(0<\delta<1\) and \(\eta>0\), define \(g_{\delta,\eta}\) to be zero for \(x\leq0\), linear from zero at \(0\) to \(\delta^{-\alpha}\) at \(\delta\), equal to \(x^{-\alpha}\) on \([\delta,1]\), linear from one at \(1\) to zero at \(1+\eta\), and zero for \(x\geq1+\eta\). The endpoint values agree, so \(g_{\delta,\eta}\in C_c(\mathbb{R})\).
On \((0,\delta)\), the linear segment has value \(g_{\delta,\eta}(x)=\delta^{-\alpha-1}x\), which is at most \(x^{-\alpha}\): this inequality is equivalent to \(x^{\alpha+1}\leq\delta^{\alpha+1}\). Thus the error there is at most \(x^{-\alpha}\). On \((1,1+\eta)\), the function \(g_{\delta,\eta}\) is a linear ramp, and its \(p\)th power integrates to \(\eta/(p+1)\). Elsewhere the functions agree almost everywhere. It follows that $$ \|f-g_{\delta,\eta}\|_p^p \leq\int_0^\delta x^{-\alpha p}\,dx+\frac{\eta}{p+1} =\frac{\delta^{1-\alpha p}}{1-\alpha p}+\frac{\eta}{p+1}. $$ Because \(1-\alpha p>0\), both terms can be made arbitrarily small by choosing \(\delta\) and \(\eta\) sufficiently small. This gives an explicit continuous approximation despite the singularity at zero.
Scope and a Common Pitfall
The theorem concerns Lebesgue measure on Euclidean space and finite exponents \(p\). Its proof has two distinct stages: simple functions approximate the original \(L^p\) function, and regularity plus continuous cutoffs approximate the finitely many indicators in the simple function. The intermediate compact and open sets are what allow a continuous transition while keeping the \(L^p\) error small.
The restriction \(p<\infty\) matters. A small error on a set of small measure has small finite-\(p\) norm, as the indicator estimate shows. In the essential-supremum norm, the size of the transition region does not reduce the maximum essential error. For instance, \(\mathbf{1}_{[0,1]}\) cannot be approximated arbitrarily closely in \(L^\infty(\mathbb{R})\) by continuous functions: continuity forces any transition from values near one inside the interval to values near zero outside it to pass through values near \(1/2\), on a region of positive measure. The finite-\(p\) density theorem should not be transferred to \(L^\infty\) without checking the norm.
Another common mistake is to infer density on an arbitrary measure space. The proof uses regularity of Lebesgue measure and the topology of \(\mathbb{R}^d\); a general measure space need not have a meaningful class of continuous functions with this approximation property. The stated setting is part of the theorem, not a technical decoration.
Check Your Understanding
Keep track of which part of the proof controls the measure of the error region and which part controls the size of the error there.
- Why does \(m(U\setminus K)<\varepsilon^p\) imply the indicator approximation has \(L^p\) error less than \(\varepsilon\)?
- How does the distance-to-\(K\) function help produce a continuous cutoff that is one on \(K\)?
- In the density theorem proof, why must the finite simple function be approximated before estimating \(\|f-g\|_p\)?
- For the interval cutoff \(\phi_\delta\), what is the exact value of \(\|\mathbf{1}_{[0,1]}-\phi_\delta\|_p^p\)?
- Why does making a transition region small help in finite \(L^p\), but not necessarily in \(L^\infty\)?