Approximation in a Stronger Norm
The previous tutorial used a fine mesh to approximate a continuous function uniformly. A function-space approximation problem can ask for more: an approximant may need to be close not only in value, but also in its derivatives. This matters when the derivative is part of the quantity being studied, as it is in differential equations, optimization, and the analysis of smooth curves.
For a continuously differentiable function, closeness of the functions alone does not control closeness of their derivatives. A sequence can converge uniformly while its derivatives fail to converge uniformly, or even fail to converge pointwise. We therefore use a norm that measures both errors. The useful construction is to approximate the derivative first and then integrate, choosing the constant of integration to match the original function at an endpoint.
The sum in this definition means that \(\|f-g\|_{C^m}\) is small only when every derivative through order \(m\) is uniformly close. The \(C^m\) norm is therefore stronger than the supremum norm: for every \(f\in C^m[a,b]\), \(\|f\|_\infty\leq\|f\|_{C^m}\). A polynomial approximation in the \(C^m\) norm is also a uniform approximation, but the converse implication need not hold.
Why Integrating Controls Several Errors
Suppose a continuous function \(r\) is the error in approximating a derivative. If an error function \(e\) satisfies \(e(a)=0\) and \(e'=r\), the Fundamental Theorem of Calculus gives $$ e(x)=\int_a^x r(t)\,dt. $$ Consequently, if \(|r(t)|\leq\delta\) throughout the interval, then \(|e(x)|\leq\delta(x-a)\). Repeated integration gives corresponding bounds for errors in lower derivatives. The following estimate records that fact.
Proof. The hypotheses give the formula $$ e^{(j)}(x)=\int_a^x\frac{(x-t)^{s-j-1}}{(s-j-1)!}r(t)\,dt,\qquad 0\leq j<s. $$ For \(j=s-1\), this is the Fundamental Theorem of Calculus applied to \(e^{(s-1)}\), whose value at \(a\) is zero. For each smaller \(j\), integrating the formula for \(e^{(j+1)}\) from \(a\) to \(x\) gives the stated formula for \(e^{(j)}\): integrating the kernel \((u-t)^{s-j-2}/(s-j-2)!\) with respect to \(u\) from \(t\) to \(x\) gives \((x-t)^{s-j-1}/(s-j-1)!\). This proves the formula by downward induction. Taking absolute values and using \(|r(t)|\leq\delta\),
This proves the estimate for every indicated \(j\) and every \(x\in[a,b]\). \(\square\)
Polynomial Approximation in the \(C^m\) Norm
We now apply the Weierstrass Approximation Theorem to the highest derivative of a function. Once that derivative has been approximated by a polynomial, integrating the polynomial \(m\) times gives a polynomial approximant to the original function. The integration constants can be selected to match the original derivatives at \(a\).
Proof. Put \(L=b-a\), and define the positive constant $$ C_m=1+\sum_{r=1}^{m}\frac{L^r}{r!}. $$ The function \(f^{(m)}\) is continuous on \([a,b]\). By the Weierstrass Approximation Theorem, choose a polynomial \(q\) such that $$ \|f^{(m)}-q\|_\infty<\delta, $$ where \(\delta>0\) will be chosen below. Let \(p\) be the polynomial obtained by integrating \(q\) \(m\) times and choosing the constants so that \(p^{(j)}(a)=f^{(j)}(a)\) for \(j=0,\ldots,m-1\). More explicitly, one may take $$ p(x)=\sum_{j=0}^{m-1}f^{(j)}(a)\frac{(x-a)^j}{j!} +\int_a^x\frac{(x-t)^{m-1}}{(m-1)!}q(t)\,dt. $$ The integral is a polynomial in \(x\), since \(q\) is a polynomial. Differentiating it \(m\) times gives \(q\), and differentiating fewer times and evaluating at \(a\) gives zero. Thus \(p^{(m)}=q\) and the required endpoint equalities hold.
Set \(e=f-p\). Then \(e^{(m)}=f^{(m)}-q\), so \(\|e^{(m)}\|_\infty<\delta\), and \(e^{(j)}(a)=0\) for \(0\leq j<m\). Apply the Repeated-Integration Error Estimate with \(s=m\). For \(j<m\), it gives $$ \|e^{(j)}\|_\infty\leq\delta\frac{L^{m-j}}{(m-j)!}, $$ while \(\|e^{(m)}\|_\infty<\delta\). Summing these bounds yields $$ \|f-p\|_{C^m} \leq\delta\left(1+\sum_{r=1}^{m}\frac{L^r}{r!}\right) =\delta C_m. $$ Choose \(0<\delta<\varepsilon/C_m\). Then \(\|f-p\|_{C^m}<\varepsilon\), as required. \(\square\)
For \(m=1\), the construction has a particularly transparent form: approximate \(f'\) by a polynomial \(q\), then set \(p(x)=f(a)+\int_a^xq(t)\,dt\). The derivative error is the approximation error for \(q\), and the function error is its integral. The theorem shows that this procedure extends to every fixed finite number of derivatives.
Worked Examples: Using the Stronger Norm
Worked Example: An Explicit \(C^1\) Approximation to the Exponential
Take \(f(x)=e^x\) on \([0,1]\), and for an integer \(N\geq1\) let $$ p_N(x)=\sum_{k=0}^{N}\frac{x^k}{k!}. $$ This polynomial has derivative \(p_N'(x)=\sum_{k=0}^{N-1}x^k/k!\). For \(0\leq x\leq1\), the tail of the exponential series satisfies $$ |e^x-p_N(x)| \leq\sum_{k=N+1}^{\infty}\frac{1}{k!} \leq\frac{e}{(N+1)!}. $$ For the last inequality, write \(k=N+1+j\) and use \((N+1+j)!\geq (N+1)!j!\); summing \(1/j!\) gives \(e\). Applying the same estimate to the derivative tail, which starts at index \(N\), gives $$ |e^x-p_N'(x)|\leq\frac{e}{N!}. $$ Therefore $$ \|e^x-p_N\|_{C^1}\leq\frac{e}{(N+1)!}+\frac{e}{N!}. $$ For \(N=5\), this bound is \(7e/720\). The two separate estimates matter: a small error in the values alone would not establish a small \(C^1\) error.
Worked Example: Approximating a Function with a Nonsmooth Derivative
Let \(f(x)=x|x|/2\) on \([-1,1]\). For \(x>0\), \(f(x)=x^2/2\), and for \(x<0\), \(f(x)=-x^2/2\). The derivative on either side of zero is \(|x|\), and the derivative at zero is also zero; hence \(f\in C^1[-1,1]\) and \(f'(x)=|x|\).
To get an explicit polynomial approximation to the derivative, put \(g(t)=|2t-1|\) on \([0,1]\) and let \(B_ng\) be its Bernstein polynomial. The function \(g\) is Lipschitz with constant \(2\). Using the Bernstein weights \(w_{n,k}(t)\), their sum \(1\), and their second-moment identity, we obtain $$ \begin{aligned} |B_ng(t)-g(t)| &\leq 2\sum_{k=0}^n w_{n,k}(t)\left|\frac{k}{n}-t\right|\\ &\leq 2\left(\sum_{k=0}^n w_{n,k}(t)\left(\frac{k}{n}-t\right)^2\right)^{1/2}\\ &=2\sqrt{\frac{t(1-t)}{n}} \leq\frac{1}{\sqrt n}. \end{aligned} $$ The second inequality follows from the Cauchy–Schwarz inequality for the nonnegative weights, and the last uses \(t(1-t)\leq1/4\).
Define the polynomial \(q_n(x)=B_ng((x+1)/2)\), and then set \(p_n(x)=\int_0^xq_n(t)\,dt\). Since \(g((x+1)/2)=|x|\), the estimate above gives \(\|q_n-|x|\|_{\infty,[-1,1]}\leq1/\sqrt n\). Also \(f(x)=\int_0^x|t|\,dt\). It follows that $$ \|p_n'-f'\|_\infty\leq\frac{1}{\sqrt n}, \qquad |p_n(x)-f(x)| \leq |x|\frac{1}{\sqrt n} \leq\frac{1}{\sqrt n}. $$ Thus \(\|p_n-f\|_{C^1}\leq2/\sqrt n\). This example shows how polynomial approximation in \(C^1\) can apply even when the derivative is continuous but not differentiable at every point.
Worked Example: Matching Endpoint Data While Controlling Both Errors
Define \(F(x)=\int_0^x\cos(t^2)\,dt\) on \([0,1]\). Then \(F(0)=0\) and \(F'(x)=\cos(x^2)\). For \(N\geq0\), set $$ q_N(x)=\sum_{k=0}^{N}\frac{(-1)^kx^{4k}}{(2k)!}, \qquad p_N(x)=\int_0^xq_N(t)\,dt =\sum_{k=0}^{N}\frac{(-1)^kx^{4k+1}}{(2k)!(4k+1)}. $$ The cosine series has decreasing term magnitudes on this interval. Its alternating-series remainder therefore gives $$ \|F'-p_N'\|_\infty =\|\cos(x^2)-q_N(x)\|_\infty \leq\frac{1}{(2N+2)!}. $$ Because \(F(0)=p_N(0)=0\), integrating this derivative error also gives $$ \|F-p_N\|_\infty\leq\frac{1}{(2N+2)!}. $$ In particular, for \(N=2\), $$ p_2(x)=x-\frac{x^5}{10}+\frac{x^9}{216}, \qquad \|F-p_2\|_{C^1}\leq\frac{2}{720}=\frac{1}{360}. $$ The endpoint value is matched exactly, while the function and derivative errors are both controlled throughout the interval.
What the Stronger Norm Guarantees
Polynomial density in \(C^m[a,b]\) is stronger than polynomial density in \(C[a,b]\): it supplies simultaneous uniform control of the function and each derivative through order \(m\). The construction also preserves the first \(m\) endpoint data used to set the integration constants. This is useful when approximants must respect initial values, slopes, or higher-order derivative conditions.
There is an important distinction between approximating in these spaces. A continuous target belongs to \(C[a,b]\), but it need not have a continuous derivative, so a \(C^1\) approximation problem is not even defined for every target in \(C[a,b]\). When the target does belong to \(C^1[a,b]\), polynomial approximation in the supremum norm alone still does not guarantee derivative approximation. The method here addresses that gap by making the derivative error the starting point of the construction.
Check Your Understanding
Use the \(C^m\) norm and the integration construction to answer the following questions.
- Why does a small \(C^1\) error imply a small uniform error in the functions?
- In the polynomial-density proof, why is the \(m\)-fold integral of a polynomial still a polynomial?
- What endpoint equalities are imposed on the approximating polynomial in the \(C^m\) theorem?
- For \(f(x)=x|x|/2\), why does approximating \(f'\) uniformly also give a bound for \(\|f-p\|_\infty\) when \(p(0)=f(0)\)?
- Why does uniform approximation of a function alone not establish approximation in the \(C^1\) norm?