From Curvature to a Taylor Approximation
The Hessian describes second-order change at a point, but a Taylor formula organizes that information together with higher derivatives into an approximation of the function near the point. The key reduction is to follow the function along the segment from the point of approximation to the point being evaluated. This turns the multivariable problem into a one-variable Taylor theorem, while retaining the direction of the segment in each derivative.
Let \(U\subseteq\mathbb{R}^m\) be open and \(f:U\to\mathbb{R}\). When the derivative map \(Df\) is differentiable repeatedly, write \(D^k f(x)\) for the \(k\)th derivative at \(x\). It is a \(k\)-linear map: it takes \(k\) vectors in \(\mathbb{R}^m\) and returns a real number. In particular, \(D^1f(x)=Df(x)\), and the second derivative \(D^2f(x)\) is the bilinear map from Second Derivatives. For a vector \(h\), the notation \(D^k f(x)[h,\ldots,h]\) means that \(h\) is inserted in all \(k\) arguments. Set \(D^0f(x)=f(x)\) as a convention.
For \(r=1\), this is the tangent-plane approximation. For \(r=2\), the Hessian Representation of the Second Derivative gives the familiar quadratic expression \(f(a)+\nabla f(a)\cdot h+\frac12 h^{\mathsf T}H_f(a)h\). The higher-order terms follow the same pattern: evaluate the \(k\)th derivative repeatedly in the displacement direction and divide by \(k!\).
Restricting the Function to a Segment
Fix \(a\in U\) and a displacement \(h\) such that the entire segment \(\{a+th:0\leq t\leq1\}\) lies in \(U\). Define the one-variable function \(g(t)=f(a+th)\). Repeated use of the chain rule gives, whenever the indicated derivatives exist, \(g^{(k)}(t)=D^k f(a+th)[h,\ldots,h]\). Each differentiation in \(t\) introduces another argument \(h\). In particular, \(g^{(k)}(0)\) is exactly the order-\(k\) term before division by \(k!\) in the Taylor polynomial.
The next theorem gives an exact remainder formula. Assume \(f\) has continuous derivatives through order \(r+1\) on an open set containing the segment. This ensures that the restriction \(g\) has the regularity needed for the one-variable Taylor theorem, including at the endpoints.
Proof. Let \(g(t)=f(a+th)\) for \(0\leq t\leq1\). Repeated application of the chain rule gives \(g^{(k)}(t)=D^k f(a+th)[h,\ldots,h]\) for \(0\leq k\leq r+1\). We establish the one-variable remainder formula directly. Set \[ C=g(1)-\sum_{j=0}^{r}\frac{g^{(j)}(0)}{j!} \] and define \[ F(t)=g(t)-\sum_{j=0}^{r}\frac{g^{(j)}(0)}{j!}t^j-Ct^{r+1}. \] By the definition of \(C\), \(F(1)=0\). Also \(F^{(j)}(0)=0\) for every \(j=0,\ldots,r\): the polynomial terms reproduce the derivatives of \(g\) at zero through order \(r\), and \(t^{r+1}\) and its first \(r\) derivatives vanish there.
Rolle’s theorem now gives a point where \(F^{(r+1)}\) vanishes. To see the repeated step, \(F(0)=F(1)=0\) first gives a zero of \(F'\) in \((0,1)\). Together with \(F'(0)=0\), Rolle’s theorem gives a zero of \(F''\) between zero and that point. Continuing, at each stage the zero at zero of the corresponding derivative and the zero obtained at the preceding stage give a zero of the next derivative. After \(r+1\) applications there is some \(\theta\in(0,1)\) with \(F^{(r+1)}(\theta)=0\). Differentiating the definition of \(F\) \(r+1\) times yields \[ F^{(r+1)}(t)=g^{(r+1)}(t)-(r+1)!C. \] Thus \(C=g^{(r+1)}(\theta)/(r+1)!\). Substituting the definition of \(C\) gives the one-variable Taylor formula. Finally, replace each \(g^{(k)}(t)\) by \(D^k f(a+th)[h,\ldots,h]\). The resulting identity is the stated multivariable formula. \(\square\)
Worked Example: A Cubic Polynomial at a Shifted Point
Let \(f(x,y)=x^2y+2xy^2\), and approximate near \(a=(1,-1)\). Write \(h=(u,v)\), so the point being evaluated is \((1+u,-1+v)\). Expanding the two terms separately gives \[ (1+u)^2(-1+v)=-1-2u-u^2+v+2uv+u^2v \] and \[ 2(1+u)(-1+v)^2=2+2u-4v-4uv+2v^2+2uv^2. \] Adding these expressions gives \[ f(1+u,-1+v)=1-3v-u^2-2uv+2v^2+u^2v+2uv^2. \]
The terms of degree at most two form \(T_{2,a}f(a+h)=1-3v-u^2-2uv+2v^2\). The remaining terms have degree three, so the Taylor polynomial of degree two is not exact for general \(u,v\). The degree-three Taylor polynomial includes \(u^2v+2uv^2\) and equals the function exactly, as expected for a polynomial of degree three. This direct expansion also verifies the constant, linear, quadratic, and cubic contributions at the shifted point.
Quadratic Approximation and an Explicit Error Bound
The order-two case is especially useful because its coefficients can be read from the gradient and Hessian. If \(f\) has continuous derivatives through order three along the segment, the Lagrange remainder says that the difference between \(f(a+h)\) and its quadratic Taylor polynomial is a third derivative evaluated at an intermediate point. If that derivative is bounded near \(a\), the error is bounded by a constant times \(\|h\|_2^3\). The directional form makes clear that the bound depends on the same displacement vector in all three arguments.
Worked Example: Exponential Approximation with a Remainder Bound
Take \(f(x,y)=e^{x+2y}\) at \(a=(0,0)\), and let \(h=(u,v)\). Put \(s=u+2v\). Along the segment, \(g(t)=f(tu,tv)=e^{ts}\), so \(g^{(k)}(t)=s^k e^{ts}\). At zero, \(g^{(k)}(0)=s^k\), and the degree-two Taylor polynomial is \[ T_{2,a}f(a+h)=1+s+\frac{s^2}{2}. \] The Lagrange remainder gives a \(\theta\in(0,1)\) such that \[ e^s=1+s+\frac{s^2}{2}+\frac{s^3e^{\theta s}}{6}. \] This also follows from the multivariable formula because \(D^3f(tu,tv)[h,h,h]=s^3e^{ts}\).
Since \(0<\theta<1\), \(|e^{\theta s}|\leq e^{|s|}\). Therefore \[ \left|e^{u+2v}-\left(1+(u+2v)+\frac{(u+2v)^2}{2}\right)\right| \leq \frac{e^{|u+2v|}}{6}|u+2v|^3. \] By the Cauchy–Schwarz Inequality, \(|u+2v|\leq\sqrt{5}\,\|(u,v)\|_2\). In particular, when \(\|(u,v)\|_2\leq1\), the error is at most \(e^{\sqrt{5}}5^{3/2}\|(u,v)\|_2^3/6\). The bound verifies that the quadratic approximation error is of cubic order in the size of the displacement.
Worked Example: A Logarithm Near the Origin
Let \(f(x,y)=\ln(1+x-y)\), defined where \(1+x-y>0\), and take \(a=(0,0)\). For \(h=(u,v)\), write \(s=u-v\). If \(|s|<1\), then \(1+ts>0\) for every \(t\in[0,1]\), so the segment from the origin to \((u,v)\) remains in the domain. The restriction is \(g(t)=\ln(1+ts)\), with \[ g'(t)=\frac{s}{1+ts},\qquad g''(t)=-\frac{s^2}{(1+ts)^2},\qquad g'''(t)=\frac{2s^3}{(1+ts)^3}. \] Thus the degree-two Taylor polynomial is \(s-\frac{s^2}{2}\). The order-two Lagrange formula gives, for some \(\theta\in(0,1)\), \[ \ln(1+s)=s-\frac{s^2}{2}+\frac{s^3}{3(1+\theta s)^3}. \] The factor \(3\) in the denominator follows from \(g'''(\theta)/3!=2s^3/(6(1+\theta s)^3)\). If \(|s|\leq\frac12\), then \(1+\theta s\geq\frac12\), so the remainder has absolute value at most \(8|s|^3/3\). This estimate also holds when \(s=0\), in which case the remainder is zero.
Continuity Gives a Little-o Remainder
The Lagrange formula requires one additional derivative and evaluates it at an intermediate point. A different conclusion is available when the highest derivative included in the polynomial is continuous: the remainder is then smaller than order \(r\), in the precise sense that its ratio to \(\|h\|_2^r\) tends to zero. This version is often useful when one wants a local expansion without requiring an \((r+1)\)st derivative.
Proof. Apply the Multivariable Taylor Theorem with Lagrange Remainder to order \(r-1\). For all sufficiently small nonzero \(h\), its segment from \(a\) to \(a+h\) lies in the neighborhood, and there is a \(\theta_h\in(0,1)\) such that \[ f(a+h)=\sum_{k=0}^{r-1}\frac{1}{k!}D^k f(a)[h,\ldots,h] +\frac{1}{r!}D^r f(a+\theta_hh)[h,\ldots,h]. \] Subtract the order-\(r\) Taylor polynomial. The difference is \[ R(h)=\frac{1}{r!}\bigl(D^r f(a+\theta_hh)-D^r f(a)\bigr)[h,\ldots,h]. \] For a \(r\)-linear map \(A\), let \(\|A\|_{\mathrm{op}}\) denote the supremum of \(|A[v_1,\ldots,v_r]|\) over vectors \(v_i\) with \(\|v_i\|_2\leq1\). By multilinearity, \[ |R(h)|\leq\frac{1}{r!}\|D^r f(a+\theta_hh)-D^r f(a)\|_{\mathrm{op}}\|h\|_2^r. \] Since \(\|\theta_hh\|_2\leq\|h\|_2\), the points \(a+\theta_hh\) tend to \(a\) as \(h\to0\). Continuity of \(D^r f\) at \(a\) makes the operator-norm difference tend to zero. Hence \(|R(h)|/\|h\|_2^r\to0\), which is exactly the asserted little-o remainder. \(\square\)
For \(r=2\), this theorem says that a function with continuous second derivative has a quadratic Taylor approximation with error \(o(\|h\|_2^2)\). The conclusion is stronger than merely saying the error is bounded by a constant times \(\|h\|_2^2\): after division by \(\|h\|_2^2\), the error tends to zero. The proof also indicates why continuity matters: it makes the order-\(r\) derivative at the intermediate point approach its value at the center.
Using the Formula with Care
A Taylor polynomial is built from derivatives at the center \(a\), but the Lagrange remainder uses a derivative at an intermediate point on the segment. Those are different roles. The formula with Lagrange remainder therefore requires the segment to stay inside the domain and requires one more continuous derivative than the polynomial degree. The Peano remainder has a different hypothesis: continuity through order \(r\) suffices, but it gives an asymptotic statement rather than an exact intermediate-point formula.
In applications, first choose the degree needed for the desired accuracy, then verify that the segment lies in the domain and that the corresponding derivatives have the required regularity. For a quadratic approximation, the Hessian contributes the term \(\frac12h^{\mathsf T}H_f(a)h\); it is not enough to include the Hessian without controlling the remainder. The worked examples show two ways to do that: compute the remainder explicitly along the segment, or bound the relevant derivative there.
Check Your Understanding
Use the definitions, proofs, and examples in this tutorial to answer the following questions.
- Why does restricting \(f\) to \(g(t)=f(a+th)\) turn each directional derivative into an evaluation of \(D^k f\) on repeated copies of \(h\)?
- In the proof of the Lagrange remainder, why must the auxiliary function include the term \(-Ct^{r+1}\), and how is \(C\) chosen?
- What regularity and segment condition are needed for the Multivariable Taylor Theorem with Lagrange Remainder?
- For \(f(x,y)=e^{x+2y}\) at the origin, what is the degree-two Taylor polynomial in \(u,v\), and what bound holds for its error?
- How does the Peano remainder conclusion differ from an exact Lagrange remainder formula?