Tutorials › Real Analysis › Second Derivatives

Multivariable Analysis · Tutorial 793 of 1000

Second Derivatives

Learn how second derivatives encode local curvature, when they form a symmetric Hessian, and why the hypotheses behind equality of mixed partials matter.

Advanced 11 min read

What You'll Learn

  • Define second partial derivatives and interpret their order of differentiation
  • Build the Hessian matrix and relate it to the second derivative as a bilinear map
  • Compute second directional derivatives using the Hessian
  • Prove equality of mixed partials under continuity hypotheses using the Mean Value Theorem
  • Recognize why existence of mixed partials alone does not guarantee their equality

Beyond the Tangent Plane

The tangent plane from the previous tutorial records the first-order change of a differentiable function. It does not, by itself, describe how the graph bends away from that plane. Second derivatives measure changes in the first derivatives and provide a way to describe this curvature. They will also supply the quadratic term in the multivariable Taylor approximation.

Let \(U\subseteq\mathbb{R}^m\) be open and let \(f:U\to\mathbb{R}\). If the first partial derivative \(\partial_j f\) is defined near \(a\), its partial derivative in coordinate \(i\), when it exists, is written \(\partial_i(\partial_j f)(a)\), or \(\partial_i\partial_j f(a)\). The rightmost index indicates the first differentiation: \(\partial_i\partial_j f\) means differentiate first in coordinate \(j\), then in coordinate \(i\). When \(i=j\), this is the pure second partial derivative \(\partial_i^2 f(a)\); when \(i\neq j\), it is a mixed partial derivative.

Definition (Second Partial Derivative): If \(\partial_j f\) is defined on a neighborhood of \(a\) and is differentiable in coordinate \(i\) at \(a\), then $$ \partial_i\partial_j f(a) = \lim_{t\to0}\frac{\partial_j f(a+te_i)-\partial_j f(a)}{t}, $$ where \(e_i\) is the \(i\)th standard-coordinate vector. The derivative \(\partial_i\partial_j f\) differentiates \(\partial_j f\) in coordinate \(i\).

For a function of two variables, \(f_{xy}\) commonly denotes \(\partial_x(\partial_y f)\), while \(f_{yx}\) denotes \(\partial_y(\partial_x f)\). The order matters in the notation, even though an important theorem will give conditions under which the two values agree. When all second partial derivatives exist, they can be arranged in a matrix.

Definition (Hessian Matrix): If all second partial derivatives of \(f\) exist at \(a\), the Hessian matrix at \(a\) is $$ H_f(a)=\bigl(\partial_i\partial_j f(a)\bigr)_{1\leq i,j\leq m}. $$ Its entry in row \(i\), column \(j\) is \(\partial_i\partial_j f(a)\).

A matrix of second partials is useful, but it is important to distinguish the existence of these coordinate derivatives from the existence of a genuine second derivative as a linear approximation to the first derivative. For that stronger notion, regard \(Df(x)\) as a linear map from \(\mathbb{R}^m\) to \(\mathbb{R}\), as in Derivative as a Linear Map.

Definition (Second Derivative): Suppose \(f\) is differentiable on a neighborhood of \(a\), so that \(x\mapsto Df(x)\) is defined there. If this map is differentiable at \(a\), its derivative is denoted \(D^2f(a)\). It is a linear map from \(\mathbb{R}^m\) into the space of linear maps from \(\mathbb{R}^m\) to \(\mathbb{R}\). Equivalently, it defines the bilinear map $$ (u,v)\longmapsto D^2f(a)[u,v] =\bigl(D(Df)(a)[u]\bigr)(v). $$

When this second derivative exists, its coordinates are precisely the second partial derivatives, and the bilinear map is represented by the Hessian. This gives the Hessian a coordinate-independent role: it evaluates the second-order change in any pair of directions.

Theorem (Hessian Representation of the Second Derivative): Suppose \(f\) is differentiable on a neighborhood of \(a\) and \(Df\) is differentiable at \(a\). Then all second partial derivatives at \(a\) exist, and for \(u,v\in\mathbb{R}^m\), $$ D^2f(a)[u,v] = \sum_{i=1}^m\sum_{j=1}^m u_i v_j\,\partial_i\partial_j f(a) = u^{\mathsf T}H_f(a)v. $$

Proof. The \(j\)th coordinate of the linear functional \(Df(x)\) is \(Df(x)[e_j]=\partial_j f(x)\). Since \(Df\) is differentiable at \(a\), its derivative in direction \(e_i\) exists. Applying the continuous coordinate evaluation \(L\mapsto L(e_j)\) to that derivative shows that the coordinate function \(x\mapsto\partial_j f(x)\) has derivative in direction \(e_i\) at \(a\). By the definition of a second partial, this derivative is \(\partial_i\partial_j f(a)\). Thus \(D^2f(a)[e_i,e_j]=\partial_i\partial_j f(a)\). Both sides of the claimed identity are bilinear in \(u\) and \(v\). Writing \(u=\sum_i u_i e_i\) and \(v=\sum_j v_j e_j\), bilinearity gives \[ D^2f(a)[u,v] =\sum_{i=1}^m\sum_{j=1}^m u_i v_j D^2f(a)[e_i,e_j] =\sum_{i=1}^m\sum_{j=1}^m u_i v_j\partial_i\partial_j f(a). \] The last expression is the matrix product \(u^{\mathsf T}H_f(a)v\), by the definition of the Hessian entries. \(\square\)

Worked Example: Computing a Hessian

Let \(f(x,y)=x^3+2xy^2+e^y\). Its first partial derivatives are \(f_x=3x^2+2y^2\) and \(f_y=4xy+e^y\). Differentiating these again gives \(f_{xx}=6x\), \(f_{xy}=4y\), \(f_{yx}=4y\), and \(f_{yy}=4x+e^y\). Therefore

$$ H_f(x,y)= \begin{pmatrix} 6x&4y\\ 4y&4x+e^y \end{pmatrix}. $$

At \(a=(1,0)\), this becomes \(\begin{pmatrix}6&0\\0&5\end{pmatrix}\). For the direction \(v=(2,-1)\), the Hessian’s value on the pair \((v,v)\) is \(v^{\mathsf T}H_f(a)v=6(2)^2+5(-1)^2=29\). The same calculation from the bilinear form gives \(D^2f(a)[v,v]=6(2)(2)+5(-1)(-1)=29\). This quantity describes the second derivative along the line in direction \(v\), as the next result explains.

Second Derivatives Along a Direction

Given \(v\in\mathbb{R}^m\), restrict \(f\) to the line through \(a\) in direction \(v\) by setting \(g(t)=f(a+tv)\) for \(t\) near zero. The first derivative of this one-variable function records the directional rate of change. When \(Df\) itself is differentiable at \(a\), the second derivative of this line restriction at zero is the Hessian evaluated twice in direction \(v\).

Theorem (Second Directional Derivative Formula): Suppose \(f\) is differentiable on a neighborhood of \(a\) and \(Df\) is differentiable at \(a\). For \(v\in\mathbb{R}^m\), let \(g(t)=f(a+tv)\) for \(t\) near zero. Then $$ g''(0)=D^2f(a)[v,v]=v^{\mathsf T}H_f(a)v. $$

Proof. The one-variable chain rule gives \(g'(t)=Df(a+tv)[v]\) for \(t\) sufficiently close to zero. For nonzero \(t\), subtract \(g'(0)=Df(a)[v]\) and divide by \(t\): \[ \frac{g'(t)-g'(0)}{t} = \left(\frac{Df(a+tv)-Df(a)}{t}\right)[v]. \] Differentiability of the map \(Df\) at \(a\) implies that the expression in parentheses tends, as a linear map, to \(D(Df)(a)[v]\), because the displacement is \(tv\). Evaluating at \(v\), the limit is \(\bigl(D(Df)(a)[v]\bigr)(v)=D^2f(a)[v,v]\). The Hessian Representation of the Second Derivative identifies this with \(v^{\mathsf T}H_f(a)v\). This limit is the definition of \(g''(0)\), proving the formula. \(\square\)

Worked Example: Curvature Along a Line

For \(f(x,y)=x^2+3xy+2y^2\), the Hessian is the constant matrix \[ H_f(x,y)= \begin{pmatrix} 2&3\\ 3&4 \end{pmatrix}. \] At \(a=(1,1)\), take \(v=(1,-2)\). The line restriction is \(g(t)=f(1+t,1-2t)\). Expanding each term, \[ (1+t)^2=1+2t+t^2,\qquad 3(1+t)(1-2t)=3-3t-6t^2,\qquad 2(1-2t)^2=2-8t+8t^2. \] Adding gives \(g(t)=6-9t+3t^2\), so \(g''(0)=6\). Independently, the Hessian formula gives \(H_f(a)v=(2-6,3-8)=(-4,-5)\), and hence \(v^{\mathsf T}H_f(a)v=(1)(-4)+(-2)(-5)=6\). The direct restriction and the Hessian calculation agree.

When Mixed Partial Derivatives Agree

For many familiar functions, \(f_{xy}\) and \(f_{yx}\) have the same value. Equality is guaranteed by continuity conditions on the mixed partials, but it should not be assumed solely because both derivatives exist. The following proof uses rectangular increments and the one-variable Mean Value Theorem. In particular, it does not assume that the first partial derivatives are continuous.

Theorem (Equality of Mixed Partial Derivatives): Let \(f\) be defined on an open neighborhood of \(a=(a_1,a_2)\). Suppose \(f_{xy}\) and \(f_{yx}\) exist throughout that neighborhood and are continuous at \(a\). Then $$ f_{xy}(a)=f_{yx}(a). $$

Proof. Write \(a=(x_0,y_0)\). Choose nonzero \(h,k\) small enough that the rectangle with corners \((x_0,y_0)\) and \((x_0+h,y_0+k)\) lies in the neighborhood. Define its rectangular increment by \[ \Delta=f(x_0+h,y_0+k)-f(x_0+h,y_0)-f(x_0,y_0+k)+f(x_0,y_0). \] First fix \(k\), and consider \(G(x)=f(x,y_0+k)-f(x,y_0)\) on the interval with endpoints \(x_0\) and \(x_0+h\). The function is continuous on that interval and differentiable in its interior, since the relevant \(x\)-partial derivatives exist. The Mean Value Theorem gives a point \(\xi\) strictly between the endpoints such that \[ \frac{\Delta}{h}=f_x(\xi,y_0+k)-f_x(\xi,y_0). \] For this fixed \(\xi\), the function \(y\mapsto f_x(\xi,y)\) is continuous on the interval between \(y_0\) and \(y_0+k\) and differentiable in its interior because \(f_{yx}\) exists there. A second application of the Mean Value Theorem gives a point \(\eta\) between those endpoints such that \[ \frac{\Delta}{hk}=f_{yx}(\xi,\eta). \] The points \((\xi,\eta)\) lie in the rectangle, so they approach \(a\) as \(h,k\to0\). Continuity of \(f_{yx}\) at \(a\) therefore implies \[ \lim_{\substack{h\to0\\k\to0}}\frac{\Delta}{hk}=f_{yx}(a). \]

Now apply the Mean Value Theorem in the opposite order. First, for fixed \(h\), apply it in \(y\) to \(y\mapsto f(x_0+h,y)-f(x_0,y)\); then apply it in \(x\) to \(x\mapsto f_y(x,\eta')\), where \(\eta'\) is the intermediate \(y\)-coordinate from the first application. This yields intermediate points \(\eta'\) and \(\xi'\) in the same rectangle with \[ \frac{\Delta}{hk}=f_{xy}(\xi',\eta'). \] As \(h,k\to0\), these points also approach \(a\). Continuity of \(f_{xy}\) at \(a\) shows that the same quotient tends to \(f_{xy}(a)\). Since a real-valued expression cannot have two different limits, \(f_{xy}(a)=f_{yx}(a)\). The Mean Value Theorem steps require differentiability of the functions to which it is applied; their continuity follows from that differentiability, not from an assumption that the first partial derivatives are continuous as functions of both variables. \(\square\)

If all the second partial derivatives are continuous on a neighborhood, the theorem applies at every point there. In that setting, the Hessian is symmetric: its \((i,j)\) and \((j,i)\) entries agree. In two variables this means \(f_{xy}=f_{yx}\). The continuity condition in the theorem is important; existence of both mixed partials alone does not suffice.

Worked Example: A Function with Unequal Mixed Partials

Define \[ f(x,y)= \begin{cases} \dfrac{xy(x^2-y^2)}{x^2+y^2},&(x,y)\neq(0,0),\\[4pt] 0,&(x,y)=(0,0). \end{cases} \] Along the coordinate axes, \(f(x,0)=f(0,y)=0\), so \(f_x(0,0)=f_y(0,0)=0\). For fixed \(y\neq0\), the difference quotient for \(f_x(0,y)\) is \[ \frac{f(t,y)-f(0,y)}{t} =y\frac{t^2-y^2}{t^2+y^2}, \] which tends to \(-y\) as \(t\to0\). The same formula gives \(f_x(0,0)=0\) when \(y=0\), so \(f_x(0,y)=-y\) for every \(y\). Consequently, \[ f_{yx}(0,0)=\lim_{s\to0}\frac{f_x(0,s)-f_x(0,0)}{s}=-1. \]

For fixed \(x\neq0\), the quotient for \(f_y(x,0)\) is \[ \frac{f(x,t)-f(x,0)}{t} =x\frac{x^2-t^2}{x^2+t^2}, \] which tends to \(x\) as \(t\to0\). At \(x=0\), \(f_y(0,0)=0\), so \(f_y(x,0)=x\) for every \(x\). It follows that \[ f_{xy}(0,0)=\lim_{t\to0}\frac{f_y(t,0)-f_y(0,0)}{t}=1. \] Thus \(f_{xy}(0,0)=1\) while \(f_{yx}(0,0)=-1\). Both mixed partials exist at the origin, but they are unequal. The continuity hypothesis in the equality theorem cannot simply be omitted.

What Second Derivatives Do—and Do Not—Guarantee

The Hessian packages second-order information into a matrix, and \(v^{\mathsf T}H_f(a)v\) gives the second derivative along a line in direction \(v\), when the second derivative exists as defined above. But merely listing second partial derivatives at a point is weaker than proving that \(Df\) is differentiable there. Nor does the existence of both mixed partials at a point imply they agree. Use the hypotheses of the relevant result: the Hessian Representation requires differentiability of \(Df\), while the Equality of Mixed Partial Derivatives theorem requires continuity of both mixed partials at the point, as well as their existence throughout a neighborhood.

These distinctions matter in the next step of the course. A Taylor approximation with a quadratic term needs coherent second-order information, not just a collection of formal partial derivatives. Under suitable regularity assumptions, the Hessian supplies that quadratic term; directional evaluation then explains how the approximation changes along each line through the point.

Check Your Understanding

Use the definitions and results in this tutorial to answer the following questions.

  1. In the notation \(\partial_i\partial_j f(a)\), which coordinate differentiation is performed first?
  2. What does \(v^{\mathsf T}H_f(a)v\) represent when \(Df\) is differentiable at \(a\)?
  3. What hypotheses in the Equality of Mixed Partial Derivatives theorem allow its Mean Value Theorem proof?
  4. Why does the existence of both mixed partials at a point not, by itself, guarantee they are equal?
  5. For \(f(x,y)=x^2+3xy+2y^2\), compute the second derivative along the line through \((1,1)\) in direction \((1,-2)\).