Tutorials › Real Analysis › The Hessian Matrix

Multivariable Analysis · Tutorial 798 of 1000

The Hessian Matrix

Define and compute the Hessian matrix, relate it to directional second derivatives, and see how it behaves under coordinate changes and for functions of separate variables.

Advanced 10 min read

What You'll Learn

  • Define the Hessian matrix and identify its entries and coordinate order
  • Compute Hessians for polynomial and exponential examples
  • Connect the Hessian quadratic form to second derivatives along lines
  • Transform a Hessian under an invertible affine change of variables
  • Recognize block-diagonal Hessians for sums of functions in separate variables
  • Distinguish symmetry of the Hessian from assumptions that merely guarantee its entries exist

From Critical Points to Second-Order Change

At a differentiable critical point, the gradient vanishes, so the first-order approximation does not show how the function changes nearby. The next layer of information is contained in the second partial derivatives. Arranged in a matrix, they form the Hessian. Its entries describe how the first partial derivatives vary, while its associated quadratic form records second-order change along directions.

This tutorial introduces the Hessian as a matrix and connects it to results established in earlier tutorials. The Hessian Representation of the Second Derivative identifies second derivatives with a bilinear map, and the Second Directional Derivative Formula relates that map to changes along lines. Here we make those ideas concrete in coordinates. The Hessian is important in the analysis of critical points, but it does not by itself classify a point unless the relevant hypotheses and a later test are applied.

Definition and Coordinate Convention

Let \(U\subseteq\mathbb{R}^n\) be open, and suppose \(f:U\to\mathbb{R}\) has second partial derivatives at \(a\in U\). The Hessian is an \(n\)-by-\(n\) matrix. We use the convention that row \(i\), column \(j\) contains the derivative first taken with respect to coordinate \(j\), then with respect to coordinate \(i\).

Definition: The Hessian matrix of \(f\) at \(a\), when all the indicated second partial derivatives exist, is \[ H_f(a)=\bigl(\partial_i\partial_j f(a)\bigr)_{i,j=1}^n = \begin{pmatrix} \partial_1\partial_1 f(a)&\cdots&\partial_1\partial_n f(a)\\ \vdots&\ddots&\vdots\\ \partial_n\partial_1 f(a)&\cdots&\partial_n\partial_n f(a) \end{pmatrix}. \]

For two variables \(x,y\), this convention gives \[ H_f(x,y)= \begin{pmatrix} f_{xx}(x,y)&f_{xy}(x,y)\\ f_{yx}(x,y)&f_{yy}(x,y) \end{pmatrix}. \] When the second partial derivatives are continuous on a neighborhood of the point, the Equality of Mixed Partial Derivatives from the earlier tutorial gives \(f_{xy}=f_{yx}\), so the Hessian is symmetric. Continuity is a convenient sufficient hypothesis for this conclusion; the mere existence of the entries should not be confused with that hypothesis.

If \(Df\) is differentiable at \(a\), then the Hessian represents the second derivative as a bilinear map: \(D^2f(a)[u,v]=u^{T}H_f(a)v\). In particular, setting \(u=v\) gives \(D^2f(a)[v,v]=v^{T}H_f(a)v\). By the Second Directional Derivative Formula, this is the second derivative at zero of the one-variable function \(t\mapsto f(a+tv)\), whenever the stated differentiability assumptions hold. The expression \(v^{T}H_f(a)v\) is called the Hessian quadratic form evaluated at \(v\).

Worked Example: Computing a Hessian from a Polynomial

Let \(f(x,y)=2x^2+3xy-y^2+4x-5y\). First take the partial derivatives: \[ f_x=4x+3y+4,\qquad f_y=3x-2y-5. \] Taking the partial derivatives once more gives \[ f_{xx}=4,\qquad f_{xy}=3,\qquad f_{yx}=3,\qquad f_{yy}=-2. \] Therefore \[ H_f(x,y)=\begin{pmatrix}4&3\\3&-2\end{pmatrix}. \] The entries are constant, so this is the Hessian at every point of \(\mathbb{R}^2\). The equality of the off-diagonal entries is consistent with the continuous second partial derivatives of this polynomial.

Directional Meaning of the Hessian

The matrix is not just a list of derivatives. It provides a way to evaluate second-order change in any direction. If \(v\in\mathbb{R}^n\), the scalar \(v^{T}H_f(a)v\) is the second derivative along the line through \(a\) in direction \(v\), under the differentiability assumptions above. This direction need not be a unit vector; replacing \(v\) by \(cv\) multiplies the quadratic form by \(c^2\).

Worked Example: A Directional Second Derivative

Consider \(f(x,y)=x^2y+e^y\) at \(a=(1,0)\), and take \(v=(1,-1)\). The first derivatives are \[ f_x=2xy,\qquad f_y=x^2+e^y, \] so \(\nabla f(1,0)=(0,2)\). The second partial derivatives are \[ f_{xx}=2y,\qquad f_{xy}=2x,\qquad f_{yx}=2x,\qquad f_{yy}=e^y. \] Hence \[ H_f(1,0)=\begin{pmatrix}0&2\\2&1\end{pmatrix}. \] Multiplying by \(v\) gives \(H_f(1,0)v=(-2,1)\), and therefore \[ v^{T}H_f(1,0)v=(1,-1)\cdot(-2,1)=-2-1=-3. \]

We can check this value directly on the line. Define \(\phi(t)=f(1+t,-t)\). Substitution gives \[ \phi(t)=-(1+t)^2t+e^{-t}=-t-2t^2-t^3+e^{-t}. \] Thus \(\phi'(0)=-1-1=-2\), and \(\phi''(0)=-4+1=-3\). The first derivative also agrees with the directional derivative: \(\nabla f(1,0)\cdot v=(0,2)\cdot(1,-1)=-2\). The second derivative agrees with the Hessian quadratic form. This calculation illustrates why a Hessian describes second-order change, but also why that information should not be mistaken for the whole behavior of the function near a point.

At a critical point, the first-order directional change is zero in every direction. The quadratic form then captures the second-order term in the Taylor approximation. In particular, the Second-Order Necessary Condition for a Local Extremum from earlier in the course implies that, under its hypotheses, a local minimum has a nonnegative Hessian quadratic form in every direction, while a local maximum has a nonpositive one. Those are necessary conditions; classification using second derivatives is the subject of the next tutorial.

How the Hessian Changes with Coordinates

The entries of a Hessian depend on the coordinates used. Under an affine coordinate change, however, the relationship is systematic: the Hessian changes by multiplication on the left and right by the coordinate matrix and its transpose. This is a congruence transformation. It also follows directly from the chain rule.

Theorem (Hessian Under an Affine Change of Variables): Let \(U,V\subseteq\mathbb{R}^n\) be open, let \(T:V\to U\) be \(T(z)=b+Az\) for a fixed \(n\)-by-\(n\) matrix \(A\), and let \(f:U\to\mathbb{R}\). Suppose \(Df\) is differentiable at \(T(z)\), and set \(g=f\circ T\). Then \(Dg\) is differentiable at \(z\), and \[ H_g(z)=A^{T}H_f(T(z))A. \]

Proof. The Multivariable Chain Rule gives \(Dg(z)=Df(T(z))\circ A\). In gradient form, this says \(\nabla g(z)=A^{T}\nabla f(T(z))\). Differentiate this identity with respect to \(z\). The matrix \(A^{T}\) is constant, and differentiating \(\nabla f(T(z))\) gives \(H_f(T(z))A\), again by the chain rule. Consequently, \[ H_g(z)=A^{T}\bigl(H_f(T(z))A\bigr)=A^{T}H_f(T(z))A. \] This proves the formula. \(\square\)

The theorem does not require \(A\) to be invertible for the displayed identity, although invertibility is needed if the coordinate change is to be a bijective reparametrization. The formula also explains why Hessian entries generally change after a rotation or rescaling. For a direction \(w\) in the new coordinates, \[ w^{T}H_g(z)w=(Aw)^{T}H_f(T(z))(Aw), \] so the second-order change in the new direction is the change in the corresponding direction in the original coordinates.

Worked Example: A Hessian After a Linear Change

Let \(f(u,v)=u^2+4uv+3v^2\), and introduce coordinates \(u=x+y\), \(v=x-y\). The matrix taking \((x,y)\) to \((u,v)\) is \[ A=\begin{pmatrix}1&1\\1&-1\end{pmatrix}. \] The original Hessian is \[ H_f(u,v)=\begin{pmatrix}2&4\\4&6\end{pmatrix}. \] First calculate \[ H_fA=\begin{pmatrix}2&4\\4&6\end{pmatrix} \begin{pmatrix}1&1\\1&-1\end{pmatrix} =\begin{pmatrix}6&-2\\10&-2\end{pmatrix}. \] Since \(A^T=A\), the affine transformation theorem gives \[ H_g(x,y)=A^TH_fA =\begin{pmatrix}1&1\\1&-1\end{pmatrix} \begin{pmatrix}6&-2\\10&-2\end{pmatrix} =\begin{pmatrix}16&-4\\-4&0\end{pmatrix}. \] As a check, direct substitution gives \(g(x,y)=f(x+y,x-y)=8x^2-4xy\). Its second partial derivatives are \(g_{xx}=16\), \(g_{xy}=g_{yx}=-4\), and \(g_{yy}=0\), yielding the same Hessian.

Separate Variables and Block Structure

A common simplification occurs when a function is a sum of two functions that depend on separate groups of variables. Mixed second partial derivatives between the groups then vanish. The Hessian has a block-diagonal form, which can make later calculations more manageable.

Theorem (Hessian of a Sum in Separate Variables): Let \(U_1\subseteq\mathbb{R}^p\) and \(U_2\subseteq\mathbb{R}^q\) be open. Suppose \(r:U_1\to\mathbb{R}\) and \(s:U_2\to\mathbb{R}\) have continuous second partial derivatives, and define \(F(x,y)=r(x)+s(y)\) on \(U_1\times U_2\). Then \[ H_F(x,y)= \begin{pmatrix} H_r(x)&0\\ 0&H_s(y) \end{pmatrix}. \]

Proof. Write \(x=(x_1,\ldots,x_p)\) and \(y=(y_1,\ldots,y_q)\). If both differentiations are with respect to \(x\)-coordinates, then the corresponding second partial derivative of \(F\) is the same as that of \(r\), because \(s(y)\) is constant with respect to \(x\). Similarly, two differentiations with respect to \(y\)-coordinates give the corresponding second partial derivative of \(s\), because \(r(x)\) is constant with respect to \(y\). For a mixed derivative, first differentiating \(F\) with respect to any \(x_i\) gives \(\partial_i r(x)\), which does not depend on \(y\); differentiating that with respect to any \(y_j\) gives zero. Reversing the order gives zero as well, since the \(y_j\)-derivative is \(\partial_j s(y)\), independent of \(x\). Thus the two diagonal blocks are \(H_r(x)\) and \(H_s(y)\), and both off-diagonal blocks consist of zeros. \(\square\)

Worked Example: A Block-Diagonal Hessian

Let \(F(x,y,z)=x^2+\sin y+e^z\), where the variables occur in separate terms. The second derivatives within each variable are \(2\), \(-\sin y\), and \(e^z\), respectively. Every mixed second partial derivative is zero because differentiating one term with respect to a variable absent from that term gives zero. Therefore \[ H_F(x,y,z)= \begin{pmatrix} 2&0&0\\ 0&-\sin y&0\\ 0&0&e^z \end{pmatrix}. \] For a vector \(w=(w_1,w_2,w_3)\), its Hessian quadratic form is \(2w_1^2-(\sin y)w_2^2+e^zw_3^2\). Each coordinate contributes separately; there are no cross terms.

What the Hessian Tells You—and What It Does Not

The Hessian organizes second partial derivatives and provides the quadratic form governing second-order directional change. It can reveal interactions between coordinates through off-diagonal entries, and its form can simplify under separate-variable decompositions or transform predictably under affine coordinates. These features make it a central tool for studying local behavior.

A frequent mistake is to infer an extremum merely from the existence of a Hessian, or to interpret one entry in isolation. The Hessian must be considered together with the point and the directions being studied. A zero or mixed-sign quadratic form may leave second-order information inconclusive; higher-order terms or direct comparisons can matter. Likewise, the symmetry of the Hessian is guaranteed by standard regularity assumptions such as continuity of the second partial derivatives, not simply by writing down a matrix of second partials. The next tutorial develops the appropriate second derivative test and explains when the Hessian does decide the behavior of a critical point.

Check Your Understanding

Use the definitions, formulas, and examples in this tutorial to answer the following questions.

  1. Under the stated row-and-column convention, what does the \((i,j)\)-entry of the Hessian represent?
  2. How is the Hessian quadratic form \(v^{T}H_f(a)v\) related to the second derivative of the restriction of \(f\) to a line?
  3. For \(g=f\circ T\) with \(T(z)=b+Az\), what is the formula for \(H_g(z)\), and which matrices are transposed?
  4. Why do the off-diagonal blocks vanish when \(F(x,y)=r(x)+s(y)\)?
  5. What regularity assumption stated here guarantees equality of the two mixed partial derivatives?
  6. Why does knowing the Hessian alone not always settle whether a critical point is a local extremum?