Tutorials › Real Analysis › Multivariable Analysis Mastery

Multivariable Analysis · Tutorial 800 of 1000

Multivariable Analysis Mastery

Learn how a nonlinear change of coordinates transforms the gradient and Hessian, and use the result to classify critical points without confusing coordinate effects with genuine changes in local behavior.

Advanced 12 min read

What You'll Learn

  • Derive the gradient and Hessian of a function composed with a nonlinear coordinate map
  • Identify the extra Hessian term that appears away from a critical point
  • Use an invertible derivative to preserve definiteness and indefiniteness of the Hessian
  • Apply coordinate transformations to classify minima and saddle points
  • Distinguish a pointwise Hessian calculation from a claim about local invertibility

Connecting Derivatives, Hessians, and Coordinates

The preceding tutorial used the Hessian at a critical point to classify strict local minima, strict local maxima, and saddle points. A natural question is how much that classification depends on the coordinates used to describe the function. A linear change of variables transforms the Hessian by a matrix congruence. A nonlinear change of variables has an additional term, but that term vanishes at a critical point. This gives a useful synthesis of the chain rule, the Hessian, and the Second Derivative Test.

Let \(U,V\subseteq\mathbb{R}^n\) be open, let \(f:U\to\mathbb{R}\) have continuous second partial derivatives, and let \(\phi:V\to U\) also have continuous second partial derivatives. Write \(g=f\circ\phi\). At a point \(b\in V\), set \(a=\phi(b)\), and let \(A=D\phi(b)\) be the Jacobian matrix of \(\phi\) at \(b\). We will first calculate the derivatives of \(g\), then examine what the formulas say when \(a\) is critical for \(f\).

The Hessian Formula for a Composition

The first derivative follows from the Multivariable Chain Rule: the gradient of a scalar-valued composition is obtained by multiplying by the transpose of the Jacobian. In second derivatives, the Jacobian alone does not tell the whole story. The coordinate map itself may curve, and its second derivatives contribute to the Hessian.

Theorem (Gradient and Hessian of a Composition): With \(f,\phi,g,a,b\), and \(A\) as above, $$ \nabla g(b)=A^T\nabla f(a). $$ For \(i,j\in\{1,\ldots,n\}\), the Hessian entries satisfy $$ \partial_{ij}g(b) = \sum_{r=1}^{n}\sum_{s=1}^{n} \partial_{rs}f(a)\,\partial_i\phi_r(b)\,\partial_j\phi_s(b) + \sum_{r=1}^{n}\partial_r f(a)\,\partial_{ij}\phi_r(b). $$ Equivalently, $$ H_g(b)=A^T H_f(a)A+ \sum_{r=1}^{n}\partial_r f(a)\,H_{\phi_r}(b). $$ In particular, if \(\nabla f(a)=0\), then \(H_g(b)=A^T H_f(a)A\).

Proof. For each coordinate \(i\), the one-variable chain rule applied to the \(i\)-th coordinate direction gives $$ \partial_i g(b)=\sum_{r=1}^{n}\partial_r f(a)\,\partial_i\phi_r(b). $$ These component equations are exactly \(\nabla g(b)=A^T\nabla f(a)\).

Differentiate the displayed formula for \(\partial_i g\) with respect to the \(j\)-th coordinate. The product and chain rules give $$ \partial_{ij}g(b) = \sum_{r=1}^{n} \left( \sum_{s=1}^{n}\partial_{sr}f(a)\,\partial_j\phi_s(b) \right)\partial_i\phi_r(b) + \sum_{r=1}^{n}\partial_r f(a)\,\partial_{ij}\phi_r(b). $$ Since \(f\) has continuous second partial derivatives, its mixed partials agree, so \(\partial_{sr}f(a)=\partial_{rs}f(a)\). Reordering the finite sums yields the stated entry formula. The first double sum is the \((i,j)\)-entry of \(A^TH_f(a)A\), and the second sum is the \((i,j)\)-entry of \(\sum_r\partial_r f(a)H_{\phi_r}(b)\). If \(\nabla f(a)=0\), every coefficient \(\partial_r f(a)\) in the second matrix sum is zero, proving the final formula. \(\square\)

The extra term is the main feature to keep in view. For an affine coordinate map, every Hessian \(H_{\phi_r}\) is zero, so the term disappears everywhere. For a genuinely nonlinear map, it generally does not disappear away from critical points. At a critical point of \(f\), however, it vanishes regardless of how nonlinear \(\phi\) is.

What an Invertible Coordinate Derivative Preserves

When \(A\) is invertible, the formula at a critical point does more than provide a way to compute the new Hessian. It says the two Hessian quadratic forms take the same values after a bijective relabeling of directions. Consequently, their definiteness and indefiniteness agree.

Theorem (Preservation of Hessian Classification at a Critical Point): Suppose the hypotheses above hold, \(\nabla f(a)=0\), and \(A=D\phi(b)\) is invertible. Then \(b\) is a critical point of \(g=f\circ\phi\), and for every \(w\in\mathbb{R}^n\), $$ w^T H_g(b)w=(Aw)^T H_f(a)(Aw). $$ Thus \(H_g(b)\) and \(H_f(a)\) are positive definite at the same time, negative definite at the same time, and indefinite at the same time. The same correspondence holds for positive and negative semidefiniteness.

Proof. The gradient formula gives \(\nabla g(b)=A^T\nabla f(a)=0\), so \(b\) is critical for \(g\). The Hessian formula at the critical point gives \(H_g(b)=A^TH_f(a)A\). Multiplying on the left by \(w^T\) and on the right by \(w\) gives $$ w^T H_g(b)w=w^TA^TH_f(a)Aw=(Aw)^TH_f(a)(Aw). $$

Because \(A\) is invertible, \(w\ne0\) if and only if \(Aw\ne0\), and every vector \(v\in\mathbb{R}^n\) has the form \(v=Aw\) for exactly one \(w\). Therefore, \(w^TH_g(b)w\) is positive for every nonzero \(w\) exactly when \(v^TH_f(a)v\) is positive for every nonzero \(v\). This proves the equivalence of positive definiteness; the same reasoning with negative values proves the equivalence of negative definiteness. If \(H_f(a)\) is indefinite, choose nonzero \(u,v\) giving a positive and a negative quadratic-form value. Since \(A\) is onto, there are \(w_1,w_2\) with \(Aw_1=u\) and \(Aw_2=v\); both are nonzero, and the corresponding values for \(H_g(b)\) have opposite signs. The converse follows by applying the same argument to \(A^{-1}\). Finally, replacing “positive” or “negative” by “nonnegative” or “nonpositive” proves the semidefinite statements. \(\square\)

This is a statement about the Hessian and the Second Derivative Test. It does not claim that a map with an invertible derivative at one point has already been proved to have a local inverse. Establishing local invertibility is a separate matter. Here, invertibility of \(A\) is used only to show that its action gives a one-to-one correspondence between directions at the two points.

Worked Examples

Worked Example: A Nonlinear Change of Coordinates at a Minimum

Let \(f(x,y)=x^2+2y^2\), and define \(\phi(u,v)=(u+v^2,v+u^2)\). At the origin, \(\phi(0,0)=(0,0)\), and $$ D\phi(0,0)= \begin{pmatrix}1&0\\0&1\end{pmatrix}, \qquad \nabla f(0,0)=(0,0), \qquad H_f(0,0)= \begin{pmatrix}2&0\\0&4\end{pmatrix}. $$ The Jacobian is invertible, and \(H_f(0,0)\) is positive definite. The composition is $$ g(u,v)=(u+v^2)^2+2(v+u^2)^2. $$ The Hessian transformation theorem gives $$ H_g(0,0)=D\phi(0,0)^T H_f(0,0)D\phi(0,0) = \begin{pmatrix}2&0\\0&4\end{pmatrix}. $$ Thus \(g\) has a positive definite Hessian at its critical point \((0,0)\). By the Second Derivative Test, it has a strict local minimum there. The nonlinear terms in \(\phi\) do not contribute to the Hessian formula at this point because \(\nabla f(0,0)=0\).

Worked Example: A Saddle Point in New Coordinates

Take \(f(x,y)=x^2-y^2\), with \(\phi(u,v)=(u+v^2,v+u^2)\) again. The point \(a=(0,0)\) is critical, and $$ H_f(0,0)= \begin{pmatrix}2&0\\0&-2\end{pmatrix}, \qquad D\phi(0,0)= \begin{pmatrix}1&0\\0&1\end{pmatrix}. $$ The Hessian is indefinite. For \(g=f\circ\phi\), the composition formula gives $$ H_g(0,0)= \begin{pmatrix}2&0\\0&-2\end{pmatrix}. $$ Indeed, direct substitution yields $$ g(u,v)=(u+v^2)^2-(v+u^2)^2. $$ Along the \(u\)-axis, \(g(t,0)=t^2-t^4\), which is positive for all sufficiently small nonzero \(t\). Along the \(v\)-axis, \(g(0,t)=t^4-t^2\), which is negative for all sufficiently small nonzero \(t\). Hence the origin is a saddle point for \(g\), as the indefinite Hessian predicts. The coordinate transformation has not changed the second-order classification.

Worked Example: Why the Extra Term Cannot Be Dropped

Let \(f(x,y)=x+y\) and \(\phi(u,v)=(u,v+u^2)\). Then \(g=f\circ\phi\) is $$ g(u,v)=u+v+u^2. $$ At the origin, \(\nabla f(0,0)=(1,1)\), so the point is not critical for \(f\). Also, $$ D\phi(0,0)= \begin{pmatrix}1&0\\0&1\end{pmatrix}, \qquad H_f(0,0)= \begin{pmatrix}0&0\\0&0\end{pmatrix}, \qquad H_{\phi_1}(0,0)= \begin{pmatrix}0&0\\0&0\end{pmatrix}, \qquad H_{\phi_2}(0,0)= \begin{pmatrix}2&0\\0&0\end{pmatrix}. $$ The full composition formula gives $$ H_g(0,0) = D\phi(0,0)^TH_f(0,0)D\phi(0,0) + 1\cdot H_{\phi_1}(0,0) + 1\cdot H_{\phi_2}(0,0) = \begin{pmatrix}2&0\\0&0\end{pmatrix}. $$ Direct differentiation of \(g(u,v)=u+v+u^2\) confirms that its Hessian is exactly this matrix. If the second term in the composition formula were omitted, the result would incorrectly be the zero matrix. The critical-point hypothesis is essential for simplifying the formula to \(A^TH_f(a)A\).

A Reliable Coordinate-Change Workflow

The results above provide a compact method for analyzing a scalar function after a change of variables. In particular, they prevent a common mistake: treating every nonlinear coordinate change as though it were affine. When the point is not critical, the second derivatives of the coordinate map can affect the Hessian.

1
Locate the corresponding points.
For a map \(\phi:V\to U\), choose \(b\in V\) and calculate \(a=\phi(b)\). The derivatives of \(g=f\circ\phi\) at \(b\) involve the derivatives of \(f\) at \(a\).
2
Check criticality before simplifying.
Use \(\nabla g(b)=D\phi(b)^T\nabla f(a)\). Do not discard the coordinate-curvature term in the Hessian formula unless \(\nabla f(a)=0\), or unless the coordinate map is affine.
3
Apply the full Hessian formula.
At a critical point of \(f\), use \(H_g(b)=A^TH_f(a)A\). Otherwise, include \(\sum_r\partial_r f(a)H_{\phi_r}(b)\).
4
Use invertibility for classification.
If \(A\) is invertible, the two Hessian quadratic forms take corresponding values in every direction. Apply the Second Derivative Test to the resulting definiteness or indefiniteness.

The broader lesson is that a derivative calculation and a geometric conclusion are related but distinct tasks. The chain rule supplies the derivative formulas; the critical-point condition removes the nonlinear correction in the Hessian; and invertibility ensures that no direction is lost when comparing the quadratic forms. Keeping these roles separate makes the method dependable, especially when a problem introduces new coordinates to simplify its algebra.

Check Your Understanding

Use the composition formulas and the role of criticality to answer the following questions.

  1. What additional term appears in the Hessian of \(f\circ\phi\) beyond \(A^TH_f(a)A\)?
  2. Why does that additional term vanish when \(a\) is a critical point of \(f\)?
  3. If \(A\) is invertible, why does positive definiteness of \(H_f(a)\) imply positive definiteness of \(H_g(b)\) at a critical point?
  4. Why is invertibility of \(A\) needed to conclude that indefiniteness is preserved in both directions?
  5. In the third worked example, which term accounts for the nonzero Hessian of the composition?
  6. Does an invertible derivative at one point, by itself, establish that \(\phi\) has a local inverse? Explain what the Hessian comparison uses instead.