Connecting Derivatives, Hessians, and Coordinates
The preceding tutorial used the Hessian at a critical point to classify strict local minima, strict local maxima, and saddle points. A natural question is how much that classification depends on the coordinates used to describe the function. A linear change of variables transforms the Hessian by a matrix congruence. A nonlinear change of variables has an additional term, but that term vanishes at a critical point. This gives a useful synthesis of the chain rule, the Hessian, and the Second Derivative Test.
Let \(U,V\subseteq\mathbb{R}^n\) be open, let \(f:U\to\mathbb{R}\) have continuous second partial derivatives, and let \(\phi:V\to U\) also have continuous second partial derivatives. Write \(g=f\circ\phi\). At a point \(b\in V\), set \(a=\phi(b)\), and let \(A=D\phi(b)\) be the Jacobian matrix of \(\phi\) at \(b\). We will first calculate the derivatives of \(g\), then examine what the formulas say when \(a\) is critical for \(f\).
The Hessian Formula for a Composition
The first derivative follows from the Multivariable Chain Rule: the gradient of a scalar-valued composition is obtained by multiplying by the transpose of the Jacobian. In second derivatives, the Jacobian alone does not tell the whole story. The coordinate map itself may curve, and its second derivatives contribute to the Hessian.
Proof. For each coordinate \(i\), the one-variable chain rule applied to the \(i\)-th coordinate direction gives $$ \partial_i g(b)=\sum_{r=1}^{n}\partial_r f(a)\,\partial_i\phi_r(b). $$ These component equations are exactly \(\nabla g(b)=A^T\nabla f(a)\).
Differentiate the displayed formula for \(\partial_i g\) with respect to the \(j\)-th coordinate. The product and chain rules give $$ \partial_{ij}g(b) = \sum_{r=1}^{n} \left( \sum_{s=1}^{n}\partial_{sr}f(a)\,\partial_j\phi_s(b) \right)\partial_i\phi_r(b) + \sum_{r=1}^{n}\partial_r f(a)\,\partial_{ij}\phi_r(b). $$ Since \(f\) has continuous second partial derivatives, its mixed partials agree, so \(\partial_{sr}f(a)=\partial_{rs}f(a)\). Reordering the finite sums yields the stated entry formula. The first double sum is the \((i,j)\)-entry of \(A^TH_f(a)A\), and the second sum is the \((i,j)\)-entry of \(\sum_r\partial_r f(a)H_{\phi_r}(b)\). If \(\nabla f(a)=0\), every coefficient \(\partial_r f(a)\) in the second matrix sum is zero, proving the final formula. \(\square\)
The extra term is the main feature to keep in view. For an affine coordinate map, every Hessian \(H_{\phi_r}\) is zero, so the term disappears everywhere. For a genuinely nonlinear map, it generally does not disappear away from critical points. At a critical point of \(f\), however, it vanishes regardless of how nonlinear \(\phi\) is.
What an Invertible Coordinate Derivative Preserves
When \(A\) is invertible, the formula at a critical point does more than provide a way to compute the new Hessian. It says the two Hessian quadratic forms take the same values after a bijective relabeling of directions. Consequently, their definiteness and indefiniteness agree.
Proof. The gradient formula gives \(\nabla g(b)=A^T\nabla f(a)=0\), so \(b\) is critical for \(g\). The Hessian formula at the critical point gives \(H_g(b)=A^TH_f(a)A\). Multiplying on the left by \(w^T\) and on the right by \(w\) gives $$ w^T H_g(b)w=w^TA^TH_f(a)Aw=(Aw)^TH_f(a)(Aw). $$
Because \(A\) is invertible, \(w\ne0\) if and only if \(Aw\ne0\), and every vector \(v\in\mathbb{R}^n\) has the form \(v=Aw\) for exactly one \(w\). Therefore, \(w^TH_g(b)w\) is positive for every nonzero \(w\) exactly when \(v^TH_f(a)v\) is positive for every nonzero \(v\). This proves the equivalence of positive definiteness; the same reasoning with negative values proves the equivalence of negative definiteness. If \(H_f(a)\) is indefinite, choose nonzero \(u,v\) giving a positive and a negative quadratic-form value. Since \(A\) is onto, there are \(w_1,w_2\) with \(Aw_1=u\) and \(Aw_2=v\); both are nonzero, and the corresponding values for \(H_g(b)\) have opposite signs. The converse follows by applying the same argument to \(A^{-1}\). Finally, replacing “positive” or “negative” by “nonnegative” or “nonpositive” proves the semidefinite statements. \(\square\)
This is a statement about the Hessian and the Second Derivative Test. It does not claim that a map with an invertible derivative at one point has already been proved to have a local inverse. Establishing local invertibility is a separate matter. Here, invertibility of \(A\) is used only to show that its action gives a one-to-one correspondence between directions at the two points.
Worked Examples
Worked Example: A Nonlinear Change of Coordinates at a Minimum
Let \(f(x,y)=x^2+2y^2\), and define \(\phi(u,v)=(u+v^2,v+u^2)\). At the origin, \(\phi(0,0)=(0,0)\), and $$ D\phi(0,0)= \begin{pmatrix}1&0\\0&1\end{pmatrix}, \qquad \nabla f(0,0)=(0,0), \qquad H_f(0,0)= \begin{pmatrix}2&0\\0&4\end{pmatrix}. $$ The Jacobian is invertible, and \(H_f(0,0)\) is positive definite. The composition is $$ g(u,v)=(u+v^2)^2+2(v+u^2)^2. $$ The Hessian transformation theorem gives $$ H_g(0,0)=D\phi(0,0)^T H_f(0,0)D\phi(0,0) = \begin{pmatrix}2&0\\0&4\end{pmatrix}. $$ Thus \(g\) has a positive definite Hessian at its critical point \((0,0)\). By the Second Derivative Test, it has a strict local minimum there. The nonlinear terms in \(\phi\) do not contribute to the Hessian formula at this point because \(\nabla f(0,0)=0\).
Worked Example: A Saddle Point in New Coordinates
Take \(f(x,y)=x^2-y^2\), with \(\phi(u,v)=(u+v^2,v+u^2)\) again. The point \(a=(0,0)\) is critical, and $$ H_f(0,0)= \begin{pmatrix}2&0\\0&-2\end{pmatrix}, \qquad D\phi(0,0)= \begin{pmatrix}1&0\\0&1\end{pmatrix}. $$ The Hessian is indefinite. For \(g=f\circ\phi\), the composition formula gives $$ H_g(0,0)= \begin{pmatrix}2&0\\0&-2\end{pmatrix}. $$ Indeed, direct substitution yields $$ g(u,v)=(u+v^2)^2-(v+u^2)^2. $$ Along the \(u\)-axis, \(g(t,0)=t^2-t^4\), which is positive for all sufficiently small nonzero \(t\). Along the \(v\)-axis, \(g(0,t)=t^4-t^2\), which is negative for all sufficiently small nonzero \(t\). Hence the origin is a saddle point for \(g\), as the indefinite Hessian predicts. The coordinate transformation has not changed the second-order classification.
Worked Example: Why the Extra Term Cannot Be Dropped
Let \(f(x,y)=x+y\) and \(\phi(u,v)=(u,v+u^2)\). Then \(g=f\circ\phi\) is $$ g(u,v)=u+v+u^2. $$ At the origin, \(\nabla f(0,0)=(1,1)\), so the point is not critical for \(f\). Also, $$ D\phi(0,0)= \begin{pmatrix}1&0\\0&1\end{pmatrix}, \qquad H_f(0,0)= \begin{pmatrix}0&0\\0&0\end{pmatrix}, \qquad H_{\phi_1}(0,0)= \begin{pmatrix}0&0\\0&0\end{pmatrix}, \qquad H_{\phi_2}(0,0)= \begin{pmatrix}2&0\\0&0\end{pmatrix}. $$ The full composition formula gives $$ H_g(0,0) = D\phi(0,0)^TH_f(0,0)D\phi(0,0) + 1\cdot H_{\phi_1}(0,0) + 1\cdot H_{\phi_2}(0,0) = \begin{pmatrix}2&0\\0&0\end{pmatrix}. $$ Direct differentiation of \(g(u,v)=u+v+u^2\) confirms that its Hessian is exactly this matrix. If the second term in the composition formula were omitted, the result would incorrectly be the zero matrix. The critical-point hypothesis is essential for simplifying the formula to \(A^TH_f(a)A\).
A Reliable Coordinate-Change Workflow
The results above provide a compact method for analyzing a scalar function after a change of variables. In particular, they prevent a common mistake: treating every nonlinear coordinate change as though it were affine. When the point is not critical, the second derivatives of the coordinate map can affect the Hessian.
For a map \(\phi:V\to U\), choose \(b\in V\) and calculate \(a=\phi(b)\). The derivatives of \(g=f\circ\phi\) at \(b\) involve the derivatives of \(f\) at \(a\).
Use \(\nabla g(b)=D\phi(b)^T\nabla f(a)\). Do not discard the coordinate-curvature term in the Hessian formula unless \(\nabla f(a)=0\), or unless the coordinate map is affine.
At a critical point of \(f\), use \(H_g(b)=A^TH_f(a)A\). Otherwise, include \(\sum_r\partial_r f(a)H_{\phi_r}(b)\).
If \(A\) is invertible, the two Hessian quadratic forms take corresponding values in every direction. Apply the Second Derivative Test to the resulting definiteness or indefiniteness.
The broader lesson is that a derivative calculation and a geometric conclusion are related but distinct tasks. The chain rule supplies the derivative formulas; the critical-point condition removes the nonlinear correction in the Hessian; and invertibility ensures that no direction is lost when comparing the quadratic forms. Keeping these roles separate makes the method dependable, especially when a problem introduces new coordinates to simplify its algebra.
Check Your Understanding
Use the composition formulas and the role of criticality to answer the following questions.
- What additional term appears in the Hessian of \(f\circ\phi\) beyond \(A^TH_f(a)A\)?
- Why does that additional term vanish when \(a\) is a critical point of \(f\)?
- If \(A\) is invertible, why does positive definiteness of \(H_f(a)\) imply positive definiteness of \(H_g(b)\) at a critical point?
- Why is invertibility of \(A\) needed to conclude that indefiniteness is preserved in both directions?
- In the third worked example, which term accounts for the nonzero Hessian of the composition?
- Does an invertible derivative at one point, by itself, establish that \(\phi\) has a local inverse? Explain what the Hessian comparison uses instead.