Tutorials › Real Analysis › Applications of the Inverse Function Theorem

Multivariable Analysis · Tutorial 804 of 1000

Applications of the Inverse Function Theorem

Use local inverses to solve nonlinear equations, understand how nearby targets affect solutions, and compare critical points in different smooth coordinates.

Advanced 10 min read

What You'll Learn

  • Apply the Inverse Function Theorem to obtain unique local solutions of nonlinear systems
  • Interpret the derivative of a local inverse as the first-order response of a solution to changes in its target
  • Prove that a continuously differentiable map sends open sets of regular points to open sets
  • Track Hessians through a nonlinear change of coordinates at a critical point
  • Use coordinate changes to preserve definite and indefinite second-derivative classifications

What a Local Inverse Lets Us Do

The Inverse Function Theorem is more than a criterion for when an inverse exists. It turns a nonsingular derivative at one point into practical conclusions nearby: nonlinear equations have unique local solutions, those solutions vary regularly with the target, and the map sends suitable neighborhoods to open neighborhoods. It also lets us change coordinates near a critical point without changing its second-derivative classification.

Throughout, let \(f:U\to\mathbb{R}^n\) be continuously differentiable on an open set \(U\subseteq\mathbb{R}^n\). When \(Df(a)\) is invertible, the Inverse Function Theorem supplies open neighborhoods \(U_0\) of \(a\) and \(V_0\) of \(f(a)\) such that \(f:U_0\to V_0\) is bijective and its inverse is continuously differentiable. We will use that theorem and the derivative formula for the inverse established earlier, rather than repeat their proofs.

Solving a Nonlinear System Near a Known Solution

Suppose a system \(f(x)=y\) has a known solution \(x=a\) when the target is \(y=b=f(a)\). If \(Df(a)\) is invertible, the theorem gives a neighborhood of \(b\) in which every target has exactly one solution near \(a\). This is a local existence-and-uniqueness statement: it does not claim that there are no other solutions elsewhere in \(U\).

There is also a first-order description of how the solution changes. Write the local inverse as \(g\), so \(g(b)=a\). The derivative formula gives $$ Dg(b)=\bigl(Df(a)\bigr)^{-1}. $$ Consequently, for a sufficiently small target change \(k\), $$ g(b+k)=a+\bigl(Df(a)\bigr)^{-1}k+o(\|k\|_2). $$ The inverse Jacobian therefore converts a small change in the target into the corresponding first-order change in the solution. This is useful in sensitivity analysis: a large operator norm of \(\bigl(Df(a)\bigr)^{-1}\) indicates that a small target perturbation may cause a relatively large solution change.

Worked Example: A Two-Equation System Near the Origin

Consider $$ f(x,y)=(x+y+x^2,\ x-y). $$ At \((0,0)\), \(f(0,0)=(0,0)\), and $$ Df(0,0)= \begin{pmatrix} 1&1\\ 1&-1 \end{pmatrix}, \qquad \det Df(0,0)=-2\neq0. $$ The Inverse Function Theorem gives neighborhoods of the origin on which every sufficiently small target \((s,t)\) has exactly one preimage near \((0,0)\).

The derivative of the local solution map at the origin is $$ \bigl(Df(0,0)\bigr)^{-1} = \begin{pmatrix} 1/2&1/2\\ 1/2&-1/2 \end{pmatrix}. $$ Thus, to first order, changing the target by \((\delta s,\delta t)\) changes the solution by $$ \left(\frac{\delta s+\delta t}{2},\frac{\delta s-\delta t}{2}\right). $$ For a direct check of the local solution, the second equation gives \(y=x-t\). Substituting into the first gives \(x^2+2x-(s+t)=0\). The root near zero is \(x=-1+\sqrt{1+s+t}\), when the square root is defined near \(1\), and \(y=x-t\). At \((s,t)=(0,0)\), these formulas give \(x=0\) and \(y=0\), as required. The other quadratic root is \(-1-\sqrt{1+s+t}\), which equals \(-2\) at the origin and is not the solution near \((0,0)\).

Regular Points Give Open Images

A point \(x\in U\) is called a regular point of \(f\) if \(Df(x)\) is invertible. A direct application of the Inverse Function Theorem shows that \(f\) sends open sets consisting of regular points to open sets. This is a local property; it does not require \(f\) to be one-to-one across the entire open set.

Theorem (Open Images of Regular Points): Let \(U\subseteq\mathbb{R}^n\) be open, let \(f:U\to\mathbb{R}^n\) be continuously differentiable, and let \(O\subseteq U\) be open. If \(Df(x)\) is invertible for every \(x\in O\), then \(f(O)\) is open in \(\mathbb{R}^n\).

Proof. Fix \(y\in f(O)\). By definition of the image, there exists \(x\in O\) with \(f(x)=y\). Since \(O\) is open and \(f\) is continuously differentiable on \(U\), the restriction of \(f\) to \(O\) satisfies the hypotheses of the Inverse Function Theorem at \(x\): its derivative \(Df(x)\) is invertible. The theorem supplies an open neighborhood \(V_x\) of \(y\) and an open neighborhood \(W_x\) of \(x\), with \(W_x\subseteq O\), such that \(f(W_x)=V_x\). Hence \(V_x\subseteq f(O)\). We have found an open neighborhood contained in \(f(O)\) for every \(y\in f(O)\), so \(f(O)\) is open. \(\square\)

The same argument shows that the set of regular points is open. Indeed, if \(x\) is regular, apply the Inverse Function Theorem to obtain a neighborhood on which \(f\) is a diffeomorphism. At every point in that neighborhood the derivative is invertible: the chain rule applied to the local inverse and \(f\) shows that the derivative of the inverse is a two-sided inverse for the derivative of \(f\). Thus nearby points are regular as well.

Worked Example: An Open Map with a Simple Nonlinear Coordinate

Define \(f:\mathbb{R}^2\to\mathbb{R}^2\) by $$ f(x,y)=(x,\ y+y^3). $$ Its derivative matrix is $$ Df(x,y)= \begin{pmatrix} 1&0\\ 0&1+3y^2 \end{pmatrix}, \qquad \det Df(x,y)=1+3y^2>0. $$ Every point is regular. The open-images theorem therefore implies that \(f(O)\) is open for every open set \(O\subseteq\mathbb{R}^2\).

For example, if \(O\) is an open ball around \((0,0)\), then \(f(O)\) contains an open neighborhood of \(f(0,0)=(0,0)\). The derivative calculation gives this local conclusion without needing an explicit formula for the inverse of the cubic coordinate \(y\mapsto y+y^3\). The conclusion is about openness of the image, not about a fixed-radius ball being mapped to a ball of the same radius.

Changing Coordinates at a Critical Point

A smooth coordinate change can make a problem easier to express, but it changes the matrix representing the Hessian. At a critical point, the transformation has a particularly useful form: the Hessian changes by multiplication on the left by a matrix transpose and on the right by the matrix itself. This preserves whether the Hessian is positive definite, negative definite, or indefinite.

Let \(f:U\to\mathbb{R}\) be twice continuously differentiable, and let \(\Phi:W\to U\) be a twice continuously differentiable local diffeomorphism. Here “local diffeomorphism” means that near the point under consideration \(\Phi\) is a bijection onto an open set with a twice continuously differentiable inverse. Write \(a=\Phi(c)\), \(A=D\Phi(c)\), and \(q=f\circ\Phi\). The matrix \(A\) is invertible by the chain rule applied to \(\Phi^{-1}\circ\Phi\).

Theorem (Hessian Under a Nonlinear Coordinate Change at a Critical Point): If \(\nabla f(a)=0\), then \(\nabla q(c)=0\) and $$ H_q(c)=A^T H_f(a)A. $$ In particular, \(H_q(c)\) is positive definite if and only if \(H_f(a)\) is positive definite, and it is negative definite if and only if \(H_f(a)\) is negative definite. If \(H_f(a)\) is indefinite, then \(H_q(c)\) is indefinite.

Proof. Write \(\Phi=(\Phi_1,\ldots,\Phi_n)\), and use subscripts for partial derivatives. The chain rule gives, for each coordinate \(r\), $$ \partial_r q(c)=\sum_{i=1}^n (\partial_i f)(a)\,\partial_r\Phi_i(c). $$ Since \(\nabla f(a)=0\), every summand is zero, so \(\nabla q(c)=0\). Differentiating once more, for coordinates \(r,s\), gives $$ \partial_{rs}q(c) = \sum_{i,j=1}^n(\partial_{ij}f)(a)\,\partial_r\Phi_i(c)\,\partial_s\Phi_j(c) + \sum_{i=1}^n(\partial_i f)(a)\,\partial_{rs}\Phi_i(c). $$ The second sum vanishes because \(\nabla f(a)=0\). The first sum is the \((r,s)\) entry of \(A^T H_f(a)A\). This proves the matrix identity.

For any \(v\in\mathbb{R}^n\), the identity yields $$ v^T H_q(c)v=(Av)^T H_f(a)(Av). $$ Because \(A\) is invertible, \(v\neq0\) if and only if \(Av\neq0\). Thus positivity of the quadratic form for every nonzero vector is preserved in both directions, and the same reasoning applies to negativity. If the form for \(H_f(a)\) takes both a positive and a negative value, choose vectors \(u_+\) and \(u_-\) giving those values. The vectors \(A^{-1}u_+\) and \(A^{-1}u_-\) give the same respective signs for the form of \(H_q(c)\), so that form is indefinite too. \(\square\)

The vanishing-gradient hypothesis is essential to this simple formula. Away from a critical point, the second sum in the coordinate calculation need not vanish; it records the curvature of the coordinate change itself. At a critical point, the second derivative test therefore gives the same definite or indefinite classification in either coordinate system.

Worked Example: A Saddle Remains a Saddle

Let \(f(x,y)=x^2-y^2\), which has a critical point at \((0,0)\) and Hessian $$ H_f(0,0)= \begin{pmatrix} 2&0\\ 0&-2 \end{pmatrix}. $$ Use the coordinate map \(\Phi(u,v)=(u+v^2,v)\). Its inverse is \(\Phi^{-1}(x,y)=(x-y^2,y)\), so it is a smooth change of coordinates; also \(D\Phi(0,0)\) is the identity matrix.

The function in the new coordinates is $$ q(u,v)=f(\Phi(u,v))=(u+v^2)^2-v^2 =u^2+2uv^2+v^4-v^2. $$ Its Hessian at the origin is $$ H_q(0,0)= \begin{pmatrix} 2&0\\ 0&-2 \end{pmatrix}. $$ The transformed function has both positive and negative second-order directions, just as the original does. For instance, \(q(u,0)=u^2>0\) for \(u\neq0\), while \(q(0,v)=v^4-v^2<0\) whenever \(0<|v|<1\). Thus the origin remains a saddle in the new coordinates.

Worked Example: A Strict Minimum in New Coordinates

Let \(f(x,y)=x^2+y^2\), and take the invertible linear coordinate map $$ \Phi(u,v)=(2u+v,\ u+v). $$ Its derivative matrix is \(A=\begin{pmatrix}2&1\\1&1\end{pmatrix}\), whose determinant is \(1\), so it is invertible. The origin is a critical point of \(f\), and $$ q(u,v)=f(\Phi(u,v))=(2u+v)^2+(u+v)^2 =5u^2+6uv+2v^2. $$ The Hessian is $$ H_q(0,0)= \begin{pmatrix} 10&6\\ 6&4 \end{pmatrix} =A^T \begin{pmatrix} 2&0\\ 0&2 \end{pmatrix} A. $$ For every nonzero \((u,v)\), \(q(u,v)=(2u+v)^2+(u+v)^2>0\), because the invertibility of \(A\) means its two displayed linear coordinates cannot both vanish unless \((u,v)=(0,0)\). The strict minimum is therefore preserved, even though the Hessian matrix has changed.

Using the Applications Carefully

These consequences answer different questions. The local inverse solves equations and quantifies first-order sensitivity. The open-images result concerns the image of an entire open set made of regular points. The Hessian formula concerns a scalar function at a critical point under a smooth change of variables. Keeping those conclusions distinct prevents several common overstatements.

  • A local inverse gives uniqueness only within the neighborhoods supplied by the theorem; it does not establish global injectivity.
  • Openness of \(f(O)\) follows when every point of \(O\) is regular. A single regular point gives a neighborhood with open image, not a conclusion about an arbitrary larger set.
  • The Hessian transformation \(H_q=A^TH_fA\) requires a critical point. Without it, the second derivatives of the coordinate map contribute additional terms.

The shared idea is local control. An invertible derivative makes nearby behavior stable enough to transfer conclusions between the original map, its inverse, and a new coordinate description.

Check Your Understanding

Use the local inverse and coordinate-change results to answer the following questions.

  1. If \(Df(a)\) is invertible and \(g\) is the local inverse near \(b=f(a)\), what linear map describes the first-order change in \(g(b+k)\)?
  2. Why does the open-images theorem require every point of \(O\) to be regular, rather than just one point?
  3. In the Hessian coordinate-change formula, where does the invertibility of \(D\Phi(c)\) enter the proof that definiteness is preserved?
  4. What extra term appears in the second-derivative chain rule, and why does it vanish at a critical point?
  5. Does a local inverse imply that the original map is one-to-one on its entire domain? Explain.