Tutorials › Real Analysis › Statement of the Inverse Function Theorem

Multivariable Analysis · Tutorial 802 of 1000

Statement of the Inverse Function Theorem

Learn the precise local hypotheses and conclusions of the Inverse Function Theorem, and see how they guarantee unique, smoothly varying solutions near a point.

Advanced 9 min read

What You'll Learn

  • State the Inverse Function Theorem with its open-set, differentiability, and invertible-derivative hypotheses
  • Distinguish a local inverse from a global inverse
  • Read the derivative formula for the inverse at nearby points
  • Use the theorem to obtain unique nearby solutions of nonlinear equations
  • Derive a local Lipschitz bound for the inverse
  • Recognize what the theorem does not conclude when its hypotheses fail

The Local Conclusion We Need

The previous tutorial identified local injectivity as an important step toward local inversion. It also explained why comparing a function with its invertible linear approximation can prevent distinct nearby points from having the same image. Injectivity, however, is only part of the desired conclusion: we also want the image to contain a neighborhood of the output point, and we want the inverse to be differentiable.

The Inverse Function Theorem supplies all of these conclusions at once. Its hypotheses are local: the function must be continuously differentiable near the point, and its derivative at that point must be an invertible linear map. The theorem does not require the function to be one-to-one throughout its original domain.

Statement of the Inverse Function Theorem

Theorem (Inverse Function Theorem): Let \(U\subseteq\mathbb{R}^n\) be open, let \(f:U\to\mathbb{R}^n\) be continuously differentiable, and let \(a\in U\). Suppose \(Df(a)\) is invertible. Then there are open sets \(U_0\subseteq U\) and \(V_0\subseteq\mathbb{R}^n\), with \(a\in U_0\) and \(f(a)\in V_0\), such that \(f\) maps \(U_0\) bijectively onto \(V_0\). The inverse \(g=f^{-1}:V_0\to U_0\) is continuously differentiable, and for every \(y\in V_0\), $$ Dg(y)=\bigl(Df(g(y))\bigr)^{-1}. $$ In particular, $$ Dg(f(a))=\bigl(Df(a)\bigr)^{-1}. $$

For a map \(f:\mathbb{R}^n\to\mathbb{R}^n\), invertibility of \(Df(a)\) is equivalent to the nonvanishing of its Jacobian determinant, \(\det Df(a)\neq 0\). The theorem therefore says that a continuously differentiable map with a nonzero Jacobian determinant at one point has a differentiable inverse after restricting to suitable neighborhoods.

The sets \(U_0\) and \(V_0\) are not prescribed in advance. The theorem guarantees that some such neighborhoods exist. In particular, it does not claim that the original domain \(U\) maps bijectively onto its entire image, or that the image \(f(U)\) is all of \(\mathbb{R}^n\). The local neighborhoods are part of the conclusion.

What the Hypotheses and Conclusion Say

1
Begin with an open domain.
The point \(a\) must lie in an open set \(U\). This allows the theorem to study the function on a full neighborhood around \(a\), rather than only along a boundary or a lower-dimensional subset.
2
Require continuous differentiability.
The derivative must exist and vary continuously near the point. Differentiability at \(a\) alone is not the hypothesis in this version of the theorem.
3
Check the derivative at the base point.
The linear map \(Df(a):\mathbb{R}^n\to\mathbb{R}^n\) must be invertible. In matrix language, the Jacobian matrix must have nonzero determinant.
4
Restrict to neighborhoods.
The conclusion gives open neighborhoods \(U_0\) and \(V_0\) on which \(f\) is a bijection, with a continuously differentiable inverse.

The derivative formula is consistent with the chain rule. Since \(g(f(x))=x\) on \(U_0\), differentiating gives \(Dg(f(x))Df(x)=I\). The earlier result “Derivative of a Differentiable Local Inverse” explains why the derivative of the inverse must be the inverse linear map. The theorem goes further: it guarantees that the local inverse exists and is continuously differentiable in the first place.

The pointwise formula at \(f(a)\) is only one instance of the full formula. At any nearby output \(y\in V_0\), the corresponding input is \(g(y)\), so the derivative to invert is \(Df(g(y))\), not necessarily \(Df(a)\). This distinction matters when computing how the inverse changes away from the base point.

Consequences for Nearby Equations

Corollary (Local Existence and Uniqueness for Nearby Target Values): Under the hypotheses of the Inverse Function Theorem, write \(b=f(a)\). There are neighborhoods \(U_0\) of \(a\) and \(V_0\) of \(b\) such that, for every \(y\in V_0\), the equation $$ f(x)=y $$ has exactly one solution \(x\in U_0\). That solution is \(x=g(y)\), and it depends continuously differentiably on \(y\).

Proof. By the Inverse Function Theorem, the restriction \(f:U_0\to V_0\) is a bijection and has inverse \(g:V_0\to U_0\). Surjectivity means that for every \(y\in V_0\), there is an \(x\in U_0\) with \(f(x)=y\). Injectivity means that there cannot be two distinct such \(x\). By the definition of the inverse, the unique solution is \(x=g(y)\). The theorem also gives that \(g\) is continuously differentiable, proving the stated dependence on \(y\). \(\square\)

This corollary is often the practical form of the theorem. An equation may be difficult to solve explicitly, but the theorem can guarantee that for every target value sufficiently close to \(b\), there is exactly one solution near \(a\). It does not rule out other solutions far from \(a\); uniqueness is asserted only inside \(U_0\).

Worked Examples

Worked Example: A Nonlinear Map Near the Origin

Consider \(f:\mathbb{R}^2\to\mathbb{R}^2\) defined by $$ f(x,y)=(x+y+x^2,\;x-y). $$ Its Jacobian matrix is $$ Df(x,y)= \begin{pmatrix} 1+2x&1\\ 1&-1 \end{pmatrix}. $$ At the origin, $$ Df(0,0)= \begin{pmatrix} 1&1\\ 1&-1 \end{pmatrix}, \qquad \det Df(0,0)=(1)(-1)-(1)(1)=-2. $$ The determinant is nonzero, and \(f\) is continuously differentiable on the open set \(\mathbb{R}^2\). The Inverse Function Theorem therefore gives open neighborhoods \(U_0\) of \((0,0)\) and \(V_0\) of \(f(0,0)=(0,0)\) such that \(f:U_0\to V_0\) is bijective with a continuously differentiable inverse.

For every \((u,v)\in V_0\), the system $$ u=x+y+x^2,\qquad v=x-y $$ has exactly one solution \((x,y)\in U_0\). The theorem also gives the derivative of the inverse at the origin: $$ Dg(0,0)=\bigl(Df(0,0)\bigr)^{-1} = \begin{pmatrix} \frac12&\frac12\\ \frac12&-\frac12 \end{pmatrix}. $$ Indeed, multiplying the two matrices in either order gives the identity matrix. For example, the first row of \(Df(0,0)Dg(0,0)\) is \((1,1)\) times the displayed inverse, giving \((1,0)\); its second row is \((1,-1)\) times the inverse, giving \((0,1)\). The theorem guarantees a local inverse even though the nonlinear equations have not been solved explicitly.

Worked Example: A Map with a Periodic Global Obstruction

Define \(F:\mathbb{R}^2\to\mathbb{R}^2\) by $$ F(x,y)=(e^x\cos y,\;e^x\sin y). $$ The Jacobian matrix is $$ DF(x,y)= \begin{pmatrix} e^x\cos y&-e^x\sin y\\ e^x\sin y&e^x\cos y \end{pmatrix}, $$ so $$ \det DF(x,y)=e^{2x}\cos^2 y+e^{2x}\sin^2 y=e^{2x}>0. $$ Thus \(F\) is continuously differentiable and has an invertible derivative at every point. In particular, at \((0,0)\), \(F(0,0)=(1,0)\), and the Inverse Function Theorem gives a differentiable inverse between suitable neighborhoods of these points.

Nevertheless, \(F\) is not globally injective. For every \((x,y)\), $$ F(x,y+2\pi)= \bigl(e^x\cos(y+2\pi),e^x\sin(y+2\pi)\bigr) =(e^x\cos y,e^x\sin y)=F(x,y). $$ For example, \((0,0)\neq(0,2\pi)\), but both map to \((1,0)\). There is no conflict with the theorem: it guarantees bijectivity only after restricting the domain to a sufficiently small neighborhood of a chosen point.

Worked Example: A Vanishing Derivative Does Not Settle Local Invertibility

Let \(q:\mathbb{R}\to\mathbb{R}\) be \(q(t)=t^3\). It is continuously differentiable, but $$ q'(0)=3(0)^2=0. $$ The Inverse Function Theorem does not apply at \(0\), because the derivative there is not invertible as a linear map from \(\mathbb{R}\) to itself. This failure of the hypothesis means the theorem gives no conclusion; it does not, by itself, prove that a local inverse is impossible.

In this particular example, \(q\) is strictly increasing: if \(s<t\), then $$ t^3-s^3=(t-s)(t^2+ts+s^2)>0. $$ To verify positivity of the second factor, write $$ t^2+ts+s^2=(t+s/2)^2+3s^2/4. $$ It is nonnegative and can equal zero only if \(s=t=0\), which contradicts \(s<t\). Thus \(q\) is injective, and its inverse is \(g(x)=\sqrt[3]{x}\). However, at zero, $$ \frac{g(x)-g(0)}{x-0}=\frac{\sqrt[3]{x}}{x} =\frac{1}{|x|^{2/3}}\qquad (x\neq 0). $$ This quotient is unbounded as \(x\to0\), so \(g\) is not differentiable at zero. The example shows that a local inverse can exist without the differentiability conclusion supplied by the theorem; the nonzero-derivative hypothesis is what ensures the stronger conclusion.

Local Stability of the Inverse

The theorem says more than that nearby target values have unique nearby solutions. Because the inverse is continuously differentiable, small changes in the target produce controlled changes in the solution. The following consequence makes that control quantitative on a sufficiently small ball.

Corollary (A Local Lipschitz Bound for the Inverse): Under the hypotheses of the Inverse Function Theorem, let \(g:V_0\to U_0\) be the local inverse and let \(b=f(a)\). There are \(r>0\) and \(M>0\) such that the ball \(B_r^{(2)}(b)\) lies in \(V_0\) and $$ \|g(y)-g(z)\|_2\leq M\|y-z\|_2 $$ for all \(y,z\in B_r^{(2)}(b)\).

Proof. Since \(V_0\) is open and contains \(b\), choose \(r_0>0\) such that the closed ball \(\overline{B_{r_0}^{(2)}(b)}\) is contained in \(V_0\). The derivative \(Dg\) is continuous, so it is bounded on this closed ball: it is compact by the Heine–Borel Theorem, and a continuous real-valued function such as \(y\mapsto\|Dg(y)\|_{\mathrm{op}}\) attains a finite maximum there. Let that maximum be \(M\), increasing \(M\) to a positive number if necessary.

Take \(r=r_0\), or any smaller positive radius. The ball \(B_r^{(2)}(b)\) is convex, and the segment joining any \(y,z\) in it lies inside it. At every point \(w\) on that segment, \(\|Dg(w)\|_{\mathrm{op}}\leq M\). The Corollary “Derivative Bound Gives a Lipschitz Bound,” which follows from the Mean Value Estimate Along a Segment, now gives $$ \|g(y)-g(z)\|_2\leq M\|y-z\|_2. $$ This proves the claimed local bound. \(\square\)

The bound describes stability: within the chosen target neighborhood, a change of size \(\|y-z\|_2\) in the output value changes the corresponding solution by at most \(M\|y-z\|_2\). The constant \(M\) may depend on the neighborhoods; the theorem does not provide one universal bound for every target value in the original image.

Reading the Theorem Carefully

A common mistake is to treat a nonzero determinant at one point as a global guarantee. The determinant condition is checked at \(a\), while the conclusion concerns carefully chosen neighborhoods around \(a\) and \(f(a)\). The theorem provides neither a global inverse nor a statement that the inverse is defined for every point in the original codomain.

A second mistake is to omit continuous differentiability. In the stated theorem, the derivative must be continuous on the open domain, not merely exist at \(a\). There are other versions of inverse-function results with different assumptions, but they are not the statement used here.

Finally, when applying the derivative formula at a general \(y\), first identify its preimage \(g(y)\). The correct formula is \(Dg(y)=[Df(g(y))]^{-1}\). Substituting \(Df(a)\) at every \(y\) would be valid only in special cases, such as when the derivative of \(f\) is constant. The next tutorial develops the proof architecture that establishes the theorem’s local existence, bijectivity, and regularity conclusions.

Check Your Understanding

Use the theorem’s hypotheses and conclusions to answer the following questions.

  1. What must be true of \(Df(a)\) for the Inverse Function Theorem to apply?
  2. Why does the theorem guarantee a local inverse without asserting that \(f\) is globally injective?
  3. For a target \(y\) near \(f(a)\), at which input point is \(Df\) evaluated in the formula for \(Dg(y)\)?
  4. What existence and uniqueness statement does the theorem give for the equation \(f(x)=y\) when \(y\) is sufficiently close to \(f(a)\)?
  5. In the periodic-map example, why does failure of global injectivity not contradict the theorem?
  6. If \(Df(a)\) is not invertible, what can and cannot be concluded from the theorem’s hypotheses?