Tutorials › Real Analysis › Inverse Function Theorem

Multivariable Analysis · Tutorial 801 of 1000

Inverse Function Theorem

Learn how invertible derivatives control local behavior, what they do and do not establish, and how to differentiate an inverse once it exists.

Advanced 10 min read

What You'll Learn

  • Distinguish local invertibility from global invertibility
  • Use a quantitative derivative estimate to prove local injectivity
  • Explain why an invertible derivative alone does not yet construct a local inverse
  • Derive the derivative formula for a differentiable inverse
  • Recognize examples where injectivity and differentiability of the inverse behave differently

Why an Invertible Derivative Matters

The previous tutorial examined how an invertible derivative transforms directions under a change of coordinates. That discussion used invertibility to compare Hessians; it did not establish that the coordinate map itself has an inverse nearby. The Inverse Function Theorem addresses this separate question. It says, under suitable hypotheses, that if the derivative at a point is an invertible linear map, then the nonlinear function can be inverted on sufficiently small neighborhoods of that point and its image.

The word local is essential. A function may fail to be one-to-one on its entire domain while still being one-to-one near a particular point. Conversely, an invertible derivative is a pointwise piece of information; to obtain an inverse on a neighborhood, we must also control how the derivative behaves nearby. We will develop that distinction before stating the full theorem in the next tutorial.

Local Inverses and the Derivative of an Inverse

Definition: Let \(U\subseteq\mathbb{R}^n\) be open, let \(f:U\to\mathbb{R}^n\), and let \(a\in U\). We say that \(f\) has a local inverse at \(a\) if there are open sets \(U_0\subseteq U\) and \(V_0\subseteq\mathbb{R}^n\), with \(a\in U_0\) and \(f(a)\in V_0\), such that \(f\) maps \(U_0\) bijectively onto \(V_0\). The inverse map \(g:V_0\to U_0\) is defined by \(g(f(x))=x\) for every \(x\in U_0\).

This definition requires a bijection between neighborhoods, not necessarily between the entire original domain and range. It also does not, by itself, say that the inverse \(g\) is differentiable. The next theorem identifies what its derivative must be if it is differentiable.

Theorem (Derivative of a Differentiable Local Inverse): Suppose \(f:U_0\to V_0\) is a bijection between open subsets of \(\mathbb{R}^n\), \(g=f^{-1}:V_0\to U_0\), and both \(f\) and \(g\) are differentiable. For every \(a\in U_0\), $$ Dg(f(a))\,Df(a)=I. $$ Consequently, \(Df(a)\) is invertible and $$ Dg(f(a))=\bigl(Df(a)\bigr)^{-1}. $$

Proof. Since \(g(f(x))=x\) for all \(x\in U_0\), the composition \(g\circ f\) is the identity map on \(U_0\). The Multivariable Chain Rule, applied at \(a\), gives $$ D(g\circ f)(a)=Dg(f(a))\,Df(a). $$ The derivative of the identity map is the identity linear map, so \(Dg(f(a))Df(a)=I\). In finite-dimensional linear algebra, a square matrix with a left inverse is invertible: its kernel is \(\{0\}\), since \(Df(a)v=0\) implies \(v=Dg(f(a))Df(a)v=0\); an injective linear map from \(\mathbb{R}^n\) to itself is also surjective. Thus \(Df(a)\) is invertible, and the left inverse \(Dg(f(a))\) must equal its inverse. \(\square\)

This result explains why invertibility of the derivative is necessary when both a function and its local inverse are differentiable. The Inverse Function Theorem gives a useful converse: an invertible derivative, together with the theorem’s regularity assumptions, guarantees that a differentiable local inverse exists. A proof of that existence needs more than the chain rule.

A Quantitative Test for Local Injectivity

A first step toward local invertibility is proving local injectivity. The following estimate makes precise how a function whose derivative stays close to an invertible linear map cannot send two nearby points to the same value. For a linear map \(A\), its operator norm is denoted by \(\|A\|_{\mathrm{op}}\). If \(A\) is invertible, then $$ \|Av\|_2\geq \frac{1}{\|A^{-1}\|_{\mathrm{op}}}\|v\|_2 $$ for every \(v\), because \(\|v\|_2=\|A^{-1}Av\|_2\leq\|A^{-1}\|_{\mathrm{op}}\|Av\|_2\).

Theorem (Derivative Closeness Gives Injectivity): Let \(U\subseteq\mathbb{R}^n\) be open and convex, and let \(f:U\to\mathbb{R}^n\) be differentiable at every point of \(U\). Suppose \(A:\mathbb{R}^n\to\mathbb{R}^n\) is invertible and, for some constant \(c\) with \(0\leq c<1/\|A^{-1}\|_{\mathrm{op}}\), $$ \|Df(x)-A\|_{\mathrm{op}}\leq c $$ for every \(x\in U\). Then for all \(x,y\in U\), $$ \|f(x)-f(y)\|_2\geq \left(\frac{1}{\|A^{-1}\|_{\mathrm{op}}}-c\right)\|x-y\|_2. $$ In particular, \(f\) is injective on \(U\).

Proof. Define \(h:U\to\mathbb{R}^n\) by \(h(x)=f(x)-Ax\). Its derivative satisfies \(Dh(x)=Df(x)-A\), so \(\|Dh(x)\|_{\mathrm{op}}\leq c\) throughout \(U\). Since \(U\) is convex, the segment joining any \(x,y\in U\) lies in \(U\). The Corollary “Derivative Bound Gives a Lipschitz Bound,” which follows from the Mean Value Estimate Along a Segment, therefore gives $$ \|h(x)-h(y)\|_2\leq c\|x-y\|_2. $$ Also, the operator-norm inequality for \(A^{-1}\) yields $$ \|A(x-y)\|_2\geq \frac{1}{\|A^{-1}\|_{\mathrm{op}}}\|x-y\|_2. $$ Because \(f(x)-f(y)=A(x-y)+h(x)-h(y)\), the reverse triangle inequality gives $$ \begin{aligned} \|f(x)-f(y)\|_2 &\geq \|A(x-y)\|_2-\|h(x)-h(y)\|_2\\ &\geq \left(\frac{1}{\|A^{-1}\|_{\mathrm{op}}}-c\right)\|x-y\|_2. \end{aligned} $$ The coefficient is strictly positive by hypothesis. If \(f(x)=f(y)\), the left side is zero, so the inequality forces \(\|x-y\|_2=0\), hence \(x=y\). This proves injectivity. \(\square\)

The estimate also shows why a bound on the derivative throughout a neighborhood is useful. If \(Df\) is continuous at \(a\) and \(A=Df(a)\) is invertible, then sufficiently close to \(a\), \(Df(x)\) stays as close to \(A\) as desired. A small convex ball around \(a\) can then satisfy the theorem’s derivative bound, giving injectivity on that ball. This is one important ingredient in proving local invertibility.

Worked Examples

Worked Example: An Invertible Linear Map

Consider \(f:\mathbb{R}^2\to\mathbb{R}^2\) given by \(f(x,y)=(x+2y,3x-y)\). Its derivative is the constant matrix $$ Df(x,y)= \begin{pmatrix}1&2\\3&-1\end{pmatrix}, $$ whose determinant is \(-7\), so it is invertible. To solve \(u=x+2y\) and \(v=3x-y\), multiply the first equation by \(1\) and the second by \(2\) and add to obtain \(u+2v=7x\). Then \(x=(u+2v)/7\); substituting into \(u=x+2y\) gives \(y=(3u-v)/7\). Thus $$ f^{-1}(u,v)=\left(\frac{u+2v}{7},\frac{3u-v}{7}\right). $$ For verification, substituting these expressions into the first component gives \((u+2v)/7+2(3u-v)/7=u\), and the second gives \(3(u+2v)/7-(3u-v)/7=v\). Here the inverse exists globally because the map is linear and its matrix is invertible. The Inverse Function Theorem is valuable because it obtains a corresponding local conclusion for nonlinear maps.

Worked Example: A Nonlinear Map with an Explicit Local Inverse

Let \(f:\mathbb{R}^2\to\mathbb{R}^2\) be \(f(u,v)=(u,v+u^2)\). Its derivative matrix is $$ Df(u,v)= \begin{pmatrix}1&0\\2u&1\end{pmatrix}, $$ with determinant \(1\) at every point. Given \((x,y)=f(u,v)\), the first coordinate gives \(u=x\), and the second gives \(v=y-x^2\). Hence the inverse is \(g(x,y)=(x,y-x^2)\). Direct substitution verifies both compositions: $$ g(f(u,v))=(u,v+u^2-u^2)=(u,v), \qquad f(g(x,y))=(x,y-x^2+x^2)=(x,y). $$ The derivative formula is also confirmed by calculation: $$ Dg(x,y)= \begin{pmatrix}1&0\\-2x&1\end{pmatrix}, \qquad Dg(f(u,v))Df(u,v) = \begin{pmatrix}1&0\\-2u&1\end{pmatrix} \begin{pmatrix}1&0\\2u&1\end{pmatrix} =I. $$ This example has a global inverse, but the theorem being developed is local and does not require global behavior to be so simple.

Worked Example: Injectivity Does Not Guarantee a Differentiable Inverse

The function \(f:\mathbb{R}\to\mathbb{R}\) defined by \(f(t)=t^3\) is injective: if \(s<t\), then $$ t^3-s^3=(t-s)(t^2+ts+s^2)>0. $$ The second factor is positive for \(s<t\), since it equals \((t+s/2)^2+3s^2/4\) and can be zero only when \(s=t=0\), which is incompatible with \(s<t\). The inverse is \(g(x)=\sqrt[3]{x}\). At \(0\), its difference quotient is $$ \frac{g(x)-g(0)}{x-0}=\frac{\sqrt[3]{x}}{x}=\frac{1}{|x|^{2/3}}\quad (x\ne0), $$ which is unbounded as \(x\to0\). Thus \(g\) is not differentiable at \(0\), and indeed \(f'(0)=0\) is not invertible. This example does not contradict the Inverse Function Theorem: it illustrates why the derivative condition matters for a differentiable inverse.

What the Derivative Condition Does—and Does Not—Say

The quantitative injectivity theorem should not be mistaken for the full Inverse Function Theorem. It establishes one-to-one behavior under a derivative bound, but injectivity alone does not show that the image contains a neighborhood of \(f(a)\). Nor does it, by itself, show that the inverse is differentiable. The full theorem supplies these additional conclusions under appropriate hypotheses, including continuity of the derivative near the point.

There are two useful cautions. First, global injectivity is stronger than the conclusion usually needed: local inversion concerns suitably small neighborhoods. Second, a nonzero or invertible derivative is not a statement about the entire domain. For instance, the map \(f(t)=e^t\) has nonzero derivative everywhere but its image is only \((0,\infty)\), not all of \(\mathbb{R}\). It is locally invertible, and its inverse on its image is \(\log x\); it is not a bijection from \(\mathbb{R}\) onto \(\mathbb{R}\).

The main proof strategy is now visible. Near a point with invertible derivative, compare the function with its invertible linear approximation. Continuity of the derivative keeps the difference in derivatives small nearby, and the quantitative estimate prevents distinct nearby points from having the same image. Establishing that the image contains an open neighborhood and that the inverse is differentiable requires the further argument developed alongside the theorem’s formal statement.

Check Your Understanding

Use the definitions, derivative formula, and quantitative estimate to answer the following questions.

  1. In the definition of a local inverse, which two sets must be related bijectively, and where must the point \(f(a)\) lie?
  2. Why does differentiability of both \(f\) and its local inverse imply that \(Df(a)\) is invertible?
  3. In the derivative-closeness theorem, why must \(c\) be strictly less than \(1/\|A^{-1}\|_{\mathrm{op}}\)?
  4. Which earlier result turns a bound on \(\|Dh(x)\|_{\mathrm{op}}\) into a bound on \(\|h(x)-h(y)\|_2\)?
  5. Why does injectivity alone not establish that a function has a differentiable local inverse?
  6. What feature of \(t\mapsto t^3\) at zero shows the relevance of the invertible-derivative condition?