Why an Invertible Derivative Matters
The previous tutorial examined how an invertible derivative transforms directions under a change of coordinates. That discussion used invertibility to compare Hessians; it did not establish that the coordinate map itself has an inverse nearby. The Inverse Function Theorem addresses this separate question. It says, under suitable hypotheses, that if the derivative at a point is an invertible linear map, then the nonlinear function can be inverted on sufficiently small neighborhoods of that point and its image.
The word local is essential. A function may fail to be one-to-one on its entire domain while still being one-to-one near a particular point. Conversely, an invertible derivative is a pointwise piece of information; to obtain an inverse on a neighborhood, we must also control how the derivative behaves nearby. We will develop that distinction before stating the full theorem in the next tutorial.
Local Inverses and the Derivative of an Inverse
This definition requires a bijection between neighborhoods, not necessarily between the entire original domain and range. It also does not, by itself, say that the inverse \(g\) is differentiable. The next theorem identifies what its derivative must be if it is differentiable.
Proof. Since \(g(f(x))=x\) for all \(x\in U_0\), the composition \(g\circ f\) is the identity map on \(U_0\). The Multivariable Chain Rule, applied at \(a\), gives $$ D(g\circ f)(a)=Dg(f(a))\,Df(a). $$ The derivative of the identity map is the identity linear map, so \(Dg(f(a))Df(a)=I\). In finite-dimensional linear algebra, a square matrix with a left inverse is invertible: its kernel is \(\{0\}\), since \(Df(a)v=0\) implies \(v=Dg(f(a))Df(a)v=0\); an injective linear map from \(\mathbb{R}^n\) to itself is also surjective. Thus \(Df(a)\) is invertible, and the left inverse \(Dg(f(a))\) must equal its inverse. \(\square\)
This result explains why invertibility of the derivative is necessary when both a function and its local inverse are differentiable. The Inverse Function Theorem gives a useful converse: an invertible derivative, together with the theorem’s regularity assumptions, guarantees that a differentiable local inverse exists. A proof of that existence needs more than the chain rule.
A Quantitative Test for Local Injectivity
A first step toward local invertibility is proving local injectivity. The following estimate makes precise how a function whose derivative stays close to an invertible linear map cannot send two nearby points to the same value. For a linear map \(A\), its operator norm is denoted by \(\|A\|_{\mathrm{op}}\). If \(A\) is invertible, then $$ \|Av\|_2\geq \frac{1}{\|A^{-1}\|_{\mathrm{op}}}\|v\|_2 $$ for every \(v\), because \(\|v\|_2=\|A^{-1}Av\|_2\leq\|A^{-1}\|_{\mathrm{op}}\|Av\|_2\).
Proof. Define \(h:U\to\mathbb{R}^n\) by \(h(x)=f(x)-Ax\). Its derivative satisfies \(Dh(x)=Df(x)-A\), so \(\|Dh(x)\|_{\mathrm{op}}\leq c\) throughout \(U\). Since \(U\) is convex, the segment joining any \(x,y\in U\) lies in \(U\). The Corollary “Derivative Bound Gives a Lipschitz Bound,” which follows from the Mean Value Estimate Along a Segment, therefore gives $$ \|h(x)-h(y)\|_2\leq c\|x-y\|_2. $$ Also, the operator-norm inequality for \(A^{-1}\) yields $$ \|A(x-y)\|_2\geq \frac{1}{\|A^{-1}\|_{\mathrm{op}}}\|x-y\|_2. $$ Because \(f(x)-f(y)=A(x-y)+h(x)-h(y)\), the reverse triangle inequality gives $$ \begin{aligned} \|f(x)-f(y)\|_2 &\geq \|A(x-y)\|_2-\|h(x)-h(y)\|_2\\ &\geq \left(\frac{1}{\|A^{-1}\|_{\mathrm{op}}}-c\right)\|x-y\|_2. \end{aligned} $$ The coefficient is strictly positive by hypothesis. If \(f(x)=f(y)\), the left side is zero, so the inequality forces \(\|x-y\|_2=0\), hence \(x=y\). This proves injectivity. \(\square\)
The estimate also shows why a bound on the derivative throughout a neighborhood is useful. If \(Df\) is continuous at \(a\) and \(A=Df(a)\) is invertible, then sufficiently close to \(a\), \(Df(x)\) stays as close to \(A\) as desired. A small convex ball around \(a\) can then satisfy the theorem’s derivative bound, giving injectivity on that ball. This is one important ingredient in proving local invertibility.
Worked Examples
Worked Example: An Invertible Linear Map
Consider \(f:\mathbb{R}^2\to\mathbb{R}^2\) given by \(f(x,y)=(x+2y,3x-y)\). Its derivative is the constant matrix $$ Df(x,y)= \begin{pmatrix}1&2\\3&-1\end{pmatrix}, $$ whose determinant is \(-7\), so it is invertible. To solve \(u=x+2y\) and \(v=3x-y\), multiply the first equation by \(1\) and the second by \(2\) and add to obtain \(u+2v=7x\). Then \(x=(u+2v)/7\); substituting into \(u=x+2y\) gives \(y=(3u-v)/7\). Thus $$ f^{-1}(u,v)=\left(\frac{u+2v}{7},\frac{3u-v}{7}\right). $$ For verification, substituting these expressions into the first component gives \((u+2v)/7+2(3u-v)/7=u\), and the second gives \(3(u+2v)/7-(3u-v)/7=v\). Here the inverse exists globally because the map is linear and its matrix is invertible. The Inverse Function Theorem is valuable because it obtains a corresponding local conclusion for nonlinear maps.
Worked Example: A Nonlinear Map with an Explicit Local Inverse
Let \(f:\mathbb{R}^2\to\mathbb{R}^2\) be \(f(u,v)=(u,v+u^2)\). Its derivative matrix is $$ Df(u,v)= \begin{pmatrix}1&0\\2u&1\end{pmatrix}, $$ with determinant \(1\) at every point. Given \((x,y)=f(u,v)\), the first coordinate gives \(u=x\), and the second gives \(v=y-x^2\). Hence the inverse is \(g(x,y)=(x,y-x^2)\). Direct substitution verifies both compositions: $$ g(f(u,v))=(u,v+u^2-u^2)=(u,v), \qquad f(g(x,y))=(x,y-x^2+x^2)=(x,y). $$ The derivative formula is also confirmed by calculation: $$ Dg(x,y)= \begin{pmatrix}1&0\\-2x&1\end{pmatrix}, \qquad Dg(f(u,v))Df(u,v) = \begin{pmatrix}1&0\\-2u&1\end{pmatrix} \begin{pmatrix}1&0\\2u&1\end{pmatrix} =I. $$ This example has a global inverse, but the theorem being developed is local and does not require global behavior to be so simple.
Worked Example: Injectivity Does Not Guarantee a Differentiable Inverse
The function \(f:\mathbb{R}\to\mathbb{R}\) defined by \(f(t)=t^3\) is injective: if \(s<t\), then $$ t^3-s^3=(t-s)(t^2+ts+s^2)>0. $$ The second factor is positive for \(s<t\), since it equals \((t+s/2)^2+3s^2/4\) and can be zero only when \(s=t=0\), which is incompatible with \(s<t\). The inverse is \(g(x)=\sqrt[3]{x}\). At \(0\), its difference quotient is $$ \frac{g(x)-g(0)}{x-0}=\frac{\sqrt[3]{x}}{x}=\frac{1}{|x|^{2/3}}\quad (x\ne0), $$ which is unbounded as \(x\to0\). Thus \(g\) is not differentiable at \(0\), and indeed \(f'(0)=0\) is not invertible. This example does not contradict the Inverse Function Theorem: it illustrates why the derivative condition matters for a differentiable inverse.
What the Derivative Condition Does—and Does Not—Say
The quantitative injectivity theorem should not be mistaken for the full Inverse Function Theorem. It establishes one-to-one behavior under a derivative bound, but injectivity alone does not show that the image contains a neighborhood of \(f(a)\). Nor does it, by itself, show that the inverse is differentiable. The full theorem supplies these additional conclusions under appropriate hypotheses, including continuity of the derivative near the point.
There are two useful cautions. First, global injectivity is stronger than the conclusion usually needed: local inversion concerns suitably small neighborhoods. Second, a nonzero or invertible derivative is not a statement about the entire domain. For instance, the map \(f(t)=e^t\) has nonzero derivative everywhere but its image is only \((0,\infty)\), not all of \(\mathbb{R}\). It is locally invertible, and its inverse on its image is \(\log x\); it is not a bijection from \(\mathbb{R}\) onto \(\mathbb{R}\).
The main proof strategy is now visible. Near a point with invertible derivative, compare the function with its invertible linear approximation. Continuity of the derivative keeps the difference in derivatives small nearby, and the quantitative estimate prevents distinct nearby points from having the same image. Establishing that the image contains an open neighborhood and that the inverse is differentiable requires the further argument developed alongside the theorem’s formal statement.
Check Your Understanding
Use the definitions, derivative formula, and quantitative estimate to answer the following questions.
- In the definition of a local inverse, which two sets must be related bijectively, and where must the point \(f(a)\) lie?
- Why does differentiability of both \(f\) and its local inverse imply that \(Df(a)\) is invertible?
- In the derivative-closeness theorem, why must \(c\) be strictly less than \(1/\|A^{-1}\|_{\mathrm{op}}\)?
- Which earlier result turns a bound on \(\|Dh(x)\|_{\mathrm{op}}\) into a bound on \(\|h(x)-h(y)\|_2\)?
- Why does injectivity alone not establish that a function has a differentiable local inverse?
- What feature of \(t\mapsto t^3\) at zero shows the relevance of the invertible-derivative condition?