Tutorials › Real Analysis › Proof Architecture for the Inverse Function Theorem

Multivariable Analysis · Tutorial 803 of 1000

Proof Architecture for the Inverse Function Theorem

Follow the proof from a normalized contraction to a continuously differentiable local inverse, with each neighborhood and regularity step justified.

Advanced 12 min read

What You'll Learn

  • Normalize the function so its derivative at the base point is the identity
  • Use derivative control to make the nonlinear remainder a contraction
  • Construct a local inverse by solving a fixed-point equation
  • Prove the inverse is Lipschitz before establishing its differentiability
  • Derive the inverse derivative formula and continuity of that derivative

From a Linear Approximation to an Actual Inverse

The Inverse Function Theorem says that if \(Df(a)\) is invertible, then \(f\) has a continuously differentiable inverse after restricting to suitable neighborhoods. The central proof challenge is to turn the approximate statement “\(f\) behaves like its derivative near \(a\)” into an exact solution of \(f(x)=y\) for every nearby target \(y\).

The proof has a useful architecture. First, change coordinates so the base point and its image are both the origin and the derivative becomes the identity. Next, write the normalized function as the identity plus a small remainder. Continuity of the derivative makes that remainder contractive on a sufficiently small ball. A fixed-point argument then gives a unique preimage for each nearby target. Finally, a quantitative estimate proves that the inverse is continuous before its differentiability and derivative formula are established.

Normalize the Map

Let \(U\subseteq\mathbb{R}^n\) be open, let \(f:U\to\mathbb{R}^n\) be continuously differentiable, and suppose \(A=Df(a)\) is invertible. Write \(b=f(a)\). For \(h\) near \(0\), define the normalized map $$ F(h)=A^{-1}\bigl(f(a+h)-b\bigr). $$ Then \(F(0)=0\), and the chain rule gives \(DF(0)=A^{-1}Df(a)=I\). Thus the nonlinear part of \(F\) is the remainder \(R(h)=F(h)-h\), whose derivative vanishes at \(0\).

Lemma (Small Remainder on a Ball): There are \(r>0\) and \(q\) with \(0<q<1\) such that the closed ball \(\overline{B_r^{(2)}(0)}\) lies in the domain of \(F\) and $$ \|R(h)-R(k)\|_2\leq q\|h-k\|_2 $$ for all \(h,k\in\overline{B_r^{(2)}(0)}\). In particular, \(\|R(h)\|_2\leq q\|h\|_2\) on this ball.

Proof. Since \(U\) is open and contains \(a\), the translated domain of \(F\) contains a ball around \(0\). Also \(DR(0)=DF(0)-I=0\). The derivative \(DR\) is continuous, so there is a radius \(r>0\), small enough that the closed ball lies in the domain, for which $$ \|DR(h)\|_{\mathrm{op}}\leq q $$ whenever \(\|h\|_2\leq r\), with, for example, \(q=1/2\). The ball is convex, so the segment joining any \(h\) and \(k\) in it remains in the ball. By the Mean Value Estimate Along a Segment, $$ \|R(h)-R(k)\|_2\leq q\|h-k\|_2. $$ Taking \(k=0\) and using \(R(0)=0\) gives \(\|R(h)\|_2\leq q\|h\|_2\). \(\square\)

This lemma is the quantitative step that makes the proof work. Differentiability at the base point alone gives a remainder that is small compared with \(\|h\|_2\) as \(h\to0\). Continuous differentiability gives the stronger control needed here: throughout a whole sufficiently small ball, the remainder changes by at most a fixed fraction of the change in its input.

Find Preimages by a Fixed-Point Equation

To solve \(F(h)=z\), rearrange the equation as $$ h=z-R(h). $$ For each target \(z\), define \(T_z(h)=z-R(h)\). If \(\|z\|_2\leq(1-q)r\) and \(\|h\|_2\leq r\), then $$ \|T_z(h)\|_2\leq\|z\|_2+\|R(h)\|_2 \leq (1-q)r+qr=r. $$ Thus \(T_z\) maps the closed ball into itself. The small-remainder lemma also gives $$ \|T_z(h)-T_z(k)\|_2=\|R(k)-R(h)\|_2\leq q\|h-k\|_2. $$ It is a contraction, so the Fixed-Point Theorem from “Complete Spaces and Fixed Points” supplies a unique fixed point in the closed ball. A fixed point is exactly a solution to \(F(h)=z\).

Theorem (Local Existence and Uniqueness for the Normalized Map): With \(F\), \(r\), and \(q\) as above, for every \(z\) with \(\|z\|_2<(1-q)r\), there is exactly one \(h\in B_r^{(2)}(0)\) such that \(F(h)=z\).

Proof. The fixed-point argument gives a unique solution in the closed ball whenever \(\|z\|_2\leq(1-q)r\). If \(h\) is that solution, the equation \(h=z-R(h)\) and the remainder bound imply $$ \|h\|_2\leq\|z\|_2+q\|h\|_2, \qquad (1-q)\|h\|_2\leq\|z\|_2. $$ For \(\|z\|_2<(1-q)r\), it follows that \(\|h\|_2<r\), so the solution lies in the open ball. Any solution in the open ball is also in the closed ball and is a fixed point of \(T_z\); uniqueness of the fixed point therefore gives uniqueness there as well. \(\square\)

The strict inequality in the target condition matters: it ensures the fixed point lies in the open ball, not merely on its boundary. To return to the original coordinates, put \(Z=B_{(1-q)r}^{(2)}(0)\), \(V_0=b+AZ\), and $$ U_0=\{a+h:h\in B_r^{(2)}(0),\ F(h)\in Z\}. $$ The set \(V_0\) is open because \(A\) is invertible. The set \(U_0\) is open because \(F\) is continuous, and it contains \(a\). The theorem just proved shows that \(f:U_0\to V_0\) is bijective: \(f(a+h)=b+AF(h)\), and each target in \(V_0\) corresponds to exactly one \(z\in Z\) and exactly one such \(h\).

Prove Continuity Before Differentiability

Existence and uniqueness do not by themselves prove that the inverse is continuous. The key estimate comes from the same remainder bound. For any \(h,k\) in the closed ball, $$ \|F(h)-F(k)\|_2 =\|(h-k)+(R(h)-R(k))\|_2 \geq (1-q)\|h-k\|_2. $$ Here the reverse triangle inequality and the bound on \(R(h)-R(k)\) give the final inequality. Thus the inverse of \(F\), wherever defined on these sets, satisfies $$ \|h-k\|_2\leq \frac{1}{1-q}\|F(h)-F(k)\|_2. $$ This lower bound on the change in \(F\) is also a direct proof of injectivity on the ball.

Let \(g:V_0\to U_0\) be the inverse in the original coordinates. Since \(F(h)=A^{-1}(f(a+h)-b)\), the estimate becomes $$ \|g(w_1)-g(w_2)\|_2 \leq \frac{1}{1-q}\|A^{-1}\|_{\mathrm{op}}\|w_1-w_2\|_2 $$ for \(w_1,w_2\in V_0\). In particular, \(g\) is Lipschitz and therefore continuous. Establishing this continuity explicitly is important: bijectivity alone does not guarantee continuity of an inverse.

Derive the Inverse Derivative

Fix \(y\in V_0\), write \(x=g(y)\), and take a small increment \(k\) such that \(y+k\in V_0\). Set \(\Delta=g(y+k)-g(y)\). The Lipschitz estimate gives \(\|\Delta\|_2\leq C\|k\|_2\), where \(C=(1-q)^{-1}\|A^{-1}\|_{\mathrm{op}}\). In particular, \(\Delta\to0\) as \(k\to0\). Differentiability of \(f\) at \(x\) gives $$ k=f(x+\Delta)-f(x)=Df(x)\Delta+\rho(\Delta), \qquad \frac{\|\rho(\Delta)\|_2}{\|\Delta\|_2}\longrightarrow 0 $$ as \(\Delta\to0\), with the remainder taken as \(0\) when \(\Delta=0\).

The derivative \(Df(x)\) is invertible. Indeed, \(Df(x)=A\,DF(x-a)=A(I+DR(x-a))\), and for every vector \(v\), $$ \|(I+DR(x-a))v\|_2 \geq (1-q)\|v\|_2. $$ Thus \(I+DR(x-a)\) is injective and hence invertible as a linear map from \(\mathbb{R}^n\) to itself. The expansion can therefore be rearranged as $$ \Delta=Df(x)^{-1}k-Df(x)^{-1}\rho(\Delta). $$ The remainder term is \(o(\|k\|_2)\): when \(\Delta\neq0\), its norm divided by \(\|k\|_2\) is at most $$ \|Df(x)^{-1}\|_{\mathrm{op}} \frac{\|\rho(\Delta)\|_2}{\|\Delta\|_2} \frac{\|\Delta\|_2}{\|k\|_2}, $$ where the last factor is bounded by \(C\) and the middle factor tends to zero. When \(\Delta=0\), the remainder term is zero. This proves differentiability of \(g\) at \(y\) and gives the formula below.

Theorem (Regularity of the Constructed Local Inverse): The inverse \(g:V_0\to U_0\) is continuously differentiable, and for every \(y\in V_0\), $$ Dg(y)=\bigl(Df(g(y))\bigr)^{-1}. $$

Proof. The preceding increment argument proves differentiability at each \(y\) and the stated formula. It remains to check continuity of the derivative. The map \(g\) is continuous by the Lipschitz estimate, and \(Df\) is continuous by hypothesis. Hence \(Df(g(y))\) varies continuously with \(y\). Its inverses also vary continuously: for invertible matrices \(B_1,B_0\), $$ B_1^{-1}-B_0^{-1}=B_1^{-1}(B_0-B_1)B_0^{-1}. $$ In the neighborhood under consideration, the estimate above gives a uniform bound on \(\|Df(x)^{-1}\|_{\mathrm{op}}\), namely \((1-q)^{-1}\|A^{-1}\|_{\mathrm{op}}\). The identity therefore shows that \(Df(g(y))^{-1}\) depends continuously on \(y\). By the derivative formula, \(Dg\) is continuous. \(\square\)

Worked Examples

Worked Example: A Cubic Perturbation of the Identity

Let \(f:\mathbb{R}\to\mathbb{R}\) be \(f(t)=t+t^3\), and consider the base point \(a=0\). Here \(b=0\), \(A=f'(0)=1\), and the normalized map is already \(F(t)=t+t^3\). Thus \(R(t)=t^3\) and \(R'(t)=3t^2\). On the closed interval \([-1/4,1/4]\), $$ |R'(t)|\leq \frac{3}{16}=q<1. $$ The proof gives a unique solution \(t\in(-1/4,1/4)\) to \(t+t^3=z\) for every \(|z|<(1-q)/4=13/64\). It also gives a Lipschitz local inverse and, at \(y=f(t)\), the derivative $$ g'(y)=\frac{1}{1+3g(y)^2}. $$ The denominator is positive, consistent with the derivative being invertible throughout this interval.

Worked Example: A Triangular Map in Two Variables

Consider \(f(x,y)=(x,y+x^2)\) near \((0,0)\). Its derivative at the origin is the identity, and the remainder is \(R(x,y)=(0,x^2)\). Its derivative acts by $$ DR(x,y)(u,v)=(0,2xu), $$ so \(\|DR(x,y)\|_{\mathrm{op}}=2|x|\). On the closed ball of radius \(r=1/4\), this is at most \(q=1/2\). Therefore every target \(z=(u,v)\) with \(\|z\|_2<1/8\) has a unique preimage in the open ball of radius \(1/4\).

In this example the inverse can also be written explicitly: $$ f^{-1}(u,v)=(u,v-u^2). $$ Substitution verifies both identities: \(f(u,v-u^2)=(u,v)\), and \(f(x,y)=(x,y+x^2)\) followed by the displayed inverse returns \((x,y)\). The contraction proof supplies the same local conclusion without relying on an explicit solution, which is the useful feature for more complicated maps.

Worked Example: A Nonidentity Derivative at the Base Point

Define \(f:\mathbb{R}^2\to\mathbb{R}^2\) by \(f(x,y)=(2x+y+x^2,\ x+3y)\). At the origin, $$ A=Df(0,0)= \begin{pmatrix} 2&1\\ 1&3 \end{pmatrix}, \qquad \det A=6-1=5, \qquad A^{-1}=\frac15 \begin{pmatrix} 3&-1\\ -1&2 \end{pmatrix}. $$ After normalization, \(R(x,y)=A^{-1}(x^2,0)=(3x^2/5,-x^2/5)\). Its derivative sends \((u,v)\) to \((6xu/5,-2xu/5)\), so $$ \|DR(x,y)\|_{\mathrm{op}} =\frac{\sqrt{40}}{5}|x| =\frac{2\sqrt{10}}{5}|x|. $$ On the ball of radius \(r=1/4\), this is at most \(q=\sqrt{10}/10<1\). The normalized contraction argument therefore constructs a continuously differentiable local inverse for targets \(w\) satisfying $$ \|A^{-1}w\|_2<(1-q)r. $$ The target neighborhood in the original coordinates is transformed by \(A\); it need not be a Euclidean ball centered at the origin. This illustrates why normalization is useful even when the original derivative is not the identity.

Why the Proof Is Organized This Way

Each stage supplies something needed by the next. Derivative control gives a contraction; the contraction gives existence and uniqueness; the lower bound on increments gives continuity of the inverse; and continuity allows the differentiability argument to control how preimages change with targets. Omitting the continuity step can leave a gap, because a bijection need not have a continuous inverse.

The proof also separates local conclusions from global ones. The fixed-point argument works on a chosen small ball and for targets in a corresponding small neighborhood. It does not imply that \(f\) is injective on its entire domain. This is precisely why the Inverse Function Theorem is local: it turns reliable first-order behavior at one point into a well-controlled inverse nearby.

Check Your Understanding

Use the proof architecture to answer the following questions.

  1. Why does the normalization \(F(h)=A^{-1}(f(a+h)-f(a))\) make \(DF(0)\) the identity?
  2. How does a bound on \(\|DR(h)\|_{\mathrm{op}}\) lead to a contraction on a closed ball?
  3. Why does the strict target condition \(\|z\|_2<(1-q)r\) ensure that the fixed point lies in the open ball?
  4. What lower bound on \(\|F(h)-F(k)\|_2\) proves injectivity and yields a Lipschitz inverse?
  5. Why is continuity of the inverse established before the differentiability argument?
  6. At a target \(y\), which input point appears in the formula for \(Dg(y)\), and why is \(Df\) evaluated there invertible?