Tutorials › Real Analysis › Proof Architecture for the Implicit Function Theorem

Multivariable Analysis · Tutorial 807 of 1000

Proof Architecture for the Implicit Function Theorem

Learn how augmenting a system with its independent variables turns a local implicit-graph problem into a local inversion problem.

Advanced 10 min read

What You'll Learn

  • Build the augmented map that encodes both the independent variables and the system of equations
  • Verify invertibility of its block derivative from invertibility of the dependent-variable derivative
  • Track neighborhoods carefully when restricting the local inverse to target values with zero equation component
  • Distinguish uniqueness inside a chosen neighborhood from global uniqueness of solutions
  • Recognize why a singular dependent-variable derivative does not rule out an implicit graph

The Proof’s Central Change of View

The Implicit Function Theorem describes solutions of \(F(x,y)=0\) as a graph \(y=\phi(x)\). The proof becomes more transparent when this is viewed as an inversion problem. Instead of trying to solve the equations directly for \(y\), retain \(x\) as part of the output and define a map that records both \(x\) and \(F(x,y)\).

Specifically, set \(G(x,y)=(x,F(x,y))\). If \(G(x,y)=(x,0)\), then \(F(x,y)=0\). Thus solving the system at a given \(x\) is equivalent to finding the point whose image under \(G\) is \((x,0)\). The Inverse Function Theorem can do this locally if the derivative of \(G\) is invertible at the point under consideration.

The proof architecture has three distinct tasks: verify that derivative is invertible, apply the Inverse Function Theorem, and choose neighborhoods so that the targets \((x,0)\) are within the range where the local inverse is defined. Keeping these tasks separate helps prevent errors about matrix dimensions and the scope of uniqueness.

Why the Augmented Derivative Is Invertible

At a point \((a,b)\), the derivative of \(G\) sends an increment \((h,k)\in\mathbb{R}^m\times\mathbb{R}^n\) to

$$ DG(a,b)(h,k) = \bigl(h,\ D_xF(a,b)h+D_yF(a,b)k\bigr). $$

The first output component is exactly \(h\), so it preserves the independent-variable increment. Once \(h\) has been specified, solving for the second input component \(k\) requires solving an equation involving \(D_yF(a,b)\). That is why invertibility of this particular \(n\)-by-\(n\) block is the relevant condition.

Lemma (Invertibility of the Augmented Derivative): Let \(A:\mathbb{R}^m\to\mathbb{R}^n\) and \(B:\mathbb{R}^n\to\mathbb{R}^n\) be linear maps. Define \(T:\mathbb{R}^m\times\mathbb{R}^n\to\mathbb{R}^m\times\mathbb{R}^n\) by \(T(h,k)=(h,Ah+Bk)\). Then \(T\) is invertible if and only if \(B\) is invertible.

Proof. Suppose first that \(B\) is invertible. Given any \((r,s)\in\mathbb{R}^m\times\mathbb{R}^n\), the equation \(T(h,k)=(r,s)\) forces \(h=r\). Its second component then requires \(Ar+Bk=s\), whose unique solution is \(k=B^{-1}(s-Ar)\). Thus every output has exactly one preimage, so \(T\) is invertible.

Conversely, suppose \(B\) is not invertible. Since \(B\) maps \(\mathbb{R}^n\) to itself, noninvertibility means that \(B\) has a nonzero vector \(k\) with \(Bk=0\). Then \(T(0,k)=(0,0)=T(0,0)\), although \((0,k)\neq(0,0)\). Therefore \(T\) is not injective and cannot be invertible. This proves both directions. \(\square\)

For \(G(x,y)=(x,F(x,y))\), take \(A=D_xF(a,b)\) and \(B=D_yF(a,b)\). The lemma shows directly that \(DG(a,b)\) is invertible exactly when \(D_yF(a,b)\) is invertible. This is the linear-algebra hinge of the proof: the independent variables are carried unchanged, while the dependent-variable block controls whether the remaining output can be solved for uniquely.

From a Local Inverse to a Local Graph

Here is the second part of the architecture. It isolates the neighborhood issue that can be hidden by the phrase “apply the Inverse Function Theorem.” An inverse is available only on a particular open target neighborhood. We need to ensure that a whole neighborhood of independent-variable values, paired with zero, lies inside it.

Lemma (Zero-Slice Construction): Suppose \(G\) is a continuously differentiable map defined near \((a,b)\), with \(G(a,b)=(a,0)\). Suppose the Inverse Function Theorem gives open neighborhoods \(W\) of \((a,b)\) and \(Z\) of \((a,0)\) such that \(G:W\to Z\) is bijective with a continuously differentiable inverse. Then there is an open neighborhood \(X\) of \(a\) and a continuously differentiable function \(\phi:X\to\mathbb{R}^n\) such that \(\phi(a)=b\), \(G(x,\phi(x))=(x,0)\) for \(x\in X\), and \((x,\phi(x))\) is the unique point of \(W\) mapped to \((x,0)\).

Proof. Because \(Z\) is open and contains \((a,0)\), there is some \(\delta>0\) such that the open ball of radius \(\delta\) around \((a,0)\) is contained in \(Z\). Choose \(r>0\) with \(r<\delta\), and let \(X=B_r^{(2)}(a)\subseteq\mathbb{R}^m\). For every \(x\in X\), the distance from \((x,0)\) to \((a,0)\) is \(\|x-a\|_2<r<\delta\). Hence \((x,0)\in Z\).

Define \(\phi(x)\) to be the last \(n\) coordinates of \(G^{-1}(x,0)\). The inverse is continuously differentiable, and the map \(x\mapsto(x,0)\) is continuously differentiable, so their composition is continuously differentiable. Taking its last \(n\) coordinates preserves this property. By definition, \(G(x,\phi(x))=(x,0)\).

At \(x=a\), the point \((a,b)\) is in \(W\) and maps to \((a,0)\). Since \(G\) is injective on \(W\), it is the unique point of \(W\) with that image. Thus \(G^{-1}(a,0)=(a,b)\), giving \(\phi(a)=b\). Finally, if another point \((x,y)\in W\) maps to \((x,0)\), injectivity of \(G\) on \(W\) implies \((x,y)=G^{-1}(x,0)=(x,\phi(x))\). This proves the uniqueness claim. \(\square\)

For the augmented map, \(G(x,y)=(x,F(x,y))\). The identity \(G(x,\phi(x))=(x,0)\) then says exactly that \(F(x,\phi(x))=0\). The uniqueness in the lemma is uniqueness among points in \(W\), not among all points in the domain of \(F\).

Putting the Architecture Together

Suppose \(F(a,b)=0\) and \(D_yF(a,b)\) is invertible. The augmented map satisfies \(G(a,b)=(a,0)\). By the invertibility lemma, \(DG(a,b)\) is invertible. Apply the Inverse Function Theorem to obtain \(W\), \(Z\), and a continuously differentiable local inverse. Then apply the zero-slice construction to obtain \(X\) and \(\phi\). Since \(G(x,\phi(x))=(x,0)\), the last \(n\) coordinates give \(F(x,\phi(x))=0\), with \(\phi(a)=b\) and local uniqueness.

This is a proof architecture for the Implicit Function Theorem stated in the previous tutorial: the theorem is not proved by an explicit formula for \(\phi\), but by converting the equation-solving problem into local inversion. The derivative formula follows after the graph is constructed: differentiating \(F(x,\phi(x))=0\) gives \(D_xF+D_yF\,D\phi=0\), and the invertibility of \(D_yF\) allows the equation to be solved for \(D\phi\).

1
Augment the output.
Set \(G(x,y)=(x,F(x,y))\), so the equation \(F(x,y)=0\) corresponds to the target \((x,0)\).
2
Check the derivative.
Use the block form of \(DG\). Its invertibility is equivalent to invertibility of \(D_yF(a,b)\).
3
Invert locally.
Apply the Inverse Function Theorem near \((a,b)\), where \(G(a,b)=(a,0)\).
4
Restrict to the zero slice.
Choose \(X\) so that every \((x,0)\), \(x\in X\), lies in the inverse’s target neighborhood. Read the last coordinates of \(G^{-1}(x,0)\) as \(\phi(x)\).

Worked Examples

Worked Example: Checking a Two-Equation Augmented Derivative

Consider \(F(x,u,v)=(u+v-x,\ u-v-x^2)\) at \((x,u,v)=(0,0,0)\). Substitution gives \(F(0,0,0)=(0+0-0,\ 0-0-0^2)=(0,0)\). The dependent-variable derivative is

$$ D_{(u,v)}F(x,u,v) = \begin{pmatrix} 1&1\\ 1&-1 \end{pmatrix}, \qquad \det D_{(u,v)}F(0,0,0)=1(-1)-1(1)=-2. $$

The determinant is nonzero, so this block is invertible. The augmented map is \(G(x,u,v)=(x,u+v-x,u-v-x^2)\). At the origin its derivative sends \((h,k,\ell)\) to \((h,k+\ell-h,\ k-\ell)\). To check invertibility directly, set this equal to \((r,s,t)\). The first equation gives \(h=r\), and the remaining equations become \(k+\ell=s+r\) and \(k-\ell=t\). Adding and subtracting gives the unique solution \(k=(s+t+r)/2\), \(\ell=(s-t+r)/2\). The augmented derivative is therefore invertible, as the block criterion predicts.

The local inverse, restricted to targets \((x,0,0)\), produces the nearby solution graph. In fact, adding the two equations \(u+v=x\) and \(u-v=x^2\) gives \(u=(x+x^2)/2\), while subtracting the second from the first gives \(v=(x-x^2)/2\). At \(x=0\), both values are zero, verifying the base point.

Worked Example: The Neighborhood Slice Must Be Chosen

Let \(F(x,y)=y+x\), with \((a,b)=(0,0)\). Then \(F(0,0)=0\), and the augmented map is \(G(x,y)=(x,y+x)\). Its inverse is \(G^{-1}(r,s)=(r,s-r)\), as substitution verifies: \(G(r,s-r)=(r,(s-r)+r)=(r,s)\).

Suppose we use the target neighborhood \(Z=B_1^{(2)}((0,0))\). A point on the zero slice is \((x,0)\), whose distance from the center is \(\sqrt{x^2+0^2}=|x|\). Thus \((x,0)\in Z\) whenever \(|x|<1\). Restricting the inverse to that slice gives \(\phi(x)\) equal to the second coordinate of \(G^{-1}(x,0)=(x,-x)\), so \(\phi(x)=-x\). Direct substitution verifies the equation: \(F(x,\phi(x))=F(x,-x)=-x+x=0\).

The neighborhood choice is part of the proof. The local inverse can only be evaluated at targets in its target neighborhood; openness supplies a sufficiently small set of \(x\)-values for which \((x,0)\) stays there.

Worked Example: A Singular Block Does Not Disprove a Graph

Consider \(F(x,y)=y^2\) at \((0,0)\). The equation \(F(x,y)=0\) has the local graph \(y=\phi(x)=0\), since \(F(x,0)=0\) for every \(x\). But \(F_y(x,y)=2y\), so \(F_y(0,0)=0\), which is not invertible as a one-by-one matrix.

The augmented map \(G(x,y)=(x,y^2)\) has derivative at the origin given by \(DG(0,0)(h,k)=(h,0)\). It is not injective: \((0,1)\) and \((0,-1)\) both map to \((0,0)\) under this linear map. Consequently, the Inverse Function Theorem cannot be applied there. The example does not contradict the Implicit Function Theorem: its invertibility condition is sufficient for the local conclusion, not necessary for every possible graph to exist.

What the Architecture Does—and Does Not—Give

The key algebraic check concerns the dependent-variable block, not the full derivative of \(F\). The full derivative maps \(\mathbb{R}^{m+n}\) to \(\mathbb{R}^n\), so when \(m>0\) it cannot be an invertible map between spaces of the same dimension. The augmented map has \(m+n\) inputs and \(m+n\) outputs, making the Inverse Function Theorem applicable.

A second common pitfall is to overstate uniqueness. The construction yields one solution for each nearby \(x\) among points \((x,y)\) in a chosen neighborhood \(W\). It does not rule out other solutions with the same \(x\) outside \(W\). This local qualification is essential, even when a particular example happens to have a globally unique solution.

The proof also explains the role of the hypotheses. Continuous differentiability lets us apply the Inverse Function Theorem to \(G\), while invertibility of \(D_yF(a,b)\) makes its derivative invertible. Once the local graph exists, the derivative formula records how its dependent variables must change to keep all equations equal to zero.

Check Your Understanding

Use the proof architecture to answer the following questions.

  1. Why is the first component of the augmented map chosen to be \(x\) itself?
  2. For \(T(h,k)=(h,Ah+Bk)\), explain why invertibility of \(B\) lets one solve uniquely for \(k\) after \(h\) is known.
  3. Why must the target neighborhood contain \((x,0)\) for every \(x\) in some neighborhood of \(a\)?
  4. What neighborhood restriction is attached to uniqueness in the construction?
  5. What does the example \(F(x,y)=y^2\) show about the role of the invertibility condition?