Tutorials › Real Analysis › Implicit Function Theorem

Multivariable Analysis · Tutorial 805 of 1000

Implicit Function Theorem

Use the Inverse Function Theorem to recognize when an equation locally determines one variable as a continuously differentiable function of the others.

Advanced 10 min read

What You'll Learn

  • Recognize an implicit equation and its locally dependent variable
  • Derive the local graph condition from the Inverse Function Theorem
  • Prove local existence and uniqueness of the implicit solution
  • Compute the derivative of that solution using the chain rule
  • Interpret the derivative as a tangent slope and a rate of change
  • Identify why a vanishing partial derivative prevents this criterion from applying

From an Equation to a Local Graph

An equation such as \(F(x,y)=0\) describes a set of points, but it may not give \(y\) explicitly in terms of \(x\). Near some points, however, the equation does determine exactly one value of \(y\) for each nearby \(x\). In that neighborhood, the equation defines a function \(y=\phi(x)\), even if no convenient explicit formula for \(\phi\) is available.

The Inverse Function Theorem gives a way to establish this local graph property. The key condition is that the partial derivative with respect to the variable to be solved for is nonzero. We will first derive the local graph result for one dependent variable. This makes the role of the condition visible and leads directly to a formula for the derivative of the implicit function.

Throughout, \(x\) may be a vector in \(\mathbb{R}^m\), while \(y\) is a real number. The same reasoning includes the familiar two-variable equation by taking \(m=1\). We assume \(F\) is continuously differentiable on an open set in \(\mathbb{R}^{m+1}\).

Definition: Suppose \(F\) is defined near \((a,b)\) and \(F(a,b)=0\). A function \(\phi\), defined for \(x\) near \(a\), is called a local implicit function for this equation at \((a,b)\) if \(\phi(a)=b\) and \(F(x,\phi(x))=0\) for every \(x\) in its domain. The equation \(F(x,y)=0\) then describes the graph \(y=\phi(x)\) locally, provided the nearby solutions are unique.

The Local Graph Criterion

To apply the Inverse Function Theorem, package the input \(x\) together with the value of the equation. Define a map \(G\) by \(G(x,y)=(x,F(x,y))\). If we can locally invert \(G\), then the output \((x,0)\) determines the point \((x,y)\) satisfying \(F(x,y)=0\). The derivative of \(G\) is invertible precisely when the partial derivative \(F_y\) is nonzero.

Theorem (Local Implicit Graph Criterion): Let \(U\subseteq\mathbb{R}^{m+1}\) be open, let \(F:U\to\mathbb{R}\) be continuously differentiable, and suppose \((a,b)\in U\) satisfies \(F(a,b)=0\) and \(F_y(a,b)\neq 0\). Then there are an open neighborhood \(X\) of \(a\) and an open neighborhood \(W\) of \((a,b)\) such that for every \(x\in X\), there is exactly one \(y\) with \((x,y)\in W\) and \(F(x,y)=0\). These values define a continuously differentiable function \(\phi:X\to\mathbb{R}\) with \(\phi(a)=b\) and \(F(x,\phi(x))=0\).

Proof. Define \(G:U\to\mathbb{R}^{m+1}\) by \(G(x,y)=(x,F(x,y))\). The derivative at \((a,b)\), written in block form, is

$$ DG(a,b)= \begin{pmatrix} I_m & 0\\ D_xF(a,b) & F_y(a,b) \end{pmatrix}. $$

This matrix is invertible because its determinant is \(F_y(a,b)\), which is nonzero. Also, \(G(a,b)=(a,0)\). The Inverse Function Theorem therefore supplies open neighborhoods \(W\) of \((a,b)\) and \(Z\) of \((a,0)\) such that \(G\) maps \(W\) bijectively onto \(Z\), with a continuously differentiable inverse.

Since \(Z\) is open and contains \((a,0)\), there is an open neighborhood \(X\) of \(a\) such that \((x,0)\in Z\) for every \(x\in X\). Write the inverse of \(G\) on \(Z\) as \(G^{-1}(u,t)\). For each \(x\in X\), define \(\phi(x)\) to be the last coordinate of \(G^{-1}(x,0)\). Because \(G(G^{-1}(x,0))=(x,0)\), the first \(m\) coordinates of \(G^{-1}(x,0)\) are \(x\), and its last coordinate satisfies \(F(x,\phi(x))=0\). At \(x=a\), the unique inverse image of \((a,0)\) is \((a,b)\), so \(\phi(a)=b\).

For uniqueness, suppose \((x,y)\in W\) and \(F(x,y)=0\), where \(x\in X\). Then \(G(x,y)=(x,0)\). Since \(G\) is one-to-one on \(W\), \((x,y)=G^{-1}(x,0)\), so \(y=\phi(x)\). Finally, \(\phi\) is continuously differentiable because it is obtained by composing the continuously differentiable map \(G^{-1}\) with \(x\mapsto(x,0)\), then taking the last coordinate. This proves the claim. \(\square\)

The theorem is local in two ways. It guarantees uniqueness only among solutions whose points lie in \(W\), and it concerns \(x\) only in a neighborhood \(X\) of \(a\). It does not rule out other solutions far from \((a,b)\).

Worked Example: The Upper Half of a Circle

Consider \(F(x,y)=x^2+y^2-1\) near \((0,1)\). The point lies on the zero set because \(F(0,1)=0^2+1^2-1=0\). The partial derivative with respect to \(y\) is \(F_y(x,y)=2y\), so \(F_y(0,1)=2\neq0\). The local implicit graph criterion guarantees a unique continuously differentiable solution \(y=\phi(x)\) near \(x=0\), with \(\phi(0)=1\).

In this case the equation can also be solved explicitly. Near \((0,1)\), the relevant solution is \(\phi(x)=\sqrt{1-x^2}\), rather than the lower branch. Direct substitution verifies the equation:

$$ F(x,\phi(x))=x^2+\left(\sqrt{1-x^2}\right)^2-1 =x^2+1-x^2-1=0. $$

The other solution, \(-\sqrt{1-x^2}\), is near \((0,-1)\), not \((0,1)\). The local uniqueness conclusion is therefore consistent with the two branches of the whole circle: near the chosen point, only the upper branch is relevant.

Finding the Derivative Without Solving Explicitly

Once the local graph exists, its derivative follows from differentiating the equation it satisfies. For \(x\) in the neighborhood \(X\), the identity \(F(x,\phi(x))=0\) holds. The chain rule differentiates this identity with respect to \(x\), producing a linear relation between the derivative of \(F\) in the \(x\)-variables and the derivative of \(\phi\).

Theorem (Derivative Formula for an Implicit Function): Under the hypotheses of the Local Implicit Graph Criterion, the implicit function satisfies $$ D\phi(x)=-\frac{D_xF(x,\phi(x))}{F_y(x,\phi(x))} $$ for \(x\) sufficiently near \(a\). Here \(D_xF\) is a row vector and \(F_y\) is a scalar.

Proof. The identity \(F(x,\phi(x))=0\) holds throughout the neighborhood where \(\phi\) is defined. Applying the multivariable chain rule at \(x\) gives

$$ D_xF(x,\phi(x))+F_y(x,\phi(x))D\phi(x)=0. $$

The denominator \(F_y(x,\phi(x))\) is nonzero after possibly shrinking the neighborhood. Indeed, \(F_y\) is continuous and \(F_y(a,b)\neq0\); since \(\phi\) is continuous and \(\phi(a)=b\), the points \((x,\phi(x))\) remain close to \((a,b)\) when \(x\) is close to \(a\). We may therefore divide the displayed identity by \(F_y(x,\phi(x))\), obtaining the stated formula. \(\square\)

When \(m=1\), \(D_xF\) is just \(F_x\), and the formula becomes

$$ \phi'(x)=-\frac{F_x(x,\phi(x))}{F_y(x,\phi(x))}. $$

This is the slope of the curve \(F(x,y)=0\) at the point \((x,\phi(x))\). The numerator measures how the equation changes with \(x\), while the denominator measures how it changes with \(y\). The minus sign expresses how \(y\) must adjust to offset a change in \(x\) while keeping \(F\) equal to zero.

Worked Example: A Solution Curve Without an Explicit Formula

Let \(F(x,y)=y-1+xe^y\). At \((0,1)\), \(F(0,1)=1-1+0e^1=0\). Also,

$$ F_y(x,y)=1+xe^y, \qquad F_x(x,y)=e^y, \qquad F_y(0,1)=1. $$

The nonzero partial derivative implies that \(y-1+xe^y=0\) defines a unique continuously differentiable function \(y=\phi(x)\) near \(x=0\), with \(\phi(0)=1\). The derivative formula gives

$$ \phi'(0)=-\frac{F_x(0,1)}{F_y(0,1)} =-\frac{e^1}{1}=-e. $$

Thus the first-order approximation to the solution near \(x=0\) is \(\phi(x)\approx 1-ex\). This approximation describes the local response of the solution even though the equation has not been rearranged into a simple explicit formula for \(y\).

Worked Example: The Tangent to an Implicit Curve

Consider \(F(x,y)=x^2+xy+y^2-3\) at \((1,1)\). Substitution gives \(F(1,1)=1+1+1-3=0\). The partial derivatives are \(F_x(x,y)=2x+y\) and \(F_y(x,y)=x+2y\). At the point, \(F_y(1,1)=3\neq0\), so the curve is locally a graph \(y=\phi(x)\) through \((1,1)\). Its slope there is

$$ \phi'(1)=-\frac{F_x(1,1)}{F_y(1,1)} =-\frac{2(1)+1}{1+2(1)}=-1. $$

The tangent line at \((1,1)\) is therefore \(y-1=-(x-1)\), or \(x+y=2\). This tangent direction is consistent with the derivative of the defining equation: the gradient at \((1,1)\) is \((3,3)\), and the tangent direction \((1,-1)\) satisfies \(3(1)+3(-1)=0\). Thus the tangent direction makes the first-order change in \(F\) equal to zero.

What the Nonzero Partial Derivative Does—and Does Not—Say

The condition \(F_y(a,b)\neq0\) is a sufficient condition for solving locally for \(y\) as a continuously differentiable function of \(x\). It also identifies which variable is being treated as dependent. If instead a different partial derivative is nonzero, a different variable may be the one that can be solved for locally.

The condition is not necessary for every possible graph representation. Rather, it is the condition that lets this direct application of the Inverse Function Theorem work. When it fails, the conclusion may fail, or a graph may still exist for some other reason. The equation \(y^2-x^2=0\) illustrates a genuine failure of a single local graph describing the entire zero set over \(x\).

Worked Example: Why the Condition Matters

Set \(F(x,y)=y^2-x^2\). At \((0,0)\), \(F(0,0)=0\), but \(F_y(0,0)=2(0)=0\), so the local implicit graph criterion does not apply. In fact, for every nonzero \(x\), the equation \(F(x,y)=0\) becomes

$$ y^2=x^2, \qquad\text{so}\qquad y=x\ \text{or}\ y=-x. $$

There are two distinct solutions \(y\) for each such \(x\), both arbitrarily close to \(0\) when \(x\) is small. Consequently, the entire zero set near the origin cannot be represented as one graph with exactly one \(y\) for each nearby \(x\). Each branch separately is a graph, but the combined set is not a single-valued graph over the \(x\)-axis.

The implicit function viewpoint is useful whenever a model is naturally expressed as a constraint rather than an explicit formula. It establishes that a solution persists under small changes in the input and gives the derivative of that solution directly from the original equation. The central checks are to verify the equation at the point and to identify a nonzero partial derivative in the variable to be solved for.

Check Your Understanding

Use the local graph criterion and derivative formula to answer the following questions.

  1. For an equation \(F(x,y)=0\) at a point \((a,b)\), which partial derivative must be nonzero to use this criterion to solve locally for \(y\)?
  2. Why does the map \(G(x,y)=(x,F(x,y))\) have an invertible derivative when \(F_y(a,b)\neq0\)?
  3. If \(F_x(a,b)=4\) and \(F_y(a,b)=-2\), what is the slope of the local implicit function at \(a\)?
  4. Does local uniqueness rule out solutions of \(F(x,y)=0\) far from the point under consideration? Explain.
  5. What does the equation \(y^2-x^2=0\) show about the role of the nonzero-partial-derivative condition?