Tutorials › Real Analysis › Statement of the Implicit Function Theorem

Multivariable Analysis · Tutorial 806 of 1000

Statement of the Implicit Function Theorem

Learn when a system of equations determines several dependent variables locally, and how to calculate the derivative of that solution function.

Advanced 10 min read

What You'll Learn

  • State the multivariable Implicit Function Theorem with independent and dependent variables
  • Identify the Jacobian block that must be invertible
  • Understand the theorem’s local existence and uniqueness conclusion
  • Calculate the derivative of a vector-valued implicit function
  • Apply the theorem to systems with multiple equations and unknowns
  • Distinguish local uniqueness from uniqueness of all solutions

From One Equation to a System

The previous tutorial treated one equation \(F(x,y)=0\), where \(x\) may be a vector but \(y\) is a single real number. When there are several equations and several unknowns to solve for, the same local question arises: can the equations determine a vector of dependent variables as a function of the remaining variables?

Let \(x\in\mathbb{R}^m\) denote the independent variables and \(y\in\mathbb{R}^n\) the dependent variables. A system of \(n\) equations can be written as \(F(x,y)=0\), where \(F\) takes values in \(\mathbb{R}^n\). The central condition is that the derivative of \(F\) with respect to the dependent variables is an invertible \(n\)-by-\(n\) matrix at the point of interest.

Definition: Suppose \(F\) is defined near \((a,b)\in\mathbb{R}^m\times\mathbb{R}^n\) and \(F(a,b)=0\in\mathbb{R}^n\). A function \(\phi\), defined for \(x\) near \(a\) and taking values in \(\mathbb{R}^n\), is a local implicit function for the system at \((a,b)\) if \(\phi(a)=b\) and \(F(x,\phi(x))=0\) throughout its domain. The system locally determines the graph \(y=\phi(x)\) when the relevant nearby solutions are unique.

The Multivariable Implicit Function Theorem

Write \(D_xF(a,b)\) for the \(n\)-by-\(m\) derivative matrix with respect to \(x\), and \(D_yF(a,b)\) for the \(n\)-by-\(n\) derivative matrix with respect to \(y\). The theorem requires \(D_yF(a,b)\) to be invertible. It does not require the full derivative of \(F\) to be invertible: that derivative has \(n\) rows and \(m+n\) columns.

Theorem (Implicit Function Theorem): Let \(U\subseteq\mathbb{R}^{m+n}\) be open, let \(F:U\to\mathbb{R}^n\) be continuously differentiable, and suppose \((a,b)\in U\) satisfies \(F(a,b)=0\). If the matrix \(D_yF(a,b)\) is invertible, then there are open neighborhoods \(X\subseteq\mathbb{R}^m\) of \(a\) and \(W\subseteq U\) of \((a,b)\) such that for every \(x\in X\), exactly one \(y\in\mathbb{R}^n\) satisfies \((x,y)\in W\) and \(F(x,y)=0\). These values define a continuously differentiable function \(\phi:X\to\mathbb{R}^n\) with \(\phi(a)=b\) and \(F(x,\phi(x))=0\).

Proof. Define \(G:U\to\mathbb{R}^{m+n}\) by \(G(x,y)=(x,F(x,y))\). At \((a,b)\), its derivative acts on \((h,k)\in\mathbb{R}^m\times\mathbb{R}^n\) as

$$ DG(a,b)(h,k) = \bigl(h,\ D_xF(a,b)h+D_yF(a,b)k\bigr). $$

This derivative is invertible. In fact, given any output \((r,s)\in\mathbb{R}^m\times\mathbb{R}^n\), the only possible first input component is \(h=r\). Since \(D_yF(a,b)\) is invertible, the equation \(D_xF(a,b)r+D_yF(a,b)k=s\) has exactly one solution,

$$ k = \bigl(D_yF(a,b)\bigr)^{-1} \bigl(s-D_xF(a,b)r\bigr). $$

Thus \(DG(a,b)\) is a bijective linear map. Also \(G(a,b)=(a,0)\). Apply the Inverse Function Theorem to \(G\). It gives open neighborhoods \(W\) of \((a,b)\) and \(Z\) of \((a,0)\) such that \(G\) maps \(W\) bijectively onto \(Z\), with a continuously differentiable inverse \(G^{-1}:Z\to W\).

Because \(Z\) is open and contains \((a,0)\), we can choose an open neighborhood \(X\) of \(a\) such that \((x,0)\in Z\) for every \(x\in X\). Define \(\phi(x)\) to be the last \(n\) coordinates of \(G^{-1}(x,0)\). The first \(m\) coordinates of that inverse image are \(x\), since the first component of \(G\) is unchanged. Its remaining coordinates therefore satisfy \(F(x,\phi(x))=0\). At \(x=a\), the inverse image of \((a,0)\) is \((a,b)\), so \(\phi(a)=b\).

If \((x,y)\in W\), \(x\in X\), and \(F(x,y)=0\), then \(G(x,y)=(x,0)\). The injectivity of \(G\) on \(W\) implies \((x,y)=G^{-1}(x,0)\), so \(y=\phi(x)\). This proves the stated uniqueness. Finally, \(\phi\) is continuously differentiable because it is obtained by composing \(G^{-1}\) with \(x\mapsto(x,0)\) and then taking the last \(n\) coordinates. \(\square\)

The result is local: uniqueness is guaranteed only for solutions whose full points \((x,y)\) lie in \(W\). It does not assert that the system has no other solutions outside that neighborhood.

The Derivative of the Implicit Function

Once the local function exists, its derivative is determined by differentiating the identity \(F(x,\phi(x))=0\). The result is a matrix formula: changes in the independent variables are balanced by changes in the dependent variables so that all \(n\) equations remain satisfied.

Theorem (Derivative Formula for an Implicit Function): Under the hypotheses of the Implicit Function Theorem, after possibly shrinking \(X\), \(D_yF(x,\phi(x))\) is invertible for every \(x\in X\), and $$ D\phi(x) = -\bigl(D_yF(x,\phi(x))\bigr)^{-1}D_xF(x,\phi(x)). $$ Here \(D\phi(x)\) is an \(n\)-by-\(m\) matrix.

Proof. Since \(D_yF\) is continuous and \(D_yF(a,b)\) is invertible, its determinant remains nonzero for points sufficiently close to \((a,b)\). The function \(\phi\) is continuous and \(\phi(a)=b\), so by shrinking \(X\) if necessary, every \((x,\phi(x))\) lies close enough to \((a,b)\) that \(D_yF(x,\phi(x))\) is invertible.

Differentiate \(F(x,\phi(x))=0\) using the multivariable chain rule. The derivative of the left side is the zero matrix, giving

$$ D_xF(x,\phi(x)) + D_yF(x,\phi(x))D\phi(x)=0. $$

Multiplying on the left by the inverse of \(D_yF(x,\phi(x))\) and rearranging yields the formula. \(\square\)

The order of the matrix factors matters. \(D_yF\) is \(n\)-by-\(n\), while \(D_xF\) is \(n\)-by-\(m\), so \(\bigl(D_yF\bigr)^{-1}D_xF\) is \(n\)-by-\(m\), as required for \(D\phi\). Reversing the factors is generally not even a defined matrix product.

Worked Applications

Worked Example: Two Equations and One Independent Variable

Consider the system \(F(x,u,v)=(u^2+v-x,\ u+v^2-1)=0\) at \((x,u,v)=(1,1,0)\). Substitution gives \(F(1,1,0)=(1+0-1,\ 1+0-1)=(0,0)\). The matrix of derivatives with respect to the dependent variables \((u,v)\) is

$$ D_{(u,v)}F(x,u,v) = \begin{pmatrix} 2u&1\\ 1&2v \end{pmatrix}, \qquad D_{(u,v)}F(1,1,0) = \begin{pmatrix} 2&1\\ 1&0 \end{pmatrix}. $$

Its determinant is \(2(0)-1(1)=-1\), so it is invertible. The theorem gives continuously differentiable functions \(u=\phi_1(x)\) and \(v=\phi_2(x)\) near \(x=1\), with \((\phi_1(1),\phi_2(1))=(1,0)\). Since \(D_xF=(-1,0)^T\) and

$$ \begin{pmatrix} 2&1\\ 1&0 \end{pmatrix}^{-1} = \begin{pmatrix} 0&1\\ 1&-2 \end{pmatrix}, $$

the derivative formula gives

$$ \begin{pmatrix} \phi_1'(1)\\ \phi_2'(1) \end{pmatrix} = - \begin{pmatrix} 0&1\\ 1&-2 \end{pmatrix} \begin{pmatrix} -1\\ 0 \end{pmatrix} = \begin{pmatrix} 0\\ 1 \end{pmatrix}. $$

Thus \(u\) has zero first-order change at \(x=1\), while \(v\) has derivative \(1\) there.

Worked Example: Two Independent Variables and Two Dependent Variables

Let \(F(x_1,x_2,u,v)=(u+v-x_1,\ u-v-x_2)\). At the origin, \(F(0,0,0,0)=(0,0)\), and the derivative blocks are

$$ D_{(u,v)}F = \begin{pmatrix} 1&1\\ 1&-1 \end{pmatrix}, \qquad D_{(x_1,x_2)}F = \begin{pmatrix} -1&0\\ 0&-1 \end{pmatrix}. $$

The determinant of the first matrix is \(-2\), so the theorem applies. In this example the equations can also be solved directly: adding them gives \(2u=x_1+x_2\), while subtracting the second equation from the first gives \(2v=x_1-x_2\). Hence

$$ \phi(x_1,x_2) = \begin{pmatrix} (x_1+x_2)/2\\ (x_1-x_2)/2 \end{pmatrix}, \qquad D\phi = \begin{pmatrix} 1/2&1/2\\ 1/2&-1/2 \end{pmatrix}. $$

The derivative formula agrees: the inverse of \(D_{(u,v)}F\) is \(\begin{pmatrix}1/2&1/2\\1/2&-1/2\end{pmatrix}\), and multiplying it by \(D_{(x_1,x_2)}F=-I_2\), then taking the negative, gives exactly the displayed derivative.

Worked Example: A Scalar Equation as a Special Case

For a single dependent variable, take \(F(x,y)=y^2+xy-4\) at \((0,2)\). Direct substitution gives \(F(0,2)=4+0-4=0\), and \(F_y(x,y)=2y+x\), so \(F_y(0,2)=4\neq0\). The theorem guarantees a local function \(y=\phi(x)\) with \(\phi(0)=2\). Since \(F_x(x,y)=y\), its derivative at zero is

$$ \phi'(0) = -\frac{F_x(0,2)}{F_y(0,2)} = -\frac{2}{4} = -\frac12. $$

For a direct check, the branch through \(y=2\) is \(\phi(x)=\frac{-x+\sqrt{x^2+16}}{2}\) near zero. It satisfies the equation because it is a root of \(y^2+xy-4=0\), and differentiating this explicit expression at zero gives \(\phi'(0)=-1/2\). The general theorem obtains the same local information without requiring such a formula.

What to Check and What the Theorem Does Not Claim

For a system with \(n\) equations, the dependent-variable derivative must be an \(n\)-by-\(n\) invertible matrix. Checking only that some entries, or even some individual partial derivatives, are nonzero is not enough. The relevant test is whether the entire matrix has nonzero determinant. For \(n=1\), this reduces to the nonzero partial derivative condition from the previous tutorial.

Invertibility is a sufficient condition for the stated local conclusion. If it fails, this theorem gives no conclusion; a local graph may still exist, but it requires other reasoning. Likewise, the theorem says nothing about solutions far from \((a,b)\). A system can have many solutions overall while having exactly one solution in the specified neighborhood for each nearby \(x\).

The theorem is useful when equations describe constraints, equilibria, or linked quantities more naturally than explicit formulas do. It provides both local persistence of solutions and a precise sensitivity formula. In applications, the practical sequence is to verify \(F(a,b)=0\), form \(D_yF(a,b)\), check its invertibility, and then use the derivative formula to study how the solution responds to changes in \(x\).

Check Your Understanding

Use the theorem and derivative formula to answer the following questions.

  1. For \(F:\mathbb{R}^{m+n}\to\mathbb{R}^n\), what size is the matrix \(D_yF(a,b)\), and what condition must it satisfy?
  2. Why does the theorem require invertibility of \(D_yF(a,b)\), rather than invertibility of the full derivative \(DF(a,b)\)?
  3. What does local uniqueness mean in the Implicit Function Theorem, and what does it not rule out?
  4. If \(D_xF\) is \(n\)-by-\(m\), what is the size of \(D\phi\), and why must the inverse matrix appear on the left in the derivative formula?
  5. For a scalar equation, which condition in the multivariable theorem becomes the nonzero partial derivative condition from the previous tutorial?