Tutorials › Real Analysis › Jacobian Matrices

Multivariable Analysis · Tutorial 785 of 1000

Jacobian Matrices

See how partial derivatives assemble into a matrix representing the derivative, and use that matrix to interpret first-order changes in several variables.

Advanced 10 min read

What You'll Learn

  • Define the Jacobian matrix of a differentiable map between Euclidean spaces
  • Identify output coordinates with rows and input coordinates with columns
  • Recover the derivative’s action on any displacement by matrix multiplication
  • Interpret Jacobian matrices for scalar-valued and vector-valued functions
  • Bound a matrix’s action using the sizes of its entries
  • Recognize why existing partial derivatives alone do not guarantee differentiability

From a Linear Map to a Matrix

The previous tutorial treated the derivative at a point as a linear map: it takes a small input displacement and predicts the first-order change in the output. A matrix gives a coordinate description of that same map. Its columns record what the derivative does to the standard basis displacements, and its rows record the output coordinates of those responses.

Let \(U\) be open in \(\mathbb{R}^m\), let \(f:U\to\mathbb{R}^n\), and suppose \(f\) is differentiable at \(a\in U\). Write \(f=(f_1,\ldots,f_n)\), where each \(f_i\) is a real-valued component function. The derivative \(Df(a)\) is a linear map from \(\mathbb{R}^m\) to \(\mathbb{R}^n\). The Jacobian matrix is the \(n\)-by-\(m\) matrix that represents this map in the standard coordinates.

Definition (Jacobian Matrix): If \(f:U\to\mathbb{R}^n\) is differentiable at \(a\in U\subseteq\mathbb{R}^m\), its Jacobian matrix at \(a\), denoted \(J_f(a)\), is the \(n\)-by-\(m\) matrix $$ J_f(a)= \begin{pmatrix} \partial_1 f_1(a) & \partial_2 f_1(a) & \cdots & \partial_m f_1(a)\\ \partial_1 f_2(a) & \partial_2 f_2(a) & \cdots & \partial_m f_2(a)\\ \vdots & \vdots & \ddots & \vdots\\ \partial_1 f_n(a) & \partial_2 f_n(a) & \cdots & \partial_m f_n(a) \end{pmatrix}. $$ Here \(\partial_j f_i(a)\) is the partial derivative of the \(i\)th output component with respect to the \(j\)th input coordinate. Rows correspond to output components; columns correspond to input coordinates.

Differentiability guarantees that all the partial derivatives in this definition exist. The entry in row \(i\), column \(j\) describes the \(i\)th coordinate of the derivative’s response to the \(j\)th standard basis vector \(e_j\). This placement is important: changing the row-column convention would change how the matrix acts on a displacement vector.

The Jacobian Represents the Derivative

The definition is useful because the matrix does not merely collect partial derivatives. It acts on every displacement in exactly the same way as the derivative does. The following result makes that connection precise.

Theorem (Matrix Representation of the Derivative): Suppose \(f:U\to\mathbb{R}^n\) is differentiable at \(a\in U\subseteq\mathbb{R}^m\). For every \(h\in\mathbb{R}^m\), $$ Df(a)(h)=J_f(a)h. $$ The \(j\)th column of \(J_f(a)\) is \(Df(a)(e_j)\), written in standard coordinates.

Proof. Let \(L=Df(a)\). By the theorem on directional derivatives from differentiability, the directional derivative in direction \(e_j\) is \(L(e_j)\). The Coordinate-Slice Characterization from the tutorial on partial derivatives identifies that directional derivative with the vector of partial derivatives of the component functions in coordinate \(j\). Thus the \(i\)th coordinate of \(L(e_j)\) is \(\partial_j f_i(a)\). This says exactly that the \(j\)th column of \(J_f(a)\) is \(L(e_j)\).

Now write any \(h=(h_1,\ldots,h_m)\) as \(h=\sum_{j=1}^m h_j e_j\). By linearity of \(L\), $$ Df(a)(h) =L\left(\sum_{j=1}^m h_j e_j\right) =\sum_{j=1}^m h_j L(e_j). $$ Matrix multiplication by \(J_f(a)\) forms precisely this linear combination of its columns. Therefore \(Df(a)(h)=J_f(a)h\), as claimed. \(\square\)

The matrix has \(m\) columns because there are \(m\) input coordinates, and \(n\) rows because each response has \(n\) output coordinates. Consequently, multiplying an \(n\)-by-\(m\) Jacobian matrix by an \(m\)-coordinate displacement produces an \(n\)-coordinate predicted output change.

Worked Example: A Map from the Plane to the Plane

Define \(f:\mathbb{R}^2\to\mathbb{R}^2\) by \(f(x,y)=(xe^y,x^2-y)\), and consider \(a=(1,0)\). Its component functions are \(f_1(x,y)=xe^y\) and \(f_2(x,y)=x^2-y\). Their partial derivatives are \(\partial_1f_1=e^y\), \(\partial_2f_1=xe^y\), \(\partial_1f_2=2x\), and \(\partial_2f_2=-1\). Substituting \(a=(1,0)\) gives $$ J_f(1,0)= \begin{pmatrix} 1 & 1\\ 2 & -1 \end{pmatrix}. $$

For a displacement \(h=(u,v)\), the matrix acts by $$ J_f(1,0) \begin{pmatrix}u\\v\end{pmatrix} = \begin{pmatrix}u+v\\2u-v\end{pmatrix}. $$ Thus the first-order predicted change in the two output coordinates is \(u+v\) and \(2u-v\), respectively. In particular, the first column \((1,2)^\mathsf{T}\) is the response to an input displacement in the \(x\)-direction, and the second column \((1,-1)^\mathsf{T}\) is the response to an input displacement in the \(y\)-direction.

Reading Rows, Columns, and Special Cases

A column answers the question, “What first-order output change results from moving in this one input coordinate?” A row answers a different question: “How does this one output component respond to changes in all input coordinates?” These are complementary ways to read the same matrix.

When \(n=1\), the output is scalar-valued, so the Jacobian has one row and \(m\) columns. It records the linear functional that approximates the change in the scalar function. When \(m=1\), the input has one coordinate, so the Jacobian is an \(n\)-by-\(1\) column: its entries give the rates of change of the output components along that input coordinate. These are dimension conventions, not different definitions.

Worked Example: A Scalar-Valued Function

Let \(g:\mathbb{R}^3\to\mathbb{R}\) be \(g(x,y,z)=x^2y+yz\), and take \(a=(2,-1,3)\). The partial derivatives are \(\partial_1g=2xy\), \(\partial_2g=x^2+z\), and \(\partial_3g=y\). Evaluating each one at \(a\) gives \(-4\), \(7\), and \(-1\), respectively. Therefore the Jacobian is a one-row matrix: $$ J_g(2,-1,3)= \begin{pmatrix}-4&7&-1\end{pmatrix}. $$

For \(h=(u,v,w)\), its action is $$ J_g(2,-1,3) \begin{pmatrix}u\\v\\w\end{pmatrix} =-4u+7v-w. $$ For example, the displacement \(h=(1,0,-2)\) has predicted first-order change \(-4(1)+7(0)-(-2)=-2\). The Jacobian is a row because there is one output coordinate, even though the input has three coordinates.

Worked Example: A Map from Three Inputs to Two Outputs

Define \(G:\mathbb{R}^3\to\mathbb{R}^2\) by \(G(x,y,z)=(xy+z^2,xz-y^2)\), and evaluate its Jacobian at \(a=(1,-1,2)\). For the first component, the partial derivatives in input-coordinate order are \(y,x,2z\), which become \(-1,1,4\) at \(a\). For the second component, they are \(z,-2y,x\), which become \(2,2,1\). Hence $$ J_G(1,-1,2)= \begin{pmatrix} -1&1&4\\ 2&2&1 \end{pmatrix}. $$

For the displacement \(h=(u,v,w)\), the predicted output change is $$ J_G(1,-1,2) \begin{pmatrix}u\\v\\w\end{pmatrix} = \begin{pmatrix}-u+v+4w\\2u+2v+w\end{pmatrix}. $$ For instance, if \(h=(1,0,-1)\), then the matrix gives \(\bigl(-1+0-4,\ 2+0-1\bigr)=(-5,1)\). The result has two coordinates because \(G\) has two output components.

A Bound from the Entries of a Matrix

The operator-norm bound from the previous tutorial controls a linear map by its operator norm. For a Jacobian, there is also a direct estimate using its entries. This estimate is often convenient when the entries are known but the exact operator norm is not.

Definition (Frobenius Norm): For an \(n\)-by-\(m\) real matrix \(A=(a_{ij})\), its Frobenius norm is $$ \|A\|_{\mathrm F} = \left(\sum_{i=1}^n\sum_{j=1}^m a_{ij}^2\right)^{1/2}. $$ It is the square root of the sum of the squares of all matrix entries.
Theorem (Frobenius Bound for Matrix Action): For every real \(n\)-by-\(m\) matrix \(A\) and every \(h\in\mathbb{R}^m\), $$ \|Ah\|_2\leq\|A\|_{\mathrm F}\|h\|_2. $$ In particular, \(\|A\|_{\mathrm{op}}\leq\|A\|_{\mathrm F}\).

Proof. Write \(A=(a_{ij})\) and \(h=(h_1,\ldots,h_m)\). The \(i\)th coordinate of \(Ah\) is \(\sum_{j=1}^m a_{ij}h_j\). By the Cauchy–Schwarz Inequality, $$ \left(\sum_{j=1}^m a_{ij}h_j\right)^2 \leq \left(\sum_{j=1}^m a_{ij}^2\right) \left(\sum_{j=1}^m h_j^2\right). $$ Summing this inequality over \(i\) yields $$ \begin{aligned} \|Ah\|_2^2 &=\sum_{i=1}^n\left(\sum_{j=1}^m a_{ij}h_j\right)^2\\ &\leq\sum_{i=1}^n\left(\sum_{j=1}^m a_{ij}^2\right)\|h\|_2^2\\ &=\|A\|_{\mathrm F}^2\|h\|_2^2. \end{aligned} $$ Both sides are nonnegative, so taking square roots proves the first inequality. Applying it to every unit vector \(h\) and taking the supremum in the definition of the operator norm proves \(\|A\|_{\mathrm{op}}\leq\|A\|_{\mathrm F}\). \(\square\)

For a differentiable map, the Matrix Representation of the Derivative allows this estimate to be applied with \(A=J_f(a)\). It gives a bound on the size of every predicted first-order response using only the partial derivatives at the point.

Partial Derivatives Alone Are Not Enough

The Jacobian matrix represents the derivative when the function is differentiable. It is not, by itself, a test that differentiability holds. In particular, existence of all partial derivatives at a point does not guarantee that one linear map approximates the function there in every direction. The remainder condition in the definition of differentiability remains essential.

Worked Example: Partial Derivatives Without Differentiability

Define \(q:\mathbb{R}^2\to\mathbb{R}\) by $$ q(x,y)= \begin{cases} \dfrac{xy}{\sqrt{x^2+y^2}}, & (x,y)\neq(0,0),\\ 0, & (x,y)=(0,0). \end{cases} $$ Along the \(x\)-axis, \(q(t,0)=0\), so \(\partial_1q(0,0)=\lim_{t\to0}(q(t,0)-q(0,0))/t=0\). Along the \(y\)-axis, \(q(0,t)=0\), giving \(\partial_2q(0,0)=0\). Thus both partial derivatives exist and the matrix formed from them would be the one-row zero matrix.

But for \(t\neq0\), substituting \((t,t)\) gives \(q(t,t)=t^2/\sqrt{2t^2}=|t|/\sqrt{2}\). The only possible linear approximation suggested by the partial derivatives is the zero map. Its remainder along this diagonal has size \(|t|/\sqrt{2}\), while the displacement has size \(\sqrt{2}|t|\). Their ratio is $$ \frac{|q(t,t)|}{\|(t,t)\|_2} = \frac{|t|/\sqrt{2}}{\sqrt{2}|t|} =\frac12. $$ This ratio does not tend to zero. Therefore \(q\) is not differentiable at \((0,0)\), despite having both partial derivatives there. The entries exist, but they do not form a Jacobian matrix representing a derivative at that point.

Why the Jacobian View Matters

The Jacobian organizes many partial derivatives into the single linear approximation already provided by the derivative. Its dimensions keep track of the input and output spaces, and matrix multiplication applies the approximation to any displacement at once. This makes it possible to calculate and interpret first-order changes without treating each direction as a separate problem.

A common mistake is to assume that a table of partial derivatives automatically proves differentiability. The table is the Jacobian only when differentiability has been established; partial derivatives at a point, by themselves, do not ensure the required approximation in all directions. The distinction is between collecting coordinate rates and proving that one linear map approximates the function as the displacement shrinks.

Check Your Understanding

Use the definitions and results in this tutorial to answer each question.

  1. For a map from \(\mathbb{R}^m\) to \(\mathbb{R}^n\), how many rows and columns does its Jacobian have, and what does each index represent?
  2. Why is the \(j\)th column of the Jacobian the derivative’s response to the standard basis vector \(e_j\)?
  3. How does matrix multiplication express the derivative’s action on an arbitrary displacement?
  4. What does the Frobenius Bound say about the size of \(Ah\)?
  5. Why do existing partial derivatives at a point not, by themselves, establish differentiability there?