How First-Order Changes Combine
A differentiable function has a linear map that approximates its change near a point. When one function is applied after another, the first function changes the input to the second, and the second function then transforms that change. The multivariable chain rule says that these two linear effects compose: the derivative of the overall function is the derivative of the outer function applied after the derivative of the inner function.
We use the derivative as a linear map, as in the tutorial on Derivative as a Linear Map. When convenient, we represent that map by its Jacobian matrix, as in the Jacobian Matrices tutorial. Matrix order matters: the Jacobian of the inner function acts first on an input displacement, and the Jacobian of the outer function acts on the resulting displacement.
The theorem requires differentiability at the relevant points: \(f\) at \(a\), and \(g\) at the intermediate value \(f(a)\). It does not require either derivative to be continuous near those points. The conclusion is a derivative of the composition at \(a\), not a claim about the behavior of the derivatives throughout a neighborhood.
In matrix form, the derivative of \(f\) is an \(n\)-by-\(m\) matrix, and the derivative of \(g\) is a \(p\)-by-\(n\) matrix. Their product is a \(p\)-by-\(m\) matrix, the correct size for a linear map from \(\mathbb{R}^m\) to \(\mathbb{R}^p\):
This dimension check is a useful safeguard. Reversing the matrices would generally make the product undefined, and even when both products happen to be defined, the order must reflect which map acts first.
Worked Examples: Composing Derivatives
Worked Example: A Composition from the Plane to the Plane
Define \(f:\mathbb{R}^2\to\mathbb{R}^2\) and \(g:\mathbb{R}^2\to\mathbb{R}^2\) by \(f(u,v)=(u^2-v^2,2uv)\) and \(g(x,y)=(e^x\cos y,e^x\sin y)\). At \(a=(1,0)\), direct substitution gives \(f(1,0)=(1,0)\).
The Jacobian matrices are \(J_f(u,v)=\begin{pmatrix}2u&-2v\\2v&2u\end{pmatrix}\) and \(J_g(x,y)=\begin{pmatrix}e^x\cos y&-e^x\sin y\\e^x\sin y&e^x\cos y\end{pmatrix}\). Thus \(J_f(1,0)=\begin{pmatrix}2&0\\0&2\end{pmatrix}\) and \(J_g(f(1,0))=J_g(1,0)=\begin{pmatrix}e&0\\0&e\end{pmatrix}\). The chain rule gives
The intermediate value used in the outer Jacobian is \(f(1,0)=(1,0)\), not the original input \((1,0)\) by coincidence alone: it is the value produced by the inner function. Here the two points happen to be equal, but in general they need not be.
Worked Example: A Scalar Function of a Scalar Inner Quantity
Consider \(H:\mathbb{R}^2\to\mathbb{R}\) defined by \(H(x,y)=\ln(x^2+y^2)\), on the open domain \(U=\{(x,y):x^2+y^2>0\}\). Write \(q(x,y)=x^2+y^2\) and \(g(t)=\ln t\) for \(t>0\), so \(H=g\circ q\). At \(a=(1,2)\), \(q(a)=5\), \(Dq(1,2)\) is the row matrix \(\begin{pmatrix}2&4\end{pmatrix}\), and \(Dg(5)\) is the one-by-one matrix \(\begin{pmatrix}1/5\end{pmatrix}\). Therefore
The chain rule gives the two partial components of the derivative at once. The intermediate value \(5\) is positive, as required for \(g\) to be defined and differentiable there. Direct differentiation also gives \(\partial H/\partial x=2x/(x^2+y^2)\) and \(\partial H/\partial y=2y/(x^2+y^2)\); at \((1,2)\), these are \(2/5\) and \(4/5\), in agreement with the matrix calculation.
Worked Example: Differentiating Along a Parameterized Curve
Let \(\gamma:\mathbb{R}\to\mathbb{R}^2\) be \(\gamma(t)=(t^2,\sin t)\), and let \(F:\mathbb{R}^2\to\mathbb{R}^2\) be \(F(x,y)=(x+y,xy)\). At \(t=0\), \(\gamma(0)=(0,0)\), \(D\gamma(0)=\begin{pmatrix}0\\1\end{pmatrix}\), and \(DF(0,0)=\begin{pmatrix}1&1\\0&0\end{pmatrix}\). Consequently,
For a direct check, \((F\circ\gamma)(t)=(t^2+\sin t,t^2\sin t)\). The derivative at \(0\) of the first component is \(2(0)+\cos(0)=1\), and that of the second is \(2(0)\sin(0)+(0)^2\cos(0)=0\). Thus the composed derivative is indeed \((1,0)\). This calculation shows how the chain rule turns a derivative along a curve into a matrix-vector product.
Affine Changes of Variables
An affine change of variables is a map of the form \(A(z)=b+T(z-c)\), where \(T\) is linear and \(A(c)=b\). Its derivative is the same linear map \(T\) at every point. The next result can be verified directly from the definition of differentiability, so it also provides a useful special case of the chain rule.
Proof. Write \(L=Df(b)\). By differentiability of \(f\) at \(b\), for \(k\) near \(0\) we can write \(f(b+k)=f(b)+L(k)+r(k)\), where \(\|r(k)\|_2/\|k\|_2\to0\) as \(k\to0\), with the remainder taken to be \(0\) at \(k=0\). For an increment \(h\) at \(c\), the affine formula gives \(A(c+h)=b+T(h)\). Hence $$ (f\circ A)(c+h) =f(b)+L(T(h))+r(T(h)). $$ It remains to check that the last term is negligible compared with \(\|h\|_2\). If \(T=0\), then \(T(h)=0\) and \(r(T(h))=0\) for every \(h\), so the claim holds. If \(T\neq0\), the operator-norm bound gives \(\|T(h)\|_2\leq\|T\|_{\mathrm{op}}\|h\|_2\). For \(T(h)\neq0\), $$ \frac{\|r(T(h))\|_2}{\|h\|_2} = \frac{\|r(T(h))\|_2}{\|T(h)\|_2} \frac{\|T(h)\|_2}{\|h\|_2} \leq \|T\|_{\mathrm{op}} \frac{\|r(T(h))\|_2}{\|T(h)\|_2}. $$ As \(h\to0\), \(T(h)\to0\), so the final factor tends to \(0\). When \(T(h)=0\), the remainder is \(0\) as well. Thus \(r(T(h))/\|h\|_2\to0\), and the linear part of the approximation is \(L\circ T\). This proves the stated derivative formula. \(\square\)
This result covers translations, rescalings, and restrictions to affine parameterizations. It also explains why a change of coordinates contributes its own linear factor to a derivative calculation.
Worked Example: Rescaling and Translating the Input
Let \(f:\mathbb{R}^2\to\mathbb{R}\) be \(f(x,y)=x^2+3y\), and set \(A(s,t)=(2+s,1-2t)\). At \((s,t)=(0,0)\), the image is \((2,1)\). The derivative of \(f\) at \((2,1)\) is the row matrix \(\begin{pmatrix}4&3\end{pmatrix}\), while the linear part of \(A\) is \(\begin{pmatrix}1&0\\0&-2\end{pmatrix}\). Therefore the derivative of \(f\circ A\) at \((0,0)\) is
Indeed, direct substitution gives \((f\circ A)(s,t)=(2+s)^2+3(1-2t)=7+4s+s^2-6t\). Its linear terms at \((0,0)\) are \(4s-6t\), confirming the computed derivative.
Matrix Composition and a Common Pitfall
Proof. The Multivariable Chain Rule gives \(D(g\circ f)(a)=Dg(f(a))\circ Df(a)\). By the Matrix Representation of the Derivative, the matrix representing a composition of linear maps is the product of their representing matrices, in the same order as the maps are applied. The matrix for \(Df(a)\) is \(J_f(a)\), and the matrix for \(Dg(f(a))\) is \(J_g(f(a))\). Thus the matrix of the composed derivative is \(J_g(f(a))J_f(a)\). Its size is \(p\)-by-\(m\), as required. \(\square\)
For component functions, the same multiplication says that each entry is a sum over the intermediate coordinates. If \(f=(f_1,\ldots,f_n)\) and \(g=(g_1,\ldots,g_p)\), then for \(1\leq r\leq p\) and \(1\leq j\leq m\),
The intermediate index \(i\) runs over the coordinates of the output of \(f\), which are also the input coordinates of \(g\). This formula is not a license to use the chain rule when differentiability is missing: the hypotheses of the theorem must first be checked. Nor should one evaluate \(Dg\) at \(a\) unless \(a=f(a)\); the outer derivative belongs at the intermediate point \(f(a)\).
The chain rule matters because it lets complicated changes be analyzed in stages. One can differentiate a coordinate transformation, then a physical or geometric map, and multiply the resulting linear maps. In applications, this supports change-of-variable calculations, sensitivity analysis, and differentiation along curves. Keeping track of the intermediate point and the order of the factors is as important as computing the entries themselves.
Check Your Understanding
Use the chain rule and the derivative formulas in this tutorial to answer the following questions.
- For \(f:U\to\mathbb{R}^n\) and \(g:V\to\mathbb{R}^p\), at which point is \(Dg\) evaluated when differentiating \(g\circ f\) at \(a\)?
- If \(J_f(a)\) is \(4\)-by-\(2\) and \(J_g(f(a))\) is \(3\)-by-\(4\), what is the size of the Jacobian matrix of \(g\circ f\) at \(a\)?
- What derivative formula results when the inner function is affine with linear part \(T\)?
- In the component chain rule, what does the intermediate index sum over?
- Why is it necessary to check differentiability of both functions at the relevant points before applying the theorem?