The Derivative Acts on Displacements
In the previous tutorial, differentiability at \(a\) was defined by the existence of a linear map that approximates the change in the function for every sufficiently small displacement. The derivative is therefore not merely a collection of rates in selected directions. It is one linear rule, applied to the displacement \(h\), that predicts the first-order change \(f(a+h)-f(a)\).
Let \(U\) be open in \(\mathbb{R}^m\), let \(f:U\to\mathbb{R}^n\), and suppose \(f\) is differentiable at \(a\in U\). Write \(Df(a)\) for its unique derivative. The approximation can be written as \(f(a+h)-f(a)=Df(a)(h)+r(h)\), where the remainder satisfies \(\|r(h)\|_2/\|h\|_2\to0\) as \(h\to0\), \(h\neq0\).
The input to \(Df(a)\) is a vector in \(\mathbb{R}^m\), and its output is a vector in \(\mathbb{R}^n\). In particular, \(Df(a)(h)\) is the predicted change in the output caused by the small input change \(h\). Linearity means that combining displacements combines their predicted changes: \(Df(a)(h+k)=Df(a)(h)+Df(a)(k)\), and \(Df(a)(ch)=cDf(a)(h)\) for real \(c\).
The standard basis vectors \(e_1,\ldots,e_m\) determine the action of any linear map: every \(h\in\mathbb{R}^m\) has the representation \(h=\sum_{j=1}^m h_je_j\), so \(Df(a)(h)=\sum_{j=1}^m h_jDf(a)(e_j)\). Thus, once the derivative’s action on those basis displacements is known, its action on every displacement is fixed. The next tutorial will express this information in matrix form; here we focus on the map itself and on quantitative control of its action.
Measuring the Size of a Linear Map
A linear map can send some unit vectors to small vectors and others to large ones. To describe its largest possible effect per unit of input, we use the operator norm. Throughout, the input and output spaces carry their Euclidean norms.
The supremum is finite for every linear map between these finite-dimensional spaces. To see the relevant estimate, write \(h=\sum_{j=1}^m h_je_j\). By linearity and the triangle inequality, $$ \|L(h)\|_2 \leq \sum_{j=1}^m |h_j|\,\|L(e_j)\|_2 \leq \left(\sum_{j=1}^m\|L(e_j)\|_2\right)\|h\|_2. $$ The last inequality uses \(|h_j|\leq\|h\|_2\) for each coordinate. This bound already shows that a linear map cannot make arbitrarily large outputs from unit inputs.
Proof. The coordinate estimate above shows that for every unit vector \(v\), \(\|L(v)\|_2\leq\sum_{j=1}^m\|L(e_j)\|_2\). Hence the supremum defining \(\|L\|_{\mathrm{op}}\) is finite. If \(h=0\), then \(L(h)=0\), so the claimed inequality holds. If \(h\neq0\), then \(v=h/\|h\|_2\) is a unit vector. By linearity, $$ \|L(h)\|_2 = \left\|L\left(\|h\|_2v\right)\right\|_2 = \|h\|_2\|L(v)\|_2 \leq \|h\|_2\|L\|_{\mathrm{op}}. $$ This proves the bound for every \(h\). \(\square\)
The definition also implies \(\|L(v)\|_2\leq\|L\|_{\mathrm{op}}\) for every unit vector \(v\). The theorem extends that unit-input estimate to arbitrary displacements by scaling. The operator norm is the smallest constant that works in the theorem’s inequality: any constant \(C\) with \(\|L(h)\|_2\leq C\|h\|_2\) for all \(h\) must satisfy \(\|L(v)\|_2\leq C\) for every unit \(v\), and thus \(\|L\|_{\mathrm{op}}\leq C\).
Worked Example: The Operator Norm of a Scaling Map
Define \(L:\mathbb{R}^2\to\mathbb{R}^2\) by \(L(u,v)=(3u,-2v)\). For a unit vector, \(u^2+v^2=1\), and $$ \|L(u,v)\|_2^2=9u^2+4v^2 \leq 9u^2+9v^2=9. $$ Therefore \(\|L(u,v)\|_2\leq3\) for every unit vector. The unit vector \((1,0)\) satisfies \(L(1,0)=(3,0)\), whose norm is \(3\). The upper bound is attained, so \(\|L\|_{\mathrm{op}}=3\). For a general displacement, the theorem gives \(\|L(u,v)\|_2\leq3\sqrt{u^2+v^2}\).
The Derivative Controls Local Changes
The operator norm of the derivative gives a quantitative version of the local approximation. Although the remainder is smaller than the displacement by a factor tending to zero, the linear part itself has a fixed bound proportional to the displacement’s size. Together, these facts bound the entire change in the function.
Proof. By differentiability, for the chosen \(\varepsilon>0\) there is a \(\delta>0\) such that if \(0<\|h\|_2<\delta\), then \(\|f(a+h)-f(a)-L(h)\|_2/\|h\|_2<\varepsilon\). Put \(r(h)=f(a+h)-f(a)-L(h)\). The triangle inequality and the Operator-Norm Bound give $$ \begin{aligned} \|f(a+h)-f(a)\|_2 &=\|L(h)+r(h)\|_2\\ &\leq\|L(h)\|_2+\|r(h)\|_2\\ &\leq\|L\|_{\mathrm{op}}\|h\|_2+\varepsilon\|h\|_2\\ &=\bigl(\|L\|_{\mathrm{op}}+\varepsilon\bigr)\|h\|_2. \end{aligned} $$ This holds for every nonzero \(h\) with \(\|h\|_2<\delta\), as required. \(\square\)
This is a local estimate: it applies for displacements sufficiently close to zero, not necessarily for every point in the domain. The constant can be taken arbitrarily close to \(\|Df(a)\|_{\mathrm{op}}\) by choosing a correspondingly small neighborhood. The estimate complements the earlier theorem that differentiability implies continuity: it not only shows that the change tends to zero, but also gives a linear bound on the change near \(a\).
Worked Example: A Nonlinear Map with a Linear First-Order Part
Let \(f:\mathbb{R}^2\to\mathbb{R}^2\) be \(f(x,y)=(x^2+y,xy)\), and consider \(a=(1,2)\). For a displacement \(h=(u,v)\), direct expansion gives $$ \begin{aligned} f(1+u,2+v)-f(1,2) &=\bigl((1+u)^2+(2+v),\,(1+u)(2+v)\bigr)-(3,2)\\ &=(2u+v+u^2,\,2u+v+uv). \end{aligned} $$ The candidate linear map is \(L(u,v)=(2u+v,2u+v)\), and the remainder is \(r(u,v)=(u^2,uv)\). Its size satisfies $$ \|r(u,v)\|_2 =\sqrt{u^4+u^2v^2} =|u|\sqrt{u^2+v^2} \leq u^2+v^2. $$ Consequently, for \(h\neq0\), \(\|r(h)\|_2/\|h\|_2\leq\|h\|_2\to0\). Thus \(f\) is differentiable at \(a\) with derivative \(L\).
To find the operator norm, write \(s=2u+v\), so that \(\|L(u,v)\|_2=\sqrt{2}|2u+v|\). The Cauchy–Schwarz inequality gives \(|2u+v|\leq\sqrt{5}\sqrt{u^2+v^2}\), and hence \(\|L(u,v)\|_2\leq\sqrt{10}\sqrt{u^2+v^2}\). For the unit input \((2,1)/\sqrt{5}\), the value of \(2u+v\) is \(\sqrt{5}\), so equality holds. Therefore \(\|Df(1,2)\|_{\mathrm{op}}=\sqrt{10}\).
Directional Changes Are Governed by the Same Map
The previous tutorial proved that if \(f\) is differentiable at \(a\), then the directional derivative in direction \(v\) is \(Df(a)(v)\). The operator norm sharpens the interpretation: it bounds the size of every directional response, and it controls how that response changes when the direction changes.
Proof. By the theorem on directional derivatives from differentiability, \(D_vf(a)=L(v)\) and \(D_wf(a)=L(w)\). Linearity gives \(L(v)-L(w)=L(v-w)\). Applying the Operator-Norm Bound to \(v-w\) yields $$ \|D_vf(a)-D_wf(a)\|_2 =\|L(v-w)\|_2 \leq\|L\|_{\mathrm{op}}\|v-w\|_2. $$ Taking \(w=0\), and using \(D_0f(a)=L(0)=0\), gives the final bound. \(\square\)
This result says more than that directional derivatives exist. Their values vary in a controlled way as the direction varies, because they are all produced by one bounded linear map. When comparing unit directions, each directional output has norm at most \(\|Df(a)\|_{\mathrm{op}}\). The theorem also rules out abrupt changes in directional response: directions close to one another must have outputs close to one another.
Worked Example: A Scalar Derivative in Three Variables
Consider \(g:\mathbb{R}^3\to\mathbb{R}\), defined by \(g(x,y,z)=x^2+2yz\), at \(a=(1,-1,2)\). Here \(g(a)=1-4=-3\). For \(h=(u,v,w)\), $$ \begin{aligned} g(1+u,-1+v,2+w)-g(1,-1,2) &=(1+u)^2+2(-1+v)(2+w)-(-3)\\ &=2u+4v-2w+u^2+2vw. \end{aligned} $$ Define \(L(u,v,w)=2u+4v-2w\). The remainder satisfies $$ |u^2+2vw| \leq u^2+v^2+w^2 =\|(u,v,w)\|_2^2, $$ because \(2|vw|\leq v^2+w^2\). Dividing by \(\|h\|_2\) shows that the remainder ratio is at most \(\|h\|_2\), which tends to zero. Thus \(Dg(a)(u,v,w)=2u+4v-2w\).
The Cauchy–Schwarz inequality gives \(|L(u,v,w)|\leq\sqrt{24}\sqrt{u^2+v^2+w^2}\), since \(2^2+4^2+(-2)^2=24\). Equality holds for the unit vector \((2,4,-2)/\sqrt{24}\): substituting it into \(L\) gives \(\sqrt{24}\). Hence \(\|Dg(a)\|_{\mathrm{op}}=\sqrt{24}\). For any direction \(v\), the directional derivative is the value of this same linear functional on \(v\), not a separately chosen rate.
Why the Linear-Map View Matters
Thinking of \(Df(a)\) as a map clarifies what information a derivative contains. It accepts any displacement, combines displacements linearly, and predicts the leading change in the output. Its operator norm records the largest first-order output size generated by a unit input. The remainder condition then ensures that this linear prediction becomes accurate relative to the displacement as the displacement shrinks.
A common pitfall is to treat a list of directional rates as if it automatically defined a derivative. Differentiability requires those rates to fit together as one linear map and also requires the approximation error to be small for displacements approaching zero in every possible direction. The derivative-as-a-map viewpoint keeps both requirements visible: the map supplies a consistent linear response, while the remainder condition verifies that it genuinely approximates the function.
Check Your Understanding
Use the definitions and results in this tutorial to answer each question.
- What does the operator norm of a linear map measure?
- Why is the operator norm finite for a linear map between finite-dimensional Euclidean spaces?
- How does the derivative’s operator norm bound the change in a differentiable function near its base point?
- Why do directional derivatives of a differentiable function vary in a controlled way with the direction?
- For a proposed derivative \(L\), what must be shown about the remainder to establish differentiability?