Change Along a Chosen Line
A partial derivative measures change along one coordinate line. But a point in \(\mathbb{R}^m\) can be approached along many other lines, and the rate of change can depend on which line is chosen. A directional derivative records that rate in a specified direction. It is obtained by restricting the multivariable function to a one-variable slice and differentiating that slice.
Let \(U\) be open in \(\mathbb{R}^m\), let \(f:U\to\mathbb{R}\), and fix \(a\in U\). For any vector \(v\in\mathbb{R}^m\), openness ensures that \(a+tv\in U\) for all \(t\) sufficiently close to zero. This includes \(v=0\), though the zero vector does not specify a geometric direction. We allow \(v\) to have any length; when \(v\) is a unit vector, the parameter \(t\) measures signed distance along that direction.
The limit is two-sided: both positive and negative values of \(t\) are considered. For a nonzero \(v\), replacing \(v\) by \(-v\) reverses the orientation of the line. Also, \(v\) need not be a unit vector. Consequently, \(D_v f(a)\) incorporates both the line’s orientation and the speed at which \(a+tv\) traverses it as \(t\) changes.
Directional Derivatives as One-Variable Derivatives
For fixed \(a\) and \(v\), define the slice \(\phi_v(t)=f(a+tv)\) for \(t\) in a sufficiently small open interval about zero. The directional difference quotient is exactly the ordinary difference quotient of \(\phi_v\). This gives a direct way to use one-variable differentiation along a line.
Proof. Since \(U\) is open and \(a\in U\), there is an \(r>0\) such that \(B_r^{(2)}(a)\subseteq U\). If \(v\neq0\), then \[ \|a+tv-a\|_2=|t|\|v\|_2<r \] whenever \(|t|<r/\|v\|_2\). Thus the slice is defined on an open interval about \(0\). If \(v=0\), it is defined for every \(t\). In either case, \(\phi_v(0)=f(a)\), and for every sufficiently small \(t\neq0\), \[ \frac{\phi_v(t)-\phi_v(0)}{t} = \frac{f(a+tv)-f(a)}{t}. \] The two difference quotients are identical wherever they are defined near zero. Therefore one limit exists exactly when the other does, and their values agree. \(\square\)
Taking \(v=e_j\), the \(j\)th standard coordinate vector, gives \(a+te_j\), the coordinate slice used to define \(\partial_j f(a)\) in Partial Derivatives. Thus \(D_{e_j}f(a)=\partial_j f(a)\) whenever that partial derivative exists. Directional derivatives extend the same one-variable idea to lines that need not be coordinate lines.
Worked Example: A Polynomial Along a Noncoordinate Line
Let \(f(x,y)=x^2+3xy-2y^2\), take \(a=(1,-1)\), and choose \(v=(3,4)\). First, \[ f(1,-1)=1+3(1)(-1)-2(-1)^2=1-3-2=-4. \] Along the line \(a+tv=(1+3t,-1+4t)\), direct expansion gives \[ \begin{aligned} f(1+3t,-1+4t) &=(1+3t)^2+3(1+3t)(-1+4t)-2(-1+4t)^2\\ &=(1+6t+9t^2)+3(-1+t+12t^2)-2(1-8t+16t^2)\\ &=-4+25t+13t^2. \end{aligned} \] In particular, the product in the middle line is \((1+3t)(-1+4t)=-1+t+12t^2\). The difference quotient is therefore \[ \frac{f(a+tv)-f(a)}{t} =\frac{(-4+25t+13t^2)-(-4)}{t} =25+13t, \] so \(D_{(3,4)}f(1,-1)=25\). For the unit vector \(u=(3/5,4/5)=v/5\), the rate is \(D_u f(1,-1)=5\), as also follows from the direction-scaling theorem below.
How Rescaling the Direction Changes the Derivative
The length of the vector in the definition matters. Multiplying a direction vector by a positive constant makes the line parameter traverse the same oriented line more quickly; multiplying it by a negative constant also reverses orientation. The following result makes the effect exact.
Proof. If \(c=0\), then \(a+t(cv)=a\) for every \(t\), so the difference quotient defining \(D_{cv}f(a)\) is zero. The formula follows. If \(c\neq0\), then for \(t\neq0\), \[ \frac{f(a+t(cv))-f(a)}{t} = c\,\frac{f(a+(ct)v)-f(a)}{ct}. \] As \(t\to0\), the parameter \(ct\to0\). The quotient on the right tends to \(D_v f(a)\), so the left tends to \(cD_v f(a)\). This proves both existence and the stated value. \(\square\)
In particular, if \(v\neq0\) and \(u=v/\|v\|_2\), then \(v=\|v\|_2u\), and \[ D_v f(a)=\|v\|_2D_u f(a). \] Thus a directional derivative along a unit vector gives the rate per unit distance, while an arbitrary vector scales that rate by its length. The theorem concerns multiplication of a single vector; it does not claim that derivatives add when direction vectors are added.
Worked Example: A Directional Derivative of an Exponential Expression
Let \(f(x,y)=e^{x-y}+xy^2\), let \(a=(0,1)\), and let \(v=(2,-1)\). The line through \(a\) in direction \(v\) is \((2t,1-t)\). Substitution gives \[ f(2t,1-t)=e^{-1+3t}+2t(1-t)^2 =e^{-1}e^{3t}+2t-4t^2+2t^3, \] while \(f(0,1)=e^{-1}\). Hence \[ \frac{f(a+tv)-f(a)}{t} =e^{-1}\frac{e^{3t}-1}{t}+2-4t+2t^2. \] Since \(\lim_{t\to0}(e^{3t}-1)/t=3\), it follows that \[ D_{(2,-1)}f(0,1)=\frac{3}{e}+2. \] For the unit direction \(u=(2/\sqrt{5},-1/\sqrt{5})\), scaling gives \(D_u f(0,1)=(3/e+2)/\sqrt{5}\).
A Chain Rule Along a Line
A directional derivative can also be passed through a differentiable scalar function. The domain condition matters: the outer function must be defined near the value \(f(a)\). Existence of the directional derivative ensures that the values of \(f\) along the relevant slice remain near \(f(a)\) for sufficiently small parameters.
Proof. Write \(\phi(t)=f(a+tv)\), and let \(L=D_v f(a)\). By the Line-Slice Characterization, \(\phi\) is differentiable at zero with \(\phi'(0)=L\). Differentiability implies continuity at zero, so \(\phi(t)\to f(a)\). Since \(I\) is open and contains \(f(a)\), \(\phi(t)\in I\) for all sufficiently small \(t\). Thus the composition is defined along the line near \(t=0\).
Set \(s_t=\phi(t)-f(a)\). For \(s\neq0\), write \[ h(f(a)+s)-h(f(a))=h'(f(a))s+r(s)s, \] where \[ r(s)=\frac{h(f(a)+s)-h(f(a))}{s}-h'(f(a)). \] Differentiability of \(h\) at \(f(a)\) says \(r(s)\to0\) as \(s\to0\). Set \(r(0)=0\); the displayed identity then holds also when \(s=0\). For \(t\neq0\) sufficiently small, this gives \[ \frac{h(\phi(t))-h(\phi(0))}{t} =h'(f(a))\frac{s_t}{t}+r(s_t)\frac{s_t}{t}. \] The first quotient \(s_t/t\) tends to \(L\), so it is bounded near zero. Also \(s_t\to0\), hence \(r(s_t)\to0\). The second term therefore tends to zero, and the first tends to \(h'(f(a))L\). This proves the claimed directional derivative formula. \(\square\)
Worked Example: Applying the Directional Chain Rule
Let \(F(x,y)=\ln(2+x^2+y)\), on the open set where \(2+x^2+y>0\). At \(a=(1,0)\), the logarithm’s argument is \(3\), and choose \(v=(1,2)\). Along the line \(a+tv=(1+t,2t)\), \[ F(1+t,2t)=\ln\bigl(2+(1+t)^2+2t\bigr) =\ln(3+4t+t^2). \] The inner function \(g(x,y)=2+x^2+y\) has slice \(g(1+t,2t)=3+4t+t^2\), so \(D_v g(a)=4\). Applying the directional chain rule with \(h(s)=\ln s\), for which \(h'(3)=1/3\), yields \[ D_v F(a)=\frac{1}{3}\cdot4=\frac{4}{3}. \] Directly, the difference quotient is \[ \frac{\ln(3+4t+t^2)-\ln 3}{t} =\frac{\ln\bigl(1+(4t+t^2)/3\bigr)}{t}, \] whose limit is \(4/3\), confirming the result. The logarithm is defined along the line for sufficiently small \(t\), since \(3+4t+t^2\) tends to \(3>0\).
Directional Derivatives Need Not Depend Linearly on Direction
It is tempting to expect that knowing directional derivatives in coordinate directions determines the derivative in every direction by adding the coordinate contributions. That expectation is valid under stronger hypotheses, but the definition of directional derivative alone does not imply it. In particular, existence of directional derivatives in every direction does not force the map \(v\mapsto D_v f(a)\) to be linear.
Worked Example: Directional Derivatives in Every Direction, Without Linearity
Define \(f:\mathbb{R}^2\to\mathbb{R}\) by \[ f(x,y)= \begin{cases} \dfrac{x^2y}{x^2+y^2},&(x,y)\neq(0,0),\\ 0,&(x,y)=(0,0). \end{cases} \] Fix \(v=(p,q)\neq(0,0)\). For every \(t\neq0\), substitution gives \[ f(tp,tq) =\frac{(tp)^2(tq)}{(tp)^2+(tq)^2} =\frac{t^3p^2q}{t^2(p^2+q^2)} =\frac{tp^2q}{p^2+q^2}. \] Since \(f(0,0)=0\), the difference quotient is \(p^2q/(p^2+q^2)\), independent of \(t\). Therefore \[ D_{(p,q)}f(0,0)=\frac{p^2q}{p^2+q^2}. \] For \(v=(0,0)\), the derivative is zero, so directional derivatives exist in every direction.
Nevertheless, the map of directions is not additive. Its values at \(e_1=(1,0)\) and \(e_2=(0,1)\) are \[ D_{e_1}f(0,0)=\frac{1^2\cdot0}{1^2+0^2}=0, \qquad D_{e_2}f(0,0)=\frac{0^2\cdot1}{0^2+1^2}=0, \] whereas \[ D_{e_1+e_2}f(0,0)=D_{(1,1)}f(0,0) =\frac{1^2\cdot1}{1^2+1^2}=\frac12. \] If the map were additive, its value at \(e_1+e_2\) would be \(0+0=0\), contrary to this calculation. Thus existence in every direction does not itself produce a linear rule for combining directions.
Directional derivatives are useful because they answer a precise question: what is the instantaneous rate of change along this chosen line? They agree with partial derivatives in coordinate directions, scale predictably with the direction vector, and obey the scalar composition rule proved above. But separate line-by-line rates do not, by themselves, capture all the information needed to describe change under arbitrary small perturbations. In particular, the last example warns against inferring a linear first-order approximation solely from the existence of every directional derivative. The next topic will examine the stronger notion that controls changes in all directions together.
Check Your Understanding
Use the definitions and results in this tutorial to answer each question.
- How does the Line-Slice Characterization express a directional derivative as a one-variable derivative?
- If \(D_v f(a)=L\), what is \(D_{-2v}f(a)\), and which theorem justifies the calculation?
- Why must the outer function in the directional chain rule be defined on an open interval containing \(f(a)\)?
- For \(f(x,y)=x^2+3xy-2y^2\) at \((1,-1)\), what is the directional derivative in direction \((3,4)\)?
- In the final example, which three directional derivatives show that \(v\mapsto D_v f(0,0)\) is not additive?