From a Gradient to Directional Rates
The gradient packages the coordinate partial derivatives into a vector, while a directional derivative measures the rate of change along a chosen line. The connection between them is direct when the function is differentiable: the rate in direction \(v\) is the dot product of the gradient with \(v\). This formula lets us read directional behavior from one vector, but it also depends on differentiability, not merely on the existence of partial derivatives.
Throughout, let \(U\subseteq\mathbb{R}^m\) be open and \(f:U\to\mathbb{R}\). We use the Euclidean inner product and norm. Following the convention in the previous tutorial on Directional Derivatives, the direction \(v\) need not be a unit vector. Thus \(D_v f(a)\) measures the rate for the parametrized line \(a+tv\), and scaling \(v\) scales the rate.
Proof. If \(v=0\), then \(a+tv=a\) for every \(t\), so the difference quotient is zero and the formula holds. If \(v\neq 0\), the Theorem (Directional Derivatives from Differentiability) gives \(D_v f(a)=Df(a)(v)\). By the Theorem (Gradient Representation of the Derivative), \(Df(a)(v)=\nabla f(a)\cdot v\). Combining these equalities proves the formula. \(\square\)
The dot product makes two features explicit. First, changing the magnitude of a direction vector changes the rate by the same factor, as in the Theorem (Scaling the Direction Vector). Second, the rate depends on how the direction is aligned with the gradient: a direction pointing with the gradient has a positive rate, while a direction pointing against it has a negative rate. These interpretations are most directly compared using unit direction vectors.
Worked Example: Calculating a Directional Derivative
Let \(f(x,y)=xe^y+y^2\) and take \(a=(1,0)\). The partial derivatives are \(\partial_1 f(x,y)=e^y\) and \(\partial_2 f(x,y)=xe^y+2y\), so \(\nabla f(1,0)=(1,1)\). For \(v=(3,4)\), the formula gives
This is the rate for the parametrization \((1+3t,4t)\), not for a unit-speed parametrization. The corresponding unit vector is \(u=(3/5,4/5)\), and its rate is
The factor of \(5\) between these rates agrees with \(v=5u\) and the scaling rule for directional derivatives.
Greatest Increase and Decrease
A directional rate is a signed number. To compare rates without allowing the choice of a longer direction vector to dominate, restrict attention to unit vectors. The next theorem identifies the largest and smallest signed rates among all unit directions. Its proof uses the Cauchy–Schwarz inequality and checks separately the case in which the gradient vanishes.
Proof. For every unit vector \(u\), the Directional Derivative Formula for the Gradient and the Cauchy–Schwarz inequality give
If \(g\neq0\), the vector \(u=g/\|g\|_2\) has norm \(1\), and \(g\cdot u=\|g\|_2\). The vector \(-u\) also has norm \(1\), and \(g\cdot(-u)=-\|g\|_2\). Thus the upper and lower bounds are attained. If \(g=0\), then \(g\cdot u=0\) for every unit \(u\), so every directional rate is zero. \(\square\)
The gradient's magnitude is therefore the greatest first-order rate of increase per unit distance, while its negative is the greatest first-order rate of decrease. If the gradient is nonzero, its direction specifies the direction of greatest increase, and the opposite direction specifies the direction of greatest decrease. These are statements about the derivative at \(a\); they describe local first-order behavior, not necessarily how the function behaves over a long path.
Worked Example: Finding the Steepest Directions
Let \(f(x,y)=3x-4y+x^2y\) at \(a=(0,0)\). Its partial derivatives are \(\partial_1 f(x,y)=3+2xy\) and \(\partial_2 f(x,y)=-4+x^2\). Hence \(\nabla f(0,0)=(3,-4)\), whose norm is \(5\). The unit direction of greatest increase is \((3/5,-4/5)\), and its rate is
The opposite unit direction \((-3/5,4/5)\) gives rate \(-5\). For comparison, the perpendicular unit direction \((4/5,3/5)\) gives
Thus this perpendicular direction has no first-order change at the origin, even though the function need not remain constant along the line in that direction.
Directions Orthogonal to the Gradient
When \(\nabla f(a)\neq0\), a vector \(v\) orthogonal to the gradient satisfies \(\nabla f(a)\cdot v=0\), so \(D_v f(a)=0\). This says that the linear, first-order contribution to the change vanishes in that direction. It does not say that the function is constant on the line, nor that its actual change is exactly zero for nonzero steps. Terms of smaller order than the step length can still produce a change.
Worked Example: Zero Directional Rate Does Not Mean No Change
Consider \(f(x,y)=x^2+3y^2\) at \(a=(1,1)\). Its gradient is \(\nabla f(x,y)=(2x,6y)\), so \(\nabla f(1,1)=(2,6)\). The vector \(v=(3,-1)\) is orthogonal to this gradient because
Consequently \(D_v f(1,1)=0\). To compare this rate with an actual finite change, take the point \((1,1)+t(3,-1)=(1+3t,1-t)\). Direct substitution gives
Since \(f(1,1)=4\), the change is \(12t^2\), which is generally nonzero for \(t\neq0\). Its difference quotient is \(12t^2/t=12t\), which tends to zero as \(t\to0\). The directional derivative records this limiting rate, not every finite change along the line.
Directional Derivatives Do Not Guarantee a Gradient
For a differentiable function, the directional derivative as a function of \(v\) is linear: it is \(v\mapsto\nabla f(a)\cdot v\). The converse implication is not available merely from the existence of directional derivatives. Even if the two-sided directional derivative exists in every direction, those rates may fail to depend linearly on the direction. In that case, no single gradient can represent them.
Worked Example: Directional Derivatives Without a Gradient Representation
Define \(h:\mathbb{R}^2\to\mathbb{R}\) by \(h(0,0)=0\) and, for \((x,y)\neq(0,0)\), by \(h(x,y)=x^2y/(x^4+y^2)\). Fix a direction \((a,b)\). If \(b\neq0\), then for \(t\neq0\),
The denominator tends to \(b^2>0\), so the two-sided limit exists and equals \(a^2/b\). If \(b=0\), then \(h(ta,0)=0\) for every \(t\): this holds directly when \(ta=0\), and otherwise follows because the numerator is zero. Thus the difference quotient is zero and \(D_{(a,0)}h(0,0)=0\). The zero direction also has derivative zero. Every directional derivative therefore exists.
However, the rates in the directions \((1,0)\) and \((0,1)\) are both zero, whereas the rate in direction \((1,1)\) is \(1\). If a vector \(g=(g_1,g_2)\) represented these rates by dot products, the first two rates would force \(g_1=0\) and \(g_2=0\). Then \(g\cdot(1,1)=0\), contradicting the actual rate \(1\). Hence the directional derivatives cannot be represented by a gradient.
In fact, this function is not continuous at the origin. Along the points \((t,t^2)\), for \(t\neq0\),
These points approach the origin while the function values stay equal to \(1/2\). Since differentiability implies continuity, \(h\) is not differentiable there. This example shows why having directional derivatives in every direction is weaker than having a derivative that gives a linear approximation.
How to Use the Formula Carefully
When applying the Directional Derivative Formula for the Gradient, first check that differentiability has been established at the point. Then compute the gradient there and take its dot product with the specified direction. If the question asks for a rate per unit distance, make sure the direction vector is a unit vector; otherwise the rate scales with the length of that vector. Finally, interpret a zero directional derivative as zero first-order change, not as a claim that the function is constant along the line.
Use the dot-product formula when \(f\) is differentiable at the point in question.
Evaluate the coordinate partial derivatives at the point to obtain \(\nabla f(a)\).
Compute \(\nabla f(a)\cdot v\); normalize \(v\) first only when a unit-direction rate is requested.
Positive, negative, and zero values describe first-order increase, decrease, and no first-order change, respectively.
The gradient gives a complete description of directional rates at a differentiability point: one dot product computes each rate, and the gradient's norm gives the largest absolute rate among unit directions. But directional derivatives considered individually do not by themselves establish differentiability or guarantee that their rates fit together into one linear map. Keeping these two statements distinct is essential when using directional derivatives to analyze multivariable functions.
Check Your Understanding
Use the formula and examples in this tutorial to answer the following questions.
- For a differentiable scalar-valued function, how is \(D_v f(a)\) calculated from the gradient and \(v\)?
- Why can the directional derivative for a nonunit vector differ from the rate for the corresponding unit vector?
- When the gradient is nonzero, which unit directions attain the greatest and smallest signed directional rates?
- What does a zero directional derivative say, and what does it not necessarily say, about changes along that direction?
- In the final worked example, why do the rates in directions \((1,0)\), \((0,1)\), and \((1,1)\) rule out a dot-product representation?