Tutorials › Real Analysis › The Gradient

Multivariable Analysis · Tutorial 790 of 1000

The Gradient

Learn how the gradient packages the partial derivatives of a scalar-valued function into the vector that represents its derivative.

Advanced 9 min read

What You'll Learn

  • Define the gradient from the coordinate partial derivatives of a scalar-valued function
  • Prove that the derivative acts on an increment by taking its dot product with the gradient
  • Relate the Euclidean norm of the gradient to the operator norm of the derivative
  • Compute gradients and use them to describe first-order changes
  • Distinguish existence of partial derivatives from differentiability

A Vector That Represents a Scalar Derivative

For a scalar-valued function on \(\mathbb{R}^m\), the derivative at a point is a linear map from \(\mathbb{R}^m\) to \(\mathbb{R}\). In coordinates, such a map can be represented by a vector using the Euclidean dot product. The vector that represents the derivative is the gradient. It gathers the coordinate partial derivatives into one object and lets us evaluate the first-order change in any input direction with a single dot product.

Throughout, \(U\subseteq\mathbb{R}^m\) is open and \(f:U\to\mathbb{R}\) is scalar-valued. We use the Euclidean inner product \(x\cdot y\) and norm \(\|\cdot\|_2\). As established in the tutorials on Partial Derivatives and Differentiability in \(\mathbb{R}^n\), a partial derivative measures change along one coordinate direction, while differentiability means that a linear map approximates the function's total change to first order.

Definition (Gradient): Suppose the coordinate partial derivatives of \(f\) exist at \(a\in U\). The gradient of \(f\) at \(a\) is the vector $$ \nabla f(a)=\left(\partial_1 f(a),\ldots,\partial_m f(a)\right)\in\mathbb{R}^m. $$ When \(f\) is differentiable at \(a\), this vector represents its derivative through the Euclidean inner product: \(Df(a)(h)=\nabla f(a)\cdot h\) for every \(h\in\mathbb{R}^m\). The representation is established in the theorem below.

The gradient is written as a vector in the input space, even though \(Df(a)\) is a linear map from the input space to \(\mathbb{R}\). These are two ways to encode the same first-order information when the Euclidean inner product has been chosen. The gradient's \(j\)-th coordinate is the partial derivative in the \(j\)-th coordinate direction; its dot product with an arbitrary increment combines all those coordinate contributions.

The Derivative Is the Dot Product with the Gradient

Theorem (Gradient Representation of the Derivative): Let \(U\subseteq\mathbb{R}^m\) be open, let \(f:U\to\mathbb{R}\), and suppose \(f\) is differentiable at \(a\in U\). Then every coordinate partial derivative \(\partial_j f(a)\) exists, and for every \(h\in\mathbb{R}^m\), $$ Df(a)(h)=\nabla f(a)\cdot h. $$

Proof. Differentiability at \(a\) means that there is a linear map \(L:\mathbb{R}^m\to\mathbb{R}\) and a remainder \(r(h)\) such that, for \(h\) near \(0\),

$$ f(a+h)-f(a)=L(h)+r(h), \qquad \lim_{h\to 0,\ h\neq 0}\frac{|r(h)|}{\|h\|_2}=0. $$

Fix a coordinate index \(j\), and let \(e_j\) be the \(j\)-th standard basis vector. Since \(U\) is open, \(a+te_j\in U\) for all sufficiently small \(t\). For nonzero such \(t\), apply the differentiability formula with \(h=te_j\) and divide by \(t\):

$$ \frac{f(a+te_j)-f(a)}{t} = L(e_j)+\frac{r(te_j)}{t}. $$

Because \(\|te_j\|_2=|t|\), the remainder term satisfies \(\left|r(te_j)/t\right|=|r(te_j)|/\|te_j\|_2\to 0\) as \(t\to 0\). Thus the partial derivative exists and \(\partial_j f(a)=L(e_j)\). By the Unique Standard-Coordinate Representation of vectors, \(h=\sum_{j=1}^m h_j e_j\). Linearity of \(L\) now gives

$$ Df(a)(h)=L(h) =\sum_{j=1}^m h_j L(e_j) =\sum_{j=1}^m h_j\,\partial_j f(a) =\nabla f(a)\cdot h. $$

This proves both the existence of the partial derivatives and the asserted representation. \(\square\)

The theorem shows why differentiability is the key hypothesis connecting the gradient to the total first-order change. The gradient's coordinates are obtained by testing the derivative on the standard basis vectors, and linearity then determines its action on every input vector.

Worked Example: Computing a Gradient and Its Derivative

Let \(f:\mathbb{R}^2\to\mathbb{R}\) be \(f(x,y)=x^2y+\sin y\). The coordinate partial derivatives are \(\partial_1 f(x,y)=2xy\) and \(\partial_2 f(x,y)=x^2+\cos y\). Therefore,

$$ \nabla f(x,y)=(2xy,\ x^2+\cos y). $$

At \(a=(1,0)\), this gives \(\nabla f(1,0)=(0,2)\). For an increment \(h=(u,v)\), the derivative is

$$ Df(1,0)(u,v)=(0,2)\cdot(u,v)=2v. $$

For instance, the first-order change in the direction \(h=(3,-1)\) is \(-2\). This is a first-order prediction, not generally the exact change for a finite increment. Differentiability says the error in the linear approximation becomes small compared with the size of the increment as that increment tends to zero.

The Size of the Gradient

The operator norm of \(Df(a)\) measures the largest absolute first-order response to a unit input. For a scalar-valued function, the gradient makes that norm especially simple: it equals the Euclidean length of the gradient. This gives a precise link between a vector's length and the size of the linear map it represents.

Theorem (Operator Norm of a Scalar Derivative): If \(f:U\to\mathbb{R}\) is differentiable at \(a\), then $$ \|Df(a)\|_{\mathrm{op}}=\|\nabla f(a)\|_2. $$

Proof. By the Gradient Representation of the Derivative, \(Df(a)(h)=\nabla f(a)\cdot h\). For every \(h\) with \(\|h\|_2=1\), the Cauchy–Schwarz inequality gives

$$ |Df(a)(h)| = |\nabla f(a)\cdot h| \leq \|\nabla f(a)\|_2\|h\|_2 = \|\nabla f(a)\|_2. $$

Taking the supremum over all unit \(h\) shows \(\|Df(a)\|_{\mathrm{op}}\leq\|\nabla f(a)\|_2\). If \(\nabla f(a)=0\), then \(Df(a)(h)=0\) for every \(h\), so both norms are zero. If \(\nabla f(a)\neq 0\), choose the unit vector \(h=\nabla f(a)/\|\nabla f(a)\|_2\). Then

$$ |Df(a)(h)| = \left|\nabla f(a)\cdot \frac{\nabla f(a)}{\|\nabla f(a)\|_2}\right| = \|\nabla f(a)\|_2. $$

Thus the supremum over unit vectors is at least \(\|\nabla f(a)\|_2\), proving the reverse inequality and hence equality. \(\square\)

The equality depends on the Euclidean norm and inner product being used. It also covers the zero-gradient case: when \(\nabla f(a)=0\), the derivative is the zero linear map and has no nonzero first-order response in any direction.

Worked Example: The Gradient Norm and the Largest Linear Response

Consider \(f(x,y)=x^2+2y^2\) at \(a=(1,1)\). Its gradient is \(\nabla f(x,y)=(2x,4y)\), so \(\nabla f(1,1)=(2,4)\) and \(\|\nabla f(1,1)\|_2=\sqrt{20}=2\sqrt{5}\). The derivative is the linear map

$$ Df(1,1)(u,v)=2u+4v. $$

For the unit vector \(h=(1/\sqrt{5},2/\sqrt{5})\), which equals \(\nabla f(1,1)/\|\nabla f(1,1)\|_2\), direct substitution yields

$$ Df(1,1)(h) = \frac{2}{\sqrt{5}}+\frac{8}{\sqrt{5}} = \frac{10}{\sqrt{5}} = 2\sqrt{5}. $$

For every unit vector \((u,v)\), Cauchy–Schwarz gives \(|2u+4v|\leq\sqrt{20}\sqrt{u^2+v^2}=2\sqrt{5}\). The chosen vector attains this bound, verifying the operator-norm equality in this example.

Using the Gradient for First-Order Changes

The derivative approximation can now be written directly with the gradient. As \(h\to 0\), with \(a+h\in U\), differentiability and the representation theorem give

$$ f(a+h)-f(a)=\nabla f(a)\cdot h+r(h), \qquad \frac{|r(h)|}{\|h\|_2}\longrightarrow 0. $$

The dot product is the linear prediction; the remainder is negligible relative to the size of the input increment. In particular, Cauchy–Schwarz bounds the magnitude of the linear prediction by \(\|\nabla f(a)\|_2\|h\|_2\). The operator-norm theorem shows this is the best possible uniform bound for unit increments.

Worked Example: Checking a First-Order Prediction

Let \(f(x,y)=x^2+3xy-y^2\), let \(a=(1,-1)\), and take the increment \(h=(s,t)\). The gradient is \(\nabla f(x,y)=(2x+3y,3x-2y)\), so \(\nabla f(1,-1)=(-1,5)\). The derivative predicts the linear change \(-s+5t\).

To check the prediction, compute both function values: \(f(1,-1)=1-3-1=-3\), and

$$ \begin{aligned} f(1+s,-1+t) &=(1+s)^2+3(1+s)(-1+t)-(-1+t)^2\\ &=1+2s+s^2-3-3s+3t+3st-1+2t-t^2\\ &=-3-s+5t+s^2+3st-t^2. \end{aligned} $$

Consequently, \(f(a+h)-f(a)=-s+5t+s^2+3st-t^2\). The first two terms are exactly \(\nabla f(a)\cdot h\). For the remainder,

$$ |s^2+3st-t^2| \leq s^2+3|s||t|+t^2 \leq \frac{5}{2}(s^2+t^2) = \frac{5}{2}\|h\|_2^2. $$

Here we used \(2|s||t|\leq s^2+t^2\). Dividing this bound by \(\|h\|_2\) shows that the remainder-to-increment ratio tends to zero as \(h\to0\), in agreement with the first-order approximation.

What the Gradient Does Not Guarantee

The gradient is defined from partial derivatives, but the existence of all partial derivatives at a point does not by itself guarantee differentiability there. Without differentiability, the partial derivatives may fail to combine into a linear approximation, and the gradient need not represent the total change. The representation theorem has differentiability as a hypothesis for this reason.

Worked Example: Partial Derivatives Without Differentiability

Define \(g:\mathbb{R}^2\to\mathbb{R}\) by \(g(0,0)=0\) and \(g(x,y)=xy/(x^2+y^2)\) when \((x,y)\neq(0,0)\). Along either coordinate axis, \(g\) is zero. Therefore both partial derivatives at the origin exist and equal zero, so the vector of partial derivatives is \((0,0)\).

However, for \(t\neq0\),

$$ g(t,t)=\frac{t^2}{2t^2}=\frac12. $$

As \(t\to0\), the points \((t,t)\) approach the origin but the function values remain \(1/2\), not \(0\). Thus \(g\) is not continuous at the origin. Since differentiability implies continuity, \(g\) is not differentiable there. Its partial-derivative vector therefore cannot be used as a derivative representation.

When \(f\) is differentiable at every point of a region, the gradient also gives a convenient way to express a derivative bound: the Operator Norm of a Scalar Derivative theorem identifies \(\|Df(z)\|_{\mathrm{op}}\) with \(\|\nabla f(z)\|_2\). Consequently, the Mean Value Estimate Along a Segment from the previous tutorial applies whenever this gradient norm is bounded along the segment. That estimate controls the total change in \(f\) by the bound times the segment length.

The central distinction is therefore precise: partial derivatives supply the coordinates of a candidate gradient, while differentiability ensures that this vector represents the full derivative. Once that representation is justified, the dot product describes the linear response and the gradient's norm measures the largest such response among unit increments.

Check Your Understanding

Use the definition and results in this tutorial to answer the following questions.

  1. For a differentiable scalar-valued function, how does the gradient determine the derivative applied to an arbitrary vector \(h\)?
  2. Why does differentiability at a point imply that all coordinate partial derivatives exist there?
  3. How does the proof of the operator-norm identity handle the cases where the gradient is zero and nonzero?
  4. What does the counterexample show about the relationship between partial derivatives and differentiability?
  5. If the gradient is bounded in norm along a segment, which result from the previous tutorial can be used to bound the change in the function?