From Derivatives to Estimates on a Segment
The Multivariable Chain Rule lets us differentiate a function after restricting its input to a line. This simple restriction turns a question about changes in \(\mathbb{R}^m\) into a one-variable question: how much can the function change as we move from one endpoint of a segment to the other? The resulting mean value estimates are useful even when the derivative is not constant. A bound on the derivative along the segment gives a bound on the total change.
We use the Euclidean norm and the operator norm for linear maps established earlier in this course. If \(f:U\to\mathbb{R}^n\) is differentiable at \(z\), then \(Df(z)\) is a linear map from \(\mathbb{R}^m\) to \(\mathbb{R}^n\), and \(\|Df(z)(h)\|_2\leq\|Df(z)\|_{\mathrm{op}}\|h\|_2\). The estimate below uses a bound on this operator norm along the segment, not merely at its endpoints.
Proof. Put \(h=y-x\), and define the path \(\gamma(t)=x+th\) for \(0\leq t\leq1\). The segment hypothesis ensures that every \(\gamma(t)\) lies in \(U\). By the Multivariable Chain Rule, the function \(t\mapsto f(\gamma(t))\) is differentiable for \(0<t<1\), with derivative \(Df(\gamma(t))(h)\). It is continuous on \([0,1]\), because differentiability of \(f\) implies continuity and \(\gamma\) is continuous.
First suppose \(n=1\) and \(x\neq y\). The ordinary one-variable Mean Value Theorem, applied to \(t\mapsto f(\gamma(t))\), gives a \(t_0\in(0,1)\) such that
Taking absolute values and using the derivative bound gives \(|f(y)-f(x)|\leq M\|h\|_2\), which is the asserted estimate when \(n=1\). If \(x=y\), the estimate holds because both sides are zero.
Now suppose \(n\) is any positive integer. If \(f(y)=f(x)\), the estimate holds immediately. Otherwise, define \(u=(f(y)-f(x))/\|f(y)-f(x)\|_2\), so \(\|u\|_2=1\). Apply the one-variable Mean Value Theorem to the scalar function \(\psi(t)=u\cdot f(\gamma(t))\). For some \(t_0\in(0,1)\),
By the Cauchy–Schwarz inequality and the operator-norm bound,
Since \(h=y-x\), this proves the estimate. \(\square\)
The scalar conclusion is an exact formula: one point on the segment gives the full change in the function. For a vector-valued function, the proof instead selects a scalar projection in the direction of the total change. It gives a norm bound, but it does not generally produce one point where the vector difference equals the derivative applied to the segment direction.
Worked Examples: Using the Segment Estimate
Worked Example: An Exact Scalar Mean Value Point
Let \(f:\mathbb{R}^2\to\mathbb{R}\) be \(f(x,y)=x^2+xy+2y^2\). Take \(a=(1,-1)\) and \(b=(2,1)\), so \(b-a=(1,2)\). The endpoint values are \(f(a)=2\) and \(f(b)=8\), hence \(f(b)-f(a)=6\).
Parameterize the segment by \(\gamma(t)=(1+t,-1+2t)\). Substitution gives
Its derivative with respect to \(t\) is \(-5+22t\). At \(t_0=1/2\), this derivative is \(6\), exactly the endpoint difference. The corresponding point is \(c=(3/2,0)\). Directly, \(Df(x,y)(u,v)=(2x+y)u+(x+4y)v\), so \(Df(c)(1,2)=3(1)+\frac32(2)=6\). This verifies the scalar mean value formula at the point found along the segment.
Worked Example: A Vector Estimate Without an Exact Formula
Consider \(f:\mathbb{R}\to\mathbb{R}^2\) given by \(f(t)=(t^2,t^3)\) on the segment from \(0\) to \(1\). Its derivative is \(Df(t)(h)=(2th,3t^2h)\), whose operator norm is \(\sqrt{4t^2+9t^4}\leq\sqrt{13}\) for \(0\leq t\leq1\). The Mean Value Estimate therefore gives
There is no \(t\in(0,1)\) for which \(f(1)-f(0)=Df(t)(1)\): that equality would require both \(2t=1\) and \(3t^2=1\). The first equation forces \(t=1/2\), but then \(3t^2=3/4\neq1\). This shows why the general vector conclusion is an estimate rather than an exact mean value formula.
Worked Example: A Lipschitz Bound on a Square
Let \(f:\mathbb{R}^2\to\mathbb{R}^2\) be \(f(x,y)=(x^2+y,xy)\), and consider the square \(Q=[-1,1]^2\). Its Jacobian matrix is
For \((x,y)\in Q\), the square of the Frobenius norm is \(4x^2+1+y^2+x^2=5x^2+y^2+1\leq7\). The Frobenius bound for matrix action therefore gives \(\|Df(x,y)\|_{\mathrm{op}}\leq\sqrt{7}\). Since \(Q\) is convex, the segment between any two of its points stays in \(Q\). Thus, for all \(p,q\in Q\),
For example, \(p=(1,0)\) and \(q=(-1,1)\) give \(f(p)=(1,0)\), \(f(q)=(2,-1)\), and \(\|f(p)-f(q)\|_2=\sqrt{2}\). The input distance is \(\sqrt{5}\), and the estimate holds since \(\sqrt{2}\leq\sqrt{35}\).
Consequences for Lipschitz Bounds and Linear Approximation
A function is Lipschitz with constant \(M\) on a set if the distance between any two output values is at most \(M\) times the distance between the corresponding inputs. The segment estimate gives a convenient sufficient condition: on a convex open domain, a uniform bound on the derivative gives a Lipschitz bound. Convexity matters because it keeps every segment between points of the domain inside the region where the derivative bound is known.
Proof. For any \(x,y\in U\), convexity implies \([x,y]\subseteq U\). The derivative bound holds at every point of that segment, so the Mean Value Estimate Along a Segment applies and gives the stated inequality. \(\square\)
The same method estimates how far a function differs from a chosen linear map. Fix a linear map \(L:\mathbb{R}^m\to\mathbb{R}^n\), and define \(g(z)=f(z)-L(z)\). Linearity gives \(Dg(z)=Df(z)-L\). Applying the segment estimate to \(g\) bounds the difference between the change in \(f\) and the change predicted by \(L\).
Proof. Define \(g(z)=f(z)-L(z)\). Then \(g\) is differentiable and \(Dg(z)=Df(z)-L\). The hypothesis gives \(\|Dg(z)\|_{\mathrm{op}}\leq\varepsilon\) on the segment. By the Mean Value Estimate,
Since \(g(y)-g(x)=f(y)-f(x)-L(y-x)\), this is the required bound. \(\square\)
What the Estimate Requires
The derivative bound must hold along the entire segment, not just at the endpoints. Derivatives can vary between endpoints, and the one-variable Mean Value Theorem samples a point in the interior. If a bound is available only on a smaller region, the segment must stay within that region for the estimate to apply.
Nor does differentiability by itself provide a uniform derivative bound on an arbitrary domain. The theorem explicitly assumes such a bound on the segment. A common way to obtain one is to assume that \(Df\) is continuous on a compact convex set: continuity and compactness then ensure that its operator norm is bounded there. This additional continuity assumption is useful for finding a bound, but it is not needed in the segment estimate once the bound is already known.
Finally, do not infer an exact vector-valued mean value formula from the scalar theorem. Projecting onto the direction of \(f(y)-f(x)\) proves the norm inequality, but different components can have their one-variable mean value points at different locations. The vector estimate is the robust conclusion needed for Lipschitz bounds and error estimates.
Check Your Understanding
Use the line-segment argument and its hypotheses to answer the following questions.
- Why is the path \(t\mapsto f(x+t(y-x))\) useful when estimating \(f(y)-f(x)\)?
- In the vector-valued proof, how is the scalar function \(\psi\) chosen, and what does its endpoint difference equal?
- Why must the derivative bound be known along the segment rather than only at \(x\) and \(y\)?
- What role does convexity play in deriving a Lipschitz bound on a domain?
- How does applying the segment estimate to \(g(z)=f(z)-L(z)\) produce a linearization-error bound?