Local Behavior Near a Point
The Taylor formulas developed in the preceding tutorials describe how a function changes near a point. Local extrema are one important setting in which those formulas become especially informative: near a local maximum, the function cannot rise above its value at the point, and near a local minimum, it cannot fall below it. By restricting the function to lines through the point, these inequalities impose conditions on its derivatives.
Throughout, let \(U\subseteq\mathbb{R}^n\) be open, let \(f:U\to\mathbb{R}\), and let \(a\in U\). We use open Euclidean balls \(B_r^{(2)}(a)\) to describe neighborhoods. The openness of \(U\) ensures that sufficiently small balls centered at \(a\) lie in the domain.
The definition concerns only points sufficiently close to \(a\). A global maximum, in contrast, satisfies \(f(x)\leq f(a)\) for every \(x\in U\), and a global minimum satisfies the reverse inequality throughout \(U\). A global extremum is a local extremum when \(U\) is open, but a local extremum need not be global. Also, a local extremum need not be strict: the function may have the same value at other nearby points.
The First-Order Necessary Condition
At an interior local extremum, every directional derivative that exists must vanish. In particular, if the function is differentiable, then its derivative is the zero linear map and its gradient is the zero vector. The line-restriction argument makes clear why both positive and negative directions matter.
Proof. Fix \(v\in\mathbb{R}^n\), and define \(g(t)=f(a+tv)\) for \(t\) sufficiently close to zero. Such a neighborhood of zero exists because \(U\) is open. By differentiability of \(f\) and the directional derivative formula, \(g\) is differentiable at zero and \(g'(0)=Df(a)[v]\).
Suppose first that \(a\) is a local maximum. For all sufficiently small \(t\), \(g(t)\leq g(0)\). If \(t>0\), then \((g(t)-g(0))/t\leq0\). If \(t<0\), then \((g(t)-g(0))/t\geq0\), since division by a negative number reverses the inequality. Both one-sided difference quotients have the same limit \(g'(0)\), so that limit must be both nonpositive and nonnegative. Thus \(g'(0)=0\).
If \(a\) is a local minimum, the inequalities reverse: the quotient is nonnegative for positive \(t\) and nonpositive for negative \(t\). Again, differentiability forces \(g'(0)=0\). Since \(v\) was arbitrary, \(Df(a)[v]=0\) for every \(v\in\mathbb{R}^n\), so \(Df(a)\) is the zero linear map. The Gradient Representation of the Derivative gives \(Df(a)[v]=\nabla f(a)\cdot v\); hence \(\nabla f(a)=0\). \(\square\)
The differentiability hypothesis is essential to this conclusion. A function can have a local extremum and fail to have a derivative there; for example, \(f(x)=|x|\) has a local minimum at zero but is not differentiable at zero. The theorem is a necessary condition, not a guarantee: having zero gradient alone does not imply a local extremum.
Worked Example: A Strict Local Maximum
Consider \(f(x,y)=5-(x+1)^2-4(y-2)^2\) on \(\mathbb{R}^2\). At \(a=(-1,2)\), direct substitution gives \(f(a)=5\). For every \((x,y)\), $$ f(x,y)-f(-1,2)=-(x+1)^2-4(y-2)^2\leq0. $$ Equality holds only if both \(x+1=0\) and \(y-2=0\), which means \((x,y)=(-1,2)\). Thus \(a\) is a strict global maximum and therefore a strict local maximum. The gradient is \(\nabla f(x,y)=(-2(x+1),-8(y-2))\), so \(\nabla f(-1,2)=(0,0)\), as the first-order necessary condition requires. The Hessian matrix is $$ H_f(x,y)= \begin{pmatrix} -2&0\\ 0&-8 \end{pmatrix}. $$ For every nonzero \(v=(v_1,v_2)\), its quadratic form is \(-2v_1^2-8v_2^2<0\), consistent with the function decreasing to second order in every nonzero direction.
Worked Example: A Strict Local Minimum
Let \(f(x,y)=x^2+3y^2\) and \(a=(0,0)\). Then \(f(a)=0\), and for any \((x,y)\neq(0,0)\), $$ f(x,y)=x^2+3y^2\geq x^2+y^2=\|(x,y)\|_2^2>0. $$ Consequently, \(f(x,y)>f(a)\) for every point other than \(a\), so the origin is a strict global minimum. Its gradient is \((2x,6y)\), which vanishes at the origin, and its Hessian is the diagonal matrix with entries \(2\) and \(6\). In particular, the second directional derivative at the origin in direction \(v=(v_1,v_2)\) is \(2v_1^2+6v_2^2\), which is positive whenever \(v\neq0\).
What the Second Derivatives Must Do
At a differentiable local extremum the linear Taylor term vanishes. When second derivatives are continuous near the point, the quadratic term therefore gives the first possible nonzero change in the function. Its sign is constrained by whether the point is a maximum or a minimum.
Proof. A function with continuous second derivatives is differentiable at \(a\). The First-Order Necessary Condition for a Local Extremum therefore gives \(Df(a)=0\). By the Taylor Formula with Peano Remainder, for \(h\) near zero, $$ f(a+h)=f(a)+Df(a)[h]+\frac{1}{2}D^2f(a)[h,h]+r(h), \qquad \frac{r(h)}{\|h\|_2^2}\longrightarrow0 \quad\text{as }h\to0. $$ Fix \(v\neq0\) and take \(h=tv\), where \(t\) is a nonzero real number sufficiently close to zero. Since \(Df(a)=0\) and \(D^2f(a)\) is bilinear, $$ f(a+tv)-f(a)=\frac{t^2}{2}D^2f(a)[v,v]+r(tv). $$ Moreover, \(\frac{r(tv)}{t^2}=\|v\|_2^2\frac{r(tv)}{\|tv\|_2^2}\longrightarrow0\) as \(t\to0\).
If \(a\) is a local maximum, the left side is nonpositive for all sufficiently small \(t\). Dividing by \(t^2>0\) preserves this inequality, and taking the limit as \(t\to0\) gives \(\frac12D^2f(a)[v,v]\leq0\). Hence \(D^2f(a)[v,v]\leq0\). For a local minimum, the left side is nonnegative, and the same division and limit give \(D^2f(a)[v,v]\geq0\). Finally, when \(v=0\), bilinearity gives \(D^2f(a)[0,0]=0\), so both conclusions hold in that case as well. \(\square\)
In matrix notation, \(D^2f(a)[v,v]=v^{\mathsf T}H_f(a)v\), where \(H_f(a)\) is the Hessian matrix. Thus at a local maximum this quadratic form must be nonpositive in every direction; at a local minimum it must be nonnegative in every direction. The condition is necessary, not sufficient. If the Hessian has both positive and negative values on directions, the point cannot be either type of local extremum under these hypotheses. But if the quadratic form is only nonnegative or only nonpositive, it may not determine what the function does in directions where the quadratic term vanishes.
Worked Example: A Saddle Point Is Not a Local Extremum
Let \(f(x,y)=x^2-4y^2\) and \(a=(0,0)\). The gradient at the origin is zero, and the Hessian there is $$ H_f(0,0)= \begin{pmatrix} 2&0\\ 0&-8 \end{pmatrix}. $$ To test the actual local behavior, consider the two lines through the origin. Along \((t,0)\), for every \(t\neq0\), \(f(t,0)=t^2>0=f(0,0)\). Along \((0,t)\), for every \(t\neq0\), \(f(0,t)=-4t^2<0=f(0,0)\). These points can be chosen arbitrarily close to the origin. Every neighborhood therefore contains values both above and below \(f(0,0)\), so the origin is neither a local maximum nor a local minimum. This is the basic saddle-point behavior: different directions have opposite effects.
Worked Example: A Zero Hessian Does Not Decide the Question
Compare the functions \(f(x,y)=x^4+y^4\) and \(g(x,y)=x^4-y^4\), both at the origin. For \(f\), direct calculation gives \(f(0,0)=0\) and \(f(x,y)>0\) whenever \((x,y)\neq(0,0)\). Thus the origin is a strict local minimum. For \(g\), the origin also has value zero, but \(g(t,0)=t^4>0\) and \(g(0,t)=-t^4<0\) whenever \(t\neq0\). The origin is not a local extremum of \(g\).
Both functions have zero gradient and zero Hessian at the origin: their second partial derivatives are multiples of \(x^2\), \(y^2\), or zero, and hence all vanish there. The second-order necessary condition is satisfied by both, but it does not distinguish a strict minimum from a point with values on both sides. Higher-order terms, or direct comparisons, are needed when the quadratic term gives no conclusion.
Using the Conditions Carefully
These results provide a useful sequence of tests. First, verify that the point is interior to the domain and that the function is differentiable there. If it is a local extremum, the gradient must vanish. If continuous second derivatives are available, inspect the Hessian quadratic form in every direction. A positive value in one direction and a negative value in another rules out both kinds of local extremum. If the form has the required one-sided sign, the necessary condition is met, but further analysis may still be needed.
A common mistake is to treat a necessary condition as a sufficient one. The equation \(\nabla f(a)=0\) does not establish an extremum, as the saddle example shows. Likewise, a Hessian that is nonnegative in every direction does not by itself establish a local minimum, as the fourth-degree comparison illustrates. The Taylor expansion explains the limitation: when its linear and quadratic terms vanish in some directions, the remainder or higher-order terms can govern the local behavior.
The interior-point hypothesis also matters. At a boundary point, nearby points in the domain may approach from only some directions, so the positive-and-negative line comparison used in the first-order proof may no longer apply. This tutorial concerns points in open domains; boundary extrema require separate attention to the geometry of the domain.
Check Your Understanding
Use the definitions and necessary conditions in this tutorial to answer the following questions.
- What inequality defines a local maximum, and how does the strict version differ?
- Why does the line-restriction proof of the first-order necessary condition use both positive and negative values of \(t\)?
- At a twice continuously differentiable local minimum, what sign must \(D^2f(a)[v,v]\) have for every direction \(v\)?
- Why does a Hessian with both positive and negative directional values rule out a local extremum?
- What do the two fourth-degree functions in the final worked example show about the limits of a zero Hessian test?