Critical Points as Candidates
The previous tutorial established necessary conditions for local extrema at interior points of open domains: differentiability forces the gradient to vanish. This makes the gradient an efficient first test when investigating where a function might attain a local maximum or minimum. The resulting points are called critical points. They are candidates, not conclusions: a critical point may be an extremum, a point with values on both sides, or part of an entire set of such points.
Throughout, let \(U\subseteq\mathbb{R}^n\) be open and let \(f:U\to\mathbb{R}\). For differentiable functions, critical points are found by solving a system of \(n\) equations. It is also useful to include points where the function is defined but not differentiable, since extrema can occur there too. We use the following standard convention.
The two parts of the definition matter for different reasons. At a differentiable point, a zero gradient means that the first-order linear approximation has no change in any direction. At a point where differentiability fails, that first-order approximation is unavailable, but the point can still be an extremum. If \(f\) is differentiable throughout \(U\), the critical-point search reduces simply to solving \(\nabla f(a)=0\).
The First-Order Necessary Condition for a Local Extremum from the previous tutorial immediately gives a useful candidate principle: every local extremum in \(U\) is a critical point. If the function is differentiable at the extremum, its gradient vanishes; if it is not differentiable there, the other clause of the definition applies. This principle is a one-way implication. It does not say that every critical point is an extremum.
Finding Critical Points
For a differentiable scalar-valued function, the gradient equations are \(\partial_1 f(a)=\cdots=\partial_n f(a)=0\). A complete search must solve all these equations and retain only solutions in the domain. When the function is not differentiable everywhere, one must also examine its nondifferentiability points in the domain. Points outside the domain are not critical points under this definition.
Worked Example: Several Isolated Critical Points
Let \(f(x,y)=x^3-3x+y^2\) on \(\mathbb{R}^2\). Its gradient is \(\nabla f(x,y)=(3x^2-3,2y)\). A critical point must satisfy \(3x^2-3=0\) and \(2y=0\). The first equation is equivalent to \(x^2=1\), so \(x=-1\) or \(x=1\); the second gives \(y=0\). Thus the complete critical set is \(\{(-1,0),(1,0)\}\).
The equations find candidates but do not by themselves classify them. At \((-1,0)\), write \(x=-1+s\). Direct expansion gives \[ f(-1+s,y)=2-3s^2+s^3+y^2. \] For arbitrarily small nonzero \(s\), the choice \(y=0\) gives a value below \(2=f(-1,0)\), while \(s=0\) and nonzero \(y\) gives a value above \(2\). Hence \((-1,0)\) is not a local extremum. At \((1,0)\), writing \(x=1+s\) gives \[ f(1+s,y)=-2+3s^2+s^3+y^2. \] If \(|s|<1\), then \(3s^2+s^3=s^2(3+s)\geq2s^2\geq0\), and \(y^2\geq0\). Equality holds only when \(s=y=0\). Therefore \((1,0)\) is a strict local minimum. The two critical points illustrate why solving the gradient equations is only the first step.
Worked Example: A Critical Set That Is a Circle and a Point
Consider \(f(x,y)=(x^2+y^2-1)^2\). Differentiation gives \[ \nabla f(x,y)=\bigl(4x(x^2+y^2-1),\,4y(x^2+y^2-1)\bigr). \] Both components vanish if \(x^2+y^2-1=0\), which describes the unit circle. If this factor is not zero, both components can vanish only if \(x=y=0\). The origin also satisfies the gradient equations, since the common factor there is \(-1\) while both coordinates are zero. Thus the critical set consists of the unit circle together with the origin.
Every point on the unit circle has function value zero, and \(f(x,y)\geq0\) everywhere, so every point of the circle is a global minimum. At the origin, \(f(0,0)=1\). If \(0<x^2+y^2<2\), then \[ f(x,y)=(1-(x^2+y^2))^2<1, \] because \(-1<1-(x^2+y^2)<1\). Every sufficiently small punctured neighborhood of the origin therefore has function values below \(1\), making the origin a strict local maximum. The critical set need not be finite: it can contain a whole curve, and all points in that curve can share the same critical value.
Critical Points and Lack of Differentiability
The gradient equations alone cannot find critical points at which the derivative does not exist. Such points deserve separate attention, especially when absolute values, roots, or piecewise formulas occur. The definition includes them because differentiability is not necessary for an extremum. But nondifferentiability by itself does not imply an extremum either.
Worked Example: A Nondifferentiable Critical Line
Let \(f(x,y)=|x|+y^2\) on \(\mathbb{R}^2\). At every point with \(x\neq0\), the function is differentiable and \[ \nabla f(x,y)=(1,2y)\quad\text{if }x>0, \qquad \nabla f(x,y)=(-1,2y)\quad\text{if }x<0. \] Neither gradient can be zero because its first component is \(1\) or \(-1\). At a point \((0,y_0)\), hold \(y=y_0\) fixed. The difference quotient in the \(x\)-direction is \[ \frac{f(h,y_0)-f(0,y_0)}{h}=\frac{|h|}{h}. \] It equals \(1\) for \(h>0\) and \(-1\) for \(h<0\), so the partial derivative with respect to \(x\) does not exist there. Every point on the line \(x=0\) is therefore critical under the nondifferentiability clause.
Only the origin is a local minimum. Indeed, \(f(x,y)\geq0=f(0,0)\), with equality only at the origin. For \(y_0\neq0\), points \((0,y_0+t)\) with small \(t\) of opposite signs to \(y_0\) have \((y_0+t)^2<y_0^2\), while points \((x,y_0)\) with \(x\neq0\) have \(|x|+y_0^2>y_0^2\). Thus every neighborhood of \((0,y_0)\) contains values both below and above \(f(0,y_0)\); those critical points are not local extrema.
Changing Coordinates
Critical points can often be found more easily after a translation, rotation, or other invertible affine change of coordinates. The next result confirms that such a change does not lose or invent differentiable critical points: it simply gives them new coordinates. The chain rule supplies the derivative, and invertibility ensures that a nonzero derivative cannot be turned into the zero map by the coordinate change.
Proof. First suppose \(f\) is differentiable at \(T(z)\). The Multivariable Chain Rule applies to \(g=f\circ T\), since \(T\) is differentiable, and gives \[ Dg(z)=Df(T(z))\circ A. \] Thus \(g\) is differentiable at \(z\). Conversely, \(T^{-1}\) is also an affine differentiable map. Since \(f=g\circ T^{-1}\), differentiability of \(g\) at \(z\) implies differentiability of \(f\) at \(T(z)\) by the chain rule.
Now assume the derivatives exist. If \(Df(T(z))=0\), the displayed chain rule gives \(Dg(z)=0\). If \(Dg(z)=0\), then \(Df(T(z))\circ A=0\). Given any \(v\in\mathbb{R}^n\), invertibility of \(A\) gives a vector \(w=A^{-1}v\). Therefore \[ Df(T(z))[v]=Df(T(z))[Aw]=Dg(z)[w]=0. \] Since this holds for every \(v\), \(Df(T(z))=0\). This proves the equivalence for differentiable points. The differentiability equivalence also shows that a point is nondifferentiable for one function exactly when its corresponding point is nondifferentiable for the other. The full critical-point equivalence follows. \(\square\)
This result justifies solving a problem in convenient affine coordinates and then translating the answers back. Invertibility is essential to the argument: if a coordinate map collapses directions, composing with it can conceal changes that the original function has in those directions.
Scalar Transformations and New Critical Points
A different operation is to transform the output values of a function. Let \(\phi\) be a differentiable real-valued function and consider \(\phi\circ f\). The chain rule shows that the gradient is multiplied by the scalar derivative of \(\phi\) evaluated at \(f(a)\). This gives both a way to track existing critical points and a warning: a transformation with zero derivative can create new ones.
Proof. The scalar Chain Rule gives \[ D(\phi\circ f)(a)[v]=\phi'(f(a))Df(a)[v] \] for every \(v\in\mathbb{R}^n\). By the Gradient Representation of the Derivative, \(Df(a)[v]=\nabla f(a)\cdot v\). Consequently, \[ D(\phi\circ f)(a)[v] =\bigl(\phi'(f(a))\nabla f(a)\bigr)\cdot v. \] The vector representing this linear map is unique, so \(\nabla(\phi\circ f)(a)=\phi'(f(a))\nabla f(a)\). If the scalar factor is nonzero, this gradient vanishes exactly when \(\nabla f(a)\) vanishes. \(\square\)
Worked Example: A Transformation Creates a Critical Line
Let \(f(x,y)=x+2y\) and \(\phi(t)=t^2\). The gradient of \(f\) is \((1,2)\), so \(f\) has no critical points. But \(g=\phi\circ f\) is \(g(x,y)=(x+2y)^2\), and the composition rule gives \[ \nabla g(x,y)=2(x+2y)(1,2). \] This gradient vanishes precisely when \(x+2y=0\). Thus every point on that line is critical for \(g\), even though none was critical for \(f\). The scalar factor \(\phi'(f(x,y))=2(x+2y)\) vanishes on the line, which is exactly why the nonzero-factor equivalence does not apply there. In fact \(g\geq0\) everywhere and equals zero on the line, so its points are global minima.
How to Use a Critical-Point Search
A systematic search separates the task of finding candidates from the task of determining their behavior. On an open domain, use the following sequence.
Identify the points under consideration and note whether the domain is open. Boundary points of a non-open optimization region require separate treatment.
Include points in the domain where the derivative does not exist; do not expect gradient equations to detect them.
Where \(f\) is differentiable, solve \(\nabla f(a)=0\) and check that every solution lies in the domain.
Use local comparisons or further derivative information to decide whether a candidate is a maximum, a minimum, or neither.
The main pitfall is to stop after step three and call every solution an extremum. The first-order condition from the previous tutorial only says that differentiable local extrema must be among the points with zero gradient. It does not assert the reverse. Likewise, a point where the derivative fails to exist is a candidate, not an automatic extremum. Later tests using second derivatives can sharpen the analysis, but critical-point identification comes first.
Critical points also depend on the function being studied. An invertible affine reparametrization preserves them in corresponding coordinates, as proved above, while a scalar transformation can add critical points where its derivative vanishes. Keeping these distinctions clear makes the critical-point search useful without giving it more force than its hypotheses warrant.
Check Your Understanding
Use the definition and results in this tutorial to answer the following questions.
- What are the two ways a point in the domain can qualify as a critical point?
- Why does the first-order necessary condition make critical points candidates for extrema rather than guarantee that they are extrema?
- For a differentiable function on an open subset of \(\mathbb{R}^n\), what system of equations is used to find its differentiable critical points?
- Why does an invertible affine change of coordinates preserve critical points?
- In the scalar composition rule, what can happen at a point where \(\phi'(f(a))=0\), even if \(\nabla f(a)\neq0\)?
- Why must nondifferentiability points be checked separately from the gradient equations?