From Tangent Spaces to Constrained Extrema
The previous tutorial identified the tangent space to a regular level set as the kernel of the derivative of its defining map. Its normal space is spanned by the rows of that derivative. These descriptions turn a geometric question—where can a function have an extremum while remaining on a constraint set?—into a first-order condition involving gradients.
The central idea is that, at a constrained local extremum, the objective cannot change to first order in any feasible tangent direction. Consequently, its gradient is perpendicular to the tangent space, so it belongs to the normal space generated by the constraint gradients. This gives the Lagrange multiplier equations. The condition is necessary, not by itself a test that a point is an extremum.
The neighborhood is restricted to \(M\): points off the constraint set are irrelevant to this definition. For a level-set constraint, \(M\) will be the set of all points satisfying the prescribed equations.
First-Order Change Along Feasible Directions
Suppose \(M\) is a regular level set and \(a\in M\). By the Tangent Space to a Regular Level Set Theorem, each vector \(v\in T_aM\) is the velocity at \(a\) of a continuously differentiable curve in \(M\). If \(f\) is differentiable at \(a\), the chain rule lets us measure the first-order change of \(f\) along that curve.
Proof. Fix \(v\in T_aM\). By the definition of tangent vector, there are \(\varepsilon>0\) and a continuously differentiable curve \(\gamma:(-\varepsilon,\varepsilon)\to M\) such that \(\gamma(0)=a\) and \(\gamma'(0)=v\). Since \(a\) is a constrained local extremum, \(f(\gamma(t))\) has a local extremum at \(t=0\): for sufficiently small \(t\), the point \(\gamma(t)\) lies in \(M\) and is close enough to \(a\) for the defining extremum inequality to apply.
Differentiability of \(f\) at \(a\) and differentiability of \(\gamma\) at \(0\) ensure that the composition is differentiable there, by the chain rule. Therefore its derivative must vanish at its one-variable local extremum. Applying the chain rule gives $$ 0=\left.\frac{d}{dt}f(\gamma(t))\right|_{t=0} =Df(a)\gamma'(0)=Df(a)v. $$ Since \(v\) was arbitrary, the condition holds for every tangent vector. \(\square\)
Differentiability of the objective is essential in this argument. Having a constrained extremum alone does not guarantee that the derivative along a feasible curve exists. For instance, \(f(x)=|x|\) has a minimum at \(0\) on \(M=\mathbb{R}\), but along \(\gamma(t)=t\), the composition \(|t|\) is not differentiable at \(0\). The theorem's hypothesis that \(f\) is differentiable at \(a\) prevents this problem.
The Lagrange Multiplier Theorem
Write a system of \(k\) constraints as \(F(x)=c\), where \(F:U\to\mathbb{R}^k\), \(U\subseteq\mathbb{R}^n\) is open, and \(c\in\mathbb{R}^k\). The point \(a\) is regular for this constraint if \(DF(a)\) has rank \(k\). When \(k=1\), write the constraint as \(g(x)=c\); regularity then means \(\nabla g(a)\neq 0\).
Proof. The Constrained Fermat Condition gives \(Df(a)v=0\) for every \(v\in T_aM\). The Tangent Space to a Regular Level Set Theorem identifies \(T_aM=\ker DF(a)\). Since \(f\) is scalar-valued, the Gradient Representation of the Derivative gives \(Df(a)v=\nabla f(a)\cdot v\). Thus \(\nabla f(a)\) is perpendicular to every vector in \(\ker DF(a)\), so $$ \nabla f(a)\in(\ker DF(a))^\perp. $$ By the Normal Space to a Regular Level Set Theorem, \((\ker DF(a))^\perp=\operatorname{im}DF(a)^T\). Membership in this image means precisely that some \(\lambda\in\mathbb{R}^k\) satisfies \(\nabla f(a)=DF(a)^T\lambda\). For \(k=1\), the matrix \(DF(a)\) has the single row \(\nabla g(a)^T\), giving the stated scalar equation. \(\square\)
The multiplier equation says that the objective gradient has no component tangent to the constraint set. For several constraints, the multipliers are the coefficients expressing that gradient as a combination of the constraint gradients. Some conventions put a minus sign before the multipliers; replacing \(\lambda\) by \(-\lambda\) makes the equations equivalent.
Worked Examples
Worked Example: Product on the Unit Circle
Find the points satisfying the Lagrange multiplier equations for \(f(x,y)=xy\) subject to \(g(x,y)=x^2+y^2=1\). The constraint gradient is \(\nabla g(x,y)=(2x,2y)\), which is nonzero at every point of the unit circle. The constraint is therefore regular there, and any constrained local extremum must satisfy $$ (y,x)=\lambda(2x,2y),\qquad x^2+y^2=1. $$
The equations are \(y=2\lambda x\) and \(x=2\lambda y\). Neither coordinate can be zero: if \(x=0\), the constraint gives \(y=\pm1\), contradicting \(y=2\lambda x=0\); if \(y=0\), it gives \(x=\pm1\), contradicting \(x=2\lambda y=0\). Substituting \(y=2\lambda x\) into \(x=2\lambda y\) gives \(x=4\lambda^2x\). Since \(x\neq0\), \(\lambda^2=1/4\).
If \(\lambda=1/2\), then \(y=x\). The constraint gives \(2x^2=1\), so \((x,y)=(1/\sqrt{2},1/\sqrt{2})\) or \((-1/\sqrt{2},-1/\sqrt{2})\); at both points, \(xy=1/2\). If \(\lambda=-1/2\), then \(y=-x\). The two points are \((1/\sqrt{2},-1/\sqrt{2})\) and \((-1/\sqrt{2},1/\sqrt{2})\); at both, \(xy=-1/2\). These are the candidates supplied by the necessary condition.
Worked Example: Two Constraints in Three Variables
Consider \(f(x,y,z)=z\) subject to $$ x^2+y^2+z^2=1,\qquad x+y=0. $$ Use \(F(x,y,z)=(x^2+y^2+z^2,\ x+y)\) and \(c=(1,0)\). At a feasible point, the rows of \(DF\) are \((2x,2y,2z)\) and \((1,1,0)\). They cannot be dependent: dependence would force \(z=0\) and \(x=y\); the second constraint would then give \(x=y=0\), contradicting \(x^2+y^2+z^2=1\). Thus the constraints are regular throughout their feasible set.
The multiplier equations are $$ (0,0,1)=\lambda(2x,2y,2z)+\mu(1,1,0). $$ The first two coordinates give \(0=2\lambda x+\mu\) and \(0=2\lambda y+\mu\). Subtracting yields \(2\lambda(x-y)=0\). The third coordinate gives \(1=2\lambda z\), so \(\lambda\neq0\). Hence \(x=y\), and the constraint \(x+y=0\) forces \(x=y=0\). The sphere constraint now gives \(z^2=1\), so \(z=1\) or \(z=-1\). Both points satisfy the equations: at \((0,0,1)\), take \(\lambda=1/2,\mu=0\); at \((0,0,-1)\), take \(\lambda=-1/2,\mu=0\). The necessary condition identifies these two candidates.
Worked Example: A Stationary Point That Is Not an Extremum
Let \(f(x,y,z)=x^2-y^2\) on the unit sphere \(g(x,y,z)=x^2+y^2+z^2=1\). At \(a=(0,0,1)\), \(\nabla f(a)=(0,0,0)\), so the multiplier equation holds with \(\lambda=0\); also \(\nabla g(a)=(0,0,2)\neq0\), so the constraint is regular.
Nevertheless, \(a\) is not a constrained local extremum. The two curves \(\gamma_1(t)=(\sin t,0,\cos t)\) and \(\gamma_2(t)=(0,\sin t,\cos t)\) lie on the sphere, since \(\sin^2t+\cos^2t=1\), and both pass through \(a\) at \(t=0\). Along them, $$ f(\gamma_1(t))=\sin^2t,\qquad f(\gamma_2(t))=-\sin^2t. $$ For every sufficiently small nonzero \(t\), the first value is positive and the second is negative, while \(f(a)=0\). Thus the multiplier condition is necessary but not sufficient for an extremum.
Regularity and the Limits of the Method
The rank hypothesis ensures that the constraint set has the tangent and normal spaces described in the previous tutorial. Without it, the multiplier conclusion may fail—even when the constrained extremum is genuine. Consider the constraint \(g(x,y)=x^2+y^2=0\). Its feasible set consists only of \((0,0)\). The function \(f(x,y)=x\) has both a constrained local maximum and a constrained local minimum there, because there are no other feasible points. But \(\nabla f(0,0)=(1,0)\) and \(\nabla g(0,0)=(0,0)\), so no scalar \(\lambda\) satisfies \(\nabla f(0,0)=\lambda\nabla g(0,0)\). The constraint is singular at the feasible point, and the theorem does not apply.
A practical use of the method is therefore to check the constraint gradients before solving the equations. At a regular feasible point, the equations produce candidates for constrained extrema; further reasoning is needed to decide whether a candidate is actually a maximum, a minimum, or neither. The saddle example illustrates this distinction, while the singular example shows why regularity is a substantive hypothesis rather than a technical convenience.
Use \(F(x)=c\), and identify the objective \(f\).
At the feasible points under consideration, verify that \(DF\) has rank equal to the number of constraints.
Every constrained local extremum at a regular point must satisfy \(\nabla f=DF^T\lambda\) and \(F(x)=c\).
The equations are necessary conditions; they do not by themselves classify a candidate as an extremum.
Check Your Understanding
Use the definitions and results in this tutorial to answer the following questions.
- Why must the objective be differentiable at the candidate point in the Constrained Fermat Condition?
- For a regular level set \(F(x)=c\), what is the relationship between its tangent space and \(DF(a)\)?
- Write the Lagrange multiplier equation for one regular constraint \(g(x)=c\).
- In the two-constraint example, why could the rows of \(DF\) not be dependent at a feasible point?
- Why does satisfying the multiplier equation not guarantee a constrained local extremum?
- What fails in the constraint \(x^2+y^2=0\) at the origin, and how does that affect the multiplier conclusion?