From Directional Rates to a Tangent Plane
The gradient describes how a differentiable function changes in each direction at a point. Those directional rates combine into a linear approximation, and the graph of that linear approximation is a plane. This plane captures the local first-order geometry of the graph: near the point of tangency, the function’s values differ from the plane’s values by an amount that becomes negligible compared with the distance moved in the input.
Let \(U\subseteq\mathbb{R}^m\) be open, let \(f:U\to\mathbb{R}\), and suppose \(f\) is differentiable at \(a\in U\). The graph of \(f\) lies in \(\mathbb{R}^{m+1}\), and the point on the graph above \(a\) is \((a,f(a))\). The tangent plane is an affine hyperplane in \(\mathbb{R}^{m+1}\). For \(m=2\), it is the familiar plane tangent to a surface in three-dimensional space.
When \(m=2\), write \(a=(a_1,a_2)\) and \(x=(x,y)\). The equation becomes \(z=f(a_1,a_2)+f_x(a_1,a_2)(x-a_1)+f_y(a_1,a_2)(y-a_2)\). This is an equation for a plane, not for the original surface. It agrees with the surface at the point of tangency and records the function’s first-order change there.
Proof. By the definition of differentiability, there is a linear map \(L:\mathbb{R}^m\to\mathbb{R}\) and a remainder \(r(h)\) such that \(f(a+h)=f(a)+L(h)+r(h)\), with \(|r(h)|/\|h\|_2\to0\) as \(h\to0\). The Gradient Representation of the Derivative gives \(L(h)=Df(a)(h)=\nabla f(a)\cdot h\). Substitution gives the displayed formula. Because \(U\) is open, \(a+h\in U\) for every sufficiently small \(h\), so these expressions are defined near \(h=0\). \(\square\)
The remainder condition is the essential meaning of “tangent.” It says that the error in the function value, divided by the size of the input displacement, tends to zero. It does not say that the error is always zero, or even that it is bounded by a particular multiple of \(\|h\|_2^2\). That stronger estimate requires additional information.
Worked Example: A Tangent Plane to a Quadratic Graph
Let \(f(x,y)=x^2+xy+2y^2\), and find the tangent plane at the graph point above \(a=(1,-1)\). The function value is \(f(1,-1)=1-1+2=2\). Its partial derivatives are \(f_x(x,y)=2x+y\) and \(f_y(x,y)=x+4y\), so \(\nabla f(1,-1)=(1,-3)\). The tangent plane is therefore
At \((x,y)=(1,-1)\), the right side is \(1+3-2=2\), as required. To see the approximation error explicitly, set \(x=1+s\) and \(y=-1+t\). Direct expansion gives
The tangent plane’s value at this input is \(2+s-3t\), so the error is \(s^2+st+2t^2\). Its absolute value is at most \(s^2+|st|+2t^2\leq \tfrac32s^2+\tfrac52t^2\), since \(|st|\leq(s^2+t^2)/2\). Dividing this bound by \(\sqrt{s^2+t^2}\) gives a quantity that tends to zero as \((s,t)\to(0,0)\). The plane is a first-order approximation, although the quadratic graph does not generally lie in it.
Normal Vectors and Uniqueness
A plane in \(\mathbb{R}^3\) can also be specified by a point on the plane and a normal vector, meaning a vector perpendicular to every direction in the plane. For the graph of \(f\), the tangent directions have the form \((h,\nabla f(a)\cdot h)\). Taking their dot product with \((-\nabla f(a),1)\) gives zero. Thus a normal vector to the graph’s tangent plane is \((-\nabla f(a),1)\).
Proof. A direction vector within the tangent plane has the form \(w=(h,\nabla f(a)\cdot h)\). Its dot product with \(n=(-\nabla f(a),1)\) is \[ n\cdot w=(-\nabla f(a))\cdot h+\nabla f(a)\cdot h=0. \] Thus \(n\) is perpendicular to every direction vector of the plane. It is nonzero because its last coordinate is \(1\). A point \((x,z)\) lies in the plane exactly when its displacement \((x-a,z-f(a))\) from the point of tangency is perpendicular to this normal, giving the stated equation. \(\square\)
The tangent plane is also the unique affine plane that gives a first-order approximation to the graph. The next result makes that statement precise by comparing two possible linear parts.
Proof. Differentiability gives \(f(a+h)-f(a)=\nabla f(a)\cdot h+r(h)\), where \(|r(h)|/\|h\|_2\to0\). Subtract the assumed approximation to obtain \[ (\nabla f(a)\cdot h)-M(h) =\bigl(f(a+h)-f(a)-M(h)\bigr)-r(h). \] Fix any \(v\in\mathbb{R}^m\) with \(v\neq0\), and set \(h=tv\) for nonzero \(t\) tending to zero. By linearity, divide the equality by \(t\). The absolute value of the first error term after division is \[ \frac{|f(a+tv)-f(a)-M(tv)|}{|t|} =\|v\|_2\frac{|f(a+tv)-f(a)-M(tv)|}{\|tv\|_2}, \] which tends to zero. Likewise, \(|r(tv)|/|t|\to0\). Hence \(\nabla f(a)\cdot v-M(v)=0\). This holds for every nonzero \(v\), and it holds for \(v=0\) by linearity. Therefore the two linear maps agree everywhere. \(\square\)
Worked Example: Finding a Plane and Its Normal
Let \(f(x,y)=e^{x-y}+xy\) at \(a=(0,0)\). The function value is \(f(0,0)=1\). Its partial derivatives are \(f_x(x,y)=e^{x-y}+y\) and \(f_y(x,y)=-e^{x-y}+x\). Thus \(\nabla f(0,0)=(1,-1)\), and the tangent plane is
The normal vector supplied by the theorem is \((-1,1,1)\). Indeed, the plane’s direction vectors are \((h,k,h-k)\), and \[ (-1,1,1)\cdot(h,k,h-k)=-h+k+h-k=0. \] The point \((0,0,1)\) lies on the plane because substituting \(x=0\), \(y=0\) gives \(z=1\). These checks verify both the point of tangency and the perpendicularity of the proposed normal.
Parametrized Surfaces
Some surfaces are more naturally described by two parameters than as graphs of one coordinate over the other two. Let \(V\subseteq\mathbb{R}^2\) be open and \(r:V\to\mathbb{R}^3\) be differentiable. At \(q=(u_0,v_0)\), the derivative sends a parameter displacement \(k=(s,t)\) to a vector in \(\mathbb{R}^3\). The two columns of this derivative are the parameter-direction vectors \(r_u(q)\) and \(r_v(q)\).
Proof. The definition of differentiability for a vector-valued function gives a linear map \(Dr(q):\mathbb{R}^2\to\mathbb{R}^3\) and a remainder \(\rho(k)\) such that \(r(q+k)=r(q)+Dr(q)k+\rho(k)\), with \(\|\rho(k)\|_2/\|k\|_2\to0\). For \(k=(s,t)\), linearity gives \(Dr(q)k=s\,r_u(q)+t\,r_v(q)\), where these vectors are the columns of the derivative. This lies in their span, which is the direction space of the tangent plane when they are independent. The remainder condition shows that the discrepancy from this plane-based first-order model is negligible relative to the parameter displacement. \(\square\)
Worked Example: A Tangent Plane from a Parametrization
Consider the surface parametrization \(r(u,v)=(u,v,u^2+2v^2)\) and the parameter point \(q=(1,-1)\). The point on the surface is \(r(1,-1)=(1,-1,3)\). Differentiating gives \(r_u(u,v)=(1,0,2u)\) and \(r_v(u,v)=(0,1,4v)\), so at \(q\) the direction vectors are \((1,0,2)\) and \((0,1,-4)\). They are linearly independent: if \(c(1,0,2)+d(0,1,-4)=(0,0,0)\), the first two coordinates force \(c=d=0\).
A normal vector is their cross product \((-2,4,1)\), as can be checked by dot products: \((-2,4,1)\cdot(1,0,2)=0\) and \((-2,4,1)\cdot(0,1,-4)=4-4=0\). Using the point \( (1,-1,3)\), the plane equation is
Substitution of the point of tangency into the simplified equation gives \(3=2+4-3\), so it lies on the plane. The equation describes the linearized surface at that point; it does not assert that the entire parametrized surface lies in the plane.
Level Surfaces and a Common Pitfall
A surface may also be given as a level set \(F(x,y,z)=c\). At a point \(p\) where \(F\) is differentiable, its first-order change is \(\nabla F(p)\cdot h\). Thus the linearized level equation is \(\nabla F(p)\cdot(x-p)=0\). When \(\nabla F(p)\neq0\), this equation defines a plane with normal vector \(\nabla F(p)\). For a smooth regular level surface, this is its tangent plane.
The linearized equation is a first-order condition, not an exact description of the level set. For example, take \(F(x,y,z)=x^2+2y^2+3z^2\) and the level \(F=6\). At \(p=(1,1,1)\), \(F(p)=6\) and \(\nabla F(p)=(2,4,6)\). The tangent plane is therefore
The level surface is curved: the point \((\sqrt{6},0,0)\) also has \(F=6\), but it does not satisfy this plane equation, since \(\sqrt{6}\neq6\). The plane captures the level surface locally at \(p\), not globally. Moreover, if \(\nabla F(p)=0\), the linearized equation supplies no nonzero normal and does not determine a tangent plane by this method. Checking that the gradient is nonzero is therefore essential when using a level-set equation to identify a regular tangent plane.
For a graph, differentiability is what justifies the tangent-plane approximation. Existing partial derivatives alone do not suffice: they need not provide a linear approximation with an error negligible relative to the displacement. When differentiability is known, the gradient gives the plane directly; for a parametrization, the derivative’s columns give its tangent directions; and for a regular level surface, the gradient of the defining function gives a normal.
Check Your Understanding
Use the definitions and results in this tutorial to answer the following questions.
- What condition on the approximation error expresses that a tangent plane gives a first-order approximation to a graph?
- For the graph of \(f\) at \(a\), why is \((-\nabla f(a),1)\) perpendicular to each tangent direction?
- Why must the two parameter-direction vectors be linearly independent to define a tangent plane to a parametrized surface?
- What is the tangent-plane equation for a regular level surface \(F=c\) at \(p\), and what vector is normal to it?
- Why does a tangent plane generally not describe the whole surface exactly?