Tutorials › Real Analysis › Proof of the Intermediate Value Theorem

Connectedness · Tutorial 312 of 1000

Proof of the Intermediate Value Theorem

See how the least-upper-bound property proves the Intermediate Value Theorem, and how bisection turns that existence proof into increasingly precise root brackets.

Intermediate 10 min read

What You'll Learn

  • Prove the Intermediate Value Theorem using the supremum property of the real numbers
  • Identify how continuity forces a supremum boundary to have the target function value
  • Handle both increasing and decreasing endpoint-value configurations
  • Prove that bisection produces nested intervals containing a root
  • Use interval lengths to obtain arbitrarily precise root brackets

A Proof from Completeness

The Intermediate Value Theorem was established in the previous tutorial as a consequence of connectedness. Here we prove it by a different route: the least-upper-bound property of \(\mathbb{R}\), together with continuity. This proof shows exactly where completeness enters. When a continuous function starts below a target and ends above it, the supremum of the inputs where the function is still below the target must be an input where the function equals the target.

We first isolate the boundary argument used in the proof. Continuity controls function values near a point; the supremum property supplies points from the relevant set arbitrarily close to that boundary. Together, these facts prevent the function value at the boundary from lying strictly to either side of the target.

Lemma (Supremum Boundary Lemma): Let \(S\subseteq\mathbb{R}\) be nonempty and bounded above, and let \(c=\sup S\). Suppose \(g\) is continuous at \(c\), \(g(x)\leq 0\) for every \(x\in S\), and there is a sequence \((z_n)\) with \(z_n>c\), \(z_n\to c\), and \(g(z_n)\geq 0\) for every \(n\). Then \(g(c)=0\).

Proof. By the definition of supremum, for every positive integer \(n\) there is \(x_n\in S\) such that \(c-1/n<x_n\leq c\). Otherwise \(c-1/n\) would be an upper bound for \(S\), contradicting \(c=\sup S\). Thus \(x_n\to c\). Continuity at \(c\) gives \(g(x_n)\to g(c)\) and \(g(z_n)\to g(c)\). Since \(g(x_n)\leq0\) for every \(n\), the limit satisfies \(g(c)\leq0\). Since \(g(z_n)\geq0\) for every \(n\), the same limit satisfies \(g(c)\geq0\). Therefore \(g(c)=0\). \(\square\)

The sequence approaching \(c\) from above is important: it gives the inequality in the other direction. A supremum alone provides points of \(S\) approaching \(c\) from at or below; it does not, by itself, establish that the function value at \(c\) is zero.

Proving the Intermediate Value Theorem

Theorem (Intermediate Value Theorem): Let \(J\subseteq\mathbb{R}\) be an interval, and let \(f:J\to\mathbb{R}\) be continuous. If \(u,v\in J\) and \(y\) lies between \(f(u)\) and \(f(v)\), inclusive, then there exists \(c\in J\) such that \(f(c)=y\).

Proof. If \(y=f(u)\) or \(y=f(v)\), the corresponding input is already a point \(c\) with \(f(c)=y\). It remains to consider the case in which \(y\) is strictly between the two function values. Let \(a=\min\{u,v\}\) and \(b=\max\{u,v\}\). Since \(J\) is an interval and \(a,b\in J\), every point of \([a,b]\) belongs to \(J\). The restriction of \(f\) to \([a,b]\) is therefore continuous.

First suppose \(f(a)<y<f(b)\), and define \(g(x)=f(x)-y\) for \(x\in[a,b]\). Then \(g\) is continuous, \(g(a)<0\), and \(g(b)>0\). Let \(S=\{x\in[a,b]:g(x)<0\}\). This set is nonempty because \(a\in S\), and it is bounded above by \(b\). Set \(c=\sup S\). Since \(g(b)>0\) and \(g\) is continuous at \(b\), there is \(\delta>0\) such that \(g(x)>0\) whenever \(x\in[a,b]\) and \(|x-b|<\delta\). In particular, points of \(S\) cannot lie sufficiently close to \(b\) from the left, so \(c<b\).

For each positive integer \(n\), the supremum property gives \(x_n\in S\) with \(c-1/n<x_n\leq c\). Thus \(x_n\to c\). Also, because \(c<b\), there are points \(z_n\in(c,b]\) with \(z_n\to c\); for example, for sufficiently large \(n\), take \(z_n=c+\min\{1/n,(b-c)/2\}\). No such \(z_n\) can belong to \(S\), since every member of \(S\) is at most its upper bound \(c\). Consequently \(g(z_n)\geq0\). We have \(g(x_n)<0\), so in particular \(g(x_n)\leq0\). The Supremum Boundary Lemma now gives \(g(c)=0\), which means \(f(c)=y\).

If instead \(f(a)>y>f(b)\), apply the argument just given to \(h=-f\) and the target \(-y\). Then \(h(a)<-y<h(b)\), and \(h\) is continuous on \([a,b]\). The result gives \(c\in[a,b]\) such that \(h(c)=-y\), equivalently \(f(c)=y\). In either case \(c\in[a,b]\subseteq J\), proving the theorem. \(\square\)

This proof uses no compactness theorem or connectedness theorem as a step in the argument. Its key ingredients are the completeness of the real numbers, through the existence of \(\sup S\), and continuity, through control of function values near \(c\). Connectedness gives another explanation of the same conclusion, but the supremum proof makes the real-line boundary mechanism explicit.

Worked Example: A Root of a Cubic

Consider \(f(x)=x^3+x-1\) on \([0,1]\). The function is continuous, and direct substitution gives \(f(0)=0^3+0-1=-1\) and \(f(1)=1^3+1-1=1\). Thus \(0\) lies strictly between the endpoint values. The Intermediate Value Theorem, proved above, gives \(c\in[0,1]\) such that \(f(c)=0\), or \(c^3+c-1=0\).

Neither endpoint is a root, since \(-1\ne0\) and \(1\ne0\), so \(c\in(0,1)\). The supremum proof can be applied directly to \(g(x)=x^3+x-1\): take the supremum of the inputs in \([0,1]\) where \(g(x)<0\). Continuity forces the function value at that boundary to be zero. No formula for \(c\) is needed for this existence conclusion.

Worked Example: A Decreasing Function Still Takes Intermediate Values

Let \(f(x)=2-x^2\) on \([1,2]\), and consider the target \(y=-1\). Substitution gives \(f(1)=2-1^2=1\) and \(f(2)=2-2^2=-2\), so \(-1\) lies strictly between the endpoint values, in the decreasing order. The theorem gives \(c\in[1,2]\) with \(2-c^2=-1\). Rearranging yields \(c^2=3\).

This is the reverse-sign case in the proof. Equivalently, \(h=-f\) has \(h(1)=-1<1<h(2)=2\); applying the increasing-endpoint-value argument to \(h\) gives \(h(c)=1\), and hence \(f(c)=-1\). The proof does not require the function to be increasing; it handles either order of the endpoint values.

Bisection Gives Arbitrarily Tight Root Brackets

The supremum proof establishes existence, but it does not specify how to locate the root. A related argument gives a systematic approximation procedure. Begin with an interval whose endpoint values have strictly opposite signs. At each stage, test the midpoint and keep a half-interval whose endpoint values still have opposite signs. If a midpoint is itself a root, the procedure stops; otherwise the intervals become arbitrarily short.

Theorem (Bisection Root-Bracketing Theorem): Let \(f:[a,b]\to\mathbb{R}\) be continuous, where \(a<b\), and suppose \(f(a)<0<f(b)\). Then either a midpoint chosen during repeated bisection is a root, or there are nested closed intervals \([a_n,b_n]\subseteq[a,b]\) such that \(f(a_n)<0<f(b_n)\), each interval contains a root, and \(b_n-a_n=(b-a)/2^n\). In particular, for every \(\varepsilon>0\), a root is contained in an interval of length less than \(\varepsilon\).

Proof. Set \(a_0=a\) and \(b_0=b\). At any stage \(n\), suppose \(f(a_n)<0<f(b_n)\). Let \(m_n=(a_n+b_n)/2\). If \(f(m_n)=0\), a root has been found and the procedure stops. If \(f(m_n)<0\), set \(a_{n+1}=m_n\) and \(b_{n+1}=b_n\); then \(f(a_{n+1})<0<f(b_{n+1})\). If \(f(m_n)>0\), set \(a_{n+1}=a_n\) and \(b_{n+1}=m_n\); the same strict sign inequalities hold. These are all possibilities for \(f(m_n)\) if it is not zero.

If the procedure never stops, the intervals are nonempty, closed, and nested, and their lengths satisfy \(b_{n+1}-a_{n+1}=(b_n-a_n)/2\). Induction gives \(b_n-a_n=(b-a)/2^n\), which tends to zero. By the Nested Interval Theorem, there is a point \(c\) belonging to every \([a_n,b_n]\). Since the interval lengths tend to zero, \(a_n\to c\) and \(b_n\to c\). Continuity of \(f\) at \(c\) and the endpoint inequalities give \(f(c)\leq0\) from \(f(a_n)<0\), and \(f(c)\geq0\) from \(f(b_n)>0\). Hence \(f(c)=0\). Finally, because \((b-a)/2^n\to0\), for every \(\varepsilon>0\) some \(n\) has \(b_n-a_n<\varepsilon\). The interval \([a_n,b_n]\) then gives the claimed root bracket. \(\square\)

Worked Example: Three Bisections for \(x^2-2\)

Let \(f(x)=x^2-2\) on \([1,2]\). The endpoint values are \(f(1)=1^2-2=-1\) and \(f(2)=2^2-2=2\), so the bisection theorem applies. The first midpoint is \(3/2\), and \(f(3/2)=(3/2)^2-2=9/4-8/4=1/4>0\). Keep \([1,3/2]\), whose endpoint values are negative and positive.

The next midpoint is \(5/4\), and \(f(5/4)=(5/4)^2-2=25/16-32/16=-7/16<0\). Keep \([5/4,3/2]\). Its midpoint is \(11/8\), and \(f(11/8)=(11/8)^2-2=121/64-128/64=-7/64<0\). Keep \([11/8,3/2]\). After these three bisections, the interval length is \(1/8\), and the theorem guarantees that it contains a root. Further bisections make the bracket as short as desired.

Reading the Proof Correctly

The supremum argument is an existence proof, not an explicit root-finding formula. The set \(S\) records inputs where the function remains below the target; its supremum is a boundary between inputs in \(S\) and inputs beyond \(S\). Continuity prevents the function from jumping across the target at that boundary. Bisection adds a way to narrow the location of a root, but it still does not generally produce an exact expression for the root.

A frequent gap in informal proofs is to say, “take the supremum, so the function equals the target.” The equality does not follow from the supremum property alone. The proof must use continuity and must establish the relevant behavior on both sides of the boundary. The Supremum Boundary Lemma makes those two sides explicit. In bisection, a different common omission is to keep halving without checking the midpoint sign. The choice of half-interval is justified precisely because its endpoint values must retain opposite signs.

Check Your Understanding

Use the supremum and bisection arguments to answer the following questions.

  1. In the Intermediate Value Theorem proof, why is the set \(S=\{x\in[a,b]:f(x)-y<0\}\) nonempty and bounded above?
  2. Why does the supremum property give points of \(S\) arbitrarily close to \(c=\sup S\) from at or below?
  3. In the decreasing-endpoint case, why can the proof be applied to \(-f\) and \(-y\)?
  4. If a bisection midpoint has function value zero, what should the procedure do?
  5. After \(n\) bisections, what is the length of the retained interval, and why can it be made less than any prescribed positive number?