Tutorials › Real Analysis › Proof That Convergent Sequences Are Cauchy

Sequences · Tutorial 212 of 1000

Proof That Convergent Sequences Are Cauchy

See how convergence controls tail pairs, and learn to coordinate several approximation requirements with explicit indices.

Intermediate 9 min read

What You'll Learn

  • Identify the two estimates that make the triangle-inequality argument work
  • Choose one tail index that meets several accuracy requirements
  • Prove a common-tail estimate for finitely many sequences with the same limit
  • Replace an irregular convergence modulus with a nondecreasing one
  • Check explicit indices in reciprocal and alternating-sequence examples

The Proof Idea and Its Quantifiers

The Cauchy property asks for a bound on every pair of sufficiently late terms. Convergence, by contrast, controls each sufficiently late term relative to one fixed number, its limit. The proof that convergent sequences are Cauchy connects these two statements by comparing both terms to that same limit.

Tutorial 209 established the theorem Every Convergent Sequence Has the Cauchy Property, and Tutorial 211 developed its quantitative form using a convergence modulus. We will not repeat that theorem’s proof as a new result. Instead, we will examine the choices that make its proof work, then prove two useful extensions: one for coordinating finitely many sequences with a common limit, and one for making a convergence modulus nondecreasing.

Definition (Convergence Modulus): Suppose \(a_n\to L\). A function \(M:(0,\infty)\to\mathbb{N}_0\) is a convergence modulus for \((a_n)\) and \(L\) if, for every \(\delta>0\) and every \(n\geq M(\delta)\), \(|a_n-L|<\delta\).

The key proof choice is to ask for more accuracy from the limit than the final pairwise estimate requires. If the desired pairwise distance is less than \(\varepsilon\), it is enough to put each term within \(\varepsilon/2\) of \(L\). The triangle inequality then adds two errors, each less than \(\varepsilon/2\). In this way, a single index controls both terms, however far apart their indices may be.

The quantifier order matters. The proof first fixes \(\varepsilon>0\), then chooses an index \(N\) using convergence, and finally verifies the estimate for every \(m,n\geq N\). Choosing an index that works only for one selected pair would not establish the Cauchy property.

1
Fix the required pairwise accuracy.
Let \(\varepsilon>0\) be arbitrary.
2
Allocate the error to two terms.
Use \(\varepsilon/2\) as the allowed distance of each term from the common limit.
3
Choose one tail index.
Convergence supplies an index that works for every term in that tail.
4
Compare an arbitrary pair.
Insert the limit between the terms and apply the triangle inequality.

This is a proof blueprint, not a substitute for its quantifiers: the index must be chosen before the arbitrary pair \(m,n\), and the final strict inequality must follow from the two strict error bounds. A common error is to know only that each term is “near” \(L\) without specifying a tolerance that adds up to the requested \(\varepsilon\).

One Index for Several Requirements

A proof often has several estimates to satisfy at once. For example, one may need every term to be close to a limit and also need a second sequence to be close to that same limit. If each requirement holds beyond its own index, the largest of the finitely many indices works for all of them. The result below formalizes this useful coordination step.

Theorem (Common-Tail Estimate for Finitely Many Sequences): Let \(r\) be a positive integer. Suppose \(x_n^{(j)}\to L\) for each \(j\in\{1,\ldots,r\}\), and choose a convergence modulus \(M_j\) for each sequence. For any \(\delta_1,\ldots,\delta_r>0\), set \(N=\max_{1\leq j\leq r}M_j(\delta_j)\). Then for every \(j\in\{1,\ldots,r\}\) and every \(n\geq N\), \(|x_n^{(j)}-L|<\delta_j\).

Proof. Since \(N\) is the maximum of the listed indices, \(N\geq M_j(\delta_j)\) for every \(j\). If \(n\geq N\), then \(n\geq M_j(\delta_j)\), so the definition of \(M_j\) gives \(|x_n^{(j)}-L|<\delta_j\). This holds for every \(j\), which proves the claim. \(\square\)

The theorem’s conclusion can be used to verify several estimates on one shared tail, rather than repeatedly changing the index. When the requirements involve distances between two sequences, apply the triangle inequality after the shared index has been chosen.

Worked Example: Two Sequences with a Common Limit

Let \(x_n=4+\frac{1}{n+1}\) and \(y_n=4-\frac{2}{n+2}\), for \(n\in\mathbb{N}_0\). Both converge to \(4\). Given \(\delta>0\), valid moduli are \(M_x(\delta)=\lfloor1/\delta\rfloor\) and \(M_y(\delta)=\lfloor2/\delta\rfloor\). For the first, \(M_x(\delta)+1>1/\delta\); hence, for \(n\geq M_x(\delta)\),

$$ |x_n-4|=\frac{1}{n+1} \leq\frac{1}{M_x(\delta)+1} <\delta. $$

For the second, \(M_y(\delta)+2>2/\delta\), and therefore, for \(n\geq M_y(\delta)\),

$$ |y_n-4|=\frac{2}{n+2} \leq\frac{2}{M_y(\delta)+2} <\delta. $$

To make both errors less than \(1/10\), the common-tail theorem gives \(N=\max\{M_x(1/10),M_y(1/10)\}=\max\{10,20\}=20\). Thus, for every \(n\geq20\), both \(|x_n-4|<1/10\) and \(|y_n-4|<1/10\). In particular,

$$ |x_n-y_n| \leq |x_n-4|+|y_n-4| <\frac{1}{10}+\frac{1}{10} =\frac{1}{5}. $$

The shared index is essential: it ensures both estimates hold at the same \(n\). Choosing separate indices and then comparing terms at an index that meets only one of them would not justify the displayed bound.

Making a Convergence Modulus Monotone

A convergence modulus need not increase when the requested accuracy becomes stricter. Its definition requires that it work, but does not require it to be the smallest suitable index or to vary regularly with \(\delta\). Nevertheless, it is often convenient to have a modulus with the natural monotonicity property: a smaller tolerance should not require an earlier index.

Theorem (A Nondecreasing-in-Accuracy Modulus Can Be Chosen): Suppose \(a_n\to L\) and \(M\) is a convergence modulus. There is a convergence modulus \(\widetilde M\) such that \(0<\delta_1\leq\delta_2\) implies \(\widetilde M(\delta_1)\geq\widetilde M(\delta_2)\).

Proof. For each \(\delta>0\), let \(k(\delta)\) be the least nonnegative integer \(k\) such that \(2^{-k}<\delta\). Such an integer exists because \(2^{-k}\to0\), as follows from the theorem on powers below one established earlier in this course. Define

$$ \widetilde M(\delta)=\max_{0\leq j\leq k(\delta)}M(2^{-j}). $$

This is a maximum of a nonempty finite set of nonnegative integers, so it is defined. Since \(\widetilde M(\delta)\geq M(2^{-k(\delta)})\), every \(n\geq\widetilde M(\delta)\) also satisfies \(n\geq M(2^{-k(\delta)})\). The modulus property gives

$$ |a_n-L|<2^{-k(\delta)}<\delta. $$

Thus \(\widetilde M\) is a convergence modulus. Now suppose \(0<\delta_1\leq\delta_2\). Every dyadic number strictly below \(\delta_1\) is also strictly below \(\delta_2\), so the least eligible integer satisfies \(k(\delta_1)\geq k(\delta_2)\). The maximum defining \(\widetilde M(\delta_1)\) is therefore taken over a set of indices containing the set used for \(\widetilde M(\delta_2)\). Hence \(\widetilde M(\delta_1)\geq\widetilde M(\delta_2)\), as required. \(\square\)

Worked Example: A Monotone Modulus for a Reciprocal Sequence

Consider \(a_n=3+\frac{1}{n+1}\), which converges to \(3\). The direct estimate \[ |a_n-3|=\frac{1}{n+1} \] shows that \(M(\delta)=\lfloor1/\delta\rfloor\) is a convergence modulus: if \(n\geq M(\delta)\), then \(n+1\geq M(\delta)+1>1/\delta\), and so \(1/(n+1)<\delta\). This particular modulus already has the desired monotonicity: when \(\delta_1\leq\delta_2\), \(1/\delta_1\geq1/\delta_2\), and therefore \(\lfloor1/\delta_1\rfloor\geq\lfloor1/\delta_2\rfloor\).

The regularization theorem is useful even when a given modulus does not have this property. Its construction replaces one arbitrary choice of indices with a finite maximum of indices at standard dyadic tolerances. For instance, since \(k(1/3)=2\), the regularized index at tolerance \(1/3\) is

$$ \widetilde M(1/3) =\max\{M(1),M(1/2),M(1/4)\}. $$

The final dyadic tolerance is \(1/4<1/3\), so any index at least this maximum guarantees error less than \(1/4\), and hence less than \(1/3\). The extra accuracy is harmless; the advantage is that the construction behaves monotonically across tolerances.

Worked Example: Alternating Terms and a Shared Tail

Let \(u_n=(-1)^n/(n+3)\) and \(v_n=2(-1)^{n+1}/(n+4)\). Both sequences converge to \(0\). For \(u_n\), a valid modulus is \(M_u(\delta)=\lfloor1/\delta\rfloor\), since \(n\geq M_u(\delta)\) implies \[ |u_n|=\frac{1}{n+3}\leq\frac{1}{M_u(\delta)+3}<\delta. \] For \(v_n\), take \(M_v(\delta)=\lfloor2/\delta\rfloor\); then \(n\geq M_v(\delta)\) gives \[ |v_n|=\frac{2}{n+4}\leq\frac{2}{M_v(\delta)+4}<\delta. \]

For tolerance \(\delta=1/8\), the common index is \(N=\max\{8,16\}=16\). If \(n\geq16\), both magnitudes are less than \(1/8\), regardless of the signs. Consequently,

$$ |u_n-v_n| \leq |u_n|+|v_n| <\frac18+\frac18 =\frac14. $$

The sign changes do not require separate cases because absolute values bound both sequences’ distances from their common limit. This is the same reason the proof that a convergent sequence is Cauchy works for sequences that alternate in sign.

Common Pitfalls in the Proof

One pitfall is to choose the full tolerance \(\varepsilon\) for each term’s distance from the limit. That only yields a pairwise bound less than \(2\varepsilon\), which is not the requested bound less than \(\varepsilon\). The remedy is to divide the error allowance between the two terms, for example by using \(\varepsilon/2\) for each.

Another pitfall is to use an index that depends on \(m\) or \(n\). The Cauchy property requires a single tail threshold, chosen from \(\varepsilon\), that works for every pair of indices beyond it. The convergence definition supplies exactly this uniformity: once the index is chosen, every later term is close to the same limit.

Finally, strict inequalities should be tracked through the last line. If each distance from the limit is strictly less than \(\varepsilon/2\), then their sum is strictly less than \(\varepsilon\). A proof that establishes only a bound of at most \(\varepsilon\) has not yet met the usual Cauchy condition, which uses a strict inequality.

Check Your Understanding

Use the proof structure and the two extensions in this tutorial to answer the following questions.

  1. Why must the tail index be chosen before the arbitrary pair of indices is selected?
  2. How does taking the maximum of finitely many indices ensure that all corresponding estimates hold on one tail?
  3. In the modulus regularization theorem, why is \(2^{-k(\delta)}<\delta\) needed?
  4. What property of the function \(k(\delta)\) ensures that the regularized modulus is nondecreasing as accuracy becomes stricter?
  5. Why can alternating signs be handled without splitting into even and odd cases when comparing terms to a common limit?