Tutorials › Real Analysis › Why Metric Spaces?

Metric Spaces · Tutorial 651 of 1000

Why Metric Spaces?

See how distance-based language captures familiar normed-space ideas and extends them to settings where subtraction, scaling, or finite distances are unavailable.

Advanced 10 min read

What You'll Learn

  • Understand why a general notion of distance is useful beyond vector spaces.
  • Verify that every norm supplies a natural distance between points.
  • Relate norm convergence and norm-Cauchy sequences to distance-based language.
  • See how the discrete distance describes convergence on arbitrary sets.
  • Understand how a bounded change of distance can preserve convergence and Cauchy behavior.
  • Recognize how distances can compare functions even when their differences are unbounded.

From Norms to Distances

The Approximation Mastery Examination used norms to measure errors between functions. A norm does this by measuring the size of a vector, and the difference \(x-y\) turns that into a measure of separation between two points. This works well in a vector space: points can be subtracted, and the norm of their difference gives a distance. But many objects studied in analysis do not naturally form vector spaces. A set of shapes, a collection of probability distributions, or an arbitrary set of functions may have no useful subtraction operation.

The purpose of metric spaces is to keep the idea of distance while dispensing with the extra algebraic structure of a vector space. We will use \(d(x,y)\) for the distance between two points. For now, the guiding requirements are the familiar ones: distances are nonnegative, only identical points have distance zero, reversing the order of the points does not change the distance, and traveling through an intermediate point cannot shorten the direct distance. The next tutorial will state the formal definition precisely.

These requirements are not arbitrary. They are the properties that make distance useful for describing nearness. Once distance is available, we can define a ball around a point, say which sequences approach a point, and express what it means for a sequence to be Cauchy. Those ideas need no addition or scalar multiplication. The central question here is why that change of language is worth making.

What a Norm Already Gives Us

If \(V\) is a normed vector space, define the distance from \(x\) to \(y\) by taking the norm of their difference. The norm axioms immediately give all four guiding distance requirements. This is the basic bridge between the earlier study of normed spaces and the broader framework now beginning.

Theorem (A Norm Induces a Distance): Let \(V\) be a real normed vector space. Define \(d(x,y)=\|x-y\|\) for \(x,y\in V\). Then \(d\) is nonnegative, \(d(x,y)=0\) exactly when \(x=y\), \(d(x,y)=d(y,x)\), and \(d(x,z)\leq d(x,y)+d(y,z)\) for all \(x,y,z\in V\).

Proof. A norm is nonnegative and vanishes only at the zero vector. Therefore \(d(x,y)=\|x-y\|\geq0\), and \(d(x,y)=0\) exactly when \(x-y=0\), which is equivalent to \(x=y\). Since \(y-x=-(x-y)\), homogeneity of the norm gives $$ d(y,x)=\|y-x\|=\|-(x-y)\|=\|x-y\|=d(x,y). $$ Finally, write \(x-z=(x-y)+(y-z)\). The triangle inequality for the norm gives $$ d(x,z)=\|x-z\| =\|(x-y)+(y-z)\| \leq \|x-y\|+\|y-z\| =d(x,y)+d(y,z). $$ Thus the distance properties follow from the norm axioms. \(\square\)

The induced distance lets us describe norm convergence without mentioning vector operations: \(x_n\) approaches \(x\) exactly when \(d(x_n,x)\) tends to zero. Likewise, the norm-Cauchy condition is the statement that the distances between terms eventually become arbitrarily small. The underlying notions have not changed; distance simply makes clear which parts of the theory depend on measuring nearness and which depend on linear structure.

Worked Example: Distance in the Euclidean Plane

On \(\mathbb{R}^2\), use the Euclidean norm \(\|(a,b)\|_2=\sqrt{a^2+b^2}\). The induced distance between \(u=(1,2)\) and \(v=(4,6)\) is $$ d(u,v)=\|u-v\|_2 =\|(1-4,2-6)\|_2 =\sqrt{(-3)^2+(-4)^2} =5. $$ For a third point \(w=(1,6)\), the two distances through \(w\) are $$ d(u,w)=\sqrt{(1-1)^2+(2-6)^2}=4, \qquad d(w,v)=\sqrt{(1-4)^2+(6-6)^2}=3. $$ Thus \(d(u,v)=5\leq4+3=7\), as required by the triangle inequality. In this familiar setting, the distance measures geometric separation, while the norm measures the length of a vector.

Distance Does Not Require Linear Structure

The main gain is that distance can be defined on a set where there are no vectors to subtract. One extreme example assigns distance \(1\) to every pair of distinct points and distance \(0\) to a point and itself. This is called the discrete distance. It satisfies the distance rules on any set, regardless of what its elements are. It is therefore available even when the points are books, graphs, or other objects with no prescribed arithmetic.

Worked Example: The Discrete Distance on an Arbitrary Set

Let \(X\) be any set and put \(d(x,y)=0\) if \(x=y\), and \(d(x,y)=1\) if \(x\ne y\). Nonnegativity and symmetry follow directly from the definition, and \(d(x,y)=0\) exactly when \(x=y\). To check the triangle inequality, fix \(x,y,z\in X\). If \(x=z\), then \(d(x,z)=0\leq d(x,y)+d(y,z)\). If \(x\ne z\), then \(d(x,z)=1\). In this case \(x\) and \(z\) cannot both equal \(y\), so at least one of \(x\ne y\) or \(y\ne z\) holds. Hence at least one of \(d(x,y)\) and \(d(y,z)\) equals \(1\), giving \(d(x,y)+d(y,z)\geq1=d(x,z)\).

For this distance, a sequence \(x_n\) approaches \(x\) precisely when it is eventually equal to \(x\). Indeed, if \(d(x_n,x)\) tends to zero, then eventually \(d(x_n,x)<1/2\). Since each such distance is either \(0\) or \(1\), it must be \(0\), so \(x_n=x\) from then on. Conversely, an eventually constant sequence at \(x\) has distance zero from \(x\) from that point onward. This example shows how the choice of distance determines what “approaching” means.

The discrete distance is not usually a way to express geometric closeness. Its value is that it demonstrates how little structure the general framework requires. By contrast, distances on function spaces can express uniform closeness without treating functions as vectors under a norm. For instance, for any nonempty set \(E\), the formula $$ D(f,g)=\sup_{x\in E}\min\{1,|f(x)-g(x)|\} $$ gives a finite distance between any two real-valued functions on \(E\), even if their difference is unbounded. The truncation at \(1\) ensures \(0\leq D(f,g)\leq1\). The triangle inequality follows from \(\min\{1,a+b\}\leq\min\{1,a\}+\min\{1,b\}\) for nonnegative \(a,b\), followed by taking suprema. The distance is zero only when the functions agree at every point. This is useful when an untruncated supremum of differences would be infinite and so could not serve as a finite distance.

Changing the Scale Without Changing Nearness

A distance need not record separation on a particular numerical scale. For example, one may want all distances to be bounded even if the original distances can be arbitrarily large. The following result shows that a common transformation accomplishes this while preserving convergence and the Cauchy property.

Theorem (A Bounded Distance with the Same Convergence): Suppose \(d\) satisfies the distance properties described above. Define $$ \rho(x,y)=\frac{d(x,y)}{1+d(x,y)}. $$ Then \(\rho\) is also a distance, with \(0\leq\rho(x,y)<1\). A sequence converges with respect to \(\rho\) exactly when it converges with respect to \(d\), and it is Cauchy with respect to \(\rho\) exactly when it is Cauchy with respect to \(d\).

Proof. The function \(\phi(t)=t/(1+t)\) is nonnegative for \(t\geq0\), and it equals zero exactly at \(t=0\). Thus \(\rho\) is nonnegative and \(\rho(x,y)=0\) exactly when \(d(x,y)=0\), or equivalently \(x=y\). Symmetry follows from the symmetry of \(d\). Also, \(0\leq\phi(t)<1\) for every finite \(t\geq0\), which gives the stated bound.

For the triangle inequality, first note that for \(s,t\geq0\), $$ \phi(s)+\phi(t)-\phi(s+t) =\frac{2st+st(s+t)}{(1+s)(1+t)(1+s+t)} \geq0. $$ Therefore \(\phi(s+t)\leq\phi(s)+\phi(t)\). Since \(d(x,z)\leq d(x,y)+d(y,z)\) and \(\phi\) is increasing on \([0,\infty)\), it follows that $$ \rho(x,z) =\phi(d(x,z)) \leq\phi(d(x,y)+d(y,z)) \leq\phi(d(x,y))+\phi(d(y,z)) =\rho(x,y)+\rho(y,z). $$ So \(\rho\) has the triangle inequality as well.

To compare convergence, observe that for \(0<\varepsilon<1\), $$ \phi(t)<\varepsilon \quad\Longleftrightarrow\quad \frac{t}{1+t}<\varepsilon \quad\Longleftrightarrow\quad t<\frac{\varepsilon}{1-\varepsilon}. $$ Also, for every \(\eta>0\), \(t<\eta\) implies \(\phi(t)<\eta\), because \(t/(1+t)<t\) when \(t>0\), with the zero case immediate. These comparisons show that \(d(x_n,x)\) tends to zero if and only if \(\rho(x_n,x)\) tends to zero. Applying the same comparisons to \(d(x_m,x_n)\) and \(\rho(x_m,x_n)\), uniformly for all sufficiently large \(m,n\), proves the corresponding equivalence for Cauchy sequences. \(\square\)

Worked Example: Compressing the Usual Distance on the Real Line

On \(\mathbb{R}\), begin with \(d(x,y)=|x-y|\), and define $$ \rho(x,y)=\frac{|x-y|}{1+|x-y|}. $$ For \(x=2\) and \(y=7\), \(d(2,7)=5\), whereas $$ \rho(2,7)=\frac{5}{1+5}=\frac56. $$ If \(|x-y|=100\), then \(\rho(x,y)=100/101<1\), so even very large separations are compressed below \(1\). But this does not alter which sequences converge: for \(x_n=1+1/n\), we have $$ d(x_n,1)=\frac1n\longrightarrow0, \qquad \rho(x_n,1)=\frac{1/n}{1+1/n}=\frac1{n+1}\longrightarrow0. $$ The theorem ensures the same convergence correspondence for every sequence, not just this example.

Why the General Framework Matters

Distance-based language lets analysis separate two ideas that are often combined in normed spaces: how close two points are, and what algebraic operations the points support. The norm-induced distance shows that all the familiar normed-space examples fit into this broader setting. The discrete distance and bounded function-space distance show why the broader setting is not merely a change of notation: it can describe nearness in situations where a norm is unavailable or unsuitable.

This perspective is especially useful when studying continuity and convergence on spaces of objects. A function can be continuous between spaces once distances are available in both its domain and range; neither space must be a vector space. Compactness, completeness, and continuity can likewise be developed around distance alone. Their statements then apply to many settings, while familiar results about norms remain available as special cases.

There is an important caution: the distance is part of the structure, not something automatically determined by the underlying set. The discrete distance makes only eventually constant sequences converge, while the usual distance on the real line admits many nonconstant convergent sequences. Even two different distances can have the same convergent sequences, as the bounded transformation theorem demonstrates. Therefore, whenever a set is equipped with a distance, conclusions about convergence refer to that chosen distance.

Takeaway: A norm measures vectors, and hence induces a distance between points; a metric space keeps the useful notion of distance while allowing the points to have no vector-space structure at all.

Check Your Understanding

Use the examples and results above to explain what distance adds to the language of analysis.

  1. Which norm properties give symmetry and the triangle inequality for the distance induced by a norm?
  2. Why does convergence in the discrete distance force a sequence to be eventually constant?
  3. How does truncating pointwise differences make a distance between functions finite even when their difference is unbounded?
  4. Why does the transformation \(t\mapsto t/(1+t)\) preserve convergence to zero?
  5. Can two different distances on the same set give the same convergent sequences? Explain using the bounded distance theorem.