Tutorials › Real Analysis › Integrability of Step Functions

Riemann Integration · Tutorial 472 of 1000

Integrability of Step Functions

A finite step function is Riemann integrable, and its integral is the sum of each constant value multiplied by the length of the interval where it applies.

Advanced 9 min read

What You'll Learn

  • Define step functions using a finite partition and constant values between partition points
  • Distinguish interval values from potentially exceptional values at partition points
  • Prove integrability by bounding the Darboux gap on cells adjacent to breakpoints
  • Derive the integral formula from upper and lower sums
  • Calculate integrals when the function has jumps or exceptional endpoint values
  • Recognize why partition alignment simplifies step-function calculations

From Constant Pieces to a Finite Step

A constant function has no variation anywhere on its interval, so every partition gives equal upper and lower sums. A step function is built from finitely many constant pieces. Its values can change at a finite collection of breakpoints, and those changes mean that some upper and lower sums may differ. The key is that a partition can include all the breakpoints: then every partition cell lies within one constant piece, apart from possible exceptional values at its endpoints. By making the cells adjacent to breakpoints short, we can control the total Darboux gap.

We will allow the function to have values at the breakpoints that differ from the constant values on either side. This convention includes common examples such as functions defined using half-open intervals, and it makes clear that integrability does not require continuity at the jumps. The endpoint values of the whole domain may also differ from the values immediately inside the interval.

Definition: Let \(a<b\). A function \(f:[a,b]\to\mathbb{R}\) is a step function if there is a finite partition \(a=t_0<t_1<\cdots<t_n=b\) and constants \(c_1,\ldots,c_n\in\mathbb{R}\) such that \(f(x)=c_i\) whenever \(t_{i-1}<x<t_i\). The values \(f(t_0),\ldots,f(t_n)\) may be any real numbers.

The points \(t_1,\ldots,t_{n-1}\) are the interior breakpoints. Between consecutive breakpoints, the function is constant. The definition does not require the constants on adjacent intervals to be different, and it does not require the value at a breakpoint to match either neighboring constant. A partition satisfying the definition need not be unique, but any one such partition is enough for the arguments below.

Bounding the Darboux Gap

Fix a step function and a partition \(T=\{t_0,\ldots,t_n\}\) as in the definition. Write

$$ S=\sum_{i=1}^{n}c_i(t_i-t_{i-1}). $$

This is the candidate value of the integral: on the \(i\)th constant piece, the value \(c_i\) is multiplied by the piece's length. To justify this formula, we will compare both Darboux sums with \(S\) for partitions that include every point of \(T\).

Let \(P\) be any partition that contains all of \(t_0,\ldots,t_n\), and suppose its mesh is at most \(\delta\). Every cell of \(P\) has its interior inside a single interval \((t_{i-1},t_i)\), since no breakpoint can lie strictly inside a cell. Call a cell exceptional if one of its endpoints is a point of \(T\). There are at most \(2(n+1)\) exceptional cells: each of the \(n+1\) points of \(T\) can be an endpoint of at most two cells. Each such cell has length at most \(\delta\), so the sum of the lengths of the exceptional cells is at most \(2(n+1)\delta\).

On every nonexceptional cell, the function equals the relevant \(c_i\) throughout the closed cell, and its infimum and supremum are both \(c_i\). An exceptional cell still has a constant value \(c_i\) at all its interior points, but its endpoint values may change its infimum or supremum. We can bound the resulting error because all the values involved are bounded.

Theorem (Integrability and Integral of a Step Function): Every step function \(f:[a,b]\to\mathbb{R}\) is Riemann integrable. If \(a=t_0<\cdots<t_n=b\) and \(f(x)=c_i\) for \(t_{i-1}<x<t_i\), then $$ \int_a^b f(x)\,dx=\sum_{i=1}^{n}c_i(t_i-t_{i-1}). $$ The values at the partition points do not affect this formula.

Proof. Let \(M\) be the maximum of the finite set \[ \{|c_1|,\ldots,|c_n|,|f(t_0)|,\ldots,|f(t_n)|\}. \] Then \(|f(x)|\leq M\) for every \(x\in[a,b]\): at interior points of the constant pieces this follows from the \(c_i\), and at the points of \(T\) it follows from the values included in the maximum. If \(M=0\), then \(f\) is identically zero and is integrable. Suppose henceforth that \(M>0\).

Take a partition \(P\) containing \(T\), with mesh at most \(\delta\). Each cell \(J\) has interior inside one constant piece, say the piece with value \(c_i\). Its contribution to \(S\) is \(c_i|J|\), where \(|J|\) is the length of \(J\). If \(J\) is nonexceptional, its contribution to each of \(L(f,P)\) and \(U(f,P)\) is exactly \(c_i|J|\).

If \(J\) is exceptional, its infimum \(m_J\) and supremum \(M_J\) both lie in \([-M,M]\), as does \(c_i\). Consequently, \[ |m_J-c_i|\leq 2M \quad\text{and}\quad |M_J-c_i|\leq 2M. \] Thus the lower-sum contribution \(m_J|J|\) and the upper-sum contribution \(M_J|J|\) each differ from the corresponding contribution \(c_i|J|\) to \(S\) by at most \(2M|J|\). Summing these estimates over exceptional cells, and using the bound on their total length, gives

$$ |L(f,P)-S|\leq 4M(n+1)\delta, \qquad |U(f,P)-S|\leq 4M(n+1)\delta. $$

In particular,

$$ U(f,P)-L(f,P) \leq |U(f,P)-S|+|L(f,P)-S| \leq 8M(n+1)\delta. $$

For every \(\varepsilon>0\), choose \(\delta>0\) so that \(8M(n+1)\delta<\varepsilon\). We can take a partition containing \(T\) with mesh at most \(\delta\): subdivide each interval \([t_{i-1},t_i]\) into sufficiently many equal parts and combine those subdivisions. The displayed estimate then gives \(U(f,P)-L(f,P)<\varepsilon\). The function is bounded, so the Darboux Criterion implies that \(f\) is Riemann integrable.

It remains to identify the integral. For partitions containing \(T\) with mesh tending to zero, the estimates show that both \(L(f,P)\) and \(U(f,P)\) tend to \(S\). Since the integral lies between the lower and upper sums for every partition, it must equal \(S\). More explicitly, given any \(\eta>0\), choose such a partition with \(4M(n+1)\delta<\eta\). Then \[ S-\eta<L(f,P)\leq\int_a^b f(x)\,dx\leq U(f,P)<S+\eta. \] This holds for every \(\eta>0\), which forces \(\int_a^b f(x)\,dx=S\). \(\square\)

The estimate also explains why exceptional breakpoint values do not alter the integral. Their effects on the upper and lower sums are confined to cells adjacent to breakpoints. Those cells can be made to have arbitrarily small total length, while all other cells contribute exactly as if the function were constant on each piece right up to its endpoints.

Worked Example: Two Constant Pieces with a Jump

Define \(f:[0,5]\to\mathbb{R}\) by \(f(x)=2\) for \(0<x<2\) and \(f(x)=-1\) for \(2<x<5\). Let the breakpoint values be \(f(0)=4\), \(f(2)=7\), and \(f(5)=-3\). The interval lengths are \(2-0=2\) and \(5-2=3\). The theorem gives

$$ \int_0^5 f(x)\,dx=2(2-0)+(-1)(5-2)=4-3=1. $$

The value \(7\) at the jump, as well as the unusual endpoint values \(4\) and \(-3\), can affect upper and lower sums on cells that meet those points. None changes the integral formula, because those cells can be made arbitrarily short.

Worked Example: A Step Function with a Negative Piece

Suppose \(g:[-1,4]\to\mathbb{R}\) equals \(3\) on \((-1,1)\) and \(-2\) on \((1,4)\), with any real values assigned at \(-1\), \(1\), and \(4\). The piece lengths are \(1-(-1)=2\) and \(4-1=3\). Therefore

$$ \int_{-1}^{4}g(x)\,dx=3(2)+(-2)(3)=6-6=0. $$

This integral is zero even though the function is not identically zero: the positive contribution \(6\) and the negative contribution \(-6\) cancel. The sign of each piece is retained in the formula.

Worked Example: An Indicator of an Interval

Let \(h:[0,6]\to\mathbb{R}\) equal \(1\) for \(2<x<5\) and \(0\) for \(0<x<2\) and \(5<x<6\). Assign any real values at \(0\), \(2\), \(5\), and \(6\). A partition for this step function is \(\{0,2,5,6\}\), and the three constant values are \(0\), \(1\), and \(0\). Hence

$$ \int_0^6h(x)\,dx=0(2-0)+1(5-2)+0(6-5)=0+3+0=3. $$

The result is the length of the interval on which the function takes the value \(1\). The assigned values at the endpoints of that interval do not change the calculation.

Why Aligning the Partition Matters

If a partition cell straddles a breakpoint, its infimum and supremum may come from different constant pieces. For example, a cell that includes points on both sides of a jump from \(2\) to \(-1\) has an oscillation of at least \(3\), regardless of how its extrema are attained. Refining without regard to the breakpoints can leave such a cell in place. Including every breakpoint removes this problem: each cell's interior sees only one constant value, and only cells touching the breakpoints require an error estimate.

This is a useful distinction between a bound on the height of the jumps and a bound on their contribution to the Darboux gap. The jump heights need not be small. What makes the gap small is that there are only finitely many breakpoints and the total length of the cells next to them can be made small. In the proof, \(M\) controls the possible size of the oscillation, while \(\delta\) controls the lengths on which that oscillation can occur.

1
List the constant pieces.
Choose breakpoints \(t_0,\ldots,t_n\) and record the value \(c_i\) on each open interval \((t_{i-1},t_i)\).
2
Align the partition.
Include every breakpoint in the partition, then subdivide the pieces so that the mesh is as small as needed.
3
Control the exceptional cells.
Only cells meeting breakpoints can have endpoint values that differ from the relevant constant; their total length becomes small with the mesh.
4
Calculate the integral.
Add \(c_i(t_i-t_{i-1})\) over the constant pieces, retaining negative values when a piece lies below zero.

A common pitfall is to use the formula while overlooking what the step-function definition requires. The formula uses the values on the open pieces and their lengths; it does not say that the function must take those values at the breakpoints. Conversely, the integral formula does require a finite partition into constant pieces. A function that changes value at infinitely many points is not covered by this theorem merely because it looks locally piecewise constant.

The proof is a model for handling functions with finitely many jumps using Darboux sums. It does not rely on continuity at the breakpoints. Instead, it uses a partition that isolates all possible irregularities and makes the intervals containing them short. For step functions, the remaining intervals have no oscillation at all, which is why the integral reduces to a finite sum.

Check Your Understanding

Use the definition and the partition-aligned estimates to answer the following questions.

  1. In the definition of a step function, what condition must hold on each open interval between consecutive partition points?
  2. Why can a step function have arbitrary values at its finitely many partition points and still be Riemann integrable?
  3. A function equals \(4\) on \((0,2)\) and \(-3\) on \((2,6)\). What is its integral on \([0,6]\), regardless of its values at \(0\), \(2\), and \(6\)?
  4. Why does a partition that contains every breakpoint leave possible variation only on cells adjacent to those breakpoints?
  5. In the integrability proof, what roles do the bound \(M\) and the mesh bound \(\delta\) play in controlling the Darboux gap?