From Product Integrals to Practical Calculations
Fubini’s Theorem is more than a way to evaluate the same integral in two orders. It lets us extract one-variable information from a function of two variables, calculate probabilities by slicing a region, and combine functions through convolution. In each application, the central task is to represent the quantity of interest as an integral on a product space and verify that Tonelli’s or Fubini’s hypotheses apply.
Throughout, Lebesgue measure on the real line is denoted by \(\lambda\), and product integrals on \(\mathbb{R}^2\) use \(\lambda\otimes\lambda\). We use Tonelli’s Theorem for nonnegative measurable functions and Fubini’s Theorem for absolutely integrable functions. In particular, a signed integrand cannot be rearranged merely because its iterated integrals appear to exist: absolute integrability is the condition that justifies the interchange.
Marginal Densities
Suppose a pair of real-valued random quantities has a joint probability density \(p(x,y)\). The density describes probability on the product space: the probability that the pair lies in a measurable set \(A\subseteq\mathbb{R}^2\) is the integral of \(p\) over \(A\). To find the density of the first coordinate alone, integrate out the second coordinate. Tonelli’s Theorem justifies this operation even before we know that the resulting one-variable integral is finite at every point.
Proof. Since \(p\) is nonnegative and measurable, Tonelli’s Theorem implies that the function \(x\mapsto\int_{\mathbb{R}}p(x,y)\,d\lambda(y)\) is measurable and that $$ \int_{\mathbb{R}}p_X(x)\,d\lambda(x) =\int_{\mathbb{R}}\left(\int_{\mathbb{R}}p(x,y)\,d\lambda(y)\right)d\lambda(x) =\int_{\mathbb{R}^2}p\,d(\lambda\otimes\lambda) =1. $$ For a Borel set \(B\), apply Tonelli’s Theorem to the nonnegative function \(\mathbf{1}_B(x)p(x,y)\). It gives $$ \int_Bp_X(x)\,d\lambda(x) =\int_{\mathbb{R}}\left(\int_{\mathbb{R}}\mathbf{1}_B(x)p(x,y)\,d\lambda(y)\right)d\lambda(x) =\int_{B\times\mathbb{R}}p\,d(\lambda\otimes\lambda). $$ The final expression is the probability that the first coordinate lies in \(B\), by the definition of the joint density. This proves the claim.
The theorem guarantees that the marginal integral is finite for almost every \(x\), not necessarily for every \(x\). A density can be changed on a set of measure zero without changing any probabilities, so values on an exceptional null set do not affect the resulting distribution.
Worked Example: Marginals of a Density on a Rectangle
Consider $$ p(x,y)= \begin{cases} \frac{1}{6},&0\leq x\leq2,\ 0\leq y\leq3,\\ 0,&\text{otherwise}. \end{cases} $$ The integral over the rectangle is \((1/6)(2)(3)=1\), so this is a joint probability density. For \(0\leq x\leq2\), integrating in \(y\) gives $$ p_X(x)=\int_0^3\frac{1}{6}\,dy=\frac{3}{6}=\frac{1}{2}. $$ For \(x\) outside \([0,2]\), the integrand is zero for all \(y\). Thus \(p_X(x)=1/2\) on \([0,2]\) and \(p_X(x)=0\) elsewhere. Its integral is \((1/2)(2)=1\). In the other coordinate, $$ p_Y(y)=\int_0^2\frac{1}{6}\,dx=\frac{2}{6}=\frac{1}{3} $$ for \(0\leq y\leq3\), and \(p_Y(y)=0\) elsewhere. The marginal densities are uniform on intervals of different lengths, although the joint density is constant on the rectangle.
Probabilities from Sections of a Region
A joint density also turns a geometric event into an integral over a region. The bounds in an iterated integral should come from the event’s defining inequalities, just as in iterated integration over a measurable region. Choosing the order that gives simpler sections can make the calculation substantially shorter. Because a density is nonnegative, Tonelli’s Theorem applies; when the density is integrable, Fubini’s Theorem also allows either order.
Worked Example: A Corner Event for Two Uniform Quantities
Let \((X,Y)\) have the joint density \(p(x,y)=1\) on the unit square \([0,1]\times[0,1]\) and zero elsewhere. This has total integral \(1\). To find the probability that \(X+Y>3/2\), fix \(x\). If \(0\leq x\leq1/2\), no \(y\in[0,1]\) satisfies the inequality. If \(1/2<x\leq1\), the allowed values are \(3/2-x<y\leq1\). Therefore $$ \mathbb{P}(X+Y>3/2) =\int_{1/2}^1\int_{3/2-x}^1 1\,dy\,dx =\int_{1/2}^1\left(x-\frac{1}{2}\right)dx. $$ Evaluating the last integral, $$ \left[\frac{x^2}{2}-\frac{x}{2}\right]_{1/2}^1 =0-\left(\frac{1}{8}-\frac{1}{4}\right) =\frac{1}{8}. $$ The event is a triangular corner of the square. The bounds include precisely the points satisfying the event, apart from boundary points, which have probability zero under this density.
This calculation illustrates a useful distinction. The geometry determines the integration limits, while the density determines the integrand. A region’s area gives its probability only when the density is \(1\) there, as it is in this example. With a nonconstant density, the same region must be integrated against that density rather than measured by its area alone.
Convolution and the Integral of a Sum
A further application combines two functions by integrating over all ways their arguments can add to a fixed value. For functions \(f\) and \(g\) on \(\mathbb{R}\), their convolution is formally $$ (f*g)(x)=\int_{\mathbb{R}}f(y)g(x-y)\,d\lambda(y). $$ The integral may fail to be finite for some \(x\), even if both functions are integrable. Fubini’s Theorem nevertheless shows that it is finite for almost every \(x\), and gives the integral of the convolution.
Proof. The function \((x,y)\mapsto f(y)g(x-y)\) is Borel measurable: the map \((x,y)\mapsto(x-y,y)\) is continuous, and \(f\) and \(g\) are Borel measurable. Apply Tonelli’s Theorem to its absolute value. For fixed \(y\), translation invariance of Lebesgue measure gives $$ \int_{\mathbb{R}}|g(x-y)|\,d\lambda(x)=\int_{\mathbb{R}}|g(t)|\,d\lambda(t). $$ Consequently, $$ \int_{\mathbb{R}^2}|f(y)g(x-y)|\,d(\lambda\otimes\lambda)(x,y) =\int_{\mathbb{R}}|f(y)|\left(\int_{\mathbb{R}}|g(x-y)|\,d\lambda(x)\right)d\lambda(y) =\left(\int_{\mathbb{R}}|f|\,d\lambda\right)\left(\int_{\mathbb{R}}|g|\,d\lambda\right) <\infty. $$ Tonelli’s Theorem implies that \(\int_{\mathbb{R}}|f(y)g(x-y)|\,d\lambda(y)\) is finite for almost every \(x\). Fubini’s Theorem now applies to the signed function \((x,y)\mapsto f(y)g(x-y)\), so its iterated integrals agree. In the order that integrates in \(x\) first, translation invariance again yields $$ \int_{\mathbb{R}}\left(\int_{\mathbb{R}}f(y)g(x-y)\,d\lambda(y)\right)d\lambda(x) =\int_{\mathbb{R}}f(y)\left(\int_{\mathbb{R}}g(x-y)\,d\lambda(x)\right)d\lambda(y) =\left(\int_{\mathbb{R}}f\,d\lambda\right)\left(\int_{\mathbb{R}}g\,d\lambda\right). $$ The convolution agrees almost everywhere with the inner integral in this expression. Changing its values on the exceptional null set does not change its integral. The absolute-integrability estimate also ensures that the convolution is integrable. This proves the theorem.
For probability densities, the convolution has a direct interpretation. If two independent real-valued random quantities have densities \(f\) and \(g\), then their joint density is \(f(y)g(z)\). To describe their sum, fix a sum value \(x\): the possible pairs have the form \((y,x-y)\). Integrating the joint density along these sections gives the convolution \(f*g\), which is therefore a density of the sum. The integral identity confirms that its total integral is \(1\) when both original densities integrate to \(1\).
Worked Example: Convolving Two Uniform Densities
Let \(f\) be the uniform density on \([0,2]\), so \(f(y)=1/2\) there and is zero elsewhere. Let \(g\) be the uniform density on \([0,1]\), so \(g(x-y)=1\) exactly when \(0\leq x-y\leq1\). The convolution is \(1/2\) times the length of the intersection of \([0,2]\) with \([x-1,x]\). It vanishes outside \([0,3]\). If \(0\leq x\leq1\), this intersection has length \(x\); if \(1\leq x\leq2\), it has length \(1\); and if \(2\leq x\leq3\), it has length \(3-x\). Hence $$ (f*g)(x)= \begin{cases} x/2,&0\leq x\leq1,\\ 1/2,&1\leq x\leq2,\\ (3-x)/2,&2\leq x\leq3,\\ 0,&\text{otherwise}. \end{cases} $$ The total integral is $$ \int_0^1\frac{x}{2}\,dx+\int_1^2\frac{1}{2}\,dx+\int_2^3\frac{3-x}{2}\,dx =\frac{1}{4}+\frac{1}{2}+\frac{1}{4} =1. $$ This is the density of the sum of independent quantities with the two specified uniform densities.
Worked Example: Convolving Two Exponential Functions
Take \(f(x)=e^{-x}\) for \(x\geq0\), zero otherwise, and \(g(x)=2e^{-2x}\) for \(x\geq0\), zero otherwise. Their integrals are \(1\) and \(1\), respectively. For \(x<0\), the convolution is zero. For \(x\geq0\), both factors can be nonzero only when \(0\leq y\leq x\), and therefore $$ (f*g)(x)=\int_0^x e^{-y}\,2e^{-2(x-y)}\,dy =2e^{-2x}\int_0^x e^y\,dy =2e^{-2x}(e^x-1) =2(e^{-x}-e^{-2x}). $$ The result is nonnegative for \(x\geq0\), since \(e^{-x}\geq e^{-2x}\) there. Its integral is $$ \int_0^\infty 2(e^{-x}-e^{-2x})\,dx =2\left(1-\frac{1}{2}\right) =1, $$ in agreement with the convolution theorem. This function is the density of the sum of independent quantities with the two given exponential densities.
What Fubini Does—and Does Not—Permit
These applications use Fubini in distinct ways: marginalization integrates out a coordinate; region probabilities integrate over sections; convolution integrates over decompositions of a fixed sum. The same logical check underlies each calculation: establish measurability and use the theorem appropriate to the integrand. For nonnegative functions, Tonelli permits the iterated integrals even if their value is infinite. For signed functions, the convolution argument first established absolute integrability on the product space and only then interchanged the integrals.
A common pitfall is to treat a plausible iterated calculation as proof that changing the order is valid. If a signed function is not absolutely integrable, Fubini’s Theorem does not justify that change; the two orders can behave differently. In probability applications, densities are nonnegative, so Tonelli is often the simplest justification. In applications involving signed functions, retain the absolute-integrability check rather than relying on a formal rearrangement.
Check Your Understanding
Use the definitions and results in this tutorial to answer the following questions.
- Why does Tonelli’s Theorem apply when deriving a marginal density from a nonnegative joint density?
- For a joint density on a region, what determines the bounds in an iterated probability integral?
- What product-space function is used to prove the integral identity for a convolution?
- Why does the convolution integral exist absolutely for almost every point under the hypotheses of the convolution theorem?
- What condition must be checked before using Fubini to interchange integrals of a signed function?