From Measures to Probability
A measure assigns a nonnegative size to each measurable set, and that size may be infinite. Probability measures add one important condition: the entire space has measure one. This normalization makes the measure of any event a number between zero and one, and it lets measure-theoretic results be read as rules about probabilities.
We use \(\Omega\) for the sample space and \(\mathcal{F}\) for its sigma-algebra. The elements of \(\Omega\) are possible outcomes, while the elements of \(\mathcal{F}\) are the events to which probabilities are assigned. The sigma-algebra matters: a probability measure is defined on the measurable events, not automatically on every subset of the sample space.
Because a probability measure is a measure, it has all the properties established in the tutorials on measures: in particular, \(\mathbb{P}(\varnothing)=0\), countable additivity for pairwise disjoint measurable events, and monotonicity. The new normalization condition makes these properties especially easy to interpret. The event \(\Omega\) is certain, and the empty event is impossible.
Complement Rule and Probability Bounds
If \(A\) is an event, its complement \(A^c=\Omega\setminus A\) occurs precisely when \(A\) does not. The complement rule follows from decomposing the entire sample space into these two disjoint events. It also gives a convenient way to calculate the probability that an event fails to occur.
Proof. Since \(A\) and \(A^c\) are measurable, disjoint, and have union \(\Omega\), finite additivity gives $$ 1=\mathbb{P}(\Omega)=\mathbb{P}(A)+\mathbb{P}(A^c). $$ Rearranging proves the complement rule. Nonnegativity follows from the definition of a measure. Also \(A\subseteq\Omega\), so monotonicity of measures gives \(\mathbb{P}(A)\leq\mathbb{P}(\Omega)=1\). These two inequalities establish the stated bounds. \(\square\)
The subtraction in the complement rule is always legitimate because the probability of every event is finite. In contrast, for a general measure it would not be valid to subtract an infinite measure from another infinite measure. The total-mass-one condition is what makes this elementary-looking calculation safe.
Worked Example: A Fair Coin
Let \(\Omega=\{H,T\}\), let \(\mathcal{F}=\mathcal{P}(\Omega)\), and assign probability \(1/2\) to each outcome. Thus \(\mathbb{P}(\{H\})=\mathbb{P}(\{T\})=1/2\), and \(\mathbb{P}(\Omega)=1\). If \(A=\{H\}\) is the event of heads, then \(A^c=\{T\}\), and the complement rule gives $$ \mathbb{P}(A^c)=1-\mathbb{P}(A)=1-\frac12=\frac12. $$ For the event \(C=\Omega\), the probability is \(1\), while the probability of \(\varnothing\) is \(0\). This example has two equally likely outcomes, but equal likelihood is not part of the definition of a probability measure.
Adding Probabilities When Events Overlap
Countable additivity applies directly to disjoint events. When two events overlap, simply adding their probabilities counts the overlap twice. Removing one copy of the intersection gives the correct formula.
Proof. The sets \(A\) and \(B\setminus A\) are disjoint and have union \(A\cup B\). Finite additivity therefore gives $$ \mathbb{P}(A\cup B)=\mathbb{P}(A)+\mathbb{P}(B\setminus A). $$ Also, \(B\) is the disjoint union of \(B\setminus A\) and \(A\cap B\), so $$ \mathbb{P}(B)=\mathbb{P}(B\setminus A)+\mathbb{P}(A\cap B). $$ Rearranging the second equality yields \(\mathbb{P}(B\setminus A)=\mathbb{P}(B)-\mathbb{P}(A\cap B)\). Substituting this into the first equality proves the formula. Every term is finite because it is the probability of an event. \(\square\)
If \(A\) and \(B\) are disjoint, their intersection is empty and has probability zero, so the formula reduces to additivity. If the events overlap, the intersection term corrects for double counting. In particular, the formula implies \(\mathbb{P}(A\cup B)\leq\mathbb{P}(A)+\mathbb{P}(B)\).
Worked Example: A Weighted Six-Sided Die
Let the outcomes be \(\Omega=\{1,2,3,4,5,6\}\), with the full power set as the sigma-algebra. Assign the outcomes probabilities $$ \mathbb{P}(\{1\})=\frac13,\quad \mathbb{P}(\{2\})=\frac14,\quad \mathbb{P}(\{3\})=\frac16,\quad \mathbb{P}(\{4\})=\mathbb{P}(\{5\})=\mathbb{P}(\{6\})=\frac1{12}. $$ These weights are nonnegative and sum to one: $$ \frac13+\frac14+\frac16+\frac1{12}+\frac1{12}+\frac1{12} =\frac4{12}+\frac3{12}+\frac2{12}+\frac1{12}+\frac1{12}+\frac1{12} =1. $$ Thus they define a probability measure on this finite space. Let \(A=\{2,4,6\}\) be the event of an even result and \(B=\{4,5,6\}\) the event of a result at least four. Then $$ \mathbb{P}(A)=\frac14+\frac1{12}+\frac1{12}=\frac5{12}, \qquad \mathbb{P}(B)=\frac1{12}+\frac1{12}+\frac1{12}=\frac14. $$ Their intersection is \(A\cap B=\{4,6\}\), so \(\mathbb{P}(A\cap B)=1/12+1/12=1/6\). The addition formula gives $$ \mathbb{P}(A\cup B)=\frac5{12}+\frac14-\frac16 =\frac5{12}+\frac3{12}-\frac2{12} =\frac12. $$ The outcomes are not equally likely, but the measure is still a probability measure because the weights are nonnegative and their total is one.
Partitions and Countable Sample Spaces
A partition of \(\Omega\) into measurable events consists of pairwise disjoint events whose union is all of \(\Omega\). By countable additivity, the probabilities in any countable such partition sum to one. For a countable sample space with every subset measurable, the singleton outcomes form a partition. Their probabilities therefore behave like a list of nonnegative weights with total one.
Proof. First suppose \(\mathbb{P}\) is a probability measure. Each \(p_n\) is nonnegative. The singleton sets \(\{n\}\) are pairwise disjoint and their union is \(\mathbb{N}\), so countable additivity gives $$ 1=\mathbb{P}(\mathbb{N})=\sum_{n=1}^{\infty}\mathbb{P}(\{n\}) =\sum_{n=1}^{\infty}p_n. $$ For any subset \(A\subseteq\mathbb{N}\), its singleton sets also form a disjoint countable union (or a finite union if \(A\) is finite). Countable additivity gives \(\mathbb{P}(A)=\sum_{n\in A}p_n\), with the empty set corresponding to the empty sum, which is zero.
Conversely, let \(p_n\geq0\) and \(\sum_{n=1}^{\infty}p_n=1\). By the Weighted Atomic Measure theorem from the tutorial “Examples of Measures,” the assignment \(A\mapsto\sum_{n\in A}p_n\) is a measure on \(\mathcal{P}(\mathbb{N})\). Its value on the whole space is \(\sum_{n=1}^{\infty}p_n=1\); hence it is a probability measure. This proves both directions. \(\square\)
Worked Example: A Probability Measure on the Positive Integers
For each positive integer \(n\), set \(p_n=2^{-n}\). The geometric series has sum $$ \sum_{n=1}^{\infty}2^{-n}=\frac{1/2}{1-1/2}=1, $$ so these weights define a probability measure on \((\mathbb{N},\mathcal{P}(\mathbb{N}))\). The probability of an even outcome is $$ \mathbb{P}(\{2,4,6,\ldots\}) =\sum_{k=1}^{\infty}2^{-2k} =\sum_{k=1}^{\infty}\left(\frac14\right)^k =\frac{1/4}{1-1/4} =\frac13. $$ The probability of an odd outcome is therefore \(1-1/3=2/3\). The probability of an outcome at least four can also be calculated directly: $$ \mathbb{P}(\{4,5,6,\ldots\}) =\sum_{n=4}^{\infty}2^{-n} =\frac{2^{-4}}{1-1/2} =\frac18. $$ The singleton outcomes have different probabilities, but their probabilities sum to one across the whole sample space.
A Continuous Example and a Common Pitfall
Probability measures need not be built from point weights. For a continuous example, take \(\Omega=[0,1]\), let \(\mathcal{F}\) be the Borel subsets of \([0,1]\), and restrict Lebesgue measure to this interval. Since the interval has length one, this restriction assigns total measure one and is a probability measure. For an interval \([a,b]\subseteq[0,1]\), its probability is its length, \(b-a\); a singleton has probability zero.
Worked Example: Uniform Probability on the Unit Interval
Let \(A=[1/5,3/5]\). Under the probability measure obtained by restricting Lebesgue measure to \([0,1]\), the probability of \(A\) is its length: $$ \mathbb{P}(A)=\frac35-\frac15=\frac25. $$ Its complement in \([0,1]\) consists of the portions \([0,1/5)\) and \((3/5,1]\). Their lengths are \(1/5\) and \(2/5\), respectively, so \(\mathbb{P}(A^c)=3/5\), which agrees with \(1-\mathbb{P}(A)=1-2/5=3/5\). Any individual point \(x\in[0,1]\) has probability zero because its Lebesgue measure is zero. This does not mean that the point is outside the sample space; it means that the probability measure assigns that singleton zero mass.
A useful distinction is that probability zero does not, by itself, mean logically impossible. In the unit-interval example, every point is a possible outcome, yet each individual point has probability zero. Conversely, probability one means an event has full measure; it need not literally equal the whole sample space, since its complement may be nonempty but have probability zero. These distinctions are important whenever a probability model has infinitely many possible outcomes.
Another common pitfall is to assume that all outcomes in a sample space must be equally likely. A probability measure only requires nonnegative values and total mass one. On a finite sample space, equal probabilities give the familiar uniform model, but the weighted die and the countable example show that unequal probabilities are equally valid. The sigma-algebra and the measure together specify the model: the sigma-algebra identifies the events under discussion, and the measure assigns their probabilities.
Check Your Understanding
Use the definitions and results in this tutorial to answer the following questions.
- What single condition distinguishes a probability measure from a general measure?
- If an event has probability \(3/8\), what is the probability of its complement?
- Why must the intersection be subtracted in the addition formula for two events that may overlap?
- On a countable sample space, what conditions must the singleton probabilities satisfy to define a probability measure?
- In the uniform probability measure on \([0,1]\), why does a singleton have probability zero even though it is an outcome in the sample space?