This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. Numerical results on this page are restricted to the ClinicalTrials.gov record.
1. Trial at a Glance
PALOMA-4 was a randomized, parallel-group, quadruple-masked phase 3 trial evaluating palbociclib plus letrozole versus placebo plus letrozole for first-line treatment of Asian postmenopausal women with ER+/HER2- advanced breast cancer.
| Feature | PALOMA-4 |
|---|---|
| Trial | PALOMA-4 |
| NCT ID | NCT02297438 |
| Phase | Phase 3 |
| Condition | Breast Neoplasms |
| Population | Asian postmenopausal women with ER+/HER2- advanced breast cancer |
| Allocation | Randomized |
| Design model | Parallel |
| Masking | Quadruple |
| Primary purpose | Treatment |
| Enrollment | 340 |
| Registered primary endpoint | Progression-Free Survival (PFS) Based on Investigator's Assessment: Up to Primary Completion Date |
| Primary endpoint type | Time-to-event |
| Hypothesis type | Superiority |
| Trial status | Completed |
| Start | March 23, 2015 |
| Primary completion | August 31, 2020 |
| Lead sponsor | Pfizer |
2. Clinical Question
The central statistical question was whether adding palbociclib to letrozole changes progression-free survival compared with letrozole accompanied by placebo in the registered study population.
Population
Asian postmenopausal women with ER+/HER2- advanced breast cancer receiving first-line treatment.
Intervention
Palbociclib plus letrozole.
Comparator
Placebo plus letrozole.
Primary question
Does palbociclib plus letrozole improve progression-free survival relative to placebo plus letrozole?
This is a comparative efficacy question rather than a question about whether an individual patient will or will not progress. Randomization establishes the framework for comparing the treatment groups, while the time-to-event analysis accounts for the fact that not every participant necessarily experiences progression or death during the observation period.
3. Trial Design
Palbociclib + Letrozole
- Palbociclib
- Letrozole
Placebo + Letrozole
- Placebo
- Letrozole
The ClinicalTrials.gov record identifies 340 participants as the enrollment target or total enrollment. The safety results separately provide affected/at-risk counts of 168 and 171 for the two treatment groups; these denominators should not be used to reconstruct an allocation ratio that is not explicitly reported in the ClinicalTrials.gov record.
4. Randomization, Masking, and Analysis Population
The primary PFS analysis was conducted in the intention-to-treat (ITT) population. The registry describes this as including all participants who were randomized, with study-drug assignment designated according to initial randomization, regardless of whether participants received study drug as intended.
| Feature | Registry-supported description |
|---|---|
| Randomization | Participants were randomized to palbociclib + letrozole or placebo + letrozole. |
| Primary efficacy population | ITT population. |
| Treatment assignment | Based on initial randomization. |
| Masking | Quadruple. |
| Primary analysis comparison | Palbociclib + Letrozole vs Placebo + Letrozole. |
| Stratification | Disease site, visceral versus non-visceral, per randomization. |
The ITT principle is important because it preserves the treatment comparison created by randomization. If participants discontinue treatment, deviate from the assigned treatment, or otherwise differ in exposure after randomization, analyzing them according to their randomized assignment avoids redefining the treatment groups based on events that occurred after randomization.
5. Primary Endpoint
| Endpoint | Registry definition / time frame | Analysis |
|---|---|---|
| Progression-Free Survival (PFS) Based on Investigator's Assessment | Randomization up to 65 months. PFS was based on Kaplan-Meier estimates. PFS was defined as the time from the date of randomization to the date of the first documentation of objective progression of disease (PD) or death due to any cause in the absence of documented PD, whichever occurred first. | Stratified log-rank test; hazard ratio |
The registry definition therefore combines two possible event pathways: objective disease progression and death in the absence of documented progression. The first qualifying event determines the PFS event time.
Kaplan-Meier estimation is appropriate because participants can have different lengths of follow-up and some may not have experienced the defined event by the time their observation ends.
6. Primary PFS Result
Investigator-assessed progression-free survival
95% CI: 0.529–0.867 · P = 0.0012
Two-sided superiority analysis; ITT population; stratified by disease site (visceral vs non-visceral).
| Primary endpoint | Estimate | 95% CI | P-value | Hypothesis |
|---|---|---|---|---|
| Investigator-assessed PFS | HR 0.677 | 0.529–0.867 | 0.0012 | Superiority |
The hazard ratio of 0.677 means that, under the proportional-hazards interpretation reported by the registry analysis, the estimated instantaneous rate of the PFS event was lower in the palbociclib + letrozole group than in the placebo + letrozole group. Expressed as a simple relative interpretation, 1 − 0.677 = 0.323, so the estimated hazard is about 32.3% lower under that model.
The HR does not mean that 32.3% of participants were prevented from progressing, nor does it mean that every participant had exactly a 32.3% reduction in individual risk. A hazard ratio is a relative time-to-event measure describing the comparison of event rates over follow-up.
The 95% confidence interval, 0.529 to 0.867, describes statistical uncertainty around the estimated hazard ratio. It does not describe the range of effects that individual participants experienced. Its location below 1 is consistent with the direction of the point estimate.
The p-value of 0.0012 addresses the strength of evidence against the relevant null hypothesis under the specified testing framework. It is not a measure of effect size, clinical magnitude, or the probability that the treatment works.
Finally, the registry analysis notes that the interpretation assumes proportional hazards. If the relative event rates vary materially over time rather than maintaining an approximately stable ratio, a single hazard ratio can compress a more complicated time-varying treatment effect into one summary number.
7. Secondary Time-to-Event Result: BICR PFS
PALOMA-4 also reported progression-free survival based on blinded independent central review (BICR). This was a secondary endpoint using the same broad randomization-to-progression-or-death time-to-event framework, with the registry specifying a follow-up period of randomization up to 65 months.
BICR progression-free survival
95% CI: 0.651–1.139 · P = 0.14778
Two-sided superiority analysis; ITT population; stratified by disease site (visceral vs non-visceral).
The BICR estimate of 0.861 points in the same numerical direction as the investigator-assessed PFS estimate, but the estimate is closer to 1 and its 95% confidence interval, 0.651 to 1.139, includes 1.
This result illustrates why different assessment procedures can matter statistically. The primary endpoint was based on investigator assessment, whereas this secondary analysis used blinded independent central review. The two analyses therefore address closely related outcomes through different assessment processes.
The p-value of 0.14778 should not be interpreted as a probability that the treatment effect is absent, nor should it be used to conclude that the two groups are identical. Rather, it quantifies the evidence against the specified null hypothesis under the reported analysis.
8. Secondary Response and Disease-Control Results
The registry posted several binary outcomes for objective response and disease control/clinical benefit response. These analyses used the Fisher exact test and reported odds ratios, with analyses stratified by disease site and conducted in the ITT population or the corresponding ITT subset with measurable disease at baseline.
| Endpoint | Population | OR | 95% CI | P-value |
|---|---|---|---|---|
| Objective response, investigator assessment | ITT | 1.301 | 0.805–2.100 | 0.154 |
| Objective response, investigator assessment | ITT with measurable disease at baseline | 1.255 | 0.762–2.066 | 0.206 |
| Objective response, BICR | ITT | 1.315 | 0.825–2.095 | 0.135 |
| Objective response, BICR | ITT with measurable disease at baseline | 1.392 | 0.825–2.346 | 0.117 |
| Disease control / clinical benefit response, investigator assessment | ITT | 0.945 | 0.533–1.673 | 0.471 |
| Disease control / clinical benefit response, investigator assessment | ITT with measurable disease at baseline | 0.878 | 0.474–1.621 | 0.383 |
| Disease control / clinical benefit response, BICR | ITT | 1.227 | 0.725–2.082 | 0.248 |
| Disease control / clinical benefit response, BICR | ITT with measurable disease at baseline | 1.349 | 0.731–2.509 | 0.189 |
How to read the odds ratios
For the objective-response analyses, the registry states that an odds ratio greater than 1 means better response in favor of the palbociclib + letrozole group. An odds ratio is a comparison of odds, not a direct comparison of response probabilities and not a hazard ratio.
The odds ratio can be useful for comparing binary outcomes, but its numerical value should not be read as a percentage increase in response probability.
Across these eight secondary binary analyses, the confidence intervals all span 1. That is important descriptive information about the uncertainty of these individual estimates. It does not erase the primary PFS result, because these endpoints answer different questions and occupy different roles in the trial's statistical structure.
9. Secondary Overall Survival Result
Overall survival
95% CI: 0.698–1.286 · P = 0.36502
Time frame: randomization up to 65 months; ITT population; stratified by disease site.
The OS hazard ratio of 0.947 is below 1, indicating a numerically lower estimated hazard of death in the palbociclib + letrozole group under the reported model. The corresponding 95% confidence interval is 0.698 to 1.286.
Because the interval includes 1, the estimate is compatible with a range of relative effects that includes no difference in the hazard ratio scale. The p-value of 0.36502 likewise should be understood as evidence against the specified null under the reported analysis, rather than as a probability that one treatment is or is not beneficial.
OS is also a different endpoint from PFS. PFS records the first qualifying progression or death, whereas OS records death from any cause. Consequently, the two endpoints need not produce identical treatment-effect estimates.
10. Patient-Reported Quality-of-Life Analyses
The registry also reported three continuous outcomes analyzed using repeated-measures mixed-effects models. These analyses evaluated model-estimated mean change from baseline through Cycle 65 Day 1 and used a subset of the ITT population.
| Endpoint | Mean difference | 95% CI | P-value |
|---|---|---|---|
| FACT-B Total Score | 0.476 | -2.97 to 3.92 | 0.7862 |
| EQ-5D Index Scores | 0.031 | -0.02 to 0.08 | 0.1914 |
| EQ Visual Analog Scale (VAS) Scores | 3.358 | 0.88 to 5.83 | 0.0078 |
The mixed-effects model included an intercept term, treatment, time, treatment-by-time, and baseline as a covariate. The registry states that a positive change indicates improvement from baseline and a negative change indicates deterioration.
FACT-B
The estimated mean difference was 0.476, with a 95% CI from -2.97 to 3.92 and P = 0.7862.
EQ-5D Index
The estimated mean difference was 0.031, with a 95% CI from -0.02 to 0.08 and P = 0.1914.
EQ VAS
The estimated mean difference was 3.358, with a 95% CI from 0.88 to 5.83 and P = 0.0078.
Why mixed models?
Repeated assessments from the same participant are correlated. A longitudinal mixed-effects model is designed to analyze that repeated-measures structure rather than treating every observation as independent.
11. Additional Overall Survival Analysis
The registry also posted an Other_Pre_Specified overall-survival analysis. It used the ITT population, a stratified log-rank analysis, and a hazard ratio as the effect measure.
| Analysis role | Endpoint | HR | 95% CI | P-value |
|---|---|---|---|---|
| Other pre-specified | Overall Survival: Up to Secondary Completion Date | 0.819 | 0.631–1.064 | 0.06764 |
The analysis was stratified by disease site (visceral versus non-visceral) per randomization. The registry notes that, assuming Cox proportional hazards, a hazard ratio less than 1 indicates a reduction in hazard rate in favor of palbociclib + letrozole.
The ClinicalTrials.gov record contains both a secondary OS analysis with HR 0.947 and an other pre-specified OS analysis with HR 0.819. They should not be silently combined into one estimate. Their registry labels and reported time-frame descriptions differ, so the appropriate interpretation is to preserve each analysis as its own reported result.
This is a general clinical-trial principle: an analysis estimate is inseparable from its endpoint definition, analysis population, time frame, statistical method, and role in the prespecified hierarchy.
12. Statistical Methodology
Kaplan-Meier estimation
The primary PFS endpoint was based on Kaplan-Meier estimates. Kaplan-Meier estimation is designed for time-to-event data in which some participants may not have experienced the event by the end of their available follow-up.
Here di represents the number of events at a given event time and ni represents the number at risk immediately before that time.
The practical value of Kaplan-Meier estimation is that it uses information from participants up to their censoring time instead of requiring every participant to experience an event. This is particularly important when follow-up duration differs across participants.
Stratified log-rank test
The reported PFS and OS analyses used the log-rank test, with the analyses stratified by disease site: visceral versus non-visceral. Stratification allows the treatment comparison to account for the factor used in randomization without treating that factor as though it were a post-randomization subgroup selected after seeing the results.
Hazard ratio
The hazard ratio summarizes the relative instantaneous event rate between the treatment groups. In the registry's interpretation, a hazard ratio below 1 indicates a lower hazard in favor of palbociclib + letrozole.
A hazard ratio is not a relative risk, an absolute risk difference, a difference in median survival, or the proportion of participants who benefit.
Fisher exact test
Fisher exact testing was used for the binary objective-response and disease-control/clinical-benefit outcomes. This method evaluates the treatment-group association in a two-by-two categorical framework and is particularly useful when exact inference is preferred for binary comparisons.
Repeated-measures mixed-effects model
The quality-of-life outcomes were analyzed with a repeated-measures mixed-effects model containing an intercept, treatment, time, treatment-by-time interaction, and baseline as a covariate.
The treatment-by-time term is important because repeated measurements can evolve differently over the course of follow-up. A longitudinal model can therefore represent the fact that the treatment comparison may depend on when the measurement is taken.
Covariate adjustment
Baseline was included as a covariate in the mixed-effects analyses. Conceptually, baseline adjustment compares longitudinal outcomes while accounting for the participants' starting measurements. This can improve statistical efficiency when baseline values are related to subsequent measurements.
13. Statistical Methods Explained
Why was the primary PFS endpoint analyzed with Kaplan-Meier methods?
PFS is a time-to-event endpoint, not simply a binary outcome observed at one fixed date. Kaplan-Meier estimation accommodates variable follow-up and censoring, allowing the analysis to use the timing of progression or death rather than reducing every participant to a single yes/no status.
Why use a hazard ratio rather than simply comparing the number of progressions?
A participant who progresses early and a participant who progresses much later are not equivalent observations for a time-to-event endpoint. The hazard-ratio framework incorporates the timing of events and censoring. It therefore summarizes more information than a simple event proportion, although it also requires careful interpretation and an appropriate model framework.
What does HR 0.677 mean?
Under the registry's proportional-hazards interpretation, HR 0.677 corresponds to a lower estimated instantaneous PFS event rate in the palbociclib + letrozole group. The simple transformation 1 − 0.677 = 0.323 gives an approximately 32.3% lower estimated hazard. It does not mean that 32.3% of participants avoided progression.
Why is the 95% confidence interval important?
The confidence interval provides information about precision around the estimated effect. For the primary PFS analysis, the interval was 0.529 to 0.867. A point estimate alone can hide substantial uncertainty; the interval shows the range of hazard-ratio values compatible with the specified statistical framework at the stated confidence level.
Why is the p-value not a measure of effect size?
The p-value depends on both the observed treatment contrast and the amount of information available. A small p-value does not tell us that the treatment effect is large, while a larger p-value does not prove that the treatment effect is exactly zero. Effect magnitude is better described with the estimated hazard ratio or odds ratio together with its confidence interval.
Why did the analysis use stratification by disease site?
The registry states that analyses were stratified by visceral versus non-visceral disease site per randomization. Stratification aligns the analysis with the randomized structure and can account for baseline differences in event experience associated with the stratification factor.
Why use a mixed-effects model for quality-of-life measurements?
Quality-of-life measurements were repeated over time within the same participants. Those observations are correlated because they come from the same individuals. A mixed-effects model provides a framework for modeling that longitudinal structure while incorporating treatment, time, treatment-by-time, and baseline terms.
14. Understanding the Primary Result in Context
The primary PFS result and the other posted outcomes should be read as a collection of related but distinct statistical questions rather than as a single number.
| Question | Reported analysis | What it tells us |
|---|---|---|
| Does treatment affect progression-free survival? | Investigator-assessed PFS; HR 0.677 | Relative time-to-event comparison for progression or death. |
| Does the same pattern appear under central review? | BICR PFS; HR 0.861 | A secondary time-to-event assessment using blinded independent central review. |
| How do binary response outcomes compare? | Fisher exact tests; odds ratios | Comparisons of objective response and disease-control/clinical-benefit response. |
| Does overall survival differ? | OS; HR 0.947 | A separate time-to-death comparison. |
| How do longitudinal patient-reported measures compare? | Mixed-effects models | Model-estimated mean changes from baseline over repeated assessments. |
The distinction matters because no single statistical measure captures every dimension of a randomized trial. PFS, OS, tumor response, disease control, and quality-of-life outcomes have different endpoint structures and therefore require different statistical tools.
15. Safety Results
The ClinicalTrials.gov record reports serious adverse events by treatment arm as affected participants over participants at risk.
| Safety measure | Palbociclib + Letrozole | Placebo + Letrozole |
|---|---|---|
| Serious adverse events, affected / at risk | 34/168 | 19/171 |
These are exposure-related safety counts rather than efficacy endpoints. The denominators also differ from the overall enrollment figure of 340, and the ClinicalTrials.gov record does not provide enough information to reconstruct why those denominators differ. They should therefore be reported as given rather than converted into an inferred treatment allocation.
16. Limitations and Interpretation Issues
- Single primary endpoint: the registry identifies one registered primary endpoint, investigator-assessed PFS. Other outcomes should not be silently promoted to primary status.
- Different PFS assessment methods: investigator-assessed PFS and BICR PFS produced different estimates. These should be interpreted as distinct analyses rather than averaged or reconciled into a single effect.
- Proportional-hazards assumption: the registry's interpretation of the hazard ratio depends on the proportional-hazards framework. If hazards vary substantially over time, a single HR may not describe the complete treatment-effect pattern.
- Censoring: time-to-event analyses rely on censoring information. The statistical interpretation therefore depends on the assumptions underlying how censored observations relate to the event process.
- Secondary endpoints: the response, disease-control, OS, and quality-of-life analyses are secondary analyses and should not be interpreted as interchangeable with the primary PFS result.
- Multiple analyses: the registry contains 15 statistical analyses. The ClinicalTrials.gov record does not provide a complete multiplicity-adjustment hierarchy, so individual secondary p-values should not automatically be interpreted as independent confirmatory tests.
- Different analysis populations: most efficacy analyses use the ITT population, while the quality-of-life analyses use an ITT subset and the safety results use arm-specific at-risk denominators.
- Endpoint scale: odds ratios, hazard ratios, and mean differences are different effect measures. Their numerical values cannot be compared directly as though they represented the same quantity.
- Registry scope: the ClinicalTrials.gov record does not provide baseline characteristics, median PFS, median OS, subgroup forest plots, or detailed safety tables. Those results are therefore not included here.
17. Why This Trial Matters Statistically
PALOMA-4 is a useful teaching case because the registry combines a randomized time-to-event primary endpoint with several different secondary endpoint structures. That creates a compact example of how endpoint type should determine statistical methodology.
| Concept | How it appears in PALOMA-4 |
|---|---|
| Randomization | Participants were randomized to palbociclib + letrozole or placebo + letrozole. |
| Quadruple masking | The registry classifies the study as quadruple-masked. |
| ITT analysis | The primary PFS analysis was conducted in the ITT population. |
| Kaplan-Meier estimation | Used for the primary PFS endpoint. |
| Hazard ratio | Used for PFS and OS treatment-effect estimation. |
| Stratified analysis | Time-to-event and binary analyses were stratified by disease site, visceral versus non-visceral. |
| Fisher exact test | Used for objective-response and disease-control/clinical-benefit response outcomes. |
| Odds ratio | Used as the effect measure for the binary response analyses. |
| Mixed-effects model | Used for repeated quality-of-life measurements. |
| Covariate adjustment | Baseline was included as a covariate in the repeated-measures mixed-effects model. |
| Multiple endpoint types | The registry contains time-to-event, binary, and continuous outcomes, each analyzed with a method appropriate to its structure. |
18. Relative Effects and Absolute Meaning
A central statistical lesson from PALOMA-4 is that a hazard ratio should not be treated as a complete description of treatment benefit.
The primary hazard ratio was 0.677. Under the registry's proportional-hazards interpretation, this represents a lower estimated instantaneous rate of progression or death in the palbociclib + letrozole group.
It does not tell us the median PFS, the proportion of participants who remained progression-free at a particular time point, or the treatment effect for every individual participant. Those quantities require corresponding absolute or distributional results.
The 95% CI of 0.529–0.867 describes uncertainty around the primary hazard-ratio estimate. The interval provides more information than the point estimate alone because it indicates how precisely the treatment effect was estimated within the reported statistical framework.
The p-value of 0.0012 concerns evidence against the null hypothesis. It does not quantify the size of the treatment effect. The HR and confidence interval are the appropriate quantities for describing the magnitude and precision of the estimated relative effect.
19. Reading the Secondary Results Without Overinterpreting Them
The secondary results demonstrate why statistical interpretation should be endpoint-specific.
PFS by BICR
HR 0.861 with 95% CI 0.651–1.139. This is a time-to-event estimate from a different assessment process than the primary investigator-assessed endpoint.
Objective response
The posted odds ratios range from 1.255 to 1.392 for the four objective-response analyses, with confidence intervals that include 1.
Disease control / clinical benefit
The posted odds ratios range from 0.878 to 1.349 across the four analyses, again with confidence intervals that include 1.
Overall survival
The secondary OS analysis reports HR 0.947, 95% CI 0.698–1.286, and P = 0.36502, while the other pre-specified OS analysis reports HR 0.819, 95% CI 0.631–1.064, and P = 0.06764.
The appropriate conclusion is therefore not to collapse all secondary results into a single "positive" or "negative" label. Each result has its own endpoint definition, population, estimator, uncertainty interval, and hypothesis-testing context.
20. Statistical Interpretation vs Clinical Interpretation
Statistical interpretation
The primary investigator-assessed PFS analysis reported HR 0.677 with a 95% CI of 0.529–0.867 and P = 0.0012 using a stratified log-rank analysis in the ITT population. The estimate was below 1 and the confidence interval excluded 1.
Clinical interpretation
The statistical result describes a relative difference in the time-to-event process. Determining its clinical importance requires considering the absolute magnitude and duration of benefit, patient experience, competing outcomes, safety, and other clinical evidence; those additional quantities are not fully reported in the ClinicalTrials.gov record.
This distinction is important. Statistical significance and clinical importance are related but not synonymous. A p-value answers a hypothesis-testing question, whereas clinical interpretation asks whether the magnitude, duration, and consequences of an effect are meaningful in the context of patient care.
21. Overall Statistical Takeaway
PALOMA-4 provides a clear example of a randomized phase 3 trial centered on a time-to-event endpoint. The primary investigator-assessed PFS analysis used an ITT population and a stratified log-rank approach, producing a hazard ratio of 0.677 with a two-sided 95% confidence interval of 0.529–0.867 and a p-value of 0.0012.
The broader statistical picture is more nuanced than the primary result alone. A secondary BICR PFS analysis reported HR 0.861; binary response and disease-control analyses used Fisher exact tests and odds ratios; OS was analyzed with a stratified log-rank framework and reported separately; and quality-of-life outcomes were analyzed with repeated-measures mixed-effects models incorporating treatment, time, treatment-by-time, and baseline.
For statistical interpretation, the most important lesson is that the endpoint determines the analysis. Time-to-event endpoints require methods that account for censoring and event timing. Binary outcomes require categorical-data methods. Repeated continuous measurements require methods that recognize within-participant correlation. Treating all of these outcomes as though they were the same statistical object would discard important information about the trial design.
22. Related Tutorials
Learn more about the methods used in this trial:
23. Related Calculators
24. Sources
- ClinicalTrials.gov: PALOMA-4, NCT02297438.
- PubMed: Publication indexed under PMID 39694573.
- PubMed: Publication indexed under PMID 36155117.
Continue through the Clinical Biostats knowledge graph
Connect the endpoints and statistical methods in this trial to deeper tutorials, calculators, and other clinical-trial analyses.
25. Record Summary
PALOMA-4 is a randomized phase 3, parallel, quadruple-masked trial with 340 participants enrolled to compare palbociclib + letrozole with placebo + letrozole. Its registered primary endpoint was investigator-assessed progression-free survival from randomization up to 65 months. The primary analysis used an ITT population and a stratified log-rank framework, reporting HR 0.677, 95% CI 0.529–0.867, and P = 0.0012.
The trial also illustrates the importance of matching statistical methods to endpoint structure. Secondary PFS and OS outcomes used time-to-event methods, response and disease-control outcomes used Fisher exact tests with odds ratios, and repeated quality-of-life outcomes used mixed-effects models with baseline covariate adjustment. Interpreting the trial correctly therefore requires more than reading a single p-value: the endpoint, estimand, analysis population, effect measure, confidence interval, and statistical model all contribute to the meaning of the result.