What are Effect Sizes?
Effect sizes quantify the magnitude of a difference between groups or the strength of a relationship. A p-value summarizes how compatible the observed result is with a null hypothesis under a statistical model; it does not describe the effect's magnitude or clinical importance.
Continuous Outcome Effect Sizes
Mean Difference (MD)
Simple difference between group means:
- Formula: Mean_intervention - Mean_control
- Units: Same as original measurement
- Use when: All studies use identical scale
- Example: If intervention group mean HbA1c = 7.0% and control = 7.8%, MD = -0.8%
- Interpretation: Direct, clinically meaningful difference
Standardized Mean Difference (SMD)
Difference expressed in standard deviation units:
- Also called Cohen's d or Hedges' g
- Formula: (Mean_intervention - Mean_control) / Pooled_SD
- Units: Standard deviation units (unit-free)
- Use when: Studies use different scales for same construct
- Example: Combining BDI, HAM-D, and PHQ-9 depression scores
Interpreting SMD (Cohen's d)
- 0.2 = Small effect (subtle difference)
- 0.5 = Medium effect (noticeable difference)
- 0.8 = Large effect (obvious difference)
- These are guidelines, not rules; consider clinical context
- Negative values favor intervention (if lower is better)
SMD is unit-free but harder to interpret clinically. When possible, prefer MD and provide clinical context for interpreting the magnitude.
Binary Outcome Effect Sizes
Risk Ratio (RR) / Relative Risk
Ratio of event probabilities:
- Formula: (Events_intervention / N_intervention) / (Events_control / N_control)
- Range: 0 to infinity
- RR = 1: The event is equally likely in both groups
- RR < 1: The event is less likely in the intervention group
- RR > 1: The event is more likely in the intervention group
- Example: RR = 0.75 means the event risk in the intervention group is 25% lower than in the control group
- Whether a lower or higher RR is favorable depends on the outcome and comparison order
Odds Ratio (OR)
Ratio of odds of event:
- Formula: (Odds_intervention) / (Odds_control)
- Where Odds = Events / Non-events
- Range: 0 to infinity
- OR = 1: No difference
- OR < 1: Intervention reduces odds
- Approximates RR when events are rare (<10%)
- Less intuitive than RR but mathematically convenient
Risk Difference (RD)
Absolute difference in event rates:
- Formula: Risk_intervention - Risk_control
- Range: -1 to +1 (or -100% to +100%)
- RD = 0: No difference
- Negative RD: Intervention reduces risk
- Example: RD = -0.10 means 10% absolute risk reduction
- Most clinically interpretable
- Needed to calculate Number Needed to Treat (NNT)
Number Needed to Treat (NNT)
- Formula: 1 / |RD|
- Interpretation: Number of patients to treat to prevent one event
- Example: NNT = 10 means treating 10 patients prevents 1 adverse event
- Lower NNT = more effective intervention
- Highly clinically meaningful
- Can also calculate Number Needed to Harm (NNH) for adverse effects
Choosing the Right Effect Size
For Continuous Outcomes
- Same scale across studies? → Use MD
- Different scales for same construct? → Use SMD
- Want clinically meaningful interpretation? → Prefer MD when possible
- Consider minimal clinically important difference (MCID)
For Binary Outcomes
- Cohort/RCT data? → RR or RD preferred
- Case-control data? → OR (only option)
- Clinical interpretation priority? → RD and calculate NNT
- Rare events (<10%)? → RR and OR will be similar
- Common events? → RR and RD more interpretable than OR
Confidence Intervals
Effect sizes should always be reported with 95% confidence intervals (CI):
- CI shows range of plausible values
- Wide CI = imprecise estimate (small sample or high variance)
- Narrow CI = precise estimate
- If CI excludes null, effect is statistically significant (p < 0.05)
- Clinical significance depends on whether CI includes clinically meaningful values
Statistical vs Clinical Significance
An effect can be:
- Statistically significant but clinically trivial (large sample, small effect)
- Clinically important but not statistically significant (small sample, promising effect)
- Both statistically and clinically significant (ideal)
- Neither (no effect)
- Always interpret effect sizes in clinical context, not just p-values
Calculating Effect Sizes in Scholara
- 1Ask the Analysis Chat to extract outcome data from your studies (means, SDs, sample sizes for continuous; event counts for binary)
- 2Request a meta-analysis specifying your preferred effect size metric
- 3The AI calculates effect sizes, handles variance estimation, and generates results
- 4Forest plot and summary statistics appear as interactive assets in the chat
- 5Ask follow-up questions to explore different metrics or sensitivity analyses
Handling Missing Data
When studies don't report all needed data:
- Calculate from other statistics (t-test, F-test, p-values)
- Estimate from graphs (not ideal but sometimes necessary)
- Impute from similar studies (sensitivity analysis)
- Contact authors for missing data
- Exclude from meta-analysis if insufficient data
- Document all calculations and assumptions
Converting Between Effect Sizes
Sometimes you need to convert between metrics:
- OR to RR: RR = OR / [(1 - P0) + (P0 × OR)] where P0 = control event rate
- SMD to OR: ln(OR) ≈ π/√3 × SMD (rough approximation)
- Use conversion formulas cautiously and document
- Scholara provides conversion tools when needed
Reporting Effect Sizes
In your systematic review, report:
- Effect size metric used and justification
- Point estimate with 95% CI
- Statistical significance (p-value)
- Clinical interpretation in context
- Heterogeneity measures (I², tau²)
- Certainty of evidence (GRADE)
- Example: 'SMD = -0.42, 95% CI [-0.68, -0.16], p = 0.001, indicating a moderate reduction in depressive symptoms favoring CBT (moderate certainty evidence)'
Effect sizes are more informative than p-values alone. Always interpret both statistical and clinical significance. Consider the minimal clinically important difference (MCID) for your outcome when determining whether an effect is meaningful.