Understanding Effect Sizes

What are Effect Sizes?

Effect sizes quantify the magnitude of a difference between groups or the strength of a relationship. A p-value summarizes how compatible the observed result is with a null hypothesis under a statistical model; it does not describe the effect's magnitude or clinical importance.

Continuous Outcome Effect Sizes

Mean Difference (MD)

Simple difference between group means:

  • Formula: Mean_intervention - Mean_control
  • Units: Same as original measurement
  • Use when: All studies use identical scale
  • Example: If intervention group mean HbA1c = 7.0% and control = 7.8%, MD = -0.8%
  • Interpretation: Direct, clinically meaningful difference

Standardized Mean Difference (SMD)

Difference expressed in standard deviation units:

  • Also called Cohen's d or Hedges' g
  • Formula: (Mean_intervention - Mean_control) / Pooled_SD
  • Units: Standard deviation units (unit-free)
  • Use when: Studies use different scales for same construct
  • Example: Combining BDI, HAM-D, and PHQ-9 depression scores

Interpreting SMD (Cohen's d)

  • 0.2 = Small effect (subtle difference)
  • 0.5 = Medium effect (noticeable difference)
  • 0.8 = Large effect (obvious difference)
  • These are guidelines, not rules; consider clinical context
  • Negative values favor intervention (if lower is better)

SMD is unit-free but harder to interpret clinically. When possible, prefer MD and provide clinical context for interpreting the magnitude.

Binary Outcome Effect Sizes

Risk Ratio (RR) / Relative Risk

Ratio of event probabilities:

  • Formula: (Events_intervention / N_intervention) / (Events_control / N_control)
  • Range: 0 to infinity
  • RR = 1: The event is equally likely in both groups
  • RR < 1: The event is less likely in the intervention group
  • RR > 1: The event is more likely in the intervention group
  • Example: RR = 0.75 means the event risk in the intervention group is 25% lower than in the control group
  • Whether a lower or higher RR is favorable depends on the outcome and comparison order

Odds Ratio (OR)

Ratio of odds of event:

  • Formula: (Odds_intervention) / (Odds_control)
  • Where Odds = Events / Non-events
  • Range: 0 to infinity
  • OR = 1: No difference
  • OR < 1: Intervention reduces odds
  • Approximates RR when events are rare (<10%)
  • Less intuitive than RR but mathematically convenient

Risk Difference (RD)

Absolute difference in event rates:

  • Formula: Risk_intervention - Risk_control
  • Range: -1 to +1 (or -100% to +100%)
  • RD = 0: No difference
  • Negative RD: Intervention reduces risk
  • Example: RD = -0.10 means 10% absolute risk reduction
  • Most clinically interpretable
  • Needed to calculate Number Needed to Treat (NNT)

Number Needed to Treat (NNT)

  • Formula: 1 / |RD|
  • Interpretation: Number of patients to treat to prevent one event
  • Example: NNT = 10 means treating 10 patients prevents 1 adverse event
  • Lower NNT = more effective intervention
  • Highly clinically meaningful
  • Can also calculate Number Needed to Harm (NNH) for adverse effects

Choosing the Right Effect Size

For Continuous Outcomes

  • Same scale across studies? → Use MD
  • Different scales for same construct? → Use SMD
  • Want clinically meaningful interpretation? → Prefer MD when possible
  • Consider minimal clinically important difference (MCID)

For Binary Outcomes

  • Cohort/RCT data? → RR or RD preferred
  • Case-control data? → OR (only option)
  • Clinical interpretation priority? → RD and calculate NNT
  • Rare events (<10%)? → RR and OR will be similar
  • Common events? → RR and RD more interpretable than OR

Confidence Intervals

Effect sizes should always be reported with 95% confidence intervals (CI):

  • CI shows range of plausible values
  • Wide CI = imprecise estimate (small sample or high variance)
  • Narrow CI = precise estimate
  • If CI excludes null, effect is statistically significant (p < 0.05)
  • Clinical significance depends on whether CI includes clinically meaningful values

Statistical vs Clinical Significance

An effect can be:

  • Statistically significant but clinically trivial (large sample, small effect)
  • Clinically important but not statistically significant (small sample, promising effect)
  • Both statistically and clinically significant (ideal)
  • Neither (no effect)
  • Always interpret effect sizes in clinical context, not just p-values

Calculating Effect Sizes in Scholara

  1. 1Ask the Analysis Chat to extract outcome data from your studies (means, SDs, sample sizes for continuous; event counts for binary)
  2. 2Request a meta-analysis specifying your preferred effect size metric
  3. 3The AI calculates effect sizes, handles variance estimation, and generates results
  4. 4Forest plot and summary statistics appear as interactive assets in the chat
  5. 5Ask follow-up questions to explore different metrics or sensitivity analyses

Handling Missing Data

When studies don't report all needed data:

  • Calculate from other statistics (t-test, F-test, p-values)
  • Estimate from graphs (not ideal but sometimes necessary)
  • Impute from similar studies (sensitivity analysis)
  • Contact authors for missing data
  • Exclude from meta-analysis if insufficient data
  • Document all calculations and assumptions

Converting Between Effect Sizes

Sometimes you need to convert between metrics:

  • OR to RR: RR = OR / [(1 - P0) + (P0 × OR)] where P0 = control event rate
  • SMD to OR: ln(OR) ≈ π/√3 × SMD (rough approximation)
  • Use conversion formulas cautiously and document
  • Scholara provides conversion tools when needed

Reporting Effect Sizes

In your systematic review, report:

  • Effect size metric used and justification
  • Point estimate with 95% CI
  • Statistical significance (p-value)
  • Clinical interpretation in context
  • Heterogeneity measures (I², tau²)
  • Certainty of evidence (GRADE)
  • Example: 'SMD = -0.42, 95% CI [-0.68, -0.16], p = 0.001, indicating a moderate reduction in depressive symptoms favoring CBT (moderate certainty evidence)'

Effect sizes are more informative than p-values alone. Always interpret both statistical and clinical significance. Consider the minimal clinically important difference (MCID) for your outcome when determining whether an effect is meaningful.

References