Posters · Glossary
What is statistical power?
Statistical power is the probability that a study will detect an effect of a stated size if that effect really exists, usually set at 80 or 90 percent. It depends on the size of the effect, the sample size, the variability of the data and the significance level. Low power means a real effect is easily missed.
Power is decided before a study begins, which is why it is so often forgotten by the time results reach a poster. Yet it is the number that tells a reader whether a null result means anything and whether a surprising positive result should be trusted.
Nuwan Madhusanka · Co-founder
5 min read · Published
| Participants per group | Total participants | Power | Chance of missing a true effect of this size |
|---|---|---|---|
| 20 | 40 | 34 percent | 66 percent |
| 30 | 60 | 48 percent | 52 percent |
| 50 | 100 | 70 percent | 30 percent |
| 64 | 128 | 80 percent | 20 percent |
| 86 | 172 | 90 percent | 10 percent |
| 100 | 200 | 94 percent | 6 percent |
The four quantities that move together
UCLA's statistical consulting seminar on power analysis defines power as the probability of detecting an effect, given that the effect is really there, which is the probability of rejecting the null hypothesis when it is false. It also sets out the relationship that drives planning: effect size, sample size, significance level and power are linked, so if three are fixed the fourth is determined. Larger effects are easier to detect, larger samples give more power, and a stricter significance level needs more participants for the same power. Most recommendations for power fall between 0.8 and 0.9.
Reading the table
The table shows power for a two group comparison of means with a standardised effect of 0.5, calculated with the noncentral t distribution. With 20 people per group, a real effect of that size would be missed about two times in three. Reaching 80 percent power takes 64 per group, and 90 percent takes 86. The gains shrink as the sample grows: moving from 64 to 100 per group adds only 14 points of power. A smaller expected effect changes everything; halving the effect to 0.25 roughly quadruples the sample needed for the same power. For yes or no outcomes the same logic applies with proportions: detecting a drop from 20 to 15 percent needs far more participants than a drop from 20 to 10 percent, because the absolute difference is half as large.
Why underpowered studies mislead
Low power does not only cause missed effects. Button and colleagues, in a 2013 analysis in Nature Reviews Neuroscience titled Power failure, estimated that the median power of neuroscience studies they examined was low, and showed two further consequences. A significant result from a low powered study is less likely to reflect a true effect, and when an underpowered study does find a true effect, it tends to overestimate its size, because only the lucky large estimates cross the significance line. That is one reason small, striking studies so often fail to replicate at the same effect size.
Planning and reporting a sample size
The CONSORT statement for randomised trials asks authors to report how the sample size was determined, and similar items appear in other reporting guidelines. A complete statement names the primary outcome, the difference the study was designed to detect and why it matters, the assumed standard deviation or control group rate, the significance level, the power, and any allowance for dropouts. On a poster this fits in one methods entry. Choosing the target difference is the hard part: it should be the smallest effect worth knowing about, not the effect researchers hope to see. If recruitment fell short, say so and give both the planned and the achieved numbers, because readers judge a result differently when a study designed for 90 percent power finished with half its intended sample.
Common mistakes
Calculating power after the study from the observed effect is the most common error. The UCLA seminar notes that power computed this way is directly related to the p value and adds no new information; a non significant result will always show low observed power. Other mistakes include assuming an optimistic effect size to justify a small sample, ignoring clustering in designs that randomise clinics or schools, forgetting dropouts, and describing a pilot or feasibility study as if it were powered to test effectiveness.
Where it shows up in the poster builder
Power belongs in the methods, and the protocol block, present on 38 of the 40 layouts, holds up to five labelled entries such as design, participants, intervention, outcome and analysis, so the sample size calculation fits as one line under analysis or participants. The generator writes methods from the brief, so give it the target difference, the power and the planned and achieved sample sizes. Where a board has a participant flow, on ten layouts, the analysed counts should match the numbers the power calculation assumed. The political science example reports a large randomised sample in its methods entries.
Questions people ask
Why is 80 percent power the convention?
It is a trade off inherited from the work of Jacob Cohen, who suggested that a one in five chance of missing a real effect was an acceptable balance against the cost of larger studies, alongside a one in twenty chance of a false positive. Funders and ethics committees increasingly expect 90 percent for confirmatory trials, where missing a real effect is costly.
What is the difference between power and significance level?
The significance level, usually 0.05, is the chance of a false positive: declaring an effect when none exists. Power is the chance of a true positive: detecting an effect that does exist. One minus power is the chance of a false negative. Both are set in planning, and tightening the significance level lowers power unless the sample grows.
Can a study be too powerful?
A very large study can detect effects too small to matter, which is not a flaw in itself but can produce significant results that are practically trivial. The remedy is to judge results by effect size and its interval against a meaningful threshold. Enrolling far more participants than needed also raises cost and, in trials, exposes more people to an unproven intervention.
Does power apply to qualitative research?
No. Power is a concept for hypothesis tests. Qualitative studies plan sample size around information power, depth of data or saturation, the point at which new interviews add no new themes. A mixed methods poster should describe each strand's sampling logic on its own terms rather than applying a power calculation to the qualitative component.
How do I estimate the effect size for a power calculation?
Use the smallest effect that would change practice or be worth knowing, informed by a clinically important difference, previous studies or a pilot. Effects from small published studies are often inflated, so treat them cautiously. Avoid choosing the effect by working backwards from the sample you can afford, which produces a calculation that justifies the budget rather than the question.
What should I do if my finished study was underpowered?
Say so plainly and interpret the result through its confidence interval. If the interval is wide and includes both no effect and important effects, the study is inconclusive, not negative. Report what sample would be needed to settle the question, and consider whether the data can contribute to a future meta analysis alongside similar studies.
Make one with posters
The button opens the generator with this use case already described. Change the wording to match your own.
Create a poster with OneCraftRelated questions
- What is a p value?What is a p value: how incompatible data are with a null model, not the chance a finding is true. What 0.05 means and how to report one on a poster.
- How to present a null result on a posterHow to present a null result: give the estimate and its interval, say which effects it rules out, and word it honestly on a research poster.
- What is effect size?What is effect size: how big a difference or association is, not whether it passed a significance test. Cohen's d, risk and odds ratios, and how to report them.
- What is an RCT (randomised controlled trial)?What is an RCT: a study that randomly allocates people to an intervention or a control so groups differ only by chance. How it limits bias, and what it costs.
Step by step in the builder: Add charts and diagrams.
Written and checked by the OneCraft team. Last checked .