This calculator determines the sample size needed for a chi-square test of independence to detect an association between two categorical variables with the target power, based on Cohen's \(w\) and the dimensions of the table.
Calculator
Enter Cohen's w, the table's degrees of freedom, the significance level and the desired power.
Effect size of the association between the categorical variables. Indicative values: 0.1 small, 0.3 medium (the most common), 0.5 large.
Degrees of freedom of the contingency table: (no. of rows − 1) × (no. of columns − 1). For a 2×2 table, df = 1.
Significance level of the test. The standard value is 0.05.
Probability of detecting the association if it really exists. 0.80 is the usual minimum; 0.90 for confirmatory studies.
Wizard for 2×2 tables: calculate w from p₁ and p₂
Enter the proportions of the two groups to automatically calculate w and update it in the field above.
Proportion observed or expected in the first group, for 2×2 tables.
Proportion observed or expected in the second group, for 2×2 tables.
Explanation
The chi-square test of independence answers this question: is the distribution of one categorical variable the same across all categories of another? The data are arranged in a contingency table with \(r\) rows and \(c\) columns, and each observation falls into exactly one cell. The null hypothesis \(H_0\) is that the two variables are independent: knowing an observation's row does not change the probability of each column (for example, that product preference is split the same way among men and women). The alternative \(H_1\) is that some association exists.
If the variables were independent, the expected proportion in cell \((i, j)\) would be the product of the marginal proportions, \(p_{0,ij} = p_{i\cdot} \cdot p_{\cdot j}\), and the expected frequency is obtained from the row and column totals. The statistic compares each observed frequency \(O_{ij}\) with the expected one \(E_{ij}\):
\( \chi^2 = \sum_{i=1}^{r}\sum_{j=1}^{c} \frac{(O_{ij} - E_{ij})^2}{E_{ij}}, \qquad E_{ij} = \frac{(\text{row } i \text{ total})\,(\text{column } j \text{ total})}{N} \)
Each term measures one cell's departure from independence: the difference is squared so that deviations in both directions accumulate, and it is divided by \(E_{ij}\) to weight it by the size of the cell. Under \(H_0\) the statistic follows a chi-square distribution with \(df = (r-1)(c-1)\) degrees of freedom. The reason: once the row and column totals are fixed, only \((r-1)(c-1)\) cells can vary freely; the rest are determined by subtraction. In a \(2 \times 2\) table, for instance, knowing one cell is enough to reconstruct the other three, which is why \(df = 1\).
The logic of the power calculation
Before collecting the data we do not know what value \(\chi^2\) will take, but we can reason about the two possible scenarios:
- If \(H_0\) is true, the statistic follows the central chi-square distribution with \(df\) degrees of freedom. We fix the significance level \(\alpha\) and with it the critical value \(\chi^2_{1-\alpha,\,df}\): the point that the central chi-square exceeds with probability \(\alpha\) only. We will reject \(H_0\) if the statistic lands above it. With \(df = 2\) and \(\alpha = 0.05\), that value is \(5.99\); with \(df = 1\), \(3.84\).
- If \(H_1\) is true (the variables are associated), the observed frequencies drift systematically away from those expected under independence and the statistic tends to be larger. The power is the probability that, in this scenario, the statistic exceeds the critical value and we therefore detect the association.
The critical value depends only on \(\alpha\) and \(df\); it does not change with the sample size. What does change with \(N\) is how far the statistic shifts to the right when \(H_1\) is true. To compute the power we therefore need to describe the distribution of the statistic under \(H_1\). We do this in three steps: measure the strength of the association (\(w\)), translate it into a shift of the distribution (\(\lambda\)), and compute the area that lies to the right of the critical value.
Step 1: measuring the association with Cohen's \(w\)
We need a number that says how far from independence the real proportions of the table are, and that does not depend on \(N\). Cohen's \(w\) applies the recipe of the statistic to the cell proportions: it compares the proportions you expect under \(H_1\), \(p_{1,ij}\), with those an independent table with the same margins would have, \(p_{0,ij} = p_{i\cdot} \cdot p_{\cdot j}\):
\( w = \sqrt{\sum_{i=1}^{r}\sum_{j=1}^{c} \frac{(p_{1,ij} - p_{0,ij})^2}{p_{0,ij}}} \)
If a sample of \(N\) observations reproduced the proportions \(p_{1,ij}\) exactly, its chi-square statistic would be exactly \(N \cdot w^2\). That is why it helps to think of \(w^2\) as the chi-square statistic per observation or, equivalently, \(w = \sqrt{\chi^2 / N}\): it is a property of the table of proportions, not of the sample. In \(2 \times 2\) tables, \(w\) coincides with the \(\phi\) coefficient, the correlation between the two binary variables. Cohen's (1988) benchmarks: \(w = 0.1\) (weak association), \(0.3\) (medium), \(0.5\) (strong).
Special case: a \(2 \times 2\) table with two groups of equal size. This is the most common situation (treatment versus control, with success proportions \(p_1\) and \(p_2\)) and the one the wizard solves. If each group contributes half of the sample, the proportions of the four cells under \(H_1\) are \(p_1/2\), \((1-p_1)/2\), \(p_2/2\) and \((1-p_2)/2\); under independence, both rows share the average proportion \(\bar{p} = (p_1 + p_2)/2\). Substituting into the general formula and simplifying (the four terms pair up two by two) gives:
\( w = \frac{|p_1 - p_2|}{2\sqrt{\bar{p}\,(1-\bar{p})}}, \qquad \bar{p} = \frac{p_1 + p_2}{2} \)
Read it as "half the difference between the two groups, measured in standard deviations of the average proportion", since \(\sqrt{\bar{p}(1-\bar{p})}\) is the standard deviation of a 0/1 variable with probability \(\bar{p}\). If the groups will not be of equal size, compute \(w\) with the general formula using the real row proportions.
Step 2: from effect to shift, \(\lambda = N \cdot w^2\)
When \(H_1\) is true, the statistic no longer follows the central chi-square. It follows a noncentral chi-square: the same family of distributions, but shifted to the right and more spread out. The size of the shift is governed by the noncentrality parameter \(\lambda\), which for this test is:
\( \lambda = N \cdot w^2 \qquad\text{and therefore}\qquad E\!\left[\chi^2 \mid H_1\right] = df + \lambda \)
A useful way to see it is to decompose the expected value of the statistic under \(H_1\), which is \(df + \lambda\). The term \(df\) is the "noise" part: what the statistic is worth on average even when \(H_0\) is true, purely from sampling chance. The term \(\lambda\) is the "signal" part: what the real association contributes. And that signal grows in proportion to \(N\), because with twice as many observations the systematic distance between observed and expected doubles, while the critical value stays where it is. This is exactly the mechanism by which a larger sample gives more power: the distribution under \(H_1\) moves away from the critical value and more and more area is left to its right.
This also yields the most important practical rule: for a given power, the required \(\lambda\) is nearly constant, so \(N \approx \lambda / w^2\). The sample size is inversely proportional to the square of the effect: detecting an association half as strong requires four times as many observations.
Step 3: power as the area to the right of the critical value
\( \text{Power} = P\!\left(\chi^2 > \chi^2_{1-\alpha,\,df} \mid H_1\right) = 1 - F_{\chi^2_{df,\,\lambda}}\!\left(\chi^2_{1-\alpha,\,df}\right) \)
Reading from right to left: \(F_{\chi^2_{df,\lambda}}\) is the cumulative distribution function of the noncentral chi-square with \(df\) degrees of freedom and noncentrality \(\lambda\). Evaluated at the critical value, it gives the probability that the statistic does not exceed it even though \(H_1\) is true, i.e. the probability of a type II error, \(\beta\). The power is its complement, \(1 - \beta\).
How the noncentral chi-square is computed
The noncentral chi-square has no simple closed-form formula, but it does have a representation that the calculator exploits: it is a mixture of central chi-squares with Poisson weights. Imagine first drawing an integer \(J\) from a Poisson distribution with mean \(\lambda/2\), and then generating a central chi-square with \(df + 2J\) degrees of freedom. The resulting variable is exactly a noncentral chi-square with parameters \(df\) and \(\lambda\). Its cumulative distribution function is therefore the weighted average of central cumulative distribution functions:
\( F_{\chi^2_{df,\lambda}}(x) = \sum_{j=0}^{\infty} \underbrace{\frac{e^{-\lambda/2}(\lambda/2)^j}{j!}}_{\text{Poisson weight}} \cdot \underbrace{F_{\chi^2_{df+2j}}(x)}_{\text{central chi-square}} \)
Each term combines two familiar ingredients: the Poisson probability that \(J = j\) and the cumulative distribution function of a central chi-square with \(df + 2j\) degrees of freedom. With \(df = 2\) and \(\lambda = 9.72\) (the values in the example below), the largest weights belong to \(j = 4\) and \(j = 5\) (\(0.18\) and \(0.17\)), i.e. to central chi-squares with 10 and 12 degrees of freedom, which exceed the critical value \(5.99\) easily; the weights become negligible from \(j \approx 15\) onwards. The calculator adds terms until the weight drops below \(10^{-14}\), which gives a result that is exact for all practical purposes without resorting to normal approximations.
How the calculator finds the sample size
There is no formula that solves for \(N\) directly, because \(N\) sits inside the noncentral distribution. The calculator runs a search:
- It obtains the critical value \(\chi^2_{1-\alpha,\,df}\) from \(\alpha\) and the table's degrees of freedom.
- It starts with \(N = 2\), computes \(\lambda = N \cdot w^2\) and, with the Poisson mixture, the corresponding power.
- If the power does not reach the target, it tries \(N + 1\) and repeats.
- It stops at the first \(N\) whose power equals or exceeds the target. That is why the result is the minimum integer \(N\) and the "achieved power" is slightly above the requested one.
The following table, computed with this same procedure for \(\alpha = 0.05\) and power \(0.80\), shows the two regularities worth remembering: \(N\) roughly quadruples every time \(w\) is halved, and for the same \(w\) more observations are needed the larger the table (the critical value is larger and the distribution more spread out).
| Degrees of freedom | w = 0.1 | w = 0.2 | w = 0.3 | w = 0.5 | Required λ |
|---|---|---|---|---|---|
| df = 1 (2×2 table) | 785 | 197 | 88 | 32 | ≈ 7.9 |
| df = 2 (2×3 table) | 964 | 241 | 108 | 39 | ≈ 9.6 |
| df = 4 (3×3 or 2×5 table) | 1194 | 299 | 133 | 48 | ≈ 11.9 |
Quick setup
- w: 0.1 for subtle associations, 0.3 for moderate associations (the most common value), 0.5 for strong associations. For 2×2 tables with equal-sized groups, use the wizard.
- df: (number of rows − 1) × (number of columns − 1). For 2×2 tables, df = 1; for 2×3, df = 2; for 3×3, df = 4.
- α: 0.05 is the standard. Lowering it to 0.01 raises the critical value and demands more observations.
- Power: 0.80 as the usual minimum; 0.90 in confirmatory studies.
- Expected frequencies: check that \(N \cdot p_{0,ij} \ge 5\) in every cell, or at least in 80% of them with none below 1.
Worked example, step by step
Example 1: a 2×3 table. A researcher wants to study whether there is an association between gender (2 categories) and preference for three types of product (3 categories). They have no prior data, so they plan for a medium effect, \(w = 0.3\), with \(\alpha = 0.05\) and 80% power.
Step 1: the effect. \(w = 0.3\), hence \(w^2 = 0.09\): each observation contributes on average 0.09 units of "signal" to the statistic.
Step 2: the critical value. \(df = (2-1)(3-1) = 2\). For \(\alpha = 0.05\), the critical value is \(\chi^2_{0.95;\,2} = 5.99\).
Step 3: search for \(N\). For each candidate \(N\), \(\lambda = N \cdot 0.09\), and the power is the area of the noncentral chi-square that lies above \(5.99\):
| N | λ = N·w² | Mean under \(H_1\) (df + λ) | Power |
|---|---|---|---|
| 30 | 2.70 | 4.70 | 29.3% |
| 50 | 4.50 | 6.50 | 46.0% |
| 70 | 6.30 | 8.30 | 60.6% |
| 90 | 8.10 | 10.10 | 72.3% |
| 107 | 9.63 | 11.63 | 79.98% |
| 108 | 9.72 | 11.72 | 80.37% |
| 130 | 11.70 | 13.70 | 87.5% |
With 30 observations, the mean of the statistic under \(H_1\) (4.70) does not even reach the critical value. With 107 the power stays at 79.98%, just below the target, and with 108 it exceeds it. The calculator therefore returns N = 108 observations in total (achieved power: 80.37%). With that sample, if the real association has an effect size \(w \ge 0.3\), the test will detect it at least 80% of the time. For reference, 90% power would require \(N = 141\), and \(\alpha = 0.01\) (critical value \(9.21\)) would require \(N = 155\).
Check: with 6 cells, the average expected frequency is \(108/6 = 18\). The smallest cell depends on the margins: if one of the product categories attracts only 10% of the answers, its smaller cell would have about \(108 \times 0.5 \times 0.10 \approx 5.4\) expected observations, right at the limit of the rule of 5.
Example 2: a 2×2 table with the wizard. Two groups of equal size are compared, with expected success proportions \(p_1 = 0.40\) and \(p_2 = 0.60\). With the wizard's formula, \(\bar{p} = 0.50\) and \(w = |0.40 - 0.60| / (2\sqrt{0.5 \times 0.5}) = 0.20 / 1.00 = 0.20\). You can check that the general formula gives the same: the four cells under \(H_1\) have proportions 0.20, 0.30, 0.30 and 0.20; under independence, all four are 0.25; each term is \((0.05)^2 / 0.25 = 0.01\) and \(w^2 = 4 \times 0.01 = 0.04\), hence \(w = 0.20\). With \(df = 1\), critical value \(3.84\), \(\alpha = 0.05\) and power 0.80, the search stops at N = 197 (\(\lambda = 7.88\); with 196 the power is 79.96% and with 197 it is 80.16%). Split between the two groups, that is 99 subjects per group, rounding up.
Model assumptions
- Random sample of independent observations.
- Two categorical variables crossed in a contingency table; each unit contributes to a single cell.
- Expected frequencies \(\ge 5\) in most cells, so that the chi-square distribution is a good approximation.
- The proportions under \(H_1\) used to compute \(w\) are a working hypothesis: if the true association is weaker, the real power will be lower than planned.
- Effect measured with Cohen's \(w\) and \((\text{rows}-1)(\text{columns}-1)\) degrees of freedom on the noncentral chi-square.
How to interpret the result
The value \(N\) is the total number of observations (sum of all cells in the contingency table) needed for the chi-square test of independence to detect the specified association with the desired power and significance level. Always round up. In a design where the size of each row is fixed (e.g. treatment and control groups), divide \(N\) by the number of rows to get the size of each group; in unrestricted sampling, \(N\) is simply the total number of subjects. Add a margin for dropout by dividing by \((1 - \text{dropout rate})\).
Check the expected frequencies: the most important applicability condition is that the expected frequencies in each cell under \(H_0\), computed as \(N \cdot p_{0,ij}\), are \(\ge 5\). If this is not met, the statistic may not follow the chi-square distribution and the p-value would be unreliable. When several cells have low expected frequencies, merge substantively meaningful categories and recompute \(w\) and the degrees of freedom \((r-1)(c-1)\).
Keep in mind that the sensitivity of \(N\) to \(w\) is quadratic: if the real association is 30% weaker than you assumed, \(N\) should be roughly double \((1/0.7^2 \approx 2.04)\). A very weak association (\(w = 0.1\)) may require close to a thousand observations, so it is worth checking whether the effect you are planning for is realistic.
If the resulting \(N\) is impractical, consider whether the range of variation of one of the variables could be widened, or whether some categories could be merged to concentrate the effect. When the table is \(2 \times 2\) and \(N\) is small, Fisher's exact test is more appropriate than chi-square; in that case, use the sample size calculator for Fisher's exact test. Once the data have been collected, run the analysis with the chi-square test of independence calculator and interpret the standardised residuals of each cell to locate the source of the association.
References
- Cohen, J. (1988). Statistical Power Analysis for the Behavioral Sciences (2nd ed.). Lawrence Erlbaum Associates.
- Agresti, A. (2013). Categorical Data Analysis (3rd ed.). Wiley.