Sample size

Sample size calculator for chi-square goodness-of-fit test

Calculate the sample size needed for a chi-square goodness-of-fit test at a target power, using Cohen's \(w\) and the exact noncentral chi-square distribution.

Use this tool to plan a chi-square goodness-of-fit test: calculate how many observations you need to detect, with the desired power, that your data depart from a theoretical distribution. The effect is expressed with Cohen's \(w\).

Calculator

Enter Cohen's w, the number of categories k, the significance level and the desired power.

Result pending…
Wizard: calculate w from observed and theoretical proportions

Enter the proportions under H1 (expected observed) and under H0 (theoretical), separated by commas. The number of values must equal k.

Explanation

The chi-square goodness-of-fit test answers a simple question: is the frequency of each category in my data compatible with the proportions predicted by a theoretical distribution? That theoretical distribution is the null hypothesis \(H_0\): it assigns to each of the \(k\) categories a proportion \(p_{0,i}\) (for example, \(0.20\) to each of five colours if the hypothesis is that all of them are equally frequent). The alternative hypothesis \(H_1\) simply says that at least one of those proportions is different.

To test it, the procedure compares the observed frequencies \(O_i\) with the frequencies expected under \(H_0\), which are obtained by splitting the \(N\) observations according to the theoretical proportions: \(E_i = N \cdot p_{0,i}\). The discrepancy is summarised in a single number, the chi-square statistic:

\( \chi^2 = \sum_{i=1}^{k} \frac{(O_i - E_i)^2}{E_i}, \qquad E_i = N \cdot p_{0,i} \)

Each term measures how far one category departs from what was expected: the difference is squared so that excesses and shortfalls do not cancel out, and it is divided by \(E_i\) so that a difference of 10 observations weighs more in a category where 20 were expected than in one where 200 were expected. If \(H_0\) is true, the deviations are due to sampling chance alone and the statistic follows a chi-square distribution with \(df = k - 1\) degrees of freedom. It is \(k - 1\) rather than \(k\) because the observed frequencies always add up to \(N\): once \(k - 1\) of them are known, the last one is determined. If, in addition, \(p\) parameters of the theoretical distribution have to be estimated from the data, one degree of freedom is lost for each (\(df = k - 1 - p\)); we come back to this below.

The logic of the power calculation

Before collecting the data we do not know what value \(\chi^2\) will take, but we can reason about the two possible scenarios:

  • If \(H_0\) is true, the statistic follows the central chi-square distribution with \(df\) degrees of freedom. We fix the significance level \(\alpha\) and with it the critical value \(\chi^2_{1-\alpha,\,df}\): the point that the central chi-square exceeds with probability \(\alpha\) only. We will reject \(H_0\) if the statistic lands above it. With \(df = 4\) and \(\alpha = 0.05\), that value is \(9.49\).
  • If \(H_1\) is true (the real proportions are different, \(p_{1,i}\)), the observed frequencies drift systematically away from those expected under \(H_0\) and the statistic tends to be larger. The power is the probability that, in this scenario, the statistic exceeds the critical value and we therefore detect the departure.

The critical value depends only on \(\alpha\) and \(df\); it does not change with the sample size. What does change with \(N\) is how far the statistic shifts to the right when \(H_1\) is true. To compute the power we therefore need to describe the distribution of the statistic under \(H_1\). We do this in three steps: measure the size of the departure (\(w\)), translate it into a shift of the distribution (\(\lambda\)), and compute the area that lies to the right of the critical value.

Step 1: measuring the departure with Cohen's \(w\)

We need a number that says how far the real proportions \(p_{1,i}\) are from the theoretical ones \(p_{0,i}\), and that does not depend on the sample size. Cohen's \(w\) is built with the same recipe as the statistic, but on proportions instead of frequencies:

\( w = \sqrt{\sum_{i=1}^{k} \frac{(p_{1,i} - p_{0,i})^2}{p_{0,i}}} \)

In the calculator's wizard, \(p_{1,i}\) are the "H1 proportions" (\(p_{\text{obs}}\)) and \(p_{0,i}\) the "H0 proportions" (\(p_{\text{teo}}\)). The interpretation is direct: if a sample of \(N\) observations reproduced the proportions \(p_{1,i}\) exactly, its chi-square statistic against \(H_0\) would be exactly \(N \cdot w^2\). That is why it helps to think of \(w^2\) as the chi-square statistic per observation: it is a property of the two distributions being compared, not of the sample. Cohen proposed \(w = 0.1\) (small departure), \(0.3\) (medium) and \(0.5\) (large) as benchmarks, although it is always better to compute it from the proportions you actually expect.

Step 2: from effect to shift, \(\lambda = N \cdot w^2\)

When \(H_1\) is true, the statistic no longer follows the central chi-square. It follows a noncentral chi-square: the same family of distributions, but shifted to the right and more spread out. The size of the shift is governed by the noncentrality parameter \(\lambda\), which for this test is:

\( \lambda = N \cdot w^2 \qquad\text{and therefore}\qquad E\!\left[\chi^2 \mid H_1\right] = df + \lambda \)

A useful way to see it is to decompose the expected value of the statistic under \(H_1\), which is \(df + \lambda\). The term \(df\) is the "noise" part: what the statistic is worth on average even when \(H_0\) is true, purely from sampling chance. The term \(\lambda\) is the "signal" part: what the real departure contributes. And that signal grows in proportion to \(N\), because with twice as many observations the systematic distance between observed and expected doubles, while the critical value stays where it is. This is exactly the mechanism by which a larger sample gives more power: the distribution under \(H_1\) moves away from the critical value and more and more area is left to its right.

This also yields the most important practical rule: for a given power, the required \(\lambda\) is nearly constant, so \(N \approx \lambda / w^2\). The sample size is inversely proportional to the square of the effect: detecting a departure half as large requires four times as many observations.

Step 3: power as the area to the right of the critical value

\( \text{Power} = P\!\left(\chi^2 > \chi^2_{1-\alpha,\,df} \mid H_1\right) = 1 - F_{\chi^2_{df,\,\lambda}}\!\left(\chi^2_{1-\alpha,\,df}\right) \)

Reading from right to left: \(F_{\chi^2_{df,\lambda}}\) is the cumulative distribution function of the noncentral chi-square with \(df\) degrees of freedom and noncentrality \(\lambda\). Evaluated at the critical value, it gives the probability that the statistic does not exceed it even though \(H_1\) is true, i.e. the probability of a type II error, \(\beta\). The power is its complement, \(1 - \beta\).

How the noncentral chi-square is computed

The noncentral chi-square has no simple closed-form formula, but it does have a representation that the calculator exploits: it is a mixture of central chi-squares with Poisson weights. Imagine first drawing an integer \(J\) from a Poisson distribution with mean \(\lambda/2\), and then generating a central chi-square with \(df + 2J\) degrees of freedom. The resulting variable is exactly a noncentral chi-square with parameters \(df\) and \(\lambda\). Its cumulative distribution function is therefore the weighted average of central cumulative distribution functions:

\( F_{\chi^2_{df,\lambda}}(x) = \sum_{j=0}^{\infty} \underbrace{\frac{e^{-\lambda/2}(\lambda/2)^j}{j!}}_{\text{Poisson weight}} \cdot \underbrace{F_{\chi^2_{df+2j}}(x)}_{\text{central chi-square}} \)

Each term combines two familiar ingredients: the Poisson probability that \(J = j\) and the cumulative distribution function of a central chi-square with \(df + 2j\) degrees of freedom. With \(\lambda = 12\) (the value in the example below), the largest weights belong to \(j = 5\) and \(j = 6\) (\(0.16\) each), i.e. to central chi-squares with 14 and 16 degrees of freedom, which exceed the critical value \(9.49\) easily; the weights become negligible from \(j \approx 20\) onwards. The calculator adds terms until the weight drops below \(10^{-14}\), which gives a result that is exact for all practical purposes without resorting to normal approximations.

How the calculator finds the sample size

There is no formula that solves for \(N\) directly, because \(N\) sits inside the noncentral distribution. The calculator runs a search:

  1. It obtains the critical value \(\chi^2_{1-\alpha,\,df}\) from \(\alpha\) and \(df = k - 1\).
  2. It starts with \(N = 2\), computes \(\lambda = N \cdot w^2\) and, with the Poisson mixture, the corresponding power.
  3. If the power does not reach the target, it tries \(N + 1\) and repeats.
  4. It stops at the first \(N\) whose power equals or exceeds the target. That is why the result is the minimum integer \(N\) and the "achieved power" is slightly above the requested one.

The following table, computed with this same procedure for \(\alpha = 0.05\) and power \(0.80\), shows the two regularities worth remembering: \(N\) roughly quadruples every time \(w\) is halved, and for the same \(w\) more observations are needed the more degrees of freedom the test has (the critical value is larger and the distribution more spread out).

Degrees of freedomw = 0.1w = 0.2w = 0.3w = 0.5Required λ
df = 1 (k = 2)7851978832≈ 7.9
df = 2 (k = 3)96424110839≈ 9.6
df = 4 (k = 5)119429913348≈ 11.9

Degrees of freedom when parameters are estimated

The calculator uses \(df = k - 1\), which assumes that the theoretical distribution is fully specified in advance. If you are going to estimate \(p\) parameters from the data themselves (for example, the mean of a Poisson before checking whether the counts are Poisson), the test has \(df = k - 1 - p\) degrees of freedom. To reflect this, enter \(k - p\) in the "Number of categories" field: \(w\) does not change and the degrees of freedom will be correct. Do, however, carry out the minimum-frequency-per-cell check with the real \(k\), dividing \(N\) by the true number of categories.

Quick setup

  • w: 0.1 (small), 0.3 (medium), 0.5 (large) according to Cohen. If you know the proportions you expect, use the wizard: the computed value is always better than a generic benchmark.
  • k: total number of categories. The degrees of freedom are df = k − 1 (no estimated parameters).
  • α: 0.05 is the standard. Lowering it to 0.01 raises the critical value and demands more observations.
  • Power: 0.80 as the usual minimum; 0.90 in confirmatory studies.
  • Expected frequencies: check that \(N \cdot p_{0,i} \ge 5\) in every category (and, at the very least, that \(N/k \ge 5\)) so that the chi-square approximation is valid.

Worked example, step by step

A biologist wants to test whether the colour of certain insects (5 categories: yellow, red, green, blue and black) follows a uniform distribution, i.e. \(p_{0,i} = 0.20\) for each colour. Based on prior data, they believe the real distribution is (0.30, 0.25, 0.20, 0.15, 0.10). They want \(\alpha = 0.05\) and 80% power.

Step 1: compute \(w\). Evaluate each term of the sum and take the square root of the total:

Colour\(p_{1,i}\) (real)\(p_{0,i}\) (under \(H_0\))\(p_{1,i} - p_{0,i}\)\((p_{1,i} - p_{0,i})^2 / p_{0,i}\)
Yellow0.300.20+0.100.0500
Red0.250.20+0.050.0125
Green0.200.2000
Blue0.150.20−0.050.0125
Black0.100.20−0.100.0500
Sum1.001.000\(w^2 = 0.125\)

Therefore \(w = \sqrt{0.125} \approx 0.354\), an effect between medium and large. Notice that the two categories that deviate most (yellow and black) contribute 80% of \(w^2\): this is typical, a few categories concentrate almost all of the signal.

Step 2: fix the critical value. With \(k = 5\) categories, \(df = 4\). For \(\alpha = 0.05\), the critical value is \(\chi^2_{0.95;\,4} = 9.49\): if \(H_0\) were true, only 5% of samples would give a larger statistic.

Step 3: search for the \(N\) that places the distribution under \(H_1\) far enough to the right. For each candidate \(N\), \(\lambda = N \cdot 0.125\), and the power is the area of the noncentral chi-square that lies above \(9.49\):

Nλ = N·w²Mean under \(H_1\) (df + λ)Power
303.757.7530.1%
506.2510.2548.8%
698.6312.6364.2%
8010.0014.0071.6%
9511.8815.8879.78%
9612.0016.0080.25%
12015.0019.0089.1%

With 30 insects, the mean of the statistic under \(H_1\) (7.75) does not even reach the critical value, and the departure would be detected only 30% of the time. With 95 the power stays at 79.78%, a hair below the target, and with 96 it exceeds it. The calculator therefore returns N = 96 insects (achieved power: 80.25%). You can reproduce this by entering the two lists of proportions in the wizard, clicking "Calculate w and update" and then "Calculate sample size".

Check: \(N/k = 96/5 = 19.2 \ge 5\), so the expected frequencies per cell (all equal to 19.2 under \(H_0\)) are sufficient for the chi-square approximation.

Variants: if the biologist wanted 90% power, they would need \(N = 124\); if they kept 80% but lowered \(\alpha\) to \(0.01\) (critical value \(13.28\)), they would need \(N = 134\). In both cases the requirement goes up, and \(\lambda\), and with it \(N\), must grow to compensate.

Model assumptions

  • Random sample of independent observations.
  • Categorical variable with \(k\) mutually exclusive and exhaustive categories.
  • Sufficiently large expected frequencies (usual rule: \(E_i \ge 5\)) so that the chi-square distribution is a good approximation.
  • The proportions under \(H_1\) used to compute \(w\) are a working hypothesis: if the true effect is smaller, the real power will be lower than planned.
  • The effect size is expressed with Cohen's \(w\) and power is calculated using the noncentral chi-square distribution.

How to interpret the result

The value \(N\) is the total number of observations needed (sum across all categories) for the chi-square goodness-of-fit test to have the specified power and significance level. Always round up. If you expect some records to be discarded due to errors or missing values, divide \(N\) by \((1 - \text{dropout rate})\) to get the number you actually need to collect.

Check the expected frequencies: the expected number of observations in each category \(i\) under \(H_0\) is \(N \cdot p_{0,i}\), and all of them should be \(\ge 5\). If any theoretical proportion is very small, that category may fall short even when \(N/k\) is comfortable; in that case, merge categories with theoretical justification until expected frequencies of \(\ge 5\) are reached and recompute \(k\), \(w\) and the degrees of freedom.

Keep in mind that the sensitivity of \(N\) to \(w\) is quadratic: if the real effect is 30% smaller than you assumed, \(N\) should be roughly double \((1/0.7^2 \approx 2.04)\). That is why it pays to compute \(w\) from realistic proportions and, when in doubt, to plan with a somewhat conservative value.

When the calculated \(N\) turns out to be impractical, there are three levers: (1) accept a larger \(w\), i.e. settle for detecting only larger departures; (2) reduce the target power; or (3) reduce the number of categories by merging the smallest ones, which lowers \(df\) and the required \(\lambda\) (recompute \(w\), because it changes too). If any expected frequency is still below 5 even with the calculated \(N\), fall back on an exact multinomial test. Once the data have been collected, run the test with the chi-square goodness-of-fit test calculator and examine the standardised residuals per category to locate where the differences from the theoretical distribution are concentrated.

References

  • Cohen, J. (1988). Statistical Power Analysis for the Behavioral Sciences (2nd ed.). Lawrence Erlbaum Associates.
  • Agresti, A. (2013). Categorical Data Analysis (3rd ed.). Wiley.