Simulation parameters
Distribution the individual observations are drawn from before being averaged. Its parameters are fixed (for example, Exponential with λ=1) and determine the mean μ and variance σ² shown in the "Theoretical" column. The CLT guarantees that the sample mean approaches a normal distribution whatever this starting distribution is, but the speed of convergence depends on its skewness: try a symmetric one (Uniform) and a skewed one (Exponential, Chi-square) with the same n.
Number of independent observations drawn and averaged to obtain each sample mean. With n = 1 there is no averaging and the two histograms coincide; as n grows, the right-hand histogram becomes more symmetric (the CLT effect) and narrower, because the standard deviation of the mean is σ/√n.
How many times the process "draw n observations and compute their mean" is repeated. Each repetition contributes one point to the right-hand histogram. More simulations do not change the shape of the distribution of the mean (that depends on n), but they make the histogram smoother and the figures in the "Empirical" column more stable. In total, n × simulations individual observations are generated.
Original distribution
Histogram of many individual observations drawn from the chosen distribution (bars) together with its theoretical density (red line). It shows the starting shape: symmetric, skewed, bounded or discrete.
Sampling distribution of the mean
Histogram of the sample means obtained in the simulations (bars). The red line is the normal N(μ, σ²/n) predicted by the CLT; the dashed green line is the normal fitted to the simulated means. The closer the bars follow the curves, the better the normal approximation works.
How the samples are generated and how to read the charts
Every time you click "Simulate", the browser runs a Monte Carlo experiment with the chosen parameters. The random numbers are generated on your own device, so each click produces a different simulation. The procedure is as follows:
- A sample is drawn of \(n\) independent observations from the chosen base distribution. For example, with Exponential (λ=1) and \(n = 30\), 30 random waiting times \(x_1, x_2, \ldots, x_{30}\) are generated.
- The mean is computed from those \(n\) observations, \(\bar{x} = (x_1 + \cdots + x_n)/n\). That single number is one sample mean.
- Steps 1 and 2 are repeated as many times as "Number of simulations" indicates. With 1000 simulations you get 1000 sample means, which amounts to \(1000 \times n\) individual observations in total. Those 1000 means are the data behind the right-hand chart and the "Empirical" column.
- For the left-hand chart, a separate, large sample of individual observations (up to 50,000) is also drawn from the same base distribution. These are not the observations used to compute the means, but they come from the same distribution and serve to display its shape clearly.
Left-hand chart: original distribution
The bars are the histogram of the individual observations. Their height is on a density scale, meaning that each bar's height is the proportion of observations falling in that interval divided by the interval width, so the total area of the bars equals 1 and can be compared directly with a density function. The red line is the theoretical density \(f(x)\) of the base distribution: if the bars follow it, the sample is representative of that distribution. For the discrete distributions (Bernoulli and Poisson) no curve is drawn: the bars show the relative frequency of each integer value and the gaps between them are expected. The purpose of this chart is to see what shape you start from: a bell (Normal), a plateau (Uniform), a long right tail (Exponential, Chi-square) or just two values (Bernoulli). The less it looks like a bell, the more striking the right-hand chart becomes.
Right-hand chart: sampling distribution of the mean
The bars are the histogram of the sample means, one per simulation, also on a density scale: each bar's height shows what proportion of the simulations produced a mean within that interval. The number of bins grows with the number of simulations so that the histogram is finer when there is more data. Two normal curves are overlaid on the bars:
- Solid red line, theoretical normal \(\mathcal{N}(\mu, \sigma^2/n)\): this is the distribution the CLT predicts for the sample mean. It is computed from the known mean \(\mu\) and variance \(\sigma^2\) of the base distribution and the chosen \(n\); it does not use the simulated data. It is the "prediction" being tested.
- Dashed green line, empirical normal: the normal whose mean and standard deviation are those of the simulated means. It represents the best bell curve that can be fitted to the data obtained.
To interpret this chart, look at three things:
- Do the bars follow the red curve? If so, the normal approximation is already good for that \(n\). If the bars show a longer tail on one side than the other while the curve is symmetric, convergence has not yet been reached: increase \(n\) and watch the skewness disappear.
- Do the red and green curves coincide? Almost always yes, even for small \(n\). This is because the mean of \(\bar{X}_n\) is exactly \(\mu\) and its standard deviation exactly \(\sigma/\sqrt{n}\) for any \(n\); what the CLT adds is that the shape also becomes normal. A small gap between the two curves is just Monte Carlo noise and shrinks as the number of simulations increases.
- How wide is the histogram? Compare the horizontal axis scale with that of the left-hand chart: the one for the means is much narrower, because averaging reduces the spread by a factor of \(\sqrt{n}\).
With discrete distributions and small \(n\), the histogram of the means can look like a "comb" with gaps. This is not an error: the mean of \(n\) Bernoulli values can only take the \(n + 1\) values \(0, 1/n, 2/n, \ldots, 1\), and the mean of \(n\) Poisson values only multiples of \(1/n\). As \(n\) grows, the possible values multiply and the histogram fills in.
Theoretical vs. empirical table
The Theoretical column is computed from formulas, without using the simulated data: the mean \(\mu\) of the base distribution, the standard error \(\sigma/\sqrt{n}\), the variance \(\sigma^2/n\) and the skewness \(\gamma_1/\sqrt{n}\), where \(\gamma_1\) is the skewness coefficient of the base distribution. The Empirical column is obtained from the simulated means: their average, their sample standard deviation, its square and their skewness coefficient. Mean and standard deviation should match the theoretical values for any \(n\), up to simulation noise. The skewness row is the one that measures convergence to the normal: a normal distribution has skewness 0, so the closer this value is to 0, the more normal the distribution of the mean. Since the theoretical skewness is divided by \(\sqrt{n}\), it becomes clear why highly skewed distributions (Exponential, Chi-square) need a larger \(n\) than symmetric ones (Normal, Uniform), whose skewness is already 0 to begin with.
Suggested experiments
- Exponential with n = 1, 2, 5, 30: the long right tail shrinks until it forms a bell.
- Uniform with n = 2: the plateau becomes a triangle; with n = 5 it is already almost a bell.
- Bernoulli with n = 5 and n = 100: from a histogram of a few separate bars to a continuous bell.
- Chi-square with n = 10 and n = 100: compare the skewness row of the table in both cases.
- 100 vs. 10,000 simulations with the same n: the shape does not change, only the histogram noise.
What does the central limit theorem state?
The central limit theorem (CLT) is one of the fundamental results of mathematical statistics. It states that, under general conditions, the distribution of the sum (or the mean) of a large number of independent and identically distributed (i.i.d.) random variables tends to the normal distribution, regardless of the original distribution of the variables.
Formally, let \(X_1, X_2, \ldots, X_n\) be i.i.d. random variables with mean \(\mu\) and finite variance \(\sigma^2 < \infty\). Then the sample mean \(\bar{X}_n = \frac{1}{n}\sum_{i=1}^{n} X_i\) satisfies:
Equivalently, \(\bar{X}_n \approx \mathcal{N}\!\left(\mu,\, \frac{\sigma^2}{n}\right)\) for \(n\) sufficiently large. The speed of convergence depends on the skewness of the original distribution: the more symmetric and "well-behaved" the distribution is, the faster the sample mean converges to normal. The Berry-Esseen theorem quantifies this speed with a uniform bound on the approximation error:
where \(\rho = \mathbb{E}[|X-\mu|^3]\) is the third absolute centered moment and \(C \approx 0.4748\). This implies that distributions with greater skewness require larger samples for the normal approximation to be accurate.
What sample size n does the CLT need?
There is no single threshold: the speed of convergence depends on the underlying distribution. As a practical guide:
- Symmetric distributions (Normal, Uniform): with \(n = 2\) to \(5\) the distribution of the mean is already practically normal. The Normal is, trivially, exact for any \(n\).
- Moderately skewed distributions (Exponential, Poisson): the usual rule of thumb of \(n \geq 30\) is typically enough. With the exponential (skewness coefficient \(= 2\)) the approximation is already very good at \(n = 30\).
- Bounded discrete distributions (Bernoulli): with probabilities close to 0 or 1 the skewness is high, but because the variable is bounded, convergence is relatively fast. At \(n = 30\) the approximation is excellent.
- Heavy-tailed distributions (Chi-square with few degrees of freedom): the skewness is more pronounced and \(n \geq 50\) may be needed for a satisfactory approximation.
- Distributions without finite variance (Cauchy): the CLT does not apply in its standard form, since the variance does not exist. In these cases the law of large numbers does not converge either.
The practical rule \(n \geq 30\) found in many textbooks is a conservative heuristic that is valid for the distributions commonly encountered in applied sciences. However, for highly skewed distributions or distributions with very heavy tails, it may be insufficient.
Practical implications of the CLT
The CLT has far-reaching consequences in applied statistics:
- Justifies the use of z and t tests: mean tests (z-test, t-test) assume that the sample mean follows a normal distribution. Thanks to the CLT, this holds even when the individual data are not normal, provided \(n\) is sufficiently large.
- Confidence intervals: the construction of confidence intervals for the population mean relies directly on the CLT, making \(\bar{X} \pm z_{\alpha/2}\,\sigma/\sqrt{n}\) a valid approximation regardless of the original distribution.
- Ubiquity of the normal distribution: many natural and social phenomena result from the accumulation of many small independent effects (measurement errors, genetic variation, economic fluctuations…), which explains why the normal distribution appears so frequently in nature.
- Extension to sums: since \(S_n = n\bar{X}_n\), the CLT applies equally to the sum of i.i.d. variables: \((S_n - n\mu)/(\sigma\sqrt{n}) \xrightarrow{d} \mathcal{N}(0,1)\). This is the basis of the normal approximation to the binomial.
- Experimental design and resampling: the CLT underpins techniques such as the bootstrap and power analysis, by guaranteeing that statistics based on means have asymptotically normal distributions.
Frequently asked questions
Does the CLT say that individual data points are normally distributed?
No. The CLT states that the sample mean \(\bar{X}_n\) converges to a normal distribution when \(n \to \infty\). The individual data \(X_i\) can follow any distribution with finite mean and variance: exponential, Bernoulli, Poisson, etc. Confusing the distribution of the data with the distribution of the statistic is one of the most common mistakes when interpreting the CLT.
Why does the Bernoulli distribution with n=30 look so normal?
The Bernoulli distribution with \(p = 0.3\) has finite positive skewness (\(\gamma_1 = (1-2p)/\sqrt{p(1-p)} \approx 0.873\)). Although it is not symmetric, because it is a variable bounded between 0 and 1, all of its moments exist and are finite. This makes convergence in the CLT relatively fast. The averaging effect over \(n = 30\) independent variables is strong enough that the histogram of the means looks practically normal. With a more extreme \(p\) (close to 0 or 1), larger values of \(n\) would be needed.
What happens with very heavy-tailed distributions?
If the variance of the distribution is infinite, the standard CLT does not apply. The best-known example is the Cauchy distribution, whose density function is \(f(x) = 1/[\pi(1+x^2)]\). The Cauchy has no finite mean or variance, and the sample mean of \(n\) independent Cauchy observations has exactly the same Cauchy distribution as a single observation: it does not converge to the normal for any \(n\). For these distributions there is a generalized version of the CLT based on \(\alpha\)-stable distributions (Lévy), but it is beyond the scope of everyday statistical use.
Do the variables need to be identically distributed?
The classical version of the CLT requires i.i.d. (independent and identically distributed) variables. However, more general extensions exist: the Lindeberg-Lévy theorem relaxes the requirement of identical distribution and only requires that no single variable dominate the sum (Lindeberg's condition). This extension is key in econometrics and time series analysis.
How does sample size affect the variance of the sample mean?
The variance of the sample mean is exactly \(\text{Var}(\bar{X}_n) = \sigma^2/n\), so the standard deviation (\(= \sigma/\sqrt{n}\), also called the standard error) shrinks with the square root of the sample size. Doubling the sample size reduces the standard error by a factor of \(\sqrt{2} \approx 1.41\). To halve the standard error, you need to quadruple the sample size. This relationship is visible in the simulator's histograms: as \(n\) grows, the bell curve of the sample mean becomes narrower.