The chi-square test checks whether observed counts match what you expected, such as dice rolls or survey responses. Enter both sets of counts to get the chi-square statistic and p-value.
Enter observed counts.
Enter expected counts in the same order.
Read χ² and the p-value.
Statistic: χ² = Σ (O − E)² ÷ E
The chi-square goodness-of-fit test checks whether observed counts in categories match the counts you expected under some theory. Is a die fair? Do customers choose four product colours equally? Do offspring of a genetic cross follow Mendel's 9:3:3:1 ratio? This calculator takes the observed and expected counts, shows each category's contribution (O − E)² ÷ E, and returns the chi-square statistic, the degrees of freedom and the p-value.
It is one of the first tests taught in college statistics and biology, and it works on counts rather than measurements, so it suits survey answers, category data and genetics experiments. Seeing the contribution of each category also shows where any mismatch comes from.
1. Write the observed count O for each category.
2. Work out the expected count E for each category from your hypothesis: expected proportion × total observed.
3. For each category compute (O − E)² ÷ E.
4. Add them up: χ² = Σ (O − E)² ÷ E.
5. Degrees of freedom: number of categories minus 1, and minus one more for each parameter estimated from the data.
6. Find the p-value, the upper-tail area of the chi-square distribution beyond your χ², or compare χ² with the table's critical value.
7. If p is below your significance level, reject the hypothesis that the data follow the expected pattern.
A difference of 5 means a lot if you expected 10 and very little if you expected 1000. Counts that arise by chance have a natural spread that grows with the expected count; for random counts the variance is roughly equal to E. So (O − E) ÷ √E is a standardised deviation, like a z-score, and squaring and adding these across categories gives χ². When the hypothesis is true and expected counts are not too small, this sum follows the chi-square distribution, which is the distribution of a sum of squared standard normal values.
The observed counts must add up to the sample size. Once you know all but one of them, the last is fixed, so k categories give k − 1 degrees of freedom. If you estimate anything from the data to build the expected counts, such as a mean for a Poisson fit, subtract one more for each estimate. The calculator uses k − 1, which is right when the expected proportions come from a theory fixed in advance. The chi-square test is one-tailed: only large values signal a poor fit.
The chi-square distribution is an approximation that becomes unreliable when expected counts are small. A common rule of thumb asks for every expected count to be at least 5; merge sparse categories if needed. The same statistic is also used for the test of independence in a contingency table, where expected counts are row total × column total ÷ grand total and degrees of freedom are (rows − 1) × (columns − 1). That version needs a table layout, which this goodness-of-fit calculator does not handle.
For a probability project, Aman rolled a die 60 times and got the faces 1 to 6 in counts of 8, 12, 9, 14, 6 and 11. A fair die would give 10 of each, and he wants to test whether his die looks fair.
Contribution (O − E)² ÷ E: (8 − 10)² ÷ 10 = 0.4 (12 − 10)² ÷ 10 = 0.4 (9 − 10)² ÷ 10 = 0.1 (14 − 10)² ÷ 10 = 1.6 (6 − 10)² ÷ 10 = 1.6 (11 − 10)² ÷ 10 = 0.1
Chi-square statistic: χ² = 0.4 + 0.4 + 0.1 + 1.6 + 1.6 + 0.1 = 4.2
Degrees of freedom: k − 1 = 5
p-value: P(χ²₍5₎ ≥ 4.2) = 0.520995
Answer: p-value 0.520995; χ² statistic 4.2; Degrees of freedom 5
Entering expected proportions or percentages instead of expected counts; they must add up to the same total as the observed counts.
Running the test on percentages or averages rather than raw counts.
Ignoring the rule that expected counts should usually be at least 5.
Forgetting to reduce the degrees of freedom when parameters are estimated from the data.
Concluding that a large p-value proves the hypothesis; it only means no significant misfit was found.
Testing whether a die, coin or random number generator is fair.
Checking Mendelian ratios in genetics practicals.
Comparing survey responses with known population shares.
Testing whether customer arrivals or complaints are evenly spread across weekdays.
Checking whether data fit an assumed distribution, such as Poisson or normal, in grouped form.
What if an expected count is below 5?
The approximation becomes unreliable; merge small categories.
Can I test independence in a table?
Yes, with expected = row total × column total ÷ grand total and df = (rows − 1)(cols − 1); this calculator handles goodness of fit.