T-Test Calculator

A t-test checks whether a sample mean differs significantly from a value, or whether two groups differ. Enter summary statistics to get the t statistic, degrees of freedom, p-value and a decision at your chosen significance level.

How it is calculated

Choose one-sample or two-sample.

Enter the means, SDs and sizes.

Read the p-value and decision.

Formula

One sample: t = (x̄ − μ₀) ÷ (s ÷ √n)

Welch: t = (x̄₁ − x̄₂) ÷ √(s₁²/n₁ + s₂²/n₂)

What is the T-Test Calculator?

A t-test asks whether a difference in averages is real or could easily be due to chance. The one-sample version compares a sample mean with a fixed value, such as a packet weight stated on the label. The two-sample version compares the means of two independent groups, such as two teaching methods or two fertilisers. This calculator works from summary statistics, the means, standard deviations and sample sizes, and returns the t statistic, the degrees of freedom, the two-tailed p-value and a decision at 10%, 5% or 1% significance.

T-tests are among the most used tests in dissertations, clinical and agricultural trials, psychology experiments and business A/B tests. They suit the common situation where the population standard deviation is unknown and has to be estimated from the data.

How to calculate it by hand

1. State the hypotheses. H₀: the means are equal (or the mean equals μ₀). H₁: they differ.

2. One sample: compute the standard error SE = s ÷ √n and the statistic t = (x̄ − μ₀) ÷ SE, with n − 1 degrees of freedom.

3. Two samples (Welch): SE = √(s₁²/n₁ + s₂²/n₂) and t = (x̄₁ − x̄₂) ÷ SE.

4. Welch degrees of freedom: df = (s₁²/n₁ + s₂²/n₂)² ÷ [(s₁²/n₁)²/(n₁ − 1) + (s₂²/n₂)²/(n₂ − 1)].

5. Find the two-tailed p-value from the t distribution with those degrees of freedom, or compare |t| with the critical value from a t-table.

6. If p is below the significance level α, reject H₀; otherwise, do not reject it.

Signal divided by noise

The t statistic is the observed difference divided by its standard error, the amount the difference would typically wobble from sample to sample. A t of 2 says the difference is about twice its typical random wobble. The p-value converts that into a probability: if H₀ were true, how often would chance alone produce a t at least this far from zero in either direction? A small p-value means the data would be surprising under H₀. It does not measure the size or importance of the effect.

Why t and not z

If the population SD were known, the statistic would follow the standard normal curve. In practice s is estimated from the same data, and that estimate is itself uncertain, especially for small samples. William Gosset, publishing as 'Student', worked out that the resulting statistic follows the t distribution, which has heavier tails than the normal and depends on the degrees of freedom. With 5 degrees of freedom the two-tailed 5% critical value is about 2.57; with 30 it is about 2.04; with very large samples it approaches 1.96.

Welch's version and the assumptions

The classic two-sample test pools the variances and assumes both groups have the same spread. Welch's test, used here, does not; it uses each group's own variance and adjusts the degrees of freedom with the Welch–Satterthwaite formula, which can be a fractional number. It performs well whether or not the variances are equal, so many statisticians use it by default. Both versions assume independent observations and roughly normal data or reasonably large samples. For paired data, such as before and after on the same people, run a one-sample test on the differences against zero.

Worked example, step by step

A coaching institute in Kota compares two batches on the same mock test. Batch A has 35 students with mean 72.4 and SD 8.2; batch B has 40 students with mean 68.1 and SD 9.6. It tests for a difference at the 5% level.

Standard error: SE = √(s₁²/n₁ + s₂²/n₂) = √(1.921143 + 2.304) = 2.055515

t statistic: t = (72.4 − 68.1) ÷ 2.055515 = 2.0919

Welch degrees of freedom: df = 72.96

Two-tailed p-value: p = 0.039925 0.039925 < 0.05 → reject H₀ (significant)

Answer: p-value (two-tailed) 0.039925; t statistic 2.0919; Degrees of freedom 72.96

Common mistakes to avoid

Reading 'do not reject H₀' as proof that the means are equal; it only means the evidence is not strong enough.

Using a two-sample test on paired data, which throws away the pairing and loses power.

Halving the p-value for a one-tailed test after seeing which direction the data went.

Confusing statistical significance with practical importance; a large sample can make a trivial difference significant.

Entering the standard error in place of the standard deviation.

Where it is used

Comparing average scores of two teaching methods or batches.

Checking whether a machine's mean fill weight matches the labelled weight.

Agricultural trials comparing yields of two varieties or fertilisers.

Clinical and psychology studies comparing treatment and control groups.

Website A/B tests comparing average order values.

Frequently asked questions

What about a paired t-test?

Take the difference for each pair and run a one-sample test of the differences against 0.

Is the p-value one- or two-tailed?

Two-tailed. Halve it for a one-tailed test in the predicted direction.