Enter a Z-score to get every p-value at once: left tail, right tail, two tails, from the centre and between −Z and Z. Or choose any one of those probabilities and get the Z-score back. Each area is shaded on a normal curve, and the result is checked against your significance level.
Choose what you know: the Z-score or one of the probabilities.
Enter its value.
Choose your significance level α.
Read the p-values, the Z-score and the significance verdict.
Left tail: P(x < Z) = Φ(Z)
Right tail: P(x > Z) = 1 − Φ(Z)
Two tails: p = 2 × (1 − Φ(|Z|))
Between: P(−Z < x < Z) = 2Φ(|Z|) − 1
A p-value tells you how surprising a result is. It is the probability of getting a result at least as extreme as the one observed, assuming the null hypothesis, the default claim of no effect, is true. The p-value calculator converts a Z-score into its p-values and also works backwards from any p-value to the Z-score.
For one Z-score it reports five areas under the standard normal curve: the left tail, the right tail, both tails together, the area from the centre out to Z, and the area between minus Z and plus Z. Each is shaded on its own small curve, so it is clear which area a one-tailed or two-tailed test needs. The two-tailed value is then compared with the significance level you choose, giving a plain verdict for assignments, lab reports and research summaries.
1. State the null hypothesis and choose a significance level, usually 0.05, before looking at the data.
2. Compute the test statistic: Z = (sample mean − claimed mean) ÷ (standard deviation ÷ √n).
3. Decide whether the test is one-tailed or two-tailed.
4. Find the tail area beyond Z under the standard normal curve. For two tails, double the area beyond |Z|.
5. Compare the p-value with the significance level.
6. If p is at or below the level, reject the null hypothesis; otherwise do not reject it.
The standard normal curve is a bell shape centred on zero with a total area of 1. The area to the left of a Z-score is written Φ(Z) and is the chance that a value falls below it. The right tail is 1 − Φ(Z). Because the curve is symmetric, the area beyond minus Z on the left equals the area beyond plus Z on the right, which is why a two-tailed p-value is simply twice one tail.
If the p-value is below the chosen level, the result is called statistically significant: data this extreme would be rare if the null hypothesis were true. It does not give the probability that the null hypothesis is true, and it says nothing about how large or useful the effect is. With a very large sample, a tiny and unimportant difference can still be significant, so report the size of the effect alongside the p-value.
A Z-test fits when the data are roughly normal or the sample is large, about 30 or more, and the population standard deviation is known or well estimated. For small samples where the standard deviation is estimated from the sample itself, the t-distribution has heavier tails and gives a larger p-value for the same statistic, so a t-test is the right tool. Counts in categories call for a chi-square test instead.
A factory claims its bulbs last 1,000 hours. A sample of 64 bulbs averages 985 hours, and the standard deviation is known to be 50 hours, giving Z = (985 − 1000) ÷ (50 ÷ 8) = −2.4. Is the difference significant at the 5% level?
Left tail: P(x < -2.4) = Φ(Z) = 0.0082
Right tail: P(x > Z) = 1 − Φ(Z) = 0.9918
Two tails: 2 × P(x > |Z|) = 0.0164
Between −Z and Z: 1 − two tails = 0.9836
Compare with α = 0.05: Two-tailed p = 0.0164 ≤ 0.05 Statistically significant: reject the null hypothesis.
Answer: Two-tailed p-value 0.0164; Z-score -2.4; Left tail, P(x < Z) 0.0082
Reading the p-value as the probability that the null hypothesis is true.
Using a one-tailed p-value when the question allows a difference in either direction.
Choosing the significance level after seeing the result.
Forgetting to double the tail area for a two-tailed test.
Using a Z-test for a small sample with an estimated standard deviation, where a t-test is needed.
Hypothesis tests in statistics courses and competitive exams.
Checking whether an A/B test or survey difference is likely to be real.
Quality control, such as testing whether a batch meets its stated average.
Converting between Z-scores, confidence levels and tail probabilities.
Reporting results in lab reports and research papers.
What p-value is significant?
Most fields use p ≤ 0.05. Stricter work uses 0.01 or 0.001. The level should be chosen before looking at the data.
One-tailed or two-tailed?
Use two-tailed when a difference in either direction matters, which is the usual case. Use one-tailed only when you predicted the direction in advance.
What Z-score gives p = 0.05?
For a two-tailed test, Z = ±1.96. For a one-tailed test, Z = 1.645.
Does a small p-value prove my hypothesis?
No. It says the data would be unusual if the null hypothesis were true. It does not measure the size or importance of the effect.