t-test, ANOVA or Mann–Whitney: which test should you use?

About 8 min read · Updated 2026-09-30

A statistical test is chosen almost entirely by four questions: how many groups, are the data paired, are the values roughly normal, and do you compare with a control only or all pairs? This guide walks through them in that order.

Try it with the tool

Paste values and answer two questions — the right test is chosen and explained, with a graph showing asterisks.

Open the group comparison tool →

1. The whole choice in one table

SituationRoughly normalNot normal
2 groups, different samplesWelch t-test (recommended)Mann–Whitney U
2 groups, same subjects before/afterpaired t-testWilcoxon signed-rank
3+ groups, vs control onlyANOVA + DunnettKruskal–Wallis + Dunn
3+ groups, all pairsANOVA + TukeyKruskal–Wallis + Dunn
3+ groups, very unequal variancesWelch ANOVA + Games–HowellKruskal–Wallis + Dunn

Running several t-tests across three or more groups multiplies chance “significant” results by the number of comparisons. That is why ANOVA with a post-hoc test (Tukey, Dunnett) is used.

Choosing a statistical testHow many groups?23 or moreSame subjects before/after?Compare with control only?yesnoyesnopaired t-testnon-normal: WilcoxonWelch t-testnon-normal: Mann–WhitneyANOVA + Dunnettnon-normal: Kruskal + DunnANOVA + Tukeynon-normal: Kruskal + DunnVery unequal variances → Welch ANOVA + Games–Howell · never run many t-tests across 3+ groups
Two questions (number of groups, what to compare) decide the test. If values are far from normal, use the nonparametric test in the same position.

2. Why Welch’s t-test is the default

The classic Student t-test assumes both groups spread equally (equal variances). Treated groups often spread more than controls, so this assumption breaks often. Welch’s t-test stays accurate when variances differ and gives nearly the same answer when they are equal, so it is the recommended default.

3. Cautions at n = 3

4. Tukey or Dunnett?

Dunnett compares each treatment with a single control. Fewer comparisons mean more power on the same data — use it when the question is “effect versus control”, such as dose series. Tukey compares every pair; use it when differences between treatments (is drug A stronger than drug B?) matter too.

5. Common mistakes

Paste values and answer two questions — the right test is chosen and explained, with a graph showing asterisks.

Open the group comparison tool →