Skip to main content

Mann-Whitney U Test Calculator

The Mann-Whitney U test asks whether two independent samples were drawn from the same distribution, and it answers by ranking: pool every observation from both groups, sort them, and see how the two groups share out the ranks. There is no assumption of normality and no mean to be dragged around by outliers, so it works for ordinal scores, skewed measurements and awkward sample sizes alike.

One observation per entry, or any mix of spaces, commas and new lines.

The second independent group. Both samples may be different sizes.

Set the direction before looking at the result; choosing it afterwards invalidates the p-value.

The exact distribution is used when both samples have 20 or fewer observations and there are no ties; otherwise the normal approximation is used.

Shifts the normal approximation by half a count toward the null mean, which makes it slightly more conservative.

Calculated Result
0

Mann-Whitney U (smaller of the two)

Mann-Whitney U (smaller)

0

U for sample 1

64

U for sample 2

0

Rank sum W for sample 1

100

Pairs compared

64

Pairs favouring sample 1

64

Pairs favouring sample 2

0

Pairs exactly tied

0

Mean U under H₀

32

Standard deviation of U under H₀

9.521905

Critical U at this level

13.33741

z (continuity corrected)

3.360672

Probability of superiority (A)

100%

Rank-biserial correlation (r)

1

Tied groups

0

p-value method

exact null distribution

Samples differ significantly — U = 0, p < 0.001 (exact), r = 1.

Calculation Breakdown

  1. Rank every observation togetherAll 16 values from both samples are pooled and ranked, ties sharing the average of the ranks they span. Sample 1's ranks sum to W = 100.W1=∑i=1n1RiW_1 = \sum_{i=1}^{n_1} R_i
  2. Turn the rank sum into UU₁ = W₁ − n₁(n₁ + 1)/2 = 64, and U₂ = n₁n₂ − U₁ = 0. Counting the pairs directly gives the same figure: 64 favour sample 1, 0 favour sample 2 and 0 tie, so U₁ = 64 + ½ × 0 = 64.U1=W1−n1(n1+1)2,U1+U2=n1n2U_1 = W_1 - \frac{n_1(n_1+1)}{2}, \qquad U_1 + U_2 = n_1 n_2
  3. Compare U with its null distributionUnder H₀ the two samples are indistinguishable, U is centred on n₁n₂/2 = 32 with standard deviation 9.521905, giving z = 3.360672 and a p-value of 0.000155.E[U]=n1n22,σU=n1n2(n+1−∑(t3−t)n(n−1))12E[U] = \frac{n_1 n_2}{2}, \quad \sigma_U = \sqrt{\frac{n_1 n_2\left(n+1-\frac{\sum (t^3-t)}{n(n-1)}\right)}{12}}

Pairwise comparisons behind U

Interactive visualization based on your current inputs

Pairs
0.016324864Favours sample 1Favours sample 2TiedComparisonNumber of pairs

What Is the Mann-Whitney U Test Calculator?

The Mann-Whitney U test is a non-parametric test for comparing two independent samples without assuming normality. It asks whether one sample tends to sit above the other once both are pooled and ranked, which is a statement about the whole distribution rather than about means.

It is the test of choice for ordinal scales such as exam grades, satisfaction scores or pain ratings, for heavily skewed measurements such as incomes or waiting times, and for small samples where the normality assumption behind a t-test cannot be checked.

How Does the Mann-Whitney U Test Calculator Work?

Every observation from both samples is ranked together, ties sharing the average of the ranks they span. Nothing else about the values is used, which is exactly why outliers cannot distort the result.

The ranks of sample 1 are summed to W₁ and converted to U₁ = W₁ − n₁(n₁ + 1)/2. Equivalently, U₁ counts the cross-group pairs where sample 1 ranks higher, plus half of the ties — the calculator shows both routes and checks they agree.

Under the null hypothesis the two groups are indistinguishable, so U is centred on n₁n₂/2 with a standard deviation of roughly √(n₁n₂(N + 1)/12), shrunk slightly when ties are present.

The p-value comes from counting every rank arrangement consistent with the observed U where that is possible, and from the normal curve with tie and continuity corrections otherwise.

Mann-Whitney U Test Calculator Formula & Variables

The core mathematical equation utilized by this calculator is expressed as:

U1=W1−n1(n1+1)2,U2=n1n2−U1,E[U]=n1n22U_1 = W_1 - \frac{n_1(n_1+1)}{2}, \quad U_2 = n_1 n_2 - U_1, \quad E[U] = \frac{n_1 n_2}{2}

Variable Definitions

SymbolVariable Meaning & Units
W₁sum of the ranks of the observations in sample 1
n₁number of observations in sample 1
n₂number of observations in sample 2
Nn₁ + n₂, the total number of observations
Uthe smaller of U₁ and U₂, the usual quoted statistic
tthe size of a group of tied values, summed over tie groups

Pooling and ranking both samples gives the rank sum W₁, from which U₁ follows. Because U₁ counts the pairs where sample 1 ranks higher plus half of the ties, U₁ + U₂ = n₁n₂ always holds. If the two samples are drawn from the same distribution U sits near n₁n₂/2 with the spread shown above, so the p-value measures how far the observed U has strayed from that centre.

How to Use the Mann-Whitney U Test Calculator

  1. Paste one group into each box. Commas, spaces and new lines all work, the groups may be different sizes, and the order within a group is irrelevant.
  2. Choose the alternative hypothesis before reading the result. Two-sided is the honest default unless a direction was specified in advance.
  3. Leave the exact p-value on. The calculator switches to the normal approximation by itself, and the Method row tells you which one produced the p-value you are reading.
  4. Compare U against the critical U for your level, then read the probability of superiority and the rank-biserial correlation to judge whether a significant result is also a substantial one.

Step-by-Step Example Calculation

Scores from two independent groups of eight, with no ties

Input Values:

sample1:45, 48, 51, 55, 58, 62, 66, 70
sample2:12, 19, 24, 27, 31, 34, 38, 41
alternative:two_sided
alpha:0.05
exact:yes
continuity:yes
Worked Steps: Every value of sample 1 sits above every value in sample 2, so U = 0 out of 64 pairs: complete separation. The exact two-sided p-value is 2/12870 ≈ 0.00016, well under 5%, the probability of superiority is 100% and the rank-biserial correlation is 1.

Understanding Your Result

U has no units; it is a count of pairwise outcomes out of n₁n₂ possible. Half the pairs going each way is what "no difference" looks like, and it gives U = n₁n₂/2.

The mean and standard deviation rows describe where U would sit if the null hypothesis were true. They are what make a raw U interpretable across different sample sizes.

The pairwise table is the intuition behind the statistic: it shows how many comparisons favoured each group and how many tied, so a significant U can be read as a real shift rather than an abstraction.

A non-significant p-value is not proof of equality. Small samples give the test little power, so failing to reject the null is common even when the medians differ noticeably.

Factors That Affect the Result

  • Sample size. The spread of U shrinks roughly as √(n₁n₂), so doubling both samples makes the test meaningfully sharper.
  • Ties. Repeated values inflate the variance and rule out the exact distribution, which is why the tie correction and the Method row both exist.
  • The separation of the groups. Complete separation gives U = 0, the extreme value, but with eight per group the exact p-value is still only about 0.03.
  • The shape of the two distributions. The test detects any distributional difference; it reads as a shift in medians only when the shapes match.
  • Outliers. Ranks limit the damage relative to a mean-based test, though one wild value still pushes a whole block of ranks upward.

When Should You Use This Calculator?

  • Comparing two independent groups when normality is doubtful or the measurement is only ordinal.
  • Small clinical or field samples where the normal approximation behind a t-test would be unreliable.
  • Skewed measures such as income, response times, waiting times or bacterial counts, where a mean is a poor summary anyway.
  • Teaching, as the clearest non-parametric test of two independent samples.

Assumptions & Limitations

  • Observations must be independent within and between groups. Clustered or repeated measurements make the test anti-conservative.
  • Under the null the distributions must have the same shape; otherwise a significant result does not locate the difference.
  • The measurement must be at least ordinal — a nominal scale with no meaningful order cannot be ranked.
  • Exact p-values need no ties and small samples; beyond that the normal approximation is used and is less reliable in the tails.
  • The test does not adjust for multiple comparisons, so running it on many endpoints will produce a significant result by chance.

Frequently Asked Questions

Calculation Accuracy & Reference Note

Rank sums and U are exact arithmetic on the inputs, and exact p-values come from counting rank arrangements, so neither carries rounding error. Only the normal approximation introduces an approximation, and the Method row identifies it.

Standard Reference: Mann, H. B. & Whitney, D. R. (1947), On a problem of rank significance. Annals of Mathematical Statistics 18(1), 50–60.