Grubbs test: is that reading really an outlier?

The G statistic against its critical value, the p-value, and the point implicated — plus a straight answer about what to do next, which is almost never to delete the reading.

Nothing you paste leaves this calculation. It is not stored, not logged and not sent anywhere else — the numbers are computed and returned, and that is all.

One reading per line. At least three, and the test assumes they are otherwise normal.

What Grubbs' test actually tests

Whether the most extreme reading in your sample is further from the mean than a normal distribution would plausibly produce. The statistic is simply the largest deviation expressed in standard deviations:

G = max|xᵢ − x̄| / s

and it is compared against a critical value that depends on the sample size and the significance level. Note what that means: the test is a statement about your distributional assumption at least as much as about the datum. A reading is "an outlier" only relative to a model of what ordinary looks like.

Do not delete the point

This is the whole reason to be careful with this test, and it is worth being blunt about.

A significant Grubbs result is not permission to remove a reading. In a manufacturing process, an extreme value is very often the most informative thing that happened all week: the moment the tool chipped, the batch of material that was different, the setup somebody did unusually. Deleting it removes the evidence and leaves the cause in the process, where it will happen again.

The defensible sequence is: the test flags a point, you go and find out what happened to that part, and then you decide. A reading you can trace to a transcription error or a probe that had come loose is a measurement failure and can go, with the reason recorded. A reading you cannot explain stays, because "it did not fit my model" is not a reason.

Three ways the test misleads

It assumes normality, and punishes skew

On a naturally skewed characteristic — anything bounded at zero, like a flatness, a roundness or an impurity level — the long tail is real and the test will flag perfectly ordinary readings from it. Check the shape before trusting the verdict. A histogram with a normality test is the right thing to look at first.

It tests one outlier at a time, and they can hide each other

This is called masking. Two extreme values at the same end both inflate the standard deviation, which is the denominator of G, so each makes the other look less extreme — and neither reaches significance. The sample plainly has two odd readings in it and the test reports none. Look at the plot, not only at the p-value.

It says nothing about time

Grubbs treats your readings as an unordered sample. If they came off a process in sequence, the far more useful question is not "is this value extreme?" but "when did the process change?" — and a control chart answers that, while also catching the shifts and trends that never produce a single extreme value at all.

What to use instead, often

If the readings are in time order, chart them. A control chart flags the same extreme point as rule 1, and it additionally catches a run of nine on one side, a drift over six points, and the rest — patterns that are entirely invisible to an outlier test and which usually matter more than a single spike.

Questions people ask about this

What is Grubbs test?

A test for whether the most extreme reading in a sample is further from the mean than a normal distribution would plausibly produce. The statistic is the largest deviation expressed in standard deviations, compared against a critical value that depends on the sample size.

Should I delete an outlier that Grubbs flags?

Almost never on statistical grounds alone. In a manufacturing process an extreme value is often the most informative thing that happened all week — the tool chipping, the odd batch of material, the unusual setup. Investigate what happened to that part first. A reading traceable to a measurement failure can go, with the reason recorded; one you cannot explain stays, because "it did not fit my model" is not a reason.

What if I have two outliers?

Grubbs tests one at a time, and two extreme values at the same end can hide each other. Both inflate the standard deviation, which is the denominator of the statistic, so each makes the other look less extreme and neither reaches significance. This is called masking, and it is why you look at the plot rather than only at the p-value.

Does Grubbs test assume normality?

Yes, and that matters. On a naturally skewed characteristic — anything bounded at zero, like flatness or an impurity level — the long tail is real, and the test will flag perfectly ordinary readings from it. Check the shape with a histogram before trusting the verdict.

Is there something better than an outlier test?

If the readings are in time order, yes: chart them. A control chart flags the same extreme point and additionally catches runs, shifts and trends — patterns invisible to an outlier test that usually matter more than a single spike.

Want this to keep itself up to date?

A calculator answers for the data you pasted. A control chart answers for the data your line produced this morning — limits frozen at a baseline you locked, rules evaluated on every new reading, an alert when one trips.

More analysis tools

  • Scatter diagram and correlation — Paired readings plotted, with Pearson r, Spearman rho, r², the least-squares line and its p-value — and the two reported together, because when they disagree the disagreement is the finding.
  • Standard deviation calculator — The usual summary statistics, and the thing a spreadsheet will not tell you: the within-subgroup sigma a control chart uses, beside the overall sigma a textbook means. The gap between them is why Cpk and Ppk disagree.

Everything else

  • Cp / Cpk calculator — Paste measurements — or type a mean and a sigma — with your tolerance, and get Cp, Cpk, Pp, Ppk, the sigma level and the expected parts per million out of spec.
  • Control limit calculator — Give it subgroups or individual readings and it returns the control limits for the chart, the range or sigma chart beneath it, and every constant it used to get there.
  • Control chart generator — Paste a column of numbers, or rows of subgroups, and get a real control chart: limits from the data, Nelson rules 1–4 evaluated, out-of-control points marked.
  • Nelson & Western Electric rules checker — Every run rule evaluated on your data, each violation named in plain English with what it usually indicates — a shift, a trend, tool wear, two machines mixed.
  • Cpk confidence interval & sample size — A Cpk of 1.33 from 30 pieces has a 95% interval of roughly 0.97 to 1.69. See the uncertainty in your own number, and how many parts would settle it.
  • Cpk ↔ PPM ↔ sigma level converter — What PPM is a Cpk of 1.33? What Cpk does 3.4 PPM need? Both conventions shown side by side, because the 1.5 sigma shift is why two sources disagree by a factor of ten.
  • Control chart constants — The whole table, n = 2 to 25, with the formula each constant belongs to. The same values the charts on this site are computed from.
  • X̄-R chart calculator — Subgroups in, X̄ and R charts out — limits from A₂, D₃ and D₄, run rules evaluated, and the range chart shown first because it decides whether the averages chart can be trusted.
  • X̄-s chart calculator — For subgroups big enough that the range wastes them. Limits from A₃, B₃ and B₄, sigma recovered with the c₄ correction, and the s chart read first because it decides whether the averages chart can be trusted.
  • I-MR chart calculator — For processes that give you one number at a time — a batch, an oven, a destructive test. Limits from the mean moving range, run rules evaluated, and an honest note about what an individuals chart cannot see.
  • Levey-Jennings chart with Westgard rules — Paste your QC results with the assigned mean and SD from the package insert, and get the chart with every sigma band drawn and all six Westgard rules evaluated — each one named, and labelled random or systematic error.
  • Bland-Altman plot — Paired readings from two methods, plotted as difference against average — with the bias, the limits of agreement, their confidence intervals, and a test for whether the disagreement depends on the magnitude.
  • Pareto chart generator — Categories and counts in, ranked bars and the cumulative line out — with the vital few named, an "Other" bucket that always sorts last, and an honest warning when the distribution is flat and there is no dominant cause to attack.
  • Box and whisker plot generator — One column per group, boxes side by side. Quartiles, the 1.5×IQR fences, whiskers that stop at real readings and outliers drawn individually — with the quartile method stated, because that is why your plot and Excel's disagree.
  • Histogram and normality test — A histogram with the fitted normal curve, both standard bin rules with a reasoned recommendation, and an Anderson-Darling test that refuses to tell you your data is normal — because no test can.
  • Gage R&R calculator (ANOVA) — Paste a crossed study — part, operator, reading — and get the full ANOVA: repeatability and reproducibility separated, the part-by-operator interaction tested rather than assumed away, %GRR, %Tolerance and ndc against AIAG's bands.
  • p, np, u and c chart calculator — Counted data rather than measured: proportion defective, number defective, defects per unit. Limits that step with the sample size instead of pretending it never changed, and an honest warning when the counts are too low for three-sigma limits to mean anything.