Gage R&R: is it the part, or the measurement?

Paste a crossed study — part, operator, reading — and get the full ANOVA: repeatability and reproducibility separated, the part-by-operator interaction tested rather than assumed away, %GRR, %Tolerance and ndc against AIAG's bands.

Nothing you paste leaves this calculation. It is not stored, not logged and not sent anywhere else — the numbers are computed and returned, and that is all.

One reading per line: part, operator, measurement. Separated by commas or tabs, in any order, with an optional header row. Every operator must measure every part the same number of times.

Needed for %Tolerance. Leave both blank for %Study Variation only.

What a Gage R&R actually asks

Whether the numbers coming off your measuring equipment are mostly the part, or mostly the measurement. Every reading is the true dimension plus measurement error, and this study splits the variation you observed into those pieces:

  • Repeatability (EV) — the same operator, the same part, twice. This is the equipment: its resolution, its stability, how firmly it grips.
  • Reproducibility (AV) — different operators, the same part. This is method and training: where they take the reading, how hard they seat the part, how they round.
  • Part × operator interaction — operators who disagree about some parts and not others. One reads soft material differently; a fixture only misaligns on the long ones. A real and specific failure mode, which is why it is tested before being dropped.
  • Part-to-part (PV) — the actual differences between your parts. This is the signal; everything above it is noise.

Why ANOVA and not the range method

The older average-and-range method is arithmetic you can do on paper, and it was the right answer when that mattered. It cannot estimate the part-by-operator interaction at all — it folds that variation into reproducibility and never tells you it was there.

ANOVA separates it, and gives every component a significance test. The two methods usually land close on well-behaved data and diverge exactly where the interaction is real, which is the case you most needed to know about.

%Study Variation and %Contribution are not the same number

This confusion is worth more scrapped gauges than any other in MSA.

  • %Study Variation is a ratio of standard deviations: σ_GRR / σ_total. AIAG's 10% and 30% bands are written against this one. These percentages do not sum to 100, because standard deviations add in quadrature.
  • %Contribution is a ratio of variances: σ²_GRR / σ²_total. These do sum to 100.

Because squaring a fraction makes it smaller, %Contribution is always the flattering number. A gauge at 30% study variation is 9% contribution. Quote the second against the first's acceptance bands and an unacceptable gauge passes. Both are shown above so the difference is impossible to miss.

The three things a %GRR does not tell you

1. It depends on the parts you chose

%Study Variation divides the gauge's variation by the total, and the total is dominated by how different your parts are from each other. Ten parts pulled from one good hour understate part-to-part variation, and the same gauge scores far worse than it deserves. Parts deliberately spread wider than the process ever runs flatter it.

AIAG asks for parts spanning the expected process variation for exactly this reason. A study that ignored it has measured your part selection as much as your gauge.

2. %Tolerance answers a different question

%Tolerance asks whether the gauge can accept and reject parts against the drawing. %Study Variation asks whether it can see the process move. A generous tolerance forgives a gauge that could never detect a shift; a tight one condemns a gauge that tracks the process perfectly well. Which matters depends on what the measurement is for — and a gauge used for both has to pass both.

3. A small study is a wide estimate

Ten parts, three operators, three trials is ninety readings, and the %GRR computed from them carries a confidence interval spanning several percentage points either side. A result of 9.4% is not meaningfully on the good side of 10%. Treating the point estimate as the answer claims more than the study showed.

ndc, and what to do with a failing gauge

The number of distinct categories is 1.41 × PV/GRR, floored — roughly how many separate groups the gauge can actually tell apart across your parts. AIAG asks for five or more. An ndc of 1 means the measurement system cannot distinguish your parts from each other at all, and any control chart built on it is charting the gauge.

When a study fails, the split above tells you where to go. Repeatability dominating is equipment — resolution, wear, fixturing, clamping force. Reproducibility dominating is people and method, which is usually cheaper to fix: a written procedure, a datum that is not ambiguous, an hour of training. A significant interaction is more specific still, and worth chasing: find which parts the operators disagree about and the physical cause is usually obvious once you are looking at the right ten parts.

Questions people ask about this

What is a Gage R&R study?

A designed experiment that separates the variation in your measurements into the equipment (repeatability), the operators (reproducibility) and the parts themselves. It answers whether the numbers coming off your gauge are mostly the part or mostly the measurement.

What is an acceptable %GRR?

AIAG's bands are under 10% acceptable, 10 to 30% conditionally acceptable depending on what the measurement is for and what fixing it would cost, and over 30% unacceptable. Those bands are written against %Study Variation, which is a ratio of standard deviations — not against %Contribution.

What is the difference between %Study Variation and %Contribution?

%Study Variation divides standard deviations and is what AIAG's bands refer to. %Contribution divides variances and sums to 100. Because squaring a fraction makes it smaller, %Contribution always looks better — a gauge at 30% study variation is 9% contribution. Quoting the second against the first's bands lets an unacceptable gauge pass, which is the commonest error in MSA reporting.

Why ANOVA rather than the average and range method?

Because the range method cannot estimate the part-by-operator interaction at all — it folds that variation into reproducibility and never mentions it. ANOVA separates it and tests it. The two agree closely on well-behaved data and diverge exactly where the interaction is real, which is the case you needed to know about.

What is ndc in a Gage R&R?

The number of distinct categories, 1.41 × PV/GRR rounded down: roughly how many separate groups the measurement system can tell apart across your parts. AIAG asks for five or more. An ndc of one means the gauge cannot distinguish your parts from each other, and any control chart built on it is charting the gauge.

How many parts and operators do I need?

Ten parts, three operators and three trials is the AIAG standard. The parts matter more than the count: they must span the variation the process actually produces, because %GRR divides the gauge variation by the total and parts pulled from one good hour understate the denominator, making a sound gauge look far worse than it is.

My gauge failed. What do I fix?

Look at which component dominates. Repeatability points at the equipment — resolution, wear, fixturing, clamping force. Reproducibility points at people and method, which is usually the cheaper fix: an unambiguous datum, a written procedure, an hour of training. A significant interaction is more specific again — find which parts the operators disagree about and the physical cause is usually visible.

Want this to keep itself up to date?

A calculator answers for the data you pasted. A control chart answers for the data your line produced this morning — limits frozen at a baseline you locked, rules evaluated on every new reading, an alert when one trips.

Other free tools

  • Cp / Cpk calculator — Paste measurements — or type a mean and a sigma — with your tolerance, and get Cp, Cpk, Pp, Ppk, the sigma level and the expected parts per million out of spec.
  • Control limit calculator — Give it subgroups or individual readings and it returns the control limits for the chart, the range or sigma chart beneath it, and every constant it used to get there.
  • Control chart generator — Paste a column of numbers, or rows of subgroups, and get a real control chart: limits from the data, Nelson rules 1–4 evaluated, out-of-control points marked.
  • Nelson & Western Electric rules checker — Every run rule evaluated on your data, each violation named in plain English with what it usually indicates — a shift, a trend, tool wear, two machines mixed.
  • Cpk confidence interval & sample size — A Cpk of 1.33 from 30 pieces has a 95% interval of roughly 0.97 to 1.69. See the uncertainty in your own number, and how many parts would settle it.
  • Cpk ↔ PPM ↔ sigma level converter — What PPM is a Cpk of 1.33? What Cpk does 3.4 PPM need? Both conventions shown side by side, because the 1.5 sigma shift is why two sources disagree by a factor of ten.
  • Control chart constants — The whole table, n = 2 to 25, with the formula each constant belongs to. The same values the charts on this site are computed from.
  • X̄-R chart calculator — Subgroups in, X̄ and R charts out — limits from A₂, D₃ and D₄, run rules evaluated, and the range chart shown first because it decides whether the averages chart can be trusted.
  • X̄-s chart calculator — For subgroups big enough that the range wastes them. Limits from A₃, B₃ and B₄, sigma recovered with the c₄ correction, and the s chart read first because it decides whether the averages chart can be trusted.
  • I-MR chart calculator — For processes that give you one number at a time — a batch, an oven, a destructive test. Limits from the mean moving range, run rules evaluated, and an honest note about what an individuals chart cannot see.
  • Levey-Jennings chart with Westgard rules — Paste your QC results with the assigned mean and SD from the package insert, and get the chart with every sigma band drawn and all six Westgard rules evaluated — each one named, and labelled random or systematic error.
  • Bland-Altman plot — Paired readings from two methods, plotted as difference against average — with the bias, the limits of agreement, their confidence intervals, and a test for whether the disagreement depends on the magnitude.
  • Pareto chart generator — Categories and counts in, ranked bars and the cumulative line out — with the vital few named, an "Other" bucket that always sorts last, and an honest warning when the distribution is flat and there is no dominant cause to attack.
  • Box and whisker plot generator — One column per group, boxes side by side. Quartiles, the 1.5×IQR fences, whiskers that stop at real readings and outliers drawn individually — with the quartile method stated, because that is why your plot and Excel's disagree.
  • Histogram and normality test — A histogram with the fitted normal curve, both standard bin rules with a reasoned recommendation, and an Anderson-Darling test that refuses to tell you your data is normal — because no test can.