Westgard rules are a set of decision rules applied to laboratory quality control results, plotted on a Levey-Jennings chart. Each rule detects a different kind of analytical error, and the point of using several together — a multirule strategy — is to catch real problems without rejecting good runs.
James Westgard published the scheme in 1981. The notation looks cryptic and is
not: the number is how many control results, the subscript is how many standard
deviations. 1₃s is one result beyond 3 SD. 2₂s is two results beyond 2 SD.
Where the mean and SD come from
Before any rule means anything, you need a centre line and a sigma. This is the step most often done wrong.
- The package insert value is a starting point, not your SD. Manufacturer ranges are deliberately wide — they have to cover every instrument, every lab, every operator. Using them makes your chart nearly incapable of signalling.
- Your own SD comes from your own data. Run the control material for at least 20 days under routine conditions, on your instrument, with your operators, and compute the mean and SD from that. Twenty points is a minimum, not a target.
- Recalculate when something changes — new lot, new instrument, major service. Not otherwise, and never continuously. A limit that recalculates with every new point drifts along behind a degrading assay and never trips.
That last one is the same trap as a spreadsheet whose control limits reference the whole growing column — see what is an SPC chart. Frozen limits are the entire mechanism.
The six rules
1₂s — warning, not rejection
One control result falls outside ±2 SD.
Do not reject the run on this alone. On a perfectly healthy assay, about 5% of results land outside 2 SD by chance. With two controls per run, roughly one run in ten trips 1₂s for no reason at all. A lab that rejects on 1₂s spends its week repeating good runs.
Its job is to be a trigger: 1₂s says "now check the other rules." Everything below is what you check.
1₃s — reject. Random error.
One control result falls outside ±3 SD.
Under 0.3% probable by chance, so it is almost always real. The signature is random error — imprecision. A bubble in the sample, a partly blocked probe, a pipetting failure, an unstable reagent, a temperature excursion.
2₂s — reject. Systematic error.
Two consecutive control results exceed the same 2 SD limit — both above +2 SD or both below −2 SD.
Two flavours, and they point different directions:
- Within-run: both control levels in the same run breach the same side. The problem affects the whole analytical range right now. Usually calibration.
- Across-run: the same control level breaches the same side on two consecutive runs. A drift that has now crossed a threshold — reagent deterioration, a lamp ageing, calibration slipping.
Either way this is systematic error: a shift in accuracy, not a loss of precision. Every patient result in that run is biased in the same direction.
R₄s — reject. Random error.
The range between two control results within the same run exceeds 4 SD — one control at +2 SD or beyond and another at −2 SD or beyond, in the same run.
This one is within-run only, and only across different control levels. It catches a sudden loss of precision that neither control alone would flag, because neither one is far enough out on its own.
4₁s — reject. Systematic error.
Four consecutive control results exceed the same 1 SD limit, all on the same side.
A small bias, too small for any single point to look wrong. Four in a row on one side of a 1 SD line is roughly a 1-in-1000 event on a centred assay. Classic slow-drift signature: calibration ageing, reagent lot settling in.
Counted either across four consecutive runs on one control level, or across two runs with two levels.
10ₓ — reject. Systematic error.
Ten consecutive control results fall on the same side of the mean — regardless of how far.
The most sensitive rule for small persistent bias, and the slowest to trigger. Distance does not matter; only the sign does. Ten coin flips the same way is about 1 in 500.
Variants exist for labs running different numbers of controls: 8ₓ with four controls, 12ₓ with two. Same idea, adjusted so the false-rejection rate stays comparable.
The two error types, and why the distinction is practical
| Rule | Detects | Typical cause |
|---|---|---|
| 1₃s | Random error | Bubble, blocked probe, pipetting failure |
| R₄s | Random error | Sudden imprecision, unstable conditions |
| 2₂s | Systematic error | Calibration shift, reagent change |
| 4₁s | Systematic error | Slow drift — reagent or calibration ageing |
| 10ₓ | Systematic error | Small persistent bias |
Random error means the assay got noisier. Results scatter further from truth in both directions. Look at the mechanics: sample handling, probes, bubbles, temperature stability.
Systematic error means the assay moved. Results are all wrong in the same direction by roughly the same amount. Look at calibration, reagent lots, standards.
Knowing which one you have is the difference between checking the calibrator and checking the probe. The rule that fired tells you which — which is the entire argument for multirule over a single 1₃s rule.
Choosing a rule set
More rules catch more real errors and reject more good runs. That trade-off is the whole design problem.
- A high-precision assay with good sigma metrics may only need 1₃s. Adding rules to a method that is comfortably meeting its allowable total error buys false rejections and nothing else.
- A marginal method needs the full multirule set, because it will produce real errors that a single 1₃s misses.
The honest version: rule selection should follow from the method's sigma against its allowable total error, not from copying what the lab next door uses. A method running at 6 sigma and a method running at 3 sigma should not have the same QC strategy.
And whatever you choose, write down what happens on a rejection before you need it. A rule that fires and produces a shrug is worse than no rule, because it trains everyone to shrug.
Levey-Jennings and control charts are the same object
A Levey-Jennings chart is an individuals control chart with one difference: the limits are fixed from the QC material's assigned mean and SD rather than calculated from the plotted data.
That is the only difference. The Westgard rules are run rules in the same family as the Nelson and Western Electric rules that manufacturing uses — 10ₓ is recognisably the same test as "nine points on one side of the centre line", and 2₂s is the same shape as "two of three beyond 2 sigma". Two fields, same statistics, different vocabulary.
Worth knowing if you have ever read a manufacturing SPC page and wondered whether any of it applied. It does. See what is an SPC chart.
Run your QC data
The Levey-Jennings chart tool takes your QC results plus the assigned mean and SD, draws every sigma band, and evaluates all six rules — each violation named, and labelled random or systematic error.
Free, no signup, nothing stored.