EQAWarden · the statistics

How a result becomes a score, and a score becomes a suggestion.

EQAWarden's arithmetic is the arithmetic RCPAQAP and ISO 13528 already use, applied consistently and shown on screen with its numbers. This page sets out the scoring, the reliability of the target, the series and cycle statistics, the lot and IQC comparisons, and the fixed logic that turns findings into a suggested outcome. The full scientific basis, with its references, ships with the application.

The single result

The APS score says how much of the allowed error a result has used.

RCPAQAP states most analytical performance specifications as a fixed amount up to a decision point and a percentage above it. The limit is evaluated at the target, not at the laboratory's result.

The arithmetic

APS at the targetThe fixed amount when the target is at or below the decision point, otherwise the percentage of the target
DifferenceResult minus target, positive when the laboratory reads high; also shown as a percentage of the target
APS scoreDifference divided by the APS at the target. A score of 1 means the whole allowance has been used
BoundaryRCPAQAP flags a result that is more than the APS from the target, so a difference exactly equal to the APS is within. EQAWarden compares the reported decimals so boundary cases classify exactly as RCPAQAP does
Censored resultsA result reported as less than or greater than is not assessed, as RCPAQAP does not assess it; it is shown as not scored

The zones

  • WithinAPS score up to two thirdswithin
  • Warningtwo thirds to one: the ISO 13528 "questionable" bandwarning
  • Outsideabove one: RCPAQAP flags the result for reviewoutside
  • Not scoredcensored, or no target or APSnot scored

ISO 13528 derives its standard deviation for proficiency assessment as a third of the allowed error, so EQAWarden also shows an ISO-equivalent z, three times the APS score, always labelled as derived from the APS and never as an RCPAQAP z-score. The divisor is configurable and shown.

What one result can show. A laboratory performing exactly at that standard deviation puts about one result in twenty in the warning zone and about one in 370 outside by chance. A laboratory with an SD of half the APS has about one result in twenty outside. One result outside the APS warrants investigation, but does not by itself show whether the error is random or systematic.

The expected result and its reliability

A score is only as good as its target.

RCPAQAP's expected result is an all-method median, a category median (analytical principle, measurement system, reagent or calibrator) or a specified target. EQAWarden judges how much to trust it before it interprets the score.

Small groups

RCPAQAP computes group statistics only for six or more results. Below six, EQAWarden flags a small target group and reads the score as descriptive.

Target uncertainty

Where a group SD and size are printed, the target's standard uncertainty follows ISO 13528's formula for a median. Up to 0.3 of the assessment SD it is negligible; above it, a provisional adjusted z is shown; above 0.7, following IUPAC, EQAWarden does not interpret z at all.

Commutability

RCPAQAP's lyophilised material must not be assumed commutable. All-method comparisons on non-commutable material are treated as descriptive; agreement with same-method peers is the conclusion that can be drawn. A not-standardised analyte against an all-method target is flagged, and a persistent offset is then read as the method group's difference, not as calibration.

Clerical screens

Blunders are screened first, on results outside the APS only, and never alter a result.

Transcription is the most frequent cause of EQA failures in the published troubleshooting reviews, and step two of RCPAQAP's own interpretation flowchart. At most three candidates are reported per result, ranked by the candidate's APS score.

ScreenWhat it testsStrength
Sample swapFor the two samples of a survey, whether the results fit better exchanged: fires only when the two targets are far enough apart to tell, at least one result is outside, and both land within once swapped. A likelihood ratio is reported. Two or more swap candidates in one survey, with every other discriminable analyte also fitting better exchanged, point to a specimen swap rather than an entry error.strong
Unit errorThe result divided or multiplied by one of the analyte's conversion factors (computed from IUPAC atomic weights; HbA1c by the IFCC to NGSP master equation) lands within the APS. Factors within 10% of one are skipped.strong
Decimal errorThe result times a power of ten, from a thousandth to a thousand, lands within the APS.strong
Digit errorAn adjacent-digit transposition, or a dropped or duplicated digit, lands within the APS. Worded as likely only when the candidate uses half the allowance or less, because at low concentrations such changes can land inside a wide APS by chance.weak to moderate
Wrong fieldThe result is within the APS of another analyte's target on the same sample and unit.moderate
Series and cycle

A single survey cannot separate bias from random error. A series can.

A series is one analyte on one instrument with one target source, in order of analysis date. The two samples of a survey share a run, so tests that assume independence run on survey summaries, and a method category change restarts the regression, as RCPAQAP does.

StatisticMethodThreshold owner
RepeatsTwo or more of the last six scored results outside the APS.EQAWarden default
BiasFrom six surveys: the mean survey score with a t interval, which does not assume the laboratory's SD equals the assessment SD, and a sign test. Fires below p 0.05 when the mean score is at least a quarter of the APS: small to a half, material to one, at or beyond the APS above one.Test standard; thresholds EQAWarden default
RunsA current run of six or more same-signed survey summaries, and the ten-times rule.IUPAC and Westgard, adapted
Westgard-style rulesThe 2 of 2, 4 of 1 and range rules on the ISO-equivalent z. Together they give about one false signal in eight over 24 results for a laboratory performing at the assessment SD, which is why they only suggest monitoring.Westgard 1981 definitions
DriftMann-Kendall against time (exact for ten or fewer surveys) with a Theil-Sen slope, and Pettitt's step estimate beside it. Step and drift are not chosen automatically with fewer than ten surveys.Published tests
CUSUMTwo-sided tabular CUSUM with reference value 0.5 and decision limit 5; the estimated start is the last zero of the signalling sum. With that limit a false signal within 24 results occurs about 4% of the time, and a shift of one assessment SD is detected within 12 results about 73% of the time.NIST
Cycle regressionRCPAQAP's own method: least squares of results on targets over at least six samples, giving the standard error of the estimate, CV, average bias over the low, mid and high fitted points, and MPS as twice the standard error plus the absolute bias, divided by the APS at the mid target. An atypical result is excluded once and the line refitted. Deming regression is the primary estimate when target uncertainty would attenuate the slope by more than 1%; Passing-Bablok is a robustness check only, because its intervals are unreliable below about 40 samples.RCPAQAP
Proportional, constant, curvatureA slope interval excluding one with an effect of at least a quarter of the APS at the mid target; an intercept interval excluding zero of at least a quarter of the APS at the lowest target; and, with eight or more targets, a significant quadratic term or too few runs of residual signs, each test at half the alpha so the rule's false-alarm rate is kept.EQAWarden default
ImprecisionThe CV compared three ways (meets, does not meet, inconclusive) against half the percentage APS and, when configured, the desirable CV from biological variation. The sigma metric is shown descriptively and never flags, because it depends heavily on the allowable error chosen.RCPAQAP; Fraser

Many analytes, one report. A report may carry 30 to 60 analytes; at a 5% false-alarm rate per test the chance of at least one false alarm across 20 analytes is 64%. RCPAQAP's single-result flags are never adjusted. For tests on cycle statistics EQAWarden applies Benjamini-Hochberg at a 10% false discovery rate across the report's analytes; a finding that fails it is worded "possible pattern, not significant after allowing for the number of analytes reviewed" and counts half in the cause ranking.

Lot steps

Before and after each lot change, with a permutation test.

LotWarden supplies the lots put into use and their dates; EQAWarden works out which lots were in use on each analysis date and how confident that link is: high, medium when a changeover fell within seven days or an earlier lot may still have been in use, low when the date is after LotWarden's data end, undated, or has several candidates.

The step test compares the survey summaries before and after each reagent or calibrator lot change, up to six each side: the difference in means, a Hodges-Lehmann estimate, a Welch interval and an exact permutation p, with a fixed seed where the splits are too many to enumerate so the result is reproducible. It fires below p 0.05 when the step is at least half the APS. One calibrator or reagent serving several analytes, followed by same-direction steps in two or more of them, is strong evidence.

Power is low with few surveys: about 0.16 for a one-SD step with three surveys a side and 0.35 with six. With three or fewer surveys on a side a non-significant result is never read as "no effect". Lot-based cause weights are scaled by the link's confidence.

Internal QC beside EQA

Does the control material tell the same story?

For each control level with the same control lot on both sides, IQC values from the 30 days before the analysis date (at least ten, else the lot's imported laboratory statistics) are compared with the day of analysis and the two days after: the difference, its percentage, a Welch interval with an autocorrelation-adjusted sample size, and the change in SD units.

Consistency with the EQA deviation is tested with a single statistic on the two percentage shifts and their standard errors; a value below two is read as "consistent with the same shift". IQC that moved with the EQA result points to a genuine analytical change to confirm with patient samples. IQC that did not move points to something specific to the EQA samples. IQC that stepped at a lot change while EQA did not is a control-material-specific response, which one large study found in about two in five reagent-lot QC events.

Unity rule violations within a day of analysis, a calibration or maintenance event in the three days before, the peer SDI with Bio-Rad's bands, and a control lot change inside the window are each reported in their own right.

Ranking and outcome

Points for causes, a fixed order for outcomes.

How causes are ranked

Causes follow CLSI QMS24's categories: clerical, material, analytical, target and chance. Each finding adds or subtracts points for the causes it supports or counts against: strong 3, moderate 2, weak 1, against 1 or 2. Lot findings are scaled by lot confidence, and findings that fail the false-discovery adjustment count half.

  • Strongly supported5 points or morecheck first
  • Supported3 to 4 pointscheck
  • Possible1 to 2 pointsconsider
  • Not supported0 or fewerlisted

One pattern is never counted twice under different headings: after a lot step, the calibration weight of the series findings that restate it goes to the lot; after a standardisation flag, the persistent-offset weight goes to commutability.

The four outcome rules, in order

Corrective actionBoth samples of the survey outside the APS; repeats; a bias of at least the APS surviving the adjustment; or a lot step with the post-change mean at or beyond the APS
InvestigateAny result outside the APS; MPS above 1; a bias of at least half the APS; drift or a CUSUM signal with the current mean at least half the APS; a lot step; a sample-level shift on this analyte's sample
Acceptable, monitorA warning-zone result; a run; a Westgard-style signal; a bias between a quarter and a half of the APS; imprecision not met; a small target group; an RCPAQAP colour; a method-group pattern; a correlation factor on the assay
AcceptableNone of the above

Every result outside the APS is investigated and recorded, as RCPAQAP's flowchart, NPAAC and ISO 15189:2022 clause 7.3.7.3 expect. A strongly supported clerical cause does not lower the outcome: the transcription or swap is itself a nonconformity to record.

Wording

The words scale with the evidence.

Every sentence gives the estimate, the interval and the count, names the target source, separates observation from interpretation, lists the consistent explanations with the check that discriminates them, and never states a cause as fact.

LevelCriterionWording
0No rule or test triggered"No evidence of ... in these n results (a bias of up to ... cannot be excluded)"
1One observation, a non-significant test, or too few results"is consistent with", "may indicate"
2p below 0.05 after adjustment, or a sequential rule signal"suggests"
3Level 2 plus independent corroboration: an IQC step on the same date, several analytes, or p below 0.001"provides strong evidence of"; never "proves"
n/aBelow the minimum data"There is insufficient information to assess ..."
What small numbers cannot show

Two to twenty-four results a cycle. EQAWarden says what that cannot prove.

One survey

Two results cannot separate bias from random error, and no sign pattern reaches p 0.05 with five or fewer surveys.

A lot change

"The lot change had no effect" cannot be concluded with four or fewer surveys a side. Power for a one-SD step is 0.16 to 0.22.

An SD, a line

An SD from ten degrees of freedom has a 95% interval from 0.70 to 1.75 times the estimate. Linearity cannot be judged below eight targets; Passing-Bablok slopes are not definitive below about 40 samples.

Thresholds and their owners. Every threshold sits in one parameters file with its owner. RCPAQAP owns the outside-the-APS rule, the group size of six, atypical results and MPS above 1. ISO 13528 owns the assessment SD, the z bands and the target-uncertainty limits. IUPAC, NIST, Westgard, Bio-Rad and Fraser own theirs. The rest, 61 of 101 parameters, are EQAWarden defaults: design choices listed in the laboratory review pack for the chemical pathologist to ratify or change before real review records are made.

Bring your statistician, or your scepticism.

Every number on an assessment page opens the rule and the reference behind it. Ask us to show the working on the walkthrough.