Advancing QE Skills (6) | Gauge Calibration May Be Qualified, but the Measurement System May Not: The Relationship Between Calibration and MSA
1. Comprehensive Certificates, Yet Inconsistent Judgments
A manufacturing company was releasing a batch of aluminum die-cast parts for critical dimensions. The third-party calibration certificates for the micrometers were all within the valid period, and the indication errors were within the allowable range. The "metrology management" item scored full marks during the audit. However, after three months of production, the same batch of parts was judged as qualified during the day shift and as nonconforming during the night shift. The customer also returned parts due to "complaints about conforming products exceeding tolerance limits." The quality manager reviewed the entire gauge ledger, and every micrometer was calibrated and qualified, but the problem seemed to have no clear source.
Later, a GR&R (Gauge Repeatability and Reproducibility) test was conducted: the combined repeatability and reproducibility accounted for 43% of the tolerance, far exceeding the standard. In other words, the gauges themselves were "accurate," but using them to judge the product was not up to standard. This is the most common misalignment between calibration and MSA—calibration confirms whether the gauge's indication is accurate, while MSA assesses the risk of using the measurement system to judge the product. These are not the same thing and cannot replace each other.
2. Principle: One Manages Accuracy, the Other Manages Variability Composition
Calibration traces the indication of a gauge to national or international metrological standards, outputting the indication error, correction value, and measurement uncertainty at specified conditions and specified measurement points. It focuses on the inherent systematic bias of the gauge, which is static and checked at a single point (or a limited number of points).
MSA deals with a complete system: gauge + inspector + operation method + environment + part to be measured + clamping and positioning. It breaks down the data measured by this system into several variability components and then determines whether these variabilities will cause misjudgments in actual judgment scenarios.
To summarize in statistical terms: the measurement result of a product follows
σ² observed = σ² true + σ² measurement
Calibration only suppresses the most visible part of σ² measurement (the inherent bias of the gauge), while repeatability (the spread of multiple measurements by the same person on the same part), reproducibility (differences caused by changing personnel or clamping methods), linearity (inconsistent biases across different measurement ranges), and stability (drift over time) are not addressed at all in the calibration certificate.
The consequences are quantifiable. If the variability of the measurement system (GR&R) reaches 30% of the total process variability, the observed capability will be significantly lower than the true capability. A process with a true Cp of 1.33 may only show a capability of around 1.27 when measured with this system. More critically, the judgment itself—when the measurement variability σ measurement is of the same magnitude as the "margin between the actual measurement and the specification limit," a natural "gray area" forms around the specification line. The same product falling into this area can be judged as conforming or nonconforming depending on who measures it. A practical safety condition is: the measurement variability should not exceed one-third of the margin between the actual measurement and the nearest specification limit; otherwise, the judgment conclusion is unreliable.
3. Five Practical Steps: Ensuring Both Calibration and MSA Are Done Properly
Step 1: Classify Gauges and Match Calibration Requirements to Usage. Classify gauges by characteristic importance into A/B/C grades. A-grade (key characteristic) gauges must meet TUR ≥ 4:1 (the accuracy ratio of the calibration standard to the tolerance of the gauge being inspected, with a minimum of 3:1). The calibration points must be no fewer than 5 and cover the actual usage range (including points close to the upper and lower limits). The judgment rule is: if the indication error at each point | ≤ maximum permissible error (MPE), the gauge is qualified. Calibration certificates that only cover the full-scale endpoints have no binding force on actual usage points.
Step 2: Pass the Resolution Test First. Resolution criterion: ≤ 10% of the tolerance and ≤ 1/10 of the process variability. If the resolution does not meet the standard, the GR&R value will inevitably be inflated and meaningless—what should be done is to replace the gauge or modify the measurement method, rather than continue calculating the GR&R.
Step 3: Conduct Full MSA, Not Just GR&R. Standard procedure: 10 parts × 3 people × 3 times (total measurements ≥ 90, single person single part repetition ≥ 3 times). Sampling has strict requirements: must be taken from the actual process and cover the process variability, with at least one part close to the upper specification limit and one close to the lower limit (the range should ideally cover the actual process ±3σ). Judgment rule: GR&R% < 10% is acceptable; 10%~30% is conditionally acceptable** (requires written customer approval or stricter judgment rules); **> 30% is unacceptable, and the measurement system must be improved and retested.
Step 4: Complete Bias, Linearity, and Stability Tests.
- Bias: Use traceable standard parts, repeat ≥ 10 times, with the criterion |bias| ≤ 5% of the tolerance (or ≤ 5% of the process variability), or use a t-test, p < 0.05 to determine significant bias, which requires correction or adjustment of the judgment rules.
- Linearity: Select ≥ 5 measurement ranges, with ≥ 5 repetitions per range, and use the regression slope or bias at each range to determine p < 0.05 for poor linearity—this is the typical source of "small sizes can be measured, but large sizes cannot."
- Stability: Regularly monitor using the same standard part, plotting points in time order; if points exceed control limits, stop and verify, and only resume after identifying the cause.
Step 5: Set Separate Cycles and Trigger Conditions. A-grade gauges retest MSA every 12 months, B-grade every 24 months, and C-grade can only be calibrated without MSA. MSA must be immediately retested under any of the following conditions: replacement or major repair of the gauge, replacement of the inspector, changes in process parameters or specifications, customer requests, or misjudgment complaints. Note: the calibration cycle and the MSA cycle are two separate schedules, do not use one date to manage both.
4. Common Misconceptions
- Treating "calibration qualified" as "measurement system qualified." Certificates only prove the accuracy of the gauge's indication at specified points, while personnel, methods, environment, and clamping are not within the scope of the proof.
- Treating the uncertainty on the calibration certificate as the GR&R result. The former is the traceable uncertainty of the calibration standard and process, typically at the micrometer level; the latter covers the variability of the entire measurement system under real usage conditions, and the scales and scopes do not correspond.
- Conducting MSA only before PPAP or customer audits. A new batch of personnel, major repairs to the gauge, or changes in the process can significantly alter the measurement system, making old reports merely historical documents.
- Selecting "clean" samples. Choosing 10 parts with almost identical dimensions will make the GR&R look good, but the conclusion will have no predictive power for mass production—samples that do not cover process variability are equivalent to not conducting the test.
- Calibration points and usage points are disconnected. Calibration is qualified near the full scale, but only 10% of the range is actually used, and the linearity and bias in that segment are completely unverified.
- Conducting bias and linearity tests but not checking p-values. Judging as qualified based solely on "small values" ignores statistical significance and can easily lead to misjudgments with small sample sizes.
5. Self-Check List
- Key characteristic gauges have been classified, and A-grade gauges meet TUR ≥ 4:1 and cover the actual usage range (≥ 5 points)
- The resolution of each in-use key gauge is ≤ 10% of the tolerance and ≤ 1/10 of the process variability
- GR&R (10 parts × 3 people × 3 times) has been completed, with samples taken from the actual process and covering variability, and the criteria are applied according to the 10%/30% classification
- Bias, linearity, and stability tests have been conducted and p-values or control limit conclusions provided
- MSA retest cycles and trigger conditions are written into the control plan, with separate schedules for calibration and MSA cycles
Calibration is about "calibrating the gauge," while MSA is about "proving that the gauge can provide reliable judgments in this context." Clearly defining the boundaries between the two ensures that quality data can be confidently used to draw conclusions.
A calibrated gauge does not guarantee a reliable judgment; calibration and MSA each manage a different aspect.
Knowledge code: 6.2.1
Version: v20260916
Author: QTank QTank is dedicated to providing systematic professional knowledge, methodologies, and practical tools for quality management practitioners, helping companies continuously improve their quality capabilities.