Advancing QE Skills (8) | How to Conduct MSA for Online Measurement and Automatic Inspection Equipment

By: QTank Published: 9/18/2026 Views: 18
Current rating: ★★★☆☆ Rate this Equivalent to 8 ratings

An electronics company entrusted the final inspection of its assembly line to an online automatic inspection device: parts are measured as they pass through the station, and the device automatically alarms if the measurements exceed the tolerance, directly driving the release of the parts, while the data simultaneously feeds into the SPC dashboard. On the day of equipment acceptance, the supplier conducted a "repeatability test": a standard part was measured 30 times, with a range of 0.004 mm, and a message appeared on the screen stating "equipment accuracy is better than 1% of the specification." The acceptance form was signed on the spot, and the measurement system analysis (MSA) section read "standard part repeatability is qualified." Six months later, customer complaints arose: a batch of parts with dimensions near the lower tolerance limit were judged as qualified and released, causing interference during assembly at the customer's site. Upon reviewing all records, the company discovered that apart from the standard part test, there was no evidence to show whether the device measured the actual dimensions of the parts, nor had anyone calculated whether its judgment threshold had any margin. The root of the problem lies in the fact that good performance on standard parts only proves the device's stability with its own standard parts, not its reliability with external parts—these are two entirely different statistical issues.

1. Key Principles: Standard Part Method and Part Method Answer Different Questions

Good repeatability of a device is often equated directly with "a qualified measurement system." However, the variation in a measurement system consists of several components: device repeatability, bias, workpiece positioning and orientation, part-to-part resolution, and environmental drift. The standard part method (repeatedly measuring a standard part that does not change over time) can only provide two components: repeatability and bias. It answers whether the device is stable and unbiased, but it cannot determine whether the device can distinguish subtle differences between parts—this is precisely the foundation of online inspection.

To assess resolution, the part method must be used: a set of sample parts covering the actual process variation must be selected to observe whether the device can distinguish the differences between these parts from the noise. The two methods cannot be substituted for each other, and the standard part method can never calculate the %GRR.

There is a specific quantification method for the scenario of conducting a repeatability test with a standard part—equipment capability index:

  • Cg = 0.1T / (3s), where T is the tolerance band width and s is the standard deviation of repeated measurements; it reflects the inherent precision of the equipment in the shortest time window.
  • Cgk = (0.1T − |bias|) / (3s), which deducts the system bias from the equipment capability, measuring the actual measurement capability of the equipment.

Converting these: a requirement of Cg ≥ 1.33 is equivalent to the repeatability standard deviation s not exceeding about 1/40 (2.5%) of the tolerance band width; the commonly used Cg ≥ 1.67 for new equipment acceptance is equivalent to s not exceeding 1/50 (2.0%) of the tolerance band width.

Another often overlooked statistical logic: attribute judgment (conforming/nonconforming) is essentially a binary decision, with the false negative rate and false positive rate changing inversely as the judgment threshold moves. Tightening the threshold reduces false negatives but increases false positives; loosening the threshold has the opposite effect. Therefore, there is no such thing as a "very good judgment capability" in general terms; the detection capability must be stratified according to the severity of defects.

2. Practical Steps: Five Steps to Complete MSA for Equipment

Step 1: Define the Output Type and Pre-Set the Criteria. First, determine whether the equipment outputs numerical data or attribute judgments. For numerical outputs, set the Cg/Cgk threshold and the %GRR threshold; for attribute outputs, set the false negative rate, false positive rate target, and the minimum consistency. The criteria must be included in the control plan or equipment management procedures, not just stored on the quality engineer's computer—otherwise, when the criteria are exceeded, someone will inevitably request "let it pass this time."

Step 2: Measure Repeatability and Bias Using the Standard Part Method. Take a standard part that has been traced, fix it in the working position, and measure it continuously 25 times, recording the mean and standard deviation.

  • Repeatability Criteria: s ≤ T/40 (corresponding to Cg ≥ 1.33); for new equipment acceptance, use s ≤ T/50 (corresponding to Cg ≥ 1.67).
  • Bias Criteria: Bias = measurement mean − reference value of the standard part. |Bias| ≤ 10% of the tolerance band width is acceptable; simultaneously use a one-sample t-test, |t| < t(0.05, n−1) to determine that the bias is not significant.
  • It is better to prepare at least three standard parts: one close to the specification center, and two close to the upper and lower limits, to roughly detect the linearity performance of the equipment.

Step 3: Calculate Resolution Using the Part Method. Take 10 sample parts covering the actual process variation (at least 2 parts close to the upper and lower limits, and 1 part outside the tolerance band but close to it), measure each part 2-3 times at different times, and calculate %GRR = 6σ_测量 / T. Criteria:

  • Less than 10% is acceptable.
  • 10% to 30% is conditionally acceptable, requiring the setting of a guard band and clear risk identification.
  • Greater than 30% is not allowed for the judgment and release of key characteristics.
  • Additionally, check the ratio: if the repeatability variation is significantly less than the part-to-part variation (engineers often use less than one-third as the criterion for resolution), it indicates that the equipment can distinguish differences between parts.

Step 4: Verify Detection Capability for Attribute-Type Equipment. Establish a gold sample set stratified by defect type: at least 20-30 samples for each type of defect, with at least 10 critical defects near the judgment threshold; and at least 30 good parts. Criteria:

  • The target false negative rate is 0, but the upper confidence limit must also be reported. For a sample size of 30 parts with zero false negatives, the 95% upper confidence limit is approximately 9.5%; for 50 parts, it is about 5.8%; to claim a false negative rate below 1%, the sample size must be over 100. The sample size determines the level of capability you can claim.
  • The false positive rate (overkill rate) should be controlled within 2%, adjusted according to the cost of line stoppage and scrap.
  • Use Kappa to compare with the referee judgment. Criteria: Kappa ≥ 0.75. Special note: Kappa is highly sensitive to the base rate of good parts; in scenarios with extremely high pass rates, "actual consistency rate 95%, Kappa only 0.50" can occur. In such cases, also consider the actual consistency rate or imbalance correction metrics to avoid being misled by a single number.

Step 5: Verify Consistency with Manual Reinspection and Regularize the Process. Both the equipment and manual reinspection should judge the same batch of parts, forming a cross-tabulation, and all inconsistent parts should be sent for arbitration—arbitration methods should be destructive analysis, higher-grade measuring instruments, or customer-recognized judgment methods, not just "asking a senior technician to take another look." If the inconsistency rate exceeds 3%, a root cause analysis must be initiated. Daily tasks include: measuring a standard part once and plotting a trend chart (using control chart rules for abnormality, not just checking if it is within the tolerance); periodic verification using the gold sample set; and re-analysis after changes such as light source replacement, lens cleaning or replacement, algorithm threshold adjustment, software version upgrades, or major equipment repairs.

3. Common Misconceptions

Misconception 1: Good Repeatability on Standard Parts Equals MSA Qualification. Confusing "device variation" with "measurement system variation." The part-to-part resolution and positioning and orientation variations are not examined, and the device may measure all parts as the same value under good repeatability.

Misconception 2: Treating Unverified Manual Judgments as the Only Benchmark. Calibrating inspection equipment with inspection equipment can result in both being biased. If the consistency (Kappa) of the attribute inspector has not been verified, using their judgment as a benchmark introduces error into the benchmark.

Misconception 3: Reporting a General Overall False Negative Rate. Severe defects are usually easy to detect, while minor defects are the main area of false negatives. Mixing the two into one number conceals the most dangerous part. The false negative rate for critical defects must be reported separately.

Misconception 4: Claiming Equipment Qualification Based on Zero False Negatives. Zero false negatives does not equal zero risk; with insufficient sample size, it only indicates "no false negatives in this batch." A zero false negative rate for a sample size of 30 parts does not support any false negative rate commitment below 9.5%.

Misconception 5: Setting Judgment Thresholds at the Specification Limits. Any measurement system has uncertainty; if the judgment limits are not set inward (no guard band), measurement errors can push some out-of-tolerance parts into the qualified zone. Setting the judgment limits inward according to measurement uncertainty is a rule that must be clearly stated in the equipment release guidelines.

Misconception 6: Not Repeating MSA After Changes. Adjusting the algorithm threshold or replacing the light source can have a greater impact on the judgment results than years of wear on the equipment, and such changes often lack version records.

4. Self-Check List

  • Each online/automatic inspection device has a ledger, indicating its purpose, the product characteristics involved, the version of key parameters, and the calibration status.
  • Numerical indicators have been provided according to Cg/Cgk or %GRR, and the criteria are included in the control plan, not just stored on personal computers.
  • The gold sample set for attribute-type equipment is stratified by defect type, includes critical defects, and the sample size is sufficient to support the claimed false negative rate.
  • Whether the judgment threshold is set inward according to measurement uncertainty and whether it is clearly stated in the release rules.
  • Consistency verification with manual reinspection is based on the true value of arbitration, with a handling threshold and root cause analysis records for inconsistencies.
  • Daily standard part inspections, periodic interim verifications, and re-analysis triggered by changes are solidified in the equipment management procedures.

The value of online automatic inspection equipment lies in its speed and labor-saving capabilities, but its reliability does not automatically increase just because no human is involved. MSA for such equipment still needs to answer only two questions: does it measure the true state of the parts, and does it correctly identify the parts you do not want?


No human involvement does not mean no variation—equipment must also pass the MSA test.

Knowledge code: 6.2.1

Version: v20260918

Author: QTank QTank is dedicated to providing systematic professional knowledge, methodologies, and practical tools for quality management practitioners, helping companies continuously improve their quality capabilities.