Measurement agreement

Reconciling Two Labs: Mean Agreement Does Not Prove Part-by-Part Agreement

Compare paired ceramic measurements by individual differences and their spread, not only laboratory averages or correlation.

Send Drawings7 min read
Bench resistance-measurement instruments with front-panel controls and connection sockets.
On this page

Two laboratories can report the same average resistance while disagreeing enough on individual ceramic circuits to reverse acceptance decisions. The mean difference measures systematic offset; it does not describe the size of a typical part-by-part disagreement. A useful comparison preserves matched specimens, examines the distribution of their differences and judges that spread against a predeclared engineering allowance.

Measurement purpose

Evaluate whether two measurement methods can be substituted for individual ceramic-circuit decisions within a defined disagreement allowance.

Specimens and conditions

Pairs
Same identified circuit element and comparable physical state
Range
Specimens spanning the intended use range with temperature and handling history retained

Equipment and records required

  • Two methods: Defined contacts, excitation, ranges and routine reading procedures
  • Analysis: Paired-data calculations that distinguish mean uncertainty from individual spread

Method sequence

  1. Alignment

    Agree the quantity and permitted discrepancy

    Record: Comparison scope

  2. Measurement

    Collect paired and appropriate within-method repeats

    Record: Hierarchical raw data

  3. Evaluation

    Examine difference patterns and uncertainty before judging interchangeability

    Record: Bounded agreement conclusion

Decision and uncertainty

Approve substitution only when the supported part-level disagreement meets the intended-use requirement.

Include uncertainty in estimated agreement limits and distinguish dependent repeats from independent pairs.

The measurement and acceptance authorities define the allowable discrepancy and approve any correction.

Traceable outputs

Measurement records and required contents
RecordRequired contents
Paired comparisonRaw values, signed differences, range plots, bias and spread
Use boundarySingle or averaged readings, represented conditions, correction validation and unresolved states

Method review decisions

  • Pair the same physical part and measurement state.
  • Separate uncertainty in the mean from spread of individual differences.
  • Define acceptable disagreement before calculating agreement limits.

Match the actual measurement boundaries

Pair results by specimen identity, circuit element and process state. A before-shipment measurement and an after-conditioning measurement on the same part may differ because the part changed, not because laboratories disagree. Align temperature, contact locations, excitation and stabilization requirements before evaluating interchangeability. Preserve any unavoidable differences as study conditions.

Do not pair sorted numerical lists. Matching the smallest result from one laboratory with the smallest from another manufactures agreement by discarding physical identity. If a part cannot be identified reliably, keep that pair unresolved. Include specimens across the range in which both methods are intended to be used; a narrow cluster near nominal cannot establish behaviour near all relevant limits.

Calculate each signed difference before summarizing

Define one subtraction direction and keep it throughout the report. For example, laboratory B minus laboratory A gives a positive difference when B reads higher. Plot or tabulate that difference against the pair’s average and against relevant conditions such as resistance range, temperature or measurement order. This reveals patterns that disappear in two group averages.

A correlation coefficient is not a direct agreement criterion. If one method always reads twenty percent higher, the values can still track each other closely across a wide range. The practical question is whether substituting one result for the other changes the engineering decision by an unacceptable amount. That allowance comes from the measurement task, not from the apparent neatness of a scatterplot.

Do not use the standard error as the part-level spread

Let the paired differences have mean d-bar and sample standard deviation s-d. The standard error s-d divided by the square root of the number of independent pairs describes the estimated mean’s precision. It becomes smaller as more pairs are collected. The individual differences do not become less variable merely because the study contains more parts.

For an illustrative set of twenty-five pairs with mean difference 0.2 ohm and standard deviation 1.0 ohm, the standard error is 0.2 ohm. Treating that standard error as the likely disagreement between laboratories would understate the observed part-level spread. Preserve both quantities with clear labels and avoid the ambiguous expression measurement variation when the report actually means only uncertainty in the average offset.

Use agreement limits with their assumptions visible

When independent paired differences are approximately normally distributed with a stable spread over the range, estimated limits formed from the mean plus or minus 1.96 standard deviations describe a conventional approximate ninety-five-percent range of differences. For the illustrative values above, the limits are minus 1.76 to plus 2.16 ohms. These are estimated population limits, not guaranteed bounds for every future part.

The limits themselves have sampling uncertainty, especially with a small study. A narrow confidence interval around the mean does not make the agreement limits precise. If the tails are heavy, the spread changes with magnitude or the observations are dependent, use an appropriate alternative analysis. Do not retain a simple normal model merely because its output fits the desired allowance.

d_i = B_i − A_i; estimated limits = d̄ ± 1.96 s_d

  • A_i and B_i are paired measurements of the same defined quantity.
  • d̄ is the mean signed difference.
  • s_d is the sample standard deviation of independent paired differences.

Approximately normal differences with no material magnitude-dependent bias or spread are assumed. Limits are estimates requiring uncertainty assessment; repeated measurements need an appropriate clustered model.

Compare the observed pattern with the intended use

Agreement is a fitness-for-purpose decision. A difference acceptable for a broad incoming screen may be unacceptable for a tightly controlled trim target. Establish the allowable discrepancy in the same units or relative scale as the analysis.

Different agreement patterns require different responses
PatternWhat it establishesAction before interchangeability
Small mean, wide individual spreadLittle average offset but substantial part disagreementInvestigate repeatability and compare spread with the use allowance
Stable nonzero mean, narrow spreadA consistent offset may be presentInvestigate cause and validate any authorized correction independently
Difference grows with magnitudeOne constant additive allowance may be unsuitableAssess proportional behaviour and appropriate range-specific analysis
Scatter increases at low signalNoise or measurement-range effects may dominateEvaluate the affected operating range separately
Agreement only after averaging many repeatsThe averaged methods agree better than single readingsDo not claim interchangeability of single readings

Separate repeatability from the comparison between methods

Repeated readings within each laboratory help locate the source of disagreement. If one method varies strongly when the same part is removed and reseated, the fixture or contact procedure may contribute. A mean across many repeats can reduce that variation, but the resulting comparison then represents averaged readings, not the routine single-reading workflow.

Preserve the hierarchy of specimen, laboratory, session and repeat. Treating every repeat as a new independent part exaggerates the information available. Include transport or conditioning checks when the physical part may change between laboratories. A disagreement study cannot identify instrument error uniquely when the specimen itself is unstable and its change is not independently observed.

Validate an offset correction on new paired data

Subtracting the observed mean difference can centre the study data by construction. It does not demonstrate that the correction remains valid on another day, range or specimen set. If a correction is technically justified, record its origin and apply it to reserved pairs collected under the intended routine conditions. Preserve both raw and corrected values.

Do not correct unexplained nonlinear behaviour with a convenient polynomial without a measurement model and independent verification. A fitted curve can hide swapped identities, temperature mismatch or fixture faults. Correct the evidenced cause where possible, and state the limited range of any accepted numerical adjustment.

Report the agreement boundary, not a laboratory ranking

The deliverable contains paired values, difference plots or tables, bias, spread, model assumptions, uncertainty and the predeclared acceptable discrepancy. State whether the conclusion concerns single readings or averages and which configurations were represented. If interchangeability is not established, identify the specific unresolved pattern instead of labelling one laboratory generally better.

For thick-film circuit review, connect the conclusion to the actual characteristic and acceptance decision. Agreement between two methods does not prove either is accurate against a higher-level reference, and disagreement does not prove the ceramic part is defective. ChipSimple can review the agreed measurement interface; calibration claims and laboratory accreditation require their own verified evidence.

Define the measurement comparison question

Provide paired evidence and the decision that laboratory substitution must support.

  • Same-part paired results with element and process-state IDs
  • Temperature, excitation, contact and stabilization procedures
  • Routine use of single readings or averages
  • Predeclared allowable absolute or relative disagreement
  • Within-method repeat data and any correction history

The drawing-upload form loads as you reach this section.