See the Authors

When Is a Measurement Change Real? Repeatability and Accuracy in Compression Garment Fitting

When Is a Measurement Change Real? Repeatability and Accuracy in Compression Garment Fitting

A compression garment patient returns for reassessment. The calf circumference recorded today is 4mm smaller than the measurement taken at the previous visit six weeks ago. The question is not whether the number changed. It is whether that change reflects a real reduction in limb volume or whether it falls within the normal variability of the measurement system being used.

If the system’s measurement noise is ±5mm, a 4mm difference is not a finding. It is an artefact. Acting on it, adjusting the garment specification, documenting the treatment response, and modifying the compression class means acting on data that does not contain the information it seems to contain.

This problem is not specific to tape measurement. It applies to every body measurement system used in compression garment fitting and monitoring, including digital systems. And it is systematically underaddressed in how measurement technology is evaluated and procured by compression garment manufacturers and clinical teams.

Five Terms That Are Not Interchangeable

The vocabulary around measurement quality is used imprecisely in both clinical and commercial contexts. For R&D and medical directors, the distinctions matter operationally.

Accuracy refers to how close a measurement is to the true value of the quantity being measured. A system that consistently reads 3 mm above the actual circumference is inaccurate; it has a systematic bias.

Repeatability refers to the consistency of measurements taken under the same conditions: same operator, same patient, same session. A system can be repeatable without being accurate if the same error recurs consistently.

Reproducibility refers to consistency across different conditions: different operators, different sessions, and different sites. A system with good intra-rater repeatability may show significantly lower inter-rater reproducibility when a different clinician performs the measurement.

Measurement error is the total deviation of an observed measurement from the true value. It encompasses both systematic bias (accuracy) and random variation (repeatability and reproducibility).

Clinically meaningful change is the minimum change in a measurement that corresponds to a real change in the patient’s clinical status that matters for treatment decisions. This is a clinical judgement, not a statistical property of the measurement system, and it must be established separately from the measurement error threshold.

A compression garment manufacturer relying on measurement data for longitudinal garment adjustment needs to know all five and specifically needs to know whether the observed change exceeds both the measurement error threshold and the clinically meaningful change threshold before making a production or clinical decision.

Why a 5mm Difference Is Not Automatically Meaningful

Every measurement system: tape measure, structured-light scanner, smartphone-based platform has a noise floor: the level of variation that occurs in repeated measurements of the same patient under controlled conditions, independent of any real physiological change.

The relevant metric for longitudinal comparison is the smallest detectable change (SDC), also referred to in clinical literature as the smallest real difference (SRD). This is the minimum numerical difference between two measurements that exceeds the expected measurement noise with a defined level of statistical confidence, typically 95%.

A change smaller than the SDC cannot be interpreted as real. It may represent a genuine physiological event, but the measurement system cannot distinguish it from noise. Interpreting sub-SDC differences as findings produces false signals in both directions: apparent improvement where there is none and apparent deterioration where the patient’s condition is stable.

The SDC is the minimum threshold for interpreting longitudinal measurement data. Any change reported below that threshold is statistically indistinguishable from measurement noise — regardless of how precise the measurement system appears to be.

Research on upper limb circumference measurement in breast cancer survivors with lymphoedema has established minimum detectable change values at the segmental level ranging from approximately 3.4% to 7.6% of limb volume, corresponding to 2.7 to 14.6 mL depending on the limb segment. At the global arm level, the MDC was approximately 2.4%. These figures were established for intra-rater tape measurement under controlled conditions; inter-rater variability, which is more representative of real-world clinical settings, produces larger uncertainty intervals.

The practical implication is that making a garment adjustment decision based on a 4mm or 5mm circumferential change, without knowing the SDC of the measurement system being used, lacks a valid evidentiary basis.

The Anatomical Correspondence Problem

Measurement error is not only a function of the instrument. It is also a function of the specific anatomical location where the measurement is taken.

For longitudinal comparison to be valid, measurements must be extracted at anatomically equivalent locations on every occasion. This requirement is harder to satisfy than it appears, and it is the single most underestimated source of longitudinal measurement error in compression garment fitting.

In tape measurement, landmark placement is operator-dependent. A measurement described as “mid-calf” is not a fixed anatomical point; it is a relative position that varies with how the operator interprets and locates it. Small differences in landmark placement produce real differences in circumference measurement, particularly in anatomically complex regions such as the ankle or the thigh base, where the limb cross-section changes rapidly over short distances.

Digital systems address the issue differently. Machine learning-based measurement platforms apply consistent landmark detection algorithms across captures, reducing the operator-dependent component of landmark variability. However, anatomical correspondence across sessions still depends on consistent patient positioning, consistent capture conditions, and the model’s ability to reliably locate the same anatomical position in presentations that may have changed; for example, a lymphedematous limb that has significantly reduced in volume between visits presents differently to the model than it did at the initial capture.

For any system used for longitudinal monitoring, the validation evidence should address anatomical landmark consistency explicitly, not only overall measurement accuracy.

The Metrics That Actually Matter for Validation

Compression garment manufacturers and clinical teams evaluating measurement technology for longitudinal use should require validation evidence expressed in the appropriate metrics. The following table maps the key metrics, what they measure, and what they do not tell you.

Metric What it measures What it does not tell you
Mean absolute error (MAE) Average magnitude of deviation from a reference value across a sample The distribution of errors — whether they are clustered or spread across the range
Bias and limits of agreement Systematic directional offset and the spread of individual differences between two methods Whether the system is accurate for a specific anatomical region or patient subgroup
Standard error of measurement (SEM) The expected variation in repeated measurements of the same patient by the same system Whether a given numerical change exceeds that variation
Coefficient of variation (CV) Measurement variability as a proportion of the mean, enabling comparison across anatomical sites The absolute size of the error, which matters more than the proportion in some clinical contexts
Smallest detectable change (SDC) / smallest real difference (SRD) The minimum change that can be distinguished from measurement noise with statistical confidence Whether that change is clinically meaningful — a separate and equally important threshold
Intra-rater repeatability Consistency of the same operator measuring the same patient on the same occasion Inter-rater consistency — how much results change when a different operator is involved
Inter-rater reproducibility Consistency of measurement across different operators Whether the system is suitable for a specific clinical or manufacturing use case

Two methodological points apply across all these metrics. First, correlation alone is not sufficient for validation. A high correlation coefficient between two methods indicates that they tend to move together, but it does not confirm that they agree numerically. A system that consistently reads 10mm above the reference will correlate strongly with it while being systematically biased. Limits of agreement and bias analysis are required to detect these discrepancies.

Second, validation evidence should be specific to the patient population and anatomical regions relevant to the intended use. A system validated for healthy volunteers may perform differently in oedematous limbs. A system validated for calf circumference may not have been evaluated at the ankle or thigh. The evidence must match the application.

What Recent Research Shows About Digital Measurement Systems

The clinical literature on digital measurement systems for lymphoedema and compression applications has moved significantly in the direction of rigorous psychometric evaluation of reliability, validity, measurement error, and smallest real difference rather than simple correlation with tape measurement.

A 2026 study evaluating a structured light scanner for leg volume measurement in patients with lymphoedema reported excellent intra- and inter-rater reliability, with small measurement error and the smallest real difference values that the authors considered clinically usable. Total leg volume showed a strong correlation and concurrent validity with both circumference measurement and optoelectronic volumetry. The authors also assessed clinical feasibility through timing and a purpose-designed questionnaire.

A 2025 study evaluating smartphone LiDAR measurement in patients with lower extremity lymphoedema reported excellent inter-rater reliability at the ankle, calf, and knee. Performance at the foot and thigh showed greater discrepancy with tape measurement findings that are consistent with the anatomical complexity of those regions and the challenges of landmark correspondence at sites where limb cross-section changes rapidly.

Both studies illustrate a maturing evidence base that is moving away from simple accuracy claims and toward the measurement science framework that R&D and medical directors should require from any system under evaluation. Neither study, however, resolves all open questions, particularly around oedematous presentations that fall outside the studied populations, home-based or unsupervised capture conditions, and longitudinal tracking across extended times.

Implications for Compression Garment Manufacturers

Measurement uncertainty has three direct operational consequences for manufacturers who use body measurement data for garment specification, longitudinal reordering, and treatment monitoring claims.

 

Reorder decisions. A garment reorder triggered by a measured change that falls within the system’s measurement noise is a reorder based on an artefact. Establishing the SDC of the measurement system and using it as a decision threshold reduces unnecessary remakes caused by measurement variability instead of genuine limb change.

 

Longitudinal garment adjustment. Patients with chronic conditions requiring periodic garment adjustment depend on measurement data that can distinguish real progression or improvement from noise. If the measurement system’s SDC has not been established for the relevant patient population and anatomical regions, longitudinal adjustment decisions lack a valid evidence base.

 

Treatment monitoring claims. Any claim that a measurement system supports treatment monitoring, tracking the response to compression therapy and documenting lymphoedema progression or reduction, implies that the system can detect clinically meaningful changes above its noise floor. That claim requires explicit validation evidence: SDC values for the patient population, anatomical regions, and capture conditions in which the system will be used. Without it, the claim cannot be supported.

Conclusion: The Threshold Question

Before any measurement system is used to track patient change over time, one question must be answered: what is the smallest change this system can reliably detect in this patient population, at these anatomical regions, under these capture conditions?

That question has a specific, quantifiable answer: the smallest detectable change that must be established through appropriately designed validation studies. It cannot be inferred from accuracy data alone, assumed from correlation coefficients, or generalised from studies conducted in different populations or anatomical regions.

For compression garment manufacturers, this threshold is not a regulatory technicality. It is the operational boundary between measurement data that supports clinical and production decisions and measurement data that creates noise in both. A system that cannot answer the threshold question for your specific use case, which has not been validated for longitudinal use regardless of how it performs in a demonstration environment.

The technology to achieve reliable, repeatable body measurement for compression applications exists and is improving. The evidence base is developing appropriately. Manufacturers and clinical teams must have the discipline to ask for the right evidence and apply it correctly.

Frequently Asked Questions

How do we establish the SDC for a measurement system we’re already using?

Run a repeated-measures study: measure the same patients twice under the conditions you actually work in, with no real physiological change expected between captures. Calculate the standard error of measurement from that data, then derive the smallest detectable change as 1.96 × √2 × SEM for a 95% confidence level. The result is specific to that system, that population, those anatomical sites, and that capture protocol — it does not transfer to a different setup.

Can we just use the SDC figure a vendor publishes?

Only if the validation conditions match your use case. A figure established with healthy volunteers, supervised capture, and intra-rater repeat measurements will understate the uncertainty you’ll see with oedematous limbs, unsupervised capture, or measurements taken by different staff. Treat a published SDC as evidence that the vendor has done the work, then ask whether the study population and anatomical regions correspond to yours. Where they don’t, the number is indicative rather than applicable.

How many patients does a repeatability study need?

Commonly cited methodological guidance for reliability studies puts the floor around 50 subjects, with smaller samples producing confidence intervals too wide to be operationally useful. What matters as much as the count is whether the sample reflects the range of presentations you measure — limb volumes, oedema severity, body types, and any conditions that change how the limb presents to the system.

 

Does the SDC differ between anatomical sites?

Yes, and reporting a single figure across the whole limb obscures this. Regions where the cross-section changes rapidly over a short distance — the ankle, the foot, the thigh base — carry more landmark uncertainty than the mid-calf, and the published research on digital systems reflects that pattern consistently. Validation evidence should be broken out by site, and decision thresholds should be set per site rather than globally.

Isn’t a high intraclass correlation coefficient enough?

No. ICC describes how well the system separates individuals within the sample it was tested on, which means it rises with the heterogeneity of that sample — a study with a wide range of limb sizes will produce a flattering ICC even where absolute measurement error is substantial. It also says nothing about agreement in millimetres. Reliability coefficients and measurement-error metrics answer different questions, and longitudinal decisions depend on the second.

Does this apply to initial garment specification, or only to monitoring over time?

The SDC governs longitudinal comparison — deciding whether a change between two captures is real. Initial specification is a different question, answered by accuracy and bias against a reference method: whether the measurement is close enough to the true value to select the correct size or compression class the first time. Both matter, but conflating them leads to specifying a system on accuracy data and then using it for monitoring it was never validated for.

What changes when measurements are captured at home rather than in clinic?

Unsupervised capture removes the controls that most validation studies rely on: positioning, lighting, clothing, time of day, and limb state. Each introduces variance on top of the system’s baseline noise, which widens the SDC. If home capture is part of your workflow, the validation evidence needs to have been generated under those conditions — a clinic-validated threshold applied to home-captured data will flag changes that aren’t there.

How often should a system be revalidated?

Whenever something upstream of the measurement changes: a model update or algorithm revision, a change to the capture protocol or hardware, or extension into a patient population the original study didn’t cover. Machine learning-based systems in particular can shift behaviour between versions in ways that aren’t visible from the output alone, so a validated threshold shouldn’t be assumed to carry forward across releases.

Does any of this mean tape measurement is the safer option?

No. Tape has a noise floor too, and inter-rater variability in landmark placement makes it larger in most real clinical settings than intra-rater studies suggest. The point isn’t that one method is reliable and the other isn’t — it’s that any method used for longitudinal decisions needs its measurement error quantified for the specific use case, and tape is usually the method for which that has never been done in-house.

About Esenca Sizing

Esenca Sizing is a machine learning-powered body, hand, and foot measurement platform used by manufacturers and distributors in medical garments, workwear, and footwear. The platform is browser-based, requires no hardware or dedicated app, and is ISO 8559 certified and GDPR audited. With over 500,000 measurements processed across 21 supported languages, Esenca Sizing provides consistent, body measurement data for production and fit optimisation workflows.

Put an end to the
sizing guesswork!

With Esenca Sizing, your shoppers can obtain hundreds of precise body measurements in less than 30 seconds. Fast, precise, and simple!

Related Articles

When Is a Measurement Change Real? Repeatability and Accuracy in Compression Garment Fitting

When Is a Measurement Change Real? Repeatability and Accuracy in Compression Garment Fitting

A compression garment patient returns for reassessment. The calf circumference recorded today is 4mm smaller than the measurement taken at the previous visit six weeks ago. The question is not whether the number changed. It...

by

Clinical validation vs. marketing claims: what evidence to demand from a measurement vendor

Clinical validation vs. marketing claims: what evidence to demand from a measurement vendor

Every digital measurement vendor will tell you their system is accurate. Fewer will tell you accurate at what, compared to what, on how many people, or under whose independent review. In compression garment distribution and...

by

Why Standardised Body Measurement Matters in Compression Garment Production

Why Standardised Body Measurement Matters in Compression Garment Production

In compression therapy, the garment is the treatment. Unlike a drug whose active ingredient is fixed at the point of manufacture, a compression garment delivers its therapeutic effect through a precise mechanical relationship between the...

by

Best 3D Body Measurement Tools That Use AI and Computer Vision (2026) #2

Top AI Body Measurement Tools for Workwear & Uniforms (2026)

Getting the right garment size for a workforce of hundreds or thousands used to mean fitting events, physical tape measures, and weeks of logistical coordination. Today, machine learning and computer vision have compressed that process...

by

Best 3D Body Measurement Tools That Use AI and Computer Vision (2026)

Best 3D Body Measurement Tools That Use AI and Computer Vision (2026)

1. Introduction What “3D Body Measurement” Actually Means in 2026 The phrase “3D body measurement” is used loosely across the sizing technology industry — and that looseness creates confusion when procurement teams, HR managers, or...

by

What Is 3D Body Scanning? A Complete Guide for Workwear & Uniform Buyers

What Is 3D Body Scanning? A Complete Guide for Workwear & Uniform Buyers

What 3D Body Scanning Actually Is 3D body scanning captures a person’s measurements using a camera and computer vision instead of a tape measure. Point a phone, capture a few images, and the software extracts...

by

Home

Solutions
Industries We Serve

Case Studies

About Us

Blog