Can it be measured? Forty-five minutes is enough to find out.Talk to us about your processLanguageENFR

Sampling error is reduced by the sampling plan

Your analyser settles one half of the result. The other half is settled before it, and that is the half you have most hold on.

Could you say where your product is sampled, at which point in the cycle, and how many times? Most teams answer without hesitating on the detector and the spectral resolution. On that question, the answer is rarely written down anywhere.

That is what puts the gain within reach: it calls for no extra instrument and no chemometric development. It is settled on a sampling plan, it is quantified, and it stands up in front of an auditor.

Where the two errors of a measurement are born, and what reduces each The chain runs from the batch to the displayed value. Sampling error is born upstream, in the sampling, and is reduced by the sampling plan alone. Analytical error is born downstream, in the instrument and the model. Sampling error Born here, and answered by the sampling plan. Analytical error Born here, and the one the instrument reduces. The batch heterogeneous by nature The sampling where, when, how many The test portion grams or milligrams The acquisition detector, source, noise The model and its reference The value displayed What reduces it The mass taken, the number of increments, their spread in time and in space. That is the list. Doubling the signal to noise of the analyser leaves it where it was. What reduces it A better instrument, a better calibration, the averaging of several acquisitions on the same portion. Repeatability says nothing about accuracy: the portion taken is what decides it.
The two errors are born in different places. The one on the left is settled before the product reaches the instrument. The one on the right is settled after. A better analyser acts on the second.

The case of the sampling thief

The reference tool disturbs what it samples

A thief pushed into a powder bed displaces material at the moment it samples it. It drags fines, packs locally, and favours the entry of some particle sizes into the cavity over others. The bias follows the depth, the angle, the cohesion of the powder and the operator.

The thief often stays the only tool available, and its results serve as the reference for whole models. The caution is about the method rather than the tool: when an in-line measurement and a thief sample diverge, the gap does not automatically indict the in-line measurement. Working up both is part of a serious method validation.


Two errors that are constantly confused

The result you read carries two uncertainties of different natures, and they are not dealt with in the same place.

ErrorWhere it comes fromHow it is reduced
Analytical errorDetector noise, source drift, model quality, the laboratory referenceBetter instrument, better calibration, averaging several acquisitions on the same portion
Sampling errorThe fact that the material analysed is not representative of the whole: constitution heterogeneity, distribution heterogeneity, segregation, biased samplingOnly through the design of the sampling plan: mass taken, number of increments, distribution in time and space

The theory of sampling (as formalised by Pierre Gy and then disseminated by Kim Esbensen) establishes that the second far outweighs the first in most heterogeneous materials. The literature of this field commonly reports a ratio of 10 to 50 between sampling error and analytical error. That statement belongs to that literature.

The sources, for anyone who wants to check

  • DS 3077, Representative sampling — Horizontal standard, 3rd edition, DS 3077:2024, published in October 2024, Dansk Standard. The only standard devoted to the theory of sampling: guiding principles, unit operations and management of sampling errors.
  • Esbensen K. H. & Abu-Khalaf N., Before reliable near infrared spectroscopic analysis — the critical sampling proviso, parts 1 and 2. The second part deals with the requirements specific to near infrared.
  • Esbensen K. H., Materials Properties: Heterogeneity and Appropriate Sampling Modes, Journal of AOAC International, vol. 98, no. 2, 2015, pp. 249-251. Open access.
  • ICH Q13, Continuous Manufacturing of Drug Substances and Drug Products, Step 4 of 16 November 2022. Its § 3.1.5 lists among the points to consider the impact of physical sampling on the material flow, which may affect the state of control. What this text asks of measurement.
  • USP <1097>, Bulk Powder Sampling Procedures. The compendial chapter on bulk powder sampling: theory, sampling plan, tools — and it names segregation error and sampling method error. An informational chapter, which therefore creates no requirement; but it is the text that allows this reasoning to be held in a pharmacopoeia’s vocabulary. Where it sits among the texts.

Kim Esbensen is part of our network of experts. This is not an appeal to authority: when a project stumbles on the representativeness of sampling rather than on instrument performance, that is the competence we bring in. The people who work on our projects.

The central idea, and it is unwelcome for an instrument buyer. No gain in signal-to-noise ratio, no detector cooling, no extra spectral fineness reduces a sampling error. These improvements act on how well the material presented to the sensor is interrogated. They have no hold on whether that material resembled the batch.

In other words: beyond a certain level of instrument performance, moving further up the range buys nothing useful as long as the sampling plan has not been dealt with.

Sampling also carries risks that have nothing to do with representativeness. Opening the process exposes the product to contamination. Diluting a sample for analysis can change what was to be measured, a particle size or a state of aggregation. Transporting the sample lets it evolve before analysis: settling, moisture pick-up, cooling. An in-line measurement removes these three risks. It does not remove the question of representativeness: it moves it to the measurement point.

Three words missing from most specifications

The increment

The elementary portion: material taken at one point and one instant. A spectroscopic acquisition is one, and the mass concerned is small. Its value comes from belonging to a set.

The composite sample

The aggregation of several increments taken at different places and instants. The one construction that estimates a batch mean with a known uncertainty. Multiplying increments reduces sampling error. Repeating on the same increment reduces noise.

Correct sampling of a stream

On a moving product, a portion is correct when the whole section of the stream has the same chance of contributing. That is what lets an in-line sensor leave single-grab territory: seeing the full section, and letting the product travel past it. The section covered gives the delimitation, the travel gives the increments. A sensor watching one zone of a motionless volume stays a single grab, repeated.

A heterogeneous lot, as they all areA single grabSame lot. Three grabs. Three results.Seven distributed increments, joined into one sample16 %24 %36 %matrixanalyteEach window really holds the stated content: the bed was built that way, and it can be counted.composite sampletrue content of the lot25 %
Three single grabs in the same lot give 16 %, 24 % and 36 %. Seven distributed increments give 25 %, the true content of the lot. Each window really holds the stated content: what the figure shows can be counted. Construction after Kim H. Esbensen, KHE Consulting — by the courtesy of Kim H. Esbensen.

The four concrete answers

None calls for chemometric development. All four belong to the design of the measurement and to how the process is run.

  • Multiply and distribute the increments. Triggering acquisition at a known position of the equipment, once per revolution of a blender for instance, makes the increments sweep the volume rather than return to the same place. You then average over a block of measurements and compare successive blocks rather than isolated points. The number of increments rises with the heterogeneity encountered: a design choice to document, not a universal constant. This is the only lever that acts on representativeness itself, which is why it comes first.
  • Increase the mass of each increment. A wider spot, a multi-point probe, a geometry that exposes more product. Once the distribution is secured, the mass settles the residual fundamental error, and it is calculated. It comes second, and the order matters: the mass of a sample is not what makes it representative, its distribution is. It is decided at the design of the installation.
  • Match the measurement time and the spot size to the linear speed of the process. On a stream, an integration time that is too long averages portions you meant to distinguish. Too short, it interrogates a small fraction of the flow. The point most often left behind when a laboratory trial becomes an in-line installation.
  • Fit the section into the field of view, rather than the other way round. On a free-falling stream, narrowing the flow locally — a constriction, a chute of reduced section — is often enough for the spot to cover the whole width. It is a piece of mechanical design, and it is what moves a sensor from a single grab to sampling the theory calls approximately correct — it writes “~TOS-correct”, tilde included — and “fit-for-purpose representative”: what remains is no longer the principle but the residual imperfections — window fouling, edge effects.

The four levers combine, and they are quantified. The number of increments needed follows the real variability of the product and the precision sought on the mean. That is a calculation.

A last lever exists, and it belongs to one family of measurement only: sampling no more. An in-line camera sees every unit pass, and the question of the number of increments gives way to the question of what the image really shows. The problem moves rather than disappearing, from a mass taken to a surface seen, and it holds only where the criterion is visible, which leaves out most composition questions. Where imaging replaces a sampling, and where it does not.

Pairing the spectrum with its reference

A chemometric model does not learn from a spectrum. It learns from a couple: a spectrum, and the reference value that corresponds to it. The whole question sits in the word “corresponds”.

This is the direct extension of sampling, and it is where most projects are lost. Sampling decides what the reference measures. Pairing decides what the model can learn.

Three ways to get it wrong

The scale, and it is almost never a choice. One spectrum per unit, one reference value for a group. It is often seen as a laboratory economy. Most of the time it is a physical constraint of the reference method.

A Karl Fischer titration calls for a minimum mass: on gelatin shells, it takes about ten of them to obtain a single value. A composition analysis on seeds commonly calls for a hundred grams, that is several thousand seeds. The spectrum, for its part, is taken on one unit, and could be taken on each of them.

The gap between the mass the sensor probes and the mass the reference calls for commonly reaches a factor of a thousand. The two measurements then do not bear on the same object, and no recording convention corrects that. How much material does a sensor really analyse?

The consequence is mathematical. A mean of thirty varies about five and a half times less than the individuals composing it: the dispersion of a mean falls as the square root of the number of units averaged, not as that number. The model is then asked to explain a nearly constant quantity, from spectra that carry the whole individual variability.

The object, and it is the most serious. The spectrum and the reference do not bear on the same material. Two forms: one portion is scanned and another is sent to the laboratory, or the probe looks at one zone while the thief goes into another. In both cases, the model learns the relation between two samples, not between a spectrum and a content.

On a powder drying, it is the typical error. Both portions carry the same batch number and the same hour, and nothing in the data says they are two different objects. The defect leaves no trace in the calibration set.

The hidden assumption, and why it turns against itself

Pairing two distinct samples works on one condition only: that the matrix is perfectly homogeneous. That assumption is almost always implicit, and it is almost never written down.

Yet it refutes itself. If the material were perfectly homogeneous, the uniformity question would not arise, and the measurement would have no reason to exist. Assuming homogeneity to build the model amounts to assuming solved the very difficulty the model has to handle.

The consequence is direct: the dispersion between the two samples adds to the model error, and no indicator separates it from the measurement error. You believe you are assessing an instrument; you are assessing a sampling plan.

The time. The spectrum is acquired at one instant, the sample travels to the laboratory, and the material carries on evolving. On a reaction or a drying, a few minutes are enough to offset the reference from what the spectrum saw.

The symptom reads on two figures, and it is counter-intuitive: an acceptable RMSEP with a coefficient of determination close to zero.

That combination does not signal a bad model. It signals a model that predicts the mean and nothing else. Its error stays low because it never moves away from the centre, and its explanatory power is nil for the same reason. A report that publishes only the RMSEP does not show it.

No mathematical treatment gets out of a model information that the reference does not contain. The unpaired variability stays in the residual, whole.

Four answers, and one has to be chosen

The choice is made before the first acquisition. Afterwards, the trials have to be run again.

  • Model at the scale of the reference. Where the reference bears on a group, the object modelled is the group. The ten spectra are averaged into one, and the couple becomes coherent again. The model is valid and returns a group value. You lose the individual, you gain a method that holds.
  • Look for a reference of smaller support. Not all methods have the same mass requirement. On water, a gravimetry on a single unit is sometimes practicable where the titration is not. It is worth working up before giving up on the individual measurement.
  • Predict the individual, and write it down. A model built on means technically applies to individual spectra. That is legitimate, on one condition: stating that its individual accuracy is not demonstrated, for want of a reference able to demonstrate it. Doing it without saying so is the fault.
  • Reformulate as a state measurement. The dispersion between units of one group reads in the spectra themselves, with no reference at all. It does not give a content, it gives a homogeneity, and that is often the real question. Following the convergence of a dispersion.

All four are answers. None of them consists of changing instrument.

What is worth preparing on your side

  • An equipment drawing and the view of an operator: where the material circulates, where it stagnates, where it segregates. Ten minutes of discussion often beat a study.
  • The scale at which uniformity is asked of you: the mass of the dosage unit, or the size of the sub-batch you have to be able to defend. It sets the whole sizing.
  • Your historical results, increment by increment rather than as means. The current dispersion is the honest basis for comparison.
  • A position or phase signal from the equipment, to trigger acquisitions on something other than a fixed interval. Often a minor change to the automation, and it is what makes a sampling plan defensible.
  • The possibility of varying the process deliberately, where the regulation and the material allow it. A sampling plan that has met a heterogeneous batch is a plan that has been tried.

Conditions for success

A correct increment cuts the full cross-section of the stream

The theory of sampling sets one criterion ahead of any calculation: an increment is correct only where it intercepts the entire cross-section of the moving stream — the full width of the belt, the full bore of the pipe. A probe observing one point of that cross-section through a window does not meet it: it interrogates a fraction of the stream, and always the same one. What follows is not noise but a bias, and a bias does not average out. This is why the sampling interface is designed before the instrument is chosen.

Representativeness comes from varied positions

An in-line measurement produces acquisitions in very large numbers, and that is its decisive advantage over any manual sampling. Where the sensor always looks at the same zone of a bed that does not move, those thousands of values are one increment repeated. Varying the positions and the instants is what makes them representative.

A correct sampling plan reports the batch faithfully

It secures that what you measure resembles what the batch holds, which is exactly the point. Conversely, an accurate mean can sit over a local heterogeneity that matters at the scale of the dose, so the dispersion between increments counts as much as their mean. Developed on content uniformity.

Some configurations call for moving the measurement point

A nozzle exposing stagnant product only, a position downstream of a systematic segregation zone, an access intercepting one edge of the stream: these are structural sampling biases. Moving the measurement point is the answer, and saying so early is what keeps the project on solid ground.

What it changes, in practice

The first effect often surprises: part of the effort usually put into instrument performance shifts to the design of the measurement point, mechanical access and trigger automation. The second is methodological. Gaps between in-line measurement and laboratory stop being mysteries, since you know which share comes from sampling.

The third is regulatory. An explicit, justified and documented sampling strategy is a dossier element. “We measure continuously” is not. It is one of the deliverables of scoping, and it precedes the choice of equipment rather than following it.

Frequently asked questions

Where does the factor of ten to fifty come from?

From the sampling theory literature, out of the work of Pierre Gy and widely carried by Kim Esbensen. It is an order of magnitude observed on heterogeneous materials rather than a constant applicable as it stands to your process. We quote it for what it is: a reminder of proportion between two sources of error, to be established on your own case where the stake justifies it.

Does a continuous measurement solve the problem by construction?

In part, and where the increments vary. Continuous measurement brings a number of increments no manual plan can reach, which is a real advance. It brings nothing against a position bias, since a poorly placed sensor produces a large number of unrepresentative values. The question of where stays whole.

Can poor sampling be recovered by data processing?

No, and that is what sets this error apart. Noise averages out, a drift corrects, a scattering effect pre-processes. A sampling bias is a property of the material that was presented to the sensor, so the missing information is nowhere in the data set. No algorithm restores what was never measured.

Should we stop comparing instruments then?

Not at all, and the order matters. Instrument performance decides feasibility: where the signal does not separate, nothing follows. Once that bar is cleared, the next gains come almost entirely from the sampling. Comparing two optical configurations makes sense where one exposes a markedly larger mass of product than the other. Much less where the gap is signal to noise alone.

How is a sampling plan demonstrated to be correct?

By a written argument and by variability trials: replicate samples under the same conditions, and a comparison of the dispersion between positions with the dispersion between repetitions in the same place. The ratio between the two says which dominates. A modest exercise in cost, and it is what makes the file defensible in front of an auditor. Both carry a name in sampling theory: the replication experiment and variographics. They are what allow a single-point sample to be accepted, as a demonstrated exception.

Before comparing instruments, let us compare your sampling points.

Forty-five minutes is enough to place the main source of error in your measurement, say whether it belongs to the sampling or to the analysis, and sketch the increment plan that would reduce it. And where your sampling is already sound, you will leave with that confirmation.