It is the first question asked in almost every project. It has an answer, and that answer starts with a list rather than a number.
How many distinct batches, how many raw material suppliers and how many seasons do your samples cover today? That list is what sets the size of the calibration set. The number follows from it, and it is calculated. Announced ahead of it, it rests on nothing from your process.
Coverage determines the quality of a model. Not the count.
A model calibrated on samples learns a statistical relation on the population it is shown, and that relation only holds on that population. The useful question is therefore not “how many samples”, but “which sources of variation do your samples contain, and which are missing”.
Thirty well-spread samples are worth more than two hundred taken on the same day from the same batch. Statistically, the two hundred are little more than a single sample measured two hundred times.
In short
Six dimensions have to be covered: the range of the quantity, ends included, the batches, the raw material suppliers, the seasons for a natural material, the operators and the equipment. These dimensions multiply together: it is their product that sets the size of the calibration set, not an absolute number.
The six dimensions a calibration set must cover
| Dimension | What has to have been seen | What happens if it has not been seen |
|---|---|---|
| The range of the quantity | Values spread over the whole useful range, ends included, and not bunched around the target. | The model is precise at the centre and wrong at the bounds, where decisions are taken. |
| Batches | Several distinct production batches, made on different dates. | The model learns the peculiarities of one batch and takes them for the general law. |
| Raw material suppliers | Each qualified supplier, and if possible several grades per supplier. | A change of source takes the model out of its domain without warning. |
| Seasons | For any material of natural origin: several harvests, several geographical origins. | A model built on a single campaign drops off the following year. |
| Operators | The ways of presenting the sample, positioning the probe, packing a powder. | Handling variability lodges in the prediction error, without anyone knowing where it comes from. |
| Equipment | The lines, vessels or instruments on which the method will really be used. | The model works on the line where it was born, and nowhere else. |
These six dimensions multiply with each other, and it is that product, not an absolute value, that sets the size of the calibration set. A single-supplier, single-line process on a stable synthetic material will call for far less than a food process with three origins and two seasons.
Does your calibration set contain non-conforming batches?
Almost all calibration sets recovered from production contain only conforming batches. That is logical, they are the ones being made, and it is precisely where the model is expected to perform that it has learnt least.
A model trained only on conforming batches has learnt to describe normality. Faced with a batch that departs from it, it does not say “I don’t know”. It returns the value closest to what it knows. In other words, it pulls the result towards conformity. That is the worst possible behaviour for a device meant to protect the patient or the consumer.
A range that is too narrow shows in the indicators, and it flatters
A model calibrated on a tight range often displays excellent correlation coefficients. Statistics reward spread, not accuracy. A high coefficient on a narrow range is not proof of performance. What has to be looked at is the error in absolute value, compared with the tolerance of the specification.
What a cross-validation actually establishes
Cross-validation removes part of the samples, recalculates the model without them, then predicts their values. It is a useful tool, for choosing a number of latent variables for instance. Proof of performance comes from an independent validation set.
Where your samples all come from the same batch, the ones removed resemble the ones left in every respect, and the model predicts them perfectly. The indicator obtained describes an ability to interpolate inside one batch rather than to handle next week’s batch.
- Split by batch rather than by sample. Removing a whole batch, recalculating, predicting that batch: the one cross-validation that simulates the real situation.
- Keep a genuinely independent validation set, made of batches the model has never seen and never taken back into the calibration set to improve the result. That reintegration removes the independence.
- Report the error on that set, in units of the measured quantity, against the tolerance of your specification. It is what a method validation expects, and what an assessor looks for first.
This is not a school of thought. USP <1039> puts it plainly: any batches used in calibration and cross-validation must not be considered or reused as an independent dataset for method validation. And Ph. Eur. 5.21 names the split by batch — leave-block-out cross-validation — then requires three partitions: a training set, an internal test set, and an independently generated external test set.
Building the variability production has not yet supplied
When production does not supply the necessary spread, two routes exist. And you need to know which one is open to you before committing to a calendar.
You can make deliberately offset batches
It is the comfortable situation. A design of experiments on development batches quickly builds a coverage that production would take years to supply. Content deliberately low and high, moisture off target, blending time shortened.
A single precaution: that the offset batches are offset the way a real incident would offset them. An under-dose obtained by removing active ingredient does not always look like an under-dose resulting from segregation at transfer.
Production runs as it runs, a frequent case
Sterile product, qualified continuous process, saturated capacity: making a batch on purpose is not practicable. You then work with what you have: archive samples, historically rejected batches, start-up and end-of-campaign drifts, reduced-scale trials.
And you accept the consequence: a narrow validity domain at the start, reinforced monitoring, a progressive enrichment plan. It is a documented choice, not a hidden defect.
Reference values have their own uncertainty
A model calibrated on samples does not predict the truth. It predicts what the laboratory method it was calibrated on would have given. Its error, measured against that method, cannot be shown to be lower than the method’s own: it contains it (Ph. Eur. 5.21). And if the reference values are dispersed, doubling the number of samples makes up for nothing. Before counting samples, you need to know what the method that will give them a value is worth.
That error can be determined, and the protocol is short. Two analysts run the reference method on the same samples, one on the first day, the other on the second. The mean difference between them is the between-operator bias; a paired test says whether it is significant. If it is not, the standard deviation of those differences gives the error of the method, up to a factor: it bears on the gap between two measurements, so it is that standard deviation divided by √2 that has to be kept. The figure obtained includes day and operator, and that is exactly the ceiling that matters to a model. About ten samples are enough, and the result is written into the validation protocol.
One consequence, and one nuance that counts. The consequence: a model whose error falls clearly below that of the reference has not done better than the laboratory, it has run into something that deserves to be examined. The nuance: that ceiling concerns the error measured against the reference. The in-line measurement’s own precision can be better than the laboratory’s, precisely because it avoids the sub-sampling the laboratory is forced to do.
A reference sample must correspond to what the sensor saw
The spectrum is acquired on a few milligrams of material. The reference value, for its part, comes from a sample of several grams analysed in the laboratory. If the product is heterogeneous at that scale, the pairing is wrong and the model learns sampling noise. It is the question of representativeness, and it comes before any question of quantity.
On the pure-component route, the question changes nature
Everything above holds for the model calibrated on samples, the one where a spectrum is regressed on reference values. There is another approach, where the question of the number of samples gives way.
In a pure-component calibration, no model is built on a population of products. Each spectrum is decomposed onto a library of pure component spectra. No batches to collect, no seasons to wait for, no design of experiments to negotiate with production. The effort moves to the qualification of the library: traceable standards, spectra acquired in the real solvent, a dilution series covering the useful range, selectivity verified between components.
The approach carries strict conditions: a declarable composition, a quantity that really is a concentration, a fixed optical path. Its own limit sits elsewhere, since the model sees what it has been told. Where the conditions are met, the answer to how many samples becomes: none of your batches. It is the first thing to work up in a scoping phase. What calibration burden your measurement carries.
The question to put to a supplier
One sentence is enough to sort the people you are dealing with, and it is not about a number:
“On how many batches, on how many raw material suppliers, and with what independent validation set?”
Whoever answers “sixty samples” without mentioning batches or independent validation is describing a lead time, not a method. Whoever answers “it depends on the variability of your material, and here is how we characterise it before costing” is describing a piece of work.
One question comes before that of the number, by the way: will each spectrum have its own reference value? Sixty spectra paired with two group values do not make sixty calibration samples. Pairing the spectrum with its reference.
Frequently asked questions
Is there really an absolute minimum?
There are technical minima tied to the statistical method: a model with several latent variables is not calculated on a handful of samples. What limits, though, is the number of distinct batches and the width of the range covered. You meet unusable models at two hundred samples and robust models at forty.
Do the generic models supplied with an instrument save all this?
They save time at start-up, and that is real. They sit alongside a verification of accuracy on your own materials, since they were built on a population that is not yours. The minimal approach is to measure a set of your own samples, compare with the reference values and correct any bias. That verification is itself a campaign, and a shorter one.
Can we start small and enrich later?
It is often the best strategy, provided it is owned. An initial model with a narrow domain, residual monitoring and a dated enrichment plan is worth more than a wide model built on poorly spread samples. What stays consistent is that the declared domain widens once the corresponding samples have been added. Keeping a model alive is continuous work.
How long does a calibration campaign take?
The measurement time is negligible. What takes time is waiting for batches. If you make two a month and need about twenty distinct ones, the calendar is set by production rather than by chemometrics. So the point to establish early is whether you can make offset development batches, and, even more, which calibration burden applies. Building the model itself is described on the chemometrics page.
And if the variability is simply not there?
It happens, and it is good news. If every batch gives the same value, a quantitative model has little to learn. The need becomes a departure detection rather than an assay, which belongs to another approach and costs less. Saying so is more useful than selling a calibration.
Describe your variability. You will know what to collect, and over how long.
Forty-five minutes is enough to identify the sources of variation to cover, to know whether your historical batches suffice, and to check whether your case falls to a calibration burden that would save the whole campaign.