The model was validated, it works, everyone has moved on to something else. Eighteen months later it is wrong. And nothing in the system said so.
The first year goes well: the model was built on recent batches, the raw material has not changed, the instrument is new. Then the excipient supplier changes site. The lamp ages. The window fouls. A new shift presents the sample differently.
None of those events raises an alarm. The model carries on returning a value, in the right unit, at the right rate. And the value has moved.
The question to ask before buying: what happens on the day the model drifts, and who notices?
In short
A monitoring plan written before start-up, in four answers: which statistics (at least Hotelling’s T² and the Q residual at every measurement), how often they are reviewed, who looks at them, by name, and which threshold triggers what. Drift shows in the residuals, not in the predicted value: a model outside its domain returns perfectly plausible values.
A model does not age, its world changes
A model calibrated on samples is a relation learnt on a population of samples, at a given moment. It stays accurate as long as the process resembles what it was during calibration. Whatever moves it away makes it wrong without warning.
- The raw material. A change of supplier, or of manufacturing site at the same supplier, alters the particle size, crystallinity or incoming moisture. The spectrum changes with it.
- The season. Workshop humidity, the temperature of the chilled water network, the quality of an agricultural input. Slow variables that twelve months of calibration do not always cover.
- The light source. A lamp ages, its emission spectrum distorts, its intensity falls. A continuous, regular drift, the most predictable of all. And one of the least monitored.
- The window or the probe. Fouling alters the received intensity and sometimes the baseline, batch after batch, without ever crossing a visible threshold.
- The operator. Where manual presentation is involved, packing, thickness, positioning, a change of shift shifts the measurement.
- The equipment. A blade replaced, a screen changed, an agitation speed modified after optimisation: the process improves, the model drops off.
The guidance puts it for continuous manufacturing, and the observation holds everywhere: ICH Q13 (16 November 2022, § 3.1.3) states that input materials may call for the evaluation and control of attributes beyond those a batch manufacturing specification usually includes — particle size, cohesiveness, adhesiveness, hygroscopicity, electrostatic charge, specific surface area. What continuous manufacturing asks of measurement.
Three drifts, three different treatments
The word “drift” covers three situations that do not call for the same actions. Confusing them leads to recalibrating a model when a window needed cleaning, or the reverse.
| What has moved | How it is recognised | What it calls for |
|---|---|---|
| The instrument | The drift also appears on a stable control material. It is continuous and affects all the products measured on that instrument | Maintenance, cleaning, source replacement, wavelength scale check. No recalibration |
| The process | The control material stays stable, but the product spectra move away from the calibration domain. Often linked to a datable event | A process investigation. The model is not at fault: it signals that it is being shown something else |
| The model | Instrument and process are stable, but the gap between predicted values and reference values widens or becomes biased | Enriching the calibration set, or even rebuilding it. The only case that justifies a recalibration |
Diagnosis goes through the residuals, not through the predicted value. It is the most useful technical point on this page, and the least intuitive.
A predicted value that stays within limits proves nothing: a model outside its domain returns perfectly plausible values. What betrays drift are the indicators describing the position of the sample relative to the model. The distance to the centre, called Hotelling’s T², which says whether the sample resembles those of the calibration. And the spectral residual, called the Q residual, which says whether the model was able to explain this spectrum. A rising T² signals an atypical but recognisable product. A rising Q residual signals something the model has never seen. These two statistics are monitored continuously. The predicted value cannot be monitored on its own.
What the software supplies, and what it leaves to you
The building blocks exist. The trigger is where the work sits
We have gone through the product documentation of industrial chemometric suites. The monitoring blocks are there: Hotelling’s T squared and Q residuals in real time, influence plots to identify the variables responsible, alarms passed to supervision, moving blocks to set the thresholds over a reference period. Good tooling.
The recalibration trigger and the model performance follow-up in production are built around these tools. The software gives you the instruments for measuring the health of the model. The decision rule — when that health has changed, on which criterion, and what to do next — is written into your procedure, and we build it with you.
What the draft Annex 22 would turn into an obligation
Everything above reads today as good practice. A text in preparation would give it a different standing.
From 7 July to 7 October 2025 the European Commission consulted on a new EudraLex Annex 22 on artificial intelligence, drafted jointly by the EMA Inspectors’ Working Group and PIC/S. Nothing is adopted. But its scope covers models that obtained their functionality “through training with data, rather than being explicitly programmed“, static and with a deterministic output, used in critical applications “e.g. to predict or classify data“. A chemometric model ticks those criteria, and the draft provides no exclusion for chemometrics. The detail, and what stays with you.
What the draft would require overlaps this page almost entirely: regular monitoring of performance against defined metrics, monitoring that input data stays inside the model’s domain — exactly the role of Hotelling’s T squared and the Q residual — change and configuration control, and human review with records kept. A site that has already written its monitoring plan would have almost nothing to change. A site that has not would discover the requirement and the drift at the same time.
Two points would be exceptions, and they are worth naming now. Test data independency: the draft asks that the people who developed and trained the model “have never had access to the test data“, with access control and an audit trail on that set — yet it is almost always the same chemometrician who builds and tests. And explainability: the draft asks that the variables contributing to a given prediction be recorded, which puts further pressure on the non-interpretable models discussed below. Status verified on 20 September 2026: consultation closed, nothing adopted.
A monitoring plan holds in four questions
A monitoring plan fits on one page and answers four questions. It is written before commissioning rather than after the first incident.
- Which statistics? At minimum Hotelling’s T squared and the Q residual at every measurement. Then, at a defined interval, the gap between predicted values and laboratory reference values on a small number of samples.
- At what frequency? The per-measurement statistics are continuous. Their review carries a written periodicity, monthly, quarterly or per campaign. Monitoring with a scheduled review is monitoring. Without one it is a history.
- Who looks at them? The question that decides whether the plan lives. A dashboard with no named owner stays closed. It takes a name, a deputy and allotted time.
- Which threshold triggers what? Three levels are enough. An alert threshold that triggers closer observation, an action threshold that triggers an investigation, and an exit rule: a proportion of out-of-limit batches above which a recalibration is worked up. Each threshold is justified over a reference period.
The plan is built during integration rather than in operation. It is also the moment to provide for the upkeep of the sample set that will serve the updates, and for the transfer of the model to a second instrument. And the earlier a drift is detected, the shorter the production period to re-examine. Monitoring carries economic value as well as documentary value.
The reference analysis effort is not constant. It is concentrated on building and validating the model. In routine, only the monitoring analyses set by this plan remain, tightened only when an indicator raises an alert. The laboratory method keeps its role: it is what confirms the model, not the reverse.
Recalibrating in a regulated environment is more than a technical gesture
Modifying a validated model modifies an analytical procedure. Three points are settled before, rather than during.
Who decides
The decision belongs to the change management system rather than to the chemometrician who observes the drift. Know in advance who proposes, who assesses the impact, who approves and who authorises the return to service. Quality control, production and quality assurance are all concerned, and a decision taken by one of the three alone is read at the next audit.
Which documentation
A protocol written before the operation, the data that motivated the decision, traceability of the samples added, a comparison of performance before and after, and an identified and archived model version. The ALCOA+ principles apply to a model as to any data: the previous version stays retrievable.
What impact on the validation
Not everything brings a full revalidation. Adding samples inside the domain already covered carries a different weight from extending the range or changing the pre-processing. ICH Q2(R2) and Q14 give the frame, ICH Q12 the logic of post-approval change management. That classification is prepared at the initial validation. And the update period itself is prepared: what takes over while a model changes version.
“If I recalibrate, do I have to refile?” — the answer reads on two axes
The FDA guidance of August 2021 on NIR analytical procedures answers with a table. The reporting route reads at the crossing of the impact of the change on the procedure’s performance and its impact on product quality.
| What it takes | In which case |
|---|---|
| Nothing to report, the quality system covers it | The change bears little on the procedure’s performance, and product quality is not at stake. |
| Annual report | It bears moderately on the procedure while product quality is not at stake — or lightly while it is. |
| Thirty-day notification | It bears heavily on the procedure while product quality is not at stake — or moderately while it is. |
| Prior approval | It bears heavily on the procedure, and product quality is at stake. |
What that gives in practice: replacing an analyser, where it forces a new model on a procedure used for release, calls for prior approval. The same replacement on a steering measurement stays inside your quality system. The level of ambition chosen at the start is what sets the cost of every later recalibration, which is a good reason to choose it knowingly.
The particular case of non-linear models
Strong performers, and two points to settle before purchase
Support vector machines (SVM) and boosted decision trees appear more and more in software offers, and they keep their promise: on frankly non-linear relations they do better than a multivariate linear regression. Two reservations carry weight in a regulated environment.
First, they are not interpretable: which spectral band carries the prediction stays out of view. Specificity remains demonstrable — ICH Q2(R2) accepts either absence of interference or comparison with an orthogonal procedure, and neither requires opening the model. What is lost is the third route the same text allows: technology-inherent justification, available where the measurement physics itself ensures specificity. The demonstration becomes wholly experimental. Second, they produce neither Hotelling’s T squared nor a Q residual natively, and those are the two statistics everything above rests on. Without domain detection added explicitly, such a model cannot signal that it is being shown a sample it has never seen. The point is worth naming before the purchase rather than in front of an inspector.
Frequently asked questions
How often should a model be recalibrated?
There is no universal periodicity, and a fixed one works against you: it recalibrates the models that are fine and leaves the others to drift. The rule is conditional. Recalibration happens when the monitoring triggers it, against thresholds defined in advance. What is periodic is the review.
How do we know whether the drift comes from the instrument or the process?
Through a stable control material, independent of the product, measured at regular intervals. If it drifts, the instrument is the cause. If it stays stable while the product spectra move away, it is the process or the raw material. That control costs little and settles the question in minutes, and it is set up before it is needed.
Can the model update be automated?
Technically yes, and in a regulated environment a model that updates itself calls for governance. The version used to release a batch is identified, approved and retrievable. An ungoverned automatic update changes an analytical procedure outside change control. The useful automation bears on the detection rather than on the correction.
Does this call for a chemometrician in house?
Not necessarily full time, and it calls for an identified, available and trained competence, in house or by contract. The arrangement worth avoiding: a model delivered by a third party, with the reading skill left off site, and a maintenance contract covering the hardware alone. A point to settle at scoping, and it belongs as much to training as to technique. Training and autonomy.
Ask the question rarely asked at purchase: what happens on the day the model drifts, and who notices?
Forty-five minutes is enough to know what your installation watches today, and what it would take to detect a drift before it costs a batch.