The spectrum is the easy half.

A detector returns a noisy vector with a fluorescence slope and the occasional cosmic ray. Turning that into a number a process engineer can act on is a modelling problem — and it is where most of our engineering sits.

StackCNN + classical chemometrics
Calibrationclient's own samples
Updates12 months in contract

Scroll

From raw counts to a process value

Five stages. The first two are housekeeping that decides whether anything downstream can work; the middle two are the model itself; the last one is what keeps a number trustworthy six months after commissioning.

300 1650 cm⁻¹ intensity fluorescence baseline cosmic ray baseline removed · spike rejected intensity normalised every channel is an input — no hand-picked peak windows CNN — non-linear, full-spectrum viscosity · multi-component mixtures PLS · PCA — linear, interpretable scores, 2 components small calibration sets loadings a chemist can read Acid value — alkyd resin 12.4 mg KOH/g prediction error ±0.2 mg KOH/g · R² > 0.99 reference method agreement validated probe reference · self-diagnostics nominal
01 — Raw

What the detector actually hands over

A vector of pixel counts carrying the Raman bands, a fluorescence slope from the medium, shot noise, and now and then a cosmic ray that looks exactly like a very sharp peak. None of it is a concentration yet.

~1000+measurement points in a multi-hour process 8 cm⁻¹resolution per channel
02 — Preprocessing

Remove what is not chemistry

Baseline correction, spike rejection, normalisation. Every measurement is validated against the probe's built-in reference, and disturbances are compensated automatically rather than left for the model to guess at.

Every measurement is validated and disturbances are compensated automatically. Self-diagnostic algorithms detect contamination and inform the operator about deviations. dr inż. Maciej Jaworski · PIPC
03 — CNN

The network reads the whole spectrum

A convolutional network takes the full vector rather than a handful of chosen peak windows. That is what makes non-linear behaviour and heavily overlapping bands tractable — and why parameters that are not concentrations at all, viscosity among them, can be predicted.

Non-linearphenomena classical models miss Multi-componentmixtures, not single analytes
04 — Classical chemometrics

PLS and PCA are part of the stack, not the competition

Classical algorithms are chosen where they win: small calibration sets, interpretable loadings, fast validation against a reference method. The mix of both families is deliberate, and it generalises better than either one on its own.

HybridCNN + advanced classical chemometrics Per processselected for the chemistry at hand
05 — Output and guardrails

A number with a stated uncertainty

The value leaves with its error bar, an agreement check against the reference method, and the analyser's own diagnosis of the optical path. Recipes drift; models are updated for the first twelve months under the contract, and re-calibrated when the process changes.

±0.2 mg KOH/galkyd acid value, feasibility ±0.25 Pa·salkyd viscosity, feasibility Fractions of a percentuncertainty, tuned per line
Model pipeline · step 01 / 05

Why a mix, rather than a single method

Purely classical models are cheap to validate and blind to non-linearity. Purely neural models absorb non-linearity and demand data. Real process chemistry needs both, chosen case by case.

What the CNN is for

Full-spectrum input, non-linear response, overlapping bands, multi-component mixtures, and derived properties such as viscosity that no single band encodes.

What PLS and PCA are for

Compact calibrations from a limited number of samples, loadings a chemist can inspect and argue with, and quick validation against titration or HPLC.

What the combination buys

Faster, more accurate models that generalise better than a single-family approach. In a feasibility study on alkyd resins, a CNN over the whole spectrum outperformed linear regression on the same data, reaching R² above 0.99 for both acid value and viscosity.

The model has a lifecycle, not a delivery date

Calibration, validation, maintenance, re-calibration. A model that is never touched again is a model that quietly stops being right.

Calibration

Starts from feasibility samples — usually the client's R&D material plus Gekko laboratory work — and the reference values that go with them.

Validation

Predictions checked against the established reference method on the client's own process, not on a public data set.

Reinforcement

Further training on production data once the analyser is running, which is when the awkward cases show up.

Re-calibration

Triggered by what actually changes in a plant: a new recipe, a different feedstock supplier, modified process parameters. Model updates are covered for the first twelve months under the agreement, and they roll out without stopping the analyser or the line. The hardware warranty period is a separate commercial term.

Accuracy actually measured

Figures from feasibility studies on real process samples. Each one is a specific matrix and a specific reference method — not a platform-wide claim.

Alkyd resin — acid value
R² > 0.99 · ±0.2 mg KOH/g
Alkyd resin — viscosity
R² > 0.99 · ±0.25 Pa·s
PF resin — phenol
mean absolute error ~0.03 p.p.
PF resin — formaldehyde
mean absolute error ~0.21 p.p.
Silicone in PA66 recyclate
calibrated 0.3–1.0 %
Isomer impurity, FMCG synthesis
tracked from ~8.5 % to trace level
Agrochemical product growth
~90 % → > 99 % across reaction stages

Full matrices, bands and methods: feasibility results.

±0.25 Pa·s

Prediction uncertainty for alkyd resin viscosity during synthesis — a rheological property, read from a Raman spectrum. This is the number that tends to end the argument about whether Raman only measures concentrations.

Bring us a measurement that classical methods cannot handle

Non-linear response, overlapping bands, a property that is not a concentration. Those are the cases the hybrid stack exists for.

Request a feasibility study Feasibility results