Machine learning has improved the ability of Raman and infrared spectroscopy to classify biological samples, but inconsistent analytical practices remain a larger barrier to clinical adoption than model performance, according to a new review.
Published in the Microchemical Journal, the review examines the complete analytical pipeline behind machine learning-enabled vibrational spectroscopy, from sample preparation and spectral acquisition to preprocessing, modeling, interpretation, and validation.
FTIR, Raman, and surface-enhanced Raman spectroscopy can provide label-free information about proteins, lipids, nucleic acids, and metabolites. Researchers have investigated their use in applications including tumor classification, pathogen identification, biofluid screening, and treatment monitoring.
Deep learning can identify nonlinear patterns within the resulting high-dimensional datasets and has sometimes outperformed conventional chemometric methods. However, the authors warn that high classification accuracy on a small, internally validated dataset does not necessarily indicate that a model will work with different patients, instruments, or clinical sites.
Preprocessing is one source of uncertainty. Baseline correction, smoothing, normalization, spectral alignment, and derivative transformations can all alter the information presented to a model. Different combinations can produce substantially different classification results from the same raw spectra, yet preprocessing decisions are not always fully reported or justified.
The authors also identify small, institution-specific datasets as a recurring weakness. Complex neural networks trained on limited datasets are particularly vulnerable to overfitting, while differences in sample handling, patient populations, acquisition parameters, and instrument performance may confound apparent disease-related signals.
Moving between instruments introduces further problems. Portable Raman and infrared systems generally offer lower resolution and signal-to-noise ratios than benchtop instruments and may be more susceptible to ambient conditions. Models developed using laboratory systems may therefore require recalibration, transfer learning, or domain adaptation before deployment at the point of care.
The review proposes hybrid approaches that combine machine learning with chemically meaningful inputs, such as established vibrational band assignments or chemometric latent variables. Such models could retain nonlinear modeling capabilities while making predictions easier to connect with underlying biochemistry.
Other priorities include standardized reporting of spectral ranges, resolution, accumulation numbers, baseline correction, and normalization; openly available reference datasets; and external validation using geographically or institutionally independent patient cohorts. Validation should also test performance under inter-instrument variation, longitudinal drift, and realistic clinical conditions.
The authors argue that the most valuable applications may not be those reporting the highest accuracy, but those addressing specific unmet needs. Potential examples include rapid antimicrobial resistance testing, minimally invasive biofluid screening, and real-time tumor-margin assessment.
Ultimately, the review concludes that clinical progress will depend less on increasingly complex algorithms than on reproducibility, transparency, and evidence that the complete analytical system performs reliably outside the laboratory where it was developed.
