This study investigates how Raman peak features influence the accuracy of pollutant concentration quantification in Raman spectroscopy using feature-based machine learning. A Raman peak feature-based approach is applied to data acquired from laboratory Raman equipment for inorganic anion mixtures, with the objective of identifying processing strategies suitable for low-resolution Raman instruments. A dataset comprising nitrate, nitrite, and sulphate dissolved in water at varying concentrations was used to develop a pre-processing pipeline, train and validate models, and perform linear regression using peak- and area-based features extracted from analyte-specific spectral fingerprint regions. Signal downsampling was employed to simulate reduced spectral resolution and evaluate feature robustness under peak broadening conditions. Results indicate that peak-based regression is more sensitive to reduced peak resolution for nitrate and sulphate, whereas area-based regression produces more stable prediction errors across downsampling factors. Conversely, for nitrite, peak-based regression is less affected by downsampling, while area-based regression exhibits degraded performance. Area-based regression achieved mean absolute percentage errors below 7% and 16% for nitrate concentrations above 3833 mg/L and 1916 mg/L, respectively, and generally below 16% for sulphate concentrations above 1916 mg/L. For nitrite, peak-based regression yielded errors within 4%.

Raman Peaks Feature-Based Machine Learning for Raman Spectroscopy Quantification of Inorganic Pollutants

Luciani, Lorenzo;Galassi, Rossana
2026-01-01

Abstract

This study investigates how Raman peak features influence the accuracy of pollutant concentration quantification in Raman spectroscopy using feature-based machine learning. A Raman peak feature-based approach is applied to data acquired from laboratory Raman equipment for inorganic anion mixtures, with the objective of identifying processing strategies suitable for low-resolution Raman instruments. A dataset comprising nitrate, nitrite, and sulphate dissolved in water at varying concentrations was used to develop a pre-processing pipeline, train and validate models, and perform linear regression using peak- and area-based features extracted from analyte-specific spectral fingerprint regions. Signal downsampling was employed to simulate reduced spectral resolution and evaluate feature robustness under peak broadening conditions. Results indicate that peak-based regression is more sensitive to reduced peak resolution for nitrate and sulphate, whereas area-based regression produces more stable prediction errors across downsampling factors. Conversely, for nitrite, peak-based regression is less affected by downsampling, while area-based regression exhibits degraded performance. Area-based regression achieved mean absolute percentage errors below 7% and 16% for nitrate concentrations above 3833 mg/L and 1916 mg/L, respectively, and generally below 16% for sulphate concentrations above 1916 mg/L. For nitrite, peak-based regression yielded errors within 4%.
2026
Raman spectroscopy; machine learning; quantification; water analysis; nitrate; nitrite; sulphate
262
File in questo prodotto:
Non ci sono file associati a questo prodotto.

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11581/503504
 Attenzione

Attenzione! I dati visualizzati non sono stati sottoposti a validazione da parte dell'ateneo

Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
social impact