Sensor Fusion and Machine Learning for Real-world Emissions Monitoring and Climate-related Environmental Assessment in the United States: A Critical Review
Emmanuel Kiplagat Teigong *
College of Computing, Michigan Technological University, Houghton, USA.
Yeboah Mary Magdalene
University of Ghana Business School, Accra, Ghana.
*Author to whom correspondence should be addressed.
Abstract
Emissions monitoring in the United States is being reconstructed around dense, heterogeneous sensing and algorithmic inference. Low-cost electrochemical and optical sensors, mobile platforms, aircraft-mounted imaging spectrometers, point-in-space continuous monitors, and polar-orbiting and geostationary satellites now generate observations at scales and frequencies that regulatory reference networks were never designed to provide, and machine learning increasingly mediates the step from raw signal to reported emission or exposure. This review critically appraises what that reconstruction has demonstrated, what remains uncertain, and where confidence is currently misplaced. Evidence was assembled through registry-based bibliographic searching and verified source-by-source, with emphasis on studies conducted in United States settings and published from 1995 onwards. Four findings recur across otherwise separate literatures. First, reported model performance is dominated by co-location conditions rather than deployment conditions, and cross-site, cross-season and cross-instrument transfer degrades in ways that headline coefficients of determination conceal. Second, blind controlled-release testing has become the strongest evidentiary design available for methane sensing and has no established equivalent in ambient sensor calibration, producing a marked asymmetry in the quality of validation evidence between adjacent fields. Third, heavy-tailed and intermittent emission distributions break the sampling assumptions embedded in both survey design and supervised learning, so that platform-specific estimates diverge for structural rather than analytical reasons. Fourth, fusion across platforms with different detection thresholds is frequently treated as additive when it is in fact a partial-detection problem requiring explicit statistical treatment. Errors in fused products are spatially structured and correlate with monitor density, so that inferential accuracy is lowest where policy attention is greatest. Priorities identified include standardised blind evaluation for ambient sensing, detection-probability-aware fusion, distribution-shift-resilient calibration, and reporting practices that expose spatial error structure rather than aggregate it away.
Keywords: Sensor fusion, machine learning, methane emissions, low-cost air quality sensors, imaging spectrometry, emissions inventory verification, environmental monitoring policy