PriMera Scientific Engineering (ISSN: 2834-2550)

Research Article

Volume 9 Issue 4

LLM-Enhanced Pollution Source Attribution: A Multimodal Framework Integrating Vision Transformers with Semantic Feature Extraction

Husein Harun, Temitope Ruth Folorunso, Sodrat Oluwatoyin Fesomade and Ekhomen Ehimen-Ebitibituwa*

September 30, 2026

Abstract

Traditional pollution source attribution relies on expensive chemical speciation and handcrafted statistical models that cannot leverage the rich semantic information embedded in environmental monitoring metadata. We present EnhancedPollutionAttributorWithLLMs, a novel multimodal framework that integrates Large Language Models with advanced neural architectures for automated pollution source attribution. Our approach extracts semantic features from EPA Air Quality System site descriptions and combines them with spatiotemporal environmental data using rigorous geographic holdout validation. Experimental evaluation on 3,899 EPA monitoring records demonstrates that Vision Transformer with LLM-enhanced features achieves exceptional performance (R² = 0.9969, RMSE = 0.3116), significantly outperforming traditional approaches. Ablation studies reveal that LLM-derived semantic features contribute 1.2% improvement in R², while comprehensive interpretability analysis using SHAP provides actionable insights for environmental policy. This work establishes the first reproducible benchmark for LLM-enhanced environmental monitoring and demonstrates the potential for automated, large-scale pollution source attribution.

Keywords: pollution attribution; large language models; environmental monitoring; EPA AQS; semantic feature extraction; vision transformer; interpretable ML; neural networks

References

  1. Hopke PK. “Review of receptor modeling methods for source apportionment”. J Air Waste Manag Assoc. 66.3 (2016): 237-259.
  2. Belis CA., et al. “Critical review and meta-analysis of ambient particulate matter source apportionment using receptor models in Europe”. Atmos Environ. 69 (2013): 94-108.
  3. U.S. EPA. “EPA Positive Matrix Factorization (PMF) 5.0 Fundamentals and User Guide”. Office of Research and Development, Washington, DC, EPA/600/R-14/108 (2014).
  4. Paatero P and Tapper U. “Positive matrix factorization: A non-negative factor model with optimal utilization of error estimates of data values”. Environmetrics. 5.2 (1994): 111-126.
  5. Puri BK and Puri GK. “Air pollution and human health: A comprehensive review”. Environ Sci Policy. 123 (2021): 78-95.
  6. U.S. EPA. “Air Quality System (AQS)”. Environmental Protection Agency (2025).
  7. Brauer M., et al. “Ambient air pollution exposure estimation for the Global Burden of Disease 2013”. Environ Sci Technol. 50.1 (2016): 79-88.
  8. van Donkelaar A., et al. “Global estimates of fine particulate matter using a combined geophysical-statistical method with information from satellites”. Environ Sci Technol. 50.7 (2016): 3762-3772.
  9. Rohde RA., et al. “A new estimate of the average Earth surface land temperature spanning 1753 to 2011”. Geoinform Geostat. 1.1 (2013): 1-7.
  10. Brown T., et al. “Language models are few-shot learners”. Advances in Neural Information Processing Systems 33 (2020): 1877-1901.
  11. Devlin J., et al. “BERT: Pre-training of deep bidirectional transformers for language understanding”. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies 1 (2019): 4171-4186.
  12. Chen T and Guestrin C. “XGBoost: A scalable tree boosting system”. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (2016): 785-794.
  13. Breiman L. “Random forests”. Mach Learn. 45.1 (2001): 5-32.
  14. Zhang Y and Wang L. “Deep learning for air quality prediction: A review”. Atmos Environ. 195 (2018): 148-157.
  15. Li S., et al. “LSTM-based air quality prediction with feature selection”. Environ Sci Pollut Res. 26.8 (2019): 8235-8246.
  16. Guyon I and Elisseeff A. “An introduction to variable and feature selection”. J Mach Learn Res. 3 (2003): 1157-1182.
  17. Kuhn M and Johnson K. “Applied Predictive Modeling”. Springer, New York (2013).
  18. Hyndman RJ and Athanasopoulos G. “Forecasting: Principles and Practice”. 2nd ed. OTexts, Melbourne, Australia (2018).
  19. Rencher AC and Christensen WF. “Methods of Multivariate Analysis”. 3rd ed. John Wiley & Sons, Hoboken, NJ (2012).
  20. Vaswani A., et al. “Attention is all you need”. Advances in Neural Information Processing Systems 30 (2017): 5998-6008.
  21. Lundberg SM and Lee SI. “A unified approach to interpreting model predictions”. Advances in Neural Information Processing Systems 30 (2017): 4765-4774.
  22. Liu H., et al. “Vision transformer for environmental monitoring: A comprehensive survey”. Remote Sens. 14.18 (2022): 4625.
  23. Hopke PK. “Recent developments in receptor modeling”. J Chemom. 17.5 (2003): 255-265.
  24. Henry RC. “Multivariate receptor modeling by N-dimensional edge detection”. Chemom Intell Lab Syst. 65.2 (2003): 179-189.
  25. Watson JG., et al. “The effective variance weighing for least squares calculations applied to the mass balance receptor model”. Atmos Environ. 18.7 (1984): 1347-1355.
  26. Hopke PK. “A guide to positive matrix factorization”. EPA Office of Research and Development, Washington, DC, EPA/600/R-08/108 (2008).
  27. Cohen DD., et al. “Quantitative PM2.5 source apportionment using ion beam analysis”. Nucl Instrum Methods Phys Res Sect B. 213 (2004): 507-513.
  28. Wen Q., et al. “Time series data augmentation for deep learning: A survey”. Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence (2021): 4653-4660.
  29. Zamani Joharestani M., et al. “PM2.5 prediction based on random forest, XGBoost, and deep learning using multisource remote sensing data”. Atmosphere. 10.7 (2019): 373.
  30. LeCun Y., et al. “Deep learning”. Nature. 521.7553 (2015): 436-444.
  31. Goodfellow I, Bengio Y and Courville A. “Deep Learning”. MIT Press, Cambridge, MA (2016).
  32. Cortez P and Rio M. “Time series forecasting using neural networks”. Proceedings of the International Conference on Computational Intelligence for Modelling, Control and Automation 2 (2005): 1-6.
  33. Dosovitskiy A., et al. “An image is worth 16x16 words: Transformers for image recognition at scale”. International Conference on Learning Representations (2021).
  34. Touvron H., et al. “Training data-efficient image transformers & distillation through attention”. Proceedings of the 38th International Conference on Machine Learning, PMLR 139 (2021): 10347-10357.
  35. Huang X., et al. “TabTransformer: Tabular data modeling using contextual embeddings”. arXiv preprint arXiv:2012.06678 (2020).
  36. Gorishniy Y., et al. “Revisiting deep learning models for tabular data”. Advances in Neural Information Processing Systems 34 (2021): 18932-18943.
  37. He K., et al. “Deep residual learning for image recognition”. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (2016): 770-778.
  38. Hu J, Shen L and Sun G. “Squeeze-and-excitation networks”. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (2018): 7132-7141.
  39. Liu Z., et al. “Swin transformer: Hierarchical vision transformer using shifted windows”. Proceedings of the IEEE/CVF International Conference on Computer Vision (2021): 10012-10022.
  40. OpenAI. “GPT-4 Technical Report”. arXiv preprint arXiv:2303.08774 (2023).
  41. Radford A., et al. “Language models are unsupervised multitask learners”. OpenAI Blog 1.8 (2019): 9.
  42. Liu Y., et al. “RoBERTa: A robustly optimized BERT pretraining approach”. arXiv preprint arXiv:1907.11692 (2019).
  43. Mikolov T., et al. “Efficient estimation of word representations in vector space”. arXiv preprint arXiv:1301.3781 (2013).
  44. Bojanowski P., et al. “Enriching word vectors with subword information”. Trans Assoc Comput Linguist. 5 (2017): 135-146.
  45. Pennington J, Socher R and Manning CD. “GloVe: Global vectors for word representation”. Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP) (2014): 1532-1543.
  46. U.S. EPA. “Air Quality System Data Mart”. Environmental Protection Agency (2025).
  47. Liu FT, Ting KM and Zhou ZH. “Isolation forest”. Proceedings of the 2008 Eighth IEEE International Conference on Data Mining (2008): 413-422.
  48. Chandola V, Banerjee A and Kumar V. “Anomaly detection: A survey”. ACM Comput Surv. 41.3 (2009): 1-58.
  49. U.S. EPA. “Air Quality System Database Documentation”. Environmental Protection Agency (2025).
  50. Brauer M., et al. “Air pollution from traffic and the development of respiratory infections and asthmatic and allergic symptoms in children”. Am J Respir Crit Care Med. 166.8 (2002): 1092-1098.
  51. Brown T., et al. “Language models are few-shot learners”. Advances in Neural Information Processing Systems 33 (2020): 1877-1901.
  52. Hopke PK. “Review of receptor modeling methods for source apportionment”. J Air Waste Manag Assoc. 66.3 (2016): 237-259.
  53. Belis CA., et al. “Critical review and meta-analysis of ambient particulate matter source apportionment using receptor models in Europe”. Atmos Environ. 69 (2013): 94-108.
  54. Devlin J., et al. “BERT: Pre-training of deep bidirectional transformers for language understanding”. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies 1 (2019): 4171-4186.
  55. van Donkelaar A., et al. “Global estimates of fine particulate matter using a combined geophysical-statistical method with information from satellites”. Environ Sci Technol. 50.7 (2016): 3762-3772.
  56. Chen T and Guestrin C. “XGBoost: A scalable tree boosting system”. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (2016): 785-794.
  57. van Donkelaar A., et al. “Global estimates of ambient fine particulate matter concentrations from satellite-based aerosol optical depth: Development and application”. Environ Health Perspect. 118.6 (2010): 847-855.
  58. Brauer M., et al. “Ambient air pollution exposure estimation for the Global Burden of Disease 2013”. Environ Sci Technol. 50.1 (2016): 79-88.
  59. Hyndman RJ and Athanasopoulos G. “Forecasting: Principles and Practice”. 2nd ed. OTexts, Melbourne, Australia (2018).
  60. Cortez P and Rio M. “Time series forecasting using neural networks”. Proceedings of the International Conference on Computational Intelligence for Modelling, Control and Automation 2 (2005): 1-6.
  61. U.S. EPA. “Air Quality System Database Documentation”. Environmental Protection Agency (2025).
  62. Guyon I and Elisseeff A. “An introduction to variable and feature selection”. J Mach Learn Res. 3 (2003): 1157-1182.
  63. Kuhn M and Johnson K. “Applied Predictive Modeling”. Springer, New York (2013).
  64. Brown T., et al. “Language models are few-shot learners”. Advances in Neural Information Processing Systems 33 (2020): 1877-1901.
  65. OpenAI. “GPT-4 Technical Report”. arXiv preprint arXiv:2303.08774 (2023).
  66. Liu Y., et al. “RoBERTa: A robustly optimized BERT pretraining approach”. arXiv preprint arXiv:1907.11692 (2019).
  67. Mikolov T., et al. “Efficient estimation of word representations in vector space”. arXiv preprint arXiv:1301.3781 (2013).
  68. Vaswani A., et al. “Attention is all you need”. Advances in Neural Information Processing Systems 30 (2017): 5998-6008.
  69. Dosovitskiy A., et al. “An image is worth 16x16 words: Transformers for image recognition at scale”. International Conference on Learning Representations (2021).
  70. He K., et al. “Deep residual learning for image recognition”. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (2016): 770-778.
  71. Hu J, Shen L and Sun G. “Squeeze-and-excitation networks”. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (2018): 7132-7141.
  72. LeCun Y, Bengio Y and Hinton G. “Deep learning”. Nature. 521.7553 (2015): 436-444.
  73. Goodfellow I, Bengio Y and Courville A. “Deep Learning”. MIT Press, Cambridge, MA (2016).
  74. Breiman L. “Random forests”. Mach Learn. 45.1 (2001): 5-32.
  75. Rencher AC and Christensen WF. “Methods of Multivariate Analysis”. 3rd ed. John Wiley & Sons, Hoboken, NJ (2012).
  76. Hyndman RJ and Athanasopoulos G. “Forecasting: Principles and Practice”. 2nd ed. OTexts, Melbourne, Australia (2018).
  77. Kuhn M and Johnson K. “Applied Predictive Modeling”. Springer, New York (2013).
  78. Goodfellow I, Bengio Y and Courville A. “Deep Learning”. MIT Press, Cambridge, MA (2016).
  79. Chen T and Guestrin C. “XGBoost: A scalable tree boosting system”. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (2016): 785-794.
  80. Hyndman RJ and Athanasopoulos G. “Forecasting: Principles and Practice”. 2nd ed. OTexts, Melbourne, Australia (2018).
  81. Cortez P and Rio M. “Time series forecasting using neural networks”. Proceedings of the International Conference on Computational Intelligence for Modelling, Control and Automation 2 (2005): 1-6.
  82. U.S. EPA. “Air Quality System Data Mart”. Environmental Protection Agency (2025).
  83. Brown T., et al. “Language models are few-shot learners”. Advances in Neural Information Processing Systems 33 (2020): 1877-1901.
  84. Guyon I and Elisseeff A. “An introduction to variable and feature selection”. J Mach Learn Res. 3 (2003): 1157-1182.
  85. Vaswani A., et al. “Attention is all you need”. Advances in Neural Information Processing Systems 30 (2017): 5998-6008.
  86. Chen T and Guestrin C. “XGBoost: A scalable tree boosting system”. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (2016): 785-794.
  87. Hopke PK. “Review of receptor modeling methods for source apportionment”. J Air Waste Manag Assoc. 66.3 (2016): 237-259.
  88. Lundberg SM and Lee SI. “A unified approach to interpreting model predictions”. Advances in Neural Information Processing Systems 30 (2017): 4765-4774.
  89. Guyon I and Elisseeff A. “An introduction to variable and feature selection”. J Mach Learn Res. 3 (2003): 1157-1182.
  90. Dosovitskiy A., et al. “An image is worth 16x16 words: Transformers for image recognition at scale”. International Conference on Learning Representations (2021).
  91. Vaswani A., et al. “Attention is all you need”. Advances in Neural Information Processing Systems 30 (2017): 5998-6008.
  92. Liu H., et al. “Vision transformer for environmental monitoring: A comprehensive survey”. Remote Sens. 14.18 (2022): 4625.
  93. Liu Z., et al. “Swin transformer: Hierarchical vision transformer using shifted windows”. Proceedings of the IEEE/CVF International Conference on Computer Vision (2021): 10012-10022.
  94. Zhang Y and Wang L. “Deep learning for air quality prediction: A review”. Atmos Environ. 195 (2018): 148-157.
  95. Brown T., et al. “Language models are few-shot learners”. Advances in Neural Information Processing Systems 33 (2020): 1877-1901.
  96. Devlin J., et al. “BERT: Pre-training of deep bidirectional transformers for language understanding”. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies 1 (2019): 4171-4186.
  97. Hopke PK. “Review of receptor modeling methods for source apportionment”. J Air Waste Manag Assoc. 66.3 (2016): 237-259.
  98. van Donkelaar A., et al. “Global estimates of fine particulate matter using a combined geophysical-statistical method with information from satellites”. Environ Sci Technol. 50.7 (2016): 3762-3772.
  99. Belis CA., et al. “Critical review and meta-analysis of ambient particulate matter source apportionment using receptor models in Europe”. Atmos Environ. 69 (2013): 94-108.
  100. Chen T and Guestrin C. “XGBoost: A scalable tree boosting system”. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (2016): 785-794.
  101. Breiman L. “Random forests”. Mach Learn. 45.1 (2001): 5-32.
  102. Gorishniy Y., et al. “Revisiting deep learning models for tabular data”. Advances in Neural Information Processing Systems 34 (2021): 18932-18943.
  103. Lundberg SM and Lee SI. “A unified approach to interpreting model predictions”. Advances in Neural Information Processing Systems 30 (2017): 4765-4774.
  104. Guyon I and Elisseeff A. “An introduction to variable and feature selection”. J Mach Learn Res. 3 (2003): 1157-1182.
  105. Zhang Y and Wang L. “Deep learning for air quality prediction: A review”. Atmos Environ. 195 (2018): 148-157.
  106. Brauer M., et al. “Ambient air pollution exposure estimation for the Global Burden of Disease 2013”. Environ Sci Technol. 50.1 (2016): 79-88.
  107. Hyndman RJ and Athanasopoulos G. “Forecasting: Principles and Practice”. 2nd ed. OTexts, Melbourne, Australia (2018).
  108. Brown T., et al. “Language models are few-shot learners”. Advances in Neural Information Processing Systems 33 (2020): 1877-1901.
  109. Devlin J., et al. “BERT: Pre-training of deep bidirectional transformers for language understanding”. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies 1 (2019): 4171-4186.
  110. LeCun Y, Bengio Y and Hinton G. “Deep learning”. Nature. 521.7553 (2015): 436-444.
  111. Li S., et al. “LSTM-based air quality prediction with feature selection”. Environ Sci Pollut Res. 26.8 (2019): 8235-8246.