Predictive Modeling of Urban Air Quality Using Environmental Data
Keywords:
PM₂.₅ prediction, machine learning, Extra Trees, urban air quality, environmental monitoringAbstract
Precise forecasting of fine particulate matter (PM₂.₅) is crucial for the proper management of urban air quality and environmental decision-making. The purpose of this research is to build a machine learning model to predict the concentration of PM₂.₅ by applying atmospheric pollution, meteorological data, temporal characteristics, and monitoring station data, using the Beijing Multi-Site Air Quality dataset. Three coefficient of determination (R2), root mean square error (RMSE), mean absolute error (MAE), and mean squared error (MSE) were developed and evaluated for three ensemble learning algorithms: Random Forest, Extra Trees, and Extreme Gradient Boosting (XGBoost). Extra Trees model had the highest predictive accuracy with an R2 of 0.9503 and the smallest prediction errors among the evaluated models, showing its best generalization ability in highly complex environmental data. The feature importance analysis showed that PM10, CO, and NO₂ were the major features as they explain most of the variation in the data, meteorological factors were found to be boosting the robustness of the model by accounting for seasonal and atmospheric variation. The proposed system provides an efficient and reliable solution for intelligent air quality forecast, which aids the pollution control strategies, early warning system, and data-driven urban environment management.
