2 min readfrom Frontiers in Marine Science | New and Recent Articles

Data-driven modelling of coastal water quality dynamics

Our take

Long-term coastal monitoring offers a unique opportunity to assess water quality predictability, yet existing machine learning studies often lack comprehensive scope. Our analysis of 37 years of data from 94 stations across four Hong Kong Bay systems reveals significant regional variations in predictability and key predictors. Tree-based models demonstrated robust performance, particularly for temperature and salinity, while chlorophyll-a proved consistently challenging to forecast.
Data-driven modelling of coastal water quality dynamics

The increasing availability of long-term coastal monitoring data is revolutionizing our capacity to understand and predict water quality dynamics, a critical aspect of ocean health and human well-being. This recent study, analyzing 37 years of data from Hong Kong’s bay systems, exemplifies this trend, moving beyond the typical focus on individual stations or short timeframes. The researchers’ rigorous application of machine learning models, including Random Forest and Gradient Boosting, to a substantial dataset—94 stations across four distinct bay systems—provides a valuable benchmark for empirical water-quality predictability. This work builds upon similar efforts exploring the use of machine learning in environmental science; for instance, [Machine learning predictions for microbial eukaryotic plankton: implications from unevenly structured data] demonstrates the challenges and opportunities in predicting plankton biodiversity from complex datasets. Similarly, the need for robust data-driven approaches to understand and mitigate pollution is highlighted in [Editorial: Strategies for remediating marine and coastal pollution towards a sustainable development], underscoring the importance of research like this.

The study’s findings are particularly noteworthy for illuminating the regional variations in water quality predictability and the dominant influencing factors. The observation that salinity, temperature, and dissolved oxygen are key predictors in the more mixing-dominated Central, Eastern, and Southern bays, while nitrate and ammonia become more important in the Pearl River-influenced Western Bays, reinforces the crucial role of hydrography in shaping coastal ecosystems. The detection of a declining pH trend in urbanized areas, contrasting with the relative stability of Eastern Bays, is a concerning yet vital indicator of anthropogenic impact. This highlights the necessity of localized, data-driven assessments to inform targeted mitigation strategies. The fact that chlorophyll-a, a critical indicator of primary productivity, remained among the least predictable variables underscores the complexity of biological processes and the need for more sophisticated modeling approaches. The authors' recommendation to incorporate lagged hydrometeorological, watershed, and discharge variables represents a logical and necessary step towards improved forecasting accuracy.

Furthermore, the application of SHAP analysis to identify the relative importance of different predictors adds a layer of interpretability often lacking in machine learning studies. Being able to understand *why* a model makes a particular prediction is just as crucial as the prediction itself, allowing for a more nuanced understanding of the underlying ecological processes. This approach aligns with the broader trend towards explainable AI (XAI) within the environmental sciences, moving beyond “black box” models to provide actionable insights for decision-makers. The focus on empirical predictability, validated through a rigorous chronological framework, lends significant weight to the findings, making them directly applicable to coastal management and policy development. The ability to calibrate and validate models against such a long-term dataset is a powerful tool for assessing the effectiveness of implemented interventions.

Looking ahead, the integration of increased spatial resolution data, coupled with improved understanding of the complex interplay between physical, chemical, and biological processes, will be essential for refining coastal water quality models. The insights gained from this study highlight the need for a holistic, integrated data ecosystem—a concept we champion—to deliver robust "ocean intelligence" for informed decision-making. As we continue to grapple with the escalating impacts of climate change and increasing human pressures on coastal environments, how can we best leverage these longitudinal datasets and advanced modeling techniques to proactively safeguard the health and resilience of our oceans?

Long-term coastal monitoring networks provide the opportunity to evaluate how water-quality predictability varies across contrasting coastal bays. However, most machine learning studies focused on individual stations, optically active variables or short validation periods. We analyze 37 years (1986-2023) of monthly in-situ data from 94 stations across four Hong Kong Bay systems Central Harbours, Eastern Bays, Southern Bays, and Western Bays to test whether empirical predictability and highest-ranked water-quality predictors vary systematically with regional hydrography. Eight parameters: ammonia, nitrate, and nitrite nitrogen; dissolved oxygen; pH; salinity; temperature; and chlorophyll-a were assessed. Five machine-learning models Random Forest (RF), Gradient Boosting (GB), XGBoost (XGB), Support Vector Machine (SVM), and Artificial Neural Network (ANN) were evaluated under chronological 70/30 split with Bayesian-optimised hyperparameters. Tree-based ensembles produced the most stable performance overall, with strong performance for temperature and salinity, moderate-to-variable performance for selected nutrient species, and weak transferability for dissolved oxygen and pH. Chlorophyll-a remained among the least predictable variables across most bay systems under the chronological validation framework. SHAP analysis revealed regional transition in predictor structure: salinity, temperature and dissolved oxygen dominated mixing-structured Central, Eastern and Southern systems, while nitrate and ammonia became more important in Pearl River-influenced Western Bays. Long-term analyses indicated a declining pH trend in the urbanised Central Harbours and Southern and Western Bays, whereas Eastern Bays remained comparatively stable. These findings provide regional benchmark for empirical water-quality predictability under chronological validation and identify the need to incorporate lagged hydrometeorological, watershed and discharge variables for future forecasting applications.

Read on the original site

Open the publisher's page for the full experience

View original article