Abstract
Machine learning (ML) algorithms have become pivotal for merging multi-source data to generate high-quality precipitation datasets. However, several critical challenges in feature engineering remain underexplored, limiting both the accuracy and interpretability of ML-based methods. This study systematically investigates three overlooked aspects of feature engineering: data independence in gauge-calibrated datasets, the generalization of predictors across diverse regions, and the interaction between feature combinations and gauge density. We conducted controlled experiments in mainland China, focusing on spatial and temporal prediction scenarios using four ML models, including random forest, artificial neural network, convolutional neural network, and self-attention modules. The results reveal that gauge data leakage caused by data dependence problems can lead to overestimated merging performance, especially for spatial prediction scenario. Although high-quality precipitation datasets typically outperform others, their contribution varies regionally in ML applications, with other predictors in lower performance, such as infrared-based precipitation datasets, gaining increasing importance in areas like the Tibetan Plateau and Xinjiang. Moreover, our analysis shows that as gauge density decreases, simpler feature combinations excluding auxiliary variables outperform more complex ones. Overall, our findings provide actionable insights for refining feature engineering practices and developing robust and accurate precipitation datasets.
| Original language | English |
|---|---|
| Article number | 134185 |
| Journal | Journal of Hydrology |
| Volume | 663 |
| DOIs | |
| State | Published - Dec 2025 |
| Externally published | Yes |
Fingerprint
Dive into the research topics of 'Optimization of feature inputs in machine learning-based multi-source precipitation merging'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver