Abstract:
Improving the calculation accuracy of watershed hydrological similarity is of great significance for hydrological simulation in ungauged and data-insufficient regions. This study constructed a comprehensive calculation model of watershed hydrological similarity that integrates static characteristics and dynamic time series. A pseudo-label-based machine learning method was adopted to improve the traditional static identification framework for similar watersheds, in which the Shapley Additive Explanations (SHAP) was used to reduce the dimension of physical attribute indicators and quantify indicator weights. Meanwhile, a dynamic analysis dimension of rainfall time series similarity was introduced, and the Numerical Symbolic and Shape Feature Measure (NSM) distance was employed to measure dynamic similarity. Finally, the Copula function was used to dynamically assign weights to heterogeneous similarities under the two dimensions, realizing the synergistic fusion of dynamic and static hydrological similarities. Taking 82 medium and small-scale sub-basins in the Fenhe River Basin of Shanxi Province as the study area, the comprehensive similarity of watershed pairs in the upper, middle and lower reaches was calculated respectively. On this basis, the Soil and Water Assessment Tool (SWAT) model was constructed to simulate annual runoff of each watershed, and the watershed runoff similarity was calculated to verify and analyze the identification results of similar watersheds. The results demonstrated that: 1) The core drivers of hydrological similarity identified by the pseudo-label-based machine learning feature selection method are highly consistent with the climatic conditions, underlying surface patterns and runoff mechanisms of each sub-region. Among 26 indicators, the key factors for the upper, middle and lower Fenhe River basin are mean annual precipitation, terrain relief and grassland coverage (cumulative weight 0.66); sand content, mean annual precipitation and watershed area (0.56); and elevation coefficient of variation, sand content and impervious surface coverage (0.42), respectively. 2) Static dominance and balanced matching of static-dynamic dual dimensions for basin pairs are the core drivers of high hydrological similarity. The cumulative proportion of static-dominant and static-dynamic balanced basin pairs in the high-similarity group reaches 81%, which is significantly higher than 39% in the low-similarity group. Meanwhile, dynamic dominance of basin pairs is the key cause of low hydrological similarity. The proportion of dynamic-dominant basins in the low-similarity group is 61%, remarkably higher than 19% in the high-similarity group. 3) Using a difference threshold of 0.05 against the SWAT-derived runoff similarity benchmark, the accuracy rates of the proposed comprehensive similarity model for the upper, middle, and lower reaches were 54%, 84%, and 69%, respectively. These rates substantially exceeded those of the static similarity model alone, which were 31%, 55%, and 29%. The overall accuracy of the comprehensive model reached 71%, representing an 82% relative improvement over the 39% overall accuracy of the static model. In conclusion, the developed comprehensive similarity fusion method effectively addresses critical bottlenecks in traditional assessment approaches, such as strong subjectivity in indicator selection and the isolation of dynamic and static information. It provides a more objective means to capture multi-dimensional similarities between watersheds, offering a new and effective technical pathway for improving the accuracy of hydrological similar-watershed identification in ungauged or data-scarce regions.