基于深度神经网络的水稻产量预测模型构建与变量解释分析

    Construction and variable interpretation analysis of a rice yield prediction model based on deep neural networks

    • 摘要:
      目的 探究不同建模方法对水稻农艺性状与产量关系的刻画能力,构建基于深度神经网络(Deep neural network, DNN)的水稻产量预测模型,评估深度学习方法在水稻产量预测与高产性状筛选中的应用潜力。
      方法 以442个水稻杂交组合的12项农艺性状数据为基础,在标准化处理与缺失值修复后构建预测模型。设置DNN、线性回归(Linear regression)、随机森林(Random forest)与极端梯度提升(XGBoost)4种算法进行建模比较。利用均方根误差(Mean squared error,MSE)、平均绝对误差(Mean absolute error,MAE)和决定系数(Coefficient of determination,R2)评估模型性能,并基于SHAP(SHapley Additive exPlanations)方法解析模型特征贡献度,同时结合Pearson相关分析与差异显著性检验验证解释结果的统计可靠性。
      结果 DNN模型在测试集上的R2为0.947 6,MSE为1 094.08,MAE为25.96;虽然绝对精度略低于线性回归模型(R2=0.956 7),但交叉验证显示DNN的预测波动更小,表现出更优的稳健性。SHAP分析结果显示,每穗实粒数、有效穗数、千粒质量和单株穗质量是影响产量预测的核心性状,其贡献度排序与相关性检验结果高度一致,符合水稻“有效穗数—粒数—千粒质量”三要素的产量构成规律。
      结论 相较于传统统计模型,DNN模型展现了更强的抗干扰能力与预测稳健性,DNN结合SHAP的分析框架在可视化非线性特征贡献与阈值效应方面具有独特优势,为水稻产量预测与高产性状筛选提供了具备深度解释性的新工具。

       

      Abstract:
      Objective To investigate the ability of different modeling methods to characterize the relationship between rice agronomic traits and yield, construct a rice yield prediction model based on deep neural networks (DNN), and evaluate the application potential of deep learning methods for rice yield prediction and high-yield trait screening.
      Method Based on the agronomic trait data of 12 traits from 442 rice hybrid combinations, a predictive model was constructed after standardization and of missing value imputation. Four algorithms were employed for comparative modeling: DNN, linear regression, random forest, and extreme gradient boosting (XGBoost). Model performance was evaluated using mean squared error (MSE), mean absolute error (MAE), and the coefficient of determination (R2). The feature contributions of the models were analyzed using the SHapley Additive exPlanations (SHAP) method, while Pearson correlation analysis and difference significance tests were conducted to validate the statistical reliability of the interpretations.
      Result The DNN model achieved the best performance on the test set with an R2 of 0.947 6, MSE of 1 094.08, and MAE of 25.96. Although its absolute accuracy was slightly lower than that of the linear regression model (R2=0.956 7), cross-validation showed that the DNN exhibited smaller prediction fluctuations and demonstrated superior robustness. SHAP analysis revealed that filled grains per panicle, effective panicle number, 1 000-grain weight and panicle weight per plant were the core traits affecting yield prediction, with their contribution rankings highly consistent with correlation test results, consistent with the yield formation principle that rice yield was determined by “effective panicle number-grain number-1 000-grain weight” triad.
      Conclusion Compared to traditional statistical models, the DNN model demonstrates superior resistance to interference and predictive robustness. The analysis framework combining DNN and SHAP offers unique advantages in visualizing nonlinear feature contributions and threshold effects, providing a new deeply interpretable tool for rice yield prediction and high-yield trait screening.

       

    /

    返回文章
    返回