欢迎访问《昆明冶金职业大学学报》,

昆明冶金职业大学学报 ›› 2026, Vol. 42 ›› Issue (3): 81-.DOI: 10.3969/j.issn.1009-0479.2026.03.012

• 电子信息技术 • 上一篇    下一篇

基于机器学习插补的偏正态空间自回归模型贝叶斯估计研究

  

  1. (昆明冶金职业大学a.通识与素质教育学院;b.体育与康养学院,云南昆明650033)
  • 出版日期:2026-08-31 发布日期:2026-09-04
  • 作者简介:杨 燕(1997-),女,云南昆明人,助教,理学硕士,主要从事应用统计、大数据分析与高等职业教育
    研究。

Research on Bayesian Estimation of Skew‑Normal Spatial Autoregressive Models Based on Machine Learning Interpolation

  1. (a. Faculty of General and Quality Education; b. Faculty of Physical Education and Health, Kunming Metallurgy University, Kunming 650033, China)

  • Online:2026-08-31 Published:2026-09-04

摘要:

在数据收集过程中,数据隐私、人为疏忽、流程缺陷等因素的影响常常造成数据缺失,不仅会增加分析复杂度,还会造成系统偏差甚至错误的统计推断,故对缺失数据的处理至关重要。现实中的系统内在机制决定大多数真实数据分布呈现偏态分布而非严格正态分布。本文针对随机缺失偏正态空间数据,探究了样本量与缺失率变化对随机插补、均值插补、支持向量机、决策树、随机森林以及BP神经网络6种插补方法精度的影响,利用MSE和MAE进行量化评价;针对插补后的数据集,建立偏正态空间自回归模型,利用MCMC算法对其进行贝叶斯参数估计。模拟研究和实例分析结果表明:支持向量机和随机森林在拟合缺失空间偏态数据具有更强的适应能力和稳健表现,其对应的贝叶斯参数估计值较为准确。

关键词: 空间自回归模型, 偏正态, 缺失数据, 机器学习

Abstract:

 In the process of data collection, data are often missing due to the influence of various factors such as data privacy, human negligence, process defects, etc., which can increase the complexity of the analysis, resulting in systematic bias or even erroneous statistical inference. The treatment of missing data is therefore essential. The intrinsic mechanisms of the system dictate that most real data distributions exhibit skewed distributions rather than strictly normal distributions. In this paper, for random‑missing Skew‑Normal spatial data, the effects of sample size and missing rate variation on the accuracy of six interpolation methods, namely, random interpolation, mean value interpolation, support vector machine, decision tree, random forest, and BP neural network, are explored and quantitatively evaluated using MSE and MAE. For the interpolated dataset, a Skew‑Normal Spatial Autoregressive Model is built to estimate the Bayesian parameters using the MCMC algorithm. The results of simulation studies and example analysis show that support vector machines and random forests have better fitting ability and robustness to missing spatial skewed data, and the resulting Bayesian parameter estimates are more accurate.

Key words: spatial autoregressive modeling, partial normality, missing data, machine learning

中图分类号: