电子科技 ›› 2025, Vol. 38 ›› Issue (3): 7-15.doi: 10.16180/j.cnki.issn1007-7820.2025.03.002

• • 上一篇    下一篇

一种非线性表征的概率潜在因子张量模型

董佳英1, 宋燕2(), 李明3   

  1. 1.上海理工大学 理学院,上海 200093
    2.上海理工大学 光电信息与计算机工程学院,上海 200093
    3.江苏海洋大学 计算机工程学院,江苏 连云港 222005
  • 收稿日期:2023-08-28 修回日期:2023-09-11 出版日期:2025-03-15 发布日期:2025-03-11
  • 通讯作者: 宋燕(1979-),女,E-mail:sonya@usst.edu.cn,博士,教授。研究方向:模式识别、数据分析和预测控制等。
  • 作者简介:董佳英(1999-),女,硕士研究生。研究方向:数据挖掘。
  • 基金资助:
    国家自然科学基金(62073223);上海市自然科学基金(22ZR1443400)

A Nonlinear Representation-Based Probabilistic Latent Factorization Tensor Model

DONG Jiaying1, SONG Yan2(), LI Ming3   

  1. 1. College of Science,University of Shanghai for Science and Technology,Shanghai 200093,China
    2. School of Optical-Electrical and Computer Engineering,University of Shanghai for Science and Technology,Shanghai 200093,China
    3. School of Computer Engineering,Jiangsu Ocean University,Lianyungang 222005,China
  • Received:2023-08-28 Revised:2023-09-11 Online:2025-03-15 Published:2025-03-11
  • Supported by:
    National Natural Science Foundation of China(62073223);Natural Science Foundation of Shanghai(22ZR1443400)

摘要:

针对具有极度稀疏和不平衡的非负不完整数据的填补问题,文中提出了一种非线性表征的概率潜在因子张量模型。通过合理假设数据的概率分布作为先验信息,缓解了数据的稀疏性。利用非线性映射实现对数据中每一非负元素的非线性表征,提高了模型的表征能力。考虑到数据的不平衡性,对传统正则化项添加基于实例频率的权重,增加了正则化项的有效性和针对性。实验结果表明,所提模型在补全精度和时间成本方面较现有模型具有明显提升。

关键词: 非线性表征, 概率潜在因子张量模型, 实例频率, 非线性映射, 数据稀疏性, CP分解, 不平衡分布, 正则项

Abstract:

In view of the filling problem of non-negative incomplete data with extremely sparse and unbalanced data, a probabilistic potential factor tensor model is proposed. The data sparsity is mitigated by reasonably assuming the probability distribution of the data as a priori information. Nonlinear mapping is used to realize nonlinear characterization of each non-negative element in the data and improve the characterization ability of the model. Considering the unbalance of data, the weights based on instance frequency are added to the traditional regularization terms to increase the effectiveness and pertinence of regularization terms. The experimental results show that the proposed model has obvious improvement over the existing model in terms of completion accuracy and time cost.

Key words: nonlinear representation, probabilistic factorization tensor model, frequency of known entries, nonlinear mapping, data sparsity, CP decomposition, unbalanced distribution, regular term

中图分类号: 

  • TP393