J4 ›› 2011, Vol. 38 ›› Issue (4): 32-37.doi: 10.3969/j.issn.1001-2400.2011.04.006

• 研究论文 • 上一篇    下一篇

异构无线网络中基于强化学习的频谱管理算法

张文柱;邵丽娜   

  1. (西安电子科技大学 综合业务网理论及关键技术国家重点实验室,陕西 西安  710071)
  • 收稿日期:2010-12-14 出版日期:2011-08-20 发布日期:2011-09-28
  • 通讯作者: 张文柱
  • 作者简介:张文柱(1970-),男,副教授,博士,E-mail: wzzhang1@mail.xidian.edu.cn.
  • 基金资助:

    国家杰出青年科学基金资助项目(60725105);国家重点基础研究发展计划(973计划)课题资助项目(2009CB320404);长江学者和创新团队发展计划资助项目(IRT0852);国家自然科学基金资助项目(61072068,60872045);中央高校基本科研业务费专项资助项目(JY10000901031)

Dynamic spectrum allocation algorithm for heterogeneous radio networks based on reinforcement learning

ZHANG Wenzhu;SHAO Lina   

  1. (State Key Lab. of Integrated Service Networks, Xidian Univ., Xi'an  710071, China)
  • Received:2010-12-14 Online:2011-08-20 Published:2011-09-28
  • Contact: ZHANG Wenzhu

摘要:

提出了一种基于归一化径向基函数的自适应启发评价强化学习算法,用于异构无线网络系统中自主的动态频谱分配.该算法利用归一化径向基函数自适应构建状态空间,加快学习速度;利用自适应启发评价机制减少不必要的探索,提高学习效率.通过与无线环境交互,算法学会为不同接入网内的各个会话动态分配合适的频段.仿真结果表明,在同等网络条件下,该算法能获取更好的频谱利用率和服务质量,性能优于确定性频谱分配策略和一般的动态频谱分配策略.

关键词: 异构无线网络, 动态频谱分配, 强化学习, 归一化径向基函数

Abstract:

An adaptive heuristic critic (AHC) Reinforcement Learning algorithm is presented for the dynamic spectrum allocation in an autonomously deciding mode in heterogeneous radio networks based on the normalized radial basis function (NRBF). The algorithm accelerates the learning speed by utilizing the NRBF when constructing the state space, and improves the learning efficiency by using the AHC scheme to reduce the unnecessary exploration. Through interactions with the radio environment, it learns to allocate the proper frequency band for each session in multiple radio access networks. Simulation results show that the proposed algorithm can lead to a better spectrum efficiency and quality of service compared with to the fixed frequency planning scheme or general dynamic spectrum allocation policy.

Key words: heterogeneous radio networks, dynamic spectrum allocation, reinforcement learning, normalized radial basis function

中图分类号: 

  • TP393