ENGINEERING Information Technology & Electronic Engineering  2026 Vol.27 No.8 P.1-12

http://doi.org/10.1631/ENG.ITEE.2026.0047


An adaptive meta-reinforcement learning scheme for maritime dynamic spectrum access and service scheduling


Author(s):  Zhongyang MAO, Jiahuan GENG, Faping LU, Wenbiao TIAN, Jiafang KANG, Yaozong PAN

Affiliation(s):  1. Naval Aviation University, Yantai 264001, China more

Corresponding email(s):   echo_HAU@163.com

Key Words:  Dynamic spectrum access, Adaptive meta-reinforcement learning, Service scheduling, Maritime communication


Zhongyang MAO, Jiahuan GENG, Faping LU, Wenbiao TIAN, Jiafang KANG, Yaozong PAN. An adaptive meta-reinforcement learning scheme for maritime dynamic spectrum access and service scheduling[J]. Journal of Zhejiang University Science C, 2026, 27(8): 1-12.

@article{title="An adaptive meta-reinforcement learning scheme for maritime dynamic spectrum access and service scheduling",
author="Zhongyang MAO, Jiahuan GENG, Faping LU, Wenbiao TIAN, Jiafang KANG, Yaozong PAN",
journal="Journal of Zhejiang University Science C",
volume="27",
number="8",
pages="1-12",
year="2026",
publisher="Zhejiang University Press & Springer",
doi="10.1631/ENG.ITEE.2026.0047"
}

%0 Journal Article
%T An adaptive meta-reinforcement learning scheme for maritime dynamic spectrum access and service scheduling
%A Zhongyang MAO
%A Jiahuan GENG
%A Faping LU
%A Wenbiao TIAN
%A Jiafang KANG
%A Yaozong PAN
%J Frontiers of Information Technology & Electronic Engineering
%V 27
%N 8
%P 1-12
%@ 1869-1951
%D 2026
%I Zhejiang University Press & Springer
%DOI 10.1631/ENG.ITEE.2026.0047

TY - JOUR
T1 - An adaptive meta-reinforcement learning scheme for maritime dynamic spectrum access and service scheduling
A1 - Zhongyang MAO
A1 - Jiahuan GENG
A1 - Faping LU
A1 - Wenbiao TIAN
A1 - Jiafang KANG
A1 - Yaozong PAN
J0 - Frontiers of Information Technology & Electronic Engineering
VL - 27
IS - 8
SP - 1
EP - 12
%@ 1869-1951
Y1 - 2026
PB - Zhejiang University Press & Springer
ER -
DOI - 10.1631/ENG.ITEE.2026.0047


Abstract: 
The rapid proliferation of marine economic activities has led to explosive growth in maritime communication demands. Consequently, the scarcity of spectrum resources, the highly dynamic spectrum environment, and diverse service requirements have become increasingly critical issues. While conventional deep reinforcement learning (DRL) methods perform well for specific training scenarios, they fail to generalize effectively to unknown environments. To address this challenge, this paper investigates a maritime dynamic spectrum access and service scheduling scheme based on adaptive meta-reinforcement learning. First, we construct a channel occupancy model that encompasses both Markov frequency-hopping mode and spread-spectrum frequency-hopping mode. By incorporating a dual-queue mechanism for urgent and normal data packets, we formulate the multi-agent cooperative decision-making problem as a decentralized partially observable Markov decision process (POMDP). Second, to overcome the limited generalization capability of traditional DRL algorithms, we propose a spectrum access and service scheduling algorithm based on a meta-learning paradigm (MetaSASS), which facilitates rapid policy transfer through meta-training across a distribution of tasks. Furthermore, to resolve the adaptation efficiency bottleneck caused by fixed inner-loop learning rates, we design an adaptive meta-reinforcement learning SASS algorithm (AMRLSASS). This algorithm employs a task encoding module to dynamically adjust learning rates, thereby balancing convergence speeds across tasks of varying complexities. Simulation results demonstrate that the proposed AMRLSASS algorithm outperforms baseline algorithms across a series of metrics within unknown environments and effectively validate the superiority of the proposed method in complex maritime environments. The algorithm proposed in this paper provides an effective solution for the rapid adaptation of intelligent spectrum access and service scheduling to unknown scenarios in maritime wireless communication systems.

基于自适应元强化学习的海上动态频谱接入与业务调度方法

毛忠阳1,2,3,耿家欢1,2,3,陆发平1,2,3,田文飚1,2,3,康家方1,2,3,潘耀宗1,2,3
1海军航空大学,中国烟台市,264001
2毫米波与太赫兹遥感技术全国重点实验室,中国烟台市,264001
3山东省海空信息感知与处理技术重点实验室,中国烟台市,264001
摘要:随着海洋经济活动的激增,海上通信需求呈爆发式增长,有限的频谱资源、高度动态的频谱环境以及多样化业务需求问题日益凸显。传统深度强化学习方法在特定训练场景下表现良好,但难以泛化到未知网络环境。针对此问题,本文研究了一种基于自适应元强化学习的海上动态频谱接入与业务调度技术。首先,构建涵盖马尔可夫跳频与扩频跳频模式的信道占用模型,并结合紧急与普通数据包的双队列机制,将多智能体协同决策问题建模为去中心化部分可观测马尔可夫决策过程。其次,针对传统深度强化学习算法泛化能力弱的缺陷,提出一种基于元学习范式的频谱接入与业务调度算法,该算法通过在多任务分布上的元训练实现策略的快速迁移。此外,为解决固定内循环学习率导致的适应效率瓶颈,设计了自适应元强化学习算法,通过任务编码模块动态调整学习率,平衡了不同复杂度任务下的收敛速度。仿真结果表明,所提算法在未知环境中平均奖励、紧急包成功率及收敛速度等指标上均优于基线算法,有效验证了该方法在海上复杂环境下的优越性。本研究为海上无线通信系统快速适配未知场景进行智能频谱接入与业务调度提供了有效的解决方案。

关键词:动态频谱接入;自适应元强化学习;服务调度;海上通信

Darkslateblue:Affiliate; Royal Blue:Author; Turquoise:Article

Reference

[1]Albinsaid H, Singh K, Biswas S, et al., 2022. Multi-agent reinforcement learning-based distributed dynamic spectrum access. IEEE Trans Cogn Commun Netw, 8(2):1174-1185.

[2]Antoniou A, Edwards H, Storkey A, 2019. How to train your MAML. https://arxiv.org/abs/1810.09502

[3]Atimati E, Crawford D, Stewart R, 2023. Intelligent shared spectrum coordination in heterogeneous networks. IEEE Virtual Conf on Communications, p.252-257.

[4]Atimati E, Nyasulu T, Crawford D, et al., 2025. Resource management in dynamic shared spectrum networks. IEEE Int Symp on Dynamic Spectrum Access Networks, p.13-19.

[5]Dong L, Qian Y, Xing Y, 2022. Dynamic spectrum access and sharing through actor-critic deep reinforcement learning. EURASIP J Wirel Commun Netw, 2022:48.

[6]Feng MJ, Zhang WH, Krunz M, 2023. Dynamic spectrum access in non-stationary environments: a DRL-LSTM integrated approach. Int Conf on Computing, Networking and Communications, p.159-164.

[7]Finn C, Abbeel P, Levine S, 2017. Model-agnostic meta-learning for fast adaptation of deep networks. Proc 34th Int Conf on Machine Learning, p.1126-1135.

[8]Huang K, Luo ZZ, Liang L, et al., 2022. Fast spectrum sharing in vehicular networks: a meta reinforcement learning approach. IEEE 96th Vehicular Technology Conf, p.1-5.

[9]ITU, 2009. Characteristics of VHF Radio Systems and Equipment for the Exchange of Data and Electronic Mail in the Maritime Mobile Service RR Appendix 18 Channels. ITU-R M.1842-1-2009.

[10]ITU, 2012. Interim Solutions for Improved Efficiency in the Use of the Band 156–174 MHz by Stations in the Maritime Mobile Service. ITU-R M.1084-5-2012.

[11]ITU, 2024. Table of Transmitting Frequencies in the VHF Maritime Mobile Band. ITU Radio Regulations, Appendix 18.

[12]Jia XY, Wang T, Du X, 2024. Federated multi-objective meta-reinforcement learning for adaptive edge task offloading. IEEE Int Conf on High Performance Computing and Communications, p.482-489.

[13]Kai H, Le L, Shi J, et al., 2025. Meta reinforcement learning for fast spectrum sharing in vehicular networks. China Commun, 22(9):320-332.

[14]Ke ZY, Wang XM, Du ZY, et al., 2025. Intelligent frequency reuse for dynamic spectrum anti-jamming: a hybrid-reward-based multi-agent deep reinforcement learning approach. IEEE Wirel Commun Lett, 14(3):771-775.

[15]Khisa S, Elhattab M, Assi C, et al., 2025. Optimizing multi-user uplink cooperative rate-splitting multiple access: efficient user pairing and resource allocation with gradient-based meta learning. IEEE Trans Commun, 73(9):7366-7380.

[16]Li XH, Zhang YL, Ding HC, et al., 2024. Intelligent spectrum sensing and access with partial observation based on hierarchical multi-agent deep reinforcement learning. IEEE Trans Wirel Commun, 23(4):3131-3145.

[17]Li YH, Wang Y, Li Y, et al., 2023. Multi-user dynamic spectrum access based on LRQ deep reinforcement learning network. 25th Int Conf on Advanced Communication Technology, p.79-84.

[18]Li YZ, Zhang WS, Wang CX, et al., 2020. Deep reinforcement learning for dynamic spectrum sensing and aggregation in multi-channel wireless networks. IEEE Trans Cogn Commun Netw, 6(2):464-475.

[19]Liu ZY, Wang XJ, Zhang Y, et al., 2023. Meta reinforcement learning for generalized multiple access in heterogeneous wireless networks. 21st Int Symp on Modeling and Optimization in Mobile, Ad Hoc, and Wireless Networks, p.570-577.

[20]Liu ZY, Wang XJ, Guo K, et al., 2024. Federated meta-RL based multiple access protocol for diverse heterogeneous wireless networks. Int Conf on Future Communications and Networks, p.1-6.

[21]Liu ZY, Wang XJ, Feng CY, et al., 2026. Meta-reinforcement learning with mixture of experts for generalizable multi access in heterogeneous wireless networks. IEEE Trans Commun, 74:870-885.

[22]Lu ZY, Gursoy MC, 2021. Dynamic channel access via meta-reinforcement learning. IEEE Global Communications Conf, p.1-6.

[23]Lyu T, Xu HT, Liu FF, et al., 2024. Computing offloading and resource allocation of NOMA-based UAV emergency communication in marine Internet of Things. IEEE Int Things J, 11(9):15571-15586.

[24]Niu LW, Chen XF, Zhang N, et al., 2023. Multiagent meta-reinforcement learning for optimized task scheduling in heterogeneous edge computing systems. IEEE Int Things J, 10(12):10519-10531.

[25]Nomikos N, Gkonis PK, Bithas PS, et al., 2023. A survey on UAV-aided maritime communications: deployment considerations, applications, and future challenges. IEEE Open J Commun Soc, 4:56-78.

[26]Rao N, Xu H, Qi ZS, et al., 2024. Fast adaptive jamming resource allocation against frequency-hopping spread spectrum in wireless sensor networks via meta-deep-reinforcement-learning. IEEE Trans Aerosp Electron Syst, 60(6):7676-7693.

[27]Rao N, Xu H, Qi ZS, et al., 2025. Adaptive jamming decision-making against FHSS communications via inexpert demonstrations assisted meta reinforcement learning. IEEE Commun Lett, 29(1):105-109.

[28]Saggese F, Pasqualini L, Moretti M, et al., 2021. Deep reinforcement learning for URLLC data management on top of scheduled eMBB traffic. IEEE Global Communications Conf, p.1-6.

[29]Saggese F, Moretti M, Popovski P, 2022. NOMA power minimization of downlink spectrum slicing for eMBB and URLLC users. IEEE Wireless Communications and Networking Conf, p.1725-1730.

[30]Sheng HM, Zhou WJ, Zheng JJ, et al., 2024. Transfer reinforcement learning for dynamic spectrum environment. IEEE Trans Wirel Commun, 23(2):1447-1458.

[31]Sheng TQ, Zhang WS, Ding WJ, et al., 2022. Dynamic spectrum sharing and aggregation scheme based on deep reinforcement learning. Int Wireless Communications and Mobile Computing, p.290-294.

[32]So H, Soya H, 2023. Excluded channel selection scheme for dynamic spectrum sharing in various interference environments. IEEE Access, 11:69798-69806.

[33]Upadhyay D, Upadhyay A, Venu N, et al., 2025. Deep learning-based spectrum sharing for dynamic resource allocation in 6G cognitive radio networks. 4th Int Conf on Power, Control and Computing Technologies, p.1-6.

[34]Wang QF, Xu WQ, Chen HH, 2026. A heterogeneous-agent deep reinforcement learning approach for dynamic spectrum access in cognitive wireless networks. IEEE Trans Cogn Commun Netw, 12:2221-2235.

[35]Yang TT, Gao S, Li JB, et al., 2022. Multi-armed bandits learning for task offloading in maritime edge intelligence networks. IEEE Trans Veh Technol, 71(4):4212-4224.

[36]Yang TT, Zhang WS, Bo YL, et al., 2023. Dynamic spectrum sharing based on federated learning and multi-agent actor-critic reinforcement learning. Int Wireless Communications and Mobile Computing, p.947-952.

[37]Yuan L, Zhou FH, Wu QH, et al., 2024. Channel prediction-enhanced intelligent resource allocation for dynamic spectrum-sharing networks. IEEE Int Conf on Communications, p.2767-2772.

[38]Zhang SG, Wang Z, Gao GY, et al., 2023. Deep reinforcement learning for UAV-assisted spectrum sharing under partial observability. IEEE 98th Vehicular Technology Conf, p.1-6.

[39]Zhang X, Bhuyan A, Kasera SK, et al., 2023. Distributed power allocation for 6-GHz unlicensed spectrum sharing via multi-agent deep reinforcement learning. IEEE Int Conf on Industrial Technology, p.1-6.

[40]Zhang YL, Li XH, Ding HC, et al., 2023. A joint scheme on spectrum sensing and access with partial observation: a multi-agent deep reinforcement learning approach. IEEE/CIC Int Conf on Communications in China, p.1-6.

Open peer comments: Debate/Discuss/Question/Opinion

<1>

Please provide your name, email address and a comment





Full Text:   <5>

Summary:  <12>

Suppl. Mater.: 

CLC number: TN929.5

On-line Access: 2026-06-02

Received: 2026-02-13

Revision Accepted: 2026-06-04

Crosschecked: 2026-06-12

Cited: 0

Clicked: 24

Citations:  Bibtex RefMan EndNote GB/T7714

 ORCID:

Zhongyang MAO

0000-0001-6279-1627

Jiahuan GENG

0009-0004-9031-0690

Faping LU

0000-0002-8172-0311

Wenbiao TIAN

0000-0002-3558-7767

Jiafang KANG

0000-0003-4177-5622

Yaozong PAN

0000-0002-4442-6332

Journal of Zhejiang University-SCIENCE, 38 Zheda Road, Hangzhou 310027, China
Tel: +86-571-87952783; E-mail: cjzhang@zju.edu.cn
Copyright © 2000 - 2026 Journal of Zhejiang University-SCIENCE