
Zhongyang MAO, Jiahuan GENG, Faping LU, Wenbiao TIAN, Jiafang KANG, Yaozong PAN. An adaptive meta-reinforcement learning scheme for maritime dynamic spectrum access and service scheduling[J]. Journal of Zhejiang University Science C, 2026, 27(8): 1-12.
@article{title="An adaptive meta-reinforcement learning scheme for maritime dynamic spectrum access and service scheduling",
author="Zhongyang MAO, Jiahuan GENG, Faping LU, Wenbiao TIAN, Jiafang KANG, Yaozong PAN",
journal="Journal of Zhejiang University Science C",
volume="27",
number="8",
pages="1-12",
year="2026",
publisher="Zhejiang University Press & Springer",
doi="10.1631/ENG.ITEE.2026.0047"
}
%0 Journal Article
%T An adaptive meta-reinforcement learning scheme for maritime dynamic spectrum access and service scheduling
%A Zhongyang MAO
%A Jiahuan GENG
%A Faping LU
%A Wenbiao TIAN
%A Jiafang KANG
%A Yaozong PAN
%J Frontiers of Information Technology & Electronic Engineering
%V 27
%N 8
%P 1-12
%@ 1869-1951
%D 2026
%I Zhejiang University Press & Springer
%DOI 10.1631/ENG.ITEE.2026.0047
TY - JOUR
T1 - An adaptive meta-reinforcement learning scheme for maritime dynamic spectrum access and service scheduling
A1 - Zhongyang MAO
A1 - Jiahuan GENG
A1 - Faping LU
A1 - Wenbiao TIAN
A1 - Jiafang KANG
A1 - Yaozong PAN
J0 - Frontiers of Information Technology & Electronic Engineering
VL - 27
IS - 8
SP - 1
EP - 12
%@ 1869-1951
Y1 - 2026
PB - Zhejiang University Press & Springer
ER -
DOI - 10.1631/ENG.ITEE.2026.0047
Abstract: The rapid proliferation of marine economic activities has led to explosive growth in maritime communication demands. Consequently, the scarcity of spectrum resources, the highly dynamic spectrum environment, and diverse service requirements have become increasingly critical issues. While conventional deep reinforcement learning (DRL) methods perform well for specific training scenarios, they fail to generalize effectively to unknown environments. To address this challenge, this paper investigates a maritime dynamic spectrum access and service scheduling scheme based on adaptive meta-reinforcement learning. First, we construct a channel occupancy model that encompasses both Markov frequency-hopping mode and spread-spectrum frequency-hopping mode. By incorporating a dual-queue mechanism for urgent and normal data packets, we formulate the multi-agent cooperative decision-making problem as a decentralized partially observable Markov decision process (POMDP). Second, to overcome the limited generalization capability of traditional DRL algorithms, we propose a spectrum access and service scheduling algorithm based on a meta-learning paradigm (MetaSASS), which facilitates rapid policy transfer through meta-training across a distribution of tasks. Furthermore, to resolve the adaptation efficiency bottleneck caused by fixed inner-loop learning rates, we design an adaptive meta-reinforcement learning SASS algorithm (AMRLSASS). This algorithm employs a task encoding module to dynamically adjust learning rates, thereby balancing convergence speeds across tasks of varying complexities. Simulation results demonstrate that the proposed AMRLSASS algorithm outperforms baseline algorithms across a series of metrics within unknown environments and effectively validate the superiority of the proposed method in complex maritime environments. The algorithm proposed in this paper provides an effective solution for the rapid adaptation of intelligent spectrum access and service scheduling to unknown scenarios in maritime wireless communication systems.
[1]Albinsaid H, Singh K, Biswas S, et al., 2022. Multi-agent reinforcement learning-based distributed dynamic spectrum access. IEEE Trans Cogn Commun Netw, 8(2):1174-1185.
[2]Antoniou A, Edwards H, Storkey A, 2019. How to train your MAML. https://arxiv.org/abs/1810.09502
[3]Atimati E, Crawford D, Stewart R, 2023. Intelligent shared spectrum coordination in heterogeneous networks. IEEE Virtual Conf on Communications, p.252-257.
[4]Atimati E, Nyasulu T, Crawford D, et al., 2025. Resource management in dynamic shared spectrum networks. IEEE Int Symp on Dynamic Spectrum Access Networks, p.13-19.
[5]Dong L, Qian Y, Xing Y, 2022. Dynamic spectrum access and sharing through actor-critic deep reinforcement learning. EURASIP J Wirel Commun Netw, 2022:48.
[6]Feng MJ, Zhang WH, Krunz M, 2023. Dynamic spectrum access in non-stationary environments: a DRL-LSTM integrated approach. Int Conf on Computing, Networking and Communications, p.159-164.
[7]Finn C, Abbeel P, Levine S, 2017. Model-agnostic meta-learning for fast adaptation of deep networks. Proc 34th Int Conf on Machine Learning, p.1126-1135.
[8]Huang K, Luo ZZ, Liang L, et al., 2022. Fast spectrum sharing in vehicular networks: a meta reinforcement learning approach. IEEE 96th Vehicular Technology Conf, p.1-5.
[9]ITU, 2009. Characteristics of VHF Radio Systems and Equipment for the Exchange of Data and Electronic Mail in the Maritime Mobile Service RR Appendix 18 Channels. ITU-R M.1842-1-2009.
[10]ITU, 2012. Interim Solutions for Improved Efficiency in the Use of the Band 156–174 MHz by Stations in the Maritime Mobile Service. ITU-R M.1084-5-2012.
[11]ITU, 2024. Table of Transmitting Frequencies in the VHF Maritime Mobile Band. ITU Radio Regulations, Appendix 18.
[12]Jia XY, Wang T, Du X, 2024. Federated multi-objective meta-reinforcement learning for adaptive edge task offloading. IEEE Int Conf on High Performance Computing and Communications, p.482-489.
[13]Kai H, Le L, Shi J, et al., 2025. Meta reinforcement learning for fast spectrum sharing in vehicular networks. China Commun, 22(9):320-332.
[14]Ke ZY, Wang XM, Du ZY, et al., 2025. Intelligent frequency reuse for dynamic spectrum anti-jamming: a hybrid-reward-based multi-agent deep reinforcement learning approach. IEEE Wirel Commun Lett, 14(3):771-775.
[15]Khisa S, Elhattab M, Assi C, et al., 2025. Optimizing multi-user uplink cooperative rate-splitting multiple access: efficient user pairing and resource allocation with gradient-based meta learning. IEEE Trans Commun, 73(9):7366-7380.
[16]Li XH, Zhang YL, Ding HC, et al., 2024. Intelligent spectrum sensing and access with partial observation based on hierarchical multi-agent deep reinforcement learning. IEEE Trans Wirel Commun, 23(4):3131-3145.
[17]Li YH, Wang Y, Li Y, et al., 2023. Multi-user dynamic spectrum access based on LRQ deep reinforcement learning network. 25th Int Conf on Advanced Communication Technology, p.79-84.
[18]Li YZ, Zhang WS, Wang CX, et al., 2020. Deep reinforcement learning for dynamic spectrum sensing and aggregation in multi-channel wireless networks. IEEE Trans Cogn Commun Netw, 6(2):464-475.
[19]Liu ZY, Wang XJ, Zhang Y, et al., 2023. Meta reinforcement learning for generalized multiple access in heterogeneous wireless networks. 21st Int Symp on Modeling and Optimization in Mobile, Ad Hoc, and Wireless Networks, p.570-577.
[20]Liu ZY, Wang XJ, Guo K, et al., 2024. Federated meta-RL based multiple access protocol for diverse heterogeneous wireless networks. Int Conf on Future Communications and Networks, p.1-6.
[21]Liu ZY, Wang XJ, Feng CY, et al., 2026. Meta-reinforcement learning with mixture of experts for generalizable multi access in heterogeneous wireless networks. IEEE Trans Commun, 74:870-885.
[22]Lu ZY, Gursoy MC, 2021. Dynamic channel access via meta-reinforcement learning. IEEE Global Communications Conf, p.1-6.
[23]Lyu T, Xu HT, Liu FF, et al., 2024. Computing offloading and resource allocation of NOMA-based UAV emergency communication in marine Internet of Things. IEEE Int Things J, 11(9):15571-15586.
[24]Niu LW, Chen XF, Zhang N, et al., 2023. Multiagent meta-reinforcement learning for optimized task scheduling in heterogeneous edge computing systems. IEEE Int Things J, 10(12):10519-10531.
[25]Nomikos N, Gkonis PK, Bithas PS, et al., 2023. A survey on UAV-aided maritime communications: deployment considerations, applications, and future challenges. IEEE Open J Commun Soc, 4:56-78.
[26]Rao N, Xu H, Qi ZS, et al., 2024. Fast adaptive jamming resource allocation against frequency-hopping spread spectrum in wireless sensor networks via meta-deep-reinforcement-learning. IEEE Trans Aerosp Electron Syst, 60(6):7676-7693.
[27]Rao N, Xu H, Qi ZS, et al., 2025. Adaptive jamming decision-making against FHSS communications via inexpert demonstrations assisted meta reinforcement learning. IEEE Commun Lett, 29(1):105-109.
[28]Saggese F, Pasqualini L, Moretti M, et al., 2021. Deep reinforcement learning for URLLC data management on top of scheduled eMBB traffic. IEEE Global Communications Conf, p.1-6.
[29]Saggese F, Moretti M, Popovski P, 2022. NOMA power minimization of downlink spectrum slicing for eMBB and URLLC users. IEEE Wireless Communications and Networking Conf, p.1725-1730.
[30]Sheng HM, Zhou WJ, Zheng JJ, et al., 2024. Transfer reinforcement learning for dynamic spectrum environment. IEEE Trans Wirel Commun, 23(2):1447-1458.
[31]Sheng TQ, Zhang WS, Ding WJ, et al., 2022. Dynamic spectrum sharing and aggregation scheme based on deep reinforcement learning. Int Wireless Communications and Mobile Computing, p.290-294.
[32]So H, Soya H, 2023. Excluded channel selection scheme for dynamic spectrum sharing in various interference environments. IEEE Access, 11:69798-69806.
[33]Upadhyay D, Upadhyay A, Venu N, et al., 2025. Deep learning-based spectrum sharing for dynamic resource allocation in 6G cognitive radio networks. 4th Int Conf on Power, Control and Computing Technologies, p.1-6.
[34]Wang QF, Xu WQ, Chen HH, 2026. A heterogeneous-agent deep reinforcement learning approach for dynamic spectrum access in cognitive wireless networks. IEEE Trans Cogn Commun Netw, 12:2221-2235.
[35]Yang TT, Gao S, Li JB, et al., 2022. Multi-armed bandits learning for task offloading in maritime edge intelligence networks. IEEE Trans Veh Technol, 71(4):4212-4224.
[36]Yang TT, Zhang WS, Bo YL, et al., 2023. Dynamic spectrum sharing based on federated learning and multi-agent actor-critic reinforcement learning. Int Wireless Communications and Mobile Computing, p.947-952.
[37]Yuan L, Zhou FH, Wu QH, et al., 2024. Channel prediction-enhanced intelligent resource allocation for dynamic spectrum-sharing networks. IEEE Int Conf on Communications, p.2767-2772.
[38]Zhang SG, Wang Z, Gao GY, et al., 2023. Deep reinforcement learning for UAV-assisted spectrum sharing under partial observability. IEEE 98th Vehicular Technology Conf, p.1-6.
[39]Zhang X, Bhuyan A, Kasera SK, et al., 2023. Distributed power allocation for 6-GHz unlicensed spectrum sharing via multi-agent deep reinforcement learning. IEEE Int Conf on Industrial Technology, p.1-6.
[40]Zhang YL, Li XH, Ding HC, et al., 2023. A joint scheme on spectrum sensing and access with partial observation: a multi-agent deep reinforcement learning approach. IEEE/CIC Int Conf on Communications in China, p.1-6.
CLC number: TN929.5
On-line Access: 2026-06-02
Received: 2026-02-13
Revision Accepted: 2026-06-04
Crosschecked: 2026-06-12
Cited: 0
Clicked: 25
Open peer comments: Debate/Discuss/Question/Opinion
<1>