Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Multi-Agent Reinforcement Learning for Joint Handover Management and Power Allocation in Multi-Orbit Satellite Networks

Domain:

digital infrastructure

Record type:

paper
Creator:
AfiBalMarAlm
Publisher:
arXiv
Host:avatar
Future sixth-generation non-terrestrial networks are expected to combine low Earth orbit (LEO), medium Earth orbit (MEO), and geostationary Earth orbit (GEO) satellites, whose complementary layers must be coordinated through joint user association, power allocation, and handover management under fast LEO dynamics. This paper studies this problem by formulating it as a mixed-integer nonlinear program and decomposing it into a multi-agent reinforcement learning (MARL) policy that selects the associations and a convex power-allocation subproblem solved exactly at each time slot that defines the reward of the MARL part. The association policy is trained with multi-agent proximal policy optimization (MAPPO) and the targeted multi-agent communication (TarMAC) mechanism, and is made aware of the orbital layer through a state that encodes layer-dependent handover penalties. Evaluated on a realistic multi-constellation scenario built from real two-line element data over Nairobi, Kenya, the proposed policy reaches 92% of the throughput of a greedy signal-to-noise ratio (SNR) maximizing scheme while triggering more than four times fewer handovers, and improves throughput by roughly 14% over a conservative stay heuristic. Compared to an LEO-only learned policy of identical architecture, it attains slightly higher throughput with fewer handovers by offloading a fraction of the users to the MEO and GEO layers, an emergent multi-orbit behavior that drives its favorable throughput and handover trade-off.

Visit

doi.org

Tags

Signal Processing (eess.SP)Systems and Control (eess.SY)FOS: Electrical engineering, electronic engineering, information engineering

Licenses

arXiv.org perpetual, non-exclusive licensehttp://arxiv.org/licenses/nonexclusive-distrib/1.0/

Similar

Dynamic preference allocation for multi-objective, multi-agent reinforcement learningOff-The-Grid Multi-Agent Reinforcement LearningUniversally Expressive Communication in Multi-Agent Reinforcement LearningMava: Fast Parallel Multi-agent Reinforcement Learning in JAXEmergent Communication in Multi-Agent Reinforcement Learning for Flying Base StationsFormation Strategy Optimization Using Multi-Agent Reinforcement Learning (MARL)

Dynamic preference allocation for multi-objective, multi-agent reinforcement learning

Dynamic preference allocation for multi-objective, multi-agent reinforcement learning

Poster presented at the Deep Learning Indaba 2023 by Asad Jeewa

Off-The-Grid Multi-Agent Reinforcement Learning

Off-The-Grid Multi-Agent Reinforcement Learning

Poster presented at the Deep Learning Indaba 2022 by Claude Formanek

Universally Expressive Communication in Multi-Agent Reinforcement Learning

Universally Expressive Communication in Multi-Agent Reinforcement Learning

Poster presented at the Deep Learning Indaba 2022 by Matthew Morris

Mava: Fast Parallel Multi-agent Reinforcement Learning in JAX

Mava: Fast Parallel Multi-agent Reinforcement Learning in JAX

Poster presented at the Deep Learning Indaba 2023 by Ruan de Kock

Emergent Communication in Multi-Agent Reinforcement Learning for Flying Base Stations

Formation Strategy Optimization Using Multi-Agent Reinforcement Learning (MARL)

Formation Strategy Optimization Using Multi-Agent Reinforcement Learning (MARL)

Poster presented at the Deep Learning Indaba 2023 by Abdel Mfougouon Njupoun