Abstract
Keywords
Introduction
Unmanned aerial vehicles (UAVs) are increasingly deployed as mobile edge-computing (MEC) servers that bring computation closer to ground-based Internet-of-Things (IoT) devices in areas where fixed infrastructure is sparse, damaged, or overloaded [1]. A large number of papers using DTs combined with federated deep reinforcement learning (FDRL) so that multiple UAV cells can jointly learn an offloading policy without sharing raw user data. These works are generally referred to as Digital Twin-Federated Deep Reinforcement Learning (DT-FDRL).
The positive results reported for DT-FDRL and related UAV-MEC strategies are usually obtained in simulators that simplify one or more of the three dynamics that actually govern offloading decisions in the field: UAV and user mobility, the air-to-ground wireless channel, and UAV battery consumption [2,3]. Free-space or fixed-loss channels remove small-scale fading and line-of-sight variability while straight-line or pure random-walk mobility removes the temporal correlation of real flight controllers and energy-agnostic simulators remove the feedback between battery state and offloading. When any of these simplifications is present, it is difficult to know whether an algorithmic improvement reported in the literature reflects the algorithm itself or an artifact of an easy simulated environment [4,5].
This paper addresses that gap by proposing DT-FDRL+SemCom, a digital-twin-assisted federated reinforcement learning strategy augmented with semantic-communication (SemCom) task compression, and by validating it on a purpose-built, open, and fully reproducible evaluation stack rather than on a simplified simulator. The proposed strategy is benchmarked against six representative baselines under identical seeds, identical mobility, identical channel conditions, and identical energy dynamics, so that any observed advantage is attributable to the algorithm rather than to the environment.
The paper makes four contributions:
A UAV-MEC offloading strategy that combines federated Q-learning, digital-twin-assisted Dyna-style planning, and semantic-compression task shortage into a single policy.
A modular, open-source evaluation stack that fuses a Gauss-Markov UAV mobility model, a probabilistic line-of-sight air-to-ground channel model with Rician/Rayleigh fading, and a rotary-wing propulsion plus DVFS compute-energy model into a single Gym-style UAV-MEC offloading environment.
A reproducible seeding and configuration protocol in which every strategy is evaluated under the same per-seed random state, so that cross-strategy comparisons are not confounded by environment stochasticity.
A disciplined ablation across seven strategies including the proposed DT-FDRL+SemCom, its DT-FDRL and FDRL ancestors, and four baselines with mean, 95% confidence interval, and Cohen’s d effect-size reporting, isolating the individual contribution of federation, twin-assisted planning, and semantic compression to the overall result.
The remainder of the paper proceeds as follows. Section 2 reviews mobility, channel, energy, and DT-FDRL modeling choices in the recent literature. Section 3 defines the proposed DT-FDRL+SemCom strategy and the evaluation stack used to validate it. Section 4 describes the experimental protocol. Section 5 reports and discusses the results, Section 6 discusses limitations and future works, and Section 7 concludes.
Complete Article
The complete article, including all figures, tables, equations and algorithms, is available in the official publication PDF.
Conclusion
This paper proposed DT-FDRL+SemCom, a DT-assisted federated reinforcement learning strategy augmented with SemCom task compression for UAV-MEC offloading. It is validated on an open, modular, and fully reproducible mobility channel and energy evaluation stack against six baselines. Naive offloading heuristics were found to underperform a no-offload baseline once realistic channel fading and propulsion cost were present while FL alone produced only a modest gain over independent learning. DT-assisted Dyna-style planning produced a further, separable improvement and semantic-compression task shorten produced the largest single gain. The complete DT-FDRL+SemCom strategy cut latency by 66.2% and raised task success from 69.4% to 90.0% relative to plain FDRL, with a Cohen’s d of +27.26 in mean reward relative to the local-only.
References
- S. Ahmad and Z. Awais, “Ethical implications of artificial intelligence: A review, taxonomy, governance framework, and future research agenda, ” Journal of Emerging Paradigms in Computing Systems, vol. 1, no. 01, 2025. Available: https://jepcs.org/index.php/jepcs/article/view/19.
- Z. ul Huda, “An enhanced hybrid PSO-LSTM framework for energy-aware task scheduling in heterogeneous cloud computing environments, ” Journal of Emerging Paradigms in Computing Systems, vol. 1, no. 01, 2025. Available: https://jepcs.org/index.php/jepcs/article/view/11.
- A. Al-Hourani, S. Kandeepan, and S. Lardner, “Optimal LAP altitude for maximum coverage, ” IEEE Wireless Communications Letters, vol. 3, no. 6, pp. 569-572, 2014. doi: 10.1109/LWC.2014.2342736
- Y. Zeng and R. Zhang, “Energy-efficient UAV communication with trajectory optimization, ” IEEE Transactions on Wireless Communications, vol. 16, no. 6, pp. 3747-3760, 2017. doi: 10.1109/TWC.2017.2688328
- B. Liang and Z. J. Haas, “Predictive distance-based mobility management for PCS networks, ” in Proceedings of IEEE INFOCOM, vol. 3, New York, NY, USA, 1999, pp. 1377-1384. doi: 10.1109/INFCOM.1999.752157
- R. S. Sutton, “Integrated architectures for learning, planning, and reacting based on approximating dynamic programming, ” in Proceedings of the Seventh International Conference on Machine Learning, Austin, TX, USA, 1990, pp. 216-224. doi: 10.1016/B978-1-55860-141-3.50030-4
- H. B. McMahan, E. Moore, D. Ramage, S. Hampson, and B. Agüera y Arcas, “Communication-efficient learning of deep networks from decentralized data, ” in Proceedings of the 20th International Conference on Artificial Intelligence and Statistics (AISTATS), vol. 54, Fort Lauderdale, FL, USA, 2017, pp. 1273-1282.
- D. Gündüz, Z. Qin, I. E. Aguerri, H. S. Dhillon, Z. Yang, A. Yener, K.-K. Wong, and C.-B. Chae, “Beyond transmitting bits: Context, semantics, and task-oriented communications, ” IEEE Journal on Selected Areas in Communications, vol. 41, no. 1, pp. 5-41, 2023. doi: 10.1109/JSAC.2022.3223408
- Y. Zeng, J. Xu, and R. Zhang, “Energy minimization for wireless communication with rotary-wing UAV, ” IEEE Transactions on Wireless Communications, vol. 18, no. 4, pp. 2329-2345, 2019. doi: 10.1109/TWC.2019.2902559
- W. Khawaja, I. Guvenc, D. W. Matolak, U.-C. Fiebig, and N. Schneckenberger, “A survey of air-to-ground propagation channel modeling for unmanned aerial vehicles, ” IEEE Communications Surveys & Tutorials, vol. 21, no. 3, pp. 2361-2391, 2019. doi: 10.1109/COMST.2019.2915069