Energy-Aware Adaptive AI Inference in Edge–Cloud Computing

Authors

  • Nimesh Yadav
  • Mehtab Alam
  • Chandra Kanta Samal

DOI:

https://doi.org/10.46610/RRMLCC.2026.v05i03.004

Keywords:

Collaborative inference, Early exit, Edge computing, Energy-aware AI, Large language models, Model routing, Split inference

Abstract

Adaptive AI inference distributes the computation across devices, edge servers, and cloud infrastructure to optimize task quality, latency, and energy. This study provides a critical overview of the literature on offloading, split inference, early exits, model routing, hardware adaptation, and speculative decoding selected from 2018 to 2026. Targeted public-source searches and full-text inspection enable comparison of mechanisms, experimental settings, and energy boundaries; exhaustive systematic coverage and meta-analysis are not claimed. The synthesis differentiates between measurements of client-energy and wider node measurements, analytical estimates, and communication and monetary proxies. Qualified numerical examples demonstrate the importance of having baseline choices and differences to go with reported savings. The evaluation requirements go beyond traditional DNN inference, as the models are based on language and multimodal models introduce repeated verification, rejected computation, and response-length dependence. The primary recommendation is to consider policies at specified quality points and latency goals and to factor in communication, monitoring, transitions, and escalation. It is important to have a clear view of accounting and reproducible workloads to determine if adaptation ultimately improves overall energy efficiency or moves energy between tiers.

References

Z. Zhou, X. Chen, E. Li, L. Zeng, K. Luo and J. Zhang, “Edge intelligence: Paving the last mile of artificial intelligence with edge computing,” in Proceedings of the IEEE, vol. 107, no. 8, pp. 1738–1762, Aug. 2019.

Y. Matsubara, M. Levorato, and F. Restuccia, “Split computing and early exiting for deep learning applications: Survey and research challenges,” ACM Computing Surveys, vol. 55, no. 5, pp. 1–30, Dec. 2022.

W.-Q. Ren et al., “A survey on collaborative DNN inference for edge intelligence,” Machine Intelligence Research, vol. 20, pp. 370–395, May 2023.

E. Li, Z. Zhou, X. Chen, “Edge intelligence: On-demand deep learning model co-inference with device-edge synergy,” Proceedings of the 2018 Workshop on Mobile Edge Communications, pp. 31–36, Aug. 2018.

A. E. Eshratifar, M. S. Abrishami, and M. Pedram, “JointDNN: An efficient training and inference engine for intelligent mobile cloud computing services,” in IEEE Transactions on Mobile Computing, vol. 20, no. 2, pp. 565–576, Feb. 2021.

A. E. Eshratifar, A. Esmaili, and M. Pedram, “BottleNet: A deep learning architecture for intelligent mobile cloud computing services,” 2019 IEEE/ACM International Symposium on Low Power Electronics and Design (ISLPED), Lausanne, Switzerland, Jul. 2019, pp. 1–6.

J. Shao and J. Zhang, “BottleNet++: An end-to-end approach for feature compression in device-edge co-inference systems,” 2020 IEEE International Conference on Communications Workshops (ICC Workshops), Dublin, Ireland, 2020, pp. 1–6.

A. A. Majeed, P. Kilpatrick, I. Spence, and B. Varghese, “NEUKONFIG: Reducing edge service downtime when repartitioning DNNs,” 2021 IEEE International Conference on Cloud Engineering (IC2E), San Francisco, CA, USA, 2021, pp. 118–125.

T. Niu, Y. Teng, Z. Han, and P. Zou, “An adaptive device-edge co-inference framework based on soft actor-critic,” 2022 IEEE Wireless Communications and Networking Conference (WCNC), Austin, TX, USA, 2022, pp. 2571–2576.

S. Laskaridis, S. I. Venieris, M. Almeida, I. Leontiadis, and N. D. Lane, “SPINN: Synergistic progressive inference of neural networks over device and cloud,” Proceedings of the 26th Annual International Conference on Mobile Computing and Networking, Sept. 2020, pp. 1–15.

M. Almeida, S. Laskaridis, S. I. Venieris, I. Leontiadis, and N. D. Lane, “Dyno: Dynamic onloading of deep neural networks from cloud to device,” ACM Transactions on Embedded Computing Systems, Oct. 2022, pp. 1–24.

E. Tang, X. Guo, and T. Stefanov, “The effects of partitioning strategies on energy consumption in distributed CNN inference at the edge,” arXiv, Oct. 2022.

X. Li et al., “Predictive exit: Prediction of fine-grained early exits for computation-and energy-efficient inference,” Proceedings of the AAAI Conference on Artificial Intelligence, vol. 37, no. 7, pp. 8657–8665, 2023.

Z. Zhang, Y. Zhao, M.-C. Chang, C. Lin, and J. Liu, “E4: Energy-efficient DNN inference for edge video analytics via early exiting and DVFS,” Proceedings of the AAAI Conference on Artificial Intelligence, vol. 39, no. 1, 2025.

D. May, A. Tundo, S. Ilager, and I. Brandic, “DynaSplit: A hardware-software co-design framework for energy-aware inference on edge,” arXiv, Oct. 2024.

X. Li, H. Li, C. Sun, Q. Fan, Z. Han, and V. C. M. Leung, “Edge-enhanced intelligence: A comprehensive survey of large language models and edge-cloud computing synergy,” in IEEE Communications Surveys & Tutorials, vol. 28, pp. 1248–1284, 2026.

S. Li et al., “Collaborative inference and learning between edge slms and cloud LLMs: A survey of algorithms, execution, and open challenges,” ACM Computing Surveys, Dec. 2025.

J. Liu et al., “Edge-Cloud collaborative computing on distributed intelligence and model optimization: A survey,” in IEEE Communications Surveys & Tutorials, vol. 28, pp. 5049–5080, 2026.

I. Ong et al., “RouteLLM: Learning to route LLMs from preference data,” International Conference on Learning Representations, 2025, pp. 3443–34448.

A. Šabanović, P. J. Maliakel, and I. Brandić, “INAR-VL: Input-aware routing for edge–cloud vision–language inference,” Proceedings of the 24th Annual International Conference on Mobile Systems, Applications and Services Workshops, Jun. 2026, pp. 328–335.

X. Li, S. Ghafouri, J. Fan, B. Ali, H. Vandierendonck, and D. S. Nikolopoulos, “ConfigSpec: Profiling-based configuration selection for distributed edge-cloud speculative LLM serving,” Proceedings of the 4th International Workshop on Testing Distributed Internet of Things Systems, Apr. 2026, pp. 1–6.

Y. Chen, R. Li, X. Yu, Z. Zhao, H. Zhang, “Adaptive layer splitting for wireless LLM inference in edge computing: A model-based reinforcement learning approach,” arXiv, Jun. 2024.

X. Li et al., “Sled: A speculative LLM decoding framework for efficient edge serving,” Proceedings of the Tenth ACM/IEEE Symposium on Edge Computing, Dec. 2025, pp. 1–8.

Y. Venkatesha, S. Kundu, and P. Panda, “Fast and cost-effective speculative edge-cloud decoding with early exits,” arXiv, May 2025.

J. Park, S. Oh, and S.-L. Kim, “Energy-efficient wireless LLM inference via uncertainty and importance-aware speculative decoding,” arXiv, Aug. 2025.

M. Wang, P. Liu, A. Taherkordi, and J. Guitart, “GreenPipe: Power modeling for containerized DNN inference on Kubernetes edge nodes,” arXiv, Sept. 2026.

M. Zhang, Y. Wang, P. Yu, X. Qiu, and S. Guo, “Eco-efficient task scheduling for MLLMs in edge-cloud continuum,” Computer Networks, vol. 282, Jun. 2026.

M. J. Page et al., “The PRISMA 2020 statement: an updated guideline for reporting systematic reviews,” BMJ, Mar. 2021.

M. L. Rethlefsen et al., “PRISMA-S: an extension to the PRISMA statement for reporting literature searches in systematic reviews,” Systematic Reviews, vol. 10, Jan. 2021.

R. G. Pacheco, R. S. Couto, and O. Simeone, “On the impact of deep neural network calibration on adaptive edge offloading for image classification,” Journal of Network and Computer Applications, vol. 217, Aug. 2023.

Published

2026-10-05

Issue

Section

Articles