Edge-Optimized Vision Transformer Architecture for Ultra-Low-Latency Object Detection in 6G IoT Devices

Author
Keywords
Abstract

The advancements in the high demand for ultra-low-latency intelligence in large-scale IoT networks have enhanced the demands of edge-deployable vision models that can detect objects in real-time. In order to satisfy this need, an Edge-Optimized Vision Transformer design is created based on spectral token compression, entropy-directed patch refinement, latency-constrained multi-head attention, as well as hierarchical feature fusion that is optimized to run on limited edge devices. The model is trained with a balanced and high-variability object detection data set, which ensures that it can be resistant to illumination variations, motion blur, and small-object cases. The proposed system is experimentally found to have 96.8 % detection accuracy, 94.7% mAP@ 0.5, and 4.3 ms end-to-end latency on edge-class IoT hardware. Also, the architecture can save 41 % FLOPs of computational loads, 38 % of the energy consumption, and 22 % of the small-object recall of lightweight transformer baselines. The findings verify that the proposed.

Year of Conference
2026
Conference Name
4th IEEE International Conference on Power Electronics and IoT Applications in Renewable Energy and its Control, PARC 2026
Number of Pages
228-233,
Publisher
Institute of Electrical and Electronics Engineers Inc.
ISBN Number
979-833159183-0 (ISBN)
URL
https://ieeexplore.ieee.org/document/11453596
DOI
10.1109/PARC68365.2026.11453596
Short Title
IEEE Int. Conf. Power Electron. IoT Appl. Renew. Energy its Control, PARC
Conference Proceedings
Download citation
Cits
0
CIT

For admissions and all other information, please visit the official website of

Cambridge Institute of Technology

Cambridge Group of Institutions

Contact

Web portal developed and administered by Dr. Subrahmanya S. Katte, Dean - Academics.

Contact the Site Admin.