HTE-YOLO: A Hybrid Transformer-Enhanced Deep Learning Framework for AI-Driven Real-Time Object Detection
| Author | |
|---|---|
| Keywords | |
| Abstract |
The functions of real-time object detection on mobile and edge devices are still considered a challenging task because of lack of computational power, memory management and energy conservation needs. In spite of the high performance of modern deep learning models, several of the state-of-the-art detectors have high computational costs and are not suitable to be deployed in lightweight settings. This paper presents a Hybrid Transformer-Enhanced YOLO (HTE-YOLO) model to be used in a mobile environment. It has an architecture that combines a lightweight YOLOv8 backbone and a Vision Transformer-based attention module to achieve a better grasp of the world around and higher representation of multi-scales features. Further, knowledge distillation strategy is used to enhance the generalization of the models without adding to the computation cost. The model is trained on a personal dataset of 8,000 images gathered in the real-world setting and tested on a part of the COCO dataset. Experimental performance shows better results of 99.2mAP0.5, 89.5mAP0.5:0.95 and an inference rate of 58 FPS using only 6.5M parameters. The suggested framework balances between accuracy and efficiency effectively, thus it is applicable to real-time mobile and edge. |
| Year of Conference |
2026
|
| Conference Name |
2026 4th International Conference on Artificial Intelligence and Machine Learning Applications: Healthcare and Internet of Things, AIMLA 2026
|
| Publisher |
Institute of Electrical and Electronics Engineers Inc.
|
| ISBN Number |
979-831950634-4 (ISBN)
|
| URL |
https://ieeexplore.ieee.org/document/11522482
|
| DOI |
10.1109/AIMLA67915.2026.11522482
|
| Short Title |
Int. Conf. Artif. Intell. Mach. Learn. Appl.: Healthc. Internet Things, AIMLA
|
Conference Proceedings
|
|
| Download citation | |
| Cits |
0
|
