| Issue |
ITM Web Conf.
Volume 88, 2026
The 2026 International Conference on Artificial Intelligence, Big Data and Computer Science (AIBDCS 2026)
|
|
|---|---|---|
| Article Number | 01041 | |
| Number of page(s) | 5 | |
| Section | Artificial Intelligence, Big Data and Computer Science | |
| DOI | https://doi.org/10.1051/itmconf/20268801041 | |
| Published online | 27 July 2026 | |
A Cascaded YOLOv8s-ResNet-18 Framework for Facial Expression Recognition
School of Computer Science and Technology, Harbin Institute of Technology, Weihai, Shandong, China
* This email address is being protected from spambots. You need JavaScript enabled to view it.
Abstract
In intelligent human-computer interaction and monitoring device applications, facial expression recognition is an important technique. The recognition accuracy can directly influence the success rate of practical systems and the user experience. However, in uncontrolled scenes, the background may be cluttered and the face position may be unknown; under this condition, a single convolutional neural network is easily affected by noise, and the recognition accuracy can drop clearly. Therefore, a model with better resistance to interference in complex environments has practical value. Based on the public FER2013 dataset, this study first trained ResNet-18 as the basic model; then the pre-trained YOLOv8s model was connected with the trained ResNet-18 classifier, so a combined framework of detection first and classification later was built. To imitate real complex environments, this study also constructed a synthetic complex-scene test set by processing images from the original test set in several ways. On the original test set, ResNet-18 performed normally, but on the complex-scene test set, the accuracy of the single model was only 23.3%. After the YOLOv8s detector was added, the combined model reached an end-to-end accuracy of 51.99% under the same complex-scene setting; this result was much better than that of the single model. The experimental results show that the pre-detection module can improve the model’s ability to resist background interference in complex environments. This work provides a useful reference framework and possible improvement direction for real-time facial expression recognition systems in practical applications.
© The Authors, published by EDP Sciences, 2026
This is an Open Access article distributed under the terms of the Creative Commons Attribution License 4.0, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.
Current usage metrics show cumulative count of Article Views (full-text article views including HTML views, PDF and ePub downloads, according to the available data) and Abstracts Views on Vision4Press platform.
Data correspond to usage on the plateform after 2015. The current usage metrics is available 48-96 hours after online publication and is updated daily on week days.
Initial download of the metrics may take a while.

