Business Faculty Publications

Document Type

Article

Publication Date

10-6-2026

Publication Title

Information

Keywords

facial-expression recognition; FER2013; DINOv2; EfficientNetB3; vision transformer; feature fusion; ensemble learning; test-time augmentation

Disciplines

Artificial Intelligence and Robotics | Business

Abstract

Facial-expression recognition (FER) on FER2013 remains challenging because of low-resolution images, class imbalance, and label ambiguity. This study presents a global–local feature-fusion framework that integrates complementary representations with validation-based ensemble refinement. A frozen DINOv2 ViT-Base captures global facial semantics, while EfficientNetB3 extracts complementary local texture features. Their fused representation is used for seven-class facial-expression classification. The classification head is first trained with targeted feature-space SMOTE, and the EfficientNetB3 branch is then partially fine-tuned. Five-view test-time augmentation (TTA) is further incorporated at inference, together with an independently trained ConvNeXt-Tiny branch to provide additional architectural diversity. Ensemble weights are selected using a held-out validation set, while the test partition is reserved for final evaluation. On the 7178-image FER2013 test partition, the resulting ensemble achieves 76.51% accuracy, 75.77% macro F1, 76.32% weighted F1, 0.7158 Cohen’s kappa, and 0.9574 macro ROC AUC. The results demonstrate the potential of combining complementary global–local representations, inference augmentation, and ensemble refinement for facial-expression recognition on FER2013.

DOI

https://doi.org/10.3390/info17100982

Version

Publisher's PDF

Creative Commons License

Creative Commons Attribution 4.0 International License
This work is licensed under a Creative Commons Attribution 4.0 International License.

Volume

17

Issue

10

Share

COinS