2026-06-04
2026-04-30
2026-02-27
Manuscript received December 23, 2025; revised January 27, 2026; accepted March 26, 2026; published July 21, 2026.
Abstract—Immersive Virtual Reality (VR) is increasingly used to support experiential learning and social presence in higher education, while 360° video offers lightweight immersive experiences accessible from standard browsers and mobile devices. However, few platforms integrate deep learning-based Facial Expression Recognition (FER) with VR and 360° e-learning in a unified architecture tailored for distance education. This paper presents an integrated e-learning platform that combines: a Unity-based multi-userVR classroom, a web-accessible 360° video learning module,and a hybrid Convolution Neural Network-Deep NeuralNetwork (CNN-DNN) FER service that fuses appearancecues from grayscale face images with geometric descriptorsfrom facial landmarks. We describe a proof-of-conceptimplementation and evaluate the FER component on publicdatasets (Cohn-Kanade dataset (CK+), Japanese FemaleFacial Expression (JAFFE), University of Oulu-Institute ofAutomation, Chinese Academy of Sciences (OULU-CASIA))using subject-independent splits, and we report anexploratory pilot study with 65 university instructors whoexperienced a short VR lecture scenario with FER-drivenavatar facial animation and provided usability andperception feedback. The fused CNN-DNN achieved 100%accuracy on CK+ (6- and 7-class settings), 88.89% onJAFFE (6 classes), and 81.25% on OULU-CASIA(6 classes), while cross-dataset performance dropped to47.89% in the OULU→JAFFE setting, highlightingremaining generalization challenges. In the pilot study, 74%of instructors agreed that the VR scenario supportedemotional communication, and 72% reported that avatarfacial cues helped them notice confusion or engagement,although 93% also identified substantial integrationbarriers related to cost, infrastructure, and training. Theseresults support the technical feasibility of the platform whileclarifying the practical and ethical conditions required forresponsible affect-aware immersive learning. Keywords—virtual reality, 360° video, e-learning, facial expression recognition, deep learning, Convolution Neural Network-Deep Neural Network (CNN-DNN), affective computing
Cite: Anass Touima and Mohamed Moughit, "Hybrid CNN-DNN Facial Expression Recognition in an Affective Virtual Reality (VR) and 360° E-learning Platform," Journal of Image and Graphics, Vol. 14, No. 4, pp. 670-686, 2026.
Copyright © 2026 by the authors. This is an open access article distributed under the Creative Commons Attribution License which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited (CC BY 4.0).