Home > Articles > All Issues > 2026 > Volume 14, No. 4, 2026 >
JOIG 2026 Vol.14(4):706-719
doi: 10.18178/joig.14.4.706-719

Multi-view Integration with View-specific Self-attention for Sign Language Recognition

Hoang Quang Huy and and Tran Anh Vu*
Department of Electronic Engineering School of Electrical and Electronic Engineering (SEEE), Hanoi University of Science and Technology (HUST), Hanoi, Vietnam
Email: huy.hoangquang@hust.edu.vn (H.Q.H.); vu.trananh@hust.edu.vn (T.A.V.) (S.P.P.)
*Corresponding author

Manuscript received July 29, 2025; revised August 22, 2025; accepted November 5, 2025; published July 23, 2026.

Abstract—Sign Language Recognition (SLR) plays a pivotal role in mitigating communication barriers between the deaf community and broader society. Recently, the Video Transformer Network (VTN), an extension of the transformer architecture incorporating a multi-head self-attention mechanism, has demonstrated efficacy in video processing tasks and SLR. However, relying solely on Red Green Blue (RGB)frame information may render the model susceptible to redundant data and environmental factors such as lighting variations and complex backgrounds. Additionally, most existing SLR datasets provide only frontal viewpoints, which constrain the model’s generalizability, particularly in real-world settings where diverse perspectives are prevalent. In this study, we propose Video Transformer Network with 3 Graph Convolutional Network(VTN3GCN), a multi-view and multi-stream framework that integrates RGB data, skeleton coordinates, and pose flow from three distinct viewpoints: left, right, and center. This architecture enhances VTN by incorporating Graph Convolutional Networks (GCN) to learn skeletal frame features and employs an early fusion mechanism between the RGB and skeleton streams. Experiments conducted on the Multi-VSL200 dataset—a newly curated dataset for Vietnamese Sign Language (VSL) featuring three viewpoints per video demonstrate that the VTN3GCN framework achieves a top-1 accuracy of up to 92.92%, achieving state-of-the-art accuracy. Our code and data are available at: https://github.com/fossbk/MultiView-ISLR/tree/main/VTN3 GCN.
 
Keywords—Vietnamese sign language, sign language recognition, deep learning, video transformer network

Cite: Hoang Quang Huy and Tran Anh Vu, "Multi-view Integration with View-specific Self-attention for Sign Language Recognition," Journal of Image and Graphics, Vol. 14, No. 4, pp. 706-719, 2026.


Copyright © 2026 by the authors. This is an open access article distributed under the Creative Commons Attribution License which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited (CC BY 4.0).

Article Metrics in Dimensions