2026-06-04
2026-04-30
2026-02-27
Manuscript received February 27, 2026; revised March 18, 2026; accepted April 21, 2026; published September 10, 2026.
Abstract—Object recognition in low-resolution images remains a challenging task due to limited spatial detail and reduced discriminative features. While Convolutional Neural Networks (CNNs) effectively capture local patterns, they are limited in modeling long-range dependencies. In contrast, Vision Transformers (ViTs) provide strong global contextual modeling but often require higher computational resources. This study proposes an adaptive augmentation-based hybrid CNN-vision Transformer architecture designed to enhance feature representation for small-resolution image classification. The proposed method integrates hierarchical convolutional feature extraction with transformer-based contextual encoding and introduces an adaptive data augmentation mechanism that applies adaptive augmentation strategies based on training feedback. Experimental results on a controlled synthetic benchmark dataset demonstrate that the proposed model achieves stable convergence and consistent classification performance, reaching an accuracy of 89.15%. The results indicate that the integration of local and global feature modeling improves representation capability under constrained resolution conditions. However, the evaluation is limited to a controlled dataset, and further validation on diverse real-world datasets is required to fully assess generalization performance. Overall, the proposed approach provides a promising direction for improving classification performance in low-resolution image recognition tasks. Keywords—small-resolution object recognition, hybrid Convolutional Neural Networks (CNN)-vision Transformer, adaptive data augmentation, image classification, deep learning Cite: Syarifah Fadillah Rezky, Muhammad Dahria, Siti Julianita Siregar, Zaimah Panjaitan, and Rita Hamdani, "Adaptive Augmentation-based Hybrid CNN-vision Transformer for Small-resolution Object Recognition," Journal of Image and Graphics, Vol. 14, No. 5, pp. 747-755, 2026.
Copyright © 2026 by the authors. This is an open access article distributed under the Creative Commons Attribution License which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited (CC BY 4.0).