Fine-Grained Visual Classification Using a Convolutional Sparse Auto-Encoder

Document Type : Original Article

Authors

1 Department of Mechanical, Electrical and Computer Engineering, SR.C, Islamic Azad University, Tehran, Iran

2 Department of Electrical Engineering, Sharif University of Technology, Tehran, Iran.

Abstract
In this study, we propose a novel multi-layer architecture, termed the Convolutional Sparse Network (CSN), designed to enhance fine-grained classification accuracy. This architecture is founded upon Dictionary Learning (DL) and Convolutional Sparse Coding (CSC). Fine-grained classification is considered one of the most challenging tasks in this domain due to significant intra-class variance and subtle inter-class similarities, necessitating the extraction of highly discriminative features. To achieve optimal performance, deep neural network-based models are typically engineered to localize informative regions at both the image and part levels, thereby extracting the discriminative features essential for precise classification. The proposed model maintains the capacity for high-level feature extraction, analogous to deep convolutional neural networks, while simultaneously mitigating information redundancy across the layer depth. These dual mechanisms facilitate a form of implicit localization, enabling the extraction of discriminative features that are highly effective for fine-grained tasks. Within this framework, we introduce a Sparse Autoencoder (SAE) which, when integrated with a standard pre-trained deep neural network, improves classification accuracy on fine-grained datasets. Experimental results demonstrate that the proposed hybrid model increases classification accuracy on the CUB-200 dataset by 3.5% compared to the base model and achieves a 28% reduction in the misclassification error rate.

Keywords


Volume 5, Issue 2
Summer 2026

  • Receive Date 26 June 2026
  • Revise Date 23 July 2026
  • Accept Date 27 July 2026