طبقه بندی ریزدانه دیداری با استفاده از اتوانکدرکانولوشنال تنک

نوع مقاله : مقاله پژوهشی

نویسندگان

1 دانشجوی دکتری، دانشکده مهندسی برق و کامپیوتر، دانشگاه آزاد اسلامی، واحد علوم و تحقیقات، تهران، ایران

2 دانشیار، دانشکده مهندسی برق و کامپیوتر، دانشگاه آزاد اسلامی، واحدعلوم و تحقیقات، تهران، ایران

3 استاد، دانشکده مهندسی برق، دانشگاه صنعتی شریف، تهران، ایران.

چکیده
مقدمه: در این تحقیق یک معماری متفاوت چند لایه بر مبنای یادگیری لغت ­نامه و کدگذاری کانولوشن بنام شبکۀ کد گذاری تنک جهت ارتقاء دقت طبقه ­بندی ریزدانه معرفی است. طبقه­ بندی ریزدانه به دلیل تنوع درون­کلاسی و شباهت­ های ظریف بین­ کلاسی یکی از چالش ­برانگیزترین مأموریت­های این حوزه محسوب شده و نیازمند ویژگی­ های تفکیک ­کننده می ­باشد.
روش: برای رسیدن به نتایج مطلوب در این نوع طبقه ­بندی، مدل­ های مبتنی بر شبکه ­­های عصبی عمیق به گونه­ای طراحی می­شوند که بتوانند با مکان­یابی بخش­های حاوی اطلاعات مهم در سطح تصویر و در سطح ناحیه، ویژگی ­های تفکیک ­کننده را استخراج و در طبقه ­بندی استفاده نمایند. مدل پیشنهادی قابلیت استخراج ویژگی­­های سطح بالا را همچون شبکه ­های عصبی عمیق کانولوشنال دارد و از طرفی همزمان از افزونگی اطلاعات در عمق لایه ­ها جلوگیری می­کند. این دو عامل منجر به نوعی مکانیابی ضمنی می­شود که استخراج ویژگی­های تفکیک­ کننده از این نواحی می­تواند در مأموریت ریزدانه موثر و کارآمد باشد.
یافته ­ها: ما در این پژوهش بر مبنای معماری یاد شده، یک اتوانکدر تنک معرفی کرده­ایم که ترکیب آن با یک شبکۀ عمیق عصبی از پیش آموزش داده شده استاندارد، دقت طبقه ­بندی در مجموعه داده­ های ریزدانه را بهبود بخشید.
نتیجه ­گیری: مدل ترکیبی پیشنهادی دقت طبقه­ بندی را به میزان 5/3 درصد در مجموعه دادۀ CUB-200 نسبت به مدل پایۀ مورد استفاده در این معماری افزایش می­ دهد و نرخ خطای تشخیص کلاس­های نادرست را در این مجموعه داده به میزان 28 درصد کاهش می­دهد.

کلیدواژه‌ها

موضوعات

عنوان مقاله English

Fine-Grained Visual Classification Using a Convolutional Sparse Auto-Encoder

نویسندگان English

Farhad Sadeghi Almalou 1
Farbod Razzazi 2
Arash Amini 3
1 Ph.D. Student, Department of Electrical and Computer Engineering, SR.C, Islamic Azad University, Tehran, Iran.
2 Associate Professor, Department of Electrical and Computer Engineering, SR.C, Islamic Azad University, Tehran, Iran.
3 Professor, Department of Electrical Engineering, Sharif University of Technology, Tehran, Iran
چکیده English

In this study, we propose a novel multi-layer architecture, termed the Convolutional Sparse Network (CSN), designed to enhance fine-grained classification accuracy. This architecture is founded upon Dictionary Learning (DL) and Convolutional Sparse Coding (CSC). Fine-grained classification is considered one of the most challenging tasks in this domain due to significant intra-class variance and subtle inter-class similarities, necessitating the extraction of highly discriminative features. To achieve optimal performance, deep neural network-based models are typically engineered to localize informative regions at both the image and part levels, thereby extracting the discriminative features essential for precise classification. The proposed model maintains the capacity for high-level feature extraction, analogous to deep convolutional neural networks, while simultaneously mitigating information redundancy across the layer depth. These dual mechanisms facilitate a form of implicit localization, enabling the extraction of discriminative features that are highly effective for fine-grained tasks. Within this framework, we introduce a Sparse Autoencoder (SAE) which, when integrated with a standard pre-trained deep neural network, improves classification accuracy on fine-grained datasets. Experimental results demonstrate that the proposed hybrid model increases classification accuracy on the CUB-200 dataset by 3.5% compared to the base model and achieves a 28% reduction in the misclassification error rate.

کلیدواژه‌ها English

Convolutional sparse network
Multi-layer dictionary learning
Feature extraction
Fine-grained visual classification
Sparse auto encoder
 
[1]    C. Wah, S. Branson, P. Welinder, P. Perona, and S. Belongie, “The Caltech-UCSD Birds-200-2011 Dataset,” 2011.
[2]    S. Maji, E. Rahtu, J. Kannala, M. Blaschko, and A. Vedaldi, “Fine-Grained Visual Classification of Aircraft,” Jun. 2013, [Online]. Available: http://arxiv.org/abs/1306.5151
[3]    A. Khosla, N. Jayadevaprakash, B. Yao, and F.-F. Li, “Novel dataset for fine-grained image categorization,” CVPR Work. Fine-Grained Vis. Categ., pp. 806–813, 2011.
[4]    J. Krause, M. Stark, J. Deng, and L. Fei-Fei, “3D object representations for fine-grained categorization,” in Proceedings of the IEEE International Conference on Computer Vision, 2013, pp. 554–561. doi: 10.1109/ICCVW.2013.77.
[5]    K. Simonyan and A. Zisserman, “Very Deep Convolutional Networks for Large-Scale Image Recognition,” arXiv:1409.1556v6 [cs.CV], pp. 1–14, Sep. 2014, [Online]. Available: http://arxiv.org/abs/1409.1556
[6]    K. Zhang, M. Sun, S. Member, and T. X. Han, “Residual Networks of Residual Networks : Multilevel Residual Networks,” vol. 14, no. 8, pp. 1–12, 2016.
[7]    T. Y. Lin, A. Roychowdhury, and S. Maji, “Bilinear Convolutional Neural Networks for Fine-Grained Visual Recognition,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 40, no. 6, pp. 1309–1322, 2018, doi: 10.1109/TPAMI.2017.2723400.
[8]    Y. Xie, Q. Gong, X. Luan, J. Yan, and J. Zhang, “A Survey of Fine-Grained Visual Categorization Based on Deep Learning,” J. Syst. Eng. Electron., vol. 35, no. 6, pp. 1337–1356, 2023, doi: 10.23919/jsee.2022.000155.
[9]    W. Luo, H. Zhang, J. Li, and X. S. Wei, “Learning semantically enhanced feature for fine-grained image classification,” IEEE Signal Process. Lett., vol. 27, pp. 1545–1549, 2020, doi: 10.1109/LSP.2020.3020227.
[10]  F. V. Categorization et al., “Attentional Kernel Encoding Networks for Fine-Grained Visual Categorization,” IEEE Trans. Circuits Syst. Video Technol., vol. 31, no. 1, pp. 301–314, 2021, doi: 10.1109/TCSVT.2020.2978115.
[11]  P. Zhuang, Y. Wang, and Y. Qiao, “Learning attentive pairwise interaction for fine-grained classification,” AAAI 2020 - 34th AAAI Conf. Artif. Intell., no. Bruner, pp. 13130–13137, 2020, doi: 10.1609/aaai.v34i07.7016.
[12]  G. Sun, H. Cholakkal, S. Khan, F. S. Khan, and L. Shao, “Fine-grained recognition: Accounting for subtle differences between similar classes,” AAAI 2020 - 34th AAAI Conf. Artif. Intell., pp. 12047–12054, 2020, doi: 10.1609/aaai.v34i07.6882.
[13]  C. Liu, H. Xie, Z. Zha, L. Ma, L. Yu, and Y. Zhang, “Filtration and Distillation : Enhancing Region Attention for Fine-Grained Visual Categorization,” Proc. AAAI Conf. Artif. Intell. Vol. 34. No. 07. 2020., 2020, [Online]. Available: https://scholar.google.com/scholar?cluster=11663869378030505297&hl=en&as_sdt=0,5
[14]  R. Du, J. Xie, Z. Ma, D. Chang, Y.-Z. Song, and J. Guo, “Progressive Learning of Category-Consistent Multi-Granularity Features for Fine-Grained Visual Classification,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 44, no. 12, pp. 9521–9535, Dec. 2022, doi: 10.1109/TPAMI.2021.3126668.
[15]  M. Wang, P. Zhao, X. Lu, F. Min, and X. Wang, “Fine-Grained Visual Categorization: A Spatial-Frequency Feature Fusion Perspective,” IEEE Trans. Circuits Syst. Video Technol., vol. 33, no. 6, pp. 2798–2812, 2023, doi: 10.1109/TCSVT.2022.3227737.
[16]  Y. Pu, Y. Han, Y. Wang, J. Feng, C. Deng, and G. Huang, “Fine-Grained Recognition with Learnable Semantic Data Augmentation,” IEEE Trans. Image Process., vol. 33, pp. 3130–3144, 2024, doi: 10.1109/TIP.2024.3364500.
[17]  Z. Miao, X. Zhao, J. Wang, Y. Li, and H. Li, “Complemental Attention Multi-Feature Fusion Network for Fine-Grained Classification,” IEEE Signal Process. Lett., vol. 28, pp. 1983–1987, 2021, doi: 10.1109/LSP.2021.3114622.
[18]  K. L. and L. F.-F. Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, “ImageNet: A Large-Scale Hierarchical Image Database,” IEEE Conf. Comput. Vis. pattern recognition. Ieee, 2009, no. IEEE, pp. 248–255, 2009.
[19]  A. Lou, S. Guan, and M. Loew, “CFPNet-M: A light-weight encoder-decoder based network for multimodal biomedical image real-time segmentation,” Comput. Biol. Med., vol. 154, 2023, doi: 10.1016/j.compbiomed.2023.106579.
[20]  N. Ibtehaz and M. S. Rahman, “MultiResUNet: Rethinking the U-Net architecture for multimodal biomedical image segmentation,” Neural Networks, vol. 121, pp. 74–87, 2020, doi: 10.1016/j.neunet.2019.08.025.
[21]  H. Xie, Z. Chen, F. Hong, and Z. Liu, “CityDreamer: Compositional Generative Model of Unbounded 3D Cities,” 2024 IEEE/CVF Conf. Comput. Vis. Pattern Recognit., pp. 9666–9675, Sep. 2023, doi: 10.1109/CVPR52733.2024.00923.
[22]  P.-Y. Chou, Y.-Y. Kao, and C.-H. Lin, “Fine-grained Visual Classification with High-temperature Refinement and Background Suppression,” no. March, pp. 1–9, 2023, [Online]. Available: http://arxiv.org/abs/2303.06442
[23]  L. Liu et al., “Computing Systems for Autonomous Driving: State of the Art and Challenges,” IEEE Internet Things J., vol. 8, no. 8, pp. 6469–6486, 2021, doi: 10.1109/JIOT.2020.3043716.
[24]  G. Sen Xie, X. Y. Zhang, S. Yan, and C. L. Liu, “Hybrid CNN and Dictionary-Based Models for Scene Recognition and Domain Adaptation,” IEEE Trans. Circuits Syst. Video Technol., vol. 27, no. 6, pp. 1263–1274, 2017, doi: 10.1109/TCSVT.2015.2511543.
[25]  Y. LeCun, J. S. Denker, and S. A. Solla, “Optimal Brain Damage (Pruning),” Adv. Neural Inf. Process. Syst., pp. 598–605, 1990.
[26]  H. Tang et al., “CSC-Unet: A Novel Convolutional Sparse Coding Strategy Based Neural Network for Semantic Segmentation,” IEEE Access, vol. 12, no. January, pp. 35844–35854, 2024, doi: 10.1109/ACCESS.2024.3373619.
[27]  H. Cunningham, A. Ewart, L. Riggs, R. Huben, and L. Sharkey, “Sparse Autoencoders Find Highly Interpretable Features in Language Models,” 12th Int. Conf. Learn. Represent. ICLR 2024, pp. 1–20, 2024.
[28]  A. Makelov, G. Lange, and N. Nanda, “Towards Principled Evaluations of Sparse Autoencoders for Interpretability and Control,” no. 1, 2024, [Online]. Available: http://arxiv.org/abs/2405.08366
[29]  Z. Zhao et al., “Deep Convolutional Sparse Coding Networks for Interpretable Image Fusion,” IEEE Comput. Soc. Conf. Comput. Vis. Pattern Recognit. Work., vol. 2023-June, pp. 2369–2377, 2023, doi: 10.1109/CVPRW59228.2023.00234.
[30]  X. Sun, N. M. Nasrabadi, and T. D. Tran, “Supervised deep sparse coding networks,” Proc. - Int. Conf. Image Process. ICIP, pp. 346–350, 2018, doi: 10.1109/ICIP.2018.8451701.
[31]  V. Papyan, J. Sulam, and M. Elad, “Working Locally Thinking Globally - Part II: Stability and Algorithms for Convolutional Sparse Coding,” pp. 1–13, 2016, [Online]. Available: http://arxiv.org/abs/1607.02009
[32]  V. Papyan, J. Sulam, and M. Elad, “Working locally thinking globally: Theoretical guarantees for convolutional sparse coding,” IEEE Trans. Signal Process., vol. 65, no. 21, pp. 5687–5701, 2017, doi: 10.1109/TSP.2017.2733447.
[33]  J. Mairal, F. Bach, and J. Ponce, “Sparse modeling for image and vision processing,” Found. Trends Comput. Graph. Vis., vol. 8, no. 2–3, pp. 85–283, 2014, doi: 10.1561/0600000058.
[34]  H. Bristow, A. Eriksson, and S. Lucey, “Fast convolutional sparse coding,” Proc. IEEE Comput. Soc. Conf. Comput. Vis. Pattern Recognit., no. 2, pp. 391–398, 2013, doi: 10.1109/CVPR.2013.57.
[35]  M. Elad, Sparse and Redundant Representations. New York, NY: Springer New York, 2010. doi: 10.1007/978-1-4419-7011-4.
[36]  W. Wen, C. Wu, Y. Wang, Y. Chen, and H. Li, “Learning structured sparsity in deep neural networks,” Adv. Neural Inf. Process. Syst., no. Nips, pp. 2082–2090, 2016.
[37]  Y. Zhang et al., “Advancing Model Pruning via Bi-level Optimization,” Adv. Neural Inf. Process. Syst., vol. 35, no. NeurIPS, pp. 1–18, 2022.
[38]  J. Frankle and M. Carbin, “The lottery ticket hypothesis: Finding sparse, trainable neural networks,” 7th Int. Conf. Learn. Represent. ICLR 2019, pp. 1–42, 2019.
[39]  S. Liu and Z. Wang, “Ten Lessons We Have Learned in the New ‘Sparseland’: A Short Handbook for Sparse Neural Network Researchers,” pp. 1–16, 2023, [Online]. Available: http://arxiv.org/abs/2302.02596
[40]  Y. He and L. Xiao, “Structured Pruning for Deep Convolutional Neural Networks: A Survey,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 46, no. 5, pp. 2900–2919, 2024, doi: 10.1109/TPAMI.2023.3334614.
[41]  S. Rajamanoharan et al., “Improving Dictionary Learning with Gated Sparse Autoencoders,” pp. 1–37, Apr. 2024, [Online]. Available: http://arxiv.org/abs/2404.16014
[42]  V. Monga, Y. Li, and Y. C. Eldar, “Algorithm Unrolling: Interpretable, Efficient Deep Learning for Signal and Image Processing,” IEEE Signal Process. Mag., vol. 38, no. 2, pp. 18–44, 2021, doi: 10.1109/MSP.2020.3016905.
[43]  H. Sreter and R. Giryes, “Learned Convolutional Sparse Coding,” ICASSP, IEEE Int. Conf. Acoust. Speech Signal Process. - Proc., vol. 2018-April, pp. 2191–2195, 2018, doi: 10.1109/ICASSP.2018.8462313.
[44]  K. Gregor and Y. LeCun, “Learning Fast Approximations of Sparse Coding,” Proc. 27th Int. Confer- ence Mach. Learn. Haifa, Isr. 2010., pp. 399–406, May 2010.
[45]  F. Gao, X. Deng, M. Xu, J. Xu, and P. L. Dragotti, “Multi-Modal Convolutional Dictionary Learning,” IEEE Trans. Image Process., vol. 31, pp. 1325–1339, 2022, doi: 10.1109/TIP.2022.3141251.
[46]  T. Liu, A. Chaman, D. Belius, and I. Dokmanic, “Learning Multiscale Convolutional Dictionaries for Image Reconstruction,” IEEE Trans. Comput. Imaging, vol. 8, pp. 425–437, 2022, doi: 10.1109/TCI.2022.3175309.
[47]  C. Garcia-Cardona and B. Wohlberg, “Convolutional Dictionary Learning: A Comparative Review and New Algorithms,” IEEE Trans. Comput. Imaging, vol. 4, no. 3, pp. 366–381, Sep. 2017, doi: 10.1109/TCI.2018.2840334.
[48]  B. Wohlberg, “Convolutional sparse representation of color images,” in Proceedings of the IEEE Southwest Symposium on Image Analysis and Interpretation, 2016, pp. 57–60. doi: 10.1109/SSIAI.2016.7459174.
[49]  B. Wohlberg, “Efficient algorithms for convolutional sparse representations,” IEEE Trans. Image Process., vol. 25, no. 1, pp. 301–315, 2016, doi: 10.1109/TIP.2015.2495260.
[50]  Robert Tibshirani, “Regression Shrinkage and Selection via the Lasso,” J. R. Stat. Soc. Ser. B, vol. 58, no. 1, pp. 267–288, 1996, [Online]. Available: https://www.jstor.org/stable/2346178
[51]  A. Aberdam, J. Sulam, and M. Elad, “Multi-Layer Sparse Coding: The Holistic Way,” SIAM J. Math. Data Sci., vol. 1, no. 1, pp. 46–77, Apr. 2019, doi: 10.1137/18m1183352.
[52]  J. Sulam, A. Aberdam, A. Beck, and M. Elad, “On Multi-Layer Basis Pursuit, Efficient Algorithms and Convolutional Neural Networks,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 42, no. 320649, pp. 1–13, 2019, doi: 10.1109/TPAMI.2019.2904255.
[53]  V. Papyan, Y. Romano, J. Sulam, and M. Elad, “Theoretical foundations of deep learning via sparse representations: A multilayer sparse model and its connection to convolutional neural networks,” IEEE Signal Process. Mag., vol. 35, no. 4, pp. 72–89, 2018, doi: 10.1109/MSP.2018.2820224.
[54]  J. Sulam, V. Papyan, Y. Romano, and M. Elad, “Multilayer convolutional sparse modeling: Pursuit and dictionary learning,” IEEE Trans. Signal Process., vol. 66, no. 15, pp. 4090–4104, 2018, doi: 10.1109/TSP.2018.2846226.
[55]  V. Papyan, Y. Romano, and M. Elad, “Convolutional neural networks analyzed via convolutional sparse coding,” J. Mach. Learn. Res., vol. 18, no. 1, pp. 1–52, 2017.
[56]  O. Ronneberger, P. Fischer, and T. Brox, “U-Net: Convolutional Networks for Biomedical Image Segmentation,” IEEE Access, vol. 9, pp. 16591–16603, May 2015, doi: 10.1109/ACCESS.2021.3053408.
[57]  F. S. Almalou, F. Razzazi;, and A. Amini, “A Multi‐Layer Convolutional Sparse Network for Pattern Classification Based on Sequential Dictionary Learning,” IET Comput. Vis., vol. 20, no. 1, 2026, doi: https://doi.org/10.1049/cvi2.70055.
[58]  S. Chen and D. . Donoho, “Basis Pursuit,” Proc. 1994 28th Asilomar Conf. Signals, Syst. Comput., pp. 41–44, 1994.
[59]  S. Boyd, N. Parikh, E. Chu, B. Peleato, and J. Eckstein, “Distributed Optimization and Statistical Learning via the Alternating Direction Method of Multipliers,” Found. Trends® Mach. Learn., vol. 3, no. 1, pp. 1–122, 2010, doi: 10.1561/2200000016.
[60]  O. Russakovsky et al., “ImageNet Large Scale Visual Recognition Challenge,” Int. J. Comput. Vis., vol. 115, no. 3, pp. 211–252, 2015, doi: 10.1007/s11263-015-0816-y.
[61]  M. Tan and Q. V. Le, “EfficientNet: Rethinking model scaling for convolutional neural networks,” 36th Int. Conf. Mach. Learn. ICML 2019, vol. 2019-June, pp. 10691–10700, 2019.
[62]  A. Dubey, O. Gupta, R. Raskar, and N. Naik, “Maximum Entropy Fine-Grained Classification,” 32nd Conf. Neural Inf. Process. Syst. (NeurIPS 2018), Montréal, Canada, no. iii, 2018.
[63]  R. Wightman and H. Touvron, “ResNet strikes back : An improved training procedure in timm,” pp. 1–22, 2021.
[64]  Y. Wang, V. I. Morariu, and L. S. Davis, “Learning a Discriminative Filter Bank Within a CNN for Fine-Grained Recognition,” Proc. IEEE Comput. Soc. Conf. Comput. Vis. Pattern Recognit., pp. 4148–4157, 2018, doi: 10.1109/CVPR.2018.00436.
[65]  X. He, Y. Peng, and J. Zhao, “Fast Fine-Grained Image Classification via Weakly Supervised Discriminative Localization,” IEEE Trans. Circuits Syst. Video Technol., vol. 29, no. 5, pp. 1394–1407, May 2019, doi: 10.1109/TCSVT.2018.2834480.
[66]  Q. Xu, S. Li, J. Wang, B. Jiang, and J. Tang, “Context-Semantic Quality Awareness Network for Fine-Grained Visual Categorization,” pp. 1–13, Mar. 2024, [Online]. Available: http://arxiv.org/abs/2403.10298
[67]  Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner, “Gradient-based learning applied to document recognition,” Proc. IEEE, vol. 86, no. 11, pp. 2278–2323, 1998, doi: 10.1109/5.726791.
[68]  K. He, X. Zhang, S. Ren, and J. Sun, “Deep Residual Learning for Image Recognition,” Front. Psychol., vol. 4, no. APR, pp. 428–429, Dec. 2015, doi: 10.3389/fpsyg.2013.00124.
[69]  G. Huang, Z. Liu, L. Van Der Maaten, and K. Q. Weinberger, “Densely connected convolutional networks,” Proc. - 30th IEEE Conf. Comput. Vis. Pattern Recognition, CVPR 2017, vol. 2017-Janua, pp. 2261–2269, 2017, doi: 10.1109/CVPR.2017.243.
دوره 5، شماره 2
تابستان 1405
صفحه 1-27

  • تاریخ دریافت 11 اسفند 1404
  • تاریخ بازنگری 01 اردیبهشت 1405
  • تاریخ پذیرش 05 مرداد 1405