مجله ماشین بینایی و پردازش تصویر

مجله ماشین بینایی و پردازش تصویر

تشخیص نمادهای ناهنجار در داده‌های تصویری با استفاده از مدل‌های یادگیری عمیق

نوع مقاله : مقاله پژوهشی

نویسندگان
گروه مهندسی کامپیوتر، دانشکده برق و کامپیوتر، دانشگاه صنعتی همدان
چکیده
با گسترش روزافزون محتوای تصویری در فضای مجازی، نظارت دستی بر نمادهای ناهنجار عملاً غیرممکن شده و نیاز به سامانه‌های هوشمند تشخیص خودکار بیش از پیش احساس می‌شود. در این پژوهش، یک مجموعه‌داده اختصاصی شامل 3659 تصویر از سه نماد «چشم جهان‌بین»، «ستاره واژگون» و «گونیا و پرگار» گردآوری و به‌صورت دقیق برچسب‌گذاری شد. داده‌ها با نسبت 80٪ برای آموزش و 20٪ برای آزمون تقسیم گردیدند و به‌منظور افزایش تنوع بصری و بهبود تعمیم‌پذیری، تکنیک‌های افزایش داده بر مجموعه آموزشی اعمال شد. عملکرد مدل‌های Faster R-CNN، SSD300 و خانواده YOLO (نسخه‌های YOLOv8 تا YOLOv12) مورد ارزیابی و مقایسه قرار گرفت. نتایج نشان داد مدل Faster R-CNN در مجموعه آزمون به مقادیر Precision برابر با 0.901، Recall برابر با 0.909 و mAP50 برابر با 0.939 دست یافته و نسبت به مدل‌های تک‌مرحله‌ای عملکرد برتری ارائه کرده است. یافته‌ها حاکی از آن است که معماری‌های دو‌مرحله‌ای در تشخیص نمادهای دارای جزئیات هندسی ظریف و شباهت ساختاری بالا، از قابلیت تعمیم‌پذیری بهتری برخوردارند. نتایج این پژوهش می‌تواند مبنایی برای توسعه سامانه‌های هوشمند پایش و فیلترینگ خودکار محتوای تصویری در کاربردهای نظارتی و تحلیل رسانه‌ای باشد.
کلیدواژه‌ها

   [1]      S. Routray, A. K. Ray, and C. Mishra, "Analysis of various image feature extraction methods against noisy image: SIFT, SURF and HOG," in 2017 Second International Conference on Electrical, Computer and Communication Technologies (ICECCT), 2017: IEEE, pp. 1–5.
   [2]      K. O'shea and R. Nash, "An introduction to convolutional neural networks," arXiv preprint arXiv:1511.08458, 2015.
   [3]      W. Zhou, S. Gao, L. Zhang, and X. Lou, "Histogram of oriented gradients feature extraction from raw bayer pattern images," IEEE Transactions on Circuits and Systems II: Express Briefs, vol. 67, no. 5, pp. 946–950, 2020.
   [4]      M. A. Chandra and S. Bedi, "Survey on SVM and their application in image classification," International Journal of Information Technology, vol. 13, no. 5, pp. 1–11, 2021.
   [5]      A. Bosch, A. Zisserman, and X. Munoz, "Image classification using random forests and ferns," in 2007 IEEE 11th international conference on computer vision, 2007: Ieee, pp. 1–8.
   [6]      J. Redmon, S. Divvala, R. Girshick, and A. Farhadi, "You only look once: Unified, real-time object detection," in Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 779–788.
   [7]      J. Solawetz. "What is YOLOv8? A Complete Guide." https://blog.roboflow.com/what-is-yolov8/ (accessed Oct 23, 2024.
   [8]      P. Potrimba. "What is New in YOLOv9? An Architecture Deep Dive." Roboflow Blog. https://blog.roboflow.com/yolov9-deep-dive/ (accessed May 20, 2024.
   [9]      P. Potrimba. "What is YOLOv10? An Architecture Deep Dive." https://blog.roboflow.com/what-is-yolov10/ (accessed Jun 14, 2024.
[10]      R. Khanam and M. Hussain, "Yolov11: An overview of the key architectural enhancements," arXiv preprint arXiv:2410.17725, 2024.
[11]      "YOLOv12." https://roboflow.com/model/yolov12 (accessed Feb 18, 2025.
[12]      Y. Tian, Q. Ye, and D. Doermann, "Yolov12: Attention-centric real-time object detectors," arXiv preprint arXiv:2502.12524, 2025.
[13]      W. Liu et al., "Ssd: Single shot multibox detector," in European conference on computer vision, 2016: Springer, pp. 21–37.
[14]      S. Ren, K. He, R. Girshick, and J. Sun, "Faster r-cnn: Towards real-time object detection with region proposal networks," Advances in neural information processing systems, vol. 28, 2015.
[15]      J. Hui. "SSD object detection: Single Shot MultiBox Detector for real-time processing." Medium. https://jonathan-hui.medium.com/ssd-object-detection-single-shot-multibox-detector-for-real-time-processing-9bd8deac0e06 (accessed Mar 14, 2018.
[16]      C. Guo, B. Fan, Q. Zhang, S. Xiang, and C. Pan, "Augfpn: Improving multi-scale feature learning for object detection," in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 12595–12604.
[17]      A. Mikołajczyk and M. Grochowski, "Data augmentation for improving deep learning in image classification problem," in 2018 international interdisciplinary PhD workshop (IIPhDW), 2018: IEEE, pp. 117–122.
[18]      Torralba, "Context-based vision system for place and object recognition," in Proceedings Ninth IEEE International Conference on Computer Vision, 2003: IEEE, pp. 273–280 vol. 1.
[19]      J. Li, X. Liang, Y. Wei, T. Xu, J. Feng, and S. Yan, "Perceptual generative adversarial networks for small object detection," in Proceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 1222–1230.
[20]      B. Singh and L. S. Davis, "An analysis of scale invariance in object detection snip," in Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 3578–3587.
[21]      B. Singh, M. Najibi, and L. S. Davis, "Sniper: Efficient multi-scale training," Advances in neural information processing systems, vol. 31, 2018.
[22]      S. Gidaris and N. Komodakis, "Object detection via a multi-region and semantic segmentation-aware cnn model," in Proceedings of the IEEE international conference on computer vision, 2015, pp. 1134–1142.
[23]      X. Zeng, W. Ouyang, B. Yang, J. Yan, and X. Wang, "Gated bi-directional cnn for object detection," in European conference on computer vision, 2016: Springer, pp. 354–369.
[24]      Z. Chen, Q. Xu, R. Cong, and Q. Huang, "Global context-aware progressive aggregation network for salient object detection," in Proceedings of the AAAI conference on artificial intelligence, 2020, vol. 34, no. 07, pp. 10599–10606.
[25]      S. Ren, K. He, R. Girshick, and J. Sun, "Faster R-CNN: Towards real-time object detection with region proposal networks," IEEE transactions on pattern analysis and machine intelligence, vol. 39, no. 6, pp. 1137–1149, 2016.
[26]      M. Sohan, T. Sai Ram, and C. V. Rami Reddy, "A review on yolov8 and its advancements," in International Conference on Data Intelligence and Cognitive Informatics, 2024: Springer, pp. 529–545.
[27]      C.-Y. Wang, I.-H. Yeh, and H.-Y. Mark Liao, "Yolov9: Learning what you want to learn using programmable gradient information," in European conference on computer vision, 2024: Springer, pp. 1–21.
[28]      A. Wang, H. Chen, L. Liu, K. Chen, Z. Lin, and J. Han, "Yolov10: Real-time end-to-end object detection," Advances in Neural Information Processing Systems, vol. 37, pp. 107984–108011, 2024.
[29]      "Pentagram." https://en.wikipedia.org/wiki/Pentagram (accessed.
[30]      "Square and Compasses." https://en.wikipedia.org/wiki/Square_and_Compasses accessed.
[31]      "Eye of Providence." https://en.wikipedia.org/wiki/Eye_of_Providence accessed.
[32]      "Satanism." https://en.wikipedia.org/wiki/Satanism accessed.
[33]      "Sigil of Baphomet." https://en.wikipedia.org/wiki/Sigil_of_Baphomet accessed.