مجله ماشین بینایی و پردازش تصویر

مجله ماشین بینایی و پردازش تصویر

داده افزایی با استفاده از شبکه‌های مولد تخاصمی جهت بهبود باز تشخیص افراد

نوع مقاله : مقاله پژوهشی

نویسندگان
1 دانشجوی دکترای رشته برق - مخابرات، واحد علوم و تحقیقات، دانشگاه آزاد اسلامی، تهران، ایران
2 عضو هیات علمی پژوهشگاه ارتباطات و فناوری اطلاعات
3 دانشکده مهندسی برق و کامپیوتر، واحد علوم و تحقیقات، دانشگاه آزاد اسلامی، تهران، ایران
چکیده
در سال‌های اخیر، پژوهش‌ها در زمینه بازتشخیص افراد به‌طور پیوسته پیشرفت کرده‌اند. در این راستا، شبکه‌های عصبی متخاصم به موفقیت چشمگیری دست یافته‌اند و به عنوان یک رویکرد برجسته در این حوزه شناخته می‌شوند. در این مقاله، با بررسی کاربرد انواع شبکه‌های متخاصم در بازتشخیص افراد، از ترکیب شبکه ATNet و مدل Pix2Pix برای داده‌افزایی استفاده شده است. مدل Pix2Pix که در تبدیل تصویر به تصویر در زمینه‌های مختلف موفقیت‌آمیز بوده، از الگوریتم PatchGAN به عنوان متمایزکننده و U-Net به عنوان تولیدکننده بهره می‌گیرد. روش پیشنهادی برای داده‌افزایی مبتنی بر وجود تصاویر فرد مورد نظر از چهار جهت (روبه‌رو، پشت، سمت چپ و راست) است. پس از دسته‌بندی تصاویر در مجموعه‌داده‌های Market-1501 و CUHK03، تصاویر نماهای ناموجود با استفاده از شبکه ATNet تولید شده‌اند. مقایسه های انجام‌شده نشان‌دهنده بهبود عملکرد روش‌های پیشرو در بازتشخیص افراد با داده افزایی پیشنهادی در این مقاله است.
کلیدواژه‌ها

[1]
Bhuiyan, Md Roman, et al. "Video analytics using deep learning for crowd analysis: a review." Multimedia Tools and Applications 81.19 (2022): 27895-27922.
[2]
Zheng, Zhedong, Liang Zheng, and Yi Yang. "A discriminatively learned cnn embedding for person reidentification." ACM transactions on multimedia computing, communications, and applications (TOMM) 14.1 (2017): 1-20.
[3]
Alqahtani, Hamed, Manolya Kavakli-Thorne, and Gulshan Kumar. "Applications of generative adversarial networks (gans): An updated review." Archives of Computational Methods in Engineering 28 (2021): 525-552
[4]
Creswell, Antonia, et al. "Generative adversarial networks: An overview." IEEE signal processing magazine 35.1 (2018): 53-65.
[5]
Isola, Phillip, et al. "Image-to-image translation with conditional adversarial networks." Proceedings of the IEEE conference on computer vision and pattern recognition. 2017.
[6]
Liu, Ming-Yu, and Oncel Tuzel. "Coupled generative adversarial networks." Advances in neural information processing systems 29 (2016).
[7]
Zhu, Jun-Yan, et al. "Unpaired image-to-image translation using cycle-consistent adversarial networks." Proceedings of the IEEE international conference on computer vision. 2017.
[8]
Cao, Chengzhi, et al. "Event-guided person re-identification via sparse-dense complementary learning." Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2023.
[9]
Asperti, Andrea, Salvatore Fiorilla, and Lorenzo Orsini. "A generative approach to person reidentification." Sensors 24.4 (2024): 1240.
[10]
Chen, Weihua, et al. "Beyond appearance: a semantic controllable self-supervised learning framework for human-centric visual tasks." Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2023.
[11]
Zhong, Zhun, et al. "Camera style adaptation for person re-identification." Proceedings of the IEEE conference on computer vision and pattern recognition. 2018.
[12]
Wei, Longhui, et al. "Person transfer gan to bridge domain gap for person re-identification." Proceedings of the IEEE conference on computer vision and pattern recognition. 2018.
[13]
Liu, Jiawei, et al. "Adaptive transfer network for cross-domain person re-identification." Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2019.
[14]
Liu, Chong, Xiaojun Chang, and Yi-Dong Shen. "Unity style transfer for person re-identification." Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2020.
[15]
Pan, Xingang, et al. "Two at once: Enhancing learning and generalization capacities via ibn-net." Proceedings of the European Conference on Computer Vision (ECCV). 2018.
[16]
Zhang, Taiping, et al. "Face recognition under varying illumination using gradientfaces." IEEE Transactions on image processing 18.11 (2009): 2599-2606.
[17]
Ledig, Christian, et al. "Photo-realistic single image super-resolution using a generative adversarial network." Proceedings of the IEEE conference on computer vision and pattern recognition. 2017.
[18]
Zheng, Liang, et al. "Scalable person re-identification: A benchmark." Proceedings of the IEEE international conference on computer vision. 2015.
[19]
Isola, Phillip, et al. "Image-to-image translation with conditional adversarial networks." Proceedings of the IEEE conference on computer vision and pattern recognition. 2017.
[20]
Zhu, Jun-Yan, et al. "Unpaired image-to-image translation using cycle-consistent adversarial networks." Proceedings of the IEEE international conference on computer vision. 2017.
[21]
Ming, Zhangqiang, et al. "Deep learning-based person re-identification methods: A survey and outlook of recent works." Image and Vision Computing 119 (2022): 104394
[22]
Zheng, Zhedong, Liang Zheng, and Yi Yang. "Unlabeled samples generated by gan improve the person re-identification baseline in vitro." Proceedings of the IEEE international conference on computer vision. 2017.
[23]
Qian, Xuelin, et al. "Pose-normalized image generation for person re-identification." Proceedings of the European conference on computer vision (ECCV). 2018.
[24]
Liu, Jinxian, et al. "Pose transferrable person re-identification." Proceedings of the IEEE conference on computer vision and pattern recognition. 2018.
[25]
Huang, Yan, et al. "Multi-pseudo regularized label for generated data in person re-identification." IEEE Transactions on Image Processing 28.3 (2018): 1391-1403.
[26]
Zheng, Zhedong, et al. "Joint discriminative and generative learning for person re-identification." proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2019.
[27]
Ge, Yixiao, et al. "Fd-gan: Pose-guided feature distilling gan for robust person re-identification." Advances in neural information processing systems 31 (2018).