مجله ماشین بینایی و پردازش تصویر

مجله ماشین بینایی و پردازش تصویر

تطبیق صحنه با ترکیب SIFT و ORB : بهبود کارایی و مقاوم‌سازی در پردازش تصویر

نوع مقاله : مقاله پژوهشی

نویسندگان
1 گروه مهندسی برق ،واحد اصفهان (خوراسگان)، دانشگاه آزاد اسلامی، اصفهان، ایران.
2 گروه مهندسی برق ،واحد اصفهان (خوراسگان)، دانشگاه آزاد اسلامی، اصفهان، ایران
3 گروه مهندسی کامپیوتر ، دانشکده مهندسی، دانشگاه بوعلی سینا، همدان، ایران. پژوهشکده علوم کامپیوتر، پژوهشگاه دانشهای بنیادی (IPM)،
چکیده
الگوریتم‌های تطبیق ویژگی زیربنای کلیدی بسیاری از کاربردهای بینایی ماشین به شمار می‌روند. با توجه به تنوع زیاد زمینه‌های کاربردی این الگوریتم‌ها و دامنه وسیع تغییرات بصری تصاویر، ارزیابی عملکرد الگوریتم‌ها یک جنبه نسبی داشته و انتخاب الگوریتم مناسب به تعادل بین دقت، سرعت، مقاومت و منابع محاسباتی بستگی دارد. در این تحقیق تلاش خواهد شد که ساختار الگوریتم‌های تطبیق کلاسیک و مبتنی بر یادگیری تحلیل شده و برخی از چالش‌های این الگوریتم‌ها بررسی شود. همچنین تلفیقی از این الگوریتم‌های SIFT و ORB در راستای بهینه‌سازی زمان پردازش در عین حفظ دقت تطبیق ارائه گردیده و SCO نامگذاری می‌شود. به منظور تحلیل عملکرد الگوریتم‌های تطبیق، در یک سخت‌افزار دارای قابلیت پردازشی لازم به صورت بهینه شبیه‌سازی شده و عملکرد الگوریتم‌ها در رابطه با تصاویر چالش برانگیز مجموعه داده اکسفورد و با معیارهای استاندارد ارزیابی می‌گردد. نتایج ارزیابی نشان می‌دهد که زمان اجرای روشSCO به طور متوسط حدود 0.28 زمان اجرای SIFT و 2.67 برابر زمان اجرای ORBاست در حالیکه تعداد تطبیق‌های درستSCO به طور متوسط حدود 0.73 تطبیق‌های صحیح SIFT و حدود 3.3 برابر تطبیق‌های ORB می‌باشد.
کلیدواژه‌ها

[1]           R. Arandjelovic, P. Gronat, A. Torii, T. Pajdla, and J. Sivic, "NetVLAD: CNN architecture for weakly supervised place recognition," in Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 5297–5307.
[2]           H. Noh, A. Araujo, J. Sim, T. Weyand, and B. Han, "Large-scale image retrieval with attentive deep local features," in Proceedings of the IEEE international conference on computer vision, 2017, pp. 3456–3465.
[3]           S. Agarwal et al., "Building rome in a day," Communications of the ACM, vol. 54, no. 10, pp. 105–112, 2011.
[4]           S. Zhu et al., "Very large-scale global sfm by distributed motion averaging," in Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 4568–4577.
[5]           L. Cui and C. Ma, "SDF-SLAM: Semantic depth filter SLAM for dynamic environments," IEEE Access, vol. 8, pp. 95301–95311, 2020.
[6]           Y. Wang, Q. Huang, C. Jiang, J. Liu, M. Shang, and Z. Miao, "Video stabilization: A comprehensive survey," Neurocomputing, vol. 516, pp. 205–230, 2023.
[7]           Y. Wang, Q. Huang, J. Liu, C. Jiang, and M. Shang, "Adaptive video stabilization based on feature point detection and full-reference stability assessment," Multimedia Tools and Applications, vol. 83, no. 11, pp. 32497–32524, 2024.
[8]           Z. Zhang, Y. Zhang, Z. Shi, and S. Huang, "Image matching based on region template correlation," Infrared Laser Eng, vol. 33, no. 6, pp. 615–618, 2004.
9]            Z. Diao, G. Peng, and L. Fu, "Fast high-precision reliable image matching algorithm based on shape," J. Comput. Appl, vol. 30, no. 2, p. 441, 2010.
[10]         Q. Huang, X. Guo, Y. Wang, H. Sun, and L. Yang, "A survey of feature matching methods," IET Image Processing, vol. 18, no. 6, pp. 1385–1410, 2024.
[11]         H. P. Moravec, "Rover Visual Obstacle Avoidance," in IJCAI, 1981, vol. 81, pp. 785–790.
[12]         C. Schimd and R. Mohr, "Local gray value invariants for image retrieval," IEEE Trans. on Pattern Analysis and Machine Intelligence, pp. 530–534, 1997.
[13]         D. G. Lowe, "Object recognition from local scale-invariant features," in Proceedings of the seventh IEEE international conference on computer vision, 1999, vol. 2: Ieee, pp. 1150–1157.
[14]         H. Bay, T. Tuytelaars, and L. Van Gool, "Surf: Speeded up robust features," in European conference on computer vision, 2006: Springer, pp. 404–417.
[15]         J.-M. Morel and G. Yu, "ASIFT: A new framework for fully affine invariant image comparison," SIAM journal on imaging sciences, vol. 2, no. 2, pp. 438–469, 2009.
[16]         C. Lin, J. Liu, and L. Cao, "Image matching by affine speed-up robust features," in MIPPR 2011: Pattern Recognition and Computer Vision, 2011, vol. 8004: SPIE, pp. 120–124.
[17]         E. Rublee, V. Rabaud, K. Konolige, and G. Bradski, "ORB: An efficient alternative to SIFT or SURF," in 2011 International conference on computer vision, 2011: Ieee, pp. 2564–2571.
[18]         S. Leutenegger, M. Chli, and R. Y. Siegwart, "BRISK: Binary robust invariant scalable keypoints," in 2011 International conference on computer vision, 2011: Ieee, pp. 2548–2555.
[19]         P. F. Alcantarilla, A. Bartoli, and A. J. Davison, "KAZE features," in European conference on computer vision, 2012: Springer, pp. 214–227.
[20]         D. G. Lowe, "Distinctive image features from scale-invariant keypoints," International journal of computer vision, vol. 60, no. 2, pp. 91–110, 2004.
[21]         D. DeTone, T. Malisiewicz, and A. Rabinovich, "Superpoint: Self-supervised interest point detection and description," in Proceedings of the IEEE conference on computer vision and pattern recognition workshops, 2018, pp. 224–236.
[22]         P. F. Alcantarilla and T. Solutions, "Fast explicit diffusion for accelerated features in nonlinear scale spaces," IEEE Trans. Patt. Anal. Mach. Intell, vol. 34, no. 7, pp. 1281–1298, 2011.
[23]         M. Pultar, "Improving the hardnet descriptor," arXiv preprint arXiv:2007.09699, 2020.
[24]         M. A. Fischler and R. C. Bolles, "Random sample consensus: a paradigm for model fitting with applications to image analysis and automated cartography," Communications of the ACM, vol. 24, no. 6, pp. 381–395, 1981.
[25]         D. Barath, J. Noskova, M. Ivashechkin, and J. Matas, "MAGSAC++, a fast, reliable and accurate robust estimator," in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 1304–1312.
[26]         R. Hartley and A. Zisserman, Multiple view geometry in computer vision. Cambridge university press, 2003.
[27]         P. Lindenberger, P.-E. Sarlin, and M. Pollefeys, "Lightglue: Local feature matching at light speed," in Proceedings of the IEEE/CVF international conference on computer vision, 2023, pp. 17627–17638.
[28]         U. o. O. Visual Geometry Group. "Affine Image Transformations." http://www.robots.ox.ac.uk/vgg/affine.html (accessed October 10, 2025).