مجله ماشین بینایی و پردازش تصویر

مجله ماشین بینایی و پردازش تصویر

مروری بر مجموعه داده‌های صوتی و تصویری در بازشناسی تصویری گفتار پیوسته

نوع مقاله : مقاله مروری

نویسندگان
گروه مهندسی کامپیوتر، واحد کرج، دانشگاه آزاد اسلامی، کرج، ایران
چکیده
بررسی مجموعه داده‌ها‌ی لب‌خوانی اولین چالش مهم در زمینه بازشناسی تصویری گفتار پیوسته است. گروهی از محققان برای بازشناسی گفتار و لب‌خوانی به جمع‌آوری مجموعه داده‌هایی جهت ارزیابی الگوریتم‌های پیشنهادی خود پرداخته‌اند. به گونه‌ای که به تناسب نیاز خود تنها برخی از ویژگی‌های داده‌ها را در نظر گرفته‌اند. چرا که داده‌های صوتی تصویری برای بازشناسی گفتار و لب‌خوانی دارای ویژگی‌های متفاوتی چون گفتار پیوسته و یا گفتار گسسته، زبان پایگاه داده‌ها، زاویه تصویربرداری از گویندگان است. محققان جهت پردازش گفتار و آغاز به کار نیاز به آمارهایی در رابطه با پایگاه داده‌های موجود دارند. چالش نداشتن آمار در زمینه داده‌های گفتار پیوسته انگیزه‌ای شد تا در این مقاله مجموعه داده‌های صوتی تصویری گفتار پیوسته معرفی شوند و نحوه جمع‌آوری آنها، تنظیمات ضبط، محیط ضبط و ویژگی‌های اصلی آنها مانند تعداد گویندگان، تعداد تکرار گفتارها، رزولوشن تصویر بررسی ‌شوند. داده‌های مربوط به پایگاه داده‌های موجود بر حسب ویژگی‌های کمی و کیفی آنها دسته‌بندی و متناسب با این ویژگی‌ها آمارهای متفاوتی همچون درصد بومی بودن گویندگان، نسبت جنسیت گویندگان، میانگین سنی گویندگان، تعداد زوایای تصویربرداری از گویندگان و مدت زمان ضبط داده‌ها ارائه شده است. در پایان مزایای پایگاه داده‌های مورد مطالعه به همراه آدرس دسترسی به آنها لیست شده است.
کلیدواژه‌ها

   [1]      مهسا هدایتی‌پور، یاسر شکفته، محسن ابراهیمی‌مقدم، مروری بر پژوهش‌های لب‌خوانی خودکار: دادگان و روش‌ها، مجله ماشین بینایی و پردازش تصویر، سال نهم، شماره چهارم، زمستان ۱۴۰1
   [2]      D. L. Howell, Confusion Modelling for Lip-Reading, Ph.D. thesis, University of East Anglia, 2015.
   [3]      J. S. Chung, A. Senior, O. Vinyals and A. Zisserman, “Lip reading sentences in the wild”, in: Proc. Conference on Computer Vision and Pattern Recognition, 2017, pp. 3444–3453. Son Chung, J.; Senior, A.; Vinyals, O.; Zisserman, A. Lip Reading Sentences in the Wild. In Proceedings of the The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, 21–26 July 2017.
   [4]      E. K. Patterson, S. Gurbuz, Z. Tufekci and J. N. Gowdy, “CUAVE: A new audio-visual database for multimodal human-computer interface research”, in: Proc. International Conference on Acoustics, Speech, and Signal Processing, Vol. 2, 2002, pp. 2017–2020.
   [5]      E. Bailly-Bailliere, S. Bengio, F. Bimbot, M. Hamouz, J. Kittler, J. Mariethoz, J. Matas, K. Messer, V. Popovici, F. Poree and et al., “The BANCA database and evaluation protocol”, in: Proc. International Conference on Audio- and Video-Based Biometric Person Authentication, (AVBPA) 2003. Lecture Notes in Computer Science, vol 2688. Springer, pp. 625–638.
   [6]      J. Huang, G. Potamianos, J. Connell and C. Neti, “Audio-visual speech recognition using an infrared headset”, Speech Communication, Vol. 44, No.1-4,2004, pp. 83–96.
   [7]      M. Cooke, J. Barker, S. Cunningham and X. Shao, “An audio-visual corpus for speech perception and automatic speech recognition”, Journal of the Acoustical Society of America, Vol. 120, No. 5, 2006, pp. 2421–2424.
   [8]      G. Zhao, M. Barnard and M. Pietikainen, “Lipreading with local spatiotemporal descriptors”, IEEE Transactions on Multimedia, Vol.11, No.7, 2009, pp. 1254–1265.
   [9]      C. Neti, G. Potamianos, J. Luettin, I. Matthews, H. Glotin, D. Vergyri, J. Sison and A. Mashari, “Audio visual speech recognition”, Tech. rep., IDIAP, 2000.
[10]      C. Sanderson, “The VidTIMIT database”, Tech. rep., IDIAP, 2002.
[11]      T. J. Hazen, K. Saenko, C.-H. La and J. R. Glass, “A segment-based audio-visual speech recognizer: Data collection, development, and initial experiments”, in: Proc. International Conference on Multimodal Interfaces, 2004, pp. 235–242.
[12]      J. Trojanova, M. Hruz, P. Campr and M. Zelezn y, “Design and recording of Czech audio-visual database with impaired conditions for continuous speech recognition”, in: Proc. International Conference on Language Resources and Evaluation, 2008.
[13]      D. Petrovska-Delacretaz, S. Lelandais, J. Colineau, L. Chen, B. Dorizzi, M. Ardabilian, E. Krichen, M.-A. Mellakh, A. Chaari, S. Guerfi and et al., “The IV 2 multimodal biometric database (including iris, 2D, 3D, stereoscopic, and talking face data), and the IV 2-2007 evaluation campaign”, in: Proc. International Conference on Biometrics: Theory, Applications and Systems, 2008, pp. 1–7. 2008 IEEE Second International Conference on Biometrics: Theory, Applications and Systems, 2008, pp. 1-7, doi: 10.1109/BTAS.2008.4699323.
[14]      Y. Lan, B.-J. Theobald, R. Harvey, E.-J. Ong and R. Bowden, “Improving visual features for lip-reading”, in: Proc. International Conference on Auditory-Visual Speech Processing, 2010
[15]      A. Vorwerk, X. Wang, D. Kolossa, S. Zeiler and R. Orglmeister, “WAPUSK20 - A database for robust audiovisual speech recognition”, in: Proc. International Conference on Language Resources and Evaluation, 2010.
[16]      Y. Benezeth, G. Bachman, G. Le-Jan, N. Souviraa-Labastie and F. Bimbot, BL-Database: A French Audiovisual Database for Speech Driven Lip Animation Systems”, Ph.D. thesis, INRIA, 2011.
[17]      Y. W. Wong, S. I. Chng, K. P. Seng, L.-M. Ang, S. W. Chin, W. J. Chew and K. H. Lim, “A new multi-purpose audiovisual UNMC-VIER database with multiple variabilities”, Pattern Recognition Letters, Vol. 32, No. 13, 2011, pp.1503–1510.
[18]      C. McCool, S. Marcel, A. Hadid, M. Pietikainen, P. Matejka, ¨ J. Cernocky, N. Poh, J. Kittler, A. Larcher, C. Levy and et al., “Bi-modal person recognition on a mobile phone: using mobile phone data”, in: Proc. International Workshop on Multimedia and Expo, 2012, pp. 635–640.
[19]      N. Harte and E. Gillen, “TCD-TIMIT: An audio-visual corpus of continuous speech”, IEEE Transactions on Multimedia, Vol. 17, No. 5, 2015, pp. 603–615.
[20]      V. Verkhodanova, A. Ronzhin, I. Kipyatkova, D. Ivanko, A. Karpov and M. Zelezny, “HAVRUS corpus: high-speed recordings of audio-visual Russian speech”, in: Proc. International Conference on Speech and Computer, 2016, pp. 338–345.
[21]      J. S. Chung and A. Zisserman, “Lip reading in profile”, in: Proc. British Machine Vision Conference, 2017.
[22]      A. Fernandez-Lopez, O. Martinez and F. M. Sukno, “Towards estimating the upper bound of visual-speech recognition: The visual lip-reading feasibility database”, in: Proc. International Conference on Automatic Face and Gesture Recognition,
2017, pp. 208–215.
[23]      R. Goecke and J. B. Millar, “The audio-video australian english speech data corpus AVOZES”, in: Proc. International Conference on Spoken Language Processing, 2004, pp. 2525–2528.
[24]      A. Ortega, F. Sukno, E. Lleida, A. F. Frangi, A. Miguel, L. Buera and E. Zacur, “AV@CAR: A Spanish multichannel multimodal corpus for in-vehicle automatic audio-visual speech recognition”, in: Proc. International Conference on Language Resources and Evaluation, 2004, pp. 763–767.
[25]      B. Lee, M. Hasegawa-Johnson, C. Goudeseune, S. Kamdar, S. Borys, M. Liu and T. S. Huang, “AVICAR: audio-visual speech corpus in a car environment”, in: Proceedings of Interspeech, 2004.
[26]      A.Bastanfard, M.Fazel, A.A. Kelishami and M.Aghaahmadi, “The Persian Linguistic Based Audio-Visual Data Corpus: AVA II, Considering Coarticulation”, S. Boll et al. (Eds.): MMM 2010, LNCS 5916, 2010, pp. 284–294.
[27]      A. G. Chitu, K. Driel and L. J. Rothkrantz, “Automatic lip reading in the Dutch language using active appearance models on high speed recordings”, in: Proc. International Conference on Text, Speech and Dialogue, 2010, pp. 259–266.
[28]      A. Rekik, A. Ben-Hamadou and W. Mahdi, “A new visual speech recognition approach for RGB-D cameras”, in: Proc. International Conference on Image Analysis and Recognition, 2014, pp. 21–28.
[29]      D. Estival, S. Cassidy, F. Cox and D. Burnham, “AusTalk: an audiovisual corpus of Australian English”, in: Proc. International Conference on Language Resources and Evaluation, 2014
[30]      I. Anina, Z. Zhou, G. Zhao and M. Pietikainen, “OuluVS2: A multi-view audiovisual database for non-rigid mouth motion analysis”, in: Proc. International Conference on Automatic Face and Gesture Recognition, IEEE, Vol. 1, 2015, pp. 1–5.
[31]      S. Petridis, J. Shen, D. Cetin and M. Pantic, “Visual-only recognition of normal, whispered and silent speech”, in: Proc. International Conference on Acoustics, Speech and Signal Processing (in press), 2018.
[32]      T. Afouras, J. S. Chung, A. Senior, O. Vinyals and A. Zisserman, “Deep Audio-Visual Speech Recognition”, IEEE Transactions, 22 December 2018. arXiv:1809.02108
[33]      V. Zue, S. Seneff, and J. Glass, “Speech database development: TIMIT and beyond. Speech Communication”, vol. 9, no. 4, 1990, pp. 351-356.
[34]      W. Fisher, G. Doddington, and K. Goudie-Marshall, “The DARPA speech recognition research database: specifications and status”, In Proceedings of the DARPA Speech Recognition Workshop, 1986.
[35]      P. Price, W. M. Fisher, J. Bernstein, and D. S. Pallett, “The darpa 1000-word resource management database for continuous speech recognition”, International Conference on IEEE In Acoustics, Speech, and Signal Processing, ICASSP-88, 1988, pp. 651-654.
[36]      A. A. Karpov and A.L. Ronzhin, “Information enquiry kiosk with multimodal user interface”, Pattern Recogn. Image Analy. Vol. 19, No. 3, 2009, pp. 546–558.
[37]      D. E. King, “Dlib-ml: A machine learning toolkit”. The Journal of Machine Learning Research, Vol. 10, 2009, pp. 1755–1758.
[38]      C. Tomasi and T. Kanade, “Selecting and tracking features for
image sequence analysis”, Robotics and Automation, 1992.
[39]      P. Boersma and et al., “Praat, a system for doing phonetics by computer”, Glot international, vol. 5, 2002, pp. 341–345.
[40]      J. C. Wells and et al., “Sampa computer readable phonetic alphabet”, Handbook of standards and resources for spoken language systems, vol. 4, 1997.
[41]      I. Matthews, T. F. Cootes, J. A. Bangham, S. Cox and R. Harvey,Extraction of visual features for lipreading”, IEEE Transactions on Pattern Analysis and Machine Intelligence, Vol. 24, No. 2, 2002, pp. 198–213.
[42]      S. J. Cox, R. Harvey, Y. Lan, J. L. Newman and B. J. Theobald,The challenge of multispeaker lip-reading”, in: Proc. International Conference on Auditory-Visual Speech Processing, 2008, pp. 179–184.
[43]      J. Wells, “Sampa computer readable phonetic alphabet”, 2003.
[44]      I. Shdaifat,R. Grigat and D. Langmann, “A System for Automatic Lip Reading”, in International Conference on Audio-Visual Speech Processing St. Jorios (ISCA) AVSP, France, 4-7 September 2003.
[45]      A. Czyzewski, B. Kostek, P. Bratoszewski, J. Kotus and
M. Szykulski, “An audio-visual corpus for multimodal automatic speech recognition”, Journal of Intelligent Information Systems, 2017, pp. 1–26.
[46]      J. S. Chung and A. Zisserman, “Lip reading in the wild”, in: Proc. Asian Conference on Computer Vision, 2016, pp. 87–103.
[47]      Dominic Howell, Stephen Cox and Barry Theobald, “Visual Units and Confusion Modelling for Automatic Lip-reading”, Image and Vision Computing Journal Elsevier, vol. 51, no. C, pp. 1-12, July 2016.
[48]      Joon Son Chung and Andrew Zisserman, “Learning to lip read words by watching videos”, Computer Vision and Image Understanding Journal Elsevier, Vol. 173, August 2018, pp. 76-85
[49]      Fernandez-Lopez and Adriana and Federico Sukno. “Survey on automatic lip-reading in the era of deep learning”, Image Vision Comput, Vol. 78, 2018, pp.53-72.
[50]      P. Cısar, M. Zelezny, Z. Krnoul, J. Kanis, J. Zelinka, and L. Muller, “Design and recording of czech speech corpus for audio-visual continuous speech recognition”, In Auditory Visual Speech Processing Workshop, 2005.
[51]      Garofolo, John S. and et al., “TIMIT Acoustic-Phonetic Continuous Speech Corpus”, LDC93S1. Web Download. Philadelphia: Linguistic Data Consortium, 1993.
[52]      S. J. Young and et al., The HTK Book, Version 3.4. Cambridge, U.K.: Cambridge Univ. Press, 2006.
[53]      A.A. Karpov and A.L. Ronzhin, “Information enquiry kiosk with multimodal user interface”, Pattern Recogn. Image Analy, Vol. 19, No. 3, 2009, pp. 546–558.
[54]      K. Messer, J. Matas, J. Kittler, J. Luettin and G. Maitre,
“XM2VTSDB: The extended M2VTS database”, in: Proc. International Conference on Audio and Video-based Biometric Person Authentication, Vol. 964, 1999, pp. 965–966.
[55]      The M2VTS database; http://www.tele.ucl.ac.be/
M2VTS/m2fdb.html.
[56]      M. Igras, B. Ziołko and T. Jadczyk, “Audiovisual database of Polish speech recordings”, Studia Informatica, Vol. 33, No. 2B, 2012, pp. 163–172.
[57]      N. A. Fox, B. A. OMullane and R. B. Reilly, “VALID: A new practical audio-visual database, and comparative results”, in: Proc. International Conference on Audio-and Video-Based Biometric Person Authentication, 2005, pp. 777–786.
[58]      P. J. Lucey, G. Potamianos and S. Sridharan, “Patch-based analysis of visual speech from multiple views, in: Proc. International Conference on Auditory-Visual Speech Processing, 2008.
[59]      A. Torfi, S. M. Iranmanesh and N. NasrAbadi, “3D Convolutional Neural Networks for Cross Audio-Visual Matching Recognition”, IEEE Computer Vision and Pattern Recognition Journal, Vol. 5, Issue: 99, 18 Jun 2017, pp. 22081 – 22091.
[60]      S. Tamura, C. Miyajima, N. Kitaoka, T. Yamada, S. Tsuge, T. Takiguchi, K. Yamamoto, T. Nishiura, M. Nakayama, Y. Denda and et al., “CENSREC-1-AV: An audio-visual corpus for noisy bimodal speech recognition”, in: Proc. International Conference on Auditory-Visual Speech Processing, 2010.
[61]      H. Hermansky, “Perceptual linear predictive (PLP) analysis of speech”, the Journal of the Acoustical Society of America, Vol. 87, No. 4, 1990, pp. 1738–1752.
[62]      J. Yuan and M. Liberman, “Speaker identification on the scotus corpus”, Journal of the Acoustical Society of America, Vol. 123, No. 5, 2008, pp. 38-78.
[63]      Zhang, X.; Xu, Y.; Abel, A.K.; Smith, L.S.; Watt, R.; Hussain, A.; Gao, C. "Visual Speech Recognition with Lightweight Psychologically Motivated Gabor Features". Entropy 2020, 22, 1367. https://doi.org/10.3390/e22121367
[64]      Assael, Y.M.; Shillingford, B.; Whiteson, S.; De Freitas, N. LipNet: End-to-End Sentence-level Lipreading. arXiv 2016, arXiv:1611.01599.
[65]      Wand, M.; Koutník, J.; Schmidhuber, J. Lipreading with long short-term memory. In Proceedings of the 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Shanghai, China, 20–25 March 2016; pp. 6115–6119.
[66]      Wand, M.; Schmidhuber, J.; Vu, N.T. Investigations on End-to-End Audiovisual Fusion. In Proceedings of the 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Calgary, AB, Canada, 15–20 April 2018; pp. 3041–3045.
[67]      Xu, K.; Li, D.; Cassimatis, N.; Wang, X. LCANet: End-to-end lipreading with cascaded attention-CTC. In Proceedings of the 2018 13th IEEE International Conference on Automatic Face & Gesture Recognition (FG 2018), Xi’an, China, 15–19 May 2018; pp. 548–555.
[68]      Shao, X.; Barker, J. Stream weight estimation for multistream audio–visual speech recognition in a multispeaker environment. Speech Commun. 2008, 50, 337–353. [CrossRef]
[69]      Lan, Y.; Harvey, R.; Theobald, B.; Ong, E.J.; Bowden, R. Comparing visual features for lipreading. In Proceedings of the International Conference on Auditory-Visual Speech Processing 2009, Norwich, UK, 10–13 September 2009; pp. 102–106.
[70]      Rekik, Ahmed & Ben-Hamadou, Achraf & Mahdi, Walid. (2015). An adaptive approach for lip-reading using image and depth data. Multimedia Tools and Applications. 75. 10.1007/s11042-015-2774-3
[71]      Pingping Wu, Hong Liu, Xiaofei Li, Ting Fan, Xuewu Zhang. A Novel Lip Descriptor for Audio-Visual Keyword Spotting Based on Adaptive Decision Fusion. IEEE Transactions on Multimedia, Institute of Electrical and Electronics Engineers, 2016, 18 (3), pp.326-338. ff10.1109/TMM.2016.2520091ff. ff
[72]      Iain Matthews, Gerasimos Potamianos, Chalapathy Neti, Juergen Luettin, A COMPARISON OF MODEL AND TRANSFORM-BASED VISUAL FEATURES FOR AUDIO-VISUAL LVCSR, IEEE International Conference on Multimedia and Expo, ICME 2001,
[73]      Cappelletta, L.; Harte, N. Viseme definitions comparison for visual-only speech recognition. In Proceedings of the 2011 19th European Signal Processing Conference, Barcelona, Spain, 29 August–2 September 2011; pp. 2109–2113.
[74]      Tim Sheerman-Chase, Eng-Jon Ong and Richard Bowden, "Feature Selection of Facial Displays for Detection of Non Verbal Communication in Natural Conversation", in natural conversation, IEEE International Workshop on Human-Computer Interaction, Kyoto, 2009
[75]      R. Bowden, LILir Twotalk Corpus, University of Surrey, 20108 http://www.ee. surrey.ac.uk/Projects/LILiR/twotalk_corpus.
[76]      Hatice Çınar Akakın, Bülent Sankur, Robust classification of face and head gestures in video, Image and Vision Computing, Volume 29, Issue 7, 2011, Pages 470-483, ISSN 0262-8856,
[77]      Eng-Jon Ong, Richard Bowden, Robust Facial Feature Tracking Using Shape-Constrained Multiresolution-Selected Linear Predictors, IEEE Transactions on Pattern Analysis and Machine Intelligence 33(9):1844 – 1859, October 2011
[78]      Philip Tresadern, Chris McCool, Norman Poh, Pavel Matejka, Abdenour Hadid, Christophe Levy, Tim Cootes, and Sebastien Marcel. Mobile biometrics (mobio): Joint face and voice verification for a mobile platform. IEEE pervasive computing, 2012.
[79]      George Sterpu, Christian Saam, and Naomi Harte. 2018. Attention-based Audio-Visual Fusion for Robust Automatic Speech Recognition. In Proceedings of the 20th ACM International Conference on Multimodal Interaction (ICMI '18). Association for Computing Machinery, New York, NY, USA, 111–115.DOI:https://doi.org/10.1145/3242969.3243014
[80]      A. Gabbay, A. Ephrat, T. Halperin and S. Peleg, "Seeing Through Noise: Visually Driven Speaker Separation And Enhancement," 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2018, pp. 3051-3055, doi: 10.1109/ICASSP.2018.8462527.
[81]      G. Morrone, S. Bergamaschi, L. Pasa, L. Fadiga, V. Tikhanoff and L. Badino, "Face Landmark-based Speaker-independent Audio-visual Speech Enhancement in Multi-talker Environments," ICASSP 2019 - 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2019, pp. 6900-6904, doi: 10.1109/ICASSP.2019.8682061.
[82]      Ariel Ephrat, Tavi Halperin, Shmuel Peleg, “Improved Speech Reconstruction From Silent Video”, Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2017, pp. 455-462
[83]      G. Sterpu, C. Saam and N. Harte, "Can DNNs Learn to Lipread Full Sentences?", 2018 25th IEEE International Conference on Image Processing (ICIP), 2018, pp. 16-20, doi: 10.1109/ICIP.2018.8451388.
[84]      George SterpuNaomi Harte, “Towards Lipreading Sentences with Active Appearance Models”, Presented at The 14th International Conference on Auditory-Visual Speech Processing (AVSP 2017),
Image and Video Processing (eess.IV); Audio and Speech Processing (eess.AS), arXiv:1805.11688v1
[85]      Kwanchiva ThangthaiHelen L BearRichard Harvey, “Comparing phonemes and visemes with DNN-based lipreading”, BMVC Lipreading Workshop 2017, arXiv:1805.02924v1 
[86]      Ivanko, D. & Ryumin, D. (2021). A NOVEL TASK-ORIENTED APPROACH TOWARD AUTOMATED LIP-READING SYSTEM IMPLEMENTATION. The International Archives of the Photogrammetry, Remote Sensing and Spatial Information Sciences. XLIV-2/W1-2021. 85-89. 10.5194/isprs-archives-XLIV-2-W1-2021-85-2021.
[87]      A. M. Sarhan, N. M. Elshennawy and D. M. Ibrahim, "Hlr-net: a hybrid lip-reading model based on deep convolutional neural networks," Computers, Materials & Continua, vol. 68, no.2, pp. 1531–1549, 2021.
[88]      Stavros Petridis, Yujiang Wang, Pingchuan Ma, Zuwei Li, Maja Pantic, "End-to-end visual speech recognition for small-scale datasets”, Pattern Recognition Letters,Volume 131, 2020, Pages 421-427, ISSN 0167-8655, ELSEVIER
[89]      BBC and Oxford University. 2017. The BBC-Oxford Multi-View Lip Reading Sentences 2 (LRS2) Dataset. http://www.robots.ox.ac.uk/~vgg/data/lip_reading_ sentences/. (2017). Online, Accessed: 11 August 2018.
[90]      Triantafyllos AfourasJoon Son ChungAndrew Zisserman, “Deep Lip Reading: a comparison of models and an online application”, Computer Vision and Pattern Recognition (cs.CV), 15 Jun 2018, arXiv:1806.06053v1
[91]      Triantafyllos AfourasJoon Son ChungAndrew Zisserman, “ASR is all you need: cross-modal distillation for lip reading”, Computer Vision and Pattern Recognition (cs.CV); Sound (cs.SD); Audio and Speech Processing (eess.AS), 28 Nov 2019 arXiv:1911.12747v2 [cs.CV]
[92]      Ya Zhao, Rui Xu, Xinchao Wang, Peng Hou, Haihong Tang, Mingli Song1, “Hearing Lips: Improving Lip Reading by Distilling Speech Recognizers”, in Proc. AAAI Conf. Artif. Intell., vol. 34, no. 04, 2020, pp. 6917–6924.
[93]      Courtney L., Sreenivas R. (2020) Using Deep Convolutional LSTM Networks for Learning Spatiotemporal Features. In: Palaiahnakote S., Sanniti di Baja G., Wang L., Yan W. (eds) Pattern Recognition. ACPR 2019. Lecture Notes in Computer Science, vol 12047. Springer, Cham. https://doi.org/10.1007/978-3-030-41299-9_24
[94]      Lee Y-H, Jang D-W, Kim J-B, Park R-H, Park H-M. Audio–Visual Speech Recognition Based on Dual Cross-Modality Attentions with the Transformer Model. Applied Sciences. 2020; 10(20):7263. https://doi.org/10.3390/app10207263
[95]      Goecke, R., & Millar, J. (2004). A Detailed Description of the AVOZES data corpus. In S. Cassidy, F. Cox, R. Mannell, & S. Palethorpe (Eds.), Proceedings of the 10th Australian International Conference on Speech Science and Technology (pp. 486-491). ASSTA. https://assta.org/proceedings/sst/2004/proceedings/papers/sst2004-354.pdf
[96]      Saragih, Jason & Goecke, Roland. (2006). Learning active appearance models from image sequences. Proc. VisHCI 2006, Volume 56 of CRPIT. 56.
[97]      Chetty, G & Wagner, M 2007, Spatiotemporal person authentication based on multilevel fusion. in M Cree (ed.), Proceedings of the Image and Vision Computing New Zealand Conference. Proceedings of Image and Vision Computing New Zealand, University of Canterbury, New Zealand, pp. 248-253
[98]      Chetty, Girija & Wagner, Michael. (2008). A robust spatio-temporal face modelling approach using 3D multimodal fusion for biometric security applications. Proc SPIE. 6944. 10.1117/12.778631.
[99]      K. Messer et al., "Face authentication test on the BANCA database," Proceedings of the 17th International Conference on Pattern Recognition IEEE, 2004. ICPR 2004., 2004, pp. 523-532 Vol.4, doi: 10.1109/ICPR.2004.1333826.
[100]      Navarathna, Rajitha & Dean, David & Lucey, Patrick & Sridharan, Sridha & Fookes, Clinton. (2010). Recognising Audio-Visual Speech in Vehicles using the AVICAR Database. Australasian International Conference on Speech Science and Technology (SST).
[101]      R. Navarathna, P. Lucey, D. Dean, C. Fookes and S. Sridharan, "Lip detection for audio-visual speech recognition in-car environment," 10th International Conference on Information Science, Signal Processing and their Applications (ISSPA 2010), 2010, pp. 598-601, doi: 10.1109/ISSPA.2010.5605429.
[102]      Kleinschmidt, Tristan and Dean, David and Sridharan, Sridha and Mason, Michael (2007) A Continuous Speech Recognition Evaluation Protocol for the AVICAR Database. In Proceedings International Conference On Signal Processing and Communication Systems, Gold Coast, Australia.
[103]      Bastanfard, Azam & Fazel, Maryam & Abdi, Alireza & Aghaahmadi, Mohammad. (2009). A comprehensive audio-visual corpus for teaching sound Persian phoneme articulation. Conference Proceedings - IEEE International Conference on Systems, Man and Cybernetics. 169 - 174. 10.1109/ICSMC.2009.5346591.
[104]      Dhairya Desai, Priyesh Agrawal, Priyansh Parikh, Piyush Kumar Soni, 2020, Visual Speech Recognition, INTERNATIONAL JOURNAL OF ENGINEERING RESEARCH & TECHNOLOGY (IJERT) Volume 09, Issue 04 (April 2020)
[105]      Dharin ParekhAnkitesh GuptaShharrnam ChhatparAnmol Yash KumarManasi Kulkarni, “Lip Reading Using Convolutional Auto Encoders as Feature Extractor”, Computer Vision and Pattern Recognition, 31 May 2018, arXiv:1805.12371v1
[106]      Alghowinem, S., Wagner, M., & Göcke, R. (2013). AusTalk — The Australian speech database: Design framework, recording experience and localisation. 2013 8th International Conference on Information Technology in Asia (CITA), 1-7.
[107]      Chung, J.S.; Zisserman, A. Out of time: Automated lip sync in the wild. In Asian Conference on Computer Vision; Springer: Berlin/Heidelberg, Germany, 2016; pp. 251–263. 23.
[108]      Petridis, S.; Wang, Y.; Li, Z.; Pantic, M. End-to-End Multi-View Lipreading. In Proceedings of the British Machine Vision Conference, London, UK, 4–7 September 2017.
[109]      S. Petridis, Y. Wang, Z. Li, and M. Pantic, ‘‘End-to-end audiovisual fusion with LSTMs,’’ 2017, arXiv:1709.04343. [Online]. Available: http://arxiv.org/abs/1709.04343 [118] S.
[110]      Fung, I.; Mak, B. End-to-end low-resource lip-reading with maxout CNN and LSTM. In Proceedings of the 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Calgary, AB, Canada, 15–20 April 2018; pp. 2511–2515.
[111]      Lee, D.; Lee, J.; Kim, K.E. "Multi-view automatic lip-reading using neural network". In Asian Conference on Computer Vision; Springer: Berlin/Heidelberg, Germany, 2016; pp. 290–302.
[112]      T. Saitoh, Z. Zhou, G. Zhao, and M. Pietikäinen, "Concatenated frame image based cnn for visual speech recognition", in Proc. Asian Conf. Comput. Vis. Cham, Switzerland: Springer, 2016, pp. 277–289.
[113]      Aghaahmadi, Mohammad & Dehshibi, Mohammad Mahdi & Bastanfard, Azam & Fazlali, Mahmood. (2013). Clustering Persian viseme using phoneme subspace for developing visual speech application. Multimedia Tools and Applications. 65. 521-541. 10.1007/s11042-012-1128-7.
[114]      Bastanfard, Azam & Aghaahmadi, Mohammad & Abdi, Alireza & Fazel, Maryam & Moghadam, Maedeh. (2009). Persian Viseme Classification for Developing Visual Speech Training Application. 1080-1085. 10.1007/978-3-642-10467-1_104.
[115]      Moghadam, Maedeh, Azam Bastanfard, and Mohammad Mahdi Dehshibi. (2011). "Toward Clustering Persian Vowel Viseme: A New Clustering Approach based on HMM.", International Conference on Signal Acquisition and Processing, Volume 2, IEEE
[116]      Bastanfard A., Rezaei N.A., Mottaghizadeh M., Fazel M. (2010) A Novel Multimedia Educational Speech Therapy System for Hearing Impaired Children. In: Qiu G., Lam K.M., Kiya H., Xue XY., Kuo CC.J., Lew M.S. (eds) Advances in Multimedia Information Processing - PCM 2010. PCM 2010. Lecture Notes in Computer Science, vol 6298. Springer, Berlin, Heidelberg. https://doi.org/10.1007/978-3-642-15696-0_65
[117]      R. Mahdavi, A. Bastanfard and D. Amirkhani, "Persian Accents Identification Using Modeling of Speech Articulatory Features", 2020 25th International Computer Conference, Computer Society of Iran (CSICC), 2020, pp. 1-9, doi: 10.1109/CSICC49403.2020.9050139.
[118]      Minsu Kim, Joanna Hong, Se Jin Park, and Yong Man Ro, “Multi-modality associative bridging through memory: Speech sound recollected from face video”, In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 296306, 2021.
[119]      D. Kumar Margam, R. Aralikatti, T. Sharma, A. Thanda, P. A K, S. Roy and S. M. Venkatesan. (2019). Lip Reading with 3D-2D-CNN BLSTMHMM and word-CTC models. arXiv:1906.12170
[120]      Minsu Kim, Hyung-Il Kim, and Yong Man Ro, “Prompt tuning of deep neural networks for speaker-adaptive visual speech recognition”, arXiv preprint arXiv:2302.08102, 2023.
[121]      S. Petridis, M. Pantic. (2016). Deep complementary bottleneck features for visual speech recognition, Proc. International Conference on Acoustics, Speech and Signal Processing.
[122]      A. Mesbah, A. Berrahou, H. Hammouchi, H. Berbia, H. Qjidaa and M. Daoudi, (2019). Lip reading with hahn convolutional neural networks. Image and Vision Computing.
[123]      D.-W. Jang, H.-I. Kim, C. Je, R.-H. Park, and H.-M. Park. (2019). Lip reading using committee networks with two different types of concatenated frame images. IEEE Access, vol. 7.
[124]      T. Shirakata and T. Saitoh, "Lip Reading Experiments for Multiple Databases using Conventional Method," 2019 58th Annual Conference of the Society of Instrument and Control Engineers of Japan (SICE), Hiroshima, Japan, 2019, pp. 409-414, doi: 10.23919/SICE.2019.8859932.
[125]      K. Thangthai, R. Harvey, S. Cox and B.-J. Theobald. (2015). Improving lip-reading performance for robust audiovisual speech recognition using DNNs. Proc. International Conference on Auditory-Visual Speech Processing
[126]      Abdel-Hamid O, Mohamed AR, Jiang H, Deng L, Penn G, Yu D (2014) Convolutional neural networks for speech recognition. IEEE/ACMTransAudioSpeechLangProcess22(10):1533–154
[127]      Sailor HB, Patil HA (2016) Novel unsupervised auditory Filterbank learning using convolutional RBM for speech recognition. IEEE/ACMTransAudioSpeechLangProcess24(12):2341–2353
[128]      Ravanelli M, Serdyuk D, Bengio Y (2018) Twin Regularization for online speech recognition. arXiv:180405374
[129]      Bhowmik T, Mandal SKD (2016) Deep neural network based phonological feature extraction for Bengali continuous speech. In: Signal and information processing (IConSIP), pp 1–5
[130]      Hegde RM, Murthy HA, Gadde VRR (2007) Significance of the modified group delay feature in speech recognition. IEEE Trans Audio Speech Lang Process 15(1):190–202
[131]      Souheil Fenghour, Daqing chen and Perry Xiao, Decoder-Encoder LSTM for Lip Reading, (2019) ICSIE '19: Proceedings of the 2019 8th International Conference on Software and Information Engineering, April 2019, Pages 162–166
[132]      Saad Naeem, Omer Beg, STRATA: Word Boundaries & Phoneme Recognition From Continuous Urdu Speech using Transfer Learning, Attention, & Data Augmentation, arXiv - EE - Audio and Speech Processing (IF), Pub Date : 2022-04-16, DOI: arxiv-2204.07848
[133]      Shiyang ChengPingchuan MaGeorgios TzimiropoulosStavros PetridisAdrian BulatJie ShenMaja Pantic, Towards Pose-invariant Lip-Reading, Computer Vision and Pattern Recognition (cs.CV),
2019, arXiv:1911.06095
[134]      A. Garg, J. Noyola, and S. Bagadia. (2016). Lip reading using CNN and LSTM. Technical report Stanford University - CS231n project report.
[135]      S. Nadeem Hashmi, H. Gupta, D. Mittal, K. Kumar, A. Nanda and S. Gupta. (2018). A Lip Reading Model Using CNN with Batch Normalization. 11th International Conference on Contemporary Computing.
[136]      Aripin, Aripin and Setiawan, Abas, Indonesian Lip-Reading Recognition Using Long-Term Recurrent Convolutional Network. Available at SSRN: https://ssrn.com/abstract=4444973 or http://dx.doi.org/10.2139/ssrn.4444973
[137]      S. Petridis, J. Shen, D. Cetin and M. Pantic. (2018). Visual-only recognition of normal, whispered and silent speech. Proc. International Conference on Acoustics, Speech and Signal Processing
[138]      Minsu Kim, Joanna Hong, and Yong Man Ro, “Lip-to-speech synthesis in the wild with multi-task learning”, In ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 15. IEEE,2023.
[139]      Minsu Kim, Jeong Hun Yeo, and Yong Man Ro, "Distinguishing homophenes using multi-head visual-audio memory for lip reading," In Proceedings of the AAAI Conference on Artificial Intelligence, volume 36, pages 11741182, 2022.
[140]      Pingchuan Ma, A. Haliassos, A. Fernandez-Lopez, H. Chen, S. Petridis and M. Pantic, "Auto-AVSR: Audio-Visual Speech Recognition with Automatic Labels," ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Rhodes Island, Greece, 2023, pp. 1-5, doi: 10.1109/ICASSP49357.2023.10096889
[141]      Jeongsoo Choi, Minsu Kim, and Yong Man Ro, “Intelligible lip-to-speech synthesis with speech units”, arXiv preprint arXiv:2305.19603, 2023
[142]      KR Prajwal, Triantafyllos Afouras, and Andrew Zisserman. “Sub-word level lip reading with visual attention”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 51625172, 2022.
[143]      Joanna Hong, Minsu Kim, Daehun Yoo, and Yong Man Ro. “Visual context-driven audio feature enhancement for robust end-to-end audio-visual speech recognition”, arXiv preprint arXiv:2207.06020, 2022.
[144]      Qiushi Zhu, Long Zhou, Ziqiang Zhang, Shujie Liu, Binxing Jiao, Jie Zhang, Lirong Dai, Daxin Jiang, Jinyu Li, and FuruWei. "Vatlm: Visual-audio-text pre-training with unified masked prediction for speech representation learning". IEEE Transactions on Multimedia, 2023.
[145]      Hareesh Mandalapu and Aravinda Reddy P N and Raghavendra Ramachandra and K Sreenivasa Rao and Pabitra Mitra and S R Mahadeva Prasanna and Christoph Busch, "Multilingual Audio-Visual Smartphone Dataset And Evaluation", mandalapu 2021, arXiv eprint arXiv:2109.04138, 2021
[146]      Setyaningsih ER, Handayani AN, Irianto WSG, Kristian Y, Chen CTSL. LUMINA: Linguistic unified multimodal Indonesian natural audio-visual dataset. Data Brief. 2024; 54:110279. Published 2024 Mar 1. doi:10.1016/j.dib. 2024.110279
[147]      Gerald Schwiebert, Cornelius Weber, Leyuan Qu, Henrique Siqueira, Stefan Wermter, "A Multimodal German Dataset for Automatic Lip Reading Systems and Transfer Learning", arXiv:2202.13403v3 [cs.CV] 11 May 2022
[148]      Weicong Chen, Xu Tan, Yingce Xia,Tao Qin, Yu Wang, Tie-Yan Liu, "DualLip: A System for Joint Lip Reading and Generation," in MM '20: Proceedings of the 28th ACM International Conference on Multimedia, New York, NY, USA, 2020.
[149]      Xue, Feng and Li, Yu and Liu, Deyin and Xie, Yincen and Wu, Lin and Hong, Richang, "LipFormer: Learning to Lipread Unseen Speakers Based on Visual-Landmark Transformers," IEEE Transactions on Circuits and Systems for Video Technology, vol. 33, no. 9, pp. 4507-4517, 2023
[150]      Feng Xue, Peng Li, Yu Li, Shujie Li,, "WPELip: enhance lip reading with word-prior information," Multimedia Systems, vol. 31, no. 2, 28 January 2025
[151]      Leyuan Qu, Cornelius Weber, Stefan Wermter, "LipSound2: Self-Supervised Pre-Training for Lip-to-Speech Reconstruction and Lip Reading," IEEE Transactions on Neural Networks and Learning Systems, vol. 35, no. 2, pp. 2772-2782, 2024
[152]      David Gimeno‑Gómez and Carlos‑D. Martínez‑Hinarejos, "Continuous lipreading based on acoustic temporal alignments," EURASIP Journal on Audio, Speech, and Music Processing, vol. 25, 2024.
[153]      Triantafyllos Afouras, Joon Son Chung, Andrew Senior, Oriol Vinyals, Andrew Zisserman, "Deep Audio-Visual Speech Recognition," IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 44, no. 12, pp. 8717-8727, 2022.
[154]      Braun, V. and Clarke, V. (2006) ‘Using thematic analysis in psychology’, Qualitative Research in Psychology, vol. 3, no.2 pp. 77–101. doi: 10.1191/1478088706qp063oa.
[155]      Huijuan Wang, Boyan Cui, Quanbo Yuan, Gangqiang Pu, Xueli Liu, Jie Zhu, "Mini-3DCvT: a lightweight lip-reading method based on 3D convolution visual transformer," The Visual Computer, vol. 41, no. 3, pp. 1957 - 1969, 11 June 2024.
[156]      Martinez, Brais and Ma, Pingchuan and Petridis, Stavros and Pantic, Maja, "Lipreading Using Temporal Convolutional Networks", 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), p.p. 6319-6323, 10.1109/ICASSP40776.2020.9053841
[157]      Ma, Pingchuan and Martinez, Brais and Petridis, Stavros and Pantic, Maja, " Towards Practical Lipreading with Distilled and Efficient Models", 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), p.p 7608-7612, 10.1109/ICASSP39728.2021.9415063
[158]      Tian LanJun SongZhiwei HouKang ChenShuping HeHai Wang, "Event-Triggered Fixed-Time Sliding Mode Control for Lip-Reading-Driven UAV: Disturbance Rejection Using Wind Field Optimization", in IEEE Transactions on Automation Science and Engineering, vol. 22, pp. 9090 - 9103, 19 November 2024, 10.1109/TASE.2024.3496939
[159]      Ya Zhao, Rui Xu, and Mingli Song. 2020. A Cascade Sequence-to-Sequence Model for Chinese Mandarin Lip Reading. In Proceedings of the 1st ACM International Conference on Multimedia in Asia (MMAsia '19). Association for Computing Machinery, New York, NY, USA, Article 32, 1–6. https://doi.org/10.1145/3338533.3366579
[160]      Dawei Luo, Dongliang Xie, Yuqing Zhang, Wanpeng Xie, Baosheng Sun, "Hard sample semantic reconstruction for mandarin visual speech recognition," Digital Signal Processing, vol. 160, p. 105066, 2025
 
[161]      P. Ma, S. Petridis and M. Pantic, "End-To-End Audio-Visual Speech Recognition with Conformers," ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Toronto, ON, Canada, 2021, pp. 7613-7617, doi: 10.1109/ICASSP39728.2021.9414567.
[162]      Bowen ShiWei-Ning HsuKushal LakhotiaAbdelrahman Mohamed, "Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction," 2022, arXiv:2201.02184
[163]      Pingchuan MaStavros PetridisMaja Pantic, "Visual Speech Recognition for Multiple Languages in the Wild," 2022, arXiv:2202.13084
[164]      Baosheng Sun, Dongliang Xie, Tiantian Duan, Variable Structure and Modeling Units for Chinese Lipreading, International Journal of Pattern Recognition and Artificial Intelligence, Vol. 36, No. 15, pp. 2256021, ;kvg aniT jovdf
[165]      G. Tan, Z. Wan, Y. Wang, Y. Cao and Z. -J. Zha, "Tackling Event-Based Lip-Reading by Exploring Multigrained Spatiotemporal Clues," in IEEE Transactions on Neural Networks and Learning Systems, vol. 36, no. 5, pp. 82 79-8291, May 2025, doi: 10.1109/TNNLS.2024.3440495.
[166]      Yang, S., Zhang, Y., Feng, D., Yang, M.,Wang, C., Xiao, J., Long, K., Shan, S., Chen, X.: LRW-1000: a naturally-distributed largescale benchmark for lip reading in the Wild. In: 2019 14th IEEE International Conference on Automatic Face and Gesture Recognition (FG 2019)
[168]      Xuejuan Chen, Jixiang Du1, Hongbo Zhang, "Lipreading with DenseNet and resBi-LSTM", 2020, Signal, Image and Video Processing, Springer, Vol. 14, No. 5, p.p. 981-989, 10.1007/s11760-019-01630-1
[169]      Marzieh Oghbaie, Arian Sabaghi, Kooshan Hashemifard, Mohammad Akbari, "When deep learning deciphers silent video: a survey on automatic deep lip reading," Multimedia Tools and Applications, 22 March 2025
[170]      Dmitry Ryumin and Alexandr Axyonov and Elena Ryumina and Denis Ivanko and Alexey Kashevnik and Alexey Karpov, "Audio–visual speech recognition based on regulated transformer and spatio–temporal fusion strategy for driver assistive systems", Expert Systems with Applications, Vol. 252, 2024, p.p. 124159
[171]      C. Chen, D. Wang and T. F. Zheng, "CN-CVS: A Mandarin Audio-Visual Dataset for Large Vocabulary Continuous Visual to Speech Synthesis," ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Rhodes Island, Greece, 2023, pp. 1-5, doi: 10.1109/ICASSP49357.2023.10095796
[172]      Z. Zhou, G. Zhao and M. Pietikäinen, "Towards a practical lipreading system," CVPR 2011, Colorado Springs, CO, USA, 2011, pp. 137-144, doi: 10.1109/CVPR.2011.5995345.
[173]      Y. Fu, X. Zhou, M. Liu, M. Hasegawa-Johnson, and T. Hunag, “Lipreading by locality discriminant graph,” in International Conference on Image Processing, pp. 325–328, San Antonio, TX, 2007.
[174]      Javad Peymanfard and Samin Heydarian and Ali Lashini and Hossein Zeinali and Mohammad Reza Mohammadi and Nasser Mozayani, "A multi-purpose audio-visual corpus for multi-modal Persian speech recognition: The Arman-AV dataset", Expert Systems with Applications, vol. 238, pp. 121648, 2024.
[175]      Y. Hu, S.-C. Liu, and T. Delbruck, "v2e: From video frames to realistic DVS events," IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Workshops (CVPRW), p. 1312–1321, Jun. 2021.