CJK1_70b6c9e31.png Jungkyu Choi 2021.10.15

Speech Recognition Dataset for the Infinite Development of Speech-Based AI

The role and importance of speech recognition dataset

Speech-based AI service that includes speech recognition, natural language processing and speech synthesis is one of the most commercially successful AI technologies. It not only has the convenience as the most natural interface, but it also has the potential as the new portal where each company can expand its platform. The good example is Amazon’s Echo speaker, which is mounted with Alexa to enable music service as well as the original service of Amazon such as item purchase.

As the AI technology develops, the importance of speech technology will increase as well. In order to dominate the market in advance, a competition between the major IT companies including Google, Facebook and Amazon is expected to get more fierce. And at the center of this technological development, there lies the benchmark dataset. It is the base for competition as it measures the performance with the common standard to assess the accuracy of AI. Benchmark dataset enables researchers to easily identify the direction of technology through application of various AI technologies while playing a pivotal role in identifying the stages of technological development through long-term operation.

The figure below[1] demonstrates the time line of LibriSpeech,[2] the most widely used benchmark dataset in the field of speech recognition from 2010 to 2020, which includes the release of speech-based AI service and paradigm in speech recognition technology.



As shown in the figure above, the point of release of speech-based AI service from each IT company can be identified, starting from Apple’s Siri. In addition, the point of appearance of each model can be checked including Kaldi,[3] a speech recognition open source most widely used prior to the release of end-to-end speech recognition, followed by Hybrid Neural Architecture,[4] DeepSpeech,[5] an end-to-end speech recognition model mostly researched recently, encoder-decoder model,[6] Transducer[7] model of Google and the most widely used LibriSpeech.

In the meantime, the word error rate of speech recognition by year in LibriSpeech and Hub5’00, the representative benchmark datasets in the field of speech recognition, is highly significant. Let’s look at the figures[1] below.



In each figure, the blue dotted line shows the test result of its performance when conducted to Human Level, meaning to people. If we take a look at error rate of benchmark dataset, it goes down below human level after 2017 and even goes beyond human level. However, this is the result of benchmark dataset, meaning the performance may be lower in the actual environment.

Let’s also take a look at the table that can easily identify the methodology of the current end-to-end speech recognition result and the SOTA (State-Of-The-Art) technology. The figure below[8] shows the experiment result of the research paper on LibriSpeech dataset of papers with code in a leaderboard form.




Status of Korean and overseas speech dataset

As the awareness on the importance of global speech recognition dataset is rising, researchers in Korea are putting more efforts to establish speech recognition dataset. The table below compares the global speech recognition dataset with Korean speech recognition dataset.

What is noticeable from the above table is the AI Hub established by the National Information Society Agency (NIA). With the diversity and the volume of established speech recognition retained by the AI Hub, it can be used as speech recognition benchmark data in various areas. As the data on the customer reception, vehicle command, conference speech, lecture, and dialect are made by meeting the requirements of organizations and companies that want to develop the service using speech recognition technology, its performance is excellent compared to the existing global speech recognition data.

If there is anything that needs to be improved, there is an issue of personal information, which makes it hard to use the recorded data in the actual customer reception service. It still needs to find a solution in that the actual data cannot be used and the speech recognition data has to be collected through re-recording or other possible ways. In addition, as the recent example of Wav2Vec[14] shows, there is a possibility of improving acoustic model performance through unsupervised learning of audio data Just like the large-volume Libri-Light of Facebook, it is expected that the audio data establishment system will be developed further.


Expectations for Korean speech dataset establishment project

As we have seen so far, considering its diversity, speech recognition dataset is expected to be more widely used in various fields. Still, there are many people who wish that there are more research papers on the dataset or benchmark leaderboard in the domestic research field.

Fortunately, meaningful and relevant contests are held to quench the thirst for the said research. To name one, the “2021 Korean Language Speech and Natural Language AI Contest — Artificial Intelligence With Hunminjeongeum” hosted by the Ministry of Science & ICT and the NIA(National Information society Agency) is held in celebration of Hangeul Day, a national Korean commemorative day that commemorates the invention of the Korean alphabet.[15] Various speech recognition technologies will be developed and related researchers will be cultivated through this contest, thereby causing positive influence on developing Korean speech recognition. As mentioned earlier, it will greatly contribute to the establishment and use of speech recognition data while taking into consideration of the issue related to personal information so that the data of actual users can be collected in the longer term.

참고
[1] The History of Speech Recognition to the Year 2030, Awni Hannun https://arxiv.org/pdf/2108.00084v1.pdf
[2] Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur. Librispeech: an asr corpus based on public domain audio books. ICASSP, pages 5206-5210. IEEE, 2015.
[3] Daniel Povey, Arnab Ghoshal, Gilles Boulianne, Lukas Burget, Ondrej Glembek, Nagendra Goel, Mirko Hannemann, Petr Motlicek, Yanmin Qian, Petr Schwarz, et al. The kaldi speech recognition toolkit. In IEEE 2011 workshop on automatic speech recognition and understanding, number CONF. IEEE Signal Processing Society, 2011.
[4] Geoffrey Hinton, Li Deng, Dong Yu, George E Dahl, Abdel-rahman Mohamed, Navdeep Jaitly, Andrew Senior, Vincent Vanhoucke, Patrick Nguyen, Tara N Sainath, et al. Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups. IEEE Signal processing magazine, 29(6):82-97, 2012.
[5] Awni Hannun, Carl Case, Jared Casper, Bryan Catanzaro, Greg Diamos, Erich Elsen, Ryan Prenger, Sanjeev Satheesh, Shubho Sengupta, Adam Coates, et al. Deep speech: Scaling up end-to-end speech recognition. https://arxiv.org/pdf/1412.5567.pdf
[6] Jan Chorowski, Dzmitry Bahdanau, Dmitriy Serdyuk, Kyunghyun Cho, and Yoshua Bengio. Attention-based models for speech recognition. NIPS Volume 1, pages 577-585, 2015.
[7] Yanzhang He, Tara N Sainath, Rohit Prabhavalkar, Ian McGraw, Raziel Alvarez, Ding Zhao, David Rybach, Anjuli Kannan, Yonghui Wu, Ruoming Pang, et al. Streaming end-to-end speech recognition for mobile devices. ICASSP, pages 6381-6385. IEEE, 2019.
[8] https://paperswithcode.com/sota/speech-recognition-on-librispeech-test-clean
[9] https://arxiv.org/pdf/1912.07875.pdf
[10] https://catalog.ldc.upenn.edu/LDC97S62
[11] https://catalog.ldc.upenn.edu/LDC93S6A
[12] https://catalog.ldc.upenn.edu/LDC93S1
[13] https://aihub.or.kr/aihub-data/natural-language/about
[14] https://arxiv.org/abs/1904.05862
[15] http://aihub-competition.or.kr/hangeul