Cross-Lingual Voice Search Systems using Transformer Models

  • Authors

    • Arvind Menon Senior Project Manager, Infosys Ltd., India. Author
    • Karthik Raman Software Architect, Tata Consultancy Services, India. Author

    DOI:

    https://doi.org/10.67228/3142788X/IJMLPA-2018PII1M4P

    Published 12-04-2018

  • Cross-Lingual Voice Search, Transformer Models, Speech Recognition, Multilingual Systems, Natural Language Processing, Speech Processing, Voice Search Optimization

    Issue

    Section

    Articles

    How to Cite

    [1]
    A. Menon and K. Raman, “Cross-Lingual Voice Search Systems using Transformer Models”, IJMLPA, vol. 1, no. 2, pp. 01–08, Dec. 2018, doi: 10.67228/3142788X/IJMLPA-2018PII1M4P.
  • Abstract

    The evolution of Transformer models has significantly impacted various domains of artificial intelligence, particularly in natural language processing and speech recognition. Their ability to capture long-range dependencies and process sequential data efficiently makes them ideal candidates for cross-lingual voice search systems. This paper investigates the application of Transformer architectures in developing voice search systems capable of seamlessly operating across multiple languages. We explore the challenges inherent in cross-lingual voice search, review existing Transformer-based models tailored for this purpose, and propose methodologies to enhance their effectiveness. Through comprehensive analysis and experimentation, we aim to provide insights into the potential of Transformer models to bridge linguistic barriers in voice search applications.​

  • References

    [1] Ashish Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention Is All You Need,” in Proceedings of the 31st International Conference on Neural Information Processing Systems (NIPS), Long Beach, CA, USA, 2017, pp. 6000–6010.

    [2] Shubham Toshniwal, T. N. Sainath, R. J. Weiss, B. Li, P. J. Moreno, E. Weinstein, and K. Rao, “Multilingual Speech Recognition with a Single End-to-End Model,” arXiv preprint arXiv:1711.01694, 2017.

    [3] Tom Sercu, G. Saon, J. Cui, X. Cui, B. Ramabhadran, B. Kingsbury, and A. Sethy, “Network Architectures for Multilingual Speech Representation Learning,” in Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), New Orleans, LA, USA, 2017, pp. 5625–5629.

    [4] P. Matejka, L. Burget, O. Glembek, P. Schwarz, V. Hubeika, M. Fapšo, T. Mikolov, and J. Černocký, “Multilingually Trained Bottleneck Features in Spoken Language Recognition,” Computer Speech & Language, vol. 46, pp. 252–267, 2017.

    [5] G. Saon, H. Soltau, D. Nahamoo, and M. Picheny, “Speaker Adaptation of Neural Network Acoustic Models Using i-Vectors,” in IEEE Workshop on Automatic Speech Recognition and Understanding (ASRU), 2013.

    [6] H. Soltau, H. Liao, and H. Sak, “Neural Speech Recognizer: Acoustic-to-Word LSTM Model for Large Vocabulary Speech Recognition,” in Proceedings of Interspeech, 2017.

    [7] D. Yu and L. Deng, Automatic Speech Recognition: A Deep Learning Approach. London, UK: Springer, 2015.

    [8] X. He and L. Deng, “Deep Learning Methods for Speech Recognition,” in Springer Handbook of Speech Processing, Springer, 2017, pp. 1–28.

    [9] M. J. F. Gales and S. J. Young, “The Application of Hidden Markov Models in Speech Recognition,” Foundations and Trends in Signal Processing, vol. 1, no. 3, pp. 195–304, 2008.

    [10] I. Szöke, M. Fapšo, M. Karafiát, and J. Černocký, “BUT System for NIST Open Keyword Search 2014,” in Proceedings of the IEEE Spoken Language Technology Workshop (SLT), 2014.

    [11] K. Yu, M. Gales, L. Wang, and P. C. Woodland, “Unsupervised Training and Directed Manual Transcription for LVCSR,” Speech Communication, vol. 52, no. 7–8, pp. 652–663, 2010.

    [12] H. Bourlard and N. Morgan, Connectionist Speech Recognition: A Hybrid Approach. Boston, MA, USA: Kluwer Academic Publishers, 1994.

  • Downloads