Multimodal Machine Learning Approaches for Voice and Visual Search

  • Authors

    • Sarah Wilson Human Resources DirectorCountry. Enterprise Group, USA. Author

    DOI:

    https://doi.org/10.67228/3142788X/IJMLPA-2020PII2Q9K

    Published 10-03-2020

  • Multimodal Machine Learning, Voice Search, Visual Search, Cross-Modal Retrieval, Deep Learning, CLIP, Visual BERT, AVA, L3-Net, Data Fusion, Representation Learning

    Issue

    Section

    Articles

    How to Cite

    [1]
    S. Wilson, “Multimodal Machine Learning Approaches for Voice and Visual Search”, IJMLPA, vol. 3, no. 2, pp. 01–08, Oct. 2020, doi: 10.67228/3142788X/IJMLPA-2020PII2Q9K.
  • Abstract

    The convergence of voice and visual data has led to significant advancements in search technologies, enabling more intuitive and efficient user interactions. This paper explores multimodal machine learning approaches that integrate voice and visual inputs to enhance search functionalities. We examine various architectures, such as vision-language models like CLIP and VisualBERT, as well as audio-visual models like AVA and L3-Net, highlighting their strengths and applications in search scenarios. Additionally, we address the challenges inherent in aligning and fusing heterogeneous data sources, emphasizing the importance of representation learning and co-learning strategies. Through a comprehensive analysis, we provide insights into the current state of multimodal search systems and propose directions for future research to overcome existing limitations.​

  • References

    [1] Baltrušaitis, T., Ahuja, C., & Morency, L.-P. (2017). Multimodal Machine Learning: A Survey and Taxonomy. arXiv preprint arXiv:1705.09406.

    [2] Morency, L.-P., & Baltrušaitis, T. (2017). Multimodal Machine Learning: Integrating Language, Vision and Speech. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics: Tutorial Abstracts (pp. 3–5). Association for Computational Linguistics.

    [3] Sanguineti, V., Morerio, P., Pozzetti, N., Greco, D., Cristani, M., & Murino, V. (2020). Leveraging acoustic images for effective self-supervised audio representation learning. In Proceedings of the 16th European Conference on Computer Vision (pp. 119–135). Springer.

    [4] Chauhan, A., & Singh, A. (2020). Deep learning-based multimodal emotion recognition from audio, visual, and text modalities: A systematic review of recent advancements and future prospects. ScienceDirect.

    [5] Baltrušaitis, T., Ahuja, C., & Morency, L. P. (2019). Multimodal Machine Learning: A Survey and Taxonomy. IEEE Transactions on Pattern Analysis and Machine Intelligence, 41(2), 423–443. https://doi.org/10.1109/TPAMI.2018.2798607

    [6] Chibelushi, C. C., Deravi, F., & Mason, J. S. D. (2002). A Review of Speech-Based Bimodal Recognition. IEEE Transactions on Multimedia, 4(1), 23–37. https://doi.org/10.1109/6046.985551

    [7] Amir, A., Iyengar, G., Lin, C. Y., Naphade, M., Natsev, A., Neti, C., Nock, H., Smith, J. R., & Tseng, B. (2004). Multimodal Video Search Techniques: Late Fusion of Speech-Based Retrieval and Visual Content-Based Retrieval. Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), 1048–1051.

    [8] de Vries, A. P., Westerveld, T. H. W., & Ianeva, T. (2004). Combining Multiple Representations on the TRECVID Search Task. Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP).

    [9] Zhang, D., & Chen, S. (2000). AVIS: A Connectionist-Based Framework for Integrated Auditory and Visual Information Processing. Information Sciences, 123(1–2), 127–148. https://doi.org/10.1016/S0020-0255(99)00114-0

    [10] Brown, M. G., Foote, J. T., Jones, G. J. F., Spärck Jones, K., & Young, S. J. (1997). Open-Vocabulary Speech Indexing for Voice and Video Mail Retrieval. Proceedings of the ACM International Conference on Multimedia, 307–316. https://doi.org/10.1145/244130.244232

    [11] Morency, L. P., & Baltrušaitis, T. (2017). Multimodal Machine Learning: Integrating Language, Vision and Speech. Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (ACL) – Tutorial Abstracts.

  • Downloads