Deep Learning Models for Document Classification
-
DOI:
https://doi.org/10.67228/30713498/IJADSMC-2021PII3D5VPublished 09-04-2021
Document Classification, Deep Learning, Natural Language Processing, CNN, RNN, Transformer, Text Mining Issue
Section
ArticlesHow to Cite
[1]S. A. Rahman, “Deep Learning Models for Document Classification”, IJADSMC, vol. 4, no. 2, pp. 01–17, Sep. 2021, doi: 10.67228/30713498/IJADSMC-2021PII3D5V.Abstract
Document classification is an essential process of natural language processing (NLP) which presupposes assigning textual documents to predefined categories depending on their contents. As the amount of digital text that needs to be classified has grown exponentially through the sources of social media, scholarly repositories, law archives, news portals and enterprise document management systems, effective and correct document classification has become more important. Most of the common machine learning models such as Naïve Bayes, Support Vector Machines, and k-Nearest Neighbors have proven to be of acceptable performance but they heavily depend on manually crafted features and are not as accurate in detecting semantic and contextual information in text. The latest technology in deep learning greatly altered the methods of document classification as it allowed extracting features and learning representations based on the context. Convolutional neural networks (CNNs) models, recurrent neural networks (RNNs), Long Short Memory networks (LSTMs), Gated Recurrent Units (GRUs) and transformer-based have been used to set the state of the art on benchmark datasets. Such models employ dense word encodings, attention, and hierarchical models to represent document-level semantic structures on both local and global levels. In the present paper, the systematic investigation of the document classification frameworks using deep learning models is offered. It analyzes background information, architectural design, learning process and optimization plans. Moreover, it evaluates the advantages and weaknesses of the different deep learning methods in processing long texts, multi-label classification, domain adaptation, and scalability. The socio-economic metrics are also standard performance metrics, and a single approach to the methodology is suggested that incorporates preprocessing, embedding learning, model training, and evaluation using these metrics. The results of the experiment when exploring representative datasets are addressed to emphasize the trends of the comparative performance. The paper will end by presenting some of the current challenges and the direction of future research and stressing the aspects of explainability, efficiency, and domain robustness.
References
[1] T. Joachims, “Text categorization with support vector machines: Learning with many relevant features,” Proc. 10th European Conf. Machine Learning (ECML), pp. 137–142, 1998.
[2] A. McCallum and K. Nigam, “A comparison of event models for Naïve Bayes text classification,” AAAI Workshop on Learning for Text Categorization, pp. 41–48, 1998.
[3] Y. Yang and X. Liu, “A re-examination of text categorization methods,” Proc. 22nd Int. ACM SIGIR Conf., pp. 42–49, 1999.
[4] G. Salton, A. Wong, and C. S. Yang, “A vector space model for automatic indexing,” Communications of the ACM, vol. 18, no. 11, pp. 613–620, 1975.
[5] R. Johnson and T. Zhang, “Effective use of word order for text categorization with convolutional neural networks,” Proc. NAACL-HLT, pp. 103–112, 2015.
[6] Y. Kim, “Convolutional neural networks for sentence classification,” Proc. EMNLP, pp. 1746–1751, 2014.
[7] T. Mikolov, K. Chen, G. Corrado, and J. Dean, “Efficient estimation of word representations in vector space,” Proc. ICLR, 2013.
[8] J. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural Computation, vol. 9, no. 8, pp. 1735–1780, 1997.
[9] K. Cho et al., “Learning phrase representations using RNN encoder–decoder for statistical machine translation,” Proc. EMNLP, pp. 1724–1734, 2014.
[10] Z. Yang, D. Yang, C. Dyer, X. He, A. Smola, and E. Hovy, “Hierarchical attention networks for document classification,” Proc. NAACL-HLT, pp. 1480–1489, 2016.
[11] A. Vaswani et al., “Attention is all you need,” Proc. Advances in Neural Information Processing Systems (NeurIPS), pp. 5998–6008, 2017.
[12] J. Devlin, M. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” Proc. NAACL-HLT, pp. 4171–4186, 2019.
[13] Y. Liu et al., “RoBERTa: A robustly optimized BERT pretraining approach,” arXiv preprint arXiv:1907.11692, 2019.
[14] Q. Le and T. Mikolov, “Distributed representations of sentences and documents,” Proc. ICML, pp. 1188–1196, 2014.
[15] S. Minaee et al., “Deep learning-based text classification: A comprehensive review,” IEEE Access, vol. 8, pp. 12956–13013, 2020.
Downloads
How to Cite
[1]S. A. Rahman, “Deep Learning Models for Document Classification”, IJADSMC, vol. 4, no. 2, pp. 01–17, Sep. 2021, doi: 10.67228/30713498/IJADSMC-2021PII3D5V.