ISSN 2979-8582 · Article No. 011
Bello Muriana: Information Technology and Resources Center Prince Abubakar Audu University, Anyigba, Nigeria
Ogba Paul: Computer Science Department Prince Abubakar Audu University, Anyigba, Nigeria
In natural language processing (NLP), topic modeling has come to be an effective technique for identifying and extracting hidden topics from big textual datasets. However, the quality and consistency of the generated topics are significantly luenced by the preprocessing techniques performed on the text data prior to model training. This study compares Latent Dirichlet Allocation (LDA) with Non-Negative Matrix infFactorization (NMF) to see how preprocessing techniques affect topic modeling performance. We show that preprocessing has a significant impact on topic coherence using an evaluation metric. Our findings indicate that NMF, when combined with TF-IDF transformation, achieves a higher coherence score (0.5750) than LDA (0.4345), underscoring the significance of preprocessing in enhancing topic quality and interpretability.
Keywords
This article is published under the Creative Commons Attribution 4.0 International License . Free to read, share, and adapt with attribution.
British Journal of Contemporary Research
Open Access · Peer Reviewed · Published by Bexford Publishing Ltd
Browse All Issues