Aplicação de PLN para classificação de comentários de jogos em língua portuguesa.

dc.contributor.advisorRodríguez, Luís Cuevas
dc.contributor.advisor-latteshttp://lattes.cnpq.br/0083210163583491
dc.contributor.authorLima, Eduardo Peres de
dc.contributor.author-latteshttp://lattes.cnpq.br/6318142415085697
dc.contributor.referee1Rodríguez, Luís Cuevas
dc.contributor.referee1Latteshttp://lattes.cnpq.br/0083210163583491
dc.contributor.referee2Tamayo, Sergio Cleger
dc.contributor.referee2Latteshttp://lattes.cnpq.br/1042990074656681
dc.contributor.referee3Silva, Eduardo Jorge Lira Antunes da
dc.contributor.referee3Latteshttp://lattes.cnpq.br/9569775142466663
dc.date.accessioned2026-09-14T18:26:26Z
dc.date.issued2026-09-19
dc.description.abstractThis work investigates the application of Natural Language Processing (NLP) techniques to the multi-label classification of Portuguese-language digital game reviews, in the context of automatic user opinion analysis. The proposed approach consists of a hybrid pipeline that combines topic modeling (Latent Dirichlet Allocation), automatic labeling using Large Language Models (LLMs), and consensus strategies among multiple models to improve label consistency. Based on this dataset, supervised learning models are evaluated under different configurations of text preprocessing and vector representations. Initially, real user comments were collected and processed using text normalization and cleaning techniques, and subsequently enriched with automatically generated labels, reducing the dependence on manual annotation. Then, classification algorithms such as Support Vector Machines (SVM), Random Forest, and Logistic Regression were applied in multi-label settings. The results indicate significant performance variations across the evaluated configurations, highlighting the impact of preprocessing strategies and the quality of automatically generated labels. Overall, the integration of supervised learning with modern data construction methods shows potential for performance improvement, although challenges remain regarding label consistency and model generalization. As future work, we propose expanding the dataset, adopting more robust consensus strategies among LLMs, and incorporating additional supervised learning models.
dc.description.resumoEste trabalho investiga a aplicação de técnicas de Processamento de Linguagem Natural (PLN) na classificação multirrótulo de comentários de jogos digitais em língua portuguesa, no contexto de análise automática de opiniões de usuários. A proposta consiste na construção de um pipeline híbrido que combina modelagem de tópicos (Latent Dirichlet Allocation), rotulagem automática com Modelos de Linguagem (LLMs) e estratégias de consenso entre modelos para aumentar a consistência dos rótulos gerados. A partir desse conjunto de dados, são avalia dos modelos de aprendizado supervisionado com diferentes configurações de pré-processamento textual e representações vetoriais. Inicialmente, comentários reais foram coletados e processados por técnicas de normalização e limpeza textual, sendo posteriormente enriquecidos com rótulos automáticos, reduzindo a dependência de anotação manual. Em seguida, foram aplicados algoritmos como Support Vector Machines (SVM), Random Forest e Regressão Logística em cenários multirrótulo. Os resultados indicam variações significativas de desempenho entre as configurações avaliadas, evidenciando a influência do pré-processamento e da qualidade dos rótulos gerados automaticamente. De forma geral, observa-se que a integração entre aprendizado supervisionado e métodos modernos de construção de dados apresenta potencial para melhoria de desempenho, embora ainda existam desafios relacionados à consistência dos rótulos e à generalização dos modelos. Como trabalhos futuros, propõe-se a ampliação do conjunto de dados, a adoção de estratégias mais robustas de consenso entre LLMs e a inclusão de novos modelos de aprendizado supervisionado.
dc.identifier.citationLIMA, Eduardo Peres de. Aplicação de PLN para classificação de comentários de jogos em língua portuguesa. Manaus, 2026. 74f. TCC- (Graduação em engenharia de Computação) –Universidade do Estado do Amazonas. Escola Superior de Tecnologia.
dc.identifier.urihttps://ri.uea.edu.br/handle/riuea/8624
dc.language.isopt
dc.publisherUniversidade do Estado do Amazonas
dc.publisher.initialsUEA
dc.relation.referencesABRAGAMES. Relatório da Indústria de Jogos Digitais no Brasil. [S.l.], 2023. Acesso em: 10 jun. 2025. Disponível em: <https//ww.abragame.org/uploads/5/6/8/0/56005537/abragames-pt.pdf>. AppBrain. AppBrain Android Statistics. 2026. <https://www.appbrain.com> . Acesso em: 19 mar. 2026. BERGSTRA, J.; BENGIO, Y. Random search for hyper-parameter optimization. In: Journal of Machine Learning Research. [S.l.: s.n.], 2012. v. 13, p. 281–305. BIAU, G.; SCORNET, E. Theory of random forests. Annual Review of Statistics and Its Application, v. 11, p. 141–163, 2024. BLEI, D. M.; NG, A. Y.; JORDAN, M. I. Latent dirichlet allocation. Journal of Machine Learning Research, v. 3, p. 993–1022, 2003. BROWN, T. B. et al. Language models are few-shot learners. NeurIPS, 2020. BROWNE, C. et al. Cross-validation methods in machine learning: a survey. arXiv preprint, 2020. CHICHE, A.; YITAGESU, B. Part of speech tagging: a systematic review of deep learning and machine learning approaches. Journal of Big Data, Springer, v. 9, n. 1, 2022. CORTES, C.; VAPNIK, V. Support-vector networks. Machine Learning, v. 20, n. 3, p. 273–297, 1995. DEVELOPERS, S. learn. Cross-validation: evaluating estimator performance. 2024. https://scikit-learn.org. GUIDO, R. et al. An overview on the advancements of support vector machine models in healthcare applications: A review. Information, MDPI, v. 15, n. 4, p. 235, 2024. HAN, M. et al. A survey of multi-label classification based on supervised and semi-supervised learning. International Journal of Machine Learning and Cybernetics, v. 14, p. 697–724, 2023. JURAFSKY, D.; MARTIN, J. H. Speech and Language Processing: An Introduction to Natural Language Processing, Computational Linguistics, and Speech Recognition with Language Models. 3. ed. Prentice Hall, 2025. Disponível em: <https://web.staford.edu/~jurafsky/slp3/>. KAMEN, A.; KAMEN, Y. Majority Rules: LLM Ensemble is a Winning Approach for Content Categorization. 2025. Preprint / Technical Report. Work on LLM ensemble voting for content classification. KOHAVI, R. A study of cross-validation and bootstrap for accuracy estimation and model selection. IJCAI, 1995. KREBS, F. et al. Applying multi-class support vector machines: One-vs-one vs. one-vs-all on the uwf-zeekdatafall22 dataset. Electronics, v. 13, n. 19, p. 3916, 2024. Disponível em: https://www.pdpi.com/2079-9292/13/19/3916. LI, Q. et al. A survey on text classification: From traditional to deep learning. ACM Transactions on Intelligent Systems and Technology, v. 13, n. 2, p. 1–41, 2022. LIU, P. et al. Pre-train, prompt, and predict: A systematic survey of prompting methods in nlp. ACM Computing Surveys, 2023. LIU, W. et al. The emerging trends of multi-label learning. IEEE Transactions on Pattern Analysis and Machine Intelligence, v. 44, n. 11, p. 7955–7974, 2022. LUGER, G. F. Artificial Intelligence: Principles and Practice. Cham: Springer Nature, 2024. ISBN 9783031574368. MANNING, C. D.; RAGHAVAN, P.; SCHüTZE, H. Introduction to Information Retrieval. [S.l.]: Cambridge University Press, 2022. MELO, T. de. Análise de comentários das plataformas online de restaurantes michelin no brasil. In: A Produção do Conhecimento nas Ciências da Comunicação, p. 226–238, 2021. MINAEE, S. et al. Deep learning based text classification: A comprehensive review. ACM Computing Surveys, v. 54, n. 3, p. 1–40, 2021. MITCHELL, T. M. Machine Learning. New York: McGraw-Hill, 1997. MURPHY, K. P. Probabilistic Machine Learning: An Introduction. Cambridge, MA: The MIT Press, 2022. ISBN 9780262046824. Disponível em:<https://probml.github.io/pml-book/book1.html>. NETO, J. A. A.; MELO, T. de. Exploring supervised learning models for multi-label text classification in brazilian restaurant reviews. arXiv preprint, 2023. Disponível em: . . . PALANIVINAYAGAM, A.; EL-BAYEH, C. Z.; DAMASEˇ VIČIUS, R. Twenty years of machine-learning-based text classification: A systematic review. Algorithms, MDPI, v. 16, n. 5, p. 236, 2023. PAWARA, P. et al. One-vs-one classification for deep neural networks. Pattern Recognition, v. 108, p. 107528, 2020. ISSN 0031-3203. Disponível em:<https://doi.org/10.1016/j.patcog. 2020.107528> . PETERS, M. E. et al. Modern natural language processing: A survey. arXiv preprint arXiv:2301.12345, 2023. Disponível em: https://arxiv.org/abs/2301.12345 PGB. Pesquisa Game Brasil 2026. 2026. Disponível em:<https://www.pesquisagamebrasil.com.br> . PROKHORENKOVA, L. et al. Catboost: Unbiased boosting with categorical features. Advances in Neural Information Processing Systems, v. 31, 2021. RIFKIN, R.; KLAUTAU, A. In defense of one-vs-all classification. Journal of Machine Learning Research, v. 5, p. 101–141, 2004. RÖDER, M.; BOTH, A.; HINNEBURG, A. Exploring the space of topic coherence measures. Proceedings of WSDM, 2015. RODRIGUEZ, J. D.; PEREZ, A.; LOZANO, J. A. Cross-validation strategies for data with temporal, spatial, hierarchical, or phylogenetic structure. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2010. RUSSELL, S. J.; NORVIG, P. Artificial Intelligence: A Modern Approach. 4. ed. Hoboken, NJ: Pearson, 2021. ISBN 9780134610993. SCHÖLKOPF, B.; SMOLA, A. J. Learning with Kernels: Support Vector Machines, Regularization, Optimization, and Beyond. Cambridge, MA: MIT Press, 2002. SEBASTIANI, F. Machine learning in automated text categorization. ACM Computing Surveys, 2022. SECHIDIS, K.; TSOUMAKAS, G.; VLAHAVAS, I. On the stratification of multi-label data. ECML PKDD, 2011. SIINO, M.; TINNIRELLO, I.; La Cascia, M. Is text preprocessing still worth the time? A comparative survey on the influence of popular preprocessing methods on Transformers and traditional classifiers. Information Systems, Elsevier, v. 121, p. 102342, 2024. SNOEK, J.; LAROCHELLE, H.; ADAMS, R. P. Practical bayesian optimization of machine learning algorithms. In: Advances in Neural Information Processing Systems (NeurIPS). [S.l.: s.n.], 2012. spaCy. Industrial-strength Natural Language Processing in Python. 2025. <https://spacy.io/>. Accessed: 2025-06-07. STEVENS, K. Modern perspectives on cross-validation in machine learning. Machine Learning Review, 2023. STEVENS, K. et al. Exploring topic coherence measures. In: Proceedings of NAACL. [S.l.: s.n.], 2012. SYED, M.; SPRUIT, M. A survey of topic modeling and its applications. IEEE Access, 2020. SZELISKI, R. Computer Vision: Algorithms and Applications. 2. ed. Cham: Springer, 2022. ISBN 9783030343712. TAREKEGN, A. N.; GIACOBINI, M.; MICHALAK, K. A review of methods for imbalanced multi-label classification. Pattern Recognition, Elsevier, v. 118, p. 107965, 2021 TIITTANEN, H. et al. Novel split quality measures for multilabel stratified cross validation. Applied Computing and Intelligence, 2022. TREVISO, M. et al. Efficient methods for natural language processing: A survey. Transactions of the Association for Computational Linguistics, MIT Press, v. 11, p. 826–860, 2023. Disponível em:<https://aclanthology.org/2023.tacl-1.48/> . TURING, A. M. Computing machinery and intelligence. Mind, v. 59, n. 236, p. 433–460, 1950. VAPNIK, V. N. Statistical Learning Theory. New York: Wiley, 1998. VOGEL, L. Natural Language Processing with Python and Transformers. [S.l.]: O’Reilly Media, 2021. WANG, X. et al. Self-consistency improves chain of thought reasoning in language models. ICLR, 2023. WEI, J. et al. Chain-of-thought prompting elicits reasoning in large language models. NeurIPS, 2022. WEI, J. et al. Finetuned language models are zero-shot learners. ICLR, 2022. ZHANG, H. et al. Large language model ensembles for robust text classification. arXiv preprint, 2023. ZHANG, Y. et al. A survey on prompt engineering. arXiv preprint, 2023
dc.rightsAttribution-NonCommercial-NoDerivs 3.0 United Statesen
dc.rights.urihttp://creativecommons.org/licenses/by-nc-nd/3.0/us/
dc.subjectProcessamento de linguagem natural
dc.subjectClassificação multirrótulo
dc.subjectJogos digitais
dc.subjectAprendizado de máquina
dc.subjectModelos de linguagem
dc.subjectRotulagem automática
dc.titleAplicação de PLN para classificação de comentários de jogos em língua portuguesa.
dc.title.alternativeApplication of PLN for Classifying Portuguese-Language Game Comments.
dc.typeTrabalho de Conclusão de Curso

Arquivos

Pacote original

Agora exibindo 1 - 1 de 1
Carregando...
Imagem de Miniatura
Nome:
Aplicação_PLN_classificação_comentários_jogos.pdf
Tamanho:
1.8 MB
Formato:
Adobe Portable Document Format