Reconhecimento de expressões faciais em tempo real com YOLO: estudo comparativo e desenvolvimento de um protótipo web.

dc.contributor.advisorMelo, Tiago Eugenio de
dc.contributor.advisor-latteshttp://lattes.cnpq.br/9912454472927669
dc.contributor.authorSilva, Paulo Ricardo Ferreira Da
dc.contributor.author-latteshttp://lattes.cnpq.br/8743536671042116
dc.contributor.referee1Melo, Tiago Eugenio de
dc.contributor.referee1Latteshttp://lattes.cnpq.br/9912454472927669
dc.contributor.referee2Costa, Elloá Barreto Guedes da
dc.contributor.referee2Latteshttp://lattes.cnpq.br/6466781778573760
dc.contributor.referee3Figueiredo, Mario Augusto Bessa de
dc.contributor.referee3Latteshttp://lattes.cnpq.br/4299614694581532
dc.date.accessioned2026-09-14T18:41:12Z
dc.date.issued2026-09-19
dc.description.abstractThis work addresses automatic facial expression recognition under uncontrolled conditions, a relevant challenge for applications in human-computer interaction, security, healthcare, and real-time interactive systems. In this context, this study investigated the use of models based on the You Only Look Once (YOLO) architecture, which performs detection and classification in a single stage, combining speed and accuracy. The methodology was conducted in two complementary stages. First, four versions of the YOLO family were trained and comparatively evaluated, specifically YOLOv8n, YOLOv10n, YOLOv11n, and YOLOv12n, using the public dataset “9 Facial Expressions for YOLO”, composed of images captured in uncontrolled scenarios and covering nine facial expression categories. The evaluation considered metrics such as precision, recall, mAP, FPS, GFLOPs, and confusion matrices. In the second stage, the selected model was integrated into a web prototype for real-time facial expression recognition, including webcam capture, backend inference, prediction display in the interface, user feedback registration, and data visualization through a dashboard. The results indicated a slight predictive advantage of YOLOv12n in the experiments performed, while YOLOv10n presented the highest inference speed. However, YOLOv8n was selected for the applied stage because it provided the best balance between accuracy, speed, and practical feasibility, achieving an mAP50 of 0.837, an mAP50–95 of 0.681, and 56.17 FPS. The developed prototype demonstrated adequate functionality in the complete flow of capture, inference, prediction display, feedback registration, and aggregated result analysis. Nevertheless, it was observed that more subtle or ambiguous expressions, pose variations, lighting conditions, and camera quality still influence the perceived performance of the system. Therefore, this work contributes both with a comparative analysis of YOLO architectures applied to facial expression recognition and with the implementation of a functional prototype for real-time inference.
dc.description.resumoEste trabalho aborda o reconhecimento automático de expressões faciais em condições não controladas, um desafio relevante para aplicações em interação humano-computador, segurança, saúde e sistemas interativos em tempo real. Neste contexto, investigou-se o uso de modelos baseados na arquitetura You Only Look Once (YOLO), capazes de realizar detecção e classificação em uma única etapa, combinando rapidez e precisão. A metodologia foi conduzida em duas etapas complementares. Na primeira, foram treinadas e avaliadas comparativamente quatro versões da família YOLO, especificamente YOLOv8n, YOLOv10n, YOLOv11n e YOLOv12n, utilizando o conjunto de dados público “9 Facial Expressions for YOLO”, composto por imagens em cenários não controlados e nove categorias de expressões faciais. A avaliação considerou métricas como precisão, revocação, mAP, FPS, GFLOPs e matrizes de confusão. Na segunda etapa, o modelo selecionado foi integrado a um protótipo web para reconhecimento de expressões faciais em tempo real, com captura por webcam, inferência no backend, exibição das predições na interface, registro de feedbacks do usuário e visualização dos dados em um dashboard. Os resultados mostraram que o YOLOv12n apresentou leve vantagem preditiva nos experimentos realizados, enquanto o YOLOv10n apresentou a maior velocidade de inferência. Entretanto, o YOLOv8n foi selecionado para a etapa aplicada por apresentar o melhor equilíbrio entre acurácia, velocidade e viabilidade prática, alcançando mAP50 de 0,837, mAP50–95 de 0,681 e 56,17 FPS. O protótipo desenvolvido demonstrou funcionamento adequado no fluxo de captura, inferência, exibição das predições, registro de feedbacks e análise agregada dos resultados. Observou-se, contudo, que expressões mais sutis ou ambíguas, variações de pose, iluminação e qualidade da câmera ainda influenciam o desempenho percebido do sistema. Assim, o trabalho contribui tanto com uma análise comparativa de arquiteturas YOLO aplicadas ao reconhecimento de expressões faciais quanto com a implementação de um protótipo funcional para inferência em tempo real.
dc.identifier.citationSILVA, Paulo Ricardo Ferreira Da. Reconhecimento de expressões faciais em tempo real com YOLO: estudo comparativo e desenvolvimento de um protótipo web, Manaus, 2026. 70 f. TCC- (Graduação em Engenharia de Computação) – Universidade do Estado do Amazonas. Escola Superior de Tecnologia.
dc.identifier.urihttps://ri.uea.edu.br/handle/riuea/8635
dc.language.isopt
dc.publisherUniversidade do Estado do Amazonas
dc.publisher.initialsUEA
dc.relation.referencesALBAWI, S.; MOHAMMED, T. A.; AL-ZAWI, S. Understanding of a convolutional neural network. In: IEEE. 2017 International Conference on Engineering and Technology (ICET). 2017. p. 1–6. Disponível em: ⟨https://doi.org/10.1109/ICEngTechnol.2017.8308186⟩. ALI, J.; NAEEM, M. et al. The yolo framework: A comprehensive review of its advances and applications. Computers, v. 13, n. 12, p. 281, 2024. Dispon´ıvel em:⟨https://www.mdpi.com/2073-431X/13/12/281⟩. ALSHAMMARI, A.; ALSHAMMARI, M. E. Emotional facial expression detection using YOLOv8. Engineering, Technology & Applied Science Research, v. 14, n. 5, p.16619–16623, 2024. EKMAN, P. An Argument for Basic Emotions. [S.l.]: Cognition and Emotion,1992. v. 6. 169–200 p. GOODFELLOW, I.; BENGIO, Y.; COURVILLE, A. Deep learning. [S.l.]: MIT press, 2016. JEGHAM, N. et al. Yolo evolution: A comprehensive benchmark and architectural review of yolov12, yolo11, and their previous versions. ResearchGate, February 2025. Preprint. Disponível em: ⟨https://doi.org/10.13140/RG.2.2.15952.83201⟩. JOCHER, G.; CHAURASIA, A.; QIU, J. Ultralytics YOLOv8. 2023. Disponível em:⟨https://github.com/ultralytics/ultralytics⟩. JOCHER, G.; QIU, J. Ultralytics YOLO11. 2024. Disponível em: ⟨https://github.com/ultralytics/ultralytics⟩. KANNA, K. et al. Yolo deep learning algorithm for object detection in agriculture: a review. Journal of Agricultural Engineering, v. 55, n. 4, 2024. Disponível em: ⟨https://www.agroengineering.org/jae/article/view/1641⟩. KUCUKAYAN, O.; YAVUZ, M. et al. Yolo-ihd: Improved real-time human detection system for indoor drones. Sensors, v. 24, n. 3, p. 922, 2024. Disponível em: ⟨https://www.mdpi.com/1424-8220/24/3/922⟩. KUMARI, J.; RAJESH, R.; POOJA, K. M. Facial expression recognition: A survey. Procedia Computer Science, Elsevier, v. 58, p. 486–491, 2015. LECUN, Y.; BENGIO, Y.; HINTON, G. Deep learning. Nature, Nature Publishing Group, v. 521, n. 7553, p. 436–444, 2015. Disponível em: ⟨https://doi.org/10.1038/nature14539⟩. LI, S.; DENG, W. Deep facial expression recognition: A survey. arXiv preprint arXiv:1804.08348, 2018. LI, W.; ZHANG, H.; CHEN, Y. Research on facial expression recognition based on improved yolov8. In: Proc. SPIE 13230, Third International Conference on Computer Vision, Image and Deep Learning. [s.n.], 2024. Disponível em: ⟨https://www.spiedigitallibrary.org/conference-proceedings-of-spie/13230/1323018/⟩. LIAO, J. et al. Facial expression recognition methods in the wild based on fusion feature of attention mechanism and lbp. Journal of Healthcare Engineering, v. 2023, p. 1–21, 2023. Disponível em: ⟨https://www.ncbi.nlm.nih.gov/pmc/articles/PMC10180539/⟩. LIU, P.; QI, B.; BANERJEE, S. Edgeeye: An edge service framework for real-time intelligent video analytics. In: Proceedings of the 1st International Workshop on Edge Systems, Analytics and Networking. New York, NY, USA: Association for Computing Machinery, 2018. (EdgeSys ’18), p. 1–6. LU, H.; XU, J.; ZHENG, H. Explore the facial expression recognition in different complex environment. In: Proceedings of CONF-CDS 2025 Symposium: Data Visualization Methods for Evaluation. [s.n.], 2025. Open access under CC BY 4.0. Disponível em: ⟨https://doi.org/10.54254/2755-2721/2025.PO24884⟩. MAHMOUD, H.; ABOZARIBA, R. A systematic review on webrtc for potential applications and challenges beyond audio video streaming. Multimedia Tools and Applications, Springer, v. 84, p. 2909–2946, 2025. MATSUMOTO, D. More evidence for the universality of a contempt expression. Motivation and Emotion, v. 16, n. 3, p. 363–368, 1992. MOLLAHOSSEINI, A.; HASANI, B.; MAHOOR, M. H. Affectnet: A database for facial expression, valence, and arousal computing in the wild. IEEE Transactions on Affective Computing, IEEE, v. 10, n. 1, p. 18–31, 2017. NAIM, I. et al. Automated analysis and prediction of job interview performance. arXiv preprint arXiv:1504.03425, 2015. Disponível em: ⟨https://arxiv.org/abs/1504.03425⟩. PANTIC, M.; ROTHKRANTZ, L. J. M. Automatic analysis of facial expressions: The state of the art. IEEE Transactions on Pattern Analysis and Machine Intelligence, v. 22, n. 12, p. 1424–1445, 2000. PINTO-COELHO, L. How artificial intelligence is shaping medical imaging technology: A survey of innovations and applications. Bioengineering, v. 10, n. 12, p. 1435, 2023. Disponível em: ⟨https://pmc.ncbi.nlm.nih.gov/articles/PMC10740686/⟩. PLESTED, J.; GEDEON, T. Deep transfer learning for image classification: a survey. arXiv preprint arXiv:2205.09904, 2022. Disponível em: ⟨https://arxiv.org/abs/2205.09904⟩. RAWAT, W.; WANG, Z. Deep convolutional neural networks for image classification: A comprehensive review. Neural computation, MIT Press, v. 29, n. 9, p. 2352–2449, 2017. REDMON, J. et al. You only look once: Unified, real-time object detection. arXiv preprint arXiv:1506.02640, 2016. SAREEN, V.; SEEJA, K. R. Video-based facial emotion recognition using YOLO and Vision Transformer. In: EPJ Web of Conferences. Les Ulis, France: EDP Sciences, 2025. v. 328, p. 01040. Published online: 18 June 2025. Disponível em: ⟨https://doi.org/10.1051/epjconf/202532801040⟩. SASSU, A.; SAENZ-COGOLLO, J. F.; AGELLI, M. Deep-framework: A distributed, scalable, and edge-oriented framework for real-time analysis of video streams. Sensors, MDPI, v. 21, n. 12, p. 4045, 2021. SHAO, J.; QIAN, Y. Three convolutional neural network models for facial expression recognition in the wild. Neurocomputing, Elsevier, v. 355, p. 82–92, 2019. TAWARE, S.; THAKARE, A. Critical analysis on multimodal emotion recognition in meeting the requirements for next generation human computer interactions. International Journal on Recent and Innovation Trends in Computing and Communication, v. 11, n. 9s, p. 523–531, 2023. ISSN 2321-8169. TIAN, Y.; YE, Q.; DOERMANN, D. Yolo12: Attention-centric real-time object detectors. arXiv preprint arXiv:2502.12524, 2025. WANG, A. et al. Yolov10: Real-time end-to-end object detection. arXiv preprint arXiv:2405.14458, 2024. ZHONG, H. et al. Research on real-time teachers’ facial expression recognition based on yolov5 and attention mechanisms. EURASIP Journal on Advances in Signal Processing, v. 2023, n. 55, p. 1–17, 2023. Disponível em: ⟨https://asp-eurasipjournals.springeropen.com/articles/10.1186/s13634-023-01019-w⟩.
dc.rightsAttribution-NonCommercial-NoDerivs 3.0 United Statesen
dc.rights.urihttp://creativecommons.org/licenses/by-nc-nd/3.0/us/
dc.subjectYOLO
dc.subjectReconhecimento de expressões faciais
dc.subjectTempo real
dc.subjectProtótipo web.
dc.subjectVisão computacional
dc.titleReconhecimento de expressões faciais em tempo real com YOLO: estudo comparativo e desenvolvimento de um protótipo web.
dc.title.alternativeReal-Time facial expression recognition using YOLO: a comparative study and development of a Web Prototype.
dc.typeTrabalho de Conclusão de Curso

Arquivos

Pacote original

Agora exibindo 1 - 1 de 1
Carregando...
Imagem de Miniatura
Nome:
Reconhecimento_expressões_faciais_yolo_visão.pdf
Tamanho:
4.23 MB
Formato:
Adobe Portable Document Format