Identifying Sources and Participants of Propaganda in TikTok Using Machine Learning
DOI:
https://doi.org/10.32515/2664-262X.2025.12(43).1.90-98Keywords:
disinformation, propaganda sources, dataset, RoBERTa model, clustering, potential propaganda participants, set of criteria for identifying propaganda participantsAbstract
The paper presents an approach to identifying sources of disinformation, fakes, and propaganda on the TikTok social network using modern natural language processing (NLP) and artificial intelligence methods. The main goal of the research is to create a system capable of automatically analyzing comments on videos with propaganda content, as well as the source of their distribution and potential participants in propaganda. As part of the research, a corpus of comments in Ukrainian and Russian was manually collected, which were classified as propaganda or neutral. Based on the analysis of the dataset, certain criteria were identified for identifying sources of disinformation and its potential participants, in particular through the use of the Russian language, repeated propaganda narratives, as well as the fakeness of accounts. The steps of the comment preprocessing algorithm are given. Two approaches to analysis were developed: a classification model based on RandomForestClassifier and a clustering model using the KMeans algorithm. Both models use RoBERTa transformers for the respective languages, as well as an additional manually generated set of comment features. A Telegram bot and a graphical interface were built for the convenient use of the system, which allows receiving comments from TikTok, classifying them, and providing the user with analysis results. The proposed system is a relevant tool for information security and combating propaganda in the digital environment.
References
List of References
1. Liu Y., Ott M., Goyal N., Du J., Joshi M., Chen D., Levy O., Lewis M., Zettlemoyer L., Stoyanov V. Roberta: A robustly optimized Bert pretraining approach. 2019. arXiv preprint arXiv:1907.11692. URL: https://doi.org/10.48550/arXiv.1907.11692.
2. Arif M., Tonja A. L., Ameer I., Kolesnikova O., Gelbukh A., Sidorov G., Meque A. G. M. CIC at CheckThat! 2022: Multi-class and Cross-lingual Fake News Detection. CEUR Workshop Proceedings 2022. Pp. 434 – 443.
3. Kalraa S., Vermas P., Sharma Y., Chauhan G. S. Ensembling of Various Transformer Based Models for the Fake News Detection Task in the Urdu Language. Proceedings of the Forum for Information Retrieval Evaluation. 2021.
4. Prytula M. Fine‑tuning BERT, DistilBERT, XLM‑RoBERTa, and Ukr‑RoBERTa models for sentiment analysis of Ukrainian language reviews. Artificial Intelligence. 2024. № 29(2). Pp. 85–97. URL: https://doi.org/10.15407/jai2024.02.085.
5. Panchenko D., Tytarenko S., et al. Evaluation and Analysis of the NLP Model Zoo for Ukrainian Text Classification. In Information and Communication Technologies in Education, Research, and Industrial Applications. Springer. 2022. Pp. 109–123. URL: https://doi.org/10.1007/978-3-031-20834-8_6.
6. Dementieva D., Babakov N., Fraser A. EmoBench‑UA: A Benchmark Dataset for Emotion Detection in Ukrainian. arXiv preprint. 2025. URL: https://doi.org/10.48550/arXiv.2505.23297.
7. Saif M. Mohammad. Ethics Sheet for Automatic Emotion Recognition and Sentiment Analysis.
Computational Linguistics. 2022. № 48(2). Pp. 239–278. URL: https://doi.org/10.48550/arXiv.2109.08256.
8. Shynkarov Y., Solopova V., Schmitt V. Improving Sentiment Analysis for Ukrainian Social Media Code‑Switching Data. In Proceedings of UNLP‑2025 Workshop (COSMUS benchmark). 2025.
9. Haltiuk M., Smywiński‑Pohl A. LiBERTa: Advancing Ukrainian Language Modeling through Pre‑training from Scratch. In Proceedings of the Third Ukrainian Natural Language Processing Workshop. 2024. Pp. 120–128.
10. Доренський О.П., Улічев О.С., Задорожний К.О., Коваленко А.С., Дрєєва Г.М. Концептуальна модель системи інформаційного протиборства координаційного центру з питань національної безпеки і оборони. Центральноукраїнський науковий вісник. Технічні науки. 2024. Вип. 10(41). Ч. 2. С. 23-31. URL: https://doi.org/10.32515/2664-262X.2024.10(41).2.23-31.
11. Лозинська О.В., Марків О.О., Висоцька В.А. Метод виявлення джерел дезінформації на основі ансамблевих моделей машинного навчання. Біоніка інтелекту. 2025. № 1 (102). С. 11-19. URL: https://doi.org/10.30837/ bi.2025.1(102).02.
References
1. Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., & Stoyanov,
V. (2019). Roberta: A robustly optimized Bert pretraining approach. arXiv preprint arXiv:1907.11692. DOI: 10.48550/arXiv.1907.11692.
2. Arif, M., Tonja, A. L., Ameer, I., Kolesnikova, O., Gelbukh, A., Sidorov, G., & Meque, A. G. M. (2022). CIC at CheckThat! 2022: Multi-class and Cross-lingual Fake News Detection. CEUR Workshop Proceedings (pp. 434 – 443).
3. Kalraa, S., Verma, P., Sharma, Y., & Chauhan, G. S. (2021). Ensembling of Various Transformer Based Models for the Fake News Detection Task in the Urdu Language. Proceedings of the Forum for Information Retrieval Evaluation.
4. Prytula, M. (2024). Fine‑tuning BERT, DistilBERT, XLM‑RoBERTa, and Ukr‑RoBERTa models for sentiment analysis of Ukrainian language reviews. Artificial Intelligence, 29(2), 85–97. https://doi.org/10.15407/jai2024.02.085
5. Panchenko, D., Tytarenko, S., et al. (2022). Evaluation and Analysis of the NLP Model Zoo for Ukrainian Text Classification. In Information and Communication Technologies in Education, Research, and Industrial Applications (pp. 109–123). Springer. https://doi.org/10.1007/978-3-031-20834-8_6.
6. Dementieva, D., Babakov, N., & Fraser, A. (2025, May 29). EmoBench‑UA: A Benchmark Dataset for Emotion Detection in Ukrainian. arXiv preprint. https://doi.org/10.48550/arXiv.2505.23297.
7. Saif M. Mohammad (2022). Ethics Sheet for Automatic Emotion Recognition and Sentiment Analysis.
Computational Linguistics, 48(2), 239–278. https://doi.org/10.48550/arXiv.2109.08256.
8. Shynkarov, Y., Solopova, V., & Schmitt, V. (2025). Improving Sentiment Analysis for Ukrainian Social Media Code‑Switching Data. In Proceedings of UNLP‑2025 Workshop (COSMUS benchmark).
9. Haltiuk, M., & Smywiński‑Pohl, A. (2024). LiBERTa: Advancing Ukrainian Language Modeling through Pre‑training from Scratch. In Proceedings of the Third Ukrainian Natural Language Processing Workshop (pp. 120–128).
10. Dorenskyi, O.P., Ulichev, O.S., Zadorozhnyi, K.O., Kovalenko, A.S., & Dreeva, G.M. (2024). The conceptual model of the information counteraction system of the coordination center for national security and defense issues. Central Ukrainian Scientific Bulletin. Technical Sciences, 10(41), Part 2, 23-31. DOI: 10.32515/2664-262X.2024.10(41).2.23-31.
11. Lozynska, O.V., Markiv, O.O., & Vysotska, V.A. (2025). Method for detecting sources of disinformation based on ensemble machine learning models. Bionics of Intelligence, 1(102), 11–19. DOI: 10.30837/ bi.2025.1(102).02.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2025 Olga Lozynska, Oksana Markiv, Victoria Vysotska

This work is licensed under a Creative Commons Attribution 4.0 International License.