聚类分析
计算机科学
假新闻
情报检索
自然语言处理
人工智能
互联网隐私
出处
期刊:Digital Scholarship in the Humanities
[Oxford University Press]
日期:2024-04-03
卷期号:39 (2): 790-804
被引量:1
摘要
Abstract Since the Internet is a breeding ground for unconfirmed fake news, its automatic detection and clustering studies have become crucial. Most current studies focus on English texts, and the common features of multilingual fake news are not sufficiently studied. Therefore, this article uses English, Russian, and Chinese as examples and focuses on identifying the common quantitative features of fake news in different languages at the word, sentence, readability, and sentiment levels. These features are then utilized in principal component analysis, K-means clustering, hierarchical clustering, and two-step clustering experiments, which achieved satisfactory results. The common features we proposed play a greater role in achieving automatic cross-lingual clustering than the features proposed in previous studies. Simultaneously, we discovered a trend toward linguistic simplification and economy in fake news. Furthermore, fake news is easier to understand and uses negative emotional expressions in ways that real news does not. Our research provides new reference features for fake news detection tasks and facilitates research into their linguistic characteristics.
科研通智能强力驱动
Strongly Powered by AbleSci AI