清晨好,您是今天最早来到科研通的研友!由于当前在线用户较少,发布求助请尽量完整地填写文献信息,科研通机器人24小时在线,伴您科研之路漫漫前行!

Classifying Malicious Domains using DNS Traffic Analysis

网络钓鱼 恶意软件 僵尸网络 计算机科学 域名系统 计算机安全 领域(数学分析) 审查 互联网 黑名单 域名 万维网 数学 数学分析
作者
Samaneh Mahdavifar,Nasim Maleki,Arash Habibi Lashkari,Matt Broda,Amir H. Razavi
标识
DOI:10.1109/dasc-picom-cbdcom-cyberscitech52372.2021.00024
摘要

Malicious domains are one of the major threats that have jeopardized the viability of the Internet over the years. Threat actors usually abuse the Domain Name System (DNS) to lure users to be victims of malicious domains hosting drive-by-download malware, botnets, phishing websites, or spam messages. Each year, many large corporations are impacted by these threats, resulting in huge financial losses in a single attack. Thus, detecting and classifying a malicious domain in a timely manner is essential. Previously, filtering the domains against blacklists was the only way to detect malicious domains, however, this approach was unable to detect newly generated domains. Recently, Machine Learning (ML) techniques have helped to enhance the detection capability of domain vetting systems. A solid feature engineering mechanism plays a pivotal role in boosting the performance of any ML model. Therefore, we have extracted effective and practical features from DNS traffic categorizing them into three groups of lexical-based, DNS statistical-based, and third party-based features. Third party features are biographical information about a specific domain extracted from third party APIs. The benign to malicious domain ratio is also critical to simulate the real-world scheme where approximately 99% of the traffic is devoted to benign. In this paper, we generate and release a large DNS features dataset of 400,000 benign and 13,011 malicious samples processed from a million benign and 51,453 known-malicious domains from publicly available datasets. The malicious samples span between three categories of spam, phishing, and malware. Our dataset, namely CIC-Bell-DNS2021 replicates the real-world scenarios with frequent benign traffic and diverse malicious domain types. We train and validate a classification model that, unlike previous works that focus on binary detection, detects the type of the attack, i.e., spam, phishing, and malware. Classification performance of various ML algorithms on our generated dataset proves the effectiveness of our model, where we achieved the best results for $k$ -Nearest Neighbors $k$ -NN) with 94.8% and 99.4% F1-Score for balanced data ratio (60/40%) and imbalanced data ratio (97/3%), respectively. Finally, we have gone through feature evaluation using information gain analysis to get the merits of each feature in each category, proving the third party features as the most influential one among the top 13 features. keywords- Malicious Domain, DNS, Feature Engineering, Lexical, Statistical, Third Party, Classification
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
科研通AI6.2应助hll采纳,获得20
1秒前
5秒前
5秒前
欧阳懿完成签到 ,获得积分10
7秒前
ztl完成签到 ,获得积分10
8秒前
hll完成签到,获得积分20
8秒前
蔡龙杰发布了新的文献求助30
9秒前
优雅雪珊发布了新的文献求助10
9秒前
lamb发布了新的文献求助10
15秒前
滕皓轩完成签到 ,获得积分20
22秒前
耍酷季节完成签到,获得积分10
26秒前
Frankie完成签到,获得积分10
30秒前
oleskarabach完成签到,获得积分20
36秒前
37秒前
尹依依发布了新的文献求助10
40秒前
牛牛的马完成签到,获得积分10
41秒前
眯眯眼的安雁完成签到 ,获得积分10
47秒前
热爱科研的小海豹完成签到 ,获得积分10
51秒前
安详的灰狼完成签到 ,获得积分10
51秒前
英姑应助尹依依采纳,获得10
56秒前
xinbadake应助拉长的寒松采纳,获得10
56秒前
003发布了新的文献求助20
59秒前
1分钟前
颖宝老公完成签到,获得积分0
1分钟前
田田完成签到 ,获得积分10
1分钟前
Tonald Yang完成签到 ,获得积分10
1分钟前
lamb完成签到 ,获得积分10
1分钟前
樵木完成签到,获得积分10
1分钟前
噗愣噗愣地刚发芽完成签到 ,获得积分10
1分钟前
迷你的金鱼完成签到,获得积分10
1分钟前
daisy完成签到 ,获得积分10
1分钟前
lt0217完成签到,获得积分10
1分钟前
慈祥的寻芹完成签到,获得积分10
1分钟前
柒柒球完成签到 ,获得积分10
1分钟前
失眠的青寒完成签到,获得积分10
1分钟前
小蘑菇应助慈祥的寻芹采纳,获得10
1分钟前
宇文雨文完成签到 ,获得积分10
1分钟前
奔跑应助科研通管家采纳,获得10
1分钟前
奔跑应助科研通管家采纳,获得10
1分钟前
粗心的语芹完成签到,获得积分10
1分钟前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
Rosenblum, Global Change Biology 800
自動車の空力技術 800
Essentials of Carbohydrate Chemistry and Biochemistry, 4th Edition 800
Organizational Behavior 510
Management and the Arts 510
Matrix Methods in Data Mining and Pattern Recognition Second Edition 510
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 计算机科学 化学工程 工程类 有机化学 物理 复合材料 生物化学 内科学 细胞生物学 基因 遗传学 免疫学 冶金 光电子学 癌症研究
热门帖子
关注 科研通微信公众号,转发送积分 7778308
求助须知:如何正确求助?哪些是违规求助? 9318778
关于积分的说明 20365940
捐赠科研通 7365435
什么是DOI,文献DOI怎么找? 3319203
关于科研通互助平台的介绍 2467070
邀请新用户注册赠送积分活动 2334608