Network Traffic Classification Using Machine Learning, Transformer, and Large Language Models

Antari, Ahmad; Abo-Aisheh, Yazan; Shamasneh, Jehad; Ashqar, Huthaifa I.

Computer Science > Machine Learning

arXiv:2503.02141 (cs)

[Submitted on 4 Mar 2025]

Title:Network Traffic Classification Using Machine Learning, Transformer, and Large Language Models

Authors:Ahmad Antari, Yazan Abo-Aisheh, Jehad Shamasneh, Huthaifa I. Ashqar

View PDF

Abstract:This study uses various models to address network traffic classification, categorizing traffic into web, browsing, IPSec, backup, and email. We collected a comprehensive dataset from Arbor Edge Defender (AED) devices, comprising of 30,959 observations and 19 features. Multiple models were evaluated, including Naive Bayes, Decision Tree, Random Forest, Gradient Boosting, XGBoost, Deep Neural Networks (DNN), Transformer, and two Large Language Models (LLMs) including GPT-4o and Gemini with zero- and few-shot learning. Transformer and XGBoost showed the best performance, achieving the highest accuracy of 98.95 and 97.56%, respectively. GPT-4o and Gemini showed promising results with few-shot learning, improving accuracy significantly from initial zero-shot performance. While Gemini Few-Shot and GPT-4o Few-Shot performed well in categories like Web and Email, misclassifications occurred in more complex categories like IPSec and Backup. The study highlights the importance of model selection, fine-tuning, and the balance between training data size and model complexity for achieving reliable classification results.

Subjects:	Machine Learning (cs.LG); Computation and Language (cs.CL); Cryptography and Security (cs.CR)
Cite as:	arXiv:2503.02141 [cs.LG]
	(or arXiv:2503.02141v1 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.2503.02141

Submission history

From: Huthaifa I. Ashqar [view email]
[v1] Tue, 4 Mar 2025 00:18:58 UTC (302 KB)

Computer Science > Machine Learning

Title:Network Traffic Classification Using Machine Learning, Transformer, and Large Language Models

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Network Traffic Classification Using Machine Learning, Transformer, and Large Language Models

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators