Network traffic classification from handcrafted features to foundation models: A comprehensive survey
Computer Networks, cilt.287, 2026 (SCI-Expanded, Scopus)
- Yayın Türü: Makale / Kısa Makale
- Cilt numarası: 287
- Basım Tarihi: 2026
- Doi Numarası: 10.1016/j.comnet.2026.112578
- Dergi Adı: Computer Networks
- Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus, ABI/INFORM, Aerospace Database, Applied Science & Technology Source, Compendex, INSPEC, Library, Information Science & Technology Abstracts (LISTA), zbMATH, Information Science & Technology Abstracts (LISTA), EBSCO Communication Source, Business Source Ultimate (EBSCO), Communication Source (EBSCO), Engineering Source (EBSCO), Technology Collection (ProQuest)
- Anahtar Kelimeler: Concept drift, Deep learning, Encrypted traffic, Foundation models, Network traffic classification, Pre-trained models, QUIC, TLS 1.3
- Hacettepe Üniversitesi Adresli: Evet
Özet
Network traffic classification underpins network management, security enforcement, and Quality of Service (QoS) provisioning. The near-universal adoption of encryption has eliminated the observable signals on which traditional inspection techniques relied. Today, over 95% of web traffic is served over HTTPS, QUIC powers more than 35% of websites, and Encrypted Client Hello looms on the horizon. These shifts have driven a paradigm shift toward learning-based methods and, most recently, foundation models (FM). This systematic review synthesizes 157 primary studies selected through a structured literature search across five major databases (2015–2025, with an update sweep through January 2026). We propose a dual-dimensional taxonomy that organizes the literature along two axes. The method axis spans port-based inspection, statistical machine learning (ML), deep learning (DL), and traffic FMs. The challenge axis covers concept drift, open-world classification, label scarcity, privacy preservation, adversarial robustness, explainability, and deployment efficiency. Our analysis reveals three central findings: (i) FMs consistently outperform supervised DL by 2–10 percentage points while requiring far less labeled data; (ii) the gap between closed-world and open-world performance remains the most critical barrier to operational deployment; and (iii) no existing model jointly addresses protocol heterogeneity, concept drift, and privacy preservation. We conclude with a structured research roadmap identifying eight open problems and actionable recommendations for both researchers and practitioners.