Exploring Spam Classification with Open-Source Language Models and Real-World Gmail Data

Publication Date

1-1-2026

Document Type

Conference Proceeding

Publication Title

Communications in Computer and Information Science

Volume

2933 CCIS

DOI

10.1007/978-3-032-22205-3_20

First Page

272

Last Page

286

Abstract

Unwanted emails, widely known as spam, pose a significant and persistent problem in daily digital lives. Spam can carry security risks such as phishing attacks, making effective detection crucial. While machine learning(ML) has driven advancements in spam filtering, a key challenge remains: most publicly available datasets for training these filters are outdated. These datasets do not reflect the complex mix of “ham” (legitimate) and spam emails encountered today. To address this, a current dataset was built from scratch using real Gmail data. To truly understand the effectiveness of traditional ML models, which have evolved over the years, they need to be tested against real-world scenarios. Simultaneously, recent breakthroughs in artificial intelligence, particularly with Large Language Models (LLMs), are fundamentally changing how information is interacted with. These powerful models offer new possibilities for understanding and classifying text. This paper presents a direct comparison that evaluates the performance of several established traditional ML models, including Naive Bayes, Support Vector Machines (SVM), and XGBoost. The capabilities of these models are then compared against three distinct LLMs. This work aims to provide clear insights into the capabilities of open-source LLMs in detecting spam in contemporary email environments.

Keywords

Gemma, Gmail, LlaMa, LLM, Mistral, Spam

Department

Computer Science

Share

COinS