Abstract
Large Language Models (LLMs) enable powerful language and task automation but create scalable avenues for abuse, including jailbreaks and fine tuning-driven safety collapse. Small Language Models (SLMs) are increasingly favored for cost, latency, and deployment flexibility, yet exhibit weaker safety priors and heightened vulnerability to adversarial adaptation. Existing defenses are piecemeal and often degrade task utility or efficiency. This paper proposes a modular security alignment framework that enables fine tuned SLMs to approach LLM level robustness while preserving domain performance and low latency. The framework integrates multi model dataset auditing, adversarial artifact conditioning (e.g., toxic token extraction), and an SLM based guardrail wrapper that performs multi step chain of thought intent and output judging. We outline an empirical evaluation agenda focused on reducing attack success rates and unsafe compliance while maintaining high precision and recall on benign tasks and sub second response times, enabling safer SLM deployment in security sensitive domains. Making the SLM capable to perform as an LLM in a specific domain knowledge and similar safety/security levels, this will open the door to scalable solutions with low budgets.
| Original language | English |
|---|---|
| Title of host publication | SoutheastCon 2026 |
| Publisher | Institute of Electrical and Electronics Engineers Inc. |
| ISBN (Electronic) | 9798331546427 |
| DOIs | |
| State | Published - 2026 |
| Event | 2026 IEEE SoutheastCon, SoutheastCon 2026 - Hybrid, Huntsville, United States Duration: Feb 20 2026 → Mar 15 2026 |
Publication series
| Name | Conference Proceedings - IEEE SOUTHEASTCON |
|---|---|
| ISSN (Print) | 1091-0050 |
| ISSN (Electronic) | 1558-058X |
Conference
| Conference | 2026 IEEE SoutheastCon, SoutheastCon 2026 |
|---|---|
| Country/Territory | United States |
| City | Hybrid, Huntsville |
| Period | 2/20/26 → 3/15/26 |
Bibliographical note
Publisher Copyright:© 2026 IEEE.
ASJC Scopus Subject Areas
- Software
- Control and Systems Engineering
- Signal Processing
- Computer Networks and Communications
- Electrical and Electronic Engineering
Keywords
- Fine Tuning
- Jailbreak
- Security
- Small Language Model
Fingerprint
Dive into the research topics of 'Small Language Model Jailbreak Defender Framework'. Together they form a unique fingerprint.Cite this
- APA
- Standard
- Harvard
- Vancouver
- Author
- BIBTEX
- RIS