Skip to main navigation Skip to search Skip to main content

Small Language Model Jailbreak Defender Framework

Research output: Chapter in Book/Report/Conference proceedingConference contribution

Abstract

Large Language Models (LLMs) enable powerful language and task automation but create scalable avenues for abuse, including jailbreaks and fine tuning-driven safety collapse. Small Language Models (SLMs) are increasingly favored for cost, latency, and deployment flexibility, yet exhibit weaker safety priors and heightened vulnerability to adversarial adaptation. Existing defenses are piecemeal and often degrade task utility or efficiency. This paper proposes a modular security alignment framework that enables fine tuned SLMs to approach LLM level robustness while preserving domain performance and low latency. The framework integrates multi model dataset auditing, adversarial artifact conditioning (e.g., toxic token extraction), and an SLM based guardrail wrapper that performs multi step chain of thought intent and output judging. We outline an empirical evaluation agenda focused on reducing attack success rates and unsafe compliance while maintaining high precision and recall on benign tasks and sub second response times, enabling safer SLM deployment in security sensitive domains. Making the SLM capable to perform as an LLM in a specific domain knowledge and similar safety/security levels, this will open the door to scalable solutions with low budgets.

Original languageEnglish
Title of host publicationSoutheastCon 2026
PublisherInstitute of Electrical and Electronics Engineers Inc.
ISBN (Electronic)9798331546427
DOIs
StatePublished - 2026
Event2026 IEEE SoutheastCon, SoutheastCon 2026 - Hybrid, Huntsville, United States
Duration: Feb 20 2026Mar 15 2026

Publication series

NameConference Proceedings - IEEE SOUTHEASTCON
ISSN (Print)1091-0050
ISSN (Electronic)1558-058X

Conference

Conference2026 IEEE SoutheastCon, SoutheastCon 2026
Country/TerritoryUnited States
CityHybrid, Huntsville
Period2/20/263/15/26

Bibliographical note

Publisher Copyright:
© 2026 IEEE.

ASJC Scopus Subject Areas

  • Software
  • Control and Systems Engineering
  • Signal Processing
  • Computer Networks and Communications
  • Electrical and Electronic Engineering

Keywords

  • Fine Tuning
  • Jailbreak
  • Security
  • Small Language Model

Fingerprint

Dive into the research topics of 'Small Language Model Jailbreak Defender Framework'. Together they form a unique fingerprint.

Cite this