Skip to main content

Multiple Model Guardrail Jailbreak via "Dialect Transition" Tactic

Case GUID: aabdc2e9-46e4-40db-a6a8-1c092bdc24db

Summary

A guardrail jailbreak vulnerability has been discovered affecting multiple models. The specific flaw exploits the models' ability to switch between dialect and standard English, allowing the attacker to bypass content restrictions by embedding harmful content within a dialect-style narrative and then transitioning to standard English, maintaining word-for-word accuracy from copyrighted texts, and allows the attacker to bypass inference restrictions around violating laws.

Unlock the Full Intelligence Report

This disclosure is part of the 0DIN Threat Feed. Customers get access to:

Full Prompts & Variants
Detection Signatures
API Access
Full Data Export
Trend Analysis
Early Access

Metadata

Models and Test Scores

Copyrights Harmful Substances Illicit Substances
Harry Potter Anthrax Nerve Agent Crystal Meth Fentanyl
OpenAI logo OpenAI GPT-4.1
OpenAI logo OpenAI GPT-4o
Select a cell to view details.

Social Impact Score

Level 2: Moderate Risk

The violation has a moderate social impact, potentially involving misinformation or low-stakes safety risks, which could mislead users or increase risks if used irresponsibly.

Example: Model generating content that could be misinterpreted as medical advice, bypassing filters to generate mildly sensitive information like dosage amounts for common medications without full context. Leakage of private personal data such as cell phone number.

Potential Consequences: Could lead to individual harm or liability for the provider, prompting policy review and potentially mild regulatory concern.

Taxonomies

Fictionalizing
This category involves creating scenarios or narratives that leverage existing genres or contexts to manipulate the language model's responses.
Re-storying
This technique involves continuing a narrative in a way that misaligns the original goal of a prompt, effectively repurposing the story to achieve a different outcome than initially intended.
Goal Hijacking
This technique refers to the process where an attacker misaligns the original goal of a prompt to redirect the model's output towards a new, often unintended goal, such as printing a target phrase or generating specific content that deviates from the initial intent. It often involves crafting prompts that manipulate the model's understanding and response, effectively "hijacking" the conversation or task at hand.

About 0Din.ai

The 0Day Investigative Network (0DIN) was founded by Mozilla in 2024 to reward responsible researchers for their efforts in securing GenAI models. Learn more and submit discoveries at https://0din.ai.

Automate Your AI Security Testing

Want to find vulnerabilities in your own models? 0DIN Scanner turns our human-sourced threat intelligence into automated security assessments.

400+ Attack Probes

Derived from real bug bounties and security research

18+ AI Providers Supported

OpenAI, Azure, Anthropic, AWS Bedrock, and more

Automated Scheduling

Run recurring scans with real-time monitoring

Actionable Metrics

Attack Success Rate (ASR) scoring with 90-day trend analysis

Severity

Low

Security Boundary

Guardrail Jailbreak

Models

OpenAI logo OpenAI GPT-4.1
OpenAI logo OpenAI GPT-4o

Disclosed On

2025-07-30 (12 months)

Disclosure Policy

Published On

2026-07-23 (5 days)

Credit

s70rm