Papers
arxiv:2604.19461

Involuntary In-Context Learning: Exploiting Few-Shot Pattern Completion to Bypass Safety Alignment in GPT-5.4

Published on Jun 3
Authors:
,

Abstract

The study introduces Involuntary In-Context Learning, an attack that uses abstract operator framing and few-shot examples to override safety training in large language models, achieving high bypass rates through specific structural patterns.

Safety alignment in large language models relies on behavioral training that can be overridden when sufficiently strong in-context patterns compete with learned refusal behaviors. We introduce Involuntary In-Context Learning (IICL), an attack class that uses abstract operator framing with few-shot examples to force pattern completion that overrides safety training. Through 3479 probes across 10 OpenAI models, we identify the attack's effective components through a seven-experiment ablation study. Key findings: (1)~semantic operator naming achieves 100% bypass rate (50/50, p < 0.001); (2)~the attack requires abstract framing, since identical examples in direct question-and-answer format yield 0%; (3)~example ordering matters strongly (interleaved: 76%, harmful-first: 6%); (4)~temperature has no meaningful effect (46-56% across 0.0--1.0). On the HarmBench benchmark, IICL achieves 24.0% bypass [18.6%, 30.4%] against GPT-5.4 with detailed 619-word responses, compared to 0.0% for direct queries.

Community

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2604.19461
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2604.19461 in a model README.md to link it from this page.

Datasets citing this paper 1

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2604.19461 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.