jargon

Applied AI·Security and safety

someone talked your assistant out of its own rules with a role-play framing and it produced exactly what it was built to refuse.

Jailbreak

Draft summary, pending review

Techniques for talking a model out of its safety behaviour: roleplay framing, encoding tricks, many-turn manipulation. Providers patch continuously and the cycle continues. Application-layer implication: never rely on the model's refusals as your only control.

Commonly confused with