❗️ "Stopping every AI jailbreak is impossible. So the real question is what to block first."
That's how Yang Jong-heon, who leads S2W's TALON, framed the issue in a recent interview with Seoul Economic Daily.
Jailbreaking — coaxing a model into saying or doing what it was built to refuse — is an old problem. What's changed is where the consequences fall.
"Agentic AI doesn't raise how often a jailbreak succeeds. It raises what a successful one costs, and that grows beyond comparison."
Here are three key takeaways from Yang's interview. 👇
━━━━━━━━━━━━
✔️ The threat has hands now
Earlier attacks were largely limited to extracting a forbidden answer from a chatbot. Today's models connect to real systems through MCP and APIs, reading files, running code, sending emails, calling other services. A jailbroken agent can now take the action itself.
✔️ A 1% gap is not a small gap
Success rates vary by model, and they run higher than most people expect. Nature Communications put GPT-4o at 61%, Gemini 2.5 Flash at 71%, DeepSeek-V3 at 90%. Yang's point isn't the headline figure. It's that a single reliable crack, however narrow, will eventually be exploited.
✔️ Prioritize instead of chasing perfection
Many teams now conduct months-long red-team exercises before launch, and even financial institutions are commissioning outside red teams to probe for weak points. You still can't close every gap that surfaces. The work is deciding which ones matter most.
━━━━━━━━━━━━
Threats move fast. Defensive resources are finite. What closes the distance is an accurate diagnosis, and that's where S2W works — mapping where an enterprise AI service is actually exposed, so security effort goes where it counts.
Why a jailbreak can escalate from stolen data to runaway token costs and full service downtime: Yang walks through it in the full interview. 🧐
📰 [Seoul Economic Daily] "As agentic AI raises jailbreak risk, defend by priority"
🔗 https://bit.ly/3TsnSRj