Fable 5 Cybersecurity Classifier Taxonomy: What Every Operator Must Know Before Routing
Anthropic has published the full taxonomy of Fable 5 cybersecurity classifiers — four categories from Prohibited to Benign — plus a formal jailbreak severity scale. Here is what it means for your routing policy, fallback chain, and false-positive budget.
Archive item produced with AI assistance from the cited source and published without individual review. Editor of record: Joe Werner.

When your AI gateway routes a request to Fable 5 and gets back a refusal, the question that matters for engineering teams is not "was this a jailbreak?" — it is "which of the four classifier buckets fired, and what should my fallback chain do next?" Anthropic's July 2 publication of the full cybersecurity classifier taxonomy and the Cyber Jailbreak Severity (CJS) framework finally gives operators the vocabulary to answer that question precisely.
What happened
Anthropic published a detailed technical post at anthropic.com/news/fable-safeguards-jailbreak-framework that does two things:
-
Classifier taxonomy disclosure. For the first time, Anthropic explicitly maps what types of cybersecurity prompts are Prohibited (always blocked), High-risk dual-use (blocked until better access controls exist), Low-risk dual-use (usually allowed, sometimes blocked as part of the safety margin), or Benign (allowed with monitoring).
-
Cyber Jailbreak Severity (CJS) framework. A proposed four-dimension scoring scale — Capability gain, Breadth of capability gain, Ease of weaponization, and Discoverability — producing a band from CJS-0 (Informational) to CJS-4 (Critical). The framework is being co-developed with Glasswing partners including Amazon, Microsoft, and Google, with a HackerOne bug bounty open for submissions.
Why it matters for AI engineering teams
Before this post, operators routing to Fable 5 were working from a black box: requests returned either a response or a refusal, with no documented taxonomy explaining why. Teams building security tooling — vulnerability scanners, log-analysis pipelines, SOC enrichment workflows, CI/CD code auditors — faced unpredictable refusal rates with no principled way to design fallback logic.
The four-category breakdown changes that:
| Category | Classifier behavior | Operator signal |
|---|---|---|
| Prohibited use | Always blocked | Don't route here; use a different model |
| High-risk dual use | Blocked until access controls improve | Route to a non-Fable model; no exception path |
| Low-risk dual use | Mostly allowed; safety margin may block | Expect false positives; build fallback |
| Benign use | Allowed with monitoring | Safe to route; occasional false positives only |
The practical implication: Benign use covers the bulk of legitimate security engineering work — secure coding, debugging, translating code to safer languages, general IT and cloud administration, log analysis, SOC enrichment, threat hunting, incident response, and malware reverse engineering. These are now officially documented as expected pass-throughs. If you are seeing refusals on these workloads, you are hitting the safety margin and should report them to Anthropic's HackerOne program.
High-risk dual use is the operationally important boundary. It explicitly covers penetration testing, bug bounties, exploit development, red-team engagements, and high-uplift vulnerability finding. These are blocked "until we have better controls to limit access to known good actors." That phrase tells operators: this gate is not permanent, but it is policy-gated, not prompt-tunable. No system prompt configuration will reliably unblock it today.
The Fable 5 cybersecurity classifier operator routing angle
For operators, the taxonomy has three direct routing policy implications:
1. False-positive budget calibration. The safety margin for Fable 5 is explicitly set larger than for prior models. Low-risk dual-use workloads — OSINT, public vulnerability scanning, SSL/TLS testing — will see a higher false-positive rate than on Sonnet 5 or Opus 4.8. If your SLA requires consistent responses on those workloads, your fallback chain should trigger on a 400 classifier block and route to an alternate model automatically, rather than surfacing the refusal to users.
2. Routing policy segmentation. The taxonomy gives you enough signal to pre-segment request types: route known-benign workloads (log analysis, code review, secure coding) directly to Fable 5 with a lightweight fallback; route known-high-risk workloads (pentest automation, exploit scaffolding) to a model without cyber classifiers from the start. This avoids burning tokens on Fable 5 for requests it will always block.
3. CJS severity as a routing health signal. As the CJS framework becomes an industry standard, gateway operators should expect Anthropic to tighten or loosen the safety margin in response to reported jailbreaks. A CJS-4 finding that goes public will likely result in a classifier update that may temporarily increase false-positive rates while the new boundary is calibrated. Building observability into your Fable 5 refusal rate — tracking it as a routing health metric alongside latency and error rate — lets you detect these shifts before users do.
What to watch
- The HackerOne program at hackerone.com/anthropic-cyber-jailbreak is now the official channel for reporting false positives on legitimate security use cases. If you are running a regulated security operation and Fable 5 is consistently blocking benign work, submitting through this channel is the documented path to improvement.
- High-risk dual-use access controls. Anthropic's statement that high-risk dual use is blocked "until we have better controls" is an explicit signal that operator-level access grants for verified security teams are on the roadmap. Watch the Anthropic platform release notes for policy-gated access announcements.
- CJS framework standardization. If this scale becomes industry-adopted (alongside Amazon, Microsoft, and Google), it will affect how all AI gateway providers classify and route cybersecurity requests — not just Fable 5.
Teams routing to Fable 5 for security-adjacent workloads now have the classifier map they needed. The next step is building the refusal-rate observability to know which bucket you are hitting in production.

Anthropic Inference Hooks Put a Pre-Inference Gate at the Provider Layer: What It Means for Your Routing Architecture
Anthropic's new Inference Hooks let enterprise organizations intercept every governed Claude prompt before the model runs. For teams already filtering at the gateway layer, this creates a dual-gate architecture that changes where enforcement belongs.

The Jailbreak Severity Framework That Will Reshape AI Gateway Routing Policy
Anthropic, Amazon, Microsoft, and Google are jointly developing a four-dimension jailbreak severity scoring standard. Here is what the new framework means for operator routing fallback policy and safety governance at the API gateway layer.

Configure Claude Code Workload Identity Federation with OIDC
Step-by-step Claude Code workload identity federation setup: connect an OIDC issuer, create Anthropic service accounts and federation rules, then replace static sk-ant keys in CI/CD, gateways, and agent runtimes with short-lived tokens.