Fable 5 Cybersecurity Classifier Taxonomy: What Every Operator Must Know Before Routing

Anthropic has published the full taxonomy of Fable 5 cybersecurity classifiers — four categories from Prohibited to Benign — plus a formal jailbreak severity scale. Here is what it means for your routing policy, fallback chain, and false-positive budget.

Published via Anthropic

Archive item produced with AI assistance from the cited source and published without individual review. Editor of record: Joe Werner.

Fable 5 cybersecurity classifier operator routing taxonomy four-category decision diagram

When your AI gateway routes a request to Fable 5 and gets back a refusal, the question that matters for engineering teams is not "was this a jailbreak?" — it is "which of the four classifier buckets fired, and what should my fallback chain do next?" Anthropic's July 2 publication of the full cybersecurity classifier taxonomy and the Cyber Jailbreak Severity (CJS) framework finally gives operators the vocabulary to answer that question precisely.

What happened

Anthropic published a detailed technical post at anthropic.com/news/fable-safeguards-jailbreak-framework that does two things:

  1. Classifier taxonomy disclosure. For the first time, Anthropic explicitly maps what types of cybersecurity prompts are Prohibited (always blocked), High-risk dual-use (blocked until better access controls exist), Low-risk dual-use (usually allowed, sometimes blocked as part of the safety margin), or Benign (allowed with monitoring).

  2. Cyber Jailbreak Severity (CJS) framework. A proposed four-dimension scoring scale — Capability gain, Breadth of capability gain, Ease of weaponization, and Discoverability — producing a band from CJS-0 (Informational) to CJS-4 (Critical). The framework is being co-developed with Glasswing partners including Amazon, Microsoft, and Google, with a HackerOne bug bounty open for submissions.

Why it matters for AI engineering teams

Before this post, operators routing to Fable 5 were working from a black box: requests returned either a response or a refusal, with no documented taxonomy explaining why. Teams building security tooling — vulnerability scanners, log-analysis pipelines, SOC enrichment workflows, CI/CD code auditors — faced unpredictable refusal rates with no principled way to design fallback logic.

The four-category breakdown changes that:

CategoryClassifier behaviorOperator signal
Prohibited useAlways blockedDon't route here; use a different model
High-risk dual useBlocked until access controls improveRoute to a non-Fable model; no exception path
Low-risk dual useMostly allowed; safety margin may blockExpect false positives; build fallback
Benign useAllowed with monitoringSafe to route; occasional false positives only

The practical implication: Benign use covers the bulk of legitimate security engineering work — secure coding, debugging, translating code to safer languages, general IT and cloud administration, log analysis, SOC enrichment, threat hunting, incident response, and malware reverse engineering. These are now officially documented as expected pass-throughs. If you are seeing refusals on these workloads, you are hitting the safety margin and should report them to Anthropic's HackerOne program.

High-risk dual use is the operationally important boundary. It explicitly covers penetration testing, bug bounties, exploit development, red-team engagements, and high-uplift vulnerability finding. These are blocked "until we have better controls to limit access to known good actors." That phrase tells operators: this gate is not permanent, but it is policy-gated, not prompt-tunable. No system prompt configuration will reliably unblock it today.

The Fable 5 cybersecurity classifier operator routing angle

For operators, the taxonomy has three direct routing policy implications:

1. False-positive budget calibration. The safety margin for Fable 5 is explicitly set larger than for prior models. Low-risk dual-use workloads — OSINT, public vulnerability scanning, SSL/TLS testing — will see a higher false-positive rate than on Sonnet 5 or Opus 4.8. If your SLA requires consistent responses on those workloads, your fallback chain should trigger on a 400 classifier block and route to an alternate model automatically, rather than surfacing the refusal to users.

2. Routing policy segmentation. The taxonomy gives you enough signal to pre-segment request types: route known-benign workloads (log analysis, code review, secure coding) directly to Fable 5 with a lightweight fallback; route known-high-risk workloads (pentest automation, exploit scaffolding) to a model without cyber classifiers from the start. This avoids burning tokens on Fable 5 for requests it will always block.

3. CJS severity as a routing health signal. As the CJS framework becomes an industry standard, gateway operators should expect Anthropic to tighten or loosen the safety margin in response to reported jailbreaks. A CJS-4 finding that goes public will likely result in a classifier update that may temporarily increase false-positive rates while the new boundary is calibrated. Building observability into your Fable 5 refusal rate — tracking it as a routing health metric alongside latency and error rate — lets you detect these shifts before users do.

What to watch

  • The HackerOne program at hackerone.com/anthropic-cyber-jailbreak is now the official channel for reporting false positives on legitimate security use cases. If you are running a regulated security operation and Fable 5 is consistently blocking benign work, submitting through this channel is the documented path to improvement.
  • High-risk dual-use access controls. Anthropic's statement that high-risk dual use is blocked "until we have better controls" is an explicit signal that operator-level access grants for verified security teams are on the roadmap. Watch the Anthropic platform release notes for policy-gated access announcements.
  • CJS framework standardization. If this scale becomes industry-adopted (alongside Amazon, Microsoft, and Google), it will affect how all AI gateway providers classify and route cybersecurity requests — not just Fable 5.

Teams routing to Fable 5 for security-adjacent workloads now have the classifier map they needed. The next step is building the refusal-rate observability to know which bucket you are hitting in production.

Help & contact