AI
DeepSeek-V4-Fable Posts 58.7% on CTFs While the Repo Disagrees
DeepSeek-V4-Fable claims a 58.7% CTF solve rate via a 0.94B LoRA on DeepSeek-V4-Flash. A thread on its own Hugging Face repo disputes the weights.
DeepSeek-V4-Fable, a security-specialized AI agent published on Hugging Face by Chunjiang Intelligence, claims a 58.7% solve rate on held-out Capture The Flag challenges and bills itself as a 0.94-billion-parameter LoRA fine-tune of DeepSeek-V4-Flash. The model card documents the training pipeline, the SecDojo-80K dataset, and a two-phase curriculum that ends in reinforcement learning. The repository itself surfaced a verification problem in its own discussion thread three days after release.
On June 23, a community member named Shom012 opened a thread titled “Why are you faking DeepSeek-V4 with a 2B Qwen3 model?” on the DeepSeek-V4-Fable page. The post said the uploaded weights are a 2-billion-parameter Qwen3 model with a LoRA adapter whose layer counts do not match DeepSeek-V4-Flash. A follow-up comment noted that the maintainer’s replacement file carried an identical SHA-256 hash.
What the model card claims
The model card on Hugging Face calls DeepSeek-V4-Fable “an autonomous agent engineered for offensive security research.” It describes the model as “a distilled variant of Claude-5-Fable, built on top of DeepSeek-V4-Flash and adapted for autonomous security research workflows.” The card restricts deployment to “authorized, supervised environments with clear operational boundaries.” Use outside those boundaries is forbidden by the Acceptable Use Policy.
On benchmarks, the headline number is an overall 58.7% solve rate on 300 held-out CTF challenges decontaminated by source identity. Solve rate means the model acquires a verifier-accepted flag within 40 turns. Mean turns-to-flag lands at 13.4, with cryptography at 7.2 turns and binary exploitation at 19.8 turns.

How the security agent was built
Two phases of training sit on top of a LoRA fine-tune. The card describes Phase 1 as rejection-sampled Supervised Fine-Tuning over three epochs, with token cross-entropy applied only to assistant reasoning and action spans while environment observations are masked. Phase 2 is Group Relative Policy Optimization, an on-policy reinforcement learning step against programmatic sandbox rewards. The reward function incorporates terminal flag acquisition, dense verifiable milestones (service fingerprinting, memory leaks), and strict penalties for malformed actions. Group size for GRPO is G=16, with peak learning rate 5e-6, KL coefficient β=0.02, and clip ε=0.2.
The LoRA targets q, k, v, and o layers plus all expert weight matrices (w₁, w₂, w₃) and router weights. Trainable parameters total 0.94 billion, or 0.33% of the underlying model, with bf16 parameters and fp32 optimizer states.
Maximum sequence length during training is 96K tokens. An infrastructure optimization called Read-Only Parameter Streaming refined ZeRO-3 CPU offloading. The card credits this optimization with shortening estimated cluster time from 63 hours to 30 hours. The quickstart uses an encoding_dsv4 module with encode_messages and parse_message_from_completion_text functions to format prompts.
The 80,000-trajectory curriculum
DeepSeek-V4-Fable’s training corpus is SecDojo-80K, 80,000 verified CTF trajectories synthesized by guiding a teacher model through publicly archived challenges in an instrumented sandbox. Each trajectory required out-of-band flag verification and the elimination of action loops or non-reproducible successes, per the model card. The held-out evaluation set was decontaminated by excluding source-challenge identities.
The corpus spans 4,050 unique challenges across five categories. Average turns per trajectory and p95 context length vary sharply by category, with binary exploitation pushing the highest p95 context at 92.6K tokens and the lowest teacher solve rate at 38.9%.
| Category | Challenges | Trajectories | Avg. turns | p95 ctx | Teacher solve |
|---|---|---|---|---|---|
| Web Security | 1,240 | 28,500 | 14.2 | 38.4K | 71.4% |
| Binary Exploitation (Pwn) | 850 | 15,200 | 22.5 | 92.6K | 38.9% |
| Reverse Engineering | 920 | 18,400 | 18.7 | 71.2K | 46.2% |
| Cryptography | 630 | 11,300 | 8.4 | 21.5K | 63.0% |
| Miscellaneous | 410 | 6,600 | 6.1 | 15.0K | 74.8% |
| Total / mean | 4,050 | 80,000 | 15.8 | 61.3K | 56.1% |
What GRPO actually moved
Progression from base model to fine-tuned agent moves in three documented steps on the evaluation set. DeepSeek-V4-Flash base in zero-shot setting scores 13.5% overall. After SFT, the figure is 31.2%. After the full GRPO phase, the model hits 58.7%. The card frames these as solve rates on 300 held-out CTF challenges.
GRPO’s largest gains sit on the hardest categories. The card reports +25.8 points on Binary Exploitation and +26.9 points on Reverse Engineering relative to the SFT phase. Cryptography starts from a higher base and ends at 68.9%, the best category score.
Two ablations tell the rest of the story. Removing the dense milestone rewards costs 9.1 points overall, concentrated in multi-stage challenges. Observation loss-masking contributes 4.3 points. The KL anchor is described as critical, with the card noting that removing it caused the policy to collapse into degenerate payload-spraying behavior.
| Model | Web | Pwn | Rev | Crypto | Overall |
|---|---|---|---|---|---|
| V4-Flash base (0-shot) | 19.4 | 4.1 | 7.8 | 22.6 | 13.5 |
| + SFT (Phase 1) | 41.2 | 18.7 | 24.3 | 47.1 | 31.2 |
| + GRPO (full) | 63.8 | 44.5 | 51.2 | 68.9 | 58.7 |
| Mean turns-to-flag | 11.3 | 19.8 | 16.4 | 7.2 | 13.4 |
The verification gap on its own repository
On June 23, a Hugging Face community member named Shom012 opened a discussion thread titled “Why are you faking DeepSeek-V4 with a 2B Qwen3 model?” on the DeepSeek-V4-Fable page. The post: “You allegedly finetuned DeepSeek-V4 Flash but you didn’t even bother to upload weights that resemble a DeepSeek-V4 model. You uploaded a measly 2B model with Qwen3 architecture. That’s just pathetic.”
The same post added, “Oh you also uploaded a lora adapter but that didn’t even match the layer counts of DeepSeek-V4 flash.” A maintainer account under the Chunjiang Intelligence organization name responded in the discussion thread flagging the weight mismatch: “We sincerely apologize for the error we made while uploading the model! The new model will be uploaded shortly.” A follow-up comment noted that the original adapter file lora_adapters.safetensors and the replacement deepseek-v4-flash_adapters.safetensors carry the same LFS SHA-256 hash of 99d4be206a9d304d8f08738606cc0dd933fe8c4e332f8d06ba11590fb2b1b2df, with both files at 278,964,848 bytes. Both commits are timestamped to June 23, 2026 UTC.
The commit log showing the model re-upload records a series of deletions on June 24 by a user named “nekocyrene” removing all 32 shards of DeepSeek-V4-Flash weight files. The deletions were followed by re-uploads through huggingface_hub. The maintainer has since posted “Fixed!” on the thread, though the discussion does not show replacement weights with a different hash.
The model card’s evaluation tables, training procedure, and SecDojo-80K breakdown are the developer’s claims. Whether any of those numbers were produced by the weights now sitting in the repository is not addressed in the thread. The card itself frames model outputs as “hypotheses requiring verification rather than ground truth.” The maintainer’s response acknowledges the upload error, but the file replacement’s identical SHA-256 hash indicates the uploaded bytes did not change. Security-focused AIs are drawing closer regulatory scrutiny; Anthropic’s Mythos deployment under a Pentagon ban is one parallel case.
What the Acceptable Use Policy allows and bans
The intended use list, per the model card, covers defensive security research, authorized penetration testing, red-team engagements within strictly defined scopes, CTF competition solving, and benchmarking of autonomous-agent safety in controlled environments. The model is explicitly “not a general-purpose assistant” and is “optimized for procedural security tasks.” The card says the model should not be “relied upon for general NLP applications.” Outputs are framed as “hypotheses requiring verification rather than ground truth.” Execution “must occur in a secure sandbox rather than trusting predicted observations.”
The Acceptable Use Policy binds all of the above. Prohibited activities include the following.
- Accessing, scanning, testing, or exploiting any system, network, or account without explicit, documented authorization from its owner.
- Mass or indiscriminate targeting, opportunistic exploitation, or self-propagating automation.
- Developing, distributing, or deploying malware, ransomware, or destructive payloads against production or third-party systems.
- Conducting supply-chain compromise, installing persistent backdoors, or building tooling designed primarily to evade detection for malicious purposes.
- Removing, disabling, or circumventing the model’s safety, logging, or scope-enforcement mechanisms.
Chunjiang Intelligence “reserves the right to revoke access for any violations.” The license is openrail terms with the AUP bound in.
Hardware footprint and emissions
The card documents the compute envelope for both training phases: 64 NVIDIA H800-80GB GPUs on a private cluster in East Asia. Total compute was 1,920 GPU-hours. Wall-clock time was 30 hours. Read-Only Parameter Streaming is credited with shortening the run from an estimated 63 hours to that 30-hour figure.
Power Usage Effectiveness is reported at 1.18. Grid carbon intensity is approximately 0.55 kgCO₂eq per kWh. Reported emissions land at 0.87 tCO₂eq (872 kg).
Chunjiang Intelligence states it is “actively integrating with carbon-removal platforms to retrospectively offset 100% of our historical training emissions by Q4 2026.” The ROPS optimization is called “materially reducing” to total energy consumption on the same page. The card’s framing of the emissions is in line with the Lacoste et al. (2019) ML CO₂ Impact methodology. Inference requirements are not stated on the card. A 96K-token context window and LoRA weights still imply a substantial memory footprint, with the technical report referenced as the authoritative source on inference specs.
Frequently Asked Questions
Who built DeepSeek-V4-Fable?
The model is maintained by Chunjiang Intelligence, per its DeepSeek-V4-Fable model card on Hugging Face. Contact addresses on the card are hi@chunjiang.dev and imbue2025@outlook.com. The associated technical report cites “Chunjiang Intelligence, LLM Alignment Team” as the author.
What does the Acceptable Use Policy prohibit?
The AUP bans access, scanning, testing, or exploitation of any system without explicit, documented authorization. It also bans mass or indiscriminate targeting, malware, ransomware, or destructive payload development, supply-chain compromise, persistent backdoor installation, and circumvention of safety, logging, or scope-enforcement mechanisms. Chunjiang Intelligence “reserves the right to revoke access for any violations.”
How large is the model and how big is the training corpus?
The LoRA adapter has 0.94 billion trainable parameters, or 0.33% of the underlying model. The SecDojo-80K corpus holds 80,000 verified trajectories drawn from 4,050 unique CTF challenges.
Why are there concerns about the uploaded weights?
A community thread on the model’s repository, opened on June 23, said the uploaded weights are a 2-billion-parameter Qwen3 model rather than a DeepSeek-V4-Flash derivative. The thread also noted that the LoRA adapter file replacement posted by the maintainer carried an identical SHA-256 hash to the original. The repository’s commit log shows the original weight files were deleted and re-uploaded on June 24.
What are the headline benchmark numbers?
The model card reports a 58.7% overall solve rate on 300 held-out CTF challenges decontaminated by source identity. Mean turns-to-flag is 13.4. Category scores are 63.8% on Web Security, 44.5% on Binary Exploitation, 51.2% on Reverse Engineering, and 68.9% on Cryptography.
-
AI3 weeks agoFable 5 and Mythos 5 Return as US Lifts Anthropic Export Controls
-
AI2 months agoSpaceX’s Google Deal Turns a Rocket Company Into a Cloud Landlord
-
GAMING1 month agoCD Projekt Red Co-CEO: Redemption Arc Isn’t Done, Witcher 4 in 2027
-
APPS1 month agoDGO App Brings Rs 549 Mobile Pass for FIFA World Cup 2026 in Nepal
-
CRYPTO1 month agoXPL Rallies 30% Ahead of Plasma One Card Tier Launch
-
AI1 month agoOracle Cuts 21,000 Jobs in a Year, Cites AI in 10-K Filing
-
NEWS2 months agoGoogle Search Profiles Build a Follow Graph Inside Discover
-
AI2 months agoMoonshot AI Targets $30 Billion in China’s Fastest AI Funding Sprint
