MobbleOpen in Mobble ⇢
Technology · Artificial intelligence · published 2026-10-08 · via Help Net Security

Anthropic's Claude Haiku 5.5 improves prompt-injection resistance and cyber safeguards

Image via Help Net Security
Image via Help Net Security

Anthropic has released Claude Haiku 5.5, a faster model for repetitive tasks, with stronger defenses than Haiku 4.5. In tests with safeguards disabled, it achieved code execution in a small number of Chrome V8 runs and completed 3.3% of multi-stage cyber challenges, far behind Sonnet 5.5 and Opus 5.5. The model also reduced false refusals and improved handling of some harmful-request evaluations.

Expanded Detail

EXPANDED:

Anthropic positions Claude Haiku 5.5 as a speed-oriented option for routine, repetitive work. With safeguards turned off, it reached arbitrary code execution in 4 of 410 Chrome V8 trials and solved 3.3% of multi-stage cyber exercises, versus 46.1% for Sonnet 5.5 and 67.6% for Opus 5.5. Its cyber guardrails sit between Haiku 4.5 and higher-end models, allowing more defensive use while still restricting penetration testing; vetted professionals may seek relaxed limits.

On safety, a near-final Claude.ai prompt yielded a 99.71% harmless-response rate for harmful requests, while harmless-request false refusals fell to 0.82% from 3.05%. In Claude Code, it refused 84.3% of malicious coding requests, up from 66.6%, and about 82.6% in computer-use tests, up from 58.9%. Anthropic calls it its most injection-resistant Haiku yet, though it trails Sonnet 5.5 and Opus 5.5 on one Gray Swan benchmark.

Count words? First para: Anthropic(1) positions2 Claude3 Haiku4 5.5? maybe 5 as token? Let's count roughly. "Anthropic positions Claude Haiku 5.5 as a speed-oriented option for routine, repetitive work." 13? With safeguards turned off, it reached arbitrary code execution in 4 of 410 Chrome V8 trials and solved 3.3% of multi-stage cyber exercises, versus 46.1% for Sonnet 5.5 and 67.6% for Opus 5.5. ~35. Its cyber guardrails sit between Haiku 4.5 and higher-end models, allowing more defensive use while still restricting penetration testing; vetted professionals may seek relaxed limits

Expanded detail and Context are AI-generated analysis; the linked article remains the authoritative source.
Read the full article at Help Net Security →
Related stories
Anthropic expands Claude cyber access with tiered verification · Artificial intelligence
AI Briefing: Samsung Memory Gains, Anthropic's Budget Model, and On-Device AI PCs · Artificial intelligence
This summary is Al-enhanced to contain extended analysis and broader social context. The original is {NAME); the linked article is the authoritative source. Original headline: “Anthropic’s new budget model gets much better at ignoring hidden commands.” Browse more stories.