OpenAI Flags Jailbreak Gains and Safety Trade-Offs in October GPT-6 Models

OpenAI's October GPT-6 release brings Sol to paid ChatGPT users and Luna to Free and Go users, replacing the earlier GPT-5.6 versions in the consumer product. The accompanying system card says the models resist jailbreaks better overall but show statistically significant regressions on some safety benchmarks, including self-harm for Sol and self-harm, gore, and sexual content for Luna. OpenAI rates both models as High capability in cybersecurity and biological/chemical domains, while describing the reviewed failures as generally low severity and noting additional system-level safeguards.
OpenAI’s October GPT-6 rollout splits by tier: Sol reaches paying ChatGPT customers, while Luna serves Free and Go. These consumer models supersede GPT-5.6 Sol and Luna, though Codex and ChatGPT Work still use September builds. The system card rates both as High capability for cybersecurity and biological/chemical risks, below Critical cybersecurity and below High for AI self-improvement.
Testing found improved jailbreak resistance overall, yet Production Benchmarks showed significant declines for Sol on self-harm and for Luna on self-harm, gore, and sexual content. U18 evaluations also declined on age-restricted material, sexual content, and emotional reliance, with Luna additionally on gore. OpenAI calls reviewed failures generally low severity and cites extra runtime safeguards.
The rollout may affect millions of ChatGPT users, especially minors and people discussing self-harm or sexual content, if benchmark regressions translate into real interactions. Paid and free tiers could see different risk profiles. Businesses and developers relying on Codex or Work may face less immediate change because those remain on September builds. OpenAI’s added safeguards may reduce exposure, but independent testing could shape trust and regulatory scrutiny.