DeepSeek unveils efficient V4.1 Flash model claiming superior performance in coding and security tests

DeepSeek has introduced its V4.1 Flash model, which uses a mixture-of-experts architecture to reduce computational demands. The company reports that it outperforms its earlier flagship and competitors on tasks such as coding and cybersecurity, achieving a score of 90.6 on Terminal-Bench 2.1. The release underscores the ongoing race among Chinese AI developers to deliver high-performance models at low operational costs.
The V4.1 Flash utilizes a selective routing architecture, engaging only a small fraction of its 552 billion total parameters for each query. It activates 8 billion parameters for input processing and 16 billion for output generation, significantly reducing computational demands. The model also features native multimodal visual understanding.
On the Terminal-Bench 2.1 evaluation, the new model scored 90.6, surpassing OpenAI's GPT-5.6 Sol and Moonshot AI's Kimi K3. This launch highlights the broader Chinese AI sector's drive to deliver high-performance capabilities while keeping operational expenses minimal, a strategy shaped by hardware costs and export restrictions.
This efficiency-focused release could intensify price competition across China's AI landscape, potentially lowering entry barriers for small firms and independent developers seeking advanced coding or security tools. As inference costs drop, businesses may integrate these models more broadly, though benchmark dominance does not guarantee real-world reliability. Consumers could ultimately see faster, cheaper AI services, while smaller developers might struggle to match the rapid iteration pace of major labs.