AI Coding Value Leaderboard for September 23
The September 23 ranking of AI coding models by value again shows Mistral's mistral-nemo in first place with a score of 2275.2. The list includes models from OpenAI, DeepSeek, and Google, with pricing and capability scores determining the order. The value score is calculated by dividing capability by a blended cost per million tokens.
The value metric divides each model's capability score—drawn from benchmarks like SWE-bench and HumanEval—by a blended per-million-token cost weighted 25% input and 75% output, reflecting coding workloads. This favors cheaper models; mistral-nemo's $0.0272 rate yields a commanding lead despite a modest 62 capability rating.
High-capability models don't necessarily top the list. DeepSeek's v4-flash scores 91 in capability but ranks seventh at $0.1551 per million tokens, while OpenAI's gpt-oss-120b, rated 93, sits 35th. The 61-model field spans Amazon, Nvidia, Cohere, and IBM, with free-tier models receiving a substantial boost.
This ranking could reshape how developers and companies choose coding assistants, potentially pushing demand toward smaller, cheaper models for routine tasks while reserving premium models for complex work. Organizations may recalibrate budgets as value scores highlight cost-efficiency gaps, though capability benchmarks don't capture every real-world nuance. Individual developers could benefit from lower-cost options, while providers face pressure to price competitively or differentiate on specialized strengths.