Daily AI Coding Value Leaderboard for September 25
The daily ranking evaluates AI models for coding by combining capability scores from benchmarks like SWE-bench and HumanEval with API pricing. Mistral's mistral-nemo leads with a value score of 2275.2, driven by a low cost of $0.0272 per million tokens. The list includes models from OpenAI, DeepSeek, Google, and others.
The leaderboard blends benchmark performance from SWE-bench, HumanEval, and LiveCodeBench with a blended cost metric that weights output tokens more heavily than input tokens. Mistral’s nemo model tops the list largely because its per-million-token price of $0.0272 is dramatically lower than rivals, despite a mid-tier capability score of 62. The full ranking includes 61 models from providers such as OpenAI, DeepSeek, Google, Meta, and Amazon, with capability scores ranging from 30 to 96 and costs spanning roughly $0.027 to over $0.90 per million tokens. Free-tier models receive a significant value boost under this formula.
This daily ranking could influence developer tool choices, especially for startups and independent coders sensitive to API costs. By emphasizing value over raw capability, it may push providers to compete on pricing efficiency rather than just benchmark scores. Smaller teams could benefit from affordable models that perform adequately for routine coding tasks, while enterprises might still prioritize higher-capability models despite higher costs. The methodology’s weighting of output tokens also reflects real-world coding workloads, potentially shaping how future models are priced and marketed.