Cost vs score
Each point shows one No Skills agent run. The X axis shows percent of monthly Go usage. The Y axis shows the soft score (0–100). Point size shows output tokens. Hover for details.
What we test
Each agent implements one feature in a real Next.js 16 / React 19 app. The feature includes a catalogue, a detail page, server actions, REST handlers, optimistic UI, and auth. Hidden probes check extra cases. The probes follow Vercel's React and Next.js best-practices guide. The code-quality score uses react-doctor static analysis. The methodology page documents both sources. This matches our product work in the Reading Advantage monorepo.
Leaderboard snapshot
OpenCode Go High Usage models, week 2026w35. The host had load during grading. Treat totals as diagnostic. Ranking uses the No Skills score. If no No Skills score exists, ranking uses the Skills score. See results for tokens, cost, and run time.
| # | Model | No Skills | Skills | Passes |
|---|