Open vs closed-source gap closed to single digits in April; Gemini 4 Argon retook the lead in October
Open-weight vs closed-source gap narrowed to single digits in April. Gemini 4 Argon and Claude Opus 5.5 retook the lead in October, but 80% of tasks are still served well by top open models.
Source: thorstenmeyerai.com
Open vs closed-source gap closed to single digits in April; Gemini 4 Argon retook the lead in October
The story is not "open-source won" — it's that for 80% of tasks, open weights are already enough.
October state
| Model | Vendor | Type | Status |
|---|---|---|---|
| Gemini 4 Argon | Closed | 10/1 launch · #1 Text Arena | |
| Claude Sonnet 5.5 | Anthropic | Closed | 10/2 · 30% faster, 30% cheaper |
| GPT-6.1 Sol | OpenAI | Closed | 10/4 · $2/$10 |
| Claude Opus 5.5 | Anthropic | Closed | April |
| Qwen 3.7 | Alibaba | Apache 2.0 | H1 flagship; Qwen 4 in training |
| DeepSeek V4.1 Flash | DeepSeek | MIT | Released · 1.6T / 49B active |
| GLM-5.3 | Z.ai | MIT | Anthropic called "most cyber-capable" |
| Llama 4.5 Maverick | Meta | Llama Community | Meta may not ship a next gen |
| Gemma 4 31B | Apache 2.0 | Clean license |
What happened in April
Thorsten Meyer's April piece "Single Digits": open vs closed on mainstream benchmarks narrowed from 20-30% to single digits (<10%).
The point is not "open won" — it's that once capability catches up, open weights give you three more things:
- ✅ Private data stays in house
- ✅ Marginal cost of calls → 0
- ✅ Tunable / offline / auditable
Gemini 4 Argon retook the closed-source lead in October — but for 80% of tasks (coding / translation / summarization / extraction) the top open-weight flagships are enough.
Three judgment points
1️⃣ Closed-source leaderboard churn ≠ what you actually use. You'll use 10% of Gemini 4 Argon's 1M output 2️⃣ Open ≠ free. Qwen 3.5 27B is ~130GB, single 4090 needs sharding, real full-local needs 64GB+ cards 3️⃣ Meta not shipping next gen — open camp loses a major player. Qwen / DeepSeek / GLM fill in