The capability-cost curve collapsed 2020–2025. Open-source models that were impossible to run in 2022 now run on a gaming laptop. The frontier is being commoditized.
15–20%Performance gap: Llama 4 vs GPT-4o on standard benchmarks
~$0Marginal cost to run Llama 4 (self-hosted)
$100M+OpenAI's reported API revenue at risk from OS alternatives
18moLag from frontier to open-source equivalent
Choose your depth. The data doesn't change — just the explanation.
Imagine if a car company spent a billion dollars building the best car in the world. Then a year and a half later, anyone could download a slightly-less-good version for free and build it themselves. That's what's happening with AI. Meta releases Llama for free. It's only 15-20% worse than the expensive ChatGPT — and it costs nothing to run if you have a decent computer.
The open-source AI model gap is closing at roughly 18 months. GPT-3.5 was state-of-the-art in late 2022; by mid-2023, open-source alternatives like Llama-2 were within 20-30%. GPT-4 launched March 2023; by early 2025, Llama 4 (405B) reaches within 15-20% on MMLU, HumanEval, and GSM8K. The cost difference is staggering: GPT-4 API costs ~$30/million tokens; self-hosted Llama costs ~$0.50/million tokens on rented hardware. For high-volume enterprise use, this is a 98% cost reduction.
Benchmark sources: MMLU (Hendrycks et al.), HumanEval (Chen et al.), MATH, GSM8K, HellaSwag. Llama 4 Scout/Maverick (2025) benchmarks from Meta AI technical report. GPT-4 benchmark scores from OpenAI system card. The 15-20% gap is real but benchmark-dependent — some tasks (coding, reasoning) show smaller gaps; long-context and multimodal tasks favor closed models. The commoditization risk for OpenAI is that enterprise customers with high query volumes will increasingly self-host. Andreessen Horowitz estimated open-source inference costs at $0.40-0.80/M tokens vs $15-30/M for GPT-4.
Benchmark data: LMSYS Chatbot Arena leaderboard (lmsys.org/leaderboard), Open LLM Leaderboard (huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard). Meta Llama 4 technical report: ai.meta.com/research. OpenAI GPT-4 system card. API cost comparison: OpenAI pricing page, Together.ai, Replicate, Fireworks.ai. Key insight: measure cost-per-benchmark-point, not raw capability. At current trajectory, open-source models will match GPT-5 performance 18-24 months after GPT-5 release.
The Capability-Cost Curve Collapse (2020–2025)
Two years after any frontier model launches, an open-source equivalent emerges at near-zero marginal cost. The timeline is accelerating.
Performance Gap: Open Source vs. Frontier (MMLU Benchmark)
Open LLM Leaderboard data. Gap = frontier score minus best open-source score at same time.
2020
GPT-3 era
35pt gap
Open source 35 MMLU points behind. No viable alternative.
2022
GPT-3.5 / Llama-1
28pt gap
Llama-1 emerges, but impractical for most tasks
2023
GPT-4 / Llama-2
22pt gap
Llama-2 70B becomes production-viable for many tasks
2024
GPT-4o / Llama-3
18pt gap
Llama-3 405B matches or beats GPT-3.5 on most tasks
2025
GPT-4.5 / Llama-4
15–20pt
Llama 4 within 15-20% of frontier on all standard benchmarks
API Cost Per Million Tokens (USD)
Frontier API vs. self-hosted open-source equivalent.
Open-Source Downloads (HuggingFace, Millions)
Monthly downloads of Llama family models.
The Commoditization Threat
OpenAI's business model depends on charging premium prices for frontier capability. But as the gap to open-source narrows, enterprise buyers will increasingly self-host. The 18-month lag gives frontier labs a shrinking window to monetize each capability breakthrough.
Sources
• Meta AI. Llama 4 technical report. ai.meta.com/research (2025).