Epoch AI Data Records 47% Quarterly Drop in Cost per Fixed AI Performance Level
AI inference costs for fixed performance levels declined 47% per quarter from 2023-2026 per Epoch AI data. The resulting 725-fold drop in 18 months accelerates deployment in cost-sensitive domains and narrows the economic gap between closed frontier systems and open models.
Epoch AI analysis by Emberson and Roodman documents a 47% average quarterly reduction in inference cost for constant benchmark performance over three years. This compounds to a 13-fold annual decline. The GPQA example illustrates a 725-fold reduction in under 18 months. Comparable historical rates for transistors or solar cells were slower and sustained over longer periods.
Data from the same report and cross-referenced with 2024-2025 MLPerf inference submissions show frontier models achieving equivalent accuracy at lower FLOPs per token. Open-weight releases lag frontier cost curves by 6-12 months even when parameter counts match. Lower inference prices directly expand viable use cases in real-time applications where token budgets previously constrained deployment.
Operational effects include expanded access to high-accuracy reasoning for non-enterprise users and reduced per-query margins for API providers. Daily workflows such as code assistance and document summarization become economically viable at higher volumes. The price trajectory compresses the window in which any single model architecture maintains a cost advantage.
Continued 40%+ quarterly declines imply that models meeting 2025 frontier thresholds will reach commodity pricing by late 2027, shifting value capture from inference hardware toward data curation and application integration.
Epoch AI: Median cost per GPQA Diamond point will fall below $0.00001 by Q4 2027.
Sources (2)
- [1]Epoch AI Report: AI Price Trends 2023-2026(https://epoch.ai/reports/price-trends-2026)
- [2]MLPerf Inference v5.0 Results(https://mlcommons.org/en/inference-results/)