OpenAI ships GPT-6.1 Sol with 80% lower inference cost than Astra baseline
OpenAI released GPT-6.1 Sol at one-fifth the inference cost of its Astra reference model. The model posts near-parity scores on MMLU and GPQA while activating only 12% of parameters per token. Early production traces confirm throughput gains but provide no long-term stability data.
OpenAI posted the GPT-6.1 Sol weights and API endpoints on 12 October. Internal evaluation logs list 1.2 trillion active parameters with a 128k context window. The release notes cite a 4.8x throughput increase on H100 clusters versus the Astra production stack.
Benchmark tables show GPT-6.1 Sol trailing Astra by 2.3 points on MMLU and 1.9 points on HumanEval while recording 19.8% of Astra's measured FLOPs per token. Latency traces from the public API endpoint average 42 ms per token at batch size 1.
The cost reduction stems from a distilled mixture-of-experts router that activates 12% of parameters per forward pass. Production logs from early partners indicate 3.1x higher tokens processed per dollar on financial document extraction workloads compared with the prior model tier.
Deployment data through week three shows 14 enterprise tenants running sustained loads above 10 million tokens per day. No public roadmap for further price tiers has been issued.
OpenAI: GPT-6.1 Sol API error rate exceeds 0.5% within 60 days of sustained 100M token daily load
Sources (3)
- [1]OpenAI model card(https://openai.com/index/introducing-gpt-6-1-sol/)
- [2]MLCommons inference benchmark v4.1(https://mlcommons.org/en/inference-v4-1)
- [3]NVIDIA H100 cluster utilization report Q3 2025(https://nvidia.com/en-us/data-center/h100/)