Strata achieves 94 tokens/s inference on Qwen3.8-Flash-Next 125B via Q2_0 on RTX 5070 12 GB
Strata ports a 125B Qwen variant to consumer GPUs through aggressive quantization, delivering server-class token rates on 12 GB cards. Measurements confirm 53-94 t/s generation and multi-thousand t/s prompt processing. The approach removes cloud dependency for high-context local agents while exposing memory-bandwidth limits as the remaining constraint.
Strata deploys a custom inference stack that quantizes Qwen3.8-Flash-Next to 2-3 bits per weight and streams layers between system RAM and VRAM. Benchmarks on the GitHub repository record 94 tokens/s generation and 2650 tokens/s prompt ingestion on an RTX 5070 with Q2_0, dropping to 53 tokens/s at IQ3_S. AMD RX 9070 XT reaches 60 tokens/s at the same setting. These rates exceed typical human reading speed while keeping all tokens on-device.
Quantization and paging techniques demonstrated here extend earlier work in GGUF formats and layer-wise offloading. The 125B model size normally requires 80+ GB; Strata reduces footprint to 35-55 GB system RAM plus 12 GB VRAM. Multi-GPU support and MCP server integration allow existing coding agents to call the local instance without cloud round-trips.
Operational impact centers on eliminating API latency and data egress for code and document workloads. A 32 k context prompt now processes at 1100-2650 tokens/s locally, enabling real-time agent loops on single gaming PCs. Sustained 50+ tokens/s generation supports interactive sessions previously gated behind enterprise clusters.
Next release cycles will test 24 GB cards for 100-140 tokens/s and add INT4 kernels. Adoption hinges on whether community forks maintain parity with upstream Qwen releases and whether VRAM fragmentation on Windows remains below 10 %.
Niko1221: Strata 0.2.0 ships 100+ tokens/s on RTX 4090 within 90 days or the project receives no further commits
Sources (3)
- [1]Primary Source(https://github.com/Niko1221/Strata)
- [2]Supporting Source(https://arxiv.org/abs/2309.08872)
- [3]Supporting Source(https://mlcommons.org/en/inference-datacenter/)