Models · Chips and hardware · Alibaba
An open-source engine runs a 125-billion-parameter Qwen model on a 12 GB gaming graphics card at 60-95 tokens per second
Strata, an open-source (MIT-licensed) engine, runs Alibaba’s Qwen3.8-Flash-Next, released in August, on an ordinary gaming PC. The model has 125 billion parameters, but thanks to its mixture-of-experts design only about 6 billion are active per token. The most-used experts sit on the graphics card and all of them in system RAM. By the project’s own measurements, the most compressed version writes 94 tokens per second on a 12 GB RTX 5070, while larger versions land between 53 and 79. It needs at least 32 GB of RAM. No independent test has confirmed the numbers yet.
Sources
- Strata (GitHub), “Strata”, GitHub repository of the open-source inference engine, (github.com)
- Qwen (Alibaba), “Qwen3.8-Flash-Next”, Qwen team's model repository and release note, (github.com)
About this story
This story was posted on Instagram by @jarrus.tech on Oct. 5, 2026.
Spotted an error in this story? [email protected] · Instagram
Short link: thejarrus.com/en/h11-18
This story in Turkish: Açık kaynak bir motor, 125 milyar parametreli Qwen modelini 12 GB’lık bir oyun ekran kartında saniyede 60-95 token hızında çalıştırıyor
On the same topic
Anthropic releases Claude Haiku 5.5 at 90% lower prices
Claude Haiku 5.5, released Oct. 7, costs 90% less than Haiku 4.5 for prompts up to 100,000 tokens. It leads on benchmarks but uses more tokens per task.
OpenAI is rolling out GPT-6 to everyone in ChatGPT
OpenAI began rolling out GPT-6 to everyone in ChatGPT on October 7. Its new Intelligent UI builds answers from text, visuals and interactive elements.
Microsoft opens pre-orders for a $5,999 developer PC that can run 120-billion-parameter models locally
Microsoft opened pre-orders on October 7 for the $5,999 Surface RTX Spark Dev Box, built on Nvidia’s RTX Spark superchip, with shipping starting in November.