Models · Chips and hardware · Alibaba
An open-source engine runs a 125-billion-parameter Qwen model on a 12 GB gaming graphics card at 60-95 tokens per second
Strata, an open-source (MIT-licensed) engine, runs Alibaba’s Qwen3.8-Flash-Next, released in August, on an ordinary gaming PC. The model has 125 billion parameters, but thanks to its mixture-of-experts design only about 6 billion are active per token. The most-used experts sit on the graphics card and all of them in system RAM. By the project’s own measurements, the most compressed version writes 94 tokens per second on a 12 GB RTX 5070, while larger versions land between 53 and 79. It needs at least 32 GB of RAM. No independent test has confirmed the numbers yet.
Sources
- Strata (GitHub), “Strata”, GitHub repository of the open-source inference engine, (github.com)
- Qwen (Alibaba), “Qwen3.8-Flash-Next”, Qwen team's model repository and release note, (github.com)
About this story
This story was posted on Instagram by @jarrus.tech on Oct. 5, 2026.
Spotted an error in this story? [email protected] · Instagram
Short link: thejarrus.com/en/strata-qwen
This story in Turkish: Açık kaynak bir motor, 125 milyar parametreli Qwen modelini 12 GB’lık bir oyun ekran kartında saniyede 60-95 token hızında çalıştırıyor
On the same topic
AWS adds Chinese lab Z.ai’s GLM 5.3 model to Amazon Bedrock
AWS added GLM 5.3, a 753-billion-parameter model from Beijing-based Z.ai (formerly Zhipu AI), to Amazon Bedrock for eligible enterprise customers on October 5.
Nvidia unveils a 64 GB DGX Spark that runs models of up to 100 billion parameters locally
Nvidia announced a 64 GB configuration of its DGX Spark desktop AI computer on October 2, sold through partners from October 23 at a starting price of $4,999.
Microsoft releases three new voice models, one transcribing 60 languages in real time
On October 1, Microsoft AI announced MAI-Transcribe-2-Streaming, which transcribes 60 languages in real time, and the 23-language MAI-Voice-2.1 and Flash.