Skip to content
Türkçe

Models · Chips and hardware · Alibaba

An open-source engine runs a 125-billion-parameter Qwen model on a 12 GB gaming graphics card at 60-⁠95 tokens per second

Published: 2 sourcesTürkçe

Strata, an open-source (MIT-licensed) engine, runs Alibaba’s Qwen3.8-Flash-Next, released in August, on an ordinary gaming PC. The model has 125 billion parameters, but thanks to its mixture-of-experts design only about 6 billion are active per token. The most-used experts sit on the graphics card and all of them in system RAM. By the project’s own measurements, the most compressed version writes 94 tokens per second on a 12 GB RTX 5070, while larger versions land between 53 and 79. It needs at least 32 GB of RAM. No independent test has confirmed the numbers yet.

Sources

  1. Strata (GitHub), “Strata”, GitHub repository of the open-source inference engine, (github.com)
  2. Qwen (Alibaba), “Qwen3.8-Flash-Next”, Qwen team's model repository and release note, (github.com)

About this story

This story was posted on Instagram by @jarrus.tech on Oct. 5, 2026.

Spotted an error in this story? [email protected] · Instagram

This story in Turkish: Açık kaynak bir motor, 125 milyar parametreli Qwen modelini 12 GB’lık bir oyun ekran kartında saniyede 60-95 token hızında çalıştırıyor

On the same topic