DeepSeek · Models · Pricing and access
DeepSeek V4.1 Flash is out: 552 billion parameters, but only 8 billion active per token
It arrives with a new causal encoder-decoder architecture and is almost twice the size of its predecessor, yet the price fell. Peak-hour rates are $0.30 per million input tokens and $1.20 output. On the retired V4 Flash those were $0.44 and $1.32. The weights were released under an MIT license.
Sources
- DeepSeek, “DeepSeek-V4.1-Flash: Smarter, Faster, More Efficient”, launch announcement in API docs, (api-docs.deepseek.com)
- DeepSeek, “Models & Pricing”, API pricing page (current) (api-docs.deepseek.com)
- DeepSeek (Internet Archive kopyası), “Models & Pricing”, archived copy of the pricing page from September 9, 2026 (V4 Flash prices), (web.archive.org)
- Hugging Face (DeepSeek), “deepseek-ai/DeepSeek-V4.1-Flash”, model weights page, (huggingface.co)
About this story
This story was posted on Instagram by @jarrus.tech on Sept. 15, 2026.
Spotted an error in this story? [email protected] · Instagram
Short link: thejarrus.com/en/deepseek-v4-1-flash
This story in Turkish: DeepSeek V4.1 Flash çıktı: 552 milyar parametre, ama token başına sadece 8 milyarı aktif
On the same topic
AWS adds Chinese lab Z.ai’s GLM 5.3 model to Amazon Bedrock
AWS added GLM 5.3, a 753-billion-parameter model from Beijing-based Z.ai (formerly Zhipu AI), to Amazon Bedrock for eligible enterprise customers on October 5.
Nvidia unveils a 64 GB DGX Spark that runs models of up to 100 billion parameters locally
Nvidia announced a 64 GB configuration of its DGX Spark desktop AI computer on October 2, sold through partners from October 23 at a starting price of $4,999.
Higgsfield launches Ads Studio, turning a website into ad images for about 20 cents each
Higgsfield’s Ads Studio, announced October 1, builds a brand kit from a website and makes up to 20 static ads per product at about $0.20 each.