Models · Image and video · Agents
Qwen3.8-Omni-Flash is out, handling text, image, audio and video in one model
Alibaba released it on September 18. Rather than stitching separate specialist systems together, it processes all four modes inside one model, with a one million token context window. Thinking mode, tool calling and web search are supported, so a single call can watch a clip and then act on what it saw. Alibaba says its own measurement shows the average across 30 evaluations improved by more than 26% over Qwen3.5-Omni-Plus. No open weights were released, so self-hosting is not an option.
Sources
- Qwen (Alibaba), “Qwen3.8-Omni-Flash: Omni Senses. Agentic Delivery.”, (qwen.ai)
- Alibaba Cloud Community, “Qwen3.8-Omni-Flash: Omni Senses. Agentic Delivery.”, (alibabacloud.com)
- TechNode, “Alibaba’s Qwen releases Qwen3.8-Omni-Flash with 1M-token context”, (technode.com)
About this story
This story was posted on Instagram by @jarrus.tech on Sept. 24, 2026.
Spotted an error in this story? [email protected] · Instagram
Short link: thejarrus.com/en/qwen-omni-flash
This story in Turkish: Qwen3.8-Omni-Flash çıktı, metin, görsel, ses ve videoyu tek modelde işliyor
On the same topic
OpenAI’s image model posts a near-perfect score on text-dense images
On UltraText Bench, a dense-text test led by Westlake University, OpenAI’s GPT Image 2 ranked first of 24 configurations with 99.35 out of 100.
HackerRank’s AI interviewer opens to all customers after more than 500,000 job interviews
HackerRank made its AI interviewer Chakra generally available to customers on October 5, after a roughly six-month beta with more than 500,000 interviews.
AWS adds Chinese lab Z.ai’s GLM 5.3 model to Amazon Bedrock
AWS added GLM 5.3, a 753-billion-parameter model from Beijing-based Z.ai (formerly Zhipu AI), to Amazon Bedrock for eligible enterprise customers on October 5.