阿里巴巴通义千问发布 Qwen3.8-27B 和 Qwen3.8-2.4T-A95B 开放权重
阿里巴巴通义千问发布了 Qwen3.8-27B(一个拥有 270 亿参数的原生多模态稠密模型)和 Qwen3.8-2.4T-A95B(Max 级别模型)的开放权重。27B 模型整体性能超越 Qwen3.7-Plus,支持 262K 原生上下文(可通过 YaRN 扩展至 100 万 token),并采用 Apache 2.0 许可证。两个模型均可在 Hugging Face 和 ModelScope 下载。 此次发布将前沿能力带入紧凑、可本地部署的模型,使开发者无需大规模 GPU 集群即可使用先进 AI。Apache 2.0 许可证以及主流推理框架(vLLM、SGLang、Ollama、LM Studio、Unsloth)的 Day-0 支持降低了商业和研究使用的门槛。这标志着高效开放权重模型可在消费级硬件和边缘设备上运行的趋势。 Qwen3.8-27B 本地运行约需 17GB 内存,使用 SGLang 配合 NVFP4 量化在单张 RTX 5090 上可实现 206 token/s 的解码速度。它可在单张 Blackwell GPU 上以 BF16、FP8 或 NVFP4 精度运行,在 100 万上下文下 GB300 可容纳约 660 万 KV token。该模型原生支持多模态,能理解图像和视频,并支持灵活的思考控制以完成智能体任务。
@Alibaba_Qwen原推文1 张图片We promised open weights for Qwen3.8. Now, time to meet them! 🎉
⚡ Qwen3.8-27B:
- A native multimodal dense model. With just 27B parameters, it outperforms Qwen3.7-Plus overall and shines in real-world coding & office workflows.
- 262K native context, easily extendable to 1M tokens via YaRN.
- Built for builders. Highly efficient, high-quality, and licensed under Apache 2.0.
🚀 The open weights for Qwen3.8-2.4T-A95B (Max-level) have also been released recently.
Whether you're shipping lightweight applications with Qwen3.8-27B locally or building agents with Qwen3.8-2.4T-A95B, they're yours now!
Download, deploy, and build something we haven't imagined yet. 👀👇
- Hugging Face:
https://huggingface.co/collections/Qwen/qwen38
- ModelScope:
https://www.modelscope.cn/collections/Qwen/Qwen38
--- From twitter ---
Laptop-size model, frontier-size leap. 🏃♀️Qwen3.8-27B is live on LM Studio. Try it! @lmstudio
> 引用 @lmstudio: Qwen3.8-27B is here! 🚀
>
> It's a leap in capabilities for a laptop size model.
>
> Requires ~17GB to run locally.
>
> Model page: https://lmstudio.ai/qwen/qwen3.8-27b
--- From twitter ---
Qwen3.8-27B is now part of your everyday life — from smartphones to vehicles. ⚡️🚀Appreciate your work! @MediaTek
> 引用 @MediaTek: Congratulations to the @Alibaba_Qwen on the launch of Qwen 3.8!
>
> MediaTek continues our long-standing collaboration with Qwen, with Day-0 support now available on Dimensity Auto Cockpit C-X1 and the latest Dimensity flagship mobile SoC, bringing smarter and more intuitive on-device AI experiences to more people, from smartphones to vehicles, and helping make agentic AI a seamless part of everyday life.
--- From twitter ---
Qwen3.8-27B is available on @ollama !Any harness you run, the model always delivers. Let's code together, and show us what you build.😎
> 引用 @ollama: Qwen 3.8 27B is now available on Ollama.
>
> It's one of the best open models at this size, and made for agentic tasks and professional work.
>
> Try it directly with the apps & harnesses you use:
>
> Claude Code:
> ollama launch claude --model qwen3.8
>
> OpenCode:
> ollama launch opencode --model qwen3.8
>
> Hermes Agent:
> ollama launch hermes --model qwen3.8
>
> Pi:
> ollama launch pi --model qwen3.8
>
> We have also optimizations for Apple Silicon!
>
> Try it with the model name: qwen3.8:27b-mlx
--- From twitter ---
27B on 17GB RAM. Are you ready to create something incredible? 😎
Thanks for highlighting it! @UnslothAI
> 引用 @UnslothAI: Qwen3.8-27B can now be run locally! ✨
>
> Run on 17GB RAM via Unsloth Dynamic GGUFs.
>
> Qwen3.8-27B is by far the strongest model for its size. We also uploaded NVFP4 quants.
>
> GGUF: https://huggingface.co/unsloth/Qwen3.8-27B-GGUF
> Guide: https://unsloth.ai/docs/models/qwen3.8
--- From twitter ---
Yes, we are back👑, with 206 tok/s on a single RTX 5090!
Amazing Day-0 work from the SGLang team. Give it a try~@sgl_project
> 引用 @sgl_project: The king of small models is back! Qwen3.8-27B from @Alibaba_Qwen is open source, and Day-0 support is live in SGLang:
> - 206.1 tok/s decode on a single RTX 5090, with our NVFP4 plus DSpark
> - 38.28 tok/s decode on DGX Spark
> Qwen3.8-27B raises the bar again for what a small model can do on agentic planning and long-horizon tasks.
>
> Long live the (small model) king! Run it locally with SGLang 👇
--- From twitter ---
One GPU, 1M context, Day-0 ready. Big props to the vLLM team for the seamless integration!👍
Try Qwen3.8-27B on vLLM: @vllm_project
https://recipes.vllm.ai/Qwen/Qwen3.8-27B
> 引用 @vllm_project: 🎉 Qwen3.8-27B is here from @Alibaba_Qwen, and the whole thing fits on a single GPU. Same hybrid backbone as the 2.4T flagship, dense instead of MoE. Day-0 support in vLLM. 🚀
>
> What is in it for serving ✨
>
> - Fits one Blackwell GPU in every precision. Qwen ships BF16 and FP8, the NVFP4 build from @inferact
> - 262K native context, stretching to 1M. At that length one GB300 still has room for roughly 6.6M KV tokens. Six full-length sequences in flight, on one GPU
> - An MTP draft head rides inside the checkpoint, so speculative decoding needs no separate speculator repo. Acceptance on short prompts measured 92.2% in BF16 and 84.8% in FP8
>
> Verified end-to-end on @NVIDIA GB300: BF16 and FP8 at TP4, NVFP4 at TP1, tool calls working, correct generations at 1M context. Two prerequisites, vLLM nightly and transformers 5.8.0+.
>
> 🔗 http://recipes.vllm.ai/Qwen/Qwen3.8-27B
--- From twitter ---
Less than 2 hours to say hi👋. It's almost time!
See you soon: 👀
https://huggingface.co/Qwen/Qwen3.8-27B
--- From twitter ---
Ready right now! 😎Incredible to see Qwen3.8-27B fly on RTX Spark. Download, deploy and create somthing new. @NVIDIARTXSpark
> 引用 @NVIDIARTXSpark: Ready to run Qwen3.8 locally? 👀
>
> Qwen3.8-27B packs powerful AI into an open model developers can download, serve locally, and build with on their own terms.
--- From twitter ---
🎁A gift for developers: Qwen3.8-27B, running locally on AMD from day zero.
Appreciate the work from the AMD team! @AMD
> 引用 @AMD: Qwen3.8 27B brings a new state-of-the-art dense model for local AI development.
> ⚡ Run it on AMD Ryzen™ AI Max+ processors or single Radeon™ AI PRO R9700 card
> ⚡ Experience it with @LMStudio
> ⚡ Turn it into an app with @lemonade_server
>
> Start building with AMD Day 0 support for @Alibaba_Qwen now: https://www.amd.com/en/blogs/2026/run-qwen-3-8-27b-on-amd-ryzen-ai-max-and-radeon-graphics-cards-day-0.html
展开原推文收起原推文
We promised open weights for Qwen3.8. Now, time to meet them! 🎉 ⚡ Qwen3.8-27B: - A native multimodal dense model. With just 27B parameters, it outperforms Qwen3.7-Plus overall and shines in real-world coding & office workflows. - 262K native context, easily extendable to 1M tokens via YaRN. - Built for builders. Highly efficient, high-quality, and licensed under Apache 2.0. 🚀 The open weights for Qwen3.8-2.4T-A95B (Max-level) have also been released recently. Whether you're shipping lightweight applications with Qwen3.8-27B locally or building agents with Qwen3.8-2.4T-A95B, they're yours now! Download, deploy, and build something we haven't imagined yet. 👀👇 - Hugging Face: https://huggingface.co/collections/Qwen/qwen38 - ModelScope: https://www.modelscope.cn/collections/Qwen/Qwen38 --- From twitter --- Laptop-size model, frontier-size leap. 🏃♀️Qwen3.8-27B is live on LM Studio. Try it! @lmstudio > 引用 @lmstudio: Qwen3.8-27B is here! 🚀 > > It's a leap in capabilities for a laptop size model. > > Requires ~17GB to run locally. > > Model page: https://lmstudio.ai/qwen/qwen3.8-27b --- From twitter --- Qwen3.8-27B is now part of your everyday life — from smartphones to vehicles. ⚡️🚀Appreciate your work! @MediaTek > 引用 @MediaTek: Congratulations to the @Alibaba_Qwen on the launch of Qwen 3.8! > > MediaTek continues our long-standing collaboration with Qwen, with Day-0 support now available on Dimensity Auto Cockpit C-X1 and the latest Dimensity flagship mobile SoC, bringing smarter and more intuitive on-device AI experiences to more people, from smartphones to vehicles, and helping make agentic AI a seamless part of everyday life. --- From twitter --- Qwen3.8-27B is available on @ollama !Any harness you run, the model always delivers. Let's code together, and show us what you build.😎 > 引用 @ollama: Qwen 3.8 27B is now available on Ollama. > > It's one of the best open models at this size, and made for agentic tasks and professional work. > > Try it directly with the apps & harnesses you use: > > Claude Code: > ollama launch claude --model qwen3.8 > > OpenCode: > ollama launch opencode --model qwen3.8 > > Hermes Agent: > ollama launch hermes --model qwen3.8 > > Pi: > ollama launch pi --model qwen3.8 > > We have also optimizations for Apple Silicon! > > Try it with the model name: qwen3.8:27b-mlx --- From twitter --- 27B on 17GB RAM. Are you ready to create something incredible? 😎 Thanks for highlighting it! @UnslothAI > 引用 @UnslothAI: Qwen3.8-27B can now be run locally! ✨ > > Run on 17GB RAM via Unsloth Dynamic GGUFs. > > Qwen3.8-27B is by far the strongest model for its size. We also uploaded NVFP4 quants. > > GGUF: https://huggingface.co/unsloth/Qwen3.8-27B-GGUF > Guide: https://unsloth.ai/docs/models/qwen3.8 --- From twitter --- Yes, we are back👑, with 206 tok/s on a single RTX 5090! Amazing Day-0 work from the SGLang team. Give it a try~@sgl_project > 引用 @sgl_project: The king of small models is back! Qwen3.8-27B from @Alibaba_Qwen is open source, and Day-0 support is live in SGLang: > - 206.1 tok/s decode on a single RTX 5090, with our NVFP4 plus DSpark > - 38.28 tok/s decode on DGX Spark > Qwen3.8-27B raises the bar again for what a small model can do on agentic planning and long-horizon tasks. > > Long live the (small model) king! Run it locally with SGLang 👇 --- From twitter --- One GPU, 1M context, Day-0 ready. Big props to the vLLM team for the seamless integration!👍 Try Qwen3.8-27B on vLLM: @vllm_project https://recipes.vllm.ai/Qwen/Qwen3.8-27B > 引用 @vllm_project: 🎉 Qwen3.8-27B is here from @Alibaba_Qwen, and the whole thing fits on a single GPU. Same hybrid backbone as the 2.4T flagship, dense instead of MoE. Day-0 support in vLLM. 🚀 > > What is in it for serving ✨ > > - Fits one Blackwell GPU in every precision. Qwen ships BF16 and FP8, the NVFP4 build from @inferact > - 262K native context, stretching to 1M. At that length one GB300 still has room for roughly 6.6M KV tokens. Six full-length sequences in flight, on one GPU > - An MTP draft head rides inside the checkpoint, so speculative decoding needs no separate speculator repo. Acceptance on short prompts measured 92.2% in BF16 and 84.8% in FP8 > > Verified end-to-end on @NVIDIA GB300: BF16 and FP8 at TP4, NVFP4 at TP1, tool calls working, correct generations at 1M context. Two prerequisites, vLLM nightly and transformers 5.8.0+. > > 🔗 http://recipes.vllm.ai/Qwen/Qwen3.8-27B --- From twitter --- Less than 2 hours to say hi👋. It's almost time! See you soon: 👀 https://huggingface.co/Qwen/Qwen3.8-27B --- From twitter --- Ready right now! 😎Incredible to see Qwen3.8-27B fly on RTX Spark. Download, deploy and create somthing new. @NVIDIARTXSpark > 引用 @NVIDIARTXSpark: Ready to run Qwen3.8 locally? 👀 > > Qwen3.8-27B packs powerful AI into an open model developers can download, serve locally, and build with on their own terms. --- From twitter --- 🎁A gift for developers: Qwen3.8-27B, running locally on AMD from day zero. Appreciate the work from the AMD team! @AMD > 引用 @AMD: Qwen3.8 27B brings a new state-of-the-art dense model for local AI development. > ⚡ Run it on AMD Ryzen™ AI Max+ processors or single Radeon™ AI PRO R9700 card > ⚡ Experience it with @LMStudio > ⚡ Turn it into an app with @lemonade_server > > Start building with AMD Day 0 support for @Alibaba_Qwen now: https://www.amd.com/en/blogs/2026/run-qwen-3-8-27b-on-amd-ryzen-ai-max-and-radeon-graphics-cards-day-0.html












