8月15日2026 · 星期六

从 48 条抓取中筛选 12 条 · twitter × 5 账号 · 08:24 UTC 生成

今日信号 · 高度即评分 · 点击直达


  1. 阿里巴巴通义千问发布 Qwen3.8-27B 和 Qwen3.8-2.4T-A95B 开放权重9.0
  2. Qwen3.8-Max:2.4万亿参数MoE模型在多平台Day 0上线9.0
  3. Anthropic 将为 Claude 实施文本水印以符合欧盟 AI 法案8.0
  4. Anthropic 根据负责任扩展政策发布第二份风险报告8.0
  5. Qwen3.8 借助 TokenSpeed 大规模运行,性能比 TP16 提升 30% 以上8.0
  6. 阿里巴巴通义千问发布开源模型 Qwen3.8-2.4T-A95B,多家平台提供 Day-0 支持8.0
  7. Tailscale 发现隐藏16年的 SQLite WAL 重置缺陷8.0
  8. Pi 作者 Armin Ronacher 称赞 DeepSeek Harness 激发设计反思7.0
  9. Qwen在Hugging Face开放模型报告中领跑本地推理7.0
  10. oil-motion:AI驱动的滚动触发动画技能7.0
  11. SkillUI 从网站提取完整设计系统供 Claude Code 使用7.0
  12. Playwright Skill 让 AI 编码代理运行浏览器测试并自动保存截图和日志7.0
019.0

阿里巴巴通义千问发布 Qwen3.8-27B 和 Qwen3.8-2.4T-A95B 开放权重

阿里巴巴通义千问发布了 Qwen3.8-27B(一个拥有 270 亿参数的原生多模态稠密模型)和 Qwen3.8-2.4T-A95B(Max 级别模型)的开放权重。27B 模型整体性能超越 Qwen3.7-Plus,支持 262K 原生上下文(可通过 YaRN 扩展至 100 万 token),并采用 Apache 2.0 许可证。两个模型均可在 Hugging Face 和 ModelScope 下载。 此次发布将前沿能力带入紧凑、可本地部署的模型,使开发者无需大规模 GPU 集群即可使用先进 AI。Apache 2.0 许可证以及主流推理框架(vLLM、SGLang、Ollama、LM Studio、Unsloth)的 Day-0 支持降低了商业和研究使用的门槛。这标志着高效开放权重模型可在消费级硬件和边缘设备上运行的趋势。 Qwen3.8-27B 本地运行约需 17GB 内存,使用 SGLang 配合 NVFP4 量化在单张 RTX 5090 上可实现 206 token/s 的解码速度。它可在单张 Blackwell GPU 上以 BF16、FP8 或 NVFP4 精度运行,在 100 万上下文下 GB300 可容纳约 660 万 KV token。该模型原生支持多模态,能理解图像和视频,并支持灵活的思考控制以完成智能体任务。

@Alibaba_Qwen原推文1 张图片We promised open weights for Qwen3.8. Now, time to meet them! 🎉 ⚡ Qwen3.8-27B: - A native multimodal dense model. With just 27B parameters, it outperforms Qwen3.7-Plus overall and shines in real-world coding & office workflows. - 262K native context, easily extendable to 1M tokens via YaRN. - Built for builders. Highly efficient, high-quality, and licensed under Apache 2.0. 🚀 The open weights for Qwen3.8-2.4T-A95B (Max-level) have also been released recently. Whether you're shipping lightweight applications with Qwen3.8-27B locally or building agents with Qwen3.8-2.4T-A95B, they're yours now! Download, deploy, and build something we haven't imagined yet. 👀👇 - Hugging Face: https://huggingface.co/collections/Qwen/qwen38 - ModelScope: https://www.modelscope.cn/collections/Qwen/Qwen38 --- From twitter --- Laptop-size model, frontier-size leap. 🏃‍♀️Qwen3.8-27B is live on LM Studio. Try it! @lmstudio > 引用 @lmstudio: Qwen3.8-27B is here! 🚀 > > It's a leap in capabilities for a laptop size model. > > Requires ~17GB to run locally. > > Model page: https://lmstudio.ai/qwen/qwen3.8-27b --- From twitter --- Qwen3.8-27B is now part of your everyday life — from smartphones to vehicles. ⚡️🚀Appreciate your work! @MediaTek > 引用 @MediaTek: Congratulations to the @Alibaba_Qwen on the launch of Qwen 3.8! > > MediaTek continues our long-standing collaboration with Qwen, with Day-0 support now available on Dimensity Auto Cockpit C-X1 and the latest Dimensity flagship mobile SoC, bringing smarter and more intuitive on-device AI experiences to more people, from smartphones to vehicles, and helping make agentic AI a seamless part of everyday life. --- From twitter --- Qwen3.8-27B is available on @ollama !Any harness you run, the model always delivers. Let's code together, and show us what you build.😎 > 引用 @ollama: Qwen 3.8 27B is now available on Ollama. > > It's one of the best open models at this size, and made for agentic tasks and professional work. > > Try it directly with the apps & harnesses you use: > > Claude Code: > ollama launch claude --model qwen3.8 > > OpenCode: > ollama launch opencode --model qwen3.8 > > Hermes Agent: > ollama launch hermes --model qwen3.8 > > Pi: > ollama launch pi --model qwen3.8 > > We have also optimizations for Apple Silicon! > > Try it with the model name: qwen3.8:27b-mlx --- From twitter --- 27B on 17GB RAM. Are you ready to create something incredible? 😎 Thanks for highlighting it! @UnslothAI > 引用 @UnslothAI: Qwen3.8-27B can now be run locally! ✨ > > Run on 17GB RAM via Unsloth Dynamic GGUFs. > > Qwen3.8-27B is by far the strongest model for its size. We also uploaded NVFP4 quants. > > GGUF: https://huggingface.co/unsloth/Qwen3.8-27B-GGUF > Guide: https://unsloth.ai/docs/models/qwen3.8 --- From twitter --- Yes, we are back👑, with 206 tok/s on a single RTX 5090! Amazing Day-0 work from the SGLang team. Give it a try~@sgl_project > 引用 @sgl_project: The king of small models is back! Qwen3.8-27B from @Alibaba_Qwen is open source, and Day-0 support is live in SGLang: > - 206.1 tok/s decode on a single RTX 5090, with our NVFP4 plus DSpark > - 38.28 tok/s decode on DGX Spark > Qwen3.8-27B raises the bar again for what a small model can do on agentic planning and long-horizon tasks. > > Long live the (small model) king! Run it locally with SGLang 👇 --- From twitter --- One GPU, 1M context, Day-0 ready. Big props to the vLLM team for the seamless integration!👍 Try Qwen3.8-27B on vLLM: @vllm_project https://recipes.vllm.ai/Qwen/Qwen3.8-27B > 引用 @vllm_project: 🎉 Qwen3.8-27B is here from @Alibaba_Qwen, and the whole thing fits on a single GPU. Same hybrid backbone as the 2.4T flagship, dense instead of MoE. Day-0 support in vLLM. 🚀 > > What is in it for serving ✨ > > - Fits one Blackwell GPU in every precision. Qwen ships BF16 and FP8, the NVFP4 build from @inferact > - 262K native context, stretching to 1M. At that length one GB300 still has room for roughly 6.6M KV tokens. Six full-length sequences in flight, on one GPU > - An MTP draft head rides inside the checkpoint, so speculative decoding needs no separate speculator repo. Acceptance on short prompts measured 92.2% in BF16 and 84.8% in FP8 > > Verified end-to-end on @NVIDIA GB300: BF16 and FP8 at TP4, NVFP4 at TP1, tool calls working, correct generations at 1M context. Two prerequisites, vLLM nightly and transformers 5.8.0+. > > 🔗 http://recipes.vllm.ai/Qwen/Qwen3.8-27B --- From twitter --- Less than 2 hours to say hi👋. It's almost time! See you soon: 👀 https://huggingface.co/Qwen/Qwen3.8-27B --- From twitter --- Ready right now! 😎Incredible to see Qwen3.8-27B fly on RTX Spark. Download, deploy and create somthing new. @NVIDIARTXSpark > 引用 @NVIDIARTXSpark: Ready to run Qwen3.8 locally? 👀 > > Qwen3.8-27B packs powerful AI into an open model developers can download, serve locally, and build with on their own terms. --- From twitter --- 🎁A gift for developers: Qwen3.8-27B, running locally on AMD from day zero. Appreciate the work from the AMD team! @AMD > 引用 @AMD: Qwen3.8 27B brings a new state-of-the-art dense model for local AI development. > ⚡ Run it on AMD Ryzen™ AI Max+ processors or single Radeon™ AI PRO R9700 card > ⚡ Experience it with @LMStudio > ⚡ Turn it into an app with @lemonade_server > > Start building with AMD Day 0 support for @Alibaba_Qwen now: https://www.amd.com/en/blogs/2026/run-qwen-3-8-27b-on-amd-ryzen-ai-max-and-radeon-graphics-cards-day-0.html原推文媒体预览展开原推文收起原推文

@Alibaba_Qwen

We promised open weights for Qwen3.8. Now, time to meet them! 🎉 ⚡ Qwen3.8-27B: - A native multimodal dense model. With just 27B parameters, it outperforms Qwen3.7-Plus overall and shines in real-world coding & office workflows. - 262K native context, easily extendable to 1M tokens via YaRN. - Built for builders. Highly efficient, high-quality, and licensed under Apache 2.0. 🚀 The open weights for Qwen3.8-2.4T-A95B (Max-level) have also been released recently. Whether you're shipping lightweight applications with Qwen3.8-27B locally or building agents with Qwen3.8-2.4T-A95B, they're yours now! Download, deploy, and build something we haven't imagined yet. 👀👇 - Hugging Face: https://huggingface.co/collections/Qwen/qwen38 - ModelScope: https://www.modelscope.cn/collections/Qwen/Qwen38 --- From twitter --- Laptop-size model, frontier-size leap. 🏃‍♀️Qwen3.8-27B is live on LM Studio. Try it! @lmstudio > 引用 @lmstudio: Qwen3.8-27B is here! 🚀 > > It's a leap in capabilities for a laptop size model. > > Requires ~17GB to run locally. > > Model page: https://lmstudio.ai/qwen/qwen3.8-27b --- From twitter --- Qwen3.8-27B is now part of your everyday life — from smartphones to vehicles. ⚡️🚀Appreciate your work! @MediaTek > 引用 @MediaTek: Congratulations to the @Alibaba_Qwen on the launch of Qwen 3.8! > > MediaTek continues our long-standing collaboration with Qwen, with Day-0 support now available on Dimensity Auto Cockpit C-X1 and the latest Dimensity flagship mobile SoC, bringing smarter and more intuitive on-device AI experiences to more people, from smartphones to vehicles, and helping make agentic AI a seamless part of everyday life. --- From twitter --- Qwen3.8-27B is available on @ollama !Any harness you run, the model always delivers. Let's code together, and show us what you build.😎 > 引用 @ollama: Qwen 3.8 27B is now available on Ollama. > > It's one of the best open models at this size, and made for agentic tasks and professional work. > > Try it directly with the apps & harnesses you use: > > Claude Code: > ollama launch claude --model qwen3.8 > > OpenCode: > ollama launch opencode --model qwen3.8 > > Hermes Agent: > ollama launch hermes --model qwen3.8 > > Pi: > ollama launch pi --model qwen3.8 > > We have also optimizations for Apple Silicon! > > Try it with the model name: qwen3.8:27b-mlx --- From twitter --- 27B on 17GB RAM. Are you ready to create something incredible? 😎 Thanks for highlighting it! @UnslothAI > 引用 @UnslothAI: Qwen3.8-27B can now be run locally! ✨ > > Run on 17GB RAM via Unsloth Dynamic GGUFs. > > Qwen3.8-27B is by far the strongest model for its size. We also uploaded NVFP4 quants. > > GGUF: https://huggingface.co/unsloth/Qwen3.8-27B-GGUF > Guide: https://unsloth.ai/docs/models/qwen3.8 --- From twitter --- Yes, we are back👑, with 206 tok/s on a single RTX 5090! Amazing Day-0 work from the SGLang team. Give it a try~@sgl_project > 引用 @sgl_project: The king of small models is back! Qwen3.8-27B from @Alibaba_Qwen is open source, and Day-0 support is live in SGLang: > - 206.1 tok/s decode on a single RTX 5090, with our NVFP4 plus DSpark > - 38.28 tok/s decode on DGX Spark > Qwen3.8-27B raises the bar again for what a small model can do on agentic planning and long-horizon tasks. > > Long live the (small model) king! Run it locally with SGLang 👇 --- From twitter --- One GPU, 1M context, Day-0 ready. Big props to the vLLM team for the seamless integration!👍 Try Qwen3.8-27B on vLLM: @vllm_project https://recipes.vllm.ai/Qwen/Qwen3.8-27B > 引用 @vllm_project: 🎉 Qwen3.8-27B is here from @Alibaba_Qwen, and the whole thing fits on a single GPU. Same hybrid backbone as the 2.4T flagship, dense instead of MoE. Day-0 support in vLLM. 🚀 > > What is in it for serving ✨ > > - Fits one Blackwell GPU in every precision. Qwen ships BF16 and FP8, the NVFP4 build from @inferact > - 262K native context, stretching to 1M. At that length one GB300 still has room for roughly 6.6M KV tokens. Six full-length sequences in flight, on one GPU > - An MTP draft head rides inside the checkpoint, so speculative decoding needs no separate speculator repo. Acceptance on short prompts measured 92.2% in BF16 and 84.8% in FP8 > > Verified end-to-end on @NVIDIA GB300: BF16 and FP8 at TP4, NVFP4 at TP1, tool calls working, correct generations at 1M context. Two prerequisites, vLLM nightly and transformers 5.8.0+. > > 🔗 http://recipes.vllm.ai/Qwen/Qwen3.8-27B --- From twitter --- Less than 2 hours to say hi👋. It's almost time! See you soon: 👀 https://huggingface.co/Qwen/Qwen3.8-27B --- From twitter --- Ready right now! 😎Incredible to see Qwen3.8-27B fly on RTX Spark. Download, deploy and create somthing new. @NVIDIARTXSpark > 引用 @NVIDIARTXSpark: Ready to run Qwen3.8 locally? 👀 > > Qwen3.8-27B packs powerful AI into an open model developers can download, serve locally, and build with on their own terms. --- From twitter --- 🎁A gift for developers: Qwen3.8-27B, running locally on AMD from day zero. Appreciate the work from the AMD team! @AMD > 引用 @AMD: Qwen3.8 27B brings a new state-of-the-art dense model for local AI development. > ⚡ Run it on AMD Ryzen™ AI Max+ processors or single Radeon™ AI PRO R9700 card > ⚡ Experience it with @LMStudio > ⚡ Turn it into an app with @lemonade_server > > Start building with AMD Day 0 support for @Alibaba_Qwen now: https://www.amd.com/en/blogs/2026/run-qwen-3-8-27b-on-amd-ryzen-ai-max-and-radeon-graphics-cards-day-0.html

背景
Qwen 是阿里云开发的大语言模型系列。开放权重意味着训练好的模型参数可公开下载和使用,不同于封闭 API。Apache 2.0 是一种允许商业使用的宽松开源许可证。YaRN 是一种无需完全重新训练即可高效扩展模型上下文窗口的技术。稠密模型对每个 token 激活所有参数,而混合专家(MoE)模型只激活一部分参数,以内存换取计算。
社区讨论
社区反响非常积极,MediaTek、Ollama、Unsloth、SGLang 和 vLLM 等合作伙伴均宣布 Day-0 支持。许多人强调该模型在其规模下性能强劲,适合智能体任务和本地部署。也有人注意到其在消费级 GPU 上令人印象深刻的 token 生成速度以及在笔记本电脑上运行的便利性。

8月14日 15:02在 X 打开#AI #Open Source #Qwen #Multimodal #Model Release

029.0

Qwen3.8-Max:2.4万亿参数MoE模型在多平台Day 0上线

阿里巴巴通义千问发布了Qwen3.8-Max,这是一个拥有2.4万亿总参数、950亿活跃参数的混合专家(MoE)模型,上下文窗口达100万token。该模型在发布当天(Day 0)即通过Together AI、DigitalOcean、Modal、Nebius Token Factory和Fireworks等合作伙伴上线。它被称为首个达到“Max”规模的开源权重模型,具备强大的编程和智能体能力。 此次发布标志着开源权重前沿模型的重要进展,在发布当天就将2.4万亿参数的MoE模型带给广泛的开发者生态。100万token的上下文窗口和对智能体的侧重使其适用于长周期编程和自主智能体工作流,有望挑战闭源模型。多平台即时可用降低了企业和开发者采用最先进模型的门槛。 Qwen3.8-Max采用混合专家架构:总参数2.4万亿,但每个token仅激活950亿参数,降低了计算成本,但内存需求仍与总参数量相关。它支持100万token的上下文,能够处理上百页文档或大型代码库。Modal使用针对工具调用密集数据训练的自定义DFlash推测器提供服务;DigitalOcean采用NVIDIA HGX B300 GPU并按使用量计费。该模型为开源权重,但具体许可条款在提供的资料中未详细说明。

@Alibaba_Qwen引用推文1 张图片Qwen3.8-Max is live on Together AI. Together AI is with us as a Day 0 launch partner, and we couldn’t ask for a better name to share Day 0 with. 2.4T parameters, 95B active, 1M context — all together now. 🧑‍🤝‍🧑@togethercompute原推文媒体预览展开原推文收起原推文

@Alibaba_Qwen

Qwen3.8-Max is live on Together AI. Together AI is with us as a Day 0 launch partner, and we couldn’t ask for a better name to share Day 0 with. 2.4T parameters, 95B active, 1M context — all together now. 🧑‍🤝‍🧑@togethercompute

@togethercompute

Qwen3.8-2.4T-A95B is live on Together AI. The Qwen team’s new open-weight MoE packs 2.4T total parameters with 95B active, a 1M context window, and strong coding & agent capabilities. Start building: https://www.together.ai/models/qwen3-8-max

背景
混合专家(MoE)模型使用多个专门的子网络(“专家”)和一个路由器,每个token只激活少数专家,从而在总参数量很大的情况下降低推理计算量。Qwen是阿里巴巴的大语言模型系列,“Max”代表其最强版本。“Day 0发布合作伙伴”意味着模型从发布那一刻起就在该平台上可用,通常还带有优化服务。开源权重模型允许开发者下载并自行托管权重,但商业使用可能受许可证约束。

8月14日 15:24在 X 打开#Qwen #MoE #AI model release #Together AI #open-weight

038.0

Anthropic 将为 Claude 实施文本水印以符合欧盟 AI 法案

Anthropic 发布了一份常见问题解答,宣布将在 Claude 中实施文本水印,以符合欧盟 AI 法案。该公司表示,其他主要模型开发商已签署相同的实践准则,也将实施水印。Anthropic 澄清,水印方法不会影响输出质量、成本或可读性。 此举标志着 AI 生成内容在监管合规方面迈出了重要一步,因为欧盟 AI 法案要求通用 AI 模型采取透明度措施。它可能为其他 AI 提供商树立先例,并影响全球检测 AI 生成文本的标准。Claude 的用户和开发者需要了解水印如何影响其工作流程和内容真实性。 Anthropic 表示,水印方法不会增加额外 token、隐藏字符或任何读者可察觉的差异。水印无法追溯到特定个人、组织或聊天。该公司强调,对 Claude 输出的质量或内容没有实际影响。常见问题解答可在 https://www.anthropic.com/news/claude-text-watermark 查看。

@AnthropicAI原推文We’ve written an FAQ to answer some of the questions we've received about watermarking. In summary: • We’re implementing watermarking to comply with the EU AI Act. Other major model developers have signed the same Code of Practice and will also be implementing watermarking; • Our watermarking method doesn’t have any practical impact on the quality or content of Claude’s outputs; • The difference between watermarked and un-watermarked text will not be distinguishable to readers; • Nothing is added to the text and there are no hidden characters; • Watermarking doesn’t require extra tokens, and will not be more expensive; • Watermarks can’t be traced to a specific person, organization, or chat. Read more: https://www.anthropic.com/news/claude-text-watermark展开原推文收起原推文

@AnthropicAI

We’ve written an FAQ to answer some of the questions we've received about watermarking. In summary: • We’re implementing watermarking to comply with the EU AI Act. Other major model developers have signed the same Code of Practice and will also be implementing watermarking; • Our watermarking method doesn’t have any practical impact on the quality or content of Claude’s outputs; • The difference between watermarked and un-watermarked text will not be distinguishable to readers; • Nothing is added to the text and there are no hidden characters; • Watermarking doesn’t require extra tokens, and will not be more expensive; • Watermarks can’t be traced to a specific person, organization, or chat. Read more: https://www.anthropic.com/news/claude-text-watermark

How Claude's text watermarking worksanthropic.com · 直连原文
背景
文本水印是一种在文本内容中嵌入隐藏信息以验证其来源或真实性的技术。随着生成式 AI 的兴起,对 AI 生成文本进行水印已成为检测合成内容的关键方法。欧盟 AI 法案分阶段生效,其中包括对通用 AI 模型的透明度义务,而实践准则为合规提供了框架。Anthropic 的公告与这些监管发展保持一致。

8月14日 19:16在 X 打开#AI #Anthropic #watermarking #EU AI Act #Claude

048.0

Anthropic 根据负责任扩展政策发布第二份风险报告

Anthropic 根据其负责任扩展政策发布了第二份风险报告,详细说明了系统风险和准备情况。该报告评估了多个类别的灾难性风险,并指出自动化 AI 研发正成为一个日益令人担忧的问题。报告还提及了一个未发布的内部模型,表明能力评估仍在持续进行。 这份报告是领先 AI 实验室在透明度方面的重要里程碑,让公众和政策制定者能够了解前沿 AI 的风险。它强调了 AI 研究加速发展的趋势,以及自动化研发可能成为重大风险的潜力,这可能会影响行业安全实践和监管讨论。 2026 年 8 月的报告指出,当前模型的整体灾难性风险较低,但自动化 AI 研发是一个日益增长的担忧。报告根据能力和防护措施两方面评估风险,并提及了一个未发布的内部模型。该报告是 Anthropic 负责任扩展政策的一部分,该政策要求定期公开风险报告。

@AnthropicAI原推文As part of our Responsible Scaling Policy, we publish regular Risk Reports. These share detailed information on the risks of our systems and how prepared we are to address them. Our second Risk Report is now available: https://www.anthropic.com/aug-2026-risk-report展开原推文收起原推文

@AnthropicAI

As part of our Responsible Scaling Policy, we publish regular Risk Reports. These share detailed information on the risks of our systems and how prepared we are to address them. Our second Risk Report is now available: https://www.anthropic.com/aug-2026-risk-report

背景
Anthropic 的负责任扩展政策(RSP)是一套技术和组织协议框架,旨在管理日益强大的 AI 系统带来的风险。该政策要求定期发布风险报告,分享有关系统风险和准备情况的详细信息。该政策已经历多个版本,最新版本扩大了负责任扩展官的职责,并在评估中增加了外部专家意见。

8月14日 18:00在 X 打开#AI safety #Anthropic #Responsible Scaling Policy #risk assessment #AI policy

058.0

Qwen3.8 借助 TokenSpeed 大规模运行,性能比 TP16 提升 30% 以上

阿里巴巴通义千问宣布其 2.4T 参数的 Qwen3.8 模型现已借助 LightSeek 的开源推理引擎 TokenSpeed 实现大规模运行。TokenSpeed 为 Qwen3.8 在多节点 NVIDIA Blackwell 系统上提供了 Day-0 支持,性能比 TP16 提升 30% 以上。优化包括跨节点的 DP/EP 扩展以及采用单 CUDA 图优化的 DSpark 推测解码。 这表明开源推理引擎能够为超大规模模型提供生产级性能,降低服务万亿参数模型的成本和复杂度。相比 TP16 提升 30% 以上的速度直接降低了延迟并提高了对延迟敏感的智能体工作负载的吞吐量。这也表明 Qwen 模型的生态系统支持日益增强,LLM 推理优化领域的竞争也在加剧。 TokenSpeed 专为智能体工作负载设计,声称具备 TensorRT-LLM 级别的性能和 vLLM 级别的易用性。性能提升来自跨节点优化数据并行(DP)和专家并行(EP),而非仅依赖张量并行(TP16)。集成了源自 DeepSeek 的 DSpark 推测解码,并通过单 CUDA 图降低开销。该公告引用了 lightseek.org 上的详细博客文章。

@Alibaba_Qwen引用推文4 张图片Excited to see Qwen3.8 running at scale with TokenSpeed! 🚀 Light on latency, big on speed. Kudos to LightSeek for the fantastic Day-0 support! @lightseekorg原推文媒体预览+3展开原推文收起原推文

@Alibaba_Qwen

Excited to see Qwen3.8 running at scale with TokenSpeed! 🚀 Light on latency, big on speed. Kudos to LightSeek for the fantastic Day-0 support! @lightseekorg

@lightseekorg

We’re proud to be the Day 0 open-source inference engine partner for @Alibaba_Qwen 3.8. To serve this 2.4T-parameter model across multi-node @NVIDIAAI Blackwell inference, we optimized DP/EP scaling across nodes, delivering 30%+ faster performance than TP16, plus DSpark speculative decoding with single CUDA graph optimization👇 https://lightseek.org/blog/tokenspeed-qwen3-8.html

背景
Qwen3.8 是阿里巴巴的 2.4 万亿参数模型,鉴于提到了专家并行,很可能是混合专家(MoE)架构。张量并行(TP)将模型权重拆分到多个 GPU 上,但会带来高昂的通信开销;TP16 表示使用 16 个 GPU 进行张量并行。数据并行(DP)在多个 GPU 上复制模型以服务更多请求,而专家并行(EP)将 MoE 模型的不同专家分布到不同 GPU 上。推测解码使用小型草稿模型提出 token,然后由大模型并行验证,从而加速生成。TokenSpeed 是 LightSeek 的开源推理引擎,旨在兼顾高性能和易用性。

8月14日 16:58在 X 打开#Qwen #LLM inference #open-source #performance optimization #Blackwell

068.0

阿里巴巴通义千问发布开源模型 Qwen3.8-2.4T-A95B,多家平台提供 Day-0 支持

阿里巴巴通义千问开源了 Qwen3.8-2.4T-A95B,这是一个拥有 2.4 万亿总参数、95B 激活参数的稀疏混合专家模型。该模型已在 SiliconFlow 和 DeepInfra 上提供 Day-0 支持,定价为每百万输入 token 2 美元、每百万输出 token 6 美元。它具备 1M token 上下文窗口,专为自主编码、深度研究和端到端智能体执行而设计。 此次发布是开源 AI 的重要进展,将前沿规模的模型能力带给社区。凭借 2.4T 参数和 1M 上下文窗口,它支持以前仅限于闭源模型的复杂智能体工作流。多家推理服务商提供 Day-0 支持,降低了开发者构建和部署高级 AI 应用的门槛。 Qwen3.8-2.4T-A95B 采用稀疏 MoE 架构,包含 512 个专家、混合注意力机制和 1M token 上下文窗口。它是 Qwen3.8 Max 的开源权重版本。SiliconFlow 上的定价为每百万输入 token 2 美元、输出 6 美元、缓存输入 0.25 美元;DeepInfra 的缓存输入价格略低,为 0.20 美元。该模型最大输出 token 数为 52,429。

@Alibaba_Qwen引用推文1 张图片From idea to implementation in one go.🏃‍♀️ Max-level intelligence, served fresh on Day 0. Qwen3.8-2.4T-A95B is live on SiliconFlow. Thanks! @SiliconFlowAI原推文媒体预览展开原推文收起原推文

@Alibaba_Qwen

From idea to implementation in one go.🏃‍♀️ Max-level intelligence, served fresh on Day 0. Qwen3.8-2.4T-A95B is live on SiliconFlow. Thanks! @SiliconFlowAI

@SiliconFlowAI

🚀 Day-0 Support! @Alibaba_Qwen has open-sourced Qwen3.8-2.4T-A95B — and it’s now live on SiliconFlow. ⚡ With 2.4T parameters and 95B active, Qwen3.8 is built to take a goal and come back with finished work — coding, researching, planning, and executing along the way. 💸 Per 1M tokens: • Input: $2.00 • Output: $6.00 • Cached input: $0.25 💻 Autonomous coding. 🔬 Deep research. 🤖 End-to-end agent execution. Built for serious workloads — from idea to implementation in one go. Time to cook → http://siliconflow.com 🍳

背景
通义千问(Qwen)是阿里巴巴的大语言模型系列,在编码、推理和多语言任务上表现优异。稀疏混合专家(MoE)模型对每个输入只激活部分参数,在保持高容量的同时降低计算成本。Day-0 支持意味着推理服务商在模型发布当天就提供托管服务,方便开发者立即试用。SiliconFlow 和 DeepInfra 是托管开源模型并提供 API 访问的云平台。

8月14日 16:50在 X 打开#AI #Open Source #Qwen #Model Release #SiliconFlow

078.0

Tailscale 发现隐藏16年的 SQLite WAL 重置缺陷

Tailscale 在调查了19起数据损坏事件后,发现了一个至少存在了16年的 SQLite WAL 重置缺陷。该缺陷是一个检查点竞态条件:页面可能被报告为已复制,但已提交的写入仍然缺失。他们使用追踪 VFS 捕获了该问题,并在 Tailscale 的博客文章中详细介绍了发现过程。 SQLite 是全球部署最广泛的数据库之一,嵌入在无数应用和设备中。其 WAL 模式中长期隐藏的损坏缺陷可能悄然影响许多用户的数据完整性。这一发现凸显了对基础软件进行严格测试和可观测性的重要性。 该缺陷仅在 WAL 模式激活、同一文件上有多个数据库连接、且写事务与 WAL 重置发生冲突时才会触发。Tailscale 修补了他们的 SQLite 驱动,在这些操作重叠时记录警告;两个月后警报触发,证明生产环境中确实存在这些条件。SQLite 团队已发布关于 WAL 重置问题的说明。调查过程中还发现了第二个过期的表达式索引缺陷。

@bibryam原推文1 张图片SQLite's WAL-reset bug hid for at least 16 years. After 19 corruption incidents, Tailscale found a rare checkpoint race: pages could be reported copied while committed writes were still missing. A tracing VFS caught it... https://tailscale.com/blog/sqlite-wal-reset-bug原推文媒体预览展开原推文收起原推文

@bibryam

SQLite's WAL-reset bug hid for at least 16 years. After 19 corruption incidents, Tailscale found a rare checkpoint race: pages could be reported copied while committed writes were still missing. A tracing VFS caught it... https://tailscale.com/blog/sqlite-wal-reset-bug

背景
SQLite 是一个自包含、无服务器、零配置的 SQL 数据库引擎,广泛用于嵌入式系统、移动应用和桌面软件。WAL(预写式日志)是一种日志模式,通过允许读写操作互不阻塞来提高并发性。WAL 重置是一种截断或重置 WAL 文件的操作,通常发生在检查点期间。追踪 VFS(虚拟文件系统)是一个自定义层,用于拦截文件系统调用以记录或分析数据库 I/O 行为。

8月14日 23:27在 X 打开#SQLite #database #bug #Tailscale #WAL

087.0

Pi 作者 Armin Ronacher 称赞 DeepSeek Harness 激发设计反思

Pi 项目作者、开发者社区中备受尊敬的人物 Armin Ronacher 在 X(原 Twitter)上公开称赞了 DeepSeek Harness。他表示,尽管该 Harness 并不完美,但这是 AI 智能体领域第一个让他受到启发、重新审视自己部分设计选择的新进展。他特别强调了开源在促成这种启发方面的价值。 这位知名工程师的认可表明,DeepSeek Harness 引入了真正值得研究的新颖理念,而不仅仅是渐进式改进。这可能会影响其他开发者探索或采用其基于插件的架构,从而可能塑造智能体 Harness 设计的方向。对开源的强调也强化了透明、社区驱动的 AI 基础设施日益增长的重要性。 DeepSeek Harness 是一个开源的智能体 Harness,其中每个智能体能力都以可替换或可重组的插件形式实现。它目前处于开发者预览阶段,源代码已在 GitHub 上公开。Ronacher 的评论侧重于概念上的启发,而非具体的技术基准或性能指标。

@dotey引用推文1 张图片Pi 作者对 DeekSeek Harness 评价原推文媒体预览展开原推文收起原推文

@dotey

Pi 作者对 DeekSeek Harness 评价

@mitsuhiko

I don't think the DeepSeek Harness is perfect but this is for sure the first time I have been looking at something new in the space and felt quite inspired to revisit some of our choices. I love that part about Open Source a lot!

背景
DeepSeek Harness 是一个用于构建和运行 AI 智能体的开源框架,强调基于插件的架构,使能力可以替换或重组。Armin Ronacher 是一位杰出的软件开发者,以创建 Flask 并参与 Jinja2、Pygments 等项目而闻名;他也与 Pi 项目有关联。在 AI 领域,“Harness”一词指的是管理智能体执行、工具集成和编排的脚手架或运行时环境。AI 领域的开源允许开发者检查、修改和学习代码,通过共享知识促进创新。

8月14日 15:11在 X 打开#DeepSeek #Open Source #AI #Harness #Community

097.0

Qwen在Hugging Face开放模型报告中领跑本地推理

阿里巴巴Qwen宣布,根据Hugging Face 2026年夏季《开放模型现状》报告,其在本地推理领域处于领先地位。报告指出,尽管前沿模型越来越大,但小模型在实际应用中仍占主导地位,Qwen在本地推理中排名第一,其次是Gemma。Qwen对此成就表示自豪,并感谢Hugging Face的工作。 这一认可凸显了小型高效模型在实际AI应用中的重要性,尤其是在大型模型资源消耗日益增加的背景下。它标志着开源AI生态系统中本地推理和边缘部署正获得更多关注,有利于寻求成本效益和隐私保护方案的开发者和企业。Qwen的领先地位可能影响采用趋势,并激励紧凑模型架构的进一步创新。 Hugging Face 2026年夏季《开放模型现状》报告指出,Qwen在本地推理中领先,Gemma位居第二。报告还提到AI智能体正在成为Hugging Face Hub上的重要力量。虽然推文未提供具体下载量或基准数据,但相关报告显示Hugging Face上下载量最高的200个模型占总下载量的近一半,凸显了使用集中度。报告对本地推理的关注表明,小模型正越来越多地部署在用户设备上,而不仅仅是在云端。

@Alibaba_Qwen引用推文Small models, big real-world impact. Proud to see Qwen leading local inference in the State of Open Models. Enjoy the sunshine from Qwen.☀️ 😎 Appreciate your work! @huggingface展开原推文收起原推文

@Alibaba_Qwen

Small models, big real-world impact. Proud to see Qwen leading local inference in the State of Open Models. Enjoy the sunshine from Qwen.☀️ 😎 Appreciate your work! @huggingface

@huggingface

The State of Open Models, Summer 2026 ☀️ frontier models are getting larger, but small models still dominate real-world usage. Qwen leads local inference, followed by Gemma. AI agents are becoming a major force on the Hub Full picture on the blog 🤗 https://huggingface.co/blog/state-of-open-models-summer-2026

背景
本地推理指的是直接在用户设备(如笔记本电脑、智能手机或边缘服务器)上运行AI模型,而不是将数据发送到远程云服务器。这种方法具有低延迟、增强隐私和降低运营成本等优势。Hugging Face是托管和分享开源机器学习模型的领先平台,其《开放模型现状》报告定期分析模型使用和开发的趋势。Qwen是阿里巴巴开发的一系列开源权重语言模型,以在各种任务和规模上的强劲表现而闻名。Gemma是谷歌推出的一系列轻量级开放模型,也常用于本地部署。

8月14日 17:05在 X 打开#open models #local inference #Qwen #Hugging Face #AI trends

107.0

oil-motion:AI驱动的滚动触发动画技能

oil-motion 是一个开源的 Agent Skill,能够根据自然语言描述和素材生成滚动触发的动画。它利用 AI 生成连续动作画面,并将页面滚动、鼠标、拖拽等操作对应到动画的每一帧。该工具先确认关键画面,中间过渡交给 AI 视频生成,最后按实际显示尺寸压缩资源。 该工具解决了网页开发者的一个常见痛点:创建滚动驱动动画通常需要手动逐帧制作或使用复杂库。通过利用 AI,oil-motion 降低了交互式叙事和产品展示页面的门槛,有可能加速原型开发并带来更具吸引力的用户体验。 GitHub 仓库约有 1.6K 星标,被描述为与代理无关,意味着它可以与各种 AI 编码代理配合使用。它支持桌面端和移动端适配,示例包括产品展开拆解、角色跟鼠标转头、拖进度条控制动画。该技能可在 skills.sh 和 SourceForge 等平台获取。

@GitHub_Daily原推文1 张图片产品官网上那种跟着滚动逐帧变化的动画,想自己做一个,不知道从哪下手。 oil-motion 是一个 Agent 交互动画 Skill,只要说清楚效果、给上素材就行。 AI 生成连续动作画面,再把页面滚动、鼠标、拖拽这些操作对应到动画的每一帧。 GitHub:http://github.com/oil-oil/oil-motion 先确认关键画面,中间过渡交给 AI 视频生成,最后按实际显示尺寸压缩资源。 产品展开拆解、角色跟鼠标转头、拖进度条控制动画,桌面端和移动端分别适配。 做产品介绍页或角色互动这类交互动画的朋友可以试试。原推文媒体预览展开原推文收起原推文

@GitHub_Daily

产品官网上那种跟着滚动逐帧变化的动画,想自己做一个,不知道从哪下手。 oil-motion 是一个 Agent 交互动画 Skill,只要说清楚效果、给上素材就行。 AI 生成连续动作画面,再把页面滚动、鼠标、拖拽这些操作对应到动画的每一帧。 GitHub:http://github.com/oil-oil/oil-motion 先确认关键画面,中间过渡交给 AI 视频生成,最后按实际显示尺寸压缩资源。 产品展开拆解、角色跟鼠标转头、拖进度条控制动画,桌面端和移动端分别适配。 做产品介绍页或角色互动这类交互动画的朋友可以试试。

背景
Agent Skills 是最近为 AI 编码助手扩展专业能力的一种标准,类似于 MCP(模型上下文协议)。滚动触发动画是一种网页设计技术,动画进度与用户的滚动位置相关联,常用于叙事和产品展示。传统实现需要 Motion for React 等库或手动编写 JavaScript。

8月15日 07:30在 X 打开#animation #AI #web development #open source #interactive design

117.0

SkillUI 从网站提取完整设计系统供 Claude Code 使用

SkillUI 是一个新工具,只需输入网站网址,即可自动提取完整的设计系统,包括配色、字体、间距、动画和组件截图。它将信息打包成 Claude Code 可直接读取的 .skill 文件。该工具还提供 Ultra 模式,使用 Playwright 进行完整视觉提取,捕捉滚动截图、交互状态和动画检测。 该工具解决了开发者和设计师在复制网站视觉风格时需手动检查 CSS 和资源的常见痛点。通过自动化设计系统提取并与 Claude Code 集成,它大幅减少了重现或适配设计所需的时间和精力。它还降低了在前端开发工作流中使用 AI 编码助手的门槛。 SkillUI 可以分析网址、本地项目或 GitHub 仓库,并生成文档和令牌文件。Ultra 模式使用 Playwright 进行高级视觉提取,包括滚动路径和交互状态。该工具可在 GitHub 上获取(github.com/amaancoderx/npxskillui),并可通过 npx 运行。它输出一个文件夹,可在 Claude Code 中打开以应用提取的样式。

@GitHub_Daily原推文1 张图片看到一个好看的网站想照着做,光提取配色、字体、间距就得翻半天 CSS。 SkillUI 给它一个网址,能把整套设计系统自动扒下来。 配色、字体、间距、动画、组件截图,打包成 Claude Code 能直接读的 Skill。 GitHub:http://github.com/amaancoderx/npxskillui 打开输出文件夹启动 Claude,说照这个风格做就行,设计细节已经全在里面了。 还支持 Ultra 模式,用 Playwright 做完整视觉提取,滚动截图、交互状态、动画检测都能抓到。 想快速复刻某个网站的视觉风格,这个工具挺省事。原推文媒体预览展开原推文收起原推文

@GitHub_Daily

看到一个好看的网站想照着做,光提取配色、字体、间距就得翻半天 CSS。 SkillUI 给它一个网址,能把整套设计系统自动扒下来。 配色、字体、间距、动画、组件截图,打包成 Claude Code 能直接读的 Skill。 GitHub:http://github.com/amaancoderx/npxskillui 打开输出文件夹启动 Claude,说照这个风格做就行,设计细节已经全在里面了。 还支持 Ultra 模式,用 Playwright 做完整视觉提取,滚动截图、交互状态、动画检测都能抓到。 想快速复刻某个网站的视觉风格,这个工具挺省事。

背景
设计系统是可复用组件和指南的集合,定义了产品的视觉语言,包括颜色、排版、间距和动画。Claude Code 是一个 AI 编码助手,可以通过“技能”(模块化的指令和资源包)进行扩展。Playwright 是一个浏览器自动化库,可以捕获截图、检查元素并模拟交互,适合进行彻底的视觉提取。

8月15日 04:00在 X 打开#design-system #web-development #AI-tools #Claude-Code #automation

127.0

Playwright Skill 让 AI 编码代理运行浏览器测试并自动保存截图和日志

一个名为 Playwright Skill 的新工具让 AI 编码代理能够按需编写并运行 Playwright 脚本来执行浏览器测试。测试过程中会自动保存截图和日志,且浏览器窗口默认可见,用户可以实时观察测试运行。该技能支持在 Claude Code、Cursor、Codex 和 Gemini CLI 中安装。 该工具解决了开发人员在编码和手动浏览器测试之间频繁切换的常见痛点。通过将浏览器测试直接集成到 AI 编码代理中,它可以节省大量时间并减少上下文切换,尤其适合需要测试多个页面或表单的团队。它支持多种流行的 AI 编码 CLI,因此具有广泛的适用性。 该技能可在 GitHub 上获取:http://github.com/lackeyjb/playwright-skill。它的工作方式是用户告诉代理要测试什么,代理随后编写并执行 Playwright 脚本。截图和日志会自动保存,方便事后查看。浏览器窗口默认打开,使测试过程透明可见。

@GitHub_Daily原推文1 张图片写代码时经常要打开浏览器,测试表单能不能提交、页面跑不跑得动、手机端显示对不对。 如果安装了 Playwright Skill 这个技能,可以让我们的 Agent 工具在浏览器上进行测试。 只需告诉它测什么,它当场写 Playwright 脚本跑起来。浏览器窗口默认是开的,过程看得到。 GitHub:http://github.com/lackeyjb/playwright-skill 除此之外,测试过程的相关截图和日志自动保存,跑完回来直接看结果就行。 支持 Claude Code、Cursor、Codex、Gemini CLI 安装,需要批量测页面的朋友试试。原推文媒体预览展开原推文收起原推文

@GitHub_Daily

写代码时经常要打开浏览器,测试表单能不能提交、页面跑不跑得动、手机端显示对不对。 如果安装了 Playwright Skill 这个技能,可以让我们的 Agent 工具在浏览器上进行测试。 只需告诉它测什么,它当场写 Playwright 脚本跑起来。浏览器窗口默认是开的,过程看得到。 GitHub:http://github.com/lackeyjb/playwright-skill 除此之外,测试过程的相关截图和日志自动保存,跑完回来直接看结果就行。 支持 Claude Code、Cursor、Codex、Gemini CLI 安装,需要批量测页面的朋友试试。

背景
Playwright 是一个开源的浏览器自动化框架,允许开发人员编写脚本来控制网络浏览器进行测试和抓取。像 Claude Code、Cursor、Codex 和 Gemini CLI 这样的 AI 编码代理是使用大型语言模型来辅助编写和运行代码的工具。这里的“技能”是指扩展 AI 代理能力的指令或功能包,通常通过提供专业知识或工具集成来实现。

8月15日 00:00在 X 打开#Playwright #AI agents #browser testing #developer tools #automation