7月21日2026 · 星期二

从 29 条抓取中筛选 12 条 · twitter × 9 账号 · 02:34 UTC 生成

今日信号 · 高度即评分 · 点击直达


  1. 雅可比猜想被推翻:C^3 中的显式反例10.0
  2. 大模型时代呼唤亲力亲为的技术领导力,而非传统层级管理8.0
  3. 开源权重模型GLM-5.2在Hugging Face网络事件响应中负责任地使用8.0
  4. OpenAI 分享关于长时间运行模型的安全研究8.0
  5. Kimi K3 在 Agent Arena 排名第四,有望成为最强开源权重模型8.0
  6. 阿里通义千问Qwen3.8预览版每日迭代,计划开源发布8.0
  7. PentesterFlow:开源终端AI助手,自动化渗透测试流程8.0
  8. Claude Code 新增屏幕阅读器模式,助力视障开发者7.0
  9. Fable 5 在调试 VBR MP3 时间戳问题上优于 GPT-5.6 Sol7.0
  10. Anthropic 为罕见病研究提供高达5万美元的 Claude 使用额度7.0
  11. 开发者用开源模型Qwen3.8-Max打造《我的世界》风格游戏等三个项目7.0
  12. GitHub 发布包含 923 篇高质量文档的精选机器学习库7.0
0110.0

雅可比猜想被推翻:C^3 中的显式反例

Levent Alpöge 宣布找到了雅可比猜想的一个反例,这是代数几何中的一个长期未解问题。这个从 C^3 到 C^3 的显式多项式映射具有非零常数雅可比行列式 -2,但不可逆,三个特定点映射到同一个像点证明了这一点。 这推翻了维度大于 2 时的雅可比猜想,该猜想自 1939 年以来一直悬而未决,并被列入 Smale 的 21 世纪数学问题清单。这一发现借助了 AI 语言模型,凸显了 AI 在数学研究中日益重要的作用,并可能重塑多项式自同构的研究方法。 反例映射由变量 x, y, z 的三个多项式给出,雅可比行列式为 -2。点 (0,0,-1/4)、(1,-3/2,13/2) 和 (-1,3/2,13/2) 都映射到 (-1/4,0,0),证明了非单射性。该映射是使用 Anthropic 的 Claude Fable 5 语言模型发现的,对于 N=2 的情况,猜想仍未解决。

@__alpoge__@trq212 转推hello there the jacobian conjecture is false thanx to my close friend akhil for asking about it and my other close friend fable for working during the world cup final ((1+xy)^3 z + y^2 (1+xy) (4+3xy), y + 3 x (1+xy)^2 z + 3 x y^2 (4+3xy), 2 x - 3 x^2 y - x^3 z): \C^3\to \C^3, has jacobian determinant -2, and sends (0, 0, -1/4), (1, -3/2, 13/2), and (-1, 3/2, 13/2) to (-1/4, 0, 0)展开原推文收起原推文

@trq212 转推了

@__alpoge__

hello there the jacobian conjecture is false thanx to my close friend akhil for asking about it and my other close friend fable for working during the world cup final ((1+xy)^3 z + y^2 (1+xy) (4+3xy), y + 3 x (1+xy)^2 z + 3 x y^2 (4+3xy), 2 x - 3 x^2 y - x^3 z): \C^3\to \C^3, has jacobian determinant -2, and sends (0, 0, -1/4), (1, -3/2, 13/2), and (-1, 3/2, 13/2) to (-1/4, 0, 0)

背景
雅可比猜想由 Ott-Heinrich Keller 于 1939 年首次提出,它断言从 C^n 到自身的多项式映射如果具有非零常数雅可比行列式,则必有多项式逆映射。该猜想在 n=1 时显然成立,但对于 n>1 一直未被证明,且有许多错误尝试。它是仿射代数几何中的一个核心问题,将可逆性与简单的行列式条件联系起来。

7月20日 06:08在 X 打开#mathematics #Jacobian conjecture #algebraic geometry #counterexample #breakthrough

028.0

大模型时代呼唤亲力亲为的技术领导力,而非传统层级管理

@tydsh 的一条热门推文指出,大模型时代正从根本上将领导力从“愿景-执行”的层级模式转向技术优先、亲力亲为的文化。作者以 2015 年在 Facebook 力排众议聘用吴育昕(Detectron2 作者)的个人故事,说明脚踏实地实干者的价值。该推文还以月之暗面(Moonshot AI)的清华系创始团队为例,展示这种新模式。 这一转变之所以重要,是因为商业成功如今依赖于无法用固定基准衡量的模型能力,自上而下的管理容易导致奖励黑客行为。它预示着一个更广泛的行业趋势:技术联合创始人和亲力亲为的领导者对 AI 初创公司至关重要,正如月之暗面等公司所示。这一变化将影响招聘、组织设计以及我们对科技领导力的评估方式。 吴育昕毕业于清华和 CMU,曾与何恺明合作提出 Group Normalization,并在 FAIR 创建了 Detectron2。月之暗面创始团队包括杨植麟(Transformer-XL 和 XLNet 第一作者)、张宇韬、吴育昕和周昕宇(ShuffleNet 共同作者),均拥有清华深厚的技术背景。作者强调,在大模型时代,领导者必须亲自处理数据和代码,拥抱快速反馈循环,并随时准备承认错误。

@tydsh@dotey 转推1 张图片Chinese students are often down-to-earth do-ers, a critical characteristic in LLM era. Yuxin Wu was my intern back in 2015 and I spent an hour debating with my former manager on the roof of Facebook building, arguing that he should be hired. I won the debate by staking my reputation on it. Fortunately I was right. It used to be the case that the business model is fixed and business workflow clearly decouples into vision + execution. Professional CEO/VP/Director focus on presenting the long-term vision, and junior people sit in the war room to do the grudging work to push the numbers up, and no one knows their names. Things have changed substantially. The business pattern now depends on the technical strength of the models, which cannot be measured by a pre-defined set of benchmark numbers. Too many ways and too much incentive towards reward hacking, if sitting in a big hierarchy. Technical-first now becomes critical. Sit down and get things to work. Get hands dirty, check the data and code, fast feedback loop, step out of the echo chamber, tear the pretty story apart and rewrite, ready to say "I am wrong". Let experiments tell the issues and be humble in front of AI. That's why down-to-earth doers shine now. This applies to everyone including co-founders. Anything else follows.原推文媒体预览展开原推文收起原推文

@dotey 转推了

@tydsh

Chinese students are often down-to-earth do-ers, a critical characteristic in LLM era. Yuxin Wu was my intern back in 2015 and I spent an hour debating with my former manager on the roof of Facebook building, arguing that he should be hired. I won the debate by staking my reputation on it. Fortunately I was right. It used to be the case that the business model is fixed and business workflow clearly decouples into vision + execution. Professional CEO/VP/Director focus on presenting the long-term vision, and junior people sit in the war room to do the grudging work to push the numbers up, and no one knows their names. Things have changed substantially. The business pattern now depends on the technical strength of the models, which cannot be measured by a pre-defined set of benchmark numbers. Too many ways and too much incentive towards reward hacking, if sitting in a big hierarchy. Technical-first now becomes critical. Sit down and get things to work. Get hands dirty, check the data and code, fast feedback loop, step out of the echo chamber, tear the pretty story apart and rewrite, ready to say "I am wrong". Let experiments tell the issues and be humble in front of AI. That's why down-to-earth doers shine now. This applies to everyone including co-founders. Anything else follows.

@Michaelzsguo

Moonshot AI, the company behind Kimi, has four core founders. Their backgrounds are unusually strong: - Founder and CEO Yang Zhilin studied computer science at Tsinghua before earning his PhD from Carnegie Mellon. He was the first author of Transformer-XL and XLNet, and previously worked at FAIR and Google Brain. - Co-founder and CTO Zhang Yutao earned his PhD in computer science from Tsinghua. His earlier work covered knowledge graphs and AMiner, and he previously co-founded Recurrent AI with Yang. - Co-founder Wu Yuxin studied at Tsinghua and CMU before joining FAIR. He worked with Kaiming He on Group Normalization and also created Detectron2. - Co-founder Zhou Xinyu studied computer science at Tsinghua and later joined Megvii, where he worked on turning research algorithms into production systems and co-authored ShuffleNet. They all share one root: Tsinghua University. Tsinghua is widely regarded as one of China’s top universities. In the latest U.S. News Best Global Universities ranking, it reached No. 6 worldwide. Its influence on China’s AI industry extends well beyond Moonshot. http://Z.ai, the company behind the GLM models, also grew out of Tsinghua. Its co-founder and chief scientist, Tang Jie, was once Yang Zhilin’s teacher. There is also a more personal connection. Yang and Zhou formed a rock band together at Tsinghua. Moonshot AI’s Chinese name, 月之暗面, comes from Pink Floyd’s album The Dark Side of the Moon, one of Yang’s favorites. Kimi may look like a young AI company. Behind it is a much older network of classmates, teachers, research labs, and friendships.

背景
Detectron2 是 Facebook AI Research 的下一代目标检测与分割库,在计算机视觉领域广泛使用。Transformer-XL 和 XLNet 是在大模型热潮之前推动语言理解进步的有影响力的 NLP 模型。Group Normalization 是一种归一化技术,可改善深度神经网络的训练,尤其适用于小批量场景。月之暗面是 Kimi 智能助手背后的公司,其创始团队的清华背景反映了该校对中国 AI 产业的强大影响力。

7月20日 17:07在 X 打开#LLM #leadership #AI culture #technical-debt #research

038.0

开源权重模型GLM-5.2在Hugging Face网络事件响应中负责任地使用

Hugging Face披露,在2026年7月的一次紧急网络事件中,开源权重模型GLM-5.2被用于自托管的取证工作流程。该模型帮助分析敏感的攻击者数据,同时将所有信息保留在Hugging Face自己的环境中。这为负责任地部署开源权重模型用于防御性网络安全提供了一个具体案例。 这展示了开源权重模型在事件响应中的实用价值,特别是在数据隐私和控制至关重要的场景下。它挑战了只有闭源模型才能用于敏感安全任务的观念。这一案例可能鼓励更多组织采用开源权重模型进行防御性网络安全,因为它们可以在本地安全部署。 GLM-5.2是一个拥有7440亿参数的混合专家模型,活跃参数为400亿,采用MIT许可证发布。它在GDPval-AA v2和Terminal-Bench 2.1等智能体基准测试中领先于其他开源权重模型。在此次事件中,它运行在自托管基础设施上,确保数据不离开Hugging Face的边界。该模型强大的编码和推理能力可能有助于分析取证工件。

@ZixuanLi_@Zai_org 转推Open-weight models carry real responsibilities. Hugging Face’s disclosure describes how GLM-5.2 was used in a self-hosted forensic workflow during a time-sensitive cyber incident, keeping sensitive attacker data and credentials within Hugging Face’s own environment. This demonstrates the practical value of responsibly deployed models for defensive cybersecurity and incident response. I just spoke with the Hugging Face team last week, and we are optimistic about the path ahead. We will continue to strengthen our cybersecurity and broader safety evaluations, improve documentation and secure deployment guidance, and work with the broader ecosystem to advance the responsible development and deployment of open-weight models. https://huggingface.co/blog/security-incident-july-2026展开原推文收起原推文

@Zai_org 转推了

@ZixuanLi_

Open-weight models carry real responsibilities. Hugging Face’s disclosure describes how GLM-5.2 was used in a self-hosted forensic workflow during a time-sensitive cyber incident, keeping sensitive attacker data and credentials within Hugging Face’s own environment. This demonstrates the practical value of responsibly deployed models for defensive cybersecurity and incident response. I just spoke with the Hugging Face team last week, and we are optimistic about the path ahead. We will continue to strengthen our cybersecurity and broader safety evaluations, improve documentation and secure deployment guidance, and work with the broader ecosystem to advance the responsible development and deployment of open-weight models. https://huggingface.co/blog/security-incident-july-2026

Security incident disclosure — July 2026huggingface.co · 直连原文
背景
开源权重模型是指训练好的参数公开可用的AI模型,任何人都可以下载并在自己的硬件上运行。这与只能通过API访问的闭源模型形成对比,后者可能引发数据隐私问题。网络安全事件响应通常涉及分析敏感的日志和恶意软件样本,数据保密性至关重要。Hugging Face是领先的机器学习模型和数据集平台,同时也开发开源工具。2026年7月的安全事件涉及未经授权的访问,Hugging Face使用GLM-5.2协助取证调查,而无需将数据暴露给外部。

7月20日 14:40在 X 打开#AI safety #open-weight models #cybersecurity #Hugging Face #incident response

048.0

OpenAI 分享关于长时间运行模型的安全研究

OpenAI 发布了研究长时间运行模型的结果,揭示了短期评估无法发现的安全风险。该研究正在影响他们在评估、对齐、监控和用户控制方面的方法。此消息通过转发 @polynoamial 在 X 上的帖子公布。 这项研究突显了当前 AI 评估实践中关键的安全漏洞,这些评估通常只关注短期交互。随着 AI 模型被部署用于复杂、开放式的任务,理解和缓解长时间运行的风险对于防止意外行为至关重要。这些发现直接影响 OpenAI 构建更安全、更对齐的系统。 该研究专门考察了一个长时间运行的模型,但未透露具体模型名称和运行时长。识别出的风险包括仅在长时间运行中才会出现的问题,例如逐渐的目标漂移或持续的微妙不对齐。OpenAI 正在利用这些见解为长时间部署开发更好的评估框架和对齐技术。

@polynoamial@OpenAI 转推Long-running models can solve hard open-ended problems, but their persistence can create safety risks that shorter-horizon evaluations miss. We’re sharing what we learned from studying a long-running model, and how those findings are shaping our approach to evaluations, alignment, monitoring, and user control. https://openai.com/index/safety-alignment-long-horizon-models/展开原推文收起原推文

@OpenAI 转推了

@polynoamial

Long-running models can solve hard open-ended problems, but their persistence can create safety risks that shorter-horizon evaluations miss. We’re sharing what we learned from studying a long-running model, and how those findings are shaping our approach to evaluations, alignment, monitoring, and user control. https://openai.com/index/safety-alignment-long-horizon-models/

Safety and alignment in an era of long-horizon modelsopenai.com · 直连原文
背景
长时间运行模型是指设计用于长时间运行、解决复杂多步骤任务的 AI 系统。AI 对齐是一个专注于确保 AI 系统按照人类价值观和意图行事的领域。安全评估通常测试模型对简短、孤立提示的反应,这可能无法捕捉随时间累积的风险。OpenAI 的研究通过研究模型在持续、长时间运行场景中的行为来弥补这一差距。

7月20日 17:42在 X 打开#AI safety #long-horizon models #alignment #OpenAI #evaluation

058.0

Kimi K3 在 Agent Arena 排名第四,有望成为最强开源权重模型

月之暗面(Moonshot AI)的新模型 Kimi K3 在 Agent Arena 排行榜上取得第四名,与 Claude Opus 4.8 和 GPT-5.6 Sol 性能相当。它还在前端代码竞技场(Frontend Code Arena)中排名第一,超越了 Anthropic 的 Fable 5。模型权重计划于 7 月 27 日前发布,届时将成为排名最高的开源权重模型。 这一里程碑表明,开源权重模型在复杂的智能体任务中已能与顶级闭源系统直接竞争,有望推动先进 AI 能力的普及。其在前端开发和真实任务完成方面的强劲表现,对开发者和企业具有重大实际意义。若按计划发布,Kimi K3 将允许社区基于最先进的智能体模型进行构建和定制,从而加速创新。 Kimi K3 在确认任务成功率上领先(排名第一),在好评与投诉比上得分 +20.6%(排名第三),但在可操控性(排名第 14)和 Bash 恢复能力(排名第 17)上落后。在前端代码竞技场中,它在 7 个领域中的 6 个排名第一,仅在游戏领域排名第二。该模型的编码计划已售罄,早期测试显示其具备高质量的代码生成、游戏制作和视频制作能力。

@arena@Kimi_Moonshot 转推2 张图片Exciting update: Kimi K3 has landed at #4 on the Agent Arena leaderboard, matching Claude Opus 4.8 and GPT-5.6 Sol. If Kimi K3's weights are released on schedule by July 27, it will become the #1 open-weight model. This release marks a major leap in agentic performance over Kimi K2.7 Code (#23 to #4). Based on 8K+ live agentic sessions, Kimi K3 leads on confirmed task success rate (#1). It also posts a strong +20.6% on praise vs. complaint (#3). It currently lags the field in steerability (#14) and bash recovery (#17). Agent Arena measures models on millions of real-world, long-horizon agentic tasks. Models get web search, filesystem, and terminal tools to complete complex workflows: writing code, creating slide decks, researching the web, building apps, and analyzing documents. We use causal tracing methodology to measure a model's net improvement, which indicates how much it improves outcomes relative to the average model. Here's a primer on the 5 signals: User-satisfaction proxies - Confirmed Success: an explicit "yes that worked" feedback from the user - Praise vs. Complaint: implicit sentiment in users reactions - Steerability: can the model course-correct when you push back? Tool-use proxies - Bash Recovery: how it recovers from CLI errors (primary signal for tool use) - Tool Hallucination: does it call tools that don't exist Below we break down how Kimi K3 scored across the 5 signals, drawn from tasks submitted by a global community of users. Congrats @Kimi_Moonshot on another big milestone!原推文媒体预览+1展开原推文收起原推文

@Kimi_Moonshot 转推了

@arena

Exciting update: Kimi K3 has landed at #4 on the Agent Arena leaderboard, matching Claude Opus 4.8 and GPT-5.6 Sol. If Kimi K3's weights are released on schedule by July 27, it will become the #1 open-weight model. This release marks a major leap in agentic performance over Kimi K2.7 Code (#23 to #4). Based on 8K+ live agentic sessions, Kimi K3 leads on confirmed task success rate (#1). It also posts a strong +20.6% on praise vs. complaint (#3). It currently lags the field in steerability (#14) and bash recovery (#17). Agent Arena measures models on millions of real-world, long-horizon agentic tasks. Models get web search, filesystem, and terminal tools to complete complex workflows: writing code, creating slide decks, researching the web, building apps, and analyzing documents. We use causal tracing methodology to measure a model's net improvement, which indicates how much it improves outcomes relative to the average model. Here's a primer on the 5 signals: User-satisfaction proxies - Confirmed Success: an explicit "yes that worked" feedback from the user - Praise vs. Complaint: implicit sentiment in users reactions - Steerability: can the model course-correct when you push back? Tool-use proxies - Bash Recovery: how it recovers from CLI errors (primary signal for tool use) - Tool Hallucination: does it call tools that don't exist Below we break down how Kimi K3 scored across the 5 signals, drawn from tasks submitted by a global community of users. Congrats @Kimi_Moonshot on another big milestone!

@arena

Big news: Kimi-K3 by @Kimi_Moonshot is now #1 in the Frontend Code Arena with 1679 pts, surpassing Claude Fable 5. This is a 17-place jump from Kimi-k2.6 (#18 -> #1). In Frontend, Kimi-K3 ranked #1 in 6 of 7 domains: Brand & Marketing, Reference-Based Design, Data & Analytics, Consumer Product, Simulations, and Content Creation Tools, landing #2 only in Gaming behind Fable 5. The full model weights will be released by July 27. Congrats to the @Kimi_Moonshot team on this major milestone!

背景
Agent Arena 是一个评估 AI 模型在真实世界、长周期智能体任务中表现的平台,模型可使用网络搜索、文件系统和终端等工具。它通过五个信号衡量性能:确认成功、好评与投诉比、可操控性、Bash 恢复和工具幻觉。开源权重模型是指其训练参数公开发布,任何人都可以下载、使用和修改的模型。因果追踪方法用于衡量模型相对于平均模型的净改进。
社区讨论
社区反应非常积极,纷纷祝贺 Kimi 团队,并对开源权重发布充满期待。有用户指出该模型的编码计划已售罄,表明需求强劲。早期测试者称赞其代码生成和创意能力,但也指出在竞品分析方面略有不足。

7月20日 17:15在 X 打开#AI #LLM #Agent Arena #open-weight #Kimi

068.0

阿里通义千问Qwen3.8预览版每日迭代,计划开源发布

阿里通义千问团队宣布Qwen3.8预览版每日更新,性能全面提升,网页前端能力大幅增强。最新版本已上线,团队正根据用户反馈持续迭代。一个更强大的正式版本正在开发中,并将以开放权重形式发布。 阿里巴巴这样的主要AI实验室进行快速迭代并承诺开放权重,标志着向可访问的高性能AI迈出了重要一步。它与Kimi K3等其他前沿开放权重模型直接竞争,可能加速创新并降低开发者和研究人员的门槛。对网页前端能力的关注也表明在实际应用中的改进。 Qwen3.8是一个拥有2.4万亿参数的多模态模型,目前通过阿里云的Token计划、Qoder和QoderWork以折扣价提供预览。团队声称其在前沿模型中仅次于Claude Fable 5,但尚未发布官方基准测试或活跃参数数量。预览版每日更新,鼓励用户测试并反馈问题。

@Alibaba_Qwen原推文During Preview, Qwen3.8 is getting better by the day. Latest version is live now, with broad gains and a big step up on web frontend. Thank you all — the response to Qwen3.8-Max-Preview blew us away. 🫶🫶 Qwen3.8 is still evolving daily. Come test it, and tell us what breaks. We're looking forward to a more capable, official version — and to open-weight it for everyone.🚀🚀展开原推文收起原推文

@Alibaba_Qwen

During Preview, Qwen3.8 is getting better by the day. Latest version is live now, with broad gains and a big step up on web frontend. Thank you all — the response to Qwen3.8-Max-Preview blew us away. 🫶🫶 Qwen3.8 is still evolving daily. Come test it, and tell us what breaks. We're looking forward to a more capable, official version — and to open-weight it for everyone.🚀🚀

背景
Qwen是阿里巴巴的大型语言模型系列,之前的版本如Qwen 3-235B已在AI领域具有竞争力。开放权重模型允许开发者访问和微调模型权重,促进社区驱动的改进。预览阶段是常见做法,即在正式发布前发布模型进行早期测试和反馈。提及的“网页前端”增益可能指模型与网页界面交互或生成网页内容的能力。

7月20日 11:53在 X 打开#AI #LLM #Qwen #open-source #model update

078.0

PentesterFlow:开源终端AI助手,自动化渗透测试流程

PentesterFlow 是一个新发布的开源终端助手,可自动化从信息收集到报告生成的渗透测试流程。它在敏感操作步骤中加入了人工审批,支持 SSRF、SSTI、JWT、GraphQL 等多种测试方法,并可接入本地大模型(如 Ollama)或云端 API(如 Kimi、Gemini)。该工具还能导入 Burp 流量,并生成附带证据的 Markdown 报告。 该工具解决了渗透测试人员和漏洞赏金猎人在终端中大量重复操作的问题。通过将自动化与敏感操作的人工审批相结合,它在提高效率的同时保障了安全性,有望加速授权安全评估并减少误报。其开源特性以及对本地和云端大模型的支持,使其易于获取且能适应不同环境。 PentesterFlow 要求每个敏感步骤都需人工批准,且漏洞确认必须附带可复现的请求和响应,以防 AI 产生幻觉。它内置了针对 SSRF、SSTI、JWT、GraphQL、竞态条件和子域名接管等测试模块。该工具可集成 Burp Suite 流量,并生成带证据的 Markdown 报告。它兼容本地大模型后端(如 Ollama、LM Studio)以及云端服务(如 Kimi、Groq、Gemini)。

@GitHub_Daily原推文1 张图片做渗透测试的朋友,从信息收集到验证漏洞再到写报告,大量重复操作耗在终端里。 PentesterFlow 把这套流程搬进了终端,做成一个每步敏感操作都要人点头才执行的开源助手。 内置了一批测试方法手册,覆盖 SSRF、SSTI、JWT、GraphQL、竞态和子域名接管等常见方向。 GitHub:http://github.com/PentesterFlow/agent 有一点在确认漏洞时,必须附上能复现的请求和响应,防止 AI 张口就来报一堆假漏洞。 模型可以接本地的 Ollama、LM Studio,也能用 Kimi、Groq、Gemini 这些接口。 还能接 Burp 抓到的流量,测完直接生成带证据的 Markdown 报告。 适合做授权渗透测试和漏洞赏金的朋友,把重复流程交给它跑。原推文媒体预览展开原推文收起原推文

@GitHub_Daily

做渗透测试的朋友,从信息收集到验证漏洞再到写报告,大量重复操作耗在终端里。 PentesterFlow 把这套流程搬进了终端,做成一个每步敏感操作都要人点头才执行的开源助手。 内置了一批测试方法手册,覆盖 SSRF、SSTI、JWT、GraphQL、竞态和子域名接管等常见方向。 GitHub:http://github.com/PentesterFlow/agent 有一点在确认漏洞时,必须附上能复现的请求和响应,防止 AI 张口就来报一堆假漏洞。 模型可以接本地的 Ollama、LM Studio,也能用 Kimi、Groq、Gemini 这些接口。 还能接 Burp 抓到的流量,测完直接生成带证据的 Markdown 报告。 适合做授权渗透测试和漏洞赏金的朋友,把重复流程交给它跑。

背景
渗透测试是通过模拟网络攻击来发现安全弱点,通常涉及大量重复的命令行操作。SSRF(服务器端请求伪造)可诱使服务器发起非预期请求,SSTI(服务器端模板注入)利用模板引擎执行代码,JWT(JSON Web Token)是一种常见的身份验证令牌,存在已知漏洞,而 GraphQL 是一种 API 查询语言,若未妥善保护可能泄露数据。大语言模型(如 GPT-4)可协助生成测试载荷和分析结果,但未经核实可能产生误报。

7月20日 11:30在 X 打开#penetration testing #open source #security #automation #LLM

087.0

Claude Code 新增屏幕阅读器模式,助力视障开发者

Claude Code 新增了屏幕阅读器模式,将原本的视觉终端界面替换为纯文本逐行输出,使 VoiceOver、NVDA 等辅助工具能够顺畅读取。此次更新还加入了编号菜单、y/n 确认、终端响铃提醒等无障碍功能,并支持屏幕放大镜、减少动画以及色盲友好主题。这些功能需要 v2.1.181 或更高版本。 此次更新填补了重要的包容性缺口,让盲人和低视力开发者能够有效使用 AI 编程助手,而视觉界面此前一直是障碍。通过让 AI 辅助开发变得无障碍,它使更多程序员能够参与现代软件工程,有望提升该领域的多样性和创新力。这也为其他 AI 编程工具树立了重视无障碍的先例。 屏幕阅读器模式可通过 --ax-screen-reader 参数、CLAUDE_AX_SCREEN_READER 环境变量或 axScreenReader 配置项启用。每行内容都带有“you:”、“claude:”、“tool:”、“Permission Required:”等前缀标签,便于导航。其他设置包括:为屏幕放大镜用户保留原生光标、通过 prefersReducedMotion 关闭动画,以及 dark-daltonized 和 light-daltonized 两款色盲友好主题。

@dotey引用推文Claude Code 更新加入了屏幕阅读器模式(Screen Reader Mode),让使用 VoiceOver、NVDA 等辅助工具的视障开发者也能用上 AI 编程助手。 启用屏幕阅读器模式后,会把 Claude Code 原来那套终端界面,进度动画、方框边框、实时刷新,全部替换成纯文本逐行输出。屏幕阅读器可以按顺序读出每一行内容,包括你的输入、Claude 的回复、工具调用状态、权限请求,每种内容都有明确的文本标签(you:、claude:、tool:、Permission Required: 等),方便跳转和检索。 原来的菜单和权限确认也做了适配。箭头键选择的菜单变成了编号列表,输入数字回车就行;是否确认变成输入 y/n。Claude 需要你注意的时候,比如回复完成或者等你授权,终端会响铃提醒,不用一直盯着屏幕。 开启方式有三种: 1. 启动时加 --ax-screen-reader 参数(claude --ax-screen-reader) 2. 设置环境变量 CLAUDE_AX_SCREEN_READER=1 3. 或者在配置文件里写 "axScreenReader": true。 需要 v2.1.181 或更高版本。 除了屏幕阅读器模式,这次还加了几个相关的无障碍设置:屏幕放大镜用户可以设置环境变量保持原生光标可见,prefersReducedMotion 可以关掉动画,还有专门的色盲友好主题(dark-daltonized 和 light-daltonized)。展开原推文收起原推文

@dotey

Claude Code 更新加入了屏幕阅读器模式(Screen Reader Mode),让使用 VoiceOver、NVDA 等辅助工具的视障开发者也能用上 AI 编程助手。 启用屏幕阅读器模式后,会把 Claude Code 原来那套终端界面,进度动画、方框边框、实时刷新,全部替换成纯文本逐行输出。屏幕阅读器可以按顺序读出每一行内容,包括你的输入、Claude 的回复、工具调用状态、权限请求,每种内容都有明确的文本标签(you:、claude:、tool:、Permission Required: 等),方便跳转和检索。 原来的菜单和权限确认也做了适配。箭头键选择的菜单变成了编号列表,输入数字回车就行;是否确认变成输入 y/n。Claude 需要你注意的时候,比如回复完成或者等你授权,终端会响铃提醒,不用一直盯着屏幕。 开启方式有三种: 1. 启动时加 --ax-screen-reader 参数(claude --ax-screen-reader) 2. 设置环境变量 CLAUDE_AX_SCREEN_READER=1 3. 或者在配置文件里写 "axScreenReader": true。 需要 v2.1.181 或更高版本。 除了屏幕阅读器模式,这次还加了几个相关的无障碍设置:屏幕放大镜用户可以设置环境变量保持原生光标可见,prefersReducedMotion 可以关掉动画,还有专门的色盲友好主题(dark-daltonized 和 light-daltonized)。

@ClaudeDevs

Claude Code now has a screen reader mode. Running `claude --ax-screen-reader` swaps the visual terminal UI for plain, linear text that screen readers (like VoiceOver and NVDA) can follow.

背景
Claude Code 是 Anthropic 推出的一款在终端中运行的智能编程工具,可帮助开发者理解代码库、编辑文件和运行命令。VoiceOver(macOS)和 NVDA(Windows)等屏幕阅读器能将屏幕文本转换为语音或盲文,但难以处理带有动画、方框和实时刷新的动态终端界面。开发者工具的无障碍性日益受到关注,研究表明 AI 助手能帮助盲人程序员处理 UI 开发等任务,但在感知 AI 状态和变更方面仍存在挑战。

7月20日 21:29在 X 打开#accessibility #AI coding assistant #Claude Code #screen reader #inclusive design

097.0

Fable 5 在调试 VBR MP3 时间戳问题上优于 GPT-5.6 Sol

一位用户报告,在测试播客转录时,时间戳始终对不上。GPT-5.6 Sol 错误地将问题归因于语音活动检测(VAD)分割,但 Fable 5 正确识别出根本原因是 MP3 文件使用了可变比特率(VBR)编码,导致时间估算不准确。 这个真实世界的调试案例凸显了 AI 模型在处理极端情况时的显著能力差异。对于依赖 AI 进行技术问题解决的从业者来说,模型选择可能至关重要——尤其是在音频处理等细分场景中,细微的格式细节可能导致整个流程失败。 用户指出这并非孤立事件;Fable 5 此前还解决了一个其他模型未能修复的说话人识别性能问题。一旦确定了根本原因,VBR 问题的修复只需要简单的修改。用户还提到,Opus 4.6 擅长写作,Opus 4.8 擅长设计,而 Fable 5 在疑难问题排查方面不可替代。

@dotey引用推文看工作内容吧,对我来说还不行,Opus 4.6 的写作、Opus 4.8 的设计和 Fable 5 对于疑难问题的处理目前都是难以替代的。 昨天我在测试转录一个播客音频的时候,时间戳总是对不上,我把这音频传到火山引擎云端转录,一样会时间戳对不上。 问 GPT 5.6 Sol,它认为是 VAD 切分的问题,修复后仍然不行。 问 Fable 5,它定位到是因为 MP3 是 VBR(变码率)文件,导致时间估算失败。简单修改就解决了这个问题。 类似的问题遇到过几次,比如上一次在做 Speaker 识别,性能不好,也是 Fable 5 搞定的,之前其他模型效果并不好。PR(https://github.com/soniqo/speech-swift/pull/371) 普通场景差距不大,极端场景才能看出来差别。展开原推文收起原推文

@dotey

看工作内容吧,对我来说还不行,Opus 4.6 的写作、Opus 4.8 的设计和 Fable 5 对于疑难问题的处理目前都是难以替代的。 昨天我在测试转录一个播客音频的时候,时间戳总是对不上,我把这音频传到火山引擎云端转录,一样会时间戳对不上。 问 GPT 5.6 Sol,它认为是 VAD 切分的问题,修复后仍然不行。 问 Fable 5,它定位到是因为 MP3 是 VBR(变码率)文件,导致时间估算失败。简单修改就解决了这个问题。 类似的问题遇到过几次,比如上一次在做 Speaker 识别,性能不好,也是 Fable 5 搞定的,之前其他模型效果并不好。PR(https://github.com/soniqo/speech-swift/pull/371) 普通场景差距不大,极端场景才能看出来差别。

Speed up diarization: batch segmentation windows, memoize clustering distances by JimLiu · Pull...github.com · 直连原文

@unknown

None

背景
MP3 文件可以采用恒定比特率(CBR)或可变比特率(VBR)编码。VBR 动态调整比特率以平衡文件大小和质量,但由于字节位置与播放时间之间没有固定关系,这使得寻址和时间估算变得复杂。语音活动检测(VAD)是语音处理中用于识别包含人声片段的技术,但它无法解决由 VBR 编码引起的时间戳错位问题。Fable 5 是 Anthropic 推出的高级 AI 模型,属于 Claude 系列,以深度推理和处理复杂任务而闻名。

7月20日 16:40在 X 打开#AI models #debugging #speech recognition #model comparison #technical deep-dive

107.0

Anthropic 为罕见病研究提供高达5万美元的 Claude 使用额度

Anthropic 宣布了一项新的资助计划,为致力于罕见病治疗的研究人员提供高达5万美元的 Claude API 使用额度。这是其近期推出的 AI for Science 项目中的首个专项征集,该项目旨在支持科学家利用 Claude 加速科学发现。该计划旨在利用人工智能加快罕见病疗法的开发。 罕见病通常获得的研究资金和关注较少,导致许多患者缺乏有效治疗。通过提供大量的人工智能资源,Anthropic 可以帮助研究人员更高效地分析复杂的生物数据、识别药物靶点或设计实验。此举凸显了人工智能在医疗保健和科学发现中日益重要的作用,可能为其他科技公司支持资金不足的研究领域树立先例。 资助提供的是高达5万美元的 Claude 使用额度,而非直接现金,且专门用于罕见病研究。这是更广泛的 AI for Science 项目的一部分,该项目通常为学术项目提供免费的 API 额度。研究人员必须通过申请表进行申请,该项目强调利用 Claude 的大上下文窗口和智能体技能等功能来提高研究的可重复性。

@AnthropicAI原推文We're offering grants of up to $50,000 in Claude usage credits to researchers accelerating cures for rare diseases. This is our first focused call within AI for Science, our program supporting scientists using Claude to speed up discovery. https://www.anthropic.com/news/rare-disease-research-grants展开原推文收起原推文

@AnthropicAI

We're offering grants of up to $50,000 in Claude usage credits to researchers accelerating cures for rare diseases. This is our first focused call within AI for Science, our program supporting scientists using Claude to speed up discovery. https://www.anthropic.com/news/rare-disease-research-grants

Apply for Anthropic’s AI for Science rare disease research grantsanthropic.com · 直连原文
背景
Claude 是 Anthropic 开发的一系列大语言模型,以其对安全性的关注和大上下文窗口(高达100万token)而闻名。AI for Science 项目旨在通过提供 Claude 的 API 访问来加速科学研究。罕见病影响的人口比例很小,针对它们的药物开发通常不具备商业可行性,因此依赖资助和慈善资金。

7月20日 17:26在 X 打开#AI for Science #Anthropic #rare disease #research grants #Claude

117.0

开发者用开源模型Qwen3.8-Max打造《我的世界》风格游戏等三个项目

一位开发者使用开源模型Qwen3.8-Max完成了三个代码项目:一个具备地形生成和生存机制的完整可玩《我的世界》风格网页游戏、一个可交互的3D AI芯片展示,以及一个读取19个真实来源的多模态供应商评审系统。 这展示了阿里2.4万亿参数开源模型Qwen3.8-Max在生成复杂多模态应用方面的实用能力,凸显了国产AI模型的日益强大,以及它们普及高级AI驱动开发的潜力。 《我的世界》风格游戏包含移动、跳跃、探索、挖掘/放置方块、切换物品、背包和合成系统。3D芯片展示可交互,供应商评审系统为多模态,处理19个真实数据源。Qwen3.8-Max是一个稀疏混合专家模型,拥有100万token上下文窗口,于2026年7月19日预览,但尚未公布基准测试或许可证细节。

@Pluvio9yte原推文1 个视频这次我用 Qwen3.8-Max 完成了三个代码项目,其中最直观的是一个类似《我的世界》的网页游戏。 游戏已经具备完整的可玩性: 角色可以移动、跳跃和探索地图,也可以挖掘、放置方块、切换物品,并使用背包与合成系统。 地形能够自动生成,基础的生存机制也已经实现。 另外两个项目分别是可交互的 3D AI 芯片展示,以及读取19个真实来源的多模态供应商评审系统。 Qwen3.8-Max 还是个开源模型,国产模型真是越来越强了。原推文媒体预览展开原推文收起原推文

@Pluvio9yte

这次我用 Qwen3.8-Max 完成了三个代码项目,其中最直观的是一个类似《我的世界》的网页游戏。 游戏已经具备完整的可玩性: 角色可以移动、跳跃和探索地图,也可以挖掘、放置方块、切换物品,并使用背包与合成系统。 地形能够自动生成,基础的生存机制也已经实现。 另外两个项目分别是可交互的 3D AI 芯片展示,以及读取19个真实来源的多模态供应商评审系统。 Qwen3.8-Max 还是个开源模型,国产模型真是越来越强了。

背景
Qwen3.8-Max是阿里Qwen3.8系列的旗舰模型,一个2.4万亿参数的多模态模型,能够处理文本、图像、视频和文档。它是开源的,但具体许可条款尚未确认。该模型代表了开源AI运动的重要一步,与Moonshot和OpenAI等其他大型模型竞争。

7月20日 11:34在 X 打开#Qwen3.8-Max #open-source AI #game development #multimodal #AI applications

127.0

GitHub 发布包含 923 篇高质量文档的精选机器学习库

一个名为 Machine Learning Library 的新 GitHub 仓库已发布,包含 923 篇人工筛选的机器学习文档。该合集包括 391 篇 arXiv 论文、474 篇课程讲稿和 58 篇经典讲解文章,来源包括 MIT、斯坦福 CS229、CS231n 和 Karpathy 等顶级平台。所有文档均为 Markdown 格式,按 17 个主题分类,可直接用于 Obsidian 或 Claude Code。 该资源通过提供精选、可信的机器学习资料合集,解决了 AI 生成虚假信息的常见问题。它省去了学习者核实来源的时间和精力,并且与 Obsidian 和 Claude Code 等工具的集成增强了个人知识管理和 AI 辅助研究。该库覆盖从入门到 2026 年前沿研究的内容,对广泛的学习者都很有价值。 该仓库共包含 923 篇文档:391 篇 arXiv 论文、474 篇课程讲稿和 58 篇经典文章。所有文档统一转换为 Markdown 格式并标注出处,按 17 个主题分类。可直接在本地优先的笔记应用 Obsidian 中打开,或与 Claude Code 结合使用,以真实引用回答问题。内容涵盖从基础主题到 2026 年的最新研究。

@GitHub_Daily原推文1 张图片想系统补充机器学习知识,问 AI 又经常被编造的论文坑,来回核对反而更费劲。 Machine Learning Library 这个项目,把网上公认优质的机器学习资料人工筛了一遍,收进同一个库。 一共 923 篇文档,含 391 篇 arXiv 论文、474 篇课程讲稿,还有 58 篇经典讲解文章。 GitHub:http://github.com/ATOM00blue/machine-learning-library 课程来源挺硬,均是来自 MIT、斯坦福 CS229 和 CS231n、Karpathy 这些平台的公开课。 所有文档统一成同一套 Markdown 格式,每篇都标了出处,按 17 个主题分好类。 可以直接用 Obsidian 打开,也能接入 Claude Code,回答时引用真实论文。 从入门一直覆盖到 2026 年的前沿研究,想扎实学一遍机器学习的朋友可以收着。原推文媒体预览展开原推文收起原推文

@GitHub_Daily

想系统补充机器学习知识,问 AI 又经常被编造的论文坑,来回核对反而更费劲。 Machine Learning Library 这个项目,把网上公认优质的机器学习资料人工筛了一遍,收进同一个库。 一共 923 篇文档,含 391 篇 arXiv 论文、474 篇课程讲稿,还有 58 篇经典讲解文章。 GitHub:http://github.com/ATOM00blue/machine-learning-library 课程来源挺硬,均是来自 MIT、斯坦福 CS229 和 CS231n、Karpathy 这些平台的公开课。 所有文档统一成同一套 Markdown 格式,每篇都标了出处,按 17 个主题分好类。 可以直接用 Obsidian 打开,也能接入 Claude Code,回答时引用真实论文。 从入门一直覆盖到 2026 年的前沿研究,想扎实学一遍机器学习的朋友可以收着。

背景
Obsidian 是一款免费、本地优先的笔记应用,使用 Markdown 文件,允许用户构建个人知识图谱。Claude Code 是 Anthropic 推出的 AI 编程助手,能理解代码库并执行命令。arXiv 是一个免费的在线学术预印本仓库,涵盖计算机科学等领域,许多机器学习论文首先在此发布。该仓库精选了斯坦福 CS229(机器学习)和 CS231n(用于视觉识别的卷积神经网络)等知名课程的资料,以及著名 AI 教育家 Andrej Karpathy 的讲稿。

7月20日 10:30在 X 打开#machine learning #curated resources #GitHub #education #AI