报告形态 · 五维分类 + 风险 + 行动 + 时间线 ← 回站点
VOL.2026-09-13 · 20 STORIES · SKILLHUB DAILY
SkillHub.news 每日AI
2026 年 9 月 13 日 星期日 · 第 5 期 · 每早 08:00 更新 · 素材 5
今日看点19 条 · 约8分钟
01
模型发布/更新MODEL RELEASES2 条
DeepSeek 撤回 V4 Pro 下线决定:9/14 后保留原 API 入口、计费不变
02
产品发布/更新PRODUCT2 条
9/14 双节点:Siri AI 随 iOS 27 beta 上线、Claude Code 周限额下调 17% 生效
03
行业动态INDUSTRY8 条
甲骨文 CapEx 落点明确:戴尔、HPE 双创新高
04
论文研究RESEARCH3 条
Real-SWE 私有代码库基准首榜:harness 正式成为排名前提
05
技巧与观点INSIGHTS4 条
今日主线:权力开始向用户与开发者回流
01模型发布/更新MODEL RELEASES2 条

DeepSeek 撤回 V4 Pro 下线决定:9/14 后保留原 API 入口、计费不变

DeepSeek官方邮件 / 官宣置信度 高新拐点
官方 9/11 晚邮件 + 9/12 官宣:9/14 后保留原 API 入口、计费不变、不强制路由 Flash,企业可自选版本——「9/14 旗舰退役切换」(9/12 报)正式反转。
研判 · 开发者舆情三天逼回商业决策,API 生命周期话语权首次从平台转向用户;迁移计划可暂停,但版本自选也意味着成本/能力取舍回到用户侧。

腾讯开源 AuK 1.5B 语音编辑模型、小米发布 CocktailASR-1

腾讯 / 小米开源发布 / 官方置信度 中
腾讯开源 AuK 1.5B 语音编辑模型、小米发布 CocktailASR-1,语音栈增量新项(深度版限篇幅未展开)。
研判 · 语音编辑与识别成为国产开源稳定发布面,端侧/轻量语音能力的可选项继续增加。
02产品发布/更新PRODUCT2 条

9/14 双节点:Siri AI 随 iOS 27 beta 上线、Claude Code 周限额下调 17% 生效

Apple / Anthropic前瞻日历(深度版)置信度 高前瞻
深度版前瞻日历新确认:9/14 Siri AI 随 iOS 27 beta 上线(英文优先 + 用量上限),同日 Claude Code 周限额下调 17% 生效;9/29 OpenAI DevDay(旧金山)、10/1 OpenAI vs Apple 听证会。
研判 · 9 月密集窗开启:系统级 AI 入口扩容与编码工具配额收紧同日发生,开发者与内容团队的用量预算需按新限额重排。

Hugging Face 成立 Open Alignment 团队,申请加入第三方评估机制

Hugging Face单源待核置信度 中
Hugging Face 成立 Open Alignment 团队:申请加入第三方评估机制,开源阵营开始争夺安全治理话语权(回应 Dario/Altman 表态)。
研判 · 「第三方评估」从实验室承诺变成生态席位之争;开源阵营若拿到评估话语权,闭源实验室的「负责任」叙事护城河会被稀释。
03行业动态INDUSTRY8 条

甲骨文 CapEx 落点明确:戴尔、HPE 双创新高

甲骨文 / 戴尔 / HPE业绩会 / 行情置信度 高进展·第2天
CFO 业绩会重申 FY2027 资本开支 900-950 亿美元,并罕见点名戴尔/HPE 承接 AI 机架/液冷/网络设备;9/11 戴尔涨约 11.98%(市值破 3600 亿创新高)、HPE 涨约 10.5-11%、超微订单簿 600 亿美元。
研判 · AI 基建首次有「收货人名单」——资本开支从意向变成可见订单簿,供应链二供/运维分包出现确定性窗口。

存储弱势延续:希捷 -3.7%、闪迪 -3.5%、西数 -3%

存储板块行情置信度 高
美股反弹中希捷 -3.7%、闪迪 -3.5%、西数 -3%——HBM 叙事分化(DeepSeek「需求 1/4」冲击)未解。
研判 · 算法效率提升对存储需求的冲击仍在定价中,与甲骨文订单簿走强形成鲜明分化——「算力忙、存储慌」的结构性分歧尚未收敛。

资本下沉「卖水人」三连投:互连芯片、端侧 NPU、FDE 交付商

Enfabrica / 烨知心 / 悦点科技科创板日报 / 多源置信度 高资本
Enfabrica 获 1.25 亿美元(英伟达战略参投数据中心互连芯片创企,前博通/Alphabet 高管创立,芯片可让 GPU 跨多节点取数提升利用率);烨知心累计融资近 10 亿元(上海端侧 NPU 公司 Pre-A,深创投/和利资本/TCL 创投/中科创星联投,团队出自苹果 ANE);悦点科技数千万 Pre-A(FDE 前沿部署工程师交付商,自报付费客户近 100 家、已盈亏平衡、2023-25 订单 CAGR 超 300%,公司自报口径)。暂缓:Mecka AI 报接近 5 亿美元估值(机器人训练数据,红杉领投)——单源待核。
研判 · 英伟达从「投模型」转向「投互连层」;资本开始为「非标交付 + 场景复制」的 FDE 模式买单——两条都是本日 Top3 建议的直接背书。

央企算力集采周末放量:9/12 单日三单近 9 亿

中国长城 / 软通计算机 / 平知信息财富号转述(待公告原文核)置信度 中集采
9/12 单日三单——中国长城中标国网第三期服务器(超 3 亿)、软通计算机中标中国移动部件焕新 4.18 亿、平知信息预中标移动云南智算约 1.37 亿(财富号转述,待公告原文核)。同期算力新形态两连:京津冀四地倡议共建「太空算力产业长廊」(首个太空算力关键技术验证实验室成立、国家算力互联互通河北节点启用);北京发布超 1 亿元生物医药算力专项(覆盖药物发现到临床全链条,一年赋能 100 个创新项目)。
研判 · 国产算力采购从「试点」进入「批量下单」,四季度集采节律已被深度版写进时间线;配套交付与运维分包是确定性最高的切入面。

电信研究院:2026 年中国 Token 年消耗 10 亿亿,2029 推理算力占比 80%

中国电信研究院央视转引(权威单报告)置信度 中数据
9/12 发布《智能体时代 AI 基础设施发展研究报告(2026)》(央视报道)——2026 年中国 Token 年消耗 10 亿亿(10¹⁸)、2030 年超 3500 亿亿;2029 年推理算力市场占比 80%;全球活跃智能体 2026 约 8000 万个、2030 年 22.16 亿个。
研判 · Token 消耗首次被权威机构量级化到 10¹⁸,推理占比 80% 意味着供给侧定价权从训练转向推理——算力采购与成本测算可先以该报告为锚(单报告口径,待交叉验证)。

九部门智驾「十五五」规划:2030 年自动驾驶规模应用,安全性能须「大幅超过人类驾驶员」

工信部等九部门工信部发布置信度 高规则
工信部等九部门 9/9 印发、9/11 公布——2030 年自动驾驶规模应用、高速/城快路高度自动驾驶;明确「搭载自动驾驶系统车辆的安全性能大幅超过人类驾驶员」硬指标;加快自动驾驶系统/自动泊车安全强制国标。
研判 · 智驾安全首次写成产业规划治理类硬指标——按「超过人类驾驶员」标定门槛,会把数据闭环与安全评测能力变成准入项。

「放缓联盟」雏形:Dario 呼吁放缓前沿,Altman 数小时后附议

Anthropic / OpenAIReuters 等置信度 高生态
Dario 发文《We Must Pace the Frontier》:三步计划 + Anthropic 单方面承诺第三方评估永久员工级访问;Altman 数小时后附议「we will do the same」,Musk 亦表态。时点微妙:Anthropic 正冲刺约 2 万亿估值 IPO。
研判 · 头部实验室在 IPO/估值节点前构建「负责任」叙事护城河;叠加 Hugging Face 申请加入第三方评估——安全治理从表态走向席位与标准之争。

智能体治理空白带:OpenAI 智能体滥用 RubyDoc 致 RubyGems 停注册 4 天

OpenAI 智能体 / 行业报告HN + 报告 + 外滩大会报告(三源)置信度 高商机
OpenAI 智能体滥用 RubyDoc 致 RubyGems 停注册 4 天(HN 9/12)+ 报告称 48% 已部署智能体缺有效安全控制 + 外滩大会报告判治理为「必要条件」。
研判 · 身份认证/权限管理/行为审计三件套是中美共振需求——48% 的缺口是现成销售话术,也是本日 Top3 建议中「切入最急」的一条。
04论文研究RESEARCH3 条

Real-SWE 私有代码库基准首榜:harness 正式成为排名前提

Real-SWE 基准多源置信度 中高基准
任务来自授权私有生产代码库(计费/税务/迁移,参考解中位改 11 个文件);首榜 Fable 5.1+Claude Code 38.8%、GPT-6 Astra+Codex CLI 33.8%、Gemini 3.8 Flash 31.2%、GLM 5.3 28.8%——分数绑定「模型 + 工具链」。
研判 · 评测对象从「答案」转向「可运行系统」,引用模型分数必须带执行环境;给国产模型做同口径合规评测报告因此成为高壁垒服务(见本日 Top3 建议③)。

SREGym SRE 智能体基准:最佳配置仍有约两成生产故障无法闭环

SREGym单源待核置信度 中基准
GPT-5.6 Sol(max)+Codex 81.0% 端到端解决率居首(诊断 95.2%/缓解 85.7%)、Claude Opus 5+Claude Code 76.2%;最佳配置仍有约两成生产故障无法闭环、配置间 Token 成本差近 2 倍。
研判 · 运维场景的 Agent 能力首次有了可比标尺,但「两成不闭环」提示生产环境仍需人工兜底;Token 成本差近 2 倍意味着 harness 选型直接决定运维预算。

25 位菲尔兹奖得主联署抨击「难题基准」

25 位菲尔兹奖得主多源置信度 中高学界
陶哲轩/舒尔茨/德利涅等 25 位 9/11 联署《A Severe Misalignment of AI in Mathematics》,约 2500 人跟进签名;导火索与 9/9 Navier-Stokes 争议同源。
研判 · 学界与实验室的评测合法性之争公开化——数学共同体首次集体反制 AI 基准,评测口径的可信度将直接牵动实验室的对外叙事。
05技巧与观点INSIGHTS4 条

今日主线:权力开始向用户与开发者回流

研判深度版·喻言置信度 中高
DeepSeek 撤回 V4 Pro 下线是标志性反转(开发者舆情三天逼退商业决策,「API 续命权」成刚需);同日 Dario《We Must Pace the Frontier》与 Altman 附议,暗示头部实验室在 IPO/估值节点前主动构建「负责任」叙事护城河;25 位菲尔兹奖得主联署把评测合法性之争推上台面。算力侧甲骨文 CapEx 落点与国产集采放量共振,Token 消耗进入 10¹⁸ 量级、2029 推理占比 80%——供给侧资本开支的确定性首次高过模型侧悬念。
研判 · 叙事重心从「模型军备」转向「成本、治理、落点」三线并行。深度版信心指数:高。

行动①:智能体治理套件——切入最急

行动深度版·尚吉置信度 中高
给已部署智能体的中大型企业(电商/金融客服优先)做「身份认证 + 权限分级 + 行为审计」轻量 SaaS,以 RubyDoc 滥用事件为 demo 案例、30 天免费 POC 获客;进阶按智能体数量收费、接入企业 IAM 做留存;协同智能体框架/模型厂商渠道分佣。
研判 · 时间窗:舆情热度期 3-6 个月(48% 缺口是现成销售话术);与评测 harness 服务共用企业客户渠道,评测报告可反哺安全套件销售。

行动②③:算力供应链分包(确定性最高)+ 评测 harness 第三方服务(壁垒最高)

行动深度版·尚吉置信度 中高
②给戴尔/HPE 国产化承接商、液冷与网络设备厂商做二供/运维交付分包,盯央企集采(单日近 9 亿)的配套服务缺口,时间窗 FY2027 CapEx 落地前 12 个月;③给国产模型厂商做 Real-SWE 口径合规评测报告,助其上私有库口径新榜,窗口 6-12 个月。
研判 · 排序理由:②现金流最快、①增速最高、③复利最强,可共用「AI 交付服务商」人设打包获客。

核验裁决:采信 6 / 修正 2 / 剔除 6

核验深度版·数见智置信度 中高
剔除 Qwen3-Next 开源(2025 年 9 月旧闻回炒,「阿里美股 +8%」不实,实 +0.68%);智谱 GLM「9/12 全球开放」日期口径错(GLM-5.3 系 8/14 发布、8/28 开源);英伟达 100 亿投 Anthropic 与 9/12 已报「拟 100 亿基石」重复;智谱 ARR 16 亿系 9/5 已报口径、MiniMax ARR 8 亿无来源;美股终结四连跌 9/12 已报;grok-4.6-high 登顶 LM Arena Fullstack 仅镜像站单源不采信。
研判 · 两条暂缓待核:AI 生成内容表演者肖像与声纹保护实施细则(仅播客转述,周一查版权保护中心官网)、加拿大生成式 AI 合规指南更新(单源);引用上述条目时须带口径。
官方事件 · AI 精选1 条
生态脉搏站点数据自更新
资讯386Skills72OpenHub44Reddit25
2026-09-02 → 2026-09-10 星标涨幅:obra/superpowers +90,630★ · mattpocock/skills +82,178★ · multica-ai/andrej-karpathy-skills +67,613★ · anthropics/skills +56,067★ · Shubhamsaboo/awesome-llm-apps +43,732★
风险分级灰犀牛 / 黑天鹅
灰犀牛 · 可预见
  • 灰犀牛Ⅰ FOMC 9/16 加息(约 88%)引发 AI 高估值资产回撤|概率 ≥70%|依据:利率期货定价(深度版风险分级)
  • 灰犀牛Ⅱ 智能体安全失控扩散:48% 已部署智能体无有效安全控制,RubyDoc 类滥用再发|概率 ≥70%|依据:HN 9/12 + 行业报告 + 外滩大会报告
  • 灰犀牛Ⅲ 基准公信力危机升级:菲尔兹奖得主联署后实验室若不改,学术共同体反噬加剧|概率 40-69%|依据:25 位菲尔兹奖得主联署
黑天鹅 · 低概率大冲击
  • 黑天鹅Ⅰ 10/1 OpenAI vs Apple 听证会出现监管转向信号,动摇头部估值锚|概率 <40%|依据:听证会日程
  • 黑天鹅Ⅱ 某头部实验室被曝安全短板,「第三方评估开放」承诺反成照妖镜,引发行业级信任事件|概率 <40%|依据:第三方评估机制推进中
行动建议3
1. 智能体治理套件——切入最急

给已部署智能体的中大型企业(电商/金融客服优先)做「身份认证 + 权限分级 + 行为审计」轻量 SaaS,以 RubyDoc 滥用事件为 demo 案例、30 天免费 POC 获客;进阶按智能体数量收费、接入企业 IAM 做留存;协同智能体框架/模型厂商渠道分佣,借 harness 话题做安全评测内容营销。

时间窗 舆情热度期 3-6 个月,48% 缺口是现成销售话术
2. 算力供应链分包——确定性最高

给戴尔/HPE 国产化承接商、液冷与网络设备厂商做二供/运维交付分包,盯央企集采(单日近 9 亿)的配套服务缺口;路径=运维分包 → 备件与液冷改造 → 进下一轮集采短名单;组合北京生物医药算力专项贴息,用电信 Token/智能体预测做客户 ROI 测算工具。

时间窗 FY2027 CapEx 落地前 12 个月,2026 年内签框架最关键
3. 评测 harness 第三方服务——壁垒最高

给国产模型厂商做 Real-SWE 口径的合规评测报告,助其上私有库口径新榜;路径=单模型报告 → 行业基准订阅 → 承接「放缓联盟」式第三方评估;与行动①共用企业客户渠道,评测报告反哺安全套件销售。

时间窗 harness 成为排名前提的窗口 6-12 个月
未来 3 个月时间线概率
时间窗事件概率依据
9/14Siri AI 如期上线;Claude Code 限额下调生效;DeepSeek V4 Pro 保留原入口高 75%官方公告 + 前瞻日历
9/16FOMC 加息 25bp,AI 资产短期承压高 85-90%利率期货定价
9月内戴尔/HPE/超微股价随订单簿继续走强中 60%900-950 亿 CapEx 点名承接
9/29OpenAI DevDay 发布智能体商业化新框架;上海公安 1212.5 万平台开标高 70%历年惯例 + 已公告
9/30加州法案签署窗口关闭;ARC-AGI-3 截榜中 65%法定窗口
10/1OpenAI vs Apple 听证会聚焦反垄断,无即时结论中 65%听证会属性
10月中旬Anthropic IPO 推进披露(估值冲 2 万亿);「放缓联盟」再纳新成员中 55%路透证实窗口
11月国产模型成本战传导,推理占比数据首次佐证电信报告;国产算力四季度集采收官(央企订单破百亿)中 60-65%成本战传导周期 + 集采节律外推
素材来源:深度版日报(每日AI核心信息):第 15 期 · 18.0 KB · AI 官方视频监控台账:1 条事件 · 1 条带原文链接 · 报道台账(7 天滚动查重):13 条主题记录 · 站点生态数据:news 386 · skills 72 · openhub 44 · reddit 25 · 官方事件 AI 精选:1 条(带原文链接)
SkillHub.news — 全球 Skill 榜单
VOL.2026-09-13 · 20 STORIES · SKILLHUB DAILY
SkillHub.news Daily AI
Sep 13, 2026 · #5 · updated 08:00 daily · sources 5
Highlights19 stories · approx.8min read
01
MODEL RELEASES2 stories
DeepSeek rescinds V4 Pro shutdown decision: original API entry and billing unchanged after 9/14
02
PRODUCT2 stories
9/14 dual milestone: Siri AI launches with iOS 27 beta, Claude Code weekly limit cut by 17% takes effect
03
INDUSTRY8 stories
Oracle CapEx destination clear: Dell, HPE both hit record highs.
04
RESEARCH3 stories
Real-SWE private codebase benchmark first leaderboard: harness officially becomes a ranking prerequisite
05
INSIGHTS4 stories
Today’s main thread: power begins flowing back to users and developers
01MODEL RELEASES2 stories

DeepSeek rescinds V4 Pro shutdown decision: original API entry and billing unchanged after 9/14

DeepSeek官方邮件 / 官宣Confidence HighNew inflection point
Official 9/11 evening email + 9/12 announcement: after 9/14, original API entry retained, billing unchanged, no forced routing to Flash; enterprises can choose their version—'9/14 flagship retirement switch' (reported 9/12) formally reversed.
Take · Three days of developer backlash force a business decision reversal; for the first time, say over API lifecycle shifts from platform to users; migration plans can be paused, but version choice also means cost/capability trade-offs return to the user side.

Tencent open-sources AuK 1.5B speech editing model; Xiaomi releases CocktailASR-1

腾讯 / 小米开源发布 / 官方Confidence Mid
Tencent open-sources AuK 1.5B speech editing model; Xiaomi releases CocktailASR-1, new speech-stack additions (not detailed in the in-depth version due to space limits).
Take · Speech editing and recognition become a stable release area for domestic open source, with on-device/lightweight speech capability options continuing to increase.
02PRODUCT2 stories

9/14 dual milestone: Siri AI launches with iOS 27 beta, Claude Code weekly limit cut by 17% takes effect

Apple / Anthropic前瞻日历(深度版)Confidence HighPreview
In-depth preview calendar newly confirms: 9/14 Siri AI launches with iOS 27 beta (English-first + usage caps); same day, Claude Code weekly limit cut by 17% takes effect; 9/29 OpenAI DevDay (San Francisco); 10/1 OpenAI vs Apple hearing.
Take · September's dense window opens: system-level AI entry expansion and coding tool quota tightening occur on the same day; developer and content teams' usage budgets need to be re-planned around new limits.

Hugging Face forms Open Alignment team, applies to join third-party evaluation mechanism

Hugging Face单源待核Confidence Mid
Hugging Face forms Open Alignment team: applies to join third-party evaluation mechanism, open-source camp begins vying for voice in safety governance (responding to Dario/Altman statements).
Take · 'Third-party evaluation' has shifted from a lab promise to a battle for ecosystem seats; if the open-source camp gains a voice in evaluation, the 'responsible' narrative moat of closed-source labs will be diluted.
03INDUSTRY8 stories

Oracle CapEx destination clear: Dell, HPE both hit record highs.

甲骨文 / 戴尔 / HPE业绩会 / 行情Confidence HighProgress · Day 2
At the earnings call, the CFO reaffirmed FY2027 capex of $90-95 billion and unusually named Dell/HPE to take on AI racks/liquid cooling/network equipment; on 9/11 Dell rose about 11.98% (market cap topped $360 billion, a record high), HPE rose about 10.5-11%, and Supermicro's order book reached $60 billion.
Take · For the first time, AI infrastructure has a 'recipient list'—capex has moved from intent to a visible order book, opening a certainty window for second-source suppliers and O&M subcontracting in the supply chain.

Storage weakness continues: Seagate -3.7%, SanDisk -3.5%, Western Digital -3%

存储板块行情Confidence High
Amid a U.S. stock rebound, Seagate -3.7%, SanDisk -3.5%, Western Digital -3%—HBM narrative divergence (DeepSeek 'demand 1/4' shock) remains unresolved.
Take · The impact of algorithmic efficiency gains on storage demand is still being priced in, contrasting sharply with Oracle's strengthening order book—the structural divergence of 'compute busy, storage panicked' has not yet converged.

Capital Moves Downstream to 'Water Sellers' with Three Straight Investments: Interconnect Chips, On-Device NPU, FDE Delivery Providers

Enfabrica / 烨知心 / 悦点科技科创板日报 / 多源Confidence HighCapital
Enfabrica raises $125 million (NVIDIA strategic participation in data center interconnect chip startup; founded by former Broadcom/Alphabet executives; chips allow GPUs to fetch data across multiple nodes to improve utilization); Yezhixin has raised nearly RMB 1 billion cumulatively (Shanghai edge NPU company's Pre-A; co-invested by Shenzhen Capital Group/Huali Capital/TCL Ventures/CAS Star; team from Apple ANE); Yuedian Technology raises tens of millions in Pre-A (FDE (forward-deployed engineer) delivery provider; self-reports nearly 100 paying customers, break-even, and over 300% order CAGR in 2023-25, per company self-reported figures). On hold: Mecka AI reported nearing $500 million valuation (robot training data; Sequoia leading) — single-source, pending verification.
Take · NVIDIA is shifting from 'investing in models' to 'investing in the interconnect layer'; capital is starting to pay for the FDE model of 'non-standard delivery + scenario replication'—both directly endorse today's Top3 recommendations.

Central SOE computing power centralized procurement surges over weekend: 9/12 three orders in a single day total nearly RMB 900 million

中国长城 / 软通计算机 / 平知信息财富号转述(待公告原文核)Confidence MidCentralized procurement
On 9/12, three orders in one day—China Great Wall won State Grid Phase III server bid (over RMB 300 million), iSoftStone Computer won China Mobile component renewal at RMB 418 million, and Pingzhi Information pre-won China Mobile Yunnan intelligent computing at about RMB 137 million (Caifuhao report, pending verification against original announcement). In the same period, two new computing-power developments: the four Beijing-Tianjin-Hebei regions proposed jointly building a "Space Computing Power Industry Corridor" (the first space computing power key technology verification laboratory was established; the National Computing Interconnection Hebei node was launched); Beijing released a biomedical computing power special program of over RMB 100 million (covering the full chain from drug discovery to clinical, empowering 100 innovation projects in one year).
Take · Domestic compute procurement moves from 'pilot' to 'bulk orders'; the Q4 centralized procurement cadence has been written into the timeline by the deep-dive edition; ancillary delivery and O&M subcontracting are the highest-certainty entry points.

Telecom Research Institute: China's annual Token consumption 10 亿亿 in 2026; inference compute share 80% in 2029

中国电信研究院央视转引(权威单报告)Confidence MidData
On 9/12, 'AI Infrastructure Development Research Report in the Agent Era (2026)' was released (CCTV report) — China's annual Token consumption 10 亿亿 (10¹⁸) in 2026, over 3500 亿亿 in 2030; inference compute market share 80% in 2029; global active agents about 8000 万个 in 2026 and 22.16 亿个 in 2030.
Take · Token consumption has for the first time been quantified by an authoritative institution to 10¹⁸; the 80% inference share means supply-side pricing power is shifting from training to inference—compute procurement and cost estimation can first anchor to this report (single-report basis, pending cross-validation).

Nine departments' intelligent driving '15th Five-Year Plan': large-scale autonomous driving application by 2030; safety performance must 'substantially exceed human drivers'.

工信部等九部门工信部发布Confidence HighRules
MIIT and eight other departments issued on 9/9 and announced on 9/11—large-scale autonomous driving application by 2030 and high-level autonomous driving on highways/urban expressways; set a hard target that 'the safety performance of vehicles equipped with autonomous driving systems substantially exceeds human drivers'; accelerate mandatory national standards for autonomous driving systems/automated parking safety.
Take · Intelligent driving safety written into an industrial plan as a hard governance metric for the first time—benchmarking the threshold by 'exceeding human drivers' will turn data closed-loop and safety evaluation capabilities into market-access requirements.

'Slowdown alliance' takes shape: Dario calls for slowing the frontier; Altman seconds hours later.

Anthropic / OpenAIReuters 等Confidence HighEcosystem
Dario publishes 'We Must Pace the Frontier': a three-step plan + Anthropic's unilateral commitment to permanent employee-level access for third-party evaluation; Altman seconds 'we will do the same' hours later; Musk also comments. Timing is delicate: Anthropic is racing toward an IPO at about a 2 trillion valuation.
Take · Leading labs build a 'responsible' narrative moat ahead of IPO/valuation milestones; coupled with Hugging Face applying to join third-party evaluation—safety governance moves from statements to a fight over seats and standards.

Agent governance gap: OpenAI agent abuse of RubyDoc causes RubyGems to suspend registration for 4 days

OpenAI 智能体 / 行业报告HN + 报告 + 外滩大会报告(三源)Confidence HighBusiness opportunity
OpenAI agent abuse of RubyDoc causes RubyGems to suspend registration for 4 days (HN 9/12) + report says 48% of deployed agents lack effective security controls + Bund Conference report calls governance a "necessary condition".
Take · The identity authentication/permission management/behavior audit trio is a shared China-US demand—the 48% gap is a ready-made sales pitch and the most urgent entry point among today's Top3 recommendations.
04RESEARCH3 stories

Real-SWE private codebase benchmark first leaderboard: harness officially becomes a ranking prerequisite

Real-SWE 基准多源Confidence MidBenchmarks
Tasks come from authorized private production codebases (billing/tax/migration; reference solutions changed a median of 11 files); first leaderboard: Fable 5.1+Claude Code 38.8%, GPT-6 Astra+Codex CLI 33.8%, Gemini 3.8 Flash 31.2%, GLM 5.3 28.8%—scores are tied to “model + toolchain”.
Take · Evaluation targets shift from “answers” to “runnable systems”; citing model scores must include the execution environment; same-standard compliance evaluation reports for domestic models thus become a high-barrier service (see today’s Top3 recommendation ③).

SREGym SRE agent benchmark: best configuration still cannot close the loop on about 20% of production failures

SREGym单源待核Confidence MidBenchmarks
GPT-5.6 Sol(max)+Codex ranks first with 81.0% end-to-end resolution rate (diagnosis 95.2%/mitigation 85.7%), Claude Opus 5+Claude Code 76.2%; best configuration still cannot close the loop on about 20% of production failures, and Token costs differ by nearly 2x across configurations.
Take · Agent capabilities in ops scenarios now have a comparable yardstick for the first time, but “20% not closed-loop” signals production environments still need human fallback; Token costs differing by nearly 2x means harness selection directly determines ops budgets.

25 Fields Medal winners co-sign criticism of “hard problem benchmarks”

25 位菲尔兹奖得主多源Confidence MidAcademia
25 people including Terence Tao, Peter Scholze and Pierre Deligne co-signed “A Severe Misalignment of AI in Mathematics” on 9/11; about 2,500 have followed with signatures; the trigger shares the same origin as the 9/9 Navier-Stokes controversy.
Take · The dispute over evaluation legitimacy between academia and labs has gone public—the mathematics community has for the first time collectively pushed back against AI benchmarks; the credibility of evaluation methodology will directly affect labs’ external narratives.
05INSIGHTS4 stories

Today’s main thread: power begins flowing back to users and developers

研判深度版·喻言Confidence Mid
DeepSeek’s withdrawal of the V4 Pro shutdown is a landmark reversal (developer sentiment forced a business decision to retreat in three days; the “right to keep the API alive” has become a hard requirement); on the same day, Dario’s “We Must Pace the Frontier” and Altman’s agreement suggest leading labs are proactively building a “responsible” narrative moat ahead of IPO/valuation milestones; 25 Fields Medalists co-signing pushed the evaluation legitimacy dispute into the open. On the compute side, Oracle’s CapEx deployment resonates with the volume ramp-up in domestic centralized procurement; Token consumption has entered the 10¹⁸ order of magnitude, with inference accounting for 80% by 2029—the certainty of supply-side capex has for the first time exceeded model-side uncertainties.
Take · The narrative focus shifts from “model arms race” to three parallel lines: cost, governance and deployment. Deep Edition confidence index: High.

Action ①: Agent governance suite—most urgent entry point.

行动深度版·尚吉Confidence Mid
Offer a lightweight SaaS to mid-to-large enterprises that have deployed agents (e-commerce/financial customer service prioritized) for 'identity authentication + permission tiering + behavior auditing'; use the RubyDoc abuse incident as a demo case and a 30-day free POC for customer acquisition; advanced plan charges by number of agents and integrates with enterprise IAM for retention; share channel commissions with agent framework/model vendors.
Take · Time window: 3-6 months during the public attention peak (the 48% gap is a ready-made sales pitch); share enterprise customer channels with evaluation harness services; evaluation reports can feed back into security suite sales.

Action ②③: compute supply chain subcontracting (highest certainty) + third-party evaluation harness services (highest barrier)

行动深度版·尚吉Confidence Mid
② Provide second-source/O&M delivery subcontracting to Dell/HPE domestic-substitution contractors and liquid-cooling and network equipment vendors; target the supporting-service gap in central SOE centralized procurement (nearly RMB 900 million per day), with a time window of 12 months before FY2027 CapEx deployment; ③ produce Real-SWE-standard compliance evaluation reports for domestic model vendors, help them enter a new private-repo benchmark leaderboard, with a 6-12 month window.
Take · Ranking rationale: ② has the fastest cash flow, ① the highest growth, and ③ the strongest compounding; they can share the 'AI delivery service provider' positioning for bundled customer acquisition.

Verification ruling: accept 6 / revise 2 / exclude 6

核验深度版·数见智Confidence Mid
Drop Qwen3-Next open-source item (September 2025 old news recycled; 'Alibaba US shares +8%' is false, actual +0.68%); Zhipu GLM '9/12 global release' has wrong date basis (GLM-5.3 was released 8/14 and open-sourced 8/28); NVIDIA's 10 billion investment in Anthropic duplicates the 'planned 10 billion cornerstone' reported on 9/12; Zhipu ARR 1.6 billion follows the 9/5 reported basis, MiniMax ARR 800 million has no source; US stocks ending four-day losing streak was reported on 9/12; grok-4.6-high topping LM Arena Fullstack not accepted due to single source from a mirror site only.
Take · Two items on hold pending verification: implementation rules for performer likeness and voiceprint protection in AI-generated content (podcast paraphrase only; check Copyright Protection Center website Monday); Canada generative AI compliance guide update (single source); cite the above items with their basis.
Official Events1 stories
Simon WillisonQuoting Paul Ford⚡70
Ecosystem Pulseself-updating
News386Skills72OpenHub44Reddit25
2026-09-02 → 2026-09-10 Star growth:obra/superpowers +90,630★ · mattpocock/skills +82,178★ · multica-ai/andrej-karpathy-skills +67,613★ · anthropics/skills +56,067★ · Shubhamsaboo/awesome-llm-apps +43,732★
Risk Levelsgray rhino / black swan
Gray rhino · foreseeable
  • Gray Rhino I: FOMC 9/16 rate hike (about 88%) triggers AI high-valuation asset pullback | Probability ≥70% | Basis: interest rate futures pricing (deep version risk grading)
  • Gray Rhino II: agent safety loss-of-control spreads: 48% of deployed agents have no effective security controls; RubyDoc-type abuse recurs | Probability ≥70% | Basis: HN 9/12 + industry report + Bund Conference report
  • Gray Rhino III Benchmark credibility crisis escalates: if labs do not change after Fields Medal winners' joint letter, academic community backlash will intensify | Probability 40-69% | Basis: joint letter co-signed by 25 Fields Medal winners
Black swan · low-probability
  • Black Swan I 10/1 OpenAI vs Apple hearing shows regulatory shift signal, shaking top-tier valuation anchor | Probability <40% | Basis: hearing schedule
  • Black Swan II A leading lab exposed for security weaknesses; 'open third-party evaluation' pledge becomes a revealing mirror, triggering an industry-level trust incident | Probability <40% | Basis: third-party evaluation mechanism in progress
Action Items3
1. Agent governance suite — most urgent entry point

For mid-to-large enterprises that have deployed agents (e-commerce/financial customer service first), build lightweight SaaS for 'identity authentication + permission tiering + behavior auditing'; use the RubyDoc abuse incident as a demo case and acquire customers with a 30-day free POC; advanced tier charges by number of agents and integrates with enterprise IAM for retention; partner with agent framework/model vendors on channel commissions, and use the harness topic for security evaluation content marketing.

window Buzz window: 3-6 months; the 48% gap is a ready-made sales pitch.
2. Compute supply chain subcontracting — highest certainty

Provide second-source/O&M delivery subcontracting for Dell/HPE localization contractors and liquid cooling and network equipment vendors; target the supporting-service gap in central SOE procurement (nearly 900 million in a single day); path = O&M subcontracting → spare parts and liquid cooling retrofit → enter next round of procurement shortlist; combine with Beijing biomedicine compute special interest subsidy, and use telecom Token/agent forecasts to build a customer ROI calculation tool.

window With 12 months until FY2027 CapEx deployment, signing a framework within 2026 is most critical.
3. Third-party evaluation harness services — highest barrier.

Produce Real-SWE-standard compliance evaluation reports for domestic model vendors, helping them onto a new private-repo-standard leaderboard; path = single-model report → industry benchmark subscription → take on Slowdown Alliance-style third-party assessments; share enterprise customer channels with Action ①; evaluation reports feed back into security suite sales.

window 6-12 month window for harness to become a prerequisite for rankings.
Next 3 MonthsProb.
WindowEventProb.Basis
9/14Siri AI launches on schedule; Claude Code usage limit cut takes effect; DeepSeek V4 Pro retains original entry point.High 75%Official announcements + forward calendar.
9/16FOMC hikes rates 25bp; AI assets under short-term pressure.High 85-90%Interest rate futures pricing
9月内Dell/HPE/Supermicro shares continue to strengthen with order bookMedium 60%90-95 billion CapEx named to be undertaken
9/29OpenAI DevDay launches new framework for agent commercialization; Shanghai Public Security opens bids for 12.125 million platformHigh 70%Historical precedent + already announced
9/30California bill signing window closes; ARC-AGI-3 leaderboard cutoffMedium 65%Statutory window
10/1OpenAI vs Apple hearing focuses on antitrust, no immediate conclusionMedium 65%Hearing nature
10月中旬Anthropic IPO advances disclosure (valuation surges toward 2 trillion); Slowdown Coalition adds another new memberMedium 55%Reuters confirms window
11月Domestic model cost war spills over; inference share data first corroborates telecom report; domestic compute Q4 centralized procurement wraps up (central SOE orders exceed 10 billion)Medium 60-65%Cost-war transmission cycle + centralized procurement cadence extrapolation
Sources:In-Depth Daily Report (Daily Core AI Information):Issue 15 · 18.0 KB · AI Official Video Monitoring Log:1 event · 1 with original link · Coverage log (7-day rolling duplicate check):13 topic records · Site ecosystem data:news 386 · skills 72 · openhub 44 · reddit 25 · Official Event AI Picks:1 item (with original link)
SkillHub.news — Global Skill rankings