报告形态 · 五维分类 + 风险 + 行动 + 时间线 ← 回站点
VOL.2026-09-14 · 21 STORIES · SKILLHUB DAILY
SkillHub.news 每日AI
2026 年 9 月 14 日 星期一 · 第 6 期 · 每早 08:00 更新 · 素材 5
今日看点20 条 · 约8分钟
01
模型发布/更新MODEL RELEASES2 条
DeepSeek 切换落地实测:V4.1 Flash 智能指数 39.5 反超 V4 Pro,但单任务输出多 62%
02
产品发布/更新PRODUCT2 条
评测底座与框架同日小步快跑:ageval 开源、DSPy 3.4.0 Beta 1 发布
03
行业动态INDUSTRY8 条
智谱 50 亿美元融资完成:20 亿配售 + 30 亿零息可转债
04
论文研究RESEARCH3 条
GPT-6 Astra 攻克 FrontierMath Tier 4:整套基准被 AI 拿下
05
技巧与观点INSIGHTS5 条
今日主线:比价逻辑从「单价」转向「单任务总成本」
01模型发布/更新MODEL RELEASES2 条

DeepSeek 切换落地实测:V4.1 Flash 智能指数 39.5 反超 V4 Pro,但单任务输出多 62%

DeepSeekAA 独立实测(多源)置信度 高进展·第4天
AA 独立实测:V4.1 Flash 智能指数 39.5(峰值档报 40),超过 V4 Pro 的 36;但单任务平均输出约 8.9 万 token、比 V4 Pro 多 62%,部分抵消降价红利。组合项为 MIT 开源权重 + KV cache 需求仅前代 1/4。
研判 · 官方「保留原入口」(9/13 已报)与实测路由行为并存,迁移决策以官方文档为准——比价逻辑从「单价」转向单任务总成本,自托管成本曲线成为企业选型新变量。

Cognition 发布 SWE-2:基于 Kimi K3(2.8T 参数)RL 后训练

Cognition官方博客置信度 高新发布
官方博客证实基于 Kimi K3(2.8T 参数)RL 后训练;自家 FrontierCode 1.1 Main 50.0%(距 Fable 5.1 不足 1 分)、成本声称低 64%。
研判 · FrontierCode 系厂商自建榜单、竞品分数为厂商自测口径,公信力打折;海外编码 Agent 采购国产开源底座,是开源权重外溢的直接样本。
02产品发布/更新PRODUCT2 条

评测底座与框架同日小步快跑:ageval 开源、DSPy 3.4.0 Beta 1 发布

浙大 ZJU-REAL / DSPy开源发布置信度 中高工具链
浙大 ZJU-REAL 开源 ageval:可插拔 Agent 评测底座,环境插件覆盖 Docker/E2B/Daytona/Local,一行配置切换运行时;DSPy 3.4.0 Beta 1 发布,LM 执行层换内置引擎。
研判 · 生态从「堆模型」转向「评测底座 + 推理经济学」;评测运行时可插拔,意味着 harness 兼容性成为 Agent 选型的隐性成本项。

数据产权登记系统 9/11 上线:北京/上海/深圳三所率先成为登记机构

国家数据局官方 / 多源置信度 高新面孔
国家数据局 9/11 系统上线,北京/上海/深圳三所数据交易所率先成为登记机构,出具数据「身份证明」,明确持有权/使用权/经营权归属;本地大模型办公(消费级 PC 可跑 27B)商业化信号同步出现(置信中)。
研判 · 数据资产入表与交易的前置基建落地——确权进入实操期,本地/私有化部署的商业化路径随之打开。
03行业动态INDUSTRY8 条

智谱 50 亿美元融资完成:20 亿配售 + 30 亿零息可转债

智谱公告(9/13)置信度 高资本·新面孔
9/13 公告:20 亿配售 + 30 亿零息可转债;MaaS ARR 16 亿美元系 8/31 半年报已披露口径(勿当新增);API 毛利率 -0.4% → 24.6% 转正;Token 调用量较年初 +40 倍。
研判 · 「小股大债」结构显示海外长线资金抢筹国产开源旗舰;两个月内再度吸金,10 月上旬资金到账放量值得跟踪。

Anthropic 选定纳斯达克,推介材料首次披露财务:Q2 营收超 115 亿美元、首次盈利

AnthropicBI / 彭博 / 路透置信度 高进展·第2天
多源证实交易所锁定纳斯达克;推介材料披露 Q2 营收超 115 亿美元、首次盈利,Q3「有望」连续盈利;英伟达基石至多 100 亿美元仍在洽谈。
研判 · 史上最大 IPO(约 2 万亿估值)进入推介阶段,与「放缓宣言」同框——资本叙事与安全叙事并行;Altman 称 OpenAI 2026 年不上市,两家分化。

优必选柳州万台级人形机器人超级工厂投产:每 10 分钟下线 1 台

优必选金台资讯 / 科创板日报置信度 高新面孔
9/12 投产:1.4 万㎡、每 10 分钟下线 1 台 WalkerS/Cruzr、年产能超万台,全球首个万台级基地(联合西门子数字孪生、零部件国产化率超 75%);另中标乐山商用服务人形机器人项目 1.5081 亿元(9/8 公示)。
研判 · 规模降本启动——宇树单价两年已降 72% 至 16.64 万元,中国厂商占全球出货约 97%(产业口径);具身量产进入兑现窗口。

人形机器人 2026H1 全球出货 1.91 万台、同比 +272%

Smart Analytics Global / 高盛单一机构口径 / 投行预测置信度 中高数据
2026H1 全球人形机器人出货约 1.91 万台、同比 +272%(单一机构口径);高盛预测 2035 年整机均价降至 2.13 万美元(-49%)。
研判 · 量产竞赛正式开跑,Q3 出货数据是下一个验证点;出货口径为单一机构,引用须带来源。

FOMC 前夜双压:加息定价 86.2% ×「AI 降速」叙事冲击暗盘

CME / 美股暗盘CME 官方 / 行情置信度 高进展
CME 9/14 口径加息概率 86.2%(9/13 盘后约 90%、一周前约 59%);9/13 暗盘 SK 海力士 -4.60%、英特尔 -3.81%、美光 -3.25%、英伟达 -2.16%(与 21 财经/雪球逐项吻合);9/16 议息 + 8 月零售销售同日公布。
研判 · 「降速」情绪冲击与甲骨文 CapEx 指引(算力开支未退潮)背离,属短线情绪、中期待 10 月底 SaaS 三季报验证;白宫「没有理由加息」原话未获独立核实,已剔除引用。

治理基建三件套:智能体身份国标立项 + 词元券 + 数据产权登记

国家数据局 / 成都 / 外滩大会多源置信度 高规则·商机
《区块链智能体身份管理要求》国家标准化指导性技术文件在外滩大会启动立项;成都 9/11 发布词元券征求意见稿(年 1 亿元、单主体最高 200 万、补贴 30-50%、小微企业 50%),重庆「词元券 + 词元贷」先行;国家数据局 9/11 数据产权登记系统上线(北京/上海/深圳三所)。
研判 · 「确权—身份—补贴」三位一体的治理基建窗口打开,「标准即采购清单」;空白带:词元代申领/核销、模型调用聚合采购、券适配服务商。

政策与趋势面:特朗普反对「减速共识」、工信部中小企业计划、赛迪报告判推理需求与训练持平

特朗普 / 工信部 / 赛迪顾问 / 外滩机构多源置信度 高规则·前瞻
9/13 特朗普在爱尔兰表态反对减速 AI(「谁赢 AI 谁赢」),称风险警告系「负面势力炒作」、只支持护栏,Sanders 同日回击并呼吁中美谈判 AI 暂停条约;工信部《人工智能中小企业创业支持计划》(8/29 印发)三年培育科技和创新型中小企业 1 万家以上、专精特新「小巨人」突破 2000 家;赛迪顾问《前沿技术预测预见研究》(9/8 发布)称大模型 2 年内产业化趋近度达 1、推理算力需求与训练持平;外滩大会五家头部机构确认「上下两层」(底层算力芯片 + 上层 Agent/新入口)投资分层。
研判 · 「减速共识」正面对撞,9/16 前后 Ban ASI Act 议程攻防白热化;中小企业侧算力与场景补贴入口打开,竞争焦点从算法单点转向「算法—系统—硬件」工程化协同。

天数智芯供字节 GPU 翻倍至 10 万颗(暂缓,待官宣)

天数智芯 / 字节知情人士口径(暂缓)置信度 中暂缓
多渠道转引(原路透 6 月口径为 5 万颗),知情人士口径、未见路透原文。
研判 · 国产卡紧缺由「涨价」转入「抢货锁量」阶段的信号,待官宣再升级置信。
04论文研究RESEARCH3 条

GPT-6 Astra 攻克 FrontierMath Tier 4:整套基准被 AI 拿下

Epoch AI / GPT-6 Astra量子位首发 / 多源印证置信度 高基准
解出该层级全部剩余题目、得分 97.6%(此前最高仅约 5%),Epoch AI 宣布 Tier 4 饱和;叠加此前 Erdős 层 2/68 的上下文,整套 FrontierMath 被 AI 完整攻克。
研判 · 科研级基准失去区分度,评测体系加速换代——私有/垂类榜单成为新旧口径切换战场,独立可复现评测出现结构性空位。

WPS AI 表格 Agent 登顶 SpreadsheetBench V2 Full:50.86%

WPS AI(金山办公)中媒一致转引置信度 中新面孔
50.86% 居首,超 Claude Opus 4.6(34.89%)、GPT-5.2(26.79%)、Gemini 3.1 Pro(23.68%);三个月内第三次国际登顶(榜单原始页未独立复核)。
研判 · 垂类办公场景成国产出奇兵赛道;引用须带「中媒转引、原始榜单页未复核」口径。

评测信任危机:出题方身兼卖题方

SemiAnalysis / DatacurveSemiAnalysis 调查置信度 中高学界·风险
SemiAnalysis 调查:DeepSWE 出题方 Datacurve 同时向前沿实验室出售专家级编码训练数据(对具体厂商的点名指控未独立核实,暂缓引用);此前审计口径:该基准存在假阳性 8.5% / 假阴性 24%。
研判 · 叠加 9/13 菲尔兹奖得主联署,「谁出题」的信任危机从数学圈扩散到工程榜——「可复现 + 无利益冲突」的独立评测验证成为结构性空位。
05技巧与观点INSIGHTS5 条

今日主线:比价逻辑从「单价」转向「单任务总成本」

研判深度版·喻言置信度 中高主线
DeepSeek 切换落地实测证明「便宜要按总账算」(单任务输出多 62% 吃掉一部分降价红利);同日 GPT-6 Astra 攻克 FrontierMath Tier 4 让科研级基准失去区分度,评测信任危机从数学圈扩散到工程榜;资本侧智谱 50 亿美元落地与 Anthropic 抢跑史上最大 IPO 并行,具身量产(万台级工厂 + 出货 +272%)进入兑现窗口。
研判 · 叙事重心从「模型军备」转向「成本、评测、落点」三线并行;风险分级显示灰犀牛集中在利率与评测公信力,黑天鹅在「减速共识」政策化与 IPO 定价。深度版信心指数:高。

行动①:切换「单任务总成本」计费框架(时间窗 2 周内)

行动深度版·尚吉(行动①)置信度 中高采购方
用 V4.1 Flash 做 POC,公开对比「token 总量 × 单价」的真实 TCO(实测单任务输出多 62% 即现成论据);路径:企业 CTO → 采购评审委员会 → 写入招标条款;打包词元券补贴申请(30-50%)+ 智能体身份国标预合规,一份方案双红利。
研判 · 时间窗最短、见效最快:比价框架的重定义只在版本切换落地后的几周内有效。

行动②:抢占「第三方评测验证」空位(时间窗 9/30 ARC-AGI-3 截榜前)

行动深度版·尚吉(行动②)置信度 中高创业者
做「可复现 + 无利益冲突」的独立评测服务,承接基准失锚后的采购决策需求(Datacurve 卖题兼出题是现成反面教材);路径:开源社区信誉 → 企业尽调订单 → 国标参编资格;打包绑定数据产权登记三所,形成「评测 + 确权」服务包。
研判 · 壁垒最高、复利最强;窗口固定在 ARC-AGI-3 截榜与智能体身份国标立项两个节点之间。

行动③:押注 Agent 治理与身份基础设施(时间窗 9/29-10/1)

行动深度版·尚吉(行动③)置信度 中高从业者
围绕 OpenAI DevDay 发布 + 上海公安 1212.5 万开标做智能体身份/审计中间件;路径:政府试点标段 → 金融等强合规行业 → 乘「1 万家中小企业」政策放量;打包 PaperCut 类事件复盘白皮书引流 + 小巨人申报辅导。
研判 · 切入最急:身份国标已立项、「标准即采购清单」,政策与招标节点挤在同一周。

核验裁决:采信 14 / 修正 5 / 剔除 1 / 暂缓 3

核验深度版·数见智置信度 中高剔除/暂缓
剔除:OpenRouter 周量「122.61 万亿」未获独立证实(12.1T 周量 vs 46.7T 日量并存的口径冲突);「开源权重占比破 50%」系口径拼接(OpenRouter 官方口径约 1/3,破 50% 的是中国模型份额);SWE-Bench Pro「掉 21.48 分」疑似拼接旧审计造新数(真实审计口径为假阳性 8.5% / 假阴性 24%)。暂缓:天数智芯 10 万颗、DeepSWE 对具体厂商点名、国产份额 40%/Epoch 十亿美元(单源/未核到)。
研判 · 引用上述条目必须带口径;「哈塞特『没有理由加息』」原话未获独立核实已剔除;Ruchir「四O」泡沫框架系旧观点回炒,仅作利率敏感期背景。
官方事件 · AI 精选1 条
生态脉搏站点数据自更新
资讯392Skills72OpenHub44Reddit25
2026-09-02 → 2026-09-10 星标涨幅:obra/superpowers +90,630★ · mattpocock/skills +82,178★ · multica-ai/andrej-karpathy-skills +67,613★ · anthropics/skills +56,067★ · Shubhamsaboo/awesome-llm-apps +43,732★
风险分级灰犀牛 / 黑天鹅
灰犀牛 · 可预见
  • 灰犀牛Ⅰ FOMC 加息落地引发 AI 估值二次压缩|概率 87%|冲击面:高估值 AI / 算力链|观察哨:9/16 点阵图 + 零售销售(深度版风险分级)
  • 灰犀牛Ⅱ 评测信任危机扩散至榜单生态|概率 60%|冲击面:基准方 / 采购方|观察哨:SemiAnalysis 后续报道(深度版风险分级)
黑天鹅 · 低概率大冲击
  • 黑天鹅Ⅰ 头部实验室「减速共识」变成政策约束|概率 15%|冲击面:前沿模型供给|观察哨:特朗普 vs 实验室表态升级(深度版风险分级)
  • 黑天鹅Ⅱ Anthropic IPO 定价失败拖累一级市场|概率 10%|冲击面:融资环境|观察哨:10 月中旬路演认购倍数(深度版风险分级)
行动建议3
1. 采购方:切换「单任务总成本」计费框架

用 V4.1 Flash 做 POC,公开对比「token 总量 × 单价」的真实 TCO(实测单任务输出多 62% 就是现成论据)。路径:企业 CTO → 采购评审委员会 → 写入招标条款。打包:词元券补贴申请(30-50%)+ 智能体身份国标预合规,一份方案双红利。

时间窗 时间窗:2 周内
2. 创业者:抢占「第三方评测验证」空位

做「可复现 + 无利益冲突」的独立评测服务,承接基准失锚后的采购决策需求(Datacurve 卖题兼出题是现成反面教材)。路径:开源社区信誉 → 企业尽调订单 → 国标参编资格。打包:绑定数据产权登记三所,形成「评测 + 确权」服务包。

时间窗 时间窗:9/30 ARC-AGI-3 截榜前
3. 从业者:押注 Agent 治理与身份基础设施

围绕 DevDay 发布 + 上海公安 1212.5 万开标做智能体身份/审计中间件。路径:政府试点标段 → 金融等强合规行业 → 乘「1 万家中小企业」政策放量。打包:PaperCut 类事件复盘白皮书引流 + 小巨人申报辅导。

时间窗 时间窗:9/29-10/1
未来 3 个月时间线概率
时间窗事件概率依据
9/15-16FOMC 议息 + 8 月零售销售同日落地87% 加息↓ 高估值(CME 定价 + 深度版时间线)
9/16 前后ASI 法案相关议程攻防(Ban ASI Act,9/3 预告)85%↑ 治理(政策议程)
9/22 前后DeepSeek V4.1 生态适配潮90%↑ 开源(适配进展外推)
9/29OpenAI DevDay + 上海公安智能体平台开标90%↑ Agent(已公告 + 临近)
9/30加州法案签署窗口 + ARC-AGI-3 截榜80%双向(法定窗口)
10/1OpenAI vs Apple 听证会90%↓ 封闭生态(听证会日程)
10 月上旬智谱 50 亿资金到账放量75%↑ 国产(公告节奏外推)
10 月中旬Anthropic IPO 路演 / 定价70%↑ 现金流 AI(路透证实窗口)
10 月下旬HBM 荒向国产卡二段传导65%↑ 算力成本(产业链传导)
11 月初词元券首批补贴发放70%↑ 中小企业(成都征求意见稿)
11 月中旬人形机器人 Q3 出货验证(+272% 是否延续)60%验证(Q3 数据披露窗口)
12 月智能体身份国标征求意见稿55%↑ 标准(立项后常规节奏)
素材来源:深度版日报(每日AI核心信息):第 16 期 · 20.3 KB · AI 官方视频监控台账:1 条事件 · 1 条带原文链接 · 报道台账(7 天滚动查重):14 条主题记录 · 站点生态数据:news 392 · skills 72 · openhub 44 · reddit 25 · 官方事件 AI 精选:1 条(带原文链接)
SkillHub.news — 全球 Skill 榜单
VOL.2026-09-14 · 21 STORIES · SKILLHUB DAILY
SkillHub.news Daily AI
Sep 14, 2026 · #6 · updated 08:00 daily · sources 5
Highlights20 stories · approx.8min read
01
MODEL RELEASES2 stories
DeepSeek switch rollout test: V4.1 Flash intelligence index 39.5 overtakes V4 Pro, but per-task output is 62% higher
02
PRODUCT2 stories
Eval infrastructure and framework make same-day incremental moves: ageval goes open source, DSPy 3.4.0 Beta 1 released
03
INDUSTRY8 stories
Zhipu completes US$5 billion financing: US$2 billion placement + US$3 billion zero-coupon convertible bonds
04
RESEARCH3 stories
GPT-6 Astra conquers FrontierMath Tier 4: entire benchmark taken by AI.
05
INSIGHTS5 stories
Today's main line: price comparison logic shifts from 'unit price' to 'total cost per task'
01MODEL RELEASES2 stories

DeepSeek switch rollout test: V4.1 Flash intelligence index 39.5 overtakes V4 Pro, but per-task output is 62% higher

DeepSeekAA 独立实测(多源)Confidence HighProgress · Day 4
AA independent test: V4.1 Flash intelligence index 39.5 (peak tier reported 40), above V4 Pro's 36; but average per-task output about 89,000 tokens, 62% more than V4 Pro, partially offsetting the price-cut benefit. Combined factors: MIT open-source weights + KV cache demand only 1/4 of previous generation.
Take · Official 'retain original entry point' (reported 9/13) coexists with observed routing behavior; migration decisions should be based on official documentation—price-comparison logic shifts from 'unit price' to total per-task cost, and the self-hosting cost curve becomes a new variable in enterprise selection.

Cognition releases SWE-2: RL post-trained on Kimi K3 (2.8T parameters)

Cognition官方博客Confidence HighNew release
Official blog confirms RL post-training based on Kimi K3 (2.8T parameters); its own FrontierCode 1.1 Main 50.0% (less than 1 point from Fable 5.1), claimed cost 64% lower.
Take · FrontierCode is a vendor-built leaderboard, and competitor scores use vendor self-testing criteria, denting credibility; overseas coding agents procuring domestic open-source bases is a direct sample of open-source weight spillover.
02PRODUCT2 stories

Eval infrastructure and framework make same-day incremental moves: ageval goes open source, DSPy 3.4.0 Beta 1 released

浙大 ZJU-REAL / DSPy开源发布Confidence MidToolchain
Zhejiang University's ZJU-REAL open-sources ageval: a pluggable Agent evaluation base, with environment plugins covering Docker/E2B/Daytona/Local and one-line config to switch runtimes; DSPy 3.4.0 Beta 1 released, with the LM execution layer switched to a built-in engine.
Take · The ecosystem is shifting from 'stacking models' to 'eval infrastructure + inference economics'; pluggable evaluation runtimes mean harness compatibility becomes a hidden cost item in Agent selection.

Data property rights registration system goes live on 9/11: the three entities in Beijing/Shanghai/Shenzhen first become registration agencies

国家数据局官方 / 多源Confidence HighNewcomer
The National Data Administration system goes live on 9/11; the three data exchanges in Beijing/Shanghai/Shenzhen are first to become registration agencies, issuing data 'identity certificates' and clarifying holding/use/operating rights; commercialization signals for local large-model office work (27B can run on consumer PCs) appear simultaneously (medium confidence).
Take · The foundational infrastructure for data asset balance-sheet inclusion and trading is in place—rights confirmation enters the implementation phase, and the commercialization path for local/private deployment opens accordingly.
03INDUSTRY8 stories

Zhipu completes US$5 billion financing: US$2 billion placement + US$3 billion zero-coupon convertible bonds

智谱公告(9/13)Confidence HighCapital · New Faces
9/13 announcement: US$2 billion placement + US$3 billion zero-coupon convertible bonds; MaaS ARR of US$1.6 billion is the figure already disclosed in the 8/31 half-year report (not new); API gross margin -0.4% → 24.6%, turning positive; Token call volume +40x from start of year.
Take · 'Small equity, large debt' structure shows overseas long-term funds scrambling for the domestic open-source flagship; attracting capital again within two months; capital arrival and volume ramp-up in early October are worth tracking.

Anthropic picks Nasdaq; pitch materials disclose finances for first time: Q2 revenue over US$11.5 billion, first profit

AnthropicBI / 彭博 / 路透Confidence HighProgress · Day 2
Multiple sources confirm Nasdaq as listing exchange; pitch materials disclose Q2 revenue over $11.5 billion, first profit, Q3 “expected” to be profitable again; NVIDIA cornerstone investment of up to $10 billion still under negotiation.
Take · History’s largest IPO (about $2 trillion valuation) enters the roadshow stage, appearing alongside a “slowdown declaration” — capital narrative and safety narrative run in parallel; Altman says OpenAI will not go public in 2026, the two diverge.

UBTECH’s Liuzhou 10,000-unit-class humanoid robot super factory begins production: one unit rolls off the line every 10 minutes

优必选金台资讯 / 科创板日报Confidence HighNewcomer
9/12: production starts—1.4 万㎡, one WalkerS/Cruzr rolls off the line every 10 minutes, annual capacity over 10,000 units, the world's first 10,000-unit-class base (with Siemens digital twin, component localization rate over 75%); also won the Leshan commercial service humanoid robot project worth 1.5081 亿元 (announced 9/8).
Take · Scale-driven cost reduction begins—Unitree's unit price has fallen 72% in two years to 16.64 万元; Chinese manufacturers account for about 97% of global shipments (industry figures); embodied mass production enters a realization window.

Humanoid robots: 2026H1 global shipments 1.91 ten-thousand units, +272% YoY

Smart Analytics Global / 高盛单一机构口径 / 投行预测Confidence MidData
2026H1 global humanoid robot shipments about 1.91 ten-thousand units, +272% YoY (single-institution basis); Goldman Sachs forecasts the average complete-unit price to drop to 2.13 ten-thousand USD in 2035 (-49%).
Take · The mass-production race has officially begun; Q3 shipment data is the next validation point; shipment figures are from a single institution, so citations must include the source.

Double pressure on FOMC eve: rate-hike pricing 86.2% × 'AI slowdown' narrative hits dark pool

CME / 美股暗盘CME 官方 / 行情Confidence HighProgress
CME 9/14 rate-hike probability 86.2% (about 90% after the 9/13 close, about 59% a week earlier); 9/13 dark pool: SK Hynix -4.60%, Intel -3.81%, Micron -3.25%, NVIDIA -2.16% (matching 21 Finance/Xueqiu item by item); 9/16 rate decision + August retail sales due same day.
Take · The 'slowdown' sentiment shock diverges from Oracle CapEx guidance (compute spending has not ebbed); it is short-term sentiment, with medium-term validation pending late-October SaaS Q3 earnings; White House's 'no reason to raise rates' exact quote not independently verified, citation removed.

Governance infrastructure trio: agent identity national standard project + token vouchers + data property rights registration

国家数据局 / 成都 / 外滩大会多源Confidence HighRules · Business Opportunities
The national standardization guiding technical document 'Blockchain Agent Identity Management Requirements' launched its project at the Bund Conference; Chengdu released a draft for comments on token vouchers on 9/11 (annual 1 亿元, max 200 万 per entity, 30-50% subsidy, 50% for micro and small enterprises), while Chongqing's 'token vouchers + token loans' takes the lead; the National Data Administration launched its data property rights registration system on 9/11 (three exchanges in Beijing/Shanghai/Shenzhen).
Take · The governance infrastructure window for the trinity of 'rights confirmation—identity—subsidies' is opening, with 'standards as procurement lists'; gaps: token voucher application/redemption on behalf, aggregated procurement of model calls, voucher adaptation service providers.

Policy and trends: Trump opposes the 'deceleration consensus'; MIIT SME plan; CCID report says inference demand on par with training

特朗普 / 工信部 / 赛迪顾问 / 外滩机构多源Confidence HighRules · Outlook
9/13 Trump in Ireland opposed slowing AI ('whoever wins AI wins'), calling risk warnings 'hype by negative forces' and supporting only guardrails; Sanders hit back the same day and called for China-US negotiations on an AI pause treaty; MIIT's 'AI SME Startup Support Program' (issued 8/29) aims to cultivate over 1 万家 technology and innovative SMEs in three years, with specialized, refined, distinctive and innovative 'little giants' exceeding 2000 家; CCID Consulting's 'Frontier Technology Forecast and Foresight Research' (released 9/8) says large models' industrialization proximity reaches 1 within 2 years, and inference computing demand is on par with training; five leading institutions at the Bund Conference confirmed 'two layers' (bottom-layer computing chips + upper-layer Agent/new entry points) investment stratification.
Take · The 'deceleration consensus' is in head-on collision, with the Ban ASI Act agenda battle intensifying around 9/16; on the SME side, entry points for computing power and scenario subsidies are opening, and the competitive focus is shifting from single-point algorithms to 'algorithm—system—hardware' engineering collaboration.

Iluvatar CoreX's GPU supply to ByteDance doubles to 100,000 units (on hold, pending official announcement).

天数智芯 / 字节知情人士口径(暂缓)Confidence MidOn hold
Multiple channels relayed it (original Reuters June figure: 50,000 units), citing people familiar with the matter; no original Reuters text seen.
Take · Signal that domestic GPU shortage is shifting from “price hikes” to “scrambling for supply and locking in volume”; confidence to be upgraded only after official announcement.
04RESEARCH3 stories

GPT-6 Astra conquers FrontierMath Tier 4: entire benchmark taken by AI.

Epoch AI / GPT-6 Astra量子位首发 / 多源印证Confidence HighBenchmarks
Solved all remaining problems at that tier, scoring 97.6% (previous high only about 5%); Epoch AI declares Tier 4 saturated; combined with earlier Erdős level 2/68 context, the entire FrontierMath has been fully conquered by AI.
Take · Research-grade benchmarks have lost discriminative power, accelerating the generational shift in evaluation systems—private/vertical leaderboards are becoming the battleground for the old-to-new standard switch, while independent, reproducible evaluation shows a structural gap.

WPS AI Spreadsheet Agent tops SpreadsheetBench V2 Full: 50.86%

WPS AI(金山办公)中媒一致转引Confidence MidNewcomer
50.86% ranks first, beating Claude Opus 4.6 (34.89%), GPT-5.2 (26.79%), and Gemini 3.1 Pro (23.68%); third international No. 1 in three months (original leaderboard page not independently reviewed).
Take · Vertical office scenarios become a surprise winning track for domestic products; citations must carry the caveat 'Chinese media relay, original leaderboard page not reviewed'.

Evaluation trust crisis: test creators double as test sellers

SemiAnalysis / DatacurveSemiAnalysis 调查Confidence MidAcademia · Risk
SemiAnalysis investigation: DeepSWE question setter Datacurve also sells expert-level coding training data to frontier labs (named allegations against specific vendors not independently verified; hold off citing); earlier audit standard: the benchmark has 8.5% false positives / 24% false negatives.
Take · On top of 9/13 Fields Medal winners co-signing, the 'who sets the questions' trust crisis has spread from the math community to engineering leaderboards—'reproducible + conflict-of-interest-free' independent evaluation verification has become a structural gap.
05INSIGHTS5 stories

Today's main line: price comparison logic shifts from 'unit price' to 'total cost per task'

研判深度版·喻言Confidence MidMain line
Real-world tests of switching to DeepSeek prove that 'cheap must be calculated on the total account' (62% more output per task eats into part of the price-cut dividend); on the same day, GPT-6 Astra conquers FrontierMath Tier 4, making research-grade benchmarks lose their discriminating power, and the evaluation trust crisis spreads from the math community to engineering leaderboards; on the capital side, Zhipu's 50 hundred million USD financing lands in parallel with Anthropic racing ahead on the largest IPO in history, while embodied mass production (ten-thousand-unit-scale factories + shipments +272%) enters a realization window.
Take · Narrative focus shifts from 'model arms race' to three parallel tracks: cost, evaluation, and deployment; risk grading shows gray rhinos concentrated in interest rates and evaluation credibility, while black swans are the 'deceleration consensus' becoming policy and IPO pricing. Deep-dive confidence index: High.

Action ①: Switch to a “per-task total cost” billing framework (within a 2-week window)

行动深度版·尚吉(行动①)Confidence MidProcurement side
Use V4.1 Flash for POC, publicly compare the real TCO of “total tokens × unit price” (actual measured per-task output is 62% higher, a ready-made argument); path: enterprise CTO → procurement review committee → written into tender terms; bundle token voucher subsidy application (30-50%) + agent identity national standard pre-compliance, one proposal, double benefits.
Take · Shortest time window, fastest impact: the redefinition of the price-comparison framework is only valid for a few weeks after the version switch lands.

Action ②: Seize the “third-party evaluation verification” gap (time window before the 9/30 ARC-AGI-3 leaderboard cutoff)

行动深度版·尚吉(行动②)Confidence MidEntrepreneurs
Build “reproducible + conflict-of-interest-free” independent evaluation services to meet procurement decision needs after benchmarks lose their anchor (Datacurve selling questions while setting them is a ready-made cautionary tale); path: open-source community reputation → enterprise due-diligence orders → eligibility to co-draft national standards; bundle with three data property rights registration agencies to form an “evaluation + rights confirmation” service package.
Take · Highest barriers, strongest compounding; the window is fixed between two milestones: the ARC-AGI-3 leaderboard cutoff and the initiation of the agent identity national standard project.

Action ③: Bet on Agent governance and identity infrastructure (time window 9/29-10/1)

行动深度版·尚吉(行动③)Confidence MidPractitioners
Build Agent identity/audit middleware around OpenAI DevDay releases + Shanghai Public Security's 1212.5 万 tender opening; path: government pilot tender lot → finance and other strongly regulated industries → ride the '1 万家 SMEs' policy for scale; bundle a PaperCut-type incident postmortem whitepaper for lead generation + Little Giant application coaching.
Take · Most urgent entry point: the national identity standard has been initiated, 'standards are procurement lists,' and policy and tender milestones are squeezed into the same week.

Verification ruling: accepted 14 / corrected 5 / excluded 1 / deferred 3

核验深度版·数见智Confidence MidExcluded/Deferred
Excluded: OpenRouter weekly volume '122.61 trillion' was not independently confirmed (conflicting definitions: 12.1T weekly vs. 46.7T daily coexisting); 'open-weight share topping 50%' is a conflation of definitions (OpenRouter's official figure is about 1/3; what topped 50% is Chinese models' share); SWE-Bench Pro's 'drop of 21.48 points' is suspected of splicing old audits to create new figures (actual audit figures: false positives 8.5% / false negatives 24%). Deferred: 天数智芯 100,000 chips, DeepSWE naming specific vendors, domestic share 40%/Epoch $1 billion (single-source/not verified).
Take · Citing the above items must include the definition/methodology; Hassett's 'no reason to raise rates' quote was not independently verified and has been excluded; Ruchir's '4O' bubble framework is a rehash of old views and serves only as background for a rate-sensitive period.
Official Events1 stories
Ecosystem Pulseself-updating
News392Skills72OpenHub44Reddit25
2026-09-02 → 2026-09-10 Star growth:obra/superpowers +90,630★ · mattpocock/skills +82,178★ · multica-ai/andrej-karpathy-skills +67,613★ · anthropics/skills +56,067★ · Shubhamsaboo/awesome-llm-apps +43,732★
Risk Levelsgray rhino / black swan
Gray rhino · foreseeable
  • Gray Rhino I: FOMC rate hike lands, triggering a second AI valuation compression | Probability 87% | Impact: high-valuation AI / compute chain | Watch: 9/16 dot plot + retail sales (deep-dive risk grading)
  • Gray Rhino II: Evaluation trust crisis spreads to leaderboard ecosystem | Probability 60% | Impact: benchmark providers / procurers | Watch: SemiAnalysis follow-up report (deep-dive risk grading)
Black swan · low-probability
  • Black Swan I: Leading labs' 'slowdown consensus' becomes a policy constraint | Probability 15% | Impact area: frontier model supply | Watchpoint: Trump vs. lab rhetoric escalates (deep-dive risk grading)
  • Black Swan II: Anthropic IPO pricing failure drags down primary market | Probability 10% | Impact area: financing environment | Watchpoint: mid-October roadshow subscription multiple (deep-dive risk grading)
Action Items3
1. Procurement side: switch to a 'total cost per task' billing framework

Use V4.1 Flash for a POC; publicly compare real TCO on 'total tokens × unit price' (measured 62% more output per task is ready-made evidence). Path: enterprise CTO → procurement review committee → write into tender terms. Package: token voucher subsidy application (30-50%) + pre-compliance with the national standard for agent identity — one proposal, dual benefits.

window Time window: within 2 weeks
2. Entrepreneurs: seize the 'third-party evaluation verification' gap

Build an independent evaluation service that is 'reproducible + free of conflicts of interest,' taking on procurement decision needs after benchmark de-anchoring (Datacurve selling questions while also setting them is a ready-made counterexample). Path: open-source community credibility → enterprise due diligence orders → national standard co-drafting qualification. Package: partner with three data property rights registration institutions to form an 'evaluation + rights confirmation' service package.

window Time window: 9/30, before ARC-AGI-3 leaderboard cutoff
3. Practitioners: Betting on Agent governance and identity infrastructure

Build agent identity/audit middleware around DevDay launch + Shanghai police 1212.5万 tender opening. Path: government pilot lots → finance and other highly regulated industries → ride '1万 SMEs' policy for scale. Package: PaperCut-type incident retrospective whitepaper for lead gen + Little Giant application coaching.

window Time window: 9/29-10/1
Next 3 MonthsProb.
WindowEventProb.Basis
9/15-16FOMC rate decision + August retail sales due same day87% rate hike↓ High valuation (CME pricing + deep-dive timeline)
9/16 前后ASI Act-related agenda fight (Ban ASI Act, 9/3 preview)85%↑ Governance (policy agenda)
9/22 前后DeepSeek V4.1 ecosystem adaptation wave90%↑ Open source (extrapolated from adaptation progress)
9/29OpenAI DevDay + Shanghai Public Security agent platform bid opening90%↑ Agent (announced + imminent)
9/30California bill signing window + ARC-AGI-3 leaderboard cutoff80%Two-way (statutory window)
10/1OpenAI vs Apple hearing90%↓ Closed ecosystem (hearing schedule)
10 月上旬Zhipu's 5 billion funding arrives, volume ramps75%↑ Domestic (announcement cadence extrapolated)
10 月中旬Anthropic IPO roadshow / pricing70%↑ Cash-flow AI (Reuters-confirmed window)
10 月下旬HBM shortage second-stage transmission to domestic cards65%↑ Compute costs (supply-chain transmission)
11 月初First batch of token voucher subsidies issued70%↑ SMEs (Chengdu draft for comments)
11 月中旬Humanoid robot Q3 shipment verification (whether +272% continues)60%Verification (Q3 data disclosure window)
12 月Agent identity national standard (draft for comments)55%↑ Standard (routine cadence after project approval)
Sources:In-Depth Daily Report (Daily Core AI Information):Issue 16 · 20.3 KB · AI Official Video Monitoring Log:1 event · 1 with original link · Coverage log (7-day rolling duplicate check):14 topic records · Site ecosystem data:news 392 · skills 72 · openhub 44 · reddit 25 · Official Event AI Picks:1 item (with original link)
SkillHub.news — Global Skill rankings