报告形态 · 五维分类 + 风险 + 行动 + 时间线 ← 回站点
VOL.2026-09-12 · 31 STORIES · SKILLHUB DAILY
SkillHub.news 每日AI
2026 年 9 月 12 日 星期六 · 第 4 期 · 每早 08:00 更新 · 素材 5
今日看点21 条 · 约12分钟
01
模型发布/更新MODEL RELEASES3 条
DeepSeek 9/14 旗舰退役式切换:V4 Pro 全量路由至 V4.1 Flash
02
产品发布/更新PRODUCT3 条
OpenAI 暂停 $200/月 Pro 20X 新订阅与升级
03
行业动态INDUSTRY8 条
AI 代理首次被规模化武器化:PaperCut 漏洞攻破 48 国 395 家组织
04
论文研究RESEARCH2 条
评测范式转向 harness:HarnessDev 基准 + 方升-Code 2.0
05
技巧与观点INSIGHTS5 条
今日主线:AI 代理被武器化之日,国家信用入场之时
01模型发布/更新MODEL RELEASES3 条

DeepSeek 9/14 旗舰退役式切换:V4 Pro 全量路由至 V4.1 Flash

DeepSeek官方公告+行情置信度 高进展·第2天
官方 API 日志证实 9/14 起 deepseek-v4-pro 全部路由至 V4.1 Flash 并按 Flash 计价(官方模型卡 Terminal-Bench 4.0 = 31.2);官方公告披露新架构 KV cache 对 HBM 需求仅为旧版 1/4、SSD 为 1/8。
研判 · 存储链应声走弱(SK 海力士美股 -5.2%、美光 -4.9%、首尔周五再跌 3%+)——算法效率首次反向冲击「HBM 荒涨价」叙事;注意评测口径分裂,AA 独立口径明显低于官方模型卡,引用必带 harness。

ElevenLabs 发布 Music v2.5

ElevenLabs官方视频置信度 高
ElevenLabs 官方视频发布 Music v2.5(Introducing Music v2.5)。
研判 · 音乐生成模型迭代提速,音频生成成为官方渠道的稳定发布面。
原文 ↗

GPT-Live-1 正式进入 API

OpenAI官方视频置信度 高
OpenAI 官方视频宣布 GPT-Live-1 已进入 API(GPT-Live-1 is now in the API)。
研判 · 语音能力从产品内功能走向可编程接口,语音 Agent 的调用链可被独立集成与计价。
原文 ↗
02产品发布/更新PRODUCT3 条

OpenAI 暂停 $200/月 Pro 20X 新订阅与升级

OpenAIX 官宣 / PCMag / 多家外媒置信度 高
核心产品负责人 Sottiaux 9/10 在 X 官宣:因 GPT-6 Astra 需求「前所未见」,暂停 $200/月 ChatGPT Pro 20X 新订阅与升级(存量不受影响)。深度版另注:第三方单源估算 Plus/Pro 5X 使用率超约 11.4% 即入负毛利区间(置信低,仅作线索)。
研判 · 旗舰订阅档被自家需求挤爆,与 Token 账本(中位研究员日推理费 >600 美元)互证——算力瓶颈从传闻变实锤,重度依赖该档位的工作流需备降级预案。

蚂蚁定调「智能体商业时刻」:支付宝设每年 1000 万元「智能体涌现奖」

蚂蚁 / 支付宝第一财经 / 新华财经置信度 高
蚂蚁 CEO 韩歆毅 9/11 外滩大会宣布「智能体商业的 ChatGPT 时刻已到来」,点名瓶颈「用户想用,但商家没有好的智能体服务」;支付宝设每年 1000 万元「智能体涌现奖」(单项最高 100 万)。
研判 · 供给端「Agent 可调用改造」(AEO)成为明确空白带——把商家服务改成 Agent 能理解、能调用,是眼下的 SEO 级机会。

小度家庭 Agent 产品线连发:从「我的搭子」到「我们家的搭子」

百度智能云官方视频(B站)置信度 高
百度智能云 9/12 连发官方视频:Agent 走进真实生活(从「我的搭子」到「我们家的搭子」)、小度添添闺蜜机 Ultra(移动随心、AI 随行)、小度智能音箱 Pro Max(有情绪、会感知),并预告 9/16 2026 智能经济论坛。
研判 · 家庭场景 Agent 从单点硬件走向「移动屏+音箱+陪伴」产品矩阵;9/16 论坛是产业智能体的下一个观测点。
原文 ↗
03行业动态INDUSTRY8 条

AI 代理首次被规模化武器化:PaperCut 漏洞攻破 48 国 395 家组织

GreyNoise / Blackpoint双源安全厂商置信度 高
双源证实:攻击者利用 PaperCut 两个漏洞(CVE-2026-81578/82078),让基于 OpenAI Codex(harness)与 DeepSeek(模型)的 AI 代理自动完成漏洞开发与执行,数小时内攻破 48 国 395 家组织,美国一高中 7 分钟拿到域管权限(440 个实例中仅 12 个全域管,引用须带口径)。
研判 · 「代理被武器化」从推演变成已发生事件:Agent 的 API 密钥与自动化权限就是新攻击面,安全审计与权限管控成下一刚需赛道。

五角大楼拟向云初创 Fluidstack 提供约 50 亿美元贷款

五角大楼 / Fluidstack华尔街日报 / 路透社置信度 高
WSJ 9/10 首发、路透跟进证实:美国国防部经战略资本办公室(OSC)洽谈向云计算初创 Fluidstack 提供约 50 亿美元贷款,用于数据中心零部件供应链(而非新建),尚在谈判未定案。
研判 · 若达成将是美政府对私营 AI 基建最直接的融资动作——AI 基建从「风投逻辑」进入「国家信贷逻辑」,与 DeepSeek 架构级降需求形成一镜像。

月之暗面 8 月 ARR 达 10 亿美元,年底冲 20 亿

月之暗面Bloomberg置信度 高进展·第1天
彭博 9/11:受 7 月 Kimi K3 发布带动,月之暗面 8 月年化经常性收入达 10 亿美元(6 月 3 亿、5 周增 2.3 倍以上),计划年底冲 20 亿美元;正按 500 亿美元投前估值推进融资、或年内赴港上市。
研判 · 为 9/11 已报「A+H 两地上市探索」补上收入端验证;P/ARR 仍是开源阵营最高一档,下一验证点是月度环比不失速。

两大 IPO 同框:Anthropic 前移 10 月 + OpenAI 秘密递表

Anthropic / OpenAI路透 / FT 转述多源置信度 中高
路透 9/12 证实:Anthropic IPO 窗口前移至 10 月中旬路演、中期选举前完成,募资至多 1000 亿、估值约 2 万亿,英伟达拟 100 亿基石投资;同日多源证实 OpenAI 已于 6/8 秘密递交 S-1。
研判 · 安全叙事与资本叙事首次同框对冲,Q4 全球 AI 估值锚重定;招股书里的成本/安全披露比发布会更值得逐页读(OpenAI 递表时间以官方披露为准)。

银行智能体底座首单落定:6489.61 万元中标

浙商银行 / 北京可利邦财联社 / 新华财经置信度 高进展·第4天
9/8 报道的浙商银行智能体集采落定:「智能体基础软硬件」由北京可利邦 6489.61 万元中标(年内金融机构智能体单笔新高),叠加蚂蚁系 946 万财富/供应链订单合计约 7436 万。
研判 · 银行采购从「买模型」切换到「买整套底座」(GPU 算力池+知识工程+模型调度+工具调用),需求主力为城农商行——金融 Agent 招标密集期开启。

8 月 CPI 落地:软件分项 +25.4%,AI 成本首次写进通胀口径

市场(BLS / CME)BLS / CME FedWatch / 财经媒体置信度 高进展·第2天
美国 8 月 CPI 9/11 落地:整体 3.4% 符合预期、核心环比 +0.3% 超预期(4 月来最大单月涨幅);9/16 FOMC 加息 25bp 概率尾盘升至约 87%-90%(CME 不同时点快照);CPI 分项「消费软件及配件」同比 +25.4% 创有记录新高。
研判 · AI 成本正式进入通胀口径,AI 股利率敏感窗口收紧;美股当日「利空落地」不跌反涨、终结四连跌。

工信部「人工智能+软件」专项行动:2028 覆盖 2 万家规上软件企业

工信部工信部发布 / 新华社置信度 高
《「人工智能+软件」专项行动实施方案》(信发〔2026〕209 号,9/2 印发、9/11 公开发布):到 2028 年覆盖 2 万家规上软件企业、100 项技改、100 个智能体标杆、5 个开源项目,配套算力券;首次系统部署 AI 生成代码安全审查与智能体身份标识/可信互联/行为管控。
研判 · 合规触角伸进「AI 写代码」的供应链:代码安全审查与智能体身份标识成为下一个标书必答题,也是新付费品类(注意印发与公开发布的日期差)。

加州禁 AI 陪伴玩具 + ASI 法案亮罚则

加州 / 美国参议院加州官方 / Sanders fact sheet置信度 高
加州签署 SB 867:全美首例禁止 16 岁以下儿童玩具搭载 AI 陪伴聊天功能(为期 5 年、无披露/认证豁免路径,功能即违规),波及玩具商-AI 供应商-零售商全链条;同期 Sanders/Casar 公布《禁止人工超级智能法案》框架:企业解散式「死刑」+个人至多 20 年刑期(9/3 提案、9/16 两党吹风,仍是提案非生效法律)。
研判 · 「软件行为」首次被按实体产品安全监管;做儿童/陪伴类 AI 产品需立即自查加州口径——功能即违规,没有「加了披露就没事」的后门。
04论文研究RESEARCH2 条

评测范式转向 harness:HarnessDev 基准 + 方升-Code 2.0

字节 Seed / 信通院encorp.ai 转述 / 腾讯新闻置信度 中高
字节 Seed 联合多机构发布 HarnessDev 基准:同一权重换执行环境,分差近 15 点(GPT-5 在 Terminus 2 下 35.2%、换 Codex CLI 升至 49.6%);信通院发布「方升-Code 2.0」(5000+ 真实项目数据,新增工程质量/稳定性维度)。
研判 · 评测对象从「答案」转向「可运行系统」——harness 本身成了考题,引用模型分数必须带执行环境。

Cohere 分享 LeVJEPA:无启发式的高效可扩展视频预训练

Cohere官方视频置信度 中
Cohere 官方分享(Lukas Kuhn):LeVJEPA——Efficient & Scalable Video Pretraining without the Heuristics。
研判 · 视频预训练「少假设、可扩展」路线的一条官方技术信号。
原文 ↗
05技巧与观点INSIGHTS5 条

今日主线:AI 代理被武器化之日,国家信用入场之时

前瞻深度版研判置信度 中高
PaperCut 代理武器化(已发生)× 五角大楼 50 亿美元贷款(国家信贷逻辑)× 蚂蚁「智能体商业时刻」三线同日:主题从能力叙事转向安全、资本与合规三条硬约束。
研判 · 深度版灰度判断:确定性叙事为「算力供需缺口与涨价、DeepSeek 架构级降本、Agent 安全监管收口(国内外同步)」;IPO 时点与 V4.1 真实能力上限仍待验证。

抢占 AEO:让企业成为「智能体可调用」的服务商

行动深度版建议置信度 中高
从有线上 API 的商家(OTA/金融产品/本地生活)切入,做接口标准化、结构化数据、Agent 信任凭证三件套;路径:免费诊断 3-5 家「Agent 可达性」→交付改造沉淀 SOP 与工具链→平台化认证服务,对标早期 SEO。
研判 · 蚂蚁定调即官方吹风;关闭信号=大厂推出官方 AEO 标准。AEO 天然适配按调用/成交计费,与「按效计费」建议协同。

做银行智能体底座的「工具调用+知识工程」分包商

行动深度版建议置信度 中高
以浙商银行 6489.61 万中标为标书模板,切底座中最缺人的模块:行内知识库 RAG 治理、工具调用网关、智能体调度监控;路径:攻 10 家股份制/城商行→2-3 单后产品化工具网关→嵌入工信部「AI+软件」专项申报拿算力券。
研判 · 银行从买模型转买底座刚开始,2026-2027 招标密集期;关闭信号=头部大行被四大集成商垄断固化。把 Agent 安全审计(PaperCut 教训)打包进每单交付——安全是银行必答题。

押注「按效计费」基础设施:工作量计量与结果归因工具

行动深度版建议置信度 中高
做智能体时代的「电表」——API 调用计量、结果归因、多方分账结算中间件;路径:单场景 MVP(Token 分成结算)→接入 1-2 个智能体平台验证→争取成为行业计量标准起草方。
研判 · 上海高金×蚂蚁研究院报告系统性背书「按效计费」,工作量计量/结果归因/多方结算工具需求变硬;它是 AEO 与银行底座两项的公共底座。

核验口径提示:今日核验 15 条,采信 6 / 修正 5 / 暂缓 2 / 降级 2

提示深度版核验置信度 中高
要点:加息概率统一表述「约 87%-90%」(CME 不同时点快照);工信部方案系 9/2 印发、9/11 公开发布,非 9/11 印发;蚂蚁为每年 1000 万元「智能体涌现奖」基金(非「1000 万奖」);《禁止 ASI 法案》罚则出自 9/3 提案;DeepSeek 评测以官方模型卡为准(AA 39.55 / LiveBench 77.3 未独立核到,不采信)。
研判 · 降级/剔除:算力大会「需求 +417% vs 供给 +128%」系 Q1 旧数据重新引用,降为背景;Discovery Loop 估值 500 亿仅 Business Insider 单源(要价非成交)暂缓;千问虹膜支付眼镜为「或将成为」宣传口径降级;OpenAI 断供 Cursor(9/1 已报)无增量不重复报道。
生态脉搏站点数据自更新
资讯299Skills72OpenHub44Reddit25
2026-09-02 → 2026-09-10 星标涨幅:obra/superpowers +90,630★ · mattpocock/skills +82,178★ · multica-ai/andrej-karpathy-skills +67,613★ · anthropics/skills +56,067★ · Shubhamsaboo/awesome-llm-apps +43,732★
风险分级灰犀牛 / 黑天鹅
灰犀牛 · 可预见
  • 灰犀牛Ⅰ FOMC 紧缩超预期,AI 估值戴维斯双杀(9/16 加息 25bp 概率约 90%)|高|观察 10Y 能否站稳 5%、甲骨文 FCF 后续
  • 灰犀牛Ⅱ Agent 安全事件常态化,触发紧急监管(PaperCut 型)|高(75%)|观察同类事件频率、国内细则落地节奏
  • 灰犀牛Ⅲ 算力缺口→价格螺旋,中小 AI 公司成本爆仓|中高(70%)|观察租赁报价、SaaS 三季报毛利率
黑天鹅 · 低概率大冲击
  • 黑天鹅Ⅰ IPO 窗口拥挤 + 信贷收紧叠加,一级估值跳水(500 亿传闻类)|低(<15%)|观察 Anthropic/OpenAI 招股书披露质量
  • 黑天鹅Ⅱ《禁止 ASI 法案》进入实体立法,全球能力竞赛降温|低(<10%)|观察 9/16 吹风会后参议院联署数
行动建议3
1. 抢占 AEO:让企业成为「智能体可调用」的服务商

从有线上 API 的商家(OTA/金融产品/本地生活)切入,做接口标准化、结构化数据、Agent 信任凭证三件套;先免费诊断 3-5 家「Agent 可达性」,交付改造沉淀 SOP 与工具链,再做平台化认证服务。

时间窗 蚂蚁定调即官方吹风起;关闭信号=大厂推出官方 AEO 标准
2. 做银行智能体底座的「工具调用+知识工程」分包商

以浙商银行 6489.61 万中标为标书模板,切行内知识库 RAG 治理、工具调用网关、智能体调度监控三个最缺人的模块;攻 10 家股份制/城商行后在 2-3 单基础上产品化,并嵌入工信部「AI+软件」专项申报拿算力券。

时间窗 2026-2027 银行招标密集期;关闭信号=头部大行被四大集成商垄断固化
3. 押注「按效计费」基础设施:工作量计量与结果归因工具

做智能体时代的「电表」——API 调用计量、结果归因、多方分账结算中间件;先做单场景 MVP(Token 分成结算),接入 1-2 个智能体平台验证,再争取成为行业计量标准起草方。

时间窗 计量工具先行者稀缺;关闭信号=云大厂内置结算能力免费开放
未来 3 个月时间线概率
时间窗事件概率依据
9/14DeepSeek 旗舰切换:V4 Pro 全量路由至 V4.1 Flash约 99%官方 API 日志与公告
9/15-16FOMC 议息:加息 25bp 落地,看点阵图对 2027 指引约 90%CME FedWatch 定价
9/16《禁止 ASI 法案》两党吹风(大概率仅框架宣示)80%Sanders/Casar 公布节奏
9/29上海公安智能体平台开标(1212.5 万标段)金融/政务 Agent 招投标是商业化最硬指标
9/30加州立法窗口关闭:SB 867 或带动 2-3 个同类州提案60%法定窗口
10月中旬Anthropic IPO 路演窗口(中期选举前完成)50%路透口径,前移系传闻级
10月底SaaS 三季报验证「AI 成本→毛利率」传导财报季,看甲骨文模式是否扩散
11/12OpenAI 断供 Cursor 生效观察国产模型承接度,开发者工具格局分水岭
11月外滩大会后首季:智能体商业化数据首次可同比蚂蚁「智能体商业时刻」定调后的首个验证窗
素材来源:深度版日报(每日AI核心信息):第 14 期 · 20.5 KB · AI 官方视频监控台账:14 条事件 · 14 条带原文链接 · 报道台账(7 天滚动查重):13 条主题记录 · 站点生态数据:news 299 · skills 72 · openhub 44 · reddit 25 · 官方事件 AI 精选:10 条(带原文链接)
SkillHub.news — 全球 Skill 榜单
VOL.2026-09-12 · 31 STORIES · SKILLHUB DAILY
SkillHub.news Daily AI
Sep 12, 2026 · #4 · updated 08:00 daily · sources 5
Highlights21 stories · approx.12min read
01
MODEL RELEASES3 stories
DeepSeek 9/14 flagship retirement-style switch: V4 Pro fully routed to V4.1 Flash
02
PRODUCT3 stories
OpenAI pauses new subscriptions and upgrades for $200/month Pro 20X
03
INDUSTRY8 stories
AI agents weaponized at scale for the first time: PaperCut vulnerabilities breached 395 organizations in 48 countries
04
RESEARCH2 stories
Evaluation paradigm shifts to harness: HarnessDev benchmark + 方升-Code 2.0
05
INSIGHTS5 stories
Today’s main thread: the day AI agents are weaponized is when state credit enters.
01MODEL RELEASES3 stories

DeepSeek 9/14 flagship retirement-style switch: V4 Pro fully routed to V4.1 Flash

DeepSeek官方公告+行情Confidence HighProgress · Day 2
Official API logs confirm that from 9/14, deepseek-v4-pro is fully routed to V4.1 Flash and priced at Flash rates (official model card Terminal-Bench 4.0 = 31.2); the official announcement discloses the new architecture's KV cache has HBM demand only 1/4 of the old version, and SSD 1/8.
Take · Memory chain weakens in response (SK Hynix US shares -5.2%, Micron -4.9%, Seoul down another 3%+ on Friday) — algorithmic efficiency for the first time strikes back at the HBM shortage price-hike narrative; note the split in evaluation methodology, AA's independent methodology is clearly lower than the official model card, citations must include harness.

ElevenLabs releases Music v2.5

ElevenLabs官方视频Confidence High
ElevenLabs official video releases Music v2.5 (Introducing Music v2.5).
Take · Music generation model iteration accelerates; audio generation becomes a stable release surface on official channels.
Original ↗

GPT-Live-1 officially enters the API

OpenAI官方视频Confidence High
OpenAI announced in an official video that GPT-Live-1 is now in the API.
Take · Voice capabilities are moving from an in-product feature to a programmable interface, allowing voice Agent call chains to be independently integrated and priced.
Original ↗
02PRODUCT3 stories

OpenAI pauses new subscriptions and upgrades for $200/month Pro 20X

OpenAIX 官宣 / PCMag / 多家外媒Confidence High
Core product lead Sottiaux announced on X on 9/10: due to 'unprecedented' demand for GPT-6 Astra, new subscriptions and upgrades for $200/month ChatGPT Pro 20X are paused (existing subscriptions unaffected). Deep version note: a third-party single-source estimate suggests Plus/Pro 5X enters negative gross margin territory once usage exceeds about 11.4% (low confidence, for lead purposes only).
Take · The flagship subscription tier is being overwhelmed by its own demand, corroborated by the Token ledger (median researcher daily inference cost >$600)—the compute bottleneck has moved from rumor to hard fact, and workflows heavily dependent on this tier need downgrade contingency plans.

Ant Group sets the tone for the 'agent commerce moment': Alipay establishes an annual 10 million yuan 'Agent Emergence Award'.

蚂蚁 / 支付宝第一财经 / 新华财经Confidence High
Ant Group CEO Han Xinyi announced at the Bund Summit on 9/11 that 'the ChatGPT moment for agent commerce has arrived,' flagging the bottleneck: 'Users want to use it, but merchants don't have good agent services'; Alipay is establishing an annual 10 million yuan 'Agent Emergence Award' (single award up to 1 million yuan).
Take · Supply-side 'Agent-callable transformation' (AEO) has become a clear white space—retrofitting merchant services so Agents can understand and call them is the SEO-level opportunity right now.

Xiaodu family Agent product line launches a series: from 'My Companion' to 'Our Family Companion'.

百度智能云官方视频(B站)Confidence High
Baidu AI Cloud released a series of official videos on 9/12: Agent enters real life (from 'My Companion' to 'Our Family Companion'), Xiaodu Tiantian Guimi Machine Ultra (move freely, AI on the go), Xiaodu Smart Speaker Pro Max (emotional, perceptive), and previewed the 9/16 2026 Smart Economy Forum.
Take · Home-scenario Agents are moving from single-point hardware to a “mobile screen + speaker + companionship” product matrix; the 9/16 forum is the next watchpoint for industry agents.
Original ↗
03INDUSTRY8 stories

AI agents weaponized at scale for the first time: PaperCut vulnerabilities breached 395 organizations in 48 countries

GreyNoise / Blackpoint双源安全厂商Confidence High
Two sources confirm: attackers exploited two PaperCut vulnerabilities (CVE-2026-81578/82078), letting AI agents based on OpenAI Codex (harness) and DeepSeek (model) automatically complete vulnerability development and execution, breaching 395 organizations in 48 countries within hours; a U.S. high school gained domain admin access in 7 minutes (only 12 of 440 instances had full domain admin; citations must include the stated scope).
Take · 'Agent weaponization' has moved from simulation to real incidents: Agents' API keys and automation permissions are the new attack surface; security audits and permission controls are becoming the next must-have category.

Pentagon plans to provide about a $5 billion loan to cloud startup Fluidstack

五角大楼 / Fluidstack华尔街日报 / 路透社Confidence High
WSJ first reported on 9/10, Reuters followed and confirmed: The U.S. Department of Defense, through the Office of Strategic Capital (OSC), is in talks to provide about a $5 billion loan to cloud computing startup Fluidstack for data center component supply chain (not new construction); still under negotiation and not finalized.
Take · If completed, it would be the U.S. government's most direct financing move for private AI infrastructure—AI infrastructure moving from 'VC logic' to 'state credit logic,' forming a mirror image of DeepSeek's architecture-level demand reduction.

Moonshot AI's August ARR reached $1 billion, targeting $2 billion by year-end

月之暗面BloombergConfidence HighProgress · Day 1
Bloomberg 9/11: Driven by the July launch of Kimi K3, Moonshot AI's August annualized recurring revenue reached $1 billion (June $300 million, up more than 2.3x in 5 weeks), with plans to reach $2 billion by year-end; it is advancing financing at a $50 billion pre-money valuation and may list in Hong Kong within the year.
Take · Adds revenue-side validation to the 'A+H dual-listing exploration' reported on 9/11; P/ARR remains the highest tier in the open-source camp, and the next checkpoint is whether month-over-month growth holds up.

Two major IPOs together: Anthropic moves up to October; OpenAI files confidentially.

Anthropic / OpenAI路透 / FT 转述多源Confidence Mid
Reuters confirmed on 9/12: Anthropic's IPO window moves up to a mid-October roadshow and completion before the midterm elections, raising up to 1000 亿 and valued at about 2 万亿, with NVIDIA planning a 100 亿 cornerstone investment; on the same day, multiple sources confirmed OpenAI confidentially filed its S-1 on 6/8.
Take · For the first time, the safety narrative and capital narrative are juxtaposed and offset each other, resetting the global AI valuation anchor in Q4; the cost/safety disclosures in the prospectus are more worth reading page by page than the launch event (OpenAI's filing date is subject to official disclosure).

First bank agent foundation order settled: 6489.61 万元 winning bid

浙商银行 / 北京可利邦财联社 / 新华财经Confidence HighProgress · Day 4
The Zheshang Bank agent centralized procurement reported on 9/8 is settled: 'agent foundational software and hardware' won by Beijing Kelibang at 6489.61 万元 (a new single-order high for financial institution agents this year), plus Ant-affiliated 946 万 wealth/supply chain orders totaling about 7436 万.
Take · Bank procurement shifts from 'buying models' to 'buying the full foundation' (GPU compute pool + knowledge engineering + model scheduling + tool calling), with demand mainly from urban and rural commercial banks—a dense period for financial Agent tenders begins.

August CPI released: software component +25.4%, AI costs enter inflation gauge for the first time

市场(BLS / CME)BLS / CME FedWatch / 财经媒体Confidence HighProgress · Day 2
US August CPI released on 9/11: headline 3.4% in line with expectations, core MoM +0.3% above expectations (largest monthly gain since April); probability of a 25bp FOMC rate hike on 9/16 rose late in session to about 87%-90% (CME snapshots at different times); CPI component 'consumer software and accessories' +25.4% YoY, a record high.
Take · AI costs formally enter the inflation gauge, tightening the rate-sensitive window for AI stocks; US stocks rose on the day as the 'bad news landed', ending a four-day losing streak.

MIIT 'AI+Software' special action: cover 20,000 above-scale software enterprises by 2028

工信部工信部发布 / 新华社Confidence High
Implementation Plan for the 'AI+Software' Special Action (信发〔2026〕209 号, issued 9/2, publicly released 9/11): by 2028, cover 20,000 above-scale software enterprises, 100 technical upgrades, 100 agent benchmarks, 5 open-source projects, with supporting computing power vouchers; first systematic deployment of security review for AI-generated code and agent identity identification/trusted interconnection/behavior control.
Take · Compliance reach extends into the 'AI coding' supply chain: code security review and agent identity become the next must-answer items in bids, and a new paid category (note the date gap between issuance and public release).

California bans AI companion toys + ASI bill sets penalties

加州 / 美国参议院加州官方 / Sanders fact sheetConfidence High
California signs SB 867: first in the U.S. to ban AI companion chat features in toys for children under 16 (for 5 years, no disclosure/certification exemption path; the feature itself is a violation), affecting the entire toymaker-AI supplier-retailer chain; meanwhile Sanders/Casar unveil the framework for the 'Ban Artificial Superintelligence Act': corporate dissolution-style 'death penalty' + up to 20 years' imprisonment for individuals (proposed 9/3, bipartisan briefing 9/16; still a proposal, not law in effect).
Take · For the first time, 'software behavior' is regulated as physical product safety; makers of children's/companion AI products must immediately self-check against California's stance—the feature itself is a violation, with no 'disclose and you're fine' backdoor.
04RESEARCH2 stories

Evaluation paradigm shifts to harness: HarnessDev benchmark + 方升-Code 2.0

字节 Seed / 信通院encorp.ai 转述 / 腾讯新闻Confidence Mid
ByteDance Seed and multiple institutions release HarnessDev benchmark: same weights, different execution environment, score gap nearly 15 points (GPT-5 scores 35.2% under Terminus 2, rises to 49.6% with Codex CLI); CAICT releases '方升-Code 2.0' (5000+ real project data, adds engineering quality/stability dimensions).
Take · Evaluation targets shift from “answers” to “runnable systems”—the harness itself becomes the test, and citing model scores must include the execution environment.

Cohere shares LeVJEPA: efficient, scalable video pretraining without heuristics.

Cohere官方视频Confidence Mid
Cohere official post (Lukas Kuhn): LeVJEPA—Efficient & Scalable Video Pretraining without the Heuristics.
Take · An official technical signal for the “fewer assumptions, scalable” route in video pretraining.
Original ↗
05INSIGHTS5 stories

Today’s main thread: the day AI agents are weaponized is when state credit enters.

前瞻深度版研判Confidence Mid
PaperCut agent weaponization (already occurred) × Pentagon $5 billion loan (state credit logic) × Ant “agent commerce moment” on the same day across three threads: the theme shifts from capability narrative to three hard constraints—safety, capital, and compliance.
Take · Deep-dive grey-area judgment: the certainty narrative is “compute supply-demand gap and price hikes, DeepSeek architecture-level cost reduction, Agent safety regulation tightening (synchronized domestically and internationally)”; IPO timing and V4.1’s true capability ceiling remain to be verified.

Seize AEO: make enterprises 'agent-callable' service providers

行动深度版建议Confidence Mid
Start with merchants that have online APIs (OTA/financial products/local life), and build the trio of interface standardization, structured data, and Agent trust credentials; path: free 'Agent reachability' diagnostics for 3-5 companies → deliver transformation and accumulate SOPs and toolchain → platform-based certification services, benchmarked against early SEO.
Take · Ant setting the tone is official signaling; shutdown signal = a big tech firm launches an official AEO standard. AEO naturally fits per-call/per-transaction billing and synergizes with the 'performance-based billing' recommendation.

Be a subcontractor for 'tool calling + knowledge engineering' in banks' agent foundation

行动深度版建议Confidence Mid
Use Zheshang Bank's 6489.61 万 winning bid as a tender template, and target the most understaffed modules in the foundation: in-bank knowledge base RAG governance, tool-calling gateway, and agent scheduling/monitoring; path: target 10 joint-stock/city commercial banks → after 2-3 deals, productize the tool gateway → embed into MIIT's 'AI+Software' special application to obtain computing power vouchers.
Take · Banks have just begun shifting from buying models to buying foundations, with 2026-2027 a dense tender period; shutdown signal = leading big banks become monopolized and entrenched by the four major integrators. Bundle Agent security audits (the PaperCut lesson) into every delivery—security is a mandatory question for banks.

Bet on 'performance-based billing' infrastructure: workload metering and result attribution tools

行动深度版建议Confidence Mid
Be the 'electric meter' of the agent era—API call metering, result attribution, and multi-party split settlement middleware; path: single-scenario MVP (Token split settlement) → integrate with 1-2 agent platforms for validation → strive to become a drafter of industry metering standards.
Take · Shanghai Advanced Institute of Finance × Ant Research Institute report systematically endorses “pay-for-performance”; demand hardens for workload metering, outcome attribution, and multi-party settlement tools; it is the shared foundation for both AEO and the banking base.

Verification note: 15 verified today, 6 accepted / 5 corrected / 2 deferred / 2 downgraded

提示深度版核验Confidence Mid
Key points: rate-hike probability uniformly stated as “about 87%-90%” (CME snapshots at different times); the MIIT plan was issued on 9/2 and publicly released on 9/11, not issued on 9/11; Ant has an annual 1000 万元 “Agent Emergence Award” fund (not a “1000 万 award”); penalties in the “Banning ASI Act” come from the 9/3 proposal; DeepSeek evaluations follow the official model card (AA 39.55 / LiveBench 77.3 not independently verified; not accepted).
Take · Downgraded/removed: the Computing Power Conference’s “demand +417% vs supply +128%” was old Q1 data re-cited, downgraded to background; Discovery Loop’s 500 亿 valuation has only Business Insider as a single source (asking price, not a completed deal), deferred; Qwen iris payment glasses downgraded due to “may become” promotional framing; OpenAI cutting off Cursor (reported 9/1) has no increment, not repeated.
Ecosystem Pulseself-updating
News299Skills72OpenHub44Reddit25
2026-09-02 → 2026-09-10 Star growth:obra/superpowers +90,630★ · mattpocock/skills +82,178★ · multica-ai/andrej-karpathy-skills +67,613★ · anthropics/skills +56,067★ · Shubhamsaboo/awesome-llm-apps +43,732★
Risk Levelsgray rhino / black swan
Gray rhino · foreseeable
  • Gray Rhino I: FOMC tightening exceeds expectations, AI valuations suffer Davis double kill (about 90% chance of a 25bp hike on 9/16) | High | Watch whether 10Y can hold 5%, Oracle FCF follow-up
  • Gray Rhino II: Agent security incidents become routine, triggering emergency regulation (PaperCut-type) | High (75%) | Watch frequency of similar incidents, pace of domestic rule implementation
  • Gray Rhino III: Compute gap → price spiral, cost blowups at small and mid-sized AI companies | Medium-high (70%) | Watch rental quotes, SaaS Q3 gross margins
Black swan · low-probability
  • Black Swan I: Crowded IPO window + tighter credit combine, primary-market valuations plunge (500 亿 rumor type) | Low (<15%) | Watch disclosure quality of Anthropic/OpenAI prospectuses
  • Black Swan II: ASI Prohibition Act enters substantive legislation, global capability race cools | Low (<10%) | Watch Senate cosponsor count after 9/16 briefing
Action Items3
1. Capture AEO: make enterprises “agent-callable” service providers

Start with merchants that have online APIs (OTA/financial products/local services) and deliver the three-piece set of interface standardization, structured data, and Agent trust credentials; first provide free “Agent reachability” diagnostics for 3-5 companies, use delivery transformations to consolidate SOPs and a toolchain, then offer platform-based certification services.

window Ant setting the tone marks the start of official signaling; shutdown signal = a big tech company launches an official AEO standard
2. Subcontractor for 'tool calling + knowledge engineering' in the bank agent foundation.

Use Zheshang Bank's 6489.61 万 winning bid as a proposal template; target the three most understaffed modules—in-bank knowledge base RAG governance, tool-calling gateway, and agent scheduling/monitoring; after landing 10 joint-stock/city commercial banks, productize on the basis of 2-3 orders, and embed into MIIT's 'AI+Software' special application to obtain compute vouchers.

window 2026-2027 is a dense period for bank tenders; shutdown signal = leading major banks monopolized and locked in by the Big Four integrators.
3. Bet on 'performance-based billing' infrastructure: workload metering and outcome attribution tools.

Build the 'electricity meter' of the agent era—API call metering, outcome attribution, and multi-party split settlement middleware; start with a single-scenario MVP (Token revenue-share settlement), integrate with 1-2 agent platforms for validation, then seek to become a drafter of industry metering standards.

window Metering tool first movers are scarce; shutdown signal = cloud giants open built-in settlement capabilities for free.
Next 3 MonthsProb.
WindowEventProb.Basis
9/14DeepSeek flagship switch: V4 Pro fully routed to V4.1 Flash.About 99%Official API logs and announcements.
9/15-16FOMC rate decision: 25bp hike delivered; focus on dot plot guidance for 2027About 90%CME FedWatch pricing
9/16Ban ASI Act bipartisan briefing (likely only a framework declaration)80%Sanders/Casar release cadence
9/29Shanghai Public Security agent platform opens bids (1212.5 万 lot)Finance/government Agent bidding is the hardest commercialization metric
9/30California legislative window closes: SB 867 may drive 2-3 similar state proposals60%Statutory window
10月中旬Anthropic IPO roadshow window (completed before midterm elections)50%Reuters framing: earlier timing is rumor-level
10月底SaaS Q3 earnings to validate 'AI cost → gross margin' transmissionEarnings season: watch whether Oracle's model spreads
11/12OpenAI cutoff to Cursor takes effectWatch domestic model uptake; a watershed for the developer tools landscape
11月First quarter after the Bund Conference: agent commercialization data becomes YoY-comparable for the first timeFirst validation window after Ant's 'Agent Commerce Moment' positioning
Sources:In-Depth Daily Report (Daily Core AI Information):Issue 14 · 20.5 KB · AI Official Video Monitoring Log:14 events · 14 with original source links · Coverage log (7-day rolling duplicate check):13 topic records · Site ecosystem data:news 299 · skills 72 · openhub 44 · reddit 25 · Official Event AI Picks:10 items (with source links)
SkillHub.news — Global Skill rankings