01MODEL RELEASES2 stories
DeepSeek switch rollout test: V4.1 Flash intelligence index 39.5 overtakes V4 Pro, but per-task output is 62% higher
DeepSeekAA 独立实测(多源)Confidence HighProgress · Day 4
AA independent test: V4.1 Flash intelligence index 39.5 (peak tier reported 40), above V4 Pro's 36; but average per-task output about 89,000 tokens, 62% more than V4 Pro, partially offsetting the price-cut benefit. Combined factors: MIT open-source weights + KV cache demand only 1/4 of previous generation.
Take · Official 'retain original entry point' (reported 9/13) coexists with observed routing behavior; migration decisions should be based on official documentation—price-comparison logic shifts from 'unit price' to total per-task cost, and the self-hosting cost curve becomes a new variable in enterprise selection.
Cognition releases SWE-2: RL post-trained on Kimi K3 (2.8T parameters)
Cognition官方博客Confidence HighNew release
Official blog confirms RL post-training based on Kimi K3 (2.8T parameters); its own FrontierCode 1.1 Main 50.0% (less than 1 point from Fable 5.1), claimed cost 64% lower.
Take · FrontierCode is a vendor-built leaderboard, and competitor scores use vendor self-testing criteria, denting credibility; overseas coding agents procuring domestic open-source bases is a direct sample of open-source weight spillover.
02PRODUCT2 stories
Eval infrastructure and framework make same-day incremental moves: ageval goes open source, DSPy 3.4.0 Beta 1 released
浙大 ZJU-REAL / DSPy开源发布Confidence MidToolchain
Zhejiang University's ZJU-REAL open-sources ageval: a pluggable Agent evaluation base, with environment plugins covering Docker/E2B/Daytona/Local and one-line config to switch runtimes; DSPy 3.4.0 Beta 1 released, with the LM execution layer switched to a built-in engine.
Take · The ecosystem is shifting from 'stacking models' to 'eval infrastructure + inference economics'; pluggable evaluation runtimes mean harness compatibility becomes a hidden cost item in Agent selection.
Data property rights registration system goes live on 9/11: the three entities in Beijing/Shanghai/Shenzhen first become registration agencies
国家数据局官方 / 多源Confidence HighNewcomer
The National Data Administration system goes live on 9/11; the three data exchanges in Beijing/Shanghai/Shenzhen are first to become registration agencies, issuing data 'identity certificates' and clarifying holding/use/operating rights; commercialization signals for local large-model office work (27B can run on consumer PCs) appear simultaneously (medium confidence).
Take · The foundational infrastructure for data asset balance-sheet inclusion and trading is in place—rights confirmation enters the implementation phase, and the commercialization path for local/private deployment opens accordingly.
03INDUSTRY8 stories
Zhipu completes US$5 billion financing: US$2 billion placement + US$3 billion zero-coupon convertible bonds
智谱公告(9/13)Confidence HighCapital · New Faces
9/13 announcement: US$2 billion placement + US$3 billion zero-coupon convertible bonds; MaaS ARR of US$1.6 billion is the figure already disclosed in the 8/31 half-year report (not new); API gross margin -0.4% → 24.6%, turning positive; Token call volume +40x from start of year.
Take · 'Small equity, large debt' structure shows overseas long-term funds scrambling for the domestic open-source flagship; attracting capital again within two months; capital arrival and volume ramp-up in early October are worth tracking.
Anthropic picks Nasdaq; pitch materials disclose finances for first time: Q2 revenue over US$11.5 billion, first profit
AnthropicBI / 彭博 / 路透Confidence HighProgress · Day 2
Multiple sources confirm Nasdaq as listing exchange; pitch materials disclose Q2 revenue over $11.5 billion, first profit, Q3 “expected” to be profitable again; NVIDIA cornerstone investment of up to $10 billion still under negotiation.
Take · History’s largest IPO (about $2 trillion valuation) enters the roadshow stage, appearing alongside a “slowdown declaration” — capital narrative and safety narrative run in parallel; Altman says OpenAI will not go public in 2026, the two diverge.
UBTECH’s Liuzhou 10,000-unit-class humanoid robot super factory begins production: one unit rolls off the line every 10 minutes
优必选金台资讯 / 科创板日报Confidence HighNewcomer
9/12: production starts—1.4 万㎡, one WalkerS/Cruzr rolls off the line every 10 minutes, annual capacity over 10,000 units, the world's first 10,000-unit-class base (with Siemens digital twin, component localization rate over 75%); also won the Leshan commercial service humanoid robot project worth 1.5081 亿元 (announced 9/8).
Take · Scale-driven cost reduction begins—Unitree's unit price has fallen 72% in two years to 16.64 万元; Chinese manufacturers account for about 97% of global shipments (industry figures); embodied mass production enters a realization window.
Humanoid robots: 2026H1 global shipments 1.91 ten-thousand units, +272% YoY
Smart Analytics Global / 高盛单一机构口径 / 投行预测Confidence MidData
2026H1 global humanoid robot shipments about 1.91 ten-thousand units, +272% YoY (single-institution basis); Goldman Sachs forecasts the average complete-unit price to drop to 2.13 ten-thousand USD in 2035 (-49%).
Take · The mass-production race has officially begun; Q3 shipment data is the next validation point; shipment figures are from a single institution, so citations must include the source.
Double pressure on FOMC eve: rate-hike pricing 86.2% × 'AI slowdown' narrative hits dark pool
CME / 美股暗盘CME 官方 / 行情Confidence HighProgress
CME 9/14 rate-hike probability 86.2% (about 90% after the 9/13 close, about 59% a week earlier); 9/13 dark pool: SK Hynix -4.60%, Intel -3.81%, Micron -3.25%, NVIDIA -2.16% (matching 21 Finance/Xueqiu item by item); 9/16 rate decision + August retail sales due same day.
Take · The 'slowdown' sentiment shock diverges from Oracle CapEx guidance (compute spending has not ebbed); it is short-term sentiment, with medium-term validation pending late-October SaaS Q3 earnings; White House's 'no reason to raise rates' exact quote not independently verified, citation removed.
Governance infrastructure trio: agent identity national standard project + token vouchers + data property rights registration
国家数据局 / 成都 / 外滩大会多源Confidence HighRules · Business Opportunities
The national standardization guiding technical document 'Blockchain Agent Identity Management Requirements' launched its project at the Bund Conference; Chengdu released a draft for comments on token vouchers on 9/11 (annual 1 亿元, max 200 万 per entity, 30-50% subsidy, 50% for micro and small enterprises), while Chongqing's 'token vouchers + token loans' takes the lead; the National Data Administration launched its data property rights registration system on 9/11 (three exchanges in Beijing/Shanghai/Shenzhen).
Take · The governance infrastructure window for the trinity of 'rights confirmation—identity—subsidies' is opening, with 'standards as procurement lists'; gaps: token voucher application/redemption on behalf, aggregated procurement of model calls, voucher adaptation service providers.
Policy and trends: Trump opposes the 'deceleration consensus'; MIIT SME plan; CCID report says inference demand on par with training
特朗普 / 工信部 / 赛迪顾问 / 外滩机构多源Confidence HighRules · Outlook
9/13 Trump in Ireland opposed slowing AI ('whoever wins AI wins'), calling risk warnings 'hype by negative forces' and supporting only guardrails; Sanders hit back the same day and called for China-US negotiations on an AI pause treaty; MIIT's 'AI SME Startup Support Program' (issued 8/29) aims to cultivate over 1 万家 technology and innovative SMEs in three years, with specialized, refined, distinctive and innovative 'little giants' exceeding 2000 家; CCID Consulting's 'Frontier Technology Forecast and Foresight Research' (released 9/8) says large models' industrialization proximity reaches 1 within 2 years, and inference computing demand is on par with training; five leading institutions at the Bund Conference confirmed 'two layers' (bottom-layer computing chips + upper-layer Agent/new entry points) investment stratification.
Take · The 'deceleration consensus' is in head-on collision, with the Ban ASI Act agenda battle intensifying around 9/16; on the SME side, entry points for computing power and scenario subsidies are opening, and the competitive focus is shifting from single-point algorithms to 'algorithm—system—hardware' engineering collaboration.
Iluvatar CoreX's GPU supply to ByteDance doubles to 100,000 units (on hold, pending official announcement).
天数智芯 / 字节知情人士口径(暂缓)Confidence MidOn hold
Multiple channels relayed it (original Reuters June figure: 50,000 units), citing people familiar with the matter; no original Reuters text seen.
Take · Signal that domestic GPU shortage is shifting from “price hikes” to “scrambling for supply and locking in volume”; confidence to be upgraded only after official announcement.
04RESEARCH3 stories
GPT-6 Astra conquers FrontierMath Tier 4: entire benchmark taken by AI.
Epoch AI / GPT-6 Astra量子位首发 / 多源印证Confidence HighBenchmarks
Solved all remaining problems at that tier, scoring 97.6% (previous high only about 5%); Epoch AI declares Tier 4 saturated; combined with earlier Erdős level 2/68 context, the entire FrontierMath has been fully conquered by AI.
Take · Research-grade benchmarks have lost discriminative power, accelerating the generational shift in evaluation systems—private/vertical leaderboards are becoming the battleground for the old-to-new standard switch, while independent, reproducible evaluation shows a structural gap.
WPS AI Spreadsheet Agent tops SpreadsheetBench V2 Full: 50.86%
WPS AI(金山办公)中媒一致转引Confidence MidNewcomer
50.86% ranks first, beating Claude Opus 4.6 (34.89%), GPT-5.2 (26.79%), and Gemini 3.1 Pro (23.68%); third international No. 1 in three months (original leaderboard page not independently reviewed).
Take · Vertical office scenarios become a surprise winning track for domestic products; citations must carry the caveat 'Chinese media relay, original leaderboard page not reviewed'.
Evaluation trust crisis: test creators double as test sellers
SemiAnalysis / DatacurveSemiAnalysis 调查Confidence MidAcademia · Risk
SemiAnalysis investigation: DeepSWE question setter Datacurve also sells expert-level coding training data to frontier labs (named allegations against specific vendors not independently verified; hold off citing); earlier audit standard: the benchmark has 8.5% false positives / 24% false negatives.
Take · On top of 9/13 Fields Medal winners co-signing, the 'who sets the questions' trust crisis has spread from the math community to engineering leaderboards—'reproducible + conflict-of-interest-free' independent evaluation verification has become a structural gap.
05INSIGHTS5 stories
Today's main line: price comparison logic shifts from 'unit price' to 'total cost per task'
研判深度版·喻言Confidence MidMain line
Real-world tests of switching to DeepSeek prove that 'cheap must be calculated on the total account' (62% more output per task eats into part of the price-cut dividend); on the same day, GPT-6 Astra conquers FrontierMath Tier 4, making research-grade benchmarks lose their discriminating power, and the evaluation trust crisis spreads from the math community to engineering leaderboards; on the capital side, Zhipu's 50 hundred million USD financing lands in parallel with Anthropic racing ahead on the largest IPO in history, while embodied mass production (ten-thousand-unit-scale factories + shipments +272%) enters a realization window.
Take · Narrative focus shifts from 'model arms race' to three parallel tracks: cost, evaluation, and deployment; risk grading shows gray rhinos concentrated in interest rates and evaluation credibility, while black swans are the 'deceleration consensus' becoming policy and IPO pricing. Deep-dive confidence index: High.
Action ①: Switch to a “per-task total cost” billing framework (within a 2-week window)
行动深度版·尚吉(行动①)Confidence MidProcurement side
Use V4.1 Flash for POC, publicly compare the real TCO of “total tokens × unit price” (actual measured per-task output is 62% higher, a ready-made argument); path: enterprise CTO → procurement review committee → written into tender terms; bundle token voucher subsidy application (30-50%) + agent identity national standard pre-compliance, one proposal, double benefits.
Take · Shortest time window, fastest impact: the redefinition of the price-comparison framework is only valid for a few weeks after the version switch lands.
Action ②: Seize the “third-party evaluation verification” gap (time window before the 9/30 ARC-AGI-3 leaderboard cutoff)
行动深度版·尚吉(行动②)Confidence MidEntrepreneurs
Build “reproducible + conflict-of-interest-free” independent evaluation services to meet procurement decision needs after benchmarks lose their anchor (Datacurve selling questions while setting them is a ready-made cautionary tale); path: open-source community reputation → enterprise due-diligence orders → eligibility to co-draft national standards; bundle with three data property rights registration agencies to form an “evaluation + rights confirmation” service package.
Take · Highest barriers, strongest compounding; the window is fixed between two milestones: the ARC-AGI-3 leaderboard cutoff and the initiation of the agent identity national standard project.
Action ③: Bet on Agent governance and identity infrastructure (time window 9/29-10/1)
行动深度版·尚吉(行动③)Confidence MidPractitioners
Build Agent identity/audit middleware around OpenAI DevDay releases + Shanghai Public Security's 1212.5 万 tender opening; path: government pilot tender lot → finance and other strongly regulated industries → ride the '1 万家 SMEs' policy for scale; bundle a PaperCut-type incident postmortem whitepaper for lead generation + Little Giant application coaching.
Take · Most urgent entry point: the national identity standard has been initiated, 'standards are procurement lists,' and policy and tender milestones are squeezed into the same week.
Verification ruling: accepted 14 / corrected 5 / excluded 1 / deferred 3
核验深度版·数见智Confidence MidExcluded/Deferred
Excluded: OpenRouter weekly volume '122.61 trillion' was not independently confirmed (conflicting definitions: 12.1T weekly vs. 46.7T daily coexisting); 'open-weight share topping 50%' is a conflation of definitions (OpenRouter's official figure is about 1/3; what topped 50% is Chinese models' share); SWE-Bench Pro's 'drop of 21.48 points' is suspected of splicing old audits to create new figures (actual audit figures: false positives 8.5% / false negatives 24%). Deferred: 天数智芯 100,000 chips, DeepSWE naming specific vendors, domestic share 40%/Epoch $1 billion (single-source/not verified).
Take · Citing the above items must include the definition/methodology; Hassett's 'no reason to raise rates' quote was not independently verified and has been excluded; Ruchir's '4O' bubble framework is a rehash of old views and serves only as background for a rate-sensitive period.
Ecosystem Pulseself-updating
News392Skills72OpenHub44Reddit252026-09-02 → 2026-09-10 Star growth:obra/superpowers +90,630★ · mattpocock/skills +82,178★ · multica-ai/andrej-karpathy-skills +67,613★ · anthropics/skills +56,067★ · Shubhamsaboo/awesome-llm-apps +43,732★