AI-Pilled 日报

AI 圈每天在变,我帮你看着

§ 01

News

AI 圈每天发生了什么
2026.09.26Hacker News (AI)
One Month Without AI
↗
2026.09.25Hacker News (AI)
Too AI; Didn't Read
↗
2026.09.25OpenAI
Proaction boosts sales 60% and saves 75+ hours with CodexWith Codex, GPT-Live-1, and GPT-6 Astra, Proaction builds, operates, and sells modern fleet management faster.
↗
2026.09.25Hacker News (AI)
Yes, Claude can do nine loops
↗
2026.09.25Hacker News (AI)
Classified estimates show the NSA is paying billions to test AI models
↗
2026.09.25Hacker News (AI)
Microsoft abandons personal AI chatbot race with Copilot reboot
↗
2026.09.25arXiv cs.AI
When Should Forecasting Agents Reason? Behavioral Stress Tests for Reliability RoutingForecasting agents increasingly combine language-model reasoning, retrieval, ensembling, and calibration, but it remains unclear when each behavior should be trusted. We study this question on ForecastBench-style binary forecasting tasks, treating the choice to retrieve, reason, defer to a market prior, or use a histo…
↗
2026.09.25arXiv cs.AI
TW3Cast: A Frozen Router of Lightly Fine-Tuned Foundation Models for Time-Series Forecasting on GIFT-Eval, Selected Entirely on the Training SplitTW3Cast is a time-series forecasting system that reaches position 3 of 130 entries on the GIFT-Eval benchmark by mean MASE rank, as of 2026-09-14. The two entries above it belong to the leaderboard's agentic category, multi-step systems that use agents or language models to reason about, generate or select forecasts.…
↗
2026.09.25arXiv cs.AI
PAWS: Policy-driven Agentic World SimulationPolicy interventions propagate through public communication, institutional decisions, and stakeholder responses, yet datasets for financial multi-agent simulation rarely connect these processes to temporally aligned historical evidence. We introduce PAWS, a Policy-driven Agentic World Simulation dataset covering 36 ve…
↗
2026.09.25arXiv cs.AI
Pistis Technical ReportWe introduce the Pistis model family, comprising 27B- and 9B-parameter multimodal large language models built on Qwen3.6 and Qwen3.5, respectively, and developed through a general and scalable post-training framework. The framework first establishes a strong foundation through large-scale multimodal supervised fine-tu…
↗
看全部 News →
§ 03

Perspectives

他者之声
“
Software 1.0 是手写代码;Software 2.0 是神经网络的权重;Software 3.0 是 prompt。
Andrej KarpathySoftware 3.0, May 2026
我的看法这个分类很顺,但它低估了转换成本。把 SOP 翻译成 prompt 比把伪代码翻译成代码难得多——它要捕捉的是隐性知识,不是显性逻辑。
“
这个时代最持久的能力,是学习任何你想学的东西。
Naval RavikantAlmanack, 2020
我的看法十年前对,今天更对。当工具的 half-life 缩短到 18 个月,可迁移的"元能力"才是真正的复利。
§ 04

Plans

正在进行

Carol's Website

用 Claude Code 把这个站做出来,两天 MVP。一次从 idea 到 product 的完整练习,你正在看的就是它。

Shipped · 2026

The Big Inventory of AI Products

2025 年底做的一份 AI 产品分类图谱,被几个朋友长期收藏。

Shipped · 2025

AI 产品周记

每周一份精选短评,记录我在 AI 产品圈看到的、想到的、被反复触发的判断。

In Progress

数字分身 Agent v2

从全文 prompt 升级到 RAG,加上对话历史。让分身真的能"代我回答"。

Exploring
§ 05

About

关于这份日报

清华本科 · 大厂 AI PM · 亿级用户产品增长 · AI Pilled

我是 Carol,白天做 AI 产品,晚上把看见的、想到的、喜欢的人和观点写下来。这个站不是博客,更像一份给自己也开放给别人的 AI-Pilled 日报。

AI 让信息变多、让判断变少。所以这份日报每天替你过一遍,然后把我看完之后真正想说的那部分留下来。每一篇都希望值得被读两遍.

如果你也在思考 AI 与产品、与表达、与生活方式的关系,欢迎留个联系方式,或者直接去跟我的数字分身聊聊。