The AI-Pilled Daily

The AI world keeps moving.I'm watching it for you

§ 01

News

What happened in AI today
Aug 13, 2026arXiv cs.AI
Dynamic Governance of Multi-LLM Agent Systems for Collaborative Conversational OutcomesWhen two LLM agents with structurally opposed objectives interact across multiple turns, the absence of a shared goal function produces not competition but collapse: the visitor capitulates, the site agent stops varying its approach, and the conversation terminates without achieving either agent's stated objective. Th…
Aug 13, 2026arXiv cs.AI
Distribird: Literature-Informed Prior Distribution Design for Bayesian Model CalibrationBayesian calibration of process-based models requires a prior distribution for each model parameter. Despite decades of methodological work, researchers almost always fall back on uniform priors. The main reason is that building informative priors from scientific literature is slow and needs both domain and statistica…
Aug 13, 2026arXiv cs.AI
A Forced-Structure Reduction and Verifiable Bounds for Conway's 99-GraphConway's 99-graph problem asks whether a strongly regular graph with parameters $\mathrm{srg}(99,14,1,2)$ exists. We report a systematic, fully reproducible attack by an autonomous AI research agent, scored under the track's partial-credit metric. Our verifiable contributions are: (1) an exhaustive proof that no circu…
Aug 13, 2026arXiv cs.AI
Detecting a Route Flip Is Easier Than Knowing Whether to Fix It: Causal Route-Mediated Damage in Quantized Mixture-of-ExpertsTop-k Mixture-of-Experts (MoE) routing is discontinuous, so a deployment-motivated numerical disturbance -- simulated 4-bit KV-cache quantization read by a protected BF16 gate -- pushes tokens across decision boundaries and flips which experts fire. This paper proposes no new mitigation; it supplies a causal apparatus…
Aug 13, 2026arXiv cs.AI
Poor Man's Agentic Modeling: Simulating Large LLM-Agent Societies on a LaptopSimulating societies of many large language model (LLM) agents is expensive, yet the questions asked of such simulations are usually macroscopic: phase behaviour, stylised facts, and scaling with the number of agents $N$, not the cognition of any single agent. We turn a statistical-physics observation into a method: r…
Aug 13, 2026arXiv cs.AI
AutoWorldModel-Bench: A State-Centric Benchmark for Automated World-Model ResearchWorld modeling is an unsettled field: architectures, training objectives, and state representations interact in complex ways, and no single recipe dominates across environments. This makes it an ideal testbed for AI coding agents acting as autonomous researchers--a setting in which the improvement direction is not spe…
Aug 13, 2026arXiv cs.AI
MaSRead: Content-Addressed Reading of Replicated Latent StoresIndependent agents that reason in latent space can share computed state as key-value cache fragments rather than text. Merged by a conflict-free replicated data type, these fragments form a store that converges under any delivery order or duplication. Yet a later query, unknown at encode time, cannot reliably read the…
Aug 13, 2026arXiv cs.AI
From Monolithic to Modular: Segment-level Automatic Prompt OptimizationAutomatic Prompt Optimization (APO) often rewrites prompts monolithically, which can improve one behavior while degrading others. We present SAPO, a segment-level APO method that decomposes prompts into role, context, tasks, and output format, then applies targeted improvements based on top-5 and bottom-5 examples. Th…
Aug 12, 2026Hugging Face
Introducing OlmoEarth embeddings: Custom embedding exports from OlmoEarth Studio for downstream analysis
Aug 12, 2026Hacker News (AI)
German advocacy group lodges criminal complaint over Meta AI glasses
See all News
§ 03

Perspectives

Voices I keep
Software 1.0 is hand-written code; Software 2.0 is the weights of neural networks; Software 3.0 is the prompt.
Andrej KarpathySoftware 3.0, May 2026
My takeA neat taxonomy, but it understates the conversion cost. Translating an SOP into a prompt is far harder than translating pseudocode into code — what you have to capture is tacit knowledge, not explicit logic.
The most durable skill of this era is learning anything you want to learn.
Naval RavikantAlmanack, 2020
My takeTrue ten years ago, even truer today. When tools have a half-life of 18 months, the transferable meta-skills are where the real compounding happens.
§ 04

Plans

In motion

Carol's Website

Built this site in two days with Claude Code — a full idea-to-product exercise. You're looking at it.

Shipped · 2026

The Big Inventory of AI Products

A taxonomy I built in late 2025. Quietly bookmarked by a handful of friends since.

Shipped · 2025

AI Product Weekly

A curated weekly note: things I noticed, things I changed my mind on, things worth re-reading.

In Progress

Digital Twin Agent v2

Upgrading from system-prompt context to RAG + conversation history. Make the twin actually answer for me.

Exploring
§ 05

About

About this daily

Tsinghua undergrad · AI PM @ major tech co · 100M+ user product growth · AI-pilled

I'm Carol. By day I build AI products; by night I write down the people, ideas and observations I want to keep. This isn't a blog — it's closer to an AI-pilled daily, kept open to whoever finds it useful.

AI gives us more information and less judgement. So this daily walks through it for you, then keeps the part I actually wanted to say after reading it. Each piece should be worth reading twice.

If you're also thinking about AI in relation to product, expression, or the shape of your life, drop a contact — or just talk to my digital twin.