§ 01
News
AI 圈每天发生了什么来源:OpenAI / Anthropic / DeepMind / 月之暗面 / arXiv 等公开 RSS。每日 06:00 / 18:00 自动更新。
2026.09.26Hacker News (AI)
One Month Without AI
↗2026.09.25Hacker News (AI)Too AI; Didn't Read
↗2026.09.25OpenAIProaction boosts sales 60% and saves 75+ hours with CodexWith Codex, GPT-Live-1, and GPT-6 Astra, Proaction builds, operates, and sells modern fleet management faster.
↗2026.09.25Hacker News (AI)Yes, Claude can do nine loops
↗2026.09.25Hacker News (AI)Classified estimates show the NSA is paying billions to test AI models
↗2026.09.25Hacker News (AI)Microsoft abandons personal AI chatbot race with Copilot reboot
↗2026.09.25arXiv cs.AIWhen Should Forecasting Agents Reason? Behavioral Stress Tests for Reliability RoutingForecasting agents increasingly combine language-model reasoning, retrieval, ensembling, and calibration, but it remains unclear when each behavior should be trusted. We study this question on ForecastBench-style binary forecasting tasks, treating the choice to retrieve, reason, defer to a market prior, or use a histo…
↗2026.09.25arXiv cs.AITW3Cast: A Frozen Router of Lightly Fine-Tuned Foundation Models for Time-Series Forecasting on GIFT-Eval, Selected Entirely on the Training SplitTW3Cast is a time-series forecasting system that reaches position 3 of 130 entries on the GIFT-Eval benchmark by mean MASE rank, as of 2026-09-14. The two entries above it belong to the leaderboard's agentic category, multi-step systems that use agents or language models to reason about, generate or select forecasts.…
↗2026.09.25arXiv cs.AIPAWS: Policy-driven Agentic World SimulationPolicy interventions propagate through public communication, institutional decisions, and stakeholder responses, yet datasets for financial multi-agent simulation rarely connect these processes to temporally aligned historical evidence. We introduce PAWS, a Policy-driven Agentic World Simulation dataset covering 36 ve…
↗2026.09.25arXiv cs.AIPistis Technical ReportWe introduce the Pistis model family, comprising 27B- and 9B-parameter multimodal large language models built on Qwen3.6 and Qwen3.5, respectively, and developed through a general and scalable post-training framework. The framework first establishes a strong foundation through large-scale multimodal supervised fine-tu…
↗2026.09.25arXiv cs.AIBaseCamp --- An Agentic AI Framework for Automating DNA Sequencing Data PipelinesDNA sequencing pipelines, spanning quality control, alignment, variant calling, and annotation, are now reliably executed by workflow management systems that orchestrate established bioinformatics tools at scale. What remains manual is the decision layer surrounding that execution: selecting quality thresholds appropr…
↗2026.09.25arXiv cs.AIDEEPO: Dual-Entropy Enhanced Policy Optimization for Hallucination in MLLMsReinforcement learning (RL) is widely used to sharpen reasoning in multimodal large language models (MLLMs), yet its effect on hallucination is uneven. We trace this to two weak points in the \emph{correction chain} from reward to parameter update. At the rollout level, hard queries---those with high semantic entropy-…
↗2026.09.25arXiv cs.AITWIST: A Proposed Benchmark for Intervention Quality in Conversational Memory, with a Human-Validated Draft-AlignmentLong-conversation memory benchmarks increasingly test recall and prompted knowledge updates, and recent work studies evolving user beliefs and memory state. TWIST is a proposed benchmark suite for a complementary, unmeasured property: intervention quality -- whether a deployed memory system, exercised through its own…
↗2026.09.25arXiv cs.AIAdversarial Closed-Loop Curriculum for Evolving Role-Playing AgentsRole-playing agents based on large language models have been widely applied in areas such as personalized assistance and social simulation. Recent RL methods typically train on a fixed scenario pool collected before learning begins. This creates a distributional bottleneck: as the agent improves, the scenarios where i…
↗2026.09.24Hacker News (AI)How I changed teaching after AI managed to do all my homework assignments
↗2026.09.24Google DeepMindIntroducing Gemini 3.8 Live with Live Avatar
↗2026.09.24Hacker News (AI)Tutoring company tells parents to save their money and 'use AI instead'
↗2026.09.24Hacker News (AI)AI safety is mostly a sex cult in Berkeley
↗2026.09.24Hacker News (AI)Best LLM for every budget, updated daily
↗2026.09.24Hugging FaceAccelerating vision-language models with LFM2.5-VL-DSpark
↗2026.09.24Hacker News (AI)'That's so AI ' What gen Alpha's biggest insult tells us
↗2026.09.24Hacker News (AI)Meta takes down a critical video about meta AI Glasses after filming at Meta
↗2026.09.24Hacker News (AI)Early rogue AI agent activity and attempts to hack found on urlquery.net
↗2026.09.24arXiv cs.AISilent Failures in Agent-Tool Interaction: An Audit of ToolUniverseAgentic AI systems are increasingly adopting automated pipelines that integrate multiple tools. While prior research and benchmarks have studied about task success and task completion of these agentic systems, the research about agent to tool interaction, specifically in biology agentic workflow is limited. This study…
↗2026.09.24arXiv cs.AIHarness as a Language: A Minimalist Agent Framework With Maximal ExpressivityModern language-model agents are built around the \textit{agent loop}, where the LLM is placed in an environment exposing a set of tools, and the LLM has full control over the workflow by alternating between tool calls and observing their output. However, certain workflows currently require additional engineering beyo…
↗2026.09.24arXiv cs.AITwinCheck: Evidence-Grounded Negative-Twin Verification for Stateful Tool AgentsA single locally plausible tool call can derail an otherwise successful agent trajectory. Suspicion alone does not justify intervention, because the replacement itself can introduce the very failure verification is meant to prevent. We introduce TwinCheck, an inference-time verification policy that considers replaceme…
↗2026.09.24arXiv cs.AIBuilding Socio-Affective Artificial Intelligence for Interactive Multi-Agent SimulationsThe objective of this article is to provide design principles and a software architecture for enabling interaction between humans and multiple agents in simulated dynamic worlds. This connects the current era of general artificial intelligence (AI/AGI) with the proliferation of transformer-based conversational agents…
↗2026.09.24arXiv cs.AIWhich Objectives Need a Dial? Predicting Objective Conflict and Covering Trade-offs in Steerable Pluralistic AlignmentPeople hold diverse, sometimes conflicting values, so no single aligned model can satisfy everyone. Pluralistic alignment therefore calls for steerable models that can balance competing objectives differently. Multi-Objective Direct Preference Optimization (MODPO) does this by using an objective weight to span a conti…
↗2026.09.24arXiv cs.AIEscaping Python Dependency Hell: A Hybrid Replay-and-Repair Pipeline for Python Dependency ResolutionDependency conflicts in Python ecosystems arise from incompatible version constraints, missing packages, and undocumented compatibility relationships, causing many real-world code snippets to fail at execution. This paper presents PLLM+, a hybrid dependency-repair pipeline evaluated on the HG2.9K benchmark of 2,891 de…
↗2026.09.24arXiv cs.AISame evidence, different judgments: Evidence noncommutative in vision/speech-text conflictsFor multimodal large language models, when images or speech conflict with accompanying text, measured text reliance can entangle modality preference with evidence position. Earlier studies of text bias often used a fixed evidence order or moved task instructions with the evidence, leaving the contribution of order unc…
↗2026.09.24arXiv cs.AIReinforcement Learning with Decomposed SubtasksGroup Relative Policy Optimization (GRPO) and related policy-gradient methods for training language model agents collapse an entire multi-turn rollout into a single scalar trajectory reward before it enters the policy update. When the task composes distinct skills, especially under sparse and delayed environmental fee…
↗2026.09.24Hacker News (AI)Feds Target AI Critics as "Foreign Agents"
↗2026.09.23Hacker News (AI)Mercury 2.5 LLM hits 770 tokens per second
↗2026.09.23Hacker News (AI)Claude's Load-Bearing Seams
↗2026.09.23Hacker News (AI)Once Claude can measure something, it can make it faster
↗2026.09.23Hugging FaceHow to Use NVIDIA Warp and MjWarp to Accelerate Robotics Simulation and Learning Workflows
↗2026.09.23Google DeepMindAdvancing Private AI Compute with secure, server-side memoryIntroducing private, server-side memory to Private AI Compute for personal AI.
↗2026.09.23OpenAITwo years of OpenAI AcademyMarking two years of OpenAI Academy and bringing AI skills to even more communities.
↗2026.09.23Google DeepMindGemini 3.8 text-to-speech says hello
↗2026.09.23Hacker News (AI)GPT-6 Astra has gained the ability to drive a car
↗2026.09.23Hacker News (AI)Stripe's Knowledge AI Platform
↗2026.09.23OpenAIOpenAI extends cyber access to Ukraine for civilian defenseOpenAI is extending access to its Daybreak program to the Government of Ukraine to support the cyber defense of civilian infrastructure.
↗2026.09.23Hacker News (AI)Claude Code reads AGENTS.md only when telemetry is on [fixed]
↗2026.09.23OpenAISam Altman’s remarks at the United Nations Security CouncilOpenAI CEO Sam Altman discusses AI safety, human control, and international cooperation in remarks to the United Nations Security Council.
↗2026.09.23OpenAIHarvey turns legal context into stronger drafts with GPT-6 AstraGPT-6 Astra produces more structured, context-aware legal documents, freeing lawyers to focus on strategy.
↗2026.09.23OpenAIHow invideo improves color grading 3x with GPT‑6 AstraWith GPT‑6 Astra, invideo plans edits with greater precision, improves color correction and grading threefold, and produces 50 custom effects in one day.
↗2026.09.23OpenAIRingg’s AI agents resolve up to 65% of customer calls with OpenAIUsing GPT-5.6, Ringg powers multilingual agents across voice, chat, WhatsApp, and web for 90% less cost vs. GPT-4.1.
↗2026.09.23OpenAIIntroducing MentalHealthBenchMentalHealthBench is an expert-informed benchmark for evaluating helpful and safe AI responses across realistic mental health conversations.
↗2026.09.23arXiv cs.AIDidactic knowledge or Clinical Cases? How Data Types Shape Medical Large Language ModelsMedical large language models are commonly trained on mixtures of didactic data (e.g., textbooks) and clinical data (e.g., patient records), yet how these data types differentially shape model capabilities remains unclear. We address this issue with token-matched experiments that vary the didactic-to-clinical ratio an…
↗2026.09.23arXiv cs.AIAn Affordable AI-Integrated Smart Cane for Multimodal Mobility Assistance of Visually Impaired UsersVisual impairment affects over 2.2 billion people worldwide, yet conventional white canes cannot detect elevated hazards or provide semantic environmental context. Existing AI-assisted navigation systems typically rely on expensive hardware or cloud connectivity, limiting accessibility in resource-constrained settings…
↗