§ 01
News
What happened in AI todaySources: OpenAI / Anthropic / DeepMind / Moonshot / arXiv and other public RSS feeds. Updated twice daily at 06:00 / 18:00.
Sep 24, 2026arXiv cs.AI
Silent Failures in Agent-Tool Interaction: An Audit of ToolUniverseAgentic AI systems are increasingly adopting automated pipelines that integrate multiple tools. While prior research and benchmarks have studied about task success and task completion of these agentic systems, the research about agent to tool interaction, specifically in biology agentic workflow is limited. This study…
↗Sep 24, 2026arXiv cs.AIHarness as a Language: A Minimalist Agent Framework With Maximal ExpressivityModern language-model agents are built around the \textit{agent loop}, where the LLM is placed in an environment exposing a set of tools, and the LLM has full control over the workflow by alternating between tool calls and observing their output. However, certain workflows currently require additional engineering beyo…
↗Sep 24, 2026arXiv cs.AITwinCheck: Evidence-Grounded Negative-Twin Verification for Stateful Tool AgentsA single locally plausible tool call can derail an otherwise successful agent trajectory. Suspicion alone does not justify intervention, because the replacement itself can introduce the very failure verification is meant to prevent. We introduce TwinCheck, an inference-time verification policy that considers replaceme…
↗Sep 24, 2026arXiv cs.AIBuilding Socio-Affective Artificial Intelligence for Interactive Multi-Agent SimulationsThe objective of this article is to provide design principles and a software architecture for enabling interaction between humans and multiple agents in simulated dynamic worlds. This connects the current era of general artificial intelligence (AI/AGI) with the proliferation of transformer-based conversational agents…
↗Sep 24, 2026arXiv cs.AIWhich Objectives Need a Dial? Predicting Objective Conflict and Covering Trade-offs in Steerable Pluralistic AlignmentPeople hold diverse, sometimes conflicting values, so no single aligned model can satisfy everyone. Pluralistic alignment therefore calls for steerable models that can balance competing objectives differently. Multi-Objective Direct Preference Optimization (MODPO) does this by using an objective weight to span a conti…
↗Sep 24, 2026arXiv cs.AIEscaping Python Dependency Hell: A Hybrid Replay-and-Repair Pipeline for Python Dependency ResolutionDependency conflicts in Python ecosystems arise from incompatible version constraints, missing packages, and undocumented compatibility relationships, causing many real-world code snippets to fail at execution. This paper presents PLLM+, a hybrid dependency-repair pipeline evaluated on the HG2.9K benchmark of 2,891 de…
↗Sep 24, 2026arXiv cs.AISame evidence, different judgments: Evidence noncommutative in vision/speech-text conflictsFor multimodal large language models, when images or speech conflict with accompanying text, measured text reliance can entangle modality preference with evidence position. Earlier studies of text bias often used a fixed evidence order or moved task instructions with the evidence, leaving the contribution of order unc…
↗Sep 24, 2026arXiv cs.AIReinforcement Learning with Decomposed SubtasksGroup Relative Policy Optimization (GRPO) and related policy-gradient methods for training language model agents collapse an entire multi-turn rollout into a single scalar trajectory reward before it enters the policy update. When the task composes distinct skills, especially under sparse and delayed environmental fee…
↗Sep 23, 2026Hacker News (AI)Once Claude can measure something, it can make it faster
↗Sep 23, 2026Hugging FaceHow to Use NVIDIA Warp and MjWarp to Accelerate Robotics Simulation and Learning Workflows
↗Sep 23, 2026Google DeepMindAdvancing Private AI Compute with secure, server-side memoryIntroducing private, server-side memory to Private AI Compute for personal AI.
↗Sep 23, 2026OpenAITwo years of OpenAI AcademyMarking two years of OpenAI Academy and bringing AI skills to even more communities.
↗Sep 23, 2026Google DeepMindGemini 3.8 text-to-speech says hello
↗Sep 23, 2026Hacker News (AI)GPT-6 Astra has gained the ability to drive a car
↗Sep 23, 2026Hacker News (AI)Stripe's Knowledge AI Platform
↗Sep 23, 2026OpenAIOpenAI extends cyber access to Ukraine for civilian defenseOpenAI is extending access to its Daybreak program to the Government of Ukraine to support the cyber defense of civilian infrastructure.
↗Sep 23, 2026Hacker News (AI)Claude Code reads AGENTS.md only when telemetry is on [fixed]
↗Sep 23, 2026OpenAISam Altman’s remarks at the United Nations Security CouncilOpenAI CEO Sam Altman discusses AI safety, human control, and international cooperation in remarks to the United Nations Security Council.
↗Sep 23, 2026OpenAIHarvey turns legal context into stronger drafts with GPT-6 AstraGPT-6 Astra produces more structured, context-aware legal documents, freeing lawyers to focus on strategy.
↗Sep 23, 2026OpenAIHow invideo improves color grading 3x with GPT‑6 AstraWith GPT‑6 Astra, invideo plans edits with greater precision, improves color correction and grading threefold, and produces 50 custom effects in one day.
↗Sep 23, 2026OpenAIRingg’s AI agents resolve up to 65% of customer calls with OpenAIUsing GPT-5.6, Ringg powers multilingual agents across voice, chat, WhatsApp, and web for 90% less cost vs. GPT-4.1.
↗Sep 23, 2026OpenAIIntroducing MentalHealthBenchMentalHealthBench is an expert-informed benchmark for evaluating helpful and safe AI responses across realistic mental health conversations.
↗Sep 23, 2026arXiv cs.AIDidactic knowledge or Clinical Cases? How Data Types Shape Medical Large Language ModelsMedical large language models are commonly trained on mixtures of didactic data (e.g., textbooks) and clinical data (e.g., patient records), yet how these data types differentially shape model capabilities remains unclear. We address this issue with token-matched experiments that vary the didactic-to-clinical ratio an…
↗Sep 23, 2026arXiv cs.AIAn Affordable AI-Integrated Smart Cane for Multimodal Mobility Assistance of Visually Impaired UsersVisual impairment affects over 2.2 billion people worldwide, yet conventional white canes cannot detect elevated hazards or provide semantic environmental context. Existing AI-assisted navigation systems typically rely on expensive hardware or cloud connectivity, limiting accessibility in resource-constrained settings…
↗Sep 23, 2026arXiv cs.AIPAANI : On Device Visual Evidence Fusion and Explainable Guidance for River Robot SimulationMobile river monitoring robots must interpret obstacles and water boundaries that geographic waypoints alone cannot describe. On resource constrained platforms, converting imperfect visual predictions into timely and inspectable guidance is a distinct challenge. An object label or steering command does not explain whi…
↗Sep 23, 2026arXiv cs.AISocial Influence and the Allocation of Scientific Attention in AI PopulationsAI systems are becoming participants in the evaluation and use of scientific research. They encounter citation counts, download statistics and lists of popular articles developed around human readers, but the collective consequences of these signals for artificial readers remain uncertain. This paper adapts the Music…
↗Sep 23, 2026arXiv cs.AILearning 3D biophysical cell properties from 2D images and cell-population statisticsInferring 3D cellular properties from 2D microscopy is difficult when a reference instrument reports only population statistics rather than labels for individual cells. Here we develop a population-supervised framework that maps single 2D red-cell images to latent biophysical quantities and aggregates them to mean cor…
↗Sep 23, 2026arXiv cs.AIGoal-driven Variant CategorizationProcess discovery rarely yields a single coherent process structure. For analysis, a common step is to cluster process variants based on structural similarity and then assign business meaning to the resulting groups. Since these partitions are not derived from the organization's goals, analysts must manually interpret…
↗Sep 23, 2026arXiv cs.AIReplication Without Persistence in Hosted LLMs: Measurement Sensitivity in Action-Time Belief EvaluationBehavioural evaluations of hosted language models can vary because the evaluated service, the measurement instrument, or both differ across runs. We separate three validation questions: whether a prior finding recurs on fresh data under its historical configuration (replication), whether the endpoint changes when the…
↗Sep 23, 2026arXiv cs.AIThe Wisdom of Artificial Deliberative CrowdsThe aggregation of many lay estimates often outperforms individual expert judgment, a phenomenon known as the wisdom of crowds. While this is usually attributed to the independence of estimates, an even stronger effect arises through deliberation: averaging the consensus estimates of small deliberating groups outperfo…
↗Sep 23, 2026OpenAIChatGPT Ads expands to Southeast Asia and TaiwanChatGPT Ads is expanding to Southeast Asia and Taiwan, giving eligible businesses new ways to reach people across more than 60 countries.
↗Sep 23, 2026OpenAIAirbnb widens access to GPT-6 Astra and OpenAI frontier modelsLearn how Airbnb is expanding access to GPT-6 Astra and OpenAI frontier models to help engineering teams solve bugs, design systems, and ship faster.
↗Sep 23, 2026AnthropicClaude discovers a novel enzyme system with CRISPR-like repeatsClaude discovers a novel enzyme system with CRISPR-like repeats
↗Sep 22, 2026OpenAIBetter prompt caching for GPT-6Learn how GPT-6 improves prompt caching with higher cache hit rates, new diagnostics, explicit breakpoints, and controls that reduce latency and costs.
↗Sep 22, 2026Hacker News (AI)LLM Ass Bench
↗Sep 22, 2026Hacker News (AI)Pentagon says overreliance on AI contributed to missile strike on Iran school
↗Sep 22, 2026Hacker News (AI)GPT-6 Sol and Luna
↗Sep 22, 2026OpenAIIntroducing GPT-6 Sol and LunaMeet GPT-6 Sol and Luna, two models that bring frontier intelligence to everyday work with different balances of capability and cost.
↗Sep 22, 2026Hacker News (AI)Claude Opus 5.5 Intelligence, Performance and Price Analysis (Max)
↗Sep 22, 2026Hacker News (AI)Claude Opus 5.5
↗Sep 22, 2026Hacker News (AI)OpenAI GPT–6 Astra breaks Enigma message that has resisted solution since 2005
↗Sep 22, 2026Hacker News (AI)AI Has No Wisdom and Neither Will You
↗Sep 22, 2026OpenAIParallel cut research time and cost in half with GPT‑6 AstraGPT‑6 Astra allowed Parallel’s agents to research and synthesize labor-market data in half the time and at half the cost vs. prior models.
↗Sep 22, 2026Kimi1.52.0: feat(cli): short-circuit entry points to a Kimi Code installer (#2666)Co-authored-by: jackfish212 jackfish212@outlook.com
↗Sep 22, 2026Hugging FaceJun Kim, oMLX creator and maintainer, joins Hugging Face to support the MLX community
↗Sep 22, 2026OpenAIPriorities and principles for effective third party assessmentsOpenAI outlines priorities and principles for rigorous, secure, and independent third-party AI safety assessments of frontier models and safeguards.
↗Sep 22, 2026Hugging FaceHow UK AISI and EvalEval Are Making Benchmark Results Reproducible
↗Sep 22, 2026Hugging FaceTransformers now runs llama.cpp quants
↗Sep 21, 2026Hacker News (AI)Amazon blocks Meta’s new Muse AI agent from shopping on amazon.com
↗Sep 21, 2026Kimi1.51.0What's Changed chore: archive kimi-cli and point users to Kimi Code CLI by @RealKai42 in #2659 chore(release): bump kimi-cli to 1.51.0 by @RealKai42 in #2660 Full Changelog: 1.50.0...1.51.0
↗