AI News — Monday, September 21, 2026

9
RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning

This research introduces a novel method for agentic reinforcement learning that allows on-policy distillation to self-retire, improving efficiency and performance.

Hugging Faceresearch
9
How OpenAI Used Its Own LLMs to Design Its Jalapeño Chip

OpenAI leveraged its own large language models to assist in the design process of a new hardware chip, showcasing a practical application of AI in engineering.

Lobste.rsindustry
8
World model companies are keeping a lot of secrets

This article discusses the increasing trend of secrecy among leading AI model development companies, raising concerns about transparency and open research.

TechCrunchindustry
8
Is the AI industry really ready to slow down?

An analysis questions the industry's willingness to decelerate AI development amidst calls for caution and regulation, highlighting the ongoing race for innovation.

TechCrunchindustry
7
WeVisDoc: From Coverage to Capability for Robust End-to-End Document Parsing

New research presents WeVisDoc, a method aimed at enhancing the robustness and capability of end-to-end document parsing systems.

Hugging Faceresearch
7
VABench: Measuring Embodied Spatial Intelligence through Visual Demonstrations, Active Perception, and Metric Control

VABench is introduced as a new benchmark designed to evaluate embodied spatial intelligence in AI agents using visual demonstrations and active perception.

Hugging Faceresearch
7
Laya — 33ms Multilingual System 1 Decision Engine

Laya is announced as a new multilingual decision engine boasting an impressive 33-millisecond response time, indicating advancements in real-time AI processing.

Lobste.rsproduct
7
CodeMidas: Scaling Agentic Coding RL Environments from Code Itself

This paper explores CodeMidas, a framework for scaling reinforcement learning environments for agentic coding directly from existing codebases.

Hugging Faceresearch
6
EvoOntology: A Self-Evolving Ontology Layer for Data Agents

Researchers propose EvoOntology, a self-evolving ontology layer designed to enhance the knowledge representation and adaptability of data agents.

Hugging Faceresearch
6
When2Think: Learning Difficulty-Aware Length Control for Efficient Hybrid Reasoning Models

This research introduces a method for hybrid reasoning models to learn difficulty-aware length control, optimizing efficiency without sacrificing accuracy.

Hugging Faceresearch
6
UFO: Chain-of-Evaluation for Omni-Condition Alignment in Multi-Modal Image Generation

UFO presents a novel chain-of-evaluation approach to achieve omni-condition alignment, significantly improving multi-modal image generation.

Hugging Faceresearch
6
Sample Count Is Not Enough: Candidate-Generation Strategy Shapes the Energy and Performance of LLM Test-Time Scaling

A study reveals that the candidate-generation strategy, not just sample count, critically influences the energy consumption and performance of LLM scaling during testing.

Hugging Faceresearch
6
Don't Mask the Environment: Observation Supervision Changes How Agents Explore Under RL

This paper investigates how observation supervision fundamentally alters exploration strategies for agents in reinforcement learning environments.

Hugging Faceresearch
6
Introducing the Australian Youth Safety Blueprint

OpenAI has unveiled the Australian Youth Safety Blueprint, outlining strategies and commitments to ensure safer AI interactions for young people.

OpenAI Blogindustry
5
Architecting a Resilient DevSecOps Pipeline for Enterprise AI Agents

This article provides guidance on building robust and secure DevSecOps pipelines specifically tailored for the deployment and management of enterprise AI agents.

Dev.toindustry