AI News — Friday, July 31, 2026

10
OpenAI Unveils GPT-5.6, Advancing Price-Performance Frontier

OpenAI announced the release of GPT-5.6, claiming significant improvements in both performance and cost efficiency for its latest large language model.

OpenAI Blogproduct
9
Anthropic's AI Models Successfully Breach Companies in Security Tests

Anthropic revealed that its own AI models managed to breach three companies during internal security assessments, highlighting advanced capabilities and potential risks.

TechCrunchindustry
8
How Enabling Two Settings Tripled ARC-AGI-3 Benchmark Scores

OpenAI detailed how a simple adjustment of two settings dramatically improved their models' performance on the challenging ARC-AGI-3 benchmark, indicating significant progress in AI reasoning.

OpenAI Blogresearch
7
Reddit Reports Solid Quarter Amidst Signs of AI's Impact

Reddit announced strong quarterly results but also indicated that the growing influence of AI is beginning to show effects on its platform.

TechCrunchindustry
7
CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization

New research introduces CoRT, a method using counterfactual replay to optimize policies for token-level rubric guidance, enhancing AI agent performance in complex tasks.

Hugging Faceresearch
7
HumanCLAW: Can Vision-Language Models Act Through a Body?

A new paper explores the capabilities of Vision-Language Models to control a physical body, pushing the boundaries of embodied AI.

Hugging Faceresearch
6
Metis: Introducing a Novel Memory Foundation Model

Researchers present Metis, a new memory foundation model designed to improve how AI systems store and retrieve information over long horizons.

Hugging Faceresearch
6
DecoEvo: Score-Decoupled Co-Evolution of Solver and Rubric-Generator Skills in Text Space

DecoEvo proposes a novel co-evolutionary approach that decouples solver and rubric-generator skills to improve text-based problem-solving.

Hugging Faceresearch
6
PhiZero: A World Model Built Around Physical Language

PhiZero introduces a new world model that integrates physical language to better understand and predict real-world dynamics.

Hugging Faceresearch
6
Writing the PHP Virtual Machine in Rust (with a lot of help from AI)

A developer shares their experience of rewriting the PHP Virtual Machine in Rust, significantly leveraging AI tools throughout the development process.

Lobste.rsindustry
5
CLBench-V: Evaluating Multimodal Context Learning from Grounding to Knowledge Acquisition

CLBench-V is introduced as a new benchmark for evaluating multimodal context learning in AI, spanning from basic grounding to complex knowledge acquisition.

Hugging Faceresearch
5
Skills vs MCP: How AI tools have evolved

This article discusses the evolution of AI tools, contrasting 'skills-based' approaches with 'MCP' (Multi-Competency Platform) models.

Dev.toindustry
5
AskChem: Claim-Centered Infrastructure for Chemistry Literature Synthesis

AskChem proposes a claim-centered infrastructure to facilitate the synthesis of chemistry literature, leveraging AI for better organization and retrieval.

Hugging Faceresearch
5
Beacon: Knowing When and How to Perform Agentic Visual Reasoning

Beacon introduces a framework that enables AI agents to intelligently decide when and how to apply visual reasoning for improved performance.

Hugging Faceresearch
5
Flux-OPD: On-Policy Distillation with Evolving Contexts

Flux-OPD presents a new on-policy distillation method that adapts to evolving contexts, enhancing the learning efficiency and robustness of AI models.

Hugging Faceresearch