AI News — Friday, July 31, 2026
OpenAI announced the release of GPT-5.6, claiming significant improvements in both performance and cost efficiency for its latest large language model.
Anthropic revealed that its own AI models managed to breach three companies during internal security assessments, highlighting advanced capabilities and potential risks.
OpenAI detailed how a simple adjustment of two settings dramatically improved their models' performance on the challenging ARC-AGI-3 benchmark, indicating significant progress in AI reasoning.
Reddit announced strong quarterly results but also indicated that the growing influence of AI is beginning to show effects on its platform.
New research introduces CoRT, a method using counterfactual replay to optimize policies for token-level rubric guidance, enhancing AI agent performance in complex tasks.
A new paper explores the capabilities of Vision-Language Models to control a physical body, pushing the boundaries of embodied AI.
Researchers present Metis, a new memory foundation model designed to improve how AI systems store and retrieve information over long horizons.
DecoEvo proposes a novel co-evolutionary approach that decouples solver and rubric-generator skills to improve text-based problem-solving.
PhiZero introduces a new world model that integrates physical language to better understand and predict real-world dynamics.
A developer shares their experience of rewriting the PHP Virtual Machine in Rust, significantly leveraging AI tools throughout the development process.
CLBench-V is introduced as a new benchmark for evaluating multimodal context learning in AI, spanning from basic grounding to complex knowledge acquisition.
This article discusses the evolution of AI tools, contrasting 'skills-based' approaches with 'MCP' (Multi-Competency Platform) models.
AskChem proposes a claim-centered infrastructure to facilitate the synthesis of chemistry literature, leveraging AI for better organization and retrieval.
Beacon introduces a framework that enables AI agents to intelligently decide when and how to apply visual reasoning for improved performance.
Flux-OPD presents a new on-policy distillation method that adapts to evolving contexts, enhancing the learning efficiency and robustness of AI models.