AI News — Sunday, July 19, 2026

9
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities

A new unified infrastructure is proposed for comprehensively evaluating the capabilities of AI agents, addressing a critical need in agent development.

Hugging Faceresearch
9
PolicyShiftGuard: Benchmarking and Improving Policy-Adaptive Image Guardrails

Researchers introduce a new benchmark and methods to improve policy-adaptive image guardrails, enhancing AI safety and alignment with ethical guidelines.

Hugging Faceresearch
9
The agent security gap: 54% of enterprises have already had an AI agent incident, and most still let agents share credentials

A new report reveals that over half of enterprises have experienced security incidents involving AI agents, largely due to agents sharing credentials, highlighting a critical vulnerability.

VentureBeatindustry
8
The agent evaluation gap: Enterprise AI organizations have a reality-alignment problem, not a coverage problem — and most are shipping to production anyway

Enterprises are struggling with evaluating AI agents effectively, leading to a disconnect between perceived and actual performance, yet many are still deploying them to production.

VentureBeatindustry
8
Kimi: Threat or menace?

TechCrunch explores the potential impact and implications of the new AI entity 'Kimi,' questioning its role in the evolving AI landscape.

TechCrunchindustry
8
A scorecard for the AI age

OpenAI introduces a new framework or 'scorecard' to assess progress and challenges in the development and deployment of AI technologies.

OpenAI Blogindustry
8
KeyFrame-Compass: Towards Comprehensive Evaluation of Keyframe-Conditioned Video Generation

A new evaluation benchmark, KeyFrame-Compass, is introduced to thoroughly assess the quality and consistency of video generation models conditioned on keyframes.

Hugging Faceresearch
7
Neil Rimer thinks the AI money is coming back out

Prominent investor Neil Rimer suggests that the significant capital influx into the AI sector may soon begin to recede, signaling a potential shift in investment trends.

TechCrunchindustry
7
Agentic orchestration: Enterprise AI organizations have a deployment problem, not a platform problem — and most are calling chatbots agents

Many enterprises are mislabeling chatbots as advanced AI agents and face significant challenges in orchestrating and deploying true agentic systems, highlighting a gap in understanding and implementation.

VentureBeatindustry
7
Your PDFs Are Eating Your LLM's Tokens for Breakfast

This article highlights how inefficient PDF processing can consume excessive tokens in Large Language Models, impacting performance and cost.

Dev.toproduct
7
MetaView: Monocular Novel View Synthesis with Scale-Aware Implicit Geometry Priors

Researchers present MetaView, a new method for generating novel views from a single image by incorporating scale-aware implicit geometry priors.

Hugging Faceresearch
7
MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation

MultiRef-Compass is introduced as a new benchmark for thoroughly evaluating AI models that generate audio-video content from multiple references.

Hugging Faceresearch
7
Self-Improvements in Modern Agentic Systems: A Survey

This survey provides a comprehensive overview of current techniques and challenges in enabling self-improvement capabilities within modern AI agentic systems.

Hugging Faceresearch
6
How Cars24 scales conversations and builds faster with OpenAI

Cars24 shares insights into how they leverage OpenAI's technologies to scale customer conversations and accelerate their development processes.

OpenAI Blogindustry
6
RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination

RxBrain introduces an embodied cognition foundation model capable of joint language-visual reasoning and imaginative capabilities, pushing boundaries in multimodal AI.

Hugging Faceresearch