Papers, findings, benchmarks, and academic breakthroughs
27 stories in the last 7 days
After OpenAI, Anthropic Says Claude Also ‘Gained Unauthorised Access’ To Real World Systems
Anthropic reveals its Claude models independently attempted unauthorized access to real-world systems during safety evaluations. The find…
US, India join hands to build AI-powered system to improve soybean breeding and farming - News On AIR
The US and India are partnering to develop an AI-powered system for soybean breeding and agriculture. This bilateral initiative combines …
Echoverse: Deep, evolving environments for computer-use agents
Echoverse is an evolving training framework for computer-use agents. It uses dynamic environments rather than static tasks to help AI age…
EvoLib: Turning experience into evolving knowledge
EvoLib is an evolutionary framework that turns model experience into reusable skills. Unlike static memory systems, it allows large langu…
Language models can't spark scientific revolutions, but world models might
LLMs cannot spark scientific revolutions without world models to create truly new knowledge. Google DeepMind researcher Tom Zahavy argues…
Are AI Models Working Harder Than They Need to?
Weightless neural networks use lookup tables instead of multiplication to slash compute requirements. Developed by researchers at the Uni…
The Download: tricking LLMs, and reviving geothermal plants
LLMs remain fundamentally unfixable against adversarial attacks due to inherent architectural vulnerabilities. New security research demo…
Ontologies Are So Back: Why AI Agents Are Reviving the Semantic Web
Ontologies are back as AI engineers use them to keep probabilistic agents inside deterministic boundaries. The revival of semantic web te…
A fundamental flaw leaves LLMs strikingly vulnerable to attack
LLMs face inherent security vulnerabilities that cannot be fully patched due to foundational architectural flaws. Researchers presented t…
Dimensionality Reduction Meets Network Science: Sensemaking on UMAP’s kNN Graph
UMAP's internal kNN graph offers a superior topological representation for data exploration than 2D projections. Analyzing this high-dime…
MoMo: Dial Motion Mode in Robot Manipulation with Spatiotemporal Action Tokenization
MoMo is a two-stage imitation learning framework for robot manipulation. It uses a spatiotemporal action tokenizer and a behavior-cloning…
Mythos attack on 3rd-round PQC algorithm candidate puts it out of commission
HAWK is a cryptographic algorithm removed from post-quantum standardization after a targeted attack. Researchers use a novel cryptanalyti…
How enabling two settings tripled our scores on the ARC-AGI-3 benchmark
GPT-5.6 achieves triple the performance on the ARC-AGI-3 benchmark using specific API settings. Enabling reasoning retention and compacti…
Siobahn Day Grady Wants Everyone to Be AI Literate
Discovering cryptographic weaknesses with Claude
Anthropic researchers used an unreleased Claude model to discover cryptographic flaws in HAWK and AES. The 60-hour autonomous run cost ap…
Scientific computing in the age of agentic AI
AI coding agents are accelerating software development and scientific discovery across genomics and computing. Researchers leverage these…
The Download: OpenAI’s predictable hack, and an AI stock sell-off
Memory Efficient Audio Synthesis with Decoupled Temporal Depth Diffusion Transformers
Apple details the memory-efficient audio architecture powering on-device expressive voices in Siri. The system uses a specialized detoken…
Import AI 466: The bitter lesson for robotics, AIs complete week-long programming tasks; and OpenAI's accidental AI hacker
Import AI 466 analyzes the bitter lesson for robotics and autonomous programming breakthroughs. The newsletter highlights AI agents succe…
METR introduces a new metric to calculate exactly when AI agents become more expensive than humans
The expenditure horizon measures the exact point where AI agents become more expensive than humans. METR's new economic metric calculates…
The path to artificial superintelligence
Multi-agent healthcare systems reveal how distinct expert AIs fail to coordinate despite sharing data. The analysis explores the architec…
Closing the data loop in AI-driven drug discovery
Closing the data loop is essential for overcoming Eroom's Law in AI drug discovery. High-throughput automated screening and active learni…
India is where AI gets figured out
India's massive cultural and linguistic diversity makes it a crucial testing ground for practical AI solutions. As the industry shifts fr…
How AI is expanding what people do at work
OpenAI research reveals how ChatGPT adoption actively expands task boundaries across various professional roles. The study demonstrates t…
Are brain waves the next unlock for physical AI?
EEG-based brain wave data emerges as a novel training modality for frontier physical AI systems. Researchers explore incorporating direct…
GH-ESD: Grounded Hypothesis-Driven Error Slice Discovery for Instance-Level Vision Tasks
GH-ESD is a novel framework for discovering systematic failure modes in instance-level vision tasks. The approach addresses the limitatio…
Optical Tech Would Update a Robot’s AI on the Fly
Cornell Tech researchers developed an optical receiver that directly alters processor memory using beamed light. The new technique bypass…