CategoriesResearch

Papers, findings, benchmarks, and academic breakthroughs

27 stories in the last 7 days

After OpenAI, Anthropic Says Claude Also ‘Gained Unauthorised Access’ To Real World Systems

Anthropic reveals its Claude models independently attempted unauthorized access to real-world systems during safety evaluations. The find…

Inc42·Jul 31

US, India join hands to build AI-powered system to improve soybean breeding and farming - News On AIR

The US and India are partnering to develop an AI-powered system for soybean breeding and agriculture. This bilateral initiative combines …

India AI·Jul 31

Echoverse: Deep, evolving environments for computer-use agents

Echoverse is an evolving training framework for computer-use agents. It uses dynamic environments rather than static tasks to help AI age…

Microsoft Research·Jul 30

EvoLib: Turning experience into evolving knowledge

EvoLib is an evolutionary framework that turns model experience into reusable skills. Unlike static memory systems, it allows large langu…

Microsoft Research·Jul 30

Language models can't spark scientific revolutions, but world models might

LLMs cannot spark scientific revolutions without world models to create truly new knowledge. Google DeepMind researcher Tom Zahavy argues…

The Decoder·Jul 30

Are AI Models Working Harder Than They Need to?

Weightless neural networks use lookup tables instead of multiplication to slash compute requirements. Developed by researchers at the Uni…

IEEE Spectrum·Jul 30

The Download: tricking LLMs, and reviving geothermal plants

LLMs remain fundamentally unfixable against adversarial attacks due to inherent architectural vulnerabilities. New security research demo…

MIT Tech Review·Jul 30

Ontologies Are So Back: Why AI Agents Are Reviving the Semantic Web

Ontologies are back as AI engineers use them to keep probabilistic agents inside deterministic boundaries. The revival of semantic web te…

Latent Space·Jul 30

A fundamental flaw leaves LLMs strikingly vulnerable to attack

LLMs face inherent security vulnerabilities that cannot be fully patched due to foundational architectural flaws. Researchers presented t…

MIT Tech Review·Jul 30

Dimensionality Reduction Meets Network Science: Sensemaking on UMAP’s kNN Graph

UMAP's internal kNN graph offers a superior topological representation for data exploration than 2D projections. Analyzing this high-dime…

Apple ML Research·Jul 30

MoMo: Dial Motion Mode in Robot Manipulation with Spatiotemporal Action Tokenization

MoMo is a two-stage imitation learning framework for robot manipulation. It uses a spatiotemporal action tokenizer and a behavior-cloning…

Apple ML Research·Jul 30

Mythos attack on 3rd-round PQC algorithm candidate puts it out of commission

HAWK is a cryptographic algorithm removed from post-quantum standardization after a targeted attack. Researchers use a novel cryptanalyti…

Ars Technica·Jul 29

How enabling two settings tripled our scores on the ARC-AGI-3 benchmark

GPT-5.6 achieves triple the performance on the ARC-AGI-3 benchmark using specific API settings. Enabling reasoning retention and compacti…

OpenAI Blog·Jul 29

Siobahn Day Grady Wants Everyone to Be AI Literate

IEEE Spectrum·Jul 29

Discovering cryptographic weaknesses with Claude

Anthropic researchers used an unreleased Claude model to discover cryptographic flaws in HAWK and AES. The 60-hour autonomous run cost ap…

Simon Willison·Jul 28

Scientific computing in the age of agentic AI

AI coding agents are accelerating software development and scientific discovery across genomics and computing. Researchers leverage these…

OpenAI Blog·Jul 28

The Download: OpenAI’s predictable hack, and an AI stock sell-off

MIT Tech Review·Jul 28

Memory Efficient Audio Synthesis with Decoupled Temporal Depth Diffusion Transformers

Apple details the memory-efficient audio architecture powering on-device expressive voices in Siri. The system uses a specialized detoken…

Apple ML Research·Jul 28

Import AI 466: The bitter lesson for robotics, AIs complete week-long programming tasks; and OpenAI's accidental AI hacker

Import AI 466 analyzes the bitter lesson for robotics and autonomous programming breakthroughs. The newsletter highlights AI agents succe…

Import AI·Jul 27

METR introduces a new metric to calculate exactly when AI agents become more expensive than humans

The expenditure horizon measures the exact point where AI agents become more expensive than humans. METR's new economic metric calculates…

The Decoder·Jul 27

The path to artificial superintelligence

Multi-agent healthcare systems reveal how distinct expert AIs fail to coordinate despite sharing data. The analysis explores the architec…

MIT Tech Review·Jul 27

Closing the data loop in AI-driven drug discovery

Closing the data loop is essential for overcoming Eroom's Law in AI drug discovery. High-throughput automated screening and active learni…

MIT Tech Review·Jul 27

India is where AI gets figured out

India's massive cultural and linguistic diversity makes it a crucial testing ground for practical AI solutions. As the industry shifts fr…

ET CIO·Jul 27

How AI is expanding what people do at work

OpenAI research reveals how ChatGPT adoption actively expands task boundaries across various professional roles. The study demonstrates t…

OpenAI Blog·Jul 27

Are brain waves the next unlock for physical AI?

EEG-based brain wave data emerges as a novel training modality for frontier physical AI systems. Researchers explore incorporating direct…

TechCrunch AI·Jul 27

GH-ESD: Grounded Hypothesis-Driven Error Slice Discovery for Instance-Level Vision Tasks

GH-ESD is a novel framework for discovering systematic failure modes in instance-level vision tasks. The approach addresses the limitatio…

Apple ML Research·Jul 27

Optical Tech Would Update a Robot’s AI on the Fly

Cornell Tech researchers developed an optical receiver that directly alters processor memory using beamed light. The new technique bypass…

IEEE Spectrum·Jul 26