Skip to main content

byEmulated

Autoresearch
Bench

Evaluating language model agents
on long-horizon autonomous research

Tasks

17 research tasks across LLM training, inference, optimization, scientific ML and classical ML. Each one gives an agent a prepared workspace and a measurable objective, and is graded by a hidden verifier. Every task page covers the background, what the agent is asked to do, how it is evaluated, and what agents have done so far.

Attribution sweep speedup

Make the backward sweep of our circuit-tracing library substantially faster on CPU, with every attribution graph exactly unchanged.

Optimization
avg 0.72 · best 0.95Claude Opus 4.8

Exact speculative decoding speedup

Make Qwen3.5-9B decode substantially faster with its own multi-token-prediction head, without changing a single emitted token.

Inference
avg 0.77 · best 0.80Claude Opus 4.8

DBSCAN clustering speedup

Build a fast torch-native 2D DBSCAN from a written spec, with hidden datasets matched exactly against a reference engine.

Optimization
avg 0.60 · best 1.00Claude Fable 5

Triton training-kernel speedup

Make a training codebase measurably faster by writing Triton kernels for its hot paths, without changing what it computes.

Optimization
avg 0.83 · best 1.00Claude Opus 4.8

Degraded-document OCR

Train an OCR model for degraded text from scratch on CAPTCHA corpora generated inside the offline environment.

Classical ML
avg 0.60 · best 0.64Claude Opus 5

Diffusion checkpoint reconstruction

Reconstruct a hidden Stable Diffusion checkpoint by merging three given checkpoints to match its image-generation behavior.

LLM training
avg 0.17 · best 0.25Claude Opus 5

rav1d decoder performance

Close the decode-performance gap between the memory-safe Rust AV1 decoder and its hand-optimized C original.

Optimization
avg 0.56 · best 0.82Claude Opus 5

Cell instance segmentation

Build a model that predicts a mask for every cell in light-microscopy images, graded on a sealed private split with a submission quota.

Scientific ML
avg 0.34 · best 0.47Claude Opus 5

Pasture biomass regression

Predict five dry-biomass components from top-view pasture photos.

Scientific ML
avg 0.61 · best 0.65Claude Opus 5

Copy-move forgery detection

Classify images as authentic or manipulated and segment every duplicated region.

Scientific ML
avg 0.60 · best 0.62Claude Opus 5

Smartphone GNSS positioning

Recover a phone's driving track from raw multi-constellation GNSS measurements.

Scientific ML
avg 0.57 · best 0.61Claude Opus 5

Freezing-of-gait detection

Detect freezing-of-gait events at every timestep of lower-back accelerometer recordings.

Scientific ML
avg 0.38 · best 0.39Claude Opus 5

Contrail segmentation

Segment aviation contrails in nine-channel infrared satellite frames.

Scientific ML
avg 0.60 · best 0.62Claude Opus 5

Single-cell modality prediction

Predict one single-cell measurement modality from another.

Scientific ML
avg 0.37 · best 0.38Claude Fable 5

RNA 3D structure prediction

Predict five candidate 3D structures per RNA sequence, scored best-of-five.

Scientific ML
avg 0.33 · best 0.48Claude Opus 5

ECG image digitization

Reconstruct twelve ECG traces from a photographed or scanned printout.

Scientific ML
avg 0.12 · best 0.13Claude Opus 5

Scientific image forgery detection

Find and segment copy-moved regions in biomedical research images.

Scientific ML
avg 0.36 · best 0.46Claude Opus 5