AGENTIC IDEA MANAGER

Better research starts
with better idea management.

AIM: Agentic Idea Management for Automated Research

Organize what you know. Decide what to try.
Make every research direction—and the evidence behind it—inspectable.

Hyeong Kyu Choi1,2, Bhavana Dalvi Mishra1, Jiefeng Chen1, Mihir Parmar1, Rui Meng1,
Chun-Liang Li1, Xiangru Tang1, Sharon Li2, Jinsung Yoon1, Tomas Pfister1

1 Google Cloud AI Research   2 University of Wisconsin–Madison

THE SEARCH SPACE
Research ideaResearch ideaResearch idea
Agentic Surrogate

Organize ideas · Estimate promise

Agentic Acquisition
EXPLOITEXPLORE

Dispatch · Solve · Expand

EXPERIMENTSImplement.
Evaluate.
Audit.
10AutoLab tasks
+1.6 ppSystem Optimization
+4.9 ppModel Development & CUDA
Up to 3.1×Faster to the best baseline score

Paper-reported results. Gains over ScientistOne; time-to-score comparison on Flash Attention. §5, Table 1 ↗ · Table 2 ↗ · Figure 4 ↗

01 / THE FRAMEWORK

A map of possibilities.
A reason for the next experiment.

AIM searches over explicit research ideas. Inspired by Bayesian optimization, it separates understanding the idea space from choosing where to spend the experimental budget.

AGENTIC SURROGATE

What looks promising?

Organize. Group ideas by semantic research direction, even when they come from different generation lineages. Rebuild the map as the pool grows.

Estimate. Rank clusters and unevaluated ideas using observed scores, evidence gaps, novelty, and implementation lessons.

Ordinal ranks express relative promise—not calibrated reward predictions.
AGENTIC ACQUISITION

What should we try next?

Dispatch. Choose explore/exploit actions at both the cluster and idea levels, then select concrete ideas for parallel solvers.

Solve & Expand. Implement and evaluate the selected ideas. Use audited evidence to refine, combine, repair, or introduce new ideas.

A promising direction can still contain an unfamiliar idea worth exploring.
SOLUTION AUDITOR

What did we actually test?

Validate. Check task validity and idea–solution alignment before results inform the next research decision.

Align. Discard invalid evidence. When the implementation differs from the proposed idea, reconstruct the idea to describe the mechanism actually evaluated.

Attribute scores and lessons to the method that was built.
RESOURCE PLANNER

How should we spend the budget?

Allocate. Choose the number of parallel solver branches for each iteration within the remaining experimental budget.

Adapt. Balance wider exploration with more sequential rounds, so new evidence can guide later experiments.

Adapt parallelism while preserving the total branch and execution budgets.
AIM framework: the search state connects the Agentic Surrogate’s Organize and Estimate operators, the Agentic Acquisition’s Dispatch, parallel Solvers and Expand operators, the Solution Auditor, and the Resource Planner.
The AIM framework: organize and estimate ideas, dispatch parallel experiments, audit outcomes, and allocate the remaining budget. Click to view the full-size figure.

02 / INTERPRETABILITY IN PRACTICE

Follow the decision.
Inspect the evidence.

Replay recorded research iterations across nine tasks and 27 runs. See the map, the rankings, the chosen actions, and what actually happened.

More to explore
Iteration
Reading this replay. Scores are the recorded score field on a 0–1 scale, not raw task accuracy. Ranks are relative; 1 is most promising. Cluster IDs are local to each iteration and may change meaning. Rationale text and lessons are agent-authored records, not independent verification. This explorer does not rerun experiments.

03 / RESULTS FROM THE PAPER

Better ideas. Better outcomes.

Mean scores across three runs. AIM leads the two task-group averages; individual task outcomes vary. Values below are transcribed from Tables 1 and 2, separately from the artifact replay.

System Optimization

AIM67.0
ScientistOne65.4

Model Development & CUDA

AIM55.8
ScientistOne50.9
Task scores (%) · mean ± standard error · 3 runs
TaskScientistOneAIMDifference (pp)

Selected comparator: ScientistOne, the strongest reported baseline by task-group average. This table is not a claim that AIM wins every task against every method; for example, AdaEvolve leads Data Selection IFEval. See the paper for the full baseline comparison.

WALL-CLOCK TIME ANALYSIS

How quickly does discovery happen?

On Flash Attention, AIM reaches ScientistOne’s best score up to 3.1× faster and reaches its own best score of 90.5% after 3.3 hours. Explore the supplied timing figures for all ten tasks.

Flash Attention: mean best-so-far score versus wall-clock time, alongside final scores and the times at which they are reached.

The left panel tracks mean best-so-far scores; the right panel marks final score levels and their attainment times. Time to a shared score threshold differs from time to each method’s own best score. Results vary by task; the 3.1× claim applies to Flash Attention. Figure 4 ↗ · Appendix D.3 ↗

RESEARCH YOU CAN FOLLOW

From a promising idea
to an inspectable experiment.

Read the method, explore the decisions, and trace the evidence.

BibTeX

to be released soon