PostTrainBench⁰: Can LLM Agents Automate LLM Post-Training Without Gradients?

Yuqiao Tan · Minzheng Wang · Shizhu He · Jun Zhao · Kang Liu
Institute of Automation, Chinese Academy of Sciences

PostTrainBench⁰ tests whether autonomous LLM agents can post-train language models without gradients, evaluating whether agents can discover generalizable optimization strategies using only black-box scalar feedback.

Benchmark setup

Each four-hour run provides one base checkpoint, eight inference workers, and seven conflicting tasks: Countdown, GSM8K, MATH-500, OlympiadBench, MBPP, ROCStories, and USPTO-50K. Gradients, evaluator edits, model replacement, and inference-time ensembles are forbidden.

Core findings

Across 13 frontier agents, all systems improve visible scores and demonstrate experimental reasoning such as multi-scale tuning and rollback. However, a simple tuned evolution strategy substantially outperforms every agent, and 9 of 13 agents fail to generalize on held-out test sets. These results reveal that current agents excel at operational execution but cannot yet achieve true algorithmic discovery.

Code and benchmark materials