Task Sample · Robotics

Isaac Lab PegInsert Reward Search

Design a better training reward so a robot learns to align and insert a peg more reliably, with the simulator and training budget held fixed.

Best score per runInsertion success rate (0–1)
00.653
Model · harnessBestSubmissionsRuntime
GPT-5.6 SolCodex · xhigh0.60453224.23 h
Claude Opus 5Claude Code · max0.60163224.37 h

Absolute terminal insertion success rate after fresh fixed-budget PPO training; higher is better. Each submission tests up to two reward recipes and takes the better result. The dashed 0.378 reference was measured during task validation. Runtime is reported agent-process wall time, including shutdown; the fixed homepage budget is 24 hours.

Teaser figure from Isaac Lab
Source project · Isaac Lab · source ↗

The climb

Both agents search for a better training reward. Each dot is the best recipe from one scored Judge submission, and the step lines show the best result reached so far. The dashed line is the original reward's validation-measured success rate. Four submissions have no official score; they remain in the tables as dashes, not zero-valued points.

0.2770.3740.470.5670.664Baseline · 0.378 success rate081624.4Elapsed time (h)Insertion success rateGPT-5.6 Sol · submission 1 · 0.5312 · 0.784 h, Judge result recordedGPT-5.6 Sol · submission 3 · 0.5029 · 1.58 h, Judge result recordedGPT-5.6 Sol · submission 4 · 0.4932 · 2.34 h, Judge result recordedGPT-5.6 Sol · submission 5 · 0.5381 · 3.12 h, Judge result recordedGPT-5.6 Sol · submission 6 · 0.5078 · 3.89 h, Judge result recordedGPT-5.6 Sol · submission 7 · 0.5166 · 4.66 h, Judge result recordedGPT-5.6 Sol · submission 8 · 0.5674 · 5.44 h, Judge result recordedGPT-5.6 Sol · submission 9 · 0.4844 · 6.2 h, Judge result recordedGPT-5.6 Sol · submission 10 · 0.5283 · 6.99 h, Judge result recordedGPT-5.6 Sol · submission 11 · 0.5459 · 7.77 h, Judge result recordedGPT-5.6 Sol · submission 12 · 0.4766 · 8.55 h, Judge result recordedGPT-5.6 Sol · submission 13 · 0.5059 · 9.33 h, Judge result recordedGPT-5.6 Sol · submission 14 · 0.5859 · 10.1 h, Judge result recordedGPT-5.6 Sol · submission 15 · 0.4834 · 10.9 h, Judge result recordedGPT-5.6 Sol · submission 16 · 0.5303 · 11.7 h, Judge result recordedGPT-5.6 Sol · submission 17 · 0.4902 · 12.5 h, Judge result recordedGPT-5.6 Sol · submission 18 · 0.5557 · 13.3 h, Judge result recordedGPT-5.6 Sol · submission 19 · 0.4697 · 14.1 h, Judge result recordedGPT-5.6 Sol · submission 20 · 0.5039 · 14.9 h, Judge result recordedGPT-5.6 Sol · submission 21 · 0.5264 · 15.7 h, Judge result recordedGPT-5.6 Sol · submission 22 · 0.5068 · 16.4 h, Judge result recordedGPT-5.6 Sol · submission 23 · 0.6045 · 17.2 h, Judge result recordedGPT-5.6 Sol · submission 24 · 0.5762 · 18 h, Judge result recordedGPT-5.6 Sol · submission 25 · 0.5312 · 18.8 h, Judge result recordedGPT-5.6 Sol · submission 26 · 0.5176 · 19.5 h, Judge result recordedGPT-5.6 Sol · submission 27 · 0.4990 · 20.3 h, Judge result recordedGPT-5.6 Sol · submission 28 · 0.4824 · 21.1 h, Judge result recordedGPT-5.6 Sol · submission 29 · 0.5342 · 21.9 h, Judge result recordedGPT-5.6 Sol · submission 30 · 0.4961 · 22.7 h, Judge result recordedGPT-5.6 Sol · submission 31 · 0.4902 · 23.4 h, Judge result recordedGPT-5.6 Sol · 0.6045Claude Opus 5 · submission 1 · 0.3936 · 0.905 h, Judge result recordedClaude Opus 5 · submission 2 · 0.4648 · 1.73 h, Judge result recordedClaude Opus 5 · submission 3 · 0.5605 · 2.51 h, Judge result recordedClaude Opus 5 · submission 4 · 0.4609 · 3.33 h, Judge result recordedClaude Opus 5 · submission 5 · 0.5107 · 4.12 h, Judge result recordedClaude Opus 5 · submission 6 · 0.4365 · 4.93 h, Judge result recordedClaude Opus 5 · submission 7 · 0.5557 · 5.7 h, Judge result recordedClaude Opus 5 · submission 8 · 0.4561 · 6.48 h, Judge result recordedClaude Opus 5 · submission 9 · 0.5010 · 7.25 h, Judge result recordedClaude Opus 5 · submission 10 · 0.5195 · 8.03 h, Judge result recordedClaude Opus 5 · submission 11 · 0.5439 · 8.82 h, Judge result recordedClaude Opus 5 · submission 12 · 0.4307 · 9.6 h, Judge result recordedClaude Opus 5 · submission 13 · 0.4951 · 10.4 h, Judge result recordedClaude Opus 5 · submission 14 · 0.6016 · 11.2 h, Judge result recordedClaude Opus 5 · submission 15 · 0.5869 · 12 h, Judge result recordedClaude Opus 5 · submission 16 · 0.6016 · 12.8 h, Judge result recordedClaude Opus 5 · submission 18 · 0.6016 · 13.6 h, Judge result recordedClaude Opus 5 · submission 19 · 0.5303 · 14.4 h, Judge result recordedClaude Opus 5 · submission 20 · 0.4404 · 15.1 h, Judge result recordedClaude Opus 5 · submission 21 · 0.5020 · 15.9 h, Judge result recordedClaude Opus 5 · submission 22 · 0.5625 · 16.7 h, Judge result recordedClaude Opus 5 · submission 23 · 0.4756 · 17.5 h, Judge result recordedClaude Opus 5 · submission 24 · 0.3359 · 18.2 h, Judge result recordedClaude Opus 5 · submission 25 · 0.4023 · 19 h, Judge result recordedClaude Opus 5 · submission 26 · 0.6016 · 19.8 h, Judge result recordedClaude Opus 5 · submission 27 · 0.6016 · 20.6 h, Judge result recordedClaude Opus 5 · submission 28 · 0.6016 · 21.3 h, Judge result recordedClaude Opus 5 · submission 29 · 0.6016 · 22.1 h, Judge result recordedClaude Opus 5 · submission 30 · 0.6016 · 22.8 h, Judge result recordedClaude Opus 5 · submission 31 · 0.6016 · 23.6 h, Judge result recordedClaude Opus 5 · 0.6016
The agent runs · 2 runs · submissionrunning bestbaseline · validation measuredTime since run start · points mark recorded Judge results

The task

A robot must align a peg with a socket and insert it. The agent does not control the robot directly: it designs the reward that guides reinforcement learning. A better recipe should teach the same policy to succeed more often within the same training budget.

Environment

Reference baseline

The dashed 0.378 line is the original Factory reward, measured during task validation: 387 successful terminal episodes out of 1,024. It rewards keypoint alignment, engagement, and successful insertion under the same fixed training and evaluation protocol used for candidates.

This task uses shorter episodes and wider reset randomization than the stock environment. The released reference checkpoint and the paper's results therefore are not comparable baselines. Neither published agent run submitted the original reward unchanged, and Judge does not rerun it alongside every candidate.

Research loop

  1. Form a hypothesis about which behavior the reward should encourage or discourage.
  2. Express it as a valid reward graph, optionally with a second recipe to compare.
  3. Submit the recipes; Judge freshly trains and evaluates each one.
  4. Compare the per-recipe success rates, then keep, revise, or replace the idea.

What the agent may change

What stays fixed

Evaluation

Judge validates each recipe, trains a fresh policy for the complete fixed budget, retains the selected checkpoint, and measures successful terminal insertions / 1,024 evaluation episodes. The score lies between zero and one. With two recipes, the submission score is the better of the two, not their average.

Feedback includes each recipe's aggregate score, completion counts, selected epoch, timing, and bounded diagnostics. Invalid or interrupted submissions have no official score. In particular, the final in-flight round in each run retained candidate metrics but closed without a submission score when the run budget expired; those diagnostics are not substituted for a scored result. A completed evaluation with no successful insertions would instead be a genuine zero.

The Isaac Lab paper explains the underlying Factory tasks; this task's shorter episodes, wider resets, reward-search interface, and fixed evaluation protocol are defined by the linked task instructions.

Agent runs

Each figure shows one agent's scored submissions and running best. Both runs contain 32 submissions, of which 30 have an official score. Open the table beneath each figure for every submission; a dash means no score was recorded.

0.4450.4910.5370.5830.6290.6045round 132Insertion success rateGPT-5.6 Sol · Codex
GPT-5.6 Sol · 32 submissions · best 0.6045
GPT-5.6 Sol — every submission, in order (32)
#Score
10.5312
2
30.5029
40.4932
50.5381
60.5078
70.5166
80.5674
90.4844
100.5283
110.5459
120.4766
130.5059
140.5859
150.4834
160.5303
170.4902
180.5557
190.4697
200.5039
210.5264
220.5068
230.6045
240.5762
250.5312
260.5176
270.4990
280.4824
290.5342
300.4961
310.4902
32
0.2880.3780.4690.5590.6490.6016round 132Insertion success rateClaude Opus 5 · Claude Code
Claude Opus 5 · 32 submissions · best 0.6016
Claude Opus 5 — every submission, in order (32)
#Score
10.3936
20.4648
30.5605
40.4609
50.5107
60.4365
70.5557
80.4561
90.5010
100.5195
110.5439
120.4307
130.4951
140.6016
150.5869
160.6016
17
180.6016
190.5303
200.4404
210.5020
220.5625
230.4756
240.3359
250.4023
260.6016
270.6016
280.6016
290.6016
300.6016
310.6016
32

Links