ACT · carrot-in-pot · EEF-delta (state 16 / action 7) — checkpoint 10k

Action Chunking Transformer trained on the real UR7e "Put carrot in pot" demonstrations (54 GELLO-teleop takes, 30 fps) in the EEF-delta action space (eef_delta_v1). This is the 10k-step checkpoint, picked as the least-overfit of 10k..50k — all were within 0.04 mm of each other on the open-loop metric.

Joint-space siblings for this task live in the sim evaluation (sim_collect/eval), the IFQL policy in Bigenlight/carrot-in-pot-ifql.

Observation / action space (eef_delta_v1)

  • observation.state 16-D = [q1..q6 (rad, UR order), tcp_x, tcp_y, tcp_z (m, base_link), r11, r21, r31, r12, r22, r32 (first two columns of the TCP rotation — continuous 6-D rep), grip_pos (0=open..1=closed)]. TCP = ur_kin.fk(q) (base_link→tool0) + 0.174 m along flange +Z.
  • action 7-D = [dx, dy, dz, drx, dry, drz, grip_cmd]: the achieved TCP motion between consecutive 30 fps frames (dp = p_{t+1}-p_t, drot = so3_log(R_{t+1} R_t^T), base frame, rotation left-multiplied), gripper = absolute recorded command 0..1. Deploy inverts it exactly (p_target = p_live + dp, R_target = so3_exp(drot) R_live, analytic IK with branch locking).
  • Cameras: observation.images.cam1 (scene), cam2 (wrist), RGB 720×1280 in the dataset, resized to 360×640 at train time (image_transforms.resize) — resize the same way at inference.
  • Backbone ResNet18 (ImageNet), chunk_size = n_action_steps = 100, MEAN_STD normalization, ~51.6M params.

Training

  • Dataset: carrot_in_pot_eef_lerobot_v3 — a local LeRobot v3 re-export of Bigenlight/carrot_in_pot_lerobot_v3 (54 episodes / 17,085 frames after dropping the stale tail; joints shifted by the recorder's per-take τ≈0.90 s cache lag and linearly re-interpolated). The EEF re-export is not yet on the Hub (train_config names it Bigenlight/carrot_in_pot_eef_lerobot_v3).
  • lerobot-train, batch 8, seed 1000, 50k steps configured (save_freq 10k), eval_split 0.111 (held-out episodes 48–53), single RTX A4000 (kanu). Job act_carrot_eef.

Held-out results (open-loop, episodes 48–53, k=30)

checkpoint pos MAE grip acc chunk-30 cumulative error
10k (this) 0.82–0.86 mm (all ckpts) 0.95 36.6–38.0 mm vs 65.6 mm zero-motion baseline

All checkpoints 10k–50k are statistically indistinguishable on this metric; 10k was chosen as least-overfit. lerobot's own eval_loss is computed on un-resized 720p and was not used.

Status

Real-robot closed-loop evaluation: not yet run (the deploy path is EEF mode of gello_policy/policy_leader_node + eef_space.apply_delta; the shipped ZMQ servers are joint-space 7/7 and refuse this checkpoint's 16-D state). Provenance: gello_software branch feat/carrot-eef-il (converter scripts/dataset/convert_carrot_to_lerobot_eef.py, validator 71/71 PASS).

Downloads last month
18
Safetensors
Model size
51.7M params
Tensor type
F32
·
Video Preview
loading