Robotics
LeRobot
Safetensors
imitation-learning
act
b-spline
so101
bspline_act

pickblock_bspline_act

An ACT policy that predicts B-spline trajectory segments instead of discrete action chunks, trained on SO-101 black-cube picking. This is the "Reg.+BSP" variant from B-spline Policy: Accelerating Manipulation Policies via B-spline Action Representations.

The policy outputs a curve, so executing a(n·t) replays the same trajectory geometry n times faster without retraining. Speed-up is an inference flag, verified exact to 0.000e+00 deviation at 2× and 4×.

This model is trained on 13 episodes and does not generalize well. Held-out error is 22.8° against 6.2° on training data — a 3.7× gap that never closed across 40,000 steps. It has never been run on hardware. Treat it as a pipeline validation artifact, not a working policy. See Evaluation and Limitations.

Download

hf download Dimios45/pickblock_bspline_act --local-dir ./pickblock_bspline_act

The repository root holds the recommended checkpoint, so --policy.path also accepts the repo id directly. Individual checkpoints live under checkpoints/<step>/.

Run it

lerobot-rollout \
    --strategy.type=base \
    --policy.path=./pickblock_bspline_act/checkpoints/010000/pretrained_model \
    --policy.speed_up=1.0 \
    --device=cuda \
    --robot.type=so101_follower \
    --robot.port=/dev/ttyACM0 --robot.id=FOLLOWER \
    --robot.cameras='{front: {type: opencv, index_or_path: 0, width: 640, height: 360, fps: 30}, ee: {type: opencv, index_or_path: 4, width: 640, height: 360, fps: 30}}' \
    --task="Grab the black cube" \
    --fps=30 --duration=60

--policy.path takes either a Hub repo id (namespace/name) or a local directory containing config.json — which is the pretrained_model/ subdirectory, not the step directory above it. There is no --policy.subfolder flag.

Two things differ from the cup-stacking model — check both

Resolution is 640×360, not 640×480. The source demonstrations were recorded at 1920×1080; they were downscaled to 640×360 for training, preserving the 16:9 aspect ratio. Your cameras must be configured to match.

The cameras are front and ee, not front and gripper:

Key Shape
observation.images.front (3, 360, 640)
observation.images.ee (3, 360, 640)
observation.state (6,)

Joint order is shoulder_pan, shoulder_lift, elbow_flex, wrist_flex, wrist_roll, gripper, in degrees. Swapping the two cameras produces confident but wrong behaviour rather than an error.

To try the speed-up, raise --policy.speed_up to 2.0 or 4.0. The paper reaches 4× on cube picking specifically, so this task should tolerate it better than precise-placement tasks — but validate at 1× first, and note this policy has not been shown to complete the task at any speed.

Requirements

This policy type is not in upstream LeRobot. You need a checkout containing the bspline_act policy, installed with the scipy extra:

uv pip install -e ".[feetech,bspline]"

scipy>=1.15 is required for interpolate.generate_knots.

Unnormalization statistics are stored as buffers inside model.safetensors, so each checkpoint is self-contained — you do not need the training dataset on the robot machine.

Checkpoints

Step Epoch Train error Held-out error Eval loss
10000 58 6.19° 22.81° 0.3903 recommended — best
20000 116 4.67° 24.15° 0.3986
30000 174 3.90° 24.77° 0.3996
40000 234 3.44° 26.22° 0.4030 final, worst held-out

Held-out error gets monotonically worse after step 10,000 while training error keeps falling. Use 10000. Later checkpoints are included only for reproducibility — they are strictly worse policies.

Error is mean absolute joint error in degrees, worst joint, over each predicted segment's horizon (~0.76 s), measured by decoding the spline and comparing against ground-truth actions.

Training

Steps 40,000
Batch size 32
Epochs 234
Dataset atharva-pantheon/pick_block_bowl → downscaled → idle-trimmed → B-spline converted
Episodes 13 train / 3 held out
Frames 6,422 (from 10,342; 37.9% boundary idle removed)
Optimizer AdamW, lr 2e-5 (backbone 1e-5), wd 1e-3
Precision bf16 autocast + channels_last
Augmentation ColorJitter (brightness/contrast/saturation/hue/sharpness), up to 3 per frame
Hardware 1× RTX 4090, ~2 h 45 min total
Seed 1000

Augmentation and 10× weight decay were added specifically to combat overfitting. They slowed the fit to training data and moved held-out loss by 0.0003. An unaugmented control run is available on request; the two are indistinguishable on unseen data.

Preprocessing

Source video is 1920×1080 AV1, which is unusable for ACT directly — a ResNet18 stride-32 feature map is 34×60, putting over 4,000 tokens into the transformer encoder across two cameras. Transcoded to 640×360 H.264, which also sped up random-access decode from 40 to 183 frames/s per worker.

Leading and trailing motionless frames were then removed: 46 s of lead-in and 27 s of trailing idle. This dataset is unusually idle-heavy — 37.9% of frames — including 20 s and 32 s dead stretches at the ends of two episodes where recording continued after the task finished.

Motion is measured as peak-to-peak joint range over a ±0.5 s window rather than frame-to-frame delta, because isolated jitter samples inside a dead stretch defeat a delta test and silently leave the idle in.

Architecture

Standard ACT with one change: the decoder emits a B-spline parameter matrix rather than an action chunk.

(n_knots, 1 + action_dim) = (16, 7) = 112 values per prediction
   column 0   knot vector, in source-frame units, 0 = "now"
   columns 1: control points, one per joint
Params 51.6 M
Vision backbone ResNet18 (ImageNet init, fine-tuned)
dim_model / chunk_size 512 / 16
VAE enabled, kl_weight 10.0
B-spline degree 3 (cubic, C² continuous)
bspline_chunk_size 10
Fitting tolerance ε 0.2 (degrees)
Max knot span 6 frames

The paper uses a DINOv2 encoder for its regression variant; this port uses LeRobot's ResNet18, which is what LeRobot's ACT supports. The paper reports no backbone ablation, and holds the encoder fixed between its baseline and its B-spline variant, so none of its claims depend on that choice.

Segments span 0.50–1.10 s (0.76 s mean). The network runs only when a segment is exhausted, which decouples policy rate from control rate.

On max_knot_span

This dataset is why the flag exists. Adaptive knot fitting covers a motionless stretch with one enormous knot span, and uncapped fitting here produced segments with a p95 of 18.4 seconds — the policy would commit to eighteen seconds of open-loop motion without looking at a camera. Capping span width at 6 frames bounds that at chunk_size × 6 = 2 s, costing 0.04× compression.

After correct idle-trimming the uncapped p95 drops to 1.17 s on its own, so the dead recording time was the real cause. The cap remains as a cheap safety rail.

Evaluation

Offline only. No hardware evaluation has been performed.

Train (13 eps) Held-out (3 eps)
Best checkpoint (10k) 6.19° 22.81°
Final (40k) 3.44° 26.22°

Reference points:

  • Inter-demonstration spread: 54.83° — how far apart two human demonstrations of this task are, phase-aligned. At 22.8° the policy is well inside that, so it is extracting real per-episode information from the cameras rather than replaying an average trajectory. This is the one genuinely positive result.
  • Hold-still baseline: 16.23°. The policy scores worse than this. Over a 0.76 s horizon on a slow task, "don't move" is a decent short-horizon predictor while being a useless policy, so this is not the damning comparison it looks like — but combined with the widening train/eval gap it is not encouraging either.

Held-out loss plateaued at ~0.39 by step 2,000 and never improved across the remaining 38,000 steps.

Why the held-out number is a pessimistic bound

The split holds out the last 3 of 16 episodes. Inspecting the raw recording:

  • 8 of 15 episode boundaries are contiguous to under 1° (several to exactly 0.00°), so episodes were largely cut from a continuous teleoperation stream. Adjacent episodes are correlated, meaning an interleaved split would leak.
  • There is a 97.67° discontinuity between episodes 9 and 10, and episodes 10, 13, 14 and 15 all end with the gripper open (8–16°) while the other twelve end closed (~2°).
  • All three held-out episodes sit after that change point, and two of them start with the gripper open — a state the model never sees at t=0 in training.

So the tail split is biased pessimistic and an interleaved split would be biased optimistic. With 16 semi-continuous episodes there is no clean holdout. 22.8° is a lower bound on quality, not a verdict.

Limitations

  • 13 training episodes is too few for this task. This is the dominant limitation and no hyperparameter fixed it. 40–50 demonstrations with varied cube placement and an independent reset between takes would address both the training and the measurement problem.
  • No hardware evaluation. Task success rate is unknown, and given the above it should not be assumed to be non-zero.
  • Held-out split is confounded — see above.
  • Resolution and camera names differ from the cup-stacking model. Check both.
  • Camera assignment is silent when wrong.

Citation

@article{han2026b,
  title={B-spline Policy: Accelerating Manipulation Policies via B-spline Action Representations},
  author={Han, Xiaoshen and Xiong, Haoyu and Chen, Haonan and Liu, Chaoqi and
          Torralba, Antonio and Zhu, Yuke and Du, Yilun},
  journal={arXiv preprint arXiv:2607.09648},
  year={2026}
}
Downloads last month
-
Safetensors
Model size
51.6M params
Tensor type
F32
·
Video Preview
loading

Dataset used to train Dimios45/pickblock_bspline_act

Papers for Dimios45/pickblock_bspline_act