RadeonVLA-Reflex SmolVLA-1K Model Card

Release status: trained checkpoint. The Physical-1K data and checkpoint are published; formal closed-loop evaluation of this checkpoint is still in progress and no final success rate is claimed in this revision.

Model

Intended use

Language-conditioned Franka fruit sorting in the project Genesis scene.

Inputs and outputs

  • Image inputs: observation.images.world + observation.images.wrist (320×240 RGB); renamed to camera1/camera2 for SmolVLA
  • State shape and ordering: 9-D qpos — panda_joint1..7, panda_finger_joint1..2
  • Language input: natural-language task instruction
  • Action type, shape, and ordering: 9-D absolute joint positions (same order as state)
  • Control frequency: 20 Hz
  • Action chunk size: determined by SmolVLA checkpoint config

Training

Item Value
Radeon GPU AMD Radeon Graphics, gfx1100, 51.5 GB VRAM
ROCm 7.2.1 / HIP runtime 7.2.53211
PyTorch 2.9.1+rocm7.2.1
Precision FP32 (use_amp=false)
Batch size 4
Gradient accumulation Not configured; one optimizer update per batch
Training steps 20,000 (80,000 sampled frames)
Training time 45m 34s
Peak training memory reported by LeRobot 2.22 GB
Final logged minibatch loss 0.091

Evaluation

The final 20-task held-out evaluation has not yet been completed for this Physical-1K checkpoint. This release therefore does not claim a closed-loop success rate. The checkpoint has passed a forced-offline SmolVLAPolicy.from_pretrained load test after vendoring its VLM tokenizer and processor assets. A later revision will bind the formal evaluation JSON, Normal-vs-Reflex comparison, inference latency, and observed failure modes without replacing the immutable training metadata above.

Limitations

Training and evaluation are limited to Genesis simulation. Coverage includes the registered language and task suite, dual RGB cameras at 320×240, and 9-D absolute joint actions. Real-robot transfer is outside the current evaluation scope. The published card lists failure modes observed in the final evaluation run.

License note

The upstream lerobot/smolvla_base repository did not declare license metadata when the base revision above was frozen. This model card therefore uses Hugging Face's other marker instead of inventing a permissive license. The Physical-1K training dataset is CC BY 4.0 and its attribution requirements remain applicable to the dataset and rendered examples. Users must review the upstream base-model terms before redistribution or commercial use.

Downloads last month
13
Safetensors
Model size
0.5B params
Tensor type
F32
·
BF16
·
Video Preview
loading

Model tree for a3124371940/radeonvla_reflex_smolvla_1k

Finetuned
(7354)
this model