Instructions to use aist-eart/toast with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use aist-eart/toast with Transformers:
# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("aist-eart/toast", device_map="auto") - Notebooks
- Google Colab
- Kaggle
TOAST: Action Tokenizer for Autoregressive VLAs
Project Page | Paper | Github
TOAST is a stochastic robot action tokenizer for autoregressive vision-language-action models such as π₀-FAST. An action chunk [time, dim] (normalized to [-1, 1]) is discretized with a DCT along time, flattened by action dimension, and segmented into a token sequence by a unigram language model.
The interface follows the FAST tokenizer: AutoProcessor loads it with trust_remote_code=True, tokenizer(actions) tokenizes, tokenizer.decode(tokens) detokenizes, and tokenizer.fit(actions) builds a new vocabulary.
This repository ships the vocabulary used in our LIBERO experiments: 512 subwords over DCT coefficients (scale 10), fitted on 1M end-effector action chunks (6D pose + gripper, 16 steps) from DROID (spm_toast_droid-eef.model).
Installation
pip install transformers sentencepiece scipy numpy
import numpy as np
from transformers import AutoProcessor
tokenizer = AutoProcessor.from_pretrained("aist-eart/toast", trust_remote_code=True)
actions = np.random.rand(8, 10, 7) * 2 - 1 # [batch, time, dim] action chunks, normalized to [-1, 1]
tokens = tokenizer(actions) # list[list[int]], the most likely segmentation
tokens = tokenizer(actions, sample=True) # a sampled segmentation (use during training)
decoded = tokenizer.decode(tokens, time_horizon=10, action_dim=7)
Segmentation sampling is controlled by sample, alpha (smoothing) and nbest_size (number of candidate
segmentations; -1 for all), which can also be set when loading:
AutoProcessor.from_pretrained(..., sample=True, alpha=0.1, nbest_size=64).
Building Tokenizer
A vocabulary is fitted on a corpus of normalized action chunks, each of shape [time, dim] (the horizons may differ):
action_data = ... # list of arrays or an array [num_chunks, time, dim], normalized to [-1, 1]
tokenizer = AutoProcessor.from_pretrained("aist-eart/toast", trust_remote_code=True)
new_tokenizer = tokenizer.fit(action_data, scale=10, vocab_size=512)
new_tokenizer.save_pretrained("my_tokenizer") # processor_config.json, the SentencePiece model and toast.py
new_tokenizer = AutoProcessor.from_pretrained("my_tokenizer", trust_remote_code=True)
fit also supports timestep-major flattening via order="timestep" (which flattens the coefficients one timestep at a time, as in FAST), BPE tokenization via model_type="bpe", and two alternative quantization schemes: binning via quantization="binning" (e.g., quantizer_kwargs={"bin_size": 256, "use_bin_center": True}) and BEAST via quantization="beast" (e.g., quantizer_kwargs={"num_basis": 5, "gripper_zero_order": True}).