Vector · The guide
The model
The one network every entry fills, vector_v1.1: what it is shown, how the demonstration and the live frames flow through it, and the actions it returns.
6 min read
One architecture
Every submission is the weights of vector_v1.1, Robotensor's demonstration-prompted diffusion policy for two arms, designed for this competition: a CLIP ViT-B/16 image encoder per camera, a six-layer transformer decoder that reads the live frames against the demonstration, and a 1-D diffusion U-Net that turns what it read into actions. It has 781 M parameters in 753 float32 tensors.
The validator builds the network with its own code and loads your weights into it, so the architecture is fixed and the weights are yours: what a duel compares is training and data.
The architecture, end to end
The network has two paths that share one frame tokenizer.
- The prompt is built once per unit, when the demonstration arrives. One frame is kept every 15 steps; each kept frame becomes nine tokens (one per camera, six for the arms' pose), is joined by an embedding of the 15 actions that followed it, and is pooled into a single prompt token. Four learned sink tokens go in front. The result is the decoder's memory.
- The policy runs at every prediction. The current frame and the one before it become 18 tokens, which a six-layer decoder passes through self-attention and cross-attention to the prompt memory. The 18 tokens and the decoder's 18 outputs, flattened, condition a 1-D U-Net that denoises 16 actions from Gaussian noise in 16 DDIM steps.
Where the parameters are
| Part | What it is |
|---|---|
| Image encoders | 3 × CLIP ViT-B/16, one per camera; the CLS token, width 768 |
| Pose projections | 6 × Linear, one per pose key, to width 768 |
| Action projection | MLP from 15 actions × 20 to width 768 |
| Prompt pool | Attention pooling, with one learned query, of a chunk's 10 tokens (9 for the frame, 1 for its actions) into one |
| Decoder | 6 pre-norm transformer decoder layers, width 768, 8 heads, MLP 3,072 |
| Diffusion head | 1-D conditional U-Net, channels 256 · 512 · 1024, kernel 5, FiLM in each of its 12 residual blocks |
| Positions, sinks, normalisation | Learned prompt and history positions, 4 sinks, per-key scale and offset |
Half of the network is the FiLM projections: each residual block maps the whole 27,776-number condition (27,648 from the tokens, 128 for the denoising step) to a scale and a bias for its channels. That is the price of conditioning the action head with no bottleneck between what the decoder read and the actions it denoises.
The normalisation is part of the weights: every pose key and the action have a per-dimension
scale and offset you set from your own data. Weights with a zero scale or a non-finite value
anywhere do not load, and a submission whose weights do not load is refused before its duel runs.
Input: the demonstration
At the start of a unit the policy is handed one demonstration of the task, recorded by the benchmark's expert in a different scene. It arrives as named arrays over T steps:
| Array | Shape | Type | What it holds |
|---|---|---|---|
frames_head_camera | T × 240 × 320 × 3 | uint8 | RGB from the head camera |
frames_left_camera | T × 240 × 320 × 3 | uint8 | RGB from the left wrist camera |
frames_right_camera | T × 240 × 320 × 3 | uint8 | RGB from the right wrist camera |
endpose | T × 16 | float64 | Per arm: position (x, y, z), orientation quaternion (w, x, y, z), gripper opening |
qpos, times, frequency | float64 | Joint positions and timing; sent, but vector_v1.1 does not read them. The expert's actions are qpos[1:] and are not sent separately |
What the network makes of it:
- State and action. Row t of
endposeis the state at step t, and row t + 1 is the action taken from it. Each arm's pose becomes position 3, a 6-D rotation (the first two rows of its rotation matrix) and gripper 1: a 20-number row for both arms. - Frames. Every stored frame goes through a JPEG round trip, as recorded episodes do, and is resized to 224 × 224. The network crops the centre 95 % (212 px), resizes back to 224 and applies CLIP's pixel normalisation before the encoder.
- Kept frames. One frame in 15 is kept, with the 15 actions that follow it; the last group is padded with zeros. A demonstration can be up to 1,005 steps (67 prompt tokens); a longer one is refused.
Input: the observation
At every call during the rollout the policy receives the live robot, in the same form without the time axis:
| Array | Shape | Type | What it holds |
|---|---|---|---|
frames_head_camera, frames_left_camera, frames_right_camera | 240 × 320 × 3 | uint8 | The three cameras now |
endpose | 16 | float64 | Both arms' pose and gripper now |
A prediction reads the current frame and the one before it; at the start of an episode the first frame stands in for the one before. Live frames are resized to 224 × 224 without the JPEG round trip. The policy never sees the scene seed, the success condition or anything else the simulator knows.
Output: an action chunk
Each prediction is 16 actions of 20 numbers: per arm, the end-effector position, a 6-D rotation
and the gripper. Each is converted to the benchmark's 16-number pose row, the same layout as
endpose: position, a unit quaternion with w ≥ 0 recovered from the 6-D rotation, and the
gripper clipped to between 0 and 1. The robot moves both arms' end effectors to that target.
| Field, per arm | Numbers | Range |
|---|---|---|
| Position (x, y, z) | 3 | metres, the frame endpose is reported in |
| Orientation (w, x, y, z) | 4 | unit quaternion, w ≥ 0 |
| Gripper | 1 | 0 closed to 1 open |
The diffusion head
The U-Net is trained as a denoising diffusion model of action chunks:
| Noise schedule | 50 steps, squared cosine |
| Prediction | ε, the noise |
| Sampling | DDIM, 16 steps (t = 45, 42, …, 0), from Gaussian noise |
| Clipping | Predicted samples are clipped to [−1, 1] in normalised action space |
Because samples are clipped, the action normalisation stored in your weights should map the actions you train on into [−1, 1]; an action outside it can never be produced.
One unit, call by call
The first 12 actions of each chunk are executed; the last 4 are replaced by the next prediction, made on the frame reached after the twelfth. The call that makes a prediction returns 11 of its 12 kept actions at once and the next call the twelfth, so the network runs once per 12 control steps and the two frames it reads are consecutive. The policy has 60 s to answer each call, and the unit ends when the task succeeds or at its step limit, 2× the expert’s steps in the scene.
Training your weights
Train the network your way, and save the result as model.safetensors holding exactly the
architecture's tensors: every name, shape and float32 dtype in the published manifest, and nothing
else. robotensor miner check holds the file's layout to the manifest; the validator also refuses,
at load, any non-finite value and any zero normalisation scale. What you train on, how long, and
with what augmentation is yours: Training an adaptive model lists the
choices that matter.
The validator hands a demonstration over the way the benchmark records an episode: three cameras,
JPEG-compressed frames, and state and action as consecutive endpose rows. Training on data in
that form means the prompt your model sees in a duel looks like the prompts it learned from.
Replaying the demonstration's actions does not solve a unit: they were taken in another scene. The tasks you are scored on are listed on The benchmark, and the scenes of a duel are never ones the benchmark publishes.
vector_v1.1 is based on the Behavior Prompting Policy architecture (Patel et al., 2026).
