Skip to content
Attune
All pages

Vector · The guide

The model

The one network every entry fills, vector_v1.1: what it is shown, how the demonstration and the live frames flow through it, and the actions it returns.

6 min read

One architecture

Every submission is the weights of vector_v1.1, Robotensor's demonstration-prompted diffusion policy for two arms, designed for this competition: a CLIP ViT-B/16 image encoder per camera, a six-layer transformer decoder that reads the live frames against the demonstration, and a 1-D diffusion U-Net that turns what it read into actions. It has 781 M parameters in 753 float32 tensors.

The validator builds the network with its own code and loads your weights into it, so the architecture is fixed and the weights are yours: what a duel compares is training and data.

The architecture, end to end

vector_v1.1 reads one demonstration and two live frames, and returns 16 actions781.0 M parameters in 753 float32 tensors · width 768 · 8-head attention throughoutFrame tokenizer · shared by both pathsPrompt encoder · once per unitPolicy decoder · every predictionDiffusion action headOutput · an action chunkDemonstration · once per unitup to 1,005 expert steps, recorded in another scene1 frame kept per 15 steps: P ≤ 67 kept framesLive observation · every predictionthe current frame and the one before it3 camera images + both arms’ pose3 camerashead · left wrist · right wristRGB 240 × 320 × 3, uint8Image preprocessingresize to 224 × 224 (demo frames: JPEG first)center crop 212 px, back to 224, CLIP normCLIP ViT-B/16, one per cameraown weights per camera, 85.8 M eachthe CLS token → 1 token × 7686 pose keysper arm: orientation 6 · position 3 · gripper 1from the 16-number pose: quaternion → 6-DAffine normalizationper-dimension scale and offsetset by the miner, stored in the weights6 linear projectionsone per key, each to 768 · 0.02 M1 token per keyFrame tokens: 9 × 768 per frame3 camera tokens + 6 pose tokensKept demonstration frames9 tokens × 768 per kept frameand the 15 actions after it: 15 × 20Action tokennormalized, flattened to 300 numbersMLP 300 → 300 → 768, GELU · 0.3 MAttention pooling, 10 tokens → 11 learned query, 8 heads, MLP × 4 · 7.1 Mone prompt token per kept framePositions and sinks+ learned prompt position (up to 67)4 learned sink tokens in frontPrompt memory(4 + P) tokens × 768Live tokens2 frames × 9 = 18 tokens × 768+ learned history positions (18 × 768)Decoder layer × 6, pre-normself-attention over the 18 tokens, 8 headscross-attention to the prompt memory, 8 headsDecoder output18 tokens × 768Condition18 live tokens (no positions) + 18 decoder outputs36 × 768, flattened: 27,648 numbersDiffusion step embeddingsinusoidal 128 → 512 → 128, Mishjoined to the condition: 27,776 numbersAction trajectoryGaussian noise at first, 16 × 20DDIM stepremoves the predicted noise ε, clips to [−1, 1]Down 12 res blocks, 20 → 256 channelslength 16, then stride-2 conv → 8Down 22 res blocks, 256 → 512length 8, then stride-2 conv → 4Down 32 res blocks, 512 → 1,024length 4Middle2 res blocks, 1,024 channels, length 4Up 1joins Down 3’s output: 2,048 → 5122 res blocks, transposed conv → length 8Up 2joins Down 2’s output: 1,024 → 2562 res blocks, transposed conv → length 16Final convolution256 → 256 (kernel 5), then 1 × 1 → 20predicts the noise ε, 16 × 20Inverse normalizationthe action’s scale and offset undone: 16 actions × 20Per arm, per actionend-effector position 3 · 6-D rotation 6 · gripper 1Benchmark pose row, 16 numbersper arm: position 3 · unit quaternion 4 with w ≥ 0 · gripper 1, clipped to 0 to 1Execute 12, then predict againthe last 4 are replaced by the next prediction, made on the frame after the twelfthfeed-forward 768 → 3072, GELU56.7 M parameters across the 6 layerskept frameslive framesskip: the 18 live tokens, without positionsrepeat × 16skipskipEach res block: 2 conv blocks (kernel 5, GroupNorm 8, Mish) with FiLM scale and biasfrom the condition between them, in all 12 res blocks. U-Net: 459.4 M parameters.after 16 steps
vector_v1.1 in full, as the validator's runtime builds and runs it: 781 M parameters in 753 float32 tensors, width 768, 8-head attention throughout. The prompt encoder runs once per unit; the decoder and the diffusion head once per prediction, every 12 control steps.

The vector_v1.1 architecture

The network has two paths that share one frame tokenizer.

  • The prompt is built once per unit, when the demonstration arrives. One frame is kept every 15 steps; each kept frame becomes nine tokens (one per camera, six for the arms' pose), is joined by an embedding of the 15 actions that followed it, and is pooled into a single prompt token. Four learned sink tokens go in front. The result is the decoder's memory.
  • The policy runs at every prediction. The current frame and the one before it become 18 tokens, which a six-layer decoder passes through self-attention and cross-attention to the prompt memory. The 18 tokens and the decoder's 18 outputs, flattened, condition a 1-D U-Net that denoises 16 actions from Gaussian noise in 16 DDIM steps.

Where the parameters are

U-Net FiLM projections398.2 M · 51.0%Image encoders, 3 × CLIP ViT-B/16257.4 M · 33.0%U-Net convolutions and step MLP61.2 M · 7.8%Decoder, 6 layers56.7 M · 7.3%Prompt attention pooling7.1 M · 0.9%Action projection MLP0.30 M · <0.1%Positions, sinks, normalisation0.07 M · <0.1%Pose projections, 6 × Linear0.02 M · <0.1%
vector_v1.1 by component, 781 M parameters in all. Half of them are the FiLM projections that inject the full 27,776-number condition into each of the U-Net's twelve residual blocks.

Parameters by component

PartWhat it is
Image encoders3 × CLIP ViT-B/16, one per camera; the CLS token, width 768
Pose projections6 × Linear, one per pose key, to width 768
Action projectionMLP from 15 actions × 20 to width 768
Prompt poolAttention pooling, with one learned query, of a chunk's 10 tokens (9 for the frame, 1 for its actions) into one
Decoder6 pre-norm transformer decoder layers, width 768, 8 heads, MLP 3,072
Diffusion head1-D conditional U-Net, channels 256 · 512 · 1024, kernel 5, FiLM in each of its 12 residual blocks
Positions, sinks, normalisationLearned prompt and history positions, 4 sinks, per-key scale and offset

Half of the network is the FiLM projections: each residual block maps the whole 27,776-number condition (27,648 from the tokens, 128 for the denoising step) to a scale and a bias for its channels. That is the price of conditioning the action head with no bottleneck between what the decoder read and the actions it denoises.

The normalisation is part of the weights: every pose key and the action have a per-dimension scale and offset you set from your own data. Weights with a zero scale or a non-finite value anywhere do not load, and a submission whose weights do not load is refused before its duel runs.

Input: the demonstration

At the start of a unit the policy is handed one demonstration of the task, recorded by the benchmark's expert in a different scene. It arrives as named arrays over T steps:

ArrayShapeTypeWhat it holds
frames_head_cameraT × 240 × 320 × 3uint8RGB from the head camera
frames_left_cameraT × 240 × 320 × 3uint8RGB from the left wrist camera
frames_right_cameraT × 240 × 320 × 3uint8RGB from the right wrist camera
endposeT × 16float64Per arm: position (x, y, z), orientation quaternion (w, x, y, z), gripper opening
qpos, times, frequencyfloat64Joint positions and timing; sent, but vector_v1.1 does not read them. The expert's actions are qpos[1:] and are not sent separately

What the network makes of it:

  • State and action. Row t of endpose is the state at step t, and row t + 1 is the action taken from it. Each arm's pose becomes position 3, a 6-D rotation (the first two rows of its rotation matrix) and gripper 1: a 20-number row for both arms.
  • Frames. Every stored frame goes through a JPEG round trip, as recorded episodes do, and is resized to 224 × 224. The network crops the centre 95 % (212 px), resizes back to 224 and applies CLIP's pixel normalisation before the encoder.
  • Kept frames. One frame in 15 is kept, with the 15 actions that follow it; the last group is padded with zeros. A demonstration can be up to 1,005 steps (67 prompt tokens); a longer one is refused.

Input: the observation

At every call during the rollout the policy receives the live robot, in the same form without the time axis:

ArrayShapeTypeWhat it holds
frames_head_camera, frames_left_camera, frames_right_camera240 × 320 × 3uint8The three cameras now
endpose16float64Both arms' pose and gripper now

A prediction reads the current frame and the one before it; at the start of an episode the first frame stands in for the one before. Live frames are resized to 224 × 224 without the JPEG round trip. The policy never sees the scene seed, the success condition or anything else the simulator knows.

Output: an action chunk

Each prediction is 16 actions of 20 numbers: per arm, the end-effector position, a 6-D rotation and the gripper. Each is converted to the benchmark's 16-number pose row, the same layout as endpose: position, a unit quaternion with w ≥ 0 recovered from the 6-D rotation, and the gripper clipped to between 0 and 1. The robot moves both arms' end effectors to that target.

Field, per armNumbersRange
Position (x, y, z)3metres, the frame endpose is reported in
Orientation (w, x, y, z)4unit quaternion, w ≥ 0
Gripper10 closed to 1 open

The diffusion head

The U-Net is trained as a denoising diffusion model of action chunks:

Noise schedule50 steps, squared cosine
Predictionε, the noise
SamplingDDIM, 16 steps (t = 45, 42, …, 0), from Gaussian noise
ClippingPredicted samples are clipped to [−1, 1] in normalised action space

Because samples are clipped, the action normalisation stored in your weights should map the actions you train on into [−1, 1]; an action outside it can never be produced.

One unit, call by call

reset(seed)seeds the noiseset_demonstrationbuilds the prompt memoryact(observation)until done or out of steps#1#2#30122440stepexecuted (12)predicted, then replaced (4)
A prediction is made every 12 steps, on the frame it is made at and the one before. The last 4 actions of each chunk are never run: the next prediction replaces them.

One unit, call by call

The first 12 actions of each chunk are executed; the last 4 are replaced by the next prediction, made on the frame reached after the twelfth. The call that makes a prediction returns 11 of its 12 kept actions at once and the next call the twelfth, so the network runs once per 12 control steps and the two frames it reads are consecutive. The policy has 60 s to answer each call, and the unit ends when the task succeeds or at its step limit, 2× the expert’s steps in the scene.

Training your weights

Train the network your way, and save the result as model.safetensors holding exactly the architecture's tensors: every name, shape and float32 dtype in the published manifest, and nothing else. robotensor miner check holds the file's layout to the manifest; the validator also refuses, at load, any non-finite value and any zero normalisation scale. What you train on, how long, and with what augmentation is yours: Training an adaptive model lists the choices that matter.

The validator hands a demonstration over the way the benchmark records an episode: three cameras, JPEG-compressed frames, and state and action as consecutive endpose rows. Training on data in that form means the prompt your model sees in a duel looks like the prompts it learned from.

Replaying the demonstration's actions does not solve a unit: they were taken in another scene. The tasks you are scored on are listed on The benchmark, and the scenes of a duel are never ones the benchmark publishes.

vector_v1.1 is based on the Behavior Prompting Policy architecture (Patel et al., 2026).

The model · Docs · Attune