Skip to content
Attune
All pages

Vector · The guide

Introduction

Vector measures action-aware in-context adaptation: a policy is shown one robot demonstration of a task and must do it in a scene it has never seen, with its weights frozen.

2 min read

What Vector measures

A policy is shown an expert do a task once, in one scene. It then has to do the same task on its own, in a second scene where the objects have moved and the other arm may have to act. Its weights are frozen from the moment they are committed on chain: whatever it learns about the task, it learns from the demonstration, inside its forward pass.

Copying the demonstration's motions does not work, because the objects are somewhere else. The policy has to read the demonstration as a description of the task, not as a trajectory.

The prompt · the expert’s demonstration

Start

Start

Grasp and carry

Grasp and carry

Done

Done

Same task · another scene · weights frozen

The test · the scene the policy is scored in

Start: objects moved

Start: objects moved

The left arm acts

The left arm acts

Success

Success

One unit of place_empty_cup. Top: the demonstration the policy is shown, recorded by the expert in one scene; the right arm moves the cup. Bottom: the scene the policy is scored in, drawn independently, where the cup and the coaster have swapped sides and the left arm must act, shown solved by the expert, which has to solve every scored scene too. The policy gets 2× the expert's steps in that scene.

One evaluation unit: a demonstration scene and a scored scene

At a glance

Tasks16 manipulation tasks on a two-arm Franka setup, each scored on its own
ContextOne demonstration: three camera streams, both arms' poses and the actions taken
Modelvector_v1.1, the same network for every entry; you train the weights
Duel160 (10 × 16 tasks), dealt from a chain block finalized after the commitment
CrownThe challenger's score beats the king's by at least +3.0 pts
Rewards30% of miner emissions, shared 40 / 30 / 20 / 10% by the four most recent champions
SubmissionOne per hotkey: a Hugging Face repository holding model.safetensors
RecordEvery duel published unit by unit, with three clips per unit

The model in one picture

Once per unit · the promptEvery prediction · the policyInput · the demonstration≤ 1,005 steps · 3 cameras · 16-D pose1 frame kept per 15 stepsInput · the observation3 cameras · 16-D pose, now+ the frame before: 2 framesFrame tokenizer · shared by both paths3 × CLIP ViT-B/16, per camera → CLS → 7686 × Linear, per pose key → 768images 224 px, centre crop 95%, CLIP norm · pose: rot6d, normalised9 tokens / framePrompt token, per kept frame9 tokens + MLP(15 × 20 actions)→ attention pool → 1 token, + positionHistory tokens2 × 9 = 18 tokens × 768+ a learned position eachPrompt memory4 sinks + P prompt tokensP ≤ 67 (≤ 1,005 steps)Transformer decoder × 6self-attn → cross-attn → MLPpre-norm, 8 heads, width 768crossskip, no positionEncoded once, when the demo arrives,reused by every prediction in the unit.Condition[18 history ; 18 decoded] = 36 × 768flattened: 27,648 + 128 stepGaussian noise16 × 20Diffusion head · 1-D U-Net256 · 512 · 1024 channels, FiLMDDIM × 16, ε, clipped to [−1, 1]then the normalisation undoneOutput · an action chunk16 × 20; the first 12 are executedEach 20-D row: per arm position 3,rotation 6-D and gripper 1, sent to therobot as a 16-D pose row.
vector_v1.1, the only architecture a submission may fill: 781 M parameters in 753 float32 tensors. The prompt memory is built once per unit; everything on the right runs once per prediction, every 12 control steps. You train the weights; the validator builds this network with its own code and loads them into it.

The vector_v1.1 architecture, in one picture

The demonstration becomes a prompt memory once per unit. At every prediction the network reads the two latest camera frames and arm poses against that memory and denoises a chunk of end-effector actions for both arms. The model takes it apart, input by input.

Where the crown stands

The crown

—average over the tasks

Crown vacant

No model holds the crown yet: the first accepted entrant is crowned by genesis.

0 duels published · the dashboard

Compete in four steps

  1. Train the weights of vector_v1.1 with your own data, curriculum and method. Training an adaptive model lists what is yours to decide.
  2. Check your model.safetensors against the pinned architecture with robotensor miner check.
  3. Submit: robotensor miner submit uploads to a new private repository, commits it on chain with the weights' sha256, then makes the repository public. One submission per hotkey.
  4. Duel: your entry is queued as soon as a validator reads the commitment, oldest first, and plays whoever holds the crown when its turn comes.