SIM XR. Bring us a failing job
What we do

You have the robot. We teach it the customer’s job.

A robot model is a starting point, not a deployment. One real customer job — reach it, grip it, carry it, place it, close the door — contains dozens of workplace-specific skills that are not in the box. SIM XR teaches those skills to the model you already run, and measures whether they actually worked.

You hand us
One job the robot fails, with the success criteria you actually care about.
A phone walkthrough of the room and photographs of the objects.
Your checkpoint — optional. Plenty of teams only want the data.
You get back
Validated demonstrations in the robot’s own action space, LeRobot-ready.
A post-trained model, if you want us to train it.
A before-and-after number on one protocol, with the failures decomposed.
Robot-hours it costs you
None. Collection runs off-robot, in parallel copies of the scene, with instant resets.
Where this stalls

Three places the work gets stuck.

01

Every customer is a new skill.

Your model generalises across a category, not across a workplace. Every deployment adds steps nobody trained for, and every one of them needs demonstrations. Collecting those demonstrations costs robot-hours you would rather spend shipping.

02

The world the robot must learn in does not exist yet.

Training in simulation needs the customer’s room, their objects, their clutter and their lighting. Scans come back noisy, meshes are not physics-ready, and rebuilding all of it by hand is nobody’s favourite sprint.

03

The policy fails and nobody can say why.

A deployed policy misses and the default answer is “collect more data”. Sometimes that is right. Often the real cause is a missing input, a training recipe or the inference loop — and more data buys nothing at all.

How it works

One loop. Enter at any step.

We run the whole chain, and we run any part of it. Some teams hand us a failing job and take back a measured policy. Some hand us a phone walkthrough and take back a scene. Some send nothing but rollouts and take back a diagnosis. All three are normal.

The SIM XR loop Five steps in sequence: teleop rig, scene and assets, demonstrations, training, measurement. Measurement feeds back into demonstrations. A customer can enter at any step. ENTER AT ANY STEP 010203 0405 Teleop rig Scene & assets Demonstrations Training Measurement Your robot becomesdrivable in VR. The customer’s roombecomes a training world. Remote operators recordthe missing skill. Your checkpoint ispost-trained on them. Before and after,on one protocol. MEASURED FAILURES BECOME THE NEXT DEMONSTRATIONS
  • 01Teleop rig
    Your robot becomes drivable in VR.
  • 02Scene & assets
    The customer’s room becomes a training world.
  • 03Demonstrations
    Remote operators record the missing skill.
  • 04Training
    Your checkpoint is post-trained on them.
  • 05Measurement
    Before and after, on one protocol — and measured failures become the next demonstrations.
Nothing in this loop consumes your robot-hours. Collection runs off-robot, in simulation, in parallel copies of the scene — which is why adding a tenth operator costs a headset rather than a tenth robot.
Capabilities

What we actually do.

01 — Teleop rig

We make your robot drivable in VR.

Most robots are not built to be demonstrated on. Getting a human’s hands onto a specific embodiment cleanly enough that the recording is trainable is its own engineering job.

  • Retargeting from consumer hand-tracking to your embodiment — absolute or relative, with joint limits and clamps that hold under fast motion
  • Grip and contact handling on hands and grippers that were never designed for teleoperation
  • Automatic episode segmentation, so a session becomes clean labelled trajectories instead of one long recording
  • Reliability instrumentation: pose-freshness gauges, stop-latch detection, session recovery when an operator drops
  • A reproduction kit that brings the whole thing up with a single command
Proof

In a pilot we stood up a customer’s entire stack on our own infrastructure and returned three defects in their teleoperation — including one where a continuous signal was being read as an event and latched the emergency stop.

From you

A robot description and your action / observation schema. Nothing else.

02 — Scene & assets

We rebuild the workplace the robot has to work in.

A skill learned on a clean grey table does not survive contact with a real kitchen, a real shelf or a real bench. The scene is part of the skill.

  • 3D reconstruction of a real room from an ordinary phone walkthrough
  • Physics-ready assets built from photographs of the customer’s actual objects
  • Repair of broken or unusable meshes — collision geometry, grasp affordances, scale and mass
  • Parallel scene instances with instant reset, so operators never wait for a human to tidy the table
Proof

A trained policy was moved into a 3D scan of a real room and kept working with no new demonstrations — in simulation.

From you

A walkthrough video of the space and photographs of the objects. CAD if you happen to have it; we do not need it.

03 — Demonstrations

Our operators record the missing skill.

Not blanket hours. Demonstrations aimed at the failure modes the measurement actually found.

  • Recorded in the robot’s own action space through VR teleoperation — not retargeted from human video
  • Targeted at measured failure states rather than sold as volume
  • Accept / reject gates and replay validation on every single trajectory. Rejects are our cost, not yours
  • Delivered as LeRobot v2.1, or in your own schema
Proof

Our export passed a customer’s own validator with no errors — including the ordering of all 51 joints.

From you

Nothing. No robot time is required for collection.

Why this scales differently

Physical teleoperation adds capacity by buying another robot and another station. Our next operator needs a headset. More than twenty million consumer VR headsets have already shipped; our operators connect from home, and their sessions run in parallel against copies of the same scene.

04 — Training

We post-train the model you already run.

We are not asking you to switch architectures. The foundation model is the starting point; we close the deployment-specific gap on top of it.

  • Fine-tuning of your existing checkpoint on the targeted demonstrations — GR00T-family models, other VLA families, or your own architecture
  • GPU infrastructure in the EU and the US, so this does not queue behind your own research team
  • Or no training at all: take the validated dataset and train it yourself
Proof

We have run the loop across two fundamentally different model types and multiple robot, hand and gripper configurations — including a model family released days before we ran it through.

From you

A checkpoint — or nothing, if you only want the data.

05 — Measurement

We tell you what actually changed.

Evaluation decides the next intervention. Sometimes the answer is more demonstrations. Sometimes it is a missing input, a recipe or the inference loop — and we would rather tell you that than sell you hours.

  • Before and after on one protocol and identical initial states — your baseline against the retrained model
  • Failure decomposition: where along the path each attempt exits, how far failures get, whether a drop is terminal, and where in space the failures cluster
  • Attempt economics — what the policy is worth at one attempt and at N, because a robot that can retry is a different business from one that cannot
  • The honest version of the number: the confidence interval at your sample size, and what the result does when the starting pose is allowed to vary
Proof

This runs on your rollouts alone — object poses and success labels. We never need your data or your weights, which is why we also sell it on its own.

From you

Rollouts.

Proof

Two things we can show you.

90 seconds · end-to-end case, run in our simulation pipeline Sound on

0 → 9 out of 10, in simulation.

GR00T N1.7 — fine-tuned by NVIDIA for its own G1 benchmark task — scored 0% on a structurally new task. One person at home recorded 50 VR demonstrations through simxr.app. Trained on those demonstrations, the model succeeded in 9 out of 10 trials, in simulation, from a fixed starting configuration.

Same policy, same protocol, in simulation — only the starting pose moves
Start-pose jitterSuccessScale
0 cm — fixed90%
±2 cm73.3%
±5 cm50%

We publish the curve, not only the headline. Knowing exactly where a skill stops holding is the difference between a demo and a deployment — and finding that edge is the part of the work we sell.

Someone else’s stack, running on ours, in days.

A pilot customer — under NDA, so not named here — handed us their full teleoperation and training stack. We stood it up on our own infrastructure, found and returned three defects in their pipeline, recorded VR teleoperation with automatic episode slicing, and delivered LeRobot v2.1 that passed their own validator with no errors, together with a kit that reproduces the entire run with one command.

The point is not that we are good guests. It is that the integration risk of working with us is measured in days, and the first thing we produced was a bug report against our own customer’s code rather than an invoice for episodes.

A pair of robot arms operating in a photorealistic reconstruction of a real room, rebuilt from a phone walkthrough.
A real room, reconstructed and made physics-ready, with the robot working inside it — in simulation.
Ways to start

Three doors, in order of commitment.

01

Offline diagnostic

Send rollouts, get a failure decomposition. No data, no weights, no integration — object poses and success labels are enough. You get where the policy exits, how far its failures get, what your number is actually worth at your sample size, and what we would fix first.

02

Scene and data package

We build the world and record the demonstrations. You train. For teams with their own training expertise and not enough hands for scenes or collection.

03

Policy Improvement Sprint

One failing task or a bounded failure distribution, at a fixed project price. Demonstrations, validation and the evaluation report always included. The success protocol is agreed before the first session.

See the sprint
Stack

Where this plugs in.

Your model
GR00T-family models, other VLA families, or your own architecture. We improve what you already run — we are not building a competing foundation model, and we are not asking you to switch.
Your robot
Humanoids, dual-arm cells, parallel grippers and dexterous hands. Demonstrations are recorded in the robot’s own action space, not retargeted from human video.
Simulation & export
NVIDIA Isaac Sim, Isaac Lab, Isaac Teleop. CloudXR streaming. Export in LeRobot format or your own schema. Compatible with your simulator and training code where integration permits.
Operators & headsets
Quest 2, Quest 3, Pico and Apple Vision Pro, plus motion capture where the task needs it. Operators connect from home across Europe and the United States; sessions run in parallel.
Infrastructure
Our own GPU servers in the EU and the US, so this does not queue behind your research team. Scenes reset instantly and run in parallel copies — which is what makes targeted collection cheap enough to aim.
Delivery
Every trajectory replay-validated before it reaches you. Rejects are our cost, not yours. A reproduction kit brings the whole run back up with one command.
NVIDIA Inception member AWS Activate member
Talk to us

Show us the job the robot fails.

Describe the task and where it breaks. We come back with what we would do first, what it needs from you, and what it does not.

One reply, from the team. No mailing list, no drip sequence.