SIM XR. Bring us a failing job
What we do

We teach robots the jobs they cannot do yet — and measure that they can.

Bring the task and the model you already run. We rebuild the job in simulation, collect human demonstrations in VR, scale them into training data, post-train your policy and return it with a before-and-after number on a protocol you can rerun yourself — without spending an hour of your robot’s time.

0 → 65 % across four tasks for a humanoid manufacturer · 0 → 90 % from fifty VR demonstrations · 68 of 100 on a walk-and-place job in a scanned room — in simulation, on protocols you can rerun.

A humanoid robot at a kitchen counter misses a bottle, which tips over beside the bowl
You hand us
One job the robot fails, with the success criteria you actually care about.
A phone walkthrough of the room and photographs of the objects.
Your checkpoint — optional. Plenty of teams only want the data.
You get back
Validated demonstrations in the robot’s own action space, LeRobot-ready.
A post-trained model, if you want us to train it.
A before-and-after number on one protocol, with the failures decomposed.
Robot-hours it costs you
None. Collection runs off-robot, in parallel copies of the scene, with instant resets.
Why it stalls

Why the next job is never free.

The same robot in a kitchen, at a warehouse shelf and at a workbench 01

Every new job is a new skill.

Your model generalises across a category, not across a workplace. Every deployment adds steps nobody trained for, and every one of them needs demonstrations. Collecting those demonstrations costs robot-hours you would rather spend shipping.

A half-built room: one half finished, the other dissolving into wireframe with missing patches 02

The world the robot must learn in does not exist yet.

Training in simulation needs the customer’s room, their objects, their clutter and their lighting. Scans come back noisy, meshes are not physics-ready, and rebuilding all of it by hand is nobody’s favourite sprint.

A robot has dropped the bottle; a magnifying glass over the failure point and three empty checkboxes 03

The policy fails and nobody can say why.

A deployed policy misses and the default answer is “collect more data”. Sometimes that is right. Often the real cause is a missing input, a training recipe or the inference loop — and more data buys nothing at all.

How it works

One loop. Your job goes in, a working policy comes out.

You bring
From you
A job the robot cannot do yet And whatever you already have. We build what is missing.
  • one scene or many
  • one task per scene or many
  • your robot description
  • scene & objects, if you have them
  • success criteria — written together
Build the data
Teach and prove

Measured failures go back into the loop — demonstrations, spread or recipe — until the number holds.

You get
From us
A policy that does the job it could not do before Measured in simulation, before and after, on one protocol you can rerun yourself.
  • the checkpoint
  • the validated dataset
  • the evaluation protocol, with seeds
  • zero robot-hours spent
Beyond the loop
A wireframe robot in simulation and a solid robot in the real world, with a dashed arrow between them
Sim-to-real, on your robot The step after the loop: the same policy and the same protocol on your hardware — planned with you from the first session.
We bring
Two server racks on a floor grid linked to a cloudCompute that is ours. GPU fleets in the EU and the US, up in hours and down to zero after — nothing queues behind your research team.
Five houses with VR headsets above them, all linked to one robot at a tableOperators who are ours. Onboarded, qualified, connecting from home; the next operator costs a headset, not a robot.
An open notebook with two measured curves and error bracketsRecipes that are measured. Every stage above has run end-to-end on customer tasks — several tasks across several scenes, every delivery accepted by the customer’s own validator.
Cases

Three jobs. Three measured results.

Case 01 · Pilot

A new skill from fifty VR demonstrations.

0 → 90%success, 27 of 30 paired trials
50VR demonstrations
1operator, from home

GR00T N1.7 — fine-tuned by NVIDIA for its own benchmark task — scored zero on a structurally new pick-and-place task in our scene. One operator recorded fifty demonstrations in VR through simxr.app. Post-trained on them, the same model completed the task in 27 of 30 paired trials on a fixed protocol — in simulation.

  • post-trained checkpoint
  • validated dataset
  • evaluation protocol
  • 90-second video
90 seconds · the full loop on one task · sound on
Case 02 · Full project

Four tasks in two scenes for a humanoid manufacturer.

0 → 65%success across all four tasks, customer’s criterion
~1,500successful demonstrations from ~20 operator-hours
~10,000generated episodes, every batch accepted

The customer brought their own robot, two scenes and a strict data contract. We made the robot drivable in VR, rebuilt the scenes, recorded the demonstrations across four tasks with our operators, multiplied them into validated episodes and delivered every batch in the customer’s own format — each one accepted by their validator. Post-trained on that data, a policy that scored zero on every task reached 65 % success across all four under the customer’s own criterion, in simulation — and 80–85 % once the acceptance-side defects we reported are corrected.

  • 4 accepted deliveries
  • customer-format data
  • post-trained checkpoint
  • evaluation protocol
Robot and scenes
Theirs — made drivable and physics-ready by us
Demonstrations
~1,500 successful, four tasks, our operators
Generated episodes
~10,000, replay-validated, in the customer’s format
Deliveries
4 of 4 accepted by the customer’s validator
Post-training
0 % → 65 % across four tasks, customer’s criterion, in simulation; 80–85 % with acceptance-side defects corrected
Case 03 · Locomanipulation

A humanoid carries breakfast across a scanned room.

68 of 100attempts succeed, on a job the robot could not do before
15VR demonstrations, one operator
500generated episodes, with the walk added

Pick up a jam jar, walk across the room and stand it on a breakfast tray: a job the robot could not do. We scanned the salon of a French château as a 3D Gaussian splat, rebuilt it as a physics-ready scene and made the jar from a single photograph. One operator at home showed the pick-and-place in VR through simxr.app.

Fifteen demonstrations became 500 training episodes, with the walk across the room added in simulation. Post-trained on them, the model now drives a Unitree G1 through the whole job on its own, and succeeds in 68 of 100 attempts from random starts — in simulation.

  • scanned scene
  • validated dataset
  • post-trained checkpoint
  • evaluation protocol
76 seconds · from the scanned room to the robot doing the job · no sound
Scene
A real room as a 3D Gaussian splat, objects from photographs
Demonstrations
15, recorded in VR by one operator at home
Generated
500 episodes; the robot walks about six metres in each
Robot and model
Unitree G1 · GR00T N1.5, post-trained · head and wrist cameras
Result
68 of 100 attempts from random starts, in simulation
NVIDIA Inception member and Powered by AWS badges

Member of NVIDIA Inception. AWS Activate credits issued through the NVIDIA Inception programme. Our stack runs on NVIDIA Isaac Sim, Isaac Lab and CloudXR, on our own GPU fleets in the EU and the US.

Ways to start

Three doors, in order of commitment.

Data cards with check and cross marks going into a clipboard, a report with a bar chart coming out
01

Offline diagnostic

Send rollouts, get a failure decomposition. No data, no weights, no integration — object poses and success labels are enough. You get where the policy exits, how far its failures get, what your number is actually worth at your sample size, and what we would fix first.

An open crate holding a miniature scene and a stack of data cards
02

Scene and data package

We build the world and record the demonstrations. You train. For teams with their own training expertise and not enough hands for scenes or collection.

A timeline with a loop of dashed arrows around a small robot, ending in a checked card
03

Policy Improvement Sprint

One failing task or a bounded failure distribution, at a fixed project price. Demonstrations, validation and the evaluation report always included. The success protocol is agreed before the first session.

See the sprint
Stack

Where this plugs in.

Your model
GR00T-family models, other VLA families, or your own architecture. We improve what you already run — we are not building a competing foundation model, and we are not asking you to switch.
Your robot
Humanoids, dual-arm cells, parallel grippers and dexterous hands. Demonstrations are recorded in the robot’s own action space, not retargeted from human video.
Simulation & export
NVIDIA Isaac Sim, Isaac Lab, Isaac Teleop. CloudXR streaming. Export in LeRobot format or your own schema. Compatible with your simulator and training code where integration permits.
Operators & headsets
Quest 2, Quest 3, Pico and Apple Vision Pro, plus motion capture where the task needs it. Operators connect from home across Europe and the United States; sessions run in parallel.
Infrastructure
Our own GPU servers in the EU and the US, so this does not queue behind your research team. Scenes reset instantly and run in parallel copies — which is what makes targeted collection cheap enough to aim.
Delivery
Every trajectory replay-validated before it reaches you. Rejects are our cost, not yours. A reproduction kit brings the whole run back up with one command.
NVIDIA Inception member AWS Activate member
Talk to us

Show us the job the robot fails.

Describe the task and where it breaks. We come back with what we would do first, what it needs from you, and what it does not.

One reply, from the team. No mailing list, no drip sequence.