Policy Improvement Sprint

Fix one policy failure. Measure the improvement.

Bring us one task your robot struggles with. SIM XR reproduces the task in simulation, collects targeted human demonstrations exactly where the policy breaks, validates every trajectory, and measures what changed — without consuming additional robot hours on your side.

Bring us one failure case
Internal benchmark · run end-to-end in our simulation pipeline
0% → 90%

Success rate on a structurally new manipulation task — before and after fine-tuning on our targeted demonstrations, same evaluation protocol and initial states.

50 demos

Recorded by a single operator. Baseline: GR00T N1.7, fine-tuned by NVIDIA for its own G1 benchmark task, scored 0% on this structurally new task.

~3 days

From task definition to a measured policy — collection, fine-tune and evaluation included.

Process

How a sprint runs.

1
Define

The failing task, success criteria, action/observation schema and evaluation protocol — agreed up front.

2
Reproduce

The task and its failure distribution in simulation — including your environment, reconstructed from a scan when the scene matters.

3
Collect

Targeted VR demonstrations aimed at the measured failure modes — not blanket hours.

4
Validate

Accept/reject gates and replay validation on every trajectory. Rejects are our cost, not yours.

5
Measure

Before/after evaluation on the same protocol — in your stack where integration permits, or in our evaluation harness.

Scope

One task or a bounded failure distribution, at a fixed project price. Success protocol agreed before the first session; demonstrations, validation and the evaluation report always included.

Deliverables

What you receive — and why it works.

What you receive
  • Targeted demonstrations in your schema — LeRobot-ready, replayable, with per-episode logs
  • Validation results — accept/reject outcomes and replay checks for every trajectory
  • Evaluation report — baseline vs. post-training on identical initial states, plus coverage of the failure distribution
  • Optional fine-tuned checkpoint and recommendations for the next iteration
Why it works
  • Demonstrations in the robot’s action space — recorded via teleop, not retargeted from video
  • Our own operator pool and infrastructure: consumer headsets, GPU servers in the EU and US, instant scene resets
  • No client robot time required for data collection — sessions run off-robot, in simulation
  • Compatible with NVIDIA Isaac and GR00T workflows, with LeRobot-format export
Scope a sprint

Show us where your policy breaks.

Describe the task and where the current policy fails. We’ll come back with a sprint scope: evaluation protocol, collection plan, and a fixed price.

One reply from the team. No mailing list.

NVIDIA Inception Member · Powered by AWS Activate Member