CI/CD · Managed runners · Failure evidence

The evaluation platform for robots.

Regression testing across field replay and simulation. Catch regressions before deployment.

Your changecandidate v42policy checkpoint · commit 4f2a
Control planeruns · gates · evidence
Runner 01Policy checks

Simulation replay

Fixed seeds, perturbed scenarios, baseline vs candidate

Runner 02Incident checks

Field replay

Rosbag replay: your incidents become repeatable tests. No sim required.

● Release check · failed

Regression found.

Evidence attached to three cases.

signal → your CI−14 pp

The problem

Testing is broken for robots.

Manual test cycles drag every release, and the failures that slip through break deployments in the field — with the evidence buried in logs and footage.

ExecutionDifferent runners, no shared record.DecisionInvalid runs blur into behavior failures.IterationEngineers rewatch entire episodes.

01 · Trigger

A change selects trusted cases.

A policy checkpoint or stack commit triggers the suite: versioned eval packs built from your environments and field incidents.

02 · Run

Cases execute through the right runner.

Simulation sweeps seeds and perturbed conditions: lighting, object poses, sensor noise. Field replay recomputes fresh outputs from a real incident.

03 · Decide

Results become a release signal.

Simulation compares baseline and candidate. Replay verifies each run is sound. The control plane owns the gate: pass or fail, with evidence.

04 · Inspect

A regression opens with evidence attached.

Every run is annotated with subtasks and sensor anomalies. Video, signals, and actions point to the failed transition, and similar failures cluster across runs.

05 · Improve

Review the failed window, not the entire run.

Each failure window becomes curated training data and a new regression case. The suite grows from what actually broke.

RoboLens / operating looprelease-42 · live record
Changecandidate v42policy checkpoint · commit 4f2a

Policy evaluation

Simulation replay

Fixed seeds, perturbed scenarios, baseline vs candidate.

Incident evaluation

Field replay

Recorded inputs, pinned ROS image, fresh outputs.

Simulation

Candidate regressed

−14 pp

Replay

Run is valid

VALID · fresh outputs

→ release gate

● Release check · failed

Regression found.

Three cases need review. Every result links to its run record.

wrist cam · ep_086 · synced record
grasp fails · 04.3s

Behavior timeline

00.0 · approach
04.0 · gripper closes
04.3 · grasp fails
07.2 · regrasp

Failure localizedGripper closes before contact and lifts empty. Video and aperture trace agree.

focused window · 04.0 – 05.6s
curated for retrainingfailure window → +4 clips

Focused review

1 of 4 subtasks complete

Approachdone
Graspfailed
Liftnot reached
Placenot reached

promote window → regression case

Sim EvalPacksYour environment, packaged and versioned.Perturbation testingLighting, poses, sensor degradation.Rosbag replayNo sim required. Start from your bags.Managed runnersOur cloud or your VPC.CI + agentsAgents launch evals and inspect failures over MCP.

Design partnerships

Bring your critical evaluation.

We will make your simulation or replay workflow repeatable, measurable, and useful in your release process.

Build with RoboLens ↗