IROS 2026: Collaboration with RobotiqRead recap

Evaluation & Annotation infra for robots.

Annotate robot runs and demonstrations. Evaluate task performance and investigate failures with video and sensor evidence.

Build with RoboLens ↗
RoboLens / shared evidenceAnnotation + evaluation
Your deployment data

Robot runs

Teleoperation recording

Autonomous runs and operator demos.

Video · actions · sensor signals
RoboLensShared evidence

Annotation, evaluation
or both.

Understand the recording

Annotation

Reviewed labels

Task steps, outcomes and interventions.

Reviewed labels, linked to the recording
Understand the behavior

Evaluation

Failure evidence

Task outcomes, comparisons and failure evidence.

“Which steps failed?”

From a question to the relevant evidence

Recording source

Evaluation

Catch regressions.
Debug faster.

See whether a change improves robot performance. Find the failing steps and curate the data for your next iteration.

See evaluation workflow

Illustrative scenario. Claude Code prompt: Did the perception update regress? Compare 120 test runs and suggest data to curate. RoboLens MCP compares the test batch, failure annotations, robot states and motion. Response: Misses rose 2/60 → 8/60, mostly under partial occlusion. The tote stays visible as its track drops, then the robot pauses. Curate occluded frames + matched successes for box and visibility review. Example counts and signals are simulated.

Evidence reviewIllustrative
Candidate · partial occlusion10.0 s
Failure annotationMissed tote detection8.0–10.0 s
Tote track
TrackedLost
Robot state
ApproachHold
TCP speed0–0.2 m/s
6 s8 s10 s
Suggested curation

Occluded tote frames + matched successes. Review boxes and visibility labels before retraining.

Illustrative workflow · example counts and signalsSwipe to see the evidence →
Catch regressions
Compare policy or software versions on matched tasks, using recorded runs or repeatable simulation evaluations.
Debug faster
Ask Codex or Claude Code through RoboLens MCP to locate failures and return the video, labels and signals.
Train on the right examples
Curate recurring failures and matched successes. Retest the same cases after training.

Annotation

Accurate labels.
At scale.

Parallelized processing delivers high throughput at low cost. Task-specific quality review prepares your deployment data for training.

Your task, your labels
Task steps, retries and human interventions, aligned to the recording.
Review included
Action boundaries and outcomes checked against your criteria. Uncertain cases flagged.
Training-ready exports
Labels, timestamps and source references in structured formats.
See annotation examples
Deployment annotationRecording → labels
One run. A useful record.Actions, outcomes and assistance on one timeline
Task steps
Outcomes
Human assistance
Keep uncertainty visible

Ambiguous boundaries, missing context and uncertain outcomes are flagged for review.

Reviewed labels + source referencesFor policy training and task evaluation
Illustrative alignment

Work with RoboLens

Start with one task.

Start with a scoped annotation or evaluation pilot. We agree on your data, success criteria and deliverables, then configure RoboLens with your team.

Build with RoboLens ↗