Evaluation

Ship more reliable robots.

Catch regressions before rollout. Find the failing steps with video and sensor evidence, then curate the data needed to improve the next version.

Where does this task fail, and what evidence explains it?

Analyze recorded runs

Your recordingsVideos, actions and signals
Your success criteriaSubtasks, outcomes and collection context

Run simulation evaluations

Baseline + candidatePolicy or software versions
Repeatable casesYour environment, same tasks and seeds
RoboLensScore and compare · Investigate · Report or curate
Recorded runTask outcomeEvidence
ep_086Incomplete12.4–13.0 s
ep_088Incomplete17.3–17.7 s

Same unsuccessful subtask in both sample runs.

Compare the runs and inspect evidence ↗

01 · Define the task and criteria

Make the outcome explicit.

Choose the behavior to evaluate, the task outcome and the relevant policies or cohorts. Keep collection conditions alongside the comparison.

RoboLens / evaluationTask → behavior → evidence
Task definition

Your task. Your acceptance criteria.

Define the required behaviorSpecify subtask and task outcomesChoose versions or comparison cohorts
Demonstration context

Put the red marker in the purple bowl.
Same session and task, different policies.

Two selected public failures, not a policy ranking.

Start with recorded runs and task criteria. For simulation, use an existing benchmark or your environment.