Source-linked AI summary

PRM-as-a-Judge 1.5: A Toolkit for Robot Process Assessment

Yuyang Liu, Yanqing Shen, Ruike Chen, Jifan Zhao, Yuxuan Tian, Yichi Zhang, Tianfeng Long, Zixuan Yin, Yipu Wang, Ziheng Qin, Wenxing Tan, Yang Shi, Mingyu Cao, Runze Xiao, Ziqi Wang, Zhixin Yin, Shiwei Chu, Yi-Fan Zhang, Yao Mu, Yuheng Ji, Yihao Wang, Jun Yan, Zhongyuan Wang, Pengwei Wang, Xiaolong Zheng

arXiv:2608.14284v1cs.ROcs.CV

TL;DR

Embodied-model evaluation needs to capture progress and failure patterns beyond final success rates or rule-based scores. PRM-as-a-Judge 1.5 converts rollout videos into progress curves, process metrics, and reports, and its assessments expose detailed behavioral signatures while showing that larger models do not necessarily perform better.

  • Problem

    Binary success rates and rule-based scores insufficiently characterize diverse progress and failure patterns in embodied-model rollouts.

  • Method

    PRM-as-a-Judge 1.5 converts rollout videos into progress curves, computes conditioned process metrics, and generates assessment reports.

  • Results

    Large-scale assessments expose detailed behavioral signatures, with no clear positive correlation between model parameter count and embodied performance.

  • Takeaways & Limitations

    Process assessment can identify specific failure patterns and guide failure-targeted data collection and training strategies.

  • Takeaways & Limitations

    The assessment framework would benefit from richer visual, state, contact, or semantic evidence to make reports more interpretable.

Abstract

from arXiv · show

Fine-grained robotic evaluation matters for understanding embodied models, going beyond binary success rates and rule-based process scores. We present PRM-as-a-Judge 1.5, a toolkit for robot process assessment that turns rollout videos into dense progress curves and derives multiple fine metrics. PRM-as-a-Judge 1.5 introduces three metrics, building on version 1.0, that characterize failure-side progress, post-drawdown recovery, and success-side execution quality, helping users understand embodied model capability. Based on the rollout videos from benchmarks, we perform a comprehensive assessment of the embodied models, providing some fine-grained metric results and key findings. We further introduce RoboPulse++ to evaluate the reliability of process reward models (PRM), providing evaluators with a more accurate testing platform. Moreover, we release a user-friendly assessment suite, including the benchmark, metric implementation, and visualization tools, to support reproducible manipulation process evaluation. We call on the community to rethink how robots are evaluated and establish transparent, procedural, and reproducible assessment as a foundation for the next generation of embodied intelligence.

1. Introduction

As embodied tasks become longer and more complex, binary success rates and manually defined rule-based scores cannot capture progress, failures, recovery, or execution quality. PRM-as-a-Judge 1.5 addresses this gap with process-curve metrics, comprehensive benchmark assessment, RoboPulse++, and an open-source evaluation suite.

  • Motivation: Long, multi-task manipulation requires evaluation that reveals progress, hesitation, regression, and recovery rather than only final task success.Increasing execution complexity produces more diverse failure patterns, making trajectory evolution important for assessing embodied models.
  • Problem: Binary success rates and manually defined rule-based scores (Mees et al., 2022; Liu et al., 2023; Li et al., 2024; Chen et al., 2025a; Fei et al., 2025; Zhang et al., 2025; Yakefu et al., 2025; Chen et al., 2026a) reduce complex trajectories to coarse outcomes and miss execution quality.Failed rollouts can fail early or complete most of a task, while successful rollouts can differ in smoothness, efficiency, hesitation, and recovery behavior.
  • Contribution: PRM-as-a-Judge 1.5 upgrades PRM-as-a-Judge 1.0 (Ji et al., 2026) by converting rollout videos into progress curves, computing metrics, and generating fine-grained assessment reports.Its expanded metric system covers reachability, efficiency, stagnation, failure-side progress, recovery, and success-side quality.
  • Contribution: The toolkit comprehensively assesses mainstream embodied-model rollouts across reachability, failure-side progress, recovery behavior, success-side quality, and failure fingerprints, providing insights beyond official success-rate rankings.The assessment uses available rollout videos from mainstream benchmarks (Chen et al., 2026a).
  • Contribution: RoboPulse++ evaluates process reward models with interval-level samples from robot rollouts, while the open-source suite provides benchmarks, metric implementations, visualization tools, and guidance for reproducible assessment.The platform is designed to compare process evaluators and support more transparent, fair, and process-aware evaluation.

2. PRM-as-a-Judge 1.5

PRM-as-a-Judge 1.5 is an end-to-end framework that uses rollout videos to assess embodied models. It estimates progress, computes process metrics, and generates assessment reports from specified task inputs and optional views.

  • 2. PRM-as-a-Judge 1.5: PRM-as-a-Judge 1.5 provides an end-to-end assessment framework that converts rollout videos into process metrics and model assessment reports.The framework takes rollout videos as input and produces both metrics and a report for embodied model assessment.
  • 2. PRM-as-a-Judge 1.5: The pipeline estimates progress curves with a process reward model, computes process metrics, and generates reports for embodied model assessment.Figure 2 presents this overall pipeline.
  • 2. PRM-as-a-Judge 1.5: Inputs include a case ID, task instruction, rollout video, and optional additional views.These inputs support the framework's assessment process.

Metrics

PRM-as-a-Judge 1.5 converts rollout videos into progress curves and a comprehensive OPD-based assessment covering outcome, process, failure diagnosis, recovery, and execution quality. Its reports synchronize videos with progress curves to expose where and why embodied models fail, stagnate, recover, or approach success.

  • Metrics: PRMs estimate task completion at each frame or timestamp, producing the process curve from paired, sequential, or multi-frame inputs.Examples include Robo-Dopamine (Tan et al., 2025), RoboReward (Lee et al., 2026), RoboMeter (Liang et al., 2026), and RynnValue (Huang et al., 2026).
  • Metrics: The toolkit derives Failure Near-Success (FNS), Drawdown Recovery Ratio (DRR), and Success Quality Score (SQS) as conditioned metrics beyond PRM-as-a-Judge 1.0 (Ji et al., 2026).These metrics assess effectiveness, failure reasons, recovery, and execution quality.
  • Metrics: Reports synchronize rollout videos with progress curves and provide failure location, rollback analysis, and trajectory exploration for fine-grained assessment.They can reveal early failure, near success, stagnation, and recovery.
  • Metrics: PRM-as-a-Judge 1.5 turns a progress curve into an actionable assessment showing how models progress, fail, recover, and complete tasks reliably.The OPD system combines Outcome, Process, and Diagnosis to quantify reachability, progress efficiency, and regression or stagnation patterns.

3. Evaluation

PRM-as-a-Judge 1.5 evaluates embodied models on real-world and simulation manipulation benchmarks using dense process metrics beyond binary success rates. The evaluation shows that model rankings vary across metrics, making success rate alone insufficient.

  • Evaluation scope: PRM-as-a-Judge 1.5 reveals how far models progress, how efficiently they execute, and whether they regress, stagnate, or recover.The evaluation goes beyond binary success rates to characterize process-level behavior.
  • Benchmark: The evaluation covers RoboDojo-RealWorld deployment tasks and RoboDojo-Sim tasks targeting generalization, memory, long-horizon results, and instruction following.The benchmark includes both real-world and simulation manipulation settings.
  • Metrics: Tables 2 and 3 report evaluation performance on RoboDojo-RealWorld and RoboDojo-Sim using milestone completion, MaxProcess, success rate, path-weighted progress length, drawdown recovery, and failure-near-success metrics.These metrics capture progress, execution directness, recovery after drawdown, and failure-side behavior.
  • Evaluation results: Model rankings vary across metrics, and binary success rate is not entirely consistent with process metrics such as SQS, DRR, or FNS.This demonstrates that relying solely on SR is insufficient for evaluating embodied models.

4. Key Findings

The findings show that VLAs generally outperform WAMs, while larger model size does not guarantee better embodied performance. Across model, task, and deployment analyses, π0.5 is generally strongest, Precision tasks perform best, and simulation correlates only weakly with real-world performance.

  • Finding 1: VLAs consistently outperform WAMs across Top-3, Top-5, and Top-10 rankings on RoboDojo-Sim.VLAs exhibit substantially higher representation at each ranking level.
  • Finding 2: Larger model size does not guarantee stronger performance, with no clear positive correlation between parameter count and embodied-model rankings.Several relatively smaller models substantially outperform larger counterparts.
  • Finding 3: π0.5 is generally the best model across different metrics and performs strongly across multiple dimensions.The comparison uses results from Table 4 and Figure 5.
  • Finding 4: Precision tasks perform best, Open-vocabulary tasks are most challenging, and Long-Horizon tasks show the largest model-dependent performance variance.Precision leads Long-Horizon and Generalization across MC@25, MP, and FNS; Open-vocabulary scores cluster at the low end, while Long-Horizon spread best distinguishes model capabilities.
  • Finding 5: Simulation and real-world performance correlate weakly, with Spearman ρ=0.18–0.58 and larger degradation for precise-alignment or complex-contact tasks.Higher-tolerance, simpler-contact tasks show smaller Sim–Real gaps, whereas Classify objects, Hang mugs, and Insert tubes show large negative gaps across most models.

5. Conclusion

PRM-as-a-Judge 1.5 enables fine-grained robot process assessment beyond binary success rates and rule-based scores through a more comprehensive conditioned metric suite. Its process diagnoses also motivate improved progress judges, richer evidence, and a closed evaluation–data–training loop.

  • Conclusion: PRM-as-a-Judge 1.5 moves robotic evaluation beyond binary success rates and rule-based scores with three conditioned metrics that make the OPD system more comprehensive.Large-scale assessments show that process assessment can expose more detailed behavioral signatures.
  • Discussion & Future Directions: Future progress judges should model temporal context because identical local motions can indicate progress or regression depending on prior states.Sequence-style modeling and offline use of past and future observations could provide more stable judgments.
  • Discussion & Future Directions: Future assessment systems could combine progress curves with visual, state, contact, or semantic evidence to make process diagnoses more interpretable.Additional evidence would provide clearer explanations for observed execution patterns.
  • Discussion & Future Directions: Process assessment can identify specific failure patterns and guide failure-targeted data collection and training, including data reweighting and objective design.This closes the loop between evaluation, data, and training.

6. Author List · Appendix · A. Related Work

The paper’s sixth section contains the author list, followed by appendices covering related work and additional materials. The supplied passage identifies Appendix A as Related Work and lists appendices B–F by topic.

  • 6. Author List: The author list appears after the paper’s conclusion and before the references.
  • Appendix: The appendix portion includes Appendix A, titled “Related Work.”
  • Appendix: Appendix B is identified as covering metrics.
  • Appendix: Appendix C is identified as covering RoboPulse++.
  • Appendix: Appendices D–F address efficiency, consistency, and visualization, respectively.

A.1. Evaluation Metrics and Paradigms in Robotics … B.1. Progress Curve Construction

Robotic evaluation methods range from binary outcome scores that collapse trajectories into success or failure to process-oriented approaches that model dense progress over time. PRM-based progress curves provide the common, smoothed input for the paper’s process metrics, capturing advancement, stagnation, regression, and recovery.

  • A.1. Evaluation Metrics and Paradigms in Robotics: Binary Success Rate (SR) compresses an entire manipulation trajectory into one success or failure label, treating smooth successes like repeated-attempt successes and obscuring near-failure progress.
  • A.2. Process Reward Models in Robotics: Process Reward Models (PRMs) convert sparse task outcomes into dense step-level rewards whose episode-wide predictions form progress curves for process evaluation.
  • A.2. Process Reward Models in Robotics: PRMs primarily differ in the temporal context used for prediction, making temporal context a central distinction among process-reward approaches.
  • B. Metric Definitions: The paper’s process metrics share a normalized task-progress curve aligned with sampled video time steps as their common input.
  • B.1. Progress Curve Construction: For each rollout, judge outputs are converted into normalized task progress and aligned with sampled video time steps to construct the progress curve.
  • B. Metric Definitions: A progress value closer to 1 indicates that a rollout is closer to task completion, enabling metric computation over the trajectory’s evolving state.
  • B.1. Progress Curve Construction: Gaussian smoothing is applied before metric computation to reduce local prediction noise in the constructed progress curve.

B.2. OPD Metrics … C.4.3. Metric Computation

The toolkit combines outcome, process, and diagnosis metrics for fine-grained rollout assessment with RoboPulse++, an interval-level benchmark for evaluating progress judges across diverse robot executions. RoboPulse++ uses human-annotated Rising/Falling intervals and a common evaluation protocol for specialized PRMs and general-purpose VLMs.

  • B.2. OPD Metrics: PRM-as-a-Judge organizes rollout assessment into Outcome, Process, and Diagnosis levels, reporting continuous metrics per eligible rollout and aggregating them at task and model levels by median.Outcome captures task reachability, Process measures efficiency over the full rollout, and Diagnosis characterizes regression, stagnation, recovery, and success quality.
  • B.2. OPD Metrics: The OPD metrics quantify reachability, efficiency, and execution failure patterns through MP, MC@q, PPL, CRA, STR, FNS, DRR, and SQS.MP and MC@q measure maximum progress and milestone reachability; PPL measures path efficiency; CRA, STR, and FNS characterize regression, stagnation, and near-success failure; DRR measures post-drawdown recovery, while SQS summarizes successful execution quality.
  • B.2. OPD Metrics: DRR measures recovery after the largest progress loss, reaches 1 for full recovery, and excludes rollouts without a drawdown from aggregation.Its numerator is the best subsequent recovery and its denominator is the largest progress loss.
  • C.1. From Pairwise Comparison to Interval-Level Progress Assessment: RoboPulse++ extends RoboPulse (Ji et al., 2026) from pairwise state comparisons to temporal interval-level progress-direction assessment because absolute completion percentages lack consistent ground truth across diverse tasks.Each interval is labeled Rising or Falling, enabling process-judge evaluation throughout complete robot trajectories.
  • C.2. Dataset Construction and Composition: RoboPulse++ contains 700 manipulation trajectories, 275 task entries, 17,052 frames, and 2,244 human-annotated intervals spanning real-world and simulated manipulation settings.The dataset includes 439 real-world trajectories (62.7%) and 261 simulation trajectories (37.3%), covering atomic, compositional, and long-horizon tasks such as grasping, pushing, tool use, and multi-stage manipulation.
  • C.3. Interval Annotation Protocol: Annotators divide complete trajectories into contiguous intervals and label each as Rising (+1) or Falling (−1) according to task-relevant state changes, using full-video context to resolve subtle fluctuations.The interface supports video scrubbing and direct interval-boundary marking.
  • C.4. Evaluation Protocol; C.4.1. PRM Evaluation; C.4.2. General-Purpose VLM Evaluation: RoboPulse++ evaluates specialized PRMs and general-purpose VLMs against human interval labels using Pair Style for adjacent observations and Sequence Style for temporally ordered sequences.The PRM evaluation covers Robo-Dopamine, RoboMeter, LRM, TOPReward, VLAC, GVL, and PRIMO, with multiple inference formulations where applicable; continuous outputs receive five-frame causal smoothing before direction classification.
  • C.4.2. General-Purpose VLM Evaluation; C.4.3. Metric Computation: Metric computation matches predictions within each annotated interval to its single Rising/Falling ground-truth label, filters the first and last two sampled steps, and reports Macro-F1, Accuracy, and class-wise metrics.General-purpose VLMs include GPT-5.4, Gemini 3.1 Pro, Qwen 3.6 Plus, and Claude Sonnet 4.6 under both evaluation styles.

C.5. Experimental Results · C.5.1. Falling Error Analysis · C.5.2. Context-Dependent Task Analysis

Specialized PRMs provide the strongest progress judgments, but falling progress remains substantially harder to detect than rising progress. Falling errors concentrate on interaction-relation failures, while sequence-level context improves recognition of regressions tied to prior subgoals.

  • C.5. Experimental Results: Robo-Dopamine (Forward) achieves the best overall Macro-F1 and Accuracy, substantially exceeding Gemini 3.1 Pro in Pair Style.
  • C.5. Experimental Results: The best Falling F1 is 0.63 versus 0.92 for Rising F1, with lower Falling recall making regression detection the principal limitation.
  • C.5.1. Falling Error Analysis: Falling progress analysis samples 155 intervals and distinguishes interaction-relation failures from task-order failures.Interaction-relation failures include dropping grasped objects or failing to place them at targets; task-order failures violate expected execution order.
  • C.5.1. Falling Error Analysis: Falling intervals are strongly concentrated in interaction-relation failures, indicating that precise real-world interaction is harder for current embodied models than high-level task logic.
  • C.5.2. Context-Dependent Task Analysis: Context-dependent progress changes can require execution history because pairwise state judgments may miss regressions that undo previously achieved subgoals.The analysis compares Robo-Dopamine (Forward), a Pair Style PRM, with RoboMeter, a Sequence Style PRM.
  • C.5.2. Context-Dependent Task Analysis: RoboMeter outperforms Robo-Dopamine (Forward) on context-dependent cases, with the largest Falling accuracy gap at 0.901 vs. 0.408.The result suggests sequence-level context is particularly useful for recognizing regressions whose meaning depends on previously achieved states or subgoals.

D. Computational Cost of Progress Judges … F.1. Case Studies

The paper profiles progress-judge efficiency, validates that progress curves recover conventional RoboDojo success rates, and illustrates how fine-grained metrics distinguish recovery, execution quality, and failure proximity. These analyses span computational budgets, model-level consistency, and representative rollout case studies.

  • D. Computational Cost of Progress Judges: Increasing batch size reduces Robo-Dopamine-4B runtime from 29.3 to 7.8 minutes and Robo-Dopamine-8B runtime from 31.7 to 9.1 minutes on one H100 GPU.Batched inference can improve throughput when sufficient GPU memory is available.
  • F. Visualization and Case Studies: Progress curves support fine-grained visualization of failure proximity, drawdown recovery, and successful-trajectory quality across representative robot rollouts.The case studies use metric-specific progress-curve behavior to distinguish trajectories that binary success rates would treat similarly.
  • E.1. Evaluation Setup: The RoboDojo analysis aligns 6,076 released rollouts across 42 simulation and 18 real-world tasks, counting a rollout as successful when maximum predicted progress reaches 1 (MP = 1).PRM-derived success rates are compared with model-level benchmark rates using MAE, signed mean difference, and Spearman correlation.
  • E.2. Success-Rate Consistency: PRM-derived success rates closely match benchmark-reported RoboDojo rates, with MAEs of 1.57 and 1.32 percentage points and Spearman correlations of 0.88 and 0.96 for Simulation and Real-World settings, respectively.This supports retaining conventional outcome-level information while enabling fine-grained process assessment.
  • F.1. Case Studies: The DRR examples show that π0.5-SF recovers beyond its pre-regression level after near-zero progress, whereas Xiaomi Robotics 0 fails to regain lost progress and achieves DRR ≈0.Figures 10 and 11 annotate drawdown and recovery intervals, distinguishing sustained post-drawdown recovery from partial recovery.
  • F.1. Case Studies: SQS favors InternVLA-A1 and X-WAM for sustained, low-regression completion, while π0.5 and GalaxeaVLA show unsuccessful attempts and regression that delay completion; FNS likewise ranks near-complete failures above trajectories stalled at initial grasping.Figures 12–15 illustrate that SQS penalizes inefficient corrective behavior and FNS rewards reaching later-stage failure states.

F.2. Evaluation Visualization Suite

PRM-as-a-Judge 1.5 provides a unified visualization suite that supports a continuous workflow from model-level result overview to behavioral inspection. Its automatically generated reports summarize evaluation results and visualize process metrics, success and failure profiles, failure progress, recovery, and individual trajectories.

  • F.2. Evaluation Visualization Suite: The unified suite supports evaluation-result examination from compact model comparison and metric visualization to behavioral inspection.The Model Leaderboard compares evaluated models, while Metric Results presents corresponding evaluation results visually.
  • F.2. Evaluation Visualization Suite: Success/Failure Conditional Profiles facilitate comparisons between different trajectory outcomes.
  • F.2. Evaluation Visualization Suite: Automatically generated reports summarize model-level results and visualize process metrics, success and failure profiles, failure progress, recovery, and individual trajectories.
Loading 2608.14284v1…