Paper deep dive
vla-eval: A Unified Evaluation Harness for Vision-Language-Action Models
Suhwan Choi, Yunsung Lee, Yubeen Park, Chris Dongjoo Kim, Ranjay Krishna, Dieter Fox, Youngjae Yu
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 99%
Last extracted: 3/22/2026, 5:03:43 AM
Summary
vla-eval is an open-source evaluation harness for Vision-Language-Action (VLA) models that addresses fragmentation in the evaluation ecosystem. It uses a WebSocket+msgpack protocol and Docker-based isolation to decouple model inference from benchmark execution, enabling parallel evaluation with a 47x throughput improvement. The framework includes a reproducibility audit of published models and a leaderboard aggregating 657 results across 17 benchmarks.
Entities (6)
Relation Signals (3)
vla-eval ā providesleaderboard ā VLA Leaderboard
confidence 100% Ā· We additionally release a VLA leaderboard aggregating 657 published results
vla-eval ā supportsbenchmark ā LIBERO
confidence 100% Ā· The framework supports 13 simulation benchmarks... LIBERO [14]
vla-eval ā supportsmodel ā OpenVLA
confidence 100% Ā· Model servers are implemented for six models: CogACT [12], OpenVLA [11]
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Vision Language Action VLA models are typically evaluated using per benchmark scripts maintained independently by each model repository, leading to duplicated code, dependency conflicts, and underspecified protocols. We present vla eval, an open source evaluation harness that decouples model inference from benchmark execution through a WebSocket msgpack protocol with Docker based environment isolation. Models integrate once by implementing a single predict() method; benchmarks integrate once via a four method interface; the full cross evaluation matrix works automatically. A complete evaluation requires only two commands: vla eval serve and vla eval run. The framework supports 13 simulation benchmarks and six model servers. Parallel evaluation via episode sharding and batch inference achieves a 47x throughput improvement, completing 2000 LIBERO episodes in about 18 minutes. Using this infrastructure, we conduct a reproducibility audit of a published VLA model across three benchmarks, finding that all three closely reproduce published values while uncovering undocumented requirements ambiguous termination semantics and hidden normalization statistics that can silently distort results. We additionally release a VLA leaderboard aggregating 657 published results across 17 benchmarks. Framework, evaluation configs, and all reproduction results are publicly available.
Tags
Links
- Source: https://arxiv.org/abs/2603.13966v1
- Canonical: https://arxiv.org/abs/2603.13966v1
Trouble viewing inline? Open PDF directly ā
Full Text
19,932 characters extracted from source content.
Expand or collapse full text
vla-eval: A Unified Evaluation Harness for Vision-Language-Action Models Suhwan Choi1*, Yunsung Lee1, Yubeen Park1, Chris Dongjoo Kim2, Ranjay Krishna2, Dieter Fox2, Youngjae Yu3 Abstract Vision-Language-Action (VLA) models are typically evaluated using per-benchmark scripts maintained independently by each model repository, leading to duplicated code, dependency conflicts, and underspecified protocols. We present vla-eval, an open-source evaluation harness that decouples model inference from benchmark execution through a WebSocket+msgpack protocol with Docker-based environment isolation. Models integrate once by implementing a single predict() method; benchmarks integrate once via a four-method interface; the full cross-evaluation matrix works automatically. A complete evaluation requires only two commands: vla-eval serve and vla-eval run. The framework supports 13 simulation benchmarks and six model servers. Parallel evaluation via episode sharding and batch inference achieves a 47Ć throughput improvement, completing 2,000 LIBERO episodes in ā¼ 18 minutes. Using this infrastructure, we conduct a reproducibility audit of a published VLA model across three benchmarks, finding that all three closely reproduce published values while uncovering undocumented requirements (ambiguous termination semantics, hidden normalization statistics) that can silently distort results. We additionally release a VLA leaderboard aggregating 657 published results across 17 benchmarks. Framework, evaluation configs, and all reproduction results are publicly available.111https://github.com/allenai/vla-evaluation-harness,222https://allenai.github.io/vla-evaluation-harness/leaderboard I Introduction The rapid development of Vision-Language-Action (VLA) models [11, 2, 12, 1] has produced a growing set of simulation benchmarks, yet the evaluation ecosystem remains fragmented: each model repository independently implements evaluation code for every benchmark it targets, creating four problems. (1) Duplicated effortāseveral model repositories maintain separate copies of benchmark-specific scripts that diverge silently when upstream interfaces change. (2) Dependency and asset conflictsāLIBERO [14] requires Python 3.8 with robosuite, ManiSkill2 [7] requires Python 3.10 with SAPIEN, and CALVIN [16] depends on PyBullet; no single environment satisfies all constraints, and each benchmark requires its own asset setup (scene files, textures, robot descriptions) with ad-hoc installation procedures. (3) Underspecified protocolsāpapers routinely omit seeds, episode counts, normalization statistics, and physics-settling steps, creating barriers to reproduction (Section I). (4) Slow evaluationāa single LIBERO evaluation (2,000 episodes) takes ā¼ 14 hours sequentially; evaluating multiple model variants and ablations across N benchmarks multiplies this cost, making routine comparative studies impractical without parallelism. We address these problems with vla-eval, a unified evaluation harness following the design philosophy of lm-evaluation-harness [6] for language models. Models integrate once, benchmarks integrate once, and the full cross-evaluation matrix works automatically. Our contributions are: ⢠An open-source evaluation harness supporting 13 benchmarks and six model servers with Docker-based isolation and a WebSocket+msgpack protocol; ⢠A reproducibility audit of a published VLA model across three benchmarks, identifying undocumented requirements that prevent independent reproduction; ⢠A parallel evaluation methodology (episode sharding + batch inference) achieving 47Ć speedup: 2,000 LIBERO episodes in ā¼ 18 minutes; ⢠A VLA leaderboard aggregating 657 published results across 17 benchmarks. I Framework Design I-A Architecture vla-eval separates model inference from benchmark execution via a client-server architecture (Fig. 1) using WebSocket with msgpack binary serialization. Each message carries a type (observation, action, episode_start/end), a benchmark-specific payload, a sequence number, and a timestamp. Host (GPU) Docker (optional GPU)ModelServer (ABC)Benchmark (ABC)PredictModelServerConnectionWebSocket clientOpenVLA / CogACT / Ļ0 _0 / ā¦SyncEpisodeRunnerStepBenchmarkLIBERO / CALVIN / SimplerEnv / ā¦obs / actionWebSocket + msgpackact()step Figure 1: System architecture. The model server runs on the host; each benchmark runs in an isolated Docker container with optional GPU access for rendering. A SyncEpisodeRunner orchestrates the observeā ā loop via a Connection (WebSocket + msgpack client). Model servers follow a layered hierarchy. The base ModelServer ABC defines a fully asynchronous interface for advanced use cases; most integrations extend PredictModelServer, which provides a blocking predict(obs, ctx) method (ā¼ 50 lines typical), automatic action chunking with configurable ensemble strategies (newest, average, EMA), and optional batched inference via max_batch_size. Listing 1 shows the complete OpenVLA integration. Dependency isolation. Each model server declares dependencies via PEP 723 inline metadata; vla-eval serve launches it through uv run, creating an isolated environment automatically. Conflicting dependencies (e.g., CogACT pinning transformers==4.40.1 vs. X-VLA requiring transformers>=4.44) coexist without interference, mirroring the Docker-based isolation used for benchmarks. Listing 1: OpenVLA model server (simplified). ⬠class OpenVLAServer(PredictModelServer): def __init__(self, model_path, **kw): super().__init__(**kw) self.model_path = model_path self._model = self._proc = None def _load_model(self): if self._model is not None: return self._proc = AutoProcessor.from_pretrained( self.model_path, trust_remote_code=True) self._model = AutoModelForVision2Seq \ .from_pretrained(self.model_path, torch_dtype=torch.bfloat16, trust_remote_code=True).to("cuda") def predict(self, obs, ctx): self._load_model() img = Image.fromarray( next(iter(obs["images"].values()))) prompt = f"In:ā£Whatā£actionā£shouldā£theā£robot" \ f"ā£takeā£toā£obs[ātask_descriptionā]? :" inp = self._proc(prompt, img) .to("cuda", dtype=torch.bfloat16) act = self._model.predict_action(**inp) return "actions": act Benchmarks follow the same two-level pattern. The base Benchmark ABC defines an async interface; StepBenchmark provides a synchronous wrapper where integrators implement four methods (reset, step, make_obs, get_step_result) that auto-bridge to the async API. Each benchmark runs inside a dedicated Docker image with pinned dependencies. No fixed observation schema is imposed; a recommended convention uses images, states, and task_description. Episode execution and error isolation. A SyncEpisodeRunner orchestrates the blocking observeā ā loop between benchmark and model server. Failures (environment crash, timeout, protocol error) are isolated at the episode level with a structured failure_reason field; the runner re-initializes the environment automatically for the next episode. Declarative configs. Two YAML configs (benchmark + model server) drive each evaluation. We publish all Docker images to ghcr.io with versioned tags and bundle all required assets (scene files, textures, robot descriptions), eliminating the ad-hoc asset installation that each benchmark otherwise requires. A complete evaluation requires only two commands: ⬠vla-eval serve --config model_server.yaml vla-eval run --config benchmark.yaml I-B Supported Benchmarks and Models TABLE I: Supported benchmarks. Docker = image size; Act. = action dimensionality; Status: Complete (reproduction verified), Integrated (adapter + Docker working). Benchmark Docker Act. St. LIBERO [14] 6.0 GB 7D C CALVIN [16] 9.5 GB 7D C SimplerEnv [13] 4.9 GB 7D C ManiSkill2 [7] 9.8 GB 7D I LIBERO-Mem [5] 11.3 GB 7D I Kinetix [15] 9.5 GB 6D I RoboCasa [17] 35.6 GB 7D I VLABench [20] 17.7 GB 7D I MIKASA-Robo [4] 10.1 GB 8D I RoboTwin 2.0 [3] 28.6 GB 14D I RLBench [9] 4.7 GB 8D I RoboCerebra [8] 6.3 GB 7D I LIBERO-Pro [22] 6.2 GB 7D I Table I lists all 13 supported benchmarks with action spaces from 6D to 14D and Docker images from 4.7 to 35.6 GB. Model servers are implemented for six models: CogACT [12], OpenVLA [11], OpenVLA-OFT [10], Ļ0 _0 [2]/Ļ0 _0-FAST [18], GR00T N1 [1], and X-VLA [21]. I-C Parallel Evaluation Environment parallelism uses episode sharding across N Docker containers; inference parallelism uses batched forward passes. We tune parallelism via a demand/supply methodology (Fig. 2): Ī»ā(N)Ī»(N) measures environment throughput as a function of shards, μā(B)μ(B) measures model throughput as a function of batch size, and the operating point satisfies Ī»ā(N)<0.8ā μā(Bā)Ī»(N)<0.8·μ(B^*) to prevent queue buildup. Figure 2: Demand/supply throughput for LIBERO + CogACT on H100. Dashed lines show supply ceilings μā(B)μ(B) at each batch size. The operating point Nā=50N^*\!=\!50 uses 78% of the supply capacity at B=16B\!=\!16, leaving headroom to absorb burst arrivals and prevent queue buildup; beyond N=80N\!=\!80, environment overhead causes throughput to drop. On LIBERO with CogACT-7B (H100 model server, separate benchmark host), episode sharding from N=1N\!=\!1 to N=50N\!=\!50 increases environment throughput by 32.6Ć (Ī»: 11.2ā 364.6 obs/s), and batch inference from B=1B\!=\!1 to B=16B\!=\!16 increases model server throughput by 2.8Ć (μ: 165.2ā 468.2 obs/s). Combined, 2,000 episodes complete in ā¼ 18 minutes versus ā¼ 14 hours sequentially, a 47Ć wall-clock speedup. The same methodology applies to CALVIN: 1,000 chained sequences with 16 shards complete in ā¼ 33 minutes (16Ć speedup), where PyBulletās CPU-bound rendering is the bottleneck, limiting the benefit of additional shards. SimplerEnv (288 episodes across 3 seeds) completes in ā¼ 8.5 minutes with 16 shards (12Ć speedup), where SAPIEN GPU rendering is the per-container bottleneck. Fig. 3 shows the wall-clock reduction for all three benchmarks. Figure 3: Wall-clock evaluation time: sequential vs. batch parallel. LIBERO: 2,000 episodes, 50 shards, B=16B\!=\!16. CALVIN: 1,000 sequences, 16 shards. SimplerEnv: 288 episodes (3 seeds), 16 shards. I Reproducibility Audit We evaluate DB-CogACT [19] (a CogACT [12] re-implementation with a modern base model) across three benchmarks using published scores as baselines. All experiments use fixed seeds and versioned Docker images from ghcr.io. LIBERO: 4 suites, 10 tasks Ć 50 episodes (2,000 total), 50 shards. CALVIN: ABCā , 1,000 chained sequences, 16 shards. SimplerEnv: 4 WidowX tasks, 24 episodes Ć 3 seeds (288 total), 16 shards. I-A Results TABLE I: Reproduction results vs. published reference values. Ī = ours ā- reference. Success rates in %. Benchmark Suite / Metric Ours Ref. Ī LIBERO Spatial 95.2 93.8 +1.4 Object 98.6 97.8 +0.8 Goal 95.2 96.2 ā-1.0 Long-Horizon 89.6 91.8 ā-2.2 CALVIN Avg Len (ABCā ) 4.051 4.063 ā-0.012 SimplerEnv Avg SR 72.22 69.45 +2.77 Table I summarizes the results. LIBERO (4 suites, 2,000 episodes) and SimplerEnv (288 episodes, 3-seed average) reproduce within ± 3 percentage points of published success rates. CALVIN (1,000 chained sequences) reproduces within 0.3% of the published average chain length. I-B Sources of Discrepancy During reproduction, we identified two classes of undocumented requirements that can silently distort evaluation results. Ambiguous termination semantics. In SimplerEnv, the Gymnasium terminated flag signals a transient success event (e.g., a block momentarily stacked) rather than episode termination. Stopping early on this flag inflates scores because the robot may disturb the object after a momentary successāthe correct protocol runs until truncated at max_episode_steps. This semantic overloading is undocumented and typically requires reading simulator source code to discover. Hidden normalization. CALVIN requires hardcoded observation normalization statistics (mean and standard deviation for 15 robot-state and 24 scene-state dimensions) derived from the training dataset. These statistics are absent from the official evaluation documentation and require tracing the training codebase to recover. Without these exact values, observation preprocessing produces incorrect inputs, degrading model performance. IV VLA Leaderboard Figure 4: VLA leaderboard (17 benchmarks, https://allenai.github.io/vla-evaluation-harness/leaderboard). Shown: models with >>10 citations. Filterable by benchmark and model; contributions via pull request. We release a VLA leaderboardLABEL:fn:leaderboard aggregating 657 results across 17 benchmarks and 509+ configurations from 2,685 papers. IV-A Curation Evaluation protocols vary across papers: SimplerEnv spans three incomparable robot configurations; CALVIN ABCā and ABCDā splits are not comparable; LIBERO papers report 4 or 5 suites. We established canonical protocol definitions for each benchmark (task subsets, metrics, splits, and comparability constraints) as contribution guidelines. An AI agent (Claude Code with Opus 4.6) reviewed 2,685 papers via MCP tool integrations (arXiv, Semantic Scholar, PDF reader) to extract and normalize results against canonical protocols. A human operator then reviewed every entry, resolving anomalies and ambiguous cases. We acknowledge that the leaderboard may contain errors despite these efforts. Each entry records its provenance (curated_by field) and the site displays all protocol caveats transparently. We accept community contributions (corrections, missing results, and protocol clarifications) via pull request with automated schema validation. IV-B Cross-Benchmark Analysis Fig. 5 shows the distribution of benchmark coverage across 509+ models and the 17 benchmarks tracked in the leaderboard: 81% are evaluated on a single benchmark, and only 6% on three or more. Cross-benchmark comparison is therefore rare, limiting our ability to assess general model capability across diverse environments and embodiments. This underscores the need for a unified evaluation framework that makes cross-benchmark evaluation practical. Figure 5: Distribution of benchmark coverage per model. 81% of the 509+ models in the leaderboard are evaluated on only one benchmark. Only 3 models (0.6%) are evaluated on 5 or more. Counts are lower bounds, as models may evaluate on benchmarks not tracked in the leaderboard. A unified evaluation framework is essential to make systematic cross-benchmark comparison feasible. V Discussion Our audit shows that even well-documented benchmarks leave critical details implicitāoverloaded termination flags (SimplerEnv), hidden normalization statistics (CALVIN)āeach a potential source of silent score distortion. vla-eval mitigates these risks by saving the full evaluation configuration (Docker image tag, seeds, episode counts, action space parameters) alongside every result, making any run reproducible from a single config file. Limitations. Our audit covers one model across three benchmarks; a broader cross-model analysis is planned. The framework currently targets simulation; real-robot evaluation is out of scope. Leaderboard results are extracted from published papers, not independently verified through re-evaluation. Finally, supported metrics are limited to task success rate; richer evaluation dimensions such as motion quality, efficiency, and safety remain future work. References [1] J. Bjorck, F. CastaƱeda, N. Cherniadev, X. Da, R. Ding, L. Fan, et al. (2025) GR00T N1: an open foundation model for generalist humanoid robots. arXiv preprint arXiv:2503.14734. Cited by: §I, §I-B. [2] K. Black, N. Brown, D. Driess, A. Esmail, M. Equi, C. Finn, et al. (2024) Ļ0 _0: A vision-language-action flow model for general robot control. arXiv preprint arXiv:2410.24164. Cited by: §I, §I-B. [3] T. Chen, Z. Chen, B. Chen, Z. Cai, Y. Liu, Z. Li, et al. (2025) RoboTwin 2.0: a scalable data generator and benchmark with strong domain randomization for robust bimanual robotic manipulation. arXiv preprint arXiv:2506.18088. Cited by: TABLE I. [4] E. Cherepanov, N. Kachaev, A. K. Kovalev, and A. I. Panov (2025) Memory, benchmark & robots: a benchmark for solving complex tasks with reinforcement learning. arXiv preprint arXiv:2502.10550. Cited by: TABLE I. [5] N. Chung, T. Hanyu, T. Nguyen, H. Le, F. Bumgarner, D. M. H. Nguyen, et al. (2025) Rethinking progression of memory state in robotic manipulation: an object-centric perspective. arXiv preprint arXiv:2511.11478. Cited by: TABLE I. [6] L. Gao, J. Tow, B. Abbasi, S. Biderman, S. Black, A. DiPofi, et al. (2024) The language model evaluation harness. Zenodo. External Links: Document Cited by: §I. [7] J. Gu, F. Xiang, X. Li, Z. Ling, X. Liu, T. Mu, et al. (2023) ManiSkill2: a unified benchmark for generalizable manipulation skills. In ICLR, Cited by: §I, TABLE I. [8] S. Han, B. Qiu, Y. Liao, S. Huang, C. Gao, S. Yan, et al. (2025) RoboCerebra: a large-scale benchmark for long-horizon robotic manipulation evaluation. In NeurIPS Datasets and Benchmarks, Cited by: TABLE I. [9] S. James, Z. Ma, D. R. Arrojo, and A. J. Davison (2020) RLBench: the robot learning benchmark & learning environment. IEEE Robotics and Automation Letters. Cited by: TABLE I. [10] M. J. Kim, C. Finn, and P. Liang (2025) Fine-tuning vision-language-action models: optimizing speed and success. arXiv preprint arXiv:2502.19645. Cited by: §I-B. [11] M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, et al. (2024) OpenVLA: an open-source vision-language-action model. In CoRL, Cited by: §I, §I-B. [12] Q. Li, Y. Liang, Z. Wang, L. Luo, X. Chen, M. Liao, et al. (2024) CogACT: a foundational vision-language-action model for synergizing cognition and action in robotic manipulation. arXiv preprint arXiv:2411.19650. Cited by: §I, §I-B, §I. [13] X. Li, K. Hsu, J. Gu, O. Mees, K. Pertsch, H. R. Walke, et al. (2024) Evaluating real-world robot manipulation policies in simulation. In CoRL, Cited by: TABLE I. [14] B. Liu, Y. Zhu, C. Gao, Y. Feng, Q. Liu, Y. Zhu, et al. (2023) LIBERO: benchmarking knowledge transfer for lifelong robot learning. In NeurIPS Datasets and Benchmarks, Cited by: §I, TABLE I. [15] M. Matthews, M. Beukman, C. Lu, and J. Foerster (2025) Kinetix: investigating the training of general agents through open-ended physics-based control tasks. In ICLR, Cited by: TABLE I. [16] O. Mees, L. Hermann, E. Rosete-Beas, and W. Burgard (2022) CALVIN: a benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks. IEEE Robotics and Automation Letters. Cited by: §I, TABLE I. [17] S. Nasiriany, A. Maddukuri, L. Zhang, A. Parikh, A. Lo, A. Joshi, et al. (2024) RoboCasa: large-scale simulation of household tasks for generalist robots. In RSS, Cited by: TABLE I. [18] K. Pertsch, K. Stachowicz, B. Ichter, D. Driess, S. Nair, Q. Vuong, et al. (2025) FAST: efficient action tokenization for vision-language-action models. arXiv preprint arXiv:2501.09747. Cited by: §I-B. [19] B. Xie, E. Zhou, F. Jia, H. Shi, H. Fan, H. Zhang, et al. (2025) Dexbotic: open-source vision-language-action toolbox. arXiv preprint arXiv:2510.23511. Cited by: §I. [20] S. Zhang, Z. Xu, P. Liu, X. Yu, Y. Li, Q. Gao, et al. (2024) VLABench: a large-scale benchmark for language-conditioned robotics manipulation with long-horizon reasoning tasks. arXiv preprint arXiv:2412.18194. Cited by: TABLE I. [21] J. Zheng, J. Li, Z. Wang, D. Liu, X. Kang, Y. Feng, et al. (2025) X-VLA: soft-prompted transformer as scalable cross-embodiment vision-language-action model. arXiv preprint arXiv:2510.10274. Cited by: §I-B. [22] X. Zhou, Y. Xu, G. Tie, Y. Chen, G. Zhang, D. Chu, et al. (2025) LIBERO-PRO: towards robust and fair evaluation of vision-language-action models beyond memorization. arXiv preprint arXiv:2510.03827. Cited by: TABLE I.