RESEARCH PAPER

RoboHarness: Memory-Driven Orchestration of Heterogeneous Robot Policies for Long-Horizon Planning

Huang, Jinbang; Hu, Yuanzhao; Li, Zhiyuan; Qi, Ran; Xiao, Yixin; Zhang, Zhanguang; Coates, Mark; Cao, Tongtong; Zhang, Yingxue

Classification

View four quadrants
Major category
Related resources
Architecture
Not applicable
Prediction paradigm
Not applicable
Source review status
Verified from primary sources

Category review. RoboHarness is a 2026 memory-and-planning orchestrator of independently developed policies. Retrieval and motion-planned handoffs do not constitute learned future/action prediction, and the framework is not a generic representation backbone. Reading evidence

AT A GLANCE

Contribution

We propose RoboHarness, a unified framework that encapsulates independently developed robot control systems as reusable agentic skills.

Abstract

Long-horizon robotic tasks require diverse capabilities that no single policy can reliably provide. Heterogeneous policies offer complementary strengths, but orchestrating them requires reasoning over uncertain capability boundaries and cross-policy distribution mismatch, which are largely overlooked by existing planning methods built on homogeneous, predefined skills with fixed applicability. We propose RoboHarness, a unified framework that encapsulates independently developed robot control systems as reusable agentic skills. Although instantiated in this work with VLAs, RL policies, and task-and-motion planning (TAMP) systems, RoboHarness is designed as a general framework compatible with a broader range of robot policies, such as navigation policies, model predictive controllers, and world-action models. RoboHarness uses multi-modal execution memory and online evidence to characterize policy capability boundaries for capability-aware decomposition and routing. To stabilize policy handoffs, its Memory Bridge retrieves execution trajectories associated with the next policy, estimates its in-distribution state region, and guides the robot toward that region without joint policy retraining. Extensive experiments on three public benchmarks, 500 customized tasks, and 135 real-robot experiments demonstrate effective capability-aware routing and stable policy orchestration, yielding substantial improvements in zero-shot long-horizon planning and out-of-distribution robustness.

Affiliations

Huawei Noah's Ark Lab; University of British Columbia; University of Toronto; McGill University,; Department of Foundation Model, 2012 Labs; Work done during the intership at Huawei Noah's Ark Lab