RESEARCH PAPER

Artificial Foveated Perception for Mitigating Shortcut Learning in Robotic Foundation Models

Sun, Xiatao; Zhuang, Yuan; Negrete, Mateo Sanchez Lopez; Coldea, Matei-Victor; Liang, Chen; Zhang, Haoyang; Liu, Che; Zeng, Ziyao; Li, Shawn; Wang, Qian; Miao, Fei; Rakita, Daniel

Classification

View four quadrants
Major category
VLA
Architecture
Not applicable
Prediction paradigm
Not applicable
Source review status
Verified from primary sources

Category review. AFP provides language-conditioned relevance masks as auxiliary supervision while fine-tuning existing robot policies. At deployment AFP is removed and the original policy directly generates actions from observations. This is a VLA perception/optimization method, not a canonical pretrained visual encoder or a new world predictor. Reading evidence

AT A GLANCE

Contribution

We propose Artificial Foveated Perception (AFP), a lightweight, policy-agnostic module that takes the same vision and language inputs as Vision-Language-Action and World Action Model pipelines and predicts task-conditioned masks over relevant objects, the robot, and other action-critical regions. We evaluate AFP across state-of-the-art robotic foundation models and show that foveated perception reduces fine-tuning time, suppresses overfitting, and improves generalization under environmental perturbations.

Abstract

Robotic foundation models have recently made substantial progress in multi-task capability, cross-embodiment transfer, and language-conditioned control. Yet robust deployment across diverse real-world settings remains difficult, in part because policies often fail to distinguish causally relevant visual structure from spurious scene-level correlations. We identify this failure mode as shortcut learning: the tendency to exploit predictive but non-causal correlations in the training distribution rather than the task-relevant visual evidence that determines successful action. Although shortcut learning has been extensively studied in computer vision and broader machine learning, its role in robotic foundation models remains comparatively underexplored. We propose Artificial Foveated Perception (AFP), a lightweight, policy-agnostic module that takes the same vision and language inputs as Vision-Language-Action and World Action Model pipelines and predicts task-conditioned masks over relevant objects, the robot, and other action-critical regions. We use these masks primarily as an auxiliary grounding signal during fine-tuning, aligning policy attention with task-relevant regions while leaving the core architecture unchanged. After fine-tuning, the policy executes on the original observation stream without requiring AFP in the control loop. We evaluate AFP across state-of-the-art robotic foundation models and show that foveated perception reduces fine-tuning time, suppresses overfitting, and improves generalization under environmental perturbations. Ablations over mask quality and grounding-loss design further show that these gains arise from directing policy learning toward task-relevant visual evidence. These results suggest that task-conditioned foveated perception is a practical mechanism for making robotic foundation models more robust, data-efficient, and scalable.

Affiliations

1 Yale University 2 University of Connecticut; 3 Peking University 4 Imperial College London 5 Digients