RESEARCH PAPER

Diagnosing and Mitigating Perception-Decision Misalignment in Omni-LLMs via Modality Subspace Activation

Jiang, Hongbo; Li, Jie; Shen, Yunhang; Xie, Tianyu; Dai, Pingyang

Classification

View four quadrants
Major category
Not assigned
Architecture
Not applicable
Prediction paradigm
Not applicable
Subcategories
Not recorded
Source review status
Verified from primary sources

Category review. MSA changes Omni-LLM hidden states for multimodal multiple-choice question answering and measures option sensitivity after modality removal. It provides no world-transition predictor, robot action interface, or general pretrained encoder. Its CMS benchmark/metric accompanies a QA intervention, so forcing the entire work into WAM Components or robot evaluation metrics would misstate its scope. Reading evidence

AT A GLANCE

Contribution

To rectify this, we propose Modality Subspace Activation (MSA), a training-free inference-time framework that uses Singular Value Decomposition (SVD) to estimate modal activation strengths.

Abstract

Omni-Large Language Models (Omni-LLMs) power complex multi-modal reasoning in applications like World Action Models and autonomous agents. However, their strong performance often masks a profound Perceptual-Decision Misalignment (PDM), where decisions remain unfaithful to multi-modal perceptions. To diagnose this, we formalize Causal Modality Sensitivity (CMS), operationalized via a dual-lens framework: Answer Retention Rate (ARR) at the macro behavioral level, and Logit Angular Discrepancy (LAD) to track microscopic distribution shifts. We also curate CausalMSBench, a diagnostic dataset isolating language priors. Benchmarking reveals that popular Omni-LLMs exhibit critically low CMS, showing negligible distribution shifts even when key modalities are removed. To rectify this, we propose Modality Subspace Activation (MSA), a training-free inference-time framework that uses Singular Value Decomposition (SVD) to estimate modal activation strengths. MSA dynamically balances modal projections in the last hidden state, effectively restoring CMS across benchmarks.

Affiliations

1Xiamen University; 2Shanghai Artificial Intelligence Laboratory; 3Tencent YouTu Laboratory