Diagnosing and Mitigating Perception-Decision Misalignment in Omni-LLMs via Modality Subspace Activation
Classification
View four quadrants- Major category
- Not assigned
- Quadrant
- Not applicable
- Architecture
- Not applicable
- Prediction paradigm
- Not applicable
- Subcategories
- Not recorded
- Source review status
- Verified from primary sources
Category review. MSA changes Omni-LLM hidden states for multimodal multiple-choice question answering and measures option sensitivity after modality removal. It provides no world-transition predictor, robot action interface, or general pretrained encoder. Its CMS benchmark/metric accompanies a QA intervention, so forcing the entire work into WAM Components or robot evaluation metrics would misstate its scope. Reading evidence
Contribution
To rectify this, we propose Modality Subspace Activation (MSA), a training-free inference-time framework that uses Singular Value Decomposition (SVD) to estimate modal activation strengths.
Abstract
Omni-Large Language Models (Omni-LLMs) power complex multi-modal reasoning in applications like World Action Models and autonomous agents. However, their strong performance often masks a profound Perceptual-Decision Misalignment (PDM), where decisions remain unfaithful to multi-modal perceptions. To diagnose this, we formalize Causal Modality Sensitivity (CMS), operationalized via a dual-lens framework: Answer Retention Rate (ARR) at the macro behavioral level, and Logit Angular Discrepancy (LAD) to track microscopic distribution shifts. We also curate CausalMSBench, a diagnostic dataset isolating language priors. Benchmarking reveals that popular Omni-LLMs exhibit critically low CMS, showing negligible distribution shifts even when key modalities are removed. To rectify this, we propose Modality Subspace Activation (MSA), a training-free inference-time framework that uses Singular Value Decomposition (SVD) to estimate modal activation strengths. MSA dynamically balances modal projections in the last hidden state, effectively restoring CMS across benchmarks.
Affiliations
1Xiamen University; 2Shanghai Artificial Intelligence Laboratory; 3Tencent YouTu Laboratory