PID-Chain: Distinguishing Heuristics from Answers in Multimodal Sequential Targets
Interpretability as a Science Workshop (NeurIPS 2026), under review, 2026
Humans naturally anticipate tasks from context. A geometry figure in a quiz primes a student to expect trigonometry before the question is read. These heuristics reflect how much (partial) information each input modality potentially carries before even fixing a task. Generalizing this intuition, interpreting how a LLM anticipates task could help inform tasks with sequential targets. We extend partial information decomposition (PID) to accommodate these structured targets and evaluate the empirical interpretations. We test this by quantifying redundant (shared), unique (exclusive), and synergistic (emergent) task-relevant information across successive targets: a domain or topic, followed by the actual task or question. Our method, PID-Chain, shows that specifying the task resolves an apparently redundancy-dominated dataset into distinct per-task structures. The first task can be purely redundant while the subsequent task becomes synergistic. Decomposing contributions in sequential targets isolates key parts of a reasoning chain that lead to emergent properties like sarcasm. Further experiments show that PID-Chain recovers the enumerated ground-truth decomposition on synthetic benchmarks and yields results consistent with previous estimators.