Yuriel Ryan
Contents
  1. Motivations
  2. Current Works
  3. Selected Publications
    1. Conference Papers
    2. Workshop Papers
    3. Under Review
    4. Intrigued but Unavailable

Motivations

My research is driven by the goal of Collaborative AI — a deliberate push to have humans and AI continuously augment each other in a positive feedback loop. To establish this, I believe that artificial intelligence should be able to continuously adapt or evolve to humans or to the problem at hand. This means possessing qualities that go beyond rote learning (e.g., social intelligence or autonomy to propose tasks to learn on its own).

Current models exhibit an “intelligence” that appears to meet these needs, but this capability often stems from scaling compute, data, and parameters. This approach is not only inefficient (relative to how humans learn), but also promotes a black-box paradigm that discourages scientific methods. For these reasons, I try to move beyond simply scaling anything and everything, and work towards a more principled understanding of how concepts — and by extension, intelligence — emerge; this usually involves adapting a theoretical framework and/or tuning one “knob” while holding others constant to isolate specific causes and effects.

I also think this human-AI symbiosis could manifest as a form of Collective Intelligence: humans and agents (or even between clusters of each side) exchanging what they know for the collective to benefit. To this end, I am looking to work on self/co-evolving agents (as a prospective PhD student) through the different lenses of Information Theory. For example, this could involve modelling a multi-agent system as communication channels (e.g., how much of the task-relevant signal is preserved/faithful when information flows from one agent to another) or considering what is accessible or useful information to an agent (e.g., V-Information or PID).

Current Works

I enjoy bridging ideas from outside machine learning — such as information theory (if we can even call this outside of ML…), the social sciences, or even philosophy — to construct research questions.

My latest work utilizes the Partial Information Decomposition (PID) framework to analyze multimodal data. This allows me to derive insights from how modalities interact — redundant interactions (overlapping information), unique interactions (exclusive information), synergistic interactions (emergent information) — to produce task-relevant information. I then operationalize these insights to systematically tune these interactions in instruction datasets. In doing so, I showed that carefully increasing redundant interactions could train vision language models that are more robust to hallucinations and ambiguous modalities.

[Ongoing] Self-evolving agents for social behavior. I’m thinking about representing how we (the humans) infer social cues (e.g., emotions) through a “lossy channel” (See figure below). This could help inform “how much information is lost/preserved” for an agent to improve itself in producing a faithful internal state.

Diagram of social inference as a lossy channel: a distressed person (internal state Z) encodes a smiling face ("I'm fine", observed behavior X), and a perceiver infers a calm state Z-hat, reading them as fine.



Selected Publications

You can also find the updated articles on my Google Scholar.

Conference Papers

2026 ICML

Self-Captioning Multimodal Interaction Tuning: Amplifying Exploitable Redundancies for Robust Vision Language Models

Yuriel Ryan, Ip Hei Man, Adriel Kuek, Paul Pu Liang, Roy Ka-Wei Lee

43rd Proceedings of the International Conference on Machine Learning, Seoul, South Korea (ICML 2026)

Details Paper (opens in a new tab)
Abstract

Current vision language models face hallucination and robustness issues against ambiguous or corrupted modalities. We hypothesize that these issues can be addressed by exploiting the shared information between modalities to compensate for the impaired one. To this end, we analyze multimodal interactions – redundant (shared), unique (exclusive), and synergistic (emergent) task-relevant information provided by the modalities – to determine their impacts on model reliability. Specifically, amplifying redundant interactions would increase this exploitable shared information to resolve these issues; yet, modern instruction datasets often eliminate redundancies to prioritize visual grounding. We bridge this gap through a self-captioning workflow featuring a Multimodal Interaction Gate: a mechanism to convert unique interactions into redundant interactions. Our findings suggest that increasing redundancy can reduce visual induced errors by 38.3% and improve consistency by 16.8%.

2026 ICWSM

Large Scale Narrative Analysis of Multimodal Memes

Jia Wang Peh, Ming Shan Hee, Bryan (Chen Zhengyu) Tan, Yuriel Wang Jun Long Ryan, Roy Ka-Wei Lee

Proceedings of the 20th International AAAI Conference on Web and Social Media, Los Angeles, USA (ICWSM 2026)

Details Paper (opens in a new tab)
Abstract

Current computational approaches to meme analysis primarily focus on individual memes, with an emphasis on tasks such as hateful content detection and sentiment analysis. In contrast, corpus-level analyses, which are necessary to reveal in-depth thematic narratives embedded in meme corpora, have remained within the purview of social science research. While qualitative methods such as inductive content analysis provide deeper insights, they are labor-intensive and lack scalability. To address this gap, we introduce MemeTopicTrees (MemeTT), a zero-shot pipeline that clusters multimodal memes based on their targets, aspects, sentiments, and opinions. In addition to clustering, MemeTT generates a descriptive narrative for each cluster to provide a nuanced understanding of meme corpora. By integrating multimodal aspect-based sentiment analysis with hierarchical clustering, MemeTT automates both semantic analysis and narrative generation at scale. Evaluated on a combined pool of three datasets spanning political, public health, and defense domains, MemeTT successfully produces fine-grained clusters and coherent narratives, with narrative relevance most pronounced at the target and aspect levels. Evaluations show that the best-performing models are highly accurate at identifying meme targets and aspects, although maintaining high accuracy for opinions and sentiments remains challenging. Furthermore, while cluster distinctiveness is robust at the target level, room for improvement remains at lower clustering levels. Despite these challenges, this approach offers a scalable solution for analyzing public sentiment and discourse in meme corpora. We lay the foundation for future research on the understudied task of automated meme corpus analysis.

2025 EMNLP

Humor in Pixels: Benchmarking Large Multimodal Models Understanding of Online Comics

Yuriel Ryan*, Rui Yang Tan*, Kenny Tsu Wei Choo, Roy Ka-Wei Lee * co-first authors

Findings of the Association for Computational Linguistics, Suzhou, China (EMNLP 2025)

Details Paper (opens in a new tab)
Abstract

Understanding humor is a core aspect of social intelligence, yet it remains a significant challenge for Large Multimodal Models (LMMs). We introduce PixelHumor, a benchmark dataset of 2,800 annotated multi-panel comics designed to evaluate LMMs’ ability to interpret multimodal humor and recognize narrative sequences. Experiments with state-of-the-art LMMs reveal substantial gaps: for instance, top models achieve only 61% accuracy in panel sequencing, far below human performance. This underscores critical limitations in current models’ integration of visual and textual cues for coherent narrative and humor understanding. By providing a rigorous framework for evaluating multimodal contextual and narrative reasoning, PixelHumor aims to drive the development of LMMs that better engage in natural, socially aware interactions.

Workshop Papers

2026 ICML

Balance Human Agency & AI Assistance in the Tussle for the ``Right’’ to Choose, Own, Work, and Learn

Zi-Yu Khoo, Yuriel Ryan, Nicole Heng Yim Oo, Hui En Pang, Eric J. W. Orlowski, Hakim Norhashim, Ruth Wan Theng Chew, Davin Choo, Rachael Hwee Ling Sim, Simon Chesterman, Jungpil Hahn, Bryan Kian Hsiang Low

Trustworthy AI for Good (AI4GOOD) Workshop, Seoul, South Korea (ICML 2026)

Details Paper (opens in a new tab)
Abstract

AI increasingly mediates and augments daily life, yet its technical properties risk reshaping the way users choose (what they consume), own (what they create), learn, and work. These risks reflect widely recognized AI principles, which unfortunately remain largely high-level and unoperationalizable. As the risks arising from AI assistance continue to subtly erode human agency, this work takes a different approach to operationalizing AI principles by translating the risks into concrete research questions that can guide the design of AI systems. This enables the research community to balance human agency and AI assistance in AI development and align it with the goal of benefiting humans.

Under Review

2026 NeurIPS

PID-Chain: Distinguishing Heuristics from Answers in Multimodal Sequential Targets

Yuriel Ryan

Interpretability as a Science Workshop (NeurIPS 2026), under review

Details
Abstract

Humans naturally anticipate tasks from context. A geometry figure in a quiz primes a student to expect trigonometry before the question is read. These heuristics reflect how much (partial) information each input modality potentially carries before even fixing a task. Generalizing this intuition, interpreting how a LLM anticipates task could help inform tasks with sequential targets. We extend partial information decomposition (PID) to accommodate these structured targets and evaluate the empirical interpretations. We test this by quantifying redundant (shared), unique (exclusive), and synergistic (emergent) task-relevant information across successive targets: a domain or topic, followed by the actual task or question. Our method, PID-Chain, shows that specifying the task resolves an apparently redundancy-dominated dataset into distinct per-task structures. The first task can be purely redundant while the subsequent task becomes synergistic. Decomposing contributions in sequential targets isolates key parts of a reasoning chain that lead to emergent properties like sarcasm. Further experiments show that PID-Chain recovers the enumerated ground-truth decomposition on synthetic benchmarks and yields results consistent with previous estimators.


Intrigued but Unavailable

I’m always intrigued by the potential of AI and the impact it can have in a variety of topics. Below is a list of projects that I wanted to explore, but am currently unable to due to existing commitments. These projects are my “hear me out” ideas. If you are interested in collaborating for any of these, please do reach out :)

Supervised Yearning: Learning the language of Love. Beyond the “5 love languages” (acts of service, quality time, gifts, touch, and affirmation) that naturally involve multiple modalities, I’m intrigued by the interplay of culture and romance. In particular,

  • Can models learn “love languages” and adapt that knowledge to different cultures (East vs West)?
  • Can we quantify romance? What makes a love letter romantic?
  • Can a model, trained on being an expert in romance, counter love scams: a type of scam that is not only common in Singapore, but also painful financially and emotionally?