Motivations
My research is driven by the goal of Collaborative AI — a symbiosis where humans and AI continuously improve or augment one another in a positive feedback loop. To establish this loop, I believe that artificial intelligence must possess social intelligence to communicate effectively with humans, and be able to adapt as the context evolves.
Current models exhibit an “intelligence” that appears to meet these needs, but this capability often stems from scaling compute, data, and parameters. This approach is not only inefficient (relative to how humans learn), but also promotes a black-box paradigm that discourages scientific methods. For these reasons, I try to move beyond simply scaling anything and everything, and work towards a more mechanistic understanding of how concepts — and by extension, intelligence — emerge; this usually involves tuning one “knob” while holding others constant to isolate specific causes and effects.
I also think this symbiosis could manifest as a form of collective intelligence: humans and agents (or even between clusters of each side) exchanging what they know so that the whole improves or adapts. To this end, I am looking to work on self/co-evolving agents (as a prospective PhD student) through an information theoretic lens. For example, modelling multi-agent system as a communication channel (e.g., how much of the task-relevant signal is accessible/preserved/is faithful when information flows from one agent to another).
Current Works
I enjoy bridging ideas from outside machine learning — such as information theory (if we can even call this outside of ML…), the social sciences, or even philosophy — to construct research questions.
My latest work utilizes the Partial Information Decomposition (PID) framework to analyze multimodal data. This allows me to derive insights from how modalities interact — redundant interactions (overlapping information), unique interactions (exclusive information), synergistic interactions (emergent information) — to produce task-relevant information. I then operationalize these insights to systematically tune these interactions in instruction datasets. In doing so, I showed that carefully increasing redundant interactions could train vision language models that are more robust to hallucinations and ambiguous modalities.
[Ongoing] Self-evolving agents for social behavior. I’m thinking about representing how we (the humans) infer social cues (e.g., emotions) through a “lossy channel” (See figure below). This could help inform “how much information is lost/preserved” for an agent to improve itself in producing a faithful internal state.

Publications
Conference Papers
Self-Captioning Multimodal Interaction Tuning: Amplifying Exploitable Redundancies for Robust Vision Language Models
43rd Proceedings of the International Conference on Machine Learning, Seoul, South Korea (ICML 2026) · 2026
@inproceedings{ ryan2026,
title = { Self-Captioning Multimodal Interaction Tuning: Amplifying Exploitable Redundancies for Robust Vision Language Models },
author = { Yuriel Ryan and Ip Hei Man and Adriel Kuek and Paul Pu Liang and Roy Ka-Wei Lee },
booktitle = { 43rd Proceedings of the International Conference on Machine Learning, Seoul, South Korea (ICML 2026) },
year = { 2026 }
}Large Scale Narrative Analysis of Multimodal Memes
Proceedings of the 20th International AAAI Conference on Web and Social Media, Los Angeles, USA (ICWSM 2026) · 2026
@inproceedings{ peh2026,
title = { Large Scale Narrative Analysis of Multimodal Memes },
author = { Jia Wang Peh and Ming Shan Hee and Bryan (Chen Zhengyu) Tan and Yuriel Wang Jun Long Ryan and Roy Ka-Wei Lee },
booktitle = { Proceedings of the 20th International AAAI Conference on Web and Social Media, Los Angeles, USA (ICWSM 2026) },
year = { 2026 }
}Humor in Pixels: Benchmarking Large Multimodal Models Understanding of Online Comics
Findings of the Association for Computational Linguistics, Suzhou, China (EMNLP 2025) · 2025
@inproceedings{ ryan2025,
title = { Humor in Pixels: Benchmarking Large Multimodal Models Understanding of Online Comics },
author = { Yuriel Ryan and Rui Yang Tan and Kenny Tsu Wei Choo and Roy Ka-Wei Lee },
booktitle = { Findings of the Association for Computational Linguistics, Suzhou, China (EMNLP 2025) },
year = { 2025 }
}Workshop Papers
Balance Human Agency & AI Assistance in the Tussle for the ``Right’’ to Choose, Own, Work, and Learn
Trustworthy AI for Good (AI4GOOD) Workshop, Seoul, South Korea (ICML 2026) · 2026
@inproceedings{ khoo2026,
title = { Balance Human Agency & AI Assistance in the Tussle for the ``Right'' to Choose, Own, Work, and Learn },
author = { Zi-Yu Khoo and Yuriel Ryan and Nicole Heng Yim Oo and Hui En Pang and Eric J. W. Orlowski and Hakim Norhashim and Ruth Wan Theng Chew and Davin Choo and Rachael Hwee Ling Sim and Simon Chesterman and Jungpil Hahn and Bryan Kian Hsiang Low },
booktitle = { Trustworthy AI for Good (AI4GOOD) Workshop, Seoul, South Korea (ICML 2026) },
year = { 2026 }
}Intrigued but Unavailable
I’m always intrigued by the potential of AI and the impact it can make in a variety of topics. Below is a list of projects that I wanted to explore, but currently unable to due to existing commitments. These projects are my “hear me out” ideas. If you are interested in collaborating for any of these, please do reach out :)
Supervised Yearning: Learning the language of Love. Beyond the “5 love languages” (acts of service, quality time, gifts, touch, and affirmation) that naturally involve multiple modalities, I’m intrigued by the interplay of culture and romance. In particular,
- Can models learn “love languages” and adapt that knowledge to different cultures (East vs West).
- Can we quantify romance? What makes a love letter romantic?
- Can a model, trained on being an expert in romance, counter love scams: a type of scam that is not only common in Singapore, but also painful financially and emotionally?
