Publications
2026
The Yōkai Learning Environment: Tracking Beliefs Over Space and Time
Constantin Ruhdorfer, Matteo Bortoletto, Johannes Forkel, Jakob Nicolaus Foerster, Andreas Bulling
Proc. Reinforcement Learning Conference (RLC), 2026.
AbstractLinksBibTeXProject
The ability to cooperate with unknown partners is a central challenge in cooperative AI and widely studied in the form of zero-shot coordination (ZSC) which evaluates an algorithm by measuring the performance of independently trained agents when paired. The Hanabi Learning Environment (HLE) has become the dominant benchmark for ZSC, but recent work has achieved near-perfect inter-seed cross-play performance, limiting its ability to track algorithmic progress. We introduce the Yokai Learning Environment (YLE) - an open-source multi-agent RL benchmark in which effective collaboration requires building common ground by tracking and updating beliefs over moving cards, reasoning under ambiguous hints, and deciding when to terminate the game based on inferred shared knowledge - features absent in the HLE, where beliefs are tied to hand slots and hints are truthful by rule. We evaluate the leading ZSC methods, including High-Entropy IPPO, Other-Play, and Off-Belief Learning, which achieve near-perfect inter-seed cross-play in the HLE, and show that in the YLE they exhibit persistent SP–XP gaps, degraded early-ending calibration, and weaker belief representations in cross-play, indicating failure to maintain consistent internal models with unseen partners. Methods that perform best in the HLE do not perform best in the YLE, indicating that progress measured on a single benchmark may not generalise. Together, these results establish YLE as a challenging new ZSC benchmark.
@inproceedings{ruhdorfer26_rlc,
title = {The Yōkai Learning Environment: Tracking Beliefs Over Space and Time},
author = {Ruhdorfer, Constantin and Bortoletto, Matteo and Forkel, Johannes and Foerster, Jakob Nicolaus and Bulling, Andreas},
year = {2026},
booktitle = {Proc. Reinforcement Learning Conference (RLC)},
}
Unsupervised Partner Design Enables Robust Ad-hoc Teamwork
Constantin Ruhdorfer, Matteo Bortoletto, Victor Oei, Anna Penzkofer, Andreas Bulling
Proc. International Conference on Machine Learning (ICML), 2026.
AbstractLinksBibTeXProject Spotlight
We introduce Unsupervised Partner Design (UPD), a population-free multi-agent reinforcement learning method for robust ad-hoc teamwork. UPD generates training partners on-the-fly and selects them adaptively based on a learnability criterion, removing the need for pre-trained partner populations or manual parameter tuning. We show that this simple mechanism enables effective partner diversity and can be extended to joint partner-environment selection when a procedural level generator is available. Across Level-Based Foraging, Overcooked-AI, and the Overcooked Generalisation Challenge, UPD consistently achieves strong performance compared to both population-based and population-free baselines. In a human-AI user study, agents trained with UPD achieve higher returns and are rated as more adaptive, more human-like, and less frustrating than all evaluated baseline methods.
@inproceedings{ruhdorfer26_icml,
title = {Unsupervised Partner Design Enables Robust Ad-hoc Teamwork},
author = {Ruhdorfer, Constantin and Bortoletto, Matteo and Oei, Victor and Penzkofer, Anna and Bulling, Andreas},
year = {2026},
booktitle = {Proc. International Conference on Machine Learning (ICML)},
shorttitle = {{UPD}},
}
MEAL: A Benchmark for Continual Multi-Agent Reinforcement Learning
Tristan Tomilin, Luka van den Boogaard, Samuel Garcin, Constantin Ruhdorfer, Bram Grooten, Fabrice Kusters, Yali Du, Andreas Bulling, Mykola Pechenizkiy, Meng Fang
Proc. International Conference on Machine Learning (ICML), 2026.
AbstractLinksBibTeXProject
Benchmarks play a central role in reinforcement learning (RL) research, yet their computational constraints often shape what is studied. Despite the motivation of lifelong learning, most continual RL papers consider only 3-10 sequential tasks, as CPU-bound environments make longer sequences impractical. Meanwhile, continual learning in cooperative multi-agent settings remains largely unexplored. To address these gaps, we introduce MEAL (Multi-agent Environments for Adaptive Learning), the first benchmark for continual multi-agent RL. By leveraging JAX and GPU acceleration, MEAL enables training on sequences of 100 tasks in a few hours on a single GPU. We find that long task sequences reveal failure modes that do not appear at smaller scales.
@inproceedings{tomilin26_icml,
title = {MEAL: A Benchmark for Continual Multi-Agent Reinforcement Learning},
author = {Tomilin, Tristan and van den Boogaard, Luka and Garcin, Samuel and Ruhdorfer, Constantin and Grooten, Bram and Kusters, Fabrice and Du, Yali and Bulling, Andreas and Pechenizkiy, Mykola and Fang, Meng},
year = {2026},
booktitle = {Proc. International Conference on Machine Learning (ICML)},
}

High Entropy Regularization Leads to Symmetry Equivariant Policies in Dec-POMDPs
Johannes Forkel, Constantin Ruhdorfer, Michael Beukman, Andreas Bulling, Jakob Nicolaus Foerster
Proc. Advances in Neural Information Processing Systems (NeurIPS), 2026.
AbstractLinksBibTeXProject
We prove that in any Dec-POMDP, sufficiently high entropy regularization ensures that the policy gradient flow with tabular softmax parametrization always converges, for any initialization, to the same joint policy, and that this joint policy is equivariant w.r.t. all symmetries of the Dec-POMDP. In particular, policies coming from different initializations will be fully compatible, in that their cross-play returns are equal to their self-play returns. Through extensive evaluation of independent PPO, arguably the standard baseline deep multi-agent policy gradient algorithm, in the Hanabi, Overcooked and Yokai environments, we find that the entropy coefficient has a massive influence on the cross-play returns between independently trained policies, and that the decrease in self-play returns coming from increased entropy regularization can often be counteracted by greedifying the learned policies after training. In Hanabi in particular we achieve a new SOTA in inter-seed cross-play this way. While we give examples of Dec-POMDPs in which one cannot learn the optimal symmetry equivariant policy this way, both our theoretical and empirical results suggest that one should consider far higher entropy coefficients during hyperparameter sweeps in Dec-POMDPs than is typically done.
@inproceedings{forkel26_neurips,
title = {High Entropy Regularization Leads to Symmetry Equivariant Policies in {{Dec-POMDPs}}},
author = {Forkel, Johannes and Ruhdorfer, Constantin and Beukman, Michael and Bulling, Andreas and Foerster, Jakob Nicolaus},
year = {2026},
booktitle = {Proc. Advances in Neural Information Processing Systems (NeurIPS)},
doi = {10.48550/arXiv.2511.22581},
}
2025

The Overcooked Generalisation Challenge: Evaluating Cooperation with Novel Partners in Unknown Environments Using Unsupervised Environment Design
Constantin Ruhdorfer, Matteo Bortoletto, Anna Penzkofer, Andreas Bulling
Transactions on Machine Learning Research (TMLR), pp. 1-25, 2025.
AbstractLinksBibTeXProject
We introduce the Overcooked Generalisation Challenge (OGC) - a new benchmark for evaluating reinforcement learning (RL) agents on their ability to cooperate with unknown partners in unfamiliar environments. Existing work typically evaluated cooperative RL only in their training environment or with their training partners, thus seriously limiting our ability to understand agents' generalisation capacity - an essential requirement for future collaboration with humans. The OGC extends Overcooked-AI to support dual curriculum design (DCD). It is fully GPU-accelerated, open-source, and integrated into the minimax DCD benchmark suite. Compared to prior DCD benchmarks, where designers manipulate only minimal elements of the environment, OGC introduces a significantly richer design space: full kitchen layouts with multiple objects that require the designer to account for interaction dynamics between agents. We evaluate state-of-the-art DCD algorithms alongside scalable neural architectures and find that current methods fail to produce agents that generalise effectively to novel layouts and unfamiliar partners. Our results indicate that both agents and curriculum designers struggle with the joint challenge of partner and environment generalisation. These findings establish OGC as a demanding testbed for cooperative generalisation and highlight key directions for future research.
@article{ruhdorfer25_tmlr,
title = {The Overcooked Generalisation Challenge: Evaluating Cooperation with Novel Partners in Unknown Environments Using Unsupervised Environment Design},
author = {Constantin Ruhdorfer and Matteo Bortoletto and Anna Penzkofer and Andreas Bulling},
year = {2025},
journal = {Transactions on Machine Learning Research (TMLR)},
pages = {1-25},
url = {https://openreview.net/forum?id=K2KtcMlW6j},
}

ToM-SSI: Evaluating Theory of Mind in Situated Social Interactions
Matteo Bortoletto, Constantin Ruhdorfer, Andreas Bulling
Proc. Empirical Methods in Natural Language Processing (EMNLP), 2025.
AbstractLinksBibTeXProject
Most existing Theory of Mind (ToM) benchmarks for foundation models rely on variations of the Sally-Anne test, offering only a very limited perspective on ToM and neglecting the complexity of human social interactions. To address this gap, we propose ToM-SSI: a new benchmark specifically designed to test ToM capabilities in environments rich with social interactions and spatial dynamics. While current ToM benchmarks are limited to text-only or dyadic interactions, ToM-SSI is multimodal and includes group interactions of up to four agents that communicate and move in situated environments. This unique design allows us to study, for the first time, mixed cooperative-obstructive settings and reasoning about multiple agents' mental state in parallel, thus capturing a wider range of social cognition than existing benchmarks. Our evaluations reveal that current models' performance is still severely limited, especially in these new tasks, highlighting critical gaps for future research.
@inproceedings{bortoletto25_emnlp,
title = {ToM-SSI: Evaluating Theory of Mind in Situated Social Interactions},
author = {Bortoletto, Matteo and Ruhdorfer, Constantin and Bulling, Andreas},
year = {2025},
booktitle = {Proc. Empirical Methods in Natural Language Processing (EMNLP)},
doi = {10.18653/v1/2025.emnlp-main.1642},
}

Brittle Minds, Fixable Activations: Understanding Belief Representations in Language Models
Matteo Bortoletto, Constantin Ruhdorfer, Lei Shi, Andreas Bulling
Findings of Empirical Methods in Natural Language Processing (EMNLP), 2025.
AbstractLinksBibTeXProject
Despite growing interest in Theory of Mind (ToM) tasks for evaluating language models (LMs), little is known about how LMs internally represent mental states of self and others. Understanding these internal mechanisms is critical - not only to move beyond surface-level performance, but also for model alignment and safety, where subtle misattributions of mental states may go undetected in generated outputs. In this work, we present the first systematic investigation of belief representations in LMs by probing models across different scales, training regimens, and prompts - using control tasks to rule out confounds. Our experiments provide evidence that both model size and fine-tuning substantially improve LMs' internal representations of others' beliefs, which are structured - not mere by-products of spurious correlations - yet brittle to prompt variations. Crucially, we show that these representations can be strengthened: targeted edits to model activations can correct wrong ToM inferences.
@inproceedings{bortoletto25_femnlp,
title = {Brittle {{Minds}}, {{Fixable Activations}}: {{Understanding Belief Representations}} in {{Language Models}}},
author = {Bortoletto, Matteo and Ruhdorfer, Constantin and Shi, Lei and Bulling, Andreas},
year = {2025},
booktitle = {Findings of Empirical Methods in Natural Language Processing (EMNLP)},
doi = {10.18653/v1/2025.findings-emnlp.1226},
shorttitle = {Brittle {{Minds}}, {{Fixable Activations}}},
}

Integration of Machine Learning in High-Enthalpy Plasma Spectroscopy
Paul Erik Hofmeyer, Hendrik Burghaus, Constantin Ruhdorfer, Johannes Oswald, Andreas Bulling, Georg Herdrich
International Conference on Flight vehicles, Aerothermodynamics and Re-entry (FAR), pp. 1--8, 2025.
AbstractLinksBibTeXProject
Optical emission spectroscopy is widely used to characterize high-enthalpy plasmas because it enables measuring a range of plasma parameters. Although the spectra are relatively easy to acquire, extracting meaningful information requires extensive analysis. In this work, a novel approach is developed to automate the analysis of broadband emission spectra by training two machine learning models on synthetic data. The first model is applied to predict plasma temperatures and species number densities in a CO2 plasma jet. The second model is designed to identify the radiation from potentially occurring species in time-resolved spectra of a titanium material sample demising in an air plasma. Developing a synthetic dataset that allows a trained machine learning model to analyze experimental spectra accurately is identified as a major challenge. Overall, these models offer a significant opportunity to automate the analysis of optical emission spectra.
@inproceedings{hofmeyer25_far,
title = {Integration of Machine Learning in High-Enthalpy Plasma Spectroscopy},
author = {Hofmeyer, Paul Erik and Burghaus, Hendrik and Ruhdorfer, Constantin and Oswald, Johannes and Bulling, Andreas and Herdrich, Georg},
year = {2025},
booktitle = {International Conference on Flight vehicles, Aerothermodynamics and Re-entry (FAR)},
pages = {1--8},
url = {https://www.researchgate.net/publication/392123052_Integration_of_Machine_Learning_in_High-Enthalpy_Plasma_Spectroscopy},
}
2024

Explicit Modelling of Theory of Mind for Belief Prediction in Nonverbal Social Interactions
Matteo Bortoletto, Constantin Ruhdorfer, Lei Shi, Andreas Bulling
Proc. 27th European Conference on Artificial Intelligence (ECAI), pp. 866--873, 2024.
AbstractLinksBibTeXProject
We propose MToMnet - a Theory of Mind (ToM) neural network for predicting beliefs and their dynamics during human social interactions from multimodal input. ToM is key for effective nonverbal human communication and collaboration, yet, existing methods for belief modelling have not included explicit ToM modelling or have typically been limited to one or two modalities. MToMnet encodes contextual cues (scene videos and object locations) and integrates them with person-specific cues (human gaze and body language) in a separate MindNet for each person. Inspired by prior research on social cognition and computational ToM, we propose three different MToMnet variants: two involving fusion of latent representations and one involving re-ranking of classification scores. We evaluate our approach on two challenging real-world datasets, one focusing on belief prediction, while the other examining belief dynamics prediction. Our results demonstrate that MToMnet surpasses existing methods by a large margin while at the same time requiring a significantly smaller number of parameters. Taken together, our method opens up a highly promising direction for future work on artificial intelligent systems that can robustly predict human beliefs from their non-verbal behaviour and, as such, more effectively collaborate with humans.
@inproceedings{bortoletto24_ecai,
title = {Explicit Modelling of Theory of Mind for Belief Prediction in Nonverbal Social Interactions},
author = {Matteo Bortoletto and Constantin Ruhdorfer and Lei Shi and Andreas Bulling},
year = {2024},
booktitle = {Proc. 27th European Conference on Artificial Intelligence (ECAI)},
pages = {866--873},
doi = {10.3233/FAIA240573},
}

Benchmarking Mental State Representations in Language Models
Matteo Bortoletto, Constantin Ruhdorfer, Lei Shi, Andreas Bulling
Proc. ICML 2024 Workshop on Mechanistic Interpretability, pp. 1--21, 2024.
AbstractLinksBibTeXProject
While numerous works have assessed the generative performance of language models (LMs) on tasks requiring Theory of Mind reasoning, research into the models' internal representation of mental states remains limited. Recent work has used probing to demonstrate that LMs can represent beliefs of themselves and others. However, these claims are accompanied by limited evaluation, making it difficult to assess how mental state representations are affected by model design and training choices. We report an extensive benchmark with various LM types with different model sizes, fine-tuning approaches, and prompt designs to study the robustness of mental state representations and memorisation issues within the probes. Our results show that the quality of models' internal representations of the beliefs of others increases with model size and, more crucially, with fine-tuning. We are the first to study how prompt variations impact probing performance on theory of mind tasks. We demonstrate that models' representations are sensitive to prompt variations, even when such variations should be beneficial. Finally, we complement previous activation editing experiments on Theory of Mind tasks and show that it is possible to improve models' reasoning performance by steering their activations without the need to train any probe.
@inproceedings{bortoletto24_icmlw,
title = {Benchmarking Mental State Representations in Language Models},
author = {Matteo Bortoletto and Constantin Ruhdorfer and Lei Shi and Andreas Bulling},
year = {2024},
booktitle = {Proc. ICML 2024 Workshop on Mechanistic Interpretability},
pages = {1--21},
url = {https://openreview.net/forum?id=yEwEVoH9Be},
}

Limits of Theory of Mind Modelling in Dialogue-Based Collaborative Plan Acquisition
Matteo Bortoletto, Constantin Ruhdorfer, Adnen Abdessaied, Lei Shi, Andreas Bulling
Proc. 62nd Annual Meeting of the Association for Computational Linguistics (ACL), pp. 1--16, 2024.
AbstractLinksBibTeXProject
Recent work on dialogue-based collaborative plan acquisition (CPA) has suggested that Theory of Mind (ToM) modelling can improve missing knowledge prediction in settings with asymmetric skill-sets and knowledge. Although ToM was claimed to be important for effective collaboration, its real impact on this novel task remains under-explored. By representing plans as graphs and by exploiting task-specific constraints we show that, as performance on CPA nearly doubles when predicting one's own missing knowledge, the improvements due to ToM modelling diminish. This phenomenon persists even when evaluating existing baseline methods. To better understand the relevance of ToM for CPA, we report a principled performance comparison of models with and without ToM features. Results across different models and ablations consistently suggest that learned ToM features are indeed more likely to reflect latent patterns in the data with no perceivable link to ToM. This finding calls for a deeper understanding of the role of ToM in CPA and beyond, as well as new methods for modelling and evaluating mental states in computational collaborative agents.
@inproceedings{bortoletto24_acl,
title = {Limits of Theory of Mind Modelling in Dialogue-Based Collaborative Plan Acquisition},
author = {Matteo Bortoletto and Constantin Ruhdorfer and Adnen Abdessaied and Lei Shi and Andreas Bulling},
year = {2024},
booktitle = {Proc. 62nd Annual Meeting of the Association for Computational Linguistics (ACL)},
pages = {1--16},
doi = {10.18653/v1/2024.acl-long.266},
}

VisRecall++: Analysing and Predicting Visualisation Recallability from Gaze Behaviour
Yao Wang, Yue Jiang, Zhiming Hu, Constantin Ruhdorfer, Mihai Bâce, Andreas Bulling
Proc. ACM on Human-Computer Interaction (PACM HCI), 8 (ETRA), pp. 1--18, 2024.
AbstractLinksBibTeXProject
Question answering has recently been proposed as a promising means to assess the recallability of information visualisations. However, prior works are yet to study the link between visually encoding a visualisation in memory and recall performance. To fill this gap, we propose VisRecall++ – a novel 40-participant recallability dataset that contains gaze data on 200 visualisations and five question types, such as identifying the title, and finding extreme values.We measured recallability by asking participants questions after they observed the visualisation for 10 seconds.Our analyses reveal several insights, such as saccade amplitude, number of fixations, and fixation duration significantly differ between high and low recallability groups.Finally, we propose GazeRecallNet – a novel computational method to predict recallability from gaze behaviour that outperforms several baselines on this task.Taken together, our results shed light on assessing recallability from gaze behaviour and inform future work on recallability-based visualisation optimisation.
@article{wang24_etra,
title = {VisRecall++: Analysing and Predicting Visualisation Recallability from Gaze Behaviour},
author = {Yao Wang and Yue Jiang and Zhiming Hu and Constantin Ruhdorfer and Mihai Bâce and Andreas Bulling},
year = {2024},
journal = {Proc. ACM on Human-Computer Interaction (PACM HCI)},
volume = {8},
number = {ETRA},
pages = {1--18},
doi = {10.1145/3655613},
}
2020

Efficient Implementation of Large-Scale Watchlists
Constantin Ruhdorfer, Stephan Schulz
PAAR+ SC^2@ IJCAR, pp. 120--133, 2020.
LinksBibTeXProject
@inproceedings{ruhdorfer20_ijcar,
title = {Efficient {Implementation} of {Large}-{Scale} {Watchlists}},
author = {Ruhdorfer, Constantin and Schulz, Stephan},
year = {2020},
booktitle = {{PAAR}+ {SC}{\textasciicircum}2@ {IJCAR}},
pages = {120--133},
url = {https://ceur-ws.org/Vol-2752/paper9.pdf},
}