Publications
2023

Facial Composite Generation with Iterative Human Feedback
Florian Strohm, Ekta Sood, Dominike Thomas, Mihai Bâce, Andreas Bulling
Proc. The 1st Gaze Meets ML workshop, PMLR, pp. 165--183, 2023.
AbstractLinksBibTeXProject
We propose the first method in which human and AI collaborate to iteratively reconstruct the human’s mental image of another person’s face only from their eye gaze. Current tools for generating digital human faces involve a tedious and time-consuming manual design process. While gaze-based mental image reconstruction represents a promising alternative, previous methods still assumed prior knowledge about the target face, thereby severely limiting their practical usefulness. The key novelty of our method is a collaborative, it- erative query engine: Based on the user’s gaze behaviour in each iteration, our method predicts which images to show to the user in the next iteration. Results from two human studies (N=12 and N=22) show that our method can visually reconstruct digital faces that are more similar to the mental image, and is more usable compared to other methods. As such, our findings point at the significant potential of human-AI collaboration for recon- structing mental images, potentially also beyond faces, and of human gaze as a rich source of information and a powerful mediator in said collaboration.
@inproceedings{strohm23_gmml,
title = {Facial Composite Generation with Iterative Human Feedback},
author = {Strohm, Florian and Sood, Ekta and Thomas, Dominike and B{\^a}ce, Mihai and Bulling, Andreas},
editor = {Lourentzou, Ismini and Wu, Joy and Kashyap, Satyananda and Karargyris, Alexandros and Celi, Leo Anthony and Kawas, Ban and Talathi, Sachin},
year = {2023},
booktitle = {Proc. The 1st Gaze Meets ML workshop, PMLR},
volume = {210},
pages = {165--183},
url = {https://proceedings.mlr.press/v210/strohm23a.html},
publisher = {PMLR},
series = {Proceedings of Machine Learning Research},
pdf = {https://proceedings.mlr.press/v210/strohm23a/strohm23a.pdf},
}

Multimodal Integration of Human-Like Attention in Visual Question Answering
Ekta Sood, Fabian Kögel, Philipp Müller, Dominike Thomas, Mihai Bâce, Andreas Bulling
Proc. Workshop on Gaze Estimation and Prediction in the Wild (GAZE), CVPRW, pp. 2647--2657, 2023.
AbstractLinksBibTeXProject Tobii Sponsor Award, Oral Presentation
Human-like attention as a supervisory signal to guide neural attention has shown significant promise but is currently limited to uni-modal integration – even for inherently multi-modal tasks such as visual question answering (VQA). We present the Multimodal Human-like Attention Network (MULAN) – the first method for multimodal integration of human-like attention on image and text during training of VQA models. MULAN integrates attention predictions from two state-of-the-art text and image saliency models into neural self-attention layers of a recent transformer-based VQA model. Through evaluations on the challenging VQAv2 dataset, we show that MULAN achieves a new state-of-the-art performance of 73.98% accuracy on test-std and 73.72% on test-dev and, at the same time, has approximately 80% fewer trainable parameters than prior work. Overall, our work underlines the potential of integrating multimodal human-like and neural attention for VQA.
@inproceedings{sood23_gaze,
title = {Multimodal Integration of Human-Like Attention in Visual Question Answering},
author = {Sood, Ekta and Fabian Kögel and Philipp Müller and Dominike Thomas and Mihai Bâce and Andreas Bulling},
year = {2023},
booktitle = {Proc. Workshop on Gaze Estimation and Prediction in the Wild (GAZE), CVPRW},
pages = {2647--2657},
url = {https://openaccess.thecvf.com/content/CVPR2023W/GAZE/papers/Sood_Multimodal_Integration_of_Human-Like_Attention_in_Visual_Question_Answering_CVPRW_2023_paper.pdf},
}

MultiMediate '23: Engagement Estimation and Bodily Behaviour Recognition in Social Interactions
Philipp M\"{u}ller, Michal Balazia, Tobias Baur, Michael Dietz, Alexander Heimerl, Dominik Schiller, Mohammed Guermal, Dominike Thomas, Fran\c{c}ois Br\'{e}mond, Jan Alexandersson, Elisabeth Andr\'{e}, Andreas Bulling
Proceedings of the 31st ACM International Conference on Multimedia, pp. 9640–9645, 2023.
AbstractLinksBibTeXProject
Automatic analysis of human behaviour is a fundamental prerequisite for the creation of machines that can effectively interact with- and support humans in social interactions. In MultiMediate'23, we address two key human social behaviour analysis tasks for the first time in a controlled challenge: engagement estimation and bodily behaviour recognition in social interactions. This paper describes the MultiMediate'23 challenge and presents novel sets of annotations for both tasks. For engagement estimation we collected novel annotations on the NOvice eXpert Interaction (NOXI) database. For bodily behaviour recognition, we annotated test recordings of the MPIIGroupInteraction corpus with the BBSI annotation scheme. In addition, we present baseline results for both challenge tasks.
@inproceedings{mueller23_mm,
title = {MultiMediate '23: Engagement Estimation and Bodily Behaviour Recognition in Social Interactions},
author = {M\"{u}ller, Philipp and Balazia, Michal and Baur, Tobias and Dietz, Michael and Heimerl, Alexander and Schiller, Dominik and Guermal, Mohammed and Thomas, Dominike and Br\'{e}mond, Fran\c{c}ois and Alexandersson, Jan and Andr\'{e}, Elisabeth and Bulling, Andreas},
year = {2023},
booktitle = {Proceedings of the 31st ACM International Conference on Multimedia},
pages = {9640–9645},
doi = {10.1145/3581783.3613851},
url = {https://doi.org/10.1145/3581783.3613851},
publisher = {Association for Computing Machinery},
address = {New York, NY, USA},
location = {Ottawa ON, Canada},
isbn = {9798400701085},
numpages = {6},
keywords = {dataset, engagement, nonverbal behaviour, challenge},
series = {MM '23},
}
2022

MultiMediate'22: Backchannel Detection and Agreement Estimation in Group Interactions
Philipp Müller, Dominik Schiller, Dominike Thomas, Michael Dietz, Hali Lindsay, Patrick Gebhard, Elisabeth André, Andreas Bulling
Proc. ACM Multimedia (MM), pp. 7109-7114, 2022.
AbstractLinksBibTeXProject
Backchannels, i.e. short interjections of the listener, serve important meta-conversational purposes like signifying attention or indicating agreement. Despite their key role, automatic analysis of backchannels in group interactions has been largely neglected so far. The MultiMediate challenge addresses, for the first time, the tasks of backchannel detection and agreement estimation from backchannels in group conversations. This paper describes the MultiMediate challenge and presents a novel set of annotations consisting of 7234 backchannel instances for the MPIIGroup Interaction dataset. Each backchannel was additionally annotated with the extent by which it expresses agreement towards the current speaker. In addition to a an analysis of the collected annotations, we present baseline results for both challenge tasks.
@inproceedings{mueller22_mm,
title = {MultiMediate'22: Backchannel Detection and Agreement Estimation in Group Interactions},
author = {M{\"{u}}ller, Philipp and Schiller, Dominik and Thomas, Dominike and Dietz, Michael and Lindsay, Hali and Gebhard, Patrick and André, Elisabeth and Bulling, Andreas},
year = {2022},
booktitle = {Proc. ACM Multimedia (MM)},
pages = {7109-7114},
doi = {10.1145/3503161.3551589},
}

MultiMediate'22: Backchannel Detection and Agreement Estimation in Group Interactions
Philipp Müller, Dominik Schiller, Dominike Thomas, Michael Dietz, Hali Lindsay, Patrick Gebhard, Elisabeth André, Andreas Bulling
arXiv:2209.09578, pp. 1--6, 2022.
AbstractLinksBibTeXProject
Backchannels, i.e. short interjections of the listener, serve important meta-conversational purposes like signifying attention or indicating agreement. Despite their key role, automatic analysis of backchannels in group interactions has been largely neglected so far. The MultiMediate challenge addresses, for the first time, the tasks of backchannel detection and agreement estimation from backchannels in group conversations. This paper describes the MultiMediate challenge and presents a novel set of annotations consisting of 7234 backchannel instances for the MPIIGroupInteraction dataset. Each backchannel was additionally annotated with the extent by which it expresses agreement towards the current speaker. In addition to a an analysis of the collected annotations, we present baseline results for both challenge tasks.
@techreport{mueller22_arxiv,
title = {MultiMediate'22: Backchannel Detection and Agreement Estimation in Group Interactions},
author = {M{\"{u}}ller, Philipp and Schiller, Dominik and Thomas, Dominike and Dietz, Michael and Lindsay, Hali and Gebhard, Patrick and André, Elisabeth and Bulling, Andreas},
year = {2022},
pages = {1--6},
doi = {10.48550/arXiv.2209.09578},
url = {http://arxiv.org/abs/2209.09578},
}
2021

MultiMediate: Multi-modal Group Behaviour Analysis for Artificial Mediation
Philipp Müller, Dominik Schiller, Dominike Thomas, Guanhua Zhang, Michael Dietz, Patrick Gebhard, Elisabeth André, Andreas Bulling
Proc. ACM Multimedia (MM), pp. 4878--4882, 2021.
AbstractLinksBibTeXProject
Artificial mediators are promising to support human group conversations but at present their abilities are limited by insufficient progress in group behaviour analysis. The MultiMediate challenge addresses, for the first time, two fundamental group behaviour analysis tasks in well-defined conditions: eye contact detection and next speaker prediction. For training and evaluation, MultiMediate makes use of the MPIIGroupInteraction dataset consisting of 22 three- to four-person discussions as well as of an unpublished test set of six additional discussions. This paper describes the MultiMediate challenge and presents the challenge dataset including novel fine-grained speaking annotations that were collected for the purpose of MultiMediate. Furthermore, we present baseline approaches and ablation studies for both challenge tasks.
@inproceedings{mueller21_mm,
title = {MultiMediate: Multi-modal Group Behaviour Analysis for Artificial Mediation},
author = {M{\"{u}}ller, Philipp and Schiller, Dominik and Thomas, Dominike and Zhang, Guanhua and Dietz, Michael and Gebhard, Patrick and André, Elisabeth and Bulling, Andreas},
year = {2021},
booktitle = {Proc. ACM Multimedia (MM)},
pages = {4878--4882},
doi = {10.1145/3474085.3479219},
}