Publications
2026
CLAD: Constrained Latent Action Diffusion for Vision-Language Procedure Planning
Lei Shi, Andreas Bulling
Proc. IEEE International Conference on Robot and Human Interactive Communication (RO-MAN), 2026.
AbstractLinksBibTeXProject
We propose CLAD, a Constrained Latent Action Diffusion model for vision-language procedure planning, the challenging task of predicting a sequence of actions that lead from a start state towards an intended goal state. Procedure planning, while critical in robot skill learning and for assistive robots, has been largely neglected so far, and existing methods have not leveraged semantic information for action generation. In contrast, CLAD exploits the fact that the latent space of diffusion models trained for procedure planning contains rich semantic information. Our method uses a Variational Autoencoder (VAE) to learn the latent representation of actions and observations as constraints and integrate them into a diffusion process. As such, our method uses these latent constraints to steer the diffusion model to generate better actions in the procedural plan. We report extensive experiments on four datasets: three covering human procedure planning and one robot learning, and show that our method outperforms state-of-the-art methods by a large margin. We demonstrate that the proposed integration of the action and observation representations learnt in the VAE latent space is key to these performance improvements.
@inproceedings{shi26_roman,
title = {{{CLAD}}: {{Constrained Latent Action Diffusion}} for {{Vision-Language Procedure Planning}}},
author = {Shi, Lei and Bulling, Andreas},
year = {2026},
booktitle = {Proc. IEEE International Conference on Robot and Human Interactive Communication (RO-MAN)},
}
2025

ActionDiffusion: An Action-aware Diffusion Model for Procedure Planning in Instructional Videos
Lei Shi, Paul-Christian Bürkner, Andreas Bulling
Proc. IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2025.
AbstractLinksBibTeXProject
We present ActionDiffusion - a novel diffusion model for procedure planning in instructional videos that is the first to take temporal inter-dependencies between actions into account. Our approach is in stark contrast to existing methods that fail to exploit the rich information content available in the particular order in which actions are performed. Our method unifies the learning of temporal dependencies between actions and denoising of the action plan in the diffusion process by projecting the action information into the noise space. This is achieved 1) by adding action embeddings in the noise masks in the noise-adding phase and 2) by introducing an attention mechanism in the noise prediction network to learn the correlations between different action steps. We report extensive experiments on three instructional video benchmark datasets (CrossTask, Coin, and NIV) and show that our method outperforms previous state-of-the-art methods on all metrics on CrossTask and NIV and all metrics except accuracy on Coin dataset. We show that by adding action embeddings into the noise mask the diffusion model can better learn action temporal dependencies and increase the performances on procedure planning.
@inproceedings{shi25_wacv,
title = {ActionDiffusion: An Action-aware Diffusion Model for Procedure Planning in Instructional Videos},
author = {Lei Shi and Paul-Christian Bürkner and Andreas Bulling},
year = {2025},
booktitle = {Proc. IEEE/CVF Winter Conference on Applications of Computer Vision (WACV)},
}
2024

Explicit Modelling of Theory of Mind for Belief Prediction in Nonverbal Social Interactions
Matteo Bortoletto, Constantin Ruhdorfer, Lei Shi, Andreas Bulling
Proc. 27th European Conference on Artificial Intelligence (ECAI), pp. 866--873, 2024.
AbstractLinksBibTeXProject
We propose MToMnet - a Theory of Mind (ToM) neural network for predicting beliefs and their dynamics during human social interactions from multimodal input. ToM is key for effective nonverbal human communication and collaboration, yet, existing methods for belief modelling have not included explicit ToM modelling or have typically been limited to one or two modalities. MToMnet encodes contextual cues (scene videos and object locations) and integrates them with person-specific cues (human gaze and body language) in a separate MindNet for each person. Inspired by prior research on social cognition and computational ToM, we propose three different MToMnet variants: two involving fusion of latent representations and one involving re-ranking of classification scores. We evaluate our approach on two challenging real-world datasets, one focusing on belief prediction, while the other examining belief dynamics prediction. Our results demonstrate that MToMnet surpasses existing methods by a large margin while at the same time requiring a significantly smaller number of parameters. Taken together, our method opens up a highly promising direction for future work on artificial intelligent systems that can robustly predict human beliefs from their non-verbal behaviour and, as such, more effectively collaborate with humans.
@inproceedings{bortoletto24_ecai,
title = {Explicit Modelling of Theory of Mind for Belief Prediction in Nonverbal Social Interactions},
author = {Matteo Bortoletto and Constantin Ruhdorfer and Lei Shi and Andreas Bulling},
year = {2024},
booktitle = {Proc. 27th European Conference on Artificial Intelligence (ECAI)},
pages = {866--873},
doi = {10.3233/FAIA240573},
}

Multi-Modal Video Dialog State Tracking in the Wild
Adnen Abdessaied, Lei Shi, Andreas Bulling
Proc. 18th European Conference on Computer Vision (ECCV), pp. 1--25, 2024.
AbstractLinksBibTeXProject
We present MST-MIXER – a novel video dialog model operating over a generic multi-modal state tracking scheme. Current models that claim to perform multi-modal state tracking fall short of two major aspects: (1) They either track only one modality (mostly the visual input) or (2) they target synthetic datasets that do not reflect the complexity of real-world in the wild scenarios. Our model addresses these two limitations in an attempt to close this crucial research gap. Specifically, MST-MIXER first tracks the most important constituents of each input modality. Then, it predicts the missing underlying structure of the selected constituents of each modality by learning local latent graphs using a novel multi-modal graph structure learning method. Subsequently, the learned local graphs and features are parsed together to form a global graph operating on the mix of all modalities which further refines its structure and node embeddings. Finally, the fine-grained graph node features are used to enhance the hidden states of the backbone Vision-Language Model (VLM). MST-MIXER achieves new state-of-the-art results on five challenging benchmarks.
@inproceedings{abdessaied24_eccv,
title = {Multi-Modal Video Dialog State Tracking in the Wild},
author = {Adnen Abdessaied and Lei Shi and Andreas Bulling},
year = {2024},
booktitle = {Proc. 18th European Conference on Computer Vision (ECCV)},
pages = {1--25},
doi = {10.1007/978-3-031-72998-0_20},
}

Benchmarking Mental State Representations in Language Models
Matteo Bortoletto, Constantin Ruhdorfer, Lei Shi, Andreas Bulling
Proc. ICML 2024 Workshop on Mechanistic Interpretability, pp. 1--21, 2024.
AbstractLinksBibTeXProject
While numerous works have assessed the generative performance of language models (LMs) on tasks requiring Theory of Mind reasoning, research into the models' internal representation of mental states remains limited. Recent work has used probing to demonstrate that LMs can represent beliefs of themselves and others. However, these claims are accompanied by limited evaluation, making it difficult to assess how mental state representations are affected by model design and training choices. We report an extensive benchmark with various LM types with different model sizes, fine-tuning approaches, and prompt designs to study the robustness of mental state representations and memorisation issues within the probes. Our results show that the quality of models' internal representations of the beliefs of others increases with model size and, more crucially, with fine-tuning. We are the first to study how prompt variations impact probing performance on theory of mind tasks. We demonstrate that models' representations are sensitive to prompt variations, even when such variations should be beneficial. Finally, we complement previous activation editing experiments on Theory of Mind tasks and show that it is possible to improve models' reasoning performance by steering their activations without the need to train any probe.
@inproceedings{bortoletto24_icmlw,
title = {Benchmarking Mental State Representations in Language Models},
author = {Matteo Bortoletto and Constantin Ruhdorfer and Lei Shi and Andreas Bulling},
year = {2024},
booktitle = {Proc. ICML 2024 Workshop on Mechanistic Interpretability},
pages = {1--21},
url = {https://openreview.net/forum?id=yEwEVoH9Be},
}

Limits of Theory of Mind Modelling in Dialogue-Based Collaborative Plan Acquisition
Matteo Bortoletto, Constantin Ruhdorfer, Adnen Abdessaied, Lei Shi, Andreas Bulling
Proc. 62nd Annual Meeting of the Association for Computational Linguistics (ACL), pp. 1--16, 2024.
AbstractLinksBibTeXProject
Recent work on dialogue-based collaborative plan acquisition (CPA) has suggested that Theory of Mind (ToM) modelling can improve missing knowledge prediction in settings with asymmetric skill-sets and knowledge. Although ToM was claimed to be important for effective collaboration, its real impact on this novel task remains under-explored. By representing plans as graphs and by exploiting task-specific constraints we show that, as performance on CPA nearly doubles when predicting one's own missing knowledge, the improvements due to ToM modelling diminish. This phenomenon persists even when evaluating existing baseline methods. To better understand the relevance of ToM for CPA, we report a principled performance comparison of models with and without ToM features. Results across different models and ablations consistently suggest that learned ToM features are indeed more likely to reflect latent patterns in the data with no perceivable link to ToM. This finding calls for a deeper understanding of the role of ToM in CPA and beyond, as well as new methods for modelling and evaluating mental states in computational collaborative agents.
@inproceedings{bortoletto24_acl,
title = {Limits of Theory of Mind Modelling in Dialogue-Based Collaborative Plan Acquisition},
author = {Matteo Bortoletto and Constantin Ruhdorfer and Adnen Abdessaied and Lei Shi and Andreas Bulling},
year = {2024},
booktitle = {Proc. 62nd Annual Meeting of the Association for Computational Linguistics (ACL)},
pages = {1--16},
doi = {10.18653/v1/2024.acl-long.266},
}

VSA4VQA: Scaling A Vector Symbolic Architecture To Visual Question Answering on Natural Images
Anna Penzkofer, Lei Shi, Andreas Bulling
Proc. Annual Meeting of the Cognitive Science Society (CogSci), 2024.
AbstractLinksBibTeXProject Oral Presentation
While Vector Symbolic Architectures (VSAs) are promising for modelling spatial cognition, their application is currently limited to artificially generated images and simple spatial queries. We propose VSA4VQA – a novel 4D implementation of VSAs that implements a mental representation of natural images for the challenging task of Visual Question Answering (VQA). VSA4VQA is the first model to scale a VSA to complex spatial queries. Our method is based on the Semantic Pointer Architecture (SPA) to encode objects in a hyper-dimensional vector space. To encode natural images, we extend the SPA to include dimensions for object’s width and height in addition to their spatial location. To perform spatial queries we further introduce learned spatial query masks and integrate a pre-trained vision-language model for answering attribute-related questions. We evaluate our method on the GQA benchmark dataset and show that it can effectively encode natural images, achieving competitive performance to state-of-the-art deep learning methods for zero-shot VQA.
@inproceedings{penzkofer24_cogsci,
title = {{VSA4VQA}: {Scaling} {A} {Vector} {Symbolic} {Architecture} {To} {Visual} {Question} {Answering} on {Natural} {Images}},
author = {Anna Penzkofer and Lei Shi and Andreas Bulling},
year = {2024},
booktitle = {Proc. Annual Meeting of the Cognitive Science Society (CogSci)},
volume = {46},
url = {https://escholarship.org/uc/item/26j7v1nf.},
}

Explaining Disagreement in Visual Question Answering Using Eye Tracking
Susanne Hindennach, Lei Shi, Andreas Bulling
Proc. International Workshop on Pervasive Eye Tracking and Mobile Gaze-Based Interaction (PETMEI), pp. 1--7, 2024.
AbstractLinksBibTeXProject
When presented with the same question about an image, human annotators often give valid but disagreeing answers indicating that their reasoning was different. Such differences are lost in a single ground truth label used to train and evaluate visual question answering (VQA) methods. In this work, we explore whether visual attention maps, created using stationary eye tracking, provide insight into the reasoning underlying disagreement in VQA. We first manually inspect attention maps in the recent VQA-MHUG dataset and find cases in which attention differs consistently for disagreeing answers. We further evaluate the suitability of four different similarity metrics to detect attention differences matching the disagreement. We show that attention maps plausibly surface differences in reasoning underlying one type of disagreement, and that the metrics complementarily detect them. Taken together, our results represent an important first step to leverage eye-tracking to explain disagreement in VQA.
@inproceedings{hindennach24_petmei,
title = {Explaining Disagreement in Visual Question Answering Using Eye Tracking},
author = {Susanne Hindennach and Lei Shi and Andreas Bulling},
year = {2024},
booktitle = {Proc. International Workshop on Pervasive Eye Tracking and Mobile Gaze-Based Interaction (PETMEI)},
pages = {1--7},
doi = {10.1145/3649902.3656356},
}

Inferring Human Intentions from Predicted Action Probabilities
Lei Shi, Paul-Christian Bürkner, Andreas Bulling
Proc. Workshop on Theory of Mind in Human-AI Interaction at CHI 2024, pp. 1--7, 2024.
AbstractLinksBibTeXProject
Inferring human intentions is a core challenge in human-AI collab-oration but while Bayesian methods struggle with complex visual input, deep neural network (DNN) based methods do not provide uncertainty quantifications. In this work we combine both approaches for the first time and show that the predicted next action probabilities contain information that can be used to infer the underlying user intention. We propose a two-step approach to human intention prediction: While a DNN predicts the probabilities of the next action, MCMC-based Bayesian inference is used to infer the underlying intention from these predictions. This approach not only allows for the independent design of the DNN architecture but also the subsequently fast, design-independent inference of human intentions. We evaluate our method using a series of experiments on the Watch-And-Help (WAH) and a keyboard and mouse interaction dataset. Our results show that our approach can accurately predict human intentions from observed actions and the implicit information contained in next action probabilities. Furthermore, we show that our approach can predict the correct intention even if only a few actions have been observed.
@inproceedings{shi24_chiw,
title = {Inferring Human Intentions from Predicted Action Probabilities},
author = {Lei Shi and Paul-Christian Bürkner and Andreas Bulling},
year = {2024},
booktitle = {Proc. Workshop on Theory of Mind in Human-AI Interaction at CHI 2024},
pages = {1--7},
}

ActionDiffusion: An Action-aware Diffusion Model for Procedure Planning in Instructional Videos
Lei Shi, Paul Burkner, Andreas Bulling
arXiv:2403.08591, pp. 1--6, 2024.
AbstractLinksBibTeXProject
We present ActionDiffusion – a novel diffusion model for procedure planning in instructional videos that is the first to take temporal inter-dependencies between actions into account in a diffusion model for procedure planning. This approach is in stark contrast to existing methods that fail to exploit the rich information content available in the particular order in which actions are performed. Our method unifies the learning of temporal dependencies between actions and denoising of the action plan in the diffusion process by projecting the action information into the noise space. This is achieved 1) by adding action embeddings in the noise masks in the noiseadding phase and 2) by introducing an attention mechanism in the noise prediction network to learn the correlations between different action steps. We report extensive experiments on three instructional video benchmark datasets (CrossTask, Coin, and NIV) and show that our method outperforms previous state-of-the-art methods on all metrics on CrossTask and NIV and all metrics except accuracy on Coin dataset. We show that by adding action embeddings into the noise mask the diffusion model can better learn action temporal dependencies and increase the performances on procedure planning
@techreport{shi24_arxiv,
title = {ActionDiffusion: An Action-aware Diffusion Model for Procedure Planning in Instructional Videos},
author = {Lei Shi and Paul Burkner and Andreas Bulling},
year = {2024},
pages = {1--6},
url = {https://arxiv.org/abs/2403.08591},
}

Neural Reasoning About Agents’ Goals, Preferences, and Actions
Matteo Bortoletto, Lei Shi, Andreas Bulling
Proc. 38th AAAI Conference on Artificial Intelligence (AAAI), pp. 456--464, 2024.
AbstractLinksBibTeXProject
We propose the Intuitive Reasoning Network (IRENE) – a novel neural model for intuitive psychological reasoning about agents’ goals, preferences, and actions that can generalise previous experiences to new situations. IRENE combines a graph neural network for learning agent and world state representations with a transformer to encode the task context. When evaluated on the challenging Baby Intuitions Benchmark, IRENE achieves new state-of-the-art performance on three out of its five tasks – with up to 48.9 % improvement. In contrast to existing methods, IRENE is able to bind preferences to specific agents, to better distinguish between rational and irrational agents, and to better understand the role of blocking obstacles. We also investigate, for the first time, the influence of the training tasks on test performance. Our analyses demonstrate the effectiveness of IRENE in combining prior knowledge gained during training for unseen evaluation tasks.
@inproceedings{bortoletto24_aaai,
title = {Neural Reasoning About Agents’ Goals, Preferences, and Actions},
author = {Matteo Bortoletto and Lei Shi and Andreas Bulling},
year = {2024},
booktitle = {Proc. 38th AAAI Conference on Artificial Intelligence (AAAI)},
volume = {38},
number = {1},
pages = {456--464},
doi = {10.1609/aaai.v38i1.27800},
}

Mindful Explanations: Prevalence and Impact of Mind Attribution in XAI Research
Susanne Hindennach, Lei Shi, Filip Miletic, Andreas Bulling
Proc. ACM on Human-Computer Interaction (PACM HCI), 8 (CSCW), pp. 1--42, 2024.
AbstractLinksBibTeXProject Best Paper Honourable Mention Award
When users perceive AI systems as mindful, independent agents, they hold them responsible instead of the AI experts who created and designed these systems. So far, it has not been studied whether explanations support this shift in responsibility through the use of mind-attributing verbs like "to think". To better understand the prevalence of mind-attributing explanations we analyse AI explanations in 3,533 explainable AI (XAI) research articles from the Semantic Scholar Open Research Corpus (S2ORC). Using methods from semantic shift detection, we identify three dominant types of mind attribution: (1) metaphorical (e.g. "to learn" or "to predict"), (2) awareness (e.g. "to consider"), and (3) agency (e.g. "to make decisions"). We then analyse the impact of mind-attributing explanations on awareness and responsibility in a vignette-based experiment with 199 participants. We find that participants who were given a mind-attributing explanation were more likely to rate the AI system as aware of the harm it caused. Moreover, the mind-attributing explanation had a responsibility-concealing effect: Considering the AI experts’ involvement lead to reduced ratings of AI responsibility for participants who were given a non-mind-attributing or no explanation. In contrast, participants who read the mind-attributing explanation still held the AI system responsible despite considering the AI experts’ involvement. Taken together, our work underlines the need to carefully phrase explanations about AI systems in scientific writing to reduce mind attribution and clearly communicate human responsibility.
@article{hindennach24_pacm,
title = {Mindful Explanations: Prevalence and Impact of Mind Attribution in XAI Research},
author = {Susanne Hindennach and Lei Shi and Filip Miletic and Andreas Bulling},
year = {2024},
journal = {Proc. ACM on Human-Computer Interaction (PACM HCI)},
volume = {8},
number = {CSCW},
pages = {1--42},
doi = {10.1145/3641009},
url = {https://medium.com/acm-cscw/be-mindful-when-using-mindful-descriptions-in-explanations-about-ai-bfc7666885c6},
}

VD-GR: Boosting Visual Dialog with Cascaded Spatial-Temporal Multi-Modal GRaphs
Adnen Abdessaied, Lei Shi, Andreas Bulling
Proc. IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), pp. 5805--5814, 2024.
AbstractLinksBibTeXProject
We propose VD-GR – a novel visual dialog model that combines pre-trained language models (LMs) with graph neural networks (GNNs). Prior works mainly focused on one class of models at the expense of the other, thus missing out on the opportunity of combining their respective benefits. At the core of VD-GR is a novel integration mechanism that alternates between spatial-temporal multi-modal GNNs and BERT layers, and that covers three distinct contributions: First, we use multi-modal GNNs to process the features of each modality (image, question, and dialog history) and exploit their local structures before performing
BERT global attention. Second, we propose hub-nodes that link to all other nodes within one modality graph, allowing the model to propagate information from one GNN (modality) to the other in a cascaded manner. Third, we augment the BERT hidden states with fine-grained multi-modal GNN features before passing them to the next VD-GR layer. Evaluations on VisDial v1.0, VisDial v0.9, VisDialConv, and VisPro show that VD-GR achieves new state-of-the-art results across all four datasets
@inproceedings{abdessaied24_wacv,
title = {VD-GR: Boosting Visual Dialog with Cascaded Spatial-Temporal Multi-Modal GRaphs},
author = {Adnen Abdessaied and Lei Shi and Andreas Bulling},
year = {2024},
booktitle = {Proc. IEEE/CVF Winter Conference on Applications of Computer Vision (WACV)},
pages = {5805--5814},
doi = {10.1109/WACV57701.2024.00570},
}
2023

Improving Neural Saliency Prediction with a Cognitive Model of Human Visual Attention
Ekta Sood, Lei Shi, Matteo Bortoletto, Yao Wang, Philipp Müller, Andreas Bulling
Proc. Annual Meeting of the Cognitive Science Society (CogSci), pp. 3639--3646, 2023.
AbstractLinksBibTeXProject
We present a novel method for saliency prediction that leverages a cognitive model of visual attention as an inductive bias. This approach is in stark contrast to recent purely data-driven saliency models that achieve performance improvements mainly by increased capacity, resulting in high computational costs and the need for large-scale training datasets. We demonstrate that by using a cognitive model, our method achieves competitive performance to the state of the art across several natural image datasets while only requiring a fraction of the parameters. Furthermore, we set the new state of the art for saliency prediction on information visualizations, demonstrating the effectiveness of our approach for cross-domain generalization. We further provide augmented versions of the full MSCOCO dataset with synthetic gaze data using the cognitive model, which we used to pre-train our method. Our results are highly promising and underline the significant potential of bridging between cognitive and data-driven models, potentially also beyond attention.
@inproceedings{sood23_cogsci,
title = {Improving Neural Saliency Prediction with a Cognitive Model of Human Visual Attention},
author = {Ekta Sood and Lei Shi and Matteo Bortoletto and Yao Wang and Philipp Müller and Andreas Bulling},
year = {2023},
booktitle = {Proc. Annual Meeting of the Cognitive Science Society (CogSci)},
pages = {3639--3646},
}

Exploring Natural Language Processing Methods for Interactive Behaviour Modelling
Guanhua Zhang, Matteo Bortoletto, Zhiming Hu, Lei Shi, Mihai Bâce, Andreas Bulling
Proc. IFIP TC13 Conference on Human-Computer Interaction (INTERACT), pp. 1--22, 2023.
AbstractLinksBibTeXProject Best Student Paper Nomination
Analysing and modelling interactive behaviour is an important topic in human-computer interaction (HCI) and a key requirement for the development of intelligent interactive systems. Interactive behaviour has a sequential (actions happen one after another) and hierarchical (a sequence of actions forms an activity driven by interaction goals) structure, which may be similar to the structure of natural language. Designed based on such a structure, natural language processing (NLP) methods have achieved groundbreaking success in various downstream tasks. However, few works linked interactive behaviour with natural language. In this paper, we explore the similarity between interactive behaviour and natural language by applying an NLP method, byte pair encoding (BPE), to encode mouse and keyboard behaviour. We then analyse the vocabulary, i.e., the set of action sequences, learnt by BPE, as well as use the vocabulary to encode the input behaviour for interactive task recognition. An existing dataset collected in constrained lab settings and our novel out-of-the-lab dataset were used for evaluation. Results show that this natural language-inspired approach not only learns action sequences that reflect specific interaction goals, but also achieves higher F1 scores on task recognition than other methods. Our work reveals the similarity between interactive behaviour and natural language, and presents the potential of applying the new pack of methods that leverage insights from NLP to model interactive behaviour in HCI.
@inproceedings{zhang23_interact,
title = {Exploring Natural Language Processing Methods for Interactive Behaviour Modelling},
author = {Zhang, Guanhua and Bortoletto, Matteo and Hu, Zhiming and Shi, Lei and B{\^a}ce, Mihai and Bulling, Andreas},
year = {2023},
booktitle = {Proc. IFIP TC13 Conference on Human-Computer Interaction (INTERACT)},
pages = {1--22},
publisher = {Springer},
}
2022
Evaluating Dropout Placements in Bayesian Regression Resnet
Lei Shi, Cosmin Copot, Steve Vanlanduit
Journal of Artificial Intelligence and Soft Computing Research, 12 (1), pp. 61--73, 2022.
AbstractLinksBibTeXProject
Deep Neural Networks (DNNs) have shown great success in many fields. Various network architectures have been developed for different applications. Regardless of the complexities of the networks, DNNs do not provide model uncertainty. Bayesian Neural Networks (BNNs), on the other hand, is able to make probabilistic inference. Among various types of BNNs, Dropout as a Bayesian Approximation converts a Neural Network (NN) to a BNN by adding a dropout layer after each weight layer in the NN. This technique provides a simple transformation from a NN to a BNN. However, for DNNs, adding a dropout layer to each weight layer would lead to a strong regularization due to the deep architecture. Previous researches [1, 2, 3] have shown that adding a dropout layer after each weight layer in a DNN is unnecessary. However, how to place dropout layers in a ResNet for regression tasks are less explored. In this work, we perform an empirical study on how different dropout placements would affect the performance of a Bayesian DNN. We use a regression model modified from ResNet as the DNN and place the dropout layers at different places in the regression ResNet. Our experimental results show that it is not necessary to add a dropout layer after every weight layer in the Regression ResNet to let it be able to make Bayesian Inference. Placing Dropout layers between the stacked blocks i.e. Dense+Identity+Identity blocks has the best performance in Predictive Interval Coverage Probability (PICP). Placing a dropout layer after each stacked block has the best performance in Root Mean Square Error (RMSE).
@article{shi22_jaiscr,
title = {Evaluating Dropout Placements in Bayesian Regression Resnet},
author = {Lei Shi and Cosmin Copot and Steve Vanlanduit},
year = {2022},
journal = {Journal of Artificial Intelligence and Soft Computing Research},
volume = {12},
number = {1},
pages = {61--73},
doi = {10.2478/JAISCR-2022-0005},
}
2021
Gaze Gesture Recognition by Graph Convolutional Networks
Lei Shi, Cosmin Copot, Steve Vanlanduit
Frontiers in Robotics and AI, 8, 2021.
AbstractLinksBibTeXProject
Gaze gestures are extensively used in the interactions with agents/computers/robots. Either remote eye tracking devices or head-mounted devices (HMDs) have the advantage of hands-free during the interaction. Previous studies have demonstrated the success of applying machine learning techniques for gaze gesture recognition. More recently, graph neural networks (GNNs) have shown great potential applications in several research areas such as image classification, action recognition, and text classification. However, GNNs are less applied in eye tracking researches. In this work, we propose a graph convolutional network (GCN)–based model for gaze gesture recognition. We train and evaluate the GCN model on the HideMyGaze! dataset. The results show that the accuracy, precision, and recall of the GCN model are 97.62%, 97.18%, and 98.46%, respectively, which are higher than the other compared conventional machine learning algorithms, the artificial neural network (ANN) and the convolutional neural network (CNN).
@article{shi21_frai,
title = {Gaze Gesture Recognition by Graph Convolutional Networks},
author = {Lei Shi and Cosmin Copot and Steve Vanlanduit},
year = {2021},
journal = {Frontiers in Robotics and AI},
volume = {8},
doi = {10.3389/frobt.2021.709952},
paper = {222},
}
GazeEMD: Detecting Visual Intention in Gaze-Based Human-Robot Interaction
Lei Shi, Cosmin Copot, Steve Vanlanduit
Robotics, 10 (2), pp. 1--18, 2021.
AbstractLinksBibTeXProject
In gaze-based Human-Robot Interaction (HRI), it is important to determine human visual intention for interacting with robots. One typical HRI interaction scenario is that a human selects an object by gaze and a robotic manipulator will pick up the object. In this work, we propose an approach, GazeEMD, that can be used to detect whether a human is looking at an object for HRI application. We use Earth Mover’s Distance (EMD) to measure the similarity between the hypothetical gazes at objects and the actual gazes. Then, the similarity score is used to determine if the human visual intention is on the object. We compare our approach with a fixation-based method and HitScan with a run length in the scenario of selecting daily objects by gaze. Our experimental results indicate that the GazeEMD approach has higher accuracy and is more robust to noises than the other approaches. Hence, the users can lessen cognitive load by using our approach in the real-world HRI scenario.
@article{shi21_robotics,
title = {{GazeEMD}: Detecting Visual Intention in Gaze-Based Human-Robot Interaction},
author = {Lei Shi and Cosmin Copot and Steve Vanlanduit},
year = {2021},
journal = {Robotics},
volume = {10},
number = {2},
pages = {1--18},
doi = {10.3390/robotics10020068},
paper = {68},
}
A Bayesian Deep Neural Network for Safe Visual Servoing in Human–Robot Interaction
Lei Shi, Cosmin Copot, Steve Vanlanduit
Frontiers in Robotics and AI, 8, pp. 1--13, 2021.
AbstractLinksBibTeXProject
Safety is an important issue in human–robot interaction (HRI) applications. Various research works have focused on different levels of safety in HRI. If a human/obstacle is detected, a repulsive action can be taken to avoid the collision. Common repulsive actions include distance methods, potential field methods, and safety field methods. Approaches based on machine learning are less explored regarding the selection of the repulsive action. Few research works focus on the uncertainty of the data-based approaches and consider the efficiency of the executing task during collision avoidance. In this study, we describe a system that can avoid collision with human hands while the robot is executing an image-based visual servoing (IBVS) task. We use Monte Carlo dropout (MC dropout) to transform a deep neural network (DNN) to a Bayesian DNN, and learn the repulsive position for hand avoidance. The Bayesian DNN allows IBVS to converge faster than the opposite repulsive pose. Furthermore, it allows the robot to avoid undesired poses that the DNN cannot avoid. The experimental results show that Bayesian DNN has adequate accuracy and can generalize well on unseen data. The predictive interval coverage probability (PICP) of the predictions along x, y, and z directions are 0.84, 0.94, and 0.95, respectively. In the space which is unseen in the training data, the Bayesian DNN is also more robust than a DNN. We further implement the system on a UR10 robot, and test the robustness of the Bayesian DNN and the IBVS convergence speed. Results show that the Bayesian DNN can avoid the poses out of the reach range of the robot and it lets the IBVS task converge faster than the opposite repulsive pose.
@article{shi21_frai_2,
title = {A Bayesian Deep Neural Network for Safe Visual Servoing in Human–Robot Interaction},
author = {Lei Shi and Cosmin Copot and Steve Vanlanduit},
year = {2021},
journal = {Frontiers in Robotics and AI},
volume = {8},
pages = {1--13},
doi = {10.3389/frobt.2021.687031},
paper = {687031},
}
2020
Visual Intention Classification by Deep Learning for Gaze-based Human-Robot Interaction
Lei Shi, Cosmin Copot, Steve Vanlanduit
IFAC-PapersOnLine, 53 (5), pp. 750-755, 2020.
AbstractLinksBibTeXProject
In this work, we propose a deep learning model to classify a human’s visual intention in gaze-based Human-Robot Interaction(HRI). We consider a scenario in which a human wears a pair of eye tracking glasses and can select an object by gaze and a robotic manipulator picks up the object. A neural network is trained as a binary classifier to classify if a human is looking at an object. The network architecture is based on Fully Convolutional Net(FCN), Convolutional Block Attention Modules(CBAM) and Residual Blocks. We evaluate our model with two experiments. In one experiment we test the performance in the scenario where only a single object exists and the other one multiple objects exist. The results show that our proposed network is accurate and it can generalize well. The F1 score on the single object is 0.971 and 0.962 on multiple objects.
@article{shi20_ifac,
title = {Visual Intention Classification by Deep Learning for Gaze-based Human-Robot Interaction},
author = {Lei Shi and Cosmin Copot and Steve Vanlanduit},
year = {2020},
journal = {IFAC-PapersOnLine},
volume = {53},
number = {5},
pages = {750-755},
doi = {10.1016/j.ifacol.2021.04.168},
}
A Deep Regression Model for Safety Control in Visual Servoing Applications
Lei Shi, Cosmin Copot, Steve Vanlanduit
Proc. IEEE International Conference on Robotic Computing (IRC), pp. 360--366, 2020.
AbstractLinksBibTeXProject
In Human-Robot Interaction scenarios, a human often needs to interact or closely working with objects and/or the robot. Hence the safety aspect needs to be taken care of in the Human-Robot Interaction scenarios. In this paper, we apply a deep learning approach to learning an optimal repulsive pose. The end effector of the robot will move the optimal repulsive pose if the human hand is too close to the end effector. We use a ResNet based deep regression model to learn the weights between the input i.e. the hand position + Tool Center Point position and output i.e. the repulsive pose. We evaluate the model with different readouts and loss functions. With the Fully Connected readout, the Mean absolute Error in the x, y and z directions are between 7.4 mm and 7.7 mm. The model inference time is also smaller than the computation time of calculating the optimal repulsive pose online.
@inproceedings{shi20_irc,
title = {A Deep Regression Model for Safety Control in Visual Servoing Applications},
author = {Lei Shi and Cosmin Copot and Steve Vanlanduit},
year = {2020},
booktitle = {Proc. IEEE International Conference on Robotic Computing (IRC)},
pages = {360--366},
doi = {10.1109/IRC.2020.00063},
}
2019
A Performance Analysis of Invariant Feature Descriptors in Eye Tracking based Human Robot Collaboration
Lei Shi, Cosmin Copot, Stijn Derammelaere, Steve Vanlanduit
Proc. International Conference on Control, Automation and Robotics (ICCAR), pp. 256--260, 2019.
AbstractLinksBibTeXProject
For eye tracking applications in Human Robot Collaboration (HRC), it is essential for the robot to be aware of where the human gaze is located in the scene. Using feature detectors and feature descriptors, the human gaze can be projected to the image from which robot could know where a human is looking at. The motion that occurs during the collaboration may affect the performance of the descriptor. In this paper, we analyse the performance of SIFT, SURF, AKAZE, BRISK and ORB feature descriptor in a real scene for eye tracking in HRC where different variances co-exist. We use a robotic arm and two cameras to test the descriptors instead of directly testing on eye tracking glasses in order that different accelerations can be tested quantitatively. Results show that BRISK, AKAZE and SURF are more favourable considering accuracy, stability and computation time.
@inproceedings{shi19_iccar,
title = {A Performance Analysis of Invariant Feature Descriptors in Eye Tracking based Human Robot Collaboration},
author = {Lei Shi and Cosmin Copot and Stijn Derammelaere and Steve Vanlanduit},
year = {2019},
booktitle = {Proc. International Conference on Control, Automation and Robotics (ICCAR)},
pages = {256--260},
doi = {10.1109/ICCAR.2019.8813478},
}
Application of Visual Servoing and Eye Tracking Glass in Human Robot Interaction: A case study
Lei Shi, Cosmin Copot, Steve Vanlanduit
Proc. International Conference on System Theory, Control and Computing (ICSTCC), pp. 515--520, 2019.
AbstractLinksBibTeXProject
Human gazes reveal a lot of information about their intention, which could be used for intuitive Human Robot Interaction(HRI). Using gaze as a modality to interact with robots has huge value in increasing freedom for human. It has a strong application potential in different domains. In this paper, we propose a system for eye tracking based HRI with which a human can select an object by his/her gaze and a robot manipulator receives the intended object and waits for the confirmation from the human before it starts the grasping task. We use Image-based Visual Servoing (IBVS) to guide the robot manipulator to perform the grasping automatically which can also bring more robustness to the system. The first hand results indicate that the use of both visual servoing features and eye tracking technique can provide successful results in term of human robot interaction.
@inproceedings{shi19_icstcc,
title = {Application of Visual Servoing and Eye Tracking Glass in Human Robot Interaction: A case study},
author = {Lei Shi and Cosmin Copot and Steve Vanlanduit},
year = {2019},
booktitle = {Proc. International Conference on System Theory, Control and Computing (ICSTCC)},
pages = {515--520},
doi = {10.1109/ICSTCC.2019.8886064},
}
Automatic Tuning Methodology of Visual Servoing System Using Predictive Approach
Cosmin Copot, Lei Shi, Steve Vanlanduit
Proc. IEEE International Conference on Control and Automation (ICCA), pp. 776--781, 2019.
AbstractLinksBibTeXProject
In this paper, a tuning methodology based on predictive model approach for visual servoing system is investigated. The proposed approach uses features prediction based method to estimate the camera velocity and thus to calculate the optimal control tuning parameter. In order to have a faster convergence and in the same time a desired behaviour of the servoing system, a new control parameter is computed every sampling time. To evaluate the designed tuning strategy, a visual servoing architecture with an eye-in-hand configuration has been considered. The experimental results showed that the proposed tuning algorithm based on prediction model has a stable and convergent behavior when dealing with visual servoing applications. To the knowledge of the authors, the proposed methodology is the first approach which enables an automatic selection of control parameter for the proportional visual control law.
@inproceedings{copot19_icca,
title = {Automatic Tuning Methodology of Visual Servoing System Using Predictive Approach},
author = {Cosmin Copot and Lei Shi and Steve Vanlanduit},
year = {2019},
booktitle = {Proc. IEEE International Conference on Control and Automation (ICCA)},
pages = {776--781},
doi = {10.1109/ICCA.2019.8899522},
}