@inproceedings {pub4209,
	title = {Spoken Language Interaction with Virtual Agents and Robots based on Transparency, Situatedness and Personalization},
	author = {Martin Ernst Heckmann},
	year = {2020},
	month = {March},
	abstract = {Spoken Language Interaction with Virtual Agents and Robots based on Transparency, Situatedness and Personalization


When two humans engage in an interaction, two independent minds with different experiences and views on the world come together.
To make this joint activity a success, they have to work together.
They need to align their mental representations to be able to form a common ground and define a joint goal for the interaction (Clark1991,Pickering2004,Garrod2004,Friston2015,Scott-Phillips2015).
The communicative signals they use for such an alignment are not limited to the words they utter but encompasses a multitude of potentially multimodal signals such as prosodic variations and gestures (Hirschberg2002,Wagner2014).
These signals also help to coordinate when the partners take their turns (Sacks1974).  
The necessary alignment between the partners affects a multitude of domains, the phonological, syntactic and semantic level as well as the situation model (Pickering2004,Garrod2004). 
With situation model, I want to refer to different aspects of the current situation such as the environment in which the interaction takes place and the progression of the interaction so far.
During the interaction they also make assumptions about their partner{\textquoteright}s mental world, her ability to perceive, understand and judge, often referred to as common sense, her prior knowledge on the topic of the interaction, her goals and intentions and her current state, traits, skills and personal preferences (Friston2015).
The larger the differences between the partners{\textquoteright} mental worlds, the more work the alignment process needs.
A true alignment can never be reached but is also not necessary as long as the joint goals of the interaction are achieved.

Unfortunately, the different building blocks fundamental to human-human communication which I have outlined above, i.e. the capabilities to perceive and interpret multimodal communicative signals, build an adequate model of the environment, keep track of the history of the interaction, reason with common sense, recruit domain knowledge, interpret the partner{\textquoteright}s goals as well as model her state, traits, skills and preferences, are at best manifested in a very different form in an artificial agent.
In today{\textquoteright}s artificial agents, they are often not present at all or only in a rudimentary form. 
Furthermore, if and how these capabilities are implemented in the artificial agent is in most cases opaque to the user, i.e. the system is not transparent to the user.
This means that the alignment process will take a lot of effort from the human and in the end will leave a large gap between the partners{\textquoteright} mental models, often too large to achieve the goal of the interaction.

A common and logical approach to improve the interaction is to improve the system{\textquoteright}s capabilities, to make them closer to that of a human. 
Many people have investigated the role of prosody (Sagisaka1997,Hirschberg2002) as well as how to integrate the information from all available modalities (Turk2014,Oviatt2018,Stiefelhagen2004a,Funakoshi2012,Kennington2017).
We have extended this by multimodal integration and personalization for prosodic analysis (Heckmann2018,Schnall2018) and reference resolution (Kleingarn2019).
In classical dialog systems, the situation model is often limited to keeping track of the dialog history (Williams2013,Henderson2014,Yoshino2019).
For human robot interaction, mainly models have been developed to visually perceive the world and link these percepts to internal concepts, words in particular.
This process is often referred to as grounding, sometimes also as anchoring, and can be based on pre-existing categories (Misra2016,Lemaignan2017,pub3646pub3696) or learned in interaction (Roy2002,Matuszek2018).
The domain knowledge of classical dialog systems is typically very narrow, based on hand-crafted domains such as bus information (Williams2013), restaurant reservation (Henderson2014), or, more recently, also technical support dialogues (Yoshino2019).
For the interaction with robots it is common to establish the domain knowledge by combining information from external ontologies with the learning of new representations in interaction with the user (Tenorth2013,Lemaignan2017).
Common sense reasoning is then implemented by reasoning on these ontologies.
We proposed an approach to automatically acquire task-specific domain knowledge from online text-based resources, e.g. in the context of a repair task (Nabizadeh2020). 

Many of the functionalities mentioned above are still nowhere near the level of a human.
In my view, the most challenging and promising domains relevant to improve the interaction with artificial agents are better environment models and domain knowledge combined with an adaptation to the individual user.

To really meet human expectations will most likely require an AI with reasoning capabilities on par with a human.
Yet such an AI does not seem to be around the corner (Grace2018).
Hence, instead of raising the system{\textquoteright}s perception and reasoning capabilities, I think a more promising approach is to increase its transparency.
If the system states, capabilities and intentions are transparent to the human, the alignment process and hence the overall interaction can be much more efficient (Moore2017a).
Humans use backchannels, facial expressions and emotional displays to give feedback on the interaction and the progress of the alignment (Eshghi2015,Truong2011,PUBA83,Ekman2003).
Backchannels, facial expressions and emotions have been also investigated in the context of human-agent interaction (Skantze2014,Becker2007,PUBA83).
However, the research on emotions in human-agent interaction typically focuses on detecting the emotions of the human to enable emphatic responses of the agent (Gunes2011,Devillers2015).
I consider using emotional displays and other non-verbal cues to provide insights into the capabilities and mental states of the agent an interesting and very relevant topic to help to align the minds of the partners and to improve the interaction.
Connected to this is the need to empower the agent to make its reasoning steps transparent, often referred to as explainable AI (Abdul2018,Adadi2018).
We took one step into this direction by suggesting an approach to create annotations with explanations that can be used to train a system to provide explanations for its reasoning (Attari2019).
In general, spoken feedback from the agent bears the risk of triggering very high expectations on the human{\textquoteright}s side with respect to the agent{\textquoteright}s understanding and reasoning capabilities.
I consider finding good models to convey the agent{\textquoteright}s limitations via its communication an important direction to increase transparency (Moore2017).

In short, I am convinced that faster progress can be made if we focus more on making the limitations of the agent transparent then on trying to overcome them.
},
	publisher = {Schloss Dagstuhl-Leibniz-Zentrum fuer Informatik},
	booktitle = {Dagstuhl Seminar Spoken Language Interaction with Virtual Agents and Robot (SLIVAR)},
	editor = {Laurence Devillers, Tatsuya Kawahara, Roger K. Moore, Matthias Scheutz}
}
