Visual Language Models Train Robots to Read Human Emotions
Researchers trained collaborative robots to read human emotions using a vision language model (VLM) that accounts for context, not just facial expressions. In experiments with 40 volunteers, the VLM matched human observers' emotion readings better than a conventional facial-analysis system. But the study found the emotional capabilities of robots only go so far: when a robot failed at its task, participants trusted it less regardless of how it apologized.
Key Takeaways
- As robots advance in terms of dexterity and other physical capabilities , it becomes more likely that humans may find themselves working alongside them.
If that happens, how will robots' emotional capabilities need to advance for them to successfully work with people?
- He notes that, while there has been a lot of hype in the advancing physical abilities of robots, this is only one piece of the puzzle.
"We need to also innovate when it comes to them actually interacting with humans, not just their physical capabilities," he says.
- For example, a person pausing to think with a furrowed brow may simply be concentrating on their task at hand, and not necessarily be angry.
Contextual factors such as drumming their fingers, pursing their lips, or other behaviors can point to the real cause of a person's furrowed brow.
- In comparison, the VLM achieved a score of 0.86.
- After collaborating with a robot that failed in its task, many participants ranked their trust in the robot as lower, regardless of how it apologized for its mistake.

In a recent study, researchers trained collaborative robots to read human emotions by not only accounting for facial expressions, but also contextual factors in the interactions as well. Through experiments with 40 volunteers, the researchers then evaluated how a robot's ability to read human emotions and adjust its behaviour in turn impacted a human's perception of the robot and its capabilities as the two collaborated on tasks. The results -which show that the emotional capabilities of robots only go so far with humans-were published 18 May in IEEE Robotics and Automation Letters .
Seung Chan Hong led the study as part of his undergraduate thesis while studying at the University of Melbourne, in Australia. He notes that, while there has been a lot of hype in the advancing physical abilities of robots, this is only one piece of the puzzle. "We need to also innovate when it comes to them actually interacting with humans, not just their physical capabilities," he says.
This prompted him to dig deeper into the emotional aspects of human-robot interactions. First, Hong and his co-authors decided to train a robot to read human emotions using a vision language model (VLM), which is similar to large language models such as ChatGPT, but which can also take visual inputs. Training VLMs for Human Emotion Recognition To train their VLM, the researchers had volunteers watch videos of robots handing over objects to humans-with varying degrees of success-and describe the emotions the humans were expressing.
For more details please read the original article at IEEE Spectrum AI.
Continue Learning
Comments
Comments appear only after moderation. Your email identifies your submission to the moderator and is never displayed here.
No approved comments yet.