# Does Emotion Simulation Make Android Robots More Accepted?
A controlled field study of android Andrea at a German museum — 73 visitors, six consecutive days, three emotion conditions — found no statistically significant difference in visitor acceptance whether the robot simulated emotions via ChatGPT 4.1, the WASABI emotion architecture, or no emotion model at all. The result is a direct challenge to a foundational assumption baked into a wide swath of social humanoid development: that expressive emotional behavior meaningfully improves human-robot interaction (HRI) acceptance scores.
The study, authored by Marcel Heisler and Christian Becker-Asano and published on arXiv on July 21, 2026, is one of the more rigorous real-world HRI field trials to appear this year. Unlike lab studies where participants know they are being evaluated, museum deployments expose robots to unscripted, demographically diverse visitors with no prior briefing — about as ecologically valid as HRI research gets outside of a factory floor.
The core finding: at a conscious, self-reported level, visitors could not distinguish the emotional conditions, and neither emotion approach improved subjective evaluations compared to the emotion-free baseline.
---
## What the Researchers Actually Tested
Andrea is a gender-ambiguous android robot with a highly humanlike physical design. The paper notes its voice was previously evaluated as "slightly artificial sounding" but rated as congruent with that gender-ambiguous design — a detail worth noting because voice-appearance congruence is a known confound in android acceptance studies.
For this second museum deployment, the researchers meaningfully upgraded Andrea's conversational capability relative to prior trials. The robot could now engage in **multi-lingual conversation** about the museum context, having been loaded with information about the museum in general and the surrounding exhibits specifically. This contextual grounding represents a step toward the kind of domain-specific [Vision-Language-Action Model](https://humanoidintel.ai/glossary/vision-language-action-model) integration that commercial humanoid teams are pursuing — though Andrea's deployment is purely conversational, not physically interactive.
Three conditions were tested:
1. **No emotion simulation** — the robot's chat architecture ran without any affective overlay
2. **ChatGPT 4.1-determined emotions** — the LLM was queried to determine emotional state alongside conversational responses
3. **WASABI architecture** — a dedicated emotion simulation system that models dynamic emotional state changes over time
Visitor acceptance was measured using an extended version of the TAM2 questionnaire — the Technology Acceptance Model, second iteration — a validated instrument that probes factors including perceived usefulness, ease of use, and subjective norms. Seventy-three visitors completed the questionnaire after interacting with Andrea across six consecutive days.
---
## The Null Result and Why It Matters
The statistical analysis found that neither emotion approach produced positive effects on visitor evaluations compared to the no-emotion baseline. The authors specifically note the effects "were not detectable on a conscious level."
That qualifier is important. The study measured self-reported, conscious perception. It does not claim that emotional simulation has no effect on behavior — physiological responses, dwell time, or return visits were not reported in the abstract. But for HRI researchers and social robot developers, TAM2-derived acceptance scores are often the primary design target.
**The industry implication is pointed.** A substantial portion of the engineering effort going into social humanoids — from expressive face actuators to affective AI layers — rests on the premise that emotion expression drives acceptance. This study doesn't disprove that premise universally, but it does add field-validated evidence that the specific implementation pathway (LLM-queried emotion state, or a dedicated affect architecture like WASABI) may not move the needle on conscious acceptance metrics in public settings.
There are at least two skeptical reads on the result worth holding simultaneously:
- **Implementation quality matters.** ChatGPT 4.1's emotion determination and WASABI's dynamics may have been producing emotional signals that never surfaced perceptibly in voice or behavior. If the affective layer doesn't change observable output, it cannot change visitor perception.
- **The Uncanny Valley may be dominating variance.** Andrea's highly humanlike design may produce such strong prior expectations — or such strong uncanny responses — that emotional nuance is simply noise. Visitors may be responding primarily to the android's appearance, not its behavioral state.
Neither interpretation is resolved by the available abstract text, and the full paper will likely address both.
---
## Broader Context for the Humanoid Industry
The commercial humanoid sector is currently making significant bets on social and conversational capability alongside physical capability. Teams across the industry are integrating large language models into robot control stacks, and some are layering affective models on top. This study suggests that the HRI research community is moving faster than the commercial sector in generating field-validated evidence about what actually works.
For humanoid developers: the finding that multi-lingual, context-grounded conversation was successfully deployed in a fully autonomous, six-day public setting is arguably more commercially relevant than the emotion null result. That's the capability threshold that matters for a museum guide, a retail assistant, or a hospital wayfinding robot — and Andrea appears to have cleared it.
The WASABI architecture result is particularly notable for researchers. WASABI is a well-established computational model of emotion dynamics in HRI; its failure to outperform a no-emotion baseline in this setting is a data point the field will need to account for in future work.
---
## Key Takeaways
- **73 visitors, 6 days, 3 emotion conditions**: the study is among the more rigorous field deployments in recent android HRI literature
- **Null result on emotion**: neither ChatGPT 4.1-driven nor WASABI-driven emotion simulation improved TAM2 acceptance scores versus no emotion at all
- **Emotions were "not detectable on a conscious level"** — the study's own framing, which leaves open the question of subconscious or behavioral effects
- **Contextual multi-lingual conversation worked**: Andrea successfully engaged visitors about museum-specific content fully autonomously — that's the baseline commercial teams should be tracking
- **Implication for developers**: affective AI layers may require perceptible behavioral output changes to register with users; invisible emotion computation doesn't appear to help
---
## Frequently Asked Questions
**What is the android robot Andrea?**
Andrea is a gender-ambiguous android robot with a highly humanlike design used in human-robot interaction research. According to the study by Heisler and Becker-Asano, its slightly artificial-sounding voice was previously evaluated as congruent with its appearance.
**What did the museum study find about emotion simulation in robots?**
The study found that neither ChatGPT 4.1-determined emotions nor the WASABI emotion simulation architecture improved visitor acceptance scores compared to a no-emotion baseline, as measured by an extended TAM2 questionnaire across 73 museum visitors.
**What is the WASABI emotion architecture?**
WASABI is a computational architecture designed to simulate dynamic emotional states in robots over time. It is a dedicated affective system, distinct from querying a general-purpose LLM for emotion determination.
**Does this mean emotional robots are no better than unemotional ones?**
Not necessarily. The study measured self-reported, conscious acceptance via TAM2. It does not address physiological responses, behavioral engagement, dwell time, or other interaction metrics. The result is specifically that emotion simulation was not detectable at a conscious level in this deployment.
**Why does this matter for commercial humanoid robots?**
Significant engineering resources across the humanoid industry are being allocated to expressive behavior and affective AI. This field study provides evidence that emotion simulation architectures need to produce perceptible behavioral changes to influence user acceptance — invisible affective computation does not appear to be sufficient.
RESEARCH
Android Andrea's Emotions Don't Move Museum Visitors
Published: July 21, 2026 at 24:00 EDTLast updated: July 21, 2026 at 07:34 EDTBy Alex Reiner, Senior EditorLast reviewed by Alex Reiner on July 21, 20267 min read
73 museum visitors rated android Andrea equally whether it used ChatGPT 4.1 emotions, WASABI dynamics, or none at all.
androidemotion-simulationhuman-robot-interactionchatgptTAM2museum-roboticsacceptance