# Does an Expressive Humanoid Robot Pay a Steeper Trust Penalty When It Fails?

**Yes — by more than half.** A Drexel University-led study published in *Science Robotics* found that a humanoid robot's measurable influence over participants' decisions fell by more than half after it began making mistakes and breaking social norms. Critically, the penalty was sharper for the expressive version of the robot — one that made eye contact, gestured, nodded, and engaged — than for a motionless counterpart delivering identical spoken content. The study, conducted with researchers from the U.S. Air Force Academy, George Mason University, and the University of Southern California's Institute for Creative Technologies, involved 50 healthy adult men interacting with Pepper, a humanoid robot, over roughly two and a half hours each. The robot was controlled by a human operator following a script, keeping verbal responses consistent across both expressive and stationary conditions. Funding came from a U.S. Department of Defense Air Force Office of Scientific Research grant. The core finding is a direct engineering problem: **expressiveness is not free**. It creates expectations that compound the cost of failure.

---

## What the Study Actually Measured

Most human-robot interaction (HRI) research on trust relies on post-session surveys alone. This study is methodologically notable for measuring trust through four simultaneous channels:

1. **fNIRS (functional near-infrared spectroscopy)** — wearable sensors monitoring activity in the prefrontal cortex
2. **Saliva samples** — measuring oxytocin levels throughout the session
3. **Self-reported surveys** — capturing subjective trust ratings
4. **Behavioral compliance** — tracking whether participants actually changed their decisions based on the robot's recommendations

"Trust is not one thing you can capture with a single measurement, so we measured several at once," said senior author Hasan Ayaz, PhD, a professor in Drexel's Howley College School of Biomedical Engineering and Science. "Each modality tells a different part of the story, and only together do they show how trust builds and how it breaks down."

This multi-modal approach is what makes the findings credible to a skeptical engineering audience. A robot's ability to move a human decision — not just produce a favorable survey score — is the operationally relevant metric for anyone deploying humanoids in collaborative or advisory roles.

---

## The Experimental Structure and Where Things Broke Down

Participants completed two interaction sessions with Pepper without deliberate errors. In the third session, the robot began interrupting, offering irrelevant comments, and providing illogical explanations — a controlled degradation of social norms rather than a mechanical failure.

This design choice matters. The study wasn't testing what happens when a robot drops something or misidentifies an object. It was testing violations of *social contract* — the implicit behavioral expectations that expressiveness itself creates. That distinction has direct relevance for [Physical AI](https://humanoidintel.ai/glossary/physical-ai) deployment contexts where humanoids work alongside humans in unstructured environments.

When the expressive robot made mistakes, participants showed increased activity in brain regions associated with social reasoning. The researchers interpreted this as evidence that people were processing the robot's failures as a social violation rather than a machine malfunction — a categorically different cognitive response. The motionless robot did not produce the same neurological pattern.

The oxytocin result is the most counterintuitive finding: oxytocin levels *increased* as the robots made mistakes, even as reported trust declined. The researchers suggest oxytocin may function as a signal of social vigilance rather than purely as a bonding indicator — a nuance that challenges the straightforward "oxytocin = trust" framing common in popular neuroscience coverage. Given that this study used fNIRS and hormonal assays simultaneously — methodologies also explored in brain-computer interface research — readers tracking the convergence of neuroscience and robotics may find [bciintel.com](https://bciintel.com) a useful parallel reference.

---

## The Engineering Implication: Expressiveness as Technical Debt

The clearest practical statement in the paper comes from corresponding author Frank Krueger, PhD, of George Mason University's School of Systems Biology: "Expressiveness is not free. It buys you engagement, and it buys you fragility at the same time, and that is a trade-off designers should be making deliberately rather than by accident."

Corresponding author Ewart J. de Visser, PhD, of the Warfighter Effectiveness Research Center at the U.S. Air Force Academy, was more direct: "An expressive robot that is unreliable is not a safer robot, it is a more disappointing one, because expressiveness raises a bar the robot then fails to clear."

For engineers and product teams, this reframes expressiveness as a form of technical debt. Every social behavior a humanoid robot exhibits — eye contact timing, nodding cadence, conversational back-channeling — is an implicit promise to the user. When the robot's downstream reliability doesn't match the social contract its expressiveness establishes, the trust collapse is steeper than it would have been with a blander design.

This has concrete implications across the humanoid stack. Companies racing to add personality layers and expressive behaviors to their platforms — often to differentiate in a crowded market — are implicitly accepting that their reliability engineering must keep pace. A humanoid with well-crafted whole-body social gestures but a 15% task failure rate may engender *less* operator trust than a functionally equivalent but expressively neutral machine.

---

## Why This Study's Timing Matters for the Industry

The humanoid sector is entering a phase where robots are moving from controlled pilots to sustained human-collaborative deployments. The question of how humans calibrate trust in real-time — not in a lab survey but through actual behavioral compliance — is directly relevant to deployment success rates and the pace at which operators cede decision authority to robotic systems.

The study's finding that **reliability ultimately mattered more than expressiveness for both robot versions** should recalibrate design priorities. The social engagement budget is not unlimited, and it cannot substitute for underlying system reliability. For teams building the behavioral layers on top of humanoid platforms — whether through [imitation learning](https://humanoidintel.ai/glossary/imitation-learning) pipelines or VLA-driven action policies — the neuroscience here is a useful constraint: users will always find out when your model fails, and the more charismatic you've made the robot, the more they'll penalize it when it does.

---

## Key Takeaways

- **Expressiveness amplifies trust loss**: A humanoid robot's decision influence fell by more than half after errors, with the expressive version suffering a sharper penalty than the stationary one.
- **Social violation, not mechanical failure**: Brain imaging showed participants processed expressive robot errors as social norm violations, not machine malfunctions — a qualitatively different cognitive response.
- **Oxytocin as vigilance signal**: Rising oxytocin during robot errors challenges the assumption that the hormone purely tracks bonding; researchers interpret it as a social vigilance indicator.
- **Reliability outranked expressiveness**: Across both robot conditions, reliability was the dominant predictor of sustained trust — expressiveness was a multiplier, not a foundation.
- **Design implication**: Social expressiveness creates implicit expectations. Engineering teams must treat reliability as a prerequisite for expressive behavior, not a parallel track.
- **Multi-modal measurement matters**: The study's simultaneous use of fNIRS, hormonal assays, surveys, and behavioral tracking provides a more defensible trust measurement framework than survey-only HRI research.

---

## Frequently Asked Questions

**What robot was used in the Drexel trust study?**
The study used Pepper, a humanoid robot, operated by a human controller following a script to keep verbal responses consistent across expressive and non-expressive conditions.

**How was trust measured in the study?**
Researchers used four simultaneous methods: wearable fNIRS brain imaging targeting the prefrontal cortex, saliva-based oxytocin measurement, participant surveys, and behavioral tracking of whether participants changed decisions after receiving the robot's advice.

**Why did the expressive robot lose more trust than the stationary robot?**
The study found that expressiveness creates social expectations. When the expressive robot made mistakes — interrupting, giving irrelevant comments, offering illogical explanations — participants' brains processed these as social violations rather than mechanical errors, producing a steeper trust penalty.

**What does oxytocin have to do with trusting robots?**
Oxytocin levels increased when both robot versions made mistakes, even as self-reported trust fell. The researchers interpret this as evidence that oxytocin can signal social vigilance, not just bonding — a finding that complicates straightforward assumptions about the hormone's role in human-robot trust.

**What does this mean for humanoid robot design?**
Designers should treat expressiveness as a commitment, not a differentiator. A socially engaging robot that fails to meet the behavioral expectations it creates generates *more* trust damage than a neutral machine making the same errors. Reliability engineering must precede — not follow — the addition of expressive social behaviors.