# Voice Agents Need to Learn Voice Acting

> Animation spent almost a century learning how voice, timing, subtext, and restraint create a character. Voice agents should study the same craft.

Voice agents are getting very good at speaking. The next frontier is learning how to perform.

**Published:** 2026-08-26  
**Tags:** voice agents, LLMs, voice acting, animation, interaction design, performance, Soul  
**Canonical URL:** https://blog.nikdesign.ca/posts/voice-agents-need-to-learn-voice-acting

---
I started watching *Soul* tonight because it felt like the right movie for where my head was.

I had been thinking about life, purpose, work, recovery, identity, and the strange way we keep postponing being alive until some future condition is met.

After I get the job.

After I make enough money.

After I build the thing.

After I prove something.

After everything finally settles down.

Then *Soul* starts asking a much more uncomfortable question: what if life is not the thing waiting for us after all of that?

What if life is the thing happening while we wait?

A few minutes into the movie, another thought hit me.

It had nothing to do with the philosophy of *Soul*.

It was the voices.

Animated movies are one of the places where the term **voice acting** really earns its meaning.

In live action, we usually just call it acting. The actor's face, eyes, body, movement, posture, environment, timing, and voice all work together.

Animation breaks that performance apart.

The animators perform through movement.

The voice actor performs through sound.

The writers create intention.

The director shapes rhythm.

The character emerges somewhere in between.

That immediately made me think about voice agents.

Right now, we are putting enormous effort into making machines speak.

I am not sure we are putting enough effort into teaching them how to perform.

There is a difference. A very big one.

## I Learned This Before I Knew What Voice Acting Was

My relationship with animated performance started long before I knew any of these terms.

Growing up in India in the 1990s, two of my earliest relationships with animation were *Tom and Jerry* and *The Jungle Book*.

They taught me almost opposite lessons.

*Tom and Jerry* barely needed dialogue. I still think it represents animation at its purest.

Put *Tom and Jerry* on today and I will watch it. My father will watch it. If there are kids in the room, they will watch it too.

Three generations can sit together and understand exactly what is happening without anyone explaining anything.

Tom does not need to tell you he has realised he has made a terrible mistake. You see it.

Jerry does not need to explain that he is three steps ahead. You know.

The anticipation before something goes wrong, the pause before impact, the expression, the impossible deformation of a body, the timing of a glance: the characters are acting.

Animation itself is performing.

That is one of the most important lessons in the history of animation:

> **You do not need realism to create believable performance.**

Sometimes you do not even need language.

## Animation Has Been Studying This Problem for Almost a Century

Then there was *The Jungle Book*.

Not the Disney film. The one a generation of Indian kids knew from television: *The Jungle Book: Shōnen Mowgli*, the Japanese animated series dubbed into Hindi.

I had absolutely no idea it was anime. Anime was not really part of my vocabulary growing up.

It was just *The Jungle Book*.

It was Mowgli.

It was Baloo, Bagheera, and Shere Khan.

Nana Patekar voiced Shere Khan in the Hindi version. When you learn that as an adult, it suddenly makes perfect sense.

But as a child, there was no Nana Patekar performing an animated tiger. There was just Shere Khan.

That voice belonged to him.

The Hindi adaptation itself became part of the show's identity. The opening song, *Jungle Jungle Baat Chali Hai*, had lyrics by Gulzar and music by Vishal Bhardwaj.

Think about what happened there.

Stories written by an English author became a Japanese animated series, which was then performed in Hindi by Indian actors and musicians, and somehow arrived in my living room as something that felt completely natural to me.

That is more than dubbing.

That is performance creating cultural identity.

Somewhere between *Tom and Jerry* and Mowgli, I was being introduced to two sides of the same idea without knowing what either was called.

*Tom and Jerry* taught me that an artificial character can perform without speaking.

Mowgli taught me that when the character does speak, the voice can become inseparable from who that character is.

Look across the history of animation and you can see the discipline evolving.

*Snow White and the Seven Dwarfs*.

*Pinocchio*.

*Dumbo*.

*Bambi*.

*Cinderella*.

*Peter Pan*.

Disney's *The Jungle Book*.

Then *The Little Mermaid*, *Beauty and the Beast*, *Aladdin*, and *The Lion King*.

By the time Robin Williams performed the Genie, it became impossible to think of voice acting as somebody simply reading lines into a microphone.

The performance shaped the character.

### The Genie Was a Generative Performance

There is a detail about that performance that matters even more now.

Disney did not hire Robin Williams and then ask him to stay inside the script. The filmmakers had imagined the Genie around Williams from the beginning, and directors Ron Clements and John Musker explicitly gave him room to improvise. Musker later said they never expected Williams to simply come in and read what they had written.

The first recording session showed them what that actually meant. A scene designed to last only a few minutes was performed again and again as Williams kept adding, changing, riffing, impersonating, and discovering new possibilities. By the 25th take, according to Clements, that material had expanded to roughly 20 minutes. Supervising animator Eric Goldberg was apparently laughing so uncontrollably during the session that he had to be removed from the recording stage because he was interfering with the takes.

And that was only the beginning. Williams did several more four-hour sessions. The filmmakers ended up with roughly **sixteen hours of recorded material for a movie that runs about ninety minutes**.

That ratio is worth sitting with.

Most of what Williams generated was never going to appear in the finished movie. That was not waste. **The abundance was part of the process.** He generated possibilities. The filmmakers selected from them. The animators then had something unusually rich to respond to.

Goldberg and the animation team did not simply animate a fixed vocal track as though Williams were reading dialogue from a page. They could take an unexpected voice, rhythm, reference, sound effect, or character transformation that emerged in the booth and invent a visual performance around it. Williams might suddenly become a game-show host, an evangelist, Walter Cronkite, Arnold Schwarzenegger, John Wayne, or somebody else entirely, and the Genie could physically transform with him.

The loop was not simply:

> **script → actor → animation**

It was closer to:

> **script → intent → improvisation → variation → selection → animation → character**

That distinction matters.

Williams was not generating randomness. He was operating inside constraints: the character, the scene, the relationship with Aladdin, the story, the audience, the emotional objective, and decades of accumulated performance knowledge. Inside those constraints, he could generate possibilities in real time.

That is remarkably close to the problem voice agents are beginning to face.

A useful agent should not merely find a sentence and pronounce it convincingly. It needs to understand context well enough to decide how this particular moment should be performed. Sometimes that may mean restraint. Sometimes hesitation. Sometimes warmth. Sometimes a joke. Sometimes an unexpected turn that was not explicitly scripted but is completely right for the moment.

The Genie also created a problem for the awards system because the performance did not fit comfortably into the categories available at the time. *Aladdin* received five Academy Award nominations, but Williams was not nominated for acting. Golden Globe voters confronted the unusual nature of the performance more directly and gave him a **Special Achievement Award** in 1993.

In retrospect, that difficulty classifying the performance feels almost appropriate. Williams was doing something that sat between established disciplines: writing, acting, improvisation, vocal performance, and animation. The character emerged from all of them interacting.

That is why *Aladdin* deserves more than a footnote in a discussion about voice agents. The Genie was, in a very literal sense, a generative performance system decades before we started using that language.

Jeremy Irons did not merely give Scar a voice. He gave him rhythm, intelligence, arrogance, threat, and theatricality.

Then you move through *Mulan*, *The Prince of Egypt*, *The Iron Giant*, *The Emperor's New Groove*, *Lilo & Stitch*, and *Spirited Away*.

Different visual traditions. Different cultures. Different storytelling languages.

The same fundamental principle.

The voice is not decoration added after the character has been created. **The voice is part of the character's creation.**

Then 3D animation made the relationship even more obvious.

*Toy Story* works because Woody and Buzz feel like people before they feel like rendered geometry.

*Shrek* works because every character has an unmistakable vocal identity.

*Finding Nemo* carries fear, grief, frustration, affection, and humour through delivery.

*The Incredibles* works because every member of the family has a different rhythm.

*Ratatouille* gives a rat interiority.

*Kung Fu Panda* lets absurd comedy and genuine vulnerability coexist in the same character.

*Up* understands restraint.

*How to Train Your Dragon* understands awkwardness and vulnerability.

## Speech Is Not Performance

*Inside Out* literally turns emotional states into characters.

*Coco* uses voice to carry memory, family, grief, and warmth.

*Spider-Man: Into the Spider-Verse* builds identity through rhythm, attitude, timing, and style.

*Puss in Boots: The Last Wish* allows a familiar character to suddenly sound genuinely afraid.

*The Wild Robot* gives us something especially interesting for the age of voice agents: a synthetic character whose vocal performance changes as the character changes.

And then there is *Soul*.

A movie about purpose, existence, mortality, and the experience of being alive makes those ideas feel intimate.

That is not just writing.

It is performance.

Take the simplest possible sentence:

> **“I understand.”**

What does it mean?

It could mean:

- I hear you.
- I believe you.
- I sympathise with you.
- I disagree, but I understand your position.
- I am impatient and want you to stop talking.
- I am frightened but trying to sound calm.
- I know more than I am telling you.
- I do not understand at all, but I do not want to admit it.
- I am reassuring you.
- I am warning you.

The words have not changed.

The performance has.

That difference is voice acting.

It is also where I think voice agents still have enormous room to grow.

Current voice systems are getting very good at speech: pronunciation, latency, turn-taking, interruption, natural pauses, emotional synthesis, voice cloning, prosody, and real-time generation.

All of that matters.

But those capabilities do not automatically produce a performance.

A natural-sounding voice can still feel empty.

A perfectly cloned voice can still sound wrong.

A voice can be technically expressive while having no idea what the emotional objective of the scene is.

Actors think differently.

A good actor is not simply thinking, “How should this sentence sound?”

They are thinking:

- What does this character want right now?
- What happened immediately before this line?
- What am I trying to make the other person feel?
- What am I hiding?
- How much should I reveal?
- Am I certain, or am I pretending to be certain?
- Do I answer immediately?
- Do I hesitate?
- Do I breathe first?
- Do I soften the ending?
- Do I let the silence do the work?

That is a different computational problem.

## Voice Agents Need Directors

I think voice-agent companies will eventually need something like a **performance direction function**.

Not only speech engineers, but voice actors, dialogue directors, animation voice directors, writers, interaction designers, linguists, behavioural designers, and people who understand timing.

People who understand subtext.

People who know why one performance feels emotionally truthful while another feels like somebody moved an “emotion” slider to 70 percent.

These teams should study animation systematically, not only as entertainment, but as research.

Study *Tom and Jerry* to understand performance without dialogue.

Study the Hindi *Shōnen Mowgli* to understand how voice can recreate a character across language and culture.

Study *Aladdin* for improvisational energy.

Study *The Iron Giant* for restraint.

Study *Up* for emotional economy.

Study *Puss in Boots: The Last Wish* for fear.

Study *The Wild Robot* for the evolution of a synthetic character's voice.

Study *Soul* for how performance can carry philosophy without sounding like a philosophy lecture.

Then study puppetry, radio drama, audiobooks, theatre, dubbing, and video games.

Study anywhere humans have learned to create a person using primarily a voice.

## The Architecture May Be Missing a Layer

The simplified architecture of a voice agent today often looks something like:

> **LLM → text → TTS → audio**

Real systems are obviously more complicated than that. But conceptually, the voice still tends to appear toward the end.

The model figures out what to say.

The speech system figures out how to say the words aloud.

I think we are going to need another layer:

> **intent → character state → performance direction → language → vocal performance**

That performance layer could contain information the user never sees:

- The user sounds discouraged.
- Lower your energy.
- Do not become artificially cheerful.
- Pause before answering.
- You are uncertain here. Let that uncertainty remain audible.
- This is reassurance, not instruction.
- The joke should sound accidental rather than performed.
- Do not fill this silence.
- You already know this person well. Do not sound like customer support.
- Let the final sentence breathe.

None of that necessarily changes the semantic meaning of the response.

It changes the performance.

You could think of it as a **digital voice director sitting between cognition and speech**.

## A Voice Is an Instrument

This becomes even more important as agents become persistent.

I have been thinking about it while working with dynamic voices in my own agent systems.

A persistent agent cannot simply have a recognisable voice. That is not enough.

It needs a recognisable way of speaking.

A rhythm.

A range.

A temperament.

A sense of restraint.

A way of responding when something is funny.

A different way of responding when something is serious.

A pattern of silence.

Maybe even imperfections.

That is how we recognise people.

A cloned voice gives you an instrument.

**It does not give you the musician.**

A highly realistic synthetic voice without performance intelligence may eventually become the voice equivalent of photorealistic graphics with terrible animation.

Technically impressive.

Emotionally dead.

Animation learned this lesson a long time ago.

The goal was never merely to draw a convincing human. Some of its greatest characters are not human at all.

They are toys, rats, robots, ogres, pandas, fish, cars, emotions, animals, monsters, and, in *Soul*, literal souls.

Animation discovered that audiences do not require realism.

They require performance.

Maybe voice agents should spend less time obsessing over the question:

> *How do we make this sound human?*

And start asking a more interesting one:

> **How would an actor perform this?**
