Text-to-speech Multi-speaker dialogue ElevenLabs Commercial use
Voice Creation

ElevenLabs v3
– Text-to-speech
model in Phygital+

ElevenLabs v3 is the most expressive model in the ElevenLabs lineup. It reads inline audio tags such as [whispers], [excited], and [sighs] as performance direction, supports more than 70 languages, and renders multi-speaker dialogue as a single coherent scene.

In Phygital+ it is the default model in Sound Creation and Dialogue Creation, so speech, sound effects, and two-speaker scripts all run from the same canvas as your video and image nodes.

ElevenLabs v3

At a glance ElevenLabs v3

Built for

Voice-over, narration, character dialogue, ads, explainers, and localised audio

Content, localisation, and game teams that need delivery, not just intelligible speech

Vendor
ElevenLabs
Category
AI audio — text-to-speech and dialogue
Modalities
Text → audio
Languages
70+
Audio tags
Inline emotion and delivery direction
Nodes
Sound Creation, Dialogue Creation, Voice Creation
Commercial use
Yes
Access
All Phygital+ plans

Why use ElevenLabs v3

Inline performance direction, 70+ languages, and multi-speaker dialogue in one model.

1

Direction written into the script

Bracketed cues like [whispers], [excited], or [sighs] go straight into the script and tell the model how to deliver the line. It is closer to directing an actor than configuring a synthesiser, and it happens inside the text rather than in a settings panel.

2

70+ languages, one voice identity

More than 70 languages, up from 29 in Multilingual v2. A cloned voice keeps its accent character across all of them, so the same speaker stays recognisable in every localised version instead of turning into a different person per market.

3

Multi-speaker dialogue

Two speakers, two scripts, one generation. The model matches prosody across the exchange, so reactions and interruptions land the way a conversation actually does rather than sounding like two monologues cut together.

4

Better with difficult text

Chemical formulas, numbers, abbreviations, and other awkward strings that earlier models mangled are handled far more reliably. Less time spent respelling words phonetically to make them come out right.

5

Reusable voice identities

Build a voice from a written brief with Voice Creation, or clone one from audio samples with Voice Clone, then reuse that Voice ID across every generation. Recurring characters and a consistent brand voice stay consistent across clips.

Built for content, localisation, and game teams

ElevenLabs v3 fits wherever the delivery matters as much as the words: narration, character dialogue, ads, and localised versions of the same script.

ElevenLabs v3
ElevenLabs
ElevenLabs
ElevenLabs
ElevenLabs v3
ElevenLabs
ElevenLabs
ElevenLabs

How to use ElevenLabs v3

Write the script, place audio tags inline, pick a voice, and generate.

soundcreation
Step 1

Pick the node

Add a Sound Creation node for narration and sound effects, or Dialogue Creation for a two-speaker scene. Both default to the eleven_v3 model.

soundcreation
Step 2

Write and direct the script

Type the script and place audio tags inline where you want direction: [whispers], [excited], [sighs]. Punctuation controls breathing and pauses, so commas and full stops matter. Pick a preset voice, or paste a Voice ID from Voice Creation or Voice Clone.

soundcreation
Step 3

Generate, iterate & export

Fix a seed to reproduce a take, and use Previous text and Next text to keep tone consistent across a series of clips. Chain the audio into a video node, or export it with commercial rights on your plan.

Generated with ElevenLabs v3

ElevenLabs v3
ElevenLabs v3
ElevenLabs v3
ElevenLabs v3
Start Generating!

Pricing and access

One subscription across 30+ AI models, with no per-tool credit balances or separate signups. Credit cost per generation is shown live in the node before you run it.

Free

Try ElevenLabs v3 free

500 weekly credits to test the model and hear the output quality. No credit card required.

  • 500 weekly credits
  • Access to 30+ models
  • Personal use
Join as Free
Pro

Professional access

45,000 monthly credits with cheaper per-credit pricing, commercial rights, and full downloads.

  • 45,000 monthly credits
  • Commercial use license
  • 15% cheaper credits
Join as Pro
Team

Team collaboration

90,000 monthly credits, up to 10 seats, shared workspace, centralized billing, and priority support.

  • 90,000 monthly credits
  • Up to 10 seats
  • Centralized billing
Join as Team
Enterprise

Enterprise scale

210,000 monthly credits, API and integration support, unlimited seats, and dedicated management.

  • API & integration support
  • Unlimited seats
  • Dedicated manager
Join as Enterprise

Technical Specifications

Full specs for ElevenLabs v3 in Phygital+.

Model ID
eleven_v3
Generation
Third-generation text-to-speech
Languages
70+
Dialogue
Native multi-speaker text-to-dialogue
Voice presets
10 in the node, plus custom Voice IDs
Reproducibility
Seed-controlled
Vendor
ElevenLabs
Released
Public alpha June 2025; general availability 2026
Audio tags
Full range of emotion, delivery, and effects
Latency
Higher than Flash and Turbo — not for real-time
Custom voices
Voice Creation (TTV) and Voice Clone (IVC)
Other models in node
Multilingual v2, Turbo v2 / v2.5, Flash v2 / v2.5

ElevenLabs v3 vs Other AI models

How ElevenLabs v3 compares with the other ElevenLabs models available in Phygital+.

Model

ElevenLabs v3

Maximum expressiveness with audio tags and multi-speaker dialogue

Eleven Multilingual v2

Broader consistency for neutral narration, 29 languages

Eleven Flash v2.5

Low-latency generation for real-time and conversational use

Eleven Turbo v2.5

Balanced speed and quality with language enforcement

Best for Performance: emotion, pacing, character Consistent neutral narration Real-time and interactive audio Balanced production work
Languages 70+ 29 70+ 70+
Audio tags Full range of emotion and delivery tags Basic pauses and breaks Basic pauses and breaks Basic pauses and breaks
Latency Higher — built for rendered output, not live calls Higher Lowest Low

Use ElevenLabs v3 Alongside other AI models In Phygital+

Chain scripting, voice design, speech, and video into one repeatable workflow.

ElevenLabs v3

Chain with a text model

Write the script with a text node, then send it straight into speech generation so wording and delivery are decided in one pass.

ElevenLabs v3

Chain with voice creation

Build a voice with Voice Creation or clone one with Voice Clone, then reuse that Voice ID across every v3 generation.

ElevenLabs v3

Chain with a video model

Generate the visuals with a video node and drop the v3 track underneath for narration, dialogue, or sound design.

ElevenLabs v3

Chain with sound effects

Generate ambience and effects in the same node, then layer them under the voice track for a finished scene.

Start Generating!

Everything you need to know about ElevenLabs v3 in Phygital+.

FAQ about ElevenLabs v3

ElevenLabs v3
What is ElevenLabs v3?

Eleven v3 is ElevenLabs' third-generation text-to-speech model. It was released as a public alpha in June 2025 and reached general availability in 2026. Its defining feature is inline audio tags: bracketed cues written directly into the script that direct emotion, pacing, and non-verbal reactions.

What are audio tags?

Audio tags are bracketed instructions you place inside the text itself, such as [whispers], [excited], [sighs], or [laughs]. They tell the model how to deliver the line rather than what to say. Earlier ElevenLabs models handled basic pauses and breaks; v3 responds to a full range of emotion and delivery direction.

How many languages does v3 support?

More than 70, up from 29 in Multilingual v2. A cloned voice keeps its accent character across languages, and the same synthetic voice stays recognisably the same speaker in each one, which matters when you are localising a series rather than a single clip.

Can v3 generate multi-speaker dialogue?

Yes. In Phygital+, the Dialogue Creation node runs the v3 text-to-dialogue workflow: two speakers, separate scripts, a semicolon between phrases, and preset voices or custom Voice IDs per speaker. The model matches prosody across the exchange instead of stitching two independent monologues together.

Is v3 suitable for real-time use?

It trades latency for fidelity. v3 uses a larger model and a higher-fidelity codec, so it takes longer to run and is not intended for live conversational agents. For real-time use, Flash v2.5 or Turbo v2.5 are the right choices, and both are available in the same Sound Creation node.

Which Phygital+ nodes run v3?

In Phygital+ you reach it through three nodes. Sound Creation covers text-to-speech and sound effects. Dialogue Creation covers two-speaker scenes. Voice Creation and Voice Clone build the voice identities that the first two use.

How should I write scripts for v3?

Punctuation carries prosody, so commas and full stops do real work. Add audio tags for direction, and use the Previous text and Next text fields when generating a series of clips so tone and pacing stay consistent across them. Results vary by voice, so it is worth testing tags on the specific voice you plan to ship.

How much does ElevenLabs v3 cost in Phygital+?

Usage is credit-based and included in every plan, from Free up to Enterprise. The credit cost per generation is shown in the node before you run it, so there is no separate ElevenLabs subscription to manage.

Can I use the audio commercially?

Outputs created on any paid plan come with commercial-use rights. Check the model card for provider-specific limits, and make sure you have permission for any voice you clone from a real person.

Explore more AI models In Phygital+

Browse the full catalog of 30+ AI models available in Phygital+.

Dialogue Creation ElevenLabs v3 FLUX Gemini Omni GPT-Image-2 Hailuo 3.0 Ideogram 3 Kling 3.0 Kling Omni Kling Omni Image Krea AI LTX Video Luma Magnific Upscale Midjourney Nano Banana Pro OmniHuman Qwen Image Recraft V4 Reve Runway 4.5 Seedance 2.5 Seedream 5 Sound Creation Topaz Upscale Upscale Video Veo 3.1 Voice Clone Voice Creation WAN 2.5 Wan Video 2.1

Ready to work Faster with ElevenLabs v3?

Join 100+ teams using Phygital+ – every model in one workspace

Try Phygital+ free
bg