Text-to-speech Sound effects ElevenLabs Commercial use
Voice Creation

Sound Creation
– Text-to-speech
model in Phygital+

Sound Creation turns text into audio two ways. In text-to-speech mode it reads your script in a chosen voice. In sound effects mode it generates a clip from a written description of a sound, from a door closing to rain on a tin roof.

It runs on ElevenLabs with six model versions to pick from, ten voice presets, and support for custom Voice IDs built elsewhere on the canvas. Continuity fields let a series of clips hold the same tone instead of drifting between takes.

ElevenLabs v3

At a glance Sound Creation

Built for

Voice-over for ads, demos, explainers, tutorials, and UI previews, plus short foley and ambience

Content, product, and localisation teams that need audio without booking a studio

Powered by
ElevenLabs
Category
AI audio — text-to-speech and sound effects
Modalities
Text → audio
Modes
Text to speech, sound effects
Default model
eleven_v3
Voice presets
10, plus custom Voice IDs
Effect duration
0.5 to 22 seconds
Access
All Phygital+ plans

Why use Sound Creation

Voice-over and sound effects from one node, with continuity controls and reusable voices.

1

Speech and effects in one node

Narration and sound effects come out of the same node. Switch Mode and the text input goes from a script to a description of a sound, so a voice track and its ambience do not need two separate tools.

2

Six models, one dropdown

eleven_v3 for expressive delivery, Multilingual v2 for steady neutral narration, Flash and Turbo v2.5 when speed and enforced language codes matter. One dropdown covers the trade-off instead of a different product per case.

3

Continuity across a series

Previous text and Next text tell the model what surrounds the current line, so a series of clips keeps its tone and pacing instead of drifting. That matters when a script is split across a dozen generations.

4

Presets or your own voice

Ten voice presets cover most narration and character needs. Beyond those, design a voice from a written brief with Voice Creation or clone one with Voice Clone, then paste the Voice ID here and reuse it indefinitely.

5

Direction inside the text

Punctuation controls breathing and pauses, and bracketed directions like [whispering] shape delivery. It is closer to marking up a script for an actor than to configuring a synthesiser.

Built for content, product, and localisation teams

Sound Creation fits wherever a video, demo, or prototype needs a voice track and booking a studio would take longer than the edit itself.

voicemodel
voicemodel6
voicemodel
voicemodel
voicemodel
voicemodel6
voicemodel
voicemodel

How to use Sound Creation

Pick a mode, write the text, choose a voice, and generate.

soundcreation
Step 1

Choose the mode

Leave Mode on text to speech for narration, or switch it to sound effects to generate a sound from a description. The same text input feeds both.

soundcreation
Step 2

Write the text and pick a voice

Type the script or the sound description. Pick a voice preset, or paste a Voice ID from Voice Creation or Voice Clone. Keep eleven_v3 for expressive delivery, or switch to Flash or Turbo v2.5 when you also need to enforce a language code.

soundcreation
Step 3

Direct, generate & chain

Use punctuation for natural pauses and bracketed directions such as [whispering] for delivery. For a series of clips, fill Previous text and Next text so tone carries across them. Fix a seed to reproduce a take, then chain the audio into a video node or export it.

Generated with Sound Creation

ElevenLabs v3
ElevenLabs v3
ElevenLabs v3
ElevenLabs v3
Start Generating!

Pricing and access

One subscription across 30+ AI models, with no per-tool credit balances or separate signups. Credit cost per generation is shown live in the node before you run it.

Free

Try Sound Creation free

500 weekly credits to test the node and hear the output quality. No credit card required.

  • 500 weekly credits
  • Access to 30+ models
  • Personal use
Join as Free
Pro

Professional access

45,000 monthly credits with cheaper per-credit pricing, commercial rights, and full downloads.

  • 45,000 monthly credits
  • Commercial use license
  • 15% cheaper credits
Join as Pro
Team

Team collaboration

90,000 monthly credits, up to 10 seats, shared workspace, centralized billing, and priority support.

  • 90,000 monthly credits
  • Up to 10 seats
  • Centralized billing
Join as Team
Enterprise

Enterprise scale

210,000 monthly credits, API and integration support, unlimited seats, and dedicated management.

  • API & integration support
  • Unlimited seats
  • Dedicated manager
Join as Enterprise

Technical Specifications

Full specs for Sound Creation in Phygital+.

Powered by
ElevenLabs
Input
Text — script or sound description
Modes
text_to_speech, sound_effects
Default model
eleven_v3
Custom voices
Voice ID from Voice Creation or Voice Clone
Language code
Supported on Flash v2.5 and Turbo v2.5 only
Reproducibility
Seed-controlled
Category
Text-to-speech and sound effects
Output
Single audio track
Models
eleven_v3, Multilingual v2, Turbo v2 / v2.5, Flash v2 / v2.5
Voice presets
10 built in
Duration
0.5 to 22 seconds
Continuity
Previous text and Next text fields

Sound Creation vs Other AI models

How Sound Creation fits alongside the other audio nodes in Phygital+.

Model

Sound Creation

Single-voice narration and text-described sound effects in one node

Dialogue Creation

Two-speaker scenes with per-speaker voices and prosody matching

Voice Creation

Builds new synthetic voices from a written description

Voice Clone

Clones an existing voice from audio samples into a reusable Voice ID

Best for Voice-over, narration, and sound effects Scripted conversations and interviews Inventing a voice that does not exist yet Keeping a real speaker's voice
Input A block of text Two scripts, one per speaker A written voice brief Audio samples
Output A single audio track Rendered dialogue audio Several voice candidates A reusable Voice ID
Models Six ElevenLabs models including v3, Turbo, and Flash v3 or default TTV v3 or multilingual TTV v2 Instant Voice Cloning

Use Sound Creation Alongside other AI models In Phygital+

Chain scripting, voice design, narration, and video into one repeatable workflow.

ElevenLabs v3

Chain with a text model

Write or translate the script in a text node, then feed it straight into Sound Creation so copy and voice-over stay in sync.

ElevenLabs v3

Chain with Voice Creation or Voice Clone

Design a voice from a written brief, or clone one from a recording, then paste the Voice ID into the custom voice field.

ElevenLabs v3

Chain with a video model

Generate the visuals in a video node and lay the narration underneath, or feed the audio in as a rhythm reference.

ElevenLabs v3

Chain speech with effects

Switch the node to sound effects mode for ambience and foley, then layer those takes under the voice track.

Start Generating!

Everything you need to know about Sound Creation in Phygital+.

FAQ about Sound Creation

ElevenLabs v3
What is Sound Creation?

Sound Creation is the Phygital+ node for turning text into audio via ElevenLabs. It has two modes. Text to speech reads your script in a chosen voice. Sound effects turns a written description of a sound into an audio clip. Both run from the same node and the same text input.

Which ElevenLabs models are available?

Six: eleven_v3, Multilingual v2, Turbo v2 and v2.5, and Flash v2 and v2.5. v3 is the default and the most expressive. Multilingual v2 is steadier for neutral narration. Turbo and Flash v2.5 are the fast options and the only ones that respect the Language code field.

How do I generate sound effects?

Switch Mode to sound effects and describe the sound instead of writing dialogue. Specificity does the work: heavy rain hitting a tin roof with thunder gives a far better result than rain. Duration runs from 0.5 to 22 seconds, and 3 to 10 seconds is the useful range for most foley and ambience.

What voices can I use?

The node ships with ten presets covering narration, news delivery, character, educational, and conversational voices. For anything beyond those, build a voice with Voice Creation or clone one with Voice Clone, then paste the resulting Voice ID into the custom voice field.

How do I make the speech sound natural?

Punctuation carries prosody, so commas and full stops control breathing and pauses. Bracketed stage directions such as [whispering] or [excitedly] work on v3, though results vary by voice. When generating a series of clips, fill in Previous text and Next text so the model knows what came before and after and keeps the tone consistent.

How do I control the language?

The Language code field enforces a specific language, but only on Flash v2.5 and Turbo v2.5. On other models the field is ignored. If you need to lock the output language rather than let the model infer it from the text, pick one of those two models.

Can I reproduce the same take twice?

Fix the seed. The same seed with the same text and voice reproduces the same take, and changing it gives you a different read of the identical script. It is the fastest way to audition several deliveries of one line.

How much does Sound Creation cost in Phygital+?

Usage is credit-based and included in every plan, from Free up to Enterprise. The credit cost per generation is shown in the node before you run it, so there is no separate ElevenLabs subscription to manage.

Can I use the audio commercially?

Outputs created on any paid plan come with commercial-use rights. Make sure you have permission for any voice cloned from a real person before publishing.

Explore more AI models In Phygital+

Browse the full catalog of 30+ AI models available in Phygital+.

Dialogue Creation ElevenLabs v3 FLUX Gemini Omni GPT-Image-2 Hailuo 3.0 Ideogram 3 Kling 3.0 Kling Omni Kling Omni Image Krea AI LTX Video Luma Magnific Upscale Midjourney Nano Banana Pro OmniHuman Qwen Image Recraft V4 Reve Runway 4.5 Seedance 2.5 Seedream 5 Sound Creation Topaz Upscale Upscale Video Veo 3.1 Voice Clone Voice Creation WAN 2.5 Wan Video 2.1

Ready to work Faster with Sound Creation?

Join 100+ teams using Phygital+ – every model in one workspace

Try Phygital+ free
bg