AI Tools Police
Reader-supported — we may earn a commission from links, at no cost to you. Rankings are never sold. How we investigate →

How to make ElevenLabs speech sound less flat

Add sighs, breaths, pauses and emphasis using ElevenLabs' documented Audio Tags and Voice Settings, and see which popular tricks never appear in the documentation at all.

You will end up with
One exported ElevenLabs audio file where a specific line carries a deliberate pause, emphasis or non-speech sound, such as a sigh, placed with a documented control rather than left to the model's default reading.
Time
About 20 to 30 minutes, most of it listening back and adjusting one line at a time
Cost
Free to learn on the ElevenLabs free plan: about 10,000 credits, roughly 10 minutes of audio a month, no commercial usage rights. Commercial use needs Starter at $6/mo, or $5/mo billed annually. Nothing in this guide requires Creator at $22/mo, which governs voice cloning, not delivery.
By Mucahit KayaAug 29, 2026Last checked Aug 29, 2026~12 min read

ElevenLabs speech usually goes flat for one of three reasons: the script never marks a sigh, a breath or a pause anywhere in the text, the Stability setting is high enough to smooth out the voice's natural variation, or the audio was generated on a model that does not read delivery direction at all. ElevenLabs documents a fix for each: bracket-notation Audio Tags that mark non-speech sounds and emotional delivery directly in the script, which is the documented way to add a sigh, the Voice Settings sliders, and model selection, since only one current model family reads Audio Tags at all. This guide works through all three, separates what ElevenLabs' own documentation confirms from the punctuation tricks and invented tag syntaxes that circulate around this tool, and flags the one real catch up front: switching models to reach Audio Tags also switches off some of those same sliders.

Every documented claim below is checked against ElevenLabs' own documentation, read directly on 2026-08-29, and cross-referenced against the plan details verified in our ElevenLabs review. Where a widely repeated technique could not be traced to ElevenLabs' documentation, it is named as such in its own section rather than folded into the steps, and the specific open questions sit under what we could not confirm at the end, following the same sourcing standard as the rest of this site. The example lines below are scripts to generate yourself; they are not audio produced for this guide or described from listening to it.

Which ElevenLabs plan you need

Every control this guide uses, Audio Tags, the Voice Settings sliders, and model selection, is available on ElevenLabs' free plan. ElevenLabs' own pricing page lists a row for the Eleven v3 model in its plan-comparison table, checked for every tier including Free, so v3 and Audio Tags are gated only by the shared credit pool, not by which plan you are on. The free plan carries roughly 10,000 credits a month, which that same page equates to about 10 minutes of generated audio, enough to work through every step below more than once. Two limits apply once you are past testing. Free-tier output carries no commercial usage rights at all, so a file you actually intend to publish or hand to a client needs the Starter plan, $6/mo, or $5/mo billed annually. And credits are spent by every generation attempt, kept or not, so a 10-minute monthly allowance disappears fast if you regenerate a whole script instead of the one line that sounds wrong, which step 6 covers how to avoid. Nothing here requires Creator at $22/mo. That plan unlocks Professional Voice Cloning, which is about how closely a voice matches a specific person, a different problem from how expressively any voice, cloned or not, delivers a line.

That is not universal in this category. Murf AI documents comparable delivery controls, named Emphasis, Variability and Say It My Way, but gates all three to its Business plan and above; they are not available on Murf's own entry Creator tier at all, per our Murf AI review. If you have not settled on ElevenLabs yet, our AI voice generator comparisons cover the field; everything from here on assumes you are already generating audio in ElevenLabs.

Step 1: switch to the model that reads delivery direction

ElevenLabs runs several model families side by side, and they are not interchangeable for this purpose. Eleven v3 is the model ElevenLabs' own documentation ties to Audio Tags, the bracket-notation directions this guide uses for sighs, breathing, whispering and similar delivery. The older Multilingual v2 and the low-latency Turbo v2.5 and Flash v2.5 models are documented as built for consistency and speed rather than tag-based direction, and are not documented as reading Audio Tags at all.

Switching to Eleven v3 is a trade, not a straightforward upgrade, and ElevenLabs documents the trade directly. ElevenLabs' Voice Settings documentation states, as three separate flat sentences, that Similarity, Speaker Boost and Speed, three of the sliders covered in step 4, are not available for the Eleven v3 model at all. Stability and Style are not addressed either way for v3 on that same page; this guide treats that as an open question rather than guessing, and step 4 covers what it means in practice. None of that changes the recommendation: Audio Tags only exist on v3, and a tag is a more direct fix for a flat line than any slider.

Even on v3, a tag is not guaranteed to land. ElevenLabs' own blog post on the feature states that v3 "will sometimes speak a tag aloud as text instead of interpreting its direction," and names the cause as a mismatch between the selected voice and the requested delivery, its own example being a soft-spoken voice asked to perform [shouts]. That is a voice-and-context problem on the right model, not a wrong-model problem, and step 5 covers picking a voice built for the range a script actually needs.

In the ElevenLabs web app, open the text-to-speech tool and select Eleven v3 from the model menu before you write a single tag. If you are integrating through the API, set model_id to the current Eleven v3 identifier, and confirm it against ElevenLabs' models documentation, since identifiers are versioned and do change between model updates.

Step 2: mark sighs, breathing and other non-speech sounds with Audio Tags

Audio Tags are the documented way to place a sigh, a laugh, a breath or a whisper at a specific point in a line. Write the direction in square brackets immediately before the words it should shape, in plain descriptive language rather than a fixed code, for example: [sighs] I suppose we should get started. ElevenLabs' own prompting guide for Eleven v3 gives dozens of worked examples, covering non-speech sounds such as [sighs], [laughs] and [exhales]; delivery and emotion such as [whispers], [shouts] and [sad]; and pacing markers such as [pause]. ElevenLabs even names a further, explicitly experimental tier it documents as less consistent across different voices, including tags like [sings], [woo] and [fart], each carrying its own instruction to test thoroughly before production use. That is the vendor's own invitation to try tags beyond the core list, not a gap the community filled in; ElevenLabs states directly that its published examples are representative rather than exhaustive.

This is a different feature from ElevenLabs' phoneme tags, which use IPA or CMU Arpabet notation on the older model family to fix how one specific word is pronounced. Audio Tags are about delivery; phoneme tags are about pronunciation, and this guide is about the former.

Two habits keep Audio Tags reliable. Put the tag next to the words it should shape, not gathered at the top of a script; a [sighs] placed at the start of a five-line paragraph will not necessarily still register by line four. And treat every tag as a suggestion the model weighs against the rest of the line rather than a command it always executes exactly as written; a sentence carrying three or four tags tends to read more erratically than the same tags spread across a longer passage, because the model is reconciling several directions in the time it takes to say a handful of words.

Bracket-style direction is not unique to ElevenLabs. LOVO AI's 2026 Pro V2 voices document a similar natural-language cue syntax, such as [sobbing] or [british accent], in place of steering delivery through sliders alone. Our LOVO AI review notes that independent coverage describes the feature without a controlled comparison against LOVO's older voices, so treat bracket-tag direction as a pattern this category is converging on rather than an ElevenLabs-only mechanic, while how well any one vendor's version performs stays a separate, less settled question.

Step 3: control pacing, pauses and emphasis in the text itself

Punctuation still does most of the pacing work, tags or no tags. ElevenLabs' own guidance for Eleven v3 states that punctuation significantly affects delivery: ellipses add pauses and weight, capitalisation increases emphasis, and standard punctuation provides natural speech rhythm. ElevenLabs' own worked example combines several of these in one line: It was a VERY long day [sigh] ... nobody listens anymore. One word capitalised for emphasis, an Audio Tag, and an ellipsis for a trailing pause, all in a single sentence. Capitalising a whole sentence rather than one word is not addressed either way in that documentation; it is genuinely untested territory rather than a documented technique, so stick to the single-word version in ElevenLabs' own example if you want a result you can trace back to something documented.

ElevenLabs documents a plainer pause option too, alongside ellipses: a simple dash (-) or an em dash (), and even a doubled dash such as -- --, for a shorter break. ElevenLabs' own help documentation calls this less consistent than break tags or Audio Tags, so treat it as a light option rather than the first one to reach for.

For a longer, controlled pause on Eleven v3, use the [pause] Audio Tag from step 2; that is the one mechanism this guide can confirm for that model. A line break between sentences is not documented anywhere as a pacing or pausing technique. This was checked directly against ElevenLabs' v3 prompting guide, which does not address line breaks, paragraph breaks or new lines in script text at all, so treat ordinary paragraph structure in your script as a readability choice for you, not an instruction to the model.

If you are not using Eleven v3

Multilingual v2, Flash v2 and Flash v2.5 support a separate, older pause mechanism that Eleven v3 does not: a limited set of SSML, specifically the break tag, written as <break time="1.5s" /> and inserted directly in the text wherever a pause of that length belongs. ElevenLabs documents an explicit ceiling on it, up to 3 seconds, and warns that too many break tags in one script can cause its own problems, speech speeding up or added noise. (Turbo v2.5 and Turbo v2 are not named on the page documenting this feature; ElevenLabs separately states they are functionally equivalent to Flash v2.5 and Flash v2 respectively, which suggests, without directly confirming, that the same break-tag support carries over.)

This is the more precise of the pause mechanisms this guide covers, since it names a duration rather than describing an effect, but ElevenLabs states plainly that Eleven v3 does not support SSML break tags. If step 1 has you on v3, pasting a break tag into the script will not pause it; use the [pause] Audio Tag, an ellipsis, or a dash instead.

Step 4: know which Voice Settings still apply once you are on Eleven v3

Voice Settings is the panel ElevenLabs documents alongside every generation, and step 1 already flagged the catch: switching to Eleven v3 removes some of it. ElevenLabs' Voice Settings documentation states plainly, as three separate sentences, that Similarity, Speaker Boost and Speed are not available for the Eleven v3 model at all. Stability and Style are not addressed either way for v3 on that same page; this guide reports that as an open question rather than guessing an answer. Open your own dashboard with Eleven v3 selected to see which sliders actually appear for your account before planning a workflow around any of them.

SettingDefaultDocumented effectOn Eleven v3
Stability0.5Lower values introduce broader emotional range; higher values can result in a monotonous voice with limited emotionNot stated either way
Similarity0.75How closely the output adheres to the reference voice's tonal signatureDocumented as not available
Style0Amplifies the style of the original speakerNot stated either way
Speaker BoostonBoosts similarity to the original speaker, at a small latency costDocumented as not available

The defaults above come from ElevenLabs' own voice settings API reference. ElevenLabs does not publish an exact minimum or maximum for Stability, Similarity or Style in its written documentation, only these defaults; it does publish an explicit numeric range elsewhere on the same Voice Settings page, 0.7 to 1.2, for the separate Speed setting, which is one more setting not available on v3. The missing range for Stability, Similarity and Style is a real gap in what ElevenLabs has published, not a step this guide skipped.

Stability is the setting most directly tied to sounding flat, and it is the one ElevenLabs describes in exactly those terms: lower values trade consistency for the broader emotional range a flat script is usually missing, and higher values are what ElevenLabs itself associates with a monotonous, limited-emotion read. Style, documented separately, defaults to 0, and raising it carries a stated cost: additional computational resources, added latency, and a documented tendency toward slightly less stability, all as soon as it moves off 0. Raise it in small steps if your dashboard shows it under Eleven v3, rather than maxing it out on a first attempt.

None of this has one correct value, and two of the four settings above will not even appear once Eleven v3 is selected. Where a slider does show up, move it alone, generate the same line again, and compare the two files directly instead of trusting a memory of how the last version sounded. Step 6 covers how to do that without spending a month's credits on one sentence.

Step 5: pick a voice built for range, not narration

Some voices in ElevenLabs' Voice Library are built and tuned for steady, even narration, which is exactly the quality that reads as flat when a script calls for a sigh, a laugh or a sharp change in tone. The Voice Library lets you preview a voice before committing it to a project; previewing a candidate against a line that actually needs range, rather than the calm sample paragraph in its default preview, is the reliable way to find out whether it has any before a script is built around it. If you are cloning a voice rather than picking one from the library, the source recording sets the ceiling: a clone trained on one steady reading performance will not invent range the source never had, whatever the Stability slider is set to afterward. That is a different problem from the one this guide solves; see how Instant and Professional Voice Cloning differ if a source recording is the actual constraint.

Common advice that is not actually documented

A few techniques get repeated constantly around this tool. Two of them are worth naming precisely because they do not appear in ElevenLabs' documentation at all. A third is more subtle: it is not the technique itself that oversteps what is documented, but one specific, elaborate version of it.

Stacking punctuation for intensity, several exclamation marks or question marks in a row instead of one, shows up often in scripts shared around this tool as a way to push emotion higher. This was checked directly, twice, against ElevenLabs' v3 prompting guide: it documents ordinary punctuation and ellipses shaping delivery, and says nothing about repeating a mark for a stronger effect. No stated mechanism explains why it would work, so treat it as unconfirmed rather than a documented control.

Typing a persona or an instruction straight into the text field, something like You are exhausted, speak wearily: followed by the actual line, borrows a habit from prompting a chat model. This was checked directly too: ElevenLabs' documentation has no section describing unbracketed persona instructions as a way to guide delivery. ElevenLabs' text-to-speech field is not an instruction-following prompt outside the bracket-tag syntax from step 2; text placed in that field gets read aloud. Type an instruction sentence without brackets and the checkable result is that the voice says the instruction out loud as part of the line, not that it quietly follows it.

Inventing a tag beyond ElevenLabs' published examples is not folklore; it is closer to the opposite. ElevenLabs' own documentation frames its tag list as representative rather than exhaustive, and names a further, explicitly experimental tier it invites people to test, so a close variation like [tired sigh] or [long sigh] is the vendor's own suggested approach, not a community workaround. What has no documented precedent either way is a long, multi-clause invented stage direction, something like [pauses dramatically for three seconds while looking away]. ElevenLabs' examples, ordinary or experimental, stay short: a word or a short phrase. Nothing in the documentation says a longer invented direction is read differently, and nothing confirms it works the same way a short one does. Keep an invented tag close to the length and shape of ElevenLabs' own examples if you want a result you can trace back to something documented, and treat a much longer invented direction as genuinely untested rather than an established trick.

Pasting a <break time="1s" /> tag into an Eleven v3 script because it worked on an older model is the most common version of applying the right idea to the wrong model. See step 3 for which pause mechanism actually applies to which model.

Step 6: generate one line at a time and isolate what you change

Generation attempts spend credits whether or not you keep the result, and unused credits do not roll over to the next month, so the fastest way to burn a testing budget is regenerating a whole paragraph every time one word sounds off. Isolate the specific line that sounds flat, change one thing, the tag, the punctuation, the dash, or a single slider, and generate just that line on its own. Once it sounds right in isolation, generate the full passage once, in context, since a tag or pause that reads correctly alone can still land differently surrounded by the rest of the script.

Keep a plain record as you go: which model, which tag, punctuation or dash choice, and where Stability and Style sat, if your dashboard showed them, for the take you kept. The History tab in ElevenLabs' web app logs each take, but it will not tell you why one attempt sounded different from another unless you already know what changed between them.

Check it worked

Play back the exported file at the exact point where you placed a tag or a pause, and confirm you hear the thing you asked for: an audible exhale or vocal break for a sigh tag, a real silence roughly the length you expected for a pause or break tag, or a clear stress shift on the one capitalised word. If the bracketed text is read aloud as words instead of performed, first check you actually generated on Eleven v3, since that is the only model ElevenLabs documents as reading Audio Tags at all; if you were already on v3, ElevenLabs documents that this can happen on a voice-and-context mismatch, for example a soft-spoken voice asked to perform [shouts], so try a different voice from step 5 before assuming the tag itself is broken. To compare two settings honestly, generate the same line under each and play them back to back rather than relying on memory of an earlier take, and check the History tab in ElevenLabs' web app to confirm which model and slider values actually produced the file you are judging.

What we could not confirm

  • What a model other than Eleven v3 actually does with a bracket tag such as [sighs]. Six separate ElevenLabs documentation pages were checked directly and none addresses it. This guide does not claim an answer; it recommends Eleven v3 for tag-based direction regardless, per step 1, since that is the only model Audio Tags are documented on at all.
  • A complete, current list of every Audio Tag ElevenLabs' documentation recognises. ElevenLabs' own pages give dozens of worked examples across ordinary and explicitly experimental tiers, but state directly that the list is representative rather than exhaustive, so no complete current list exists to confirm against.
  • Whether the break tag is supported on Turbo v2.5 and Turbo v2 specifically. ElevenLabs' pause documentation names Multilingual v2, Flash v2 and Flash v2.5 explicitly; the Turbo models are not named on that page. A separate ElevenLabs page states the Turbo models are functionally equivalent to their Flash counterparts, which suggests, without directly confirming, that the same break-tag support carries over.
  • Whether stacking extra punctuation, several exclamation or question marks together, measurably changes delivery. This was checked directly, twice, against ElevenLabs' v3 prompting documentation, and neither check found it addressed anywhere, so it is named only as unconfirmed advice in this guide, not included as a step.
  • Whether Stability or Style apply to the Eleven v3 model specifically. ElevenLabs explicitly documents that Similarity, Speaker Boost and Speed do not apply to v3; for Stability and Style, the same page is silent rather than confirming or excluding them. Check your own dashboard with v3 selected to see which sliders actually appear.
  • Whether a line break or paragraph break in the script text affects pacing or creates a pause. This was checked directly against ElevenLabs' v3 prompting documentation, which does not address line breaks, paragraph breaks or new lines in script text at all. This guide does not use line breaks as a pacing technique for that reason.
  • Whether a long, multi-clause invented tag, a full stage direction rather than a short phrase, is read the same way a short invented tag is. ElevenLabs' documented examples, ordinary or experimental, stay short, and nothing in the documentation addresses longer invented directions specifically.