Ask a model to tidy a script and it will quietly round a number, move a date or drop a name. Here every rewrite is checked against the original and thrown away if a fact went missing.
A model has no idea which sounds a voice has barely produced. Coverage is measured against the 44 sounds of British English using a real pronunciation lexicon, not a guess.
Every speech engine takes a different kind of hint. Send the wrong one and it reads the markup out loud as words. The output is shaped for the engine you actually picked.
In today's rapidly evolving landscape, we are delving into the transformative power of Kubernetes. We saw a 47% uplift on 3 March 2026, cutting costs by £12,500 and moving 2.3TB to Cloudflare R2, with Redis in front.
We saw a forty-seven percent uplift on the third of March, twenty twenty-six, cutting costs by twelve thousand five hundred pounds and moving two point three TB to /klˈaʊd flˈeə/ R2, with /ɹˈɛdɪs/ in front.
The tells that make writing sound machine-made are found and removed by rule, not by asking a model nicely. Deterministic, so the same script gives the same answer every time.
What is left goes through a rewrite that is checked afterwards. Every number, sum, percentage, date and proper noun has to survive. A rewrite that loses one is rejected and tried again, never shipped.
Figures, dates, money and percentages become the words a person would actually say. The third of March, twenty twenty-six. Twelve thousand five hundred pounds.
The script is measured against the 44 sounds of British English. Anything the voice will barely have said is reported, with extra lines to record if you are cloning a voice.
Technical and brand terms are looked up and marked so they are said correctly. Redis is reddiss, not re-deez. Cidr is cyder. Your own names are saved and beat the dictionary.
Pauses, tags and chunking are emitted in the form the engine you picked actually accepts, and split under its size limit.
Pick where the script is going and the output changes to match. Pause syntax, pronunciation syntax and the size it is split into all follow the engine, not a house preference. Limits below are each vendor's own documented figures.
| Engine | Pause tag | Pronunciation | Limit |
|---|---|---|---|
| Amazon Polly 6,000 characters per request, of which 3,000 are billed. SSML tags are not billed. |
Yes | SSML, IPA | 6,000 |
| Cartesia Sonic 3 Takes a break tag but no pronunciation tag and no <speak> wrapper. No hard limit is published, so 2,000 is a safe working chunk; the Turbo variant caps at 500. |
Yes | No | 2,000 |
| Deepgram Aura 2 Pronunciations go in as escaped JSON inside the text rather than as a tag. No pause tag, so timing comes from punctuation. Hard limit of 2,000 characters. |
No | Inline JSON, IPA | 2,000 |
| ElevenLabs Flash v2.5 The only engine here that takes ARPAbet in a phoneme tag, and the largest limit of the set. |
Yes | SSML, ARPAbet | 40,000 |
| ElevenLabs Multilingual v2 Takes pauses but no pronunciation tag, so unusual words are listed for you to check by ear. |
Yes | No | 10,000 |
| ElevenLabs v3 No break tag. Pronunciation goes inline between slashes and pauses come from punctuation. |
No | Inline IPA | 5,000 |
| Google Cloud Text-to-Speech 5,000 bytes per request, and the SSML markup counts towards it. |
Yes | SSML, IPA | 5,000 |
| Hume Octave 2 No markup at all. Delivery is steered by a separate plain-English description field, so the script stays clean. 5,000 characters per utterance. |
No | No | 5,000 |
| Inworld TTS Markup support is not confirmed, so plain text is sent. 2,000 is a safe working chunk rather than a published limit. |
No | No | 2,000 |
| LMNT Markup support is not confirmed, so plain text is sent. 2,000 is a safe working chunk rather than a published limit. |
No | No | 2,000 |
| Microsoft Azure AI Speech Azure caps a request at ten minutes of audio rather than a character count. 8,000 is a safe working chunk. |
Yes | SSML, IPA | 8,000 |
| MiniMax Audio Markup support is not confirmed, so plain text is sent. 2,000 is a safe working chunk rather than a published limit. |
No | No | 2,000 |
| Murf Markup support is not confirmed, so plain text is sent. 2,000 is a safe working chunk rather than a published limit. |
No | No | 2,000 |
| OpenAI gpt-4o-mini-tts No markup of any kind. Delivery is steered by the separate instructions field, so the script itself stays clean. |
No | No | 4,096 |
| OpenAI tts-1 and tts-1-hd No markup and no instructions field. Everything has to be carried by the words themselves. |
No | No | 4,096 |
| PlayAI Markup support is not confirmed, so plain text is sent. 2,000 is a safe working chunk rather than a published limit. |
No | No | 2,000 |
| Resemble AI Markup support is not confirmed, so plain text is sent. 2,000 is a safe working chunk rather than a published limit. |
No | No | 2,000 |
| Rime Markup support is not confirmed, so plain text is sent. 2,000 is a safe working chunk rather than a published limit. |
No | No | 2,000 |
| Speechify Markup support is not confirmed, so plain text is sent. 2,000 is a safe working chunk rather than a published limit. |
No | No | 2,000 |
| Synthesia Express-Voice Synthesia has run on its own Express-Voice engine since August 2025, with ElevenLabs voices available on Enterprise. Pronunciation is set in their editor rather than in the script, so plain text is sent. |
No | No | 2,000 |
| Something else, plain text only Safe default for any engine not listed. No markup is emitted at all. |
No | No | 5,000 |
Registration takes an email address and a one-time code. Some domains are let straight through. The rest are read by someone here, usually the same day.