The question
A text-to-speech model can drop a letter and tell no one.
So we checked every letter.
Multilingual TTS models are usually scored with one error rate per language. A script, though, is a long tail of letters: a few are common and many are rare. This site accompanies the paper , and lets you inspect every result behind it — per model, language, letter, word position and sentence frame — next to the human speech it is measured against.
Abstract
The probe
One question per cell, one floor per letter
TailProbe asks one question per cell: how often does a model pronounce consonant letter g correctly at word position p (initial, medial or final), in a word that opens the utterance or sits inside a sentence? Each word is spoken in two carrier sentences, the expected sound comes from the espeak-ng grapheme-to-phoneme rules, and a phone recogniser checks whether that sound was produced.
No recogniser is perfect, so “the model is wrong” must mean “worse than a person under the same recogniser”. Every consonant in the FLEURS test and dev recordings — human tokens — is scored with exactly the same pipeline. That gives each letter its own human floor. A cell fails when the whole 95% interval of model − floor lies more than 10 points below zero.
How common it is
Every model has failing letters
of the cells fail: –% of a model's cells, most for MMS-TTS, which was trained on one book per language. Human floors are far from 100%: averaged over a language's cells they range from to % (median %), which is why every number below is relative to the letter's own floor.
One row per model, one column per language, languages grouped by script. Blue means the model is above the human floor, red below it; with “failing cells %” darker means more failing cells. A blank cell is a language the model does not officially support, so it was not tested. Hover a cell for its numbers; click it to open that model and language in the letter explorer below.
Letter explorer
Every cell, every letter
Pick a model and a language. Each row is one consonant letter, each column one word position and sentence frame. A cell shows the model's accuracy and, below it, the human floor; a red outline marks a failing cell. Letters with no input row in that model are tagged. The probe words are listed on the right (screening words first, held-out reserve words after the bar).
Accuracy is the share of tokens whose target consonants were all recognised exactly (strict). Colour follows model − floor (blue above, red below). A dash means that cell had no probe item. Cells with fewer than 20 floor tokens use the letter's floor pooled over positions, as in the paper.
Letters without an input row
In cases a model has no input row for a letter. The tokenizer drops it, so the model never receives it: at most % of its tokens are correct, at least % are heard as deleted, and nothing reports the problem.
The start of an utterance
The first consonant of an utterance is weaker — in models, not in people
The first consonant of a word is less accurate when the word opens the utterance than when the same word follows other words. Under the main recogniser this holds in all models, by to points; in the mixed model it is significant for of models under w2v2p and of under ZIPA. Human speech shows, on average, the opposite: in FLEURS, utterance-initial consonants score to points higher.
Each row is one model, on the languages it supports. The blue pair is the model (w2v2p), the green pair is human FLEURS speech in the same languages. An arrow pointing left means accuracy drops when the word opens the utterance. The right column gives the difference with its 95% interval (w2v2p) and the same difference under ZIPA.
Listen
The same word, at the start and inside a sentence
Each card plays one MMS-TTS clip in which the probe word opens the utterance, then the same word inside a sentence. Below each clip is what the main phone recogniser heard for the target consonant and for the whole utterance. More audio, for every letter that E4 trained, is in Hear E4.
What causes it · rarity
Rarity costs accuracy where training data are scarce
Written-text frequency gives no consistent picture, so we counted every letter in MMS-TTS's own training text (one New Testament per language), under pre-registered analyses. Accuracy rises with the training count under ZIPA ( log-odds per decade, p), but not significantly under w2v2p (, p). The letters seen fewer than 100 times fall points below their floors (w2v2p: ). On held-out reserve words none of the four pre-registered replication tests survives Holm correction (MMS-TTS: under w2v2p, under ZIPA).
Each dot is one letter of one language: its accuracy minus its human floor (vertical) against how often it occurs in MMS-TTS's training text (horizontal, log scale). The shaded band holds letters seen fewer than 100 times. The line joins the pre-registered count groups. Hover a dot for the letter.
What causes it · controlled finetuning
A new letter is learned well only when its row gets a strong training signal
Released models mix frequency with everything else, so experiment E2 sets it directly. Twelve English letters each get a second spelling, a code point whose sound the model already knows, which appears in exactly 0, 10, 40 or 160 utterances of a fixed training set. With its own learning rate, a new symbol closes % of its gap to the plain letter at 160 occurrences in F5-TTS and % in VITS (E2c), where the gap falls from to points. Under the standard recipe the same 160 occurrences close only % and %.
Zero (dotted) is the plain letter for the same sound: a new symbol that reaches it is pronounced as well as the letter the model already knows. Thick lines are arm means with 95% bands; thin lines are the twelve letters. The dashed line is the model before finetuning.
Can it be repaired
The onset can be repaired without training. The rows could not.
Lead-in-and-cut (W1) speaks a short phrase before the sentence and cuts it from the audio, at least 80 ms before the probe word. Pre-registered and tested in four models on six languages, it raises word-initial accuracy in all four under both rulers; the pooled gain is points (ZIPA: ). A warm-up sentence (W2) gains ; repeating the sentence (W3) does not work in general ().
Each line is one model's change in accuracy with its 95% interval over letters; diamonds are the random-effects estimate pooled over the four models. Right of zero means the remedy helped. For MMS-TTS, W2 repeats W1 exactly: its vocabulary has no sentence-final mark, so the two inputs are the same text.
Repairs of the released rows themselves did not help, in three different ways:
Each dot is one failing letter. Horizontal: accuracy with the released row; vertical: accuracy after the edit chosen on separate reserve words. Dots on the diagonal did not change (grey: the gate kept the released row); above improved, below got worse. Rings are letters with no row, given a new one.
Listen · E4
Hear what training only the rows did to every letter
In E4 (pre-registered), the input rows of failing MMS-TTS letters in languages were trained on 10 or 40 FLEURS-train utterances per letter, with every other weight frozen. Accuracy on held-out reserve words fell from % to %, and of the letters got worse. Here you can listen to all of them: every letter and every reserve word, spoken by the released model and by the two trained versions.
Each row is one held-out reserve word, in the sentence shown. The three players are the same item and take from the released model and from the models whose rows of the failing letters were trained on 10 and on 40 utterances. Under each player is what the main phone recogniser (w2v2p) heard for the target consonant: green when it matched the expected phone, red when it did not. The readouts above give the letter's accuracy over both takes. All clips are MMS-TTS with its built-in voice.
Reproduce it
All the code, byte for byte
Every script below is the file that produced the paper's results, copied verbatim; its SHA-256 is shown next to it. From the
released per-clip outputs (Data), these commands from Appendix F.3 of the paper rebuild its tables, figures and
numbers into paper/generated/; the mixed models need R with lme4. Synthesis, scoring and training scripts for every
experiment are included with the drivers that ran them on rented GPUs.
Pre-registration
Every plan, hashed before its data
Each analysis plan was frozen before its data, and its SHA-256 recorded with a UTC timestamp in a hash log. The plans, the amendments and the log are published unchanged, together with the code files and the E5′ gate decisions that the log also hashes. Press the button to fetch every published file and recompute its hash in your browser.
One line per log entry, oldest first. A file hashed more than once (an engineering fix or an amendment) appears once per version; the published file must match its latest entry, and earlier entries document the versions before it. The deviation log lists every change and why it was made.
Documents
Data
Every output, ready to download
The probes, synthesis manifests (with seeds), per-clip scores under every ruler, the scored human floor tokens, and every aggregate, packed into compressed archives. Large groups are split into parts; unpack each part into the same folder. The archives unpack into the repository layout the code expects.
# download the archives and SHA256SUMS into one folder, then: sha256sum -c SHA256SUMS # macOS: shasum -a 256 -c SHA256SUMS for f in *.tar.gz; do tar -xzf "$f"; done
Audio hosted here: MMS-TTS clips in its built-in voice, on held-out reserve words (the listening examples and the E4 explorer). Not included: the audio of the other models, five of which clone the voice of a FLEURS speaker from one recording (the paper releases speech in a real speaker's voice only as far as needed to reproduce the scores), the rest of the synthesised audio, model weights (all seven models are public releases, named in Table 4 of the paper), FLEURS audio (public, CC BY 4.0), the scores of the earlier scorer version, superseded when every score was recomputed with the current one (Appendix F.2 of the paper), and the answer keys of the prepared but not yet run listening test, which must stay unseen.
Tables
Every table of the paper
Generated from the same files as the PDF; numbers, captions and table numbers are identical.
Figures
Every figure of the paper
Click a figure to enlarge it with its full caption.
Limits
What this does not show
- The ruler. The main ruler is one phone recogniser whose labels come from espeak-ng. It cannot output some phones and often misses others (aspiration, retroflexion). The human floor, relaxed criteria and a second recogniser (ZIPA) reduce this but do not remove it; a letter whose floor is near zero cannot be judged.
- No listeners yet. A native-listener test for Bengali and Hindi is prepared but not yet run, so no claim rests on listeners.
- The G2P reference. Expected phones come from espeak-ng 1.52, whose rules are imperfect for some languages. Finding target spans by substitution favours words with transparent spelling.
- Frequency proxies. The training data of most released models are unknown; training-text counts are used only where a proxy of the training text is public, and only E2 and E2c set frequency directly.
- Frames and words. The carrier frames were written for this study and have not yet been reviewed by native speakers; probe words are real, frequent words.
- Scope. Each model is tested only in the languages it officially supports. The onset repair was tested in four of the seven models, the controlled finetuning in two architectures.
- Unfinished experiments. Two analysis plans were amended, always before any outcome data were examined. E5′'s gate admitted no language, so its gated test-word result is fixed by construction; E7 was stopped after letter selection. Both are reported as incomplete.