← All posts

Assimilation in connected speech: why "ten bucks" sounds like "tem bucks"

You can know every word in a sentence and still not hear where one ends and the next begins, because English sounds change shape to match their neighbours.

Assimilation is the name for that change. When two words run together, the last sound of the first word and the first sound of the second have to be produced almost simultaneously, and the mouth takes the shortcut: it starts moving toward the second sound before it has finished the first. The result is that ten bucks comes out as /tem bʌks/ — the /n/ has become an /m/. Nobody decided this; it is what happens when a tongue and a pair of lips try to do two jobs at once at speaking speed. Learners who segment speech by listening for the citation form of each word are listening for something that is not there.

Place assimilation: alveolar sounds borrow the next word's position

The most common type in English. The sounds /t/, /d/ and /n/ are made with the tongue tip at the ridge behind the upper teeth. When the next word begins with a sound made at the lips or at the soft palate, those three shift to match it.

Before a lip sound — /p/, /b/, /m/ — they become lip sounds:

Before a back sound — /k/, /ɡ/ — they move to the back of the mouth:

The word and is the most extreme case, because it is already reduced to /ən/ before assimilation gets to it. So bread and butter is /bred əm bʌtə/ and rock and roll is /rɒk əŋ rəʊl/. Note that this only runs forward across a boundary in this direction: the first word adapts to the second, not the reverse.

Yod coalescence: /d/ + /j/ becomes /dʒ/

The second type is the one that hides whole words. When a word ending in /t/, /d/, /s/ or /z/ is followed by one beginning with /j/ — the y sound, which phoneticians call yod — the two fuse into a single new consonant.

There are four fusions to learn: /t/ + /j/ → /tʃ/, /d/ + /j/ → /dʒ/, /s/ + /j/ → /ʃ/, /z/ + /j/ → /ʒ/. This one causes more listening failures than any other, because the result is a sound that appears in neither original word. If you are waiting for a /d/ and then a /j/ in did you, you will hear neither, and the phrase arrives as an unfamiliar blur. Once you know the four pairings, what do you want stops being mysterious.

Voicing assimilation: voiced sounds go quiet next to voiceless ones

The third type changes whether the vocal folds are buzzing. A voiced consonant next to a voiceless one often gives up its voicing, because switching the voice on and off twice in a few hundredths of a second is more work than the mouth is willing to do.

The have in have to is a genuinely different word from the have in I have it — /hæf/ versus /hæv/. This is so established that some dictionaries list it separately. The same holds for used to, where the /z/ of used has hardened into /s/.

The reason all of this matters is not that you need to produce it. Careful, unassimilated speech is perfectly clear, and no listener will mark you down for saying /ten bʌks/. It matters for the other direction. Comprehension of fast speech is mostly a matter of matching what arrives at your ear to something you already have stored, and if the only stored forms are the citation forms, most of a natural sentence will fail to match anything. Assimilation is systematic and short — three types, a handful of rules each — and learning to expect it converts a wall of noise back into words.

Try it on your own speech before you try to hear it in someone else's. Say ten bucks twice, once with a deliberate /n/ and once letting your lips close early, and notice how much less effort the second one takes.

(CTA) Not sure how a phrase is actually pronounced once the words run together? PHONO annotates any English text with per-word IPA from an offline dictionary, and reads the full sentence aloud so you can hear the joins. Download: https://apps.apple.com/us/app/phono-english-phonetics-ipa/id6783243417 · phono.douhouse.com