Q: Why do people say āumā and āerā when hesitating in their speech?
A: This question can actually be split into two: why do people say anything at all while hesitating and why do they say āerā and āumā instead of other possible sounds?
To answer the first question, linguists known as conversation analysts have observed that people vocalise in a conversation when they think it is their turn to talk, and there are several ways of negotiating the taking of those turns. One of them is the relinquishing of a turn by the current speaker and another speaker taking the floor. Therefore, silence is often construed as a signal that the current speaker is ready to give up his or her turn.
Advertisement
So, if we wish to continue our speaking turn, we often need to fill the silences with a sound to show that we intend to carry on speaking. If we always thought out thoroughly everything we were going to say in a conversation, or memorised our lines perfectly, there would be no hesitation at all.
But, as it is, we do a lot of what is called local management, or improvisation, during conversation for many reasons ā not least because we cannot predict the reactions of our interlocutor. In order to keep the floor while we hesitate, we place dummy words in the empty spaces between our words, much as we might drape our coats on a seat at the cinema to prevent others from taking it.
The second question, as to why āerā and āumā are used instead of say, āeeā or āchooā is not as easy to answer. āEr, in British English, is a transcription of the phonetic schwa sound found in unstressed syllables of English words (such as the vowel sound in the first syllable of āpotatoā).
In traditional phonetics this was called the neutral sound because it is the vowel sound produced when the mouth is not in gear, that is, not tensed to say any of the other formed vowels such as āeā.
The āumā sound is more difficult to explain unless it is just a bad transcription of the same neutral sound with a consonant that closes the mouth in preparation for another real word.
By the way, these sounds are not universal. Many speakers of other languages hesitate in other ways. In Latin languages, for example, the pure sound of the vowel āeā is often used.
A: āErā is used as a conversation filler because it is the most easily pronounced voiced sound for an Anglophone. This is shown easily in a stress-timed language such as English in which all unstressed vowel sounds tend towards the central vowel position.
It is known as the central vowel position because it is pronounced in the centre of the mouth, irrespective of the written vowel, as in āAmericaā, ātrousersā, āferocious, prospectiveā, āpurposeā. āUmā is really only āerā with a closed mouth, as can be shown empirically.
English-speaking pupils learning foreign languages have a tendency to āumā and āerā in a way which is quite foreign to native speakers of the target language. It is also the case that using the correct alternatives gives an impression of fluency greater than that shown by pupils who avoid such utterances, but whose pronunciation is almost flawless.
A: āUmā and āerā are culturally determined. For example, Mandarin Chinese speakers often say āzhege zhege zhegeā (this this this). Some young, hip foreigners learning Mandarin soon āzhege zhege zhegeā with the best of them.
A: People donāt say um and āerā any more. Instead, they say ābasically ā¦ā