← Back to blog
7 min readTabla

Mm-hm, right, I see: why native speakers think you're not listening

You listen with everything you have and still get asked "are you still there?". In English, silence doesn't sound like attention. These six sounds do.

  • conversation
  • common mistakes
  • real life
Mm-hm, right, I see: why native speakers think you're not listening

Your boss is walking you through the new delivery process on a video call. You're listening harder than anyone: every word, every number. Forty seconds in, she stops and asks, "Are you still there? You with me?" You were there. You were more there than ever. But in English, your silence sounded like absence.

The problem isn't your comprehension. It's that you're sending no signals. An English conversation runs on a steady stream of tiny sounds from the listener: mm-hm, uh-huh, right, I see. Linguists call them backchannels, a term Victor Yngve coined in 1970. Spanish has a back channel of its own, but it runs on different pieces at a different rhythm. So when you switch languages, one of two things happens. You go quiet, because all your effort goes into decoding. Or you transplant your "sí, sí" as "yes, yes," which in English sounds like agreement or impatience. The fix is a kit of six sounds and the exact moment to drop them.

In Spanish you respond too, just with whole sentences

A native Spanish listener sends listening signals constantly. In the conversation corpus of Ana María Cestero Mancera, the so-called turnos de apoyo make up 32.3% of all turns, about 2.8 per minute of conversation. Cestero defines them as "brief utterances that recur throughout every conversation to show that the listener is following the utterance in progress". , claro, ya, ¿en serio?.

The difference from English shows up in the shape of the signal. Anne Berry recorded two hour-long dinners, one among four American women and one among four Spanish women. Where the Americans said mhm or uh-huh, the Spaniards said "sí, no, es verdad, sí" or "hombre, sí, sí." Listener responses that repeated or completed what the other person was saying showed up 5 times at the English dinner and 38 times at the Spanish one.

Two parallel streams: the upper one carries long, chained blocks; the lower one, tiny regular pulses

Each group read the other as not listening. In Berry's interviews, a short uh-huh signaled interest to the Americans, but "this same behavior in Spanish implied a lack of interest and was interpreted as 'yeah, okay, hurry up and finish'". To the Americans, the Spaniards' long phrases sounded like they had no interest in listening. Your silence in English is what's left when your Spanish pieces don't transfer and you haven't loaded the English ones yet.

"Yes, yes" doesn't mean "I'm following"

An uh-huh in English doesn't say "yes." It says "keep going, I don't want the turn." Emanuel Schegloff described it in 1982 as a continuer: "'Uh huh', etc. exhibit this understanding, and take this stance, precisely by passing an opportunity to produce a full turn at talk." The sound confirms that the other person can keep talking. It doesn't confirm that you agree or that you understood.

Yes, yes breaks that contract. In Pino Cutrone's synthesis, a "yeah yeah" in a creaky voice with a sharp downstep in pitch "is often construed as a brusque way of telling the interlocutor to stop repeating themselves and get to the point." The same paper reports native-speaker teachers misreading their students' yes and mhm "as displays of understanding, rather than simply polite expressions of attending." A study of 40 Canadians and 40 Chinese participants in simulated medical consultations found that in the mixed-culture pairs, more listening signals went with worse recall of the information: the mm turned into misleading feedback.

Wrong: your teammate in Austin explains at standup why the deploy is slipping and you fire off "yes, yes, yes" after every sentence. He hears impatience and cuts the explanation short. Right: "mm-hm... right... oh, I see," spread across the gaps. He hears that you're with him and gives you the detail you were missing.

The kit: six sounds, each with a job

Six pieces cover almost everything, each with its own function. Following Cutrone's classification and the Cambridge Dictionary's list of response tokens:

  • mm-hm and uh-huh: "keep going." Mouth closed for the first, open for the second, pitch rising on the second syllable. In White's data as cited by Gunnel Tottie, mm-hm alone is 35% of American backchannels, uh-huh another 10%.
  • yeah: the American all-rounder, 42% in that same data. One at a time, never in a burst.
  • right: "got it, that fits." Goes where you'd say claro in Spanish.
  • I see: "now I get it." Only when something has just clicked.
  • really? with rising pitch: "tell me more." It's your ¿en serio?.
  • oh no / wow: the emotional reaction, once per story.

Six distinct geometric tokens lined up on a dark background, each with its own shape and size

The dose depends on which side of the Atlantic you're on. Tottie compared two corpora and found 318 backchannels in 19 minutes of American conversation versus 270 in 53 and a half minutes of British conversation: more than three times as many per minute. In Chicago, with no backchannels, you seem absent. In Manchester, Chicago's dose is too much.

The exact moment: a 110-millisecond drop in pitch

The speaker tells you where to put the sound, with their voice rather than their punctuation. Nigel Ward and Wataru Tsukahara analyzed American English conversations and formulated the rule: when the speaker's pitch falls into its bottom quarter and stays there for at least 110 milliseconds, after at least 700 milliseconds of speech, the listener responds about 700 milliseconds later. In their words, "after the speaker produces a region of low pitch lasting 110 milliseconds the listener tends to produce back-channel feedback." That drop marks the end of an idea, not the end of a sentence.

A waveform that descends gently into a low valley and, an instant later, a small cyan pulse answers from below

The pace is higher than you think. In an analysis of 1,155 American English conversations, 19% of all utterances were backchannels: 37,096 out of 205,000. Cutrone's synthesis puts the native pace at one backchannel every 30–40 words, roughly one every 20 seconds. On the phone that pace carries everything. With no face and no nodding head, the sound is the only proof that you're still on the other end.

What you gain: the other person speaks better when you respond

A responsive listener makes the speaker better. Janet Bavelas, Linda Coates, and Trudy Johnson had 63 pairs of strangers tell a close-call story and distracted half of the listeners with a counting task. Narrators with a distracted listener told the story worse, scoring 2.15 versus 3.10, and their endings "were abrupt or choppy, or they circled around and retold the ending more than once." Attentive listeners gave almost 9 generic responses per minute. The authors' conclusion: "the relative absence of listener responses, particularly specific responses, caused the narration to falter."

The effect works in your favor when it's your turn to talk, too. In a study of 14 Japanese learners of English, the students spoke more fluently when the listener gave them mm-hm and uh-huh than when the listener only nodded, and worse still when the listener gave nothing. You send signals and get clearer explanations back. When you're the one talking, the other person's signals hold you up.

Does an mm-hm come out of you in that half-second gap without thinking, or do you arrive late, with the next sentence already playing? That reflex, the right response inside a time limit, is exactly what Tabla trains: short phrases out loud against the clock until they come out on their own.

Video call. Your manager says, "...so we pushed the release to Thursday, which gives QA two extra days," and his voice drops at the end. What do you do if you're following? And if you disagree?

If you're following, an mm-hm or a right timed to the pitch drop, and he keeps going. If you disagree, don't hum a yeah, yeah: that sounds like "yes, yes, wrap it up." Take the turn with one piece of understanding and a pivot: "Right. One thing, though: QA is already booked Thursday." First you confirm you followed, then you disagree.

The next time someone explains something to you in English, listen to their voice instead of their words, and when it drops, make the sound. It's one syllable, and it changes what the other person believes is happening in your head.

Try this today: play an English podcast for two minutes and say mm-hm or right out loud every time the host's voice drops to close an idea. Count how many you get in: fewer than six means that in a real conversation you're going quiet.