Skip to content
Language Learning

Shadowing for Language Learning: What the Research Actually Shows

You listen, you repeat, immediately. Here is whether it works, who it helps most, where it stops, and what the evidence shows specifically for Japanese and English learners.

June 24, 2026| 14 min read
Published June 24, 2026 · My Senpai Team
Shadowing is the act of repeating speech aloud as you hear it: simultaneously, with almost no delay. It started as a training tool for conference interpreters. Researchers in Japan's EFL context adapted it for general language learning in the 1990s. Three decades of research followed. This article synthesizes what that research found for any language learner, with a dedicated section on what the evidence shows specifically for Japanese.

Your first shadowing session probably felt like a malfunction. The audio moved, you tried to follow, you fell half a sentence behind, panicked, and either caught up by skipping words or just gave up and listened until the next gap. Then someone online told you that was fine, that it gets easier, and that you should do it every day. You did it twice more and forgot about it for a month.

The question is whether any of that was doing anything. Shadowing has accumulated a serious body of research over the past three decades, most of it in Japanese EFL and JFL contexts, some of it involving brain imaging. The core mechanisms apply to language learning broadly, and a later section covers what the findings mean specifically for Japanese. The short version: shadowing is good at some things, mediocre at others, and often oversold as a general pronunciation fix. What it actually trains is worth understanding before you spend time on it.

What Is Actually Happening When You Shadow

Shadowing is not just listening with extra steps. It forces simultaneous engagement of systems that normally run in sequence. You are decoding incoming speech, holding it in short-term memory, preparing articulatory output, monitoring your own production against the model, and doing all of this fast enough that none of the steps can wait for the others to finish.

Shuhei Kadota's framework, developed across decades of shadowing research, organizes this into four overlapping effects. The Input effect: repeated exposure at speed automatizes phoneme decoding, the process of recognizing where words begin and end in a stream of connected speech. The Practice effect: the constant recycling of speech through short-term memory strengthens the phonological loop, which Baddeley (1992) identified as the brain's mechanism for holding and rehearsing sound sequences. The Output effect: you are not just hearing speech, you are rehearsing the motor sequences needed to produce it, training the tongue, lips, jaw, and vocal folds to execute target-language patterns at natural speed. The Monitoring effect: you have to continuously compare what you are producing against what you are hearing, which forces a form of real-time self-correction that passive listening never requires.

The IPOM Loop: What Shadowing Trains Simultaneously

Kadota's four-effect framework for why shadowing works

Input EffectAutomatizes phonemedecoding at speedPractice EffectStrengthens thephonological loopOutput EffectRehearses articulatorymotor sequencesMonitoring EffectReal-time comparisonof self vs. modelSHADOWINGall four activesimultaneously

The four effects overlap within every session rather than running in sequence. That simultaneous load is what separates shadowing from passive listening or simple repetition drills.

On a neural level, shadowing engages the dorsal stream most intensively: a left-hemisphere pathway that maps auditory sound representations onto articulatory motor commands. Hickok and Poeppel's dual-stream model of speech processing (2007) identifies this pathway as the brain's mechanism for converting heard speech into reproducible output. Shadowing is essentially a sustained, high-frequency workout for that specific circuit.

Takeuchi et al. (2020) gave this a randomized test. 119 young adults were assigned to four weeks of intensive training: shadowing, reading aloud, listening to compressed speech, or an active control. MRI scans before and after showed that the shadowing group had measurable reductions in gray matter volume and activation in the left cerebellum, a region tied to the phonological loop. In cognitive neuroscience, training-related decreases in brain activation mean efficiency gains, not atrophy: the brain accomplishes the same task with less effort. The speaking-based groups also showed faster reaction times on working memory tasks than the listening-only group.

What Shadowing Reliably Trains (and What It Doesn't)

The most useful framing comes from the largest and most recent systematic review of the literature. Whitworth and Rose (2025) screened six databases and included 44 eligible studies. Their findings are more differentiated than most shadowing promoters acknowledge.

What Shadowing Trains vs. Where Evidence Is Weak

Based on Whitworth & Rose (2025) systematic review, 44 studies

✓ Consistent positive evidence~ Weak or mixed evidenceSpeaking fluencyAll 8 fluency studies: positive resultsGlobal comprehensibilityListener-rated gains: p < .002 (Foote & McDonough)Prosody & rhythmIntonation, weak forms, pitch controlListening comprehensionFor lower-proficiency learners(ceiling effect at higher levels)~Individual phoneme accuracyInconsistent across studiesAccentednessNo significant change (Foote & McDonough)~Vocabulary retentionSecondary effect; weaker than rote practice~Listening comprehensionFor advanced learners(bottom-up decoding already fast)

If accent reduction is your primary goal, shadowing alone is not the right tool. The evidence for fluency and rhythm is strong; the evidence for sounding less foreign is not.

Fluency, meaning speaking at an appropriate rate with appropriate pausing, is the skill that responds most directly to shadowing. This makes mechanical sense: fluency is exactly what you train when you must match a model speaker's timing. Foote and McDonough (2017) measured this directly: 16 learners shadowed short dialogues on iPods at least four times per week for eight weeks. Independent listener ratings showed significant fluency gains (p < .006, d = .35) and comprehensibility gains (p < .002, d = .25). Accentedness did not change significantly. The pattern repeated in a larger quasi-experimental study, where the shadowing group's mean fluency score rose from 54 to 83 while a control group rose from 58 to 64 (F = 22.456, p < .001, partial η² = .290).

Prosody is a more specific story. The 11 studies in Whitworth and Rose's review that measured suprasegmental features (rhythm, intonation, pitch control, the production of weak forms of function words) were generally positive. The effect is on the melodic architecture of speech rather than on individual sound accuracy. Shadowing trains you to match the flow, the stress, the reduction of unstressed syllables. It does not reliably sharpen individual phonemes.

Who Benefits Most, and When It Stops Working

The proficiency-dependent pattern in listening comprehension is one of the clearest findings in the shadowing literature, and it matters practically. Phoneme perception tends to improve across levels, but comprehension test scores show the largest gains in lower-proficiency learners. Higher-proficiency learners have already automatized the basic phoneme decoding that shadowing accelerates and have less room to improve on that dimension.

This pattern has been replicated enough across studies to be considered a defining feature of the technique rather than a quirk of individual results. It does not mean shadowing is useless for intermediate and advanced learners. It means their gains show up in different places: fluency, prosodic control, and top-down processing of content rather than bottom-up recognition of phonemes.

Practical tip: cap repetitions per passage

Shiki et al. (2010) found that reproduction accuracy plateaus after four to five repetitions of the same material. The practical ceiling for a single passage is around six to eight attempts before fatigue sets in and gains stop accruing. Rotating to new material is more effective than grinding the same passage. Shorter sessions with more material variety outperform long sessions on a single text.

The Psychological Side: Motivation, Anxiety, and the Speed Problem

Shadowing has an uncomfortable relationship with anxiety. Survey research on learner attitudes, grounded in self-determination theory, consistently finds that the majority of students rate shadowing as effective for both listening and speaking. But the motivational picture splits along proficiency lines.

High-performing students were driven by the challenge of the task and by engagement with content. Lower-performing students were motivated more by visible week-to-week gains in basic accuracy. These are different psychological mechanisms, and they respond differently to feedback design. A program that works for one group can frustrate the other.

Several studies note that learners experience something that researchers call positive anxiety during shadowing: an alert, engaged tension quite different from the debilitating foreign language anxiety documented in oral classroom tasks. That moderate arousal may actually support retention. But the same researchers flag two reliable sources of negative affect: rising audio speed in progressive shadowing programs, and accuracy-based scoring that makes even small errors salient to perfectionistic learners. The technique's psychological profile depends heavily on implementation.

Best for:

Lower-to-intermediate learners focused on fluency and listening comprehension. Also useful for intermediate and advanced learners working specifically on prosody and rhythm. Not the right primary tool if your goal is reducing your accent or memorizing vocabulary.

Shadowing works best as a private, re-recordable activity, not a live classroom performance. Letting learners re-record until satisfied significantly reduces test anxiety without sacrificing the core training effect.

Shadowing and Japanese

The majority of the research cited in this article was conducted in Japanese EFL and JFL contexts: Japanese university students learning English, and learners of Japanese as a foreign language. Kadota's IPOM framework was developed specifically within Japanese SLA research. Tamai's foundational 2002 study, which established the lower-proficiency advantage for listening comprehension, was conducted with Japanese EFL learners. Hamada (2016) replicated and extended those findings with 43 Japanese university EFL students across nine sessions, confirming that phoneme perception improved across proficiency levels while comprehension test gains appeared primarily in lower-proficiency learners. And the motivational survey discussed above was conducted by Sumiyoshi and Svetanant (2017) with 36 students in an advanced JFL course, where over 80% rated shadowing as effective and the proficiency-based motivational split was especially pronounced.

This concentration in Japanese contexts is not accidental. Japanese phonology presents a specific set of challenges that shadowing is positioned to address directly. Japanese is mora-timed rather than stress-timed, meaning each mora takes roughly equal duration to produce. English is built around patterns of strong and weak syllables where function words routinely reduce to near-nothing in connected speech. Learners who have internalized Japanese rhythm often produce English with uniform syllable weight, affecting both fluency scores and perceived comprehensibility. Shadowing at native speed forces exposure to those reduction patterns in a way that reading transcripts cannot — the mechanism behind the comprehensibility gains Foote and McDonough (2017) documented.

For learners of Japanese as a foreign language, the challenges run in the opposite direction. Pitch accent — Japanese uses pitch contour to mark word meaning and grammatical function — is one of the most commonly cited production difficulties for English-speaking learners. The suprasegmental research in Whitworth and Rose's 2025 review covers pitch control broadly with generally positive results, but studies targeting pitch accent specifically in JFL contexts remain sparse. The working consensus: shadowing builds prosodic awareness before it produces reliable production accuracy. Repeated imitation develops sensitivity to pitch contour, which is the necessary step before consistent production becomes possible.

For Japanese learners specifically

Japanese phonotactics — open CV syllable structure, rare consonant clusters — can make Japanese audio feel imitable earlier than English would, which helps beginners get started. The risk is that pitch accent errors embed before perceptual accuracy has developed enough to catch them. Use material from speakers with clearly modeled pitch contours, and where visual pitch accent annotations are available alongside the audio, use them for review rather than treating shadowing as the only input source.

What the Evidence Base Is Still Missing

The overall pattern in the literature is positive. But there are methodological constraints worth knowing before treating any of these findings as settled.

Most shadowing studies use controlled, scripted speaking tasks rather than spontaneous speech, which limits how much the results tell you about real conversation. The literature is heavily concentrated in Japanese EFL and JFL contexts, with a relatively narrow range of first languages and educational settings. The technique may behave differently for learners whose L1 phonology is less distant from Japanese. Few studies include delayed post-tests, so whether gains persist over months rather than weeks is mostly unknown. And questionnaire-based attitude research tends to use researcher-designed instruments of unreported reliability.

None of this overturns the positive findings. It means that claims about shadowing's durability, generalizability, and real-world communicative impact should be held with some caution. The research base is substantial and growing, but it is not yet the airtight case that some shadowing advocates treat it as. If you want to go deeper into the primary studies and methodology, the full academic paper is available to read or download.

References

Baddeley, A. D. (1992). Working memory. Science, 255(5044), 556–559. doi:10.1126/science.1736359

Foote, J. A., & McDonough, K. (2017). Using shadowing with mobile technology to improve L2 pronunciation. Journal of Second Language Pronunciation, 3(1), 34–56. doi:10.1075/jslp.3.1.02foo

Hamada, Y. (2016). Shadowing: Who benefits and how? Uncovering a booming EFL teaching technique for listening comprehension. Language Teaching Research, 20(1), 35–52. doi:10.1177/1362168815597504

Hickok, G., & Poeppel, D. (2007). The cortical organization of speech processing. Nature Reviews Neuroscience, 8(5), 393–402. doi:10.1038/nrn2113

Horwitz, E., Horwitz, M., & Cope, J. (1986). Foreign language classroom anxiety. The Modern Language Journal, 70(2), 125–132. doi:10.2307/327317

Kadota, S. (2019). Shadowing as a practice in second language acquisition: Connecting inputs and outputs. Routledge. doi:10.4324/9781351049108 · Routledge page

Shiki, O., Mori, Y., Kadota, S., & Yoshida, S. (2010). Exploring differences between shadowing and repeating practices. Annual Review of English Language Education in Japan, 21, 81–90. CiNii record

Sumiyoshi, H., & Svetanant, C. (2017). Motivation and attitude towards shadowing. Asian-Pacific Journal of Second and Foreign Language Education, 2, Article 16. doi:10.1186/s40862-017-0039-0

Takeuchi, H., et al. (2020). Effects of training of shadowing and reading aloud of second language on working memory and neural systems. Brain Imaging and Behavior, 15(3), 1253–1269. doi:10.1007/s11682-020-00324-4 · Free full text (PMC)

Whitworth, B., & Rose, H. (2025). A systematic review of research on the use of shadowing for second language pronunciation teaching. Language Teaching, 58(2), 239–269. doi:10.1080/29984475.2025.2546827

Frequently Asked Questions

Does shadowing actually improve speaking fluency? +
Yes, fluency is the outcome with the most consistent positive evidence. Every study in Whitworth and Rose's 2025 systematic review that measured fluency reported gains. The mechanism makes sense: you are forced to match a speaker's timing, which trains exactly what fluency requires. Effect sizes range from moderate (d = .35 in Foote and McDonough's listener-rated study) to large (η² = .290 in ANCOVA-based classroom comparisons).
Will shadowing fix my accent? +
Probably not in the way you are hoping. Shadowing improves comprehensibility and rhythm, but accentedness (how foreign you sound overall) shows weak and inconsistent improvement across studies. Foote and McDonough (2017) found no significant change in accentedness despite significant gains in fluency and comprehensibility from the same eight-week program. If sounding native is the goal, shadowing alone will not get you there.
How many times should I repeat a passage when shadowing? +
Research by Shiki et al. (2010) found that reproduction accuracy plateaus after four to five repetitions of the same material. Around six to eight repetitions per passage is the practical ceiling before fatigue sets in without further gains. Rotating to new passages is more effective than grinding a single one. Frequent practice with varied material outperforms long sessions on the same text.
Is shadowing better for beginners or advanced learners? +
Both benefit, but differently. Lower-proficiency learners tend to see larger listening comprehension gains because they have more room to improve bottom-up phoneme decoding. Higher-proficiency learners benefit more on fluency, prosody, and top-down content processing. The technique is not equally useful for all goals at all levels, and material difficulty should be matched to learner level.
Does shadowing help with vocabulary? +
It can, but vocabulary is shadowing's weakest outcome area. Gains appear in some studies but are smaller and less consistent than fluency or comprehensibility gains. A direct comparison found that shadowing produced faster initial lexical recall and higher engagement, while conventional repetition practice produced better long-term retention of isolated vocabulary items. Shadowing picks up vocabulary as a side effect, not as its core function.