Your first shadowing session probably felt like a malfunction. The audio moved, you tried to follow, you fell half a sentence behind, panicked, and either caught up by skipping words or just gave up and listened until the next gap. Then someone online told you that was fine, that it gets easier, and that you should do it every day. You did it twice more and forgot about it for a month.
The question is whether any of that was doing anything. Shadowing has accumulated a serious body of research over the past three decades, most of it in Japanese EFL and JFL contexts, some of it involving brain imaging. The core mechanisms apply to language learning broadly, and a later section covers what the findings mean specifically for Japanese. The short version: shadowing is good at some things, mediocre at others, and often oversold as a general pronunciation fix. What it actually trains is worth understanding before you spend time on it.
What Is Actually Happening When You Shadow
Shadowing is not just listening with extra steps. It forces simultaneous engagement of systems that normally run in sequence. You are decoding incoming speech, holding it in short-term memory, preparing articulatory output, monitoring your own production against the model, and doing all of this fast enough that none of the steps can wait for the others to finish.
Shuhei Kadota's framework, developed across decades of shadowing research, organizes this into four overlapping effects. The Input effect: repeated exposure at speed automatizes phoneme decoding, the process of recognizing where words begin and end in a stream of connected speech. The Practice effect: the constant recycling of speech through short-term memory strengthens the phonological loop, which Baddeley (1992) identified as the brain's mechanism for holding and rehearsing sound sequences. The Output effect: you are not just hearing speech, you are rehearsing the motor sequences needed to produce it, training the tongue, lips, jaw, and vocal folds to execute target-language patterns at natural speed. The Monitoring effect: you have to continuously compare what you are producing against what you are hearing, which forces a form of real-time self-correction that passive listening never requires.
The IPOM Loop: What Shadowing Trains Simultaneously
Kadota's four-effect framework for why shadowing works
The four effects overlap within every session rather than running in sequence. That simultaneous load is what separates shadowing from passive listening or simple repetition drills.
On a neural level, shadowing engages the dorsal stream most intensively: a left-hemisphere pathway that maps auditory sound representations onto articulatory motor commands. Hickok and Poeppel's dual-stream model of speech processing (2007) identifies this pathway as the brain's mechanism for converting heard speech into reproducible output. Shadowing is essentially a sustained, high-frequency workout for that specific circuit.
Takeuchi et al. (2020) gave this a randomized test. 119 young adults were assigned to four weeks of intensive training: shadowing, reading aloud, listening to compressed speech, or an active control. MRI scans before and after showed that the shadowing group had measurable reductions in gray matter volume and activation in the left cerebellum, a region tied to the phonological loop. In cognitive neuroscience, training-related decreases in brain activation mean efficiency gains, not atrophy: the brain accomplishes the same task with less effort. The speaking-based groups also showed faster reaction times on working memory tasks than the listening-only group.
What Shadowing Reliably Trains (and What It Doesn't)
The most useful framing comes from the largest and most recent systematic review of the literature. Whitworth and Rose (2025) screened six databases and included 44 eligible studies. Their findings are more differentiated than most shadowing promoters acknowledge.
What Shadowing Trains vs. Where Evidence Is Weak
Based on Whitworth & Rose (2025) systematic review, 44 studies
If accent reduction is your primary goal, shadowing alone is not the right tool. The evidence for fluency and rhythm is strong; the evidence for sounding less foreign is not.
Fluency, meaning speaking at an appropriate rate with appropriate pausing, is the skill that responds most directly to shadowing. This makes mechanical sense: fluency is exactly what you train when you must match a model speaker's timing. Foote and McDonough (2017) measured this directly: 16 learners shadowed short dialogues on iPods at least four times per week for eight weeks. Independent listener ratings showed significant fluency gains (p < .006, d = .35) and comprehensibility gains (p < .002, d = .25). Accentedness did not change significantly. The pattern repeated in a larger quasi-experimental study, where the shadowing group's mean fluency score rose from 54 to 83 while a control group rose from 58 to 64 (F = 22.456, p < .001, partial η² = .290).
Prosody is a more specific story. The 11 studies in Whitworth and Rose's review that measured suprasegmental features (rhythm, intonation, pitch control, the production of weak forms of function words) were generally positive. The effect is on the melodic architecture of speech rather than on individual sound accuracy. Shadowing trains you to match the flow, the stress, the reduction of unstressed syllables. It does not reliably sharpen individual phonemes.
Who Benefits Most, and When It Stops Working
The proficiency-dependent pattern in listening comprehension is one of the clearest findings in the shadowing literature, and it matters practically. Phoneme perception tends to improve across levels, but comprehension test scores show the largest gains in lower-proficiency learners. Higher-proficiency learners have already automatized the basic phoneme decoding that shadowing accelerates and have less room to improve on that dimension.
This pattern has been replicated enough across studies to be considered a defining feature of the technique rather than a quirk of individual results. It does not mean shadowing is useless for intermediate and advanced learners. It means their gains show up in different places: fluency, prosodic control, and top-down processing of content rather than bottom-up recognition of phonemes.
Practical tip: cap repetitions per passage
Shiki et al. (2010) found that reproduction accuracy plateaus after four to five repetitions of the same material. The practical ceiling for a single passage is around six to eight attempts before fatigue sets in and gains stop accruing. Rotating to new material is more effective than grinding the same passage. Shorter sessions with more material variety outperform long sessions on a single text.
The Psychological Side: Motivation, Anxiety, and the Speed Problem
Shadowing has an uncomfortable relationship with anxiety. Survey research on learner attitudes, grounded in self-determination theory, consistently finds that the majority of students rate shadowing as effective for both listening and speaking. But the motivational picture splits along proficiency lines.
High-performing students were driven by the challenge of the task and by engagement with content. Lower-performing students were motivated more by visible week-to-week gains in basic accuracy. These are different psychological mechanisms, and they respond differently to feedback design. A program that works for one group can frustrate the other.
Several studies note that learners experience something that researchers call positive anxiety during shadowing: an alert, engaged tension quite different from the debilitating foreign language anxiety documented in oral classroom tasks. That moderate arousal may actually support retention. But the same researchers flag two reliable sources of negative affect: rising audio speed in progressive shadowing programs, and accuracy-based scoring that makes even small errors salient to perfectionistic learners. The technique's psychological profile depends heavily on implementation.
Best for:
Lower-to-intermediate learners focused on fluency and listening comprehension. Also useful for intermediate and advanced learners working specifically on prosody and rhythm. Not the right primary tool if your goal is reducing your accent or memorizing vocabulary.
Shadowing works best as a private, re-recordable activity, not a live classroom performance. Letting learners re-record until satisfied significantly reduces test anxiety without sacrificing the core training effect.
Shadowing and Japanese
The majority of the research cited in this article was conducted in Japanese EFL and JFL contexts: Japanese university students learning English, and learners of Japanese as a foreign language. Kadota's IPOM framework was developed specifically within Japanese SLA research. Tamai's foundational 2002 study, which established the lower-proficiency advantage for listening comprehension, was conducted with Japanese EFL learners. Hamada (2016) replicated and extended those findings with 43 Japanese university EFL students across nine sessions, confirming that phoneme perception improved across proficiency levels while comprehension test gains appeared primarily in lower-proficiency learners. And the motivational survey discussed above was conducted by Sumiyoshi and Svetanant (2017) with 36 students in an advanced JFL course, where over 80% rated shadowing as effective and the proficiency-based motivational split was especially pronounced.
This concentration in Japanese contexts is not accidental. Japanese phonology presents a specific set of challenges that shadowing is positioned to address directly. Japanese is mora-timed rather than stress-timed, meaning each mora takes roughly equal duration to produce. English is built around patterns of strong and weak syllables where function words routinely reduce to near-nothing in connected speech. Learners who have internalized Japanese rhythm often produce English with uniform syllable weight, affecting both fluency scores and perceived comprehensibility. Shadowing at native speed forces exposure to those reduction patterns in a way that reading transcripts cannot — the mechanism behind the comprehensibility gains Foote and McDonough (2017) documented.
For learners of Japanese as a foreign language, the challenges run in the opposite direction. Pitch accent — Japanese uses pitch contour to mark word meaning and grammatical function — is one of the most commonly cited production difficulties for English-speaking learners. The suprasegmental research in Whitworth and Rose's 2025 review covers pitch control broadly with generally positive results, but studies targeting pitch accent specifically in JFL contexts remain sparse. The working consensus: shadowing builds prosodic awareness before it produces reliable production accuracy. Repeated imitation develops sensitivity to pitch contour, which is the necessary step before consistent production becomes possible.
For Japanese learners specifically
Japanese phonotactics — open CV syllable structure, rare consonant clusters — can make Japanese audio feel imitable earlier than English would, which helps beginners get started. The risk is that pitch accent errors embed before perceptual accuracy has developed enough to catch them. Use material from speakers with clearly modeled pitch contours, and where visual pitch accent annotations are available alongside the audio, use them for review rather than treating shadowing as the only input source.
What the Evidence Base Is Still Missing
The overall pattern in the literature is positive. But there are methodological constraints worth knowing before treating any of these findings as settled.
Most shadowing studies use controlled, scripted speaking tasks rather than spontaneous speech, which limits how much the results tell you about real conversation. The literature is heavily concentrated in Japanese EFL and JFL contexts, with a relatively narrow range of first languages and educational settings. The technique may behave differently for learners whose L1 phonology is less distant from Japanese. Few studies include delayed post-tests, so whether gains persist over months rather than weeks is mostly unknown. And questionnaire-based attitude research tends to use researcher-designed instruments of unreported reliability.
None of this overturns the positive findings. It means that claims about shadowing's durability, generalizability, and real-world communicative impact should be held with some caution. The research base is substantial and growing, but it is not yet the airtight case that some shadowing advocates treat it as. If you want to go deeper into the primary studies and methodology, the full academic paper is available to read or download.
References
Baddeley, A. D. (1992). Working memory. Science, 255(5044), 556–559. doi:10.1126/science.1736359
Foote, J. A., & McDonough, K. (2017). Using shadowing with mobile technology to improve L2 pronunciation. Journal of Second Language Pronunciation, 3(1), 34–56. doi:10.1075/jslp.3.1.02foo
Hamada, Y. (2016). Shadowing: Who benefits and how? Uncovering a booming EFL teaching technique for listening comprehension. Language Teaching Research, 20(1), 35–52. doi:10.1177/1362168815597504
Hickok, G., & Poeppel, D. (2007). The cortical organization of speech processing. Nature Reviews Neuroscience, 8(5), 393–402. doi:10.1038/nrn2113
Horwitz, E., Horwitz, M., & Cope, J. (1986). Foreign language classroom anxiety. The Modern Language Journal, 70(2), 125–132. doi:10.2307/327317
Kadota, S. (2019). Shadowing as a practice in second language acquisition: Connecting inputs and outputs. Routledge. doi:10.4324/9781351049108 · Routledge page
Shiki, O., Mori, Y., Kadota, S., & Yoshida, S. (2010). Exploring differences between shadowing and repeating practices. Annual Review of English Language Education in Japan, 21, 81–90. CiNii record
Sumiyoshi, H., & Svetanant, C. (2017). Motivation and attitude towards shadowing. Asian-Pacific Journal of Second and Foreign Language Education, 2, Article 16. doi:10.1186/s40862-017-0039-0
Takeuchi, H., et al. (2020). Effects of training of shadowing and reading aloud of second language on working memory and neural systems. Brain Imaging and Behavior, 15(3), 1253–1269. doi:10.1007/s11682-020-00324-4 · Free full text (PMC)
Whitworth, B., & Rose, H. (2025). A systematic review of research on the use of shadowing for second language pronunciation teaching. Language Teaching, 58(2), 239–269. doi:10.1080/29984475.2025.2546827
Frequently Asked Questions
Does shadowing actually improve speaking fluency?
Will shadowing fix my accent?
How many times should I repeat a passage when shadowing?
Is shadowing better for beginners or advanced learners?
Does shadowing help with vocabulary?
Related Articles
The Ultimate Guide to Speaking Japanese
From pronunciation basics to conversation practice: a practical roadmap for getting words out of your mouth.
Read more →Language LearningExploring the Leitner System
The flashcard method that predates Anki, and what it still does better than the algorithm.
Read more →Language LearningWhy People Quit Duolingo
The specific failure modes that end streaks, and what to do differently.
Read more →Language LearningThe Echo Fluency Cycle
How repeated listening and output reinforce each other, and how to build a practice around that loop.
Read more →