FlashPod Beta waitlist

Research

Why audio-first review can work, and what we still have to prove.

Retrieval practice doesn’t have to stop when you leave the screen. FlashPod asks a question aloud, gives you time to recall the answer, reads it back, and carries on through your scheduled review.

The claim is narrower than “audio beats reading.” For the right material, learning can still happen when watching a screen is impractical. That leads to the question behind FlashPod: what if spaced repetition could fit into more of the day without turning every review into another screen session?

Last reviewed September 2026. These are our own summaries of the studies, and we are not affiliated with the authors. Spotted an error? amanda@flashpodlearn.com

At a glance

What the research supports

  • Retrieval practice improves retention compared with restudying
  • Spacing learning episodes affects long-term retention
  • Listening can support comprehension for suitable material
  • Retrieval questions can improve later retention from audio
  • Silent retrieval can strengthen memory
  • Useful verbal learning can occur during some movement

Reasonable design inferences

  • Short verbal cards suit audio better than dense visual material
  • An eyes-free loop fits moments when a screen is impractical
  • Learners should be able to move between audio and visual as context changes
  • Distributed review may make spaced repetition usable in more of the day

What FlashPod still has to prove

  • Whether audio-first review helps people complete more scheduled reviews
  • How long-term retention compares between audio and visual review
  • Which everyday situations actually work for audio

All six beta questions

How to read the studies below

  • Several involved small samples, medical trainees, or memory tests given minutes after learning.
  • None of them tested FlashPod. Together they support a design; they don’t show that FlashPod works.
  • Audio-first won’t suit every learner or every card. Learners who are deaf or hard of hearing, noisy places, and visual material all call for Visual mode.
  • Use FlashPod only when it’s safe to do so. Some situations need your full attention, and FlashPod isn’t meant for those.

How this connects to the rest of the site. The problem described on the other pages is friction: the study session that never gets started. This page asks a different question. If you take away the chair, does learning survive? For the right material it largely does. Whether taking away the chair gets people to review more is a hypothesis the beta has to test.

Audio-first is different from adding audio.

Most flashcard apps were designed around a visual loop: read the question, reveal the answer, rate the card, move on. Adding text-to-speech makes those cards audible but keeps the same screen-centered interaction in place.

FlashPod starts from the interaction that has to work when the screen isn’t the center of attention. Hands-free, Tap misses and Visual are three ways through the same cards and the same scheduling history. See the three modes

Human-computer interaction research supports the broader distinction. Speech-only interfaces call for a different design from graphical ones,1 and mobile visual attention becomes fragmented in real-world settings such as busy streets and public transport.2 Screens aren’t bad. A system built around continuous screen attention just behaves differently when that attention isn’t available.

See how FlashPod compares with other study tools

Sources: Yankelovich, Levow & Marx, CHI 1995 [1]; Oulasvirta et al., CHI 2005 [2].

The evidence, study by study

Each card gives the takeaway first. Open it for what the study found and where it stops.

A close precedent: adaptive audio flashcards

In a small 2012 study, people walking retained about the same from audio flashcards as from text flashcards.

In 2012, researchers at Microsoft Research Asia and three universities built MemReflex, an adaptive flashcard system for mobile microlearning.3 Twelve participants learned factual material (animal lifespans) while walking, using either audio-only or text-only cued-recall flashcards.

Audio83.8%
Text83.6%

Limits. The authors report no significant difference between the two. The study was small, each task lasted about five minutes, and the memory test came only two minutes later, so it doesn’t establish long-term equivalence.

Another finding may matter as much: when the walking task added visual demands, accuracy on the repetition test dropped more for audio than for text. The researchers argued for switching between modalities as the learner’s context changes. That matches FlashPod’s principle: audio-first, not audio-only. Audio learners also walked about 15% faster than text learners.

Source: Edge, Fitchett, Whitney & Landay, MobileHCI 2012 [3].

Listening can support learning, for the right material

Across 46 studies, listening and reading produced similar comprehension overall. Reading did better on inference questions and when reading was self-paced.

A 2022 meta-analysis combined 46 studies and 4,687 participants comparing reading and listening comprehension.4 Overall comprehension wasn’t reliably different between the two formats, and literal comprehension was essentially equivalent. Reading had an advantage when inferential comprehension was tested and when reading was self-paced, though the self-paced effect was small.

Limits. The comparison is of reading and listening comprehension, not of flashcard review.

Well suited to audio

  • Vocabulary
  • Terminology
  • Definitions
  • Names and dates
  • Concise factual questions
  • Verbal concepts
  • Clearly expressed process steps

Better on screen

  • Equations
  • Diagrams, maps and graphs
  • Dense tables
  • Code and complex notation
  • Anything needing detailed visual comparison

FlashPod keeps Visual review next to audio because the best modality depends on both the card and the context.

Source: Clinton-Lisell, Review of Educational Research, 2022 [4].

Retrieval is the learning mechanism, not listening

FlashPod asks first because trying to recall the answer is what does the work. Spacing decides when to bring it back.

FlashPod doesn’t simply play answers. A 2017 meta-analysis of practice-testing research found that practice tests improved learning compared with restudying and other non-testing conditions.5 A major quantitative review of spacing reached a complementary conclusion for verbal recall tasks: when learning episodes occur matters for later retention.6

A 2013 foreign-vocabulary study compared retrieval practice with imitation, repeating after hearing the answer, and favored retrieval, with tests given right away and after a two-day delay.7 That is the choice an audio review makes on every card: ask first, or just play the answer.

Together these support the loop underneath FlashPod: retrieve, get feedback, bring the material back later.

Limits. These findings aren’t specifically about audio delivery. They support a learning loop that FlashPod applies through headphones or on a screen.

Sources: Adesope, Trevisan & Sundararajan, 2017 [5]; Cepeda et al., 2006 [6]; Kang, Gollan & Pashler, 2013 [7].

Questions turn audio from passive listening into practice

Podcast listeners who got retrieval questions scored higher two to three weeks later.

A randomized trial of 137 emergency-medicine trainees compared two versions of the same educational podcast; one included five retrieval questions.8 Within 48 hours there was no significant overall difference, although the questioned items themselves already scored higher. Two to three weeks later, the group that received questions scored 5.6 percentage points higher overall (the confidence interval was wide) and 10.1 points higher on the material that had actually been questioned.

Limits. This wasn’t a FlashPod study. It tested a podcast with embedded questions, not scheduled flashcards. It does show that adding retrieval questions to audio can change later retention, beyond making listening feel more interactive.

Source: Weinstock et al., Annals of Emergency Medicine, 2020 [8].

You don’t have to answer out loud

Recalling the answer silently still strengthens memory.

FlashPod gives you time to retrieve the answer mentally before you hear it. A 2025 meta-analysis of 18 studies and 2,560 participants found that silent retrieval produced a small but significant learning benefit.9 Producing the answer overtly (writing, typing or speaking) was somewhat more effective, but silent retrieval still helped.

Limits. Silent and overt retrieval aren’t identical. The practical point is narrower: trying to remember in your head still counts as retrieval practice.

Source: Yu et al., Educational Psychology Review, 2025 [9].

Movement has limits, but it doesn’t automatically block learning

In one trial, recall was about the same after listening during exercise as after listening while seated.

FlashPod doesn’t claim that walking improves memory or that multitasking is free. The useful question is whether light or routine movement necessarily prevents useful learning. In a 2024 randomized crossover trial, 96 emergency-medicine residents listened to a 30-minute educational podcast seated and another during 30 minutes of continuous aerobic exercise.10

Recall after podcast listening (initial test within 30 minutes; delayed test at 30 days)
TestExercisingSeated
Immediate recall74.4%76.3%
30-day recall52.3%52.5%

Neither difference was statistically significant.

Limits. The study used passive podcast listening, not flashcard retrieval, so it can’t show that FlashPod review while moving equals seated study. Combined with the MemReflex experiment it supports a narrower conclusion: useful verbal learning can occur during some forms of movement.

Context still matters. Use FlashPod only when it’s safe to do so; some situations demand your full attention, and FlashPod isn’t meant for those. It aims to make the screen optional when the situation allows, not to make every situation right for studying.

Source: Gottlieb et al., Academic Medicine, 2024 [10].

Flashcards may suit audio because they’re short

Spoken information disappears as you hear it, so short, segmented units help.

Written information stays put for you to reread; speech doesn’t. That makes length and segmentation especially important for audio. Research on the transient-information effect suggests that long spoken passages can overload working memory,11 and a meta-analysis of 56 investigations found that dividing multimedia instruction into meaningful segments improved retention and transfer and reduced cognitive load.12

Limits. This research was about multimedia instruction in general, not flashcards or audio specifically, so it would be too strong to call flashcards inherently ideal for audio. The design implication still holds: short question-and-answer units, deliberate pauses, replay and learner control fit spoken information better than long uninterrupted passages.

Sources: Leahy & Sweller, 2011 [11]; Rey et al., Educational Psychology Review, 2019 [12].

The advantage is more usable time, not faster learning

Small retrieval opportunities can be used productively when they fit naturally into the day.

FlashPod doesn’t claim audio makes learning faster; speech can be slower than reading. The opportunity is that scheduled review becomes possible in moments when continuously watching a flashcard app would be inconvenient.

In one two-week field study, 20 people used WaitChatter, a Google Chat extension that showed short vocabulary quizzes while they waited for instant-message replies. In a post-study quiz they could recall an average of 57 new words learned during those brief waiting moments.13

Limits. That study was visual, not audio. It had no control group and measured recall on a quiz afterward, not long-term retention, and it doesn’t prove FlashPod will increase study time. It shows something narrower: small retrieval opportunities work when they fit the day.

Source: Cai et al., CHI 2015 [13].

What FlashPod still has to prove.

Research can support the design. It can’t tell us whether FlashPod itself works. The private beta will test questions like these:

  1. Does audio-first review lead people to complete more of their scheduled reviews?
  2. How does long-term retention compare between FlashPod audio and visual review?
  3. Which real-world situations work well for audio review?
  4. Do Hands-free and Tap misses meaningfully reduce study friction?
  5. Which cards do learners naturally switch to Visual for?
  6. Does moving between audio and visual review within one scheduling history make spaced repetition easier to maintain?

These are product hypotheses, not established outcomes.

How the beta will measure this

Primary outcome
Weekly completion of scheduled reviews, with self-rated recall on repeat reviews as a secondary measure.
Comparison
Each tester’s habits in the month before the beta, and audio versus visual sessions within it.
Duration
Eight weeks, long enough for cards to return through several review intervals.

Results will describe what testers did; they won’t prove that audio caused it.

Not covered on this page

  • Evidence for the scheduling algorithm (FSRS)
  • How synthetic versus recorded voices affect learning
  • Long-term, head-to-head studies of audio versus visual flashcards

Help test the part the research can’t answer for us.

FlashPod’s first private beta is coming to iOS. It will test whether an audio-first review system makes spaced repetition easier to use in everyday life.

Join the iOS beta waitlist

Research sources

View research sources
  1. Yankelovich, N., Levow, G.-A., & Marx, M. (1995). Designing SpeechActs: Issues in speech user interfaces. CHI ’95, 369–376. doi.org/10.1145/223904.223952
  2. Oulasvirta, A., Tamminen, S., Roto, V., & Kuorelahti, J. (2005). Interaction in 4-second bursts: The fragmented nature of attentional resources in mobile HCI. CHI ’05, 919–928. doi.org/10.1145/1054972.1055101
  3. Edge, D., Fitchett, S., Whitney, M., & Landay, J. A. (2012). MemReflex: Adaptive flashcards for mobile microlearning. MobileHCI ’12, 431–440. doi.org/10.1145/2371574.2371641
  4. Clinton-Lisell, V. (2022). Listening ears or reading eyes: A meta-analysis of reading and listening comprehension comparisons. Review of Educational Research, 92(4), 543–582. doi.org/10.3102/00346543211060871
  5. Adesope, O. O., Trevisan, D. A., & Sundararajan, N. (2017). Rethinking the use of tests: A meta-analysis of practice testing. Review of Educational Research, 87(3), 659–701. doi.org/10.3102/0034654316689306
  6. Cepeda, N. J., Pashler, H., Vul, E., Wixted, J. T., & Rohrer, D. (2006). Distributed practice in verbal recall tasks: A review and quantitative synthesis. Psychological Bulletin, 132(3), 354–380. doi.org/10.1037/0033-2909.132.3.354
  7. Kang, S. H. K., Gollan, T. H., & Pashler, H. (2013). Don’t just repeat after me: Retrieval practice is better than imitation for foreign vocabulary learning. Psychonomic Bulletin & Review, 20(6), 1259–1265. doi.org/10.3758/s13423-013-0450-z
  8. Weinstock, M., Pallaci, M., Aluisio, A. R., et al. (2020). Effect of interpolated questions on podcast knowledge acquisition and retention: A double-blind, multicenter, randomized controlled trial. Annals of Emergency Medicine, 76(3), 353–361. doi.org/10.1016/j.annemergmed.2020.01.021
  9. Yu, Y., Zhao, W., Li, A., Shanks, D. R., Hu, X., Luo, L., & Yang, C. (2025). Is covert retrieval an effective learning strategy? Is it as effective as overt retrieval? Answers from a meta-analytic review. Educational Psychology Review, 37(2). doi.org/10.1007/s10648-025-10024-4
  10. Gottlieb, M., Cooney, R., Haas, M. R. C., King, A., Fung, C.-C., & Riddell, J. (2024). A randomized trial assessing the effect of exercise on residents’ podcast knowledge acquisition and retention. Academic Medicine, 99(5), 575–581. doi.org/10.1097/ACM.0000000000005592
  11. Leahy, W., & Sweller, J. (2011). Cognitive load theory, modality of presentation and the transient information effect. Applied Cognitive Psychology, 25(6), 943–951. doi.org/10.1002/acp.1787
  12. Rey, G. D., Beege, M., Nebel, S., Wirzberger, M., Schmitt, T. H., & Schneider, S. (2019). A meta-analysis of the segmenting effect. Educational Psychology Review, 31(2), 389–419. doi.org/10.1007/s10648-018-9456-4
  13. Cai, C. J., Guo, P. J., Glass, J. R., & Miller, R. C. (2015). Wait-learning: Leveraging wait time for second language education. CHI ’15, 3701–3710. doi.org/10.1145/2702123.2702267