← Lab notebook

The learning methods evidence supports — and the ones it doesn't

Why I researched learning methods before building the app

Rereading and highlighting were the two techniques I used most. A familiar page felt like progress while I was looking at it.

I own and build this site and the AI Maker Lab channel — this is a build log, not an independent review.

The learning app does not exist yet. This article reports research that will shape its design and claims no result for it.

Before choosing features, I wanted an evidence inventory with its conditions and limits.

In short:

  • Practice testing and distributed practice received high utility ratings in the review that anchors this article.
  • Interleaved practice, self-explanation, and elaborative interrogation received moderate utility ratings.
  • Summarization, highlighting, the keyword mnemonic, imagery for text learning, and rereading received low utility ratings.
  • The composite methods below do not share one evidence base, so I treat their components separately.
  • The learning map and mastery view are product ideas informed by this research, not measured results from an app.

Ten learning techniques and their utility ratings

Dunlosky and colleagues assigned ten techniques to three utility tiers in their 2013 review: high for practice testing and distributed practice; moderate for elaborative interrogation, self-explanation, and interleaved practice; and low for summarization, highlighting, the keyword mnemonic, imagery use for text learning, and rereading.

Ten study techniques ranked by evidence: practice testing and distributed practice high utility; interleaving, self-explanation and elaborative interrogation moderate; rereading, highlighting, summarization, keyword mnemonics and imagery low.
These are the review authors' utility ratings across the conditions they considered, not individual guarantees.

The rating framework asked whether benefits generalized across learning conditions, student characteristics, materials, and criterion tasks. It was not a league table of effect sizes.

The two high-rated techniques earned that assessment because the review found benefits across learners of different ages and abilities, many criterion tasks, and educational contexts. Even that synthesis does not establish the same effect for every learner or subject.

Moderate did not mean ineffective, and low did not mean useless. The authors also used those ratings when evidence was limited in breadth or amount, or when benefits did not generalize as widely. The tiers are a decision aid with conditions attached.

Active recall and spaced repetition: what the evidence says

I use active recall in the heading as an editorial label. The studies below use their own terms: practice testing, repeated testing, and distributed practice.

Roediger and Karpicke’s second experiment makes the timing tradeoff concrete. Five minutes after studying prose passages, four study periods produced 83% recall, three study periods plus one test produced 78%, and one study period plus three tests produced 71%.

One week later, the ordering reversed: one study period plus three tests produced 61%, three study periods plus one test produced 56%, and four study periods produced 40%. The first contrast in that delayed ordering was marginal, and the authors said the boundary conditions of the testing effect were not yet known.

Cepeda and colleagues synthesized 839 assessments from 317 experiments in 184 articles, restricted to verbal-memory tasks with recall outcomes. Across 271 massed-versus-spaced comparisons, mean final-test accuracy was 47.3% for spaced study and 36.7% for massed study, while 12 comparisons showed no effect or a negative spacing effect.

The spacing interval associated with maximal retention grew as the intended retention interval grew. The relationship was not a simple rule that more delay is always better, and some interval combinations had sparse evidence.

Murre and Dros replicated the shape of Ebbinghaus’ forgetting curve with one participant and the method of savings, and they reported that the curve was not completely smooth. The figure combines that qualitative decay shape with the spacing principle: review is placed over time, without prescribing one calendar for every learner, material, or goal.

A schematic of retention: it falls after study, and each spaced review raises it again, with the decay shape drawn after a replicated forgetting curve.
Equal-shaped segments make this a schematic, not a forgetting-rate claim.

For the app, that means reviews should be scheduled from evidence and recency rather than from a universal timetable. It also means a quick success immediately after study should not be treated as durable retention.

Interleaving, self-explanation, and elaborative interrogation: conditions and costs

Rohrer and Taylor’s second experiment had 18 undergraduates at one university complete all three sessions on four solid-volume problem types. A week later, mixed practice scored 63% against 20% after blocked practice; the authors noted the small sample.

The cost appeared during practice. The mixed group scored 60%, while the blocked group scored 89%. In that unfamiliar mathematics task, the schedule that produced the stronger delayed result looked worse during the work.

That result does not make interleaving a rule for every subject. The study covered one task, and the publisher made the abstract and figure public while the full text remained paywalled.

Self-explanation and elaborative interrogation stayed at moderate utility because the review authors found the efficacy evidence limited, including inadequate evidence from real educational contexts. These techniques remain plausible tools with narrower support, not automatic upgrades to every lesson.

Rereading, highlighting, and summarization: lower-utility habits

Rereading and highlighting are the techniques students reported using most in the literature surveyed by the review. The review found that they did not consistently improve student performance and recommended other techniques in their place.

Summarization also received a low utility rating. That label carries the same qualification as the other tiers: limited evidence or generalizability is not the same as a finding that a technique never helps.

My own experience is narrower. Rereading and highlighting were the two techniques I used most, and a familiar page felt like progress while I was looking at it. That is my study history, not a measured result about other learners.

The honest product implication is not to forbid reading, notes, or marks on a page. It is to avoid treating exposure and familiarity as sufficient evidence for advancing a topic.

Layered learning and other composite learning methods

A bounded search on 2026-08-23 found no direct peer-reviewed efficacy evaluation of the named three-pass package: core concepts, then important details, then minor details. That search covered exact-phrase queries in PubMed and ERIC plus broader web queries; it cannot establish that no study exists anywhere.

I therefore treat layered learning as a structuring strategy, not as a directly evaluated technique. The three passes give me an order for covering the same topic map, while the package itself remains unmeasured in the bounded search.

Three passes over the same topic map: core concepts first, then important details, then minor details, each pass deepening the same nodes.
Layered passes organize coverage; the diagram is not direct efficacy evidence for the package.

Keshav’s passes cover the general idea, content without details, then virtually re-implementing the paper: make its assumptions, re-create its work, and compare that with the actual paper. This is neither software implementation nor experimental replication; its support is fifteen years of personal practice, not an experiment.

The similarly named layered curriculum is a different model with C, B, and A activity levels. A 2025 meta-analysis reported 41 effect sizes from 13 Turkish postgraduate theses and a random-effects estimate of 0.80 (95% CI 0.616 to 0.988); that result does not evaluate the three-pass package described here.

Adjacent sequencing primitives have their own evidence. A meta-analysis of 135 studies found a small facilitative effect from advance organizers. In one experiment with 164 introductory-psychology undergraduates at one university, placing objectives immediately before each passage produced .55 on the test against .47 without objectives, while placing all objectives up front made no difference.

Expecting to teach is another adjacent idea, but the evidence is mixed. In the first Nestojko experiment, the expectation improved free recall from .13 to .17 of idea units and organized recall more like the source passage; nobody actually taught. The second experiment found no overall advantage, only an interaction by information type and a marginal advantage for main points, with no effect on details.

Deliberate practice has a tighter definition than repetition. In Ericsson and colleagues’ framework, tasks target current weaknesses, provide immediate informative feedback, repeat relevant work, and increase in difficulty when appropriate. The framework says improvement is minimal without adequate feedback.

Its contribution also varies by domain. A 2014 meta-analysis reported that deliberate practice explained 26% of performance variance in games, 21% in music, 18% in sports, 4% in education, and less than 1% in professions; those percentages are variance explained, not causal effect sizes.

Bloom’s mastery-learning procedure split courses into small units, used diagnostic progress tests, prescribed corrective work after nonmastery, and required thorough mastery before dependent work began. It did not prescribe one universal numeric cutoff.

Bloom’s 1984 two-sigma discussion came from four samples in grades 4, 5, and 8, two subjects, and eleven class periods over three weeks. Under those specific conditions, he reported tutoring plus formative tests and correctives at about two control-group standard deviations and mastery learning alone at roughly one.

A 2024 systematic review of preK–12 tutoring field experiments reported a pooled effect of 0.288 standard deviations, with larger effects for teacher or paraprofessional tutors, earlier grades, at least three weekly sessions, and in-school delivery; it did not measure adult self-study or an app.

A meta-analysis of 108 controlled evaluations found positive examination effects for mastery-learning programs, with stronger effects for weaker students and variation by procedure, design, and content. It also reported costs: self-paced versions often reduced college course completion.

How learning maps show what to learn first

A learning map is my name for a graph of topics and prerequisite relationships, with a selected goal highlighting one relevant path. It is a proposed way to answer what to learn next, not a claim that one ordering is objectively correct for every learner.

A prerequisite graph of skills where choosing a goal highlights one viable path of topics to learn first.
A goal marks one proposed path, not the only viable path.

Concept maps represent concepts in boxes joined by labelled relationship lines. Novak and Cañas organize them around a focus question, with general concepts above more specific ones, and stress that a map depends on context rather than expressing one universal subject order.

A 2006 meta-analysis extracted 67 effect sizes from 55 studies with 5,818 participants. Concept-map use was associated with increased knowledge retention, with mean effects ranging from small to large depending on how maps were used and what they were compared with.

Knowledge-space theory models a learner’s state as the subset of a domain’s problems that the learner can solve, rather than as one score. In its formal definition, feasible states closed under arbitrary unions form a knowledge space, and those spaces correspond to a particular kind of AND/OR graph.

One deployed Beginning Algebra structure contained roughly 60,000 knowledge states and billions of feasible paths. Its outer fringe represented material ready to learn next, while its inner fringe represented mastered material worth reviewing after difficulty.

The same authors warned that expert judgement alone does not validate a prerequisite structure. They required learner data and reported predicted-versus-observed response correlations around .7 to .8 for that Beginning Algebra structure.

A computing-curriculum case modelled desired results, topics, and courses as graph nodes with prerequisite edges, then topologically sorted topics into a teaching order. Its graph held 150 desired results, 1,232 topics, and 2,295 prerequisite edges, while 15.7% of topics were not covered by any course.

How to visualize learning progress and mastery

The proposed app view assigns a visible state to each topic and lets old evidence fade rather than presenting mastery as permanent. The state is an estimate derived from performance and time, not a direct observation of knowledge.

The same skill graph with each node shaded by mastery level from unseen to mastered, and one practised node faded to show its evidence ageing.
Five states distinguish decaying evidence from the rest.

The four-stage practitioner ladder uses the labels unconsciously unskilled, consciously unskilled, consciously skilled, and unconsciously skilled. Gordon Training International attributes the stages to Noel Burch, but the source available here is a 2016 institutional retrospective rather than the original document, so I treat the ladder as reflection vocabulary only.

The 1980 Dreyfus report names novice, competence, proficiency, expertise, and mastery, and argues that performers rely less on abstract rules and more on concrete experience as skill develops. Its authors acknowledged that their descriptive evidence might lack the objectivity of controlled experiments, and a later clinical review questioned the model’s fit for clinical problem solving without claiming it fails everywhere.

The revised Bloom taxonomy is a classification framework for objectives on two dimensions. Its cognitive-process categories are Remember, Understand, Apply, Analyze, Evaluate, and Create, and its relaxed hierarchy allows overlap; those categories are not six sequential measurements of a learner.

Knowledge tracing keeps a probability estimate for whether each rule in a tutor model has been learned. The original model uses four rule-specific parameters—prior learning, acquisition, guess, and slip—and explicitly assumes no forgetting.

In one validation across 345 required-exercise goals, expected and actual mean error rates correlated .75 with mean absolute prediction error .07. Refitting on the same group raised the correlation to .90, which the authors described as an upper bound rather than out-of-sample predictive accuracy.

Half-life regression adds decay by modelling recall probability from time since practice and an estimated memory half-life. The half-life is a hypothetical model construct, not an observed quantity.

On 12.9 million Duolingo student-word traces, two leading variants reached recall-prediction mean absolute error .128 against .235 for a Leitner schedule, a reduction of at least 45%. Those are predictive errors on one product’s noisy, range-restricted data, not measured learning outcomes.

In a two-week experiment with 3.3 million students, one variant increased next-day activity, lesson retention, and practice retention by 12.0%, 1.7%, and 9.5%. Those were product-engagement measures from a comparison between two model variants, not learning gains.

What the evidence means for the learning app

The research leaves me with five design principles, not five product results:

  1. Map first. Represent prerequisites and goal-relevant paths, then validate the structure with learner data rather than expert judgement alone.
  2. Learn in layers. Use core concepts, important details, and minor details as an organizing sequence while keeping the package’s unmeasured status visible.
  3. Test to advance. Ask learners to retrieve or apply material, diagnose gaps, provide corrective work, and avoid treating one immediate success as permanent mastery.
  4. Space reviews. Schedule another attempt from retention goals, recency, and observed performance without imposing one universal interval.
  5. Show decay-aware state honestly. Display a per-topic estimate with uncertainty and let it fade as its evidence ages.

The app remains unbuilt, so each principle is a design decision to test rather than a product outcome.

Which low-utility technique will you stop using this week?

Sources

Keep reading