Skip to content
NeoStudy

Evidence

Why NeoStudy does what it does.

Every scheduling decision in this app, every screen that refuses to show you something, and every default you can change traces back to a finding below — or to an admission that there is no finding, only a design opinion.

Each entry gives the claim, the study behind it with authors and year, the effect size where a real one exists, and the specific behaviour it produced. Sources are cited in full rather than linked: this page argues for taking evidence seriously, so it does not ship URLs it cannot verify.

Findings
40
Sections
10
Sources cited
56
With no evidence
5

Revision 2 · Last reviewed 27 July 2026

How to read the labels

Every finding is tagged with how much weight it can carry. The fourth tag is the one that matters.

Strong evidence
Replicated across many experiments and supported by at least one meta-analysis. We would be surprised to be wrong about the direction.
Moderate evidence
Good primary evidence, and a mechanism that makes sense — but fewer replications, or an effect whose size is less settled than its existence.
Mixed evidence
A real effect with real moderators. It holds in some conditions and shrinks or vanishes in others, and we say which.
Design inspiration
No experimental evidence at all. Borrowed because it is good design or because someone we respect built it — never because it was tested.

On this page

01The engine

Retrieval practice

Recalling something is not how you check that you learned it. It is how you learn it. This is the most robust finding in applied learning research, and it is why every surface in NeoStudy funnels toward an attempt rather than toward a shelf of material.

Practice testing and distributed practice were the only two techniques, out of ten in common use, that earned a rating of high utility.
Dunlosky, Rawson, Marsh, Nathan & Willingham, 2013
Strong evidence

Testing is a memory modifier, not a memory readout.

g ≈ 0.50–0.61two independent meta-analyses of testing against restudy, pooling several hundred effect sizes

Across hundreds of experiments, attempting to retrieve material beats spending the same time restudying it. The act of reconstructing an answer changes the memory it reconstructs — which is why a test is an intervention and not a measurement.

Sources

  • Meta-analysisRowland (2014). The effect of testing versus restudy on retention: A meta-analytic review of the testing effect. Psychological Bulletin.
  • Meta-analysisAdesope, Trevisan & Sundararajan (2017). Rethinking the use of tests: A meta-analysis of practice testing. Review of Educational Research.
  • ReviewDunlosky, Rawson, Marsh, Nathan & Willingham (2013). Improving students’ learning with effective learning techniques. Psychological Science in the Public Interest.

What NeoStudy does

Nothing in NeoStudy can be “reviewed” by looking at it. Every card demands an attempt before the answer exists, and revealing it is a deliberate keystroke rather than a hover or a scroll.

Strong evidence

The advantage shows up at delay — and reverses on an immediate test.

61% vs 40%free recall one week after learning: repeated testing against repeated study (Roediger & Karpicke, 2006)

In the canonical experiment, students who read a passage four times beat students who read it once and were tested three times, five minutes later. One week later the ordering had flipped, and flipped hard. Fluency during study is a systematically poor guide to retention after it.

Sources

  • ExperimentRoediger & Karpicke (2006). Test-enhanced learning: Taking memory tests improves long-term retention. Psychological Science.
  • ExperimentKarpicke & Roediger (2008). The critical importance of retrieval for learning. Science.

What NeoStudy does

Analytics report retention at delay and never in-session accuracy. There is no “cards studied today” number on the review screen, because a number that rises when you study badly is worse than no number.

Strong evidence

Retrieval improves inference and transfer, not just recall.

d ≈ 1.0retrieval practice against concept mapping, short-answer test one week later (Karpicke & Blunt, 2011)

The standard objection is that testing only drills facts. Karpicke and Blunt gave students the same science texts under matched time and compared elaborative concept mapping with a single session of free-recall practice. Retrieval won on verbatim questions, won on inference questions, and won when the outcome was itself a concept map.

Where it stops holding

The comparison is with concept mapping under equal time, not with unlimited elaboration. Retrieval is the better default use of a study hour — not a substitute for understanding the material the first time.

Sources

  • ExperimentKarpicke & Blunt (2011). Retrieval practice produces more learning than elaborative studying with concept mapping. Science.
  • ExperimentButler (2010). Repeated testing produces superior transfer of learning relative to repeated studying. Journal of Experimental Psychology: Learning, Memory, and Cognition.

What NeoStudy does

A formula’s card family is not only symbol-definition clozes. It includes “why does this assumption matter”, “what breaks without it” and “derive the next step” prompts, because those are the ones that transfer.

Strong evidence

Feedback after the attempt is what stops errors sticking.

Testing without feedback still helps, but feedback reliably amplifies the effect and is what keeps a confidently wrong answer from being strengthened by the very act of retrieving it. Delayed feedback is at least as good as immediate feedback, and often better.

Sources

  • ReviewRoediger & Butler (2011). The critical role of retrieval practice in long-term retention. Trends in Cognitive Sciences.
  • ExperimentButler & Roediger (2008). Feedback enhances the positive effects and reduces the negative effects of multiple-choice testing. Memory & Cognition.

What NeoStudy does

The answer appears the instant you rate — always, including after a failure. Rating Again offers a one-key “explain why”, answered from the anchored source span rather than from a model’s recollection, and offers to split the card if the failure looks like an atomicity problem.

02The clock

Spacing & scheduling

Two identical review sessions produce very different retention depending on when the second one happens. Spacing is the rare intervention that is free — it costs no extra study time, only a decision about timing — which makes the scheduler the highest-leverage thing this app does on your behalf.

The optimal gap between two study sessions is roughly 10–20% of the interval over which you need the memory to survive.
Cepeda, Vul, Rohrer, Wixted & Pashler, 2008
Strong evidence

Distributing the same study time beats massing it.

≈ 9 pointsmean percentage-point recall advantage of spaced over massed practice, across 254 studies (Cepeda et al., 2006)

The synthesis covers 254 studies and more than 800 effect sizes, and the direction almost never reverses at delay. The benefit is not from extra effort; it is from the reconstruction that a gap forces and massed repetition removes.

Source

  • Meta-analysisCepeda, Pashler, Vul, Wixted & Rohrer (2006). Distributed practice in verbal recall tasks: A review and quantitative synthesis. Psychological Bulletin.

What NeoStudy does

There is no cram mode and no “study ahead” button that pulls tomorrow’s cards forward for the satisfaction of clearing a queue. The scheduler owns when; you own whether you showed up.

Strong evidence

The best gap scales with how long you need the memory.

10–20%of the retention interval — the optimal inter-study gap, measured across retention intervals from 7 to 350 days (Cepeda et al., 2008)

Cepeda and colleagues mapped a ridgeline: the gap that maximises retention grows with the retention interval, at roughly a tenth to a fifth of it. A fixed “review in three days” rule is therefore wrong for a one-week exam and wrong again for a formula you want in ten years.

Sources

  • ExperimentCepeda, Vul, Rohrer, Wixted & Pashler (2008). Spacing effects in learning: A temporal ridgeline of optimal retention. Psychological Science.
  • ExperimentBahrick (1979). Maintenance of knowledge: Questions about memory we forgot to ask. Journal of Experimental Psychology: General.

What NeoStudy does

Intervals come from a per-card model of your forgetting aimed at a desired retention you choose — 0.90 by default, adjustable from 0.80 to 0.95 — not from a fixed ladder of days.

Moderate evidence

FSRS-6 reaches the same retention with materially fewer reviews.

≈ 20–30%fewer reviews at matched retention; FSRS beat SM-2 on 99.6% of roughly 10,000 benchmarked collections

The open spaced-repetition benchmark replays real Anki review histories through competing schedulers and scores how well each predicts recall. FSRS-6 beat SM-2 on essentially every collection tested, and the practical consequence is fewer reviews for the same target retention.

Where it stops holding

This is a large open benchmark on donated review logs, not a randomised trial with an independent outcome measure. It is strong engineering evidence about fit to review histories. It is not evidence about exam performance, and it does not tell you that 0.90 is the retention target you personally want.

Source

  • Benchmarkopen-spaced-repetition contributors (2023–). srs-benchmark: comparing spaced-repetition algorithms on real review logs.

What NeoStudy does

FSRS-6 through ts-fsrs, a complete record of every review you have ever done from the very first one, and per-user parameter re-optimisation unlocked at 400 reviews — because below that your own history cannot beat the population prior, and pretending otherwise is just noise fitting.

Moderate evidence

Successive relearning: a few spaced re-successes, not one good session.

≈ 3 recallsin the initial session, followed by successful relearning across roughly three further spaced sessions

Rawson and Dunlosky varied how much practice students did within a session and how many later sessions relearned the material to criterion. Reaching about three correct recalls in the first session, then relearning to criterion across a few spaced sessions, is where additional effort stopped paying for itself. It is the efficiency sweet spot rather than the ceiling.

Sources

  • ExperimentRawson & Dunlosky (2011). Optimizing schedules of retrieval practice for durable and efficient learning: How much is enough?. Journal of Experimental Psychology: General.
  • ReviewRawson, Dunlosky & Sciartelli (2013). The power of successive relearning: Improving performance on course exams and long-term retention. Educational Psychology Review.

What NeoStudy does

A new card graduates on in-session mastery, and is badged learned only after three spaced successful recalls. One lucky Good does not buy the badge.

Mixed evidence

Backlog, not forgetting, is what ends spaced-repetition habits.

Returning after a week away to eight hundred due cards is the most commonly reported quit point in every SRS community, and it is why Anki ships a leech mechanic at all. What we have here is a large body of consistent practitioner observation, not a controlled study of attrition.

Where it stops holding

We have no experimental evidence for the specific thresholds below. They are defaults chosen to fail in the direction of less work, and the dashboard reports what they cost you so you can move them.

Sources

  • PractitionerAnki manual and community reports (ongoing). Leeches, due-load management and the returning-user problem.
  • PractitionerWoźniak (ongoing). SuperMemo guru: overload, postponement and the collapse of a schedule.

What NeoStudy does

Due-load smoothing across the coming days; catch-up mode ordered by retrievability, so the most-likely-forgotten card comes first rather than the oldest; new cards throttled to zero while the backlog exceeds twice the daily target; a leech auto-suspended at eight lapses into a rewrite queue; and an explicit away mode.

03The shuffle

Interleaving

Ten problems of the same kind in a row make practice feel productive and make the delayed test harder. Interleaving restores the step that blocked practice quietly performs for you: working out which method the problem actually needs.

Interleaved practice produced 61% correct on a test one month later, against 38% for blocked practice — in a randomised trial in real classrooms.
Rohrer, Dedrick, Hartwig & Cheung, 2020
Strong evidence

Interleaving mathematics problems roughly doubles delayed scores.

61% vs 38%d = 0.83; randomised classroom trial, tested a month after the last practice (Rohrer et al., 2020). An earlier lab study found 63% vs 20% at one week

Rohrer and Taylor first showed it with volume formulas: identical problems, identical total practice, only the order differed. The 2020 replication took it into real classrooms with random assignment and a test a month after the last practice session, and the gap survived.

Sources

  • ExperimentRohrer & Taylor (2007). The shuffling of mathematics problems improves learning. Instructional Science.
  • ExperimentRohrer, Dedrick, Hartwig & Cheung (2020). A randomized controlled trial of interleaved mathematics practice. Journal of Educational Psychology.

What NeoStudy does

The review queue and every exercise set interleave across topics by default. Blocking exists in exactly one place — first-exposure learn mode, where there is nothing yet to discriminate between.

Moderate evidence

The mechanism is discrimination — choosing the method, not executing it.

Blocked practice tells you the answer type before you read the question, so the hardest part of a real problem never gets practised. Interleaving forces the comparison between superficially similar problems that need different treatments, which is the skill that survives to the exam and to the paper.

Sources

  • ReviewRohrer (2012). Interleaving helps students distinguish among similar concepts. Educational Psychology Review.
  • ExperimentKornell & Bjork (2008). Learning concepts and categories: Is spacing the “enemy of induction”?. Psychological Science.

What NeoStudy does

Coding and maths problems never arrive in same-technique blocks, and the prompt never names the technique. Identifying it is part of the task, and the attempt log records whether you picked right before you picked well.

Strong evidence

Learners reliably judge the worse schedule to be the better one.

≈ 80%of participants judged blocking the more helpful schedule, despite interleaving producing better transfer (Kornell & Bjork, 2008)

Kornell and Bjork taught painting styles blocked or interleaved. Interleaving produced better classification of new paintings; roughly four in five participants nevertheless reported that blocking had helped them more. The illusion is not ignorance — it is fluency, correctly perceived and wrongly interpreted.

Source

  • ExperimentKornell & Bjork (2008). Learning concepts and categories: Is spacing the “enemy of induction”?. Psychological Science.

What NeoStudy does

Interleaving is the default, not a checkbox with a warning label. If you think it is hurting you, the calibration chart is the place to settle it — your rating comfort is not evidence and we do not treat it as such.

Mixed evidence

The effect is real, and it is not universal.

g = 0.4259 studies; large for inductive category learning, near zero for some material types (Brunmair & Richter, 2019)

Brunmair and Richter’s meta-analysis found a solid average benefit that hides very large moderation. Interleaving is strongest when categories or methods are confusable and the learner has to tell them apart; for unrelated items — the vocabulary case — it shrinks toward the spacing effect it partly rides on, and in a few materials it disappears.

Source

  • Meta-analysisBrunmair & Richter (2019). Similarity matters: A meta-analysis of interleaved learning and its moderators. Psychological Bulletin.

What NeoStudy does

Interleaving is applied within a topic neighbourhood rather than by shuffling your whole collection, and siblings from one note are buried the same day — otherwise “interleaved” quietly degenerates into four consecutive clozes from one sentence.

04The scaffold

Worked examples & fading

For a novice, studying a worked example beats solving the equivalent problem — same content, less flailing. For someone who already knows the topic, that same worked example makes learning worse. Both halves of that sentence are experimentally established, and together they say something specific: the scaffold has to move.

Instructional support that helps a novice becomes redundant, and then harmful, as expertise grows. The method has to change with the learner, not with the subject.
Kalyuga, Ayres, Chandler & Sweller, 2003
Strong evidence

Worked examples beat equivalent problems — for novices.

Sweller and Cooper’s algebra students who alternated worked examples with problems spent less time acquiring the material and made fewer errors on later problems. Conventional problem solving spends scarce working memory on means-ends search rather than on the schema the practice was supposed to build.

Sources

  • ExperimentSweller & Cooper (1985). The use of worked examples as a substitute for problem solving in learning algebra. Cognition and Instruction.
  • ReviewAtkinson, Derry, Renkl & Wortham (2000). Learning from examples: Instructional principles from the worked examples research. Review of Educational Research.

What NeoStudy does

First exposure to a technique is a worked solution plus a self-explanation prompt, not a blank editor. The blank editor is a rung you earn.

Strong evidence

The same support hurts once you know the topic.

Expertise reversal is one of the better-replicated interactions in instructional research: guidance that produces a large gain for novices produces a null or negative effect for more advanced learners, because processing redundant explanation costs the capacity that solving would have used.

Source

  • ReviewKalyuga, Ayres, Chandler & Sweller (2003). The expertise reversal effect. Educational Psychologist.

What NeoStudy does

Fading is driven by measured per-topic expertise. “Advanced” is never a property of a person: you can be advanced in stochastic calculus and a novice in convex optimisation on the same afternoon, and the ladder has to know that.

Moderate evidence

Backward fading is the evidence-backed transition.

Renkl and Atkinson compared removing worked steps from the end of a solution first against removing them from the beginning, and against simple example-problem pairs. Fading backwards won on transfer: you keep the part you cannot yet do and give up the part you can.

Source

  • ExperimentRenkl & Atkinson (2003). Structuring the transition from example study to problem solving in cognitive skill acquisition: A cognitive load perspective. Educational Psychologist.

What NeoStudy does

The exercise ladder runs worked solution → Parsons → completion → blank, and a completion problem blanks the final steps first. Success moves you up a rung; struggle moves you down one, not to the bottom.

Moderate evidence

Parsons problems are a cheap worked-example analogue for code.

Rearranging correct but scrambled code produced learning equivalent to writing the same program, in substantially less time. The constraint removes syntax flailing while preserving the structural decision — which is the part that carries.

Sources

  • ExperimentEricson, Margulieux & Rick (2017). Solving Parsons problems versus fixing and writing code. Koli Calling.
  • ExperimentEricson, Foley & Rick (2018). Evaluating the efficiency and effectiveness of adaptive Parsons problems. ICER.

What NeoStudy does

Rung two of the coding ladder is a Parsons problem, and solving one is still recorded and rated like any other card — an exercise attempt is a review, under the same forgetting model.

05The ladder

Hints & help-seeking

Tutoring works, and step-level feedback is most of the reason. But an unconditional “show answer” button turns a tutor into a copying machine, and that behaviour is documented well enough to have a name in the literature.

Step-based tutoring systems produce effect sizes close to human tutors — d ≈ 0.76 against no tutoring, compared with d ≈ 0.79 for a human.
VanLehn, 2011
Strong evidence

Step-level help is where tutoring’s effect comes from.

d ≈ 0.76step-based intelligent tutoring against no tutoring; human tutoring d ≈ 0.79 (VanLehn, 2011)

VanLehn’s synthesis compared answer-based systems, step-based systems and human tutors. Granularity of feedback, not the humanity of the tutor, predicted the effect: systems that intervene at the level of individual solution steps land close to human tutoring.

Sources

  • ReviewVanLehn (2011). The relative effectiveness of human tutoring, intelligent tutoring systems, and other tutoring systems. Educational Psychologist.
  • ReviewVanLehn (2006). The behavior of tutoring systems. International Journal of Artificial Intelligence in Education.

What NeoStudy does

Every exercise carries a three-tier ladder — point, teach, bottom-out — rather than a single binary solution reveal. The first tier names the sub-goal you are stuck on and nothing else.

Strong evidence

Help abuse is real, measurable, and predicts worse learning.

Students who click straight through to the bottom-out hint learn less than students who request help after a genuine attempt. Rapid hint-drilling — “gaming the system” — is detectable directly from interaction logs and correlates with poorer post-tests, independent of prior ability.

Sources

  • ReviewAleven, Stahl, Schworm, Fischer & Wallace (2003). Help seeking and help design in interactive learning environments. Review of Educational Research.
  • ExperimentBaker, Corbett, Koedinger & Wagner (2004). Off-task behavior in the cognitive tutor classroom: When students “game the system”. CHI.

What NeoStudy does

Hints unlock only after a real attempt — for code, the tests must have been run at least once — and there is a short delay between tiers. Both are friction, and both are deliberate.

Moderate evidence

Knowing when to ask for help is itself a teachable skill.

Roll and colleagues built a tutor that gave feedback on help-seeking behaviour rather than on the domain step, and improved that behaviour. The implication for a tool is uncomfortable: students do not arrive knowing when a hint is warranted, so the interface has to model it.

Source

  • ExperimentRoll, Aleven, McLaren & Koedinger (2011). Improving students’ help-seeking skills using metacognitive feedback in an intelligent tutoring system. Learning and Instruction.

What NeoStudy does

Before the ladder unlocks, a metacognitive prompt asks which technique you think is closest. Taking a bottom-out hint auto-drafts a study card of the idea you were missing, so the hint becomes an item in the queue rather than an exit from it.

06The struggle

Productive failure

Attempting a problem you have not been taught how to solve feels like wasted time, and by every in-session measure it is: you fail. The transfer test afterwards disagrees — provided the instruction actually arrives.

Problem-solving before instruction beat instruction-first on transfer, g ≈ 0.36 — but the advantage depended on a consolidation phase following the failure.
Sinha & Kapur, 2021
Moderate evidence

Problem-before-instruction beats instruction-first on transfer.

g ≈ 0.36meta-analysis of productive-failure designs against direct-instruction-first comparisons (Sinha & Kapur, 2021)

Kapur’s design has learners generate and compare their own inadequate solutions before the canonical method is taught. They perform worse during the generation phase and better on transfer afterwards, and the meta-analysis across the paradigm confirms the direction.

Sources

  • ExperimentKapur (2008). Productive failure. Cognition and Instruction.
  • Meta-analysisSinha & Kapur (2021). When problem solving followed by instruction works: Evidence for productive failure. Review of Educational Research.

What NeoStudy does

New material is ordered struggle-first. “Show me how” is a second click and never the landing state, and the attempt you made is stored so the eventual explanation can address it specifically.

Moderate evidence

The failure only pays if consolidation follows.

The meta-analysis is explicit that the effect is a property of the whole design, not of the struggle. Generation has to be followed by instruction that contrasts what the learner produced with the canonical solution. Without that phase the paradigm is just failure, and failure on its own teaches nothing.

Where it stops holding

This is the finding most often quoted badly. “Let them struggle” is not the result. “Let them struggle, then explicitly contrast their attempt with the right answer” is.

Sources

  • Meta-analysisSinha & Kapur (2021). When problem solving followed by instruction works: Evidence for productive failure. Review of Educational Research.
  • ExperimentKapur & Bielaczyc (2012). Designing for productive failure. Journal of the Learning Sciences.

What NeoStudy does

Every struggle-first item ends in a worked solution that references your attempt, not a generic one. The attempt is persisted for exactly this reason.

Strong evidence

This is a desirable difficulty, and it is supposed to feel bad.

Bjork’s framing unifies most of this page: conditions that slow acquisition — spacing, interleaving, retrieval, generation — frequently improve retention and transfer, and conditions that speed acquisition frequently do not. The subjective signal points the wrong way by design.

Sources

  • ReviewBjork (1994). Memory and metamemory considerations in the training of human beings. Metacognition: Knowing about Knowing.
  • ReviewBjork & Bjork (2011). Making things hard on yourself, but in a good way: Creating desirable difficulties to enhance learning. Psychology and the Real World.

What NeoStudy does

We do not soften a first attempt to make a session feel better, and no accuracy figure is shown during a struggle-first set. The number that matters arrives later, and it is retention.

07The craft

Prompt writing

The scheduler cannot rescue a bad card. Most of the variance in whether spaced repetition works for a given person is in what the prompts ask — and this is the area where the strongest practical advice rests on practitioner experience plus cognitive load theory rather than on direct trials. That is worth saying out loud on a page like this.

The card face shows the prompt and nothing else. Every counter, badge, source and due date is one keystroke away and none of them is on the card.
NeoStudy, applying Mayer & Moreno, 2003
Strong evidence

Extraneous material on the card face costs retrieval.

Cognitive load theory’s most actionable claim is that working memory spent on presentation is not available for the task. Decorative detail, redundant text and split attention between two places on a screen all measurably reduce learning from the same content.

Sources

  • ReviewMayer & Moreno (2003). Nine ways to reduce cognitive load in multimedia learning. Educational Psychologist.
  • ExperimentSweller (1988). Cognitive load during problem solving: Effects on learning. Cognitive Science.

What NeoStudy does

The review screen shows the prompt, centred, in a serif at a generous size — and nothing else. No streak, no counter, no source, no due date, no deck name. Source, stats and the editor are one keystroke away and never on the face.

Moderate evidence

Minimum information: one prompt, one thing.

Woźniak’s rule is that simple items are easier to schedule, easier to grade and less likely to fail for irrelevant reasons. The mechanism is plausible and well-motivated — competing responses interfere, and a compound item fails whenever any part of it fails — and the rule is the single most consistent piece of advice from people who have run large collections for decades.

Where it stops holding

There is no randomised trial of card atomicity that we can point at. This is strong practitioner consensus with a coherent mechanism, not an experimental result, and we treat it as a default worth linting for rather than as a law.

Sources

  • PractitionerWoźniak (1999). Effective learning: Twenty rules of formulating knowledge. SuperMemo.
  • ReviewMayer & Moreno (2003). Nine ways to reduce cognitive load in multimedia learning. Educational Psychologist.

What NeoStudy does

The formula template explodes one equation into a card family — symbols, term clozes, assumptions, intuition, derivation steps — instead of one card that asks you to reproduce the lot. The editor flags over-long cards and offers a one-click split into cloze siblings.

Moderate evidence

Good prompts are focused, precise, consistent, tractable and effortful.

Matuschak’s five criteria are the most usable articulation of the craft anyone has written down. Precision matters most: a prompt that admits several right answers trains you to accept vagueness, and the scheduler will faithfully space the vagueness for years.

Source

  • PractitionerMatuschak (2020). How to write good prompts: Using spaced repetition to create understanding.

What NeoStudy does

The card-critique action checks a draft against those five criteria and proposes a rewrite. It arrives as a draft you approve — AI output never edits a live card, here or anywhere else in the app.

Strong evidence

Generating the answer beats recognising it.

d ≈ 0.40meta-analysis of generation against reading, across many materials (Bertsch et al., 2007)

The generation effect is one of the oldest reliable results in the memory literature: material you produce is remembered better than material you read, even when the produced and read versions are identical. Recognition formats let you succeed without generating anything.

Sources

  • ExperimentSlamecka & Graf (1978). The generation effect: Delineation of a phenomenon. Journal of Experimental Psychology: Human Learning and Memory.
  • Meta-analysisBertsch, Pesta, Wiscott & McDaniel (2007). The generation effect: A meta-analytic review. Memory & Cognition.

What NeoStudy does

Typed-answer and cloze cards are preferred wherever the material supports checking, and maths answers are compared after normalisation with a numeric tolerance rather than by string equality. Multiple choice is not a card type in NeoStudy.

08The mirror

Metacognition & calibration

Learners are systematically bad at knowing what they know, and the errors point in a consistent direction: fluency during study is mistaken for durability after it. An app that reports how studying felt is an app that reinforces the illusion.

The conditions that produce the fastest gains in performance during learning are often not the conditions that produce durable learning.
Bjork, 1994
Strong evidence

Fluency is not durability, and it is what people measure.

Ease of processing feels like knowing. Rereading a passage makes it fluent, and fluency inflates predicted recall without improving actual recall. The same misattribution explains why massed practice, blocked practice and re-highlighting all feel like the productive option.

Sources

  • ReviewBjork (1994). Memory and metamemory considerations in the training of human beings. Metacognition: Knowing about Knowing.
  • ExperimentKoriat & Bjork (2005). Illusions of competence in monitoring one’s knowledge during study. Journal of Experimental Psychology: Learning, Memory, and Cognition.

What NeoStudy does

The only numbers on the dashboard are outcome numbers: predicted against actual retention, per-topic forgetting curves, forecast load, prompts created per capture. Nothing rewards time spent, and nothing rewards volume.

Strong evidence

Judgements made after a delay are far better calibrated.

Nelson and Dunlosky found that asking people to predict later recall immediately after study produced weak accuracy, while asking after a delay produced dramatically better accuracy. The delayed judgement has to be made from a real retrieval attempt rather than from an item still sitting in mind.

Source

  • ExperimentNelson & Dunlosky (1991). When people’s judgments of learning are extremely accurate at predicting subsequent recall: The “delayed-JOL effect”. Psychological Science.

What NeoStudy does

You rate a card after the retrieval attempt and never before the reveal, and the four ratings describe the retrieval that just happened — not how well you believe you know the material.

Moderate evidence

Overconfidence makes people stop studying too early.

Dunlosky and Rawson tracked how students’ confidence governed when they stopped practising key terms. Overconfident students terminated study sooner and retained less; the metacognitive error translates directly into a scheduling error the learner makes on their own behalf.

Source

  • ExperimentDunlosky & Rawson (2012). Overconfidence produces underachievement: Inaccurate self evaluations undermine students’ learning and retention. Learning and Instruction.

What NeoStudy does

The mastery criterion belongs to the app, not to you. A card is learned after three spaced successes, and there is no “mark as known” button that skips the evidence.

Moderate evidence

A lenient grader shows up as falling retention, if you measure it.

Self-grading and conversational grading are both noisier and more generous than a typed-answer checker, and feedback quality is exactly what determines whether a review helps. Rather than guessing at a correction factor, the honest move is to instrument the difference and let the data settle it.

Where it stops holding

The specific mitigation below is our design, not a research finding. What the literature supports is that grading quality matters and that self-assessment is optimistic; the instrumentation is ours.

Sources

  • ReviewRoediger & Butler (2011). The critical role of retrieval practice in long-term retention. Trends in Cognitive Sciences.
  • ExperimentDunlosky & Rawson (2012). Overconfidence produces underachievement: Inaccurate self evaluations undermine students’ learning and retention. Learning and Instruction.

What NeoStudy does

Agent-conducted reviews are stored with the verbatim answer and rationale, tagged AGENT, excluded from FSRS optimiser training, and plotted against in-app reviews on a calibration chart. If agent grading drifts lenient, retention falls below target and the chart says so before you notice.

09The subtractionsWhat the design refuses

What does not work

A design is defined as much by what it refuses. These are the practices with the weakest support relative to their popularity — and several of them are the default behaviour of the tools NeoStudy is meant to replace.

Highlighting, underlining, rereading and summarisation were all rated low utility. They are also the four things students actually do.
Dunlosky, Rawson, Marsh, Nathan & Willingham, 2013
Strong evidence

Rereading and highlighting are low-utility.

Both were rated low utility in the review that rated practice testing and distributed practice high. They are cheap, popular, and largely ineffective for the time they consume; highlighting can even hurt on inference questions by drawing attention to isolated sentences at the expense of the structure connecting them.

Source

  • ReviewDunlosky, Rawson, Marsh, Nathan & Willingham (2013). Improving students’ learning with effective learning techniques. Psychological Science in the Public Interest.

What NeoStudy does

Capture is a funnel, never a shelf. The Library’s headline metric is prompts created per capture and time-to-first-review; unprocessed captures visibly dim at fourteen days and auto-archive at thirty. Collecting is not progress and the interface will not pretend it is.

Strong evidence

Learning styles: no credible evidence for the meshing hypothesis.

The claim requires a crossover interaction — visual learners doing better with visual instruction and verbal learners doing better with verbal instruction. Pashler and colleagues found that almost no studies used the design capable of detecting it, and the few that did failed to find it. The preference is real; the instructional consequence is not.

Sources

  • ReviewPashler, McDaniel, Rohrer & Bjork (2008). Learning styles: Concepts and evidence. Psychological Science in the Public Interest.
  • ExperimentRogowsky, Calhoun & Tallal (2015). Matching learning style to instructional method: Effects on comprehension. Journal of Educational Psychology.

What NeoStudy does

There is no learner-type setting anywhere in NeoStudy. Everything adaptive keys off measured per-topic performance, which is a fact about the topic and your history with it rather than a claim about your personality.

Strong evidence

Cramming buys the immediate test and sells the delayed one.

Massed practice is genuinely superior when the test is minutes away, which is why it survives: it is reinforced every time it is used. Over any interval that matters for a technical field, the ordering reverses.

Sources

  • Meta-analysisCepeda, Pashler, Vul, Wixted & Rohrer (2006). Distributed practice in verbal recall tasks: A review and quantitative synthesis. Psychological Bulletin.
  • ExperimentRoediger & Karpicke (2006). Test-enhanced learning: Taking memory tests improves long-term retention. Psychological Science.

What NeoStudy does

No cram mode, and no button that pulls tomorrow’s cards forward. If you want more work today, the honest lever is the new-card budget, and it is capped for a reason.

Mixed evidence

Points, streaks and confetti are a weak and short-lived motivator.

g = 0.49and highly heterogeneous: a gamification meta-analysis whose effects are dominated by short studies (Sailer & Homner, 2020)

Gamification’s meta-analytic effect on cognitive learning outcomes is positive but highly heterogeneous and concentrated in short interventions. Meanwhile tangible extrinsic rewards reliably reduce intrinsic motivation for activities people already found interesting — which describes almost exactly the person who installs a spaced-repetition app to study their own field.

Where it stops holding

This is the softest call on the page. The evidence does not say gamification never works — it says the average is unstable and the mechanism cuts against a motivated adult learner. We are making a design judgement and labelling it as one.

Sources

  • Meta-analysisSailer & Homner (2020). The gamification of learning: A meta-analysis. Educational Psychology Review.
  • Meta-analysisDeci, Koestner & Ryan (1999). A meta-analytic review of experiments examining the effects of extrinsic rewards on intrinsic motivation. Psychological Bulletin.

What NeoStudy does

One streak — sessions completed — shown small, and nothing else. No XP, no badges, no confetti, no notification that shames you for a missed day.

10The honest partNo experimental support

Borrowed, not proven

Several ideas in NeoStudy have no experimental support at all. They are here because they are good design, or because someone we respect built something that works — neither of which is evidence. Listing them is not a disclaimer; it is the thing that makes everything above worth reading. Each one ships with the measurement that would tell us it was wrong.

A page that cites research for the parts that have it, and stays quiet about the parts that do not, is marketing. The list below is the difference.
Why this section exists
Design inspiration

Incremental reading is an interaction design, not a result.

SuperMemo’s incremental reading is genuinely original, has a devoted user base, and as far as we can find has never been tested in a controlled study. Its components lean on real findings — spacing, retrieval, extraction — but the composite has never been compared against the obvious alternative of reading a paper and then making cards from it.

Source

  • PractitionerWoźniak (1999–). Incremental reading. SuperMemo.

What NeoStudy does

We borrow only the extract → cloze funnel and none of the queueing model. The claim we are willing to make is measured: prompts created per capture, and time from capture to first review, reported whether or not the numbers flatter the feature.

Design inspiration

Evergreen notes and progressive summarisation are untested.

The writing on evergreen notes and note-linking is thoughtful, influential, and entirely uncontrolled. The closest relevant evidence is Dunlosky’s review, which rates summarisation low-utility. A large, beautifully linked note collection is not known to produce learning, and the aesthetic pleasure of building one is a poor proxy for whether it did.

Sources

  • PractitionerMatuschak (2019–). Evergreen notes.
  • ReviewDunlosky, Rawson, Marsh, Nathan & Willingham (2013). Improving students’ learning with effective learning techniques. Psychological Science in the Public Interest.

What NeoStudy does

Notes exist as raw material for prompts, not as a deliverable. The Library counts prompts rather than notes, and the inbox decays on purpose so that an unprocessed capture is visibly a debt rather than an asset.

Design inspiration

FIRe-style implicit prerequisite credit is a company’s account of its own system.

Succeeding on an advanced problem is decent Bayesian evidence that its prerequisites are intact, and Math Academy report that crediting them keeps review load sane in hierarchical domains. That report is a description of a product by the people who sell it. The nearest published theory — knowledge spaces — validates inferring what a learner knows from an assessment, not silently rescheduling their reviews on that inference.

Sources

  • PractitionerMath Academy (2024). The Math Academy Way.
  • ReviewFalmagne & Doignon (2011). Learning Spaces. Springer.

What NeoStudy does

Implicit credit ships in shadow mode: computed, logged and charted with the damping factor κ set to 0, so it changes nothing at all. It is applied only if and when the calibration chart shows that predicted and actual retention agree for the cards it would have touched.

Design inspiration

Desired retention 0.90 — and every other default — is a guess.

0.90 is a defensible starting point, not a proven optimum for you. Nor is the daily new-card cap of eight, the leech threshold of eight lapses, the fourteen-day dim, the thirty-day archive, or the two-times-target backlog throttle. They are round numbers chosen to fail in the direction of less work.

Source

  • Benchmarkopen-spaced-repetition contributors (2023–). srs-benchmark and the FSRS parameter discussion on desired retention.

What NeoStudy does

Every one of these is exposed in Settings rather than buried in a constant, and the dashboard reports what each is costing you in reviews per day. A default you cannot see and cannot move is a claim you cannot check.

Design inspiration

Lab effect sizes are not a forecast for you.

Almost every number on this page comes from undergraduates learning Swahili word pairs, biology passages or unfamiliar geometry under laboratory conditions, tested days or weeks later. You are a specialist studying your own field over years, with prior knowledge that changes how every one of these mechanisms behaves. The direction of these findings is trustworthy. The magnitudes are not yours.

Source

  • ReviewDunlosky, Rawson, Marsh, Nathan & Willingham (2013). Improving students’ learning with effective learning techniques. Psychological Science in the Public Interest.

What NeoStudy does

This is why the app keeps a complete review log from day one and shows predicted against actual retention on your own material. The literature sets the prior; your data updates it. Where they disagree, your data wins.

If something here is wrong, that is a bug.

Misread a result, quoted a superseded effect size, or leaned on a study that failed to replicate — each of those is a defect in the product, not a typo on a marketing page, because the behaviour above was built on it. Findings are versioned with the design, and the sections that carry no evidence at all are listed rather than quietly omitted.

Compiled from the NeoStudy research foundation, Revision 2, last reviewed 27 July 2026.