Independent research and engineering project · Manik Maurya

Can the way a learner answers
reveal why they are wrong?

A wrong answer is where most measurement stops. This project asks whether the timing, hesitation, and confidence around that answer carry a readable signal about what memory is doing, forgetting, a missing prerequisite, a half-formed grasp, a confident misconception, and whether reading it can make revision land where it is actually needed.

Status: Exploratory Focus: memory-grounded learning diagnostics Trial: pre-registered (OSF 8PJQH) · no results yet Not a medical or clinical diagnostic

Where the question came from

I could see what a student did, and never why.

Teaching across rural classrooms, I kept hitting the same wall: I could see exactly what a student did, which question, which wrong answer, how long they took, and still had no way to see why. Was it forgetting, a misconception, or a missing prerequisite? Every system I had access to recorded the output and threw the reason away. That gap is the whole problem, and it is what this project set out to close.

The question

One primary question, three that follow from it.

Can the correctness, timing, hesitation, and confidence in how a learner answers reveal why they are wrong, forgetting, a missing prerequisite, a half-formed grasp, or a confident misconception, well enough to act on?

  • Can one kind of wrong be told from another? Can plain forgetting, a missing prerequisite, a half-formed grasp, and a confident misconception be separated from timing and confidence, reliably enough that trained raters agree?
  • Does reading the type help? Does reading the TYPE of error and remediating for it retain more, two weeks later, than adapting on correctness alone? (This is the pre-registered trial's primary question.)
  • Is the reading useful on its own? Even if scheduling shows no gain, is a per-learner memory state a useful thing to hand a teacher or a tool?

The instrument

What was built, and how it turns answers into observations.

Cognivia records three signals on every answer that passive study never produces: whether the trace held, how long retrieval took, and how confident the learner was. From those it estimates a personal decay rate per concept with uncertainty attached, sorts each miss into one of four error types, and schedules the next review. It runs in a browser on a single shared computer, so the classrooms that need it most do not need new hardware. See the mechanism or run it on sample questions.

The investigation, in order

Hypothesis, instrument, pilot, pre-registration, trial.

Hypothesis

Error type carries actionable signal

The starting bet: that hesitation and confidence around a wrong answer distinguish kinds of failure a score cannot, and that the distinction is worth acting on.

Instrument

A browser-based reader on a shared device

Built the engine that records timing and confidence, estimates a per-concept decay rate, and assigns an error type, with everything running on one low-cost computer.

Pilot

A small pilot (n = 5) to seed the defaults

The formulas, smoothing coefficients, and thresholds were set as engineering defaults from a small pilot and the literature, explicitly as starting values, not validated constants.

Pre-registration

The analysis plan locked before data

The trial design, primary outcome, and analysis were registered before any data collection (OSF 8PJQH), so results cannot be selected after the fact.

Now

Trial running; findings gated

Data is being collected. Participant-derived results are withheld until enough participants contribute, and reported either way. See the live status.

What is and is not established

The evidence, labeled honestly.

Demonstratedestablished science

Forgetting follows a steep, regular curve (Ebbinghaus 1885; replicated Murre & Dros 2015), and retrieval retains more than re-reading (Roediger & Karpicke 2006). The instrument itself runs today and produces a per-learner reading with uncertainty. Read the sources.

Preliminaryobserved, not validated

The error-type classification and the model's thresholds are seeded from a small pilot (n = 5) and the literature. They are how the system computes today, not validated constants, and the inter-rater agreement that would make the classification a measurement is still being established.

Not establishedthe open question

Whether error-type-aware remediation beats adapting on correctness alone (and standard classroom instruction) is exactly what the trial tests. No efficacy number is claimed anywhere on this site; the accuracy line stays hidden until at least twenty participants per arm have contributed. The design is public: read the pre-registered protocol.

Where the evidence changed my mind

I thought the scheduling was the product.

I started convinced the intervention, adapting the review schedule to the error type, was the whole point, and that the measurement underneath it was just plumbing. Building the instrument and watching it read individual learners reversed that. The per-learner memory state, the Genome and the Passport, is useful on its own: it tells a teacher which idea is slipping and why, whatever the scheduling result. So I now treat a null trial result as falsifying one specific claim, not the project. That changed what I consider the durable contribution here, and it changed what I would build next.

Limitations, in plain sight

What this does not show.

The pilot is small (n = 5), so the thresholds are defaults, not calibrated constants. The error-type classification is not yet validated by independent raters at scale. No efficacy result exists yet. And the instrument reads learning signals only, it is not a medical, clinical, or psychological diagnostic, and it does not claim to understand anyone's mind. These limits are on this page on purpose, not in a footnote.

What I personally did

Built by one person, on public evidence.

This is a research project run by one person. I designed the study and locked the analysis plan before collecting data, built the instrument and the engine that computes the Genome and schedules review, wrote the software that runs it on a shared classroom device, and I report the results as they come, including the ones that contradict what I predicted. It draws openly on the published literature it cites, and the work can be checked directly rather than taken on trust. Read the fuller account.

Built personally

The research question and study design, the pre-registered protocol and locked analysis plan, the confirm-or-clear diagnostic and the engine that computes the Learning Genome, the software that runs it on a shared classroom device, and this site. The decisions, the claims, and what they are allowed to say are Manik's own.

Drawn on, not claimed

Open-source software (the FSRS-6 scheduler), the published cognitive science this cites, and standard developer tooling, including AI assistants for parts of the coding and drafting. The empirical phase will run under formal institutional supervision. None of these are presented as Cognivia's own invention.

The next experiment

One bounded question next.

Reach enough participants per arm for the pre-registered comparison to be read at all, then test the narrowest version of the claim: does routing a wrong answer to the fix its error type implies retain more, two weeks later, than the same content with correctness-only adaptation? The design is fixed; the result is not.