Home / Research / Method

Study design & methodology

The protocol is fixed.
The result is not.

This describes the study design, what it is testing, and what it is not. It was written before data collection began, and the analysis plan was fixed at that point. Results are reported as they are, not selected for the ones that look best.

Pre-registration (OSF)

What this study is

Cognivia is a student-led research project. It reads why a learner's answer was wrong, telling a careless slip apart from a fading memory, a missing prerequisite, and a confident misconception, and records that diagnosis as structured data on a single shared classroom computer. Every response and session is logged, timestamped, and open, so the effect of acting on the diagnosis can be measured against what actually happens to retention.

The core question is precise, and it is not whether review helps: does knowing the type of a student's error, not merely that they erred, produce measurably better retention two weeks later than adapting to correctness alone? The study is ongoing. Results are reported as evidence becomes available, including any findings that do not support the hypothesis.

The trial randomly assigns students to three arms. Two use Cognivia, one adapting to the type of error a student makes, one adapting to correctness alone; the third is standard NCERT classroom instruction. All three cover the same content, in the same schools, over the same six weeks.

What each reading needs

Timing and confidence are not enough on their own.

A common shortcut says the four error types can be told apart from timing and confidence alone. That is not defensible for all of them. Each type has a minimum evidence bar, and when that evidence is absent the classifier returns insufficient evidence rather than guessing.

ClassificationMinimum evidence needed
Plain forgettingMastery shown before, then decline after an interval. Needs concept history, not one answer.
Missing prerequisiteA prerequisite map plus weak or absent mastery of the concepts it depends on.
Partial understandingInconsistent transfer across items, with latency and confidence that wobble.
Confident misconceptionThe same specific wrong answer or distractor, repeated, with high confidence.

Timing and confidence sharpen a reading and flag the confident-misconception case early, but forgetting needs prior mastery, and a prerequisite gap needs the concept graph. The public demo and the API both return insufficient evidence when the bar is not met, and both let you upload a distractor map and a prerequisite map to raise it.

Roadmap

What happens next, and when.

The trial has a closing date. If the scheduling result comes back null, that one claim is reported as null and set aside; the data stays public, and the Genome and the Passport, the measurement the scheduling was built on, carry on as a diagnostic in their own right.

Feb - Apr 2026

India pilot · n = 5 · closed

One rural school in Kanpur. 548 questions answered, 36 sessions logged. Feasibility study only, not a result.

Jun - Aug 2026

Three-arm RCT recruitment · n = 90 (30 per arm)

1:1:1 allocation, stratified by school and class level, concealed until after the pre-test. Pre-registered exclusion criteria. Recruiting through government schools in Kanpur Dehat.

Sep 2026

2-week delayed retention collected (primary)

Primary outcome data collected for every enrolled participant. No further intervention after this date, only the probe.

Oct 2026

Pre-registered analysis runs · OSF posted

The analysis script (already public in the repository) runs once on the locked dataset. Results posted to OSF regardless of outcome, positive, null, or negative.

Nov 2026

Journal submission

Primary target: Computers & Education or Journal of Educational Psychology. Submission goes out regardless of result; a null finding is a publishable finding.

Dec 2026

The mechanism bet resolves

If the effect appears (Cohen's d ≥ 0.3, pre-registered), planning begins for an independent replication run by someone else, somewhere else. If it does not, the scheduling claim is reported as a null and set aside, while the Genome and Passport measurement layer carries on. The repository, the data, and the protocol stay public, that is the only honest version of research.


Limitations

Stated clearly, not buried in footnotes.

Every study has limits. Stating them clearly is not weakness; it is the only way the results carry weight. These are the conditions under which anything in this study can be trusted.

What works in a laboratory may not work in a real classroom

Decades of research have shown that spaced practice and retrieval testing improve memory. Almost all of that research was done with university students sitting at individual computers in quiet rooms. Whether the same effects appear in a resource-limited school, with many students sharing one device, irregular attendance, and teachers under pressure to cover a fixed syllabus, is an open question. We do not assume the answer.

Five students is not enough to draw broad conclusions

This pilot involved five students from one school. That is too small a group to make any general claim about how students learn everywhere, or even in all classrooms like it. What it can do is establish whether the method is worth testing more rigorously, with a larger, properly randomized group spanning different countries and contexts.

The individual student profiles are theoretical at this stage

Cognivia tracks ten variables about each student's learning behaviour and builds a profile from them. Whether that profile can reliably predict who will struggle or identify the best time to intervene has not yet been proven. This study begins to generate the kind of data needed to test that.

The teacher knows which class is piloting the new tool

Both groups share the same teacher, subject, and week, which controls for a lot. What it does not control for is enthusiasm: a teacher who knows one class is using new technology may, without meaning to, teach it with more energy or attention. That expectancy effect would inflate the result independently of the tool. Full blinding is impossible here, so we treat it as a live threat to validity, plan a teacher-behaviour check, and read the two-arm technology comparison (A versus B, where both classes use Cognivia) as the cleaner test precisely because it holds that enthusiasm roughly constant.

Students will miss sessions, and some will drop out

In the classrooms we work in, irregular attendance is a fact, not an edge case, so the analysis is pre-specified for it. The primary analysis is intention-to-treat: a student is analysed in the arm they were assigned to, whether or not they completed every session. The number of sessions attended is recorded as a dosage covariate. Missing scores are handled by a pre-registered rule (multiple imputation under a missing-at-random assumption, with a complete-case sensitivity analysis alongside), and attrition is reported per arm in a CONSORT flow diagram. If dropout is heavy or unbalanced between arms, that is reported as a limitation, not smoothed over.


If the evidence holds

What becomes possible.

None of what follows is certain. Each possibility depends on the results being real, the methodology holding up under review, and the work being replicated at a larger scale. These are not promises. They are what becomes possible if the evidence keeps pointing this way.

Any teacher anywhere could replicate this study

The study design, data format, and analysis approach are all documented and public. A teacher in a different city, with different students and a different subject, could run the same experiment and compare results. The value multiplies with each independent replication.

A rare kind of data, from classrooms rarely studied

There are very few long-term, question-level datasets tracking how real students in low-resource schools actually forget and retain over time. If this study scales, even modestly, it begins to fill a gap no single experiment can address alone.

Submitting the findings to peer review

The methodology, results, and limitations are being prepared for a peer-reviewed journal. A negative result will be submitted alongside a positive one. The credibility of the work depends on being willing to publish both.

Giving teachers information they can actually use

The dashboard already shows which students are struggling with which subjects, and how performance changes week by week. Developing that into something a teacher can consult at the start of each lesson, without any technical knowledge, is a natural next step.

More classrooms, only if it actually works

If the results hold up and replication supports them, Cognivia can be deployed in additional classrooms without any additional infrastructure. It runs on a single computer, with nothing to install. The conditions that make it practical in a resource-limited school are the same ones that make it scalable, globally.

The data matters more than the tool

Cognivia is one tool. The more durable outcome is the evidence it collects about how students in underfunded schools actually learn and forget, evidence that does not currently exist in the research literature. That has the potential to influence how teaching is approached at a systemic level, in ways no single application ever could.