After fourteen of these, a pattern gets hard to ignore.
I’ve now gone through the primary literature on magnesium, omega-3, creatine, anthocyanins, choline compounds, ginkgo, tart cherry, binaural beats and caffeine. The verdicts differ in the details, but the shape recurs with a regularity that started to feel less like a series of findings and more like a property of the method.
Trial recruits healthy people. Trial measures cognition. Trial finds a small effect or no effect. Headline says the thing doesn’t work.
Somewhere around the creatine article I started wondering whether I was reading fourteen answers or one answer fourteen times. So I went and looked at what happens inside those trials that has nothing to do with the substance being tested.
The first thing I found reorganised how I read all of them.
The effect of taking the test twice
If you give someone a cognitive test, wait a few months, and give it to them again, they do better. Not because anything changed in their head, but because they’ve seen the test.
This has a name — practice effects — and a literature going back decades. What I hadn’t appreciated is the size.
In healthy adults tested seven times across a year, practice effects were clinically relevant through the first three months, with Cohen’s d ranging from 0.36 to 1.19, most pronounced early on, then plateauing. In older populations assessed three times over six to twelve months, a methodological review estimates the effect on composite cognitive measures at about 0.25. A protocol for a current trial, citing drug-trial data over three to six months, gives a maximum of 0.10 to 0.15 and calls it “quite small.”
Those three figures disagree, and the disagreement is informative — the effect depends heavily on how often you test and who you test. But hold the middle estimate, 0.25, next to what we’ve been hunting.
Caffeine with L-theanine, in the 2025 meta-analysis: SMD 0.20 to 0.33. Anthocyanins in impaired older adults: SMD 0.34. Creatine on memory, before the statistical correction that erased it: SMD 0.31. Magnesium bisglycinate on insomnia severity: d 0.20.
The artifact is the same size as the signal.
That’s the sentence I keep coming back to. In an older population tested repeatedly over a year, the improvement from simply having done the test before sits squarely inside the range of the compound effects the trial is trying to detect.
Why this doesn’t mean the trials are broken
The obvious objection is correct, and I want to state it clearly before going further, because half the internet gets this wrong in the other direction.
Randomisation handles this. Both arms take the test the same number of times. Both improve from practice. The comparison between them subtracts the artifact out. A well-run randomised trial is not fooled by practice effects, and anyone telling you that supplement trials are invalid because of them is selling something.
The problem is narrower and more specific: it shows up when you try to interpret what a null actually means.
When both arms improve and there’s no difference between them, two very different situations produce an identical-looking result. Either the compound did nothing — or both groups got better at the test, and whatever the compound did was smaller than that noise floor. From the outside, those look the same.
The MIND diet trial is the cleanest example I’ve found. Six hundred and four older adults, three years, a diet purpose-built for brain health against a control diet. No difference between groups. But both groups improved.
The authors — and this is why I like this paper — listed three possible explanations in their own discussion. The control group probably improved their diet too, given similar weight loss. Practice effects of repeated cognitive testing could account for improvement in both groups, as observed in previous randomized trials. Or these interventions simply don’t improve cognition.
They named it. In the New England Journal of Medicine. It just never left the Limitations section.
What the researchers already do about it
This is the part that made me abandon the framing I started with, which was something like “the field is missing this.”
It isn’t. The mitigations are published and specific. A 2015 methodological review lists three: massed practice in a pre-baseline period to burn off task familiarity before randomisation; tests designed to minimise item-specific improvement; and well-matched alternate forms — different word lists, different items, same difficulty. Recent modelling work proposes a single-arm placebo run-in phase before randomisation specifically to extinguish practice effects, and calculates that trials designed this way would need substantially smaller samples.
There’s even a paper titled, precisely, “Practice effects in nutrition intervention studies with repeated cognitive testing.” Someone is already on this.
And alternate forms don’t fully solve it. In 502 cognitively unimpaired adults aged 60 to 85, tested at screening and again at baseline a median of 3.5 months apart using alternate versions, significant practice effects still appeared on the total scale and on both memory indices.
That study also produced a null worth reporting alongside: visuospatial construction, language and attention showed no significant practice effect. So this isn’t uniform — verbal memory and verbally-mediated executive tasks are the vulnerable ones, visuospatial measures much less so. If a trial’s primary endpoint is a word list, practice matters more than if it’s a reaction-time task.
The part where I have to be honest about a nice-sounding idea
There’s a second explanation people reach for, and it’s intuitively appealing: the participants got better because someone was paying attention to them. They were seen, checked on, given something to do. In the MIND trial, both groups had a dietitian calling them regularly for three years.
That intuition has a famous name — the Hawthorne effect — and I went looking for its evidence expecting to use it. What I found instead is a cautionary tale about exactly the kind of confident citation this site exists to interrogate.
The original studies ran at the Western Electric works near Chicago between 1924 and 1933. They’ve been re-analysed repeatedly — in 1992, 2009 and 2011 — and the reanalyses have not been kind. As early as 1958, Landsberger concluded that the effect originated “in the bias of the creators rather than in the facts it seeks to explain.” When Levitt and List recovered the original illumination-experiment data in 2011, what they found was “some weak evidence” that workers respond more to experimental manipulations than to naturally occurring changes in light. The original design had a small volunteer sample, attrition from operators removed for insubordination, and an author who was an officer of the company.
So does research participation change behaviour? Probably yes. A 2014 systematic review in the Journal of Clinical Epidemiology found nineteen purposively designed studies, most showing some evidence of an effect, with the randomised trials among them tending to show small statistically significant effects — and some showing no effect at all.
But then the authors say the thing that settles it: “the heterogeneity of these studies means that little can be confidently inferred about the size of these effects, the conditions under which they operate, or their mechanisms.” And, more damning: “as the Hawthorne effect construct has not successfully led to important research advances in this area over a period of 60 years, new concepts are needed.”
Sixty years, and nobody can tell you how big it is. The field has largely moved to calling it measurement reactivity and studying narrower, better-defined versions.
Which leaves an honest asymmetry. Practice effects: measured, replicated, quantified, with published fixes. Attention effects: probably real, essentially uncharacterised, and named after an experiment that doesn’t support them.
Four other ways a trial arrives at nothing
Practice effects aren’t the only route. Across fourteen articles I hit four more, each with a different remedy.
The responders get averaged away. This one recurs so often it’s become the site’s most repeated finding. Magnesium’s sleep benefit was greatest in participants with low baseline dietary intake. Blueberry anthocyanins missed their primary cognitive endpoint at p = 0.23 in a mixed population while dropping inflammatory markers dramatically. Creatine’s memory effect, after correction, survived only in older adults. Recruit people with nothing to correct and you’re measuring a repletion intervention in the replete.
The dose never reached the target. The most under-appreciated one. Creatine at 20 g/day for four weeks raises brain creatine by 8.7% — and most trials use 5 g and never measure brain creatine at all. Omega-3 needs a specific carrier molecule to cross the blood-brain barrier, and standard fish and krill oils largely don’t supply it. A trial can faithfully demonstrate that a compound doesn’t work while the compound never arrived.
The statistics inflated their own precision. Two meta-analyses on creatine and memory entered multiple subtests from the same participants as independent observations, making the analysis behave as though it had more data than it had. When the first was corrected — at the authors’ own initiative, after a reader wrote in — the effect disappeared except in older adults. The second made the same error three years later.
Both arms got an intervention. Back to MIND: both groups were coached to cut 250 calories a day, both had regular dietitian contact for three years, both lost about five kilograms. That’s not “brain diet versus nothing.” It’s “brain diet versus another diet intervention,” and a null there means something much narrower than the headline suggests.
When a null is just a null
Here’s where I have to check myself, because everything above could be assembled into a machine for explaining away any inconvenient result, and that would make this site worthless.
So: ginkgo.
Two trials. The first randomised 3,069 people aged 75 and over and followed them for a median of 6.1 years. The second took 2,850 primary-care patients with memory complaints and followed them for five. Both used the best-characterised extract in existence, at the studied dose. The first was funded by a US government body. Neither found a reduction in dementia.
Nearly six thousand people, half a decade each, adequate dose, right population, no arithmetic problems. That is not an artifact. That is an answer, and I’m not going to explain it away.
The difference between that and the MIND trial isn’t sophistication — it’s that the ginkgo trials had the design features that let a null mean something: a population with something to gain, a dose known to reach its target, a control that wasn’t also receiving an intervention, and enough people and years for the effect to have shown up if it existed.
Five questions for the next headline you read
This is the part worth keeping. When you next see “study finds no benefit,” you can get most of the way to an interpretation in about two minutes.
Who was in it? Healthy, well-nourished, well-rested people are the population in which almost everything fails. If the trial recruited people with nothing to correct, a null tells you about them, not about you.
Did the compound reach where it needed to go? Did anyone measure it? For anything acting on the brain, blood levels are not brain levels, and a startling number of trials never check.
What was the control group actually doing? If they were also on a diet, also being counselled, also changing something — the comparison is narrower than it appears.
How many times did they take the test? Repeated cognitive testing produces improvement of the same order as the effects being hunted. It doesn’t invalidate a randomised comparison, but it does tell you what the noise floor looks like.
Was there a correction? Corrections, replies to letters, and failed replications almost never travel with the original finding. They’re usually one search away and they’re usually the most informative thing available.
None of those questions require statistical training. They require knowing that the Limitations section exists and is often the most honest part of the paper.
Which is the actual finding here, and it’s not about supplements at all. The researchers aren’t hiding this. They write it down, in the paper, in plain language, and then it stops. What reaches you is a headline that has stripped out every condition that made the result interpretable.
The gap isn’t in the science. It’s in the transmission.
About this article
Written by Drew Anton. Covers nootropics, stimulants, sleep and focus protocols, and wearables. Not a physician or research scientist — reads the primary literature closely and refuses to round up.
Medical review: None. NeuriFuel does not currently have a licensed clinician on the editorial team, and this article has not been medically reviewed. We state this rather than implying an authority we do not have. See our About page for our full methodology.
Sources: Built from the methodological literature on practice effects and research-participation effects, plus the primary trials covered in our own previous fourteen articles. Where three published estimates of practice-effect magnitude disagree, all three are given rather than one chosen. Nothing here is presented as a novel insight — the effects described are documented in the psychometric literature, the mitigations are published, and the MIND trial’s authors named practice effects themselves in the New England Journal of Medicine. The failure described is one of transmission, not of research conduct.
Corrections: Found an error? Write to hello@neurifuel.com with a source and we will fix it and log the correction.
Last updated: 29 July 2026
References
- Bartels C, Wegrzyn M, Wiedl A, Ackermann V, Ehrenreich H (2010). Practice effects in healthy adults: a longitudinal study on frequent repetitive cognitive testing. BMC Neurosci 11:118. DOI 10.1186/1471-2202-11-118; PMC2955045 — N=36, d 0.36–1.19 through month 3
- Practice effects due to serial cognitive assessment: implications for preclinical Alzheimer’s disease randomized controlled trials (2015). Alzheimers Dement (Amst). DOI S2352872915000068 — effect size ~0.25; three published mitigation strategies
- Practice effect of repeated cognitive tests among older adults: associations with brain amyloid pathology and other influencing factors (2022). Front Aging Neurosci. DOI 10.3389/fnagi.2022.909614; PMC9297730 — N=502, aged 60–85, alternate forms, memory indices only
- Bell L, Lamport DJ, Field DT, Butler LT, Williams CM (2018). Practice effects in nutrition intervention studies with repeated cognitive testing. DOI 10.3233/nha-170038
- Implications of practice effects for the design of Alzheimer clinical trials. PMC12420670 — single-arm placebo run-in proposal
- McCambridge J, Witton J, Elbourne DR (2014). Systematic review of the Hawthorne effect: new concepts are needed to study research participation effects. J Clin Epidemiol. PMC3969247 — 19 studies; size, conditions and mechanisms cannot be confidently inferred
- Levitt SD, List JA (2011) — reanalysis of the original Hawthorne illumination experiments
- Landsberger HA (1958) — the effect originated “in the bias of the creators”
- French DP, Sutton S (2010) — “measurement reactivity” as replacement terminology
- Barnes LL, et al. (2023). Trial of the MIND diet for prevention of cognitive decline in older persons. N Engl J Med. DOI 10.1056/NEJMoa2302368 — N=604, three years, both arms improved; authors name practice effects among three explanations
- DeKosky ST, et al. (2008). Ginkgo biloba for prevention of dementia (GEM study). JAMA. PMC2823569 — N=3,069, median 6.1 years, null
- Vellas B, et al. (2012). GuidAge trial. Lancet Neurol. — N=2,850, five years, null
- NeuriFuel’s own prior coverage, from which the effect sizes and the four additional null mechanisms are drawn — see the linked articles in text

