Over the last couple of decades, there has been a shift in thinking around education research. Once the preserve of university lecturers who read the theory and criticised the practice, it is now dominated by the dreams of politicians, who fund it to discover ‘what works’ in the classroom.
Many think they might have found it.
In 2019, England’s inspectorate Ofsted released an ‘overview of research’ which introduced the ‘Learning Sciences’, ‘a relatively new interdisciplinary field that seeks to apply understanding generated by cognitive science to classroom practice’. This ‘understanding’ included spaced practice, retrieval, interleaving, dual coding and cognitive load theory.
The same year, the government introduced a ‘minimum entitlement’ (latest version here) for initial teacher training – effectively, a curriculum for trainee teachers. It became mandatory for all teacher training providers to cover key aspects of the learning sciences.
Then there is the Education Endowment Foundation, which, like its US counterpart the Institute for Education Sciences, funds large-scale trials in an effort to discover ‘what works’ in schools. If you go onto the EEF’s webpage, you can find out exactly how much an individual intervention will help your students via the Teaching and Learning Toolkit. It’s a shopping list of improvements for school leaders.
A diverse range of developments – evidence backing up an inspection framework, a mandatory teacher training entitlement and a summary of the evidence which doesn’t force anyone to do anything. And I’m going to widen the net further as we go.
What they have in common is the fact that they are all converted into policy by schools. Some things schools buy in, some they have to follow and some they choose to implement, but when they appear in Monday morning briefing, they have the same effect and the origin is often irrelevant.
Beyond that, they share the idea that rigorous experimentation can tell us what must be done in the classroom. That through this ‘relatively new interdisciplinary field’ we are getting closer to a single ‘right way’ of doing things in the classroom. Deviation means poor practice.
There are right and wrong questions. There are good and bad methods. There are relevant and irrelevant outcomes. These can be communicated clearly and simply as a holistic and coherent model of ‘good teaching’.
They also imply that we can make predictions about the exact effects of particular instructional methods.
This is the science of learning. And note that it isn’t really about learning at all; it’s about teaching.
Why we need evidence
The move to make teaching accountable to some standards beyond simply a feeling of what works is a good thing. If we claim we want to make teaching more effective, we need evidence to support this. What about learning styles? Brain gym? The thing the deputy head read about at the weekend and now expects us all to implement in our classrooms? We should be subjecting methods and claims to critical, rigorous scrutiny. The key question, though, is what do our conclusions warrant? Do we have a science of learning? Or do we have rules of thumb about what works in the classroom?
A few weeks ago, something strange happened. With very little fanfare, an academic paper was published declaring ‘education researchers have tried but largely failed to conduct experiments capable of directly informing decisions’.1 The authors were not disgruntled Bourdieu scholars but included academics who had worked on EEF-funded projects.
Randomised controlled trials are not, on their own, the solution, these academics pointed out. Effect sizes of interventions are small and often shrink when interventions are scaled up and variation is sizeable – what increases scores for one student at one school might have the opposite effect for a student at another school (and we don’t really know why this is). For many interventions, the confidence intervals are so wide the true effect could be positive. Or negative. Or zero.
Their diagnosis is right. The ‘what works’ project has largely failed to show ‘what works’. The average EEF trial, they point out, costs around half a million pounds.2 As they point out in the pre-publication version of the paper, this has contributed significantly to the fact that ‘Rich nations have spent billions of pounds on experimental education research in the last twenty years, however much of it has proven uninformative’.
But what does it mean for the findings to be informative or uninformative? We could take the results to be a very expensive but meaningful lesson – education isn’t a science. The results of an RCT alone – without considering what we want from our schools – cannot decide policy.
Their solution is a different set of methods and a different role for experiments, but I’m not convinced that substituting a set of untested assumptions in place of randomisation is the answer. Still, they do suggest a more modest goal for the evidence generated – to inform teachers’ mental models rather than telling anyone what to do.
But what about the ‘Learning Sciences’?
The evidence for Cognitive Load Theory was initially developed through small-N lab experiments, not large-scale in-school trials. Do the criticisms still apply here?
I think that those things lumped under the umbrella term ‘Science of Learning’ can be useful, but I do not think that gives us a single ‘correct way’ to teach. As the EEF concluded after conducting a systematic review, ‘The evidence for the application of cognitive science principles in everyday classroom conditions (applied cognitive science) is limited, with uncertainties and gaps about the applicability of specific principles across subjects and age ranges’. They also argue, ‘Theories from basic cognitive science imply principles for effective teaching and learning.’ The key term, here, being principles rather than laws.
Dylan Kane embeds this discussion in his own practice and experience, pointing out that while we have some useful techniques, the classroom is a complex place and we’re still some distance from a comprehensive set of procedures that work all the time for all students.
Cognitive Load Theory in particular deserves extensive consideration, not offhand dismissal, so I’ll return to it in detail in a later post. Three things are relevant here, though. First, CLT is a theory of how information is processed during learning, not a holistic explanation of learning or teaching. It can guide a teacher on how to explain a complex idea while avoiding overloading working memory, but teaching involves considerations beyond task design, prior knowledge and working memory. Children’s learning is also affected by interest in a subject, what they had for breakfast, whether they’re suffering from hay fever, whether they think you ‘looked at them funny’ on the way in, and whether they’re wearing their new prescription glasses today or have ‘left them at home’ and so cannot see the board from where they’re sitting. These are not incidental questions; they can derail a whole lesson for a student. The only theory or algorithm that can weigh them all up simultaneously lives in the teacher’s brain.
Second, in practice, cognitive load theory is a rule of thumb, not a scientific law. An informed classroom observer judges qualitatively how well a teacher is adhering to it. We cannot make quantitative predictions about the exact cognitive load a student will experience under certain conditions.
Third, before we tackle the question ‘how best can we do it?’ we need to address another: ‘what’s it for?’ Is school as an institution a form of incarceration or a vital tool in addressing societal inequality? The debate goes deep, and CLT appears along one branch at a late stage. Before that come other questions – What counts as ‘society’? Who defines inequality? What’s a desirable outcome? What should we measure and how can we measure it, if we can even measure anything at all?
The problem with turning teaching into a science is that you get seemingly objective standards that you can hold a teacher to. You can always refer to CLT as if it tells you exactly what to do under any situation. You can tell all teachers to do the same thing as if it absolves you of the burden of explaining why. Is it because you want to control teachers and hold them to account? Because it makes inspections easier? Or do you think it leads to better outcomes for students? If you think it’ll lead to the best outcomes for students, you’ll need to do better than ‘Doug Lemov told me to do it.’
Should we all be Teaching Like a Champion?
Many of the things Doug Lemov highlights in Teach Like a Champion are effective – and the evidence points that way. But there is no evidence to say that requiring every teacher in a school to follow its prescriptions to the letter leads to better outcomes. As discussed, which outcomes are we talking about? Who for? Even if we fall back narrowly on test outcomes for a large number of students and average them for whole groups, we still can’t claim it works all the time for every student. There is no such evidence for Teach Like a Champion as a package.
Lemov derived these ideas from watching effective teachers, but not every teacher used every technique. Imagine they had. Imagine Lemov had approached each teacher and told them to use every technique with perfect fidelity – would that have improved what they were doing?
Mandating certain things in the classroom is necessary. Some things – don’t assault your students – are so obvious we don’t need to spell them out. My concern, as I’ve pointed out before, is around mandating whole packages as if they’re evidence based when the evidence is for individual elements.
We still judge effective teaching by walking into a classroom and seeing if we like what’s going on. Teachers work from rules of thumb, not algorithms. ‘Good teaching’ is based on professional judgement, how political trends affect what happens in schools, noisy exam grades and student feedback. All of these things are evidence, and we can debate the relative importance of each form, but it doesn’t make education into a science. Insisting it is one narrows the debate around effective teaching rather than opening it up.
Use this secret trick to turn your five-year-olds into thirteen-year-olds
Visit the EEF’s Teaching and Learning Toolkit and you get the impression a school can simply choose an intervention and their students’ results will shoot up by X many months. Now, the EEF do not advocate working through the strands like an eager toddler at the pick ‘n’ mix counter, adding up interventions as you go. But nor do they say you can’t. In fact, the way they’re presented invites it – so let’s do it.
If every school in England chose just one of these interventions, which have an average impact of around 3.7 months,3 we would expect students in a secondary school to make about 18 months’ extra progress over the five years of secondary school. But there is no indication from large-scale tests like PISA or TIMSS that English students are now a year and a half ahead of their international peers.
As Terry Wrigley – then visiting professor at Northumbria – pointed out back in 2017, add the strands together and your students make more than eight years of extra progress. ‘It would seem quite a ludicrous claim that, if you were to follow all these things, suddenly you would have five-year-olds behaving like 13-year-olds.’ In the same article, philosopher of education Gert Biesta was equally critical of the idea that education could work as an input-output system that can be measured through test scores: ‘This may be a logic that works with pig farming, but not with the complex endeavour of education’.
These conclusions are clearly ridiculous. The toolkit is otherwise helpful.4 The summaries are clear and the most serious methodological weaknesses are highlighted. The EEF have done, at scale, what has only been done piecemeal elsewhere.
But the ‘intervention as months of progress’ model implies that if we take two schools and do the intervention in one then we can predict with some exactitude the outcomes for their students. We cannot.5 Pretending we can undermines the nuance and reality of the application of evidence within schools – and applying evidence in the classroom is central to the EEF’s mission. Yes, teachers are busy. But giving them a headline figure rather than the messy reality is not the right solution.
Disagreement is the point
Treating teaching like there’s one form of science backing it up means there’s a danger of disappearing down a rabbit hole searching for a Grand Unified Theory of Teaching. Find it and we’re justified in insisting everyone must follow it. Science, we imagine, converges on the truth. Aristotle’s ideas were superseded by Newton’s which gave way to Einstein’s. Education, learning, teaching – these things are different. There is no single truth they can converge on which tells us how to teach; homogenising them as if we have agreement glosses over this complexity.
Instead, I think there are many useful theories, of which Cognitive Load Theory is one. The disagreement is not a sign that something’s broken but, like politics, it’s the point – it’s what strengthens the field. It points to more educational research, different kinds of educational research, better teacher understanding of the nuance, broader and more sustained teacher education and supported autonomy in the classroom.
Rather than unification, what we need is more curiosity. Why do we have disagreement? Why does someone follow this theory of learning when I subscribe to another? How can one informed commentator claim a classroom is a form of incarceration when another claims it leads to freedom?
This curiosity is what we expect of our students. These are the standards we should hold ourselves to. How else will we improve things for our students?
Thanks to my supervisor Dr Richard Brock for sending me a link to the paper. As he predicted, it was very much of interest.
While I interned there, I took part in meetings deciding, over the course of an hour, who to give the half a million pounds to, which brings a different kind of pressure to teaching thirty unruly kids about Ohm’s Law on a wet Thursday afternoon.
I’ve added these up by hand, skipping those with no stated effect size but including repeating a year (-2 months) and setting and streaming (0 months). The small group tuition page implies the figure is per year, so I’ve assumed 4 months’ additional progress per academic year and 20 months across five years of secondary school – an extension I’m not aware there’s any evidence to support. The same page also suggests 10 weeks may be the ideal length for the intervention, meaning 4 months’ progress for something lasting considerably less than a year. And ‘months’ progress’ is a conversion of effect size, which is what the studies actually measure.
I’m not opposed to meta-analyses, but as Greg Ashman highlighted (thanks to George Lilley for pointing it out) we need to think carefully about bringing together very different studies, and the extent to which a single conclusion is meaningful. Rob Coe has also argued, ‘Given two (or more) numbers, one can always calculate an average. However, if they are effect sizes from experiments that differ significantly in terms of the outcome measures used, then the result may be totally meaningless.’
Read the headlines rather than the small print and you could be forgiven for expecting a full refund on your intervention purchase of choice.


Who are you arguing with here. It comes across as a strawman.
First cognitive scientists are going to say the science of learning is the science of how we learn and would be applicable to how to teach. The science of teaching would be an overlapping domain that added items like behavior management.
Second it’s always the case that science is not values based. It doesn’t attempt to tell you what outcomes you should desire. Who is claiming that science can tell you which outcome is better?
Third it is always the case that what organizations consider- statistical results are not the same as what individuals consider- one sample result.
Fourth, what professional do you think is so much more scientific than teaching? A similar argument to your about the science of flight would go something like - each landing is different, weather airport, plane, crews state so we cannot apply the science of flight here. Further, what outcomes we care about - safety, comfort, fuel cost, time cannot be determined scientifically so we can’t use science here.
Again my point is just that most of what you say here is obvious and would apply in any profession so it is worth pointing to who thinks differently to avoid it seem you are arguing with an imaginary ignoramus.
Thank you for writing this blog. I’m not sure if the science of learning is about teaching after carrying out a replication study... There are many more aspects involved than we currently understand I think. I’ve written about it here: https://psycnet.apa.org/fulltext/2026-36551-001.html