By Jamie Scott
We have an announcement that might raise an eyebrow. At Evidence Based Education, we’ve spoken about how unreliable classroom observation can be. Our Director of Research, Professor Rob Coe, wrote a widely shared piece on exactly that more than a decade ago, and the point remains today.
So it is fair to ask: why are we now building a tool to help schools capture observations of teaching?
First, what the evidence tells us about observation
Have you sat at the back of a colleague’s lesson and formed a judgement about how good it was? It feels valid. Experienced teachers surely know good teaching when they see it.
The evidence says otherwise, and it is worth exploring. When trained observers using rigorously developed protocols watch the same lessons, their agreement is often only modest. In the Gates-funded Measures of Effective Teaching project, the reliability of observation instruments ranged from around 0.24 to 0.68. Put plainly, if one observer rates a lesson at the top of the scale, the chance a second observer watching the same lesson would disagree is between 51% and 78%, depending on the instrument. In one study where untrained observers were asked to sort effective from ineffective teachers on video, they did worse than they would have done by chance.
Why does something that feels so right so often get it wrong? A few reasons stand out. We tend to judge what we personally like, which is not the same as what helps children learn. Much of what we most want to see -learning itself- is invisible in the moment, so we fall back on proxies: a busy room, engaged faces, correct answers. Notably, none of this guarantees anything has actually been learned. Some of what we recognise as good practice is more fashion than evidence. And we simply miss a great deal, in the way that viewers famously fail to spot a gorilla walking through a basketball game when their attention is fixed elsewhere.
So why build the tool?
Observation has real weaknesses and they’ve not vanished. But they only become a real problem in one particular situation: when observation stands alone, treated as a single, authoritative verdict.
An entire approach to professional learning that rests on one observer’s judgement, and then treats that judgement as the data by which quality is gauged and improvement is claimed, is standing on foundations far weaker than they appear. A one-off impression, from one person, in one lesson, is being asked to carry a great deal more than it can bear.
That is not how the Great Teaching Toolkit has ever worked! The Toolkit is built on the idea that no single source of insight about teaching is enough. Student perceptions, gathered through validated surveys. A teacher’s own structured reflection. Feedback from a trusted colleague. Each of these is a mirror, and each has blind spots. The power is in bringing them together, cross-reading one against another, all against the same shared framework in our Model for Great Teaching.
Seen that way, observation is not something to avoid. The reason we are adding a teaching snapshot tool is precisely that it does not stand alone. Alongside student voice and self-reflection, a goal-focused snapshot of what happened in the room helps to balance, challenge and triangulate the fuller picture. On its own, it is one person’s view. In the mix, it earns its place.
Observation with a clear purpose
This matters, so it is worth being specific about how the tool is being designed. A snapshot is never a general appraisal of a lesson. The teacher being observed chooses a goal, tied to a specific Element and Dimension of the Model for Great Teaching, and the observation looks only at that. Or the observer suggests an Element or Dimension to become the teacher’s focus, as well as celebrating existing practice. The feedback then feeds into the goals and techniques in the Toolkit that help to develop a specific area of practice. The observer is not weighing up the broad, and largely unanswerable, question of whether this was a good lesson. They are looking at one area through the same shared language the teacher is already working with and helping them to improve it. It is observation and feedback in the service of development, not judgement.
What the snapshot tool is, and is not
Because the word “observation” carries such heavy baggage, it is worth being clear about what we are and are not building.
- It is developmental, designed to inform a teacher’s growth, not to grade them.
- It is one lens of several. It can be triangulated with student and self-reflection feedback.
- It is goal-focused, anchored to a specific Element and Dimension of the Model for Great Teaching, with feedback that points to the goals and techniques that help to develop it.
- It is a snapshot, not a verdict. It reflects moments to celebrate and learn from; it is not a summary of a teacher’s ability.
For the avoidance of doubt, the snapshot tool is not a grading instrument, not a compliance exercise, and not a performance-management tool wearing a friendlier name.
The bigger idea
There is a neat irony in all of this. The very reason we can add observation with a clear conscience is the same reason we spent years cautioning against it. Because we know any single source can be unreliable, we have built an approach where no single source has to carry the weight alone.
That is triangulation by design, and it is the thing we believe sets great professional learning apart. Observation, done in isolation and treated as truth, is as questionable as it ever was. Observation, done as one careful voice in a balanced conversation, is something we are glad to help schools do well.
The Great Teaching Toolkit teaching snapshot is being tested by a small number of schools now!
What next?
What next?
Your next steps in becoming a Great Teaching school


