Published On: August 5, 202619 min read

By Stuart Kime

A strong assessment result can be reassuring. But what, exactly, does it allow us to conclude? 

In this article, Céline Courenq, Head of World Languages at Bangkok Patana School (a Great Teaching Centre), reflects on how apparently successful end-of-unit assessments masked fragile learning, and how her department began distinguishing recent performance from knowledge students could retrieve and use later. Drawing on the department's evolving practice and ideas she sharpened while completing the Great Teaching Toolkit's Assessment Lead Programme, Céline explores assessment validity, cumulative checking and the vital link between evidence and action. Her central challenge is simple: before acting on a result, we must be clear about the inference it can genuinely support. 

What does this result actually tell us?

A few years ago, students in our department were producing strong results in end-of-unit assessments. The perfect tense looked accurate, consistent and secure. A few weeks later, in a longer or different piece of writing, it had almost completely disappeared. 

This happened often enough that we stopped treating it as an oddity. Students could perform well when the content was recent, the task familiar and the support still available. Once there was a time gap, less scaffolding or a less familiar context, the knowledge often did not hold. We had been reading successful performance under favourable conditions as evidence of learning. 

That raised a question. If the assessment said things were fine, and they clearly were not, what exactly had it measured? 

That question brought us back to the validity of the assessment: whether the task was giving evidence for the conclusion we wanted to draw. In our case, that meant designing checks that built in retrieval practice and spacing, so we could distinguish recent performance from durable learning. It also meant treating assessment as diagnostic rather than something separate from the learning process. The information needed to show what students could still retrieve, where knowledge was fragile, and what teaching should do next. At the same time, we had to stay alert to the limits of interpretation. No single score, task or baseline could carry more meaning than the evidence allowed. 

Assessment as a driver

Once we accepted that recent performance was giving us too much reassurance, we changed what we assessed. We identified the grammar and vocabulary students needed to secure at each key stage, not simply encounter or use successfully with help. These became our non-negotiables. 

We also had to decide what secure actually meant, and define it the same way for everyone, otherwise it just became each teacher's opinion and the tracking told us nothing we could compare. So secure meant the same thing in every class. What we differentiated was the support and how long we expected a student to take to get there, not the standard itself. Some students would take longer, and a few might not reach full security in every structure, but the bar they were working towards stayed the same. 

We built the non-negotiables into the curriculum map and tracked them through shared markbooks and short retrieval checks. We also designed the curriculum so these structures were encountered and used across different contexts, not just in the unit that introduced them. The checks were cumulative. Each year included core language from the previous year, so students could not rely on a recent burst of revision. Either they could still retrieve and use it, or they could not. 

That shift changed teaching without requiring a separate initiative. Once the assessment focused on retention rather than recent coverage, teachers had a reason to revisit material, space practice and repair weak areas. The checks made those decisions easier because they showed where knowledge had held and where it had not. 

This shift was not straightforward. Deciding what should become non-negotiable forced us to confront weaknesses in our curriculum planning. At times, we expected too much: too many structures introduced too quickly, with the assumption that exposure would lead to security. At other points, we reduced the content too far in an attempt to secure it, only to find that students were underprepared for the range of language required later. 

There were also practical tensions. Some colleagues felt pressure to move quickly through the curriculum so that all content had been "covered" in time for assessment. Slowing down to revisit and secure earlier material could feel like falling behind. In practice, however, moving on without security created a more significant delay later, when gaps resurfaced and had to be repaired under greater pressure. 

It also meant that problems appeared earlier. A student who could not retrieve core verb forms in week six was visible in week six, while the gap was still small enough to address. We no longer had to wait for a large end-of-year assessment, or for the same weakness to surface at IGCSE. The assessment was useful because it changed what happened next. 

Extended writing and end-of-unit tasks still had a place. We simply stopped asking them to do all the work. They showed how students could bring knowledge together, while the shorter cumulative checks told us whether the foundations were still available. Looking at both gave us a more honest picture than either could provide alone. 

Reporting and what counts as progress

The students who lost the perfect tense in independent writing had almost certainly received positive reports throughout KS3. Those reports were not invented or dishonest. They accurately described the things the system had been built to capture. That was the difficulty. It also gave the student the impression they were secure, so they had no reason to go back and shore up the thing that was actually fragile. 

Our reports included engagement, effort and performance goals. These are worth communicating to parents. They do not, however, tell us whether the language underneath has been secured. A student who works hard, enjoys the subject and completes scaffolded tasks successfully may receive the same positive message as a student who can retrieve and apply the language independently. 

There is also a persistent pressure within reporting systems to produce a clear, summarised judgement of future attainment at a given moment. Teachers understandably want a single, communicable outcome that can be shared with parents and students. However, that summary often requires a compression of evidence that does not neatly represent what we mean by progress. 

In many cases, what we were observing was still in development. Students were part-way through securing structures, or could demonstrate them under some conditions but not others. The reporting framework, however, encouraged us to turn that partial and conditional performance into a fixed statement. The result was a simplified picture of future progress that was easier to communicate, but less faithful to the underlying learning. 

This is also why apparently clear statements such as "I can use the perfect tense" need some qualification. Does the student mean in the current unit, after recent practice and with a model nearby? Or can the language still be used weeks later, independently and in a different context? Those are both achievements, but they are not the same achievement. Reporting systems often collapse them into one. 

In our own data, we could trace a familiar pattern: grammar gaps in Year 7, insecure tense use in Year 8, limited control of several time frames by Year 9, and a much more visible grammar weakness by Year 10. The eventual IGCSE or IB result could look like a sudden drop, even though the problem had been developing for years. 

This pattern is not just statistical; it is grammatical. The dependencies are built into the language. For example, in French, if être and avoir are not secure, students cannot reliably form the perfect tense, because those verbs are the auxiliaries it depends on. If aller is not mastered, the near future does not hold. The imperfect assumes a secure present-tense stem to build from. So a gap in a few core verbs in Year 7 or 8 is rarely a single gap. It is a foundation that several later tenses are built on, which is why early fragility does not simply add up over time but compounds. 

This matters because tense security is built into the examinations. At IGCSE, without secure control of several time frames a student is effectively capped below the higher grades, because narrating, describing and justifying across past, present and future is what those bands require. At IB Language B it costs twice: insecure tenses lose marks directly in Language, and again in Message, because when the tense is wrong the meaning blurs and the communication marks fall with it. 

One limitation of reports built in this way was that early signs were not always easy to see. A report could only reflect the evidence the assessment had gathered, so if retention had not been checked, it was difficult to comment on it confidently. As a result, parents and students were not always prompted to revisit foundations that seemed secure in class. When the same weakness became more visible in an external examination, it could then take longer to address. 

Rebuilding KS3 assessment around the non-negotiables changed the conversations we could have. We could talk about what a student had actually retained, rather than relying mainly on how well they had managed a supported task. It is worth being clear about what did and did not change here. The reporting system itself stayed as it was: it is set at school level, and we do not control its categories or the single judgement it asks for. What improved was the evidence sitting underneath our judgements. We knew more accurately what a student had secured, even where the framework we had to report it through could not fully carry that distinction. That gap, between better evidence and a reporting structure not designed to hold it, is part of the problem rather than something we had resolved. 

Data that does not lead to action

One of the clearest assessment failures we encountered did not come from a badly designed tool. It came from a system that had been designed carefully but was not being used consistently. 

We had common markbooks, shared checks, non-negotiable tracking and projected grades. In most classes, the system was doing what it was meant to do. In a small number, the information was present but the follow-up was incomplete. 

Two parallel exam classes with similar prior-attainment profiles developed sharply different grammar outcomes. In one class, grammar security fell from 38% to 13% within a term. The other class, with a near-identical starting profile, held steady. Intake alone did not explain the gap. 

The markbook helped explain what had happened. Some quizzes were missing or undated. The regular checks that should have shown the weakness early had not been maintained, and students who needed support were identified later than the available evidence allowed. 

The gaps were not invisible. They simply did not lead to action soon enough. Between the data showing fragility and any action being taken, the security drained away. For those students, the assessment system became a paper trail rather than a diagnostic tool. Where that follow-up is optional or inconsistent, the quality of the original assessment design is largely wasted. 

There is a second half to this that sits with the student. Even when the feedback is honest and the gap is named, progress depends on the student doing something with it. We saw students receive the same information about an insecure structure and respond very differently: some went back and secured it, others noted it and moved on. Feedback that asks for no action from the person receiving it changes very little. Part of the work, then, was making students responsible for acting on what the checks showed, not only making the checks more accurate. 

Using predictive baselines carefully

A different problem arises when progress is judged against a baseline that cannot support the interpretation being made from it. The difficulty comes when a broad prediction is treated as a precise statement about what an individual student, class or subject should achieve. Where a school cohort differs from the reference population in ways that matter for the outcome, the prediction may be less informative. This is particularly relevant in languages, where prior exposure varies enormously. 

Students may begin the language at different ages, arrive from different curricula, study it for very different amounts of time or have some connection to it at home. Those differences affect later attainment, but a general-ability measure may not represent them fully. A student can therefore appear to be underperforming against a prediction that does not reflect the actual starting point. Another may appear to have exceeded expectations when the original prediction was simply too low. 

At subject level, very small cohorts make value-added summaries particularly unstable, because one or two students can move the overall picture substantially. The number may look precise while the conclusion drawn from it is not. 

In our own cohorts, subject-specific prior attainment was more informative than a general-ability score when predicting later outcomes. A strong IGCSE language result told us more about likely post-16 success, because it reflected the cumulative knowledge and previous exposure on which the next course depended. 

This does not make general predictive tools useless. It means the interpretation needs to stay within what the evidence can support. A broad baseline can start a conversation. It should not settle a judgement about teaching, progress or intervention without the subject evidence alongside it. 

Task design and durable learning

The end-of-unit assessment itself can also give the wrong signal. A familiar pattern in languages is a test set after a period of intensive revision. Students know the topic, the task type and often the exact language they are expected to use. Many perform well, and the result is recorded as progress. 

That result answers a narrow question. It shows whether students can complete that task, on that topic, after recent rehearsal. Later courses and external examinations ask for something broader. Students need to retrieve and adapt language when it is less familiar, less supported and no longer recent. 

If the first result is read as evidence of the second, the assessment gives a misleading picture. A student may move through several years of high-support, recently rehearsed tasks without receiving a clear signal about independent performance. The gap becomes visible only when the scaffolding is removed, sometimes in an official examination. 

For a long time, we used National Curriculum levels, and one example shows the problem well. A student could answer comprehension questions about a text written in three tenses, and that performance earned them a level 6. On paper, a level 6 meant the student "knew" three tenses. But what does knowing mean there? What the assessment actually measured was reading: the ability to infer, decode and extract meaning from a text. Students are good at inferring, and it is a genuinely useful skill. But it is a different skill. A student could work out what a past tense sentence meant without being able to produce a past tense themselves, and the level could not tell the difference. We were crediting comprehension as grammatical knowledge. 

The contrast with what we do now is stark. The same core structure, say the perfect tense, is checked through short production tasks: no text to lean on, no options to choose from, weeks after the teaching, sometimes on mini whiteboards so every student produces an answer at once and nobody can hide behind a stronger neighbour. The question is no longer "can they work out what this means?" but "can they still make it themselves, without help, later?" Only the second predicts whether the language will hold at IGCSE. 

Running the same short checks on the same items across parallel classes, repeatedly over time, taught us something we had not expected: scores could vary hugely from one week to the next. A structure that looked secure on Tuesday could be gone ten days later. Under one-off assessment, that volatility was invisible. Repeated checking made the instability itself visible, which is exactly what a system tracking durable learning needs to see. The colour-coded markbook grew out of that: the same items, the same thresholds for what counted as secure, across every class, accumulating over time into a picture no single test could give. 

We made other mistakes too, and it is worth being honest about them. Wanting to stretch students, we added extra vocabulary and unseen structures to assessments and called it "challenge". The scores looked more demanding, and we treated them as a tougher, fairer test. But the difficulty had changed what we were measuring. A student doing well on those items was often showing that they could decipher an unfamiliar word from context or extract meaning from a text, reading and inference, rather than that they could use the language themselves. We had not raised the bar on the same construct. We had quietly switched to a different one, and gone on reading the result as if it measured productive control. Making an assessment harder is not the same as making it more valid. If the added difficulty draws on a different skill, a higher score tells you less, not more. 

Using more delay, less support and less familiar contexts is not about making assessment unnecessarily difficult. It is about matching the task to the claim. If we want to know whether language has been learned well enough to retrieve and apply later, the assessment needs to include some distance from the original teaching. 

How this thinking developed

None of this came from principle. For years, one of the main instruments was classroom impression: the feeling that, because we saw our students every lesson, we had a clear sense of where they were. At the time, it did not feel like we were guessing. It came from the normal things you pick up in lessons, for example, who answers, who writes confidently . But everything a teacher observes in a classroom happens under the most favourable conditions there are: the content is recent, the support is available, and the students you notice most are the ones performing. The impression is made of evidence, but it is exactly the evidence this article has been questioning. The levels example made this clear. A student could be rewarded for understanding a text, but the result could still be recorded as grammar. Once we had seen that happen, it became harder to trust some of the judgements we had been making. 

Some of the things that turned that question into a working system: 

cognitive science: how memory actually works, the difference between performance during learning and learning that lasts, why forgetting is normal and retrieval is what fights it. That gave us the why. It explained the vanishing perfect tense, the week-to-week volatility, the gap between classwork and exams, and it pointed directly at spaced, repeated, low-stakes checking as the response. 

Another was the Assessment Lead Programme. What it gave me, particularly as a head of faculty, was the language: validity, construct, fitness for purpose, what a result can and cannot support. Before that, we could sense that something was wrong with how we assessed; afterwards, I could name it properly, and that made it much easier to explain to others. It is one thing to tell a department "I don't think our tests tell us much." It is another to say "this task measures comprehension and we are reporting it as grammatical control, and those are different constructs." The first can sound like a personal view. The second is easier to discuss properly, because it is tied to the evidence. Leading assessment change in a team of experienced professionals needed that precision, and I did not have it until I was taught it. 

None of this required expensive tools. We wrote our own mini assessments: short, targeted at the non-negotiables, and used in the same form across every class so results meant the same thing wherever they came from. And the single best purchase we made in years was a set of mini whiteboards. 

What this asks of assessment design

The starting point is to be precise about what a result allows us to say. A high score on a familiar task completed immediately after revision supports a fairly narrow claim. It does not automatically show retention, independence or transfer. If those are the things we care about, they need to be present somewhere in the assessment model. 

The rest is less about adding more tests and more about making the existing system work. Evidence has to be reviewed while there is still time to respond. Shared checks need to be used consistently enough for comparisons to mean something. Baseline information has to be read alongside subject history, prior exposure and the uncertainty that comes with small groups. None of this is complicated in principle, but it is easy for one part of the process to become detached from the others. 

That detachment is where assessment starts to mislead. A valid task followed by no action is of little use. A carefully maintained markbook cannot rescue a task that measures recent rehearsal and is then read as durable learning. A precise-looking prediction cannot carry a judgement it was never designed to support. 

These issues are particularly noticeable in languages because learning is cumulative and exposure-dependent. Early gaps can remain quiet for years and become much harder to address once later work depends on them. The wider point applies to any subject in which new learning rests heavily on earlier foundations. 

The students who performed well on the perfect tense had shown that they could use it under the conditions of that assessment. The assessment failed to show that the learning was not yet durable. Because the result was read as secure, nothing in the system prompted a different response. 

Designing assessment that gives an honest signal of progress has been slower and less straightforward than we first expected. It requires us to examine what tasks really measure, decide what knowledge must be retained, build routines that are used across classes and make sure the information leads to action. It also requires enough restraint not to draw conclusions that the evidence cannot carry. In that sense, good assessment needs a kind of cognitive realism: it has to take seriously how learning actually behaves, including forgetting, instability and dependence on context, rather than assuming that a successful performance today means secure knowledge tomorrow. 

It is also important to recognise that this work remains in progress . We began reshaping our assessment and curriculum model several years ago, and while it has improved the clarity of what we see, it has not resolved every issue. Decisions about curriculum content, pacing and reporting continue to require adjustment. Some sequences still carry too much, others too little, and consistency across classes remains something we work towards rather than fully secure. 

The work is worth doing because assessment affects the direction of teaching whether we intend it to or not. It influences what teachers revisit, what students practise, what leaders notice and where support is directed. 

Progress depends on both: assessment that gives an honest account of learning, and a system that responds to what it shows.

Want to speak to Bangkok Patana School about their journey or to another Great Teaching Centre? Let us know and we'll make an introduction.

Your next steps in becoming a Great Teaching school

See the Great Teaching Toolkit platform and what it can do for you!

Request a quote for your school, college or Group!

Still thinking about how the Toolkit can be implemented in your context?