MSA English Practice: Building Exam Familiarity Without Pretending It Is Learning

Post 5 of 5 — MSA exam preparation

What the MSA is for

The Mittlerer Schulabschluss is the endpoint of formal English education for many of the students sitting it. That changes what preparation should aim at. This is not a stepping stone to be optimised past; for a substantial proportion of the cohort it is the certification that will represent their English competence for the rest of their working life.

The written exam assesses listening, reading and writing. The demands are functional rather than sophisticated: can this student follow a conversation, extract information from a text, and write a clear letter. That is a reasonable standard to certify, and preparation should treat it as one.

Where standard classroom practice falls short

Four specific gaps, each of which the materials try to close:

Listening practice is often not authentic. A teacher reading a text slowly and clearly bears little resemblance to a recorded speaker at natural pace with connected speech. Students who have only practised the former are surprised by the latter.

Reading texts are often pre-simplified. The exam does not simplify. Practising on controlled-vocabulary texts builds confidence that does not transfer.

Writing feedback often targets grammar rather than the actual rubric. The MSA Bewertungstabelle rewards task completion, content relevance, coherence and range alongside accuracy. A student whose feedback has only ever addressed errors does not know what else is being marked.

Nothing is timed until the mock exam. Pacing is a skill. Discovering under exam conditions that your planning habits are unaffordable is an avoidable failure.

Structure of the units

MSA units are built as three-part practice matching the exam: listening, reading, writing. They are not tied to a textbook unit, which means topic selection can follow what the cohort will engage with rather than what falls next in a sequence.

Listening audio uses the browser’s built-in speech synthesis. This is a pragmatic compromise worth being honest about: synthesised speech is more natural than a teacher reading aloud and infinitely more available than recorded native speakers, but it is not authentic audio. It handles pace and connected speech reasonably and reproduces neither regional accent nor genuine spontaneity. For building extraction skills under time pressure it works; for accent exposure it does not, and that gap has to be filled elsewhere.

A live example: listening task progression

From one of the twenty live MSA units — Weekend Job at a Café. Each unit runs Listening, Reading and Writing in that order, and each can be worked through without a login.

Exercise A carries two separate recordings, mirroring the exam: a monologue, then a dialogue. Each plays a maximum of twice, the button disables when the plays run out, and the transcript is never shown on the page — the audio is generated on the student’s own device, so there is nothing to read ahead.

Part 1 — the manager’s briefing. A single female speaker runs through a first day at Riverside Café: training nine until half past twelve, then shadowing until two; black apron and closed-toe shoes; a twenty-minute break; clock in and out on the tablet by the staff door; if you are going to be late, ring the café rather than messaging on social media; payday the last Friday of the month; tips split equally. The questions are all specific information, and the distractors are the neighbouring plausible values — 12:00 against 12:30, the first Friday against the last. Nothing rewards a general impression of the recording.

Part 2 — two coworkers. A short exchange between Jake and Mia about the Saturday shift: learn the till system first, twenty percent staff discount but only when you are not on shift, and you submit your availability though the manager sets the final rota. This is harder than Part 1 for a reason. Two voices means tracking who said what, and every answer here is qualified — “Sort of”, “only when”. A student who hears “discount” and “choose our own shifts” and stops listening gets both wrong.

Then reading, then writing, on the same world. Exercise B matches applicants to job adverts and then runs true/false items on workplace notices. Exercise C is the exam’s Writing Part 3: choose one of two prompts — describe your first day at a new weekend job, or write to a friend about a job you would love to try — and produce 80 to 100 words. By that point the student has spent twenty minutes inside the vocabulary of shifts, rotas, breaks and tills, so the writing task tests composition rather than recall.

Why this order: single voice before two voices, receptive before productive, and the same scenario throughout. Reversing any of those — opening with the dialogue, or setting the writing task on an unrelated topic — produces students who guess.

Writing: rubric transparency

Writing tasks are marked against the official MSA Bewertungstabelle rather than a classroom scale. The units say so on the page — the writing section is labelled as graded by the teacher, not automatically — so a student never mistakes the automatic Note from the listening and reading sections for a judgement on their writing. Paired model responses at contrasting quality levels are on the list below rather than in the units; the Abitur packs have them, MSA does not yet.

Timed writing practice reduces support in stages — from a structural outline and vocabulary list with extended time, through a supplied opening sentence, to exam timing with nothing. By the exam, a student has completed several successful timed writes. That, rather than reassurance, is what reduces anxiety.

Integrity, and where it belongs

The MSA written exam is invigilated. Preparation that consists entirely of unsupervised home practice therefore risks producing a score that does not predict exam performance — which is the specific failure a preparation programme exists to prevent.

The distinction I work to is between practice and evidence.

Practice units are open. Students may repeat them, use dictionaries, and use whatever tools they have. The learning is in the attempt, and restricting it achieves nothing.

Timed writing assessment runs invigilated. The assessment pages carry the client-side measures set out in full in the companion post on academic integrity: paste and shortcut blocking, right-click and DevTools blocks, focus detection with an on-screen warning, typing-speed anomaly detection, a submission time gate, word-count and vocabulary enforcement, non-selectable prompt text, four-variant prompt randomisation, progress snapshots, and an integrity record submitted with the work.

Two caveats matter especially for MSA cohorts. These are deterrents and signals, not proof — a second device defeats all of them, and a flagged anomaly is grounds for a conversation rather than a verdict. And the measures need explaining in advance, because a fifteen-year-old who discovers mid-task that their tab switching is being logged will reasonably feel ambushed.

Used honestly, what they buy is a mock score that means something. A student who scores well under those conditions has grounds for confidence; one who does not has time to act on it.

Grading

Percentage-to-Note conversion is calculated automatically on submission using the school’s Punktetabelle, and MSA writing is assessed against the official MSA Bewertungstabelle rather than the classroom scale — so practice feedback mirrors the criteria that will be applied in the exam.

Student-facing output uses whole-number Noten with a floor applied. This is deliberate: a page that returns the lowest grade to a student practising voluntarily is unlikely to see them return.

Infrastructure

Pages are static HTML on GitHub Pages with a WordPress hub. Submission originally ran through Google Apps Script with an email fallback; that backend is migrating to Excel Online and Power Automate following a school data-protection restriction on Google services. Given how much of the current German debate concerns Datenschutz in schools, building for that constraint rather than around it has proved the right decision.

What I am still working on

  • Video-based listening, for visual context and something closer to authentic speech
  • Better accent exposure, which synthesised audio cannot provide
  • Targeted remediation, using performance patterns to direct practice at specific question types

The argument

The MSA is often treated as a hurdle. It is more useful to treat it as an opportunity for students to demonstrate genuine functional competence — not sophistication, not perfection, but the ability to listen, read and write in English well enough to be certified for it.

Preparation that removes format uncertainty lets the exam measure English rather than familiarity. Preparation that also makes some practice genuinely comparable to exam conditions lets a student know, before the day, whether they are ready.


Read the complete series:

Abitur Writing Preparation: Teaching the Task, Not Just the Language

Post 4 of 5 — Abitur English preparation

This is a companion piece to the existing Abitur guides on this site — der vollständige Leitfaden, kreatives Schreiben, and Kommentar schreiben — which are written for students. This one is written for teachers, on the design logic behind the practice tasks.

Why general advanced English does not prepare students for this

The Abitur English writing component is a set of specific, formalised tasks: text analysis, creative writing, argumentative writing, summary, mediation, letter and email. Each has conventions that examiners expect and that students do not discover by becoming generally better at English.

A candidate can be genuinely proficient and still lose marks for analysing a text without naming its techniques, for writing an argument that lists rather than develops, or for mediating word-by-word instead of conveying meaning. These are task-literacy failures, not language failures, and they respond to different teaching.

So the materials I build for Abitur preparation are organised by task type rather than by theme or grammar. The unit of design is “argumentative writing”, not “technology and society”.

The design logic

Work backwards from the task

For each task type, the first step is decomposition. What does the task actually reward, where do candidates typically lose marks, and what can be practised in isolation?

For text analysis, for instance:

  • Rewards: naming techniques, explaining effect, connecting to purpose, sustaining a line of interpretation
  • Common losses: quoting without analysing, generic interpretation, discussing sophisticated texts in simple language
  • Isolable sub-skills: identifying a device; explaining an effect in one sentence; linking effect to authorial purpose

Each sub-skill gets its own short exercise before students attempt a full response. A candidate who has practised “name the device, explain the effect, connect to purpose” forty times in isolation writes a very different essay from one who has only ever written whole essays.

Teach the formulaic language explicitly

Formal analysis runs on a small set of recurring phrases — the author employs, this passage illustrates, the juxtaposition of X and Y serves to, the paradox here reveals. These are not stylistic flourishes; they are the load-bearing structures of the genre.

Teaching them directly is sometimes treated as coaching rather than education. I think that is wrong. Withholding the conventional language of a genre from students who have not encountered it at home is not neutrality; it advantages the ones who have. Making the conventions explicit lets students spend their cognitive effort on interpretation rather than on guessing at form.

A live example: guided analysis with a structural frame

From one of the sixteen Abitur packs — Text Analysis: Aims and Ambitions. The set covers four task types (text analysis, argumentative writing, summary, mediation) across four Abitur themes.

Task type: Text analysis | Target level: B2–C1

The source text is a magazine feature, The Ambition Gap: Why Young Americans Are Rewriting the Rules of Success, with five passages highlighted in place. The pack then works through the decomposition above in order. Übung 1 asks students to match each highlighted quotation to the device it demonstrates — and the instruction says explicitly that these devices are different from the ones in the reference guide above, which stops the exercise from becoming a memory test. Übung 2 is purpose, audience and tone. Übung 3 takes each device in turn and asks not what it is but what it does to the reader. The article stays available behind a fold on every screen, so nothing turns on recall.

ÜBUNG 4 — Aufbau eines Analyseaufsatzes ordnen
The five descriptions appear shuffled; students click them into order.

  Introduction: names the text, author/source, and states the overall
  purpose and main argument.

  Body paragraph 1: analyses content/structure — what claims the article
  makes and how it is organised.

  Body paragraph 2: analyses language and style — key rhetorical devices
  and their effect on the reader.

  Body paragraph 3: evaluates effectiveness — how convincingly the writer
  achieves her purpose for the intended audience.

  Conclusion: summarises the analysis and offers a final judgement on the
  text's overall impact.

ÜBUNG 5 — Vollständige Analyse schreiben (150–200 words)
  Focus: Analyse how the writer uses the extended metaphor of the "ladder"
  and the contrast between generations to convey her argument about
  changing attitudes to ambition.

Why the frame helps rather than hinders: Übung 4 gives students the shape of an analysis essay without giving them a single sentence of one. They have to know that structure comes before language, and that evaluation comes after both — which is exactly the sequencing candidates get wrong when they write straight into the language paragraph. By Übung 5 the frame is gone: one focus, 150–200 words, a live word counter, and no stems. The scaffold has moved from the page into the student’s head, which is where it needs to be on the day.

Show the rubric

Every practice task carries the assessment criteria it will be marked against. In the packs this is a ten-item self-assessment list that appears the moment the essay is submitted — “I explained the effect of each device on the reader, not just named it”, “I linked my points back to the writer’s overall purpose”, “my register is formal/academic” — followed by an annotated model paragraph with its devices marked in place. Students tick their own work against the list before they read the model, then read the model knowing what they were looking for.

This is more effective than feedback on their own work alone. Recognising the difference between a strong and an adequate response is a transferable skill; being told your own essay was adequate is not.

Progression

Support is withdrawn in a deliberate sequence rather than all at once:

StageConditionsFocus
DeconstructedSub-skills in isolation, models suppliedThe analytical or argumentative move
SupportedFull task, frame and vocabulary available, extended timeStructure and coherence
ReducedFull task, no frame, near-exam timingIndependence and pacing
Exam conditionsExact timing, no supportPerformance under constraint

The final stage exists because pacing is a distinct skill. Candidates who have only written untimed essays discover their planning habits are unaffordable on the day.

Integrity in writing preparation

Abitur preparation raises a version of the problem that is worth stating plainly: essays written at home no longer reliably evidence unaided writing capability, and the Abitur itself is written under invigilation. A preparation programme built on unsupervised home essays can therefore produce confidence that does not survive contact with the exam.

The practical response is not to ban tools during preparation. It is to make sure some practice happens under conditions resembling the exam.

For timed practice, the assessment pages carry the client-side integrity measures described in full in the companion post on academic integrity — paste and shortcut blocking, focus detection, non-selectable prompt text, a submission time gate, word-count and vocabulary enforcement, prompt randomisation, and an integrity record. The same caveats apply, and they matter more here because the stakes are higher: none of it prevents determined substitution, and a second device defeats all of it. What it does is make one category of practice genuinely comparable to exam performance, so a student’s score means something they can plan around.

The interactive Abitur practice packs are a separate product line, distributed independently and designed for self-study rather than invigilation. They carry no integrity measures at all, and should not — a student working alone to improve has no incentive to deceive themselves, and treating self-study material as an exam would defeat its purpose.

That distinction seems worth being explicit about. Not every piece of practice material needs securing. The ones that produce a number someone will act on do.

Thematic content

Abitur prompts recur across a predictable set of themes, and vocabulary and background knowledge are built accordingly: technology and society, identity and belonging, environment and sustainability, education, and human relationships. Thematic preparation matters less than task literacy but is not negligible — a candidate who has thought about a theme writes with more to say.

What I am still working on

  • Comparative analysis exercises, where students evaluate two responses rather than producing one
  • Deeper mediation practice, which is consistently the least-prepared task type
  • Peer review training, since giving useful feedback develops the analytical eye faster than receiving it

The argument

Abitur preparation is often criticised as teaching to the test. The criticism has force when preparation stops at exam technique. It has none when the technique is genuine: analysing a text properly, constructing an argument that holds, conveying meaning across languages accurately. Those are real capabilities. The exam is simply where they get measured.

What preparation should remove is the exam format as a source of uncertainty — so that on the day, what is being measured is the student’s English rather than their familiarity with a rubric.


Read the complete series:

English for Vocational IT Trainees: Teaching the Language the Job Actually Requires

Post 3 of 5 — Vocational IT English

This is a companion piece to Englisch in der IT-Ausbildung: Was Azubis wirklich können müssen — that post sets out the case for trainees and employers; this one is the practitioner’s version, for other teachers building similar material.

The mismatch

A vocational IT trainee will spend their working life reading documentation, writing tickets, and troubleshooting in English with colleagues who are also not native speakers. Very little of that is served by a general English syllabus built around social conversation and travel vocabulary.

This is not a proficiency gap. It is a domain gap. A trainee at B1 who can read a configuration guide and write a clear incident report is more employable than one at B2 who can discuss holiday plans fluently. Vocational English instruction should optimise for the former, and much of it does not.

What the job actually demands

Analysing the language tasks first, before deciding on any grammar sequence:

Reading

  • Dense procedural documentation, read for extraction rather than comprehension
  • Error messages and log output, read diagnostically
  • Documentation written by non-native speakers, requiring tolerance of inconsistent register and occasional ungrammaticality

Writing

  • Support tickets and incident reports where imprecision has direct operational cost
  • Procedural documentation intended for readers less expert than the writer
  • Email that is economical, unambiguous, and appropriately urgent

Interaction

  • Asynchronous troubleshooting where a badly framed question wastes a day
  • Clarifying questions that narrow a diagnosis rather than restating the problem

Notice how much of this is precision under constraint rather than range. That has consequences for how it should be taught and assessed.

Exercise design

Scenario-grounded rather than topic-grounded

Every exercise starts from a situation a technician would recognise. Not “read about networking” but “this ticket arrived; what do you know, what is missing, what do you ask?”

A live example: ticket comprehension and helpdesk register

Taking one of the live pages rather than a constructed illustration: IT Support & Service Requests, one of ten units in the IT English series. It is open — no login, no name required.

Domain: Service desk — first- and second-level support | Target level: B1–B2

Trainees read the history of Ticket #4821. A user in accounting cannot open files on the shared drive; an error appears every time. The text is written the way a ticket is written — timestamped entries, first-level support asking for a restart that does not help, an escalation to second level at 09:30, a workaround issued before the actual fix, closure at 11:05 inside a four-hour SLA, and a follow-up action left open. The cause turns out to be permissions changed during an overnight system update.

Exercise A — extraction. Six items on what the ticket actually states. The distractors are the plausible wrong diagnoses — a broken hard drive, a virus, a forgotten password — so the trainee has to read what this ticket says rather than what a helpdesk problem usually is. One item asks what second-level support will check next, which lives in a single clause at the end that a fast reader skips.

Exercise B — nomenclature. Seven gapped sentences fed by a word bank of service-desk terms: ticket, escalation, SLA, workaround, first-level support, incident, resolution, service desk — plus sprint and encryption, which are not needed. The surplus entries matter. A bank with exactly the right number of words can be completed by elimination; a bank with two decoys cannot.

Exercise C — register. Eight items on how a technician actually addresses a user: indirect questions (“Could you tell me … ?”), reported speech for summarising what the user said, and modal choice for softening a request. The last item drops the multiple choice entirely and asks for a rewrite — turn “What is your computer name?” into something you would send to a customer. That single open item is the one that shows whether the pattern has been internalised or merely recognised.

Why this works better than a documentation-reading exercise: all three sections come out of one artefact the job produces, and every language point is one that artefact needs. The trainee is not reading about service desks; they are reading a service desk’s own record and then being asked to speak its language. What this unit stops short of is extended production — so that lives in a separate page, Writing Task: Texts the Job Produces, which assigns one of five real workplace texts (ticket reply, incident report, procedure for a non-expert, status email to a non-technical manager, asynchronous troubleshooting message) by a hash of the student’s name, so neighbours write different things.

Vocabulary in layers

Technical English separates cleanly into layers that need different treatment:

LayerContentTeaching approach
NomenclatureDomain nouns and acronymsGlossary, always available
Process verbsdeploy, configure, escalate, patch, roll backTaught in collocation, not isolation
Causal languageas a result, due to, this triggers, stems fromTaught through log and incident analysis
RegisterPassive for documentation, active for collaborationTaught by contrast between paired examples

The glossary stays available permanently, including during assessment. Working technicians use references; removing them tests memory rather than competence.

Assessment and the integrity question

Vocational cohorts raise the integrity problem in a sharper form than school groups, because the deliverable — a written ticket, a report, a procedure — is exactly the kind of short professional text a language model produces well.

The response is to separate the two task types explicitly.

Practice tasks are open. Trainees may use whatever they like. Using a model to check a draft ticket is a legitimate professional workflow and pretending otherwise is not credible to adult learners.

Assessed writing runs under invigilated conditions — but not yet in this series. The client-side toolkit exists and is in use on the university writing assessment: paste and shortcut blocking, right-click and DevTools blocks, focus detection with an on-screen warning, typing-speed anomaly detection, a submission time gate, word-count and required-vocabulary enforcement, non-selectable prompt text, four-variant prompt randomisation by name hash, three-minute progress snapshots, and an integrity record submitted with the work. The IT units do not use any of it, for the simple reason that they have no assessed writing to secure — the writing task described above is submitted for teacher feedback, not for a mark. When an assessed IT variant is built it will inherit the same toolkit; until then, nothing in this series produces a number worth defending.

The full implementation, the caveats it requires, and the data-protection questions it raises are set out in the companion post on academic integrity. The short version: it is deterrence and evidence, not prevention — a phone beside the laptop defeats all of it — and a flagged anomaly is grounds for a conversation rather than a verdict.

One point is specific to adult vocational cohorts. They are considerably more tolerant of invigilation than of a marking scheme they suspect is unenforceable. Explaining that assessed writing is secured because the certificate is supposed to mean something lands well with trainees who intend to use it. What does not land is pretending the practice tasks are secured too.

What I am still working on

  • Multimodal input, combining configuration files and system screenshots with language tasks
  • Extended scenarios requiring sustained language use across several linked artefacts
  • Peer review of technical writing, which develops the reader’s judgement as much as the writer’s

The argument

Vocational language teaching works when learners recognise they are building professional capability rather than studying a subject. Authentic scenarios do most of that work. Clear, honestly explained assessment conditions do the rest — adult learners are considerably more tolerant of invigilation than of a marking scheme they suspect is unenforceable.


Read the complete series:

When Homework Stops Proving Anything: Rebuilding Secondary English Exercises Around Verifiable Work

Post 2 of 5 — Secondary curriculum exercises

The problem worth naming first

Ask secondary English colleagues in Germany what has changed most in the last two years and very few will start with methodology. They start with marking. The Philologenverband put it bluntly this spring: teachers report correcting themselves to death on Klassenarbeiten and Klausuren, with little time left for lesson preparation or individual support.

Layer onto that a second problem. Written work produced at home no longer reliably evidences what a student can do unaided. That is not a moral panic; it is a measurement problem. If an exercise cannot distinguish between a student’s own English and a generated paragraph, it has stopped functioning as assessment — however good the task design is.

Most of my exercise development over the past year has been an attempt to answer both problems at once: reduce marking load, and make “the student’s own work” something a task can actually establish rather than assume.

Design principles

1. Separate practice from evidence

Not every task needs to be secure. Vocabulary practice, comprehension checks and low-stakes production benefit from being open, repeatable and unsupervised — if a student uses a translation tool to get through a matching exercise, the cost is low and the habit is visible soon enough.

Evidence tasks are different. Where a mark contributes to a report, the task needs conditions that make unaided production plausible. Conflating the two categories is what makes the whole system feel untrustworthy.

So exercises are built in two distinct classes:

  • Open practice pages — unlimited attempts, immediate feedback, no monitoring
  • Timed assessment pages — single sitting, invigilated conditions, integrity signals recorded

Students are told explicitly which is which. Ambiguity is what breeds resentment.

2. Feedback where the workload actually is

The marking burden is concentrated in open written production. The pages are therefore built to return structured feedback on objective components automatically — comprehension, vocabulary use, word count, task completion — so that teacher attention goes to the part only a teacher can judge: whether the argument holds, whether the register fits, whether the writing sounds like a person.

Automated grade calculation converts percentage scores to Noten on submission using the school’s own Punktetabelle. Student-facing output uses whole-number Noten with a floor applied — a deliberate choice, since a page that returns the lowest grade to a fourteen-year-old practising voluntarily is not doing pedagogical work.

3. Thematic coherence over grammar sequencing

Curriculum objectives are met, but organised around themes rather than structures. A unit on digital society carries conditionals, evaluative vocabulary and argument structure; a unit on friendship carries reported speech, present perfect and descriptive language. The grammar is scheduled; it simply is not the organising principle the student sees.

Exercise architecture

Each thematic unit progresses through four stages, with support withdrawn deliberately:

StageSupport levelPurpose
Vocabulary primingFull — context, glossaryRemove decoding load before the main task
Guided comprehensionHigh — question scaffolds, model reasoningBuild confidence on authentic content
Supported productionModerate — vocabulary bank, structural frameMove from recognition to use
Independent productionMinimal or noneGenerate evidence of unaided capability

The last stage is the only one that needs securing.

A live example: a supported production task

This one is real rather than illustrative — it is Exercise D of California: Natural Hazards and Human Impact, currently in use with Year 9. The page is open: it can be worked through without a login or a name.

Theme: Natural hazards and human impact in California | Year 9, Gymnasium | Target level: B1

The unit builds up to it. Exercise A establishes the content — what California’s environmental problems actually are. Exercise B drills modal verbs in gapped sentences about wildfire risk. Exercise C practises reporting structures (“California is said to have the best weather in the US”). All three are auto-marked. By the time students reach Exercise D, both the vocabulary and the grammar are in place, and the writing task can ask for the thing the unit is really about: connections. How deforestation, urban sprawl and coastal building relate to mudslides, floods and water shortage.

The support is a word bank plus a reference grid of the hazards and human activities involved, and the task itself sets the constraint:

WORD BANK — cause and effect language
  to affect          to cause             to be the consequence of
  to contribute to   to influence         to lead to
  to result in       to show the link between

REFERENCE GRID
  Natural hazards            Human activities
  earthquakes                deforestation
  tsunamis                   urbanisation of the coast
  heavy rain / floods        tourism / farming
  drought                    urban sprawl
  wildfires                  increase of traffic
  mudslides                  soil erosion / water shortage

(a) Write a paragraph (100-150 words) explaining how natural hazards
    and human activities in California are connected. Use at least
    four phrases from the word bank above.
    Hint: explain at least two different connections.
    Starter shown in the box: "Deforestation contributes to mudslides
    because..."

(b) In your opinion, which of California's environmental problems is
    the most serious? Give reasons. Write 2-3 sentences.

Why the support is there: the word bank and the grid remove the retrieval problem, so the cognitive work lands on the actual target — identifying a causal relationship and stating it precisely. A student who cannot summon “to contribute to” under time pressure has not thereby failed to understand that clearing a hillside makes it slide. Requiring four phrases also forces variation: a paragraph built entirely on “because” does not meet the brief. Task (b) then withdraws the frame entirely and asks for a judgement in the student’s own words — the same content, one support level down.

Why this task stays open rather than secured: the Note the student sees comes from the auto-marked sections A to C. Exercise D carries no points — it is submitted for feedback, not for a grade. The learning happens in the attempt, and nothing rides on the mark.

Making written evidence verifiable

Where a task carries weight, the assessment pages run a set of client-side integrity measures — paste and shortcut blocking, non-selectable prompt text, tab and focus detection, a submission time gate, word-count and vocabulary enforcement, prompt randomisation, and an integrity record submitted with the work.

Rather than repeat that here, it is set out in full — along with the caveats it needs, and the Datenschutz questions it raises — in the companion post on academic integrity. Two points matter for this age group specifically.

Tell students first, in specific terms. A fifteen-year-old who discovers mid-task that their tab switching is logged will reasonably feel ambushed. The argument for the measures is one most students accept when it is made honestly; it does not survive being sprung on them.

Prompt design beats surveillance. Of everything in that toolkit, the measure that does most work is the least visible: four scenario variants distributed by a hash of the student’s name, so neighbours write different prompts. A prompt specific enough that a generic answer looks generic is worth more than any keyboard block.

Infrastructure

Task pages are static HTML on GitHub Pages, indexed from a WordPress hub. Submissions originally posted to Google Apps Script with an email fallback; that backend is being migrated to Excel Online and Power Automate following a data-protection restriction on Google services at school — a constraint German colleagues will recognise, and one worth designing for from the start rather than retrofitting.

Static hosting matters more than it sounds. There is no server to maintain, no login for students to lose, and no account data held anywhere. A page is a URL.

What I am still working on

  • Peer review protocols, which reduce marking load and build metalinguistic awareness simultaneously
  • Spaced vocabulary review carrying terms across thematic units
  • Better in-class oral assessment formats, since oral work is the one thing no tool can produce for a student

The underlying argument

The current discussion in German schools tends to frame integrity measures and good pedagogy as opposites — surveillance versus trust. I do not think that holds. Clear conditions are what make a mark meaningful, and a meaningful mark is what makes feedback worth reading.

The question is not whether students will use AI. It is which tasks are supposed to prove something, and whether those tasks are built to do it.


Read the complete series:

What Counts as the Student’s Own Work? One Teacher’s Working Answer

The question German staffrooms are actually arguing about is not whether AI belongs in school. It is narrower and more awkward than that: what should a student produce individually, what may be supported, and how can performance still be graded fairly?

I do not think there is a general answer. But there is a workable local one, and it is what I have built into the timed writing assessments I use with my own secondary classes — with the same infrastructure now being extended to vocational and university writing tasks. It is offered as a position to argue with rather than a finished solution.

Two pressures, not one

The discussion usually gets framed as an integrity problem. It is really two problems that happen to have arrived together.

Marking load. The Philologenverband put it starkly this spring: teachers report correcting themselves to death on Klassenarbeiten and Klausuren, with little time left for preparation or individual support. Anything that adds verification work on top of existing correction load will not be adopted, however principled.

Unverifiable evidence. Written work produced at home no longer reliably demonstrates what a student can do unaided. This is a measurement problem before it is a moral one. A task that cannot distinguish a student’s English from generated English has stopped functioning as assessment, regardless of how well designed it is.

Solutions that address only one of these tend to make the other worse. Requiring handwritten drafts increases marking. Abandoning written assessment for oral examination is defensible but expensive in contact time — and it is where much of the current discussion is drifting by default rather than by decision.

The distinction that does the work

The single most useful move I have made is to stop treating all tasks as the same kind of thing.

Practice is open. Vocabulary work, comprehension, drafting, revision — students may use dictionaries, translation tools, language models, each other. The learning happens in the attempt. Restricting tools here achieves nothing except teaching students to conceal what they are doing. For adult vocational learners, pretending otherwise destroys credibility immediately; using a model to check a draft ticket is a legitimate professional workflow.

Evidence is invigilated. Where a mark contributes to a report, a certificate, or a decision, the task runs under conditions that make unaided production plausible.

Students are told explicitly which category any given task belongs to. In my experience the ambiguity is what generates resentment, not the invigilation. A student who knows the rules of a task rarely objects to them; a student who suspects the rules are unenforceable learns something worse than cheating — that the mark was never real.

This distinction also solves the marking-load problem, because it turns out most tasks do not need securing. Only the ones producing a number someone will act on.

What the invigilated pages actually do

The timed writing assessment pages are static HTML with client-side integrity measures. Grouped by what they are for rather than listed flat:

Preventing pasted answers

  • The paste event is blocked
  • Ctrl/Cmd+V is blocked

Preventing the prompt being copied out

  • Scenario and vocabulary boxes are set non-selectable
  • Right-click is disabled during the exam

Detecting attention leaving the task

  • Tab and window focus changes are detected via visibilitychange and blur
  • An on-screen warning appears when the student returns

Detecting inspection or tampering

  • F12, Ctrl+Shift+I, Ctrl+Shift+J and Ctrl+U are blocked
  • A DevTools size heuristic runs every four seconds

Detecting non-human input patterns

  • Typing-speed anomaly detection flags bursts above sixty characters in under half a second

Enforcing task compliance

  • A time gate disables submission for the first ten minutes of a sixty-minute task
  • Word count is enforced — submission is disabled outside the accepted range
  • Required vocabulary is checked at submission, so target language must actually appear

Making copying from a neighbour pointless

  • Four scenario variants, assigned by a hash of the student’s name

Creating a record

  • Progress snapshots every three minutes, submitted with the final text
  • An integrity payload accompanies the work: tab switches, paste attempts, anomaly flags, DevTools detection, user agent, session ID

What this does not do

A checklist presented as a solution would be dishonest, so:

It is not prevention. A phone beside the laptop defeats every item above. Nothing implemented in a browser can prevent a student reading from a second screen. What the measures remove is the frictionless path — the tab-switch-and-paste that takes four seconds and requires no planning.

It is not proof. A typing anomaly means a burst of fast input occurred. That might be a paste through an unblocked route; it might be a student who touch-types and had a sentence ready. The flag is grounds for a conversation, never a verdict. Automated integrity decisions would be both unjust and, in a school context, indefensible.

The human stays in the loop. This is the phrase circulating in the German debate and it is the right one. The data informs a teacher’s judgement. It does not replace it, and a system that tried would fail on its first false positive.

The quiet measure is the strong one. Of everything on that list, scenario randomisation by name hash is the most effective and the least visible. Adjacent students write different prompts. Unlike the blocking measures, it cannot be circumvented by a second device, and unlike the detection measures it produces no false positives. Prompt design does more integrity work than surveillance does — a prompt concrete and specific enough that a generic answer is obviously generic is worth more than any keyboard block.

The Datenschutz question, since it will be asked first

Any German colleague reading that feature list will reach the same question immediately, and they are right to.

The integrity payload is behavioural data about a named minor. Tab switches, session identifiers and user agent strings are personal data. Collecting them requires a legal basis, a retention decision, and transparency toward students and parents — not merely a page that works.

Concretely, this has meant:

  • Telling students in advance, in specific terms, what is recorded. This is both a legal requirement and a pedagogical one: a fifteen-year-old who discovers mid-task that their tab switching is logged will reasonably feel ambushed, and will be right.
  • Collecting only what informs a judgement. Anomaly counts and flags, not keystroke logs.
  • Migrating the backend. Submissions originally posted to Google Apps Script with an email fallback. Following a data-protection restriction on Google services at school, that is moving to Excel Online and Power Automate. Given how much of the current debate concerns Datenschutz, building for that constraint rather than around it has proved the right decision, if a laborious one.

Anyone adopting measures like these should treat the data question as part of the design rather than a compliance step afterwards. It changes what you build.

What I would say to a colleague considering this

Start with prompt design, not blocking. A specific, situated prompt with a named context and a required vocabulary set does most of the work, costs nothing, and produces no privacy questions.

Then secure only what needs securing. If you find yourself invigilating everything, the category distinction has collapsed and the marking load is about to return.

Tell students exactly what is happening and why. The argument — that a mark only means something if the conditions were clear — is one most students accept when it is made honestly.

And be careful about the claim you make for it. The measures buy the ability to keep written production as a meaningful assessed component instead of retreating entirely to oral examination. That is a real gain. It is not the same as having solved anything.

The position, stated plainly

The current framing treats integrity measures and trust as opposites. I do not think they are. Clear conditions are what make a mark meaningful, and a meaningful mark is what makes feedback worth reading.

The question was never whether students will use AI. It is which tasks are supposed to prove something — and whether those tasks are actually built to do it.


Read the complete series: