Post 2 of 5 — Secondary curriculum exercises
The problem worth naming first
Ask secondary English colleagues in Germany what has changed most in the last two years and very few will start with methodology. They start with marking. The Philologenverband put it bluntly this spring: teachers report correcting themselves to death on Klassenarbeiten and Klausuren, with little time left for lesson preparation or individual support.
Layer onto that a second problem. Written work produced at home no longer reliably evidences what a student can do unaided. That is not a moral panic; it is a measurement problem. If an exercise cannot distinguish between a student’s own English and a generated paragraph, it has stopped functioning as assessment — however good the task design is.
Most of my exercise development over the past year has been an attempt to answer both problems at once: reduce marking load, and make “the student’s own work” something a task can actually establish rather than assume.
Design principles
1. Separate practice from evidence
Not every task needs to be secure. Vocabulary practice, comprehension checks and low-stakes production benefit from being open, repeatable and unsupervised — if a student uses a translation tool to get through a matching exercise, the cost is low and the habit is visible soon enough.
Evidence tasks are different. Where a mark contributes to a report, the task needs conditions that make unaided production plausible. Conflating the two categories is what makes the whole system feel untrustworthy.
So exercises are built in two distinct classes:
- Open practice pages — unlimited attempts, immediate feedback, no monitoring
- Timed assessment pages — single sitting, invigilated conditions, integrity signals recorded
Students are told explicitly which is which. Ambiguity is what breeds resentment.
2. Feedback where the workload actually is
The marking burden is concentrated in open written production. The pages are therefore built to return structured feedback on objective components automatically — comprehension, vocabulary use, word count, task completion — so that teacher attention goes to the part only a teacher can judge: whether the argument holds, whether the register fits, whether the writing sounds like a person.
Automated grade calculation converts percentage scores to Noten on submission using the school’s own Punktetabelle. Student-facing output uses whole-number Noten with a floor applied — a deliberate choice, since a page that returns the lowest grade to a fourteen-year-old practising voluntarily is not doing pedagogical work.
3. Thematic coherence over grammar sequencing
Curriculum objectives are met, but organised around themes rather than structures. A unit on digital society carries conditionals, evaluative vocabulary and argument structure; a unit on friendship carries reported speech, present perfect and descriptive language. The grammar is scheduled; it simply is not the organising principle the student sees.
Exercise architecture
Each thematic unit progresses through four stages, with support withdrawn deliberately:
| Stage | Support level | Purpose |
|---|---|---|
| Vocabulary priming | Full — context, glossary | Remove decoding load before the main task |
| Guided comprehension | High — question scaffolds, model reasoning | Build confidence on authentic content |
| Supported production | Moderate — vocabulary bank, structural frame | Move from recognition to use |
| Independent production | Minimal or none | Generate evidence of unaided capability |
The last stage is the only one that needs securing.
A live example: a supported production task
This one is real rather than illustrative — it is Exercise D of California: Natural Hazards and Human Impact, currently in use with Year 9. The page is open: it can be worked through without a login or a name.
Theme: Natural hazards and human impact in California | Year 9, Gymnasium | Target level: B1
The unit builds up to it. Exercise A establishes the content — what California’s environmental problems actually are. Exercise B drills modal verbs in gapped sentences about wildfire risk. Exercise C practises reporting structures (“California is said to have the best weather in the US”). All three are auto-marked. By the time students reach Exercise D, both the vocabulary and the grammar are in place, and the writing task can ask for the thing the unit is really about: connections. How deforestation, urban sprawl and coastal building relate to mudslides, floods and water shortage.
The support is a word bank plus a reference grid of the hazards and human activities involved, and the task itself sets the constraint:
WORD BANK — cause and effect language
to affect to cause to be the consequence of
to contribute to to influence to lead to
to result in to show the link between
REFERENCE GRID
Natural hazards Human activities
earthquakes deforestation
tsunamis urbanisation of the coast
heavy rain / floods tourism / farming
drought urban sprawl
wildfires increase of traffic
mudslides soil erosion / water shortage
(a) Write a paragraph (100-150 words) explaining how natural hazards
and human activities in California are connected. Use at least
four phrases from the word bank above.
Hint: explain at least two different connections.
Starter shown in the box: "Deforestation contributes to mudslides
because..."
(b) In your opinion, which of California's environmental problems is
the most serious? Give reasons. Write 2-3 sentences.
Why the support is there: the word bank and the grid remove the retrieval problem, so the cognitive work lands on the actual target — identifying a causal relationship and stating it precisely. A student who cannot summon “to contribute to” under time pressure has not thereby failed to understand that clearing a hillside makes it slide. Requiring four phrases also forces variation: a paragraph built entirely on “because” does not meet the brief. Task (b) then withdraws the frame entirely and asks for a judgement in the student’s own words — the same content, one support level down.
Why this task stays open rather than secured: the Note the student sees comes from the auto-marked sections A to C. Exercise D carries no points — it is submitted for feedback, not for a grade. The learning happens in the attempt, and nothing rides on the mark.
Making written evidence verifiable
Where a task carries weight, the assessment pages run a set of client-side integrity measures — paste and shortcut blocking, non-selectable prompt text, tab and focus detection, a submission time gate, word-count and vocabulary enforcement, prompt randomisation, and an integrity record submitted with the work.
Rather than repeat that here, it is set out in full — along with the caveats it needs, and the Datenschutz questions it raises — in the companion post on academic integrity. Two points matter for this age group specifically.
Tell students first, in specific terms. A fifteen-year-old who discovers mid-task that their tab switching is logged will reasonably feel ambushed. The argument for the measures is one most students accept when it is made honestly; it does not survive being sprung on them.
Prompt design beats surveillance. Of everything in that toolkit, the measure that does most work is the least visible: four scenario variants distributed by a hash of the student’s name, so neighbours write different prompts. A prompt specific enough that a generic answer looks generic is worth more than any keyboard block.
Infrastructure
Task pages are static HTML on GitHub Pages, indexed from a WordPress hub. Submissions originally posted to Google Apps Script with an email fallback; that backend is being migrated to Excel Online and Power Automate following a data-protection restriction on Google services at school — a constraint German colleagues will recognise, and one worth designing for from the start rather than retrofitting.
Static hosting matters more than it sounds. There is no server to maintain, no login for students to lose, and no account data held anywhere. A page is a URL.
What I am still working on
- Peer review protocols, which reduce marking load and build metalinguistic awareness simultaneously
- Spaced vocabulary review carrying terms across thematic units
- Better in-class oral assessment formats, since oral work is the one thing no tool can produce for a student
The underlying argument
The current discussion in German schools tends to frame integrity measures and good pedagogy as opposites — surveillance versus trust. I do not think that holds. Clear conditions are what make a mark meaningful, and a meaningful mark is what makes feedback worth reading.
The question is not whether students will use AI. It is which tasks are supposed to prove something, and whether those tasks are built to do it.
Post 1 sets out the academic integrity framework this post refers to. Post 3 covers vocational IT English; post 4, Abitur writing preparation; post 5, MSA exam practice.
