Your assistant writes the prompt builder, the retrieval filter, the tool schema — and tells you it looks fine. This track trains the judgement to know when it is wrong, then hands you the pattern that holds.
Training a whole team?
Every exercise opens with your assistant's confident security summary. It's plausible, and it's wrong.
Click the line where the boundary actually breaks — across files, not just the one you expect.
Write the guardrail that blocks it, or the eval that catches it in CI — then watch the attack replay against your fix.
Each pass produces a file you can open as a PR — not a certificate of attendance.
AI-generated diffs that pass a normal review, each with a subtle flaw baked in — a secret in the prompt string, a filter applied after retrieval, a tool argument nobody validated. You decide ship or block, then name the reason.
Submit a real guardrail, schema or policy. We run an attack suite and a benign suite — ordinary requests that must still work — against it. Blocking every attack is not enough if the feature stops working.
A validator, a context builder, a CI eval suite. Copy it, download it, or open a PR against your own repository — so the training shows up in the codebase rather than in a completion report.
The free tier is 15 labs against a live model, no card required. Paid plans open every lab and the full curriculum.