Labs.
Hands-on walk-throughs on attacking and defending LLM agents. Each one you can run on your own hardware.
- 01 Postcondition-only prompting. An attack on tool-using agents that hides the action verb in the constraint set, the four validators that miss it, and the defenses that compose against it.
- 02 Catch malicious PDFs. Train a small adapter that teaches an open 22-billion-parameter model to catch a directive hidden inside a document, then verify the defense is real and not a benchmark grading its own work.