proband.xyz labs

Labs.

Hands-on walk-throughs on attacking and defending LLM agents. Each one you can run on your own hardware.

  1. 01 Postcondition-only prompting. An attack on tool-using agents that hides the action verb in the constraint set, the four validators that miss it, and the defenses that compose against it. prompt-injection defenses agents ~25 min + labs
  2. 02 Catch malicious PDFs. Train a small adapter that teaches an open 22-billion-parameter model to catch a directive hidden inside a document, then verify the defense is real and not a benchmark grading its own work. training prompt-injection defenses ~20 min + labs