LLM Jailbreaks & Defenses
Automated jailbreaks, prompt attacks, and the defenses built to catch or contain them.
Explore
Browse research by topic, paper trail, and recurring adversarial ML themes.
Automated jailbreaks, prompt attacks, and the defenses built to catch or contain them.
How search, evolution, gradients, and black-box optimization discover failures in ML systems.
Security risks in the AI software stack, from model dependencies and agent tools to package ecosystems, plugins, and deployment pipelines.
Security risks that emerge when AI systems can plan, call tools, write code, use memory, coordinate tasks, or act with partial autonomy.
Attacks against image classifiers, perception systems, and vision-model assumptions.
How reward functions, agents, environments, and evaluators can be gamed, exploited, or misaligned under optimization pressure.
A ground-up series on why AI systems are vulnerable by design, not by accident. Starting from the decision boundary every model relies on, this series builds the theory needed to understand every major AI attack category, and what can actually be done about each one.