AI Safety Evolved: Secure-by-Design, Safe-by-Measurement
Published a self-audit of Sublime’s AI safety posture: a maturity checklist for deploying autonomous AI against attacker-authored content, and the four pillars we hold ourselves to — treating the input as the adversary, calibrated (not merely confident) output, tiered autonomy applied to our own agents, and symmetric red-teaming.
Companion writeup: project page · Sublime blog post