Four engineering pillars, a maturity checklist, and the concrete artifacts, gates, and numbers behind Sublime's AI safety program.
Projects
Frameworks, agents, and open-source work.
AI Safety Evolved: Secure-by-Design, Safe-by-Measurement
Prompt Injection in the Wild
A real-world phishing campaign carrying an adversarial prompt injection payload against AI-based email security — and how ASA caught it.
MQL Benchmark
A 30,000-example open-source benchmark for evaluating natural-language → DSL generation, with a public model leaderboard.
Trust, Then Autonomy
A framework for evaluating earned autonomy in deployed AI systems.
Secure-by-Design Agents
How we architected ASA and ADÉ for adversarial production.
Evaluating LLM-Generated Detection Rules
A benchmark and three metrics for measuring LLM-generated cybersecurity rules — CAMLIS 2025.
BabbelPhish
Accelerating Adoption of Domain-Specific Languages with Large Language Models.
BabbelPhish Dataset
Open-source natural language to domain-specific language dataset for email security.
MalwareRL
Malware bypass research using reinforcement learning.