Four engineering pillars, a maturity checklist, and the concrete artifacts, gates, and numbers behind Sublime's AI safety program.
#evaluation
Content tagged with "evaluation"
AI Safety Evolved: Secure-by-Design, Safe-by-Measurement
MQL Benchmark
A 30,000-example open-source benchmark for evaluating natural-language → DSL generation, with a public model leaderboard.
Trust, Then Autonomy
A framework for evaluating earned autonomy in deployed AI systems.
Evaluating LLM Generated Detection Rules in Cybersecurity
An open-source evaluation framework and three benchmark metrics for measuring LLM-generated cybersecurity detection rules.
CAMLIS 2025: Evaluating LLM-Generated Detection Rules
Paper accepted at CAMLIS 2025 — an open-source benchmark and three metrics (detection accuracy, economic cost of syntactic correctness, robustness of query) for measuring LLM-generated security rules.
Evaluating LLM-Generated Detection Rules
A benchmark and three metrics for measuring LLM-generated cybersecurity rules — CAMLIS 2025.