A label-free evaluation framework for continual learning harnesses in cybersecurity, using teacher-relative lift as a proxy for uplift against a held-out gold standard.
#evaluation
Content tagged with "evaluation"
Evaluating Agentic Learning Harness Capabilities Without Labels via the Scaling Hypothesis
AI Safety Evolved: Secure-by-Design, Safe-by-Measurement
Four engineering pillars, a maturity checklist, and the concrete artifacts, gates, and numbers behind Sublime's AI safety program.
MQL Benchmark
A 30,000-example open-source benchmark for evaluating natural-language → DSL generation, with a public model leaderboard.
Trust, Then Autonomy
A framework for evaluating earned autonomy in deployed AI systems.
More Than 'Plausible Nonsense': A Rigorous Eval for ADÉ, Our Security Coding Agent
A three-pillar framework — detection accuracy, robustness, and economic cost of coverage — for evaluating LLM-generated detection rules, applied to Sublime's ADÉ agent.
Evaluating LLM Generated Detection Rules in Cybersecurity
An open-source evaluation framework and three benchmark metrics for measuring LLM-generated cybersecurity detection rules.
Evaluating LLM-Generated Detection Rules
A benchmark and three metrics for measuring LLM-generated cybersecurity rules — CAMLIS 2025.