reading guardian…
guardian.runguardian.dev
Eval suites created, run and tracked — pass rates, regressions, and what changed between them.
Datasets
not deployed yet
Versions, lineage and size for every dataset a run or an eval was pointed at.
Every agent and LLM call, timed and recorded — the same watching, one level down.
Reusable skills, versioned and callable — with the metrics for how each one performs.
Add a guardian
a subdomain, a tile, one login