Public catalog
Browse tools
3 shown. Search ranks against problem language; browsing defaults to stronger public signals.
persona-lock-eval
evaluate whether an agent keeps its persona locked
Verified 2026-08-10 · 5,452 stars
deepeval
pytest style LLM evaluation metrics
Verified 2026-08-10 · 3,114 stars
NVIDIA NeMo Guardrails
Your LLM chat application must block jailbreaks and prompt injections and moderate or mask sensitive data in user input and model output before it reaches the user
Verified 2026-08-10