Public catalog
evals-testing
25 shown. Search ranks against problem language; browsing defaults to stronger public signals.
langfuse
observe LLM traces in production
Verified 2026-08-10 · 6,855 stars
persona-lock-eval
evaluate whether an agent keeps its persona locked
Verified 2026-08-10 · 5,452 stars
promptfoo
regression test LLM prompts and RAG outputs in CI
Verified 2026-08-10 · 5,452 stars
prove-it-side-effect-check
require agents to prove side effects before claiming done
Verified 2026-08-10 · 5,452 stars
calc-outside-llm
do exact math outside the LLM to avoid arithmetic errors
Verified 2026-08-10 · 5,099 stars
deepeval
pytest style LLM evaluation metrics
Verified 2026-08-10 · 3,114 stars
OpenLLMetry
OpenTelemetry traces for LLM and agent calls
Verified 2026-08-10 · 1,587 stars
AIConfig
config-driven prompts across model providers
Verified 2026-08-10 · 873 stars
Weave
trace multi-step LLM apps with W&B Weave
Verified 2026-08-10 · 545 stars
instructor
typed structured outputs from LLMs via pydantic
Verified 2026-08-10 · 0 stars
Cypress
write interactive end-to-end and component tests for web applications
Verified 2026-08-10
Eval Skills: Eval Audit
My LLM evaluation pipeline has gaps and I need a prioritized audit of likely problems.
Verified 2026-08-10
EvalScope
You need to score an OpenAI-compatible model endpoint or a local model on standard benchmarks like GSM8K, MMLU, or SWE-bench with a single command instead of writing a harness
Verified 2026-08-10
Guardrails AI
Your LLM app can return toxic text, competitor mentions, or policy-violating content to users and you need programmatic input/output checks that block or raise on failure
Verified 2026-08-10
GuideLLM
I need to load test an OpenAI-compatible LLM endpoint and find its maximum sustainable rate.
Verified 2026-08-10
instruction-compliance-docs
use instruction-compliance-docs for evals testing
Verified 2026-08-10
LightEval
You fine-tuned or quantized a model and need to score it on standard benchmarks like MMLU, GSM8K, or GPQA to compare against baselines.
Verified 2026-08-10
LM Evaluation Harness
You fine-tuned or pretrained a model and need its scores on standard academic benchmarks (MMLU, HellaSwag, GSM8K) reported the same way papers and leaderboards report them
Verified 2026-08-10
LMMs-Eval
You need to benchmark a vision, video, or audio language model on standard suites like MMMU, MME, or VideoMME without hand-wiring each dataset and its post-processing
Verified 2026-08-10
MTEB (Massive Text Embedding Benchmark)
You must choose an embedding model for retrieval, classification, or clustering and want reproducible benchmark scores on relevant tasks instead of vendor claims
Verified 2026-08-10
Needle In A Haystack (niah)
You need to measure whether a model actually retrieves facts placed at specific depths as context length grows, instead of trusting the advertised context window
Verified 2026-08-10
phoenix
local tracing eval UI for LLM agent pipelines
Verified 2026-08-10
Ragas
evaluate RAG answer faithfulness and context recall
Verified 2026-08-10
rules-hierarchy
use rules-hierarchy for evals testing
Verified 2026-08-10
skillopt
use skillopt for evals testing
Verified 2026-08-10