Public catalog
Browse tools
5 shown. Search ranks against problem language; browsing defaults to stronger public signals.
Eval Skills: Eval Audit
My LLM evaluation pipeline has gaps and I need a prioritized audit of likely problems.
Verified 2026-08-10
EvalScope
You need to score an OpenAI-compatible model endpoint or a local model on standard benchmarks like GSM8K, MMLU, or SWE-bench with a single command instead of writing a harness
Verified 2026-08-10
GuideLLM
I need to load test an OpenAI-compatible LLM endpoint and find its maximum sustainable rate.
Verified 2026-08-10
LightEval
You fine-tuned or quantized a model and need to score it on standard benchmarks like MMLU, GSM8K, or GPQA to compare against baselines.
Verified 2026-08-10
LMMs-Eval
You need to benchmark a vision, video, or audio language model on standard suites like MMMU, MME, or VideoMME without hand-wiring each dataset and its post-processing
Verified 2026-08-10