Public catalog
Browse tools
3 shown. Search ranks against problem language; browsing defaults to stronger public signals.
EvalScope
You need to score an OpenAI-compatible model endpoint or a local model on standard benchmarks like GSM8K, MMLU, or SWE-bench with a single command instead of writing a harness
Verified 2026-08-10
LM Evaluation Harness
You fine-tuned or pretrained a model and need its scores on standard academic benchmarks (MMLU, HellaSwag, GSM8K) reported the same way papers and leaderboards report them
Verified 2026-08-10
LMMs-Eval
You need to benchmark a vision, video, or audio language model on standard suites like MMMU, MME, or VideoMME without hand-wiring each dataset and its post-processing
Verified 2026-08-10