MerchantryTidbits

Public catalog

evals-testing

25 shown. Search ranks against problem language; browsing defaults to stronger public signals.

evals-testingclifree

langfuse

observe LLM traces in production

Verified 2026-08-10 · 6,855 stars

evals-testingpatternfree

persona-lock-eval

evaluate whether an agent keeps its persona locked

Verified 2026-08-10 · 5,452 stars

evals-testingclifree

promptfoo

regression test LLM prompts and RAG outputs in CI

Verified 2026-08-10 · 5,452 stars

evals-testingpatternfree

prove-it-side-effect-check

require agents to prove side effects before claiming done

Verified 2026-08-10 · 5,452 stars

evals-testingpatternfree

calc-outside-llm

do exact math outside the LLM to avoid arithmetic errors

Verified 2026-08-10 · 5,099 stars

evals-testinglibraryfree

deepeval

pytest style LLM evaluation metrics

Verified 2026-08-10 · 3,114 stars

evals-testinglibraryfree

OpenLLMetry

OpenTelemetry traces for LLM and agent calls

Verified 2026-08-10 · 1,587 stars

evals-testinglibraryfree

AIConfig

config-driven prompts across model providers

Verified 2026-08-10 · 873 stars

evals-testinglibraryfree

Weave

trace multi-step LLM apps with W&B Weave

Verified 2026-08-10 · 545 stars

evals-testinglibraryfree

instructor

typed structured outputs from LLMs via pydantic

Verified 2026-08-10 · 0 stars

evals-testinglibraryfree

Cypress

write interactive end-to-end and component tests for web applications

Verified 2026-08-10

evals-testingskillfree

Eval Skills: Eval Audit

My LLM evaluation pipeline has gaps and I need a prioritized audit of likely problems.

Verified 2026-08-10

evals-testingclifree

EvalScope

You need to score an OpenAI-compatible model endpoint or a local model on standard benchmarks like GSM8K, MMLU, or SWE-bench with a single command instead of writing a harness

Verified 2026-08-10

evals-testinglibraryfree

Guardrails AI

Your LLM app can return toxic text, competitor mentions, or policy-violating content to users and you need programmatic input/output checks that block or raise on failure

Verified 2026-08-10

evals-testingclifree

GuideLLM

I need to load test an OpenAI-compatible LLM endpoint and find its maximum sustainable rate.

Verified 2026-08-10

evals-testingpatternfree

instruction-compliance-docs

use instruction-compliance-docs for evals testing

Verified 2026-08-10

evals-testingclifree

LightEval

You fine-tuned or quantized a model and need to score it on standard benchmarks like MMLU, GSM8K, or GPQA to compare against baselines.

Verified 2026-08-10

evals-testingclifree

LM Evaluation Harness

You fine-tuned or pretrained a model and need its scores on standard academic benchmarks (MMLU, HellaSwag, GSM8K) reported the same way papers and leaderboards report them

Verified 2026-08-10

evals-testingclilocal

LMMs-Eval

You need to benchmark a vision, video, or audio language model on standard suites like MMMU, MME, or VideoMME without hand-wiring each dataset and its post-processing

Verified 2026-08-10

evals-testinglibraryfree

MTEB (Massive Text Embedding Benchmark)

You must choose an embedding model for retrieval, classification, or clustering and want reproducible benchmark scores on relevant tasks instead of vendor claims

Verified 2026-08-10

evals-testingcliapi_key

Needle In A Haystack (niah)

You need to measure whether a model actually retrieves facts placed at specific depths as context length grows, instead of trusting the advertised context window

Verified 2026-08-10

evals-testinglibraryfree

phoenix

local tracing eval UI for LLM agent pipelines

Verified 2026-08-10

evals-testinglibraryfree

Ragas

evaluate RAG answer faithfulness and context recall

Verified 2026-08-10

evals-testingpatternfree

rules-hierarchy

use rules-hierarchy for evals testing

Verified 2026-08-10

evals-testinglibraryfree

skillopt

use skillopt for evals testing

Verified 2026-08-10