MerchantryTidbits

Public catalog

Browse tools

54 shown. Search ranks against problem language; browsing defaults to stronger public signals.

llm-inferencelibraryfree

litellm

model router fail over between providers

Verified 2026-08-10 · 53,182 stars

llm-inferenceclilocal

ollama

run open source models fully offline on a laptop

Verified 2026-08-10 · 43,694 stars

rag-retrievalclilocal

open-webui

self-host ChatGPT-like UI over local models

Verified 2026-08-10 · 27,370 stars

llm-inferencelibrarylocal

vllm

high throughput local model serving

Verified 2026-08-10 · 26,080 stars

evals-testingclifree

langfuse

observe LLM traces in production

Verified 2026-08-10 · 6,855 stars

llm-inferencelibraryfree

Llama Cpp Python

run GGUF models in-process via llama.cpp

Verified 2026-08-10 · 6,823 stars

scrapingclifree

lightpanda

use lightpanda for scraping

Verified 2026-08-10 · 4,292 stars

rag-retrievalclilocal

text-embeddings-inference

use text-embeddings-inference for rag retrieval

Verified 2026-08-10 · 2,251 stars

scrapinglibraryfree

headroom

use headroom for scraping

Verified 2026-08-10 · 1,814 stars

evals-testinglibraryfree

OpenLLMetry

OpenTelemetry traces for LLM and agent calls

Verified 2026-08-10 · 1,587 stars

rag-retrievallibraryfree

FastEmbed

need fast local embeddings without heavy torch

Verified 2026-08-10 · 1,000 stars

deploy-infraclifree

mission-control

use mission-control for deploy infra

Verified 2026-08-10 · 975 stars

llm-inferenceapipaid

Langsmith

trace LangChain and custom LLM runs

Verified 2026-08-10 · 345 stars

rag-retrievallibraryfree

Batched

batch embedding requests to cut API cost

Verified 2026-08-10 · 1 stars

llm-inferencelibrarypaid

Anthropic Python SDK

integrate Claude through the supported Python client

Verified 2026-08-10

rag-retrievalappfree

AnythingLLM

stand up private document chat without assembling a RAG stack

Verified 2026-08-10

agent-orchestrationlibraryfree

AutoGen

build event-driven agents with tools and explicit team patterns

Verified 2026-08-10

deploy-infralibraryfree

BentoML

package a Python model behind a production HTTP service

Verified 2026-08-10

media-processinglibrarylocal

Chatterbox

generate natural text to speech locally for narration and voice workflows

Verified 2026-08-10

agent-orchestrationlibraryfree

CrewAI

coordinate role-based agents across a bounded workflow

Verified 2026-08-10

doc-parsinglibraryfree

Easyocr

OCR images and screenshots with EasyOCR offline

Verified 2026-08-10

rag-retrievallibraryfree

Embedchain

prototype document ingestion and question answering with minimal glue

Verified 2026-08-10

llm-inferencelibraryfree

FastChat

serve and compare chat models behind an OpenAI-compatible endpoint

Verified 2026-08-10

llm-inferencelibrarypaid

Fireworks AI

call hosted open models through a Python SDK

Verified 2026-08-10

rag-retrievallibrarylocal

flag-embedding

use flag-embedding for rag retrieval

Verified 2026-08-10

llm-inferencelibrarypaid

Google Gen AI Python SDK

integrate Gemini models using Google's current unified SDK

Verified 2026-08-10

llm-inferencelibraryfree

GPT4All

run a small local GGUF chat model from Python

Verified 2026-08-10

llm-inferencelibrarypaid

Groq Python SDK

use low-latency hosted inference through an OpenAI-style client

Verified 2026-08-10

rag-retrievallibraryfree

Haystack

compose typed retrieval and generation components into a pipeline

Verified 2026-08-10

llm-inferenceapipaid

Helicone

proxy OpenAI calls for cost and latency observability

Verified 2026-08-10

media-processingapilocal

Kokoro-FastAPI

You need an OpenAI-compatible text-to-speech endpoint running on your own hardware so audio never leaves the machine and there is no per-character bill

Verified 2026-08-10

agent-orchestrationlibraryfree

LangChain.js

compose model, tool, retrieval, and streaming steps in TypeScript

Verified 2026-08-10

llm-inferenceappfree

LM Studio

run and inspect local models through a desktop interface

Verified 2026-08-10

research-workflowapplocal

Local Deep Researcher

You need iterative multi-cycle web research on a topic, with gap analysis and follow-up queries, condensed into one cited markdown summary instead of manually running and pasting searches

Verified 2026-08-10

llm-inferencelibraryfree

MLC LLM

compile language models for laptops, phones, GPUs, or browsers

Verified 2026-08-10

deploy-infralibrarypaid

Modal

run bursty Python or GPU jobs without managing clusters

Verified 2026-08-10

llm-inferencelibrarypaid

OpenAI Node SDK

integrate the OpenAI Responses API from TypeScript

Verified 2026-08-10

auth-integrationclipaid

opencode-openai-codex-auth

You already pay for ChatGPT Plus or Pro and want OpenCode to use that subscription for GPT-5.x and Codex models instead of racking up per-token OpenAI API charges

Verified 2026-08-10

llm-inferenceapipaid

Openrouter

route chat completions across many hosted models

Verified 2026-08-10

scrapingclifree

OpenSERP

Your agent needs live Google or Bing results as structured JSON but a paid SERP API is too expensive per call

Verified 2026-08-10

doc-parsinglibraryfree

Pdfplumber

extract PDF tables and text with pdfplumber

Verified 2026-08-10

media-processingclilocal

Piper

generate speech locally without per-character API charges

Verified 2026-08-10

llm-inferenceapipaid

Portkey

AI gateway for logging caching and fallbacks

Verified 2026-08-10

rag-retrievalappfree

PrivateGPT

run document question answering without sending files to a hosted service

Verified 2026-08-10

llm-inferencelibraryfree

Pymupdf

extract text and tables from PDF with PyMuPDF

Verified 2026-08-10

llm-inferencelibrarypaid

Replicate Python

run a published image, audio, or language model without hosting it

Verified 2026-08-10

agent-orchestrationlibraryfree

Semantic Kernel

expose typed application functions to a model

Verified 2026-08-10

llm-inferencelibraryfree

SGLang

serve language or multimodal models at high throughput

Verified 2026-08-10

llm-inferenceserverfree

Text Generation Inference

serve supported Hugging Face models with continuous batching

Verified 2026-08-10

llm-inferencelibrarypaid

Together Python

call hosted open models through an OpenAI-style Python client

Verified 2026-08-10

ml-operationslibraryfree

Transformers

load, fine-tune, and evaluate pretrained transformer models

Verified 2026-08-10

deploy-infralibraryfree

Truss

package a custom model with its Python and system dependencies

Verified 2026-08-10

rag-retrievallibraryfree

txtai

add semantic search to a Python dataset with a small API

Verified 2026-08-10

llm-inferencelibraryfree

Vercel AI SDK

build a streaming AI interface in TypeScript

Verified 2026-08-10