MerchantryTidbits

Public catalog

llm-inference

39 shown. Search ranks against problem language; browsing defaults to stronger public signals.

llm-inferencelibraryfree

litellm

model router fail over between providers

Verified 2026-08-10 · 53,182 stars

llm-inferenceclilocal

ollama

run open source models fully offline on a laptop

Verified 2026-08-10 · 43,694 stars

llm-inferencelibrarylocal

vllm

high throughput local model serving

Verified 2026-08-10 · 26,080 stars

llm-inferencelibraryfree

Tiktoken

count tokens for OpenAI-compatible models

Verified 2026-08-10 · 9,224 stars

llm-inferencelibraryfree

Llama Cpp Python

run GGUF models in-process via llama.cpp

Verified 2026-08-10 · 6,823 stars

llm-inferencelibraryfree

guidance

control generation with guidance grammars and templates

Verified 2026-08-10 · 5,917 stars

llm-inferencelibraryfree

outlines

constrain LLM outputs to JSON schemas or grammars

Verified 2026-08-10 · 3,496 stars

llm-inferencelibraryfree

Tokenizers

fast Hugging Face tokenization for training and RAG

Verified 2026-08-10 · 2,502 stars

llm-inferenceapipaid

Langsmith

trace LangChain and custom LLM runs

Verified 2026-08-10 · 345 stars

llm-inferencelibrarypaid

Pinecone Client

managed vector DB client for production RAG

Verified 2026-08-10 · 250 stars

llm-inferencelibrarylocal

AirLLM

A model you want to run locally is far larger than your GPU VRAM, for example 70B on a 4GB card, and you do not want to shrink it with quantization, distillation, or pruning

Verified 2026-08-10

llm-inferencelibrarypaid

Anthropic Python SDK

integrate Claude through the supported Python client

Verified 2026-08-10

llm-inferencelibraryapi_key

Claude SDK for TypeScript

You are building a Node.js or edge service that needs to send messages to Claude models and you do not want to hand-write fetch calls, request typing, and response parsing against the raw HTTP API

Verified 2026-08-10

llm-inferenceappfree

CLI Proxy API Management Center

You run a CLI Proxy API instance routing to Claude, Gemini, Codex, and OpenAI-compatible providers and need a UI to manage provider keys, model aliases, and quotas instead of hand-editing config.yaml

Verified 2026-08-10

llm-inferencelibraryfree

FastChat

serve and compare chat models behind an OpenAI-compatible endpoint

Verified 2026-08-10

llm-inferencelibrarypaid

Fireworks AI

call hosted open models through a Python SDK

Verified 2026-08-10

llm-inferencelibrarypaid

Google Gen AI Python SDK

integrate Gemini models using Google's current unified SDK

Verified 2026-08-10

llm-inferencelibraryfree

GPT4All

run a small local GGUF chat model from Python

Verified 2026-08-10

llm-inferencelibrarypaid

Groq Python SDK

use low-latency hosted inference through an OpenAI-style client

Verified 2026-08-10

llm-inferenceapipaid

Helicone

proxy OpenAI calls for cost and latency observability

Verified 2026-08-10

llm-inferenceclilocal

llama-swap

You run several local models with llama.cpp, vllm, or tabbyAPI but cannot keep them all loaded, and want one endpoint that starts and swaps the right server per request.

Verified 2026-08-10

llm-inferenceclilocal

llamafile

You want to hand colleagues an LLM they can run on macOS, Linux, BSD, or Windows by downloading one file and executing it, with no Python, drivers, or package installs.

Verified 2026-08-10

llm-inferencelibrarylocal

LLM Compressor

Your model is too large for available GPU VRAM and you need to quantize weights, activations, or KV cache before serving with vLLM

Verified 2026-08-10

llm-inferenceappfree

LM Studio

run and inspect local models through a desktop interface

Verified 2026-08-10

llm-inferencelibraryfree

MLC LLM

compile language models for laptops, phones, GPUs, or browsers

Verified 2026-08-10

llm-inferenceappfree

Ollama App

You run Ollama on a home server or workstation and want a native phone or desktop chat client instead of a browser tab

Verified 2026-08-10

llm-inferencelibrarylocal

Ollama Python Library

Your Python application needs chat, generation, or embeddings from models served by a local Ollama instance

Verified 2026-08-10

llm-inferencelibrarylocal

OllamaSharp

You are building a C# or .NET application that must chat with local Ollama models with streamed responses and tracked conversation history

Verified 2026-08-10

llm-inferencelibrarypaid

OpenAI Node SDK

integrate the OpenAI Responses API from TypeScript

Verified 2026-08-10

llm-inferencelibraryapi_key

openai-python

official Python SDK for OpenAI chat and embeddings APIs

Verified 2026-08-10

llm-inferenceapipaid

Openrouter

route chat completions across many hosted models

Verified 2026-08-10

llm-inferenceapipaid

Portkey

AI gateway for logging caching and fallbacks

Verified 2026-08-10

llm-inferencelibraryfree

Pymupdf

extract text and tables from PDF with PyMuPDF

Verified 2026-08-10

llm-inferencelibrarypaid

Replicate Python

run a published image, audio, or language model without hosting it

Verified 2026-08-10

llm-inferencelibraryfree

SGLang

serve language or multimodal models at high throughput

Verified 2026-08-10

llm-inferenceserverfree

Text Generation Inference

serve supported Hugging Face models with continuous batching

Verified 2026-08-10

llm-inferencelibrarypaid

Together Python

call hosted open models through an OpenAI-style Python client

Verified 2026-08-10

llm-inferencelibraryfree

Transformers.js

You want text classification, embeddings, translation, object detection, or speech recognition inside a web app without standing up any inference backend

Verified 2026-08-10

llm-inferencelibraryfree

Vercel AI SDK

build a streaming AI interface in TypeScript

Verified 2026-08-10