Public catalog
llm-inference
39 shown. Search ranks against problem language; browsing defaults to stronger public signals.
litellm
model router fail over between providers
Verified 2026-08-10 · 53,182 stars
ollama
run open source models fully offline on a laptop
Verified 2026-08-10 · 43,694 stars
vllm
high throughput local model serving
Verified 2026-08-10 · 26,080 stars
Tiktoken
count tokens for OpenAI-compatible models
Verified 2026-08-10 · 9,224 stars
Llama Cpp Python
run GGUF models in-process via llama.cpp
Verified 2026-08-10 · 6,823 stars
guidance
control generation with guidance grammars and templates
Verified 2026-08-10 · 5,917 stars
outlines
constrain LLM outputs to JSON schemas or grammars
Verified 2026-08-10 · 3,496 stars
Tokenizers
fast Hugging Face tokenization for training and RAG
Verified 2026-08-10 · 2,502 stars
Langsmith
trace LangChain and custom LLM runs
Verified 2026-08-10 · 345 stars
Pinecone Client
managed vector DB client for production RAG
Verified 2026-08-10 · 250 stars
AirLLM
A model you want to run locally is far larger than your GPU VRAM, for example 70B on a 4GB card, and you do not want to shrink it with quantization, distillation, or pruning
Verified 2026-08-10
Anthropic Python SDK
integrate Claude through the supported Python client
Verified 2026-08-10
Claude SDK for TypeScript
You are building a Node.js or edge service that needs to send messages to Claude models and you do not want to hand-write fetch calls, request typing, and response parsing against the raw HTTP API
Verified 2026-08-10
CLI Proxy API Management Center
You run a CLI Proxy API instance routing to Claude, Gemini, Codex, and OpenAI-compatible providers and need a UI to manage provider keys, model aliases, and quotas instead of hand-editing config.yaml
Verified 2026-08-10
FastChat
serve and compare chat models behind an OpenAI-compatible endpoint
Verified 2026-08-10
Fireworks AI
call hosted open models through a Python SDK
Verified 2026-08-10
Google Gen AI Python SDK
integrate Gemini models using Google's current unified SDK
Verified 2026-08-10
GPT4All
run a small local GGUF chat model from Python
Verified 2026-08-10
Groq Python SDK
use low-latency hosted inference through an OpenAI-style client
Verified 2026-08-10
Helicone
proxy OpenAI calls for cost and latency observability
Verified 2026-08-10
llama-swap
You run several local models with llama.cpp, vllm, or tabbyAPI but cannot keep them all loaded, and want one endpoint that starts and swaps the right server per request.
Verified 2026-08-10
llamafile
You want to hand colleagues an LLM they can run on macOS, Linux, BSD, or Windows by downloading one file and executing it, with no Python, drivers, or package installs.
Verified 2026-08-10
LLM Compressor
Your model is too large for available GPU VRAM and you need to quantize weights, activations, or KV cache before serving with vLLM
Verified 2026-08-10
LM Studio
run and inspect local models through a desktop interface
Verified 2026-08-10
MLC LLM
compile language models for laptops, phones, GPUs, or browsers
Verified 2026-08-10
Ollama App
You run Ollama on a home server or workstation and want a native phone or desktop chat client instead of a browser tab
Verified 2026-08-10
Ollama Python Library
Your Python application needs chat, generation, or embeddings from models served by a local Ollama instance
Verified 2026-08-10
OllamaSharp
You are building a C# or .NET application that must chat with local Ollama models with streamed responses and tracked conversation history
Verified 2026-08-10
OpenAI Node SDK
integrate the OpenAI Responses API from TypeScript
Verified 2026-08-10
openai-python
official Python SDK for OpenAI chat and embeddings APIs
Verified 2026-08-10
Openrouter
route chat completions across many hosted models
Verified 2026-08-10
Portkey
AI gateway for logging caching and fallbacks
Verified 2026-08-10
Pymupdf
extract text and tables from PDF with PyMuPDF
Verified 2026-08-10
Replicate Python
run a published image, audio, or language model without hosting it
Verified 2026-08-10
SGLang
serve language or multimodal models at high throughput
Verified 2026-08-10
Text Generation Inference
serve supported Hugging Face models with continuous batching
Verified 2026-08-10
Together Python
call hosted open models through an OpenAI-style Python client
Verified 2026-08-10
Transformers.js
You want text classification, embeddings, translation, object detection, or speech recognition inside a web app without standing up any inference backend
Verified 2026-08-10
Vercel AI SDK
build a streaming AI interface in TypeScript
Verified 2026-08-10