LLM API provider directory

25 provider profiles with model lineups, pricing notes, access routes and source boundaries. Use the evidence to shortlist a provider; a profile is not an endorsement.

OpenAI openai

GPT-4o, GPT-4 Turbo, o1 reasoning models

View OpenAI profile →
Anthropic anthropic

Claude Sonnet 4.5, Opus 4.6, Haiku 4.5

View Anthropic profile →
Google Gemini google

Gemini 2.5 Pro, 2.5 Flash, Flash-8B

View Google Gemini profile →
Mistral AI mistral

Mistral Large, Codestral, Pixtral

View Mistral AI profile →
Meta Llama meta

Llama 3.3, Llama 3.2 Vision, Llama 4

View Meta Llama profile →
xAI Grok xai

Grok-4.5, Grok-4.3, Grok-4.20

View xAI Grok profile →
Groq groq

Sub-100ms inference for Llama, Mixtral, Whisper

View Groq profile →
Together AI together-ai

200+ open-weights models — Llama, Qwen, DeepSeek

View Together AI profile →
Fireworks AI fireworks

Fast inference for Llama, DeepSeek, function calling

View Fireworks AI profile →
Perplexity Sonar perplexity

Sonar Pro — web-grounded answers with citations

View Perplexity Sonar profile →
Cohere cohere

Command A, Command A+, Rerank

View Cohere profile →
AI21 Labs ai21

Jamba 1.5 Large — hybrid Mamba-Transformer

View AI21 Labs profile →
AWS Bedrock aws-bedrock

Claude, Llama, Titan, Mistral — through your AWS account

View AWS Bedrock profile →
Azure OpenAI azure-openai

GPT-4o on Azure — your Azure subscription, your data residency

View Azure OpenAI profile →
Google Vertex AI vertex-ai

Gemini + Claude + Llama on GCP — your project, your residency

View Google Vertex AI profile →
OpenRouter openrouter

300+ models across providers — universal router

View OpenRouter profile →
DeepInfra deepinfra

Cheap open-weights inference — Llama, Qwen, Mixtral

View DeepInfra profile →
Replicate replicate

Run any open-weights model — Llama, FLUX, Whisper, ComfyUI

View Replicate profile →
Cerebras cerebras

Wafer-scale inference — fastest tokens/sec on the market

View Cerebras profile →
Lambda Labs lambdalabs

GPU-cloud inference — Hermes 3, Llama 3.3

View Lambda Labs profile →
OctoAI octoai

Optimized inference — Llama, Stable Diffusion, custom

View OctoAI profile →
Lepton AI lepton

Production-grade open-weights inference

View Lepton AI profile →
NVIDIA NIM nvidia-nim

NVIDIA-optimized microservices for open-weights LLMs

View NVIDIA NIM profile →
Databricks Mosaic databricks-mosaic

DBRX, Llama, Mixtral served on Databricks

View Databricks Mosaic profile →
AI21 Jamba jamba

Hybrid Mamba-Transformer with 256K context — open weights

View AI21 Jamba profile →