The AI Gateway
for platform teams
Self-host in minutes. No credit card.

“LiteLLM gives NVIDIA engineers a single, consistent way to access more than 100 AI model endpoints.”
“LiteLLM streamlines the complexities of managing multiple LLM models.”
“LiteLLM has let my team provide the latest LLM models to our users, usually within a day of them being released… it has saved us months of work.”
“If we decide to switch the backend model, it’s a simple configuration update in the gateway; no code changes, procurement cycles, or repetitive security reviews required.”

Give your whole org access to every model, agent, and MCP, then see, control, and optimize every request.
Access: Give your whole company every model.
Put every model, agent, and MCP behind one API and one login. We handle key management and work with the secret manager you already run, so platform teams open access to the whole org without becoming the bottleneck, and developers build in minutes, not a procurement cycle.
- One OpenAI-compatible API to 140+ providers and 1,800+ models. Swap models without changing app code.
- One login with SSO, scoped by team, project, or app.
- Day-zero support for new models, so developers get the latest the day it ships.
- Reach agents and MCP servers through the same gateway, not just LLMs.
- Bring your own internal, fine-tuned, and self-hosted models behind the same key.
Visibility & Control: See every request, and cap it before it runs.
See who and what is driving usage and spend, attribute every request for chargeback, and cap budgets before they run. Set budgets and rate limits per team; when they hit the cap, requests stop.
- Usage and spend tracked per key, user, team, org, tool, agent, and MCP, across 140+ providers.
- Enterprise chargeback: attribute every request and bill teams and business units for what they use.
- Hard budgets per key, team, org, and model, with daily and monthly resets. At the cap, requests stop.
- Rate limits and leaked-key protection, so no runaway job or compromised key runs up your bill.
- Model access control and guardrails, with an audit log on every request.
Cost optimization: Maximize the ROI of your AI.
Swap models without changing a line of code, and let Auto Routing send each request to the model that should handle it, so the budget you set goes further.
- Load balancing across providers, regions, and keys.
- Lowest-cost routing to the cheapest deployment that can serve the request.
- Auto Routing that sends simple prompts to cheaper models and hard ones to stronger models.
- Response and semantic caching (Redis, S3, GCS), so you never pay twice for the same answer.
- Prompt compression, so you send fewer tokens for the same result.
Ease of deployment: Go live in your stack in an afternoon
Self-host the same open-source gateway behind 240M+ Docker pulls, in your own cloud or fully air-gapped. Simple enough that it just works.
- Deploy with an official Helm chart or Terraform module
- Official Docker images, including database-bundled and non-root variants
- Runs on your own Postgres and Redis, and scales out with Kubernetes autoscaling
- Self-host anywhere, including fully air-gapped environments
- One-click deploy into your hyperscaler (AWS, GCP, or Azure)
Sub-millisecond overhead. Read the benchmark.
The LiteLLM Rust AI Gateway is live. On our benchmarks it adds 0.66 ms at p99 — 3.5× lower overhead than the next AI gateway — measured with AI Gateway Bench, an open standard for benchmarking and comparing AI gateways. Run it yourself.
Identical hardware; every gateway pointed at the same deterministic mock upstream; single client.
Same benchmark run; resident memory with the gateway idle.
Measured on identical hardware, with every gateway pointed at the same deterministic upstream, so only the gateway’s own overhead is left. Every number, chart, and script is public.
Learn more about our plans.
- 140+ LLM provider integrations
- Langfuse, Arize Phoenix, LangSmith, and OTEL logging
- Virtual keys, budgets, and teams
- Load balancing and RPM/TPM limits
- LLM guardrails
- Everything in open source
- Enterprise support and custom SLAs
- JWT auth, SSO, and audit logs
- All enterprise features. See the docs.
Your keys, your infra, your audit trail.
Security fixes, stable releases, and what we're working on next. All public. Don't take our word for it; read the changelog and the code.
- Every release, in the open — see exactly what changed, in the public changelog.
- A stable release track — production images ship after load testing.
- Security in the open — when there's an issue, we publish the disclosure and the fix. You can read every one.
- Read the code your keys flow through — open an issue, send a PR, or fork it.
- See what we're building — the work in progress lives in our docs and engineering blog.
- No lock-in — the open-source gateway is MIT-licensed and production-grade. Enterprise adds SSO, RBAC, audit logs, and support on top, not a different core.
The security controls a review will check for:
Cosign-signed, hardened non-root images — verify provenance before deploy
Vulnerability-scanned (Grype, zero high or critical) with CodeQL on the codebase
SSO, JWT auth, and RBAC through your identity provider, with SCIM provisioning
PII masking and prompt-injection guardrails (Presidio, Lakera, and more)
Secrets pulled from AWS Secrets Manager, Vault, or Azure Key Vault, never hardcoded
Self-hosted with no telemetry, so your data never leaves your infrastructure

Run it yourself. Start today.
Deploy the open-source gateway in an afternoon, with spend tracked and capped on every request. Add SSO, audit logs, and an SLA when it goes org-wide.





