Self-hosted and air-gapped.

The Enterprise AI Gateway

Gateways
LLM GatewayMCP GatewayAgent Gateway
Controls
Budgets & Rate LimitsBuilt-in and 3rd Party GuardrailsRBAC & Usage Tracking by Key, Team, Org
Operations
Logging to Datadog & OpenTelemetryOIDC, SSO, Custom Authentication

Self-host in minutes. No credit card.

Run in production by the teams shipping AI at scale.
Testimonial
Testimonial
Testimonial

“LiteLLM gives NVIDIA engineers a single, consistent way to access more than 100 AI model endpoints.”

Ajay Dogra
Ajay DograProduct, NVIDIA
Testimonial
Testimonial

“LiteLLM streamlines the complexities of managing multiple LLM models.”

Mark Koltnuk
Mark KoltnukPrincipal Architect, Lemonade
Testimonial
Testimonial
Testimonial

“LiteLLM has let my team provide the latest LLM models to our users, usually within a day of them being released… it has saved us months of work.”

David Leen
David LeenStaff Software Engineer, Netflix
Testimonial
Testimonial

“If we decide to switch the backend model, it’s a simple configuration update in the gateway; no code changes, procurement cycles, or repetitive security reviews required.”

Dennis Henry
Dennis HenryProductivity Architect, Okta
Testimonial
LiteLLM abstract hero background
01

Deploy in any environment

Self-hosted, on-prem, or in your cloud.

On-prem
Multi-Cloud
Kubernetes
Helm
02

Sub-millisecond overhead. Read the benchmark.

The LiteLLM Rust AI Gateway is live. On our benchmarks it adds 0.66 ms at p99 — 3.5× lower overhead than the next AI gateway — measured with AI Gateway Bench, an open standard for benchmarking and comparing AI gateways. Run it yourself.

AI gateway benchmark results measured with AI Gateway Bench. LiteLLM (Rust) adds 0.66 ms at p99. Throughput 2,800+ req/s (at ~21% CPU); ~4.5× more requests per dollar. Identical hardware; every gateway pointed at the same deterministic mock upstream; single client.
AI gatewayp99 added latencyMemory at restThroughputCPU utilization
LiteLLM (Rust)0.66 ms~22 MB2,800+ req/s~21%
Portkey2.29 ms
Bifrost4.54 ms~199 MB
0.66 ms
p99 added latency
~22 MB
memory at rest
2,800+ req/s
at ~21% CPU
~4.5×
more requests per dollar
Overhead the gateway adds — p99 ms (lower is better)
LiteLLM (Rust)0.66 ms
Portkey2.29 ms
Bifrost4.54 ms

Identical hardware; every gateway pointed at the same deterministic mock upstream; single client.

Memory at rest — peak RSS (lower is better)
LiteLLM (Rust)~22 MB
Bifrost~199 MB

Same benchmark run; resident memory with the gateway idle.

Measured on identical hardware, with every gateway pointed at the same deterministic upstream, so only the gateway’s own overhead is left. Every number, chart, and script is public.

See the benchmark
140+
LLM providers
1,892
Unique models
240M+
Docker pulls
1B+
Requests served
53K+
GitHub stars
1,005+
GitHub contributors
Running in production
Single access source with consistency
LiteLLM gives NVIDIA engineers a single, consistent way to access more than 100 AI model endpoints across cloud providers, open source deployments, and internal NVIDIA services.
Ajay Dogra
AD
Ajay Dogra
Product
New models in a day, not hours of rework
LiteLLM has let my team provide the latest LLM models to our users, usually within a day of them being released… it has saved us months of work.
David Leen
David Leen
Staff Software Engineer
Multiple models, one interface
LiteLLM streamlines the complexities of managing multiple LLM models.
Mark Koltnuk
Mark Koltnuk
Principal Architect, GenAI Platform
Swap models without new security reviews
If we decide to switch the backend model, it’s a simple configuration update in the gateway; no code changes, procurement cycles, or repetitive security reviews required.
Dennis Henry
Dennis Henry
Productivity Architect
Data centre server hall
03

Run it yourself. Start today.

Deploy the open-source gateway in an afternoon, with spend tracked and capped on every request. Add SSO, audit logs, and an SLA when it goes org-wide.