Ollama
Run open LLMs locally with a single command.
Open-source self-hosted AI gateway putting 31 providers behind one endpoint
GoModel is aimed squarely at teams that want OpenRouter-style consolidation without the dependency on a third party sitting in the request path. The MIT core is unusually complete for an open-source gateway: budgets, virtual keys, audit logs, and failover are all in the free tier rather than held back. Air-gapped operation makes it viable in environments where a hosted gateway is simply not allowed. The Pro tier is expensive at 499 dollars per month, and its headline features, prompt compression and SSO, are narrower than the price suggests, so most teams will get what they need from the free edition. Being self-hosted, the operational burden is yours.
GoModel is an MIT-licensed, self-hosted AI gateway in Go that fronts 31 LLM providers behind OpenAI and Anthropic compatible endpoints with caching, failover, budgets, virtual keys, and audit logs.
GoModel positions itself as an open-source, self-hosted alternative to hosted routing services. The core is MIT licensed and deploys as a single Go binary, via Docker Compose, or through Helm, and the vendor states it works air-gapped and never phones home, which is the point of difference against a managed gateway. Functionally it consolidates 31 providers, including OpenAI, Anthropic, Gemini, Vertex AI, Bedrock, Azure OpenAI, Groq, Fireworks, DeepSeek, xAI, and local runtimes such as Ollama and vLLM, behind endpoints that are compatible with both the OpenAI and Anthropic API shapes. Around that it layers the operational controls teams actually need: virtual keys, per-key budgets and rate limits, caching, failover and load balancing, audit logs, and usage and cost tracking, with OpenTelemetry and Prometheus for observability. A commercial Pro tier adds prompt compression, OIDC SSO with group gating, per-child quota templates, and intelligent routing in beta, priced at 499 dollars per month or 4,999 dollars per year with unlimited requests, seats, nodes, and environments.
GoModel is a self-hosted, MIT-licensed AI gateway consolidating 31 LLM providers behind OpenAI and Anthropic compatible endpoints.
GoModel is built by enterpilot, Inc. and distributed as open source under the MIT licence, with a commercial Pro tier.
It positions itself as an open-source, self-hosted alternative to hosted routing services, emphasising that it never phones home and supports air-gapped deployment.
The gateway fronts 31 providers behind OpenAI and Anthropic compatible APIs and adds model aliases, virtual models, caching, failover, load balancing, budgets, rate limits, virtual keys, audit logs, and usage and cost tracking.
Deployment options include a single binary, Docker, Docker Compose, and Helm, with OpenTelemetry and Prometheus for observability and OIDC providers for Pro SSO.
Platform and infrastructure teams managing multi-provider LLM access, especially in regulated or air-gapped environments that cannot use a hosted gateway.
Application developers calling the gateway endpoint.
Platform engineering and infrastructure leads.
Open-source infrastructure and LLMOps communities.
A platform team that needs multi-provider routing, budgets, and audit logs without a third party in the request path.
Not stated; the vendor entity is enterpilot, Inc.
The core gateway is MIT licensed. The Pro tier is sold under a separate commercial licence agreement.
The vendor lists 31, including OpenAI, Anthropic, Gemini, Vertex AI, Bedrock, Azure OpenAI, Groq, DeepSeek, xAI, Ollama, and vLLM.
Yes. The vendor states it never phones home and works in air-gapped deployments.
As a single Go binary, via Docker or Docker Compose, or with a Helm chart.
Prompt compression, OIDC SSO with group gating, per-child quota templates, intelligent routing in beta, and founder-level support.
Side-by-side pages for pricing, features, and best-fit use cases.
Run open LLMs locally with a single command.
One API for hundreds of AI models across providers.
AI cloud with 200+ model APIs, serverless inference and GPU instances
Serverless GPU runtime for AI inference, training, and sandboxes