RunPod
GPU cloud for training and serverless AI inference with zero egress fees
Cheapest serverless inference for open-source LLMs, pay per token
DeepInfra competes almost entirely on price, and it delivers, with some of the cheapest per-token rates for open-source models plus batch and Flex tiers that cut costs further for non-urgent work. The OpenAI-compatible API and dedicated GPU option give teams a clean path from prototype to scale. It focuses on open-source models, so you won't find proprietary frontier models here. As with any low-cost provider, evaluate latency and reliability for your specific production needs before committing.
DeepInfra offers low-cost serverless inference for open-source LLMs with pay-per-token billing, latency tiers, discounted batch inference, and dedicated GPU rentals.
DeepInfra provides serverless inference for open-source large language models and other AI models through a simple, OpenAI-compatible API. Developers pay only for the tokens they use, with no minimums and no charge for idle time, making it a cost-efficient way to run models like Llama, Mistral, DeepSeek, and many others in production without managing GPUs. Pricing is organized into latency tiers: Flex bills at 0.8x the base per-token rate for non-production and asynchronous work, the default Standard tier bills at 1x, and Priority bills at 1.5x for faster time-to-first-token during peak demand. DeepInfra also offers serverless batch inference at roughly half price with results delivered within 24 hours, and dedicated GPU-hour rentals from A100 up to B300 for teams that want reserved capacity. Per-token rates span a wide range depending on model size. DeepInfra raised a $107M Series B in May 2026, reflecting strong growth in the serverless inference market. It suits developers and teams that want the lowest-cost path to running open-source models at scale, competing with providers like Together AI, Fireworks, and Groq on price and model coverage.
DeepInfra is a low-cost serverless inference provider for open-source LLMs, with pay-per-token billing, latency tiers, discounted batch inference, and dedicated GPU rentals.
DeepInfra is an AI infrastructure company offering serverless inference for open-source models through a simple, cost-focused API. It has built a reputation for aggressive per-token pricing.
The company raised a $107M Series B in May 2026, reflecting strong demand for affordable open-source model inference. It continues to expand model coverage and infrastructure options.
DeepInfra serves open-source LLMs and other models via an OpenAI-compatible API with pay-per-token billing and no idle charges. It offers Flex, Standard, and Priority latency tiers, plus batch inference at roughly half price for large asynchronous jobs.
For teams needing reserved capacity, it rents dedicated GPUs from A100 up to newer generations. The combination of low prices, tiered latency, and batch discounts makes it a strong cost-optimization choice.
DeepInfra targets cost-sensitive developers, startups, and teams running open-source LLMs at scale who want the lowest-cost serverless inference without managing GPU infrastructure.
Developers and ML engineers calling LLM APIs.
Technical founders and engineering leads optimizing inference costs.
AI architects and open-source model practitioners.
Cost-conscious teams running open-source LLMs in production who prioritize low per-token pricing and flexible latency and batch options.
Raised a $107M Series B in May 2026 (verify with vendor).
DeepInfra offers some of the lowest per-token rates available, starting around $0.02 per 1M input tokens for small models like Llama 3.1 8B, scaling up with model size.
Flex bills at 0.8x for non-production and async work, Standard at 1x by default, and Priority at 1.5x for faster time-to-first-token during peak demand.
Yes. DeepInfra offers serverless batch inference at roughly 50% off the per-token price, with results delivered within 24 hours.
Yes. DeepInfra offers an OpenAI-compatible API, so you can often switch by changing the base URL and API key.
Yes. DeepInfra offers dedicated GPU-hour rentals ranging from around $0.89/hr for an A100 up to higher rates for newer GPUs like the B300.
Side-by-side pages for pricing, features, and best-fit use cases.
GPU cloud for training and serverless AI inference with zero egress fees
Run open LLMs locally with a single command.
MCP server registry, in-browser inspector, and gateway with an LLM API bundled in.
Copy.ai is a GTM (Go-To-Market) AI platform designed for sales, marketing, and revenue teams to automate workflows including sales outreach, content creation, lead processing, and ABM campaigns.