Skip to main content
DeepInfra logo

DeepInfra

Cheapest serverless inference for open-source LLMs, pay per token

ai-infrastructure#serverless-inference#llm-api#open-source-models#gpu-rental
Free trial Claimed API Teams
Toolglade’s take

DeepInfra competes almost entirely on price, and it delivers, with some of the cheapest per-token rates for open-source models plus batch and Flex tiers that cut costs further for non-urgent work. The OpenAI-compatible API and dedicated GPU option give teams a clean path from prototype to scale. It focuses on open-source models, so you won't find proprietary frontier models here. As with any low-cost provider, evaluate latency and reliability for your specific production needs before committing.

About DeepInfra

DeepInfra offers low-cost serverless inference for open-source LLMs with pay-per-token billing, latency tiers, discounted batch inference, and dedicated GPU rentals.

DeepInfra provides serverless inference for open-source large language models and other AI models through a simple, OpenAI-compatible API. Developers pay only for the tokens they use, with no minimums and no charge for idle time, making it a cost-efficient way to run models like Llama, Mistral, DeepSeek, and many others in production without managing GPUs. Pricing is organized into latency tiers: Flex bills at 0.8x the base per-token rate for non-production and asynchronous work, the default Standard tier bills at 1x, and Priority bills at 1.5x for faster time-to-first-token during peak demand. DeepInfra also offers serverless batch inference at roughly half price with results delivered within 24 hours, and dedicated GPU-hour rentals from A100 up to B300 for teams that want reserved capacity. Per-token rates span a wide range depending on model size. DeepInfra raised a $107M Series B in May 2026, reflecting strong growth in the serverless inference market. It suits developers and teams that want the lowest-cost path to running open-source models at scale, competing with providers like Together AI, Fireworks, and Groq on price and model coverage.

TL;DR

DeepInfra is a low-cost serverless inference provider for open-source LLMs, with pay-per-token billing, latency tiers, discounted batch inference, and dedicated GPU rentals.

Company overview

DeepInfra is an AI infrastructure company offering serverless inference for open-source models through a simple, cost-focused API. It has built a reputation for aggressive per-token pricing.

The company raised a $107M Series B in May 2026, reflecting strong demand for affordable open-source model inference. It continues to expand model coverage and infrastructure options.

Product features

DeepInfra serves open-source LLMs and other models via an OpenAI-compatible API with pay-per-token billing and no idle charges. It offers Flex, Standard, and Priority latency tiers, plus batch inference at roughly half price for large asynchronous jobs.

For teams needing reserved capacity, it rents dedicated GPUs from A100 up to newer generations. The combination of low prices, tiered latency, and batch discounts makes it a strong cost-optimization choice.

Target market

DeepInfra targets cost-sensitive developers, startups, and teams running open-source LLMs at scale who want the lowest-cost serverless inference without managing GPU infrastructure.

Buyer personas

End users

Developers and ML engineers calling LLM APIs.

Buyers

Technical founders and engineering leads optimizing inference costs.

Key influencers

AI architects and open-source model practitioners.

Ideal customer profile

Cost-conscious teams running open-source LLMs in production who prioritize low per-token pricing and flexible latency and batch options.

Funding & performance

Raised a $107M Series B in May 2026 (verify with vendor).

Pros & cons

Pros

  • Among the lowest per-token prices
  • Pay only for tokens, no idle charges
  • OpenAI-compatible API for easy migration
  • Discounted batch inference
  • Latency tiers to trade cost vs speed
  • Dedicated GPU rentals available

Cons

  • Focused on open-source, not proprietary models
  • No free plan
  • Latency and reliability vary by tier
  • Fewer enterprise features than large clouds
  • No self-hosting

Key features

API
Team collaboration
Multi-language
Integrations
OpenAI-compatible API, Python SDK, LangChain, HTTP API
Input types
text, image
Output types
text
Best For
Low-cost LLM inference, Batch processing, Running open-source models

Compare key features

View all alternatives →
Feature
DeepInfra
RunPod
Ollama
Pricing
Paid
Paid
Freemium
Free plan
No
No
Yes
Free trial
Yes
No
No
API
Yes
Yes
Yes
Self-hosted
No
No
Yes
Team support
Yes
Yes
No

Frequently asked questions

How cheap is DeepInfra?+

DeepInfra offers some of the lowest per-token rates available, starting around $0.02 per 1M input tokens for small models like Llama 3.1 8B, scaling up with model size.

What are the latency tiers?+

Flex bills at 0.8x for non-production and async work, Standard at 1x by default, and Priority at 1.5x for faster time-to-first-token during peak demand.

Does DeepInfra support batch inference?+

Yes. DeepInfra offers serverless batch inference at roughly 50% off the per-token price, with results delivered within 24 hours.

Is the API compatible with OpenAI?+

Yes. DeepInfra offers an OpenAI-compatible API, so you can often switch by changing the base URL and API key.

Can I rent dedicated GPUs?+

Yes. DeepInfra offers dedicated GPU-hour rentals ranging from around $0.89/hr for an A100 up to higher rates for newer GPUs like the B300.

Reviews (0)

Write a review

Pick a rating
Loading reviews…
Compare

Compare DeepInfra with other AI tools

Side-by-side pages for pricing, features, and best-fit use cases.

All comparisons →

Similar tools you may like