Skip to main content
Koyeb logo

Koyeb

Serverless cloud with scale-to-zero GPUs for AI inference and apps

ai-infrastructure#serverless#gpu#inference#scale-to-zero
Free plan Free trial Claimed API Teams

About Koyeb

Koyeb is a serverless cloud with scale-to-zero autoscaling and per-second billing across CPUs and GPUs up to B200, aimed at affordable AI inference, APIs, and app deployments.

Koyeb is an AI-focused serverless cloud designed to deploy inference workloads and applications without managing infrastructure. It offers native autoscaling and scale-to-zero so GPU and CPU instances spin down when idle, and it bills by the second, making it well suited to bursty inference traffic and cost-sensitive deployments. The platform spans a broad hardware range, from CPUs to GPUs like L4, L40S, A100, H100, and B200, plus next-gen Tenstorrent accelerators in preview. Koyeb has repeatedly cut serverless GPU prices, advertising an H100 rate around $2.50/hr that undercuts several major competitors, and supports global deployments across US, EU, and Asia regions with one-click model deployments. Pricing combines a flat plan fee with per-second usage: a Pro plan at $29/month including some compute credit, a Scale plan at $299/month with more included compute, and separately priced databases, GPUs, and extra instances. Koyeb is aimed at developers and teams that want a simple, autoscaling home for AI models, APIs, and full applications.

TL;DR

Koyeb is a serverless cloud with autoscaling scale-to-zero GPUs and per-second billing, offering competitive AI inference pricing across a broad hardware range.

Company overview

Koyeb is a serverless cloud provider focused on making AI inference and application deployment simple and affordable through autoscaling and scale-to-zero. It competes with serverless GPU platforms by emphasizing per-second billing and aggressive GPU pricing.

The company regularly publishes benchmarks and price cuts across its GPU lineup and positions itself as a developer-friendly alternative to both hyperscalers and other serverless GPU startups.

Product features

Koyeb provides serverless deployment of inference endpoints, APIs, web apps, workers, and databases with native autoscaling and scale-to-zero. It supports CPUs and GPUs from RTX 4000 Ada up to B200, with Tenstorrent accelerators in preview.

Billing is per second on top of a flat plan fee, deployments are global across US, EU, and Asia, and one-click model deployments simplify getting AI workloads live.

Target market

Koyeb targets AI startups, developers, and teams that need affordable, autoscaling GPU inference and app hosting without managing infrastructure.

Buyer personas

End users

Developers deploying AI models and applications.

Buyers

Startup CTOs and engineering leads managing cloud spend.

Key influencers

ML engineers and DevOps practitioners.

Ideal customer profile

AI-first startups and small-to-mid teams that want cost-efficient, autoscaling serverless GPU inference and application hosting.

Funding & performance

Koyeb has raised venture funding; specific amounts should be verified with the vendor.

Pros & cons

Pros

  • Scale-to-zero saves idle GPU cost
  • Per-second billing
  • Competitive H100 pricing
  • Broad GPU range up to B200
  • Global multi-region deployments
  • One-click model deployments

Cons

  • Flat plan fee on top of usage
  • Not a full hyperscaler feature set
  • GPU availability can vary
  • Less mature ecosystem than AWS/GCP
  • Enterprise controls still maturing

Pricing plans

Free / Starter
$0
  • Free tier resources
  • Scale-to-zero
  • Per-second billing
Pro
$29 / month
  • $10 compute included
  • CPU and GPU access
  • Autoscaling
  • Global regions
Scale
$299 / month
  • $100 compute included
  • Higher limits
  • Team features
  • Priority support

Key features

API
Team collaboration
Multi-language
Integrations
GitHub, Docker, PostgreSQL, REST API
Input types
text
Output types
text
Best For
Serverless GPU inference, Autoscaling AI apps, Cost-sensitive deployments

Compare key features

View all alternatives →
Feature
Koyeb
RunPod
Ollama
Pricing
Freemium
Paid
Freemium
Free plan
Yes
No
Yes
Free trial
Yes
No
No
API
Yes
Yes
Yes
Self-hosted
No
No
Yes
Team support
Yes
Yes
No

Frequently asked questions

What GPUs does Koyeb offer?+

A range from RTX 4000 Ada up to B200, including L4, L40S, A100, and H100, plus Tenstorrent accelerators in preview.

How is Koyeb billed?+

A flat monthly plan fee plus per-second usage for compute, GPUs, and databases.

Does Koyeb support scale-to-zero?+

Yes, both standard CPU and GPU instances can scale to zero when idle.

How does Koyeb's H100 price compare?+

Koyeb advertises around $2.50/hr for H100, undercutting several major serverless competitors.

Can I deploy full applications?+

Yes, Koyeb runs inference, APIs, web apps, workers, and databases.

Reviews (0)

Write a review

Pick a rating
Loading reviews…
Compare

Compare Koyeb with other AI tools

Side-by-side pages for pricing, features, and best-fit use cases.

All comparisons →

Similar tools you may like