RunPod
GPU cloud for training and serverless AI inference with zero egress fees
Serverless cloud with scale-to-zero GPUs for AI inference and apps
Koyeb is a serverless cloud with scale-to-zero autoscaling and per-second billing across CPUs and GPUs up to B200, aimed at affordable AI inference, APIs, and app deployments.
Koyeb is an AI-focused serverless cloud designed to deploy inference workloads and applications without managing infrastructure. It offers native autoscaling and scale-to-zero so GPU and CPU instances spin down when idle, and it bills by the second, making it well suited to bursty inference traffic and cost-sensitive deployments. The platform spans a broad hardware range, from CPUs to GPUs like L4, L40S, A100, H100, and B200, plus next-gen Tenstorrent accelerators in preview. Koyeb has repeatedly cut serverless GPU prices, advertising an H100 rate around $2.50/hr that undercuts several major competitors, and supports global deployments across US, EU, and Asia regions with one-click model deployments. Pricing combines a flat plan fee with per-second usage: a Pro plan at $29/month including some compute credit, a Scale plan at $299/month with more included compute, and separately priced databases, GPUs, and extra instances. Koyeb is aimed at developers and teams that want a simple, autoscaling home for AI models, APIs, and full applications.
Koyeb is a serverless cloud with autoscaling scale-to-zero GPUs and per-second billing, offering competitive AI inference pricing across a broad hardware range.
Koyeb is a serverless cloud provider focused on making AI inference and application deployment simple and affordable through autoscaling and scale-to-zero. It competes with serverless GPU platforms by emphasizing per-second billing and aggressive GPU pricing.
The company regularly publishes benchmarks and price cuts across its GPU lineup and positions itself as a developer-friendly alternative to both hyperscalers and other serverless GPU startups.
Koyeb provides serverless deployment of inference endpoints, APIs, web apps, workers, and databases with native autoscaling and scale-to-zero. It supports CPUs and GPUs from RTX 4000 Ada up to B200, with Tenstorrent accelerators in preview.
Billing is per second on top of a flat plan fee, deployments are global across US, EU, and Asia, and one-click model deployments simplify getting AI workloads live.
Koyeb targets AI startups, developers, and teams that need affordable, autoscaling GPU inference and app hosting without managing infrastructure.
Developers deploying AI models and applications.
Startup CTOs and engineering leads managing cloud spend.
ML engineers and DevOps practitioners.
AI-first startups and small-to-mid teams that want cost-efficient, autoscaling serverless GPU inference and application hosting.
Koyeb has raised venture funding; specific amounts should be verified with the vendor.
A range from RTX 4000 Ada up to B200, including L4, L40S, A100, and H100, plus Tenstorrent accelerators in preview.
A flat monthly plan fee plus per-second usage for compute, GPUs, and databases.
Yes, both standard CPU and GPU instances can scale to zero when idle.
Koyeb advertises around $2.50/hr for H100, undercutting several major serverless competitors.
Yes, Koyeb runs inference, APIs, web apps, workers, and databases.
Side-by-side pages for pricing, features, and best-fit use cases.
GPU cloud for training and serverless AI inference with zero egress fees
Run open LLMs locally with a single command.
The open hub for machine learning models, datasets, and demos.
Build, deploy, and host remote MCP servers on Cloudflare's edge with OAuth built in.