Ollama
Run open LLMs locally with a single command.
Open-source unified inference platform for serving AI models and apps
BentoML is an open-source inference platform that packages and serves AI models as scalable APIs via Python decorators, with managed BentoCloud offering per-second, scale-to-zero GPU billing.
BentoML gives software engineers a Pythonic way to turn models into production inference services: you define a service class, and BentoML handles packaging (a 'Bento'), API generation, batching, and serving. It supports a wide range of workloads — LLM applications, model inference APIs, asynchronous job queues, and multi-model pipelines — and is model- and framework-agnostic. The managed BentoCloud runs these services in the cloud with autoscaling and per-second, scale-to-zero billing, so idle deployments cost nothing. It offers tiers from a pay-as-you-go Starter to SLA-backed Scale and a bring-your-own-cloud (BYOC) Enterprise option, with GPU instances metered per second and priority access to high-end accelerators on higher plans. The open-source framework remains free and widely used, while BentoCloud provides the managed serving layer. In February 2026, BentoML was acquired by Modular AI; the open-source project continues, though post-acquisition BentoCloud pricing may be re-disclosed over time, so teams should confirm current terms.
BentoML is an open-source inference platform for packaging and serving AI models as scalable APIs, with managed BentoCloud offering per-second, scale-to-zero billing; acquired by Modular AI in Feb 2026.
BentoML is the company behind the widely used open-source inference framework of the same name, focused on making model deployment simple for software engineers. It provides BentoCloud as its managed serving layer.
In February 2026, BentoML was acquired by Modular AI, aligning it with a broader AI infrastructure stack while the open-source project continues.
BentoML packages models and AI apps into deployable 'Bentos' via Python decorators, generating inference APIs with batching and autoscaling. It supports LLM apps, multi-model pipelines, and job queues across frameworks.
BentoCloud runs these services with per-second, scale-to-zero billing and tiered plans, including BYOC for enterprises, with priority access to high-end GPUs on higher plans.
BentoML targets ML and platform engineers who need to deploy models and LLM applications as scalable APIs, from startups to enterprises wanting managed or self-hosted serving.
ML and backend engineers deploying inference services.
Platform and ML infrastructure leaders selecting a serving stack.
MLOps practitioners and open-source contributors.
Engineering teams serving models or LLMs in production that want Pythonic packaging and scale-to-zero economics.
BentoML was acquired by Modular AI in February 2026; verify current corporate and pricing details with the vendor.
Yes. The BentoML open-source framework is free. BentoCloud, the managed serving platform, uses usage-based pricing with per-second billing and offers trial credits.
A Bento is the packaged artifact BentoML builds from your service code, models, and dependencies, ready to deploy as a scalable inference API.
Yes. BentoCloud compute is metered per second and scale-to-zero deployments incur no cost while idle.
Yes. BentoML was acquired by Modular AI in February 2026. The open-source framework remains free, and post-acquisition BentoCloud pricing may be re-disclosed over time.
Yes. You can self-host the open-source framework on your own Kubernetes or infrastructure, or use the managed BentoCloud.
Side-by-side pages for pricing, features, and best-fit use cases.
Run open LLMs locally with a single command.
GPU cloud for training and serverless AI inference with zero egress fees
The open hub for machine learning models, datasets, and demos.
Open-source MLOps platform for experiments, pipelines, and model management