Anyscale
Managed Ray for scaling AI and Python workloads

Serverless cloud for AI, ML, and data workloads in Python.
Modal is a favorite among engineers who want to express cloud infrastructure as Python and run AI, ML, and data workloads without managing servers. Its per-second billing, fast cold starts, autoscaling, and GPU access make it flexible for inference, training, batch jobs, and pipelines. It is best for developers comfortable writing code who need general-purpose serverless compute rather than a fixed model API. Caveats: there is a learning curve to its programming model, costs are usage-based and can grow with GPU time, and it is less turnkey than a hosted inference endpoint. Monthly free credits make it easy to try.
Modal is a serverless compute platform for running Python code on cloud CPUs and GPUs with per-second billing, autoscaling, and no infrastructure management. It is popular for AI inference, model training, batch processing, and data pipelines, appealing to developers who want to define infrastructure as Python code rather than manage clusters and orchestration.
Modal is a serverless cloud built for compute-heavy Python workloads. You define functions in Python, decorate them to run on Modal, and the platform provisions containers on CPUs or GPUs on demand, scaling from zero to many instances and billing by the second of actual usage. It removes the need to manage clusters, Dockerfiles-by-hand, or orchestration, letting developers move quickly from local code to scalable cloud execution. Modal is widely used for AI and ML: running inference endpoints, fine-tuning and training models, batch processing, embeddings generation, and data pipelines. It provides fast container start times, GPU access (including high-end accelerators), scheduled jobs, web endpoints, storage, and secrets management, all defined in code. Its programmable, developer-first model appeals to teams that want infrastructure expressed as Python rather than YAML. The platform suits engineers who are comfortable writing code and want flexible, general-purpose serverless compute rather than a fixed model API. It offers monthly free compute credits to start, then usage-based pricing tied to CPU, GPU, and memory time. For teams whose needs go beyond calling a hosted model, Modal is a powerful, flexible foundation.
Modal is a serverless compute platform for running Python code on cloud CPUs and GPUs with per-second billing, autoscaling, and no infrastructure management. It is popular for AI inference, model training, batch jobs, and data pipelines, and appeals to developers who want infrastructure expressed as Python. It is best for engineers needing flexible, general-purpose serverless compute rather than a fixed model API. Trade-offs include a learning curve, Python-only support, and usage-based GPU costs. Monthly free credits make it easy to try.
Modal (Modal Labs) was founded in 2021 by Erik Bernhardsson, former CTO of Better and creator of several well-known open-source tools at Spotify, along with co-founders. The company builds developer-first serverless infrastructure for compute-heavy workloads.
Modal has grown quickly with the rise of AI workloads that need flexible GPU compute, raising significant venture funding through 2025 and 2026.
Modal lets developers define functions in Python and run them on autoscaling cloud CPUs and GPUs, provisioning containers on demand with fast cold starts and per-second billing. It supports web endpoints, scheduled jobs, storage, secrets, and GPU access including high-end accelerators.
The platform is general-purpose but heavily used for AI and ML, including inference, training, fine-tuning, embeddings, batch processing, and data pipelines, all expressed as code.
Software and ML engineers and teams who want flexible, serverless GPU and CPU compute for AI, ML, and data workloads without managing clusters, and who prefer defining infrastructure in Python.
Engineers and ML practitioners running compute-heavy Python workloads.
Engineering leaders choosing serverless compute infrastructure.
ML infrastructure engineers and Python developer-experience advocates.
A technical team that needs flexible, autoscaling GPU/CPU compute for AI, ML, and data workloads and wants to express infrastructure as Python rather than manage clusters.
Modal raised a reported $87 million Series B in September 2025 at around a $1.1 billion valuation led by Lux Capital, and a reported round of about $355 million in 2026 that pushed its valuation to approximately $4.65 billion. Treat specific figures as reported and verify with the company.
A serverless compute platform for running Python code on cloud CPUs and GPUs with per-second billing and autoscaling, used for AI inference, training, batch jobs, and data pipelines.
No. You define functions in Python and Modal provisions and scales containers on demand, handling orchestration for you.
Modal provides monthly free compute credits to start. Beyond that, pricing is usage-based on the CPU, GPU, and memory time you consume.
General compute workloads written in Python, including model inference endpoints, training and fine-tuning, embeddings, batch jobs, scheduled tasks, and web endpoints.
No. It is general-purpose serverless compute, though AI and ML workloads are among its most common uses because of its GPU support.
Side-by-side pages for pricing, features, and best-fit use cases.
Managed Ray for scaling AI and Python workloads
Deploy and scale ML models in production inference.
Inference, fine-tuning, and GPU clusters for open models.
Query-aware prompt compression that cuts LLM input tokens by roughly 60% before inference.