Groq
Very fast LLM inference on custom LPU hardware.

Inference, fine-tuning, and GPU clusters for open models.
Together AI is a strong middle-ground AI cloud: cheaper than closed frontier APIs, broader than a pure inference endpoint, and able to scale from a serverless token call to reserved GPU clusters. It suits teams committed to open models that want both flexibility and the option to fine-tune and own weights. The caveats are a wide pricing surface (per-token vs per-GPU-hour vs fine-tuning) and the usual open-model reality that you are responsible for evaluating quality and safety. Good default for cost-conscious teams building on open models.
Together AI is an AI cloud for open and open-weight models offering serverless inference, fine-tuning, and dedicated GPU clusters via an OpenAI-compatible API. It emphasizes cost efficiency and model ownership, spanning cheap per-token usage for developers and reserved GPU capacity for larger teams.
Together AI is a full-stack AI cloud for open and open-weight models. It provides serverless inference across a large catalog (Llama, DeepSeek, Qwen, Mixtral, image models, and more), fine-tuning services, and dedicated GPU clusters for training and high-volume inference. The API is OpenAI-compatible, and pricing is per token for serverless use or per GPU-hour for dedicated capacity. The company positions itself around cost and openness: running open models on Together is typically cheaper than proprietary frontier APIs, and teams retain flexibility to fine-tune and own their weights. It also invests in research and optimized inference kernels, which feeds back into serving performance. Together spans two buyer motions: developers who just want a cheap, fast token API, and larger teams that rent reserved GPU clusters for training or dedicated inference. That range is a strength for scaling from prototype to production, though it means the pricing surface is broader than a single per-token number.
Together AI is a full-stack AI cloud for open and open-weight models, offering serverless inference, fine-tuning, and dedicated GPU clusters through an OpenAI-compatible API. It emphasizes cost efficiency and model ownership, scaling from cheap per-token usage to reserved GPU capacity. It suits cost-conscious teams committed to open models. The trade-offs are a broad pricing surface and the responsibility to evaluate open-model quality yourself. It is a strong middle ground between raw GPU rental and closed frontier APIs.
Together AI was founded in 2022 by Vipul Ved Prakash, Ce Zhang, Chris Re, and Percy Liang, with a mission to make open and open-weight AI models accessible and affordable. The company combines a research effort on efficient inference and training with a commercial cloud.
It has grown quickly alongside demand for open-model infrastructure and raised substantial venture funding through 2025 and 2026.
Together offers serverless per-token inference across a large model catalog, fine-tuning services with weight ownership, and dedicated GPU clusters for training and high-volume inference. The API is OpenAI-compatible.
The company also invests in optimized inference kernels and research, which improves serving performance and cost. Product lines span individual developers and large enterprises.
Developers and ML teams building on open and open-weight models who want lower costs than proprietary APIs, the ability to fine-tune, and a path to dedicated GPU capacity as they scale.
Developers and ML engineers running inference and fine-tuning on open models.
Engineering and platform leaders choosing an open-model cloud and reserving GPU capacity.
AI researchers and open-source advocates.
A team standardizing on open models that wants low-cost inference, fine-tuning with weight ownership, and the option to rent dedicated GPUs at scale.
Together AI raised a reported $305 million Series B in early 2025 at around a $3.3 billion valuation, and an $800 million Series C was reported in mid-2026 at approximately an $8.3 billion valuation, led by Aramco Ventures with participation from investors including Nvidia and General Catalyst. Treat specific figures as reported; verify with the company.
A large catalog of open and open-weight models including Llama, DeepSeek, Qwen, Mixtral, and various image and embedding models, updated over time.
Yes. Together offers fine-tuning services, and you can deploy and retain ownership of the resulting weights.
Yes. Together provides an OpenAI-compatible API so existing client code works with minimal changes.
Serverless is per-token pay-as-you-go inference. Dedicated GPU clusters provide reserved capacity billed per GPU-hour for training or high-volume, consistent inference.
New accounts typically receive starter credits, and pricing is pay-as-you-go after that. Verify current credit amounts with the vendor.
Side-by-side pages for pricing, features, and best-fit use cases.
Very fast LLM inference on custom LPU hardware.
Fast, production inference for open and custom models.
The open hub for machine learning models, datasets, and demos.
Run and deploy open-source AI models with one API call.