Together AI
Inference, fine-tuning, and GPU clusters for open models.

Very fast LLM inference on custom LPU hardware.
Groq is the go-to when raw inference speed matters: agents, voice, and interactive chat feel noticeably snappier on its LPU hardware, and per-token prices are competitive. The trade-off is scope. You serve models from Groq's supported catalog rather than arbitrary custom weights, and the exact model lineup shifts over time. Corporate turbulence around the 2026 Nvidia deal and a down round is worth noting for those making long-term platform bets, though day-to-day service has stayed reliable. A strong default for latency-sensitive apps on popular open models.
Groq is an inference cloud running on custom LPU chips that deliver very high tokens-per-second on supported open models. It exposes an OpenAI-compatible API with a free tier, per-token paid pricing, and batch discounts. It is best for latency-sensitive workloads that fit its curated model catalog rather than custom-model hosting.
Groq designs custom Language Processing Unit (LPU) hardware optimized for low-latency LLM inference and sells access through GroqCloud, an OpenAI-compatible API. Its headline advantage is speed: on supported open models such as Llama, it typically delivers much higher tokens-per-second than general-purpose GPU inference, which makes it attractive for chat, agents, and real-time applications. The service focuses on serving a curated set of popular open-weight models rather than letting you upload arbitrary custom models. Developers get a free tier with rate limits, a paid developer tier with higher limits and discounts, and batch pricing for large jobs. Pricing is quoted per million tokens and varies by model. Groq's business has been eventful. It raised large rounds through 2025, then a 2026 licensing arrangement with Nvidia reshaped the company, followed by a down round. For users, the practical picture is stable: a fast, low-cost inference API for a defined model catalog, best suited to latency-sensitive workloads that fit the supported models.
Groq is an inference cloud built on custom LPU chips that deliver very high tokens-per-second on popular open models via an OpenAI-compatible API. It offers a free tier and competitive per-token pricing, making it a strong choice for latency-sensitive chat, agents, and real-time apps. The main limits are its curated model catalog and lack of custom-model hosting. Its 2026 corporate changes, including an Nvidia licensing deal and a down round, are worth noting. Day-to-day the service has remained reliable.
Groq was founded in 2016 by Jonathan Ross, who previously helped create Google's Tensor Processing Unit. The company designs LPU hardware and operates GroqCloud, an inference service for open models.
In 2025 and 2026 Groq raised large financing rounds and entered a significant licensing arrangement with Nvidia in 2026, which was followed by a down-round financing. Despite the corporate changes, the inference service has continued operating.
GroqCloud provides an OpenAI-compatible API for a curated catalog of open-weight models, with a free tier, a paid Developer tier, and batch pricing. Its defining feature is inference speed driven by LPU hardware.
The platform targets latency-sensitive use cases such as chat, agents, and voice, where fast token generation improves user experience. It focuses on text generation rather than broad multimodal workloads.
Developers and teams building latency-sensitive AI applications, including chat assistants, agents, and real-time or voice products, who are comfortable using open-weight models from Groq's supported catalog.
Developers integrating fast LLM inference into chat, agent, and voice products.
Engineering leaders choosing an inference provider for latency-sensitive workloads.
AI infrastructure engineers and open-model advocates.
A product team whose app depends on fast, low-cost inference on popular open models and that does not need custom-model hosting.
Groq raised a reported $640 million Series D in 2024 at around a $2.8 billion valuation, and additional financing in 2025 was reported at a valuation near $6.9 billion. In 2026, following a licensing arrangement with Nvidia, a further round of roughly $350 million was reported at about a $3.5 billion valuation, described as a down round. Treat specific figures as reported, and verify with the company.
It runs inference on custom Language Processing Unit (LPU) hardware designed for LLM serving, which delivers much higher tokens-per-second than typical GPU inference on supported models.
Generally no. Groq serves a curated set of popular open-weight models rather than arbitrary custom weights.
Yes. GroqCloud has a free tier with rate limits and no credit card required, plus a paid Developer tier with higher limits.
Yes. It exposes an OpenAI-compatible endpoint, so most existing OpenAI client code works with minimal changes.
Per million tokens, with rates varying by model, plus batch discounts for large jobs.
Side-by-side pages for pricing, features, and best-fit use cases.
Inference, fine-tuning, and GPU clusters for open models.
Fast, production inference for open and custom models.
The open hub for machine learning models, datasets, and demos.
Run and deploy open-source AI models with one API call.