SuperCompress
Query-aware prompt compression that cuts LLM input tokens by roughly 60% before inference.

Intelligent LLM router that picks the best model per query
Not Diamond is a focused, credible entrant in the model-routing space, and its pitch (route each query to the best-fit model to save cost without sacrificing quality) is easy to reason about. The honest caveats: routing quality depends heavily on your workload and the candidate model set, published savings figures (20-40%) come from the vendor, and the exact per-token fee is not fully transparent on the pricing page. Routing is a crowded, fast-moving category, so validate on your own evals before committing.
Not Diamond is an intelligent LLM router that predicts which model in a candidate set will best answer a given prompt, then routes the request there to improve quality while cutting cost and latency. It offers a hosted routing API plus tools to train custom routers on your own evaluation data, and cites cost reductions of 20-40% with roughly 100-150ms of added latency. Not Diamond is aimed at LLM applications and coding agents that want to move beyond a single fixed model.
Not Diamond focuses on a narrow but valuable problem: no single LLM is best at everything, and choosing the right model per request can materially change quality, cost, and speed. Not Diamond trains routers that, given a prompt, predict which model in a candidate set will produce the best response, then send the request there. This can cut inference costs while holding or improving quality compared with always using one frontier model. The product supports both a hosted routing service and tooling to train custom routers on your own evaluation data, so teams can optimize routing for their specific tasks and quality bars. It is designed to add minimal latency to each request (the company cites roughly 100-150ms of added router latency) and to slot into existing LLM applications, including coding agents. Not Diamond markets itself around measurable ROI: adopters are cited as cutting inference costs by 20-40% while maintaining quality, with the router priced as a small fee per million tokens routed. It offers an early-access tier to try the service and an Enterprise plan with custom pricing. Because model routing is an emerging category with several competitors, buyers should benchmark on their own workloads.
Not Diamond is an intelligent LLM router that predicts and sends each prompt to the best-fit model in a candidate set. It aims to cut inference costs 20-40% while maintaining quality, adding only ~100-150ms of latency. It offers a hosted API and custom router training on your own eval data, with a free early-access tier and custom Enterprise pricing. It targets LLM applications and coding agents that use more than one model. Exact per-token pricing is not fully published.
Not Diamond is a company focused on AI model routing, offering both a hosted routing service and tooling for training custom routers. It publicly launched its router and has positioned itself around measurable cost savings for LLM applications and coding agents.
It operates in the emerging model-routing category alongside competitors, and cites usage by companies such as OpenRouter and Rootly to reduce inference costs.
The core product is a router that, given a prompt, predicts which model will produce the best response and routes accordingly. It supports a defined candidate set of models and can be trained into custom routers using a customer's own evaluation data.
It is designed for low added latency (cited ~100-150ms) and integrates via a Python SDK and API into existing applications, including coding agents. Provider model calls are billed separately by each provider.
LLM application developers, AI agent builders, and cost-sensitive engineering teams that run multiple models and want to optimize quality, cost, and latency per request.
Developers integrating LLM calls who want automatic model selection per query.
Engineering and product leaders focused on lowering LLM inference spend without hurting quality.
ML/AI infrastructure engineers and evaluation-focused practitioners with their own benchmark data.
Teams running multiple LLMs in production, especially coding agents and high-volume applications, that have or can build evaluation data to validate routing gains.
Specific funding rounds and totals for Not Diamond are not clearly disclosed in public sources reviewed; verify current funding details directly with the company.
It routes each LLM query to the model most likely to answer it best, improving quality while reducing cost and latency compared with always using one model.
There is a free early-access tier and a custom-priced Enterprise plan. The router charges a small per-token fee (publicly referenced around $0.05 per million tokens routed); provider model usage is billed separately. Confirm current rates with the vendor.
Not Diamond cites roughly 100-150ms of added latency per request for the routing decision.
Yes. Not Diamond provides tooling to train routers on your own evaluation data so routing is optimized for your specific tasks and quality bar.
It supports routing across major providers such as OpenAI, Anthropic, and Google, and is used alongside platforms like OpenRouter and in coding agents.
Side-by-side pages for pricing, features, and best-fit use cases.
Query-aware prompt compression that cuts LLM input tokens by roughly 60% before inference.
One API for hundreds of AI models across providers.
Open-source platform for building production-ready LLM apps and agents.
Eval-first evaluation and observability platform for AI applications.