Skip to main content
Not Diamond logo

Not Diamond

Intelligent LLM router that picks the best model per query

coding#llm-router#model-routing#cost-optimization#ai-infrastructure
Free plan API Teams
Toolglade’s take

Not Diamond is a focused, credible entrant in the model-routing space, and its pitch (route each query to the best-fit model to save cost without sacrificing quality) is easy to reason about. The honest caveats: routing quality depends heavily on your workload and the candidate model set, published savings figures (20-40%) come from the vendor, and the exact per-token fee is not fully transparent on the pricing page. Routing is a crowded, fast-moving category, so validate on your own evals before committing.

About Not Diamond

Not Diamond is an intelligent LLM router that predicts which model in a candidate set will best answer a given prompt, then routes the request there to improve quality while cutting cost and latency. It offers a hosted routing API plus tools to train custom routers on your own evaluation data, and cites cost reductions of 20-40% with roughly 100-150ms of added latency. Not Diamond is aimed at LLM applications and coding agents that want to move beyond a single fixed model.

Not Diamond focuses on a narrow but valuable problem: no single LLM is best at everything, and choosing the right model per request can materially change quality, cost, and speed. Not Diamond trains routers that, given a prompt, predict which model in a candidate set will produce the best response, then send the request there. This can cut inference costs while holding or improving quality compared with always using one frontier model. The product supports both a hosted routing service and tooling to train custom routers on your own evaluation data, so teams can optimize routing for their specific tasks and quality bars. It is designed to add minimal latency to each request (the company cites roughly 100-150ms of added router latency) and to slot into existing LLM applications, including coding agents. Not Diamond markets itself around measurable ROI: adopters are cited as cutting inference costs by 20-40% while maintaining quality, with the router priced as a small fee per million tokens routed. It offers an early-access tier to try the service and an Enterprise plan with custom pricing. Because model routing is an emerging category with several competitors, buyers should benchmark on their own workloads.

TL;DR

Not Diamond is an intelligent LLM router that predicts and sends each prompt to the best-fit model in a candidate set. It aims to cut inference costs 20-40% while maintaining quality, adding only ~100-150ms of latency. It offers a hosted API and custom router training on your own eval data, with a free early-access tier and custom Enterprise pricing. It targets LLM applications and coding agents that use more than one model. Exact per-token pricing is not fully published.

Company overview

Not Diamond is a company focused on AI model routing, offering both a hosted routing service and tooling for training custom routers. It publicly launched its router and has positioned itself around measurable cost savings for LLM applications and coding agents.

It operates in the emerging model-routing category alongside competitors, and cites usage by companies such as OpenRouter and Rootly to reduce inference costs.

Product features

The core product is a router that, given a prompt, predicts which model will produce the best response and routes accordingly. It supports a defined candidate set of models and can be trained into custom routers using a customer's own evaluation data.

It is designed for low added latency (cited ~100-150ms) and integrates via a Python SDK and API into existing applications, including coding agents. Provider model calls are billed separately by each provider.

Target market

LLM application developers, AI agent builders, and cost-sensitive engineering teams that run multiple models and want to optimize quality, cost, and latency per request.

Buyer personas

End users

Developers integrating LLM calls who want automatic model selection per query.

Buyers

Engineering and product leaders focused on lowering LLM inference spend without hurting quality.

Key influencers

ML/AI infrastructure engineers and evaluation-focused practitioners with their own benchmark data.

Ideal customer profile

Teams running multiple LLMs in production, especially coding agents and high-volume applications, that have or can build evaluation data to validate routing gains.

Funding & performance

Specific funding rounds and totals for Not Diamond are not clearly disclosed in public sources reviewed; verify current funding details directly with the company.

Pros & cons

Pros

  • Directly targets cost savings without quality loss
  • Supports custom routers trained on your own data
  • Low added router latency (cited ~100-150ms)
  • Free early-access tier to evaluate
  • Works with major providers and coding agents
  • Simple per-token fee model on top of provider costs

Cons

  • Routing benefits depend heavily on your specific workload
  • Exact per-token pricing is not fully published
  • Vendor-cited savings figures need independent validation
  • Adds an external dependency and network hop
  • Crowded, fast-moving competitive category
  • No self-hosted option for the routing service

Pricing plans

Early Access
Free
  • Access to hosted routing API
  • Route across supported models
  • Evaluate on your workloads
  • Community/docs support
Usage (Router Fee)
~$0.05 per 1M tokens routed
  • Per-token routing fee on top of provider costs
  • Custom router training on your eval data
  • Low added latency routing
  • Provider model usage billed separately
Enterprise
Custom
  • Custom routing configurations
  • Volume pricing
  • Priority support
  • Advanced integration assistance

Key features

API
Team collaboration
Integrations
OpenAI, Anthropic, Google, OpenRouter, Python SDK, Dagster
Input types
text
Output types
text
Best For
Routing queries to the best-fit LLM, Reducing inference cost, Improving response quality across models, Powering coding agents with multi-model routing

Compare key features

View all alternatives →
Feature
Not Diamond
SuperCompress
OpenRouter
Pricing
Freemium
Freemium
Freemium
Free plan
Yes
Yes
Yes
Free trial
No
No
No
API
Yes
Yes
Yes
Self-hosted
No
Yes
No
Team support
Yes
No
Yes

Frequently asked questions

What does Not Diamond do?+

It routes each LLM query to the model most likely to answer it best, improving quality while reducing cost and latency compared with always using one model.

How much does Not Diamond cost?+

There is a free early-access tier and a custom-priced Enterprise plan. The router charges a small per-token fee (publicly referenced around $0.05 per million tokens routed); provider model usage is billed separately. Confirm current rates with the vendor.

How much latency does routing add?+

Not Diamond cites roughly 100-150ms of added latency per request for the routing decision.

Can I train a custom router?+

Yes. Not Diamond provides tooling to train routers on your own evaluation data so routing is optimized for your specific tasks and quality bar.

Which models can it route between?+

It supports routing across major providers such as OpenAI, Anthropic, and Google, and is used alongside platforms like OpenRouter and in coding agents.

Reviews (0)

Write a review

Pick a rating
Loading reviews…
Compare

Compare Not Diamond with other AI tools

Side-by-side pages for pricing, features, and best-fit use cases.

All comparisons →

Similar tools you may like