Arize Phoenix
Open-source LLM and agent observability built on OpenTelemetry.
Evaluation intelligence and observability platform for GenAI apps and agents
Galileo is a GenAI evaluation and observability platform using purpose-built Luna models to deliver 20+ quality metrics with low latency, plus production monitoring and guardrails, with free, $100/month Pro, and enterprise plans.
Galileo provides 'evaluation intelligence' for generative AI: a managed platform to measure output quality, monitor production behavior, and enforce runtime guardrails. Rather than relying solely on a large model to grade outputs, Galileo uses its own purpose-built Luna evaluation models, with the Luna-2 generation delivering 20+ metrics — including groundedness, completeness, correctness, and toxicity — at sub-200ms latency and lower cost than LLM-as-judge approaches. The platform spans the lifecycle: offline evaluation and experimentation during development, production observability with tracing, and guardrails that can block or flag risky outputs at runtime. A polished console ties these together with minimal setup, making it accessible to teams that don't want to build evaluation infrastructure themselves. Galileo was founded in 2021 by ex-Google and Apple engineers and has raised roughly $68M, with customers including HP, Twilio, Reddit, and Comcast. As of 2026 it offers a Free plan (5,000 traces/month, unlimited users and custom evals, no RBAC or advanced analytics), a $100/month Pro plan, and enterprise plans that add self-hosting and advanced controls.
Galileo is a GenAI evaluation and observability platform using purpose-built Luna models for fast, research-backed quality metrics, plus monitoring and guardrails, with free, $100/month Pro, and enterprise tiers.
Galileo was founded in 2021 by ex-Google and Apple engineers to bring rigorous evaluation to generative AI. It has raised roughly $68M and counts HP, Twilio, Reddit, and Comcast among its customers.
The company differentiates with its own Luna evaluation models rather than relying on general LLMs for grading, delivering faster and cheaper quality measurement across development and production.
Galileo spans offline evaluation, production observability with tracing, and runtime guardrails in one polished console. Luna-2 models produce 20+ metrics — groundedness, completeness, correctness, toxicity — at sub-200ms latency.
Teams instrument apps via SDKs or OpenTelemetry, run custom and built-in evals, monitor production behavior, and enforce guardrails, with enterprise self-hosting for data control.
Galileo targets teams and enterprises building and operating GenAI applications and agents that need reliable evaluation, production monitoring, and safety guardrails.
AI engineers and data scientists evaluating and monitoring GenAI apps.
Engineering and AI leaders investing in evaluation and observability.
ML practitioners and quality/compliance stakeholders.
Organizations shipping production GenAI who need rigorous, low-latency evaluation, monitoring, and guardrails, optionally self-hosted.
Founded 2021; raised approximately $68M to date. Verify the latest funding details via public sources or the vendor.
Galileo is an LLM evaluation and observability platform for testing, monitoring, and guardrailing GenAI apps and agents, using purpose-built Luna models to score output quality.
Luna models are Galileo's purpose-built evaluation models; the Luna-2 generation delivers 20+ metrics like groundedness and correctness at sub-200ms latency, cheaper than LLM-as-judge.
Yes. The Free plan includes 5,000 traces per month, unlimited users, and unlimited custom evals, but omits RBAC, advanced analytics, and dedicated support.
As of 2026, Pro is $100/month, with enterprise plans adding self-hosting and advanced controls.
Yes. Galileo offers runtime guardrails that can flag or block risky outputs in production, alongside evaluation and observability.
Side-by-side pages for pricing, features, and best-fit use cases.
Open-source LLM and agent observability built on OpenTelemetry.
Open-source LLM evaluation, tracing, and observability by Comet.
Open-source LLM evaluation framework with pytest-style testing.
Open-source evaluation toolkit for RAG and LLM applications.