Skip to main content

Galileo vs Arize Phoenix

GalileoArize Phoenix

Bottom line: Galileo for teams shipping GenAI apps to production; Arize Phoenix for engineers debugging LLM and agent apps.

Evaluation intelligence and observability platform for GenAI apps and agents

Visit

Open-source LLM and agent observability built on OpenTelemetry.

Visit
Votes00
PricingFreemiumFreemium
CategoryLlm ObservabilityLlm Observability
Tags
llm-evaluationobservabilityguardrailsgenai-metricsagent-monitoring
observabilityllm-evaluationopen-sourcetracingmonitoring
Best for
  • Teams shipping GenAI apps to production
  • Organizations needing evaluation rigor
  • Agent developers wanting guardrails
  • Engineers debugging LLM and agent apps
  • Teams wanting free, self-hosted observability
  • Eval-driven development workflows
Pros
  • Purpose-built Luna evaluation models (fast, cheaper than LLM-judge)
  • 20+ research-backed metrics
  • Covers evaluation, observability, and guardrails
  • Polished console with minimal setup
  • Free plan with 5,000 traces/month
  • Free and source-available, self-hosts in one command
  • Built on open OpenTelemetry standards
  • Strong tracing and evaluation for agents and RAG
  • Works locally, good for private/dev-time debugging
  • Backed by Arize's observability expertise
Cons
  • Pro at $100/month may be steep for small teams
  • Free plan lacks RBAC and advanced analytics
  • Managed platform hosts your evaluation data (unless enterprise self-host)
  • Text/GenAI-focused
  • Advanced features gated to higher tiers
  • Production-scale features require paid Arize AX
  • Self-hosting means you run the infrastructure
  • Dynatrace acquisition may change roadmap/governance
  • Observability setup still requires instrumentation effort
  • Evaluation quality depends on judge models and config

Comparison generated from each tool's listing. Add or remove tools above to change it.