Skip to content

ObservabilityGalileo

Galileo

Evaluation and observability for GenAI apps and agents, with inline guardrails.

Pricing
FREEMIUM
Hosting
Cloud
Platforms
WebAPI
Models
Model-agnostic
Verified
Jun 8, 2026

A platform for testing, monitoring, and guardrailing LLM and agent applications. It ships 20+ out-of-the-box evals for RAG, agents, and safety, lets teams author custom evaluators, and turns those offline evals into real-time production guardrails powered by its own Luna eval models.

Capabilities 4

What it actually does — grouped by capability family.

  • LLM evaluation (primary capability)
  • LLM observability (primary capability)
  • Guardrails (secondary capability)
  • AI security scanning (secondary capability)

Pros & cons

  • 20+ out-of-the-box evals for RAG and agents
  • Inline runtime guardrails, not just offline scoring
  • Own Luna models keep eval costs low
  • Model-agnostic across providers
  • Pricing tiers gate the production guardrails
  • Proprietary eval models, not open source
  • Heavier setup than a drop-in proxy

Tags

View all Observability
  • View Arize Phoenix details
    ObservabilityFREEMIUM

    Arize Phoenix

    Arize AI

    LLM tracing and evaluation with retrieval debugging.

    Phoenix is Arize's observability platform — run locally in a notebook or as a hosted service. Especially strong for inspecting RAG pipelines, finding bad chunks, and tracking retrieval quality over time.

    Source-available, runs locally
    Less polished than hosted SaaS evals
    • tracing
    • rag
    • retrieval-debugging
  • View Langfuse details
    ObservabilityFREEMIUMOpen core

    Langfuse

    Langfuse

    Open-source LLM observability. Self-hostable, OpenTelemetry-native.

    Tracing, evals, prompt management, and dataset tooling for LLM apps — self-host on your own infra or use Langfuse Cloud. The open-source default when you want full ownership of your observability stack.

    Own your observability data
    Self-host infra cost at scale
    • open-source
    • tracing
    • evals
    • self-hosted
  • View Braintrust details
    EvalFREEMIUM

    Braintrust

    Braintrust

    Hosted eval + tracing platform for LLM apps.

    Production-grade eval orchestration with a dashboard, dataset versioning, and OpenTelemetry tracing. Useful once eval volume outgrows a CI YAML file.

    Eval workflow as the primary interface
    Closed-source SaaS
    • eval
    • tracing
    • datasets
    • production
  • View HoneyHive details
    EvalFREEMIUM

    HoneyHive

    HoneyHive

    The observability and evaluation layer for production AI agents.

    A platform that unifies monitoring and testing for LLM apps and agents into one improvement loop: distributed tracing, online evaluations and alerts, offline experiments, annotation queues for expert feedback, and CI/CD-integrated regression testing. Built OpenTelemetry-native with support for 100+ models and agent frameworks. The free Developer tier covers small teams; Enterprise adds scale, self-host, and compliance.

    Unifies tracing and evaluation
    SaaS-only (self-host = Enterprise)
    • eval
    • observability
    • tracing
    • agents
    • +1