Observability & Eval
Once an AI system runs in production you need to see inside it. These tools trace each step, evaluate output quality, track cost and latency, and help you catch regressions before your users do.
Pricing
License
Language
Last updated: Show all
Sort: A–Z
No tools match.
Tracing & Monitoring
Capture, inspect, and monitor agent runs step by step.
Arize Phoenix
Arize AI
Open-source LLM and agent observability — OpenTelemetry tracing and evals you can self-host
Helicone
Helicone
Open-source LLM observability via a one-line proxy — logging, caching, and cost tracking
Langfuse
Langfuse
Open-source tracing, evaluation, and analytics for LLM and agent applications
LangSmith
LangChain
LangChain's hosted platform for tracing, evaluating, and monitoring LLM and agent apps (works with or without LangChain)
Opik
Comet
Comet's open-source platform for tracing, evaluating, and monitoring LLM apps
Evaluation
Score and test LLM and RAG output quality.
DeepEval
Confident AI
Open-source "Pytest for LLMs" — unit-test your model and RAG output with metrics.
Ragas
Exploding Gradients
Open-source evaluation framework for RAG pipelines and LLM applications