Observability & Eval

Once an AI system runs in production you need to see inside it. These tools trace each step, evaluate output quality, track cost and latency, and help you catch regressions before your users do.

Pricing
License
Language
Last updated: Show all
Sort: A–Z

Tracing & Monitoring

Capture, inspect, and monitor agent runs step by step.

Open-source LLM and agent observability — OpenTelemetry tracing and evals you can self-host

arize-phoenix-v19.4.02026-07-22PythonElastic-2.0
10.7k Open source Free tier Paid
Helicone

Helicone

Helicone

Open-source LLM observability via a one-line proxy — logging, caching, and cost tracking

No updates in 6+ monthsv2025.08.21-12025-08-21TypeScriptApache-2.0
6.0k Open source Free tier Paid
Langfuse

Langfuse

Langfuse

Open-source tracing, evaluation, and analytics for LLM and agent applications

v3.224.02026-07-22TypeScriptMIT
31.7k Open source Free tier Paid
LangSmith

LangSmith

LangChain

LangChain's hosted platform for tracing, evaluating, and monitoring LLM and agent apps (works with or without LangChain)

SaaS / SDKProprietary
Free tier Paid
Opik

Opik

Comet

Comet's open-source platform for tracing, evaluating, and monitoring LLM apps

v2.1.322026-07-17PythonApache-2.0
20.8k Open source Free tier Paid

Evaluation

Score and test LLM and RAG output quality.

DeepEval

DeepEval

Confident AI

Open-source "Pytest for LLMs" — unit-test your model and RAG output with metrics.

v4.1.32026-07-12PythonApache-2.0
17.1k Open source Free tier Paid
Ragas

Ragas

Exploding Gradients

Open-source evaluation framework for RAG pipelines and LLM applications

No updates in 6+ monthsv0.4.32026-01-13PythonApache-2.0
15.0k Completely free Open source