The observability platform engineers actually like. Now with first-class LLM tracing.
Honeycomb has been the engineer's choice for production observability since 2016 โ Slack, HelloFresh, and LaunchDarkly run on it. In 2025 they shipped native LLM tracing that uses the same OpenTelemetry standard as everything else. Paste in your OpenAI/Anthropic calls, get full traces of every prompt, every tool call, every token cost, every latency spike. And because it's Honeycomb, you get the famous query builder that lets you slice by 50 dimensions in one query โ no PromQL, no SQL.
Who it's for: Engineering teams running LLM apps in production who already think in traces, not logs.
Auto-instrumented spans for OpenAI, Anthropic, Cohere, Bedrock. Drop-in SDK, traces show up immediately.
Click any spike in latency or cost. Honeycomb auto-finds the field that correlates. No manual query writing.
Every LLM call shows prompt tokens, completion tokens, and dollar cost. Group by user, model, feature.
No PromQL, no SQL. Drag fields, filter, groupby, visualize. The tool observability should have been all along.
See which downstream services, vector DBs, and models your LLM depends on. Spans light up red when one is slow.
Build dashboards in minutes. Share links, embed in Slack. No more "check Grafana" tickets.
If you're already an OpenTelemetry shop or you need to observe LLM apps alongside the rest of your microservices, Honeycomb is the obvious pick. The query builder is unmatched, BubbleUp genuinely saves hours, and the LLM tracing is first-class without being a separate product. For pure LLM eval workflows, Braintrust or LangSmith are more focused โ but for full-stack observability, nothing else comes close.