Everything you need to ship AI agents you can debug.
Six capabilities, built on OpenTelemetry, designed to be adopted without forcing you to change framework, vendor, or backend. Drop the SDK in, set your key, ship the agent. Trefur shows you what it does.
No reformatting your stack. No lock-in. No forced migration. Just the trace, the cost, the alert, and the fix.
What changes when you turn it on.
These are the outcomes engineering teams shipping production agents see in the first week. Each one is backed by a capability you can click into below.
Find the failing decision in seconds, not days
Search by trace_id, customer interaction, error class, or build. Open the trace, see the exact step that went wrong, see what the model saw before it answered.
Stop being surprised by your model bill
Every token, every dollar, attributed to the step that emitted it. Roll up by agent, by team, by deployment. See drift the day it starts, not the day finance asks.
Hear about regressions before customers do
Alert envelopes on retry rate, error rate, cost-per-run, p95 latency. Routed to Slack, PagerDuty, Teams, or any webhook — with the offending trace already attached.
Swap model vendors without breaking your dashboards
One canonical trace shape across OpenAI, Anthropic, Bedrock, Azure OpenAI, Gemini. Move workloads to the cheapest or fastest provider without losing observability continuity.
Keep your data inside your network
Run the collector inside your VPC. Batch, buffer, and redact on your side before anything leaves. Fan out to Trefur and to your existing observability backend in parallel.
Hand engineering a repro, not a screenshot
Reconstruct any failing run from production — exact prompt, exact tool response, exact retry. Engineering opens it, sees the problem, ships the fix.
The capabilities behind it.
Pick the one closest to your problem.
The agent trace, end to end
One root span per agent run. Every LLM call, tool call, retry, and outcome as a child span. Schema-validated, OpenTelemetry-canonical, replayable.
See exactly what every run costs
Tokens and dollars attributed to the step that emitted them. Roll up by agent, build, deployment, model vendor. Catch cost drift before the invoice.
Alerts that page humans, not noise
Retry-rate spikes, cost-per-run drift, error-rate envelopes, latency walks. Routed to Slack, PagerDuty, Teams, webhook — with the trace attached.
One trace across every model vendor
OpenAI, Anthropic, Bedrock, Azure OpenAI, Gemini, Groq, MiniMax. Move workloads between vendors without losing trace continuity or breaking dashboards.
Self-hostable collector
Single Go binary inside your network. Local batching, on-disk buffering, redaction. Fans out to Trefur and your existing backends in parallel.
Replay a failing run from production
Reconstruct the exact prompt the model saw, the exact tool response, the exact retry. Hand engineering a repro, not a screenshot.
Together, this is what changes.
Each feature is useful on its own. Run them together and the shape of the work changes. You stop reacting to invoices, screenshots, and customer escalations. You ship faster because the agent is no longer a black box.
- Ship agents to production with confidence — every decision is captured the moment it happens
- Make model-vendor choices on price and quality, not on which one your tracing supports
- Hand finance, ops, and engineering the same source of truth — by run, by agent, by team
- Adopt without forcing a re-platform — OpenTelemetry-native, drops into existing pipelines
Start free. First trace in under five minutes.
Drop the SDK in, set your key, run your agent. The trace appears.