Execution tracing

Record retrieval, inference, and tool calls on one run so you can inspect and replay it without joining three logs.

Why agent failures are hard to reconstruct

When an answer is wrong, the evidence is usually split. The model vendor shows tokens. The retrieval service shows a query. The tool API shows a 200. None of them show the order a person needs: what was retrieved, what the model saw, what it called, and what it said. Teams spend the incident joining those logs by timestamp.

What an execution trace must contain

A useful trace is the run itself. It names each step, how long it took, and the payload that mattered: sources in, tool arguments out, tool result back, model output. Aggregate latency is a health check. It does not explain a single bad decision.

How to implement tracing

  • Record the input and the step type for every stage of the run.
  • Store duration and token use on inference, not only a final latency number.
  • Attribute which sources were retrieved and what each tool returned.
  • Give a person an inspector for one execution, not only an aggregate dashboard.
  • Replay the same input with another model or instruction set and keep both trails.

How Obliq records and replays a run

Obliq stores deterministic step-by-step latency, token, and source tracking on the execution. Replay re-runs that trace with alternate parameters so you can change the model or the instructions and compare. Evaluation can then score groundedness and tool reliability on the same evidence.

Questions

What is AI execution tracing?

Execution tracing is the ordered record of one agent run: input, retrieval, assembled context, model inference, tool calls, and the final output, with latency and source attribution.

How is a trace different from application logs?

Logs are whatever each service printed. A trace is one execution with those steps already joined, so you can see retrieval, the model, and the tool call in order.

Can I replay a traced execution?

Yes. Obliq can re-run a past trace with alternate parameters, such as a different model or instructions, and you compare the original with the replay.