> ## Documentation Index
> Fetch the complete documentation index at: https://docs.agenticenv.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Token Usage

> Aggregate prompt, completion, and reasoning token counts across LLM rounds in a run

Each LLM completion can report token counts via [`interfaces.LLMUsage`](https://pkg.go.dev/github.com/agenticenv/agent-sdk-go/pkg/interfaces#LLMUsage) on [`LLMResponse.Usage`](https://pkg.go.dev/github.com/agenticenv/agent-sdk-go/pkg/interfaces#LLMResponse). OpenAI, Anthropic, and Gemini clients populate the fields when the provider returns them.

An agent run can involve **multiple LLM calls** — the initial generation, then one more per tool round as the model processes tool results and decides next steps. `AgentRunResult.LLMUsage` is the **sum across all rounds**, giving you a single number for the full cost of a run.

## Usage fields

| Field | Description |
| - | - |
| `PromptTokens` | Input tokens |
| `CompletionTokens` | Output tokens |
| `TotalTokens` | Sum when reported by provider |
| `CachedPromptTokens` | Prompt cache hits (when supported) |
| `ReasoningTokens` | Extended thinking tokens (when supported) |

`CachedPromptTokens` is populated by OpenAI automatically (server-side caching, no client option needed) and by Anthropic only when [`llm.WithPromptCaching(true)`](/getting-started/llm-providers#prompt-caching) is set — it's `0` for Anthropic otherwise, since caching is disabled by default there.

## Run

[`AgentRunResult.LLMUsage`](https://pkg.go.dev/github.com/agenticenv/agent-sdk-go/pkg/agent#AgentRunResult) is the **sum across all LLM calls** in that run — including tool rounds. Use it for cost estimates, quotas, and logging:

```go theme={null}
agentRun, err := a.Run(ctx, prompt, nil)
if err != nil {
    return err
}
result, err := agentRun.Get(ctx)
if err != nil {
    return err
}

if result.LLMUsage != nil {
    fmt.Printf("prompt=%d completion=%d total=%d\n",
        result.LLMUsage.PromptTokens,
        result.LLMUsage.CompletionTokens,
        result.LLMUsage.TotalTokens,
    )
}
```

## Stream

The same aggregate appears on `LLMUsage` in the `AgentEventTypeRunFinished` event — `Result` is an `*AgentRunResult`:

```go theme={null}
for ev := range eventCh {
    if ev.Type() != agent.AgentEventTypeRunFinished {
        continue
    }
    finished, ok := ev.(*agent.AgentRunFinishedEvent)
    if !ok || finished.Result == nil || finished.Result.LLMUsage == nil {
        continue
    }
    u := finished.Result.LLMUsage
    fmt.Printf("total tokens: %d\n", u.TotalTokens)
}
```

OpenAI streaming with `include_usage` surfaces totals on `RUN_FINISHED`. Example: [Stream](/examples/stream).

## Budget enforcement

To automatically stop or pause a run when a token or cost limit is reached, use [`WithBudget`](/features/budget). It works across all runtimes and covers sub-agent usage transparently without any hook code.

## Per-request usage

For per-LLM-call breakdowns — to log each round individually, build cost dashboards, or track iteration-level usage — use [`AfterLLMHook`](/features/hooks) and inspect `in.Response.Usage` on each iteration:

```go theme={null}
agent.WithHooks("cost", agent.AgentHooks{
    AfterLLM: []agent.AfterLLMHook{
        func(ctx context.Context, in agent.AfterLLMHookInput) (agent.AfterLLMHookOutput, error) {
            if in.Response.Usage != nil {
                log.Printf("iter=%d tokens=%d", in.RunMeta.Iteration, in.Response.Usage.TotalTokens)
            }
            return agent.AfterLLMHookOutput{Response: in.Response}, nil
        },
    },
})
```

## Example

<CardGroup cols={2}>
  <Card title="Simple Agent" icon="play" href="/examples/simple-agent" horizontal>
    SHOW\_LLM\_USAGE footer
  </Card>
</CardGroup>

## Related

<CardGroup cols={2}>
  <Card title="Budget" icon="gauge" href="/features/budget" horizontal>
    Token and cost limits with automatic enforcement
  </Card>

  <Card title="Reasoning" icon="lightbulb" href="/features/reasoning" horizontal>
    ReasoningTokens from extended thinking
  </Card>

  <Card title="Streaming" icon="wave-pulse" href="/getting-started/streaming" horizontal>
    Usage on RUN\_FINISHED events
  </Card>
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.