> ## Documentation Index
> Fetch the complete documentation index at: https://docs.agenticenv.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# LLM Providers

> Configure OpenAI, Anthropic, Gemini, DeepSeek, or Ollama — or bring your own LLM client

Every agent requires an [`interfaces.LLMClient`](https://pkg.go.dev/github.com/agenticenv/agent-sdk-go/pkg/interfaces#LLMClient). Pass it to [`NewAgent`](/getting-started/quickstart) with [`WithLLMClient`](/getting-started/configuration).

Built-in clients live under `pkg/llm/` and share option helpers from `pkg/llm`:

```go theme={null}
import "github.com/agenticenv/agent-sdk-go/pkg/llm"
```

## Supported providers

| Provider | Package | `GetProvider()` value |
| - | - | - |
| OpenAI | `pkg/llm/openai` | `openai` |
| Anthropic | `pkg/llm/anthropic` | `anthropic` |
| Gemini | `pkg/llm/gemini` | `gemini` |
| DeepSeek | `pkg/llm/deepseek` | `deepseek` |
| Ollama | `pkg/llm/ollama` | `ollama` |

They all implement the same interface:

```go theme={null}
type LLMClient interface {
    Generate(ctx context.Context, request *LLMRequest) (*LLMResponse, error)
    GenerateStream(ctx context.Context, request *LLMRequest) (LLMStream, error)
    GetModel() string
    GetProvider() LLMProvider
    IsStreamSupported() bool
}
```

## OpenAI

```go theme={null}
import (
    "github.com/agenticenv/agent-sdk-go/pkg/llm"
    "github.com/agenticenv/agent-sdk-go/pkg/llm/openai"
)

llmClient, err := openai.NewClient(
    llm.WithAPIKey(os.Getenv("OPENAI_API_KEY")),
    llm.WithModel("gpt-4o"),
    llm.WithBaseURL("https://api.openai.com/v1"), // optional — for proxies or Azure-compatible endpoints
)
if err != nil {
    return err
}
```

Common models: `gpt-4o`, `gpt-4o-mini`. The model name is sent as-is to the provider — an invalid name will return an API error at run time, not at `NewAgent`.

## Anthropic

```go theme={null}
import (
    "github.com/agenticenv/agent-sdk-go/pkg/llm"
    "github.com/agenticenv/agent-sdk-go/pkg/llm/anthropic"
)

llmClient, err := anthropic.NewClient(
    llm.WithAPIKey(os.Getenv("ANTHROPIC_API_KEY")),
    llm.WithModel("claude-sonnet-4-20250514"),
)
```

Anthropic supports extended thinking via [`WithLLMSampling`](/features/reasoning). See [Reasoning](/features/reasoning).

### Prompt caching

Anthropic supports prompt-cache breakpoints (`cache_control`) that let the provider reuse the KV-cache for the stable part of a request — system prompt, tool definitions, and everything up to the second-to-last message — instead of reprocessing it every call. Enable it with `llm.WithPromptCaching`:

```go theme={null}
llmClient, err := anthropic.NewClient(
    llm.WithAPIKey(os.Getenv("ANTHROPIC_API_KEY")),
    llm.WithModel("claude-sonnet-4-20250514"),
    llm.WithPromptCaching(true),
)
```

**Disabled by default.** Anthropic's cache writes cost more than a normal input token and have a short (default 5-minute) TTL, so it's only worth enabling once request volume is high enough that repeated calls reliably land inside that window. `WithPromptCaching` is Anthropic-only — OpenAI, Gemini, and DeepSeek ignore it (OpenAI caches automatically server-side with no client-side option needed; Gemini/DeepSeek have no equivalent in this SDK today).

See [Token Usage](/features/token-usage) for how to measure `CachedPromptTokens` and confirm caching is actually being hit.

## Gemini

```go theme={null}
import (
    "github.com/agenticenv/agent-sdk-go/pkg/llm"
    "github.com/agenticenv/agent-sdk-go/pkg/llm/gemini"
)

llmClient, err := gemini.NewClient(
    llm.WithAPIKey(os.Getenv("GOOGLE_API_KEY")),
    llm.WithModel("gemini-2.5-flash"),
)
```

<Note>
  `llm.WithAPIKey` takes precedence. If you omit it, the Gemini client falls back to the `GOOGLE_API_KEY` environment variable automatically.
</Note>

## DeepSeek

DeepSeek exposes an OpenAI-compatible API. The base URL defaults to `https://api.deepseek.com`, so `WithBaseURL` is optional.

```go theme={null}
import (
    "github.com/agenticenv/agent-sdk-go/pkg/llm"
    "github.com/agenticenv/agent-sdk-go/pkg/llm/deepseek"
)

llmClient, err := deepseek.NewClient(
    llm.WithAPIKey(os.Getenv("DEEPSEEK_API_KEY")),
    llm.WithModel("deepseek-chat"), // or "deepseek-reasoner"
)
```

Common models: `deepseek-chat` and `deepseek-reasoner`. Reasoning is not configured via the `Reasoning` field — `deepseek-reasoner` reasons automatically, and its trace surfaces via the streaming `LLMStreamChunk.ThinkingDelta` and, for non-streaming responses, `LLMResponse.Metadata["reasoning_content"]`.

<Note>
  DeepSeek supports only plain JSON mode, not schema-constrained output. A [response format](/features/response-format) with `Type: ResponseFormatJSON` maps to DeepSeek's `json_object`; any `Schema` is ignored (not enforced by the API), so prompt the model for the shape you want, and include the word "json" in the conversation.
</Note>

## Ollama

Ollama runs open models locally and exposes an OpenAI-compatible API, so no API key is needed. The base URL defaults to `http://localhost:11434/v1`; pass `WithBaseURL` to reach a remote host.

```go theme={null}
import (
    "github.com/agenticenv/agent-sdk-go/pkg/llm"
    "github.com/agenticenv/agent-sdk-go/pkg/llm/ollama"
)

llmClient, err := ollama.NewClient(
    llm.WithModel("llama3.2"), // any locally pulled model
    // llm.WithBaseURL("http://remote-host:11434/v1"), // optional
)
```

Common models: any model you have pulled with `ollama pull`, e.g. `llama3.2`, `qwen2.5`, `mistral`. Start the daemon with `ollama serve`. Reasoning models such as `deepseek-r1` and `qwen3` surface their trace via the streaming `LLMStreamChunk.ThinkingDelta` and, for non-streaming responses, `LLMResponse.Metadata["reasoning_content"]`.

### Ollama Cloud

[Ollama Cloud](https://ollama.com) hosts larger models behind an API key, for machines that can't run them locally. Supplying an API key — via `llm.WithAPIKey` or the `OLLAMA_API_KEY` environment variable — automatically switches the default base URL to Ollama Cloud's endpoint (`https://ollama.com/v1`); no other code change is needed. `llm.WithBaseURL` still overrides either default, e.g. for a self-hosted daemon behind an auth proxy.

```go theme={null}
llmClient, err := ollama.NewClient(
    llm.WithAPIKey(os.Getenv("OLLAMA_API_KEY")), // omit to fall back to the OLLAMA_API_KEY env var
    llm.WithModel("qwen3-coder:480b-cloud"),      // cloud models are suffixed "-cloud"
)
```

Generate a key at [ollama.com/settings/keys](https://ollama.com/settings/keys).

<Note>
  Not every Ollama model supports tool-calling or JSON mode, so tool use and response formats behave differently depending on which model you pull or select on Cloud. As with DeepSeek, a [response format](/features/response-format) with `Type: ResponseFormatJSON` maps to `json_object`; any `Schema` is ignored, so prompt the model for the shape you want.
</Note>

## Shared LLM options

All built-in clients accept functional options from `pkg/llm`:

| Option | Purpose |
| - | - |
| `llm.WithAPIKey(key)` | Provider API key |
| `llm.WithModel(model)` | Model name sent to the provider |
| `llm.WithBaseURL(url)` | Custom base URL (OpenAI-compatible endpoints, proxies) |

## Sampling and reasoning

Control temperature, max tokens, and reasoning per agent with [`WithLLMSampling`](/getting-started/configuration):

```go theme={null}
temp := 0.7
a, err := agent.NewAgent(
    agent.WithLLMClient(llmClient),
    agent.WithLLMSampling(&agent.LLMSampling{
        Temperature: &temp,
        MaxTokens:   4096,
    }),
)
```

Field support varies by provider:

| Field | OpenAI | Anthropic | Gemini | DeepSeek | Ollama |
| - | - | - | - | - | - |
| `Temperature` | ✓ | ✓ | ✓ | ✓ | ✓ |
| `MaxTokens` | ✓ | ✓ | ✓ | ✓ | ✓ |
| `TopP` | ✓ | — | ✓ | ✓ | ✓ |
| `TopK` | — | ✓ | — | — | — |
| `Reasoning` | Partial | ✓ | ✓ | — | — |

See [Reasoning](/features/reasoning) for extended thinking configuration per provider.

## Bring your own LLM client

Implement [`interfaces.LLMClient`](https://pkg.go.dev/github.com/agenticenv/agent-sdk-go/pkg/interfaces#LLMClient) and pass your implementation to `WithLLMClient`.

Your implementation must handle:

* **`Generate`** — synchronous completion; populate `LLMResponse.ToolCalls` when the provider returns tool calls
* **`GenerateStream`** — streaming chunks via `LLMStream.Next()` and `LLMStream.Current()` when `IsStreamSupported()` returns `true`
* **`GetModel`** / **`GetProvider`** — metadata that appears on `AgentRunResult.Model` and `AgentRunResult.AgentName`

<Tip>
  Copy the patterns from `pkg/llm/openai`, `pkg/llm/anthropic`, or `pkg/llm/gemini` as a starting point — they show request marshaling, response parsing, tool call handling, and stream chunk iteration.
</Tip>

## Related

<CardGroup cols={2}>
  <Card title="Quickstart" icon="bolt" href="/getting-started/quickstart" horizontal>
    Run your first agent end to end
  </Card>

  <Card title="Configuration" icon="sliders" href="/getting-started/configuration" horizontal>
    WithLLMClient, WithLLMSampling, and other options
  </Card>
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.