Augments LabsAugments ADK

Reasoning / Extended Thinking

Configure chain-of-thought reasoning across OpenAI, Anthropic, and Google Gemini models.

Quick Start

from augments.adk.agents import Agent
from augments.adk.llms import LLMConfig
from augments.adk.types.common import Reasoning
 
agent = Agent(
    name="Analyst",
    system_prompt="Think carefully before answering.",
    llm_config=LLMConfig(
        reasoning=Reasoning(effort="high"),
    ),
)

Configuration

The Reasoning class on LLMConfig controls three dimensions:

FieldTypeDescription
effort"none" / "minimal" / "low" / "medium" / "high" / "xhigh"How hard the model reasons
mode"auto" / "manual"Whether the model decides when to think, or the caller controls it
budgetint (tokens)Explicit token budget for reasoning (required when mode="manual")
summary"auto" / "concise" / "detailed"Level of detail for reasoning summaries
include_thoughtsboolWhether to include reasoning content in responses

Effort Only (All Providers)

The simplest configuration โ€” works across all reasoning-capable models:

Reasoning(effort="high")

Adaptive Mode (Anthropic)

Recommended for Claude Opus 4.6 and Sonnet 4.6. The model dynamically decides when and how much to think:

Reasoning(mode="auto", effort="high")

Explicit Budget (Anthropic / Gemini)

Control exactly how many tokens the model can use for reasoning:

Reasoning(mode="manual", budget=8000)
 
# Or simply set budget โ€” mode="manual" is inferred:
Reasoning(budget=8000)

Reasoning Summaries (OpenAI)

Get a summary of the model's reasoning process:

Reasoning(effort="high", summary="concise")

Include Thoughts (Gemini)

Include thought content in the response:

Reasoning(effort="high", include_thoughts=True)

Provider Support

OpenAI

FeatureSupport
effort"none", "minimal", "low", "medium", "high", "xhigh"
modeNot applicable (reasoning is always on for o-series)
budgetNot applicable (tokens reserved internally)
summary"auto", "concise", "detailed"
include_thoughtsNot applicable
Modelso3, o4-mini, o1, gpt-5, gpt-5-mini, gpt-5-nano

Notes:

  • Reserve minimum 25,000 tokens for max_output_tokens on complex tasks
  • Reasoning tokens count as output tokens and occupy context window
  • Avoid chain-of-thought prompts โ€” the model handles this internally

Anthropic

FeatureSupport
effort"low", "medium", "high" ("max" Opus 4.6 only โ€” use extra_args)
mode"auto" (adaptive) or "manual" (explicit budget)
budgetMust be less than max_output_tokens
summaryNot directly used (Claude 4+ returns summarized thinking automatically)
include_thoughtsNot directly used (thinking blocks always included when enabled)
ModelsClaude Opus 4.6, Sonnet 4.6, Opus 4.5, Sonnet 4

Notes:

  • mode="auto" (adaptive) is recommended for Opus 4.6 / Sonnet 4.6
  • budget is deprecated on Opus 4.6 / Sonnet 4.6 in favor of adaptive thinking
  • When using thinking + tools, only tool_choice: auto or none is allowed
  • Thinking blocks with signatures must be preserved between tool calls (handled automatically by the Runner)

Google Gemini

FeatureSupport
effort"minimal", "low", "medium", "high" (Gemini 3 models)
mode"manual" โ†’ thinkingBudget (Gemini 2.5)
budget0โ€“32768 tokens (Gemini 2.5); 0 disables (Flash only)
summaryNot applicable (use include_thoughts instead)
include_thoughtsMaps to thinkingConfig.includeThoughts
ModelsGemini 3.1 Pro, Gemini 3.0, Gemini 2.5 Pro/Flash

Notes:

  • "minimal" effort not supported on Gemini 3.1 Pro
  • Thinking cannot be disabled on Gemini 3.1 Pro and 2.5 Pro
  • Thought signatures are mandatory on Gemini 3 function calls (400 error if omitted โ€” handled automatically by the Runner)

Response Data

When reasoning is enabled, responses include additional fields:

result = await Runner.arun(agent, "Solve this problem...")
 
# Access via the LLM response in new_items
for item in result.new_items:
    if isinstance(item, dict) and item.get("role") == "assistant":
        # Reasoning content (unified string)
        reasoning = item.get("reasoning_content")
 
        # Thinking blocks (structured, with signatures)
        blocks = item.get("thinking_blocks")

Token Usage

Reasoning tokens are tracked in usage statistics:

usage = result.context.usage
print(f"Reasoning tokens: {usage.output_tokens_details.reasoning_tokens}")

Streaming

Reasoning content is streamed as "reasoning_delta" events:

result = Runner.run(agent, "Solve this...", stream=True, run_config=RunConfig(model="claude-opus-4-6"))
 
async for event in result.stream_events():
    if event.type == "raw_response_event":
        if event.data.type == "reasoning_delta":
            print(f"[thinking] {event.data.reasoning_delta}", end="")
        elif event.data.type == "content_delta":
            print(event.data.content_delta, end="")

Context Management

Thinking blocks accumulate tokens in conversation history. The framework provides clear_thinking_blocks() to remove older thinking content while preserving recent turns:

from augments.adk.context import ContextConfig
 
# Keep thinking blocks from the last 2 turns, clear older ones
config = ContextConfig(
    clear_thinking_blocks=True,
    thinking_turns_to_keep=2,
)

How It Works (Internal)

The Reasoning config on LLMConfig is provider-agnostic. The LiteLLM implementation resolves it into litellm-specific parameters via litellm_reasoning_resolver.py:

LLMConfig.reasoning (Reasoning)
    โ†“
resolve_reasoning_params()
    โ†“
LiteLLMReasoningParam
    โ”œโ”€โ”€ reasoning_effort: str  โ†’  litellm.acompletion(reasoning_effort=...)
    โ””โ”€โ”€ thinking: dict         โ†’  litellm.acompletion(thinking=...)

litellm then maps these to the appropriate provider-specific API parameters.