LLM Retry Policy
Framework-level retry with exponential backoff and jitter for transient LLM
failures. Complements (rather than replaces) LLMConfig.num_retries, which
is the SDK-side retry hint forwarded to the provider library.
Why two retry knobs?
| Knob | Runs where | Knobs exposed |
|---|---|---|
LLMConfig.num_retries | Inside the provider library (litellm, anthropic-sdk) | Count only |
LLMConfig.retry_policy | Outside the SDK, in the framework | Count, delay, backoff multiplier, cap, jitter, category filter |
Use num_retries for cheap transient-error recovery that the SDK can handle
itself. Use retry_policy when you need:
- Explicit backoff and jitter control (avoid thundering herds)
- A category filter (retry timeouts but not rate limits, or vice versa)
- Observability β retry attempts emit framework logs instead of disappearing into the SDK
Error categories
Each LLM implementation classifies provider-specific exceptions into one of:
"rate_limit"β HTTP 429, upstream throttling"server_error"β HTTP 5xx, provider-side faults"timeout"β connect/read timeouts
Authentication errors, invalid-request errors, and other permanent failures are deliberately not retried β no category covers them.
Basic usage
from augments.adk.llms import LLMConfig
from augments.adk.types.llms import LLMRetryPolicy
config = LLMConfig(
retry_policy=LLMRetryPolicy(
max_retries=5,
initial_delay=0.5,
max_delay=30.0,
multiplier=2.0,
jitter=True,
)
)Delay progression (no jitter, initial_delay=0.5, multiplier=2.0, max_delay=30):
attempt 0 β 0.5s
attempt 1 β 1.0s
attempt 2 β 2.0s
attempt 3 β 4.0s
attempt 4 β 8.0s
attempt 5 β 16.0s
attempt 6 β 30.0s (capped)
With jitter=True (default), each delay is randomized to
[0, computed_delay] β avoiding the thundering-herd failure mode where many
workers retry in lock-step after a shared upstream outage.
Category filtering
from augments.adk.types.llms import LLMRetryPolicy
# Retry server errors and timeouts, but give up immediately on rate limits
# (useful when upstream has a hard quota and retries are pointless).
policy = LLMRetryPolicy(
retry_on=frozenset(["server_error", "timeout"]),
)Interaction with streaming
Retries apply to non-streaming calls only. Reconnecting mid-stream would silently lose tokens or double-emit events, so streaming failures surface immediately. If you need retry semantics around a streamed call, wrap it in your own loop that discards the partial output and starts fresh.
See also
src/augments/adk/types/llms/retry_policy.pyβ dataclass definitionsrc/augments/adk/llms/litellm/litellm_retry.pyβ litellm exception classifier and retry looptests/unit/llms/test_retry_policy.pyβ tests for the policy and loop