Augments LabsAugments ADK

LLM Retry Policy

Framework-level retry with exponential backoff and jitter for transient LLM failures. Complements (rather than replaces) LLMConfig.num_retries, which is the SDK-side retry hint forwarded to the provider library.

Why two retry knobs?

KnobRuns whereKnobs exposed
LLMConfig.num_retriesInside the provider library (litellm, anthropic-sdk)Count only
LLMConfig.retry_policyOutside the SDK, in the frameworkCount, delay, backoff multiplier, cap, jitter, category filter

Use num_retries for cheap transient-error recovery that the SDK can handle itself. Use retry_policy when you need:

  • Explicit backoff and jitter control (avoid thundering herds)
  • A category filter (retry timeouts but not rate limits, or vice versa)
  • Observability β€” retry attempts emit framework logs instead of disappearing into the SDK

Error categories

Each LLM implementation classifies provider-specific exceptions into one of:

  • "rate_limit" β€” HTTP 429, upstream throttling
  • "server_error" β€” HTTP 5xx, provider-side faults
  • "timeout" β€” connect/read timeouts

Authentication errors, invalid-request errors, and other permanent failures are deliberately not retried β€” no category covers them.

Basic usage

from augments.adk.llms import LLMConfig
from augments.adk.types.llms import LLMRetryPolicy
 
config = LLMConfig(
    retry_policy=LLMRetryPolicy(
        max_retries=5,
        initial_delay=0.5,
        max_delay=30.0,
        multiplier=2.0,
        jitter=True,
    )
)

Delay progression (no jitter, initial_delay=0.5, multiplier=2.0, max_delay=30):

attempt 0 β†’ 0.5s
attempt 1 β†’ 1.0s
attempt 2 β†’ 2.0s
attempt 3 β†’ 4.0s
attempt 4 β†’ 8.0s
attempt 5 β†’ 16.0s
attempt 6 β†’ 30.0s  (capped)

With jitter=True (default), each delay is randomized to [0, computed_delay] β€” avoiding the thundering-herd failure mode where many workers retry in lock-step after a shared upstream outage.

Category filtering

from augments.adk.types.llms import LLMRetryPolicy
 
# Retry server errors and timeouts, but give up immediately on rate limits
# (useful when upstream has a hard quota and retries are pointless).
policy = LLMRetryPolicy(
    retry_on=frozenset(["server_error", "timeout"]),
)

Interaction with streaming

Retries apply to non-streaming calls only. Reconnecting mid-stream would silently lose tokens or double-emit events, so streaming failures surface immediately. If you need retry semantics around a streamed call, wrap it in your own loop that discards the partial output and starts fresh.

See also

  • src/augments/adk/types/llms/retry_policy.py β€” dataclass definition
  • src/augments/adk/llms/litellm/litellm_retry.py β€” litellm exception classifier and retry loop
  • tests/unit/llms/test_retry_policy.py β€” tests for the policy and loop