Augments LabsCrucible Code

Prompt caching

Prompt caching lets a provider reuse an identical logical prompt prefix. It is an optimization inside a provider request: crucible still sends the request, still asks the model for a new answer, and still preserves the complete logical instructions, visible tools, transcript, attachments and model parameters. It is not response caching, conversation storage, compaction, previous_response_id or connection reuse.

The default

Prompt caching is on by default through each provider's own reviewed mechanism. The default policy is prefer, session-isolated, provider-default retention, and forbids persistent cached-content resources.

ProviderDefault behaviorRequest change
OpenAIProvider-managed implicit prefix cachingAn opaque, derived routing key scoped to the session by default.
AnthropicAutomatic short-lived prefix cachingTop-level cache_control: {"type":"ephemeral"}.
Google GeminiProvider-managed implicit prefix cachingNo additional cache field or remote cached-content resource.
Moonshot/KimiProvider-managed automatic context cachingAn opaque, derived prompt_cache_key scoped to the session by default; Kimi manages cache creation and lifetime.
Meta, xAIProvider-managed automatic prefix cachingAn opaque, derived prompt_cache_key scoped to the session by default; the vendor manages cache creation and lifetime.
DeepSeek, Z.ai, Qwen, MiMo, MiniMaxProvider-managed automatic prefix cachingNone. These vendors take no cache field on the route crucible uses; crucible reads the cached-token count each answer reports.

Only exact built-in endpoint/model records advertise support. A custom baseUrl, proxy route or unreviewed model is unknown; crucible sends no speculative cache field. prefer then sends the unchanged request.

observeOnly explicitly asks crucible to add no cache control. It does not promise that a provider with unavoidable automatic caching will stop caching. require fails before a request when no reviewed, policy-permitted mechanism is eligible. prohibit also fails before sending unless the exact provider/model record has a real cache opt-out that the adapter can encode.

Configuration

The promptCaching object is policy, not a place for provider wire fields:

{
  "promptCaching": {
    "mode": "prefer",
    "allowedMechanisms": ["automaticPrefix", "explicitBreakpoints"],
    "isolationScope": "session",
    "requestedRetention": {
      "class": "ephemeral",
      "maxSeconds": 1800
    },
    "persistentResources": { "mode": "forbid" },
    "namespace": "personal-agent"
  }
}

The canonical values are:

  • mode: observeOnly, prefer, require, prohibit
  • allowedMechanisms: providerManagedUsageOnly, automaticPrefix, explicitBreakpoints, persistentContent
  • isolationScope: run, session, workspace, user
  • retention class: providerDefault, ephemeral, extended
  • persistent-resource mode: forbid, reuse, create, require

maxSeconds is a hard ceiling, not a promise that a provider offers that exact TTL. Extended retention and resource creation must originate in the user configuration. Project and local project files can only narrow inherited mode, mechanisms, isolation, retention and resource authority; they cannot turn caching on, broaden sharing, lengthen retention, choose a namespace or authorize a remote resource.

Persistent resources are a separate opt-in. The default performs no startup request, creates no cleanup worker and creates no local cache directory. When explicitly authorized, crucible stores only bounded resource metadata in an owner-only user directory. Prompt text, responses and credentials are never stored there. Provider and owner identities are retained only as redacted digests. Model, provider and active-credential switches run one bounded cleanup pass for the current run/session's exclusive resources before the switch; workspace- and user-shared resources remain until explicit cleanup or expiry.

Use /cache or /cache inspect to see the resolved policy, declared support, predicted eligibility, actual wire encoding, request disposition, provider-reported outcome, normalized usage/cost and redacted resource state. Use /cache cleanup for one bounded, cancellable cleanup pass. A cache read is called a hit only when provider-reported usage says so; a local fingerprint match or accepted request is not a hit.

Shipped capability records

Each record's official source and version are compiled into its adapter, so ordinary startup never scrapes mutable documentation.

Adapter and exact modelsMechanismsOfficial sourceRecord version
OpenAI Responses: gpt-6-astraimplicit caching and up to four explicit input-content breakpoints on the public APIOpenAI prompt cachingopenai-prompt-cache-2026-09-06
Anthropic Messages: claude-fable-5-1automatic control and up to four legal block breakpoints; 5-minute and 1-hour classesAnthropic preserved thinkinganthropic-prompt-cache-2026-09-06
Google Interactions: gemini-3.8-flash, gemini-3.7-flash, gemini-3.6-flash, gemini-3.1-pro-previewimplicit caching; no explicit cache resource or retention control on this routeGoogle cachinggoogle-interactions-cache-2026-09-06
OpenAI Responses: gpt-5.6-sol, gpt-5.6-terra, gpt-5.6-lunaimplicit automatic caching; up to four explicit input-content breakpoints on the public APIOpenAI prompt cachingopenai-prompt-cache-2026-08-31
OpenAI Responses: gpt-5.5implicit automatic caching; optional documented 24-hour retentionOpenAI prompt cachingopenai-prompt-cache-2026-08-31
Anthropic Messages: claude-fable-5, claude-opus-5, claude-sonnet-5, claude-haiku-4-5top-level automatic control and up to four explicit block breakpoints; 5-minute and 1-hour classesAnthropic prompt caching, tool useanthropic-prompt-cache-2026-08-31
OpenAI Responses: gpt-6.1-sol, gpt-6-sol, gpt-6-lunaimplicit automatic caching; up to four explicit input-content breakpoints on the public APIOpenAI prompt cachingopenai-prompt-cache-2026-10-01
Anthropic Messages: claude-opus-5-5, claude-sonnet-5-5top-level automatic control and up to four explicit block breakpoints; 5-minute and 1-hour classesAnthropic prompt cachinganthropic-prompt-cache-2026-10-01
Meta Responses: muse-spark-1.3, muse-spark-1.3-contributor, muse-spark-1.2, muse-spark-1.2-contributorautomatic prefix caching with a routing keyMeta prompt cachingmeta-prompt-cache-2026-10-01
xAI Responses: grok-4.7, grok-4.6automatic prefix caching with a routing keyxAI prompt cachingxai-prompt-cache-2026-10-01
DeepSeek Chat Completions: deepseek-flash, deepseek-v4-proprovider-managed caching, reported usage onlyDeepSeek context cachingdeepseek-prompt-cache-2026-10-01
Z.ai Chat Completions: glm-5.3, glm-5.3-flash, glm-5.2provider-managed caching, reported usage onlyZ.ai context cachingzai-prompt-cache-2026-10-01
Qwen Chat Completions: qwen3.8-max, qwen3.8-flash, qwen3.7-plus, qwen3.6-plusprovider-managed caching, reported usage onlyModel Studio context cacheqwen-prompt-cache-2026-10-01
MiMo Chat Completions: mimo-v2.6-pro, mimo-v2.6-flashprovider-managed caching, reported usage onlyMiMo pricingmimo-prompt-cache-2026-10-01
MiniMax Chat Completions: MiniMax-M3, MiniMax-M2.7provider-managed caching from 512 input tokens, reported usage onlyMiniMax prompt cachingminimax-prompt-cache-2026-10-01
Moonshot/Kimi Chat Completions: k3, k3-256k, kimi-for-coding, kimi-for-coding-highspeedprovider-managed automatic caching after the documented prefix minimum, with a stable routing key for agent sessionsKimi context caching and Chat Completions APIkimi-prompt-cache-2026-08-31

The OpenAI public pricing record uses the current OpenAI pricing and model-specific documentation. The Anthropic record uses the pricing table in the prompt-caching guide. Pricing is selected only for an exact protocol, endpoint, model, revision, date, retention class and input band. Subscription billing and Moonshot membership billing remain unknown rather than being invented as token prices.

For the seven vendors added in 0.44, what no source settled is said here rather than guessed at. None of Meta, xAI, DeepSeek, Z.ai or MiMo publishes the smallest prefix it caches, so crucible's records assume no hit below 1,024 tokens, the floor Qwen publishes for its own. Qwen lists qwen3.6-plus on neither of its implicit cache's lists while that model's own page says it takes one, so crucible claims only the reading of what an answer reports for it. Qwen does not document caching at its plan addresses, Z.ai and MiniMax state no cache lifetime, and MiMo's examples never show whether its cached count lies inside the prompt's; crucible reads it as inside. Qwen's explicit cache_control markers are not written.

Private continuation is distinct from a cache hit. Compatible Google and Astra native history is replayed in order; Fable thinking whose bound prefix was rewritten by compaction or pruning is omitted. Public text and settled tool results remain available, but prefix changes can reduce cache reuse. Only the provider's reported cached-token usage proves a hit. Native web-tool charges and account-specific allowances are not inferred from token prices.

Persistent-storage quantities, when a future lifecycle adapter reports them, are priced in token-hours rather than raw tokens. A same-currency total may therefore combine token and token-hour rate provenance; missing storage usage or pricing keeps the dependent total unknown.

Protocol references

Additional official protocol documentation is listed below. Gemini's shipped Interactions route is described above; these links do not imply that Crucible ships the other adapters.

Adapter conformance

An adapter must keep vendor fields inside its own wire module, intersect its current encoding/parsing ability with an exact endpoint/model record, and leave custom routes unknown until their upstream semantics are resolved. It must prove default and observeOnly wire behavior, legal bounded controls, inclusive versus disjoint usage accounting, unknown-preserving pricing, retry/cancellation attempt identity, and redaction of keys, handles, prompt content and credentials. Stateful response continuation remains a separate capability and never satisfies prompt-cache support.