Prompt caching
Prompt caching lets a provider reuse an identical logical prompt prefix. It is
an optimization inside a provider request: crucible still sends the request,
still asks the model for a new answer, and still preserves the complete logical
instructions, visible tools, transcript, attachments and model parameters.
It is not response caching, conversation storage, compaction,
previous_response_id or connection reuse.
The default
Prompt caching is on by default through each provider's own reviewed mechanism.
The default policy is prefer, session-isolated, provider-default retention,
and forbids persistent cached-content resources.
| Provider | Default behavior | Request change |
|---|---|---|
| OpenAI | Provider-managed implicit prefix caching | An opaque, derived routing key scoped to the session by default. |
| Anthropic | Automatic short-lived prefix caching | Top-level cache_control: {"type":"ephemeral"}. |
| Google Gemini | Provider-managed implicit prefix caching | No additional cache field or remote cached-content resource. |
| Moonshot/Kimi | Provider-managed automatic context caching | An opaque, derived prompt_cache_key scoped to the session by default; Kimi manages cache creation and lifetime. |
| Meta, xAI | Provider-managed automatic prefix caching | An opaque, derived prompt_cache_key scoped to the session by default; the vendor manages cache creation and lifetime. |
| DeepSeek, Z.ai, Qwen, MiMo, MiniMax | Provider-managed automatic prefix caching | None. These vendors take no cache field on the route crucible uses; crucible reads the cached-token count each answer reports. |
Only exact built-in endpoint/model records advertise support. A custom
baseUrl, proxy route or unreviewed model is unknown; crucible sends no
speculative cache field. prefer then sends the unchanged request.
observeOnly explicitly asks crucible to add no cache control. It does not
promise that a provider with unavoidable automatic caching will stop caching.
require fails before a request when no reviewed, policy-permitted mechanism is
eligible. prohibit also fails before sending unless the exact provider/model
record has a real cache opt-out that the adapter can encode.
Configuration
The promptCaching object is policy, not a place for provider wire fields:
{
"promptCaching": {
"mode": "prefer",
"allowedMechanisms": ["automaticPrefix", "explicitBreakpoints"],
"isolationScope": "session",
"requestedRetention": {
"class": "ephemeral",
"maxSeconds": 1800
},
"persistentResources": { "mode": "forbid" },
"namespace": "personal-agent"
}
}The canonical values are:
mode:observeOnly,prefer,require,prohibitallowedMechanisms:providerManagedUsageOnly,automaticPrefix,explicitBreakpoints,persistentContentisolationScope:run,session,workspace,user- retention
class:providerDefault,ephemeral,extended - persistent-resource
mode:forbid,reuse,create,require
maxSeconds is a hard ceiling, not a promise that a provider offers that exact
TTL. Extended retention and resource creation must originate in the user
configuration. Project and local project files can only narrow inherited mode,
mechanisms, isolation, retention and resource authority; they cannot turn
caching on, broaden sharing, lengthen retention, choose a namespace or
authorize a remote resource.
Persistent resources are a separate opt-in. The default performs no startup request, creates no cleanup worker and creates no local cache directory. When explicitly authorized, crucible stores only bounded resource metadata in an owner-only user directory. Prompt text, responses and credentials are never stored there. Provider and owner identities are retained only as redacted digests. Model, provider and active-credential switches run one bounded cleanup pass for the current run/session's exclusive resources before the switch; workspace- and user-shared resources remain until explicit cleanup or expiry.
Use /cache or /cache inspect to see the resolved policy, declared support,
predicted eligibility, actual wire encoding, request disposition,
provider-reported outcome, normalized usage/cost and redacted resource state.
Use /cache cleanup for one bounded, cancellable cleanup pass. A cache read is
called a hit only when provider-reported usage says so; a local fingerprint
match or accepted request is not a hit.
Shipped capability records
Each record's official source and version are compiled into its adapter, so ordinary startup never scrapes mutable documentation.
| Adapter and exact models | Mechanisms | Official source | Record version |
|---|---|---|---|
OpenAI Responses: gpt-6-astra | implicit caching and up to four explicit input-content breakpoints on the public API | OpenAI prompt caching | openai-prompt-cache-2026-09-06 |
Anthropic Messages: claude-fable-5-1 | automatic control and up to four legal block breakpoints; 5-minute and 1-hour classes | Anthropic preserved thinking | anthropic-prompt-cache-2026-09-06 |
Google Interactions: gemini-3.8-flash, gemini-3.7-flash, gemini-3.6-flash, gemini-3.1-pro-preview | implicit caching; no explicit cache resource or retention control on this route | Google caching | google-interactions-cache-2026-09-06 |
OpenAI Responses: gpt-5.6-sol, gpt-5.6-terra, gpt-5.6-luna | implicit automatic caching; up to four explicit input-content breakpoints on the public API | OpenAI prompt caching | openai-prompt-cache-2026-08-31 |
OpenAI Responses: gpt-5.5 | implicit automatic caching; optional documented 24-hour retention | OpenAI prompt caching | openai-prompt-cache-2026-08-31 |
Anthropic Messages: claude-fable-5, claude-opus-5, claude-sonnet-5, claude-haiku-4-5 | top-level automatic control and up to four explicit block breakpoints; 5-minute and 1-hour classes | Anthropic prompt caching, tool use | anthropic-prompt-cache-2026-08-31 |
OpenAI Responses: gpt-6.1-sol, gpt-6-sol, gpt-6-luna | implicit automatic caching; up to four explicit input-content breakpoints on the public API | OpenAI prompt caching | openai-prompt-cache-2026-10-01 |
Anthropic Messages: claude-opus-5-5, claude-sonnet-5-5 | top-level automatic control and up to four explicit block breakpoints; 5-minute and 1-hour classes | Anthropic prompt caching | anthropic-prompt-cache-2026-10-01 |
Meta Responses: muse-spark-1.3, muse-spark-1.3-contributor, muse-spark-1.2, muse-spark-1.2-contributor | automatic prefix caching with a routing key | Meta prompt caching | meta-prompt-cache-2026-10-01 |
xAI Responses: grok-4.7, grok-4.6 | automatic prefix caching with a routing key | xAI prompt caching | xai-prompt-cache-2026-10-01 |
DeepSeek Chat Completions: deepseek-flash, deepseek-v4-pro | provider-managed caching, reported usage only | DeepSeek context caching | deepseek-prompt-cache-2026-10-01 |
Z.ai Chat Completions: glm-5.3, glm-5.3-flash, glm-5.2 | provider-managed caching, reported usage only | Z.ai context caching | zai-prompt-cache-2026-10-01 |
Qwen Chat Completions: qwen3.8-max, qwen3.8-flash, qwen3.7-plus, qwen3.6-plus | provider-managed caching, reported usage only | Model Studio context cache | qwen-prompt-cache-2026-10-01 |
MiMo Chat Completions: mimo-v2.6-pro, mimo-v2.6-flash | provider-managed caching, reported usage only | MiMo pricing | mimo-prompt-cache-2026-10-01 |
MiniMax Chat Completions: MiniMax-M3, MiniMax-M2.7 | provider-managed caching from 512 input tokens, reported usage only | MiniMax prompt caching | minimax-prompt-cache-2026-10-01 |
Moonshot/Kimi Chat Completions: k3, k3-256k, kimi-for-coding, kimi-for-coding-highspeed | provider-managed automatic caching after the documented prefix minimum, with a stable routing key for agent sessions | Kimi context caching and Chat Completions API | kimi-prompt-cache-2026-08-31 |
The OpenAI public pricing record uses the current OpenAI pricing and model-specific documentation. The Anthropic record uses the pricing table in the prompt-caching guide. Pricing is selected only for an exact protocol, endpoint, model, revision, date, retention class and input band. Subscription billing and Moonshot membership billing remain unknown rather than being invented as token prices.
For the seven vendors added in 0.44, what no source settled is said here
rather than guessed at. None of Meta, xAI, DeepSeek, Z.ai or MiMo publishes
the smallest prefix it caches, so crucible's records assume no hit below 1,024
tokens, the floor Qwen publishes for its own. Qwen
lists qwen3.6-plus on neither of its implicit cache's lists while that
model's own page says it takes one, so crucible claims only the reading of
what an answer reports for it. Qwen does not document caching at its plan
addresses, Z.ai and MiniMax state no cache lifetime, and MiMo's examples never
show whether its cached count lies inside the prompt's; crucible reads it as
inside. Qwen's explicit cache_control markers are not written.
Private continuation is distinct from a cache hit. Compatible Google and Astra native history is replayed in order; Fable thinking whose bound prefix was rewritten by compaction or pruning is omitted. Public text and settled tool results remain available, but prefix changes can reduce cache reuse. Only the provider's reported cached-token usage proves a hit. Native web-tool charges and account-specific allowances are not inferred from token prices.
Persistent-storage quantities, when a future lifecycle adapter reports them, are priced in token-hours rather than raw tokens. A same-currency total may therefore combine token and token-hour rate provenance; missing storage usage or pricing keeps the dependent total unknown.
Protocol references
Additional official protocol documentation is listed below. Gemini's shipped Interactions route is described above; these links do not imply that Crucible ships the other adapters.
Adapter conformance
An adapter must keep vendor fields inside its own wire module, intersect its
current encoding/parsing ability with an exact endpoint/model record, and leave
custom routes unknown until their upstream semantics are resolved. It must
prove default and observeOnly wire behavior, legal bounded controls, inclusive
versus disjoint usage accounting, unknown-preserving pricing, retry/cancellation
attempt identity, and redaction of keys, handles, prompt content and
credentials. Stateful response continuation remains a separate capability and
never satisfies prompt-cache support.