Fast
Some vendors serve a model faster for a higher price. /fast asks the model
in force for that speed, where its vendor serves one, and says what it costs
before you choose it.
Choosing it
/fast opens a panel over the model in force: its provider, its name and the
credential it is served by, what the vendor says of fast, and two rows,
Standard and Fast, with the price and the vendor's speed beneath Fast.
The mark opens on the speed in force, so Enter changes nothing by accident;
Escape leaves it as it was.
Speed · openai · gpt-6-astra · OpenAI key
OpenAI may serve a fast request at standard speed when fast capacity is short;
it is then billed at the standard price.
Standard
The standard price and speed
› Fast
2x the price · up to 2.5x faster
enter to choose · esc to cancel/fast on and /fast off choose without the panel. A window with no room for
the panel, and a run with no keyboard, print the speed in force and those two
lines, the price beside on. While a turn runs, /fast opens as /model
does, and the speed taken is asked for once the turn ends.
A model with no fast form, and a model that is itself a fast model, answer
/fast with one line and no panel:
⎿ anthropic · claude-sonnet-5 has no fast form
⎿ moonshot · kimi-for-coding-highspeed is a fast model of its own; /model lists the othersIn /model, a model with a fast form says fast at the end of its row, after
no rung where it has both.
What you are told afterwards
The label under the box ends in · fast only after an answer the vendor says
it served fast. A vendor that serves a fast request at standard speed says so
in its answer, and the label then leaves fast off until one is served fast
again; nothing is added to the transcript.
Where it is kept
/fast writes the model in force and "fast": true together under the
provider in the configuration file in your home directory, since the speed was
chosen at that model's price, and /fast off takes fast out:
{ "providers": { "openai": { "model": "gpt-6-astra", "fast": true } } }The model written is the one in force, even where --model or a project's
file chose it, so it becomes the one a plain start asks for where no project
file names another, and /fast off leaves it there.
Only that file is read for it. A project file that sets fast stops crucible
before it draws anything, as every key a checkout may not set does: fast costs
more on every request, and a repository could otherwise spend your money.
A request carries the fast form only where all of these hold: your file says
fast for the provider beside the model in force; the model in force has a
fast form in the table below for the credential in force; and no baseUrl is
set for the provider. A fast left in the file where one of these does not
hold is left there and asks for standard. Making room in the model's window
asks at the same speed as a turn.
When it goes back to standard
The speed goes back to standard, and fast is taken out of the file, when:
- another model is chosen for the provider in
/model. Choosing the model in force again keeps it; - a credential is stored, removed or replaced for the provider and the one in
force changes by it: a
/loginon another row, a key stored where the environment served, a/logoutthat leaves the environment serving, or a second credential taken out at a start. What fast costs was shown for the credential it was chosen under. A key written again on the same row, and a change of the environment between two runs, move nothing; - the vendor refuses the fast form, as below.
When the vendor refuses it
Where a vendor answers a fast request with a refusal of fast before anything else, crucible sends the same message once more at standard speed, turns fast off, and says why in one line:
⎿ openai refused fast: Invalid service_tier argument: The requested service tier is not allowed for this project. Sent again at standard speed; fast is off.Stopping the turn while the message goes again leaves the line ending Fast is off. instead.
Any other error is reported as it would be at standard speed, and nothing is sent again. What counts as a refusal of fast:
- OpenAI with a key: a 400 whose error is
invalid_request_errorfor the parameterservice_tier, as OpenAI's error reference documents. - OpenAI with a ChatGPT sign-in: none is documented, so none is taken for one.
- Anthropic: a 400
invalid_request_errorwhose message names fast mode orspeed, and a 429 saying "Usage credits are required for fast mode.". Anthropic says only that such a request "returns an error"; this rule is taken from open harnesses that are not Anthropic's. - Google: none is documented. Congestion does not refuse; it serves the request at standard speed, which the answer then says.
- Kimi Code:
kimi-for-coding-highspeedis a model of its own and asks for nothing, so there is nothing to refuse. A plan without it answers with an error like any other, and nothing is sent again.
The fast forms
Each price, caveat and speed is the vendor's, as the panel shows it: the speed
beside the price under Fast, where the vendor states one.
| Provider | Credential | Models | Price | Caveat | Speed | Source | Read |
|---|---|---|---|---|---|---|---|
| OpenAI | key | gpt-6-astra, gpt-5.6-sol, gpt-5.6-terra, gpt-5.6-luna | 2x the price | OpenAI may serve a fast request at standard speed when fast capacity is short; it is then billed at the standard price. | up to 2.5x faster | OpenAI fast mode, OpenAI pricing | 30 Sep 2026 |
| OpenAI | key | gpt-5.5 | 2.5x the price | OpenAI may serve a fast request at standard speed when fast capacity is short; it is then billed at the standard price. | up to 2.5x faster | OpenAI fast mode, OpenAI pricing | 30 Sep 2026 |
| OpenAI | sign-in | gpt-6-astra | 2.5x your plan's usage; 2x purchased credits | Not stated | ChatGPT speed | 30 Sep 2026 | |
| OpenAI | sign-in | gpt-5.6-sol, gpt-5.6-terra, gpt-5.6-luna | 2.5x your plan's usage; 2x purchased credits | 1.5x faster | ChatGPT speed | 30 Sep 2026 | |
| OpenAI | key | gpt-6.1-sol, gpt-6-sol, gpt-6-luna | 2x the price | OpenAI may serve a fast request at standard speed when fast capacity is short; it is then billed at the standard price. | up to 2.5x faster | OpenAI fast mode, OpenAI pricing | 1 Oct 2026 |
| OpenAI | sign-in | gpt-6.1-sol, gpt-6-sol, gpt-6-luna | 2.5x your plan's usage; 2x purchased credits | Not stated | ChatGPT speed | 1 Oct 2026 | |
| Anthropic | key | claude-opus-5 | $10 / $50 per million input / output tokens | Fast mode is a research preview; Anthropic turns it on per organization. | up to 2.5x higher output tokens per second | Anthropic fast mode | 30 Sep 2026 |
| Anthropic | key | claude-opus-5-5 | $8 / $40 per million input / output tokens | Fast mode is a research preview; Anthropic turns it on per organization. | up to 2.5x higher output tokens per second | Anthropic fast mode | 1 Oct 2026 |
| key | gemini-3.8-flash, gemini-3.7-flash, gemini-3.6-flash, gemini-3.1-pro-preview | 75-100% more than Standard | For Tier 2 and Tier 3 accounts only. Google serves a priority request at standard speed when priority is congested, and bills it at the standard price. | Not stated | Gemini priority inference | 30 Sep 2026 | |
| MoonshotAI | any | kimi-for-coding-highspeed, a fast model of its own | 3x the quota | 6x the speed | Kimi Code models | 29 Sep 2026 | |
| DeepSeek | key | None | DeepSeek serves no faster tier and no faster model. | DeepSeek models and pricing | 1 Oct 2026 | ||
| Meta | key | None | Meta's reference lists a service_tier field, priority among its values, with no price, no speed and no stated behaviour, so no source settles what it costs, how it is refused or how an answer says its speed. | Meta Create Response | 1 Oct 2026 | ||
| xAI | key | None in this release | xAI serves a priority tier at 2x the price, and bills a request it served at standard speed at the standard price. No source shows how xAI refuses a priority request. | xAI priority processing | 1 Oct 2026 | ||
| Z.ai | key | None in this release | glm-5.3-flashx is a faster glm-5.3-flash, at 200 tokens/s for a higher price. No source shows how Z.ai refuses it or prints an answer from it. | Z.ai glm-5.3-flash | 1 Oct 2026 | ||
| Qwen | any | None in this release | qwen3.8-max-prime, priced on the mainland China site only, is a faster qwen3.8-max. No source shows how an answer says the speed it was served at, or how a refusal reads, and the vendor's own pages on its fast mode do not list it. | Model Studio pricing | 1 Oct 2026 | ||
| MiMo | key | None in this release | mimo-v2.6-pro-ultraspeed is a faster mimo-v2.6-pro, up to 20x faster at 10x the price. No source shows how an answer says the speed it was served at, or how a refusal reads. | MiMo pay-as-you-go pricing | 1 Oct 2026 | ||
| MiniMax | any | None in this release | MiniMax-M3 takes a priority tier and MiniMax-M2.7 has a faster MiniMax-M2.7-highspeed, each at a higher price. No source shows how a Chat Completions answer says the tier it was served at, or how either is refused. | MiniMax pay-as-you-go pricing | 1 Oct 2026 |
OpenAI states twice the price for GPT-6 Astra, GPT-6.1 Sol and GPT-5.6 Sol; for GPT-6 Sol, GPT-6 Luna, GPT-5.6 Terra and Luna, and for GPT-5.5, the multiple is worked out from the prices on its pricing page. The ChatGPT speed page states no speed for the GPT-6 models or GPT-6.1 Sol. OpenAI serves no fast form of the GPT-6 models or GPT-6.1 Sol with EU data residency.
A provider whose row says None in this release serves a faster form, but no
source settled how it refuses one, or how an answer says which speed served
it; crucible offers fast only where both are known. A row that says None
names a vendor that documents no faster form crucible could ask for.