CROWBOTDOCS
CHAT LLMS.TXT DISCORD LOG IN

Models

Everything below is rendered from the live model table, so it is what the gateway will actually charge and serve right now.

model id in cached in out list in list out context max output status
CROW 2 crow-2 $6.00 $0.60 $10.00 — — 1M 32.8K live
CROW 1 crow-1 $3.00 $0.30 $5.00 — — 262K 32.8K live
GPT 6 Astra gpt-6-astra $0.30 $0.03 $1.50 $10.00 $50.00 1M 272K live
Claude Fable 5.1 claude-fable-5.1 $1.50 $0.0375 $7.50 $10.00 $50.00 1M 128K live
GPT 6 Sol gpt-6-sol $0.18 $0.018 $0.90 $2.00 $10.00 1M 128K live
Kimi K3 kimi-k3 $0.27 $0.027 $1.35 $3.00 $15.00 1M 1M live
GPT 5.6 Sol gpt-5.6-sol $0.08 $0.008 $0.40 $4.00 $20.00 1M 128K live
DeepSeek V4.1 Flash deepseek-v4.1-flash $0.045 $0.0009 $0.18 $0.30 $1.20 1M 128K live
Gemini 3.8 Flash gemini-3.8-flash $0.0675 $0.0067 $0.3375 $0.75 $3.75 1M 65.5K live
Claude Opus 5.5 claude-opus-5-5 — — — — — 1M 128K coming_soon
Claude Opus 5 claude-opus-5 $0.15 $0.015 $0.75 $5.00 $25.00 1M 128K deprecated

Rates are US dollars per million tokens. Input, cached input and output are billed separately at their own rate — there is no single blended price, because a request's cost depends on its shape.

Cached input bills at the cached rate whenever the upstream reports a hit, at the ratio we get. Not every model reports hits, so budget for the full input rate, and keep the stable part of your prompt first: one changed token near the front invalidates everything after it.

Directions

One key reaches every direction below.

Frontier models, resold. The frontier models are sold at a fixed discount off the vendor's own published price. What you pay therefore moves only when the vendor moves its list price; what it costs us to serve you is our problem and changes our margin instead of your bill. The list in and list out columns are that published price, so you can check the discount yourself; a row without them is not resold. How the resellers are tested, and what we keep: why it's this cheap.

Uncensored models, our own. The CROW line runs abliterated weights with no policy layer at the model level, at fixed rates. Same endpoint, same key, same wallet.

There is no cached output rate, here or at any vendor. Caching stores the state of a prompt prefix; output tokens are generated fresh on every call, so there is nothing to reuse.

Statuses

status in the chat picker callable via /v1
live selectable yes
coming_soon shown, greyed out no
deprecated not shown no
hidden not shown no

A retired model keeps its row, so past usage stays attributable. Naming a non-live model over /v1 is a 404 model_not_found listing what is live; an id we do not recognise at all falls back to the default, since agentic clients routinely send other providers' names.

Choosing

  • CROW 2 (crow-2) — Zero Refusal. Deepest Reasoning. Ideal for Red Teaming.
    $6.00 in / $10.00 out per 1M, 1M context.
  • CROW 1 (crow-1) — The cheaper one. Surprisingly creative, though.
    $3.00 in / $5.00 out per 1M, 262K context.
  • GPT 6 Astra (gpt-6-astra) — OpenAI's latest baby. Good for anything.
    $0.30 in / $1.50 out per 1M, 1M context (list price $10.00 in / $50.00 out).
  • GPT 6 Sol (gpt-6-sol) — OpenAI's fast one, a generation on.
    $0.18 in / $0.90 out per 1M, 1M context (list price $2.00 in / $10.00 out).
  • Kimi K3 (kimi-k3) — Moonshot's coding specialist.
    $0.27 in / $1.35 out per 1M, 1M context (list price $3.00 in / $15.00 out).
  • GPT 5.6 Sol (gpt-5.6-sol) — OpenAI's fast one. Cheap and cheerful.
    $0.08 in / $0.40 out per 1M, 1M context (list price $4.00 in / $20.00 out).
  • DeepSeek V4.1 Flash (deepseek-v4.1-flash) — DeepSeek's fast one. Cheap enough to leave running.
    $0.045 in / $0.18 out per 1M, 1M context — cached input $0.0009, 50x cheaper (list price $0.30 in / $1.20 out).
  • Gemini 3.8 Flash (gemini-3.8-flash) — Google's fast one. A million tokens of context, quick to answer.
    $0.0675 in / $0.3375 out per 1M, 1M context (list price $0.75 in / $3.75 out).

Every live model streams. Reasoning models return their thinking trace in reasoning_content; /v1/models says per model which of tool calling, reasoning, temperature and prompt caching it has, under capabilities.

Pick on price; the lineup above says which is the default, and omitting model gets it.

Limits

context is the prompt window; max output is the ceiling on one response, clamped rather than rejected. Both are editable without a deploy, so treat /v1/models as authoritative.