Night Shift docs

Night Shift is one API in front of every model you use. It takes requests in the OpenAI or the Anthropic format, grades how hard each one is, and sends it to the cheapest model on your list that can do it. When your top model's budget runs low, more work moves down the list. When everything is out, Night Shift's own GPU model answers instead of an error.

Night Shift is paid for by burning $SHIFT. A workspace's API and MCP server work while it has credit from a burn, and need no provider keys: requests go to Night Shift's GPU model until you add your own.

Quickstart

  1. Open the app and connect a wallet, or continue without one. You get a workspace and a Night Shift key that starts with ns_.
  2. Burn $SHIFT on the Credits page. The API and the MCP server work once the workspace has credit.
  3. Copy the key from the Get started panel.
  4. Send a request with any OpenAI or Anthropic client. Use any model name; Night Shift picks the model and says which in the x-nightshift-model header.
from openai import OpenAI

client = OpenAI(base_url="https://nightshiftcode.com/v1", api_key="YOUR_NIGHTSHIFT_KEY")
reply = client.chat.completions.create(
    model="nightshift",
    messages=[{"role": "user", "content": "Hello from Night Shift"}],
)
print(reply.choices[0].message.content)

Keys and sign-in

Send the Night Shift key as Authorization: Bearer ns_... or as x-api-key: ns_.... Both work on every endpoint, so OpenAI and Anthropic clients need nothing special. Night Shift stores only a hash of each key and shows a key once, when it's made.

A workspace made with a wallet is linked to it. Signing in with that wallet again makes a new key for the same workspace, so a lost key never loses the workspace. A workspace made without a wallet depends on its key until you link one on the Credits page or from the Get started panel. Linking signs a one-time message and sends no transaction.

Make more keys, see when each was last used, and revoke them on the Keys page.

Connect your tools

Claude Code reads its address and key from environment variables. Leave ANTHROPIC_API_KEY empty so it uses the Night Shift key.

export ANTHROPIC_BASE_URL=https://nightshiftcode.com
export ANTHROPIC_AUTH_TOKEN=YOUR_NIGHTSHIFT_KEY
export ANTHROPIC_API_KEY=
claude

Anything that speaks chat completions takes Night Shift's address as its base URL: the OpenAI SDKs, Cursor, Continue, Aider, LangChain. Use https://nightshiftcode.com/v1 as the base URL, the Night Shift key as the API key and nightshift as the model. Anthropic SDKs take https://nightshiftcode.com as the base URL.

Codex uses OpenAI's Responses API, which Night Shift doesn't serve yet.

API reference

EndpointWhat it does
POST /v1/chat/completionsOpenAI chat completions, streaming or not, with tools.
POST /v1/messagesAnthropic messages, streaming or not, with tools and thinking where the model supports them.
POST /v1/messages/count_tokensToken count for a messages request. Exact with an Anthropic key on your top model, an estimate otherwise.
GET /v1/modelsThe models on your ladder, plus nightshift.

Replies come straight from the model that answered, in the format you asked in. Two headers say what happened: x-nightshift-model is the model that answered and x-nightshift-rung is its place on your ladder, 0 being the top. Request bodies can be up to 25 MB, and long streams are never cut off by Night Shift.

MCP server

Night Shift is also an MCP server, for agents that would rather call a tool than an API. Connect to https://nightshiftcode.com/mcp with the Night Shift key as a Bearer token, or to https://nightshiftcode.com/mcp/YOUR_NIGHTSHIFT_KEY for clients that can't set headers. Like the API, it needs credit from a burn.

claude mcp add --transport http nightshift https://nightshiftcode.com/mcp --header "Authorization: Bearer YOUR_NIGHTSHIFT_KEY"
ToolWhat it does
shift_askSends one prompt through your ladder and returns the answer, the model that wrote it, the cost and the time.
shift_swarmRuns a list of prompts at once on Night Shift's GPU and returns straight away with a swarm id.
shift_swarm_resultsWhere a swarm is and its answers so far, in pages.
shift_swarm_stopStops a swarm; waiting jobs are dropped and running ones cut off.
shift_creditsThe workspace's credit and whether Night Shift's GPU is running.
shift_modelsYour ladder, top first, and the fallback model.

Swarm

Swarm is for batches where each prompt is short and each answer is long: a hundred tickets to label, forty variants of a page to draft, a test to write for every function in a list. Instead of forty shift_ask calls in a row, your agent sends the whole list in one shift_swarm call and gets a swarm id back at once. Night Shift runs the jobs on its GPU 4 at a time, or 16 with Pro, while the agent carries on.

The agent collects the answers with shift_swarm_results, which waits up to 40 seconds for the swarm to finish and then returns the answers by index, in the order the prompts were sent. Long result sets come in pages; when next is a number, the agent asks again from there.

JobsUp to 100 per swarm, 3 swarms open per workspace
ModelNight Shift's GPU, or the Pro model with Pro; your ladder isn't used
Replies1,024 tokens each by default, up to 8,192
PriceEach job is charged like any GPU request, when it finishes
Kept7 days

Each prompt has to carry everything the model needs, because the GPU can't see your files or the agent's conversation. Your agent writes every prompt as its own output, so pasting whole files into a swarm usually costs more than it saves. A swarm started while the GPU sleeps waits for it to wake, so a cold start no longer runs into the agent's tool timeout. If credit runs out, the jobs still waiting stop and nothing more is charged. A deploy in the middle of a swarm puts its running jobs back in the queue.

How routing works

Your ladder is your list of models from the top, the one you'd use for everything, down to the cheapest. You set it, a daily dollar budget for the top model, and the point at which Night Shift starts moving work down, on the Ladder page.

Night Shift grades each request by rules, not with another model: words like plan, migrate or refactor push it up, words like rename, summarise or format push it down, and attached tools add a little. While more than the start point of the budget is left, everything goes to the top model. Below it, the bar for staying on top rises as the budget runs down, and easier requests go to a model further down the list.

A conversation stays on the model it started on, because thinking blocks and prompt caches belong to one model. It only moves down when that model fails or is out. When a model rejects an option it doesn't support, Night Shift retries once without the newer options. A rate-limit error pauses that model until its retry time passes. If every model on the ladder fails, the request goes to the fallback model, Night Shift's GPU by default.

Night Shift's GPU

ModelQwen3 30B A3B Instruct, FP8
Context65,536 tokens; replies are capped at 8,192
FormatsOpenAI and Anthropic, streaming, tool calls
Price$0.10 in and $0.40 out per million tokens, paid from credit
In flightUp to 4 requests at once per workspace

The GPU scales to zero when nobody is using it. The first request after a quiet spell starts a worker, which can take several minutes; Night Shift holds the request open until it's answered. Requests after that answer in about a second. Very long agent sessions can outgrow the 64K context; add your own models for those.

Credits and burns

Night Shift is paid for by burning $SHIFT, its token on Robinhood Chain. The API and the MCP server work while a workspace has credit, and answer 402 when it has none. Requests to Night Shift's GPU model draw the credit down at its token prices; requests routed to your own provider keys don't.

$SHIFT is the Robinhood Chain token (address posted at launch). Each million burned adds $5 of credit. The Credits page builds the burn for you: one transaction from your wallet that sends the tokens to the dead address 0x000000000000000000000000000000000000dEaD. Night Shift reads it back from the chain before adding credit, so a burn only counts for the wallet linked to the workspace, and only once.

Upgrades

Credit also buys more compute for your agents, on the Upgrades page. Buying more of an upgrade you already have extends it.

UpgradeWhat it doesPrice
ProNight Shift's GPU requests go to gpt-oss-120b instead of Qwen3 30B, with a 128K context, replies up to 32K tokens and 16 requests at once. Tokens cost $0.25 in and $1 out per million.$3 a day
Warm GPUKeeps a worker running on Night Shift's GPU so requests never wait for a cold start. While anyone holds warm hours, the GPU stays on for everyone. Hours on sale are limited by what Night Shift's GPU account can pay for.$2 an hour
Warm Pro GPUThe same for the Pro GPU. Needs Pro.$4 an hour

Your own models

Add provider keys on the Keys page. Night Shift checks each key with its provider before saving it, encrypted, and only uses it to call that provider for your requests.

ProviderReaches
OpenRouterEvery model OpenRouter lists, in both formats. The simplest way to reach Grok, DeepSeek, Mistral, Qwen, Llama and the rest.
AnthropicClaude models directly, in both formats.
OpenAIGPT models directly, for OpenAI-format requests.
GoogleGemini models directly, for OpenAI-format requests.

When a model can be reached both directly and through OpenRouter, Night Shift uses the direct key.

Errors

Errors come back in the format of the request: an OpenAI error object for chat completions, an Anthropic error object for messages.

StatusMeaning
400The body isn't valid, or no model on your ladder can take this kind of request with your keys.
401The Night Shift key is missing, wrong or revoked.
402The workspace has no credit. Burn $SHIFT on the Credits page.
413The request body is over 25 MB.
429Every model that could answer is rate limited. Retry after the time in the provider's error.
5xxThe provider failed or couldn't be reached, after Night Shift tried the rest of the ladder.

Privacy

Night Shift keeps the model, token counts, cost, timing and status of each request, and never the prompt or the reply. Provider keys are encrypted at rest. Prompts for Night Shift's GPU model run on GPUs Night Shift rents, and Night Shift doesn't store them; prompts for your own models go to those providers under your own agreements with them.