Skip to content
Docs · AI

AI Gateway

One endpoint and one key for Claude, GPT, Gemini, Grok, Kimi and Llama: chat with history and streaming, an OpenAI-compatible URL, embeddings, and exact cost on every answer.

View as Markdown
All topics
On this page

Call POST /api/ai/chat with a model and messages. Every answer reports the model that answered, the provider, tokens, latency, why it stopped and exactly what it cost. Usage is paid from the plan's included AI credit first, where the plan has one, then the Balance.

harakumo-1 (alias auto) is a router: it tries vendor models in order within a tier (fast, balanced or deep) until one answers, and is billed at the price of the model that answered — Claude Sonnet leads the balanced tier. Name a model to pin it; llama-3.3-70b and kimi-k2.6 run on Harakumo's edge inference. A named model that cannot be reached right now may be answered by another (Kimi K3 by Kimi K2.6, for example): modelUsed and fellBackFrom always say which model answered, and it is billed at that model's price. Some models are off until an admin turns them on in Dashboard → AI Gateway; GET /api/ai/chat (SDK hk.ai.models()) marks each with enabled and its price.

promptAI Gatewayprovider · latency · cost per callclaude-*gpt-*gemini-*llama · kimi (edge)one endpoint, every model — harakumo-1 tries the next model when one cannot answer

Chat

curl
curl https://harakumo.com/api/ai/chat \
  -H "Authorization: Bearer $HARAKUMO_API_KEY" -H "content-type: application/json" \
  -d '{ "model": "harakumo-1", "system": "Answer in one sentence.",
        "messages": [{ "role": "user", "content": "What is an edge function?" }], "max_tokens": 300 }'
# → { "output", "modelUsed", "provider", "finishReason", "truncated", "tokensIn", "tokensOut", "cost", "latencyMs" }
Streaming (server-sent events)
const res = await fetch('https://harakumo.com/api/ai/chat', {
  method: 'POST',
  headers: { authorization: 'Bearer ' + process.env.HARAKUMO_API_KEY, 'content-type': 'application/json' },
  body: JSON.stringify({ model: 'claude-sonnet-5', stream: true, messages: [{ role: 'user', content: 'Tell me a story.' }] }),
});
// Events: "meta" (the model answering), "delta" ({ text }), then "done"
// ({ finishReason, truncated, tokensIn, tokensOut, cost }) or "error".
for await (const chunk of res.body) process.stdout.write(new TextDecoder().decode(chunk));
SDK
const answer = await hk.ai.chat({
  model: 'harakumo-1', system: 'Answer in one sentence.',
  messages: [{ role: 'user', content: 'What is an edge function?' }],
});
console.log(answer.output, answer.modelUsed, answer.cost);

for await (const ev of hk.ai.stream({ prompt: 'Tell me a story.' })) {
  if (ev.event === 'delta') process.stdout.write(ev.data.text);
  if (ev.event === 'done') console.log('\n', ev.data.cost);
}

OpenAI-compatible

Point any OpenAI client at https://harakumo.com/api/ai/v1 with a Harakumo key. The request and response shapes are the same (streaming included), model is any Harakumo model id, and each reply carries a harakumo field with the provider, cost and any fallback taken. Refusals come in OpenAI's error shape, { "error": { "message", "type", "code" } }, with the codes below.

OpenAI SDK
import OpenAI from 'openai';

const ai = new OpenAI({ apiKey: process.env.HARAKUMO_API_KEY, baseURL: 'https://harakumo.com/api/ai/v1' });
const r = await ai.chat.completions.create({ model: 'harakumo-1', messages: [{ role: 'user', content: 'Hi' }] });
console.log(r.choices[0].message.content);

Request fields

FieldMeaning
modelA model id, or harakumo-1 / auto (the default when left out). GET /api/ai/chat lists the models this workspace can use and its limits.
messages or promptA chat history (system, user, assistant roles), or one prompt string.
systemA system prompt.
max_tokensDefault 1,024, up to 8,192 (2,048 on the Llama edge models).
temperature0–2 (Claude models are capped at 1).
tierFor harakumo-1: fast, balanced or deep.
streamtrue for server-sent events.
projectIdAttribute the spend to a project (or send an X-Harakumo-Project header).

An answer that stopped at max_tokens has truncated: true and a note on how to get the rest.

Embeddings

POST /api/ai/embeddings { input, model } returns vectors from bge-small (384 dimensions), bge-base (768, the default) or bge-large (1,024), up to 100 inputs per call. They pair with Vector DB collections of the same size.

Price, limits and logs

  • Paid plans include AI usage each month ($10 on pro, $100 on enterprise). Past it, and for every request on a plan that includes none (hobby: pay as you go, and vendor models need the owner's verified email), the Balance pays, at each model's per-token price: the model maker's list price, plus up to 5% on third-party models to cover buying that usage (edge models never carry it). GET /api/ai/chat lists the exact price of every model.
  • Rate limit per workspace: 60 a minute on hobby, 600 a minute on pro, 6,000 a minute on enterprise; chat and embeddings have separate buckets. Over it: 429 ai_rate_limited with Retry-After.
  • Input up to 200,000 characters per request (413 over it).
  • GET /api/ai/requests lists recent calls (model, tokens, cost, speed — never the prompt or answer), kept 30 days, with this month's gateway spend by model and project. Billing → Balance has every AI charge this month by model, embeddings, memory and vectors included.
  • The AI Gateway and embedding endpoints are workspace-wide: use a workspace key, not a project key.

How a request is paid for

Before the model is called, Harakumo sets an estimate aside, sized from your input and max_tokens (for harakumo-1, at the price of the dearest model its tier may use). When the answer is back you are charged the exact cost of the model that answered, and the rest is released. The dashboard playground shows that estimate before you press Run.

  • A stream that ends early, because you disconnected or the model stopped, is charged only for what was delivered; a failure arrives as an error event with code stream_error.
  • Auto-reload, when the owner has turned it on, is tried before a 402 for credit. Requests made with an API key set it off only if the owner allowed that (Billing → Limits & alerts).
  • If a request is lost on Harakumo's side before it is settled, the estimate set aside for it is settled automatically within about half an hour: at what it had delivered, or given back.

When a request is refused

A refused or failed request costs nothing, with one exception: a model that spends its whole token budget without answering has used those tokens, and they are charged.

AnswerWhenCharged
402 ai_credits_exhaustedIncluded AI (on a plan that has it) and the Balance are both used up.Nothing
402 ai_credits_insufficientSome credit is left, but less than this request's estimate. Lower max_tokens, shorten the prompt or choose a cheaper model.Nothing
402 spend_limit_reachedThe estimate would take the month past the workspace's spending limit.Nothing
402 ai_request_too_largeThe estimate is more than one request may reserve: $2 on hobby, $10 on pro, $10 on enterprise.Nothing
402 email_unverifiedA vendor model on Hobby before the owner verified their email. harakumo-1 answers from edge inference meanwhile.Nothing
403 ai_gateway_off / ai_model_offThe gateway, or this model, is switched off in Dashboard → AI Gateway.Nothing
429 ai_rate_limitedThe per-minute limit. Wait Retry-After seconds.Nothing
400 or 413 ai_request_refusedThe model refused the request itself (a parameter, or its size). It is not tried on other models, which would refuse it too.Nothing
422 budget_exhaustedThe model used its whole max_tokens budget without an answer. Raise max_tokens.The tokens it used
502 ai_unavailable / ai_provider_errorNo model could answer right now. Try again, or use harakumo-1, which fails over between vendors.Nothing

Reference

SDK
const answer = await hk.ai.chat({ model: 'harakumo-1', prompt: 'Summarize edge functions.' });
console.log(answer.output, answer.modelUsed, answer.cost);

for await (const ev of hk.ai.stream({ messages: [{ role: 'user', content: 'Write a haiku' }] })) {
  if (ev.event === 'delta') process.stdout.write(ev.data.text);
}

const { models, limits } = await hk.ai.models();                  // what this workspace can use, with prices
const { requests, month } = await hk.ai.requests({ limit: 20 });  // recent calls and this month's spend

// Embeddings for Vector DB (bge-small 384 · bge-base 768 · bge-large 1024):
const { embeddings, dimensions } = await hk.ai.embed(['fresh mango', 'ripe banana'], 'bge-base');
// hk.ai.generate({ prompt }) is the older name of a one-prompt chat.