Call POST /api/ai/chat with a model and messages. Every answer reports the model that answered, the provider, tokens, latency, why it stopped and exactly what it cost. Usage is paid from the plan's included AI credit first, where the plan has one, then the Balance.
harakumo-1 (alias auto) is a router: it tries vendor models in order within a tier (fast, balanced or deep) until one answers, and is billed at the price of the model that answered — Claude Sonnet leads the balanced tier. Name a model to pin it; llama-3.3-70b and kimi-k2.6 run on Harakumo's edge inference. A named model that cannot be reached right now may be answered by another (Kimi K3 by Kimi K2.6, for example): modelUsed and fellBackFrom always say which model answered, and it is billed at that model's price. Some models are off until an admin turns them on in Dashboard → AI Gateway; GET /api/ai/chat (SDK hk.ai.models()) marks each with enabled and its price.
Chat
curl https://harakumo.com/api/ai/chat \
-H "Authorization: Bearer $HARAKUMO_API_KEY" -H "content-type: application/json" \
-d '{ "model": "harakumo-1", "system": "Answer in one sentence.",
"messages": [{ "role": "user", "content": "What is an edge function?" }], "max_tokens": 300 }'
# → { "output", "modelUsed", "provider", "finishReason", "truncated", "tokensIn", "tokensOut", "cost", "latencyMs" }const res = await fetch('https://harakumo.com/api/ai/chat', {
method: 'POST',
headers: { authorization: 'Bearer ' + process.env.HARAKUMO_API_KEY, 'content-type': 'application/json' },
body: JSON.stringify({ model: 'claude-sonnet-5', stream: true, messages: [{ role: 'user', content: 'Tell me a story.' }] }),
});
// Events: "meta" (the model answering), "delta" ({ text }), then "done"
// ({ finishReason, truncated, tokensIn, tokensOut, cost }) or "error".
for await (const chunk of res.body) process.stdout.write(new TextDecoder().decode(chunk));const answer = await hk.ai.chat({
model: 'harakumo-1', system: 'Answer in one sentence.',
messages: [{ role: 'user', content: 'What is an edge function?' }],
});
console.log(answer.output, answer.modelUsed, answer.cost);
for await (const ev of hk.ai.stream({ prompt: 'Tell me a story.' })) {
if (ev.event === 'delta') process.stdout.write(ev.data.text);
if (ev.event === 'done') console.log('\n', ev.data.cost);
}OpenAI-compatible
Point any OpenAI client at https://harakumo.com/api/ai/v1 with a Harakumo key. The request and response shapes are the same (streaming included), model is any Harakumo model id, and each reply carries a harakumo field with the provider, cost and any fallback taken. Refusals come in OpenAI's error shape, { "error": { "message", "type", "code" } }, with the codes below.
import OpenAI from 'openai';
const ai = new OpenAI({ apiKey: process.env.HARAKUMO_API_KEY, baseURL: 'https://harakumo.com/api/ai/v1' });
const r = await ai.chat.completions.create({ model: 'harakumo-1', messages: [{ role: 'user', content: 'Hi' }] });
console.log(r.choices[0].message.content);Request fields
| Field | Meaning |
|---|---|
| model | A model id, or harakumo-1 / auto (the default when left out). GET /api/ai/chat lists the models this workspace can use and its limits. |
| messages or prompt | A chat history (system, user, assistant roles), or one prompt string. |
| system | A system prompt. |
| max_tokens | Default 1,024, up to 8,192 (2,048 on the Llama edge models). |
| temperature | 0–2 (Claude models are capped at 1). |
| tier | For harakumo-1: fast, balanced or deep. |
| stream | true for server-sent events. |
| projectId | Attribute the spend to a project (or send an X-Harakumo-Project header). |
An answer that stopped at max_tokens has truncated: true and a note on how to get the rest.
Embeddings
POST /api/ai/embeddings { input, model } returns vectors from bge-small (384 dimensions), bge-base (768, the default) or bge-large (1,024), up to 100 inputs per call. They pair with Vector DB collections of the same size.
Price, limits and logs
- Paid plans include AI usage each month ($10 on pro, $100 on enterprise). Past it, and for every request on a plan that includes none (hobby: pay as you go, and vendor models need the owner's verified email), the Balance pays, at each model's per-token price: the model maker's list price, plus up to 5% on third-party models to cover buying that usage (edge models never carry it).
GET /api/ai/chatlists the exact price of every model. - Rate limit per workspace: 60 a minute on hobby, 600 a minute on pro, 6,000 a minute on enterprise; chat and embeddings have separate buckets. Over it: 429
ai_rate_limitedwith Retry-After. - Input up to 200,000 characters per request (413 over it).
GET /api/ai/requestslists recent calls (model, tokens, cost, speed — never the prompt or answer), kept 30 days, with this month's gateway spend by model and project. Billing → Balance has every AI charge this month by model, embeddings, memory and vectors included.- The AI Gateway and embedding endpoints are workspace-wide: use a workspace key, not a project key.
How a request is paid for
Before the model is called, Harakumo sets an estimate aside, sized from your input and max_tokens (for harakumo-1, at the price of the dearest model its tier may use). When the answer is back you are charged the exact cost of the model that answered, and the rest is released. The dashboard playground shows that estimate before you press Run.
- A stream that ends early, because you disconnected or the model stopped, is charged only for what was delivered; a failure arrives as an
errorevent with codestream_error. - Auto-reload, when the owner has turned it on, is tried before a 402 for credit. Requests made with an API key set it off only if the owner allowed that (Billing → Limits & alerts).
- If a request is lost on Harakumo's side before it is settled, the estimate set aside for it is settled automatically within about half an hour: at what it had delivered, or given back.
When a request is refused
A refused or failed request costs nothing, with one exception: a model that spends its whole token budget without answering has used those tokens, and they are charged.
| Answer | When | Charged |
|---|---|---|
402 ai_credits_exhausted | Included AI (on a plan that has it) and the Balance are both used up. | Nothing |
402 ai_credits_insufficient | Some credit is left, but less than this request's estimate. Lower max_tokens, shorten the prompt or choose a cheaper model. | Nothing |
402 spend_limit_reached | The estimate would take the month past the workspace's spending limit. | Nothing |
402 ai_request_too_large | The estimate is more than one request may reserve: $2 on hobby, $10 on pro, $10 on enterprise. | Nothing |
402 email_unverified | A vendor model on Hobby before the owner verified their email. harakumo-1 answers from edge inference meanwhile. | Nothing |
403 ai_gateway_off / ai_model_off | The gateway, or this model, is switched off in Dashboard → AI Gateway. | Nothing |
429 ai_rate_limited | The per-minute limit. Wait Retry-After seconds. | Nothing |
400 or 413 ai_request_refused | The model refused the request itself (a parameter, or its size). It is not tried on other models, which would refuse it too. | Nothing |
422 budget_exhausted | The model used its whole max_tokens budget without an answer. Raise max_tokens. | The tokens it used |
502 ai_unavailable / ai_provider_error | No model could answer right now. Try again, or use harakumo-1, which fails over between vendors. | Nothing |
Reference
const answer = await hk.ai.chat({ model: 'harakumo-1', prompt: 'Summarize edge functions.' });
console.log(answer.output, answer.modelUsed, answer.cost);
for await (const ev of hk.ai.stream({ messages: [{ role: 'user', content: 'Write a haiku' }] })) {
if (ev.event === 'delta') process.stdout.write(ev.data.text);
}
const { models, limits } = await hk.ai.models(); // what this workspace can use, with prices
const { requests, month } = await hk.ai.requests({ limit: 20 }); // recent calls and this month's spend
// Embeddings for Vector DB (bge-small 384 · bge-base 768 · bge-large 1024):
const { embeddings, dimensions } = await hk.ai.embed(['fresh mango', 'ripe banana'], 'bge-base');
// hk.ai.generate({ prompt }) is the older name of a one-prompt chat.harakumo ai "Write a haiku about the cloud" --model harakumo-1POST /api/ai/chat { "model", "messages" | "prompt", "system"?, "max_tokens"?, "temperature"?, "tier"?, "stream"?, "projectId"? }
GET /api/ai/chat # the models and limits this workspace can use
POST /api/ai/v1/chat/completions # OpenAI-compatible (base URL https://harakumo.com/api/ai/v1)
POST /api/ai/playground { "model", "prompt" } # the one-shot alias
POST /api/ai/embeddings { "input": ["…"], "model"?: "bge-base", "projectId"? } → { embeddings, dimensions }
GET /api/ai/requests?limit=&projectId= # recent calls and this month's spend