# AI Gateway

Call `POST /api/ai/chat` with a model and messages. Every answer reports the model that answered, the provider, tokens, latency, why it stopped and exactly what it cost. Usage is paid from the plan's included AI credit first, where the plan has one, then the Balance.

`harakumo-1` (alias `auto`) is a router: it tries vendor models in order within a tier (fast, balanced or deep) until one answers, and is billed at the price of the model that answered — Claude Sonnet leads the balanced tier. Name a model to pin it; `llama-3.3-70b` and `kimi-k2.6` run on Harakumo's edge inference. A named model that cannot be reached right now may be answered by another (Kimi K3 by Kimi K2.6, for example): `modelUsed` and `fellBackFrom` always say which model answered, and it is billed at that model's price. Some models are off until an admin turns them on in Dashboard → AI Gateway; `GET /api/ai/chat` (SDK `hk.ai.models()`) marks each with `enabled` and its price.

## Chat

**curl**

```bash
curl https://harakumo.com/api/ai/chat \
  -H "Authorization: Bearer $HARAKUMO_API_KEY" -H "content-type: application/json" \
  -d '{ "model": "harakumo-1", "system": "Answer in one sentence.",
        "messages": [{ "role": "user", "content": "What is an edge function?" }], "max_tokens": 300 }'
# → { "output", "modelUsed", "provider", "finishReason", "truncated", "tokensIn", "tokensOut", "cost", "latencyMs" }
```

**Streaming (server-sent events)**

```js
const res = await fetch('https://harakumo.com/api/ai/chat', {
  method: 'POST',
  headers: { authorization: 'Bearer ' + process.env.HARAKUMO_API_KEY, 'content-type': 'application/json' },
  body: JSON.stringify({ model: 'claude-sonnet-5', stream: true, messages: [{ role: 'user', content: 'Tell me a story.' }] }),
});
// Events: "meta" (the model answering), "delta" ({ text }), then "done"
// ({ finishReason, truncated, tokensIn, tokensOut, cost }) or "error".
for await (const chunk of res.body) process.stdout.write(new TextDecoder().decode(chunk));
```

**SDK**

```js
const answer = await hk.ai.chat({
  model: 'harakumo-1', system: 'Answer in one sentence.',
  messages: [{ role: 'user', content: 'What is an edge function?' }],
});
console.log(answer.output, answer.modelUsed, answer.cost);

for await (const ev of hk.ai.stream({ prompt: 'Tell me a story.' })) {
  if (ev.event === 'delta') process.stdout.write(ev.data.text);
  if (ev.event === 'done') console.log('\n', ev.data.cost);
}
```

## OpenAI-compatible

Point any OpenAI client at `https://harakumo.com/api/ai/v1` with a Harakumo key. The request and response shapes are the same (streaming included), `model` is any Harakumo model id, and each reply carries a `harakumo` field with the provider, cost and any fallback taken. Refusals come in OpenAI's error shape, `{ "error": { "message", "type", "code" } }`, with the codes below.

**OpenAI SDK**

```js
import OpenAI from 'openai';

const ai = new OpenAI({ apiKey: process.env.HARAKUMO_API_KEY, baseURL: 'https://harakumo.com/api/ai/v1' });
const r = await ai.chat.completions.create({ model: 'harakumo-1', messages: [{ role: 'user', content: 'Hi' }] });
console.log(r.choices[0].message.content);
```

## Request fields

| Field | Meaning |
| --- | --- |
| model | A model id, or `harakumo-1` / `auto` (the default when left out). `GET /api/ai/chat` lists the models this workspace can use and its limits. |
| messages or prompt | A chat history (`system`, `user`, `assistant` roles), or one prompt string. |
| system | A system prompt. |
| max_tokens | Default 1,024, up to 8,192 (2,048 on the Llama edge models). |
| temperature | 0–2 (Claude models are capped at 1). |
| tier | For harakumo-1: fast, balanced or deep. |
| stream | true for server-sent events. |
| projectId | Attribute the spend to a project (or send an X-Harakumo-Project header). |

> An answer that stopped at max_tokens has `truncated: true` and a note on how to get the rest.

## Embeddings

`POST /api/ai/embeddings { input, model }` returns vectors from bge-small (384 dimensions), bge-base (768, the default) or bge-large (1,024), up to 100 inputs per call. They pair with Vector DB collections of the same size.

## Price, limits and logs

- Paid plans include AI usage each month ($10 on pro, $100 on enterprise). Past it, and for every request on a plan that includes none (hobby: pay as you go, and vendor models need the owner's verified email), the Balance pays, at each model's per-token price: the model maker's list price, plus up to 5% on third-party models to cover buying that usage (edge models never carry it). `GET /api/ai/chat` lists the exact price of every model.
- Rate limit per workspace: 60 a minute on hobby, 600 a minute on pro, 6,000 a minute on enterprise; chat and embeddings have separate buckets. Over it: 429 `ai_rate_limited` with Retry-After.
- Input up to 200,000 characters per request (413 over it).
- `GET /api/ai/requests` lists recent calls (model, tokens, cost, speed — never the prompt or answer), kept 30 days, with this month's gateway spend by model and project. Billing → Balance has every AI charge this month by model, embeddings, memory and vectors included.
- The AI Gateway and embedding endpoints are workspace-wide: use a workspace key, not a project key.

## How a request is paid for

Before the model is called, Harakumo sets an estimate aside, sized from your input and `max_tokens` (for `harakumo-1`, at the price of the dearest model its tier may use). When the answer is back you are charged the exact cost of the model that answered, and the rest is released. The dashboard playground shows that estimate before you press Run.

- A stream that ends early, because you disconnected or the model stopped, is charged only for what was delivered; a failure arrives as an `error` event with code `stream_error`.
- Auto-reload, when the owner has turned it on, is tried before a 402 for credit. Requests made with an API key set it off only if the owner allowed that (Billing → Limits & alerts).
- If a request is lost on Harakumo's side before it is settled, the estimate set aside for it is settled automatically within about half an hour: at what it had delivered, or given back.

## When a request is refused

A refused or failed request costs nothing, with one exception: a model that spends its whole token budget without answering has used those tokens, and they are charged.

| Answer | When | Charged |
| --- | --- | --- |
| 402 `ai_credits_exhausted` | Included AI (on a plan that has it) and the Balance are both used up. | Nothing |
| 402 `ai_credits_insufficient` | Some credit is left, but less than this request's estimate. Lower `max_tokens`, shorten the prompt or choose a cheaper model. | Nothing |
| 402 `spend_limit_reached` | The estimate would take the month past the workspace's spending limit. | Nothing |
| 402 `ai_request_too_large` | The estimate is more than one request may reserve: $2 on hobby, $10 on pro, $10 on enterprise. | Nothing |
| 402 `email_unverified` | A vendor model on Hobby before the owner verified their email. `harakumo-1` answers from edge inference meanwhile. | Nothing |
| 403 `ai_gateway_off` / `ai_model_off` | The gateway, or this model, is switched off in Dashboard → AI Gateway. | Nothing |
| 429 `ai_rate_limited` | The per-minute limit. Wait Retry-After seconds. | Nothing |
| 400 or 413 `ai_request_refused` | The model refused the request itself (a parameter, or its size). It is not tried on other models, which would refuse it too. | Nothing |
| 422 `budget_exhausted` | The model used its whole `max_tokens` budget without an answer. Raise `max_tokens`. | The tokens it used |
| 502 `ai_unavailable` / `ai_provider_error` | No model could answer right now. Try again, or use `harakumo-1`, which fails over between vendors. | Nothing |

## Reference

**SDK**

```js
const answer = await hk.ai.chat({ model: 'harakumo-1', prompt: 'Summarize edge functions.' });
console.log(answer.output, answer.modelUsed, answer.cost);

for await (const ev of hk.ai.stream({ messages: [{ role: 'user', content: 'Write a haiku' }] })) {
  if (ev.event === 'delta') process.stdout.write(ev.data.text);
}

const { models, limits } = await hk.ai.models();                  // what this workspace can use, with prices
const { requests, month } = await hk.ai.requests({ limit: 20 });  // recent calls and this month's spend

// Embeddings for Vector DB (bge-small 384 · bge-base 768 · bge-large 1024):
const { embeddings, dimensions } = await hk.ai.embed(['fresh mango', 'ripe banana'], 'bge-base');
// hk.ai.generate({ prompt }) is the older name of a one-prompt chat.
```

**CLI**

```bash
harakumo ai "Write a haiku about the cloud" --model harakumo-1
```

**REST**

```http
POST /api/ai/chat                   { "model", "messages" | "prompt", "system"?, "max_tokens"?, "temperature"?, "tier"?, "stream"?, "projectId"? }
GET  /api/ai/chat                   # the models and limits this workspace can use
POST /api/ai/v1/chat/completions    # OpenAI-compatible (base URL https://harakumo.com/api/ai/v1)
POST /api/ai/playground             { "model", "prompt" }          # the one-shot alias
POST /api/ai/embeddings             { "input": ["…"], "model"?: "bge-base", "projectId"? }   → { embeddings, dimensions }
GET  /api/ai/requests?limit=&projectId=    # recent calls and this month's spend
```
