"kv" is a key-value store with TTL, to the second (carts, sessions, flags). "conversation" is an ordered transcript: append turns and read the last N back to build the next prompt. "semantic" embeds every value as you write it and recalls by meaning, with similarity scores.
One store serves every user of your app: pass a scope (a user or session id) on each call. Values, transcripts and searches in one scope never mix with another's.
A multi-user chat
REST from your server
const base = 'https://harakumo.com/api/memory/' + storeId;
const headers = { authorization: 'Bearer ' + process.env.HARAKUMO_API_KEY, 'content-type': 'application/json' };
// Append this user's turn to their own transcript
await fetch(base + '/messages', { method: 'POST', headers, body: JSON.stringify({ role: 'user', content: text, scope: userId }) });
// Read their last 20 messages to build the next prompt
const { messages } = await (await fetch(base + '/messages?scope=' + encodeURIComponent(userId) + '&limit=20', { headers })).json();
// Forget one user on request
await fetch(base + '/messages?scope=' + encodeURIComponent(userId), { method: 'DELETE', headers });Scopes
- A scope is 1–64 characters: letters, digits and . _ - : @ + = (so a user id or an email address fits as-is).
- Pass it in the query string, or in the body of a PUT or POST.
- A semantic search only matches values written in the same scope; a search with no scope sees only unscoped values.
PATCH /api/memory/:id { action: "flush", scope }clears one scope; without scope it clears the store.
Limits
- Values up to 1 MB; keys up to 512 characters.
- A conversation keeps the 1,000 most recent messages per scope; one read returns up to 500.
- Semantic values become searchable about 30 seconds after they are written — a search right after a write can legitimately find nothing, and the response says so.
- Each plan caps items per store and bytes per workspace, for every kind of store (hobby: 10,000 items per store, 25 MB per workspace; pro: 100,000 items per store, 250 MB per workspace; enterprise: 1,000,000 items per store, 1 GB per workspace); a write past the cap gets a 402 naming it. Expired values are removed nightly.
- Semantic writes and searches are metered on AI credits (embedding), and searches share the embeddings rate limit.
- Each semantic write, delete and search is one call from the workspace's 600 per 5 minutes, shared with databases and vectors. Flushing a semantic scope removes up to 1,000 values per call and answers
partial: trueuntil it is done; run it again. - Keys: a key limited to the project works for everything here; a read-only key can read and search but not write.
Reference
SDK
const { store } = await hk.memory.create(project.id, { name: 'session', kind: 'kv' });
await hk.memory.set(store.id, 'cart', { items: 3 }, { ttlSeconds: 3600, scope: 'user-42' });
const { store: notes } = await hk.memory.create(project.id, { name: 'notes', kind: 'semantic' });
await hk.memory.set(notes.id, 'fact-1', 'Peaches cost $4 per pound.');
const { matches } = await hk.memory.search(notes.id, { query: 'what does fruit cost?', topK: 3 });
// Scoped reads, searches and transcripts: pass ?scope= over REST (below).REST
POST /api/projects/:id/memory { "name": "session", "kind": "kv" | "conversation" | "semantic" }
PUT /api/memory/:id/keys/:key { "value": {...}, "ttlSeconds"?, "scope"? } # semantic values are embedded
GET /api/memory/:id/keys/:key?scope=
DELETE /api/memory/:id/keys/:key?scope=
POST /api/memory/:id/messages { "role": "user", "content": "…", "scope"? } # conversation
GET /api/memory/:id/messages?scope=&limit=50 # oldest → newest
DELETE /api/memory/:id/messages?scope= # forget one scope
POST /api/memory/:id/search { "query": "…", "topK"?: 5, "scope"? } # semantic
PATCH /api/memory/:id { "action": "flush", "scope"? } → { removed }