Skip to content
Docs · Data

Memory

Three kinds of memory for AI apps — key-value with TTL, conversation transcripts, and semantic recall by meaning — each partitioned per end user with a scope.

View as Markdown
All topics
On this page

"kv" is a key-value store with TTL, to the second (carts, sessions, flags). "conversation" is an ordered transcript: append turns and read the last N back to build the next prompt. "semantic" embeds every value as you write it and recalls by meaning, with similarity scores.

One store serves every user of your app: pass a scope (a user or session id) on each call. Values, transcripts and searches in one scope never mix with another's.

PUT cart-42 (ttl 3600)cart-42sess-9ffeat-flagsgeo-cacheGET cart-42key-value with TTL, one of three kinds of memory

A multi-user chat

REST from your server
const base = 'https://harakumo.com/api/memory/' + storeId;
const headers = { authorization: 'Bearer ' + process.env.HARAKUMO_API_KEY, 'content-type': 'application/json' };

// Append this user's turn to their own transcript
await fetch(base + '/messages', { method: 'POST', headers, body: JSON.stringify({ role: 'user', content: text, scope: userId }) });

// Read their last 20 messages to build the next prompt
const { messages } = await (await fetch(base + '/messages?scope=' + encodeURIComponent(userId) + '&limit=20', { headers })).json();

// Forget one user on request
await fetch(base + '/messages?scope=' + encodeURIComponent(userId), { method: 'DELETE', headers });

Scopes

  • A scope is 1–64 characters: letters, digits and . _ - : @ + = (so a user id or an email address fits as-is).
  • Pass it in the query string, or in the body of a PUT or POST.
  • A semantic search only matches values written in the same scope; a search with no scope sees only unscoped values.
  • PATCH /api/memory/:id { action: "flush", scope } clears one scope; without scope it clears the store.

Limits

  • Values up to 1 MB; keys up to 512 characters.
  • A conversation keeps the 1,000 most recent messages per scope; one read returns up to 500.
  • Semantic values become searchable about 30 seconds after they are written — a search right after a write can legitimately find nothing, and the response says so.
  • Each plan caps items per store and bytes per workspace, for every kind of store (hobby: 10,000 items per store, 25 MB per workspace; pro: 100,000 items per store, 250 MB per workspace; enterprise: 1,000,000 items per store, 1 GB per workspace); a write past the cap gets a 402 naming it. Expired values are removed nightly.
  • Semantic writes and searches are metered on AI credits (embedding), and searches share the embeddings rate limit.
  • Each semantic write, delete and search is one call from the workspace's 600 per 5 minutes, shared with databases and vectors. Flushing a semantic scope removes up to 1,000 values per call and answers partial: true until it is done; run it again.
  • Keys: a key limited to the project works for everything here; a read-only key can read and search but not write.

Reference

SDK
const { store } = await hk.memory.create(project.id, { name: 'session', kind: 'kv' });
await hk.memory.set(store.id, 'cart', { items: 3 }, { ttlSeconds: 3600, scope: 'user-42' });

const { store: notes } = await hk.memory.create(project.id, { name: 'notes', kind: 'semantic' });
await hk.memory.set(notes.id, 'fact-1', 'Peaches cost $4 per pound.');
const { matches } = await hk.memory.search(notes.id, { query: 'what does fruit cost?', topK: 3 });
// Scoped reads, searches and transcripts: pass ?scope= over REST (below).