Two agents, one limit, one refused.
All posts
3 min read
by

Introducing Memory: remember() and recall() for AI agents

Long-term memory for AI agents in two function calls, remember() and recall(). No vector database, no embedding pipeline, no index tuning.

memoryannouncementai

Editor's note, 3 October 2026. This post covers an earlier namespace. It still runs on the same API. Anlyon's product today is shared limits and receipts for AI agent actions.

Every AI agent eventually needs memory. A support agent should remember what the customer said last week. A coding assistant should remember your project's conventions. A sales copilot should remember every account's quirks.

Today, building that means standing up a vector database, choosing an embedding model, writing a chunking pipeline, tuning hybrid search, and re-embedding everything when a better model ships. That's infrastructure work, not product work.

Memory went live on Anlyon in July 2026. It's two function calls:

import { Client } from '@anlyonhq/sdk';

const anlyon = new Client({ apiKey: process.env.ANLYON_API_KEY });

// Store knowledge
await anlyon.memory.remember({
  content: 'Acme Corp prefers invoices by email, net-30 terms.',
});

// Retrieve it by meaning, not keywords
const { results } = await anlyon.memory.recall({
  query: 'how does Acme want to be billed?',
});

Not a vector database

Memory is intentionally not a vector database. You never pick an embedding model, configure an index, or think about dimensions. Anlyon owns those decisions and upgrades them over time, so your code stays two function calls.

Under the hood, every recall() runs hybrid retrieval: semantic vector search and keyword full-text search, fused with reciprocal rank fusion. You can force mode: "semantic" or mode: "keyword" when you know what you want, and enable LLM reranking for higher precision.

What ships today

  • remember / recall / forget: the core loop, with metadata filters, TTLs, and idempotency keys
  • Collections: namespace memories per agent, per user, or per tenant
  • Document ingest: POST a document (text, URL, or file) and it's chunked, embedded, and recallable asynchronously
  • summarize and consolidate: LLM-powered compression: merge duplicates and distill long histories into concise memories
  • Batch operations: store 100 memories with a single embedding round-trip
  • Model migration: re-embed a collection into a newer embedding model with one API call, no downtime

Everywhere you build

Memory speaks the tools you already use:

  • TypeScript: npm install @anlyonhq/sdk
  • Python: pip install anlyon (sync and async clients)
  • MCP: npx @anlyonhq/mcp-server gives Claude, Cursor, and any MCP client memory tools out of the box
  • Frameworks: adapters for the Vercel AI SDK, LangChain, and LlamaIndex

Pricing

There is no public token charge. While Early Beta Access is active (see the pricing page for the current terms) it is free, with no credit card, and includes 2,000 stored memories, 50 MiB of memory storage, and a small monthly processing allowance for embedding and summarizing. When the allowance is used up, processing pauses until the next UTC month and your stored memories stay available. Enough to run a real agent.

Where Memory fits

Added 23 September 2026. Memory is the state layer, not the headline. Anlyon is production execution for AI agents: your agent names an action, and Anlyon holds the credential and makes the call to Stripe, GitHub or your own API itself. Memory is what the agent knows when it decides what to ask for. Start with your agent's first side effect if you have not seen that half.

Start building →

Free tier, no credit card. One command if you use Claude or Cursor.

$ claude mcp add anlyon -- npx -y @anlyonhq/mcp-server