Self-Host a Memory Server for LLM Agents with Uteke
If you searched for a self-hosted memory server for LLM agents
A self-hosted memory server for LLM agents is a service you run on your own machine or VPS that stores agent memories locally and serves them back through an API or MCP. Uteke is one: a single Rust binary with SQLite storage and local embeddings. Docker up, MCP wired in, and your agent recalls in about 45 ms with zero LLM tokens per query.
This tutorial walks through the full setup: install, first memories, the MCP hookup for coding agents, the Docker path for a shared server, and the checks that prove the setup works. Every command here comes from the current Uteke README, verified September 29, 2026 against release v0.19.0.
Why self-host agent memory at all
Agent memory is sensitive by nature. It accumulates your project decisions, internal paths, client names, half-finished thoughts. Handing that to a hosted memory API means trusting another party with the one dataset you never meant to publish. Self-hosting removes that question entirely: the database file sits on your disk, queries never leave the machine, and there is no per-call bill ticking up in the background.
There is also the availability angle. Uteke runs fully offline on CPU. No API key, no region restrictions, no rate limit imposed from outside. If your network drops, your agent keeps remembering.
Install Uteke
Pick the path that matches your setup:
# macOS and Linux (shell installer)
curl -sSL codecora.dev/uteke/install | sh
# Or run as a container
docker run -d -p 127.0.0.1:8767:8767 -v uteke-data:/data ghcr.io/codecoradev/uteke:latestWindows ships a PowerShell installer in the same repository.
First run downloads the embedding model once, around 200 MB. After that everything is local.
New to Uteke, or setting it up from inside an agent? Run uteke onboard. It detects your environment, asks which agent you use, and wires the configuration for you.
Store and recall your first memories
uteke remember "Deploy v2.1 to staging at 3pm, rollback owner: Aji"
uteke recall "when do we deploy?"Recall works by meaning and by keyword at the same time. The hybrid architecture behind this (SQLite for exact matches, a local embedding index for semantic matches, fused at query time) is covered in detail in our hybrid memory architecture post.
Connect your coding agent through MCP
One line in .mcp.json gives Claude Code, Cursor, Copilot, or any MCP-compatible agent access to the same memory store:
{ "mcpServers": { "uteke": { "command": "uteke-mcp" } } }From there the agent can save and search memories as tools. The server-side details, including what the MCP surface exposes, are documented in the Uteke MCP memory server post.
Run it as a shared memory server with Docker
A laptop install serves one person. A team usually wants the memory store on a box everyone can reach. That is the Docker path:
docker run -d -p 127.0.0.1:8767:8767 -v uteke-data:/data ghcr.io/codecoradev/uteke:latestNotes from running this in production:
- The port binding above keeps the API on localhost only. Put it behind your reverse proxy with auth if agents connect from other machines.
- The
uteke-datavolume is the entire memory store. Back it up like you would any database file; it is a single SQLite-backed volume, so a file copy is a backup. - Everything is CPU-only. A small VPS is enough; warm queries land around 45 ms.
Longer-term operational lessons (monitoring, upgrades, failure modes we hit) live in the Docker production write-up.
Verify the setup works
Three checks, in order:
curl http://localhost:8767/api/healthfrom the host (or your proxied URL) returns a healthy response.- Write a memory through the agent, then
uteke recallit from the CLI. Cross-path recall proves both the API and the local store see the same data. - Ask the agent something that depends on a memory from last week. If it answers from memory instead of guessing, the loop is closed.
How it compares with the hosted memory options
Memory servers in the hosted camp (mem0, Zep, Letta among them) trade convenience for a dependency: another account, another endpoint, another data flow leaving your perimeter. Uteke takes the opposite position: the full comparison covers the tradeoffs, and our benchmark post has the numbers, including 98.4% recall@5 on LongMemEval-S with zero LLM calls in the retrieval path.
One honest caveat: a self-hosted server means you own uptime. There is no vendor to page when the disk fills. In practice the operational surface is small (one binary, one volume), but it is yours.
Where to go next
If you want the reasoning behind local-first memory rather than the setup steps, start with getting started with Uteke or why a local LLM is not dumb, just amnesic. The landing page at uteke.app has the always-current install instructions, and the source is at github.com/codecoradev/uteke under Apache 2.0.
Frequently asked questions
Does a self-hosted memory server need a GPU?
No. Uteke runs embeddings on CPU. Warm queries take roughly 45 ms on ordinary hardware, and there are zero LLM tokens spent per query.
Can multiple agents share one Uteke instance?
Yes. The Docker deployment is designed for it: agents connect through the HTTP API or MCP and share the same memory volume. Use rooms to partition memories per project or per agent.
What happens to my memories if I go offline?
Nothing. Storage and retrieval are fully local. An internet connection is only needed for downloading the embedding model the first time.
How do I back up the memory server?
Copy the data volume. It is SQLite-backed, so a consistent file copy is a complete backup. Schedule it like any other database backup.