Self-Host a Memory Server for LLM Agents with Uteke

Self-Host a Memory Server for LLM Agents with Uteke

If you searched for a self-hosted memory server for LLM agents

A self-hosted memory server for LLM agents is a service you run on your own machine or VPS that stores agent memories locally and serves them back through an API or MCP. Uteke is one: a single Rust binary with SQLite storage and local embeddings. Docker up, MCP wired in, and your agent recalls in about 45 ms with zero LLM tokens per query.

This tutorial walks through the full setup: install, first memories, the MCP hookup for coding agents, the Docker path for a shared server, and the checks that prove the setup works. Every command here comes from the current Uteke README, verified September 29, 2026 against release v0.19.0.

Why self-host agent memory at all

Agent memory is sensitive by nature. It accumulates your project decisions, internal paths, client names, half-finished thoughts. Handing that to a hosted memory API means trusting another party with the one dataset you never meant to publish. Self-hosting removes that question entirely: the database file sits on your disk, queries never leave the machine, and there is no per-call bill ticking up in the background.

There is also the availability angle. Uteke runs fully offline on CPU. No API key, no region restrictions, no rate limit imposed from outside. If your network drops, your agent keeps remembering.

Install Uteke

Pick the path that matches your setup:

# macOS and Linux (shell installer)
curl -sSL codecora.dev/uteke/install | sh

# Or run as a container
docker run -d -p 127.0.0.1:8767:8767 -v uteke-data:/data ghcr.io/codecoradev/uteke:latest

Windows ships a PowerShell installer in the same repository.

First run downloads the embedding model once, around 200 MB. After that everything is local.

New to Uteke, or setting it up from inside an agent? Run uteke onboard. It detects your environment, asks which agent you use, and wires the configuration for you.

Store and recall your first memories

uteke remember "Deploy v2.1 to staging at 3pm, rollback owner: Aji"
uteke recall "when do we deploy?"

Recall works by meaning and by keyword at the same time. The hybrid architecture behind this (SQLite for exact matches, a local embedding index for semantic matches, fused at query time) is covered in detail in our hybrid memory architecture post.

Connect your coding agent through MCP

One line in .mcp.json gives Claude Code, Cursor, Copilot, or any MCP-compatible agent access to the same memory store:

{ "mcpServers": { "uteke": { "command": "uteke-mcp" } } }

From there the agent can save and search memories as tools. The server-side details, including what the MCP surface exposes, are documented in the Uteke MCP memory server post.

Run it as a shared memory server with Docker

A laptop install serves one person. A team usually wants the memory store on a box everyone can reach. That is the Docker path:

docker run -d -p 127.0.0.1:8767:8767 -v uteke-data:/data ghcr.io/codecoradev/uteke:latest

Notes from running this in production:

  • The port binding above keeps the API on localhost only. Put it behind your reverse proxy with auth if agents connect from other machines.
  • The uteke-data volume is the entire memory store. Back it up like you would any database file; it is a single SQLite-backed volume, so a file copy is a backup.
  • Everything is CPU-only. A small VPS is enough; warm queries land around 45 ms.

Longer-term operational lessons (monitoring, upgrades, failure modes we hit) live in the Docker production write-up.

Verify the setup works

Three checks, in order:

  1. curl http://localhost:8767/api/health from the host (or your proxied URL) returns a healthy response.
  2. Write a memory through the agent, then uteke recall it from the CLI. Cross-path recall proves both the API and the local store see the same data.
  3. Ask the agent something that depends on a memory from last week. If it answers from memory instead of guessing, the loop is closed.

How it compares with the hosted memory options

Memory servers in the hosted camp (mem0, Zep, Letta among them) trade convenience for a dependency: another account, another endpoint, another data flow leaving your perimeter. Uteke takes the opposite position: the full comparison covers the tradeoffs, and our benchmark post has the numbers, including 98.4% recall@5 on LongMemEval-S with zero LLM calls in the retrieval path.

One honest caveat: a self-hosted server means you own uptime. There is no vendor to page when the disk fills. In practice the operational surface is small (one binary, one volume), but it is yours.

Where to go next

If you want the reasoning behind local-first memory rather than the setup steps, start with getting started with Uteke or why a local LLM is not dumb, just amnesic. The landing page at uteke.app has the always-current install instructions, and the source is at github.com/codecoradev/uteke under Apache 2.0.

Frequently asked questions

Does a self-hosted memory server need a GPU?

No. Uteke runs embeddings on CPU. Warm queries take roughly 45 ms on ordinary hardware, and there are zero LLM tokens spent per query.

Can multiple agents share one Uteke instance?

Yes. The Docker deployment is designed for it: agents connect through the HTTP API or MCP and share the same memory volume. Use rooms to partition memories per project or per agent.

What happens to my memories if I go offline?

Nothing. Storage and retrieval are fully local. An internet connection is only needed for downloading the embedding model the first time.

How do I back up the memory server?

Copy the data volume. It is SQLite-backed, so a consistent file copy is a complete backup. Schedule it like any other database backup.