Self-Hosting a Jev-Compatible Decision API on Plain CPU

A CPU-only sidecar for Jev-compatible decisions: 281 to 671 ms per call, 2.1 GB RSS, 16 of 16 contract checks. What running it in production taught us.

Self-Hosting a Jev-Compatible Decision API on Plain CPU

On September 15, TypeSafe AI launched Jev, a model that returns typed decisions instead of text. Three days later, an open-source clone called Laya appeared on GitHub and cleared 27,000 stars in its first ten days. It sits at over 27,600 stars as of September 29.

We run a small HTTP sidecar in front of Laya called laya-service. It serves Jev-compatible decision calls from a plain CPU container, with no GPU anywhere in the path. The image is public on Docker Hub, the repo stays private, and every number below comes from our own deployment this week, not from a press kit.

Typed decisions in one forward pass

Laya answers three kinds of questions over any text state: choice (pick from options), score (rate on a scale), and noul (yes or no). It runs one forward pass per question, 33 ms on a T4 according to its README, or 7.2 ms per question when batched. No tokens get generated, so there is nothing to parse and no prose to hallucinate.

That shape matters for software. A support triage, a router, a gating check, a priority call: these want a typed answer with a confidence number, not a paragraph that has to be interrogated afterward.

Why we put a sidecar in front

laya-service is stateless and policy-free. The service answers questions; the caller owns every action decision. Nothing in the sidecar auto-acts, logs, or escalates on its own.

Authentication is bearer token with fail-closed behavior: a missing or wrong token gets a 401, never a degraded best effort. Confidence thresholds stay advisory too. We use 0.85 and above as the auto-act suggestion in our own consumers, and each consumer decides what to do with that signal. The sidecar itself has no opinions.

This split keeps the engine swappable. If a better checkpoint ships next month, the contract does not move.

The numbers we measured

We validate the full API contract against the live deployment on every change. The last full pass on September 28 checked 16 endpoint behaviors, all green.

Latency, measured per call including network:

  • Modal CPU lane, 1 vCPU and 4 GB: 564 to 671 ms wall time
  • VPS Docker lane, shared box: 281 to 466 ms
  • Memory: about 2.1 GB RSS per instance

TypeSafe reports 70 to 500 ms for the hosted Jev API, so we trade some latency for a container we control. For decision traffic that runs in the background of a workflow, that trade is easy to accept.

Response shapes are worth knowing before you integrate. A score answer returns a continuous float plus a legend mapping indexes to labels. noul returns a probability between 0 and 1, not a boolean. Every answer carries probabilities and a confidence value, which is what makes advisory gating possible at the caller.

Gotchas that cost us time

Long state collapses. Past roughly 1,024 tokens, quality degrades, and at 2,500 tokens a single call took about 17 seconds in our tests. Keep state short. Hydrate the model with summaries, not raw logs.

Fewer labels win. choice accepts 2 to 20 options, but a 20-way zero-shot question in Indonesian measured 0.51 accuracy in our evaluation. Three to five labels per question is the reliable range. A 5-label support triage reached 0.97 on the needs_human class.

Validation is strict and loud. Out-of-bounds choices return a 400 with a Jev-style error body. Missing state is a 422. Unknown model names are a 400. This is friction in development and a feature in production.

First start downloads weights. The VPS lane pulls about 644 MB of checkpoints on first boot, then starts instantly from a cached volume. The Modal lane bakes weights into the image, so the first request is already warm.

Run it yourself

docker run -d --name laya \
  -p 8080:8080 \
  -v laya-hf-cache:/data/hf \
  --memory 4g \
  codecoradev/laya-service:latest

That starts the sidecar on port 8080 and creates a named volume for the Hugging Face cache. The first start downloads about 644 MB of checkpoints; every start after that is instant. Readiness is a GET to /healthz, and status: "ok" means the model is loaded.

The image is multi-arch (amd64 and arm64) with latest, develop, and dated tags. Environment knobs: LAYA_CHECKPOINT (default multilingual), LAYA_DEVICE (default cpu), and AUTO_THRESHOLD (default 0.85). The Docker Hub page carries the full API reference, the contract table, and a worked request example. Upstream Laya is Apache-2.0, code and checkpoints alike.

If you prefer managed infra, a 1 vCPU Modal container scales to zero between calls and costs about $0.24 per hour while warm. The VPS lane is a single compose file with the same memory limit and a healthcheck.

What we're watching next

Three things. Confidence-gated auto-act moving into consumers now that the contract is stable. Vertical fine-tuning on top of the RLCD training recipe, which is where decision models should get weird and specialized. And whether the ecosystem converges on Jev-compatible calls as a shared interface: three separate awesome-jev lists appeared within two weeks of launch, which reads like coordination nobody planned.

The product-side summary: typed decisions are cheap enough to run anywhere now. We published the container so you can check that claim yourself.