Lloyaldocs
Run & ship

Serve to many users

Load the model once and give every reader their own session — on a laptop, an office appliance or a GPU box, with no per-token bill.

On this page

The same harness that runs as a desktop app can serve a room full of people. A host loads the model once and runs one session per browser, each with its own context, its own agents and its own memory, over the one resident copy of the weights. You pay for the machine, not per token.

When to use it

You want Use
Several people using one model on one machine — an office appliance, a GPU box The host (this page): npm run serve
One person, offline, on their own machine The desktop app — Ship a desktop app
Every connection isolated in its own OS process, with its own copy of the model @lloyal-labs/relay — see Isolate by process

Quickstart

Terminal
npm run serve

serve builds the harness and starts the host. It downloads and verifies any model it still needs before it listens, so no browser ever waits on a download, and then prints that it is listening. In another shell, start the browser app:

Terminal
npm run dev:web:client

npm run dev:web starts both together, which is what you want while developing.

Configure the box

What describes the machine comes from its environment, never from harness.yml: one build serves many machines.

Variable Default What it sets
PORT 8787 The port the host listens on
HOST 127.0.0.1 The interface it binds. The host has no authentication of its own, so serving every interface (0.0.0.0) is an explicit choice — put it behind something that does authenticate.
MAX_SESSIONS 4 How many sessions may be resident at once. The rest wait their turn.
LLOYAL_CONTENT_ORIGIN — The one web origin allowed to upload files to the host
LLOYAL_DEV — 1 turns on the trace and the dev pane per session
A GPU box on the office network
HOST=0.0.0.0 PORT=8787 MAX_SESSIONS=8 npm run serve

Size MAX_SESSIONS

A session holds its context, and its own reranker and projector, for as long as its browser tab is open — even when the reader is idle. So the cap is how many open sessions the machine can hold, not how many can work at once. Four is a cautious default for one machine, not a measurement. Measure with the model, context length and vision settings you actually deploy, with the sessions idle, and set it on the box.

What a reader sees

Sessions are admitted first come, first served. A browser moves through connecting → queued → warming → ready, and a view reads that with useAvailability() (see The interface) to say "you're in the queue" rather than look broken.

  • A session is isolated. A harness that throws, or a service that fails to bind, ends that session alone, says why in the host's log, and the host keeps serving.
  • A finished session closes its socket, so a browser can tell "the session ended" from "the network dropped".
  • Visitors never change the harness. A setting saved from a browser applies to that session, in memory only. Nothing a visitor does relaunches the host, downloads a model or writes the configuration.

Point the browser at a host

The web app connects to ws://127.0.0.1:8787 unless told otherwise:

To point it at Set
A host fixed at build time VITE_WSS_URL=wss://your.host/… when building the web app
A host chosen at run time ?server=wss://your.host/… on the page's URL

Isolate by process

@lloyal-labs/relay is the other shape: one harness process per connection, each with its own residency, so the operating system is the boundary between readers. The price is memory — every connection loads its own copy of the model.

It ships as a bridge you mount in your own server (Express, Hono, Koa…): it forks your harness's bin for each connection and relays its frames over the socket. No scaffold command wires it up yet, so choose it when process isolation matters more than density, and expect to write the server around it.

↑
Search