Reference · Smith
The Smith, for developers
What the Smith shows with developer agents hired: the model servers it keeps and the rules it keeps them by, new servers and models to try on your NPU, the audits that check the models still answer right, its files, commands and local API.
- Article
- 1403
- Applies to
- Smith 0.8.3
- Last reviewed
- For
- For developers
What changes with developer agents hired#
With one of Castellan's developer agents hired, the Smith's page names things as they are: The model servers rather than The local AI, their addresses, models and programs, and three more parts: Orphans, New for your NPU and Audits. Its settings and the tour use the same words. What it does for everyone is unchanged.
The model servers#
| Accelerator | Model server |
|---|---|
| A Snapdragon's NPU | GenieX |
| An Intel Core Ultra's NPU | OpenVINO Model Server |
| An AMD Ryzen AI 300's NPU | FastFlowLM |
| A graphics card, or the processor | llama.cpp's server |
Each server serves what Set up ticked: chat, vision, embeddings, and a reranker (Qwen3-Reranker-0.6B, through llama.cpp's --reranking) that Reeve's search uses from Reeve 0.15.0.
The rules it keeps them by#
Every minute (its round, paused while it's off duty), the Smith looks at every model server the manor runs:
| When | It |
|---|---|
| Nobody has held or waited on an accelerator for its idle time (10 minutes each, in Settings) | stops its servers. The next request from any agent starts them again. |
| A game is using a graphics card, and nobody has used the card for 2 minutes | stops the card's servers. |
| Castellan's Use the graphics card for models when there's an NPU is off | stops a card's servers as soon as nobody uses the card, and doesn't restart them. |
| A server hasn't answered for 3 minutes while nobody uses its accelerator | restarts it. A probe that times out means busy: GenieX answers nothing while it loads or answers. |
| GenieX holds 9 GB of working set while nobody uses the NPU | restarts it (it keeps about 340 MB for every model it loads, until it exits). |
| A model server on the manor's ports that no configured server is (an orphan) | stops it after 5 minutes, and logs its path and command line. Only one that's the manor's own: its program in the manor's servers folders, or one the manor started. Never by its port alone. |
Each stop or restart takes the accelerator's turn as background work that joins only an empty line, so it never takes anyone's turn. What it did is in What it did on its page, and in servers\logs\reaper.log in the accelerators folder.
New for your NPU#
Once a day the Smith looks for newer servers and models for your NPU that the manor supports. Each is listed under New for your NPU, with Try it and Not now. Look now looks at once; Stop looking stops the daily look.
Try it installs the new one beside what works, and times both on the same questions:
- Faster or as fast, and it answers right: it's used at once, and the Smith says how they compare ("GenieX 0.9.0 answers 18% faster than GenieX 0.8.0"). Remove it puts the old one back.
- Slower: what you had stays in use, and both are shown side by side, with Use it anyway or Keep what I have. A newer server that installs over the old one is taken out again at once, so the old one keeps running.
- It fails: what you had stays in use, and the try's lines say why.
How your NPU has performed lists every setup your NPU has run on this PC: how long it takes to write a short paragraph, tokens a second, how long the first answer takes (loading the model), and how long a question about a picture takes. Time it now times today's setup: after a driver update, say.
The NPU setup setting (under Advanced): The one I had before the last try goes back to it, and asks you before any switch from then on.
Audits: do the models still answer right?#
A driver or model update can leave a model answering, but wrong. The Smith checks for that, in two parts:
- The fingerprint, every hour, with no model: each accelerator's driver, model server and version, and models. A change is logged, and makes the accelerator's audit due.
- An audit asks each accelerator a few small questions whose answers code checks exactly, and compares the results with its last good audit on that toolchain (its baseline):
| Check | What it asks |
|---|---|
| Chat | Golden questions with one exact answer each: a number, a name, yes or no, a line number, a date, a letter. |
| Determinism | One question asked twice, expecting the same words; and on GenieX, whether one answer leaks into the next. |
| Vision | A few words drawn on a picture, and what the vision model reads. |
| Embeddings | One sentence embedded on the NPU, compared with a reference made on the processor. |
| Latency | Median time to an answer and tokens a second, against the baseline. |
An audit runs by itself weekly on an unchanged toolchain. After a change it waits for Audit now, or for a Set up or a try you asked for, which audits every accelerator right after.
Audits, on the page, shows the verdict, each accelerator's fingerprint, last audit and history, Audit now and Read the fingerprints now, and the checks' own settings (Audit by itself, Audit every, Read the fingerprint every, and under Advanced which checks run, the lowest embedding cosine, and how much slower than the baseline warns). Each accelerator's badge says how it stands: checked: answers right, a check is due, answered wrong: see Audits.
Editing the accelerators#
Under Set up the accelerators, Accelerators shows the list every agent reads, as JSON: each accelerator's id, slots, request cap, memory, quirks, and its chat, vision, embedding and reranking servers (address, model, and the command that starts it). Order is auto (the NPU first, then graphics cards with 2 GB or more of their own memory, then graphics that share the PC's memory, then the processor), or ids, first preferred. Save the accelerators checks it first, and refuses it if the file changed since the page read it. Plan (downloads nothing), beside Set up, says what a setup would do.
Files#
| Where | What |
|---|---|
%USERPROFILE%\.manor\accelerators\config.json | The accelerators, their order, and the idle times: every agent reads it. The Smith writes only its own keys and keeps the rest. |
%USERPROFILE%\.manor\accelerators\servers, models | The servers Set up unpacked, and the models. On a PC set up before Castellan moved them, %USERPROFILE%\.reeve\servers and models. |
%USERPROFILE%\.smith | The Smith's own data; audits\ holds the checks' settings, audits, baselines and fingerprints. |
SMITH_HOME moves its data folder, and SMITH_PORT its page.
Commands#
Run with node $HOME\.smith\app\src\cli.ts <command>:
| Command | What it does |
|---|---|
run | One look now. |
accelerators | This PC, config.json, and the recommended order. --json for scripts. |
accelerators setup [<id> ...] | Sets up the accelerators named (all by default). --serve chat,vision,embed,rerank chooses what each serves; --dry-run plans only; --yes skips the question. |
fingerprints | Reads the fingerprints once, and runs any weekly audit due. |
audit | Audits every accelerator now. |
start, stop, open, shutdown, status [--json] | Duty and its page, as every agent has them. |
uninstall [--purge] [--dry-run] | Removes it; --purge removes its data too. |
Local API#
The page answers at http://smith.localhost:20202/:
| Address | What it answers |
|---|---|
GET /api/ping | Its version and duty, with the worst audit verdict and whether an audit is due (auditDue). |
GET /api/accelerators | Its last look, as JSON. |
GET /api/checked | Whether each accelerator's models still answer right: {app, state, busy, accelerators: [{id, kind, name, toolchain, state, reason, last}]}. |
GET /api/report | Everything the Audits part of the page is made from. |
Related articles
Is this page right?
If something on it is wrong or out of date, tell us and we'll fix the page.
Still stuck? Write to support@castellan-software.com and mention article 1403. Every version of Smith, and what changed in it, is in its release notes.