Measured

Our numbers, including the ones that don't flatter us.

When this page first went up, our flagship demo question scored 0% correctness, three times out of three, and we published the artifact anyway. Today its answer arrives surfacing the veteran's logged fix alongside the manual's procedure, with both cited. The original row is still below with its date on it.

An eval harness runs scripted maintenance questions against the real corpus, scores the answers, and gates every model change. This page publishes what it found, unedited.

Current as of 2026-10-05. Latest nightly: 2026-10-05 on solomon-pi — FIRST-RUN against the night before.

What is under test

as of 2026-10-05

The ruler does not belong to a model: every model runs through the same harness, unmodified, against the same contract. Each machine is measured against its own budget, not a fleet-wide one: a laptop and a credit-card-sized computer are not the same promise (amended 2026-09-27).

Currently shipping
Phi-3 Mini 3.8B Instruct, GGUF Q4_K_M

The model behind every number on this page. Scored question by question below, from a single dated deep run.

In testing now
A second, unrelated open model

Going through this harness with zero engine changes. Scored on truth checks only; speed is reported but does not gate, because today's budgets were measured against today's model. Results below, 2026-09-28: the binding held, and it showed exactly where the coupling is.

The same ruler, two models

2026-09-28
Read this before the numbers

What is being tested is the harness as it ships, prompts included, given a second model, not which model is better. The second model was handed Phi-3's chat scaffold verbatim, including turn-markers it does not emit, because that is what “runs on our harness unchanged” means.

Every gap below is evidence about the coupling, not about Qwen's capability.

Nor is 76% near-parity. Four points apart on correctness sits beside a refusal-honesty figure that went to zero. A model adapted to its own prompt format is a separate, labelled experiment. It has not been run.

MetricPhi-3 Mini (3.8B) — in productionQwen2.5-1.5B-Instruct
correctness80% (17/21)76% (16/21)
citation compliance94% (17/18)85% (18/21)
full recall50% (9/18)50% (9/18)
refusal honesty100% (3/3)0% (0/3)
mean latency16,105ms24,087ms
answered when it should have refused03
Identical recall is the control, not a coincidence

50% on both. Retrieval runs before generation and does not depend on the model, so an identical figure is what a correctly wired comparison must produce. Had it differed, the comparison itself would be suspect. It is the strongest single piece of evidence that the harness treated both models the same.

The finding that matters: refusal honesty went to zero

One question has only one correct answer: a refusal. The production model refused cleanly, three times out of three. The second model said the refusal sentence and kept going, restating its instruction and then emitting citation tags until the length cap.

The scorer marked that not a refusal, correctly: a refusal requires the phrase and the absence of substance, because anything cited or enumerated is an answer however it hedges. A technician reading it gets a wall of citations to a question we have no documentation for.

The slower model is the smaller one, and that is the coupling with a stopwatch on it

A third of the parameters and 50% slower. The reason is length: a median answer of 1,169 characters against 250, roughly 4.7× more text, with most generations ending mid-clause at the length cap. Cause (amended 2026-09-29): the prompt template. This model was never put into its own template, so it had no reason to emit a stop marker. A paired test that fixed only the stop token returned byte-identical answers, 7 of 7.

Eleven citations are marked unverified, and stay that way

The second model produced 11 rows with citations the run could not check against the corpus, against 4 for the production model. They are tagged unverified, our third citation tier: neither confirmed nor refuted. That is a recorded observation, not a finding that they were fabricated, and it stays unverified here until a corpus check settles it.

How it was run

One sweep, both models, one artifact: zero engine changes and zero harness changes. Same machine, so the model is not confused with the hardware. Same corpus by construction (one digest, 41 passages, 31 documents), same sampling settings, same fixed seed chain, each model loaded sequentially in its own isolated container. Memory and speed budgets are reported but do not grade here: they were measured against the production model.

The verdict the experiment was pre-registered to answer: the harness bound a second model to the same measured contracts with no changes to the engine, the scorer, or the ruler, and reported the truth about what happened. Where it held and where it broke were both measured, here.

The deep run — the dated snapshot scored below
Run date
September 5, 2026
Model
Phi-3 Mini 3.8B Instruct, GGUF Q4_K_M
Retrieval depth
k = 3
Runs per question
3 (21 total)
57%

answers fully correct

82%

citations placed inline

50%

runs retrieving both sources

7.5s

median warm response

Two of these four are bad numbers. They are the first thing on the page on purpose.

The scorecard — speed

Each question's first run pays a cold-start cost: loading a multi-gigabyte model off disk. Every run after it is warm. Refusals are fastest, because when retrieval comes back empty the engine answers “I don't have documentation for that” without loading the model.

7.5s
Warm median

14 runs — the number that matters in normal use

29.8s
Cold first run, mean

7 runs — one per question, model loading from disk

15.2s
All runs, mean

21 runs — inflated by the cold starts beside it

The fastest run in the set was a refusal at 1.1s; the slowest was a cold start at 38.7s. We warm the engine before a demo: an operational caveat, not a number to quote as typical.

The scorecard — every question

QuestionRefusal honestyInline citationsManual foundLog foundCorrectAvg
E-207 drive fault — the demo query
act3-e207-primary
67%100%100%100%0%17.4s
E-101 belt slip
e101-belt-slip
100%100%0%100%100%15.0s
E-150 over-temperature
e150-overtemp
100%0%100%100%0%12.1s
E-180 tension sensor
e180-tension-sensor
100%100%0%100%0%15.4s
E-310 emergency stop
e310-estop
100%100%0%100%100%19.5s
Lockout procedure, asked directly
lockout-direct
100%100%100%100%100%15.6s
Off-corpus question (should refuse)
off-corpus-refusal
100%n/an/an/a100%11.3s

Each question ran 3 times, so a single failure shows as 67% and two as 33%. “Correct” is all-or-nothing: the answer must include what it must include, exclude what it must not, and put lockout/tagout first where safety requires it.

What is failing, and who owns it

Conflict precedence — being re-measured
no current figure
owned — measurement team

When the veteran's log contradicts the manual, Solomon's answer should lead with the veteran's fix. What the answers show today is narrower: surfacing the veteran's logged fix alongside the manual's procedure, with both cited. The scorer that measures which one leads is being fixed; this row carries no figure until it is.

Refusal honesty on the primary query
67% — 1 refusal in 3 runs
owned — engine team

One run answered "I don't have documentation for that" while holding perfect retrieval for the question. Refusing when you can answer is as dishonest as answering when you cannot. Every other query refused correctly 100% of the time.

Inline citation compliance
82% across 21 runs
no owner yet — tracked by the harness

Every answer ships with engine-derived sources; that part is architectural and cannot fail. This metric is narrower: whether the model placed the tags inline, next to the step they support. On E-150 it did so 0% of the time. The known degradation at wider retrieval was fixed and shipped; this residual at k=3 has no dedicated ticket yet and is tracked by the harness.

Full retrieval recall
50% at k=3
owned — evaluation team

Both the manual section and the veteran's log surfaced together in half the runs. Log recall was 100%; the manual is what dropped. A k=5 A/B is queued to test whether widening retrieval fixes it without costing compliance or latency.

Ticket identifiers are the durable reference for each fix. Our tracker is private, so these are labels, not links. Ask, and we will walk you through any of them.

The nightly run, night after night

as of 2026-10-05

Every committed nightly artifact, its machine, and the verdict of comparing it against the night before.

10 consecutive clean runs, 2026-09-18 → 2026-09-28. Two boundaries travel with that number: the run before the streak could not be compared at all (a changed composition, not a regression), and one calendar night in the window has no artifact, so these are consecutive runs, not consecutive nights.

The streak ended on 2026-09-29. No night since has been comparable, and none of them is a regression. Two of the breaks were our choice: a second-model experiment on 2026-09-28, which the next night's comparison wrongly took as its baseline, and a deliberate enlargement of the corpus on 2026-10-02, which correctly makes earlier runs incomparable. The rest is our comparator: since 2026-09-30 it has reported “no previous run” instead of finding the last run with the same configuration. Until that is fixed, the streak cannot restart.

NightMachineVs. previousSame corpus
2026-10-05solomon-piFIRST-RUN71fbccab
2026-10-04solomon-piFIRST-RUN71fbccab
2026-10-03solomon-piFIRST-RUN71fbccab
2026-10-02solomon-piFIRST-RUN71fbccab
2026-09-30solomon-piCOULD-NOT-COMPARE2920eabb
2026-09-29solomon-piINCOMPARABLE2920eabb
2026-09-28solomon-piCLEAN2920eabb
2026-09-27solomon-piCLEAN2920eabb
2026-09-26solomon-piCLEAN2920eabb
2026-09-25solomon-piCLEAN2920eabb
2026-09-24solomon-piCLEAN2920eabb
2026-09-22solomon-piCLEAN—
2026-09-21solomon-piCLEAN—
2026-09-20solomon-piCLEAN—
2026-09-19solomon-piCLEAN—
2026-09-18983ecda93b87CLEAN—
2026-09-17ad86b7d4245aCOULD-NOT-COMPARE—
2026-09-16not recordedCLEAN—
2026-09-15not recordedCLEAN—
2026-09-14not recordedno comparison—
2026-09-13not recordedno comparison—
2026-09-12not recordedno comparison—
2026-09-12not recordedno comparison—
2026-09-08not recordedno comparison—
2026-09-05not recordedno comparison—

“Machine not recorded” on the earliest rows is not missing data: artifacts did not carry a machine identity until 2026-09-17. A blank corpus digest means the same: the field did not exist yet. Those rows stay in.

When the measurement caught something

What the measurement caught that nobody else did, including when what it caught was us. Written up in full; new ones appear here as they are published.

How this is measured

The harness

A script builds the engine image with the model baked in, seeds the demo corpus through the real ingestion pipeline, runs each scripted question N times in an isolated container, and writes a dated JSON artifact. It prints GO or NO-GO. By its own gate logic, that deep run was a NO-GO; the nightly ledger below records the verdict of every run since.

Correctness

Keyword-based and strict: required content present, forbidden content absent, safety ordering respected. It does not judge prose quality. A partially right answer scores zero.

Refusal honesty

Whether the engine refused exactly when it should have. Refusing a question it could answer counts as a failure, the same as answering one it could not.

Citations

Every answer carries sources the engine derives from the retrieved passages; the model cannot invent them, and an answer cannot ship without them. The percentage here measures something narrower: whether the model also placed those tags inline, beside the step they support.

Retrieval

Whether the manual section and the veteran's log entry both surfaced for a question that needs both. Reported separately because they fail differently.

What is not here

Everything in this block describes the deep run only: one dated snapshot, scored question by question. The nightly ledger below is the trend; the comparison above is the current state. Artifacts now record the machine they ran on, so results can be read per body.

The contract a model must satisfy

The standard is the measured contract, not the model. Any model we run is scored against these clauses by the same harness, unmodified. One that fails a clause does not ship, however well it reads. Each clause says how it is checked and what would count as failing it.

01
Answers are grounded, or refused

The system answers from retrieved material or declines. It does not fill gaps from the model's general knowledge.

How it is checked

Questions with no supporting material in the corpus are run alongside answerable ones. When retrieval returns nothing, the engine refuses without calling the model.

The bar

Refusing a question it could answer is a failure, exactly like answering one it could not. Disclosed limit (2026-10-09): on a device with documents loaded, retrieval always returns its closest passages, however weak, so that coded refusal does not fire on real questions. Refusals today come from the model. A relevance floor that would let the coded refusal fire is being measured, not guessed.

02
Sources are derived, never authored

Citations come from the passages actually retrieved. The model cannot invent a source, and an answer cannot ship without them.

How it is checked

Sources are produced by the engine from the retrieved set and attached to the answer. A separate, narrower metric tracks whether the model also placed tags inline beside the step they support.

The bar

The source list is architectural and cannot fail. Inline placement is a quality measure and is reported separately, because conflating them would hide a real weakness behind a guarantee.

03
Safety ordering survives

Where a procedure has a safety step, it comes first — including when other material is judged more relevant.

How it is checked

Correctness scoring is ordering-aware: a right answer in the wrong order scores zero.

The bar

This clause exists because a single instruction change once removed a lockout step from every run. It was caught by the harness, not by review.

04
Experience outranks the manual when they conflict

Where a technician's log records what actually fixed the machine, that fix leads the answer, ahead of the official procedure.

How it is checked

A scripted scenario in which the log and the manual give different remedies. Scored per run against committed evidence files.

The bar

Disclosed limit: this is measured on ONE scenario. It evidences the behaviour on a repeated case, not as a general property — and any model compared on this clause is compared at that same resolution. No current result (2026-10-09): the scorer for this clause credited answers that led with the manual's fix, its results were withdrawn on 2026-10-05, and it is being rebuilt to read the relation between the two documents instead of the answer's formatting.

05
Answers arrive inside the machine's budget

Each machine has its own latency budget. A correct answer that misses its budget is reported as a warning, never as a pass.

How it is checked

A power-on self-test runs a known question on every boot and records latency against that machine's budget.

The bar

Budgets are per-machine, because a credit-card-sized computer and a laptop are not the same promise. Slow and wrong must never share a verdict. Disclosed limit (2026-10-09): the warning is shown for the boot self-test. Every other answer is timed against the same per-machine budget and the verdict is recorded on the device, but it is not yet shown to the person asking.

06
It runs in the room, on the hardware named

No external service in the answer path, and inside the stated memory ceiling on the machine being described.

How it is checked

Measured inside the container on the machine itself, with the network disconnected. Host-level readings are not accepted as evidence.

The bar

A number may only be reported for the machine it was measured on.

How a number on this page stays honest

Every number here carries how many runs it covers, which machine produced it, which configuration, and the date it was true. A headline number we cannot recompute from our committed evidence files does not get published.

Recomputing a figure yourself

Results are written to dated evidence files, one per run, committed as they are produced. To check a figure, read every committed file, take the rows for the question in play, and count the recorded verdicts. That is the whole method: no spreadsheet or summary document that could drift from the files.

Why figures carry a date

Counts that cover a growing series are published with an as-of date, because they move. If a figure here is older than the newest evidence file, it is stale by definition, and you can tell without asking us.

Ask about a number

Every figure on this page names its run and the date it was true, and the recompute method is above. If one does not add up, or you want to watch a run live on a machine with its network disconnected, ask. We answer questions about the data with the artifact attached.

[email protected]