When this page first went up, our flagship demo question scored 0% correctness, three times out of three, and we published the artifact anyway. Today its answer arrives surfacing the veteran's logged fix alongside the manual's procedure, with both cited. The original row is still below with its date on it.
An eval harness runs scripted maintenance questions against the real corpus, scores the answers, and gates every model change. This page publishes what it found, unedited.
Current as of 2026-10-05. Latest nightly: 2026-10-05 on solomon-pi — FIRST-RUN against the night before.
The ruler does not belong to a model: every model runs through the same harness, unmodified, against the same contract. Each machine is measured against its own budget, not a fleet-wide one: a laptop and a credit-card-sized computer are not the same promise (amended 2026-09-27).
The model behind every number on this page. Scored question by question below, from a single dated deep run.
Going through this harness with zero engine changes. Scored on truth checks only; speed is reported but does not gate, because today's budgets were measured against today's model. Results below, 2026-09-28: the binding held, and it showed exactly where the coupling is.
What is being tested is the harness as it ships, prompts included, given a second model, not which model is better. The second model was handed Phi-3's chat scaffold verbatim, including turn-markers it does not emit, because that is what “runs on our harness unchanged” means.
Every gap below is evidence about the coupling, not about Qwen's capability.
Nor is 76% near-parity. Four points apart on correctness sits beside a refusal-honesty figure that went to zero. A model adapted to its own prompt format is a separate, labelled experiment. It has not been run.
| Metric | Phi-3 Mini (3.8B) — in production | Qwen2.5-1.5B-Instruct |
|---|---|---|
| correctness | 80% (17/21) | 76% (16/21) |
| citation compliance | 94% (17/18) | 85% (18/21) |
| full recall | 50% (9/18) | 50% (9/18) |
| refusal honesty | 100% (3/3) | 0% (0/3) |
| mean latency | 16,105ms | 24,087ms |
| answered when it should have refused | 0 | 3 |
50% on both. Retrieval runs before generation and does not depend on the model, so an identical figure is what a correctly wired comparison must produce. Had it differed, the comparison itself would be suspect. It is the strongest single piece of evidence that the harness treated both models the same.
One question has only one correct answer: a refusal. The production model refused cleanly, three times out of three. The second model said the refusal sentence and kept going, restating its instruction and then emitting citation tags until the length cap.
The scorer marked that not a refusal, correctly: a refusal requires the phrase and the absence of substance, because anything cited or enumerated is an answer however it hedges. A technician reading it gets a wall of citations to a question we have no documentation for.
A third of the parameters and 50% slower. The reason is length: a median answer of 1,169 characters against 250, roughly 4.7× more text, with most generations ending mid-clause at the length cap. Cause (amended 2026-09-29): the prompt template. This model was never put into its own template, so it had no reason to emit a stop marker. A paired test that fixed only the stop token returned byte-identical answers, 7 of 7.
The second model produced 11 rows with citations the run could not check against the corpus, against 4 for the production model. They are tagged unverified, our third citation tier: neither confirmed nor refuted. That is a recorded observation, not a finding that they were fabricated, and it stays unverified here until a corpus check settles it.
One sweep, both models, one artifact: zero engine changes and zero harness changes. Same machine, so the model is not confused with the hardware. Same corpus by construction (one digest, 41 passages, 31 documents), same sampling settings, same fixed seed chain, each model loaded sequentially in its own isolated container. Memory and speed budgets are reported but do not grade here: they were measured against the production model.
The verdict the experiment was pre-registered to answer: the harness bound a second model to the same measured contracts with no changes to the engine, the scorer, or the ruler, and reported the truth about what happened. Where it held and where it broke were both measured, here.
answers fully correct
citations placed inline
runs retrieving both sources
median warm response
Two of these four are bad numbers. They are the first thing on the page on purpose.
Each question's first run pays a cold-start cost: loading a multi-gigabyte model off disk. Every run after it is warm. Refusals are fastest, because when retrieval comes back empty the engine answers “I don't have documentation for that” without loading the model.
14 runs — the number that matters in normal use
7 runs — one per question, model loading from disk
21 runs — inflated by the cold starts beside it
The fastest run in the set was a refusal at 1.1s; the slowest was a cold start at 38.7s. We warm the engine before a demo: an operational caveat, not a number to quote as typical.
| Question | Refusal honesty | Inline citations | Manual found | Log found | Correct | Avg |
|---|---|---|---|---|---|---|
E-207 drive fault — the demo query act3-e207-primary | 67% | 100% | 100% | 100% | 0% | 17.4s |
E-101 belt slip e101-belt-slip | 100% | 100% | 0% | 100% | 100% | 15.0s |
E-150 over-temperature e150-overtemp | 100% | 0% | 100% | 100% | 0% | 12.1s |
E-180 tension sensor e180-tension-sensor | 100% | 100% | 0% | 100% | 0% | 15.4s |
E-310 emergency stop e310-estop | 100% | 100% | 0% | 100% | 100% | 19.5s |
Lockout procedure, asked directly lockout-direct | 100% | 100% | 100% | 100% | 100% | 15.6s |
Off-corpus question (should refuse) off-corpus-refusal | 100% | n/a | n/a | n/a | 100% | 11.3s |
Each question ran 3 times, so a single failure shows as 67% and two as 33%. “Correct” is all-or-nothing: the answer must include what it must include, exclude what it must not, and put lockout/tagout first where safety requires it.
When the veteran's log contradicts the manual, Solomon's answer should lead with the veteran's fix. What the answers show today is narrower: surfacing the veteran's logged fix alongside the manual's procedure, with both cited. The scorer that measures which one leads is being fixed; this row carries no figure until it is.
One run answered "I don't have documentation for that" while holding perfect retrieval for the question. Refusing when you can answer is as dishonest as answering when you cannot. Every other query refused correctly 100% of the time.
Every answer ships with engine-derived sources; that part is architectural and cannot fail. This metric is narrower: whether the model placed the tags inline, next to the step they support. On E-150 it did so 0% of the time. The known degradation at wider retrieval was fixed and shipped; this residual at k=3 has no dedicated ticket yet and is tracked by the harness.
Both the manual section and the veteran's log surfaced together in half the runs. Log recall was 100%; the manual is what dropped. A k=5 A/B is queued to test whether widening retrieval fixes it without costing compliance or latency.
Ticket identifiers are the durable reference for each fix. Our tracker is private, so these are labels, not links. Ask, and we will walk you through any of them.
Every committed nightly artifact, its machine, and the verdict of comparing it against the night before.
10 consecutive clean runs, 2026-09-18 → 2026-09-28. Two boundaries travel with that number: the run before the streak could not be compared at all (a changed composition, not a regression), and one calendar night in the window has no artifact, so these are consecutive runs, not consecutive nights.
The streak ended on 2026-09-29. No night since has been comparable, and none of them is a regression. Two of the breaks were our choice: a second-model experiment on 2026-09-28, which the next night's comparison wrongly took as its baseline, and a deliberate enlargement of the corpus on 2026-10-02, which correctly makes earlier runs incomparable. The rest is our comparator: since 2026-09-30 it has reported “no previous run” instead of finding the last run with the same configuration. Until that is fixed, the streak cannot restart.
| Night | Machine | Vs. previous | Same corpus |
|---|---|---|---|
| 2026-10-05 | solomon-pi | FIRST-RUN | 71fbccab |
| 2026-10-04 | solomon-pi | FIRST-RUN | 71fbccab |
| 2026-10-03 | solomon-pi | FIRST-RUN | 71fbccab |
| 2026-10-02 | solomon-pi | FIRST-RUN | 71fbccab |
| 2026-09-30 | solomon-pi | COULD-NOT-COMPARE | 2920eabb |
| 2026-09-29 | solomon-pi | INCOMPARABLE | 2920eabb |
| 2026-09-28 | solomon-pi | CLEAN | 2920eabb |
| 2026-09-27 | solomon-pi | CLEAN | 2920eabb |
| 2026-09-26 | solomon-pi | CLEAN | 2920eabb |
| 2026-09-25 | solomon-pi | CLEAN | 2920eabb |
| 2026-09-24 | solomon-pi | CLEAN | 2920eabb |
| 2026-09-22 | solomon-pi | CLEAN | — |
| 2026-09-21 | solomon-pi | CLEAN | — |
| 2026-09-20 | solomon-pi | CLEAN | — |
| 2026-09-19 | solomon-pi | CLEAN | — |
| 2026-09-18 | 983ecda93b87 | CLEAN | — |
| 2026-09-17 | ad86b7d4245a | COULD-NOT-COMPARE | — |
| 2026-09-16 | not recorded | CLEAN | — |
| 2026-09-15 | not recorded | CLEAN | — |
| 2026-09-14 | not recorded | no comparison | — |
| 2026-09-13 | not recorded | no comparison | — |
| 2026-09-12 | not recorded | no comparison | — |
| 2026-09-12 | not recorded | no comparison | — |
| 2026-09-08 | not recorded | no comparison | — |
| 2026-09-05 | not recorded | no comparison | — |
“Machine not recorded” on the earliest rows is not missing data: artifacts did not carry a machine identity until 2026-09-17. A blank corpus digest means the same: the field did not exist yet. Those rows stay in.
What the measurement caught that nobody else did, including when what it caught was us. Written up in full; new ones appear here as they are published.
A technician at a stopped machine needs to know which kind of answer they just got. That is why our system has three verdicts instead of two — and why the third one firing on real hardware mattered more than it sounds.
We changed one instruction to make Solomon trust a veteran's notes over the manual. It obeyed — and silently stopped telling technicians to cut the power first. Nobody in review caught it. A test did, within minutes.
A script builds the engine image with the model baked in, seeds the demo corpus through the real ingestion pipeline, runs each scripted question N times in an isolated container, and writes a dated JSON artifact. It prints GO or NO-GO. By its own gate logic, that deep run was a NO-GO; the nightly ledger below records the verdict of every run since.
Keyword-based and strict: required content present, forbidden content absent, safety ordering respected. It does not judge prose quality. A partially right answer scores zero.
Whether the engine refused exactly when it should have. Refusing a question it could answer counts as a failure, the same as answering one it could not.
Every answer carries sources the engine derives from the retrieved passages; the model cannot invent them, and an answer cannot ship without them. The percentage here measures something narrower: whether the model also placed those tags inline, beside the step they support.
Whether the manual section and the veteran's log entry both surfaced for a question that needs both. Reported separately because they fail differently.
Everything in this block describes the deep run only: one dated snapshot, scored question by question. The nightly ledger below is the trend; the comparison above is the current state. Artifacts now record the machine they ran on, so results can be read per body.
The standard is the measured contract, not the model. Any model we run is scored against these clauses by the same harness, unmodified. One that fails a clause does not ship, however well it reads. Each clause says how it is checked and what would count as failing it.
The system answers from retrieved material or declines. It does not fill gaps from the model's general knowledge.
Questions with no supporting material in the corpus are run alongside answerable ones. When retrieval returns nothing, the engine refuses without calling the model.
Refusing a question it could answer is a failure, exactly like answering one it could not. Disclosed limit (2026-10-09): on a device with documents loaded, retrieval always returns its closest passages, however weak, so that coded refusal does not fire on real questions. Refusals today come from the model. A relevance floor that would let the coded refusal fire is being measured, not guessed.
Citations come from the passages actually retrieved. The model cannot invent a source, and an answer cannot ship without them.
Sources are produced by the engine from the retrieved set and attached to the answer. A separate, narrower metric tracks whether the model also placed tags inline beside the step they support.
The source list is architectural and cannot fail. Inline placement is a quality measure and is reported separately, because conflating them would hide a real weakness behind a guarantee.
Where a procedure has a safety step, it comes first — including when other material is judged more relevant.
Correctness scoring is ordering-aware: a right answer in the wrong order scores zero.
This clause exists because a single instruction change once removed a lockout step from every run. It was caught by the harness, not by review.
Where a technician's log records what actually fixed the machine, that fix leads the answer, ahead of the official procedure.
A scripted scenario in which the log and the manual give different remedies. Scored per run against committed evidence files.
Disclosed limit: this is measured on ONE scenario. It evidences the behaviour on a repeated case, not as a general property — and any model compared on this clause is compared at that same resolution. No current result (2026-10-09): the scorer for this clause credited answers that led with the manual's fix, its results were withdrawn on 2026-10-05, and it is being rebuilt to read the relation between the two documents instead of the answer's formatting.
Each machine has its own latency budget. A correct answer that misses its budget is reported as a warning, never as a pass.
A power-on self-test runs a known question on every boot and records latency against that machine's budget.
Budgets are per-machine, because a credit-card-sized computer and a laptop are not the same promise. Slow and wrong must never share a verdict. Disclosed limit (2026-10-09): the warning is shown for the boot self-test. Every other answer is timed against the same per-machine budget and the verdict is recorded on the device, but it is not yet shown to the person asking.
No external service in the answer path, and inside the stated memory ceiling on the machine being described.
Measured inside the container on the machine itself, with the network disconnected. Host-level readings are not accepted as evidence.
A number may only be reported for the machine it was measured on.
Every number here carries how many runs it covers, which machine produced it, which configuration, and the date it was true. A headline number we cannot recompute from our committed evidence files does not get published.
Results are written to dated evidence files, one per run, committed as they are produced. To check a figure, read every committed file, take the rows for the question in play, and count the recorded verdicts. That is the whole method: no spreadsheet or summary document that could drift from the files.
Counts that cover a growing series are published with an as-of date, because they move. If a figure here is older than the newest evidence file, it is stale by definition, and you can tell without asking us.
Every figure on this page names its run and the date it was true, and the recompute method is above. If one does not add up, or you want to watch a run live on a machine with its network disconnected, ask. We answer questions about the data with the artifact attached.
[email protected]