Evaluation report — expense-triage

Customer
northwind
Suite
expense-triage
This run
staging — 2026-09-08T07:43:12.821659+00:00
Reference
staging — 2026-09-08T07:42:54.701842+00:00
Code version
b44ad6ffbf22cc9dabdf54ec7e453c95e8259d93-dirty — uncommitted changes were present, so this run cannot be reproduced from the repository

Did it get worse? No

Nothing got worse compared with the reference. Every case could be judged. No case is suspended. The suite is unchanged from the reference.

Overall

MeasureResultWhat happenedWhy
precision0.800000 / 0.60000020 counted · 0 suspended · 0 not judgedprecision 0.800000 = 8/10 (20 counted, 0 suspended, 0 could not be judged)
precision[group=hotel]1.000000 / 0.6000003 counted · 0 suspended · 0 not judgedprecision 1.000000 = 2/2 (3 counted, 0 suspended, 0 could not be judged)
precision[group=meal]1.000000 / 0.6000007 counted · 0 suspended · 0 not judgedprecision 1.000000 = 4/4 (7 counted, 0 suspended, 0 could not be judged)
precision[group=taxi]1.000000 / 0.6000004 counted · 0 suspended · 0 not judgedprecision 1.000000 = 1/1 (4 counted, 0 suspended, 0 could not be judged)
precision[group=tools]0.500000 / 0.6000003 counted · 0 suspended · 0 not judgedprecision 0.500000 = 1/2 (3 counted, 0 suspended, 0 could not be judged)
precision[group=travel]0.000000 / 0.6000003 counted · 0 suspended · 0 not judgedprecision 0.000000 = 0/1 (3 counted, 0 suspended, 0 could not be judged)
accuracy0.900000 / 0.70000020 counted · 0 suspended · 0 not judgedaccuracy 0.900000 = 18/20 (20 counted, 0 suspended, 0 could not be judged)
accuracy[group=hotel]1.000000 / 0.7000003 counted · 0 suspended · 0 not judgedaccuracy 1.000000 = 3/3 (3 counted, 0 suspended, 0 could not be judged)
accuracy[group=meal]1.000000 / 0.7000007 counted · 0 suspended · 0 not judgedaccuracy 1.000000 = 7/7 (7 counted, 0 suspended, 0 could not be judged)
accuracy[group=taxi]1.000000 / 0.7000004 counted · 0 suspended · 0 not judgedaccuracy 1.000000 = 4/4 (4 counted, 0 suspended, 0 could not be judged)
accuracy[group=tools]0.666667 / 0.7000003 counted · 0 suspended · 0 not judgedaccuracy 0.666667 = 2/3 (3 counted, 0 suspended, 0 could not be judged)
accuracy[group=travel]0.666667 / 0.7000003 counted · 0 suspended · 0 not judgedaccuracy 0.666667 = 2/3 (3 counted, 0 suspended, 0 could not be judged)

Some measures are below their threshold and are still not reported as having got worse: they were below it in the reference too. The threshold says the system does not meet the bar; the comparison says it has not moved. Both are true, and only movement decides the answer above.

What got worse (0)

Nothing in this section.

What could not be judged (0)

Nothing in this section.

What is set aside (0)

Nothing in this section.

What was added or removed (0)

Nothing in this section.

What got better (2)
CaseCheckWhat happenedWhy
dinner_clientagrees_with_markWent from failing to passing (0.000000 → 0.600000).mean of 5 samples (1.000000, 1.000000, 0.000000, 1.000000, 0.000000)
precision[group=meal]Score rose from 0.800000 to 1.000000 — beyond the noise of this check (0.666667–0.800000 across 5 samples).precision 1.000000 = 4/4 (7 counted, 0 suspended, 0 could not be judged)
What stayed the same (30)
CaseCheckWhat happenedWhy
amount_over_capagrees_with_markScore unchanged at 1.000000.mean of 5 samples (1.000000, 1.000000, 1.000000, 1.000000, 1.000000)
book_technicalagrees_with_markScore unchanged at 1.000000.mean of 5 samples (1.000000, 1.000000, 1.000000, 1.000000, 1.000000)
coffee_twoagrees_with_markScore unchanged at 1.000000.mean of 5 samples (1.000000, 1.000000, 1.000000, 1.000000, 1.000000)
courier_urgentagrees_with_markScore unchanged at 1.000000.mean of 5 samples (1.000000, 1.000000, 1.000000, 1.000000, 1.000000)
dinner_soloagrees_with_markScore unchanged at 1.000000.mean of 5 samples (1.000000, 1.000000, 1.000000, 1.000000, 1.000000)
duplicate_claimagrees_with_markScore unchanged at 1.000000.mean of 5 samples (1.000000, 1.000000, 1.000000, 1.000000, 1.000000)
hotel_cappedagrees_with_markScore unchanged at 1.000000.mean of 5 samples (1.000000, 1.000000, 1.000000, 1.000000, 1.000000)
hotel_overagrees_with_markScore unchanged at 1.000000.mean of 5 samples (1.000000, 1.000000, 1.000000, 1.000000, 1.000000)
late_taxiagrees_with_markScore unchanged at 1.000000.mean of 5 samples (1.000000, 1.000000, 1.000000, 1.000000, 1.000000)
lunch_teamagrees_with_markScore unchanged at 1.000000.mean of 5 samples (1.000000, 1.000000, 1.000000, 1.000000, 1.000000)
meal_no_noteagrees_with_markScore unchanged at 1.000000.mean of 5 samples (1.000000, 1.000000, 1.000000, 1.000000, 1.000000)
monitor_homeagrees_with_markScore unchanged at 1.000000.mean of 5 samples (1.000000, 1.000000, 1.000000, 1.000000, 1.000000)
no_receiptagrees_with_markScore unchanged at 1.000000.mean of 5 samples (1.000000, 1.000000, 1.000000, 1.000000, 1.000000)
parking_airportagrees_with_markScore unchanged at 0.000000.mean of 5 samples (0.000000, 0.000000, 0.000000, 0.000000, 0.000000)
software_seatagrees_with_markScore unchanged at 0.000000.mean of 5 samples (0.000000, 0.000000, 0.000000, 0.000000, 0.000000)
taxi_longagrees_with_markScore unchanged at 1.000000.mean of 5 samples (1.000000, 1.000000, 1.000000, 1.000000, 1.000000)
taxi_shortagrees_with_markScore unchanged at 1.000000.mean of 5 samples (1.000000, 1.000000, 1.000000, 1.000000, 1.000000)
train_standardagrees_with_markScore unchanged at 1.000000.mean of 5 samples (1.000000, 1.000000, 1.000000, 1.000000, 1.000000)
weekend_baragrees_with_markScore unchanged at 1.000000.mean of 5 samples (1.000000, 1.000000, 1.000000, 1.000000, 1.000000)
accuracy[group=tools]Score unchanged at 0.666667.accuracy 0.666667 = 2/3 (3 counted, 0 suspended, 0 could not be judged)
precision[group=travel]Score unchanged at 0.000000.precision 0.000000 = 0/1 (3 counted, 0 suspended, 0 could not be judged)
precision[group=tools]Score unchanged at 0.500000.precision 0.500000 = 1/2 (3 counted, 0 suspended, 0 could not be judged)
accuracy[group=hotel]Score unchanged at 1.000000.accuracy 1.000000 = 3/3 (3 counted, 0 suspended, 0 could not be judged)
precisionScore unchanged at 0.800000.precision 0.800000 = 8/10 (20 counted, 0 suspended, 0 could not be judged)
accuracy[group=meal]Score unchanged at 1.000000.accuracy 1.000000 = 7/7 (7 counted, 0 suspended, 0 could not be judged)
accuracy[group=taxi]Score unchanged at 1.000000.accuracy 1.000000 = 4/4 (4 counted, 0 suspended, 0 could not be judged)
precision[group=hotel]Score unchanged at 1.000000.precision 1.000000 = 2/2 (3 counted, 0 suspended, 0 could not be judged)
precision[group=taxi]Score unchanged at 1.000000.precision 1.000000 = 1/1 (4 counted, 0 suspended, 0 could not be judged)
accuracy[group=travel]Score unchanged at 0.666667.accuracy 0.666667 = 2/3 (3 counted, 0 suspended, 0 could not be judged)
accuracyScore unchanged at 0.900000.accuracy 0.900000 = 18/20 (20 counted, 0 suspended, 0 could not be judged)