2026-09-13T00:54:19Z | TERMINAL | lane=friction-p1-closer | tickets=OMN-18233,OMN-18235,OMN-18236,OMN-18238,OMN-18239 | parent=OMN-18232 | PHASE 1 COMPLETE, 4 of 5 DONE. All 8 PRs merged. OMN-18233 DONE omnibase_infra#3465 squash 877610ee1346 (four-clause verified-supersession predicate; the closer ignores a closed cascade bump only on proven replacement; no title read, no ordering clause; every clause fails closed). OMN-18235 DONE omnibase_core#1685 squash 2b8bf73022193a6c1bb4dfd6e2875184fc886c6a (cascade provenance asserted by the generator's own test + pre-commit hook; positive control measured: heading removed -> 3 red, restored -> 8 green, workflow byte-identical). OMN-18236 DONE onex_change_control#9286 squash 6ef7427a3c11 + omnimarket#2505 squash 899d04f8596c (binds_ac claim must name a real criterion at the revision it was accepted against; new OCC-local ac_bindings record; gate wired as pre-commit hook + unconditional CI job; measured over all 67 declaring contracts against real Linear bodies: 58 clean, 9 flagged, all true). OMN-18238 DONE omnibase_infra#3469 squash d2cba4721393 + omnimarket#2507 squash 178b9cff8c25 (a proposed binding does not satisfy the closer; verifier surfaces drafts; emitter renderer cannot express an acceptance). OMN-18239 IN PROGRESS onex_change_control#9290 squash 2eb06a598b59 -- measurement and grow-only ratchet landed; AC1 is a week-over-week trend and cannot be asserted on the day the measurement is built; needs one more week of data, nothing further to build. | FINDINGS: adoption number corrected 61/67 grep -> 63 parsed (4 contracts name the field in prose while declaring none: OMN-15651, OMN-16106, OMN-17276, OMN-18073); 9 tickets bind criteria their bodies never label (OMN-15660, OMN-17083, OMN-17228, OMN-17372, OMN-17812, OMN-18067, OMN-18164, OMN-18167, OMN-18185) and are unclosable through the binding path until labelled; the closer under-reads a ticket whose criteria span two sections (OMN-18186 has 6, it reads 4) and needs the same one-line fix the OCC reader took; omnibase_infra's own dependency-cascade generator emits no provenance at all. | NO TICKETS MINTED (lane is barred); those four need minting by the orchestrator.
2026-09-13T01:54:32Z | TERMINAL | lane=orchestrator | ticket=OMN-18224 parent=OMN-18205 | consent=docs/tracking/ROLLING_WORK_LEDGER.md:7118 | Stored omninode/ci-runners/github-app in us-east-1 (created 2026-09-13T01:54:17Z, one AWSCURRENT version, keys app_id/installation_id/private_key) from the operator's downloaded pem via the rev-3 recipe (umask 077, mktemp trailing Xs, jq --rawfile, --secret-string file://, trap cleanup); pem deleted after describe-secret confirmed; no value printed; App 4926129 installation 161264140 permission organization_self_hosted_runners=write read back live. Not yet done: IAM simulation readback, first drill, canary.
2026-09-13T02:10:04Z | TERMINAL | lane=cloud-ci-first-boot | OMN-18224 DONE (child of OMN-18205) | consent=docs/tracking/ROLLING_WORK_LEDGER.md:6999 and :7118 | THE CLOUD CI RUNNER FLEET HAS NOW BOOTED, REGISTERED, RUN A REAL JOB, DEREGISTERED AND TORN ITSELF DOWN. Revision-3 residual 1 (nothing on this fleet has ever run a job or a drill) is CLOSED. IDS: instance i-07dacc73186c1217f c7a.2xlarge us-east-1c on ami-0b5358cc8c5df0b02 (al2023-ami-2023.12.20260909.0); canary run 34731958904 conclusion success; runner omni-cloud-i-07dacc73186c1217f in group omni-cloud-ci (4); CloudWatch stream i-07dacc73186c1217f in /omni/ci-runners/bootstrap. SECTION 6: 22 policy simulations, EVERY ONE as expected, each allow paired with its deny -- broker reads the exact ARN allowed, instance role implicitDeny on both read verbs, NEAR-MATCH <arn>-other implicitDeny (the revision-1 regression stays fixed), unrelated secret implicitDeny, broker describe/put/delete implicitDeny, instance terminate+complete-lifecycle allowed in its own ASG and implicitDeny in the k3s dev-system ASG, set-desired/update implicitDeny, broker invoke allowed and other-function implicitDeny, bootstrap log create-stream/put allowed with get/delete implicitDeny and the BROKER log group implicitDeny as the negative control, administrator get-secret-value allowed as the simulator positive control. No secret value read. SECTION 7a SKIPPED BY ITS OWN STEP 0: the drill's premise is that describe-secret returns ResourceNotFoundException; the credential now exists so the procedure routes itself to 7b. SECTION 7b TIMELINE: 01:57:09Z set-desired-capacity 1; 01:57:18Z launched; 01:57:46Z bootstrap start; 01:58:26Z listening for jobs; 02:00:40Z API reports online; 02:00:46Z dispatched ONCE; 02:00:49-02:00:54Z job ran all six steps green, never rerun; 02:00:56Z boot-completed rc=0; 02:00:57Z self-terminated desired 1->0; 02:06:43Z terminated. IDENTITY PROVEN FROM THE JOB LOG not the green check: runner name matched omni-cloud-i-*, and IMDS answered instance-id i-07dacc73186c1217f type c7a.2xlarge az us-east-1c, the discriminator a lab runner cannot satisfy. LOCAL DOCKER: server 25.0.16, nginx:1.27-alpine pulled fresh, published port reachable on the runner's own localhost -- first demonstration of the capability this fleet exists for. FOUR POST-JOB CHECKS ALL PASS: group 4 runners 0; fleet-filtered describe-instances 0 reservations with POSITIVE CONTROL returning 3 unfiltered; ASG 0/0/4 zero instances no suspended processes and nobody lowered the dial by hand; the forwarded stream ends [bootstrap] boot-completed rc=0 at 2026-09-13T02:00:56Z. Both guard alarms OK throughout. TWO NEW DEFECTS FOUND BY RUNNING IT, no ticket minted (none granted): D1 the terminating lifecycle hook omninode-ci-runner-deregister (300s heartbeat, DefaultResult CONTINUE) is never completed -- the exit trap terminates and nothing calls CompleteLifecycleAction, so every job pays the full 300s; measured Terminating:Wait 02:00:57-02:05:57, instance lifetime 9m25s for a 5s job. Bounded and self-clearing, cents per job, but dead time on every future job; the instance role already holds CompleteLifecycleAction so the fix is to call it from the trap or drop a hook the ephemeral runner does not need. D2 the canary's 'runner labels:' line renders empty (join(runner.labels) produced nothing) -- cosmetic, the two real assertions fired, but it reads as a passing check that asserts nothing. RUNBOOK GAPS: 7b step 2 says to dispatch once a line appears in the runner listing, but the line appeared at 01:58:32 with status offline and the API did not say online until 02:00:40 -- 2m14s after the listener was already up; the step should key on status==online. And 7a is now unrunnable as written on a credentialed fleet, so the only remaining written path to drill the FAILURE teardown is 7a's second pass (renaming the bootstrap object), which does not depend on the credential. RESIDUALS: the failure-path teardown and the pre-boot-script handler are still unexercised (the AL2023 first-boot half of that gate IS now closed); D1; the broker one-hour token, the org-wide runner-group permission, the backstop-window bill and the unavailable audit-log endpoint all unchanged; both alarm subscribers still PendingConfirmation; still NO scale-up trigger so no job class routes here yet. NOT TOUCHED: desired capacity never exceeded 1 and is back at 0 with a positive control, ASG max unchanged, runner group 4 membership and settings, every Actions routing variable, and every credential -- no value read, no create, no rotation, no revoke.
2026-09-13T02:25:00Z | TERMINAL | lane=lab-delegation-dogfood | actor=claude:opus5:subagent | closes-CLAIM=docs/tracking/ROLLING_WORK_LEDGER.md:7122 | tickets=OMN-18200 related=OMN-18185,OMN-18265,OMN-17082 | ANSWER TO THE OPERATOR QUESTION: YES on the local-bus path, NO on the customer-key path, and the customer-key NO is by design rather than a defect. | (1) CUSTOMER-KEY PATH PROVEN TO THE REFUSAL, ONE FRESH RUN, ZERO WRITES: delegation 0f7f4b09-77af-47f4-810d-c5489d6c00e9 submitted 02:16:33.643Z through the real onex cloud delegate client inside the lab onex-api pod against http://127.0.0.1:8000 as tenant 6c48114b-cfa9-4125-a390-67a0047a8fcc, tenant key piped on stdin into a 0600 in-pod file and shredded after (readback: file absent). Terminal failed at 02:16:33.848Z, 11s wall submit-to-files, terminal_failure_class=CustomerKeyRefusedError, code ONEX_MARKET_CUSTOMER_PROVIDER_KEY_ABSENT, model none, 0 tokens, terminal_latency_ms 64, projection_row_hash 9e468a65f3a35b, terminal_event_hash 8daf06fd69beb3, receipt.json + run.json written, NO result.txt because there was no content. Auth, submit, poll, terminalise, receipt fetch and file write all worked. CAUSE: the lab tenant holds exactly ONE credential row and it was revoked at 2026-09-12T22:47:50.756703Z by the C7 chain's own phase-2 withdrawal, 37s after the same chain registered it at 22:47:13.207436Z. That is required, not broken -- the chain's error leg (a) asserts a keyless tenant is refused with a typed error, so it must leave the lane keyless. I DID NOT REGISTER A KEY: doing so recreates the stale row that consent row :7025 had to be obtained to revoke, and withdrawing it again is itself approval-gated by that precedent. The OMN-18265 retry did NOT fire and is NOT proven on this lane: a typed key-absent refusal is not a retryable class. | (2) LOCAL PATH WORKS AND DID REAL WORK: three real tasks from this session delegated via $OMNI_HOME/omnibase_infra/scripts/onex delegate, all tier=local backend=local-heavy-reasoning model=Qwen3.6-35B-A3B served from the lab host's own inference endpoint, 1 attempt each, 0 escalations, acceptance quality_bar_met, cost_usd 0.0. run b0b21a44 topic-provisioner ticket body 25s wall / 942 in / 3715 out / saving $0.2928; run 7c857524 standup over 24 TERMINAL rows 41s / 7448 in / 7052 out / saving $0.6406; run 41b34ae3 friction-39 ticket body 15s / 244 in / 2107 out / saving $0.1617. | (3) THE FINDING THAT MATTERS, REPRODUCED 3 OF 3: every response leaks the model's reasoning trace -- a stray closing think tag with NO opening tag, usable answer only after it. Leaked prefix 10556/13404 (79%), 15545/18111 (86%), 8297/9660 (86%). Two of the three prompts said in plain words not to show reasoning and the model emitted a self-check claiming compliance INSIDE the trace, so this is a response-boundary defect, not a prompt defect. THE QUALITY GATE SCORED ALL THREE 1.0 -- its fallback checks are refusal/empty/length and none sees an 86% preamble, so quality_gate_passed cannot today be the acceptance signal for unattended work. Substance itself was accurate: every identifier grounded in the input, nothing invented, the standup grouped 24 rows into correct themes with right ticket and PR numbers. | (4) OTHER BLOCKERS WITH EVIDENCE: the shared lane-rules-block mentions delegation ZERO times and 1 of 6 lane briefs references the skill, so nothing routes lane work to it; a second onex build at ~/.local/bin shadows the canonical wrapper and its drift guard refuses dispatch outright (ONEX_ALLOW_OMNIMARKET_DRIFT=1 exists but self-declares results are not evidence, not used); the lab onex-api pod vendors omnimarket 0.4.20 while the runtime family on the same lane is 0.4.71 and nothing checks it (behavioural impact UNVERIFIED, not claimed); no paid fallback registered for the lab tenant. | (5) ARTIFACTS: three dogfood outputs verbatim with receipt tables at docs/drafts/2026-09-13-dogfood-ticket-body-topic-provisioner.md, -standup-terminal-rows.md, -friction-39-ticket-body.md (uncommitted, local). PLAN: knowledge-base-internal PR https://github.com/OmniNode-ai/knowledge-base-internal/pull/386, beta/plans/2026-09-13-lab-delegation-dogfooding-plan.md, scrub clean, 7 checks pass 0 fail, NOT MERGED per brief. Needs an operator ruling on the lab tenant key posture: a second lab tenant holding a standing registration the chain never grades, or a register-and-withdraw wrapper any lane may run for one task. | NO Slack, NO Linear writes, NO ticket mint, NO credential lifecycle change, no key value printed anywhere.
2026-09-13T02:37:02Z | TERMINAL | lane=omn17426-savings-views-bus-backed | ticket=OMN-17426,OMN-18264 | OMN-17426 COMPLETE AND LIVE. OMN-18264 COMPLETE. LEG 8 NOT GRADED, blocked outside both. OMN-17426: omnimarket#2496+#2495 (squash 1be6c8a4) and omnibase_infra#3459 merged; both arrival-page topics are bus_backed + tenant-scoped and now answer HTTP 200 with backing=bus from inside the onex-dev projection-api pod where they answered 503 not_yet_bus_backed before; projection-api converged to ONE pod ready=true restarts=0, old ReplicaSet gone. OMN-18264 (deploy-time topic provisioner, minted under the operator ruling): FIVE merged omninode_infra PRs - #1395 the Job + pre-rollout step, #1399 sizing + per-pod log capture, #1402 the two managed-Kafka egress rosters, #1403 applying those policies before the Job that needs them, #1404 removing the Job from the overlay because a Job pod template is immutable. Live proof: the provisioner step runs SUCCESS on staging and onex.snapshot.projection.delegation.savings.v1 now probes present=True with a known-absent control still present=False. THE ONE GRADED CHAIN RUN 34733014459 FAILED UPSTREAM OF ALL OF IT: step cli_delegation exit 1 with ONEX_CORE_164_QUOTA_EXCEEDED, the gateway refusing delegation 97244f31-2aef-47e6-bd59-a713c841f589 with 429 'rate limit exceeded'; receipt not_attempted; so leg 8 records render_unavailable - 'the pass produced no receipt, so there is no run for the dashboard to render', leg 8's subject coming from leg 7. The run used the published client 0.4.69, which already carries OMN-18222's 429-is-back-pressure fix, so the 429 here is on the submit path rather than the status poll. Not my change and not fixable in this lane.
2026-09-13T02:47:58Z | TERMINAL | lane=dogfood-findings-tickets | minted 5/5 tickets from lab-delegation-dogfood findings: OMN-18278 (reasoning-trace leak, parent OMN-18232), OMN-18279 (delegation not routed, parent OMN-18232), OMN-18280 (onex shadow binary, parent OMN-18232 - brief gave no parent, added one to satisfy the Linear admission gate OMN-17942), OMN-18281 (omnimarket version skew onex-api vs runtime family, parent OMN-18168), OMN-18282 (lab tenant paid fallback / customer-key dogfooding blocked, parent OMN-18185, awaiting operator ruling) | source: knowledge-base-internal#386, docs/drafts/2026-09-13-dogfood-*.md
2026-09-13T02:51:09Z | TERMINAL | lane=delegate-setup-guide | OMN-18185 | PR https://github.com/OmniNode-ai/knowledge-base-internal/pull/388 OPEN, checks 7/7 green, not merged | guides/delegate-skill-setup.md added, CLI+overlay+plugin setup, dated blockers cited from PR #386
2026-09-13T03:44:42Z | TERMINAL | lane=leg7-submit-429 | ticket=OMN-18283 | parent=OMN-18185 (M2 leg 7) | closes-CLAIM=docs/tracking/ROLLING_WORK_LEDGER.md:7128 | FIXED AND MERGED; LEG 7 IS GREEN FOR THE FIRST TIME. | THE BRIEF'S PREMISE WAS WRONG AND THE EVIDENCE SAYS SO: the 429 was NOT on the submit path. onex-api access log for delegation 97244f31-2aef-47e6-bd59-a713c841f589 shows the submit accepted, 21 GET /v1/workflows/97244f31/status answered 200, then a 429 on the next one, then 429 on both cleanup DELETEs. It is the same poll-path defect OMN-18222 fixed, running unfixed. | THE CAUSE, AN OFF-BY-ONE RELEASE PIN: omniweb pinned OMNIMARKET_VERSION 0.4.69 to GET the OMN-18222 fix. v0.4.69 was tagged 2026-09-12T17:30:48Z; omnimarket#2499 squash 3c314dbb merged 17:53:17Z, 22 minutes LATER. git tag --contains 3c314dbb names v0.4.70 as the first release carrying it; v0.4.69 has zero occurrences of CloudDelegationThrottledError and still polls on a flat count-based loop at 2.0s default. The pin raised to fix the defect installed the last release that still had it, and the floor assertion written to prevent exactly this accepted it because its positive control asserted the wrong direction (it asserted the comparison ACCEPTS 0.4.69). | THE METER, NAMED AND MEASURED, NOT ASSUMED: onex-api general per-tenant request limiter on dev-system i-06169517a92b45f86 namespace onex-dev; scope=tenant, metric=counted api_call rows in public.usage_events, limit 30 per 60s SLIDING window from the default discovery plan entitlement, not delegation-specific. Tenant 12ddab19-6133-4572-bdef-a0c3b74ff8f6 recorded EXACTLY 30 counted requests between 02:27:42.357Z and 02:28:31.505Z (49s); rolling 60s max = 30; the 31st was refused. Split: 9 setup calls, 1 submit at 02:27:51.457Z, 20 polls at a flat 2.0s. | THE PLATFORM WAS NOT AT FAULT: public.gateway_workflows for 97244f31 = status completed, completed_at 2026-09-13T02:28:46.128Z, result_content OK, 89 tokens, 53849ms, on the customer's own credential. The client gave up at 02:28:40Z, SIX SECONDS EARLY. Same shape on e63c4385-8633-40d3-bd38-39e884eee30d from the 01:09Z run (completed 01:10:21.528Z, 51998ms). | THE FIX, omniweb#400 squash 6f5f2fbb (OCC companion 9306 squash 0ef53571 merged FIRST): POLL_BACKOFF_CLIENT_FLOOR 0.4.69 -> 0.4.70; the positive control inverted so it asserts the comparison REJECTS 0.4.69 and accepts 0.4.70 and 0.5.0; OMNIMARKET_VERSION -> 0.4.72, the latest published release, chosen because git diff v0.4.70 v0.4.72 -- src/omnimarket/cloud/ src/omnimarket/cli/cli_cloud.py is EMPTY so no other client-side change is folded in; both comments asserting 0.4.69 is the first fixed release corrected with the tag timestamps that settle it. THE RATE LIMIT WAS NOT RAISED. | RED/GREEN: floor at 0.4.70 with the pin still 0.4.69 fails the named assertion (13 passed 1 failed); GREEN after the bump (14 passed). tsc --noEmit clean, pre-commit on the changed files clean, test:workflow-completeness 23/23, test:onex-cli-run-match passed, pre-push hooks ran with no bypass flag. | LAB-FIRST GAP STATED, NOT PAPERED OVER (rule 24d): the onex-lab overlay carries neither omniweb nor Keycloak, so there is no lab surface for a browser walk; the workflow file itself already records this. The graded chain run IS the live readback. | CHAIN DISPATCHED ONCE: omniweb run 34735936052 on 6f5f2fbb, conclusion SUCCESS. All five steps completed: sign_up, mint_dashboard_key, register_provider_key, cli_delegation, receipt. Receipt present, correlation 44530cd3-7949-46c3-8a79-eb76b5ce1000, tenant 151265de-ad0e-4f17-9da2-1f2f919c6a4f, credential_source=customer_key, zero 429, zero ONEX_CORE_164_QUOTA_EXCEEDED. Savings render bound to the same correlation, all flags true, 0.401s. | BAR DISPATCHED ONCE: omninode_infra run 34736120372, conclusion success, verdict BLOCKED. LEG 7 full_customer_pass = PASS (first time). Leg 6 customer_key_intake = PASS. Leg 8 savings_dashboard_render = FAIL, on source_scan only, three pg driver imports in the omniweb tree including lib/dashboard/source/projection-source.ts, which is OMN-17340's scope and is named as a known blocker in OMN-18185's own description; the render facts themselves are all true. Legs 1 and 3 FAIL on staging deploy run 34732307589 concluding failure, another lane's surface. Legs 2, 4, 5 PASS. | TICKET MINTED: exactly one, OMN-18283, Urgent, child of OMN-18185, Gate C9. | ZERO cluster mutation: every cluster read was kubectl get/logs or a read-only DB SELECT via SSM with --comment leg7-submit-429 naming i-06169517a92b45f86; no set/patch/apply, no compose, no --cold, no credential minted rotated or revoked, no quota changed, no secret value printed.
2026-09-13T03:45:55Z | TERMINAL | lane=cloud-ci-first-boot | OMN-18277 DONE (child of OMN-18205, follow-on to OMN-18224) | consent=docs/tracking/ROLLING_WORK_LEDGER.md:6999 | ALL FOUR POST-FIRST-BOOT DEFECTS FIXED, APPLIED AND MEASURED. MERGED: omninode_infra#1406 squash 187777aedf4e0ccd20362079e000389245c89272 at 03:39:27Z; knowledge-base-internal#387 runbook revision 4; OCC companion onex_change_control#9304 squash e18d2f818045056843e5393f989419bf64357e9c MERGED FIRST (companion-first honoured; the product PR's occ-preflight eligibility check had failed at 02:28Z because the autobind stamped the body after it ran, and CI Summary with it -- both re-run after the companion landed, both green; branch was BEHIND a strict-mode dev and was updated, full suite re-ran clean). APPLIED on merged dev: 0 add / 2 change / 0 destroy (the S3 bootstrap object and the launch template), plan clean afterwards. AC2 THE MEASUREMENT, from the ASG scaling activity and describe-instances, not inferred: BEFORE i-07dacc73186c1217f terminate requested 02:00:57Z, activity ended 02:06:40Z = 5m43s teardown, 9m25s instance lifetime for a 5-second job, Terminating:Wait sustained across nine polls. AFTER i-0e55f2277532919b3 terminate requested 03:43:06Z, activity ended 03:43:48Z = 42s teardown, 2m06s lifetime, Terminating:Proceed observed at +11s and Terminating:Wait NEVER observed. Canary run 34736156036 all six steps green. THE CONTROL PAIR is what makes it conclusive: two earlier instances on the same group with the same 300s hook both sat in Terminating:Wait for the full heartbeat; the only change is the fix. AC1 THE FIX AT CAUSE: the hook fires on EVERY termination including a self-initiated one, and both shutdown paths read the lifecycle state ONCE BEFORE calling terminate, saw InService, and exited -- so the instance entered Terminating:Wait with nothing left to release it. release_terminating_hook now runs from BOTH branches of terminate_self, plus the equivalent bounded loop in the launch template's first-stage fatal handler, because drill instance i-0e4e8e2355c774274 never reached the boot script and showed the same wait; a boot-script-only change would have left half the bug and looked complete. Bounded 60x2s because it runs in the EXIT trap; fallback is today's behaviour, not a hang. THE HOOK WAS KEPT -- deleting it would remove the wait, pass every new assertion, and take the ASG-scale-in deregistration window with it; a test guards that shortcut specifically. AC3 the canary's label line no longer asserts on a runner-context property that does not exist. AC4 runbook rev 4: 7b step 2 waits for status==online (the runner appeared at 01:58:32Z reading offline while its own log shows it listening at 01:58:26Z, and the API did not say online until 02:00:40Z), and 7a step 0 selects which pass is available instead of gating the section; revision 3's 'nothing here has run' paragraph left standing with rev 4 answering it. 7a SECOND PASS PASSED (run on the shipped code, before the fix): bootstrap object renamed so the first-stage fetch failed, RESTORED IMMEDIATELY and proven byte-identical sha256 008162c0757496e459219578b56e28bf41df02ba4e00be725a12460c52807ccd and ETag 74d3e84be117b36b1979c2e7917fe2ba before and after with cmp clean; i-0e4e8e2355c774274 wrote a single-event boot-failed in user-data at line 55, DECREMENTED THE GROUP ITSELF 29 seconds after launch, terminated; start to zero 5m43s inside the ten-minute deadline -- revision 3's F1 proven at runtime. VERIFICATION: 7 new tests in tests/aws/test_ci_runner_lifecycle_hook_release.py covering both paths, with THREE MUTATION CONTROLS each turning exactly one test red (remove the release call; rename away the hook resource; restore the empty label expression), tree green after each. The template was RENDERED WITH TERRAFORM and the output linted rather than read, which caught two defects reading could not: an arithmetic expansion written with a doubled dollar that would have reached the script as the shell's PID, and a comment of mine carrying a literal interpolation opener that made templatefile fail outright while terraform validate still reported success. Deployed-artifact readback: the S3 object carries the fix, zero unrendered escapes, hook name resolved, bash -n clean; launch template version 4 carries the first-stage fix. shellcheck clean with a positive control. AC5 CAPACITY never above 1, read back 0/0/4 zero instances no suspended processes before and after every step with a positive control returning 3 unfiltered each time; both guard alarms OK throughout. NEW RESIDUAL: the release action is INVISIBLE in the forwarded boot log, because cleanup() calls forward_log before terminate_self, so every line the release emits ships after the log has gone -- the proof of this fix is the timing and the lifecycle-state observation, never a log line. CARRIED FORWARD, NOT CLOSED: there is still NO SCALE-UP TRIGGER so nothing routes to this fleet; that plus the docker-seam canary is the next step. The first-pass 7a drill (a boot that reaches the broker and finds no credential) stays unrunnable now the credential exists, and the runbook says so rather than implying full coverage. NOT TOUCHED: runner group 4 membership and settings, every Actions routing variable, the ASG max, and every credential -- no value read, no create, no rotation, no revoke.
2026-09-13T04:23:03Z | TERMINAL | lane=friction-p3-ci-evidence | tickets=OMN-18247,OMN-18249,OMN-18251,OMN-18253,OMN-18254 | closes-CLAIM=docs/tracking/ROLLING_WORK_LEDGER.md:7084 | parent=OMN-18232 | PHASE 3 COMPLETE. All five Done, all five merged to omnibase_infra dev, verified live on origin/dev at 4deaf1b0f. | PRs and squashes: OMN-18247 #3466 b080d7eb079ad6b85147aec3c15ac96394e287b8; OMN-18253 #3470 2fa528787e301278db8d489b83ba7b96bd683efb; OMN-18249 #3473 fca1a2a3eb05e9b0dda8cda3ce5677f4ee711168; OMN-18251 #3475 963e98503abdcb7acb0ed0ec55605f538fb5c38e; OMN-18254 #3476 4deaf1b0f23a4f08206104e495e2a5b98943f30d. OCC companions 9273, 9282, 9292, 9300, 9303, 9308 all MERGED. | LIVE DEV READBACK just now: lab-load-probe carries NO continue-on-error; its uploader is if-no-files-found: error; config/ci_evidence_policy.yaml declares 6 evidence artifacts; pr-ci-zombie-detector.yml carries both jobs. | MERGE ORDER DEVIATED FROM THE TICKET TEXT, deliberately and stated in every PR body: 18247 -> 18253 -> 18249 -> 18251 -> 18254, not 18247 -> 18249 -> 18251 -> 18253. A gate cannot land green against a tree it is designed to reject, and the suppression gate rejects lab-load-probe at its pre-repair head. The red-at-current-head property both tickets require is preserved as a committed byte capture at the exact pre-repair sha b080d7eb, which OMN-18249 uses as its fixture. | THREE PEER COMMITS on the 18247 branch were disclosed in the PR body and the ledger, not absorbed: ee6aa6ab6 and d08095ae9 arrived from a concurrent actor in my worktree; I kept the parts inside OMN-18247's scope and removed in 199a0bc10 the parts belonging to 18249 and 18253 so those tickets were not left hollow. | 4 incident replays registered from REAL bytes, each with a mandatory discriminator: a zero-byte lab-load artifact from a SUCCESS run, the probe job log carrying ModuleNotFoundError under a SUCCESS run, and four omniweb check-run responses. | RESIDUALS, none blocking: (1) the boot-gate diagnostics bundle is deliberately undeclared and out of scope, pinned by a membership test; (2) the alerter reads DECLARED enforcement, so a context enforced by a mechanism nobody wrote down still reads as non-required - the fix is to write it down, and a drift test keeps the declared list honest against ci_summary_gate; (3) all gates see committed workflow files only - a step swallowing its own error inside a script is invisible and is phase 2's class. | NO ticket minted. NO Slack. NO runtime lane mutated; the only live contact was a read-only probe of the org runner registry and read-only GitHub API reads.
2026-09-13T04:50:24Z | TERMINAL | lane=friction-p2-graders | closes-CLAIM=docs/tracking/ROLLING_WORK_LEDGER.md:7117 | tickets=OMN-18240,OMN-18242,OMN-18244,OMN-18245 ALL DONE | parent=OMN-18232 | PHASE 2 COMPLETE, 4 of 4 DONE, 8 PRs MERGED (4 product + 4 OCC companions, companion-first every time). OMN-18240 omninode_infra#1405 squash 069b22656aa8 + OCC#9302 squash 59e249e4d257. OMN-18242 #1407 squash 9cc0688334f8 + OCC#9305 squash c3de394af85e. OMN-18244 #1408 squash 2fc2e79ec0d8 + OCC#9307. OMN-18245 #1409 squash 0590a6660e1e + OCC#9310. No bypass flag, no skip token, no --no-verify, no hooksPath override, no admin merge. | WHAT LANDED: scripts/ci/evidence_reader.py owns per graded fact which surface defines it and how it is read, with four statuses MEASURED/ABSENT/NOT_MEASURED/REFUSED - the collapse of NOT_MEASURED into ABSENT is the whole leg-7 defect. Producer version floors are MEASURED from omnimarket's own release tags, not asserted: credential_source first declared v0.4.62, status/correlation_id/tenant_id/terminal_model_used v0.4.11, terminal_failure_code v0.4.18. A floor is consulted ONLY on an absence, which is why the port changed zero verdicts. scripts/ci/check_graded_fact_reads.py is the forbidding gate, wired as a ci.yml job registered in STRICT_GATE_JOBS plus a pre-commit hook in the same PR, with no allowlist file, skip flag, annotation or env var and two tests asserting that over the parser and the parsed module. | FOUR FALSE VERDICTS CORRECTED, EACH WITH A BEFORE/AFTER TABLE IN ITS PR BODY: leg 8 on run 34691125889 UNPROVEN->FAIL (a caught Playwright strict-mode violation was rewritten as a missing surface); leg 7 with a step recorded completed whose own evidence is an exception PASS->FAIL; leg 8 PASS->UNPROVEN twice for an empty driver list from a source scan that walked zero files or reported no file count. 18240 changed 0 of 10 cases, 18242 changed 2 of 8, 18244 changed 2 of 6, 18245 changed none. | THREE MEASUREMENTS THAT CHANGED THE DESIGN, none of which review would have caught: (1) the first exception matcher required a qualified prefix so it could not match a bare 'Error:' - it scored ZERO on all 75 real strings and read exactly like a clean result, caught only by positive controls; (2) applying the exception rule to every evidence string fails the PASSING control run twice, because item 7's negative half exists to observe a typed refusal and its evidence quotes it, so the refusal is declared per registered fact and the negative-no-key observations are deliberately unregistered with a named control test; (3) the obvious forbidding rule -- no grader may name a graded fact's key -- produced 27 findings across six graders, almost all correlation_id and tenant_id in SQL predicates and records being constructed, so the rule matches only a read at the fact's OWN path. | THE GATE'S ZERO IS NOT VACUOUS: its positive control is the real pre-port evaluator from dev at 5e2f27e6, committed as a fixture, in which the check finds EIGHT violations including the credential_source read that produced the leg-7 defect. | FIXTURES ARE REAL PUBLISHED BYTES, not synthetic: omniweb runs 34691125889 (the incident) and 34715866868 (the control), artifacts customer-pass-latest and savings-dashboard-render-latest; step_evidence dropped at capture. | C19 TRANSPORT: probe_tenant_attribution.py ships as one gzipped file over SSM and now ships two, the reader beside it; measured 79,876 bytes of 131,072 before changing anything and the existing size test now asserts the SUM. The probe's import has two arms, same file under two layouts, and dies if neither resolves - no degraded mode. | SIX EXISTING TESTS MOVED ONTO THE REPAIRED SHAPE, none deleted and no requirement dropped: five in test_evaluate_staging_green_bar.py whose all-green fixture carried the uncontrolled zero this ticket removes, and one collector refusal narrowed to fire only when a record is silent on both halves, each with a new sibling asserting the added requirement. | NO RUNTIME SURFACE: CI script, workflow-transport, hook-config, test and fixture changes only. The bar was never run, no chain was dispatched, no lane, cluster, database or Actions routing variable was touched, and no credential was read, minted, rotated or revoked. | NO TICKETS MINTED (lane is barred); two need minting by the orchestrator, both reported.
2026-09-13T04:55:17Z | TERMINAL | lane=grader-followup-tickets | parent=OMN-18232 | Minted OMN-18285 (High, omniweb: C9 customer-pass records need a producer_versions block for the evidence-reader floor) and OMN-18286 (Medium, omninode_infra: k8s/onex-lab/provider_key_chain.sh status-endpoint reads at lines 367,552 outside the Python forbidding gate). Both bound Gate: OMN-18232 AC-2. Both cited live: scripts/ci/evidence_reader.py + check_graded_fact_reads.py at omninode_infra dev HEAD 0590a666, omniweb#400 (OMN-18283) pin history 0.4.69->0.4.70->0.4.72.
2026-09-13T06:47:42Z | TERMINAL | lane=omn17340-pg-imports | ticket=OMN-17340 related=OMN-17341,OMN-18185 | closes-CLAIM=docs/tracking/ROLLING_WORK_LEDGER.md:7136 | DASHBOARD HALF LANDED AND PROVEN; LEG 8 STILL FAILS ON OMN-17341's TWO SURFACES; OMN-17340 LEFT In Progress, NOT flipped. MERGED omniweb#401 squash 61a97a77e2c2c60e72915e1f3d7328c63d44f6be (ancestor of origin/dev verified); OCC companion onex_change_control#9309 merged FIRST 04:39Z sha 902466c8f52c19e9a8ac6acf71fa136f81a00a1b, autobind-minted (autobind has recovered since OMN-18073). REMOVED lib/dashboard/source/projection-source.ts (raw pg Pool) + projection-sql.ts + two test files; selection seam drops the 'projection' mode (fixtures/api/projection-api remain). SAFE TO DELETE ON LIVE EVIDENCE, not preference: OMN-17426 made the savings topics bus-backed and the projection-api arm is proven rendering a real run, so the arm's only retention reason (sole live-data arm) is discharged. TDD: new doctrine check 'no database driver is imported under lib/dashboard/' using the bar's own _DATABASE_DRIVER_PATTERNS plus a positive control; RED 7pass/1fail, GREEN 8pass/0fail; inside the aggregate test script CI runs, so enforced. Neighbouring check tightened from a one-entry allowlist to none. Local: tsc clean, 50/50 test scripts green in 32s, topic lint clean, production build succeeds. LIVE READBACK, scheduled bar run 34743448711 (2026-09-13T06:43:28Z, schedule event, verdict BLOCKED): leg 8 savings_dashboard_render FAIL. RENDER HALF ALL TRUE - correlation 44530cd3-7949-46c3-8a79-eb76b5ce1000, surface omniweb:/app, tenant 151265de-ad0e-4f17-9da2-1f2f919c6a4f, 0.401s, savings/model/tier rendered, transport direct_database_read=False, projection_crosscheck savings_matches_render=True. SCAN HALF FAILS on 2 findings, down from 3, files_scanned 397 down from 401 - the bar picked up the merge. REMAINING, both OMN-17341 (Backlog, Medium, whose own sequencing note asked for this PR first): app/_lib/db.ts:pg (admin query layer over authoritative OLTP tables tenants/usage_events/waitlist_entries/admin_events_log - NO projection topic exposes any of them) and app/(marketing)/waitlist/actions.ts:pg (a WRITE path; a projection read surface structurally cannot replace it, and its existing event onex.evt.omniweb.waitlist-signup.v1 is domain-only by privacy design so it cannot carry the signup). Converting either would mean inventing endpoints; stopped short per brief. pg stays in package.json for the same reason. NOT DONE per brief condition 'Done only if leg 8 passes'. Zero tickets minted, no cloud mutation, no credential work, no bar dispatched. SIDE FINDING, not acted on: OMNIDASH_ANALYTICS_DB_URL is still plumbed onto the omniweb container from a secret in k8s/onex-dev-overlay-dev/patches/omniweb-deployment.yaml and now has no reader in omniweb - a least-privilege cleanup for omninode_infra, out of this lane's repo scope.
2026-09-13T06:53:21Z | TERMINAL | lane=deploy-freshness-gate | actor=claude:opus5:subagent | ticket=OMN-18284 (minted, Urgent, child of OMN-18185) parent-criterion=C2 | closes-CLAIM=docs/tracking/ROLLING_WORK_LEDGER.md:7137 | STAGING DEPLOY GREEN AND M2 IS GREEN ON THE SCHEDULED BAR. CAUSE: the plane's OMN-15796 provenance annotation had gone stale against the images it describes. deploy-onex-staging.yml moves the images in the overlay apply and writes the annotation in a later, dispatch-only step; five dispatch deploys between 16:56Z and 23:53Z on 2026-09-12 (34706722423, 34720532616, 34722876820, 34724526297, 34726594364) failed at or after the apply, the last on the OMN-18264 provisioner Job immutability error AFTER every Deployment was already server-side-applied, so images advanced to sha256:ec2254071f12 (built from omnibase_infra d2cba472) while the annotation stayed at ce70bd1e from the last green deploy 34704568627. Push deploy 34732307589 then fell back to that stale annotation as the announced identity and refused all 13 runtime workloads with image_not_announced_candidate; every other step in it, apply and pin and OMN-15009 included, was green. FIX: no code change; one sanctioned omnibase_infra deliver-dev-candidate-to-staging workflow_dispatch (sibling_ref=dev), delivery run 34736491570 on dev head 963e9850, which dispatched deploy run 34737080968 (omninode_infra head 187777ae) — CONCLUDED SUCCESS with ZERO failed steps, stamp step ran. READBACK on i-06169517a92b45f86 ns onex-dev --comment deploy-freshness-gate: one image sha256:0ac360a8b2a32ee0a4f0a5e898466a8a55b618adca4a84d0a5266210d565ef31 and one stamp 963e98503abdcb7acb0ed0ec55605f538fb5c38e, in agreement; all 13 replica-carrying runtime Deployments Ready, 4 at replicas=0 by declaration. SCHEDULED BAR run 34743448711 (06:43Z, r10): legs 1-7 PASS, leg 8 FAIL on the omniweb pg-driver source scan only (OMN-17340, another lane). M2 GREEN 5/5 fail=0 unproven=0; M4 BLOCKED 2 pass 1 fail. RESIDUAL: the repair is operational and the defect recurs on the next dispatch that fails after the apply — OMN-18284 carries the RED test and the three fix options.
2026-09-13T07:08:32Z | TERMINAL | lane=friction-p5 | actor=claude:opus5:subagent | closes-CLAIM=docs/tracking/ROLLING_WORK_LEDGER.md:7142 | tickets=OMN-18258,OMN-18259,OMN-18260,OMN-18261 (children of OMN-18232) | ALL FOUR CLAIMED PHASE-5 ITEMS DONE AND MERGED. ZERO TICKETS MINTED: the four-ticket grant went unused because lane mint-friction-plan already minted every phase-5 child at :7068. LANDED, companions first where the repo has one: OMN-18258 omni_home#265 squash 4dd53b1bb52c189d6cb46b3833a4041d96269928 (no companion; omni_home has no receipt gate or autobind path); OMN-18259 knowledge-base-internal#389 squash 542474658f20665e6b353c71c48eabfc26906e65; OMN-18260 omniclaude#2141 squash b4be5af0627b33b9d862bb4efd8b8ee8ceace343 with OCC#9312 squash a2de93c7961f7e510391d3318ab6512264bc152e merged first, plus follow-up omniclaude#2142 squash e6dbbbbb34a25b9c98cfb51a284c9121ad574681 with OCC#9313 squash 62962ff2d9fb7f4a873ff9c2d3ea6f40b9a7e30a merged first; OMN-18261 omni_home#266 squash b664d79b5cc7c4031e726f4f04f6d5db5624b4a9. NOT DONE AND NOT CLAIMED: OMN-18262 (pre-push refusal) and OMN-18263 (CI check mirroring it) are UNSTARTED and both wait on an OPERATOR RULING that is open question 1 of the plan -- how hard the branch-claim refusal should be, i.e. whether the push-refusing hook lands alongside the pull-request check or waits until the release path has been exercised in practice. The design gives both everything else they need; only the sequencing between them is unresolved, and the plan says that is not a lane call. The worktree lease (plan item 27) is unminted and is explicitly outside this phase by the plan's own text. TWO RESIDUALS I DELIBERATELY DID NOT TAKE: the lane-identity hook is NOT installed across the canonical clones, because a commit-refusing hook installed while worktrees are unregistered refuses peer lanes mid-flight; and Lane Identity Gate is not in required_status_checks on omniclaude dev, so it reports but cannot block -- a branch-protection change outside these PRs. ONE INCIDENT, ALREADY CLOSED: a test run under a leaked GIT_DIR installed that hook into the SHARED canonical-clone hooks directory and every canonical clone refused commits from unregistered worktrees for about ten minutes; found when my own commit was refused, removed by hand, verified absent from every clone and from that shared directory, and closed at root by two mechanisms with tests (the module no longer inherits git environment variables; the installer refuses any hooks directory outside the repo's own git directory). ONE GATE MERGED RED AND IS NOW GREEN: Lane Identity Gate failed on b4be5af0 because its job installed pytest alone while this repo's conftest imports project dependencies, and it did not block because it is not a required context; #2142 syncs deps the way every other test job does and the same gate is SUCCESS on e6dbbbbb on dev. NOTHING RUNTIME TOUCHED by any of the four: no container, no cluster, no broker, no database, no credential, no Slack, no Linear state beyond these four tickets. PRE-EXISTING RED NOT MINE AND NOT FIXED: the omni_home Branch Protection Guard 'audit' job has failed on every PR since at least 2026-09-02 on onex_change_control main and dev for 'approving reviews are enforced', where live readback shows required_approving_review_count=0 with require_code_owner_reviews=true -- the auditor's expectation conflicts with the deliberate OCC codeowner-review setting; it is not a required check on omni_home main and I did not mutate another repo's branch protection to silence it.
2026-09-13T07:16:43Z | TERMINAL | lane=p5-residual-tickets | actor=claude:opus5:subagent | closes-CLAIM=docs/tracking/ROLLING_WORK_LEDGER.md:7150 | tickets=OMN-18287,OMN-18288 (children of OMN-18232) | BOTH GRANTED TICKETS MINTED, PROJECT Ready, EACH CARRIES A Gate: LINE. OMN-18287 (High, Gate OMN-18232 AC-3): Branch Protection Guard 'audit' job false-FAILs onex_change_control main+dev. TWO CONFIRMED ROOT CAUSES, not one: (1) omni_home's branch-protection-guard.yml pins ONEX_CHANGE_CONTROL_REF=3d15abea8 (OMN-17186, 2026-08-30), which predates onex_change_control#7994 (OMN-17491, squash 55cb6ccc55f0f575e2e7117ef8f2f6fc8ec766fd, merged 2026-09-01) that added the REVIEW_GATED_MAIN_REPOS=(onex_change_control) carve-out -- git merge-base --is-ancestor confirmed the ordering; (2) even current onex_change_control main only applies that carve-out when branch==main, never dev, and OCC dev carries the identical required_approving_review_count=0/require_code_owner_reviews=true model as main, live-confirmed via gh api, so bumping the pin alone leaves dev red forever. Live job log pulled (run 34741472172, job 103681780034): 'Summary: 2 failure(s) across 102 checks', exit 1, both onex_change_control main and dev FAIL 'approving reviews are enforced (blocks solo-dev merges)'. Red on every omni_home pull_request run from 2026-09-01T12:10:28Z through 2026-09-13T05:55:10Z (20 runs), one green window being the PR that landed the stale pin itself; not a required check on omni_home main. OMN-18288 (Medium, Gate OMN-18232 AC-5): Lane Identity Gate (OMN-18260, omniclaude#2141 b4be5af0/#2142 e6dbbbbb) reports SUCCESS on dev HEAD e6dbbbbb (live check-run confirmed) but is absent from required_status_checks (live-confirmed absent), and its pre-push hook (scripts/lane_identity.py, scripts/hooks/prepare-commit-msg-lane, confirmed present via gh search code) is not installed across canonical clones -- friction-p5's own TERMINAL at :7153 named both residuals deliberately, citing the 2026-09-13 leaked-GIT_DIR incident as the reason a fleet-wide commit-refusing install needs a worktree-registration sweep first. AC set (a) required-check wiring via the OCC audit path with the audit's own expectation updated same-change, (b) registration sweep before fleet install, (c) per-clone readback proving zero refusals, (d) confirm/extend the hooks-directory-outside-repo refusal guard. NOTHING I COULD NOT VERIFY: every claim in both tickets is backed by a live gh api/gh search/git command run this session; no speculation. No code, no runtime, no credentials, no Slack, no Linear state beyond these two new tickets and the required patch fix to OMN-18288's Related section (nested empty issue tag, cosmetic).
2026-09-13T08:23:13Z | TERMINAL | lane=m2-board-readback | ticket=OMN-18185 (dispatch-cited; ticket is actually the unrelated C9 producing-harness effort, see note) | M2 FLIP VERIFIED LIVE, PR OPEN AUTO-MERGE ARMED, BLOCKED ON CI QUEUE. Downloaded and read staging-green-report.json from run 34743448711 (schedule, r10, collected_at 2026-09-13T06:43:57Z) in full: milestone_verdicts.M2 GREEN, 5/5 legs PASS. Cross-checked deploy run 34737080968 (success, zero failing jobs, head 187777ae) and C3/C4 probe latest runs (34714129255, 34743467238, both success) via gh; all 19 M2 carrier tickets confirmed Done via Linear. Edited tools/beta_board/beta_board_data.py: PROBE_DEFECTS['C2'] state RED->GREEN with fresh per-leg readback (workflow exit-code defect stays documented, permanent); prepended READBACK paragraphs to MILESTONES M2 narrative and EVIDENCE['C1']/EVIDENCE['C3']. 53 beta-board tests pass, pre-commit green. Ran tools/beta_board/beta_board.py against LIVE Linear+GitHub: hero KPI read 2/5 Milestones met, Met: M1, M2; all four C1-C4 chips s-pass in the M2 row (table's stricter cleared/MET badge does not fire because C1's stability half has a recent flip-flop -- documented, does not reopen M2 per the board's own state-vs-cleared distinction). PR omninode_infra#1410 opened against dev, auto-merge armed (native, REDACTED-OPERATOR). OCC autobind minted companion onex_change_control#9315, citing Evidence-Source: OCC#9315, patched onto #1410. REAL DEFECT FOUND AND FIXED: ticket OMN-18185 (the ticket this dispatch cited, actually titled 'C9 -- the producing harness for staging-green-bar legs 7 and 8', unrelated to this M2 board work) has an unlabelled 1.-7. acceptance-criteria list; the OMN-18236 Acceptance-Criterion Binding Gate refused #9315 with ac_binding_ticket_unlabelled because nothing in the autobind-generated contract could pin to an unlabelled criterion. Relabelled the ticket's existing list to AC1.-AC7. (no wording change) via Linear save_issue, noted the fix and the ticket-citation mismatch in an editorial comment on the ticket itself; pushed an empty retrigger commit to the companion branch (auto/omninode-ai-omninode_infra-pr-1410-occ-autobind) since GitHub does not re-run a completed check on a PR-body-only Linear-independent fix. The binding gate then passed. RESIDUAL, NOT RESOLVED BY THIS LANE: both #1410 and #9315 remain OPEN, blocked purely on GitHub Actions throughput -- #9315's Tests job has been in_progress upward of 55 minutes (GitHub status page shows no incident), and #1410's own OCC Companion Merged Gate failed once already on a timeout waiting for #9315 and will need its own retrigger once #9315 merges. Auto-merge is armed on both, so they complete unattended once CI clears; no further code or data change is needed. FLAG TO ORCHESTRATOR: the dispatch brief's ticket number (OMN-18185) does not describe this task -- it is a large, unrelated, in-progress C9 ticket owned by the same assignee. This lane's evidence is now permanently attached to that ticket's PR list via the autobind mechanism, which cannot be undone without deleting the companion; a correct ticket number for the M2 board-readback work should be resolved by the orchestrator (mint or identify the right OMN id) for any follow-up.
2026-09-13T08:26:33Z | TERMINAL | lane=omn18284-stale-stamp | ticket=OMN-18284 parent=OMN-18185 gate=C2 | closes-CLAIM=docs/tracking/ROLLING_WORK_LEDGER.md:7152 | STALE PROVENANCE STAMP FIXED AT CAUSE AND PROVEN ON A LIVE PUSH DEPLOY. omninode_infra#1412 squash d18a6d450ae1007669be2d7ee6243155e40f4f46 merged 08:10:50Z, OCC companion onex_change_control#9317 merged first, 66 checks pass 0 fail. OPTION: the ticket's preferred option 1 (attribution) PLUS option 2 (reconcile) — option 1 alone leaves a correctly-attributed stale stamp still red on every push deploy, which is the state the operational repair had to be hand-triggered out of. The ~30 lines of jq that promoted the live annotation into --announced-source-sha are deleted from the gate step; the comparison moved into check_runtime_image_freshness.py, which holds both halves and resolves direction from omnibase_infra history (compare/<stamp>...<image ref>): stale_provenance_stamp (annotation lagged, this ticket), plane_behind_provenance_stamp (plane reverted, OMN-18055 preserved), unresolved_provenance_stamp (diverged/unreadable, fail-closed), ambiguous_provenance_stamp (split plane, moved out of shell and given tests). New scripts/reconcile_runtime_provenance_stamp.py runs before the gate, after OMN-15009, on every run shape with no if:, and refreshes a lagging annotation from the image's own baked infra_vcs_ref in that direction only. RED/GREEN: 24+12+10+11 new/rewritten tests, 62 existing gate tests unchanged and green; ruff+mypy clean. LIVE READ-ONLY PROOF on i-06169517a92b45f86 ns onex-dev --comment omn18284-stale-stamp: real readback (17 runtime Deployments, stamp 963e9850, image sha256:0ac360a8, baked infra_vcs_ref 963e9850) run through three legs with real GitHub compares — as-is no finding; annotation rewound to ce70bd1e gives stamp_behind + stale_provenance_stamp on 13 and a decided repair of 13; annotation set to dev head gives stamp_ahead + plane_behind_provenance_stamp on 13 and no repair. That live run caught a real defect in the change (script-dir import made a second module object and a second StampDirection enum, so every is-comparison in the repair silently missed); import fixed, test added. STAGING PROOF: push deploy 34747044328 on the merge sha, CONCLUDED SUCCESS zero failed steps — OMN-15796 stamp step SKIPPED (push run, dispatch-only, the exact condition that produced the red on 34732307589), reconcile step success 'nothing reconciled', freshness gate success 'verified across 17 deployment(s)'. Post-deploy readback with positive control: one stamp 963e9850, one image sha256:0ac360a8, 17 deployments, in agreement. NOT PROVEN: no lab lane exercises this (onex-lab carries no provenance annotation and no freshness gate; no workflow other than deploy-onex-staging.yml references either script) so rule 24 has no surface here; the repair's kubectl annotate write path has not fired live because the plane is in agreement, only its decision logic is proven against the real readback; the staleness window inside a single failed run still exists — the change makes it correctly attributed and self-healing on the next deploy of any shape, not atomic.
2026-09-13T09:10:48Z | TERMINAL | lane=omn18287-bp-audit | actor=claude:sonnet5:subagent | closes-CLAIM=docs/tracking/ROLLING_WORK_LEDGER.md:7218 | ticket=OMN-18287 DONE | onex_change_control#9318 (squash 886264540572fc60cb186e2d1388aff1ddc39d47, merged dev 09:00:20Z) extends REVIEW_GATED_MAIN_REPOS/is_review_gated_main -> REVIEW_GATED_REPOS/is_review_gated, dropping branch==main restriction so the codeowner-review carve-out applies to every audited branch; TDD RED (test_review_gated_occ_dev_requires_approving_and_code_owner_reviews) reproduced the exact dev false-FAIL first, GREEN after fix, 11/11 regression tests, 74 CI checks green; hand-authored contract+self-bound receipts for the in-repo OCC PR structurally unioned with the OCC autobind companion (no duplicate ids, no rebind needed) | omni_home#267 (squash b555a45c386577b8f88d270174c547dfa385009f) bumps ONEX_CHANGE_CONTROL_REF in branch-protection-guard.yml + scheduled-gap-detect.yml to 886264540572fc60cb186e2d1388aff1ddc39d47, adds branch-protection-policy.md section 1a documenting OCC main+dev codeowner model | VERIFIED: omni_home PR 267 audit job run 34748874941 conclusion=success, Summary: 0 failure(s) across 102 checks; local bash scripts/audit_branch_protection.sh from merged OCC worktree also 0 failures/102 checks, [dev] PASS approving and code-owner reviews are enforced (review-gated branch) | Linear OMN-18287 set Done with both PR URLs cited | NOTHING UNVERIFIED
2026-09-13T09:11:47Z | TERMINAL | lane=m2-board-landing | onex_change_control#9315 merged 56de56b3 at 08:25:56Z | omninode_infra#1410 merged 7ce0b41b at 08:46:05Z (required one companion-gate retrigger empty commit + one PR branch-update-to-dev to clear a BEHIND merge-state) | board not yet flipped 25min post-merge: publish-beta-board.yml run 34748643628 (headSha 7ce0b41b) concluded failure on a transient GitHub App installation-token 500, unrelated to PR content; board still reads Met: M1. Next: M2 as of 09:11Z
2026-09-13T09:24:04Z | TERMINAL | lane=draft-pr-sweep | actor=claude:sonnet5 | closes=CLAIM docs/tracking/ROLLING_WORK_LEDGER.md:7187 | outcome=FLIPPED_3_HELD_22. 25 open draft PRs found org-wide (omnibase_core 2, omnibase_infra 3, omnimarket 8, omninode_infra 2, onex_change_control 6, knowledge-base-internal 3, RSD 1; zero on omnibase_spi/omnibase_compat/omniclaude/omnidash/omnimemory/omniintelligence/omniweb/knowledge-base/omni_home). FLIPPED (gh pr ready, isDraft:false readback confirmed, autoMergeRequest:null on all three so none auto-armed -- ready_for_review does not fire the auto-merge.yml opened-event listener, manual arm needed): omnimarket#2464 (mergeStateStatus CLEAN, checks green, Evidence-Source OCC#9009, no hold language), omninode_infra#1310 (mergeStateStatus BEHIND 96 commits not conflicting, checks green, Evidence-Source OCC#8968, author states merge is inert/prod-safe -- deploy is separately gated workflow_dispatch), omninode_infra#1214 (mergeStateStatus BEHIND 170 commits not conflicting, checks green, Evidence-Source OCC#8523, no hold language). HELD 22, each for an explicit blocking reason (explicit author hold language, mergeStateStatus CONFLICTING/DIRTY, stacked non-dev base, red checks, or OCC append-only companion sequencing left to OCC-lane review): omnibase_core#1674 #1654; omnibase_infra#3410 #3247 #2658; omnimarket#2504 #2466 #2453 #2356 #2352 #2351 #2324; onex_change_control#9036 #8377 #8316 #8306 #8094 #7938; knowledge-base-internal#360 #355 #124; RSD#6. Full per-PR reasons in the orchestrator report. No live ledger CLAIM found for any of the 25 tickets in the prior 6h (only read-only triage/classification mentions of OMN-17376/OMN-16992/OMN-17974/OMN-16964/OMN-18086 as PARENT/EXTERNAL labels, not active-write claims). No rebase/update-branch/push/close performed on any draft. Zero tickets created, zero Slack, zero Linear writes, zero merges.
2026-09-13T09:24:50Z | TERMINAL | lane=verify-cred-401-ticket | closes=own CLAIM (this file, prior row) | ticket=OMN-18292 (Medium, parent OMN-18168, project Ready) related=OMN-16633,OMN-16578,OMN-16504,OMN-16421 | OUTCOME: onex-dev/onex-verify-client-credentials and the Keycloak onex-verify client are BOTH HEALTHY -- the 401 was never a credential regression. ROOT CAUSE proven live via same-secret A/B on dev-system i-06169517a92b45f86: identical client_secret/client_id against https://auth.omninode.ai/... -> HTTP 401 unauthorized_client (reproduces OMN-16633's finding exactly); same secret against https://dev.auth.omninode.ai/... (dev-system's own Keycloak ingress host per kubectl get ingress -n auth: hosts dev.auth.omninode.ai,staging.auth.omninode.ai) -> HTTP 200 real token issued scope=profile email; negative control GET dev.api.omninode.ai/v1/whoami no-token -> 401. kcadm GET (no regenerate) confirms onex-verify enabled=true serviceAccountsEnabled=true publicClient=false, and Keycloak-side secret sha256 == k8s Secret-side sha256 (03c409902ba1f82256b215b4e8fded88b742217d845154f1cd2249e6d36c77e5, len=32) -- no drift, nothing rotated, no secret value ever printed. CONSUMER CENSUS: kubectl get deploy,statefulset,cronjob,job -A -o json and get pods -A -o json both grep-count 0 for onex-verify-client-credentials on dev-system -- zero running workloads reference it; no omninode_infra/omnibase_infra CI workflow references onex-verify at all; only consumers are docs/runbooks/2026-08-24-onex-verify-client-credentials.md (hardcodes auth.omninode.ai) plus docs/testing/2026-08-26-staging-test-walkthrough.md and docs/tracking/2026-08-28-e2e-walkthrough-runbook.md which cite it. THIS IS A RECURRENCE of the exact defect class OMN-16504 already diagnosed+fixed for a DIFFERENT client (onex-api/onex-runtime-credentials, omninode_infra#1019, guarded by tests/k8s/test_probe_onex_dev_keycloak_check_targets_dev_issuer.py) -- that guard covers only probe-onex-dev-live-verify.yml, not this omni_home doc runbook, so the mistake recurred undetected. NOT OMN-16578 (different failure mode: SSO-expiry non-interactive retrieval, unrelated). Admission guard OMN-17942 required a Gate: line on create; used 'Gate: live-gate defect: onex-verify-token-mint' (no C-id/INV-id from the pinned PRD fit this doc-hostname defect precisely; flagging to orchestrator for review). NO credential rotation/re-issue/re-sync of any kind performed or recommended -- ticket AC4 states none is warranted. NO cluster mutation (all SSM --comment verify-cred-401-ticket on i-06169517a92b45f86 only; public cluster REDACTED-INSTANCE-ID, prod, stability-test, judge, lakshman untouched). NO SendMessage to any lane.
