FE-1573: Protect long Brunch interviews from abrupt Ledger loss and stalled responses - #9761
Conversation
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
PR SummaryMedium Risk Overview Model context and provider liveness: Brunch context projection now collapses older successful Workpiece contract and UI: Settled revisions ignore typed refused Ops and tests: Persona launcher Stop aborts the active Flue submission, records settlement, then tears down process groups; browser persona flows track continuation admissions. Reviewed by Cursor Bugbot for commit 65bd76e. Bugbot is set up for automated code reviews on this repo. Configure here. |
Consolidate the experiment-configuration planning that was spread across the Mission 7d archive, the future spine and the Mission 11 draft into one draft, `docs/mission-drafts/8-experiment-configuration-from-the-ledger.md`, taking the freed Mission 8 number. The draft records the Petrinaut experiment terrain as read at HEAD, the Ledger-condition to experiment-destination correspondence table, the design assessment (Brunch drafts a session proposal prepared by Petrinaut's own `prepareExperiment`, the user presses Run; a thin document-entity design stays an upstream option), the upstream delta list for Petrinaut, the Brunch-side prerequisite of scenario and metric operations in `mutate_petrinaut_net`, and an implementer's entry for the handoff. Point the spine and the Mission 9, 10 and 11 drafts at the new home, rename the spine's deployment follow-on heading to "Hosted deployment successor" now that the Mission 8 number is reused, and record the read-only dev/prod persistence audit (SQLite locally, Postgres deployed) in the spine and the worked-example draft. Amp-Thread-ID: https://ampcode.com/threads/T-01a0aab3-72bb-71dc-8a77-a984b3617b9c Co-authored-by: Amp <amp@ampcode.com>
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #9761 +/- ##
==========================================
+ Coverage 66.28% 66.31% +0.03%
==========================================
Files 1775 1782 +7
Lines 191861 192337 +476
Branches 7853 7856 +3
==========================================
+ Hits 127171 127552 +381
- Misses 63209 63302 +93
- Partials 1481 1483 +2 Flags with carried forward coverage won't be shown. Click here to find out more. ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
TimDiekmann
left a comment
There was a problem hiding this comment.
Infra-changes looks good, please ping me when code review is done.
🌟 What is the purpose of this PR?
Long Brunch interviews could replace most of their recoverable Ledger in one rewrite, retain every superseded net observation in model context, and remain indefinitely busy when a provider stopped mid-response. This PR protects the existing product path at those boundaries: it refuses abrupt single-revision Ledger loss, reduces superseded net reads only in model context, bounds an idle model invocation with one canonical-tail retry, and settles launcher Stop gracefully.
Correctable Brunch-owned Ledger validation is now an honest non-applied tool result rather than a false terminal failure. Pending Brunch calls render gold, applied calls green, correctable refusals compact neutral with retained detail, and actually thrown failures red. Lu accepted that sequence in the real visible local browser path.
A retained 43-minute Inventory observation showed that these mechanisms enable substantial partial elicitation but do not make whole-body Ledger mutation product-scalable. Cumulative permitted rewrites still contracted the Ledger and its evidence, and the final provenance query could not recover a verified causal chain. That negative product-scale disposition remains final. For ordinary documents, Petrinaut's Stock assistant is now the fallback; Brunch remains a configured hidden Cmd-K alternate, while Brunch-focused deployments and tests can select it for fresh profiles with
VITE_PETRINAUT_DEFAULT_ASSISTANT=brunch.🔗 Related links
🚫 Blocked by
None. Mission 7e is closed as an engineering partial; its product-scale Ledger successor is deliberately separate rather than a blocker for this PR.
🔍 What does this change?
mutate_workpiecerevisions against missing headings and unannounced reductions greater than 25%, after stale-base and replay handling and before ordinal assignment.read_petrinaut_netresults to observation identity and counts while preserving the latest full result and leaving canonical/public history unchanged.🏗️ Agent notes
Ledger boundary
The guard runs at the sole persistent workpiece-state boundary. Same-tool-call/equal-Markdown replay remains idempotent; same ID with different Markdown refuses. Heading extraction covers ATX and Setext headings and ignores fenced-code false positives. Exactly 25% reduction remains permitted. A refusal changes no revision, body, hash, evidence or ordinal.
Retractions persist the withdrawal name, exact authorization quote, authorized true-user message IDs, and unique prior-body excerpts whose non-overlapping length accounts for the destructive net reduction. Source identity or free text alone is insufficient.
Typed refusal and presentation
Expected model-correctable operation validation returns
{ disposition: "refused", applied: false, correctable: true, code, message, currentRevision }. Applied and refused outputs share one discriminated schema. Refused outputs never become settled revision state or cause a false newer-revision warning. A browser-safe predicate validates the complete refusal schema and fails partial lookalikes closed.The host presentation contract supports an optional tone and collapsed detail. Brunch uses it for pending gold, applied green, and compact neutral refusals; thrown failures remain red. Stock retains its existing presentation and hidden automatic-tool behavior. Color is not the sole signal.
Context, provider liveness and Stop
Model-only projection chooses the last structurally valid net read in each independently projected slice. Entry order, IDs, Voice correlation, browser output, freshness identity and canonical Flue history are preserved.
The watchdog treats model events—not sockets, heartbeats or synthetic stream start—as progress. Production limits are 10 seconds generally, 15 seconds only from reasoning start to its first delta/completion, and two seconds for cancellation acknowledgement. Canonical
toolcall_endwins the race. Only the first acknowledged idle in one operation scope carries Flue's retryable-interruption marker.The native OpenAI/browser fixture proves one distinct retry, no duplicate user input or completed
ping, no canonical admission of either speculative tool, terminal failure after repeated idle, and no global-fetch fallback. Launcher Stop tracks initial and browser-continuation admissions, uses public Flue abort, waits up to five seconds for settlement, recordslauncher-stop.json, and terminates owned services before slower browser cleanup.Assistant selection
VITE_PETRINAUT_DEFAULT_ASSISTANTaccepts onlystockorbrunch; unset or blank resolves to Stock and an invalid value fails at startup/build. An explicit storedstockorbrunchchoice always wins, so changing the launch fallback performs no migration. The local Brunch launcher and synthetic persona/compiler paths explicitly request Brunch. The Cmd-K switch remains hidden when no Brunch endpoint is configured and is unavailable on remote worked-model routes, which continue to force Brunch without overwriting the ordinary preference.Final observation and claim boundary
run-1L2zTlis local-only. It retained 25 user submissions, 24 accepted Ledger revisions and a useful partial net over 42m 59.561s. Two abrupt losses refused and repaired. Cached context was about 210k near minute 30, later reached 254,677, and compacted after the manual answer at 258,968 tokens.The accepted body nevertheless contracted from a 33,876-character peak to 10,487 through individually permissible revisions without a retraction; current-revision evidence relations fell from 26 to four. The final explanation disclosed that
query_workpiececould not reconcile the visible place to a verified construction chain. Native records estimate $44.1015 of Brunch usage and $0.1654 of persona usage for this run.This PR establishes narrow per-revision loss prevention, model-only net-read economy, deterministic bounded provider recovery, graceful terminal settlement, honest refusal presentation, and explicit host selection. It does not establish product-scale Ledger preservation, stable provenance across whole-body rewrites, acceptable long-session latency/cost, behavioral correctness of the Inventory model, a verified final explanation, hosted reliability, or Mission 7d semantic acceptance.
Pre-Merge Checklist 🚀
🚢 Has this modified a publishable library?
This PR:
📜 Does this require a change to the docs?
The changes in this PR:
🕸️ Does this require a change to the Turbo Graph?
The changes in this PR:
turbo.jsonfiles and generated task-dependency record have been updatedset-state-in-effectwarning, and older core lint retains its recordedno-await-in-loopwarnings.🐾 Next steps
MISSION.next.md.🛡 What tests cover this?
Passed on the final branch:
@hashintel/brunch-agentcore unit tests and the full@apps/brunch-agentunit suite.test:anthropic-toolswrapper, including both native tool catalogues through Anthropic's token-count endpoint.No paid inference was used for WP-E, Stock-default verification, or rendered acceptance.
❓ How to test this?
yarn workspace @hashintel/brunch-agent test:unitandyarn workspace @apps/brunch-agent test:unit.yarn workspace @apps/brunch-agent test:integration,test:workpiece-evidence,test:passage-policy, andtest:persona.yarn workspace @apps/petrinaut-website test:unitandyarn workspace @apps/petrinaut-website test:integration.VITE_PETRINAUT_DEFAULT_ASSISTANT=brunchin a fresh profile and confirm Brunch mounts; store Stock explicitly and confirm that stored choice still wins.📹 Demo
No portable demo artifact is attached. Lu reviewed and accepted the deterministic rendered WP-E sequence in a visible local Chrome window. The decisive 43-minute Inventory observation is retained locally as
run-1L2zTl; its reviewed negative product-scale result is recorded inMISSION.md.Stack created with GitHub Stacks CLI • Give Feedback 💬