Skip to content

FE-1573: Protect long Brunch interviews from abrupt Ledger loss and stalled responses - #9761

Merged
lunelson merged 32 commits into
mainfrom
ln/fe-1573-mission-7e-express
Sep 17, 2026
Merged

lunelson merged 32 commits into
mainfrom
ln/fe-1573-mission-7e-express

Conversation

@lunelson

@lunelson lunelson commented Sep 16, 2026 •

Copy link
Copy Markdown
Contributor

🌟 What is the purpose of this PR?

Long Brunch interviews could replace most of their recoverable Ledger in one rewrite, retain every superseded net observation in model context, and remain indefinitely busy when a provider stopped mid-response. This PR protects the existing product path at those boundaries: it refuses abrupt single-revision Ledger loss, reduces superseded net reads only in model context, bounds an idle model invocation with one canonical-tail retry, and settles launcher Stop gracefully.

Correctable Brunch-owned Ledger validation is now an honest non-applied tool result rather than a false terminal failure. Pending Brunch calls render gold, applied calls green, correctable refusals compact neutral with retained detail, and actually thrown failures red. Lu accepted that sequence in the real visible local browser path.

A retained 43-minute Inventory observation showed that these mechanisms enable substantial partial elicitation but do not make whole-body Ledger mutation product-scalable. Cumulative permitted rewrites still contracted the Ledger and its evidence, and the final provenance query could not recover a verified causal chain. That negative product-scale disposition remains final. For ordinary documents, Petrinaut's Stock assistant is now the fallback; Brunch remains a configured hidden Cmd-K alternate, while Brunch-focused deployments and tests can select it for fresh profiles with VITE_PETRINAUT_DEFAULT_ASSISTANT=brunch.

🔗 Related links

🚫 Blocked by

None. Mission 7e is closed as an engineering partial; its product-scale Ledger successor is deliberately separate rather than a blocker for this PR.

🔍 What does this change?

  • Protects mutate_workpiece revisions against missing headings and unannounced reductions greater than 25%, after stale-base and replay handling and before ordinal assignment.
  • Requires exact true-user authorization, named removed material and complete non-overlapping excerpt accounting for a destructive retraction, and retains that provenance through persistence and reconstruction.
  • Returns typed non-applied results for correctable Brunch-owned workpiece refusals while keeping schema, infrastructure, cancellation, persistence, output-contract and unknown failures thrown.
  • Renders Brunch pending/applied/refused/thrown states as gold/green/compact-neutral/red with spinner/check/dash/close signals, while leaving every canonical call separate and preserving Stock presentation defaults.
  • Projects superseded successful read_petrinaut_net results to observation identity and counts while preserving the latest full result and leaving canonical/public history unchanged.
  • Detects phase-aware provider-stream idleness, waits for cancellation acknowledgement, and permits exactly one same-model/same-reasoning recovery from Flue's canonical tail without replaying the user message or completed tools.
  • Prevents retry after complete tool input, rejects late cancelled-attempt events, and makes persona-launcher Stop abort and await the exact active submission before terminating owned services.
  • Re-points evidence and passage-policy coverage to the current workpiece contract and proves citation recovery through real compaction and fresh-process reopen.
  • Makes Stock the ordinary-document assistant fallback without migrating explicit stored choices; remote worked-model routes still force Brunch, and Brunch-focused launches can request a Brunch fallback through a validated Vite environment variable.
  • Updates user-facing AI-assistant documentation, the website assistant guide, Turborepo environment ownership, the Petrinaut patch changeset, mission evidence, and the separately gated successor spine.
🏗️ Agent notes

Ledger boundary

The guard runs at the sole persistent workpiece-state boundary. Same-tool-call/equal-Markdown replay remains idempotent; same ID with different Markdown refuses. Heading extraction covers ATX and Setext headings and ignores fenced-code false positives. Exactly 25% reduction remains permitted. A refusal changes no revision, body, hash, evidence or ordinal.

Retractions persist the withdrawal name, exact authorization quote, authorized true-user message IDs, and unique prior-body excerpts whose non-overlapping length accounts for the destructive net reduction. Source identity or free text alone is insufficient.

Typed refusal and presentation

Expected model-correctable operation validation returns { disposition: "refused", applied: false, correctable: true, code, message, currentRevision }. Applied and refused outputs share one discriminated schema. Refused outputs never become settled revision state or cause a false newer-revision warning. A browser-safe predicate validates the complete refusal schema and fails partial lookalikes closed.

The host presentation contract supports an optional tone and collapsed detail. Brunch uses it for pending gold, applied green, and compact neutral refusals; thrown failures remain red. Stock retains its existing presentation and hidden automatic-tool behavior. Color is not the sole signal.

Context, provider liveness and Stop

Model-only projection chooses the last structurally valid net read in each independently projected slice. Entry order, IDs, Voice correlation, browser output, freshness identity and canonical Flue history are preserved.

The watchdog treats model events—not sockets, heartbeats or synthetic stream start—as progress. Production limits are 10 seconds generally, 15 seconds only from reasoning start to its first delta/completion, and two seconds for cancellation acknowledgement. Canonical toolcall_end wins the race. Only the first acknowledged idle in one operation scope carries Flue's retryable-interruption marker.

The native OpenAI/browser fixture proves one distinct retry, no duplicate user input or completed ping, no canonical admission of either speculative tool, terminal failure after repeated idle, and no global-fetch fallback. Launcher Stop tracks initial and browser-continuation admissions, uses public Flue abort, waits up to five seconds for settlement, records launcher-stop.json, and terminates owned services before slower browser cleanup.

Assistant selection

VITE_PETRINAUT_DEFAULT_ASSISTANT accepts only stock or brunch; unset or blank resolves to Stock and an invalid value fails at startup/build. An explicit stored stock or brunch choice always wins, so changing the launch fallback performs no migration. The local Brunch launcher and synthetic persona/compiler paths explicitly request Brunch. The Cmd-K switch remains hidden when no Brunch endpoint is configured and is unavailable on remote worked-model routes, which continue to force Brunch without overwriting the ordinary preference.

Final observation and claim boundary

run-1L2zTl is local-only. It retained 25 user submissions, 24 accepted Ledger revisions and a useful partial net over 42m 59.561s. Two abrupt losses refused and repaired. Cached context was about 210k near minute 30, later reached 254,677, and compacted after the manual answer at 258,968 tokens.

The accepted body nevertheless contracted from a 33,876-character peak to 10,487 through individually permissible revisions without a retraction; current-revision evidence relations fell from 26 to four. The final explanation disclosed that query_workpiece could not reconcile the visible place to a verified construction chain. Native records estimate $44.1015 of Brunch usage and $0.1654 of persona usage for this run.

This PR establishes narrow per-revision loss prevention, model-only net-read economy, deterministic bounded provider recovery, graceful terminal settlement, honest refusal presentation, and explicit host selection. It does not establish product-scale Ledger preservation, stable provenance across whole-body rewrites, acceptable long-session latency/cost, behavioral correctness of the Inventory model, a verified final explanation, hosted reliability, or Mission 7d semantic acceptance.

Pre-Merge Checklist 🚀

🚢 Has this modified a publishable library?

This PR:

  • modifies an npm-publishable library and I have added a changeset file

📜 Does this require a change to the docs?

The changes in this PR:

  • require changes to docs which are made as part of this PR

🕸️ Does this require a change to the Turbo Graph?

The changes in this PR:

  • affected the execution graph, and the turbo.json files and generated task-dependency record have been updated

⚠️ Known issues

  • Whole-body Ledger mutation is not product-scalable. The per-revision guard permits cumulative contraction and cannot preserve evidence relations when unrelated rewrites move or paraphrase cited passages.
  • A current observed net may become mechanically unexplainable after layout or another intervening structural change; the final paid-run query returned unavailable reconciliation and an untrusted fallback.
  • Explicit recovering copy and the dedicated acknowledged-cancellation late-completion case remain residual acceptance leaves; the implemented browser path already exposes a separate retry and terminalizes speculative rows.
  • Persona evidence capture can freeze when the persona bridge exits even though the browser remains available for manual turns; canonical SQLite history survives, but the hashed evidence bundle then omits that tail.
  • One intermediate synthetic-browser test made two unintended low-reasoning OpenAI requests over only “Read the empty ledger.” Exact request IDs and cost were not retained. No Inventory pack, Ledger, persona material or retained-run content was sent; the fixture now installs the synthetic provider before submission and rejects every global fetch.
  • Website lint retains one pre-existing set-state-in-effect warning, and older core lint retains its recorded no-await-in-loop warnings.

🐾 Next steps

  • Separately cut proportional Ledger operations that reconstruct one Markdown authority while deterministically preserving unaffected evidence and current-basis explanation across recorded structural changes. This successor also owns evidence-refusal recovery and retraction-specific live/reopen proof.
  • Re-enter the residual lifecycle/evaluation leaves only after the Ledger successor makes the long-run path viable or for another named release consumer.
  • Re-admit Mission 7d's worked-example semantic bar only after scalable settlement and verified provenance exist.
  • Keep broader explanation evaluation, repeat/change, reviewer revision, experiment configuration, optimization, hosted deployment, Voice, and the deferred Settings → Labs assistant toggle in their named future drafts or external-owner sections of MISSION.next.md.

🛡 What tests cover this?

Passed on the final branch:

  • 75 @hashintel/brunch-agent core unit tests and the full @apps/brunch-agent unit suite.
  • 28 provider-admission controls, including reasoning grace, cancellation acknowledgement, admission races, complete-tool-input terminality and late-output rejection.
  • Native OpenAI/browser idle recovery with one completed tool, one retry, no replay and denied global fetch.
  • 10-case history-retention integration coverage across compaction and fresh-process reopen.
  • 57 passage-policy rows across 165 synthetic requests and 11 workpiece-evidence observations across 33 synthetic requests.
  • 781 website unit tests, including 39 assistant-selection/app-owner tests, and four website integration tests across three files.
  • Synthetic persona and compiler-feedback browser paths; the final Stock-default head reran the full persona path with an explicit Brunch fresh-profile fallback.
  • Brunch/core/Petrinaut/website builds, TypeScript, ESLint and repository formatting gates, with only the pre-existing warnings listed above.
  • The full test:anthropic-tools wrapper, including both native tool catalogues through Anthropic's token-count endpoint.
  • Lu's visible browser acceptance of held pending gold, applied green, compact neutral refusal and detail, separate corrected green, and red schema failure.

No paid inference was used for WP-E, Stock-default verification, or rendered acceptance.

❓ How to test this?

  1. Run yarn workspace @hashintel/brunch-agent test:unit and yarn workspace @apps/brunch-agent test:unit.
  2. Run yarn workspace @apps/brunch-agent test:integration, test:workpiece-evidence, test:passage-policy, and test:persona.
  3. Run yarn workspace @apps/petrinaut-website test:unit and yarn workspace @apps/petrinaut-website test:integration.
  4. Open an ordinary document without a stored selection and confirm Stock is mounted; configure the Brunch endpoint, choose Use Brunch from Cmd-K, reload, and confirm the explicit choice persists.
  5. Launch with VITE_PETRINAUT_DEFAULT_ASSISTANT=brunch in a fresh profile and confirm Brunch mounts; store Stock explicitly and confirm that stored choice still wins.
  6. Inspect the deterministic browser recovery and refusal-presentation cases; do not resume the ignored paid-run directories or start another paid run.

📹 Demo

No portable demo artifact is attached. Lu reviewed and accepted the deterministic rendered WP-E sequence in a visible local Chrome window. The decisive 43-minute Inventory observation is retained locally as run-1L2zTl; its reviewed negative product-scale result is recorded in MISSION.md.


Stack created with GitHub Stacks CLI • Give Feedback 💬

@vercel

vercel Bot commented Sep 16, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
petrinaut Ready Ready Preview Sep 17, 2026 12:42am UTC
petrinaut-docs Ready Ready Preview Sep 17, 2026 12:42am UTC
2 Skipped Deployments
Project Deployment Actions Updated
hash Ignored Ignored Preview Sep 17, 2026 12:42am UTC
hashdotdesign-tokens Ignored Ignored Preview Sep 17, 2026 12:42am UTC

Request Review

@cursor

cursor Bot commented Sep 16, 2026 •

Copy link
Copy Markdown

PR Summary

Medium Risk
Changes provider streaming, workpiece settlement semantics, and model-facing context projection on the main chat path; coverage is broad but misconfiguration of idle retry or refusal handling could affect live conversations.

Overview
This PR tightens long Brunch interview behavior and begins WP-E-style tool presentation without changing stock Petrinaut defaults for existing users.

Model context and provider liveness: Brunch context projection now collapses older successful read_petrinaut_net results to observation identity plus collection counts while keeping the latest full definition in each slice. Chat-agent provider admission gains phase-aware stream-idle timeouts, cancellation acknowledgement, and at most one retryable recovery per submission (via Flue’s interruption marker) for the mounted chat agent only.

Workpiece contract and UI: Settled revisions ignore typed refused mutate_workpiece outputs (including retraction metadata on applied rows). The Petrinaut demo renders pending Brunch tools with a gold pending tone and correctable refusals as compact neutral cards with the refusal message, while the Ledger pane no longer treats refused inputs as newer state.

Ops and tests: Persona launcher Stop aborts the active Flue submission, records settlement, then tears down process groups; browser persona flows track continuation admissions. VITE_PETRINAUT_DEFAULT_ASSISTANT controls the build-time fallback (stock by default; Brunch local/test scripts set brunch). Integration and passage-policy tests are realigned to text-based evidence declarations, applied/refused dispositions, and citation recovery after compaction/reopen.

Reviewed by Cursor Bugbot for commit 65bd76e. Bugbot is set up for automated code reviews on this repo. Configure here.

@github-actions github-actions Bot added area/infra Relates to version control, CI, CD or IaC (area) area/libs Relates to first-party libraries/crates/packages (area) type/eng > frontend Owned by the @frontend team area/tests New or updated tests area/apps labels Sep 16, 2026
@lunelson lunelson self-assigned this Sep 16, 2026
lunelson and others added 10 commits September 16, 2026 17:51
Consolidate the experiment-configuration planning that was spread across the Mission 7d archive, the future spine and the Mission 11 draft into one draft, `docs/mission-drafts/8-experiment-configuration-from-the-ledger.md`, taking the freed Mission 8 number. The draft records the Petrinaut experiment terrain as read at HEAD, the Ledger-condition to experiment-destination correspondence table, the design assessment (Brunch drafts a session proposal prepared by Petrinaut's own `prepareExperiment`, the user presses Run; a thin document-entity design stays an upstream option), the upstream delta list for Petrinaut, the Brunch-side prerequisite of scenario and metric operations in `mutate_petrinaut_net`, and an implementer's entry for the handoff.

Point the spine and the Mission 9, 10 and 11 drafts at the new home, rename the spine's deployment follow-on heading to "Hosted deployment successor" now that the Mission 8 number is reused, and record the read-only dev/prod persistence audit (SQLite locally, Postgres deployed) in the spine and the worked-example draft.

Amp-Thread-ID: https://ampcode.com/threads/T-01a0aab3-72bb-71dc-8a77-a984b3617b9c
Co-authored-by: Amp <amp@ampcode.com>
@github-actions github-actions Bot added the area/deps Relates to third-party dependencies (area) label Sep 16, 2026
@lunelson
lunelson deployed to pull-request September 16, 2026 17:19 — with GitHub Actions Active
@lunelson
lunelson deployed to pull-request September 16, 2026 17:19 — with GitHub Actions Active
@codecov

codecov Bot commented Sep 16, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 66.31%. Comparing base (5af427b) to head (21ced0d).
⚠️ Report is 11 commits behind head on main.

Additional details and impacted files
@@            Coverage Diff             @@
##             main    #9761      +/-   ##
==========================================
+ Coverage   66.28%   66.31%   +0.03%     
==========================================
  Files        1775     1782       +7     
  Lines      191861   192337     +476     
  Branches     7853     7856       +3     
==========================================
+ Hits       127171   127552     +381     
- Misses      63209    63302      +93     
- Partials     1481     1483       +2     
Flag Coverage Δ
blockprotocol.type-system 38.15% <ø> (ø)
local.claude-hooks 0.00% <ø> (ø)
local.harpc-client 51.49% <ø> (ø)
local.hash-backend-utils 3.27% <ø> (ø)
local.hash-graph-sdk 10.02% <ø> (ø)
local.hash-isomorphic-utils 12.22% <ø> (ø)
rust.antsi 2.36% <ø> (ø)
rust.error-stack 90.81% <ø> (ø)
rust.harpc-codec 84.70% <ø> (ø)
rust.harpc-net 96.24% <ø> (+0.03%) ⬆️
rust.harpc-tower 67.03% <ø> (ø)
rust.harpc-types 0.00% <ø> (ø)
rust.harpc-wire-protocol 92.23% <ø> (ø)
rust.hash-codec 72.76% <ø> (ø)
rust.hash-config 81.14% <ø> (ø)
rust.hash-graph-api 19.71% <ø> (ø)
rust.hash-graph-authentication 96.02% <ø> (ø)
rust.hash-graph-authorization 63.14% <ø> (ø)
rust.hash-graph-embeddings 91.88% <ø> (ø)
rust.hash-graph-postgres-store 32.15% <ø> (ø)
rust.hash-graph-store 48.41% <ø> (ø)
rust.hash-graph-temporal-versioning 50.18% <ø> (ø)
rust.hash-graph-types 0.00% <ø> (ø)
rust.hash-graph-validation 84.71% <ø> (ø)
rust.hash-middleware 90.92% <ø> (ø)
rust.hashql-ast 89.63% <ø> (ø)
rust.hashql-compiletest 28.39% <ø> (ø)
rust.hashql-core 79.03% <ø> (+0.11%) ⬆️
rust.hashql-diagnostics 72.51% <ø> (ø)
rust.hashql-eval 79.82% <ø> (ø)
rust.hashql-hir 89.09% <ø> (ø)
rust.hashql-mir 87.92% <ø> (ø)
rust.hashql-syntax-jexpr 94.04% <ø> (ø)
rust.problematic 73.88% <ø> (?)

Flags with carried forward coverage won't be shown. Click here to find out more.

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@vercel
vercel Bot temporarily deployed to Preview – petrinaut-docs September 16, 2026 21:09 Inactive
@vercel
vercel Bot temporarily deployed to Preview – petrinaut September 16, 2026 21:09 Inactive
@lunelson
lunelson deployed to pull-request September 16, 2026 21:10 — with GitHub Actions Active
@lunelson
lunelson deployed to pull-request September 16, 2026 21:10 — with GitHub Actions Active
@kostandinang
kostandinang added this pull request to stack #9763 September 16, 2026 21:20
@lunelson
lunelson requested a review from a team as a code owner September 17, 2026 00:37

@TimDiekmann TimDiekmann left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Infra-changes looks good, please ping me when code review is done.

@lunelson
lunelson added this pull request to the merge queue Sep 17, 2026
Merged via the queue into main with commit 20f84cd Sep 17, 2026
214 checks passed
@lunelson
lunelson deleted the ln/fe-1573-mission-7e-express branch September 17, 2026 06:38
@hash-release hash-release Bot mentioned this pull request Sep 17, 2026

This branch was successfully deployed

4 active (2 outdated) deployments
Preview – petrinaut-docs — 65bd76ee Deployed Sep 17, 2026 by vercel[bot]
Preview – petrinaut — 65bd76ee Deployed Sep 17, 2026 by vercel[bot]
Preview – hash — bbf838e7 Deployed Sep 17, 2026 by vercel[bot]
pull-request — bd933e7f Deployed Sep 16, 2026 by lunelson via Sourcemaps (@apps/hash-integration-worker) #36807
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area/apps area/deps Relates to third-party dependencies (area) area/infra Relates to version control, CI, CD or IaC (area) area/libs Relates to first-party libraries/crates/packages (area) area/tests New or updated tests type/eng > frontend Owned by the @frontend team

Development

Successfully merging this pull request may close these issues.

3 participants