Repository navigation
FE-1763: Let Brunch build Petrinaut nets as capably as the existing assistant - #9811
Conversation
|
The latest updates on your projects. Learn more about Vercel for GitHub.
1 Skipped Deployment
|
PR SummaryHigh Risk Overview Removed the net ledger, stale-net signal, mutation sidecar verification, reported-document-revision scope, worked-model store/routes, and many integration scripts ( Petrinaut gains Reviewed by Cursor Bugbot for commit 0e6321a. Bugbot is set up for automated code reviews on this repo. Configure here. |
69d23b5 to
18101d9
Compare
18101d9 to
dbbb6bc
Compare
…pruned CI builds keep it
…stale reads, and compare browser bindings key-order-independently
…nt canonical tools
…and derive browser tool names from the catalogue
…from one table The no-mode read_petrinaut_docs tool was the last tool whose result came back as a later client-tool-result signal. It goes, with the transport's signal send, history folding and sync/async client-tool split. Integrated mode keeps canonical readPetrinautDoc in band. Which canonical tools may change the document, and which element kinds their changes name, now come from petrinautToolEffects in plugin-sdcpn. The app catalogue, rendezvous failure wording, net-change attribution, the why selector, and the website's revision and draft checks derive from it instead of hand-kept name lists.
The draft checks required a settled, current Ledger before Brunch could put up a draft, but nothing linked the draft to the Ledger's content: the browser looked the Ledger up only to confirm one existed. With no Ledger, Brunch had to write filler or call createExperiment directly, which runs at once with no Run button. A draft now needs only a prior canonical net read, on the server and in the browser; the check that the read carries a document revision stays. The unused workpiece authority options are removed, and the browser witness drafts after a net read with no Ledger.
…'s Petrinaut guidance Construction and experiment proposals no longer need to flow from, wait for or cite a settled Ledger; they follow what the person has stated, and the Ledger stays the account Brunch keeps while it asks. With one conversation mode left, the "In I" qualifiers and the construct-only branch no longer distinguish anything and are removed. Steps and rules that repeated Stock's capability guidance or the always-on instructions are gone. Experiment guidance now triggers on the person's stated decision, and the inline createExperiment steer asks Brunch to name any stated restriction the run will not enforce, since neither request shape carries constraints.
Only one mode was left, and the schema already refused it without a construction binding, so `mode` told the server nothing the binding did not. A conversation's initial data is now just `{ binding }`: when present, Brunch mounts the Petrinaut tools, the draft and the browser executor; when absent, as on Brunch's own chat page and in the deployment smoke, it offers the instruction and modelling skill only. The one-entry tool catalogue map, the website's always-true construction gate and the old persona initial-data fallbacks go with it, and names, the browser witness and its script drop "integrated", which no longer distinguishes anything.
The handle was created in an effect, so for one render after a document opened, the repository and the Brunch binding already named the new document while the handle still edited the old one. A gate hid the browser executor and the conversation binding in that render; a first message sent then would create the new conversation without its binding, and Flue records creation data only once. Creating a handle has no side effects, so the handle is now chosen in the render that opens the document, which lets the gate go: the binding, transport and handle always name the same document. The subscription that saves the handle's changes stays in an effect.
Clearing a save failure recorded for a document other than the open one was an effect that adjusted state after the document changed. It depends only on the current values, so it now happens in render, before anything reads the failure. Saving the handle's changes stays an effect, with a note on why: it subscribes to the handle, the subscription must end with that handle, and render may create handles that React discards.
The chat agent's render wrote whether the conversation had a document binding into the per-request admission scope, so the provider wrapper could allow mixed browser/server proposals made only of allow-listed tools. A conversation without a binding mounts no browser tools and cannot propose a mixed call, so the flag was true whenever it mattered. The wrapper now applies the allow-list alone, and the render no longer writes ambient state. The registration test now proves admission scoping with a proposal refused in every Brunch scope: a workpiece query beside the net read it depends on.
…g instruction The prompt weeding removed this rule from the always-on instruction as a duplicate of the skill body, but the skill body arrives once, as a tool result on the first turn, and loses weight as the conversation grows. A persona run without it built the net on 5 of 37 turns, mostly early. The rule is back, keyed to each meaning-bearing answer rather than to a Ledger settlement.
Flue resolves models through pi-ai 0.83.0, which predates GPT-6, so Brunch could not select `openai/gpt-6-sol`. Its entry is copied from pi-ai 0.87.1's OpenAI catalogue beside GPT-6 Luna, which it matches apart from name and prices, and shares Luna's removal condition.
Admission listed `mutate_workpiece` among the draft's dependencies, so a proposal that settled the Ledger and drafted an experiment together was refused as if the draft read the Ledger's result. Drafting now needs only a prior canonical net read, so that read is the draft's one dependency.
The shim was for `@ai-sdk/openai` 3.0.63, which predated GPT-6 and dropped the reasoning effort unless forced. Its removal condition, 3.0.113 or later project-wide, is met by main's bump to 3.0.114, which treats every GPT model from version 5 as a reasoning model.
…e document unchanged A browser call ends without a result in three ways: the browser reports a failure, its lease expires, or Stop aborts it. Only the failure path checked whether the tool can change the document, so a stopped or expired read or draft was reported as an attempted document effect of unknown outcome, which the model is told never to repeat. One classifier now builds the error for all three: unstarted, document unchanged, or unknown.
The panel claimed each issued call as soon as it streamed in, before it waited behind earlier document calls, and the server takes a claimed call as possibly started. Stop during a multi-tool turn therefore reported queued, never-started calls as unknown writes, and the model would not redo them. Claiming at the start of the call's lane turn makes a claimed call a started one, so a queued call stopped before its turn is reported unstarted.
…etry that could not succeed The draft card read its authority from history and prepared the experiment before claiming the issued call, so nothing renewed the call's lease meanwhile and a slow read could expire it. The card now claims first and prepares under the renewed lease. "Retry preparation" re-ran the claim after the one-use call had already been claimed and settled, so it could never succeed; the model can redraft instead, and the card still shows the failure.
`@hashintel/petrinaut`'s changeset now covers `executePetrinautAiMutation` and when an in-band call is claimed, and `@hashintel/petrinaut-core`'s covers the stock prompt's separately exported behavioural frame and capability guidance.
…head of them Since the browser claims a call only when its document-lane turn starts, calls queued behind a running one were unclaimed and expired 25 s after issue. A renewal now extends the unclaimed calls on the same binding, so queued calls expire only once the browser stops renewing.
The five planning and review notes under `libs/@hashintel/brunch-agent/docs/refactoring/` were local working material for the restack. They now live in an ignored `_holding/` directory.
`petrinautAiModel` returns to the original assistant's `gpt-5.5-2026-04-23` at `medium`, so the Stock route matches `main` and Brunch shares its baseline. Brunch's local OpenAI catalogue declares that dated snapshot, which pi-ai 0.83.0 lists only as `gpt-5.5`, keeps GPT-6 Sol as an opt-in for testing, and drops GPT-6 Luna.
`inBandBrowserTools` exposed Brunch's handoff protocol (claim, prepare, fail, release and submit) as Petrinaut public API. Petrinaut now calls `run(call, execute)` at the call's turn in same-document order, with a signal that aborts on Stop; the host decides whether the call may start, calls `execute` at most once, and reports the output or failure itself. Claiming, lease renewal and failure dispositions stay in the website.
…ics again The canonical read and mutation tools already return a call's recorded output when it runs again after a remount or history replay; the diagnostics read now does the same.
Positions are layout, which is not part of what an element means. The position tools and applyAutoLayout still count as document changes, so a net read still goes stale, but query_workpiece no longer lists them as changes to an element; applyAutoLayout named no IDs, so it was never credited anyway.
|
Good question, and you're right that it was Brunch's protocol leaking into Petrinaut's API. Narrowed in b3d0ac4: |
There was a problem hiding this comment.
Cursor Bugbot has reviewed your changes and found 2 potential issues.
❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.
Reviewed by Cursor Bugbot for commit 0e6321a. Configure here.
hashdotai
left a comment
There was a problem hiding this comment.
Legal approval – license changes are only deleted licenses in removed packages.

🌟 What is the purpose of this PR?
Brunch used to change Petrinaut nets through its own narrow editing path. That path couldn't express the sequences of changes the existing Stock assistant makes, so Brunch struggled to build a complete model from a description. This PR gives Brunch the same editing tools the Stock assistant uses. Petrinaut decides what each edit means and does, and Brunch adds what it's for: interviewing the user, keeping the Ledger of what has been agreed, proposing experiments for review, and explaining how the net came to be.
Getting there also meant removing what no longer earned its keep. The branch had carried alternative construction strategies, evaluation modes, a matched-parity evaluator, a worked-model feature, unmounted plugins and a server-side check that re-verified what the browser had already reported; all of them are gone, and with only one way left to run Brunch, so is the mode switch. A conversation opened from the Petrinaut panel is bound to its document and gets Petrinaut's tools; one opened without a binding, as on Brunch's own chat page, gets the interview and modelling guidance only. Brunch credits each net change to the tool call Petrinaut says applied it. Every Petrinaut assistant runs on one default model, which Petrinaut defines: GPT-5.5 at medium reasoning, as the original assistant did.
Browser-visible persona runs on
inventory-purchasing(from #9833) show Brunch building an executable net alongside the interview, with timed supply, quality release, shelf life, FEFO draws, metrics and saved scenarios, and no failed tool calls beyond one expired browser lease. They also show where the time goes: rewriting the whole Ledger on every turn takes 70% of model time. Native Stock parity, live Voice and running a reviewed experiment draft in a production build are still unproved.🔗 Related links
🚫 Blocked by
🔍 What does this change?
Petrinaut libraries
petrinaut-coresplits the Stock prompt into a behavioural frame and reusable capability guidance, recomposed in the original order. It also exportspetrinautAiModel, the default model for every Petrinaut assistant:gpt-5.5-2026-04-23atmedium, the original assistant's setting.petrinautexportsexecutePetrinautAiMutation, so a host can run one canonical edit with Petrinaut's own no-op detection and output. The assistant panel claims and returns browser tool results while a response is still streaming, executes same-document calls in order, and cancels them on Stop. It claims each call when that call's turn in the document lane starts, so a call stopped while still queued is reported as never started.Brunch packages
plugin-sdcpnmounts Petrinaut's complete tool catalogue with Brunch's prompt, SDCPN modelling skill and reviewed experiment draft whenever the conversation carries a document binding (initialData: { binding }). Without one it offers the modelling instruction and skill only../flueentry for the agent hooks and skills, which only a Flue build can load. Skills use Flue's nativeSKILL.mdimport.core/src/constants.ts. The browser imports it through@hashintel/brunch-agent/constants.@dev/sourceexport condition.apps/brunch-agentpetrinautAiModelby default, so Brunch and the Stock route share one baseline. Its local OpenAI catalogue declares that dated GPT-5.5 snapshot, which pi-ai 0.83.0 lists only asgpt-5.5, and GPT-6 Sol as an opt-in for testing.apps/petrinaut-websitecreateExperiment. The draft card claims its call before preparing the experiment, so the call's lease is renewed while it prepares.petrinautAiModel, unchanged frommainby default, andPETRINAUT_AI_REASONING_EFFORTcan override its reasoning effort.brunch_asktool and custom browser tooling.Pre-Merge Checklist 🚀
🚢 Has this modified a publishable library?
This PR:
@hashintel/petrinaut: in-band browser results andexecutePetrinautAiMutation.@hashintel/petrinaut-core:petrinautAiModeland the separately exported Stock prompt parts.📜 Does this require a change to the docs?
The changes in this PR:
apps/brunch-agent/README.md,apps/petrinaut-website/README.md, BrunchARCHITECTURE.md(conversation composition, package entries and loading source in dev) and the SDCPN modelling skill.🕸️ Does this require a change to the Turbo Graph?
The changes in this PR:
turbo.json's have been updated to reflect thistest:unitbuilds core first.mutate_workpieceresubmits the whole Ledger as Markdown on every settlement. In a 45-turn persona run that was 70% of model time and 88% of output tokens, and each rewrite grew from 10 to about 90 seconds as the Ledger reached 37,000 characters. FE-1764 replaces this format.applyAutoLayoutcall this way, as far as the evidence shows; Brunch reported it as unknown and did not repeat it. Unverified; needs a probe.gpt-5.5-2026-04-23snapshot nor GPT-6, soapps/brunch-agent/src/openai-provider.tsadds that snapshot andgpt-6-sollocally, with a documented removal condition. Removing it is part of FE-1799.🐾 Next steps
🛡 What tests cover this?
yarn workspace @apps/brunch-agent test:manual:browser, manual): drives the built website in a real browser against the Brunch server with a faux provider, covering canonical construction with compiler feedback, a direct experiment, and a reviewed draft with Dismiss.plugin-sdcpnconstruction tools and the draft tool;petrinaut-coreprompt composition,executePetrinautAiMutation, and in-band results in the assistant panel, including claiming at the start of each call's lane turn.yarn brunch:persona, from FE-1768: Let any coding agent play the Brunch persona through a browser bridge CLI #9833; paid, not in CI): twoinventory-purchasingruns with Brunch ongpt-6-sol, compared against earlier runs; findings above.❓ How to test this?
yarn dev:brunch, openhttp://127.0.0.1:4915, and select Brunch.yarn workspace @apps/brunch-agent test:manual:browser.Stack created with GitHub Stacks CLI • Give Feedback 💬