Repository navigation
FE-1768: Let any coding agent play the Brunch persona through a browser bridge CLI - #9833
Conversation
|
The latest updates on your projects. Learn more about Vercel for GitHub. 4 Skipped Deployments
|
PR SummaryMedium Risk Overview Resume now reconciles against Adds unit tests for RPC framing, bridge lifecycle, CLI parsing, agent launch, and brief assembly; removes Pi extension/integration tests. Reviewed by Cursor Bugbot for commit 56bd8c0. Bugbot is set up for automated code reviews on this repo. Configure here. |
be7ff43 to
6e0fd3a
Compare
6e0fd3a to
4773467
Compare
4773467 to
4bb0f02
Compare
df3add9 to
6963863
Compare
|
Thanks, both fair, and both are addressed in 6963863.
I checked both with a live launch: the agent started in an empty temporary directory, and the bridge-log entry for its utterance carried its settings. Separately, on your caution about the model: #9811 now returns every Petrinaut assistant's default to |
kostandinang
left a comment
There was a problem hiding this comment.
Thanks, this addresses both points for me. Approving.
The merge-base changed after approval.
6963863 to
7bbe63a
Compare
|
A correction to my earlier answer on isolation: I framed it the wrong way round. The case material is the persona's own knowledge, so the persona reading it isn't a leak; the boundary that matters is that Brunch sees only what the persona types, and that holds by construction because Brunch declares no Flue sandbox and so has no file or shell tools. The temporary working directory stays, but only to keep the repository's development guidance out of the persona's context. The real problem with the persona was the opposite one: it was too aligned with the interview. It knew it was in an evaluation of Brunch, followed a scripted arc (review, why, correction) and was coached to structure answers for the elicitor. 9dd069b and 7bbe63a brief it instead as a naive domain expert who knows their work but nothing about Brunch's aims or how an interview should go, answers the question asked briefly, volunteers nothing and stops by their own judgement; the case packs now describe the person rather than the test, and the verbosity and disclosure axes are gone. None of this has had a live persona run yet. |
…er bridge CLI Replace the Pi-only brunch_turn extension with a Pi-style JSONL bridge owned by the persona launcher and a per-run persona helper CLI. The launcher can start Claude Code, Codex, Cursor Agent, Pi or a custom command with a harness-neutral brief, in a Herdr pane when available or in the current terminal. Resume reconciles against a bridge log instead of the Pi session file. Co-authored-by: Cursor <cursoragent@cursor.com>
…ed bridge turns as aborted
The lower branch replaced a conversation's `{ mode, construction: { binding } }` initial data with `{ binding }`. This fixture only exists on the persona branch, so the restack left it writing the removed shape; it now writes what the launcher records.
The launcher marked a turn admitted only after appending the utterance to `bridge-log.jsonl`, so a failed log write read as a pre-admission failure even though Brunch already had the message, and the persona could resend it. The turn is now admitted first; a log failure after that is a failed admitted turn, which the persona is told never to resend.
…sona utterance to its agent The agent started in the run directory, inside the repository, so agents that load AGENTS.md or CLAUDE.md from parent directories got Brunch's development guidance in the persona's context. It now starts in a fresh temporary directory, reaching the brief and helper by absolute path. A resume starts a fresh persona session, possibly with another agent or model, and run.json keeps only the latest. Each persona entry in bridge-log.jsonl now records the agent settings that wrote it, and the README and help describe resume as continuing the Brunch conversation.
…on participant Each situation pack now describes the person: who they are and how they talk, what they want, what they know about their operation, what they take for granted and what they don't know. The behavioural rules, gating such as "reveal only when asked", staged corrections, mission and condition notes, citations of deleted source files, and modelling knowledge a domain expert wouldn't have are gone; beliefs the case relies on are stated as the person's own. Each opening message is now just the person's first message, without an operator header.
The persona brief framed the agent as a participant in an evaluation of Brunch and coached it to serve the interview. Its role document now describes a person who knows their own work but nothing about Brunch's aims or how an interview should go: they answer the question asked, briefly, volunteer nothing, and stop by their own judgement. There is no default objective; an optional --objective is the person's own aim. The verbosity and disclosure axes are deleted, since each case describes how its person talks, and the opening message is the whole case file. The README states the one boundary that matters: Brunch sees only what the persona types, while the case is legitimately the persona's own knowledge.
…never to repair the browser Co-authored-by: Cursor <cursoragent@cursor.com>
There was a problem hiding this comment.
Cursor Bugbot has reviewed your changes and found 2 potential issues.
❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.
Reviewed by Cursor Bugbot for commit 4f6f52c. Configure here.
kostandinang
left a comment
There was a problem hiding this comment.
Non-blocking: this changes the persona and every case at once, so runs from before and after aren't comparable. Could run.json record a hash of the case and brief, like it already does for --initial-net?

🌟 What is the purpose of this PR?
Persona testing lets a coding agent play a domain expert talking to Brunch through the real Petrinaut panel in the browser. Until now the simulated user had to run inside a Pi-only tool extension, so the same persona couldn't be handed to Claude Code, Codex, Cursor or any other agent. This PR replaces that extension with a launcher-owned browser bridge and a small per-run
personacommand that any agent can run from its shell. The launcher can start one of the common agents for you, or print a prompt for any agent you start yourself.The PR also reworks how the persona is briefed. The old brief told the agent it was taking part in an evaluation of Brunch, scripted the interview's arc and coached it to structure answers for the interviewer, so it behaved like a knowing participant. The persona is now a naive domain expert: they know their own work, want the help, and know nothing about Brunch's aims or how an interview should go. They answer the question they were asked, briefly, volunteer nothing and stop by their own judgement. The one boundary that matters is unchanged and holds by construction: Brunch sees only what the persona types. Brunch declares no Flue sandbox, so it has no file or shell tools with which to reach the case material.
🔗 Related links
🚫 Blocked by
🔍 What does this change?
Harness-independent bridge (b6d4029)
<run>/bin/personahelper.persona saytypes a message into the real browser composer, waits until Brunch has completely finished, including browser-tool continuations, and prints the reply. Other commands:transcript,state,end, and rawrpc. Commands and events are typed records (prompt,get_state,get_transcript,end;turn_start,assistant_message,turn_settled), so later interaction kinds, such as answering a question card, can be added as new record types.--agent claude|codex|cursor-agent|pistarts that agent, and--agent-command '<template>'starts any other; with neither, the launcher prints the launch prompt and helper path. Inside Herdr the agent gets its own pane; otherwise it takes over the launcher's terminal.bridge-log.jsonl, which records every admitted utterance, instead of a Pi session file..pi/extensions/brunch-persona-testing*),brunch_turn, the persona configuration module and their tests.Follow-up fixes
{ binding }initial data that FE-1763: Let Brunch build Petrinaut nets as capably as the existing assistant #9811 introduced (e0cab5e).Where the agent starts and who wrote each turn (6489f9d)
AGENTS.mdorCLAUDE.mdfiles. It reaches its brief and helper by absolute path.run.jsonkeeps only the latest agent. Each persona entry inbridge-log.jsonlnow records the agent settings that wrote it. The README and help describe resume as continuing the Brunch conversation.The persona as a naive domain expert (9dd069b, 7bbe63a)
launch/brief/system.mdis the persona's whole role contract. It covers who the person is, how they answer (the literal question, briefly, without organising information for Brunch), staying in character, and when to stop. It contains no evaluation or "elicitor" framing, and it replaces the imaginary "turn budget" with a real stopping rule.--objectiveremains for giving the person a private aim, phrased as something they want rather than as interview steps.--persona-verbosityand--persona-disclosureaxes; each case describes how its person talks. Olderrun.jsonaxis keys are ignored.EVALUATIONS.mddescribe the persona and the boundary accordingly.EVALUATIONS.mdalso lists the eighth case,support-desk-staffing.Pre-Merge Checklist 🚀
🚢 Has this modified a publishable library?
This PR:
📜 Does this require a change to the docs?
The changes in this PR:
🕸️ Does this require a change to the Turbo Graph?
The changes in this PR:
support-desk-staffing, which used to rely on explicit rules to stop the persona asking for the test itself.bridge-log.jsonland says so.<run>/bin/persona endor Ctrl-C.🐾 Next steps
🛡 What tests cover this?
rpc-protocol.test.ts: LF-only record framing, and how known commands and invalid input are answered.browser-bridge.test.ts: an admitted prompt through to settlement; failures before and after admission, including a Stop wrapped by the browser turn settling as aborted; sequential turns, with a disconnect aborting its turn;endacknowledged before the launcher stops; and the generated helper.persona-cli.test.ts: a message beginning with a dash is text, and the CLI's own flags stay options.launch/agent.test.ts: custom command templates, the operator-started case, the agent's environment, and reading the Herdr pane ID.launch/brief.test.ts: brief section order, with and without an objective, and the resume notice.launch.test.ts: launcher Stop settlement, case loading with the whole opening file, resume requiring the bridge log, run metadata with role and agent settings, service readiness, and help.❓ How to test this?
yarn brunch:persona --helpandyarn brunch:persona --list-casesfrom the repository root. Both are free: they start no services or inference.BRUNCH_CHAT_PORT=4332 BRUNCH_PANEL_PORT=4926 yarn brunch:persona --case inventory-purchasing --agent claude, and watch the conversation in the Chrome window.<run>/bridge-log.jsonlrecords the agent settings for each persona message.📹 Demo
Stack created with GitHub Stacks CLI • Give Feedback 💬