Skip to content

FE-1768: Let any coding agent play the Brunch persona through a browser bridge CLI - #9833

Merged
lunelson merged 10 commits into
mainfrom
ln/fe-1768-persona-cli
Sep 30, 2026
Merged

lunelson merged 10 commits into
mainfrom
ln/fe-1768-persona-cli

Conversation

@lunelson

@lunelson lunelson commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

🌟 What is the purpose of this PR?

Persona testing lets a coding agent play a domain expert talking to Brunch through the real Petrinaut panel in the browser. Until now the simulated user had to run inside a Pi-only tool extension, so the same persona couldn't be handed to Claude Code, Codex, Cursor or any other agent. This PR replaces that extension with a launcher-owned browser bridge and a small per-run persona command that any agent can run from its shell. The launcher can start one of the common agents for you, or print a prompt for any agent you start yourself.

The PR also reworks how the persona is briefed. The old brief told the agent it was taking part in an evaluation of Brunch, scripted the interview's arc and coached it to structure answers for the interviewer, so it behaved like a knowing participant. The persona is now a naive domain expert: they know their own work, want the help, and know nothing about Brunch's aims or how an interview should go. They answer the question they were asked, briefly, volunteer nothing and stop by their own judgement. The one boundary that matters is unchanged and holds by construction: Brunch sees only what the persona types. Brunch declares no Flue sandbox, so it has no file or shell tools with which to reach the case material.

🔗 Related links

🚫 Blocked by

🔍 What does this change?

Harness-independent bridge (b6d4029)

  • The launcher owns a Pi-style JSONL bridge on a private socket and writes a per-run <run>/bin/persona helper. persona say types a message into the real browser composer, waits until Brunch has completely finished, including browser-tool continuations, and prints the reply. Other commands: transcript, state, end, and raw rpc. Commands and events are typed records (prompt, get_state, get_transcript, end; turn_start, assistant_message, turn_settled), so later interaction kinds, such as answering a question card, can be added as new record types.
  • --agent claude|codex|cursor-agent|pi starts that agent, and --agent-command '<template>' starts any other; with neither, the launcher prints the launch prompt and helper path. Inside Herdr the agent gets its own pane; otherwise it takes over the launcher's terminal.
  • Resume reconciles against bridge-log.jsonl, which records every admitted utterance, instead of a Pi session file.
  • Deletes the Pi extension (.pi/extensions/brunch-persona-testing*), brunch_turn, the persona configuration module and their tests.

Follow-up fixes

  • If the agent fails to start, its Herdr pane is closed, and a Stop is reported as an aborted turn rather than a failed one (85ea258).
  • A persona message that begins with a dash is sent as text, not parsed as an option (e2388c6).
  • A turn is marked admitted before its bridge-log entry is written, so a failed log write reads as a failed admitted turn, which the persona is told never to resend (ea5384d).
  • The resume fixture writes the flat { binding } initial data that FE-1763: Let Brunch build Petrinaut nets as capably as the existing assistant #9811 introduced (e0cab5e).
  • A composer holding only whitespace counts as empty rather than as a draft the launcher refuses to overwrite, and the brief tells the persona never to repair the browser, the bridge or Brunch, or to reach Brunch any other way (b521016).

Where the agent starts and who wrote each turn (6489f9d)

  • The agent starts in a fresh temporary directory outside the repository, so it loads none of the repository's development AGENTS.md or CLAUDE.md files. It reaches its brief and helper by absolute path.
  • A resume starts a fresh persona session, possibly with another agent or model, and run.json keeps only the latest agent. Each persona entry in bridge-log.jsonl now records the agent settings that wrote it. The README and help describe resume as continuing the Brunch conversation.

The persona as a naive domain expert (9dd069b, 7bbe63a)

  • launch/brief/system.md is the persona's whole role contract. It covers who the person is, how they answer (the literal question, briefly, without organising information for Brunch), staying in character, and when to stop. It contains no evaluation or "elicitor" framing, and it replaces the imaginary "turn budget" with a real stopping rule.
  • There is no default objective; the person's goal comes from their case. --objective remains for giving the person a private aim, phrased as something they want rather than as interview steps.
  • Deletes the --persona-verbosity and --persona-disclosure axes; each case describes how its person talks. Older run.json axis keys are ignored.
  • All eight case packs now describe the person: who they are and how they talk, what they want, what they know about their operation, what they take for granted and what they don't know. The removed material includes behavioural rules, "reveal only when asked" gating, staged corrections, mission and condition notes, citations of files deleted in FE-1652: Keep Brunch planning protocols outside the product repository #9810, and modelling knowledge a domain expert wouldn't have. Each opening message is now the person's plain first message, so the launcher sends the whole file.
  • The operator README and EVALUATIONS.md describe the persona and the boundary accordingly. EVALUATIONS.md also lists the eighth case, support-desk-staffing.

Pre-Merge Checklist 🚀

🚢 Has this modified a publishable library?

This PR:

  • does not modify any publishable blocks or libraries, or modifications do not need publishing

📜 Does this require a change to the docs?

The changes in this PR:

  • are internal and do not require a docs change

🕸️ Does this require a change to the Turbo Graph?

The changes in this PR:

  • do not affect the execution graph

⚠️ Known issues

  • No live persona run has used the reworked brief and packs yet. Earlier runs, and the probe that checked the temporary directory and bridge-log attribution, used the previous brief. The case to watch is support-desk-staffing, which used to rely on explicit rules to stop the persona asking for the test itself.
  • The persona sees only Brunch's prose, not the net in the browser, so it can't review the model the way a person watching the canvas would.
  • The agent keeps the operator's shell environment, so user-level agent instructions still load; only the repository's own files are kept out.
  • Runs made with the Pi extension can't be resumed: resume requires bridge-log.jsonl and says so.
  • Nothing enforces a turn limit. The persona stops by its own rule, and the operator or a supervising agent can end a run at any point with <run>/bin/persona end or Ctrl-C.

🐾 Next steps

  • Run a few live persona cases against the new brief and calibrate it from the transcripts.
  • Exercise the experiment-draft, Run and "why" paths in a persona run; none has reached them yet. Best done after FE-1764 (internal), which reworks the Ledger.
  • Decide whether the persona should be able to see a plain summary of the model.

🛡 What tests cover this?

  • rpc-protocol.test.ts: LF-only record framing, and how known commands and invalid input are answered.
  • browser-bridge.test.ts: an admitted prompt through to settlement; failures before and after admission, including a Stop wrapped by the browser turn settling as aborted; sequential turns, with a disconnect aborting its turn; end acknowledged before the launcher stops; and the generated helper.
  • persona-cli.test.ts: a message beginning with a dash is text, and the CLI's own flags stay options.
  • launch/agent.test.ts: custom command templates, the operator-started case, the agent's environment, and reading the Herdr pane ID.
  • launch/brief.test.ts: brief section order, with and without an objective, and the resume notice.
  • launch.test.ts: launcher Stop settlement, case loading with the whole opening file, resume requiring the bridge log, run metadata with role and agent settings, service readiness, and help.
  • A real launch needs paid Brunch and agent calls and isn't in CI.

❓ How to test this?

  1. Check out the branch and run yarn brunch:persona --help and yarn brunch:persona --list-cases from the repository root. Both are free: they start no services or inference.
  2. For a live run (paid), choose unused ports, e.g. BRUNCH_CHAT_PORT=4332 BRUNCH_PANEL_PORT=4926 yarn brunch:persona --case inventory-purchasing --agent claude, and watch the conversation in the Chrome window.
  3. Confirm that the persona answers briefly and literally, doesn't coach Brunch, and ends the run itself, and that <run>/bridge-log.jsonl records the agent settings for each persona message.

📹 Demo


Stack created with GitHub Stacks CLI • Give Feedback 💬

@lunelson
lunelson added this pull request to stack #9812 September 24, 2026 16:31
@vercel

vercel Bot commented Sep 24, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

4 Skipped Deployments
Project Deployment Actions Updated
hash Ignored Ignored Preview Sep 30, 2026 11:20am UTC
hashdotdesign-tokens Ignored Ignored Preview Sep 30, 2026 11:20am UTC
petrinaut Skipped Skipped Sep 30, 2026 11:20am UTC
petrinaut-docs Skipped Skipped Sep 30, 2026 11:20am UTC

Request Review

@cursor

cursor Bot commented Sep 24, 2026 •

Copy link
Copy Markdown

PR Summary

Medium Risk
Changes the persona evaluation harness and resume format (bridge log required), but production Brunch chat paths are untouched; risk is mainly to persona runs and operator workflows, not core auth or data handling.

Overview
Replaces the Pi-only persona harness (extension, brunch_turn, isolated Pi config, verbosity/disclosure axes) with a launcher-owned JSONL socket bridge and per-run bin/persona helper (say, transcript, state, end, rpc). Any coding agent can drive the real browser composer via --agent claude|codex|cursor-agent|pi, --agent-command '{prompt}', or a printed launch prompt; agents start outside the repo so dev AGENTS.md/CLAUDE.md stay out of context.

Resume now reconciles against bridge-log.jsonl (admitted utterances + agent settings per turn) instead of Pi session files; Pi-era runs cannot resume. Persona briefing moves to persona-brief.md / launch/brief/system.md as a naive domain expert (no evaluation framing); case opening messages are sent whole-file without --- headers. Docs move to src/evaluations/persona/README.md; Brunch role settings no longer embed a default persona model in chat-model.ts.

Adds unit tests for RPC framing, bridge lifecycle, CLI parsing, agent launch, and brief assembly; removes Pi extension/integration tests.

Reviewed by Cursor Bugbot for commit 56bd8c0. Bugbot is set up for automated code reviews on this repo. Configure here.

@github-actions github-actions Bot added area/infra Relates to version control, CI, CD or IaC (area) area/libs Relates to first-party libraries/crates/packages (area) type/eng > frontend Owned by the @frontend team area/tests New or updated tests area/apps labels Sep 24, 2026
Comment thread apps/brunch-agent/src/evaluations/persona/launch/agent.ts Fixed

@cursor cursor Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread apps/brunch-agent/src/evaluations/persona/launch.ts Outdated
Comment thread apps/brunch-agent/src/evaluations/persona/launch/agent.ts
Comment thread apps/brunch-agent/src/evaluations/persona/browser-bridge.ts Outdated
@lunelson
lunelson requested a review from a team as a code owner September 25, 2026 13:17
@lunelson
lunelson force-pushed the ln/fe-1768-persona-cli branch from be7ff43 to 6e0fd3a Compare September 25, 2026 13:17

@cursor cursor Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread apps/brunch-agent/src/evaluations/persona/persona-cli.ts

@cursor cursor Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread apps/brunch-agent/src/evaluations/persona/persona-cli.ts
@lunelson
lunelson force-pushed the ln/fe-1768-persona-cli branch from df3add9 to 6963863 Compare September 29, 2026 07:43
@lunelson

Copy link
Copy Markdown
Contributor Author

Thanks, both fair, and both are addressed in 6963863.

  1. Isolation. The trade-off is intentional to a point. The launcher now drives any coding agent, so it can't impose one agent's lockdown flags the way the Pi extension did; the persona is asked to play a role rather than sandboxed. But it was weaker than it needed to be. The agent started in the run directory, which is inside the repository, so agents that load AGENTS.md or CLAUDE.md from parent directories, such as Pi and Claude Code, had Brunch's development guidance in context without reading anything. It now starts in a fresh temporary directory outside the repository and reaches its brief and helper by absolute path. It could still read the repository by path if it ignored the brief, and the operator's user-level agent instructions still apply, so a run remains a role-play observation rather than a sealed evaluation.
  2. Resume. Agreed. Resume continues the Brunch conversation with a fresh persona session, possibly on another agent or model; the earlier agent's session isn't resumed. The README and help now say so. Because run.json keeps only the latest agent, each persona entry in bridge-log.jsonl now records the agent settings that wrote it, so the turns of a resumed run can be attributed.

I checked both with a live launch: the agent started in an empty temporary directory, and the bridge-log entry for its utterance carried its settings.

Separately, on your caution about the model: #9811 now returns every Petrinaut assistant's default to gpt-5.5-2026-04-23 at medium, the original assistant's setting (b1e9d93), so your baseline doesn't move. GPT-6 Sol stays available in Brunch as an opt-in for testing.

Comment thread apps/brunch-agent/src/evaluations/persona/launch/agent.ts Dismissed
kostandinang
kostandinang previously approved these changes Sep 29, 2026

@kostandinang kostandinang left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks, this addresses both points for me. Approving.

@lunelson

Copy link
Copy Markdown
Contributor Author

A correction to my earlier answer on isolation: I framed it the wrong way round. The case material is the persona's own knowledge, so the persona reading it isn't a leak; the boundary that matters is that Brunch sees only what the persona types, and that holds by construction because Brunch declares no Flue sandbox and so has no file or shell tools. The temporary working directory stays, but only to keep the repository's development guidance out of the persona's context.

The real problem with the persona was the opposite one: it was too aligned with the interview. It knew it was in an evaluation of Brunch, followed a scripted arc (review, why, correction) and was coached to structure answers for the elicitor. 9dd069b and 7bbe63a brief it instead as a naive domain expert who knows their work but nothing about Brunch's aims or how an interview should go, answers the question asked briefly, volunteers nothing and stops by their own judgement; the case packs now describe the person rather than the test, and the verbosity and disclosure axes are gone. None of this has had a live persona run yet.

lunelson and others added 9 commits September 29, 2026 17:38
…er bridge CLI

Replace the Pi-only brunch_turn extension with a Pi-style JSONL bridge owned by the persona launcher and a per-run persona helper CLI. The launcher can start Claude Code, Codex, Cursor Agent, Pi or a custom command with a harness-neutral brief, in a Herdr pane when available or in the current terminal. Resume reconciles against a bridge log instead of the Pi session file.

Co-authored-by: Cursor <cursoragent@cursor.com>
The lower branch replaced a conversation's `{ mode, construction: { binding } }` initial data with `{ binding }`. This fixture only exists on the persona branch, so the restack left it writing the removed shape; it now writes what the launcher records.
The launcher marked a turn admitted only after appending the utterance to `bridge-log.jsonl`, so a failed log write read as a pre-admission failure even though Brunch already had the message, and the persona could resend it. The turn is now admitted first; a log failure after that is a failed admitted turn, which the persona is told never to resend.
…sona utterance to its agent

The agent started in the run directory, inside the repository, so agents that load AGENTS.md or CLAUDE.md from parent directories got Brunch's development guidance in the persona's context. It now starts in a fresh temporary directory, reaching the brief and helper by absolute path.

A resume starts a fresh persona session, possibly with another agent or model, and run.json keeps only the latest. Each persona entry in bridge-log.jsonl now records the agent settings that wrote it, and the README and help describe resume as continuing the Brunch conversation.
…on participant

Each situation pack now describes the person: who they are and how they talk, what they want, what they know about their operation, what they take for granted and what they don't know. The behavioural rules, gating such as "reveal only when asked", staged corrections, mission and condition notes, citations of deleted source files, and modelling knowledge a domain expert wouldn't have are gone; beliefs the case relies on are stated as the person's own. Each opening message is now just the person's first message, without an operator header.
The persona brief framed the agent as a participant in an evaluation of Brunch and coached it to serve the interview. Its role document now describes a person who knows their own work but nothing about Brunch's aims or how an interview should go: they answer the question asked, briefly, volunteer nothing, and stop by their own judgement. There is no default objective; an optional --objective is the person's own aim. The verbosity and disclosure axes are deleted, since each case describes how its person talks, and the opening message is the whole case file.

The README states the one boundary that matters: Brunch sees only what the persona types, while the case is legitimately the persona's own knowledge.
…never to repair the browser

Co-authored-by: Cursor <cursoragent@cursor.com>

@cursor cursor Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes and found 2 potential issues.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 4f6f52c. Configure here.

Comment thread apps/brunch-agent/src/evaluations/persona/launch/bridge-log.ts
Comment thread apps/brunch-agent/src/evaluations/persona/browser-bridge.ts

@kostandinang kostandinang left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Non-blocking: this changes the persona and every case at once, so runs from before and after aren't comparable. Could run.json record a hash of the case and brief, like it already does for --initial-net?

This branch was successfully deployed

1 active (outdated) and 2 inactive deployments
Preview – petrinaut — 56bd8c0a Deployed Sep 30, 2026 by vercel[bot]
Preview – petrinaut-docs — 56bd8c0a Deployed Sep 30, 2026 by vercel[bot]
Preview – hash — 4f6f52ca Deployed Sep 29, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area/apps area/infra Relates to version control, CI, CD or IaC (area) area/libs Relates to first-party libraries/crates/packages (area) area/tests New or updated tests type/eng > frontend Owned by the @frontend team

Development

Successfully merging this pull request may close these issues.

3 participants