Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
19 changes: 0 additions & 19 deletions .github/actions/prune-repository/action.yml
Original file line number Diff line number Diff line change
Expand Up @@ -16,25 +16,6 @@ runs:
read -r -a scopes <<< "${SCOPE//$'\n'/ }"
turbo prune "${scopes[@]}"

# `turbo prune` copies the workspace directories and the root manifests; paths a package
# reads from outside any workspace are copied here.
- name: Copy paths outside the workspaces
shell: bash
env:
SCOPE: ${{ inputs.scope }}
run: |
copy() { mkdir -p "out/$(dirname "$1")" && cp -R "$1" "out/$1"; }
read -r -a scopes <<< "${SCOPE//$'\n'/ }"
for scope in "${scopes[@]}"; do
case "$scope" in
# The app's product tests execute evaluation runners and inspect committed evidence.
@apps/brunch-agent)
copy libs/@hashintel/brunch-agent/docs
copy libs/@hashintel/brunch-agent/evaluations
;;
esac
done

- name: Copy required files
shell: bash
run: |
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -22,7 +22,7 @@ yarn brunch:persona --case inventory-purchasing \

`--persona-verbosity` accepts `terse`, `default`, or `expansive`; `--persona-disclosure` accepts `reticent`, `default`, or `forthcoming`. Each non-default setting overrides only that axis in the situation pack. The default leaves the pack's axis unchanged. Verbosity controls answer length and response effort; disclosure controls how readily relevant knowledge is volunteered. Neither changes the person's other traits, reveals private material, merges the actor with the elicitor, or asks the actor to help the interview succeed. Reticence is not hostility, feigned ignorance, or permission to withhold a directly requested answer.

`--help` lists every flag and exact literal. Each role requires its selected provider's API key: `OPENAI_API_KEY` for OpenAI and `ANTHROPIC_API_KEY` for Anthropic. The defaults therefore require both; an all-OpenAI run does not require Anthropic credentials. The launcher transfers the selected persona credential privately to its Pi pane. Paid runs still require owner authorization under the current [mission](../../../../../libs/@hashintel/brunch-agent/MISSION.md) and [execution safety](../../../../../libs/@hashintel/brunch-agent/evaluations/README.md#execution-safety); the existence of this command grants none.
`--help` lists every flag and exact literal. Each role requires its selected provider's API key: `OPENAI_API_KEY` for OpenAI and `ANTHROPIC_API_KEY` for Anthropic. The defaults therefore require both; an all-OpenAI run does not require Anthropic credentials. The launcher transfers the selected persona credential privately to its Pi pane. A live-provider run requires explicit model and spend authorization; the existence of this command grants none. See the concise [persona-testing overview](../../../../../libs/@hashintel/brunch-agent/EVALUATIONS.md).

**Persona runs have no automatic accounting cutoff.** The launcher disables the campaign accounting wrapper even if `BRUNCH_STEP_A_ACCOUNTING` was inherited. Pi uses its native provider. There are no request reservations, budget/unknown-usage refusals, or `--budget-usd` / `--accept-unknown` flags. Usage remains observational in the native records below; missing usage is not zero cost. There is no fixed turn-count limit. Use Ctrl-C to stop the run.

Expand Down Expand Up @@ -103,11 +103,11 @@ Older runs may also contain `usage-ledger.json` and `attempt-ledger.md`. Leave t

## Verification and implementation

For Anthropic schema acceptance before paid observation after a tool/schema/adapter change, run `yarn workspace @apps/brunch-agent test:anthropic-tools` from the HASH root. It rebuilds Brunch, captures its native tool catalogues and checks acceptance through Anthropic's free token-counting API with a synthetic message. It requires the normal development credential but performs no generation, sends no case data and does not settle unknown spend. This is not OpenAI acceptance. The [schema acceptance contract](../../../../../libs/@hashintel/brunch-agent/evaluations/README.md#tool-schema-acceptance) owns coverage and limitations.
For Anthropic schema acceptance after a tool, schema, or adapter change, run `yarn workspace @apps/brunch-agent test:anthropic-tools` from the HASH root. It rebuilds Brunch, captures its native tool catalogues, and checks acceptance through Anthropic's free token-counting API with a synthetic message. It requires the normal development credential but performs no generation and sends no case data. Passing proves only that the captured schemas are accepted by that endpoint; it does not prove generation quality or OpenAI compatibility.

`test/persona-construction.integration.ts` uses the actual opening helper, registered Pi extension, local socket and ordinary composer against the built ChatAgent and real Chrome with a synthetic provider. It checks empty start, opening-tool continuation, repeated workpiece/net updates, tab switching during a continuation, cancellation and no replay on reload. It also restarts the backend after an aborted turn, reconciles without sending, retains the net/workpiece and executes a new browser-tool turn in the original conversation; mismatched utterances and browser principals refuse. It establishes mechanism viability, not persona fidelity, construction quality, crash recovery at every boundary or an accepted worked example.

Add `--openai` to `yarn workspace @apps/brunch-agent test:persona` for the same proof through the registered OpenAI provider at low effort. The native Responses serializer and SSE parser remain real; only HTTP responses are synthetic. Each request checks the mounted tools' schemas/descriptions, `strict: false`, model and effort; captured `openai-requests.json` includes browser-result history. Run under the evaluation guide's loopback-only network guard (which also permits the private persona Unix socket). Passing is synthetic wiring evidence, not OpenAI server acceptance or a live-model result.
Add `--openai` to `yarn workspace @apps/brunch-agent test:persona` for the same proof through the registered OpenAI provider at low effort. The native Responses serializer and SSE parser remain real; only HTTP responses are synthetic. Each request checks the mounted tools' schemas/descriptions, `strict: false`, model and effort; captured `openai-requests.json` includes browser-result history. Passing is synthetic wiring evidence, not OpenAI server acceptance or a live-model result.

The construction proof holds the recording pause and checks that no submission occurs before release. `test/persona-extension-lifecycle.test.ts`, enabled with `PI_PERSONA_CLI=$(command -v pi)`, crosses the installed Pi's flag hydration and tool-registration boundary with a synthetic socket reply and no inference. After building Brunch, `node --experimental-strip-types test/provider-accounting.integration.ts --disabled` checks that native requests proceed with an unusable historical ledger, preserve it untouched and retain usage in the original database. These checks do not prove live-model fidelity or successful generation with the operator's credential.

Expand All @@ -117,4 +117,4 @@ From the HASH root, build and run the synthetic browser proof:
yarn workspace @apps/brunch-agent test:persona
```

The launcher composes the [Pi extension](../brunch-persona-testing.ts), [private IPC bridge](../../../src/evaluations/persona/browser-bridge.ts), [browser turn](../../../src/evaluations/persona/browser-turn.ts) and [evidence writer](../../../src/evaluations/persona/proof-artifacts.ts). The extension requires the launcher bridge; evidence retention and browser identity belong to the launcher. Independent headless construction probes under `src/evaluations/runbook/` are not persona launch methods.
The launcher composes the [Pi extension](../brunch-persona-testing.ts), [private IPC bridge](../../../src/evaluations/persona/browser-bridge.ts), [browser turn](../../../src/evaluations/persona/browser-turn.ts) and [evidence writer](../../../src/evaluations/persona/proof-artifacts.ts). The extension requires the launcher bridge; evidence retention and browser identity belong to the launcher.
22 changes: 3 additions & 19 deletions apps/brunch-agent/AGENTS.md
Original file line number Diff line number Diff line change
@@ -1,23 +1,7 @@
# Brunch agent application

This application belongs to the Brunch context rooted at
`../../libs/@hashintel/brunch-agent/`. Read that context's `AGENTS.md` and current `MISSION.md`
before changing this application. Read `MISSION.next.md` when work affects future sequence,
cross-mission constraints, open product decisions, or re-entry gates; it is the canonical future
spine, not execution authority. Consult `CONTEXT.md` or historical design documents only when a
concrete vocabulary or rationale question requires them; ADRs and specs are hypotheses, not
implementation obligations.
HASH root guidance takes precedence.
HASH root guidance applies. The Brunch package map and implemented invariants are documented in [`../../libs/@hashintel/brunch-agent/README.md`](../../libs/@hashintel/brunch-agent/README.md) and [`../../libs/@hashintel/brunch-agent/ARCHITECTURE.md`](../../libs/@hashintel/brunch-agent/ARCHITECTURE.md).

The application composes the Flue runtime, HTTP routes, and the Brunch packages required by the
current mission. It must remain independent of Petrinaut UI (`@hashintel/petrinaut`); it may import
published catalogs from `@hashintel/petrinaut-core` (for example user-guide page ids the panel
already executes). `apps/petrinaut-website` meets the editor through the AI SDK/HTTP transport.
This application composes the Flue runtime, HTTP routes, and Brunch packages. It must remain independent of Petrinaut UI (`@hashintel/petrinaut`); it may import published catalogs from `@hashintel/petrinaut-core`. `apps/petrinaut-website` meets the editor through the AI SDK/HTTP transport.

`@earendil-works/pi-ai@0.83.0` is patched at the repo root so Anthropic
`input_schema` keeps the published tool JSON Schema. Pi's adapter still
collapses parameters to `{ type, properties, required }` by design; Brunch
construction tools need the full schema on the wire, independent of
constrained sampling. `@flue/runtime` depends on `pi-ai@^0.83.0`, which is
why the caret resolution exists. Re-evaluate the patch on any `pi-ai`
upgrade. Do not treat it as a general Pi behavior change.
The repository patches `@flue/runtime@2.0.3`, `@earendil-works/pi-agent-core@0.83.0`, and `@earendil-works/pi-ai@0.83.0`. Brunch depends on their combined context-projection, argument-validation, and complete-tool-schema behavior. Re-evaluate the full patch boundary on a Flue or Pi upgrade.
8 changes: 0 additions & 8 deletions apps/brunch-agent/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -21,14 +21,6 @@ yarn brunch:persona --case inventory-purchasing

The launcher opens a dedicated Chrome window, pauses for recording readiness, then drives the real panel with a private Pi persona. Brunch's own tool calls update the visible net and workpiece. Select any listed case or a directory containing `situation-pack.md` and `opening-message.md`. There is no automatic budget cutoff; native usage is retained. The guide owns prerequisites, stop/resume and evidence instructions; consult it before paid execution.

The independent headless runbook construction probe is not a persona launch method:

```sh
yarn workspace @apps/brunch-agent runbook:headless
```

`ANTHROPIC_API_KEY` is required. `BRUNCH_CHAT_MODEL` selects the interviewer (default `claude-sonnet-4-5` for this script only). Artifacts write under `apps/brunch-agent/.data-wipe-me/evaluations/vestera-runbook-headless/` unless `BRUNCH_RUNBOOK_OUTPUT_DIR` is set. The command prints the resulting path. Do not promote that directory into the repository.

By default outside production, conversations persist in SQLite at `apps/brunch-agent/.data-wipe-me/conversations.db`. `BRUNCH_DEV_DB_PATH` overrides that local path. The hermetic browser-transport test uses `BRUNCH_CHAT_DB_PATH` to point at its own sqlite file. Flue history is the conversation log. The panel rehydrates from the SDK's canonical conversation observation and does not resubmit or replay settled turns.

The mounted Flue URL `/agents/chat/:instanceId` requires the principal and logical conversation identity in `x-brunch-principal` and `x-brunch-conversation`. The path id is the hash of those values, not a bearer token or trusted authentication.
Expand Down
1 change: 0 additions & 1 deletion apps/brunch-agent/package.json
Original file line number Diff line number Diff line change
Expand Up @@ -18,7 +18,6 @@
"petrinaut:dev": "vite dev --config petrinaut-local.vite.config.ts",
"probe:rds-iam": "node --experimental-strip-types src/rds-iam-probe.ts",
"proof:manifest": "node --experimental-strip-types src/evaluations/persona/refresh-proof-manifest.ts",
"runbook:headless": "vite build && node --experimental-strip-types src/evaluations/runbook/construction-run.ts",
"smoke:deployment": "node --experimental-strip-types src/deployment-smoke.ts",
"start": "PORT=3002 node dist/server.mjs",
"start:healthcheck": "wait-on --timeout 1200000 http-get://localhost:3002/health",
Expand Down
2 changes: 1 addition & 1 deletion apps/brunch-agent/src/conversation/net-ledger.ts
Original file line number Diff line number Diff line change
Expand Up @@ -11,7 +11,7 @@
*
* Freshness folds over these events. Whether why moves onto them too is
* decided by a parity test against its own attribution walk, not assumed
* here (see the shared-history-projection fog-line in MISSION.md).
* here; projections must remain recomputable from canonical history.
*/
import {
canonicalContent,
Expand Down
Loading
Loading