Skip to content

Tear down a roster a boot abandons. - #24

Closed
im-ian wants to merge 1 commit into
mainfrom
tear-down-a-roster-a-boot-abandons
Closed

im-ian wants to merge 1 commit into
mainfrom
tear-down-a-roster-a-boot-abandons

Conversation

@im-ian

@im-ian im-ian commented Sep 21, 2026

Copy link
Copy Markdown
Owner

TL;DR

부팅이 로스터를 전부 열고 나서 bind 에서 실패하면, 그 로스터를 아무도 정리하지 않았어요. teardown 이 accept 루프 맨 끝에 있어서 ? 로 빠지는 경로가 거기 도달하지 못했어요.

오늘 관측되는 효과는 없어요. 이건 수정이 아니라 불변식이에요.

어떻게 확인했는지

계측 바이너리로 경쟁 시 실제 순서를 확정했어요:

d1: is_socket_live=false, exists=false → unlink(no-op) → 20 agents 로드
    → 로스터 전체 open → about to bind; exists=true  → EADDRINUSE → `?` 로 사망
d2: 동일 경로 → about to bind; exists=false → 서비스 시작

검사와 bind 가 하나의 연산이 아니어서, 두 데몬이 동시에 시작하면 둘 다 "소켓 없음"을 보고 둘 다 로스터를 열어요.

왜 오늘은 효과가 없는지 (그리고 왜 그래도 넣는지)

  • 실사용 봇은 헤드리스 — uses_pty_program 이 grok/claude/codex 를 헤드리스로 분류해요. 측정: 헤드리스 봇 6개 부팅 → registered 6줄, CLI 프로세스 0개. 진 데몬이 띄우는 게 없어요.
  • PTY 봇의 자식은 데몬과 함께 죽어요 (데몬 쪽 pty 가 닫히며 SIGHUP). 측정: 봇 6개 경쟁 후 남은 프로세스 6개(승자 몫), 12개 아님. main 바이너리로도 실패 부팅 뒤 0개.
  • 갓 부팅한 데몬의 seal_all_now() 는 아무것도 안 써요. 기존 대화 파일을 심어두고 실패 부팅을 돌린 before/after 둘 다 transcript·channel 파일 무변경.

그래서 이 PR 은 "serve 가 시작한 건 run 이 끝낸다"는 불변식을 세우는 것뿐이에요. 지금은 PTY hangup 이라는 OS 동작에 기대 무해한 상태이고, 부팅에 지속적 side effect 가 하나라도 추가되면 그때 다시 발견해야 하는 종류의 문제예요.

에이전트만 정리해요. 소켓·pid·버전 파일은 "지금 서비스하는 쪽"을 가리키고, 이 경로에서 그건 이긴 데몬이에요. 여기서 지우면 살아있는 소켓을 unlink 해요 — 이전 시도들에서 리뷰가 반복해 지적한 함정이라 명시적으로 피했어요.

테스트를 안 붙인 이유

효과가 관측되지 않으니 테스트도 관측할 게 없어요. "실패 부팅 뒤 봇 프로세스 0개" 를 단정하는 테스트는 main 에서도 통과해요 (SIGHUP 이 해 주니까) — 통과하는 척하는 테스트가 돼요. run() 을 in-process 로 돌려 내부 상태를 보는 방법은 join 되지 않는 ticker 스레드 3개를 실제 홈에 남겨서 더 나빠요.

남기는 것

레이스 자체는 README 개발 섹션에 알려진 한계로 적었어요. 닫으려면 검사와 claim 이 한 연산이어야 하는데, 크래시 복구에 필요한 stale 소켓 unlink 가 bind 를 claim 이 아니게 만들어요.

이 결론에 이르기까지 #20·#23 에서 여섯 번 시도했고, 그 과정에서 제가 넣은 결함이 넷이었어요 (kill(-1), 방금 띄운 데몬 자살, kill(0), 무한 재귀). #20·#23 은 근거와 함께 닫았어요.

검증

  • cargo test --locked -p crew — 214 passed
  • cargo fmt --check — 변경한 코드 clean
  • 실패 부팅 결정적 재현 (소켓 경로에 디렉터리 → remove_stale_socket 은 remove_file 이라 못 지움 → bind 반드시 실패): 로스터 5개 열고 실패, 남은 프로세스 0
  • 정상 기동·정지 회귀 없음

`serve` opens every agent before it binds, so a daemon that loses a
start race — two can both pass the socket check while neither has bound
— has the whole roster open by the time `bind` tells it so. That error
returned straight out of `run`, and the teardown sat at the far end of
the accept loop, so nothing killed what the boot had started.

Nothing observable, today. The boot of a real roster spawns no processes
at all: `uses_pty_program` makes `grok`, `claude` and `codex` headless,
and a headless agent is registered rather than run. A PTY roster does
spawn one process per bot, and those die with the daemon anyway when its
end of the pty closes — measured, six bots raced, six left afterwards,
not twelve. So this is the invariant, not a fix: whatever `serve`
starts, `run` ends, and the next side effect added to boot does not have
to rediscover that.

Only the agents. The socket, pid and version name whoever is serving,
and on this path that is the daemon that won — clearing them here would
unlink its live socket.

The race itself stays, documented in the README. Closing it needs the
check and the claim to be one operation, and the stale-socket unlink
that crash recovery depends on is what stops `bind` from being that.
@im-ian

im-ian commented Sep 21, 2026

Copy link
Copy Markdown
Owner Author

닫습니다. 측정으로 이 PR 이 main 에 없는 data loss 를 만든다는 걸 확인했어요.

로스터를 열고 bind 에서 실패하는 부팅 (소켓 경로에 디렉터리를 두어 결정적으로 재현)
  main : cli-sessions 파일 2개 → 2개
  #24  : cli-sessions 파일 2개 → 0개

경로: shutdown_agents() → LiveAgent::kill → headless::kill → clear_session() → fs::remove_file(cli_session_path(id)). 그 파일이 봇의 --resume 세션 id 이고, save_session 은 턴이 끝난 뒤에만 다시 써요. 그래서 다음 재시작 전에 말을 걸지 않은 봇은 대화 연속성을 영구히 잃어요.

제 PR 본문의 검증은 프로세스 수·transcript·channel 파일만 봤고 cli-sessions/ 를 안 봤어요. README 가 문서화한 상태인데 빠뜨렸어요.

리뷰의 다른 지적도 맞아요:

  • clean 종료 순서를 뒤집었어요. shutdown_agents() 를 run() 으로 옮기면서 소켓·pid 제거와 Event::Shutdown 이 에이전트 kill 보다 먼저 일어나요. main 은 kill 이 먼저였어요. 그래서 wait_dead 가 구 데몬이 로스터를 쥔 채로 반환해요 — 흔한 경로(버전 업데이트 재시작)의 회귀예요.
  • bind 이후 실패 경로(write_pid/write_daemon_version)에선 소켓·pid 가 우리 것인데 정리를 건너뛰어요. 제 주석의 "이 경로에선 남의 것" 은 bind 이전에만 참이에요.
  • README 의 "닿을 수 없는 데몬은 아무것도 안 들고 있다" 는 틀렸어요. bind 에 도달했다는 건 정의상 로스터를 이미 다 열었다는 뜻이에요. 제 측정은 "빈 로스터에서만 재현된다" 였는데, 그걸 "들고 있는 게 없다" 로 옮겨 적었어요 — 빈도와 내용물을 혼동했어요.

이 레이스에 대해 저는 #20·#23·#24 로 일곱 번 시도했고, 매번 shipped 코드보다 나쁜 걸 넣었어요. 코드는 더 안 건드립니다.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant