Skip to content

Bind the socket before spawning the roster. - #20

Closed
im-ian wants to merge 1 commit into
mainfrom
fix-daemon-binds-before-it-spawns
Closed

im-ian wants to merge 1 commit into
mainfrom
fix-daemon-binds-before-it-spawns

Conversation

@im-ian

@im-ian im-ian commented Sep 21, 2026

Copy link
Copy Markdown
Owner

TL;DR

데몬이 소켓을 bind 하기 전에 봇을 전부 spawn 해서, 그 사이에 두 번째 데몬이 들어와 같은 로스터를 또 띄울 수 있었어요. bind 를 맨 앞으로 옮기고, "준비됨" 신호를 따로 뒀어요.

원인

daemon::run() 의 순서가 이랬어요:

is_socket_live()?  -> no          ← 두 번째 데몬도 여기서 통과
remove_stale_socket()
Config::load()
open_agent() × 전체 로스터        ← 수 초. 이 구간 내내 소켓이 없음
UnixListener::bind()              ← 여기서 처음 경쟁

로스터가 큰 경우 그 구간이 몇 초예요. client::ensure_daemon 은 4초 deadline 안에 소켓이 안 보이면 실패로 치고 빠지고, 다음 호출이 또 들어와요. 두 데몬이 각자 모든 봇을 같은 cwd 에 spawn 하고 같은 transcript 에 써요.

bind 에서 진 쪽은 ? 로 바로 return 해서 shutdown_agents() 에 도달하지 못해요. 그래서 죽은 데몬의 자식 CLI 들이 그대로 남아 이긴 데몬의 파일에 계속 써요.

고친 것

bind 를 맨 앞으로 — 경로에 대한 소유권 주장이 곧 bind 예요. 두 번째 데몬은 아무것도 열기 전에 bail 해요.

"준비됨"을 별도 신호로 — bind 를 앞당기면 반대 구간이 생겨요: 묶였지만 아직 accept 안 함. 여기서 ensure_daemon 이 반환하면 호출자는 연결은 되는데 응답은 없는 소켓을 받아요 (데스크톱에선 메인 스레드 command 가 read 에서 멈춤).

  • daemon.version 은 accept 루프가 뜬 뒤 쓰고, 부팅 시작 때 지워요 → "지금 서비스 중" 을 뜻하게 돼요.
  • paths::daemon_is_ready() = 버전 파일 + 소켓.
  • ensure_daemon 은 ready 를 기다려요. 소켓은 살아 있는데 버전 파일이 없으면 = 누가 부팅 중 → 판단하지 말고 기다려요. 기존 코드는 없는 파일을 "버전 불일치" 로 읽고 부팅 중인 데몬을 죽였어요.

teardown 을 run() 으로 — accept 루프 꼬리에 있던 정리를 밖으로 빼서, 어떤 경로로 끝나든 소켓·pid·버전을 지워요. 기존엔 정상 종료에서만 돌아서, 부팅 실패가 "죽은 데몬을 가리키는 살아있는 소켓" 을 남겼어요.

검증

  • cargo test --locked -p crew — 213 passed
  • a_daemon_leaves_nothing_behind_when_it_stops — 실제 데몬을 띄우고 Shutdown 보낸 뒤 소켓/pid/버전 부재 확인. 뮤테이션 확인: teardown 3줄을 지우면 FAILED
  • a_bound_socket_stops_a_second_daemon_before_it_spawns — bind 만 된(accept 전) 소켓도 is_socket_live 가 yes 로 읽는다는, 이 수정이 기대는 전제를 고정
  • cargo fmt --check — 새 코드 clean
  • 실제 실행:
cold start (ensure_daemon 경유)  -> ready-file:YES socket:YES version=0.2.0
crew.log "listening" 1회
crew stop                        -> socket:NO pid:NO version:NO

`daemon::run` opened every agent and only then bound the socket, so the
whole boot sat inside a window where `is_socket_live` still said no
daemon was coming. A second one started in that window — by
`ensure_daemon`, or by hand — passed the same check, spawned a duplicate
of every agent into the same cwd, and raced for the same path. The loser
returned through `?` from the bind, never reaching `shutdown_agents`,
and left its children writing to the winner's transcripts.

Binding first is the claim on the path, so the second daemon now bails
before it opens anything.

That leaves a window of the opposite kind: bound, but not yet accepting.
Returning from `ensure_daemon` on the bind would hand back a socket that
takes the connection and answers nothing until the roster is open — on
the desktop that is a main-thread command blocking on a read. So
readiness is its own signal. `daemon.version` is written when the accept
loop is up and cleared at boot, which makes it mean "serving" rather
than "was serving at some point", and `daemon_is_ready` pairs it with
the socket. A caller that finds a bound socket with no version file
waits for the boot in progress instead of judging its version — the old
code read the missing file as a mismatch and stopped a daemon that was
still starting.

Teardown moved out of the accept loop's tail into `run`, so every way
out clears the socket, the pid and the version. It only ran on a clean
shutdown before, which is how a failed boot could leave a socket that
answers `is_socket_live` for a daemon that is gone.
@im-ian im-ian mentioned this pull request Sep 21, 2026
5 tasks done
@im-ian
im-ian marked this pull request as draft September 21, 2026 04:57
@im-ian

im-ian commented Sep 21, 2026

Copy link
Copy Markdown
Owner Author

리뷰 결과 머지 보류. 이 PR 의 핵심 주장("bind 가 소유권 주장")이 성립하지 않아요.

  • remove_stale_socket() 이 bind 직전에 무조건 unlink 해서, 두 번째 데몬이 첫 번째의 살아있는 소켓을 지우고 새 inode 에 bind 할 수 있어요. 경쟁 구간이 좁아질 뿐 없어지지 않아요.
  • is_socket_live() 의 의미가 "serving" → "bound" 로 바뀌었는데 ensure_daemon 만 갱신했어요. main.rs 의 ~15개 if paths::is_socket_live() 게이트가 부팅 중 온라인 분기를 타고, rpc 에 read timeout 이 없어서 영원히 블록돼요.
  • 새 테스트가 run() 을 호출해 전역 EVENTS 를 초기화하고 ticker 스레드 3개를 영구 누수시켜요. with_home 이 임시 홈을 지운 뒤 그 스레드들이 사용자의 실제 ~/Library/Application Support/crew 에 씁니다.

근본 원인은 따로 있어요: 데몬 기동에 상호배제가 없고, client::rpc 에 read timeout 이 없어요. 그거 없이 bind 순서만 바꾸면 증상만 옮겨다녀요.

@im-ian

im-ian commented Sep 21, 2026

Copy link
Copy Markdown
Owner Author

닫습니다. #23 재설계 과정에서 확인한 결과, 이 PR 의 전제(bind 순서를 바꾸면 소유권이 성립한다)가 틀렸어요 — remove_stale_socket() 이 bind 직전에 무조건 unlink 하는 한 bind 는 claim 이 아니에요.

계측으로 실제 메커니즘을 확정했고(#23 코멘트 참조), 측정된 피해가 가정보다 훨씬 작아요. 살릴 조각은 serve() 분리(부팅 실패 시 로스터 teardown)뿐이라, 그것만 별도 PR 로 냅니다.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant