Skip to content

Stop waiting forever on a daemon that stopped answering. - #22

Merged
im-ian merged 2 commits into
mainfrom
rpc-does-not-wait-forever
Sep 21, 2026
Merged

im-ian merged 2 commits into
mainfrom
rpc-does-not-wait-forever

Conversation

@im-ian

@im-ian im-ian commented Sep 21, 2026 •

Copy link
Copy Markdown
Owner

TL;DR

client::rpc 에 read timeout 이 없었어요. 연결은 받아놓고 답을 안 하는 데몬을 만나면 호출자가 영원히 멈춰요. #20/#21 리뷰에서 나온 "영원히 블록" 시나리오들이 전부 여기서 나와요.

원인

let n = reader.read_line(&mut buf)?;   // set_read_timeout 없음

코드베이스 어디에도 set_read_timeout / set_write_timeout 이 없어요. 연결이 끊기면 EOF 라도 오지만, 연결은 살아 있는데 답이 없는 경우는 영구 대기예요:

  • accept 루프가 락에서 물린 데몬
  • connect 와 응답 사이에 죽은 데몬

고친 것

요청 전체에 deadline — set_read_timeout 은 read_line 한 번만 제한해요. 루프는 이 요청의 답이 아닌 이벤트를 continue 로 건너뛰는데, 그때마다 시계가 다시 감겨요. 초당 무관한 이벤트 한 줄씩 흘리는 데몬은 호출자를 영원히 붙잡을 수 있어요. deadline 을 한 번 잡고, 매 read 전에 남은 시간만 설정해요.

요청 deadline 이유
Ask 없음 사람 답을 기다려요
Subscribe 없음 질문이 아니라 데몬이 계속 밀어주는 스트림이에요
Ping 900ms 창이 1초마다 폴링하는 probe. 그 안에 포기해야 쌓이지 않아요
나머지 15초 데몬이 자기 상태만 보고 답해요. 그 이상은 바쁜 게 아니라 물린 거예요

Request 를 exhaustive 로 매칭 — catch-all 은 나중에 추가되는 모든 변형에 15초 단두대를 물려요. Subscribe 가 실제로 그 상태였어요 (어떤 arm 에도 안 걸려서 status ticker 의 frame 이 전부 _ => continue 로 감). exhaustive 면 다음 장기 요청이 스스로 밝혀야 해요.

쓰기 경로도 이름을 붙여요 — read 경로가 설명해주는 것과 같은 상황인데 os error 35 로 나왔고, 그게 창에 그대로 들어가요. 둘 다 OS 에러를 cause 로 보존해서 EAGAIN/ETIMEDOUT 이 구분돼요.

deadline 을 connect() 로 — pub 이라 다른 소비자에게 무기한 소켓을 넘겨줬어요. rpc_within 이 요청별로 read 예산만 좁혀요.

빈 줄은 skip — 데몬의 reader 가 그렇게 해요. 기존엔 프로토콜 에러로 보고했어요.

검증

  • cargo test --locked -p crew — 214 passed
  • a_socket_that_never_answers_gives_up — accept 만 하고 침묵하는 리스너. worker + deadline 으로 실행해서, 회귀 시 멈추는 게 아니라 실패해요
  • 뮤테이션 확인: 타임아웃 설정을 빼면 → FAILED ... rpc must return, not park: Timeout, 5.00초에 실패
  • a_stream_is_not_a_question — Subscribe 는 None, PING_TIMEOUT < 1s
  • a_person_gets_as_long_as_they_need_and_nobody_else_does
  • cargo fmt --check — client.rs clean
  • with_home 의 임시 경로 단축 — 바인딩하는 소켓이 stock TMPDIR 에서 104바이트 한계에 몇 바이트 남기고 있었어요

정정: 이 PR 초판 본문에 "타임아웃을 빼면 테스트가 영원히 안 끝난다 (timeout 30 으로 강제 종료)" 라고 썼는데, macOS 에 timeout 이 없어서 그 확인 자체가 무효였어요. 위 5초 실패가 제대로 측정한 결과예요.

범위

이 PR 은 무한 대기를 없애는 것까지예요. 창이 안 멈추게 하는 건 #21 (async command) 의 일이에요 — 지금 desktop.rs 의 command 는 전부 sync 라 메인 스레드에서 돌고, 창은 200ms/1초로 폴링해요. 그쪽이 들어가야 "느린 데몬에도 창이 계속 그려진다" 가 성립해요.

선행 2건 중 하나이고, 나머지는 #23 (데몬 기동 락) 이에요.

`client::rpc` set no read timeout, so a socket that took the connection
and then went quiet parked the caller for good. That is not a
hypothetical shape: a daemon whose accept loop is wedged on a lock, or
one killed between the connect and the reply, leaves the connection open
and the read blocked.

The callers are the ones that make it matter. Desktop commands run on
the main thread, and the window polls two of them every second — a
single parked read takes the UI with it, with no error and nothing in
the log.

Requests now carry a deadline. `Ask` keeps none, because it waits on a
person and no guess at that is worth making; everything else the daemon
answers out of its own state, so silence past fifteen seconds means
wedged rather than busy. `Ping` gives up in five, since it is the one
the window polls and the one whose whole job is to notice.

Writes get a deadline too. They are one short line into a socket buffer,
so taking five seconds means nobody is draining it.
@im-ian im-ian mentioned this pull request Sep 21, 2026
6 tasks done
From review of #22.

`set_read_timeout` bounds one `read_line`, and the loop skips events
that do not answer the request — so the clock re-armed on every event
and a daemon dribbling one unrelated line a second could hold a caller
indefinitely. `Subscribe` reaches that today: it matches no arm, so
every frame the status ticker pushes goes through `_ => continue`. The
budget is now a deadline taken once, with the time left set before each
read.

`answer_timeout` matches `Request` exhaustively. The catch-all handed a
fifteen-second guillotine to anything added later, `Subscribe` included
— a stream is not a question. Making it exhaustive means the next
long-running request has to say which it is.

The write path is named the same way the read path is. It came back as a
bare `os error 35` for the same condition the read path explains, and
that message was going into the window. Both keep the OS error as the
cause, so EAGAIN and ETIMEDOUT stay apart.

Deadlines moved into `connect`, which is public and handed out untimed
sockets; `rpc_within` narrows the read budget from there.

`PING_TIMEOUT` is under a second. It is the probe the window polls every
second, and at five it stacked five deep before reporting anything.

A blank line is skipped rather than reported as a protocol error, which
is what the daemon's own reader does.

The regression test runs on a worker with a deadline. It guarded against
blocking forever by blocking forever; it now fails in five seconds with
`rpc must return, not park`. The temp home is shorter too — the socket
it binds was within a few bytes of the 104-byte limit on a stock TMPDIR.
@im-ian
im-ian merged commit c321256 into main Sep 21, 2026
2 checks passed
@im-ian
im-ian deleted the rpc-does-not-wait-forever branch September 21, 2026 05:16
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant