Repository navigation
Stop waiting forever on a daemon that stopped answering. - #22
Merged
Merged
Conversation
`client::rpc` set no read timeout, so a socket that took the connection and then went quiet parked the caller for good. That is not a hypothetical shape: a daemon whose accept loop is wedged on a lock, or one killed between the connect and the reply, leaves the connection open and the read blocked. The callers are the ones that make it matter. Desktop commands run on the main thread, and the window polls two of them every second — a single parked read takes the UI with it, with no error and nothing in the log. Requests now carry a deadline. `Ask` keeps none, because it waits on a person and no guess at that is worth making; everything else the daemon answers out of its own state, so silence past fifteen seconds means wedged rather than busy. `Ping` gives up in five, since it is the one the window polls and the one whose whole job is to notice. Writes get a deadline too. They are one short line into a socket buffer, so taking five seconds means nobody is draining it.
6 tasks done
From review of #22. `set_read_timeout` bounds one `read_line`, and the loop skips events that do not answer the request — so the clock re-armed on every event and a daemon dribbling one unrelated line a second could hold a caller indefinitely. `Subscribe` reaches that today: it matches no arm, so every frame the status ticker pushes goes through `_ => continue`. The budget is now a deadline taken once, with the time left set before each read. `answer_timeout` matches `Request` exhaustively. The catch-all handed a fifteen-second guillotine to anything added later, `Subscribe` included — a stream is not a question. Making it exhaustive means the next long-running request has to say which it is. The write path is named the same way the read path is. It came back as a bare `os error 35` for the same condition the read path explains, and that message was going into the window. Both keep the OS error as the cause, so EAGAIN and ETIMEDOUT stay apart. Deadlines moved into `connect`, which is public and handed out untimed sockets; `rpc_within` narrows the read budget from there. `PING_TIMEOUT` is under a second. It is the probe the window polls every second, and at five it stacked five deep before reporting anything. A blank line is skipped rather than reported as a protocol error, which is what the daemon's own reader does. The regression test runs on a worker with a deadline. It guarded against blocking forever by blocking forever; it now fails in five seconds with `rpc must return, not park`. The temp home is shorter too — the socket it binds was within a few bytes of the 104-byte limit on a stock TMPDIR.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
TL;DR
client::rpc에 read timeout 이 없었어요. 연결은 받아놓고 답을 안 하는 데몬을 만나면 호출자가 영원히 멈춰요. #20/#21 리뷰에서 나온 "영원히 블록" 시나리오들이 전부 여기서 나와요.원인
코드베이스 어디에도
set_read_timeout/set_write_timeout이 없어요. 연결이 끊기면 EOF 라도 오지만, 연결은 살아 있는데 답이 없는 경우는 영구 대기예요:고친 것
요청 전체에 deadline —
set_read_timeout은read_line한 번만 제한해요. 루프는 이 요청의 답이 아닌 이벤트를continue로 건너뛰는데, 그때마다 시계가 다시 감겨요. 초당 무관한 이벤트 한 줄씩 흘리는 데몬은 호출자를 영원히 붙잡을 수 있어요. deadline 을 한 번 잡고, 매 read 전에 남은 시간만 설정해요.AskSubscribePingRequest를 exhaustive 로 매칭 — catch-all 은 나중에 추가되는 모든 변형에 15초 단두대를 물려요.Subscribe가 실제로 그 상태였어요 (어떤 arm 에도 안 걸려서 status ticker 의 frame 이 전부_ => continue로 감). exhaustive 면 다음 장기 요청이 스스로 밝혀야 해요.쓰기 경로도 이름을 붙여요 — read 경로가 설명해주는 것과 같은 상황인데
os error 35로 나왔고, 그게 창에 그대로 들어가요. 둘 다 OS 에러를 cause 로 보존해서 EAGAIN/ETIMEDOUT 이 구분돼요.deadline 을
connect()로 —pub이라 다른 소비자에게 무기한 소켓을 넘겨줬어요.rpc_within이 요청별로 read 예산만 좁혀요.빈 줄은 skip — 데몬의 reader 가 그렇게 해요. 기존엔 프로토콜 에러로 보고했어요.
검증
cargo test --locked -p crew— 214 passeda_socket_that_never_answers_gives_up— accept 만 하고 침묵하는 리스너. worker + deadline 으로 실행해서, 회귀 시 멈추는 게 아니라 실패해요FAILED ... rpc must return, not park: Timeout, 5.00초에 실패a_stream_is_not_a_question—Subscribe는None,PING_TIMEOUT < 1sa_person_gets_as_long_as_they_need_and_nobody_else_doescargo fmt --check—client.rscleanwith_home의 임시 경로 단축 — 바인딩하는 소켓이 stock TMPDIR 에서 104바이트 한계에 몇 바이트 남기고 있었어요정정: 이 PR 초판 본문에 "타임아웃을 빼면 테스트가 영원히 안 끝난다 (
timeout 30으로 강제 종료)" 라고 썼는데, macOS 에timeout이 없어서 그 확인 자체가 무효였어요. 위 5초 실패가 제대로 측정한 결과예요.범위
이 PR 은 무한 대기를 없애는 것까지예요. 창이 안 멈추게 하는 건 #21 (async command) 의 일이에요 — 지금
desktop.rs의 command 는 전부 sync 라 메인 스레드에서 돌고, 창은 200ms/1초로 폴링해요. 그쪽이 들어가야 "느린 데몬에도 창이 계속 그려진다" 가 성립해요.선행 2건 중 하나이고, 나머지는 #23 (데몬 기동 락) 이에요.