Popen.communicate(timeout=0) in a retry loop never drains a PIPE: child exits but stays a zombie forever
Python 3.12 (also 3.11): a subprocess keepalive loop of the form while True: try: out, err = proc.communicate(timeout=INTERVAL); break; except subprocess.TimeoutExpired: ping_db() works fine with INTERVAL=30, but a test that monkeypatched INTERVAL to 0 hung indefinitely. The child (git clone --quiet ... with stderr=subprocess.PIPE, stdout=DEVNULL) finished in under a second and showed up in /proc as state Z (zombie) with PPid = the Python process, yet every subsequent communicate(timeout=0) kept raising TimeoutExpired and proc.returncode stayed None. No CPU spin visible from outside; pytest-xdist worker reported 'running' for 20 minutes until killed. I expected a zero timeout to behave like a non-blocking poll: read whatever is ready, reap if exited, otherwise raise TimeoutExpired.
Root cause: on POSIX, Popen._communicate computes endtime = monotonic() + timeout and then, at the top of its selector loop, checks if timeout is not None and timeout < 0: self._check_timeout(..., skip_check_and_raise=True) where timeout = self._remaining_time(endtime). With timeout=0 the remaining time is already negative by the time the check runs, so it raises TimeoutExpired before calling selector.select() even once. The PIPE'd fd (stderr here) is therefore never read, never hits EOF, never gets unregistered from the selector, and the function never reaches self.wait(...), so the exited child is never reaped. Every repeated communicate(timeout=0) repeats exactly that path: permanent TimeoutExpired, permanent zombie, returncode is None.
This is specific to communicate() with at least one PIPE. proc.wait(timeout=0) and proc.poll() do try waitpid(WNOHANG) first and would reap correctly; with no pipes at all, communicate() falls through to wait() and also works.
Fix options:
- Never pass a zero (or sub-millisecond) timeout to
communicate(). Use a floor, e.g.proc.communicate(timeout=max(interval, 0.05)). - If you need a true non-blocking poll, use
proc.poll()/proc.wait(timeout=...)for liveness and read the pipe yourself (orstderr=subprocess.DEVNULL/ a temp file instead of PIPE). - In tests, don't zero the shared interval constant to force more frequent side effects; wrap the side-effect callable instead and count calls.
In my case the same constant was both the DB keepalive interval and the communicate timeout, so monkeypatch.setattr(mod, 'KEEPALIVE_INTERVAL_S', 0) silently turned the clone loop into a zombie spin. Diagnosing it from the outside: /proc/<child>/status shows State: Z with the Python process as PPid, while the Python process is stuck in the TimeoutExpired loop.
import subprocess
p = subprocess.Popen(['sh', '-c', 'echo hi >&2'], stderr=subprocess.PIPE, text=True)
import time; time.sleep(0.5) # child has exited
for _ in range(3):
try:
print(p.communicate(timeout=0))
break
except subprocess.TimeoutExpired:
print('timeout', p.returncode) # prints 'timeout None' three times; child is a zombie
print(p.communicate(timeout=1)) # reads ('', 'hi\n') and reaps