Skip to content

Surviving restarts

A shared CFD host gets restarted: for a code update, a reboot, or a stack that has to move. A case that dies with its worker costs an hour, so the design rule is that nothing the worker does is the only thing keeping a step alive. Each part of that rule came from a failure.

Steps used to run under an attached docker exec, with the step’s output streamed back through the exec client into a log file. Stopping the worker container stopped the exec client, the stream closed, and the meshing step died of SIGPIPE on its next line of output. One deploy killed a mesh this way, mid layer addition.

Now a step runs under docker exec -d. Inside the container, a small shell redirects the step’s output to its log in the case directory and, when the step ends, writes its exit code to .mcopenfoam.<step>.exit. Nothing outside the container holds a pipe to it. The worker polls the exit file.

Tested the hard way: the whole server process group was killed forty seconds into a twenty-rank snappyHexMesh. The ranks kept running and the log kept growing. A new server re-attached, waited the mesh out, took exit code 0 from the file, ran checkMesh and finished the job.

At startup the worker reconciles the jobs on disk with the queue. A job that reads running gets its queue row back at a higher priority than anything waiting, so its re-attach takes the host before a queued job can. Without that, a queued job that won the lock sat in the host guard waiting for the running job’s container, while nobody watched the running job: neither could ever finish.

A job that reads running but has no live container is marked lost at once, instead of waiting its turn.

A re-attached worker has to decide whether the step it is watching is still alive, and both ways of getting that wrong happened.

  • Counting zombies as alive. When a disk filled, a solver died mid-write. Its twenty ranks stayed as zombies, because the container’s first process was sleep, which reaps nothing, and the liveness probe counted twenty processes. Job containers now run with --init, and the probe ignores zombies.
  • Counting a slow probe as dead. Under twenty ranks on sixteen threads, a docker exec pgrep took over a minute, timed out, and failed a healthy mesh. Liveness now comes from the daemon first: docker inspect reports whether the container runs, its exit code and whether it was killed for memory, and stays fast under load. The in-container probe runs only while the container is up, and a probe that fails counts as alive for at most twenty tries before the step is given up.

So a container that dies fails its step within a minute, with container_exited and the OOM flag in the step summary, and a step that is merely slow is left alone.

A worker killed mid-job leaves its queue row marked as in progress, still holding the host lock, so after a restart nothing behind it ran. Reconcile now closes those rows, and recreates every waiting job’s row in submit order, so a job whose row was closed does not fall behind jobs submitted after it.

Several servers can share a docker daemon: a local stdio server, a test run with a temporary jobs directory, a shared server. Cleanup used to remove every mcopenfoam container whose job it did not know, and a test run killed another server’s live mesh that way. Job containers now carry a label naming their jobs directory, and each server cleans up only its own.