Why a tool call never runs a case
The first version of this idea was a tool that ran a case and returned its result. It does not survive contact with a real case. A fine external-aerodynamics run takes an hour or two, an MCP client backgrounds a tool call after about a hundred seconds, and a client that disconnects takes the run’s only observer with it. Everything below follows from refusing to run anything inside a call.
A job is a directory
Section titled “A job is a directory”submit_case validates the case, copies it into jobs/<job_id>/case/, writes job.json and the first line of events.jsonl, puts the job on a queue and returns. That takes milliseconds. From then on the job’s state lives in its directory, and every other tool reads it. A client can disconnect, a different client can pick the job up, and the server can restart: the directory is the truth.
Copying the case matters as much as storing state. The job runs on its own copy, so the source can be edited and resubmitted while a run is in flight, and two submissions of the same case never share files.
One writer per file
Section titled “One writer per file”Once the worker starts a job, only the worker writes its job.json. The server writes it once, at submit, and again only to cancel a job that is still queued. Anything else the server needs to say to a running job goes through a marker file: cancel.requested in the job directory, which the watcher sees within a few seconds.
The rule exists because it was broken once. An early set_end_time wrote job.json from the server while the worker held its own copy, and the worker’s next save silently reverted the edit. Now the tool edits only the case’s controlDict and an event; the watcher reads endTime back from the dictionary.
A queue per host, one job at a time
Section titled “A queue per host, one job at a time”Jobs go on a Procrastinate queue named after the host, with a lock that serialises them: twenty ranks of one case should not share cores with twenty of another. Locally the queue lives in memory and the worker runs inside the server. On a shared host it lives in Postgres and the worker is its own process, so the server can restart for a code change without touching a running case.
Before a queued job starts, the worker also checks for anything else occupying the host: another job’s container still up during a re-attach, an OpenFOAM container someone started by hand, a disk under its floor, or a maintenance hold. The job stays queued, with the reason in its summary, until the host is free.
Two independent axes
Section titled “Two independent axes”Transport and executor are separate choices. The transport is how a client reaches the server: stdio for a server on your own machine, streamable HTTP for a shared one. The executor is how a step runs: today a docker exec into a container per job, with a native executor planned for hosts without docker. Remote execution is not a third mode; it is the same server, run on the remote host, reached over HTTP.
What a tool call may block on
Section titled “What a tool call may block on”wait_for is the one tool that waits, and it is capped at about 85 seconds, under the point where Claude Code backgrounds a call. It returns the current snapshot and the reason it returned, so a client loops on it without ever holding a call open for the length of a run. For clients that support MCP Tasks, submit_case can instead return a task that resolves when the job ends. Cancelling that task stops the waiting, not the job; cancel_job stops the job.