Turnstile

The content on this page was written by AI under human supervision.

Turnstile is a Bash wrapper that decides when a long, memory-hungry job may start on a Linux machine that other jobs share. You put it in front of the command you want to run and state roughly how much memory and how many CPU threads the job will use; very large jobs then run one at a time in priority order, and smaller jobs run alongside only while memory and cores allow. It returns the wrapped command's exit code and logs each run's estimated versus measured peak memory, run time and exit code. Turnstile is a repository utility kept under ops/turnstile/ in the BootLoops repository, for operating a shared machine; it is not one of the toolkit packages and computes no scientific result itself.

What it does

The usual way to make big jobs take turns on one machine is flock LOCKFILE command: whoever holds the lock file runs, everyone else waits. That has three problems. The kernel hands the lock over roughly in arrival order, so an urgent job cannot move ahead. Jobs too small to need the lock are not counted, so several starting together can exhaust memory that each alone would have fit in. And a script that tests the lock with flock -n and then starts its job as a separate step holds nothing while the job runs.

turnstile_run.sh keeps the same kernel lock and adds a queue directory beside it, <lock>.q/. In the default token mode a job registers there with a name (--lane), a priority (--class or --prio, lower number first) and its expected peak memory in GiB (--est-rss, required). Waiting wrappers poll every 15 seconds; when the lock is free only the highest-priority live waiter takes it, ties going to the longest wait. The lock's file descriptor is handed to the job itself, so it stays held for the whole run even if the wrapper is killed, and the kernel releases it when the job exits. Entries left by killed wrappers are removed at the next poll.

In exempt mode (--exempt, for jobs under 300 GiB) the job does not take the lock. It starts only when available memory covers its own estimate, the not-yet-used part of every running job's estimate, any operator reserve, and a fixed floor (150 GiB by default). A new exempt job also waits while another exempt job is less than five minutes old and still below 90% of its estimate, so that only one is growing at a time. Both modes declare a thread count with --est-threads; a new job waits while the declared threads of running jobs plus its own would exceed the cores in the CPU range all jobs are pinned to (the whole machine unless TURNSTILE_PIN narrows it). A job that declares no thread count is assumed to need every core and runs alone.

An operator steers the queue through three plain files in <lock>.q/ (listed under Routines), re-read at every poll, without touching running jobs. Every run ends with an event=done line in <lock>.q/LEDGER.log giving the job name, mode, est_rss_g, measured peak_rss_g, declared est_threads, measured peak threads peak_r, wall_s and rc; those measured peaks are where the estimates for the next similar job should come from. Exit code 75 means the start was refused (--no-wait) or timed out, 69 means the platform is not Linux, 2 is a usage error, and anything else is the wrapped command's own code.

Limits: everything is local to one machine, so this is not a cluster scheduler. Priorities order the waiting queue and never interrupt a running job, and nothing caps the memory of a job once it runs. Jobs started without the wrapper count only through the memory they already occupy: their future growth and their threads are invisible to the accounting, although a plain flock on the same file still excludes correctly. Linux only: on macOS both scripts exit with code 69 and a message, and selftest.sh prints SKIP lines and exits 0.

Examples

Run the self-test suite. From the root of a clone of the repository; it uses a scratch lock in a temporary directory and never touches a real lock or a running job:

bash ops/turnstile/selftest.sh

Each check prints a PASS or FAIL line. They cover the lock held across an entire run, three waiters served in priority rather than arrival order, a PRIORITIES.conf override, a dead entry pruned, and memory and HOLD refusals (exit code 75). Further checks run three 30-thread jobs against a 64-core cap and confirm that no process the suite started is still running when it exits. A summary line counts passes and failures; the exit status is 0 when nothing failed.

Start a large solver alone, ahead of lower-priority work. The job is expected to peak near 400 GiB with 16 worker threads (the example from the package's README):

ops/turnstile/turnstile_run.sh --lane solveA_big --class critical --est-rss 400 --est-threads 16 --est-wall 12000 -- \
    nice -n 5 python3 -u solve_big.py ...

Everything after -- is the command. The wrapper registers solveA_big at priority 10 (the critical class), waits its turn, takes the lock, passes the memory and thread checks, and starts the command under taskset, sampling its memory every 30 seconds. Its own messages go to standard error. When the solver exits, the lock is released, the event=done record is written, and the wrapper exits with the solver's code.

Run two smaller jobs alongside the lock holder, then look at the queue.

turnstile_run.sh --exempt --lane jobB_ramp --est-rss 60 --est-threads 8 -- python3 grow_table.py ...
turnstile_run.sh --exempt --lane jobB_watcher --est-rss 1 --est-threads 0 -- ./watch.sh
turnstile_status.sh

The first job starts as soon as 60 GiB fits under the memory rule and 8 more threads fit under the core cap. The second, declared with --est-threads 0, is treated as a negligible helper and skips the thread accounting. turnstile_status.sh changes nothing and prints available memory and HOLD state, the lock holder with its current and peak memory, the waiting jobs numbered in service order, the running exempt jobs, the thread reservations, and the memory and core headroom left.

Routines

Scripts (in ops/turnstile/)

turnstile_run.sh options

Environment variables (defaults in parentheses; every launcher on a machine must share TURNSTILE_LOCK)

Operator files in <lock>.q/

Requirements and source

Bash on Linux with flock, taskset, setsid, nproc, GNU ps, awk and /proc/meminfo; nothing to install and no macOS port. The code is ops/turnstile/ in the BootLoops repository, beside the tools/ tree rather than inside it. The repository's run_selftests.py covers tools/ only, so run this self-test by hand with bash ops/turnstile/selftest.sh from the repository root (TURNSTILE_TEST_DIR=<dir> chooses its scratch directory). The utility was previously published as Bigram-queue (tools/bigram-queue/, scripts bigram_run.sh and bigram_status.sh); the queue-directory and log formats did not change with the rename, but the default lock path did (/var/tmp/BIGRAM.lock became /var/tmp/TURNSTILE.lock), so set TURNSTILE_LOCK or BIGRAM_LOCK to one shared path while old and new launch lines coexist on a machine. Released under the MIT license.

← back to the tools index