Turnstile
The content on this page was written by AI under human supervision.
Turnstile is a Bash wrapper that decides when a long, memory-hungry job may start on a Linux machine that other jobs share. You put it in front of the command you want to run and state roughly how much memory and how many CPU threads the job will use; very large jobs then run one at a time in priority order, and smaller jobs run alongside only while memory and cores allow. It returns the wrapped command's exit code and logs each run's estimated versus measured peak memory, run time and exit code. Turnstile is a repository utility kept under ops/turnstile/ in the BootLoops repository, for operating a shared machine; it is not one of the toolkit packages and computes no scientific result itself.
What it does
The usual way to make big jobs take turns on one machine is flock LOCKFILE command: whoever holds the lock file runs, everyone else waits. That has three problems. The kernel hands the lock over roughly in arrival order, so an urgent job cannot move ahead. Jobs too small to need the lock are not counted, so several starting together can exhaust memory that each alone would have fit in. And a script that tests the lock with flock -n and then starts its job as a separate step holds nothing while the job runs.
turnstile_run.sh keeps the same kernel lock and adds a queue directory beside it, <lock>.q/. In the default token mode a job registers there with a name (--lane), a priority (--class or --prio, lower number first) and its expected peak memory in GiB (--est-rss, required). Waiting wrappers poll every 15 seconds; when the lock is free only the highest-priority live waiter takes it, ties going to the longest wait. The lock's file descriptor is handed to the job itself, so it stays held for the whole run even if the wrapper is killed, and the kernel releases it when the job exits. Entries left by killed wrappers are removed at the next poll.
In exempt mode (--exempt, for jobs under 300 GiB) the job does not take the lock. It starts only when available memory covers its own estimate, the not-yet-used part of every running job's estimate, any operator reserve, and a fixed floor (150 GiB by default). A new exempt job also waits while another exempt job is less than five minutes old and still below 90% of its estimate, so that only one is growing at a time. Both modes declare a thread count with --est-threads; a new job waits while the declared threads of running jobs plus its own would exceed the cores in the CPU range all jobs are pinned to (the whole machine unless TURNSTILE_PIN narrows it). A job that declares no thread count is assumed to need every core and runs alone.
An operator steers the queue through three plain files in <lock>.q/ (listed under Routines), re-read at every poll, without touching running jobs. Every run ends with an event=done line in <lock>.q/LEDGER.log giving the job name, mode, est_rss_g, measured peak_rss_g, declared est_threads, measured peak threads peak_r, wall_s and rc; those measured peaks are where the estimates for the next similar job should come from. Exit code 75 means the start was refused (--no-wait) or timed out, 69 means the platform is not Linux, 2 is a usage error, and anything else is the wrapped command's own code.
Limits: everything is local to one machine, so this is not a cluster scheduler. Priorities order the waiting queue and never interrupt a running job, and nothing caps the memory of a job once it runs. Jobs started without the wrapper count only through the memory they already occupy: their future growth and their threads are invisible to the accounting, although a plain flock on the same file still excludes correctly. Linux only: on macOS both scripts exit with code 69 and a message, and selftest.sh prints SKIP lines and exits 0.
Examples
Run the self-test suite. From the root of a clone of the repository; it uses a scratch lock in a temporary directory and never touches a real lock or a running job:
bash ops/turnstile/selftest.sh
Each check prints a PASS or FAIL line. They cover the lock held across an entire run, three waiters served in priority rather than arrival order, a PRIORITIES.conf override, a dead entry pruned, and memory and HOLD refusals (exit code 75). Further checks run three 30-thread jobs against a 64-core cap and confirm that no process the suite started is still running when it exits. A summary line counts passes and failures; the exit status is 0 when nothing failed.
Start a large solver alone, ahead of lower-priority work. The job is expected to peak near 400 GiB with 16 worker threads (the example from the package's README):
ops/turnstile/turnstile_run.sh --lane solveA_big --class critical --est-rss 400 --est-threads 16 --est-wall 12000 -- \
nice -n 5 python3 -u solve_big.py ...
Everything after -- is the command. The wrapper registers solveA_big at priority 10 (the critical class), waits its turn, takes the lock, passes the memory and thread checks, and starts the command under taskset, sampling its memory every 30 seconds. Its own messages go to standard error. When the solver exits, the lock is released, the event=done record is written, and the wrapper exits with the solver's code.
Run two smaller jobs alongside the lock holder, then look at the queue.
turnstile_run.sh --exempt --lane jobB_ramp --est-rss 60 --est-threads 8 -- python3 grow_table.py ... turnstile_run.sh --exempt --lane jobB_watcher --est-rss 1 --est-threads 0 -- ./watch.sh turnstile_status.sh
The first job starts as soon as 60 GiB fits under the memory rule and 8 more threads fit under the core cap. The second, declared with --est-threads 0, is treated as a negligible helper and skips the thread accounting. turnstile_status.sh changes nothing and prints available memory and HOLD state, the lock holder with its current and peak memory, the waiting jobs numbered in service order, the running exempt jobs, the thread reservations, and the memory and core headroom left.
Routines
Scripts (in ops/turnstile/)
turnstile_run.sh [options] -- <command ...>— the launcher; token mode by default,--exemptto run alongside the lock holder.turnstile_status.sh [--lock PATH]— read-only report of holder, waiters, running exempt jobs, thread reservations and headroom.selftest.sh— the self-test suite on a scratch lock.
turnstile_run.sh options
--lane NAME— job name shown in the queue, the status report and the log, and matched byPRIORITIES.confpatterns.--class NAME— priority by class:critical10,high20,normal30,low40,misc50 (default); lower runs first.--prio Ngives a number directly.--est-rss G— expected peak memory in GiB (required).--est-threads N— expected busy threads; omitted means the whole pinned CPU range,0means a helper with no thread accounting.--est-wall S— expected run time in seconds, stored in the queue entry.--exempt— run under the memory rule without the lock; refused when--est-rssis at or aboveTURNSTILE_EXEMPT_MAX_G.--no-wait,--timeout S— exit 75 instead of polling, immediately or after S seconds.--lock PATH,--floor G,--poll S— override the lock file, memory floor and poll interval for this call.--no-stagger,--no-gate,--no-width-gate— operator escapes: ignore other jobs' growth window, skip the memory check, or skip the wait on the thread cap (threads are still recorded).
Environment variables (defaults in parentheses; every launcher on a machine must share TURNSTILE_LOCK)
TURNSTILE_LOCK— lock file (/var/tmp/TURNSTILE.lock); the queue directory is the same path with.qin place of.lock.TURNSTILE_POLLpoll seconds, integer (15);TURNSTILE_MONmemory-sample seconds (30);TURNSTILE_FLOOR_G(150);TURNSTILE_RAMP_SECSgrowth window (300);TURNSTILE_EXEMPT_MAX_G(300;TURNSTILE_FORCE_EXEMPT=1overrides);TURNSTILE_PINCPU list fortaskset -c(all cores);TURNSTILE_WIDTH_CAPcore cap (cores inTURNSTILE_PIN);TURNSTILE_DRIFT_LOG_SECSminimum seconds betweenwidth_driftlog lines, written when a job's running threads exceed twice its declared count (600).- Each
BIGRAM_<NAME>variable from the utility's former name is used as a fallback when the matchingTURNSTILE_<NAME>is unset. - The wrapped command receives
TURNSTILE_HELD=1orTURNSTILE_EXEMPT=1, plusTURNSTILE_LANEandTURNSTILE_EST_THREADS.
Operator files in <lock>.q/
PRIORITIES.conf— one<name-glob> <priority>per line, first match wins,default Nallowed; re-orders waiting jobs.HOLD— while present, no new job starts in either mode; its text is shown as the reason in the wrapper's waiting messages and inturnstile_status.sh.COTENANT_RESERVE_G— one integer, GiB kept free for another process, counted in every memory check.
Requirements and source
Bash on Linux with flock, taskset, setsid, nproc, GNU ps, awk and /proc/meminfo; nothing to install and no macOS port. The code is ops/turnstile/ in the BootLoops repository, beside the tools/ tree rather than inside it. The repository's run_selftests.py covers tools/ only, so run this self-test by hand with bash ops/turnstile/selftest.sh from the repository root (TURNSTILE_TEST_DIR=<dir> chooses its scratch directory). The utility was previously published as Bigram-queue (tools/bigram-queue/, scripts bigram_run.sh and bigram_status.sh); the queue-directory and log formats did not change with the rename, but the default lock path did (/var/tmp/BIGRAM.lock became /var/tmp/TURNSTILE.lock), so set TURNSTILE_LOCK or BIGRAM_LOCK to one shared path while old and new launch lines coexist on a machine. Released under the MIT license.