Emitall
The content on this page was written by AI under human supervision.
Emitall checks that the numbers quoted in a report or paper still agree with the stored result files they came from. You write one specification file listing each quoted number, the file it comes from and an expression that recomputes it. The engine recomputes every value, compares it with the quoted text at the quote's own printed precision, and lists every disagreement. A companion linter reads the Python scripts that generate summary sentences and flags digits typed into the text by hand, iteration in an unstable order and unseeded random sampling. Two further scripts check a paper's LaTeX sources: battery.py for pinned and forbidden phrases, paper_seams.py for sentences broken by a misplaced % comment.
What it does
Quoted numbers go wrong in ordinary ways. A value is copied by hand and a digit slips, or the result file is regenerated after the sentence was written. A count depends on an unstated choice of subset, or a range is rounded to the nearest digit when it should have been rounded outward. None of these looks wrong on the page, so Emitall recomputes every quoted value from the files and compares mechanically.
The input is a claims spec in JSON (YAML also works if PyYAML is installed). Its receipts block (the package's word for stored result files) gives an alias to each JSON or CSV file, optionally with newer_than and max_age_days freshness conditions. Its claims list has one entry per quoted number: a section and label, the quoted string exactly as printed, and an expr. Expressions use Python syntax but are evaluated by a restricted parser, never eval. They reach the files only through accessors such as f('s','/n_sites_rated') (the value at a slash-separated path), vals() (all values matching a pattern) and cw() (count table rows satisfying a list of conditions). These combine with arithmetic and helpers such as sum, len, where and frac (exact fractions). Longer derivations belong in your own script, written into the result file and then quoted.
Each claim is compared in one of three modes. printed (the default) formats the recomputed value with the quote's number of decimals and compares the text, ignoring thousands separators and a trailing percent sign: 10.4166... matches a quoted "10.4%", while -0.6845 does not match "-0.69" because it prints as -0.68. exact compares the two strings with no normalization. outward-band is for a quoted interval: at granularity $g$ the printed edges must equal $g\lfloor \mathrm{lo}/g \rfloor$ and $g\lceil \mathrm{hi}/g \rceil$, so the printed range always contains the computed one. That test runs in exact rational arithmetic, and a failure says whether a printed edge cuts into the computed interval (inward-edge) or the band is wider than the outward rounding (outward-loose).
The output is a table with one line per claim, marked MATCH, FINDING, EMITTED (a claim with no quoted string is recorded, never compared) or MISSING-RECEIPT, then the counts and, with --out, a JSON report. A missing result file is a finding on its own; so is a stale one, older by modification time than a newer_than file or than max_age_days. An expression that cannot be evaluated becomes an eval-error finding and the rest of the spec still runs. Exit code 0 means everything matches, 1 any finding, 2 a malformed spec. Two limits: Emitall checks only the claims in the spec, so coverage of the document is up to you, and printed mode confirms agreement at the printed precision, nothing more.
The linter, lint.py, parses the scripts that write summary sentences and reports three kinds of finding. digit-literal is a digit in the literal text of a string feeding an emission variable, key, argument or function (any name containing emit, headline or summar); interpolated {...} fields and % placeholders are exempt. unordered-iteration is a set or dict view iterated into such a string without sorted(). unseeded-sampling is random.sample, choice, choices or shuffle with no seed set earlier in the file. A comment # emitall-lint: allow-literal <reason> permits one literal; the reason is required, and a bare pragma is itself reported. Detection is by name pattern, so an emission list called lines is invisible to it.
Two more scripts check the paper itself. battery.py runs a JSON spec (format in its docstring) whose rows list holds assertions of four classes over the .tex files and the rendered text. FROZEN requires a pinned number or phrase verbatim in its section file or the rendered text, optionally an exact number of times; NEVER forbids a phrasing everywhere, with optional allowed contexts; NEG requires exactly zero occurrences; ABSTRACT requires a sentinel phrase in, or dropped from, the abstract. Comments are stripped first, and rendered text, from pdftotext or a pre-rendered txt file, is rejoined across hyphen wraps and page breaks. With pdf_fresh_vs set, a PDF older than its source stops the run with exit code 2 instead of passing against a stale render. paper_seams.py takes a master .tex file, follows \input and \include one level down, and reports every % comment block dropped inside a sentence. That slip leaves the sentence's opening words on a comment line, so the built text reads "...end. lowercase fragment". Table rows and macro lines are skipped, and --allow <phrase> names a legitimate lowercase sentence opening. When the built PDF sits beside the master, rendered fragments are cross-checked against the source; each finding names file and line, and the exit code is 0 clean, 1 any finding, 2 usage.
Examples
Run the included example. The package comes with a complete spec for an invented star-rating program, specs/toy_stars.json, with synthetic result files in specs/toy_receipts/. From tools/emitall/:
python3 run.py specs/toy_stars.json
Two of its twelve claims, a percentage in printed mode and a dollar range in outward-band mode (the file stores dollars; scale converts to millions before rounding at 0.1):
{"section": "RATINGS", "label": "pct changed", "quoted": "10.4%",
"expr": "100 * f('s','/n_star_changes') / f('s','/n_sites_rated')"},
{"section": "BANDS", "label": "payout band, 0.1 $M", "mode": "outward-band",
"quoted": "[1.2-2.8]", "quoted_lo": "1.2", "quoted_hi": "2.8",
"granularity": "0.1", "scale": "1e-6",
"lo_expr": "f('s','/payout_band_usd/lo')", "hi_expr": "f('s','/payout_band_usd/hi')"}
The run prints twelve lines, eleven MATCH and one EMITTED, then 12 headlines emitted; 11 MATCH; 0 FINDINGS, and exits 0. The self-tests rerun it with edits. Quoting "6" star changes where the file says 5 gives one mismatch finding and exit code 1; quoting payout-band edges that cut into the computed interval instead of rounding outward gives two inward-edge findings. Copy this spec as the starting point for your own.
Call the comparison functions from Python. With the directory containing emitall/ on sys.path, as the tests arrange:
from emitall import compare_printed, compare_outward_band compare_printed(6.652513268810491, "6.7%") # (True, "6.7") compare_printed(-0.6845146732002659, "-0.69") # (False, "-0.68") compare_outward_band(154867600, 232301400, "154.9", "232.3", "0.1", "1e-6") compare_outward_band(154867600, 232301400, "154.8", "232.4", "0.1", "1e-6")
compare_printed returns whether the value matches and the string it printed at the quote's precision. compare_outward_band takes the computed low and high (here in dollars), the quoted edges, the granularity and a scale, and returns (ok, display, note). The computed interval is [154.8676, 232.3014] million, so the first call fails with an inward-edge note naming both edges and the second, the correct outward rounding, passes.
Lint an emission script. To check the scripts that write your summary sentences, give the linter one or more files or directories (from tools/emitall/):
python3 lint.py <script-or-dir> [more paths...] [--json report.json] [--quiet]
Directories are scanned recursively for .py files. The linter prints the findings grouped by file, one line each with the line number, kind and a short detail, then a summary count and any pragma-allowed literals with their reasons; --json writes the same as a report file. The README's motivating case is a headline string in which every number is interpolated except a typed "(max 5 cents)", while the stored maximum was 0.09 dollars; the linter reports that 5 as a digit-literal (year labels typed into the same string are reported too). Exit code 0 is clean, 1 any finding, 2 a usage error.
Routines
Command line
run.py—python3 run.py SPEC.json [--out report.json] [--quiet]: evaluate a claims spec, print the table, optionally write the JSON report.lint.py—python3 lint.py <script-or-dir> [more paths...] [--json report.json] [--quiet]: lint emission scripts.battery.py—python3 battery.py SPEC.json [--quiet]: run a spec ofFROZEN/NEVER/NEG/ABSTRACTassertions over a paper's.texfiles and rendered text (exit 0 all pass, 1 any failure, 2 stale render or malformed spec);python3 battery.py --selftestruns its built-in fixture checks and needs nopdftotext.paper_seams.py—python3 paper_seams.py <master.tex> [--snapshot <dir>] [--marker <label> ...] [--allow <phrase> ...] [--quiet](alsopython3 -m emitall.paper_seams ...withtools/on the import path): report%comment blocks dropped inside a sentence across the master and the files it inputs;--snapshotcompares against an earlier copy of the tree,--markergives the label carried by the comment blocks that comparison examines (defaultFOLD).
Python (from emitall import ...)
run_spec(spec_path, out_path=None, quiet=False)— run a spec programmatically; returns(report, exit_code).load_spec(path)— parse a JSON or YAML spec; raisesSpecErrorwhen malformed.compare_printed(emitted, quoted, decimals=None)— printed-mode comparison; returns(ok, printed_string).compare_outward_band(lo, hi, quoted_lo, quoted_hi, granularity=1, scale=1)— outward-rounding test in exact fractions; returns(ok, display, note_or_None).fmt_like(quote, value, decimals=None)— format a number with the quoted string's decimal count.evaluate(expr, env, variables=None)— evaluate one field expression against the loaded files.resolve(obj, pointer),resolve_pattern(obj, pattern)— follow a slash path, or a pattern with*and{a|b}segments, into a loaded object.Env(receipts)— the alias-to-loaded-file mapping the evaluator reads.SpecError,ExprError,ReceiptMissing— exceptions for a malformed spec, an expression that cannot be evaluated, and a file absent at load.lint_source(source, path="<string>"),lint_file(path),lint_paths(targets)— the linter as functions; the first two return(findings, allowed)andlint_pathsreturns(findings, allowed, files_scanned), each finding a dict with file, line, kind and detail.
Inside spec expressions
f(alias, pointer),vals(alias, pattern),rows(alias, pointer),cw(alias, pointer, conds),joinrows(alias, pointer, template, sep)— a value, a list of values, table rows, a count of rows meeting[field, op, value]conditions, and template-formatted rows from a result file.where(xs, op, v),split(s, sep, i),join(xs, sep),jsondump(x),zipjoin(xs, ys, template, sep),frac(x)— filtering, string pieces, joining, canonical JSON, pairing two equal-length lists, exact fractions; pluslen sum min max abs int float str round bool.
Used on this site
- Medicare star ratings — the quoted headline numbers of the star-ratings analysis were re-checked against their stored result files with Emitall specs.
Requirements and source
Python 3, standard library only; PyYAML is optional, for YAML specs. The pdftotext program is needed only for battery.py specs that read a PDF (a pre-rendered txt file avoids it) and for the rendered-text cross-check in paper_seams.py, which is skipped with a notice when the program or the PDF is absent. The report timestamp is read from the system date -u command, so a Unix-like system is assumed. Run the self-tests with python3 -m pytest tools/emitall -q from the repository root, or python3 tests/test_all.py and python3 tests/test_paper_seams.py from tools/emitall/; they also pass under python3 -O. The code is in tools/emitall/ in BootLoops' bootloops-dev repository (GitHub organization BootLoops-ai), released under the MIT license: engine.py, expr.py, run.py, lint.py, battery.py, paper_seams.py, the example spec and fixtures under specs/, and tests/.