Popcorn
The content on this page was written by AI under human supervision.
Popcorn is a Python package for the expected site-frequency spectrum of a DNA sample under natural selection, the quantity that programs for inferring the distribution of fitness effects (polyDFE, dadi and fitdadi, fastDFE) evaluate in double precision inside their likelihoods. Given a sample size, a scaled selection coefficient and a requested number of digits, it returns entries, whole spectra and their derivatives with respect to the selection coefficient. Each value is exact, or an interval guaranteed to contain the true value, or comes with a count of digits on which two independent computations agree. Further modules average the spectrum over a distribution of fitness effects with exact gradients, evaluate it under dominance, and decide in exact rational arithmetic whether a polynomial is nonnegative or a point lies inside a convex hull, returning a certificate either way. Four more modules give exact expected spectra under multiple-merger coalescents, two-locus branch-length moments under recombination, the selected spectrum through a history of population-size changes for large samples, and a join that polarizes variant sites against an ancestral-allele sequence.
What it does
In a sample of $n$ chromosomes, entry $i$ of the site-frequency spectrum is the number of variable sites at which the derived (mutant) allele is carried by exactly $i$ of them. Under the Poisson random field model of a constant-size population, mutations with scaled selection coefficient $S=4N_e s$ ($S \gt 0$ advantageous) give entry $i$ the expectation
$$E_i(S)=\theta\,\frac{n}{i(n-i)}\,\frac{1-{}_1F_1(n-i;\,n;\,-S)}{1-e^{-S}},\qquad i=1,\dots,n-1,$$
with ${}_1F_1$ the confluent hypergeometric function and $\theta$ the scaled mutation rate (Popcorn uses $\theta=1$). Evaluating it is badly conditioned: at moderate $n$ and $|S|$ the intermediate quantities cancel far beyond the 16 digits of double precision (the evaluator bundle linked at the end of this page records a condition number of $10^{85}$ already at $n=100$, $S=\pm 10$), so a double-precision value can be wrong with no warning. Popcorn computes the same object in exact or arbitrary-precision arithmetic with a check attached to every number.
popcorn.sfs works on the whole vector. The values $M_i={}_1F_1(n-i;n;-S)$ obey a three-term recurrence in $i$ with elementary end values, so the vector solves one tridiagonal linear system. For rational $S$ the solution separates as $M_i=\alpha_i+\beta_i e^{-S}$ with $\alpha_i,\beta_i$ exact fractions (solve_M_exact), and each derivative $d^kM/dS^k$ is one more solve of the same system (solve_dM_exact). assemble_M evaluates $\alpha_i+\beta_i e^{-S}$ at a precision chosen from the size of the fractions, so the cancellation cannot reach the requested digits. This exact route is practical to about $n=2000$. Above that, arb_route runs the recurrence in ball arithmeticeach number is a midpoint plus a radius that provably contains the true value (python-flint) and doubles the precision until the worst relative radius meets the target. It has been run to $n=10^5$.
popcorn.dfe integrates the spectrum against polyDFE's "model C" distribution of fitness effects: a reflected gamma density (shape $b$, mean $S_d \lt 0$) for deleterious mutations plus, with weight $p_b$, an exponential (mean $S_b \gt 0$) for beneficial ones. CertifiedKernel computes ball-arithmetic spectra once at fixed quadrature nodes in $\log|S|$; each later evaluation is a weighted sum over that cache, the gradient in $(b,S_d,p_b,S_b)$ reuses it, and each mix call reports the digits on which two Gauss–Legendre degrees agree plus a bound on the neglected tails. popcorn.dominance handles a dominance coefficient $h\neq 1/2$ (genotype fitnesses $1:1+2sh:1+2s$) with E_dominant, adaptive quadrature over a closed-form inner integral, and E_dominant_qseries, an independent series evaluator for cross-checking. Their self-consistency digit counts are strong evidence of accuracy; unlike the exact fractions and ball radii of popcorn.sfs, they are not proofs.
popcorn.certificates collects the exact deciders used to turn such spectra into proven statements about a model. positivity decides whether a sparse polynomial with rational coefficients is nonnegative on $[0,1]$ by two independent routes (Bernstein expansion with subdivision, and certified root isolation with exact sign checks). exact_lp, fast_lp, cone_lp_b and CertHull decide whether a rational point lies in the convex hull of a set of rational points, returning exact weights when it does and an exact separating (Farkas) functional when it does not. region certifies statements over a whole two-parameter box rather than a grid of samples: positivity on the box, a certified supremum of a rational function, a Cauchy–Schwarz lower bound on a $\chi^2$ distance from it, and rational reconstruction with Sturm root counting. Floating point may propose an answer in these modules; only exact rational arithmetic accepts it, and every certificate is re-substituted before it is returned.
Four further modules go beyond the equilibrium selected spectrum. popcorn.lambda_coalescent handles multiple-merger coalescentsgenealogical models in which three or more ancestral lineages may merge in a single event, as under sweepstakes reproduction at constant population size. For the Kingman, Beta$(2-\alpha,\alpha)$ and Dirac$(\psi)$ families it returns the merger rates $\lambda_{b,k}$ and the expected branch length $E[L_i]$ subtending $i$ of $n$ samples (hence the expected neutral unfolded spectrum) as exact fractions. It carries built-in exact identity checks and an exact $n=20$ reference grid. popcorn.twolocus computes the second moments $E[T_i^A T_j^B]$ of the branch lengths subtending $i$ samples at one locus and $j$ at another, under the coalescent with recombination at scaled rate $\rho$ and a piecewise-constant population size; this is the expected two-locus frequency spectrum up to scale. It enumerates the two-locus configuration chain once per $n$ and then solves sparse linear systems, with no simulation, and is practical to $n=8$. popcorn.transient computes the expected spectrum under genic selection $S$ after a piecewise-constant size history, for samples of roughly $n=1000$ to $2500$. It integrates the Kimura forward diffusion on a fixed frequency grid with a banded implicit time-stepper and projects to the sample size only at the end; it also down-samples and folds spectra. Unlike the rest of the package, popcorn.twolocus and popcorn.transient are floating point: double precision, validated by measurement over a stated range and unmeasured outside it. The guide records that range (for popcorn.transient at $n=1000$, agreement with the certified stationary values to $9.4\times10^{-5}$ relative or better for $|S|\le 100$), and their numbers should be labeled as floating point wherever they appear beside certified ones. popcorn.ancestral is a data utility with nothing certified in it. It streams a table of variant sites against an Ensembl EPO ancestral-allele FASTA file and classifies each site by the EPO confidence convention (uppercase base high, lowercase low, N, - or . no call). For biallelic SNVs it orients REF and ALT to ancestral and derived, and it writes a pos ref alt anc conf table with counters and summary rates.
In BootLoops' bootloops-dev repository (GitHub organization BootLoops-ai) the package is one flat directory, tools/popcorn/: the engine modules (sfs_engine.py, arb_route.py, dfe_layer.py, certquad.py, dominance_oracle.py, dominance_qseries.py, lambda_exact.py, twolocus_engine.py, transient_sfs_engine.py, epo_join.py), the certificate modules and a reference/ folder of stored test values sit beside the package files, and the submodules (popcorn.sfs, popcorn.dfe, popcorn.dominance, popcorn.certificates, popcorn.lambda_coalescent, popcorn.twolocus, popcorn.transient, popcorn.ancestral) import them from there. The package stores a SHA-256 digest of every engine and certificate file and of REFERENCE_VALUES.json, and popcorn.verify() names any file that changed. The command-line script re-verifies the files it uses before importing them. On a digest mismatch, a disagreement between its two routes, or a failure to reproduce a stored reference value, it exits with status 3 and an error message instead of printing an unchecked number. The same package is also distributed from this site as an archive (see Requirements and source below).
Limits. The certified modules assume a constant population size at equilibrium. A piecewise-constant size history is available only in floating point (popcorn.transient for the single-locus selected spectrum, popcorn.twolocus for neutral two-locus moments), and popcorn.lambda_coalescent is neutral and constant-size. There is no migration, no multi-population spectrum, no continuous size change and no fitting driver. Spectra are unfolded, at $\theta=1$. $S$ is the $S$ of polyDFE and fastDFE; dadi and fitdadi use $\gamma=S/2$. The exact route of popcorn.sfs, the command-line entry, check and selftest modes and popcorn.dominance need only mpmath; popcorn.dfe and the ball-arithmetic route popcorn.sfs.arb_route (loaded on first use) need python-flint. The certificate modules load on first use: gmpy2 for the linear-programming certificates, python-flint for cone_lp_b and one positivity route, sympy for root counting, numpy and scipy only as proposers. popcorn.twolocus and popcorn.transient need numpy and scipy; popcorn.lambda_coalescent and popcorn.ancestral need only the standard library.
Examples
Run the self-test suite. From the root of the bootloops-dev repository:
python3 tools/popcorn/selftest.py
The suite prints one line per check (LEG <name> PASS, SKIP or FAIL, the seconds taken and a short detail), then a count of passes, skips and failures and exactly one OVERALL PASS or OVERALL FAIL: <names> line, with exit status 0 or 1. The default checks verify the digests, run the command-line self-test and the single-entry check shown next, and compare the exact and ball-arithmetic routes with direct hypergeometric evaluation. They also test the dominance evaluator's collapse to the closed form at $h=1/2$ and the series evaluator against stored reference values, build a small mixing kernel and compare its analytic gradient with a finite difference, and run the certificate modules on toy hulls and polynomials, including deliberately wrong inputs (an exterior point, a negative polynomial) that must be refused. Further checks cover the newer modules: the exact coalescent identities and the $n=20$ reference grid; the two-locus chain against exact Kingman moments and closed forms; the transient engine against the certified stationary values and its self-convergence table; and the ancestral-join fixture through the library and the command line. With every optional dependency installed the twenty default checks take about twenty seconds. A check whose dependency is missing is skipped with a line naming the package to install, and skips do not fail the suite. --full adds three slow checks (minutes each): the two-route comparison up to $n=10007$, the dominance evaluator's full self-test and the $n=20$ kernel self-test. Nothing is written into the package directory.
Compute one entry two independent ways. Here entry $i=50$ of a sample of $n=100$ at $S=-1000$, to 40 digits:
python3 tools/popcorn/certsfs.py check 100 50 -1000 --dps 40
Three lines come back, labeled closed, series and agree: the value from the hypergeometric closed form (mpmath), the value from an exact-rational series that shares no code with it, and the number of digits on which they agree against the target of 40. If agreement falls below dps minus 5, the script still prints both values, then writes an error to standard error and exits with status 3, so a calling script cannot mistake the run for a success. Only mpmath is needed; it finishes in under a second. The other modes (entry, vector, selftest) are listed under Routines.
Mix the spectrum over a distribution of fitness effects, from Python. Put the repository's tools/ directory on PYTHONPATH so that import popcorn resolves. Then check the digests and drive the mixing layer the way the engine's own self-test does:
import popcorn
popcorn.verify()
from popcorn.dfe import CertifiedKernel
n = 20
K = CertifiedKernel(n, dps=40)
K.build(-1)
K.build(1)
Ev, sc, tail = K.mix({'b': '0.4', 'Sd': '-1000', 'pb': 0})
g = K.grad_mix({'b': 0.4, 'Sd': -1000.0, 'pb': 0.02, 'Sb': 10.0})
verify() returns [] on an intact copy and otherwise one string per changed or missing file. build(-1) and build(1) fill the cache for each sign of $S$ (the slow step; python-flint required). mix returns the mixed spectrum as a list indexed by $i$ (here purely deleterious, $p_b=0$), the digits on which the two quadrature degrees agree, and the relative tail bound; grad_mix returns a dictionary keyed 'b', 'Sd', 'pb', 'Sb', each a list of $\partial E_i/\partial(\text{parameter})$.
Exact expected spectrum under a multiple-merger coalescent, from Python. With tools/ on PYTHONPATH as above, the guide's example evaluates a Beta coalescent and a Dirac coalescent for a sample of 50:
import popcorn # with tools/ on PYTHONPATH
from fractions import Fraction
L = popcorn.lambda_coalescent # exact, standard library
h = L.expected_lengths(50, L.beta_rate(Fraction(3, 2)))[50] # {i: E[L_i]} exact Fractions
x = L.xi_hat(50, L.dirac_rate(Fraction(1, 10))) # normalized expected SFS, sums to 1 exactly
beta_rate(Fraction(3, 2)) returns the merger-rate function of the Beta$(1/2,3/2)$ coalescent ($\alpha=3/2$) and dirac_rate(Fraction(1, 10)) that of the Dirac coalescent with $\psi=1/10$. expected_lengths(n, lam) returns a dictionary indexed by the number of lineages; its entry [50] maps each $i=1,\dots,49$ to the exact expected total branch length subtending $i$ of the 50 samples, as a Fraction, which is proportional to the expected unfolded spectrum. xi_hat returns that spectrum as a list of $n-1$ fractions normalized to sum to exactly 1. Pass parameters as Fractions or strings such as '3/2', never as binary floats, or the rates are exact for the wrong rational number. The cost is $O(n^3)$ rational operations whose size grows with $n$: the guide measures about 2 seconds at $n=100$ and 17 seconds at $n=200$ for Beta$(3/2)$, about 3 minutes at $n=200$ for Dirac$(1/2)$, and $O(n^2)$ for Kingman. python3 tools/popcorn/lambda_exact.py runs the exact identity checks and prints the $n=20$ reference grid without writing anything.
Routines
Command line and checks (scripts in tools/popcorn/)
certsfs.py check N I S [--dps D]— one entry by two independent routes with digits of agreement; status 3 below $D-5$.certsfs.py entry N I S [--dps D]— one entry from the closed form to $D$ digits (default 40).certsfs.py vector N S [--dps D] [--grad]— the whole spectrum, one tab-separated line per $i$, optionally with $dE_i/dS$; exact route for $N\le 2000$, ball arithmetic above; verifies the digests of the engine files beside it before importing them.certsfs.py selftest [--dps D]— verify digests, reproduce two stored reference values, cross-check the routes, compare $D$ against $2D$ digits. Every refusal raisesCertSFSGateError, exit status 3.selftest.py [--full]— the self-test suite of the first example: twenty fast checks by default, three slow ones added by--full; one line per check and oneOVERALLline, exit status 0 unless a check fails; scratch files go to a temporary directory, never into the package.gate_phase1.py— the slow two-route comparison run byselftest.py --full: the recurrence routes against direct evaluation of ${}_1F_1$ at fixed pseudo-random test points up to $n=10007$, values and $S$-derivatives, requiring at least 40 matching digits; prints a verdict and one line per point and writes a JSON record into the current directory, so run it from a scratch directory.region_certificates.py,dominance_oracle.py,dfe_layer.pyrun as scripts — the self-tests of the region certificates (printsREGION CERTS PASS; includes inputs that must be refused), of the dominance evaluator, and of the $n=20$ mixing kernel; the last two are the other slow checks of--full.
Package
popcorn.verify(verbose=False)— compare every engine and certificate module andREFERENCE_VALUES.jsonwith its stored SHA-256;[]when clean, otherwise one string per changed or missing file.popcorn.PKG_DIR,popcorn.__version__— the package directory and version string.popcorn.sfs,popcorn.dfe,popcorn.dominance,popcorn.certificates,popcorn.lambda_coalescent,popcorn.twolocus,popcorn.transientandpopcorn.ancestralload on first access.popcorn._pins.regen()— for maintainers: print fresh digests after a deliberate change to one of the checksummed files, to paste back into_pins.pyandcertsfs.py.
popcorn.sfs (engine modules sfs_engine, arb_route; arb_route loads on first use and needs python-flint)
solve_M_exact(n, S)— exact(alpha, beta), length $n+1$, with $M_i=\alpha_i+\beta_i e^{-S}$;San integer orFraction.solve_dM_exact(n, S, alpha, beta, order=1)— exact $d^kM/dS^k$ up toorder, as a list of(alpha_k, beta_k).assemble_M(alpha, beta, S, target_dps)— evaluate $M_i$ totarget_dpsdigits at an automatically padded precision; pass the sameS.sfs_engine.sfs_from_M(n, S, M_alpha_beta_or_list, dps)— $E_i$, $i=1,\dots,n-1$, from the(alpha, beta)tuple or a list of $M_i$.arb_route.solve_vector_arb(n, S, target_dps, want_grad=True, max_prec=1 << 22, S_frac=None)— ball-arithmetic $M$ and $dM/dS$ with precision doubling; returns a dictionary withM,Mp,achieved_dps,prec.arb_route.sfs_arb(n, S, sol, i, S_frac=None)— $E_i$ and $dE_i/dS$ as balls from that result.
popcorn.dfe (engine modules dfe_layer, certquad)
CertifiedKernel(n, dps=40, t_lo_dps_frac=None, t_hi=6, node_target=None)— cache object:.build(sign, shape_min=0.05),.mix(params, dps_check=True)returning (vector, self-consistency digits, tail bound),.grad_mix(params)returning the four derivative vectors;paramshas keysb,Sd,pb,Sb.certquad.integrate(f, lnf, cuts, dps, band=12, degrees=(6, 7))— composite Gauss–Legendre quadrature with panels refined from a cheap estimatelnfof $\log|f|$; returns (integral, digits of agreement between the two degrees).
popcorn.dominance (engine modules dominance_oracle, dominance_qseries)
E_dominant(n, k, S, h, dps=50)— entry $k$ at selection $S$ and dominance $h$; returns (value, self-consistency digits). Rundominance_oracle.pyas a script for its self-test.E_dominant_qseries(n, k, S, h, dps=40)— independent series evaluator; returns the value; call at two precisions and compare.
popcorn.certificates (modules positivity, exact_lp, fast_lp, cone_lp_b, cert_lp, region, each loaded on first use)
positivity.nonneg_route_B(coeffs, touch_points=(), depth_cap=12, node_cap=4000)— decide whether the polynomial{exponent: rational}is nonnegative on $[0,1]$ by exact Bernstein expansion and subdivision, after dividing out even roots attouch_points; returns (verdict, info), verdictTrue,FalseorNone(undecided at the depth cap).positivity.nonneg_route_R(coeffs, max_tries=5)— the independent route: certified root isolation and exact sign checks between roots; returns (verdict, info, dips).exact_lp.membership(points, q)— pure-rational simplex:('FEASIBLE', weights)ifqis a convex combination ofpoints, else('INFEASIBLE', (w, w0))with an exact Farkas functional.fast_lp.membership_fast(points, q, escalate_budget_s=120)gives the same verdicts with a floating-point solve proposing and exact arithmetic deciding (raisesBudgetExceededrather than guess);cone_lp_b.membership_exact(points, q, warm=None)is the exact integer dual simplex and also returns its final basis for warm starts.CertHull(points)— hull-membership object for many queries against one point set:.decide(q)returns (status, certificate, basis, route), reusing the last verified certificate or basis when it still applies; scipy only proposes,cone_lp_bdecides.region—to_bernstein(p, ds, du)andbern_nonneg(b, depth)(positivity of{(es, eu): Fraction}on $[0,1]^2$),certified_sup(Ppoly, Tpoly, lo, hi, depth=6, iters=40)(smallest certified $c$ with $c\,T-P\ge 0$ on the box),cs_chi2_lower_bound(phi_dot_target, cert_sup, phi_S_phi),reconstruct_rational(f, pts=None, dmax=40, n_held=6)andcount_roots(coeffs, lo=1, hi=2)(Sturm count on an open interval); all exact fractions.
Coalescent, two-locus, transient and ancestral-state modules
popcorn.lambda_coalescent (engine module lambda_exact; standard library only; every value an exact Fraction)
kingman_rate(),beta_rate(alpha),dirac_rate(psi),msprime_dirac_rate(psi, c)— merger-rate functionslam(b, k)giving $\lambda_{b,k}$, the rate at which a given $k$ of $b$ lineages merge, for the Kingman, Beta$(2-\alpha,\alpha)$, Dirac$(\psi)$ and Kingman-plus-Dirac (msprime's Dirac model) coalescents; passalpha,psi,casFractions or'p/q'strings.expected_lengths(n, lam)— dictionaryh[m][i], $m=1,\dots,n$, $i=1,\dots,m-1$, of exact expected branch length subtending $i$ of $m$ lineages;h[n]is the sample's spectrum.xi_hat(n, lam)— the expected unfolded spectrum normalized to unit sum, a list of $n-1$ fractions.identity_checks(n)— the built-in exact identities as booleans (Kingman $2/i$; $\alpha=2$ and $\psi=0$ give Kingman; $\alpha=1$ gives Bolthausen–Sznitman; $\psi=1$ gives a star tree).reference_grid(),load_reference_spectra(path=None),load_witness(path=None),witness_margin(w, w0, x)— the $n=20$ grid of families (Kingman, seven Beta, six Dirac), its stored exact spectra underreference/lambda_coalescent/, the stored exact linear functional that separates that grid from variable-size Kingman spectra, and the functional's exact value $\langle w,\hat\xi\rangle+w_0$ at a spectrum.lambda_exact.py [--n N] [--out FILE]— script: runs the identity checks and prints one line per family of the reference grid (with the functional's sign at $n=20$);--outalso writes the results as JSON.
popcorn.twolocus (engine module twolocus_engine; numpy and scipy; floating point, validated as described above)
TwoLocus(n=8)— enumerate the labeled two-locus configurations for a sample of $n$ and build the sparse coalescence and recombination generators once (.Nis the number of states);.moments(rho, epochs)returns(M, m1A, m1B)withM[i-1][j-1]$=E[T_i^A T_j^B]$ and the first moments at each locus.epochsis[(tau_start, eta), ...]with time in units of $2N_{\mathrm{ref}}$ generations ascending from 0,eta$=N_{\mathrm{ref}}/N$ and the last epoch extending to infinity;rho$=4N_{\mathrm{ref}}R$.chain(n),moments(n, rho, epochs=((0.0, 1.0),))— the chain cached pern, and a one-call wrapper (default: constant size).pool_pairs(M)— the symmetrized moments on pairs $i\le j$ (diagonal doubled), normalized to unit sum; compare only with vectors pooled the same way.epochs_from_Nt(times_gen, sizes, N_ref=1.0e4)— convert a size history in generations into the[(tau_start, eta)]form.twolocus_engine.py [n] [rho]— script demo (defaults 6 and 1.0): prints the state count, $E[T_i]$ beside the Kingman $2/i$, and the moment matrix.
popcorn.transient (engine module transient_sfs_engine; numpy and scipy; floating point, validated as described above)
TransientSFSEngine(n=1000, theta=1.0, grid_kw=None)— build the frequency grid and the projection to sample sizenonce;.expected_sfs(S, epochs, dt0=4e-4)returns(E, meta), the unfolded spectrum $E_i$, $i=1,\dots,n-1$, after the historyepochs = [(nu, T), ...](relative size $N/N_{\mathrm{ref}}$ and duration in $2N_{\mathrm{ref}}$ generations, past to present, following an ancestral equilibrium;T <= 0is an instantaneous size change), and the caller must checkmeta["negative_entries"] == 0;.equilibrium_sfs(S)is the projected stationary spectrum in floating point (usepopcorn.sfsfor certified values);.run_epochs(S, epochs, dt0=4e-4, min_steps=24, rannacher=4),.project(u)and.eq_u(S, rho=1.0)are the underlying steps.project_down(E, n, m)— hypergeometric down-sampling of an unfolded spectrum from sample sizentom.fold_sfs(E)— unfolded spectrum to folded minor-allele classes.make_grid(x_min=1e-8, x_lin_lo=None, x_lin_hi=None, one_minus_max=1e-8, per_decade=160, dx_mid=8e-4),ln_g_ratio(x, S)— the composite logarithmic/uniform/logarithmic frequency grid (pass overrides throughgrid_kw) and the logarithm of the stationary density used for the equilibrium start.transient_sfs_engine.py— script: a smoke test of about a second at $n=1000$ (neutral equilibrium against $1/i$, a doubled size held long against $2/i$, an $S=-5$ equilibrium that must stay put).
popcorn.ancestral (engine module epo_join; standard library only; a data utility, nothing certified)
join_sites(fasta, sites, out=None, chrom=None, layout="auto", compresslevel=6, flush_rows=100000)— join a sites table (path,.gz,-for standard input, an open handle, or an iterable of tuples) against the first record of an ancestral FASTA and write thepos ref alt anc conftable (out=Nonefor counters only); returns the counters dictionary (confidence histogram, orientation counts by confidence, unpolarizable, out-of-range and skipped rows).polarization_rates(counts)— summary fractions from those counters (ancestral equals REF, equals ALT, third-allele mismatch, low confidence, unpolarizable);Nonewhere a denominator is zero.load_fasta(path),ancestral_base(seq, pos),classify(anc),orient(ref, alt, anc),iter_sites(source, layout="auto", chrom=None, counters=None),HIGH,LOW— the pieces: read the first FASTA record, look up a 1-based position (.when out of range), the confidence class'high','low'or'none', the orientation key ('anc_eq_ref','anc_eq_alt','anc_neither','unpolarizable', orNonefor a non-SNV) with the oriented alleles, the sites-table reader, and the uppercase and lowercase base sets.epo_join.py --fasta F --sites T --out O [--chrom C] [--layout auto|chrom|nochrom] [--summary J]— the command line: writes the table (.gzcompresses,-is standard output), prints one line of counters, and with--summarywrites counters and rates as JSON; always pass--chromwith a multi-chromosome sites table, since only the first FASTA record is read.
Used on this site
- Demography of rare mutations — the exact and ball-arithmetic spectra, the mixing layer and the dominance evaluators behind that page's audit of DFE-inference software; that page links the same downloads.
Requirements and source
Python 3 with mpmath. Optional, each enabling the parts that need it: python-flint (the ball-arithmetic route and vector mode above $n=2000$, popcorn.dfe, cone_lp_b, one positivity route), gmpy2 (the linear-programming certificates), sympy (root counting), numpy and scipy (floating-point proposers for the certificates, and required by popcorn.twolocus and popcorn.transient). popcorn.lambda_coalescent and popcorn.ancestral need only the standard library. Quick test: python3 tools/popcorn/certsfs.py selftest --dps 40. Full suite: python3 tools/popcorn/selftest.py (add --full for the slow checks), from the repository root; a check whose optional dependency is missing is skipped by name and does not fail the suite. Source: tools/popcorn/, version 1.0, in the bootloops-dev repository, released under the MIT license. The same package is distributed from this site as an archive, popcorn-package.tar.gz, with the package guide MANUAL.md and a checksum list MANIFEST.sha256 beside it (sha256sum -c MANIFEST.sha256 verifies the downloads). A separate archive, popgen-dfe-release.tar.gz, is the evaluator bundle posted with the first paper on the rare-mutations page: the engine modules as used there with their check driver, the stored outputs of the validation runs and their comparison scripts.