AccStack: Catalog of Stress and Tone (linguistics)

The content on this page was written by AI under human supervision.

Most grammars say, in a sentence or two, which syllable of a word carries the stress and whether the language uses pitch to tell words apart. The ACCSTACK catalog collects those sentences for 5,561 languages, about seven in ten of the world’s spoken languages, and lists each one beside the coding it supports, with its source and page. A large language model compiled the catalog, reading the sources under the authors’ written instructions.

accstack.org — browse the catalog: every language’s entry with its quoted source and page, the bibliography, and the data files for download

What a grammar says about stress and tone

Say the English word banana aloud and the middle syllable comes out stronger than its neighbors: that syllable carries the word’s stress. In English the place of the stress must partly be learned word by word. Many languages instead place it by rule: Hungarian and Chechen put it on the first syllable, Swahili on the next-to-last. In the standard terms the paper uses, stress is fixed when its position follows from the word’s edges and number of syllables alone. It is weight-sensitive when a heavy syllable, such as one with a long vowel or a closing consonant, can pull it away from that position. It is lexical when the position must be learned word by word, and morphological when it is fixed relative to a root, stem, or affix rather than the whole word, though the same facts are sometimes described either way.

In the paper such rules are written with one symbol $\sigma$ per syllable, following the accents of Gordon’s 2002 survey: an acute accent ($\acute{\sigma}$) for the main stress, a grave accent ($\grave{\sigma}$) for a secondary stress, and a bare $\sigma$ for an unstressed syllable. Penultimate stress, as in Swahili, looks like this for words of two to five syllables:

$$\acute{\sigma}\sigma,\qquad \sigma\acute{\sigma}\sigma,\qquad \sigma\sigma\acute{\sigma}\sigma,\qquad \sigma\sigma\sigma\acute{\sigma}\sigma$$

Each group is one word length, with the accent second from the end. Badimaya (Pama-Nyungan, Australia) and Karelian (Uralic) put the main stress on the first syllable and a secondary one on every second syllable after it, never on the last:

$$\acute{\sigma}\sigma,\qquad \acute{\sigma}\sigma\sigma,\qquad \acute{\sigma}\sigma\grave{\sigma}\sigma,\qquad \acute{\sigma}\sigma\grave{\sigma}\sigma\sigma$$

In the released data files a rule is written as a string of digits, one per syllable with 1 for a stressed one: “10 010 0010” is penultimate stress in words of two to four syllables.

Many descriptions report something else: tone, pitch that distinguishes one word from another, as in Mandarin Chinese or Yoruba, or a pitch accentA system in which each word has at most one syllable marked out by pitch; databases differ on whether to count it as tone or as a class of its own.. A description of a tone language often says nothing about stress, which can leave open whether the language has word stress at all. Such a description is entered in the catalog with its report of tone alone; a language described with both tone and a stress rule is classed by the rule, with the tonal statement quoted beside it.

Fifty years of compiling

Grammar statements like the ones above, collected across hundreds of languages, are the raw material of the typology of word stress and tone. In his 1977 survey Larry Hyman tallied the dominant position of main stress in the 444 descriptions then available to him. The 2002 survey by Matthew Gordon, one of the four authors of the present paper, covered 262 languages “with predictable quantity-insensitive stress patterns.” The StressTyp database of Goedemans and van der Hulst (2009) coded 510 languages, and the stress chapters of the World Atlas of Language Structures (WALS) map its codings. In 2014 it was merged with Heinz’s Stress Pattern Database into StressTyp2, which covers 699 languages in its archived version and over 750 online. Tone has its own databases drawn from the same literature, among them LAPSyD, chapter 13 of WALS, and PHOIBLE.

ACCSTACK takes over the fixed-stress codings and references of StressTyp2 and of Gordon’s survey, and the authors describe the catalog as continuing what those surveys began. It goes beyond them in two ways: in what an entry shows and, as far as the accessible literature allows, in coverage. A row of Gordon’s appendix has no column for page, quotation, variety, or the domain of the rule, and a StressTyp2 entry gives page and quotation only where a passage has been entered. In ACCSTACK the quoted passage, with its page, is the center of every entry. As for coverage, GlottologThe standard open catalog of the world’s languages, dialects, and families, with a bibliography of the descriptive literature on each. lists about 7,800 spoken languages, and for 5,468 of them its bibliography contains 54,586 grammars, sketches, phonologies, and works whose titles name sounds, stress, or tone. In the words of the abstract, a catalog on that scale “is beyond what a linguist or a small group could compile by hand in any reasonable time.”

How ACCSTACK was compiled

The catalog starts from the 7,803 spoken languages that Glottolog (version 5.2) lists once sign languages, pidgins, argots, ritual and artificial languages, and provisional entries are set aside. These languages are the “frame” against which coverage is measured, with one entry per language. The reading of the sources was done in August and September 2026 by large language models of Anthropic’s Claude family, run with tools for fetching, saving, and searching documents. The authors wrote the coding rules, the classes of exception, the confidence scale, and the procedure for settling disagreements; the models applied them; and every change of classification and every release is the authors’ decision.

Every entry rests on two separate readings of its source. In the first, the model located an accessible copy of the work. It found the passage on stress by searching for the words for stress, accent, prominence, and tone in whatever language the work is written in (Akzent, acento). It quoted the deciding passage verbatim, with the page and a translation where needed, and coded it from a closed list. One instruction quoted in the paper reads: “Never invent a quotation or a page: every quote comes from text you opened; if you could not open a work, say so and leave its quote empty.” The second reading was made with the first report in hand, under instruction to refute it. The second reader, again a model, opened the work at the cited page, where possible, and checked that the passage is there word for word. It then read the surrounding section and checked that the work describes this language rather than a neighbor or namesake. Disagreements between the two readings went to a third reading and were decided field by field, with both earlier readings kept.

The entry for Seimat, an Austronesian language, shows the format. Its source is Wozna and Wilson’s Seimat Grammar Essentials (2005), page 7: “Stress normally occurs on the penultimate syllable of a word, except in a small group of trisyllabic words where reduplication occurs.” The reading is fixed penultimate stress with the word as its domain and a listed set of exceptions, classed under compounds or reduplications; the quotation was found verbatim at the cited page when the entry was checked; and the confidence is “fairly confident: one first-hand description, checked against the source.”

Users may differ on what counts as an exception to a fixed-stress rule, so the classes of exception a source states are coded rather than ruled on. Three nested cuts then let the user keep every entry with a dominant rule, only those exceptionless for the core vocabulary, or only those exceptionless outright. Each entry has one of five confidence levels, from “very confident” (“clear consensus”) to “not confident” (“basically just guessing”), three of them set by rule from the evidence in the entry. When descriptions of a language disagree, no compromise is coded: either the type is “descriptions disagree,” or the preferred classification stands with the dissenting source, page, and words beside it.

What the catalog holds

The release of 29 September 2026 (version 1.0) has 5,561 entries: 5,530 of the 7,803 frame languages, or 71 percent, plus 31 entered outside the frame (described pidgins, argots, and one artificial language). The largest class is tone or pitch accent with no word-stress statement (1,725 entries), then fixed stress (1,543), with weight-sensitive, lexical, and morphologically assigned stress behind them. For 123 languages the source says there is no word stress, and 528 are assigned no type, most often because the sources do not determine one. A tone statement is recorded in 2,837 entries whatever their stress type. The 363 entries rated not confident lack a confirmed first-hand statement, and the authors ask that their classifications be treated as provisional.

Horizontal bar chart titled ACCSTACK entries by stress type, release of 29 September 2026. Seven bars with counts at their ends, colored to match the map below: tone or pitch accent with no word-stress statement 1,725 (purple); fixed quantity-insensitive stress 1,543 (blue); weight-sensitive stress 724 (orange, split into stated as a rule 339 and stated with hedges 385); lexical stress 507 (green); morphologically assigned stress 411 (green); no word stress 123 (dark gray); no type assigned 528 (dark gray, split into not determinable 365, descriptions disagree 46, segmentally conditioned 44, phrase-level only 45, variable 28). Horizontal axis: catalog entries, one per language, 5,561 in all, from 0 to 2,000.

The 5,561 entries by stress type. Each bar is one value of the catalog’s stress-type field with the count of languages at its end; the weight-sensitive bar is split into rules stated outright (339) and stated with hedges (385), and the no-type bar into its five reasons, the commonest being that the sources do not determine a type (365). Colors match the classes of the map below. Descriptions that report tone or pitch accent and say nothing about word stress are the largest class and fixed stress the second, and more than half of all entries have a word-stress type. (Drawn from Table 2 of the paper.)

Among the fixed-stress entries whose rule is complete for words of three to eight syllables there are 51 distinct patterns, 26 of them found in a single language each. Coverage runs from 63 percent of the languages of PapunesiaGlottolog’s macroarea for island Southeast Asia, New Guinea, and the Pacific. and 69 percent of Africa’s to 84 percent in North America. If each isolate is counted as a family of one, the languages with an entry belong to 413 families, led by Atlantic-Congo (946) and Austronesian (889).

World map of the 7,803 spoken languages as colored points at their Glottolog coordinates, without coastlines. Legend, every class a round dot told apart by color alone: blue fixed stress (1,530), orange weight-sensitive stress (723), green lexical or morphological stress (913), purple tone or pitch accent with no word-stress statement (1,721), dark gray no word stress or no type assigned (643), light gray no entry (2,273). Purple dots mass across sub-Saharan Africa and across East and mainland Southeast Asia; blue dots are dense in Australia, island Southeast Asia, northern Eurasia, and the Americas; green dots cluster in Europe, the Caucasus, and the Americas; light gray dots, many of them hidden beneath the colored marks, are most numerous in New Guinea and Africa.

Stress and tone types from the descriptions, on a world map. Each of the 7,803 spoken languages of the frame is a point at its Glottolog coordinates, marked by the stress type of its catalog entry in the release of 29 September 2026, one color per class: blue fixed stress (1,530 languages), orange weight-sensitive stress with hedged statements included (723), green lexical or morphological stress (913), purple tone or pitch accent with no word-stress statement (1,721), dark gray no word stress or no type assigned (643), and light gray the 2,273 languages with no entry; the legend counts languages rather than entries, and the 82 languages without coordinates in Glottolog are not shown. The purple of tone-only descriptions is densest across sub-Saharan Africa and East and Southeast Asia, and the blue of fixed stress across Australia; the light gray of missing entries, much of it hidden beneath the other marks, falls mostly in Papunesia and Africa, the two macroareas with the lowest coverage (63 and 69 percent). (Figure 1 of the paper.)

Of the 2,273 frame languages with no entry, 734 had accessible descriptions that were silent on stress at the first search, 610 a description that could not be reached, and 887 no known descriptive source. The companion bibliography lists 151,570 works on the frame languages, a median of 15 per language, each tagged with what it is known to say about stress or tone and, where that is known, how a copy can be obtained. For each language without an entry it says what is lacking.

Checks, and the earlier databases

The checks reported in the paper are themselves readings by the model. The second reading corrected about one entry in three in some field, not necessarily the stress type. When the catalog stood at 1,883 languages every entry was re-read against its source. The quotation was confirmed in the copy on file for 1,754, and for one language it was not found at the page cited.

The 394 entries taken over from StressTyp2 and Gordon’s survey were each read against the sources the database cites, at the page where one is given. On the catalog’s reading those pages describe fixed stress for 283. The other 111 are entered, from the same pages, as weight-sensitive (39), morphological (32), lexical (19), or in smaller classes, with the database’s coding shown beside the catalog’s. Across the catalog 333 entries have such a note. Each records a difference between the catalog’s reading of one cited description and a coding that may rest on other sources or on another analysis of the same data. In the paper’s words it “is recorded with the passage for the user to weigh, and is not a finding that the database coding is wrong.”

Four panels of horizontal bar charts titled Agreement between ACCSTACK and earlier databases, over the languages both code; subtitle: share of languages whose catalog value is among the database's values (Table 8 of the paper). Left column, three blue panels for four stress databases (WALS chapters 14–15, StressTyp2, Stress Pattern Database, Gordon 2002): Which syllable a fixed stress falls on, 93.8, 95.8, 94.7, 96.4 percent; Whether syllable weight moves the stress, 88.9, 89.8, 86.3, 88.5 percent; The type of stress system, 59.5, 64.7, 62.3, 69.3 percent. Right column, one pink panel, Whether the language has tone, for seven databases: WALS chapter 13 87.6, ThoTDB tonal list 86.2, Maddieson and Benedict 85.8, LAPSyD 85.1, World Phonotactics 2026 79.3, World Phonotactics 2014 76.9, PHOIBLE 68.7 percent. All axes run 0 to 100 percent.

Agreement between the catalog and earlier databases, field by field. Each bar is the share of the languages coded by both the catalog and the named database on which the catalog’s value is among the database’s values. Left, four stress databases: the syllable a fixed stress falls on (over 176, 264, 171, and 137 languages), whether syllable weight can move the stress, and the type of stress system (over 477, 510, 334, and 202 languages). Right, whether a language has tone, against seven databases (over 388 languages for WALS chapter 13 and between 504 and 1,911 for the others); PHOIBLE is lowest; its inventories list tones only when their source describes them, and in the comparison an inventory without tone segments counts as toneless. Position and weight agree for roughly nine languages in ten and the presence of tone for more than four in five against four of the seven databases shown, while the finer-grained stress type agrees for only six or seven in ten. (Drawn from Table 8 of the paper.)

Thirty-one earlier databases were matched to the catalog through Glottolog codes and compared on the fields both hold. Where both code a fixed position, the positions agree for about 94 to 96 languages in 100. On whether a language has tone, the catalog agrees with WALS chapter 13, LAPSyD, Maddieson and Benedict, and ThoTDB for more than four in five. On the type of stress system agreement drops to between one half and three quarters against every database but the Stanford Phonology Archive, whose entries describe stress in prose. Several databases code stress type in only two or three values, where the catalog’s types are finer. Against WALS and StressTyp2 the commonest disagreement is a language the database enters as fixed and the catalog as morphologically assigned, lexical, or weight-sensitive. The authors add a caution about independence: of the 510 languages the catalog and StressTyp2 both code for stress type, 349 have entries that began from StressTyp2 or Gordon’s survey. Those agree with StressTyp2 at 73.4 percent and the 161 read fresh at 46.0 percent, the two groups differing in their stress systems as well as in origin.

The model re-read the sources for a random sample of 360 disagreements across nine databases. Of the 360, 191 (53 percent) are defensible under either coding, 130 (36 percent) support the catalog’s value, 11 (3 percent) the database’s, one neither, and 27 are contested. The commonest cause, assigned 189 times, was a coding convention; an error in the catalog was assigned 13 times; nine of those were corrected on that account, one of them since reversed on a later first-hand description, so that eight have the database’s value in this release. As an example of a difference in convention, a pitch accent counts as tone in WALS chapter 13 and as a class of its own in the catalog; Somali, in which Hyman (1981) writes that “a word can have only one H tone,” is such a case. In the paper’s words, no such case is an error in either coding: “each is a choice made when the coding scheme was drawn up, and the catalog’s entry quotes the passage so that the other choice can be made from it.”

Limits, and how to use the catalog

The main limit, as the authors state it, is that the catalog was compiled by a large language model, not by linguists reading every source. The authors spot-checked entries for languages they know and state that they cannot attest that every entry is correct; the checks compare one model reading with another, or with an earlier database’s coding. Each coding is also a reading of a written description and can be no more complete or accurate than it. Variation that no description reports is absent; impressionistic statements about secondary stress can reflect a describer’s expectations as well as the acoustic signal; and a mistake shared by every description of a language is carried into its entry. Every entry does give the source passage, a translation where needed, and the interpretation the coding rests on, so that a reader can go to the page and judge it.

The catalog and the bibliography are at accstack.org, where every entry is shown with all its fields and, for 4,142 languages, links to recordings. Earlier releases are kept on the site with the full reading instructions, and corrections may be proposed through it, naming the language and, where possible, the page. The authors expect the catalog to change as more descriptions become openly available and as its users correct it.

The paper

A second project by Matthew D. Schwartz, Sharpening language trees with exact statistics and AI-curated cognate sets, is in preparation. Family relationships among the world’s languages are decided by Bayes factors computed over expert-curated word lists; that project computes those quantities exactly, with no Monte Carlo sampling, and rebuilds the word lists with AI language models.

Supplementary material

The release the paper describes is distributed from the catalog’s own site rather than from this one.

References

L. M. Hyman, On the nature of linguistic stress, in Studies in Stress and Accent, Southern California Occasional Papers in Linguistics 4 (1977) 37tallied the dominant position of main stress in the 444 descriptions then available
M. Gordon, A factorial typology of quantity-insensitive stress, Nat. Lang. Linguist. Theory 20 (2002) 491the 262-language fixed-stress survey whose codings, references, and accent notation the catalog takes over
R. Goedemans and H. van der Hulst, StressTyp: a database for word accentual patterns in the world’s languages, in The Use of Databases in Cross-Linguistic Studies, Mouton de Gruyter (2009) 235the 510-language stress database whose codings the WALS stress chapters map
R. Goedemans and H. van der Hulst, Fixed stress locations, in The World Atlas of Language Structures Online, chapter 14 (2013)the WALS map of fixed-stress positions, one of the stress codings the catalog is compared with
J. Heinz, On the role of locality in learning stress patterns, Phonology 26 (2009) 303cited for Heinz’s Stress Pattern Database, later merged with StressTyp into StressTyp2
R. Goedemans, J. Heinz and H. van der Hulst, StressTyp2, st2.ullet.net (2014), version 1 archive of April 2015the 699-language stress database whose fixed-stress codings and references the catalog reuses
H. Hammarström, R. Forkel, M. Haspelmath and S. Bank, Glottolog 5.2, Max Planck Institute for Evolutionary Anthropology (2025)the catalog of languages and descriptive sources that defines the 7,803-language frame
I. Maddieson, Tone, in The World Atlas of Language Structures Online, chapter 13 (2013)the WALS tone chapter, which classes 527 languages as toneless, simple, or complex; one of the tone codings compared
I. Maddieson, S. Flavier, E. Marsico, C. Coupé and F. Pellegrino, LAPSyD: Lyon-Albuquerque Phonological Systems Database, Proc. Interspeech (2013) 3022a 764-language database with stress and tone categories, compared with the catalog on both
I. Maddieson and K. Benedict, Global linguistic data (version 1.1.2), Zenodo (2024), doi:10.5281/zenodo.11205300tone and stress categories for 1,003 languages, compared with the catalog on both
S. Moran and D. McCloy, PHOIBLE 2.0, Max Planck Institute for the Science of Human History (2019), doi:10.5281/zenodo.26266873,020 phoneme inventories for 2,186 languages, with tones where the source describes them; lowest in tone agreement among the seven databases charted above
K. Maslinsky, V. Vydrin and D. Gerasimov, ThoT Database, thot.huma-num.fr (2025)ThoTDB, which marks each Glottolog language as tonal, toneless, or unknown; one of the tone codings compared
L. M. Hyman, Tonal accent in Somali, Stud. Afr. Linguist. 12 (1981) 169the Somali description quoted as a case where conventions for coding pitch accent differ
B. Wozna and T. Wilson, Seimat grammar essentials, Data Papers on Papua New Guinea Languages 48, SIL (2005)the grammar quoted in the sample entry: penultimate stress in Seimat, page 7

← back to the web summaries