1 · Corpus text settings ▾
No corpus loaded.
2 · Collocates of a node
Statistics (per collocate c, node n): O = tokens of c inside the span windows; W = total window tokens; N = corpus tokens; E = f(c)·W/N.
MI = log₂(O/E) · t = (O−E)/√O · LL = Dunning log-likelihood on the 2×2 window table, signed (− = repelled).
Click column headers to sort; click a collocate for its concordance.
5 · N-grams (whole corpus)
N-gram MI generalises pointwise MI: log₂( f(gram)·N⁽ⁿ⁻¹⁾ / Π f(wᵢ) ), as in the Windows version. N-grams do not cross sentence boundaries. With min range > 1 this panel implements the lexical-bundle criteria: recurrent word sequences meeting both a frequency threshold (see the per-million column) and a dispersion threshold — occurring in at least k of the loaded files.
6 · P-frames
A P-frame is an n-gram with one position opened as a slot (e.g. the * of the). Filler types counts the distinct words filling the slot — high-frequency, high-variability frames are productive constructional patterns; low-variability frames are near-fixed bundles. Frames do not cross sentence boundaries. Memory grows with corpus size; for corpora above ~5M words raise the min frequency.
7 · Batch collocates
Uses the span, min window frequency, stopword, and sentence-boundary settings from panel 2. The CSV contains the full collocate table for every node, not just the displayed top rows.