Collocate 2 web edition

collocates · n-grams & bundles · P-frames · batch — MI · t · log-likelihood

1 · Corpus text settings ▾

No corpus loaded.

2 · Collocates of a node

Statistics (per collocate c, node n): O = tokens of c inside the span windows; W = total window tokens; N = corpus tokens; E = f(c)·W/N. MI = log₂(O/E) · t = (O−E)/√O · LL = Dunning log-likelihood on the 2×2 window table, signed (− = repelled). Click column headers to sort; click a collocate for its concordance.

5 · N-grams (whole corpus)

N-gram MI generalises pointwise MI: log₂( f(gram)·N⁽ⁿ⁻¹⁾ / Π f(wᵢ) ), as in the Windows version. N-grams do not cross sentence boundaries. With min range > 1 this panel implements the lexical-bundle criteria: recurrent word sequences meeting both a frequency threshold (see the per-million column) and a dispersion threshold — occurring in at least k of the loaded files.

6 · P-frames

A P-frame is an n-gram with one position opened as a slot (e.g. the * of the). Filler types counts the distinct words filling the slot — high-frequency, high-variability frames are productive constructional patterns; low-variability frames are near-fixed bundles. Frames do not cross sentence boundaries. Memory grows with corpus size; for corpora above ~5M words raise the min frequency.

7 · Batch collocates

Uses the span, min window frequency, stopword, and sentence-boundary settings from panel 2. The CSV contains the full collocate table for every node, not just the displayed top rows.