Before a language can be translated, taught, or preserved, it must first be recorded — in the voices of those who still speak it.
The linguist records actual language use: words, sentences, conversations, stories, narratives, traditional knowledge, grammatical constructions, and audio — building a corpus from which the language can be studied.
This raw material — gathered patiently, elicited respectfully, and transcribed faithfully — becomes the corpus: the body of primary data from which a lexicon, a grammar, and eventually a Scripture translation can be built. It is the same discipline that stands behind every interlinear, every triglot, and every concordance in this collection — someone, somewhere, first had to sit with a native speaker and simply listen.
Tools and archives used by linguists documenting the world's remaining unwritten and under-resourced languages
An overview of the discipline: elicited wordlists, communicative events, and analytical discussions as the three pillars of a documentary corpus.
The standard reference on the world's roughly 7,000 languages — status, location, and number of speakers for every one, including the most endangered.
Free software for building a lexicon, interlinearizing texts, and analyzing grammar directly from field data.
A permanent digital repository at SOAS, London, for audio, video, and text documentation of endangered languages worldwide.
The Pacific and Regional Archive for Digital Sources in Endangered Cultures — a major open archive of field recordings and linguistic material.
A mobile app built for communities and field linguists to record, translate, and respeak oral narratives directly on a phone.
A collaborative online catalogue of endangered languages with samples, records, and community-contributed materials.
A nonprofit that trains community members and linguists to record and archive their own endangered languages.
"Every language holds a world. To document it is to keep that world from going silent."