A Bible translator's field guide to the science of linguistics — no degree required, just a pan, a practiced eye, and a claim worth working.
"The words of the LORD are pure words: as silver tried in a furnace of earth, purified seven times." — Psalm 12:6
Every world language is a claim waiting to be worked. Somewhere inside the sentences ordinary people use to talk about their fathers, their kings, their sicknesses, and their sky, there already exist the very words God intends to use to tell them He loves them. A Bible translator's first job is not to invent those words — it is to find them. That is prospecting. And like any prospector, the translator does better work with a little geology under the belt.
This page is that geology lesson — the fundamentals of linguistics laid out the way an old-timer would explain a claim to a greenhorn: plainly, practically, and pointed always at where the gold actually sits. No jargon for jargon's sake. Just what you need to read the ground.
Linguistics is simply the disciplined study of how human language works. A Bible translator does not need a PhD in it — but does need to recognize these five veins on sight.
Every language has its own limited set of sounds and its own rules for which sounds may sit beside which. Before you can spell a word correctly in a heart language, you have to hear it correctly — every click, tone, glottal stop, and nasal vowel the ear of a trade-language speaker is trained to miss.
Words are rarely single lumps. They are built from smaller pieces (morphemes) stacked in patterned ways — a prefix for "my," a suffix for "plural," a root that never changes. Learn to crack a word open and you can predict a hundred others like it.
Every language has its own bedrock order: Subject-Verb-Object, or Verb-first, or "the verb carries everyone else on its back" (as in many Native American and African languages). Fight the formation and your translation reads like a foreigner talking. Work with it and it reads like Scripture that was always meant to live there.
Two words can look like a perfect match in a dictionary and still carry entirely different weight. "Lamb" means something different to a herding culture than to a people who have never seen a sheep. Semantics is the discipline of testing purity before you trust the nugget.
Almost every community that needs Scripture already speaks two languages: the heart language — the one of lullabies, grief, and grandmothers — and a trade language — the one of government offices, markets, and school. Translate only into the trade language and you have handed people Scripture in a second tongue. The gold was always in the heart language.
The steps a gold prospector walks across a hillside map almost exactly onto the steps a Bible translator walks across a language. Here is the same claim, worked with a different pan.
Before you touch the ground, study the maps. Has any portion of Scripture, any hymnal, any catechism, any wordlist ever been produced in this language or a close relative of it? Old missionaries and linguists rarely worked out every vein — but they rarely worked in a place with none. Search digital Bible libraries, language archives, and university theses before assuming you start from nothing.
Find the fault line where the two languages meet in a bilingual community — where French met Abenaki, where Swahili meets a village tongue, where Spanish meets a highland Quechua dialect. That contact zone is where loanwords, borrowed theology terms, and mixed vocabulary usually cluster, for better and worse.
Get the sound system, the word-building patterns, and the basic sentence order down on paper before you draft a single verse. A translation built on guessed grammar rusts fast — native speakers notice immediately, even if they cannot say why.
In any river, gold collects in predictable places — inside bends, behind boulders, in bedrock cracks. In a living language, key vocabulary collects in predictable places too: proverbs, folk tales, praise names, funeral laments, and existing hymnody. That is where the words for "holy," "spirit," "covenant," and "king" are usually already sitting, worn smooth by centuries of use.
Yesterday's discarded work is today's paydirt. Old missionary diaries, colonial-era wordlists, government linguistic surveys, seminary theses, and heritage-language Facebook and YouTube groups are digital tailings piles — full of partially processed material nobody has assayed for Bible translation purposes. (See the research methodology below.)
Never commit a word to Scripture without testing it. Read the candidate word or phrase back to native speakers of every age and background. Watch their faces before you trust their words. A word that sounds "close enough" to an outsider can carry the wrong weight, the wrong register, or an entirely different meaning to the people who will actually pray with it.
Before assuming a language is untouched ground, dig in these six digital claims. A surprising number of "unreached" languages already have scattered vocabulary sitting in plain sight.
site: filters on known archive domainsLanguage #13 on AFII's list of the 7,000 world languages still needing Scripture — picked out for a closer look because it is far enough along to teach a method, and far enough behind to still need one.
Eastern Abenaki is an Algonquian language once spoken across what is now Maine by the Penobscot and neighboring peoples. Its last fully fluent native speaker, Madeline Tower Shay, died in 1993 — yet the language is not gone. Penobscot elders still hold pieces of it, the University of Maine and the Penobscot Nation have partnered on documentation and teaching efforts, and roughly 500 words remain in everyday community use. It is exactly the kind of "worked-but-not-exhausted" claim a prospector should not walk past.
Its heart language is Abenaki itself — the tongue of kinship terms, land, and ceremony. Its historical trade languages were French (through the colonial fur trade and Catholic missions) and later English. The very first Gospel portion — Mark, translated by Wzokhilain (Paul Pierre Osunkhirhine) — was printed in 1844, which means this claim already has an "old-timer working" on record, a full 180 years before the AFII 7,000-language initiative began surveying it again.
Below is the opening line of the Lord's Prayer in Eastern Abenaki, broken into its working pieces the way a translator would lay them out on the table before drafting further verses.
| Abenaki | Gloss | What to Notice |
|---|---|---|
| Nmitôgwesna | Our Father | N– "my/our" prefix + root for "father" + –na plural-possessor suffix, all fused into one word — typical of Algonquian polysynthesis, where a whole English phrase becomes a single Abenaki word. |
| spempik | in heaven | The –ik ending marks location — "at/in" a place — doing the job English does with the separate word "in." |
| aiian | who art | Ends in –ian, a second-person marker ("you..."). |
| sôgmowal | holy | Compare to sôgmo, a root connected to lordship/chieftainship in Abenaki — "holy" is expressed through the vocabulary of rightful rule, not abstraction. |
| negwadji | may be | An optative/hortative particle — "let it be so" — carrying the force of the prayer's request. |
| eliwisian | your name | Notice the same –ian ending seen in aiian — a second-person suffix appearing again. Spotting a repeated ending across unrelated words is exactly how a prospector begins to map a grammatical vein. |
| K'tabaldamwôgan | your kingdom | K'– is a second-person prefix — the same grammatical "your" signaled elsewhere by –ian, now shown as a prefix instead. One language, two positions for the same meaning. |
| paiômwiji | come | A verb of arrival, carrying the sense "let it come to us." |
For sixty years, this has been the standard training manual carried by nearly every Bible translator in the field.
Dr. Katharine Barnwell has served with SIL International since 1963, and her textbook — now in its fourth edition — remains the closest thing the Bible translation world has to a single shared field manual. Its central conviction is simple: a translation's job is to carry the meaning of the source text into the receptor language, not to reproduce its word order or grammatical form. Word-for-word translation, however faithful it looks on paper, frequently smuggles in the wrong meaning — or no meaning at all — to a reader whose language works on entirely different rules.
Faithful to the meaning of the source text
Understandable to the ordinary reader or hearer
Reads as if it were composed in the receptor language
Uses a style and register the community will embrace
Barnwell walks the translator through the whole process: exegeting the source text before touching the target language; drafting first with the ear rather than the eye, since most receptor communities are primarily oral; testing every draft with genuine community members who were not involved in producing it; gathering back-translations so a consultant who does not speak the language can still check the work; and revising as many times as the checking process demands. The fourth edition — reflecting the changed landscape of the field — adds substantial new material on oral Bible translation, on software and online resources such as Paratext, and on translation as team ministry, where exegetes, drafters, community testers, and consultants each carry a distinct and necessary role rather than one lone translator carrying the whole claim alone.
The book is available through SIL International's resource archive, with a full set of chapter-by-chapter supplementary materials provided free alongside it.
A pan and a practiced eye still matter most. But today's translator also has power tools — and knowing what each one actually does keeps the gold real instead of fool's gold.
A computer program trained on enormous amounts of text so that it learns the statistical patterns of how words follow words. It has not "learned" a language the way a person does — it has crushed vast quantities of linguistic ore and learned to predict, very well, what pattern usually comes next. Given enough already-translated Scripture in a language, it can draft plausible sentences in that same language — always to be checked, never to be trusted blind.
An AI model trained specifically on spoken audio rather than written text — the technology behind natural-sounding Audio Bibles. For the majority of the world's languages, which are primarily oral rather than literate, a Voice Model that can render Scripture into clear, natural-sounding speech may reach more hearts than the printed page ever will.
Faith Comes By Hearing ↗ — audio Scripture recording for oral cultures
The professional software translation teams actually work inside — the assay office and the workbench in one. It holds the source texts, keeps every draft version, runs consistency and spell checks, and lets a whole team and its consultants collaborate on the same project from anywhere in the world.
Paratext's companion tool for turning a checked draft into a beautifully typeset, print-ready PDF — booklets, trial editions, and full Bibles, laid out automatically from the same files Paratext uses.
SIL's cloud companion to Paratext, built for community checking and — since recent years — AI-assisted draft generation. Once a language project has a solidly translated New Testament (enough "assayed ore" to train on), Scripture Forge can generate a rough first draft of an untranslated book, giving the team a starting point to edit rather than a blank page to fill.
The AI never replaces Steps One through Six above — it simply hands the team a shovel-ready pile of dirt to pan through, instead of an empty hillside. SIL's own data shows teams using this drafting assistance are, on average, drafting and reviewing roughly a thousand more verses than teams without it.
Everything in the AI Tools section above quietly assumes something: that a language already has a pile of digitized text sitting somewhere for a model to learn from. For most of the world's 7,000-plus languages — the majority still without a complete Bible — that assumption simply does not hold.
Strip away the hype and a "model" is just a trained software system: a program that has been shown enormous quantities of example text (or audio) in a language and has learned, statistically, which words and sounds tend to follow which others. It hasn't learned the language the way a child does. It has compressed millions or billions of sentences into a set of patterns it can then generate more of. Feed it a well-translated New Testament and enough surrounding text, and it can draft plausible new sentences in that language — the Scripture Forge example above.
But that whole process depends entirely on the size of the pile you start with. English, Spanish, Mandarin, and a few hundred other languages have that pile already — billions of already-written, already-translated words sitting on the internet for a model to learn from. A low-resource language is simply one where that pile is thin or does not exist at all: little or no digitized text, little or no audio, sometimes no settled writing system whatsoever. Since most of the roughly 3,000+ languages still awaiting a first word of Scripture fall into exactly this category, a strategy that only works once a New Testament is already finished skips the hardest and most common claim on the whole map: the one with no ore in the ground yet.
Before a model can learn anything, someone has to put words into digital form in the first place. That is field linguistics — and it has its own fast, proven method.
Developed by SIL International's Ron Moe and refined by SIL's Dictionary & Lexicography Services, Rapid Word Collection (RWC) is a workshop method that gathers a language's vocabulary by "semantic domain" — clusters of related words (kinship terms, farming, the human body, the sky) — using prepared elicitation questions that prompt mother-tongue speakers to produce word after word from memory, rather than waiting for a linguist to stumble across each one over years of fieldwork. A two-week community workshop routinely produces 10,000 or more raw word entries, glossed and entered on the spot — precisely the kind of structured, digitized lexical data a model needs, generated directly by the people who speak the language.
rapidwords.net — SIL's Rapid Word Collection ↗ click on these links FieldWorks Language Explorer (FLEx) — the software RWC data is entered into ↗
A word list alone is ore. A model learns faster from ore that has already been sorted along several dimensions at once.
A picture-linked, native-speaker-recorded resource creates something a model desperately needs for a low-resource language: aligned data. For each entry, four things are captured together —
That is a genuine parallel corpus — image, sound, and two languages linked together. Even a few thousand entries built this way is far more structured digital data than most under-documented languages currently have in any form.
If the language has never been written down, this resource does double duty — it is the linguistic fieldwork that decides how the language will be spelled, and it is the training data, produced in the same motion.
Once a model has enough aligned audio and transcription, future oral recordings — sermons, storytelling, more Scripture — can begin to be transcribed automatically rather than by hand, one recording at a time.
Asking for a full paradigm forces a native speaker to reveal how the language actually marks tense and aspect. Many languages don't mark past/present/future at all the way English does — some mark completion versus incompletion instead. That grammatical pattern is exactly what a model can generalize from a well-designed paradigm set (compare Station Three, above).
This is the same technique projects like Meta's No Language Left Behind use to seed translation models for languages with almost no digital footprint — small, carefully aligned seed corpora doing outsized work.
Abstract theological concepts get their meaning from narrative context — the flood, the Passover lamb, the exodus — not from isolated word-picture pairs.
Chronological Bible Storying (CBS) is a Bible-teaching and translation-support method used extensively by Faith Comes By Hearing and by ministries such as the OneStory Partnership ↗, in which Scripture is told as a sequence of connected stories in their biblical order rather than through topical teaching or systematic doctrine. It is built for oral cultures — an estimated two-thirds of the world learns and retains truth through narrative, not propositional argument or the written page. Each story is told, discussed, and reinforced before the next one is added, so later stories build their meaning on earlier ones already told.
The Passover account (Exodus 11–13) shows why storying outperforms word-for-word vocabulary work for theological concepts — its meaning is built entirely out of stories that came before it in the sequence:
Then Passover itself is told with its concrete, tellable details: a lamb without blemish, blood applied to the doorposts, unleavened bread, the angel of death "passing over" the marked houses, the command to keep this as a permanent memorial. A storyteller never has to reach for the abstract word "atonement" — the story simply shows it: death that should fall on the household falls instead on a substitute, marked by blood on the door.
After the story is told — usually as a recording, since oral cultures often have low literacy — guided questions draw out the meaning without imposing borrowed Western theological vocabulary: What happened? Why did God tell them to put blood on the doorposts? What would have happened if a family disobeyed? What does this show about how God deals with sin and judgment? This lets believers articulate substitution, judgment, and deliverance through blood in their own language's natural idiom, rather than forcing a loanword for "atonement" that may carry no meaning — or the wrong one — in that culture.
CBS sets are deliberately sequenced so Passover is not the end — it is the hinge that later connects to the Last Supper and the crucifixion, "Christ our Passover lamb" (1 Corinthians 5:7). A listener who has truly absorbed the Exodus Passover story has, without a page of doctrinal vocabulary, already grasped the scaffolding needed to understand the Gospel's use of lamb, blood, and substitution — which is the entire point of storying the Old Testament before the New.
Connecting this back to the model: an aligned corpus of {story audio in the target language} + {story audio in the bridge language} + {comprehension questions and answers in both} gives a model — and its human checkers — real evidence for how a language expresses concepts like substitution, guilt, and deliverance implicitly, through narrative choices, rather than demanding an isolated word-for-word gloss that may not exist for words like "propitiation" at all.
The picture-dictionary approach is a real and valuable accelerant for the concrete layer of the problem — but it needs the oral-storying layer alongside it to reach the abstract layer Scripture actually depends on.
Neither vein alone reaches the whole claim. A picture dictionary hands a model the concrete layer — nouns, verbs, and the grammar exposed by a well-built paradigm — while Chronological Bible Storying hands it the abstract layer that concrete vocabulary can never picture. Run through a model for a first draft and then panned by hand through community testing and consultant checking (Station Six, above, is never optional here either), the two together give a Bibleless, low-resource language a real path forward — not from a finished New Testament, as the AI Tools section above assumes, but from nothing but native speakers and a recorder.
Fill in a candidate word for each essential biblical term as you find it in your target heart language. Note where you found it and how confident you are before it goes to community testing.
| Biblical Term | Heart-Language Word | Where Found | Notes / Confidence |
|---|
Click Print to generate a paper field copy, or fill the fields directly on screen.