Artists for Israel International · AFII Institute

The 7,000+ Languages of the World: An Exhibit

Using Closely Related Resourced Languages to Train the Models to Create Native-Speaker-Reviewed Drafts of Low-Resourced but Related Languages

From the genealogies of Israel to the families of the world’s 7,132 tongues: how a translation model that already knows a language’s well-resourced cousin can propose a first draft, and why only native speakers can turn it into Scripture, toward every tongue reading the Besuras HaGeulah by Shavuos, June 2, 2033.

WE PRAY IN THE NAME OF THE OYBERSHTER THAT ELOKIM HAAV WILL REVERSE THE WORLD’S BABEL, AS THE RUACH HAKODESH DID ON SHAVUOS 33 C.E., FOR THEY SAID, “HOW IS IT THAT EACH OF US HEARS THEM (PREACHING THE BESURAS HAGEULAH) IN OUR OWN NATIVE LANGUAGE?”

Acts 2:8

We believe Linda, the V.P. of AFII, received her first kidney transplant on Shavuos (Pentecost) 2011 so she might live another 22 years — until Shavuos June 2, 2033 — to see this universal reversal of Babel: all 7,000+ languages given their Triglot, every word “Triglot unpacked,” so that anyone can pronounce and understand every word of the Bible in these tongues — in the dominant global lingua franca English.

I

Sefer HaYachas: How an Israelite Knew Whose He Was

ספר היחש

An Israelite did not think of himself as a free-floating individual. He stood inside four nested circles, and Scripture names every one of them: the shevet or matteh (tribe), the mishpachah (clan), the beit av (father’s house), and finally the gever, the man himself. The census of Numbers 1 is taken “by their clans, by their fathers’ houses” (Num 1:2), and on the day of mustering the people “declared their pedigrees” (Num 1:18), the Hebrew vayityaldu, a reflexive form of the verb “to beget”: they registered their descent.

The clearest demonstration of the four circles working is the lot at Ai. When Achan hid the devoted things, Yehoshua did not search the camp tent by tent. He narrowed: the tribe of Yehudah was taken, then the clan of the Zarchi, then the house of Zavdi, and at last the man, Achan ben Karmi ben Zavdi ben Zerach (Josh 7:1, 14–18). Four sieves, each finer than the last.

CircleHebrewJoshua 7Its counterpart among languages
Tribeשבט / מטהYehudahLanguage family (Indo-European)
Clanמשפחהthe ZarchiBranch (Slavic; Romance)
Father’s houseבית אבthe house of ZavdiSubgroup (East Slavic; West Ibero-Romance)
The manגברAchan ben KarmiThe language (Ukrainian; Spanish), with its dialects as the children of the house

Ezra’s yichus, all the way back to Levi

Ezra is introduced not with a résumé but with a genealogy. Ezra 7:1–5 carries him back sixteen generations to “Aharon the chief kohen.” Exodus 6:16–20 and 1 Chronicles 6:1–3 (5:27–29 in Hebrew Bibles) carry Aharon back through Amram and Kehat to Levi, and Genesis 29:34 names Levi as the son of Ya‘akov and Leah. Every link is on the page of Scripture.

Two things about this list matter for what follows. First, “ben” can mean descendant as well as son: Serayah was the chief kohen executed at Riblah in 586 BCE (2 Kgs 25:18–21), more than a century before Ezra went up to Jerusalem, so Ezra was his grandson or later. Second, the list is telescoped: set it beside 1 Chronicles 6:3–15 and you find Ezra skipping from Merayot to a later Azaryah. The line is true and unbroken in fact, even where the written record abbreviates it. Language family trees behave exactly the same way.

Ezra ben Serayah

Ezra 7:1–5; Exod 6:16–20; 1 Chr 6:1–3; Gen 29:34

  1. 1Ezraעזרא Ezra 7:1gever — the man — the sofer mahir of Ezra 7:6
  2. 2Serayahשריה
  3. 3Azaryahעזריה
  4. 4Chilkiyahחלקיה
  5. 5Shallumשלום
  6. 6Tzadokצדוק
  7. 7Achituvאחיטוב
  8. 8Amaryahאמריה
  9. 9Azaryahעזריה
  10. … here Ezra’s list passes over roughly six generations that 1 Chronicles 6:7–10 (5:33–36 in Hebrew Bibles) supplies …
  11. 10Merayotמריות
  12. 11Zerachyahזרחיה
  13. 12Uzziעזי
  14. 13Bukkiבקי
  15. 14Avishuaאבישוע
  16. 15Pinchasפינחס
  17. 16Elazarאלעזר Ezra 7:5
  18. 17Aharon HaKohen HaRoshאהרן Ezra 7:5beit av — the priestly father’s house
  19. 18Amramעמרם Exod 6:20
  20. 19Kehatקהת Exod 6:18mishpachah — the Kohathite clan — Num 3:27; 26:57
  21. 20Leviלוי Exod 6:16shevet — the tribe that bears his name
  22. 21Ya‘akov, called Yisraelיעקב Gen 29:34

Spanish (Castilian)

Its lineage as Glottolog records it, read upward like Ezra’s: each line the daughter of the next

  1. 1Spanishgever — the language itself
  2. 2Castilic
  3. 3West Ibero-Romancebeit av — the house Spanish shares with Portuguese and Galician
  4. 4Southwestern Shifted Romance
  5. 5Shifted Western Romance
  6. 6Western Romance
  7. 7Italo-Western Romance
  8. 8Romancemishpachah — the branch: daughters of spoken Latin
  9. 9Imperial Latin
  10. 10Latinic
  11. 11Latino-Faliscan
  12. 12Italic
  13. 13Classical Indo-European
  14. 14Indo-Europeanshevet — the family; its reconstructed father is Proto-Indo-European

Where a record failed, Scripture is honest about it. After the return, the sons of Chavayah, Hakkotz and Barzillai “sought their register among those enrolled by genealogy, but it was not found,” and they were set aside from the kehunah until a kohen should stand with Urim and Tummim (Ezra 2:61–63; Neh 7:63–65). Linguists keep the same kind of ledger. A few of the languages below are marked “unclassifiable” or simply have no record yet, and a hundred or so are isolates, languages with no demonstrable relatives at all, like Malki-Tzedek, whom the Letter to the Hebrews calls agenealogetos, “without genealogy” (Heb 7:3).

Tribal memory survived into the days of the Brit Chadasha. Anna the prophetess was “of the tribe of Asher” (Luke 2:36); Zecharyah served in the division of Aviyah and Elisheva was of the daughters of Aharon (Luke 1:5); and Rav Sha’ul could say without hesitation that he was “of the tribe of Binyamin, a Hebrew of Hebrews” (Phil 3:5; Rom 11:1).

II

After Their Families, After Their Tongues

למשפחתם ללשנתם

The Table of Nations closes each of its three sections with the same formula: the sons of Yefet, Cham and Shem, “after their families, after their tongues” (Gen 10:5, 20, 31). The Torah itself joins mishpachah to lashon. Modern linguistics has, in its own way, rediscovered the pairing: languages come in families, and the families can be traced.

The tool is called the comparative method. Linguists line up everyday words (water, night, son, hand) across languages and look for sound correspondences that repeat with the regularity of a law. Where Latin has f-, Spanish usually has a silent h- and Portuguese keeps the f-; where Latin has pl- or cl-, Spanish has ll- and Portuguese has ch-. Correspondences that regular cannot be coincidence or borrowing; they are inheritance. From them one can reconstruct an ancestor no one ever wrote down. Behind Latin, Greek, Sanskrit, Gothic, Old Irish and Old Church Slavonic stands Proto-Indo-European, a language reconstructed from its daughters, whose date and homeland are still debated. It stands to the Indo-European family as Levi stands to the tribe that bears his name, with one difference: Levi is attested in Scripture, and Proto-Indo-European is attested only in its children.

The whole of the world’s linguistic genealogy is kept by three complementary registries. Glottolog (Max Planck Institute for Evolutionary Anthropology) records a conservative classification into several hundred families and isolates, accepting only relationships that have been demonstrated. Ethnologue (SIL International) catalogs every living language with its speakers, status and ISO 639-3 code, and groups them somewhat more inclusively, for example keeping a large Niger-Congo family where Glottolog separates Atlantic-Congo, Mande and others. The Automated Similarity Judgment Program (ASJP) takes a 40-word list for thousands of languages and measures, by computer, how similar their basic vocabularies are. Gallery IV below uses Glottolog’s tree, because its language-by-language lineages can be joined to this page through the ISO code.

Two pairs of cousins

Spanish and Portuguese share every ancestor down to West Ibero-Romance and then part, like two sons of Zavdi. Russian and Ukrainian share every ancestor down to East Slavic. In both pairs the resemblance is strong enough that a speaker of one can follow a good deal of the other, especially in writing, and different enough that a text carried over mechanically will be wrong in dozens of places per chapter.

Regular sound correspondences: the family resemblance a machine can learn.
LatinSpanishPortugueseWhat happened
filiumhijofilhoLatin f- became a silent h- in Spanish; -li- became j in Spanish and lh in Portuguese
noctemnochenoiteLatin -ct- became -ch- in Spanish and -it- in Portuguese
plenum / clavemlleno / llavecheio / chavepl-, cl- become ll- in Spanish and ch- in Portuguese
lunam / manumluna / manolua / mãoPortuguese drops the n between vowels and nasalizes the vowel instead
portam / terrampuerta / tierraporta / terraSpanish breaks stressed o and e into ue and ie; Portuguese does not
MeaningRussianUkrainianWhat happened
cat / noseкот / носкіт / нісUkrainian turned o into i in closed syllables
breadхлебхлібOld East Slavic ѣ (yat) became e in Russian and i in Ukrainian
waterводаводаidentical
thank youспасибодякуюnot cognate: the Ukrainian word is related to Polish dziękuję
languageязыкмоваRussian uses one word for the organ and the speech; Ukrainian keeps язик for the organ and мова for the language

Three things that make cousins dangerous

Idiom
An expression whose meaning is not the sum of its words, and so cannot be translated word by word. Spanish tomar el pelo (“to take the hair”) means to tease; Portuguese pagar o pato (“to pay the duck”) means to take the blame for someone else; Portuguese engolir sapos (“to swallow toads”) means to put up with insults in silence. Russian вешать лапшу на уши (“to hang noodles on someone’s ears”) means to deceive; Ukrainian дати гарбуза (“to give a pumpkin”) means to turn down a suitor. Scripture is full of Hebrew idioms of its own: a people “hard of neck” (Exod 32:9), a man who “girded his loins” (1 Kgs 18:46). A cousin language may carry an idiom across that the target language does not use, or render literally one that the target expresses in its own way.
False friend
A word that looks or sounds the same in two related languages but means something different. Spanish embarazada is “pregnant”; Portuguese embaraçada is “embarrassed.” Spanish exquisito is “delicious”; Portuguese esquisito is “weird.” Spanish polvo is dust; Portuguese polvo is an octopus. Spanish apellido is a surname; Portuguese apelido is a nickname. Between Russian and Ukrainian: неделя is “week” but неділя is “Sunday”; Russian родина is “homeland” but Ukrainian родина is “family”; Russian час is “hour” but Ukrainian час is “time,” the Ukrainian hour being година. Imagine a draft of John 2:4, “my hour has not yet come,” carried from Russian into Ukrainian with час left in place, or “the first day of the week” (Matt 28:1) with the week turned into a Sunday.
Spelling convention
The agreed way a language writes its sounds, which can differ even where the sounds are the same. Spanish writes ñ and ll (señor, batalla) where Portuguese writes nh and lh (senhor, batalha); Portuguese marks nasal vowels with a tilde (nação against Spanish nación); only Spanish opens a question with ¿. Ukrainian uses і, ї, є, ґ and the apostrophe, and does not use Russian ы, э, ё or ъ. Worse, the same letter can stand for different sounds: Russian и is [i], Ukrainian и is closer to [ɪ]; Russian е softens the consonant before it, Ukrainian е does not (Ukrainian writes є for that sound). Proper names follow church convention too: Spanish Juan, Portuguese João; Russian Иисус, Ukrainian Ісус. A machine sees identical characters and may assume identical words.
III

The Cousin in the Machine

שנים מקרא ואחד תרגום

Machine learning, simply put. A traditional computer program follows rules a person wrote. A machine-learning model is given examples instead and adjusts millions of internal numbers until its guesses match the examples. The old custom of shnayim mikra v’echad targum, reading the weekly parashah twice in Hebrew and once in Targum Onkelos (b. Berakhot 8a), is a fair picture of it. A student who reads enough verses in Hebrew beside Aramaic begins to predict how Onkelos will render the next verse before he looks. He has not memorized a rule book; he has absorbed patterns from parallel text. Scripture Forge’s own help pages use the Rosetta Stone for the same idea: the system learns a language by seeing the same sentence written in a language it already understands and in the language being translated into.

NLLB. No Language Left Behind is a single neural translation model released by Meta AI in 2022 and described in Nature in 2024. It covers about 200 languages in one shared network. Because the languages share one set of internal numbers, what the model learns from a well-resourced language helps its relatives: the paper describes this as transfer learning across languages. NLLB also reads text in pieces smaller than words (“subwords”), so Spanish noche and Portuguese noite, or Russian хлеб and Ukrainian хліб, partly share the same pieces. That is the family resemblance, now inside the machine.

What Scripture Forge does with it. SIL’s drafting service starts from NLLB and fine-tunes it for one translation project. In the words of Understanding Draft Generation, the process has two steps: learn the language, then translate the text. For learning, it pairs the team’s reference text with the verses the team has already translated. For drafting, it translates a new book from the chosen source text. Preparing to Generate a Draft explains that the reference project need not be the text the team translates from; a second reference, such as the team’s own back translation, can be added; and extra paired data can be uploaded as a two-column spreadsheet. SIL recommends roughly a New Testament’s worth of completed verses; teams with fewer than 6,000 are directed to the Guidelines for New Testament Draft Generation. Vernacular drafting requires onboarding by SIL’s NLP team, which tests settings and recommends the best configuration for each project.

Why a resourced cousin helps. Of the 7,132 languages listed on this page, only 192 are among NLLB’s roughly 200 (marked ◆ in the lists below). The other 6,940 are unknown to the base model. When a low-resourced language has a close relative among the 200, the model starts its fine-tuning already knowing much of the family’s grammar, word order, spelling habits and vocabulary. It has, so to speak, already sat at the family table. When the nearest resourced relative is distant, or when the whole family has none (Gallery IV shows that this is true of several of the largest families on earth, including Nuclear Trans New Guinea with more than 300 languages), the model must learn nearly everything from the team’s own verses.

Why the result is only a hypothesis. Everything that makes cousins dangerous for people makes them dangerous for the machine. A model that has absorbed Russian will be tempted to leave час where Ukrainian needs година; a model shaped by Spanish may reach for embarazada in a text where the Portuguese cousin means “embarrassed.” It will copy the cousin’s idioms and the cousin’s spelling of names. So the draft is exactly what the SIL help page says it is: a rough first draft that always contains errors, to be judged by whether it helps the team, not by whether it could stand alone. Only a native speaker can say whether a sentence sounds like home or like a relative visiting from abroad.

From resourced cousin to native-speaker-reviewed draft

  1. Find the family. Look up the target language’s lineage in Glottolog and identify its nearest relatives that have a published Bible, an NLLB model, or both. (Click any language in Gallery IV to see its lineage and nearest NLLB cousins.)
  2. Translate a first body of Scripture by hand. The model can only learn the target language from verses native speakers have already translated. A related-language Bible can speed this stage; SIL’s Adapt It was built precisely for adapting a text into a related language.
  3. Sign up and configure. Connect the Paratext project to Scripture Forge, sign up for drafting, and let the onboarding team choose the reference and source texts that give the best results.
  4. Generate the draft. The model learns from the parallel verses, then translates the next book as a working hypothesis.
  5. Native-speaker revision. Read it aloud. Hunt the cousin’s false friends and idioms. Check every name and key term against the project’s spelling conventions and Paratext’s Biblical Terms list.
  6. Back-translate and check. Scripture Forge can draft a back translation into a language the consultant reads, which the team then corrects; the consultant checks meaning against the Hebrew and Greek.
  7. Community checking. Ordinary speakers answer questions about the passage in Scripture Forge’s community checking tool, revealing what is actually understood.

Each corrected book becomes training data, so the next draft is better than the last.

Three SIL tools for related languages

Adapt It

A free, open-source editor for translating between related languages. It builds a Knowledge Base of the words and phrases the translator has already adapted and inserts them automatically when they recur. It does no linguistic analysis; the bilingual human does the thinking.

adapt-it.org

FLExTrans

A rule-based machine translation tool built on FieldWorks Language Explorer, designed to help translators move between related languages using the lexicons and grammars the team has built.

ai.sil.org/projects

Scripture Forge drafting

A fine-tuned NLLB neural model that learns from the project’s own parallel verses and drafts whole books, integrated with Paratext.

help.scriptureforge.org

The three differ in method (memory, rules, statistical learning), but they end in the same place: a draft that native speakers must own.

IV

The Mishpachot of the Earth: 7,132 Tongues by Family

כל משפחת האדמה

“In you all the families of the earth shall be blessed” (Gen 12:3). Below, every language on this page is placed in its family according to Glottolog. 6,940 of the 7,132 could be joined to a Glottolog lineage through their ISO code or name, falling into 220 families, 126 isolates, 145 sign languages and 28 pidgins, creoles and other contact languages. The remaining 192 are, like the families of Ezra 2:62, waiting for their register to be found.

The chart shows the twenty-five largest families. The length of each bar is the number of languages on this page; the red part is the share still without published Scripture; the figure on the right is how many of the family’s languages NLLB already knows. A family with a red zero has no resourced cousin inside it at all.

Atlantic-Congo33 ◆
Austronesian18 ◆
Indo-European76 ◆
Sino-Tibetan8 ◆
Afro-Asiatic18 ◆
Nuclear Trans New Guinea0 ◆
Otomanguean0 ◆
Austroasiatic3 ◆
Pama-Nyungan0 ◆
Tai-Kadai3 ◆
Dravidian4 ◆
Mande2 ◆
Tupian1 ◆
Central Sudanic0 ◆
Uto-Aztecan0 ◆
Nuclear Torricelli0 ◆
Nilotic3 ◆
Arawakan0 ◆
Quechuan1 ◆
Algic0 ◆
Athabaskan-Eyak-Tlingit0 ◆
Hmong-Mien0 ◆
Turkic11 ◆
Kru0 ◆
Uralic3 ◆
Scripture publishedUnderserved◆ = languages NLLB already knows
A Single Claim, Worked All the Way Through

★ LOOK AT LANGUAGE NUMBER #13 ★

This list can read like an endless roll call of unfamiliar syllables — 7,132 names with nothing to hold onto. So before you scroll it, meet just one of them properly: Abenaki, Eastern, entry #13, still marked underserved below. Two companion pages turn this one name into an actual people, an actual grammar, and an actual verse of Scripture worked line by line.

Jump to Abenaki, Eastern in the list below ↓
For Bible translators and volunteers

Dear Fellow Translator

Suppose you are interested in finding out about a certain language with a view to possibly volunteering to help its translation team, or even starting your own translation team. Let’s use Rabha, #5369 below, as an example. Here’s what you could do to see what work, if any, has been done for that language.

Every link below is set for Rabha (rah). Type another three-letter code and press the button to point them at your language.
  1. Find the language’s three-letter code

    Nearly every website below uses the ISO 639-3 code rather than the language name. On this page, the code appears in small type after most language names in the list below, and you can type a code into the “Find a language” box to jump to it. You can also look the language up on Wikipedia, where the code appears in the box on the right.

    While you’re there, read about dialects. A Bible in one dialect may not serve speakers of another.

    For Rabha: the code is rah. Wikipedia lists three dialects, Rongdani, Maitori and Kocha, and notes that Kocha is not understood by speakers of the other two.

  2. Check what Scripture exists

    Two directories are organized by language code:

    • ScriptureEarth: Scripture text, audio and video to read online or download.
    • Joshua Project: the people groups who speak the language, its alternate names, and which Bible resources are available.
  3. Search the web

    Type [language name] language Bible translation New Testament into any search engine. Directories can lag behind events, so don’t skip this step.

    For Rabha, this search finds a scanned copy of the 1999 Rabha New Testament (Don Bosco Publications, Guwahati) and a news report that the complete Rabha Bible was released in Goalpara, Assam, in November 2025, naming the whole translation team.

  4. Look for audio and oral resources

    Many language communities are more oral than literate. Global Recordings Network lists its recordings by language code.

    For Rabha: audio-visual Bible lessons and short audio Bible stories, free to download.

  5. See who is translating openly online

    unfoldingWord’s Door43 hosts open-licensed translation work:

    For Rabha, Door43 is empty. That does not mean Rabha has no Bible, as step 3 showed. It means the Rabha Bible is not yet openly licensed and available in the digital tools translators share.

  6. Find the people

    Scripture is translated by communities, not websites. Note the translators and publishers named in steps 2 to 5, and search for local churches and associations.

    For Rabha: rabha.org, built by the Rabha Baptist Convention in English and Rabha.

  7. Put it together

    For Rabha: a complete Bible (2025) in the Rongdani dialect, audio evangelism resources, and no open digital text on Door43. Good questions for a volunteer: are speakers of the Kocha dialect served? Could the Bible become available digitally or in audio? And what does the local church itself say it needs?

Jump to Rabha in the list below ↓

The Roll Call: 7,132 Languages

Underserved — no published Scripture yet Scripture published (portions, NT, or complete Bible) Already known to the NLLB model

7,132 languages listed · 3,066 (43%) still await a first published word of Scripture · 192 already known to NLLB