Artists for Israel International · Bible Technology Exhibit

From Midway to the Mother Tongue

How the codebreakers of Station HYPO read Japanese naval traffic from fragments, and how today's Bible translators use the same habits of mind, plus real software, to build a dictionary and a first Scripture portion in a language that has never been written down.

Pearl Harbor, May 1942

A Japanese intercept, partly read. Blanks are code groups not yet recovered.

striking force·····attackAF··········early June?carriers

A translation desk, today

Mark 4:39 in a new language, partly drafted. Blanks are words not yet established.

Yeshuastoodrebukedwind·····“be still”·····calm

The story, brieflyJoseph Rochefort, Station HYPO, and the two letters “AF”

In the spring of 1942 Commander Joseph J. Rochefort ran the U.S. Navy's codebreaking unit at Pearl Harbor, Station HYPO, from a windowless basement. His team rarely read a Japanese naval message whole. They worked from fragments: the code groups they had already recovered, the patterns of who was signalling whom, and a deep knowledge of how Japanese officers wrote.

Rochefort was a linguist before he was a cryptanalyst. The Navy had sent him to Tokyo for three years of Japanese study, and several of his key men were Japanese-language officers too. That matters for our story: the breakthrough at Midway came from people who understood a language, not from mathematics alone.

By May the traffic pointed to a major operation against a target the Japanese called AF. HYPO believed AF was Midway Atoll; Washington was not convinced. So HYPO set a test. Midway was told to report, in plain language, that its fresh-water plant had broken down. Within days a Japanese message reported that AF was short of water. One designator was now confirmed, and with it the target. Admiral Nimitz, combining this with other intelligence, placed his carriers where the Japanese fleet would arrive, and in early June the Navy sank four Japanese carriers.

At Bletchley Park in England, Alan Turing was attacking German naval Enigma with a related idea: no single clue settles anything, but many weak clues, each weighed and added up, can make a hypothesis effectively certain.

What we keep from this story. Three habits, nothing more: work from fragments toward a convergent picture; let many small pieces of evidence add up rather than trusting any single one; and when a hypothesis matters, design a simple test that a native source will answer without knowing it is being tested. The rest of this exhibit is about how field linguists and Bible translators do exactly that.

A note on what we leave out. JN-25's “additive” layer (random numbers added to every code group to disguise it) was the mathematicians' problem and has no counterpart in a human language. A language is not a disguise laid over meaning. To its speakers it is perfectly clear. It is opaque only to the outsider, and the outsider's task is to learn it, not to strip anything away.

The linguistic parallelStation HYPO's methods and the field linguist's

Modern field linguists working with low-resource languages (languages with no written grammar, no dictionary, and sometimes only a few hundred speakers) use a method remarkably like HYPO's. They gather fragments, recover the vocabulary one item at a time, keep a running record of what is confirmed and what is only suspected, and test their guesses against native speakers.

Station HYPO, 1942Field team on a low-resource language, todayThe tool that does it now
Intercepted radio traffic, collected day after dayRecorded speech: stories, conversations, procedures, prayers, elicited word listsELAN, SayMore, a phone recorder
Traffic analysis: who sends to whom, when, how oftenContext: who is speaking, to whom, about what, in what settingELAN tiers and session metadata
The codebook, rebuilt group by group on index cardsThe lexicon, built entry by entry with senses, examples, and grammar notesFLEx Lexicon
Recurring code groups recognised across many messagesRecurring words and morphemes recognised across many textsFLEx interlinear texts and concordance
Place designators such as “AF” as entry pointsProper names, loanwords, and basic vocabulary as entry pointsSwadesh-type lists; Paratext Biblical Terms
The water ruse: a planted test answered by the enemyComprehension checks and back-translation answered by speakers who never saw the sourceScripture Forge community checking; Paratext back-translation projects
IBM tabulating machines sorting thousands of groupsMachine learning ranking likely translations from verses already doneServal inside Scripture Forge
HYPO's reading cross-checked against Washington'sThe draft cross-checked by a translation consultantParatext notes and consultant checks
Where the analogy is better for us than it was for Rochefort. HYPO's codebook was locked in an enemy safe. A language's “codebook” lives in the head of every native speaker, and they are on our side. The water ruse took HYPO weeks to arrange; a translation team can run one every afternoon simply by asking, “What does this sentence tell you?”

The field toolkitFrom the first word list to an interlinear text

The Swadesh list: the first crib

In the 1950s the American linguist Morris Swadesh compiled lists of basic meanings that nearly every language has a word for and rarely borrows: I, you, we, one, two, water, fire, sun, eye, hand, blood, die, eat, drink, sleep, big, small, black, white. His best-known versions run to 100 and 207 items. A field linguist on day one sits with two or three speakers, elicits these words, records them, and writes them down phonetically.

The list does three jobs at once. It gives the first few hundred dictionary entries. It reveals the sound system, because the same consonants and vowels recur across many simple words. And it lets the linguist compare the language with its neighbours: a high share of shared basic vocabulary suggests a close relative whose existing Bible translation can help later.

Swadesh 100 / 207

Basic, rarely borrowed meanings. The classic first session.

Leipzig–Jakarta list

A 100-item list (Haspelmath and Tadmor, 2009) chosen from 41 languages for resistance to borrowing.

SIL Comparative African Wordlist

About 1,700 items for deeper comparative work, widely used across Africa.

After the basic list comes Rapid Word Collection, an SIL method in which groups of speakers brainstorm words by semantic domain (“parts of a canoe,” “kinds of rain,” “ways of speaking”). Working through SIL's list of roughly 1,800 domains, a community workshop can gather thousands of words in a couple of weeks, far faster than one linguist with one speaker.

ELAN: turning recordings into data

Most real language is spoken, not listed. ELAN, free software from The Language Archive at the Max Planck Institute for Psycholinguistics in Nijmegen, lets a team load audio or video and mark each stretch of speech on a timeline. Each stretch gets layers, called tiers: a transcription, a free translation, the speaker's name, and any notes. Click a line and you hear exactly that line again.

This is the equivalent of HYPO's filed intercepts. Nothing is lost, every claim can be traced back to the recording, and a speaker can be played a sentence months later and asked about it. Many teams start in SayMore (SIL), which manages recording sessions, speaker consent forms, and “careful speech” re-recordings, before moving files into ELAN. Phoneticians use Praat alongside both for tone and vowel measurements.

FLEx: where the codebook is rebuilt

FieldWorks Language Explorer (FLEx), from SIL, is the standard program for turning raw texts into a dictionary and a grammar. It has three areas that work together.

Texts & Words

Paste or import a transcribed text. FLEx splits it into words; you break each word into morphemes and gloss each one. This is interlinearizing. Every analysis you approve is remembered, and FLEx proposes it automatically the next time that word turns up in any text.

Lexicon

Every morpheme you gloss becomes a dictionary entry with senses, part of speech, semantic domain, and example sentences pulled from your own texts. Export as LIFT, publish online through Webonary, or build a phone dictionary with Dictionary App Builder.

Grammar

Record affixes, word classes, and phonological rules. FLEx's morphological parsers (Hermit Crab and XAmple) then try to analyse words you have not seen yet, and you correct them. Each correction teaches the parser.

Why FLEx is the closest thing to HYPO's card files. When the codebreakers recovered a code group in one message, they could suddenly read it in every other message where it appeared. FLEx does the same: approve a gloss for one morpheme and it lights up across the whole corpus, exposing the next unknown word in a now-readable context.

How the pieces connect

RecordSpeakers, stories, word lists. SayMore manages sessions and consent.
TranscribeELAN: time-aligned transcription and free translation.
AnalyseExport from ELAN as a FLEx file (.flextext); interlinearize and build the lexicon in FLEx.
DraftParatext: Scripture drafting, Biblical Terms, interlinearizer using FLEx glosses.
AccelerateScripture Forge with Serval: AI drafts trained on the verses already done.
CheckBack-translation, community checking, consultant review, back into Paratext.

ELAN reads and writes the .flextext format through File > Import > FLEx File and File > Export As > FLEx File; FLEx imports it through File > Import > FLExText Interlinear. Time alignment is kept at the phrase level in a round trip.

Building a lexicon the Rochefort wayFragments, weighted evidence, planted tests

  1. Get the basic list. Elicit a Swadesh or Leipzig–Jakarta list from more than one speaker, record it, and enter it in FLEx. Differences between speakers are data, not errors; they may reveal dialects.
  2. Record natural texts early. Word lists show words in isolation; stories show grammar in use. Ten short recorded narratives, transcribed in ELAN, will teach more about verb endings than a hundred elicited sentences.
  3. Interlinearize and let the evidence accumulate. In FLEx, each word occurrence you analyse is a small piece of evidence for a gloss. One occurrence is a guess; twenty consistent occurrences in different contexts is close to certainty. This is Turing's weighing of clues, done with a concordance instead of a Bombe.
  4. Widen by semantic domain. Run a Rapid Word Collection workshop, or work through domains yourself, to fill the gaps the texts did not cover: kinship, farming, sickness, worship, emotion.
  5. Fix the anchors. In Paratext's Biblical Terms tool, decide how names (Avraham, Moshe, Yerushalayim, Yeshua) and core terms (covenant, sin, holy, Messiah) will be rendered, and record the decision. Like “AF,” these are fixed points that every later sentence can be read against.
  6. Run the water ruse. Never trust a gloss just because it fits. Use a word in a fresh sentence and ask a speaker who has not seen your notes what it means. If the answer comes back the way you predicted, the entry is confirmed.

Piecing together a translation portionReading one verse the way HYPO read one intercept

HYPO never needed every code group to know that AF was the target. A translation team likewise drafts a verse while some of its words are still uncertain, then marks and tests the uncertain ones. Select any part of the verse below to see where that piece came from and how it is being confirmed.

Mark 4:39 · draft in progress

Select a word above.

Once a few books exist in this condition, drafted, checked, and corrected, they become training data. That is where machine learning enters.

Serval and Scripture ForgeThe machine that learns from the verses already done

Scripture Forge is SIL's free, web-based workspace for Bible translation teams. Underneath it runs Serval, an open-source AI platform built by SIL's AI & NLP team, which does the machine learning for draft generation and translation suggestions.

What Serval is

An open-source web service (a REST API) that other programs call. It accepts Paratext projects directly, understands USFM Scripture markup including versification and verse ranges, and returns drafts as USFM or JSON. Scripture Forge is its main user, but because Serval is platform-independent, other tools can use it too.

How a draft is made

Two steps: first learn the language, then translate. Serval trains on pairs of verses: a source or reference text and the team's own finished verses in the target language (a project with, say, several thousand completed verses). It then drafts books the team has not yet done. Draft quality depends almost entirely on how well the learning step goes, which means on how much good, checked text already exists.

The models inside

Neural machine translation based on Meta's NLLB-200 (“No Language Left Behind”) model, fine-tuned for each new minority language, produces full drafts. NLLB-200 also powers back-translation drafts into more than 200 languages. An older statistical model learned on the fly as a translator typed and offered word-by-word suggestions; SIL has announced it is consolidating onto the neural drafting.

What it has achieved

Hundreds of projects now use it. SIL mined actual Paratext project histories and found that teams using Scripture Forge drafts reached the “drafted” stage roughly twice as fast, and drafted and reviewed about 1,000 more verses a year on average than comparable projects.

The Rochefort lesson, stated for Serval. The machine is HYPO's tabulating room, not Rochefort. It ranks likely renderings from patterns in verses already confirmed by people. It cannot know whether “rebuked” sounds angry to a grandmother in the village. Every draft is a set of hypotheses for the team to test, and the speakers remain the final authority.

For oral languages

Many low-resource communities are oral first. Faith Comes By Hearing and SIL are piloting a chain in which Scripture is first recorded aloud (in tools such as Render or Audio Project Manager), transcribed by speech recognition, drafted through Serval in Scripture Forge, and turned back into audio with synthetic speech for checking by listeners.

Who can use what

Back-translation drafts (from the vernacular into a major language) are open to any Paratext user today. Drafts into a vernacular need setup by SIL's NLP team: connect the project in Scripture Forge, open Generate draft, and choose Sign up for drafting.

How Serval reaches desktop ParatextParatext 9.5 today, Paratext 10 Studio next

There is no Serval button inside Paratext 9. The bridge is Scripture Forge, which shares the same project through Paratext's own Send/Receive servers.

  1. One login. You sign in to Scripture Forge with your Paratext registration. Your Paratext projects, and your role in each, appear automatically.
  2. Connect the project. Scripture Forge syncs with the Paratext project in both directions. Notes written in Scripture Forge appear as Paratext notes, and edits made in Paratext appear in Scripture Forge after the next Send/Receive.
  3. Train and draft in the cloud. You choose which finished books Serval should learn from and which books to draft. Training runs on SIL's servers, not your laptop, and typically takes an hour or more.
  4. Bring the draft home. Preview the draft in Scripture Forge and add chapters to the project, or export and import them. After Send/Receive the draft text is in Paratext 9.5, where all the normal checks, Biblical Terms, interlinearizer, and consultant notes apply.

FLEx inside Paratext 9.5

In Paratext, Project Properties > Associations > Associated Lexical Project links a Paratext project to a FLEx project. Since Paratext 9.4 the Paratext Interlinearizer can take its glosses from that FLEx lexicon, and FLEx can show the Paratext books in Texts & Words for full morpheme analysis. The vernacular writing-system code must be identical in both programs or the link silently fails.

Paratext 10 Studio

The successor to Paratext 9 is built on Platform.Bible, an open framework for extensions. A Scripture Forge extension already lets a translator sign in and open AI drafts inside the Paratext 10 editor, and an Interlinearizer extension is being built to work with FLEx lexicons hosted on Lexbox. Both are early releases; Paratext 9.5 plus Scripture Forge remains the production route.

Try it now with the Yiddish Triglot

Because back-translation drafting is open to all Paratext users, AFII can test the method on its own work today. Yiddish (ydd, Eastern Yiddish) is among the languages Scripture Forge can draft into, alongside English.

  • Keep the Yiddish text as the Paratext project and an English back-translation project beside it, with a few books already back-translated by hand.
  • In Scripture Forge, connect the back-translation project, set the Yiddish as its source, and choose Generate draft, selecting the books to draft and the finished books to learn from.
  • Read the machine's English against the Triglot's English gloss. Wherever they disagree, add a Paratext note for a human reviewer. Disagreement is the signal, as “AF is short of water” was.

Who's who in low-resource translationThe agencies and what each one brings

OrganizationRoleTools and data you will meet
SILField linguistics, language survey, literacy, and most of the language software in this exhibitFLEx, SayMore, Keyman, Webonary, Dictionary App Builder, Scripture Forge, Serval, PTXPrint; co-developer of Paratext; Ethnologue and its language codes originated here
United Bible SocietiesNational Bible Societies translating, publishing, and distributing; co-developer of ParatextParatext, Digital Bible Library
Wycliffe Global AllianceAbout a hundred member organizations worldwide, including Wycliffe USA and Wycliffe UKAnnual Scripture-access statistics from ProgressBible
Every Tribe Every Nation (ETEN) and the illumiNations allianceCoordination and funding toward the 2033 Scripture-access goalsETEN Innovation Lab reports on AI drafting and new tools
Seed Company, Biblica, Lutheran Bible Translators, Pioneer Bible TranslatorsTranslation partners running projects with local churchesParatext and Scripture Forge projects
Wycliffe AssociatesChurch-led rapid drafting in workshopsMAST method, BTT Writer
Faith Comes By HearingAudio Scripture and oral-first translationBible Brain (text and audio by language)
Deaf Bible Society and sign-language partnersScripture for the world's sign languagesVideo translation workflows
Clear BibleOriginal-language data for alignment and AIMacula Hebrew and Greek datasets
Max Planck Institute for PsycholinguisticsAcademic language documentation and archivingELAN, The Language Archive
Artists for Israel InternationalJewish-vernacular Scripture: OJB, Orthodox Yiddish Triglot, interlinearsParatext; published parallel texts usable as known meaning for related work

Low-resource languages: how many, and where to find what has been doneThe resource ladder and the lookup method

“Low-resource” is a ladder, not a single category. Which rung a language stands on decides which tools can help.

~200languages
High and mid resource. Covered by large multilingual models such as NLLB-200.
Back-translation drafts into these work today in Scripture Forge.
833languages, 2023
Some Scripture in open digital form. The eBible corpus gathered 1,009 translations in 833 languages from 75 families for research.
Enough verses to fine-tune Serval for drafting the rest of the Bible.
4,414languages, Sept. 2026
Translation work in progress. Many have only a few books, or are still at the word-list stage.
FLEx and ELAN first; Serval once a few books are checked.
1,688languages, Sept. 2026
No Scripture and no work under way. Of these, 543 languages with about 38 million speakers are considered vital enough to need translation.
The Swadesh list and a recorder. Rochefort's starting position.

Out of 7,394 known languages, Wycliffe UK reported in September 2026 that 3,224 still have no Scripture. Counts change monthly; ProgressBible is the underlying source.

Where to look up any language

SourceWhat it tells youAccess
EthnologueThe ISO 639-3 code, speakers, dialects, vitality, and related languages. Start here to get the three-letter code.Basic pages free; full data by subscription
GlottologA bibliography of every known grammar, dictionary, word list, and text for each language, plus its family tree. The best single answer to “what has been done?”Free
OLACA union catalogue of archived recordings, texts, and lexicons held in many archives, searchable by language code.Free
Language archives: ELAR, PARADISEC, The Language Archive, SIL Language & Culture ArchivesThe actual recordings, transcriptions, and field notes from earlier documentation projects.Mostly free; some items restricted by the community
WebonaryPublished FLEx dictionaries online, by language.Free
ScriptureEarthScripture text, audio, and video available to download in each language.Free
Bible Brain and eBible.orgWhich Bible texts and recordings exist digitally, many under open licences.Free; API keys for developers
Joshua ProjectPeople groups using each language, with a Bible-status summary.Free
ProgressBibleThe master record of translation status, project by project, for every language.Restricted to staff of Bible agencies; ask a partner
Wycliffe Global Alliance statisticsThe public annual summary drawn from ProgressBible.Free

The lookup method, step by step

  1. Get the code. Search the language name in Ethnologue or Glottolog and note its ISO 639-3 code (Eastern Yiddish is ydd). Every other source uses it.
  2. List what exists. Open the language in Glottolog and read the references: grammars, dictionaries, word lists. Then search the code in OLAC, for example language-archives.org/language/ydd, for archived recordings and lexicons.
  3. Check the neighbours. Glottolog's family tree and Ethnologue's related-language notes show close relatives. A relative with a finished Bible is your isolog: the same message already read in a neighbouring system.
  4. Check Scripture status. Search ScriptureEarth, Bible Brain, and Joshua Project for the language. For the full project history, ask a contact in SIL, Wycliffe, or a Bible Society to consult ProgressBible.
  5. Ask who is working on it. Archive deposits and ProgressBible name the linguists and agencies involved. Contacting them first avoids duplicate work and often unlocks an existing FLEx project on Lexbox.

What Rochefort and Turing teach a translation teamFive working rules

Start with what you can be sure of.Names and basic vocabulary first, as HYPO started with designators and Ventris with place names.
Keep a confirmed / suspected record.FLEx glosses, Paratext notes, and Biblical Terms renderings should show which items are tested and which are guesses.
Let evidence add up.No single occurrence proves a meaning; many consistent ones do. That is Turing's rule, and it is how both FLEx and Serval learn.
Plant the test.Back-translation and comprehension checks are the water ruse. Ask people who have not seen the source what the text says.
Use the related language.A finished translation in a neighbouring language is a partly read message. Work in clusters of related languages, not one by one.
Remember who holds the codebook.The speakers do. Software ranks hypotheses; the community decides what its language actually says.

Sources and further reading