The story, brieflyJoseph Rochefort, Station HYPO, and the two letters “AF”
In the spring of 1942 Commander Joseph J. Rochefort ran the U.S. Navy's codebreaking unit at Pearl Harbor, Station HYPO, from a windowless basement. His team rarely read a Japanese naval message whole. They worked from fragments: the code groups they had already recovered, the patterns of who was signalling whom, and a deep knowledge of how Japanese officers wrote.
Rochefort was a linguist before he was a cryptanalyst. The Navy had sent him to Tokyo for three years of Japanese study, and several of his key men were Japanese-language officers too. That matters for our story: the breakthrough at Midway came from people who understood a language, not from mathematics alone.
By May the traffic pointed to a major operation against a target the Japanese called AF. HYPO believed AF was Midway Atoll; Washington was not convinced. So HYPO set a test. Midway was told to report, in plain language, that its fresh-water plant had broken down. Within days a Japanese message reported that AF was short of water. One designator was now confirmed, and with it the target. Admiral Nimitz, combining this with other intelligence, placed his carriers where the Japanese fleet would arrive, and in early June the Navy sank four Japanese carriers.
At Bletchley Park in England, Alan Turing was attacking German naval Enigma with a related idea: no single clue settles anything, but many weak clues, each weighed and added up, can make a hypothesis effectively certain.
A note on what we leave out. JN-25's “additive” layer (random numbers added to every code group to disguise it) was the mathematicians' problem and has no counterpart in a human language. A language is not a disguise laid over meaning. To its speakers it is perfectly clear. It is opaque only to the outsider, and the outsider's task is to learn it, not to strip anything away.
The linguistic parallelStation HYPO's methods and the field linguist's
Modern field linguists working with low-resource languages (languages with no written grammar, no dictionary, and sometimes only a few hundred speakers) use a method remarkably like HYPO's. They gather fragments, recover the vocabulary one item at a time, keep a running record of what is confirmed and what is only suspected, and test their guesses against native speakers.
| Station HYPO, 1942 | Field team on a low-resource language, today | The tool that does it now |
|---|---|---|
| Intercepted radio traffic, collected day after day | Recorded speech: stories, conversations, procedures, prayers, elicited word lists | ELAN, SayMore, a phone recorder |
| Traffic analysis: who sends to whom, when, how often | Context: who is speaking, to whom, about what, in what setting | ELAN tiers and session metadata |
| The codebook, rebuilt group by group on index cards | The lexicon, built entry by entry with senses, examples, and grammar notes | FLEx Lexicon |
| Recurring code groups recognised across many messages | Recurring words and morphemes recognised across many texts | FLEx interlinear texts and concordance |
| Place designators such as “AF” as entry points | Proper names, loanwords, and basic vocabulary as entry points | Swadesh-type lists; Paratext Biblical Terms |
| The water ruse: a planted test answered by the enemy | Comprehension checks and back-translation answered by speakers who never saw the source | Scripture Forge community checking; Paratext back-translation projects |
| IBM tabulating machines sorting thousands of groups | Machine learning ranking likely translations from verses already done | Serval inside Scripture Forge |
| HYPO's reading cross-checked against Washington's | The draft cross-checked by a translation consultant | Paratext notes and consultant checks |
The field toolkitFrom the first word list to an interlinear text
The Swadesh list: the first crib
In the 1950s the American linguist Morris Swadesh compiled lists of basic meanings that nearly every language has a word for and rarely borrows: I, you, we, one, two, water, fire, sun, eye, hand, blood, die, eat, drink, sleep, big, small, black, white. His best-known versions run to 100 and 207 items. A field linguist on day one sits with two or three speakers, elicits these words, records them, and writes them down phonetically.
The list does three jobs at once. It gives the first few hundred dictionary entries. It reveals the sound system, because the same consonants and vowels recur across many simple words. And it lets the linguist compare the language with its neighbours: a high share of shared basic vocabulary suggests a close relative whose existing Bible translation can help later.
Swadesh 100 / 207
Basic, rarely borrowed meanings. The classic first session.
Leipzig–Jakarta list
A 100-item list (Haspelmath and Tadmor, 2009) chosen from 41 languages for resistance to borrowing.
SIL Comparative African Wordlist
About 1,700 items for deeper comparative work, widely used across Africa.
After the basic list comes Rapid Word Collection, an SIL method in which groups of speakers brainstorm words by semantic domain (“parts of a canoe,” “kinds of rain,” “ways of speaking”). Working through SIL's list of roughly 1,800 domains, a community workshop can gather thousands of words in a couple of weeks, far faster than one linguist with one speaker.
ELAN: turning recordings into data
Most real language is spoken, not listed. ELAN, free software from The Language Archive at the Max Planck Institute for Psycholinguistics in Nijmegen, lets a team load audio or video and mark each stretch of speech on a timeline. Each stretch gets layers, called tiers: a transcription, a free translation, the speaker's name, and any notes. Click a line and you hear exactly that line again.
This is the equivalent of HYPO's filed intercepts. Nothing is lost, every claim can be traced back to the recording, and a speaker can be played a sentence months later and asked about it. Many teams start in SayMore (SIL), which manages recording sessions, speaker consent forms, and “careful speech” re-recordings, before moving files into ELAN. Phoneticians use Praat alongside both for tone and vowel measurements.
FLEx: where the codebook is rebuilt
FieldWorks Language Explorer (FLEx), from SIL, is the standard program for turning raw texts into a dictionary and a grammar. It has three areas that work together.
Texts & Words
Paste or import a transcribed text. FLEx splits it into words; you break each word into morphemes and gloss each one. This is interlinearizing. Every analysis you approve is remembered, and FLEx proposes it automatically the next time that word turns up in any text.
Lexicon
Every morpheme you gloss becomes a dictionary entry with senses, part of speech, semantic domain, and example sentences pulled from your own texts. Export as LIFT, publish online through Webonary, or build a phone dictionary with Dictionary App Builder.
Grammar
Record affixes, word classes, and phonological rules. FLEx's morphological parsers (Hermit Crab and XAmple) then try to analyse words you have not seen yet, and you correct them. Each correction teaches the parser.
How the pieces connect
ELAN reads and writes the .flextext format through File > Import > FLEx File and File > Export As > FLEx File; FLEx imports it through File > Import > FLExText Interlinear. Time alignment is kept at the phrase level in a round trip.
Building a lexicon the Rochefort wayFragments, weighted evidence, planted tests
- Get the basic list. Elicit a Swadesh or Leipzig–Jakarta list from more than one speaker, record it, and enter it in FLEx. Differences between speakers are data, not errors; they may reveal dialects.
- Record natural texts early. Word lists show words in isolation; stories show grammar in use. Ten short recorded narratives, transcribed in ELAN, will teach more about verb endings than a hundred elicited sentences.
- Interlinearize and let the evidence accumulate. In FLEx, each word occurrence you analyse is a small piece of evidence for a gloss. One occurrence is a guess; twenty consistent occurrences in different contexts is close to certainty. This is Turing's weighing of clues, done with a concordance instead of a Bombe.
- Widen by semantic domain. Run a Rapid Word Collection workshop, or work through domains yourself, to fill the gaps the texts did not cover: kinship, farming, sickness, worship, emotion.
- Fix the anchors. In Paratext's Biblical Terms tool, decide how names (Avraham, Moshe, Yerushalayim, Yeshua) and core terms (covenant, sin, holy, Messiah) will be rendered, and record the decision. Like “AF,” these are fixed points that every later sentence can be read against.
- Run the water ruse. Never trust a gloss just because it fits. Use a word in a fresh sentence and ask a speaker who has not seen your notes what it means. If the answer comes back the way you predicted, the entry is confirmed.
Piecing together a translation portionReading one verse the way HYPO read one intercept
HYPO never needed every code group to know that AF was the target. A translation team likewise drafts a verse while some of its words are still uncertain, then marks and tests the uncertain ones. Select any part of the verse below to see where that piece came from and how it is being confirmed.
Mark 4:39 · draft in progress
Once a few books exist in this condition, drafted, checked, and corrected, they become training data. That is where machine learning enters.
Serval and Scripture ForgeThe machine that learns from the verses already done
Scripture Forge is SIL's free, web-based workspace for Bible translation teams. Underneath it runs Serval, an open-source AI platform built by SIL's AI & NLP team, which does the machine learning for draft generation and translation suggestions.
What Serval is
An open-source web service (a REST API) that other programs call. It accepts Paratext projects directly, understands USFM Scripture markup including versification and verse ranges, and returns drafts as USFM or JSON. Scripture Forge is its main user, but because Serval is platform-independent, other tools can use it too.
How a draft is made
Two steps: first learn the language, then translate. Serval trains on pairs of verses: a source or reference text and the team's own finished verses in the target language (a project with, say, several thousand completed verses). It then drafts books the team has not yet done. Draft quality depends almost entirely on how well the learning step goes, which means on how much good, checked text already exists.
The models inside
Neural machine translation based on Meta's NLLB-200 (“No Language Left Behind”) model, fine-tuned for each new minority language, produces full drafts. NLLB-200 also powers back-translation drafts into more than 200 languages. An older statistical model learned on the fly as a translator typed and offered word-by-word suggestions; SIL has announced it is consolidating onto the neural drafting.
What it has achieved
Hundreds of projects now use it. SIL mined actual Paratext project histories and found that teams using Scripture Forge drafts reached the “drafted” stage roughly twice as fast, and drafted and reviewed about 1,000 more verses a year on average than comparable projects.
For oral languages
Many low-resource communities are oral first. Faith Comes By Hearing and SIL are piloting a chain in which Scripture is first recorded aloud (in tools such as Render or Audio Project Manager), transcribed by speech recognition, drafted through Serval in Scripture Forge, and turned back into audio with synthetic speech for checking by listeners.
Who can use what
Back-translation drafts (from the vernacular into a major language) are open to any Paratext user today. Drafts into a vernacular need setup by SIL's NLP team: connect the project in Scripture Forge, open Generate draft, and choose Sign up for drafting.
How Serval reaches desktop ParatextParatext 9.5 today, Paratext 10 Studio next
There is no Serval button inside Paratext 9. The bridge is Scripture Forge, which shares the same project through Paratext's own Send/Receive servers.
- One login. You sign in to Scripture Forge with your Paratext registration. Your Paratext projects, and your role in each, appear automatically.
- Connect the project. Scripture Forge syncs with the Paratext project in both directions. Notes written in Scripture Forge appear as Paratext notes, and edits made in Paratext appear in Scripture Forge after the next Send/Receive.
- Train and draft in the cloud. You choose which finished books Serval should learn from and which books to draft. Training runs on SIL's servers, not your laptop, and typically takes an hour or more.
- Bring the draft home. Preview the draft in Scripture Forge and add chapters to the project, or export and import them. After Send/Receive the draft text is in Paratext 9.5, where all the normal checks, Biblical Terms, interlinearizer, and consultant notes apply.
FLEx inside Paratext 9.5
In Paratext, Project Properties > Associations > Associated Lexical Project links a Paratext project to a FLEx project. Since Paratext 9.4 the Paratext Interlinearizer can take its glosses from that FLEx lexicon, and FLEx can show the Paratext books in Texts & Words for full morpheme analysis. The vernacular writing-system code must be identical in both programs or the link silently fails.
Paratext 10 Studio
The successor to Paratext 9 is built on Platform.Bible, an open framework for extensions. A Scripture Forge extension already lets a translator sign in and open AI drafts inside the Paratext 10 editor, and an Interlinearizer extension is being built to work with FLEx lexicons hosted on Lexbox. Both are early releases; Paratext 9.5 plus Scripture Forge remains the production route.
Try it now with the Yiddish Triglot
Because back-translation drafting is open to all Paratext users, AFII can test the method on its own work today. Yiddish (ydd, Eastern Yiddish) is among the languages Scripture Forge can draft into, alongside English.
- Keep the Yiddish text as the Paratext project and an English back-translation project beside it, with a few books already back-translated by hand.
- In Scripture Forge, connect the back-translation project, set the Yiddish as its source, and choose Generate draft, selecting the books to draft and the finished books to learn from.
- Read the machine's English against the Triglot's English gloss. Wherever they disagree, add a Paratext note for a human reviewer. Disagreement is the signal, as “AF is short of water” was.
Who's who in low-resource translationThe agencies and what each one brings
| Organization | Role | Tools and data you will meet |
|---|---|---|
| SIL | Field linguistics, language survey, literacy, and most of the language software in this exhibit | FLEx, SayMore, Keyman, Webonary, Dictionary App Builder, Scripture Forge, Serval, PTXPrint; co-developer of Paratext; Ethnologue and its language codes originated here |
| United Bible Societies | National Bible Societies translating, publishing, and distributing; co-developer of Paratext | Paratext, Digital Bible Library |
| Wycliffe Global Alliance | About a hundred member organizations worldwide, including Wycliffe USA and Wycliffe UK | Annual Scripture-access statistics from ProgressBible |
| Every Tribe Every Nation (ETEN) and the illumiNations alliance | Coordination and funding toward the 2033 Scripture-access goals | ETEN Innovation Lab reports on AI drafting and new tools |
| Seed Company, Biblica, Lutheran Bible Translators, Pioneer Bible Translators | Translation partners running projects with local churches | Paratext and Scripture Forge projects |
| Wycliffe Associates | Church-led rapid drafting in workshops | MAST method, BTT Writer |
| Faith Comes By Hearing | Audio Scripture and oral-first translation | Bible Brain (text and audio by language) |
| Deaf Bible Society and sign-language partners | Scripture for the world's sign languages | Video translation workflows |
| Clear Bible | Original-language data for alignment and AI | Macula Hebrew and Greek datasets |
| Max Planck Institute for Psycholinguistics | Academic language documentation and archiving | ELAN, The Language Archive |
| Artists for Israel International | Jewish-vernacular Scripture: OJB, Orthodox Yiddish Triglot, interlinears | Paratext; published parallel texts usable as known meaning for related work |
Low-resource languages: how many, and where to find what has been doneThe resource ladder and the lookup method
“Low-resource” is a ladder, not a single category. Which rung a language stands on decides which tools can help.
Out of 7,394 known languages, Wycliffe UK reported in September 2026 that 3,224 still have no Scripture. Counts change monthly; ProgressBible is the underlying source.
Where to look up any language
| Source | What it tells you | Access |
|---|---|---|
| Ethnologue | The ISO 639-3 code, speakers, dialects, vitality, and related languages. Start here to get the three-letter code. | Basic pages free; full data by subscription |
| Glottolog | A bibliography of every known grammar, dictionary, word list, and text for each language, plus its family tree. The best single answer to “what has been done?” | Free |
| OLAC | A union catalogue of archived recordings, texts, and lexicons held in many archives, searchable by language code. | Free |
| Language archives: ELAR, PARADISEC, The Language Archive, SIL Language & Culture Archives | The actual recordings, transcriptions, and field notes from earlier documentation projects. | Mostly free; some items restricted by the community |
| Webonary | Published FLEx dictionaries online, by language. | Free |
| ScriptureEarth | Scripture text, audio, and video available to download in each language. | Free |
| Bible Brain and eBible.org | Which Bible texts and recordings exist digitally, many under open licences. | Free; API keys for developers |
| Joshua Project | People groups using each language, with a Bible-status summary. | Free |
| ProgressBible | The master record of translation status, project by project, for every language. | Restricted to staff of Bible agencies; ask a partner |
| Wycliffe Global Alliance statistics | The public annual summary drawn from ProgressBible. | Free |
The lookup method, step by step
- Get the code. Search the language name in Ethnologue or Glottolog and note its ISO 639-3 code (Eastern Yiddish is
ydd). Every other source uses it. - List what exists. Open the language in Glottolog and read the references: grammars, dictionaries, word lists. Then search the code in OLAC, for example
language-archives.org/language/ydd, for archived recordings and lexicons. - Check the neighbours. Glottolog's family tree and Ethnologue's related-language notes show close relatives. A relative with a finished Bible is your isolog: the same message already read in a neighbouring system.
- Check Scripture status. Search ScriptureEarth, Bible Brain, and Joshua Project for the language. For the full project history, ask a contact in SIL, Wycliffe, or a Bible Society to consult ProgressBible.
- Ask who is working on it. Archive deposits and ProgressBible name the linguists and agencies involved. Contacting them first avoids duplicate work and often unlocks an existing FLEx project on Lexbox.
What Rochefort and Turing teach a translation teamFive working rules
Sources and further reading
- U.S. Naval History and Heritage Command, Battle of Midway overview (H-Gram 006): Station HYPO, the identification of AF, and how codebreaking fitted with other intelligence.
- SIL AI & NLP, Serval and Scripture Forge project pages; Serval source code at github.com/sillsdev/serval.
- Scripture Forge Help: Understanding Draft Generation, Generating a Draft, Connect a Paratext Project.
- ETEN Innovation Lab, Are AI Drafts Really Faster? and An Update on AI and Assisted Translation Technology.
- Paratext, Using Paratext and FieldWorks Together (Paratext 9.4 FLEx integration); Technical Notes on FieldWorks Send/Receive (Associated Lexical Project).
- Platform.Bible and Paratext 10 Studio: github.com/paranext; Scripture Forge extension; Interlinearizer extension.
- SIL, FieldWorks (FLEx); Max Planck Institute for Psycholinguistics, ELAN and its FLEx import documentation.
- Akerman et al., The eBible Corpus: Data and Model Benchmarks for Bible Translation for Low-Resource Languages (2023).
- Wycliffe UK, Bible translation statistics (monthly); Wycliffe Global Alliance, 2025 Global Scripture Access.
- AFII: Orthodox Jewish Bible, Orthodox Yiddish Triglot, 7,000 Languages.