Humanity’s filing system for the ancient world is backed up like a dishwasher in a shared flat. Museums and libraries hold hundreds of thousands of ancient inscriptions and more than a million papyrus fragments, and scholars — the people who can actually read the things — have fully deciphered only about 8 percent of the papyri. The rest sit in drawers, written in a language that takes years of training to read even when mice, sand and volcanoes have left it alone.

Which is why a coalition with an unusual membership roll — the Austrian Academy of Sciences, the AI company Mistral, SAIL Reply and Northeastern University — has built Apollo, announced by the academy on September 23 as what the team calls the world’s first large language model made specifically for Ancient Greek. Its job description is beautifully unglamorous: fill in the blanks.

Broken lines, missing words and whole damaged sections are the rule in ancient documents, not the exception. Where a papyrologist might spend days weighing what a torn sentence once said, Apollo offers candidate readings in seconds — a scholarly autocomplete for texts last edited two millennia ago.

The model was trained on about 600 million words of Ancient Greek. That is pocket lint next to the tens of billions of words fed to modern language models, but the supply of authentic Ancient Greek prose is, famously, not growing. The details appear in the team’s paper, Apollo Restore: A Foundation LLM for Historical Greek Optimized for Fill-in-the-Middle Restoration of Ancient Greek Texts, posted on arXiv.

Grading the machine

Testing a model like this required some honesty-preserving trickery. The researchers took real papyri that scholars had already reconstructed, blanked out sections, and asked Apollo to fill them. The correct reading turned up among its top suggestions in about 80 percent of cases.

More telling was a blind evaluation in which 20 specialists in papyrology, epigraphy and philology compared Apollo’s suggestions against published scholarly readings without knowing which was which. For documentary papyri — the receipts, letters and administrivia of antiquity — the researchers estimated Apollo’s reading was at least as good as the published version in about 77 percent of cases. When the machine and the published scholars flatly disagreed, the blinded experts preferred Apollo in 16 percent of documentary-papyrus cases and in about 20 percent of inscription cases. The humans grading the exam included, in effect, the humans who had set it, which makes that 20 percent sting a little.

Fieldwork

Apollo has already been let loose on the real thing. In public demonstrations it completed an ancient birth certificate, recovered important passages from a papyrus scroll damaged by the eruption of Vesuvius, and helped researchers trace evidence of Roman law reaching a city on the Black Sea.

The model also has a feel for period and register. According to co-author Dolganov, Apollo can spot when a passage matches the structure and vocabulary of something like Homer’s Odyssey and then suggest missing text in a similar style — which is either deeply impressive or the founding of history’s most erudite cover band.

None of this means the papyrologists are finished. A model that is at least as good three-quarters of the time still needs people who can argue about whether a smudge is a sigma. What it promises is triage on a grand scale: plausible readings, in seconds, for texts that would otherwise wait another century for eyes that can read them. That could pull thousands of unread papyri within reach of study, with new details of ancient life, law, literature and history in the pile.

And there is something fitting about where all this computing power has ended up. The telltale evidence in that Black Sea case was, in part, a reference to a Roman tax linked to prostitution. Two thousand years later, the empire’s most legible footprint turns out to be what it always is: paperwork, much of it receipts.