Skip to main content
Live
Main content

Austrian Academy and Mistral release Apollo, an Ancient Greek LLM trained on 600M words

The model fills gaps in papyrus fragments and will be free for academics through a chatbot interface.

Jaeden Schafer
Editor in Chief · · 5 min read
Austrian Academy and Mistral release Apollo, an Ancient Greek LLM trained on 600M words

The Austrian Academy of Science on Wednesday released Apollo, a large language model trained on roughly 600 million historical Greek words and built to fill in the gaps of damaged papyrus fragments. Developed with French AI lab Mistral and technology services firm Sail Reply, the model is billed as the first advanced LLM for Ancient Greek and will be free for academics through a chatbot interface. Academic libraries hold hundreds of thousands of Ancient Greek papyrus fragments, many too tattered to read without expert reconstruction.

Apollo's training corpus draws on manuscripts, papyri, and inscriptions spanning centuries of Greek writing. The model proposes the most statistically likely words or passages for missing sections of a document, giving scholars a shortlist to choose from rather than a single verdict. That design is a deliberate hedge against error: a probabilistic model handing scholars three candidate words is different from one silently rewriting the historical record.

Reconstruction has historically been slow and specialist work. A researcher first has to identify word divisions — Ancient Greek was written without spaces — then date the fragment, weigh the political and cultural context, and consult reference material to pick plausible fills.

There are very few people in the world who are that good at Greek history.
Stephen Colvin, Professor of classics and historical linguistics at University College London

Key facts

  • 01The Austrian Academy of Science released Apollo on Wednesday, calling it the first advanced large language model for Ancient Greek.
  • 02Apollo was trained on roughly 600 million historical Greek words drawn from manuscripts, papyri, and inscriptions.
  • 03The model was built in partnership with French AI lab Mistral and technology services firm Sail Reply, and will be free for academics via a chatbot.
  • 04Academic libraries hold hundreds of thousands of Ancient Greek papyrus fragments awaiting restoration, a task Apollo is designed to accelerate.
  • 05Apollo proposes multiple candidate words for each gap rather than a single answer, keeping human scholars in the loop on final selections.

Much of that expertise is now compressed into Apollo's weights. Dimitris Vlitas, a partner at Sail Reply, said the ability to unlock knowledge this way "was unthinkable a year ago," a reflection of how quickly domain-specific LLMs have matured on narrow corpora that would once have been considered too small or too specialised for a neural approach.

The model adapts its output to the dialect and register of the input text. Anna Dolganov, a historian and papyrologist at the Austrian Academy of Science, described the behaviour in concrete terms — Homeric Greek in, Homeric Greek out; Doric inscription in, Doric out. That register-matching is what separates a useful research tool from a generic autocomplete.

When it sees Homer, it supplements Homeric Greek. When it sees an inscription in Doric dialect, it uses Doric dialect.
Anna Dolganov, Historian and papyrologist at the Austrian Academy of Science

Oxford, home to the world's largest ancient papyrus collection, is one of the institutions that stands to benefit. Armand D'Angour, a professor of classical languages and literature at the University of Oxford, said a machine offering three candidate words for each gap "would speed up matters considerably," freeing academics to focus on the implications of documents rather than the mechanics of decoding them.

The scale of the payoff is modest by design. Most unrestored papyri are mundane — personal letters, marital contracts, civil service records — rather than lost literary works. Stephen Colvin, a professor of classics and historical linguistics at University College London, cautioned that laypeople expecting new plays by Sophocles will be disappointed. What Apollo can do is add texture: fresh detail about daily life, commerce, and administration in the ancient world, one fragment at a time.

The team has built in safeguards against the obvious failure mode — a model confidently hallucinating history. Candidate lists rather than single answers, and an explicit assumption that a trained scholar makes the final call, keep the model as a research accelerator rather than a replacement.

Related · from this week
Oxford study: warmer AI models are 60% more likely to be wrong
Jaeden Schafer · 5 min read →

If Apollo works, Vlitas said the same technique could be applied to Latin, Egyptian, or any other academic discipline where a large corpus needs to be distilled and indexed. That generalisation is the more interesting business story. Narrow-domain LLMs trained on tightly curated corpora — legal filings, medical literature, engineering standards — are a category that mainstream frontier labs have largely left to specialists.

Apollo sits alongside a broader run of AI-in-research milestones. OpenAI recently said its models had solved a 200-year-old math problem, and Google DeepMind released a large dataset mapping how genetic mutations affect molecular biology, compiled with AI assistance. Different fields, same pattern: models trained or fine-tuned on domain data producing outputs that would have required rare human expertise.

Mistral's involvement is the quiet commercial signal here. The French lab has been building a reputation for open, adaptable models that plug into vertical projects rather than competing head-on for consumer chat share. A collaboration with a national academy on a heritage-language model is precisely the kind of deployment that showcases that positioning without requiring a scaled consumer product. Expect more national institutions, in more languages and disciplines, to commission similar tools now that the template exists.

ShareXLinkedInEmail
AI Box

Every AI model. One chat.

The latest models from ChatGPT, Claude, Gemini, Sora, ElevenLabs — 80+ models in a single chat. Compare answers side by side. Pick the best one every time.

  • ChatGPT, Claude, Gemini, Grok, DeepSeek — in one chat
  • Generate images & video with Sora, Veo, Ideogram
  • Compare any two models side by side
  • From $8.99/mo · 80+ models, all included
Try AI Boxaibox.ai
Trusted by 3,000+ teams
Got a tip?

Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.

Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.

AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at

Keep reading

More from Models

Oxford Internet Institute
Models

Oxford study: warmer AI models are 60% more likely to be wrong

Fine-tuning five models including GPT-4o for empathy raised error rates 7.43 points and worsened sharply when users said they felt sad.

Jaeden Schafer5 min read
Nvidia logo
Business

Nvidia backstops $500B AI data center push by guaranteeing GPU resale values

Jensen Huang gets Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs and KKR to fund the buildout — in exchange for covering 25% of any collateral shortfall.

Jaeden Schafer5 min read
Samsung in talks to back Mistral at €20 billion valuation
Business

Samsung in talks to back Mistral at €20 billion valuation

The French AI lab would more than triple its valuation from its last round, giving Samsung a strategic foothold in European frontier models.

Jaeden Schafer4 min read