Artificial intelligence has now touched one of the most storied puzzles in modern codebreaking history: the Enigma cipher that challenged Allied cryptanalysts during World War II and helped define Alan Turing's legacy. In a development that blends wartime history, archival detective work and the accelerating capabilities of frontier AI, two cryptanalysts say they have used large language models from OpenAI and Anthropic to crack two long-unsolved Enigma messages that had resisted human researchers for years.
The result is striking not only because of the technical feat, but because of the method. Rather than simply asking a model to guess at a solution, the researchers used the systems as active collaborators in historical investigation. Carter Leffen, a developer, told OpenAI's newest model, Astra, to search a database of Enigma messages for one that remained unbroken and decode it. According to the account, the model did much more than produce a guess: it conducted archival research, identified context clues, built a simulator of the Enigma machine and ultimately recovered the plaintext of a message that had baffled researchers since 2005.
Leffen then used Astra to build an interactive website explaining the entire problem, underscoring how these models are increasingly being used not just to answer questions, but to assemble the research scaffolding around them. The breakthrough was later validated by Frode Weierud, a retired electrical engineer and longtime cryptology enthusiast who maintains the Crypto Cellar website, a repository of resources, records and message databases. Weierud said the solution left him in "awe."
That reaction carries weight in a field where precision matters and false confidence can be fatal to a claim. Weierud's validation suggests the model did not merely produce a plausible reconstruction, but one that aligned with the historical and cryptographic record. Still, the process also raised a new and somewhat unsettling question: how much did the model know, and from where?
Leffen's Astra logs reportedly included discussion of archived messages in a "private collection" not hosted by Weierud. He said he is still not sure whether the model accessed those materials directly, though he speculated they may have been shared by another researcher online or drawn from the German government's public archives. Weierud, for his part, said the model behaved like a seasoned researcher.
"GPTâ6 Astra is behaving like a very professional cryptanalyst and archive researcher," he wrote. "What it has achieved in two days would take a human researcher weeks or even months. Personally, I spent several weeks researching the Bundesarchiv files GPTâ6 Astra refers to."
The second breakthrough came on September 21, when Jack Willis, a cybersecurity executive and cryptanalyst, contacted Weierud to say he had used Anthropic's Claude Opus 5 to break a different unsolved message. In that case, Willis provided more guidance to the model, which then used the known signature of a particular officer's name to help crack the text. The contrast between the two efforts is notable: one model appears to have pursued the problem with comparatively little direction, while the other succeeded with more human steering. Together, they suggest that the frontier is not whether AI can assist cryptanalysis, but how much initiative it can take before human oversight becomes secondary.
The historical resonance is hard to miss. Turing is widely remembered for the test that bears his name, but his more consequential wartime contribution may have been his work against Enigma, including the development of the Bombe, an early machine that helped the United Kingdom translate German messages during the war. Even then, not every message fell. A handful of archival Enigma messages remain unbroken, often because of transcription errors or mistakes made by the original encoders.
Weierud says there are now just seven unbroken Enigma messages left, plus one message whose plaintext is known even though the code remains unbroken. If these AI-assisted breakthroughs are any indication, that list may not stay intact for long.
For historians, cryptographers and AI researchers alike, the significance goes beyond a single code. The episode suggests that large language models, when paired with archival databases and domain expertise, may be able to accelerate work once thought to require painstaking manual effort. In this case, the machines did not replace the human experts; they amplified them, sifting records, testing hypotheses and reconstructing context at a pace that astonished even seasoned researchers.
That may be the most important lesson of all. Turing's legacy was never only about whether machines could think. It was also about whether machines could help humans solve problems that had seemed impossible. On that measure, Astra and Opus 5 have just passed another of Turing's tests.
