Cryptography · Research
Cracking (and Not Cracking) Cicada 3301
Reproducing the solved pages of the Liber Primus, and using statistics to prove why the remaining pages have resisted the internet for over a decade.
TL;DR
- I built
cicada, a small dependency-free C++ toolkit that transliterates the 29-rune Gematria Primus and runs the ciphers known to break the solved pages (Atbash, Vigenère, totient/prime key-streams), plus a statistical search harness. - It reproduces every known solve exactly, and is validated against planted known-answer ciphers, so a null result means something.
- Pointed at the unsolved pages it recovers nothing, and the Index of Coincidence explains why: those pages are statistically flat (IoC ≈ 0.034, i.e. random over 29 symbols), which rules out the whole class of attacks that cracked the easy pages.
- The one real structural handle is a 17σ suppression of doublets (adjacent identical runes): 0.66% observed vs 3.45% expected.
- A widely shared “27×27 totient map” solution has correct arithmetic but is not a verified decryption, and I show exactly what would make it one.
What is the Liber Primus?
The Cicada 3301 Liber Primus is a corpus of roughly 74 pages written in a 29-symbol runic alphabet, the Gematria Primus. Each rune maps to a Latin letter (or digraph) and to a prime number, in a fixed order:
| idx | 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 12 | 13 | 14 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| rune | ᚠ | ᚢ | ᚦ | ᚩ | ᚱ | ᚳ | ᚷ | ᚹ | ᚻ | ᚾ | ᛁ | ᛄ | ᛇ | ᛈ | ᛉ |
| lat | F | U | TH | O | R | C | G | W | H | N | I | J | EO | P | X |
| prime | 2 | 3 | 5 | 7 | 11 | 13 | 17 | 19 | 23 | 29 | 31 | 37 | 41 | 43 | 47 |
Two properties matter throughout. Several runes are digraphs (TH, EO, NG, OE, AE, IA, EA), and several Latin letters are ambiguous (U/V, C/K, S/Z share a rune). That ambiguity is a recurring source of false confidence, it hands any proposed decoder extra degrees of freedom. Of the ~74 pages, 17 were solved within months of their 2014 release; the remaining ~57 have resisted the public for over a decade.
The toolkit
cicada is a few hundred lines of dependency-free C++. It reads the same community rune files and provides:
cicada translit <file> runes -> Latin (+ gematria sum) cicada stats <file> rune frequencies + Index of Coincidence cicada decode <file> <atbash|caesar|vigenere|totient|primes> [args] cicada vigcrack <file> ... annealed Vigenère-key search (hypothesis tester) cicada subsolve <file> ... 29-symbol substitution hillclimb (hypothesis tester) cicada selftest / cracktest validation on known answers
Crucially, it is validated on known answers. selftest reproduces the published A Warning transliteration exactly, and cracktest encrypts known English-in-runes under a hidden key and confirms the search recovers it. A search tool that can’t be trusted on a known case carries no weight on an unknown one.
Reproducing the solved pages
A Warning, Atbash
The first page is simply the 29-rune alphabet reversed:
A WARNING. BELIEVE NOTHING FROM THIS BOOK. EXCEPT WHAT YOU KNOW TO BE TRUE. TEST THE KNOWLEDGE. FIND YOUR TRUTH. EXPERIENCE YOUR DEATH. DO NOT EDIT OR CHANGE THIS BOOK, OR THE MESSAGE CONTAINED WITHIN, EITHER THE WORDS OR THEIR NUMBERS. FOR ALL IS SACRED.
(BOOC→BOOK, CNOW→KNOW, BELIEUE→BELIEVE, the C/K and U/V ambiguity in action.)
An End (p56), a totient key-stream
Here the key-stream is φ(pₙ) = pₙ − 1 for the n-th prime, taken mod 29 and subtracted from the text:
AN END. WITHIN THE DEEP WEB, THERE EXISTS A PAGE THAT HASHES TO …
The Loss of Divinity, not a cipher at all
A useful reminder not to assume encryption. Its Index of Coincidence is 0.0612 (English-level, not random), and the raw transliteration is already plain English. There is nothing to “solve.”
The frequency fingerprint of the solved corpus
If the solved pages are genuinely English, their letter and word statistics should look like English. Pooling all nine solved plaintext pages (2,979 runes, 727 words) gives exactly that, E is the most common symbol at 12.8% (vs ~12.7% in English), and the function-word skeleton (THE, TO, YOU, IS, A, WE, ARE, AND…) falls right out. The pooled IoC is 0.0614, right where real English spread across 29 runes should land.
The unsolved pages, and the Index of Coincidence
The Index of Coincidence (Friedman, 1920s) is the probability that two symbols drawn at random from a text are the same. Natural language is lumpy, a few letters dominate, so English scores ~0.0667, while uniform random over 29 symbols sits at 1/29 ≈ 0.0345. The key property: the IoC is invariant under monoalphabetic ciphers, but collapses toward the random floor under polyalphabetic / running-key ciphers. One cheap number tells you which family of cipher you can even hope for.
| page group | runes | IoC | periodic Vigenère + substitution search |
|---|---|---|---|
| p0-2 | 729 | 0.0341 | no English |
| p3-7 | 1145 | 0.0346 | no English |
| p8-14 | 1729 | 0.0345 | no English |
| p15-22 | 1903 | 0.0345 | no English |
| p23-26 | 1021 | 0.0343 | no English |
| p27-32 | 1433 | 0.0342 | no English |
| p33-39 | 1680 | 0.0344 | no English |
| p40-53 | 3008 | 0.0345 | no English |
| p54-55 | 308 | 0.0338 | no English |
Every unsolved page sits on the random baseline. This is a strong, informative negative: the statistics are flat, so there is no periodic key or monoalphabetic mapping to recover. It’s not a limitation of the tool, it’s the tool reporting the truth. The solved pages used simple ciphers and left detectable structure; the unsolved pages left none.
An essential caveat: difficulty here means resistance to statistical attack, not unsolvability. An End has the lowest IoC of all (0.0325) yet is solved, because its non-repeating totient key-stream was deduced, not found statistically. A flat IoC says “statistics won’t help here.” It says nothing about whether an insight will.
Deeper probes & a number-theoretic key-stream search
Beyond the single-symbol IoC, a battery of probes test for any exploitable regularity: chi-squared vs uniform, digraphic IoC, autocorrelation/Kasiski periods, and zlib compressibility. Across the corpus they behave like a calibrated instrument, the plaintext control lights up on every probe, and every unsolved page sits at the random baseline.
The flat statistics point to a key-stream cipher, and the solved pages reveal Cicada’s taste for primes and totients. So I brute-forced a library of deterministic integer sequences as key-streams (mod 29): primes, φ(prime), φ(n), Möbius μ, divisor counts, Fibonacci, Lucas, triangular, squares, prime gaps, Thue-Morse, digits of π…
| page | best stream | score | IoC | verdict |
|---|---|---|---|---|
| p56 an_end (control) | totient(primes) | −12.1 | 0.0552 | recovered ✓ |
| p0-2 | primes | −14.6 | 0.0344 | no |
| p8-14 | totient(primes) | −14.7 | 0.0350 | no |
| p27-32 | lucas | −14.6 | 0.0343 | no |
| p40-53 | tau(n) | −14.7 | 0.0344 | no |
The search rediscovers An End with no hints (positive control), cleanly separated from noise, but against the unsolved pages, nothing. I further ruled out repeating keys of any period ≤ 40, homophonic substitution, pure transposition, and English running keys.
The one real signal: doublet suppression
Every probe so far returns “random.” There is exactly one exception, long noted by the CicadaSolvers community: adjacent identical runes are strongly suppressed. Pooled across the unsolved corpus (12,956 runes), doublets occur 86 times where 446 are expected, 0.66% against 3.45%, a 17σ deficit:
| runes | doublets | rate | expected | binomial z | |
|---|---|---|---|---|---|
| unsolved corpus (pooled) | 12,956 | 86 | 0.66% | 3.45% (446) | −17.4 |
It holds on every individual page, and matches the community’s reported values exactly, independent confirmation from a separate codebase. Each rune is drawn almost uniformly from the 28 values that are not its predecessor.
On its own it yields neither plaintext nor key, but it’s the single genuine structural handle in the corpus, and any future attack should treat doublet-freeness as a hard constraint the true solution must satisfy.
Which cipher family fits?
To identify the family, I enciphered known English runes under each candidate cipher and compared the resulting fingerprint to the observed one. Only one family matches on every axis, including the anomalous doublet suppression:
| cipher family (applied to English) | IoC | doublet % | periodic IoC | bigram dep. z |
|---|---|---|---|---|
| monoalphabetic substitution | 0.0614 | 2.62 | 0.0624 | 81.0 |
| Vigenère (repeating key, period 10) | 0.0374 | 3.39 | 0.0613 | 18.2 |
| running key (English + English) | 0.0359 | 3.63 | 0.0383 | 0.4 |
| number stream φ(prime) (the An End cipher) | 0.0345 | 3.12 | 0.0357 | 1.3 |
| one-time pad (uniform key) | 0.0345 | 3.55 | 0.0354 | −0.1 |
| stream + anti-doublet rule | 0.0345 | 0.00 | 0.0355 | 2.6 |
| unsolved pages (observed) | 0.0343 | 0.71 | 0.0370 | 0.6 |
Identification: the unsolved pages are best described as a non-repeating additive key-stream cipher over the 29 runes, the same broad family as An End, carrying one extra deliberate property: adjacent ciphertext runes are almost never equal. They are not a simple substitution, a repeating-key Vigenère, a pure transposition, or an English running key. Fingerprint matching names the mechanism to reverse-engineer; it does not, by itself, read the pages.
Evaluating the “27×27 totient map” solution
A detailed, sincere reconstruction by GitHub user 2retooz270703 proposes that pages 0–2 (the first 729 runes = 27×27 grid) decrypt via a custom route to the text “AS I GO, THE WEATHER TURNS COLD … SEE YOU SOON,” supported by striking numerology. I verified the numerical claims independently, each holds:
| claim | verified |
|---|---|
| “AS I GO THE WEATHER TURNS COLD” = 21 runes, index-sum 233 | ✓ |
| 233 is the 51st prime; COLD = 51 | ✓ |
| Fibonacci F₇ = 13, F₁₃ = 233 | ✓ |
| φ(233) = 232 (the next block’s sum) | ✓ |
| 2163 = 3 × 7 × 103 = U · O · Y (prime values) → “YOU” | ✓ |
So the arithmetic is real. But correct arithmetic is not a verified decryption, for three reasons: (1) the method has many tunable choices, which, combined with the Gematria ambiguity, can be steered toward a pre-chosen sentence; (2) the numerology was found after fixing the plaintext, and striking coincidences are expected by chance (apophenia); (3) it doesn’t match how the genuine pages work, every confirmed solve uses one simple, deterministic cipher with built-in verification (An End literally hashes to a specific value).
What would verify it: a deterministic, parameter-free procedure reproducible by an independent implementation; a null control (the same ruleset yields English from the real ciphertext but not from random runes of identical statistics); and an independently checkable artifact like a valid hash/onion. Until then, it’s an elegant hypothesis, not a solve.
Takeaways
- Measure before you attack. The IoC tells you in milliseconds whether a page is even within reach of a method. Most of the Liber Primus is not.
- Validate on known answers. A search tool is trustworthy only if it recovers a deliberately planted plaintext. Many proposed “solvers” never are.
- Numerology is not cryptanalysis. A rich symbol system always yields striking coincidences for any short text. Proof comes from constraint and reproducibility.
- An honest negative result is a result. “These pages are statistically flat and resist this entire class of attack” is more useful, and more truthful, than a forced reading.
Reproduce it yourself
make ./cicada selftest # validates the toolkit ./cicada decode data/pages/0_warning.txt atbash ./cicada decode data/pages/p56_an_end.txt totient python3 analysis/triage.py # structure probes python3 analysis/keystream_search.py # number-theoretic keystream search (+ An End control) python3 analysis/deep_probes.py # periodic IoC, isomorph, transposition, bigram, entropy python3 analysis/identify.py # cipher-family identification python3 analysis/antidoublet.py # the anti-doublet layer
Full write-up, toolkit, and analysis scripts: github.com/0cool-design/azdecrypt @ cicada. Built on a fork of Jarl Van Eycke’s AZdecrypt; rune corpus from the community LiberPrayground project.
← Back to all posts