Back to posts

Cryptography · Research

Cracking (and Not Cracking) Cicada 3301

Reproducing the solved pages of the Liber Primus, and using statistics to prove why the remaining pages have resisted the internet for over a decade.

By 0cool (Hamood AlSalmani) 11 Oct 2026 ~15 min read Source: azdecrypt @ cicada

TL;DR

  • I built cicada, a small dependency-free C++ toolkit that transliterates the 29-rune Gematria Primus and runs the ciphers known to break the solved pages (Atbash, Vigenère, totient/prime key-streams), plus a statistical search harness.
  • It reproduces every known solve exactly, and is validated against planted known-answer ciphers, so a null result means something.
  • Pointed at the unsolved pages it recovers nothing, and the Index of Coincidence explains why: those pages are statistically flat (IoC ≈ 0.034, i.e. random over 29 symbols), which rules out the whole class of attacks that cracked the easy pages.
  • The one real structural handle is a 17σ suppression of doublets (adjacent identical runes): 0.66% observed vs 3.45% expected.
  • A widely shared “27×27 totient map” solution has correct arithmetic but is not a verified decryption, and I show exactly what would make it one.

What is the Liber Primus?

The Cicada 3301 Liber Primus is a corpus of roughly 74 pages written in a 29-symbol runic alphabet, the Gematria Primus. Each rune maps to a Latin letter (or digraph) and to a prime number, in a fixed order:

idx01234567891011121314
runeᚠᚢᚦᚩᚱᚳᚷᚹᚻᚾᛁᛄᛇᛈᛉ
latFUTHORCGWHNIJEOPX
prime23571113171923293137414347

Two properties matter throughout. Several runes are digraphs (TH, EO, NG, OE, AE, IA, EA), and several Latin letters are ambiguous (U/V, C/K, S/Z share a rune). That ambiguity is a recurring source of false confidence, it hands any proposed decoder extra degrees of freedom. Of the ~74 pages, 17 were solved within months of their 2014 release; the remaining ~57 have resisted the public for over a decade.

The toolkit

cicada is a few hundred lines of dependency-free C++. It reads the same community rune files and provides:

cicada translit <file>     runes -> Latin (+ gematria sum)
cicada stats    <file>     rune frequencies + Index of Coincidence
cicada decode   <file> <atbash|caesar|vigenere|totient|primes> [args]
cicada vigcrack <file> ...  annealed Vigenère-key search (hypothesis tester)
cicada subsolve <file> ...  29-symbol substitution hillclimb (hypothesis tester)
cicada selftest / cracktest validation on known answers

Crucially, it is validated on known answers. selftest reproduces the published A Warning transliteration exactly, and cracktest encrypts known English-in-runes under a hidden key and confirms the search recovers it. A search tool that can’t be trusted on a known case carries no weight on an unknown one.

Reproducing the solved pages

A Warning, Atbash

The first page is simply the 29-rune alphabet reversed:

A WARNING. BELIEVE NOTHING FROM THIS BOOK. EXCEPT WHAT YOU KNOW TO BE TRUE. TEST THE KNOWLEDGE. FIND YOUR TRUTH. EXPERIENCE YOUR DEATH. DO NOT EDIT OR CHANGE THIS BOOK, OR THE MESSAGE CONTAINED WITHIN, EITHER THE WORDS OR THEIR NUMBERS. FOR ALL IS SACRED.

(BOOC→BOOK, CNOW→KNOW, BELIEUE→BELIEVE, the C/K and U/V ambiguity in action.)

An End (p56), a totient key-stream

Here the key-stream is φ(pₙ) = pₙ − 1 for the n-th prime, taken mod 29 and subtracted from the text:

AN END. WITHIN THE DEEP WEB, THERE EXISTS A PAGE THAT HASHES TO …

The Loss of Divinity, not a cipher at all

A useful reminder not to assume encryption. Its Index of Coincidence is 0.0612 (English-level, not random), and the raw transliteration is already plain English. There is nothing to “solve.”

The frequency fingerprint of the solved corpus

If the solved pages are genuinely English, their letter and word statistics should look like English. Pooling all nine solved plaintext pages (2,979 runes, 727 words) gives exactly that, E is the most common symbol at 12.8% (vs ~12.7% in English), and the function-word skeleton (THE, TO, YOU, IS, A, WE, ARE, AND…) falls right out. The pooled IoC is 0.0614, right where real English spread across 29 runes should land.

Rune frequency across the solved corpus
Rune frequency across the solved corpus, it tracks English closely.

The unsolved pages, and the Index of Coincidence

The Index of Coincidence (Friedman, 1920s) is the probability that two symbols drawn at random from a text are the same. Natural language is lumpy, a few letters dominate, so English scores ~0.0667, while uniform random over 29 symbols sits at 1/29 ≈ 0.0345. The key property: the IoC is invariant under monoalphabetic ciphers, but collapses toward the random floor under polyalphabetic / running-key ciphers. One cheap number tells you which family of cipher you can even hope for.

page grouprunesIoCperiodic Vigenère + substitution search
p0-27290.0341no English
p3-711450.0346no English
p8-1417290.0345no English
p15-2219030.0345no English
p23-2610210.0343no English
p27-3214330.0342no English
p33-3916800.0344no English
p40-5330080.0345no English
p54-553080.0338no English

Every unsolved page sits on the random baseline. This is a strong, informative negative: the statistics are flat, so there is no periodic key or monoalphabetic mapping to recover. It’s not a limitation of the tool, it’s the tool reporting the truth. The solved pages used simple ciphers and left detectable structure; the unsolved pages left none.

Liber Primus pages ranked by predicted difficulty
Every page the IoC labels “easy” was in fact solved; every flat page is unsolved.

An essential caveat: difficulty here means resistance to statistical attack, not unsolvability. An End has the lowest IoC of all (0.0325) yet is solved, because its non-repeating totient key-stream was deduced, not found statistically. A flat IoC says “statistics won’t help here.” It says nothing about whether an insight will.

Deeper probes & a number-theoretic key-stream search

Beyond the single-symbol IoC, a battery of probes test for any exploitable regularity: chi-squared vs uniform, digraphic IoC, autocorrelation/Kasiski periods, and zlib compressibility. Across the corpus they behave like a calibrated instrument, the plaintext control lights up on every probe, and every unsolved page sits at the random baseline.

The flat statistics point to a key-stream cipher, and the solved pages reveal Cicada’s taste for primes and totients. So I brute-forced a library of deterministic integer sequences as key-streams (mod 29): primes, φ(prime), φ(n), Möbius μ, divisor counts, Fibonacci, Lucas, triangular, squares, prime gaps, Thue-Morse, digits of π…

pagebest streamscoreIoCverdict
p56 an_end (control)totient(primes)−12.10.0552recovered ✓
p0-2primes−14.60.0344no
p8-14totient(primes)−14.70.0350no
p27-32lucas−14.60.0343no
p40-53tau(n)−14.70.0344no

The search rediscovers An End with no hints (positive control), cleanly separated from noise, but against the unsolved pages, nothing. I further ruled out repeating keys of any period ≤ 40, homophonic substitution, pure transposition, and English running keys.

Friedman periodic-IoC test
A simulated repeating key spikes at its period; the unsolved pages stay pinned to the random floor, no repeating key exists.

The one real signal: doublet suppression

Every probe so far returns “random.” There is exactly one exception, long noted by the CicadaSolvers community: adjacent identical runes are strongly suppressed. Pooled across the unsolved corpus (12,956 runes), doublets occur 86 times where 446 are expected, 0.66% against 3.45%, a 17σ deficit:

runesdoubletsrateexpectedbinomial z
unsolved corpus (pooled)12,956860.66%3.45% (446)−17.4

It holds on every individual page, and matches the community’s reported values exactly, independent confirmation from a separate codebase. Each rune is drawn almost uniformly from the 28 values that are not its predecessor.

Bigram heatmaps
English lights up specific digraphs (THE, ER…); the unsolved corpus is a uniform field, except for a dark diagonal, the suppressed doublets.

On its own it yields neither plaintext nor key, but it’s the single genuine structural handle in the corpus, and any future attack should treat doublet-freeness as a hard constraint the true solution must satisfy.

Which cipher family fits?

To identify the family, I enciphered known English runes under each candidate cipher and compared the resulting fingerprint to the observed one. Only one family matches on every axis, including the anomalous doublet suppression:

cipher family (applied to English)IoCdoublet %periodic IoCbigram dep. z
monoalphabetic substitution0.06142.620.062481.0
Vigenère (repeating key, period 10)0.03743.390.061318.2
running key (English + English)0.03593.630.03830.4
number stream φ(prime) (the An End cipher)0.03453.120.03571.3
one-time pad (uniform key)0.03453.550.0354−0.1
stream + anti-doublet rule0.03450.000.03552.6
unsolved pages (observed)0.03430.710.03700.6
Cipher-family fingerprint heatmap
Green = close to observed, red = far. Only the anti-doublet stream matches on all four axes.

Identification: the unsolved pages are best described as a non-repeating additive key-stream cipher over the 29 runes, the same broad family as An End, carrying one extra deliberate property: adjacent ciphertext runes are almost never equal. They are not a simple substitution, a repeating-key Vigenère, a pure transposition, or an English running key. Fingerprint matching names the mechanism to reverse-engineer; it does not, by itself, read the pages.

Evaluating the “27×27 totient map” solution

A detailed, sincere reconstruction by GitHub user 2retooz270703 proposes that pages 0–2 (the first 729 runes = 27×27 grid) decrypt via a custom route to the text “AS I GO, THE WEATHER TURNS COLD … SEE YOU SOON,” supported by striking numerology. I verified the numerical claims independently, each holds:

claimverified
“AS I GO THE WEATHER TURNS COLD” = 21 runes, index-sum 233✓
233 is the 51st prime; COLD = 51✓
Fibonacci F₇ = 13, F₁₃ = 233✓
φ(233) = 232 (the next block’s sum)✓
2163 = 3 × 7 × 103 = U · O · Y (prime values) → “YOU”✓

So the arithmetic is real. But correct arithmetic is not a verified decryption, for three reasons: (1) the method has many tunable choices, which, combined with the Gematria ambiguity, can be steered toward a pre-chosen sentence; (2) the numerology was found after fixing the plaintext, and striking coincidences are expected by chance (apophenia); (3) it doesn’t match how the genuine pages work, every confirmed solve uses one simple, deterministic cipher with built-in verification (An End literally hashes to a specific value).

What would verify it: a deterministic, parameter-free procedure reproducible by an independent implementation; a null control (the same ruleset yields English from the real ciphertext but not from random runes of identical statistics); and an independently checkable artifact like a valid hash/onion. Until then, it’s an elegant hypothesis, not a solve.

Takeaways

Reproduce it yourself

make
./cicada selftest                               # validates the toolkit
./cicada decode  data/pages/0_warning.txt atbash
./cicada decode  data/pages/p56_an_end.txt totient

python3 analysis/triage.py            # structure probes
python3 analysis/keystream_search.py  # number-theoretic keystream search (+ An End control)
python3 analysis/deep_probes.py       # periodic IoC, isomorph, transposition, bigram, entropy
python3 analysis/identify.py          # cipher-family identification
python3 analysis/antidoublet.py       # the anti-doublet layer

Full write-up, toolkit, and analysis scripts: github.com/0cool-design/azdecrypt @ cicada. Built on a fork of Jarl Van Eycke’s AZdecrypt; rune corpus from the community LiberPrayground project.

← Back to all posts