Guessing a five-letter word in six tries has become a daily ritual for millions of Americans, and “what’s the best starting word?” is a debate that never dies. Underneath it sits real mathematics: counting, probability and information theory. Here the game is just a jumping-off point for those three ideas, which are useful long after the puzzle is solved. Every list and number in the examples is hypothetical, picked to keep the arithmetic clean.

Counting: how many five-letter words are there?

Start with the multiplication principle: if one choice has aa options and an independent one has bb, the pair has a⋅ba \cdot b outcomes. A five-letter string over a 26-letter alphabet, with repeats allowed, has 26 options in each slot.

26⋅26⋅26⋅26⋅26=265=11,881,37626 \cdot 26 \cdot 26 \cdot 26 \cdot 26 = 26^5 = 11{,}881{,}376
Five-letter strings with repetition allowed (multiplication principle).

If all five letters must be different, the first slot has 26 options, the second 25, and so on. That is the number of permutations of 26 objects taken 5 at a time.

P(26,5)=26!21!=26⋅25⋅24⋅23⋅22=7,893,600P(26,5) = \frac{26!}{21!} = 26 \cdot 25 \cdot 24 \cdot 23 \cdot 22 = 7{,}893{,}600
Five-letter strings with no repeated letter.

Almost none of these strings are actual words. A typical answer list for this kind of game runs to a few thousand words, a tiny slice of the 11.9 million strings. That is what makes the game playable: language has already ruled out nearly everything before your first guess.

One more count: each tile of feedback gets one of three colors (right letter, right spot; letter elsewhere; letter absent). Five tiles give 353^5 = 243 possible color patterns. Some never occur in practice, but 243 is the ceiling on how many groups a single guess can create.

Probability: the odds of a blind guess

If N candidates remain and each is equally likely, picking one at random wins with probability 1/N. With 200 candidates that is 1/200, or 0.5%. With 2 candidates, 50%. The whole game is about shrinking N as fast as possible, and every guess is a question that partitions the list, one group per color pattern.

Information: why the logarithm shows up

Suppose the answer is one of N equally likely options. A well-designed yes-or-no question cuts the list in half. How many such questions do you need? The number kk with 2k=N2^k = N, that is, k=log⁡2Nk = \log_2 N. That quantity is measured in bits.

I=log⁡2NI = \log_2 N
Bits needed to pin down one of N equally likely options.

With 200 candidates, log⁡2200\log_2 200 is about 7.64 bits. A guess in the game is not a yes-or-no question: it has up to 243 answers, so at best it could deliver log⁡2243\log_2 243, about 7.92 bits. No real word comes close, but the comparison shows the theoretical limit.

05010015020025002468candidatesbitslog₂ N
Bits = log₂ N: doubling the candidates adds just 1 bit. 200 candidates need about 7.64 bits; 243 patterns give about 7.92.

When a guess splits N candidates into groups of sizes n1,n2,…,nkn_1, n_2, \dots, n_k, the probability of landing in group ii is pi=ni/Np_i = n_i/N. Two measures grade the guess:

E[remaining]=∑ipi⋅ni=∑ini2NE[\text{remaining}] = \sum_i p_i \cdot n_i = \frac{\sum_i n_i^2}{N}
Expected number of candidates left after the guess.
H=∑ipilog⁡21piH = \sum_i p_i \log_2 \frac{1}{p_i}
Entropy: the expected information from the guess, in bits.

The first is intuitive: you land in a group with probability proportional to its size, and that size is what you’re left with. The second, Shannon entropy, measures how many “halvings” the guess delivers on average. Lower expected remaining and higher entropy mean a better guess.

Worked example: comparing two guesses

Say 200 candidates remain (hypothetical numbers). Guess A splits them into four groups of 80, 60, 40 and 20. Guess B splits them into four groups of 50. Which is better?

80A: group 160A: group 240A: group 320A: group 450B: eachgroup
Guess A leaves groups of 80, 60, 40 and 20 candidates; guess B leaves four equal groups of 50.
Guess A vs. guess B
  1. p=0.4, 0.3, 0.2, 0.1\displaystyle p = 0.4,\ 0.3,\ 0.2,\ 0.1 Probabilities for guess A: each group size divided by 200.
  2. 802+602+402+202200=12,000200=60\displaystyle \frac{80^2+60^2+40^2+20^2}{200} = \frac{12{,}000}{200} = 60 Expected candidates left after A: sum the squares and divide by 200.
  3. 4⋅502200=10,000200=50\displaystyle \frac{4 \cdot 50^2}{200} = \frac{10{,}000}{200} = 50 Expected candidates left after B.
  4. 0.4log⁡22.5+0.3log⁡2103+0.2log⁡25+0.1log⁡210≈1.85\displaystyle 0.4\log_2 2.5 + 0.3\log_2 \tfrac{10}{3} + 0.2\log_2 5 + 0.1\log_2 10 \approx 1.85 Entropy of guess A, in bits (rounded).
  5. 4⋅14log⁡24=2\displaystyle 4 \cdot \tfrac14 \log_2 4 = 2 Entropy of guess B: four equal groups give exactly 2 bits.
  6. Conclusion: B leaves fewer candidates on average and delivers more information. Same number of groups, but balanced groups are worth more.

Check. A’s groups add up to 80 + 60 + 40 + 20 = 200 and B’s to 4 × 50 = 200, so no candidate went missing. A’s probabilities add to 0.4 + 0.3 + 0.2 + 0.1 = 1. Four groups can never yield more than log⁡24\log_2 4 = 2 bits, and that maximum happens only with equal groups, which is exactly case B. Finally, after B you still have 50 candidates, about 5.64 bits to go, which is why the game takes more than one guess.

60Guess A50Guess B
Expected candidates left after the guess: 60 with A (bigger groups weigh more) versus 50 with B, which splits the list evenly.

Common mistakes

  • Confusing strings with words. 26 to the 5th counts strings; the real answer list is far smaller, and that list is what sets N.
  • Using permutations when repeats are allowed. Real words repeat letters; the 7,893,600 figure only applies when all letters are distinct.
  • Assuming more greens on guess one is always better. What matters is the expected size of the group you land in, not one lucky round.
  • Taking a plain average of group sizes. The average of 80, 60, 40 and 20 is 50, but the expected value is 60, because bigger groups are more likely. Each group is weighted by its own size.
  • Using the wrong log base. Bits mean base 2. With the natural log the unit becomes “nats”; convert by dividing by ln⁡2\ln 2.

Frequently asked questions

How many five-letter combinations can you make from 26 letters?

As strings, 26 to the 5th power, which is 11,881,376. With no repeated letters, 26 × 25 × 24 × 23 × 22 = 7,893,600. Real words are a small fraction of that.

What does it mean for a guess to be worth 2 bits?

On average it shrinks the candidate list as much as two well-chosen yes-or-no questions would, down to a quarter of its size.

Why does the expected value use the sum of squares?

You land in group i with probability n_i/N, and then n_i candidates remain. Adding n_i/N times n_i over all groups gives the sum of squares divided by N.

Can math guarantee a solve in a few guesses?

Not in general; it depends on the word list. What math gives you is a rule for picking guesses that cut the list the most on average.