Guessing a five-letter word in six tries has become a daily ritual for millions of Americans, and “what’s the best starting word?” is a debate that never dies. Underneath it sits real mathematics: counting, probability and information theory. Here the game is just a jumping-off point for those three ideas, which are useful long after the puzzle is solved. Every list and number in the examples is hypothetical, picked to keep the arithmetic clean.
Counting: how many five-letter words are there?
Start with the multiplication principle: if one choice has options and an independent one has , the pair has outcomes. A five-letter string over a 26-letter alphabet, with repeats allowed, has 26 options in each slot.
If all five letters must be different, the first slot has 26 options, the second 25, and so on. That is the number of permutations of 26 objects taken 5 at a time.
Almost none of these strings are actual words. A typical answer list for this kind of game runs to a few thousand words, a tiny slice of the 11.9 million strings. That is what makes the game playable: language has already ruled out nearly everything before your first guess.
One more count: each tile of feedback gets one of three colors (right letter, right spot; letter elsewhere; letter absent). Five tiles give = 243 possible color patterns. Some never occur in practice, but 243 is the ceiling on how many groups a single guess can create.
Probability: the odds of a blind guess
If N candidates remain and each is equally likely, picking one at random wins with probability 1/N. With 200 candidates that is 1/200, or 0.5%. With 2 candidates, 50%. The whole game is about shrinking N as fast as possible, and every guess is a question that partitions the list, one group per color pattern.
Information: why the logarithm shows up
Suppose the answer is one of N equally likely options. A well-designed yes-or-no question cuts the list in half. How many such questions do you need? The number with , that is, . That quantity is measured in bits.
With 200 candidates, is about 7.64 bits. A guess in the game is not a yes-or-no question: it has up to 243 answers, so at best it could deliver , about 7.92 bits. No real word comes close, but the comparison shows the theoretical limit.
When a guess splits N candidates into groups of sizes , the probability of landing in group is . Two measures grade the guess:
The first is intuitive: you land in a group with probability proportional to its size, and that size is what you’re left with. The second, Shannon entropy, measures how many “halvings” the guess delivers on average. Lower expected remaining and higher entropy mean a better guess.
Worked example: comparing two guesses
Say 200 candidates remain (hypothetical numbers). Guess A splits them into four groups of 80, 60, 40 and 20. Guess B splits them into four groups of 50. Which is better?
- Probabilities for guess A: each group size divided by 200.
- Expected candidates left after A: sum the squares and divide by 200.
- Expected candidates left after B.
- Entropy of guess A, in bits (rounded).
- Entropy of guess B: four equal groups give exactly 2 bits.
- Conclusion: B leaves fewer candidates on average and delivers more information. Same number of groups, but balanced groups are worth more.
Check. A’s groups add up to 80 + 60 + 40 + 20 = 200 and B’s to 4 × 50 = 200, so no candidate went missing. A’s probabilities add to 0.4 + 0.3 + 0.2 + 0.1 = 1. Four groups can never yield more than = 2 bits, and that maximum happens only with equal groups, which is exactly case B. Finally, after B you still have 50 candidates, about 5.64 bits to go, which is why the game takes more than one guess.
Common mistakes
- Confusing strings with words. 26 to the 5th counts strings; the real answer list is far smaller, and that list is what sets N.
- Using permutations when repeats are allowed. Real words repeat letters; the 7,893,600 figure only applies when all letters are distinct.
- Assuming more greens on guess one is always better. What matters is the expected size of the group you land in, not one lucky round.
- Taking a plain average of group sizes. The average of 80, 60, 40 and 20 is 50, but the expected value is 60, because bigger groups are more likely. Each group is weighted by its own size.
- Using the wrong log base. Bits mean base 2. With the natural log the unit becomes “nats”; convert by dividing by .
Frequently asked questions
How many five-letter combinations can you make from 26 letters?
As strings, 26 to the 5th power, which is 11,881,376. With no repeated letters, 26 × 25 × 24 × 23 × 22 = 7,893,600. Real words are a small fraction of that.
What does it mean for a guess to be worth 2 bits?
On average it shrinks the candidate list as much as two well-chosen yes-or-no questions would, down to a quarter of its size.
Why does the expected value use the sum of squares?
You land in group i with probability n_i/N, and then n_i candidates remain. Adding n_i/N times n_i over all groups gives the sum of squares divided by N.
Can math guarantee a solve in a few guesses?
Not in general; it depends on the word list. What math gives you is a rule for picking guesses that cut the list the most on average.