Random words

Available under Creative Commons-NonCommercial-ShareAlike 4.0 International License.

To choose a random word from the histogram, the simplest algorithm is to build a list with multiple copies of each word, according to the observed frequency, and then choose from the list:

def random_word(h):
t = []
for word, freq in h.items():
t.extend([word] * freq)
 
return random.choice(t)

The expression [word] * freq creates a list with freq copies of the string word. Theextend method is similar to append except that the argument is a sequence.

Exercise 13.7.This algorithm works, but it is not very efficient; each time you choose a randomword, it rebuilds the list, which is as big as the original book. An obvious improvement is to buildthe list once and then make multiple selections, but the list is still big.

An alternative is:

Use keys to get a list of the words in the book.
Build a list that contains the cumulative sum of the word frequencies (see Exercise 10.3 in Map, ﬁlter and reduce). The last item in this list is the total number of words in the book, n.
Choose a random number from 1 to n. Use a bisection search (See Exercise 10.11 in Exercises) to find the index where the random number would be inserted in the cumulative sum.
Use the index to find the corresponding word in the word list.

Write a program that uses this algorithm to choose a random word from the book. Solution: http:// thinkpython. com/ code/ analyze_ book3. py.

3153 reads

You are here

Random words