P(w|h): Probability of word w, given some history h Example: P(because| today I am happy) w: because h: today I am happy # Approach 1: Relative Frequency Count Step 1: Take a text corpus Step 2: Count the number of times 'today I am happy' appears Step 3: Count the number of times it is followed by 'because' P(because| today I am happy) = Count(today I am happy because)/ Count(today I am happy) # In essence, we seek to answer the question, Out of the N times we saw the history h, how many times did the word w follow it? # Disadvantages of the approach: When the size of the text corpus is large, then this approach has to traverse the entire corpus. Not scalable and is clearly suboptimal in performance. # Approach 2: Bigram Model Bigram model approximates the probability of a word given all the previous words by using only the conditional probability of the preceding word.