Scoring pipeline
The text is split into tokens, lowercased, and common stop words (the, and, of) are removed. Each surviving term receives a term frequency, its count of occurrences, and an inverse document frequency, the log of the total number of documents divided by the number containing it, where a document is usually one sentence or paragraph of your input. The product of the two is the term's score.