Words are matched as Unicode letters and numbers
The counter recognizes letter and number sequences across scripts and permits internal apostrophes or hyphens. Language-specific compounds and scripts without spaces can still require a dedicated tokenizer.