Normalizing Tokens

Breaking text into tokens is only half the job. To make those tokens more easily searchable, they need to go through a normalization process to remove insignificant differences between otherwise identical words, such as uppercase versus lowercase. Perhaps we also need to remove significant differences, to make esta, ésta, and está all searchable as the same word. Would you search for déjà vu, or just for deja vu?

This is the job of the token filters, which receive a stream of tokens from the tokenizer. You can have multiple token filters, each doing its particular job. Each receives the new token stream as output by the token filter before it.

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

00_Intro.asciidoc

00_Intro.asciidoc

Normalizing Tokens

Files

00_Intro.asciidoc

Latest commit

History

00_Intro.asciidoc

File metadata and controls

Normalizing Tokens