テキストアナライザーのカスタマイズ
特定のフィールドにカスタムアナライザーを作成して適用し、検索時のテキストの処理、ステミング、インデックス作成方法を制御します。
「テキストアナライザーのカスタマイズ」はCoddyKit上の無料Elasticsearch & Full Text Search Systemsレッスンです。 これはレッスン2/4です。 下記で完全なレッスンを無料で読むことができます。その後、ブラウザ内の組み込みコードエディタと24時間対応のAIチューターでハンズオン演習できます。 これはElasticsearch & Full Text Search Systems学習パスの一部であり、ウェブとCoddyKitアプリ全体で進捗が同期されます。 Elasticsearch & Full Text Search Systemsコースには全4レッスンが含まれています。
このレッスンの一部はまだ翻訳されておらず、英語で表示されています。
Why Custom Analyzers?
Elasticsearch comes with powerful default text analyzers, but sometimes your data needs a special touch. This is where custom analyzers shine!
They allow you to precisely control how your text fields are processed for search, ensuring optimal relevancy and accuracy for your specific use case.
The Analyzer Recipe
Recall that every analyzer, custom or built-in, follows a three-step process to transform raw text into searchable tokens:
- Character Filters: Clean up the raw input string (e.g., remove HTML tags).
- Tokenizer: Breaks the processed string into individual words or tokens.
- Token Filters: Modifies, adds, or removes tokens (e.g., lowercase, remove stop words, apply stemming).
A custom analyzer lets you pick and choose these ingredients!
Defining Custom Analyzers
You define custom analyzers within an index's settings block, under analysis. This tells Elasticsearch how to process text for that index.
Here's the basic structure for creating a custom analyzer:
PUT /my_custom_index
{
"settings": {
"analysis": {
"analyzer": {
"my_custom_analyzer": {
"type": "custom",
"char_filter": [],
"tokenizer": "standard",
"filter": []
}
}
}
}
}Custom Character Filters
Character filters are the first step, acting on the raw text. They can remove or replace characters before tokenization. You can define your own or use built-in ones.
html_strip: Removes HTML tags.mapping: Replaces specified characters or strings.
Here's how to define a custom mapping filter:
PUT /my_index_with_char_filter
{
"settings": {
"analysis": {
"char_filter": {
"ampersand_to_and": {
"type": "mapping",
"mappings": ["& => and "]
}
},
"analyzer": {
"my_analyzer": {
"type": "custom",
"char_filter": ["ampersand_to_and"],
"tokenizer": "standard",
"filter": ["lowercase"]
}
}
}
}
}Selecting a Tokenizer
The tokenizer breaks the stream of characters from the character filters into individual tokens (words). Your choice here is crucial for how words are identified.
Common built-in tokenizers you can use in custom analyzers include:
standard: Good for most languages, grammar-based.whitespace: Splits text only on whitespace.keyword: Treats the entire input as a single token (useful for exact values).pattern: Splits text based on a regular expression.
Custom Token Filters
Token filters refine the tokens generated by the tokenizer. This is where most of the search logic resides, like handling synonyms or stemming.
You can define custom versions of filters or use built-in ones:
lowercase: Converts tokens to lowercase.stop: Removes common, less meaningful words (stop words).synonym: Replaces tokens with their synonyms.stemmer: Reduces words to their root form.
Let's define a custom stop word filter:
PUT /my_index_with_token_filter
{
"settings": {
"analysis": {
"filter": {
"my_custom_stop_words": {
"type": "stop",
"stopwords": ["a", "the", "is", "and", "are"]
}
},
"analyzer": {
"my_analyzer": {
"type": "custom",
"tokenizer": "standard",
"filter": ["lowercase", "my_custom_stop_words"]
}
}
}
}
}Building a Full Custom Analyzer
Now, let's combine these concepts to create a practical custom analyzer for blog post content. It will:
- Remove HTML tags.
- Tokenize standard text.
- Lowercase all tokens.
- Remove common English stop words.
This analyzer is then applied to the content field.
PUT /blog_posts_index
{
"settings": {
"analysis": {
"char_filter": {
"html_strip_char_filter": {
"type": "html_strip"
}
},
"filter": {
"english_stop_words": {
"type": "stop",
"stopwords": ["the", "a", "an", "is", "are"]
}
},
"analyzer": {
"blog_content_analyzer": {
"type": "custom",
"char_filter": ["html_strip_char_filter"],
"tokenizer": "standard",
"filter": ["lowercase", "english_stop_words"]
}
}
}
},
"mappings": {
"properties": {
"content": {
"type": "text",
"analyzer": "blog_content_analyzer"
}
}
}
}Applying to Field Mappings
Once your custom analyzer is defined in the index settings, you apply it to a text field within your index's mapping. This tells Elasticsearch to use your custom logic when indexing and searching that specific field.
You simply specify the analyzer parameter with the name of your custom analyzer:
PUT /blog_posts_index/_mapping
{
"properties": {
"content": {
"type": "text",
"analyzer": "blog_content_analyzer"
},
"title": {
"type": "text",
"analyzer": "standard"
}
}
}Testing with _analyze API
How can you be sure your custom analyzer works as expected? Use the _analyze API! It lets you simulate how text will be processed by any analyzer.
This is an indispensable tool for debugging and validating your text analysis setup.
GET /blog_posts_index/_analyze
{
"analyzer": "blog_content_analyzer",
"text": "The <b>quick</b> brown fox jumps over the lazy dog."
}Quiz Time!
You've learned about the components of a custom analyzer and how to define them. Let's test your knowledge!
Custom Analyzers: Your Search Superpower
Congratulations! You've learned how to harness the power of custom analyzers in Elasticsearch.
- You can now define custom character filters, tokenizers, and token filters.
- You know how to combine these components to create a tailor-made analyzer for your data.
- You understand how to apply this analyzer to specific fields in your mappings.
- And importantly, you know how to test your analyzer using the
_analyzeAPI.
This skill is crucial for building highly relevant and accurate search experiences!
よくある質問
「テキストアナライザーのカスタマイズ」レッスンは無料ですか?
はい。「テキストアナライザーのカスタマイズ」の完全なテキストはこのウェブで無料で読めます。インタラクティブに演習し(組み込みコードエディタと24時間対応のAIチューター)、Elasticsearch & Full Text Search Systemsコースの残りをアンロックするには、CoddyKit PROにアップグレードしてください。 Elasticsearch & Full Text Search Systemsコースには全4レッスンが含まれています。
「テキストアナライザーのカスタマイズ」で何を学びますか?
特定のフィールドにカスタムアナライザーを作成して適用し、検索時のテキストの処理、ステミング、インデックス作成方法を制御します。 ブラウザで直接実行するハンズオンコードでElasticsearch & Full Text Search Systemsを演習し、24時間対応のAIチューターがレッスンを進める中での質問に答えます。
Elasticsearch & Full Text Search Systemsを始めるのに経験は必要ですか?
事前経験は必要ありません。CoddyKitのElasticsearch & Full Text Search Systemsは初級者から上級者向けに構成されているため、ここから始めるか最初から始めて、自分のペースで進むことができます。 これはレッスン2/4です。
「テキストアナライザーのカスタマイズ」レッスンにはどのくらい時間がかかりますか?
ほとんどのCoddyKitレッスンは約5~10分かかります。各レッスンはコンパクトでインタラクティブなので、着実に進歩し、ウェブとアプリ全体で正確に前回の場所から再開できます。
このElasticsearch & Full Text Search Systemsレッスンでコードを書いて実行できますか?
はい。すべてのElasticsearch & Full Text Search Systemsレッスンに組み込みコードエディタが含まれているため、ブラウザでリアルコードを書いて実行し、即座のAIフィードバックを取得できます。ローカル設定は不要です。
このコースのすべてのレッスン
- アナライザー、トークナイザー、フィルター
- テキストアナライザーのカスタマイズ
- ブーストと関連性スコアリング
- 同義語とステミング