0Pricing
Vector Databases: Pinecone, Weaviate & pgvector · درس

البحث الدلالي والبحث الهجين

نفّذوا تقنيات بحث متقدمة تجمع بين تشابه المتجهات ومطابقة الكلمات المفتاحية للحصول على نتائج أفضل.

البحث الدلالي والبحث الهجين درس مجاني في Vector Databases: Pinecone, Weaviate & pgvector على CoddyKit. هذا هو الدرس 1 من أصل 4. يمكنك قراءة الدرس كاملاً أدناه مجاناً — ثم تمرن عليه مباشرة في المتصفح باستخدام محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7. هذا الدرس جزء من مسار التعلم في Vector Databases: Pinecone, Weaviate & pgvector، وتقدمك يتزامن عبر الويب وتطبيق CoddyKit. تتضمن دورة Vector Databases: Pinecone, Weaviate & pgvector 4 دروس في المجموع.

بعض أجزاء هذا الدرس لم تُترجم بعد وتظهر باللغة الإنجليزية.

Beyond Basic Searches

Welcome! In this lesson, we'll dive into advanced search techniques in Weaviate. Moving past simple vector searches, we'll explore how to combine different methods for incredibly precise results.

We'll cover:

  • Pure semantic search
  • Traditional keyword (BM25) search
  • The power of hybrid search

Semantic Search: Meaning First

Semantic search finds items based on their meaning, not just exact words. It uses vector embeddings to represent data, measuring "distance" to find similar concepts. Weaviate uses .with_near_text() for this.

Try this example:

import weaviate
import os

# Connect to your Weaviate instance
# Ensure WEAVIATE_URL is set (e.g., "http://localhost:8080")
client = weaviate.Client(
    url=os.getenv("WEAVIATE_URL", "http://localhost:8080")
)

# Make sure you have an 'Article' class with 'title' and 'content' properties
# and some data imported for this to work!

query_concept = "latest advancements in technology"

response = client.query.get(
    "Article", # Your class name
    ["title", "content"]
).with_near_text(
    {"concepts": [query_concept]}
).with_limit(2).do()

print("Semantic Search Results:")
for item in response["data"]["Get"]["Article"]:
    print(f"- {item['title']}")

Keyword Search Fundamentals

While semantic search is powerful, sometimes you need to find exact keywords. This is where traditional keyword search comes in. Weaviate supports this using the BM25 algorithm.

BM25 (Best Match 25) is a ranking function used by search engines to estimate the relevance of documents to a given search query. It's great for precision when you know exactly what words you're looking for.

Keyword Search with BM25

You can perform keyword searches in Weaviate by combining a .with_where() filter with a text search, and asking for the _additional {score} to see BM25 relevance.

Here's how to search for articles containing specific keywords:

import weaviate
import os

# Connect to your Weaviate instance
client = weaviate.Client(
    url=os.getenv("WEAVIATE_URL", "http://localhost:8080")
)

# Make sure you have an 'Article' class with 'title' and 'content' properties
# and some data imported for this to work!

keyword_query = "AI" # Search for articles containing "AI"

response = client.query.get(
    "Article",
    ["title", "content", "_additional {score}"] # Request BM25 score
).with_where({
    "path": ["content"], # Search in the 'content' field
    "operator": "Like",
    "valueText": f"*{keyword_query}*" # Wildcard search
}).with_limit(2).do()

print("Keyword Search Results:")
for item in response["data"]["Get"]["Article"]:
    print(f"- {item['title']} (BM25 Score: {item['_additional']['score']:.2f})")

Why Hybrid? Limitations

Both semantic and keyword searches have strengths and weaknesses:

  • Semantic: Great for conceptual understanding, but can miss exact terms.
  • Keyword: Excellent for exact matches, but struggles with synonyms or nuanced meaning.

Imagine searching for "best car for family trips." Semantic search might show SUVs, while keyword search might only show articles with "family" and "trip." What if you want both?

Introducing Hybrid Search

Hybrid search combines the strengths of semantic (vector) search and keyword (BM25) search. It retrieves results based on both conceptual similarity and exact term matching, then intelligently fuses them.

This leads to more comprehensive and relevant results, especially for complex or ambiguous queries.

Weaviate's `with_hybrid`

Hybrid search combines semantic and keyword strengths. Weaviate's .with_hybrid() operator makes this easy. It takes both a query and an alpha parameter to control the blend:

  • alpha = 0: Pure keyword
  • alpha = 1: Pure semantic
  • alpha = 0.5: Equal blend (default)

Experiment with this:

import weaviate
import os

# Connect to your Weaviate instance
client = weaviate.Client(
    url=os.getenv("WEAVIATE_URL", "http://localhost:8080")
)

# Make sure you have an 'Article' class with 'title' and 'content' properties
# and some data imported for this to work!

query = "AI tools for data analysis" # Hybrid query text
alpha_value = 0.7 # 0.7 for more semantic weighting

response = client.query.get(
    "Article",
    ["title", "content", "_additional {score, id}"] # Request score & ID
).with_hybrid(
    query=query,
    alpha=alpha_value
).with_limit(3).do()

print(f"Hybrid Search Results (alpha={alpha_value}):")
for item in response["data"]["Get"]["Article"]:
    # The 'score' here is the hybrid score
    print(f"- {item['title']} (Score: {item['_additional']['score']:.2f})")

Understanding Result Fusion (RRF)

When you perform a hybrid search, Weaviate needs a way to combine the rankings from both the semantic and keyword searches into a single, unified list. This is often done using an algorithm like Reciprocal Rank Fusion (RRF).

RRF is a clever method that assigns a score to each document based on its rank in the individual search results. Documents that rank highly in both semantic and keyword searches will get a significantly boosted final score.

Benefits of Hybrid Search

Hybrid search offers several advantages:

  • Improved Relevance: Catches both exact matches and conceptually similar items.
  • Robustness: Performs well even with short, ambiguous, or rare queries.
  • User Satisfaction: Leads to more comprehensive and helpful search results.

It's a crucial technique for building advanced search experiences in AI applications.

Test Your Knowledge

Hybrid search combines semantic and keyword search. Which parameter in Weaviate's .with_hybrid() operator controls the balance between these two search types?

Summary of Advanced Search

Great job! You've mastered advanced search techniques in Weaviate. We explored:

  • Semantic Search: Based on meaning and vector similarity.
  • Keyword Search: Using BM25 for exact term matching.
  • Hybrid Search: Combining both for superior relevance, controlled by the alpha parameter.

These powerful tools will help you build more intelligent and robust search applications. Keep experimenting with different query types and alpha values!

الأسئلة الشائعة

هل درس «البحث الدلالي والبحث الهجين» مجاني؟

نعم — نص درس «البحث الدلالي والبحث الهجين» كامل متاح مجاناً هنا على الويب. لتمرينه بشكل تفاعلي (محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7) وفتح باقي دورة Vector Databases: Pinecone, Weaviate & pgvector، انتقل إلى CoddyKit PRO. تتضمن دورة Vector Databases: Pinecone, Weaviate & pgvector 4 دروس في المجموع.

ماذا ستتعلم في «البحث الدلالي والبحث الهجين»؟

نفّذوا تقنيات بحث متقدمة تجمع بين تشابه المتجهات ومطابقة الكلمات المفتاحية للحصول على نتائج أفضل. تتمرن على Vector Databases: Pinecone, Weaviate & pgvector مع أكواد عملية تشغلها مباشرة في المتصفح، ومدرس ذكاء اصطناعي متاح 24/7 يجيب على أسئلتك أثناء عملك.

هل أحتاج إلى خبرة سابقة لأبدأ Vector Databases: Pinecone, Weaviate & pgvector؟

لا تُشترط خبرة سابقة. Vector Databases: Pinecone, Weaviate & pgvector على CoddyKit منظم للمبتدئين حتى المتقدمين، لذا يمكنك البدء من هنا أو من البداية والتقدم بسرعتك الخاصة. هذا هو الدرس 1 من أصل 4.

كم من الوقت يستغرق درس «البحث الدلالي والبحث الهجين»؟

معظم دروس CoddyKit تستغرق حوالي 5–10 دقائق. كل منها موجز وتفاعلي، لذا تحرز تقدماً مستمراً وتستأنف من حيث توقفت عبر الويب والتطبيق.

هل يمكنني كتابة وتشغيل أكواد في درس Vector Databases: Pinecone, Weaviate & pgvector هذا؟

نعم. كل درس في Vector Databases: Pinecone, Weaviate & pgvector يتضمن محرر أكواد مدمج، لذا تكتب وتشغل أكواداً حقيقية مباشرة في متصفحك وتحصل على تعليقات فورية من الذكاء الاصطناعي — بدون إعداد محلي.

جميع الدروس في هذه الدورة

  1. البحث الدلالي والبحث الهجين
  2. استخدام وحدات Weaviate
  3. استراتيجيات النسخ الاحتياطي والاستعادة
  4. تعدد المستأجرين في Weaviate
← العودة إلى Vector Databases: Pinecone, Weaviate & pgvector