0Pricing
LLM Apps in Production (RAG + Vector DB + Caching) · درس

اختبار تطبيق RAG وتقييمه

عزّز الثقة في أول تطبيق RAG لديك بإنشاء مجموعة اختبار وقياس جودة الاسترجاع والإجابات باستخدام مقاييس عملية قبل إطلاقه.

اختبار تطبيق RAG وتقييمه درس مجاني في LLM Apps in Production (RAG + Vector DB + Caching) على CoddyKit. هذا هو الدرس 4 من أصل 4. يمكنك قراءة الدرس كاملاً أدناه مجاناً — ثم تمرن عليه مباشرة في المتصفح باستخدام محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7. هذا الدرس جزء من مسار التعلم في LLM Apps in Production (RAG + Vector DB + Caching)، وتقدمك يتزامن عبر الويب وتطبيق CoddyKit. تتضمن دورة LLM Apps in Production (RAG + Vector DB + Caching) 4 دروس في المجموع.

بعض أجزاء هذا الدرس لم تُترجم بعد وتظهر باللغة الإنجليزية.

Why Evaluate RAG

A RAG app can look fine on a few queries and fail badly on others. Without measurement you cannot tell if a change helped or hurt.

Evaluation gives you a repeatable score to guide improvements.

Two Things to Measure

RAG quality has two parts:

  • Retrieval: did we fetch the right documents?
  • Generation: did the answer use them correctly?

A bad answer can come from either, so measure both.

Building a Test Set

Create a small set of questions with known correct answers and the documents that contain them. Even 20 to 50 examples are enough to start.

testset = [
    {'q': 'What is the refund window?',
     'answer': '30 days',
     'source': 'policy.md'}
]

Retrieval Metric: Hit Rate

Hit rate (or recall@k) checks whether the correct source appears in the top-k retrieved chunks. High hit rate means retrieval is doing its job.

def hit(retrieved, expected_source):
    return any(d.metadata['source'] == expected_source
               for d in retrieved)

Faithfulness

Faithfulness asks: is the answer supported by the retrieved context, or did the model make things up? An LLM judge can score this automatically.

Answer Relevance

Answer relevance measures whether the response actually addresses the question, regardless of sources. A faithful answer can still be off-topic.

LLM as a Judge

You can use a strong model to grade outputs against the expected answer, returning a pass or score with a reason.

judge_prompt = (
    'Question: {q}\nExpected: {gold}\n'
    'Got: {pred}\nIs it correct? Answer yes or no.'
)

Running the Evaluation

Loop over the test set, run your pipeline, and aggregate scores into a single report you can compare across versions.

scores = []
for case in testset:
    pred = rag.invoke(case['q'])
    scores.append(grade(case, pred))
print(sum(scores) / len(scores))

Comparing Configurations

Change one variable — chunk size, k, prompt, model — rerun the same test set, and compare scores. This turns guesswork into evidence-based tuning.

Watching for Regressions

Keep the test suite in CI. When a change drops a metric, you catch the regression before users do. Treat evaluation like unit tests for AI quality.

Improving From Results

Use failures to guide fixes:

  • Low hit rate? Adjust chunking or retrieval
  • Low faithfulness? Strengthen grounding instructions
  • Low relevance? Improve the prompt

Quick Check

Test your evaluation knowledge.

Recap

You learned to evaluate your RAG app:

  • Measure both retrieval and generation
  • Build a small test set with known answers
  • Use hit rate, faithfulness, and answer relevance
  • Let an LLM judge grade outputs
  • Compare configs and guard against regressions in CI

Evaluation turns RAG improvement into a measurable, repeatable process.

الأسئلة الشائعة

هل درس «اختبار تطبيق RAG وتقييمه» مجاني؟

نعم — نص درس «اختبار تطبيق RAG وتقييمه» كامل متاح مجاناً هنا على الويب. لتمرينه بشكل تفاعلي (محرر أكواد مدمج ومدرس ذكاء اصطناعي متاح 24/7) وفتح باقي دورة LLM Apps in Production (RAG + Vector DB + Caching)، انتقل إلى CoddyKit PRO. تتضمن دورة LLM Apps in Production (RAG + Vector DB + Caching) 4 دروس في المجموع.

ماذا ستتعلم في «اختبار تطبيق RAG وتقييمه»؟

عزّز الثقة في أول تطبيق RAG لديك بإنشاء مجموعة اختبار وقياس جودة الاسترجاع والإجابات باستخدام مقاييس عملية قبل إطلاقه. تتمرن على LLM Apps in Production (RAG + Vector DB + Caching) مع أكواد عملية تشغلها مباشرة في المتصفح، ومدرس ذكاء اصطناعي متاح 24/7 يجيب على أسئلتك أثناء عملك.

هل أحتاج إلى خبرة سابقة لأبدأ LLM Apps in Production (RAG + Vector DB + Caching)؟

لا تُشترط خبرة سابقة. LLM Apps in Production (RAG + Vector DB + Caching) على CoddyKit منظم للمبتدئين حتى المتقدمين، لذا يمكنك البدء من هنا أو من البداية والتقدم بسرعتك الخاصة. هذا هو الدرس 4 من أصل 4.

كم من الوقت يستغرق درس «اختبار تطبيق RAG وتقييمه»؟

معظم دروس CoddyKit تستغرق حوالي 5–10 دقائق. كل منها موجز وتفاعلي، لذا تحرز تقدماً مستمراً وتستأنف من حيث توقفت عبر الويب والتطبيق.

هل يمكنني كتابة وتشغيل أكواد في درس LLM Apps in Production (RAG + Vector DB + Caching) هذا؟

نعم. كل درس في LLM Apps in Production (RAG + Vector DB + Caching) يتضمن محرر أكواد مدمج، لذا تكتب وتشغل أكواداً حقيقية مباشرة في متصفحك وتحصل على تعليقات فورية من الذكاء الاصطناعي — بدون إعداد محلي.

جميع الدروس في هذه الدورة

  1. اختيار مزوّد نماذج اللغة الكبيرة
  2. أساسيات تحميل البيانات وتقسيم النصوص
  3. بناء مسار RAG بسيط
  4. اختبار تطبيق RAG وتقييمه
← العودة إلى LLM Apps in Production (RAG + Vector DB + Caching)