LLM Apps in Production (RAG + Vector DB + Caching) · Lekcja

Obserwowalność: logi, miary i śledzenie

Zintegruje Pan/Pani kompleksowe rejestrowanie logów, gromadzenie miar i rozproszone śledzenie, aby uzyskać szczegółowy wgląd w działanie aplikacji LLM.

Lekcja 2 z 411 kroki

Obserwowalność: logi, miary i śledzenie to bezpłatna lekcja LLM Apps in Production (RAG + Vector DB + Caching) na CoddyKit. To lekcja 2 z 4. Możesz przeczytać całą lekcję poniżej za darmo — a potem ćwiczyć ją interaktywnie w przeglądarce z wbudowanym edytorem kodu i tutorem AI dostępnym 24/7. To część ścieżki edukacyjnej LLM Apps in Production (RAG + Vector DB + Caching), a Twój postęp synchronizuje się między webem a aplikacją CoddyKit. Kurs LLM Apps in Production (RAG + Vector DB + Caching) zawiera 4 lekcji w sumie.

Części tej lekcji nie zostały jeszcze przetłumaczone i są wyświetlane po angielsku.

What is Observability?

In this lesson, we'll explore observability, a crucial concept for managing complex software systems, especially LLM applications.

Observability means understanding the internal state of a system by examining the data it produces. Think of it as having X-ray vision into your application's behavior.

For LLM apps, this helps us answer critical questions like:

  • Why is a request slow?
  • Is the RAG retrieval working as expected?
  • Are we incurring unexpected costs?

Logs: Recording Events

Logs are timestamped records of events that happen within your application. They are like a diary of your system's activities.

For LLM applications, logs are essential for:

  • Tracking incoming user prompts.
  • Storing responses from the LLM.
  • Recording intermediate steps in a RAG pipeline (e.g., documents retrieved).
  • Capturing errors or warnings.

They provide detailed contextual information for debugging and post-mortem analysis.

Logging LLM Interactions

Here's a simple Python example demonstrating how to log an LLM interaction. We're using Python's built-in logging module.

This helps you see exactly what prompts were sent and what responses were received, which is vital for debugging and improving your application.

import logging

logging.basicConfig(
    level=logging.INFO,
    format='%(asctime)s - %(levelname)s - %(message)s'
)

def call_llm(prompt):
    logging.info(f"LLM Request: '{prompt[:40]}...' ")
    # Simulate LLM processing
    response = f"Simulated response to: {prompt}"
    logging.info(f"LLM Response: '{response[:40]}...' ")
    return response

if __name__ == "__main__":
    user_prompt = "Explain observability simply."
    result = call_llm(user_prompt)
    print(f"Application output: {result}")

Metrics: Measuring Performance

Metrics are numerical measurements collected over time, providing aggregated insights into your system's health and performance.

Unlike logs, which are individual events, metrics are typically quantitative values that can be visualized as graphs and dashboards. Key metrics for LLM apps include:

  • Latency: How long it takes for the LLM to respond.
  • Token Usage: Input/output tokens consumed per request.
  • Error Rate: Percentage of failed LLM calls or RAG retrievals.
  • Cache Hit Rate: How often cached responses are used.

Collecting Custom Metrics

You can collect custom metrics to understand specific aspects of your LLM application. This example shows how to track the number of LLM calls and their average latency.

In a real-world scenario, you'd send these metrics to a monitoring system like Prometheus or Datadog.

import time

class LLMMetrics:
    def __init__(self):
        self.total_calls = 0
        self.total_latency = 0.0

    def record_call(self, duration):
        self.total_calls += 1
        self.total_latency += duration

    def get_avg_latency(self):
        if self.total_calls == 0:
            return 0.0
        return self.total_latency / self.total_calls

metrics_store = LLMMetrics()

def call_llm_with_metrics(prompt):
    start_time = time.time()
    # Simulate LLM processing
    time.sleep(0.05) # simulate 50ms work
    response = f"Simulated reply to: {prompt}"
    end_time = time.time()
    metrics_store.record_call(end_time - start_time)
    return response

if __name__ == "__main__":
    print("Collecting LLM call metrics...")
    call_llm_with_metrics("Hi")
    call_llm_with_metrics("How are you?")
    print(f"Total calls: {metrics_store.total_calls}")
    print(f"Avg latency: {metrics_store.get_avg_latency():.3f}s")

Tracing: Following Request Paths

Tracing is about following a single request as it flows through multiple services and components in a distributed system. This is especially vital for RAG applications that involve many steps: user input, embedding generation, vector DB lookup, LLM call, etc.

A trace visualizes the entire journey of a request, showing the exact path it took and the time spent in each operation.

Traces, Spans, and Context

A trace is a complete end-to-end journey of a request. It's composed of multiple spans.

  • A span represents a single operation or unit of work within a trace (e.g., 'retrieve documents', 'call embedding model', 'invoke LLM').
  • Spans have a parent-child relationship, forming a tree structure that shows dependencies.
  • Context propagation ensures that a unique trace ID follows the request across different services, linking all related spans together.

This helps pinpoint bottlenecks or failures across microservices.

OpenTelemetry for Tracing

While implementing tracing from scratch is complex, tools like OpenTelemetry (an open-source observability framework) provide standardized ways to instrument your code.

You'd use OpenTelemetry SDKs to:

  • Start a new trace when a request comes in.
  • Create new spans for each significant operation (e.g., a function call to a vector database or an LLM API).
  • Propagate the trace context to downstream services.

This allows you to visualize the full request flow in a tracing UI.

The Observability Triangle

Logs, metrics, and traces are often called the "observability triangle" because they offer complementary views of your system:

  • Logs: The granular details and events.
  • Metrics: The aggregated numbers and trends.
  • Traces: The end-to-end journey of a request.

Together, they provide a comprehensive understanding of your LLM application's behavior, making it easier to diagnose issues, optimize performance, and ensure reliability in production.

Quick Check: Observability

You've learned about the three pillars of observability. Let's see if you can distinguish their primary uses.

Recap: Deep Insights

Congratulations! You've explored the world of observability for LLM applications.

  • We defined observability as understanding internal system state from external data.
  • We learned about logs for detailed event recording.
  • We covered metrics for aggregated performance measurements.
  • We understood traces for visualizing end-to-end request flows.

By integrating these three pillars, you gain powerful insights, enabling you to build more reliable, performant, and cost-efficient LLM systems.

Bezpłatny start

Ucz się LLM Apps in Production (RAG + Vector DB + Caching) dzięki korepetycjom AI — za darmo

Pisz i uruchamiaj kod w przeglądarce, otrzymuj natychmiastową pomoc od korepetytora AI dostępnego 24/7 i kontynuuj naukę w sieci lub w aplikacji.

Kursy
12
Lekcje
48

Często zadawane pytania

Czy lekcja „Obserwowalność: logi, miary i śledzenie” jest bezpłatna?

Tak — pełny tekst „Obserwowalność: logi, miary i śledzenie” jest dostępny za darmo tutaj w sieci. Aby ćwiczyć ją interaktywnie (wbudowany edytor kodu i tutor AI dostępny 24/7) i odblokować resztę kursu LLM Apps in Production (RAG + Vector DB + Caching), przejdź na CoddyKit PRO. Kurs LLM Apps in Production (RAG + Vector DB + Caching) zawiera 4 lekcji w sumie.

Co nauczysz się w „Obserwowalność: logi, miary i śledzenie”?

Zintegruje Pan/Pani kompleksowe rejestrowanie logów, gromadzenie miar i rozproszone śledzenie, aby uzyskać szczegółowy wgląd w działanie aplikacji LLM. Ćwiczysz LLM Apps in Production (RAG + Vector DB + Caching) z praktycznym kodem, który uruchamiasz bezpośrednio w przeglądarce, a tutor AI dostępny 24/7 odpowiada na Twoje pytania podczas pracy nad lekcją.

Czy potrzebuję doświadczenia, aby zacząć LLM Apps in Production (RAG + Vector DB + Caching)?

Nie wymagamy żadnego doświadczenia. LLM Apps in Production (RAG + Vector DB + Caching) w CoddyKit jest strukturyzowany dla początkujących i zaawansowanych użytkowników, więc możesz zacząć tutaj lub od początku i uczyć się w swoim tempie. To lekcja 2 z 4.

Ile czasu zajmuje lekcja „Obserwowalność: logi, miary i śledzenie”?

Większość lekcji CoddyKit trwa około 5–10 minut. Każda lekcja to mały, interaktywny krok, dzięki czemu robisz systematyczne postępy i zawsze wracasz dokładnie do tego samego miejsca — na webie i w aplikacji.

Czy mogę pisać i uruchamiać kod w tej lekcji LLM Apps in Production (RAG + Vector DB + Caching)?

Tak. Każda lekcja LLM Apps in Production (RAG + Vector DB + Caching) zawiera wbudowany edytor kodu, więc piszesz i uruchamiasz prawdziwy kod bezpośrednio w przeglądarce i od razu otrzymujesz sprzężenie zwrotne od AI — bez konfiguracji na komputerze.

Wszystkie lekcje w tym kursie

  1. Skalowanie horyzontalne komponentów RAG
  2. Obserwowalność: logi, miary i śledzenie
  3. Alerty i reagowanie na incydenty w operacjach LLM
  4. Testy obciążeniowe i planowanie przepustowości
← Powrót do LLM Apps in Production (RAG + Vector DB + Caching)