Asynchroniczne frameworki agentów: LangChain i nie tylko
ainvoke(), astream() i asynchroniczne łańcuchy w LangChain i LangGraph.
Asynchroniczne frameworki agentów: LangChain i nie tylko to bezpłatna lekcja AI Agents na CoddyKit. To lekcja 4 z 4. Możesz przeczytać całą lekcję poniżej za darmo — a potem ćwiczyć ją interaktywnie w przeglądarce z wbudowanym edytorem kodu i tutorem AI dostępnym 24/7. To część ścieżki edukacyjnej AI Agents, a Twój postęp synchronizuje się między webem a aplikacją CoddyKit. Kurs AI Agents zawiera 4 lekcji w sumie.
Asynchroniczne wykonywanie w LangChain
LangChain udostępnia asynchroniczne wersje wszystkich swoich interfejsów. Każdy komponent, który ma invoke(), ma również ainvoke(), a każdy komponent z stream() ma astream(). Asynchroniczność jest zalecanym podejściem w agentach produkcyjnych.
ainvoke() dla asynchronicznych wywołań LLM
ainvoke() jest asynchronicznym odpowiednikiem invoke(). Należy używać go wewnątrz funkcji asynchronicznych, aby wykonywać nieblokujące wywołania LLM. Pozwala to wielu agentom lub żądaniom współdzielić pętlę zdarzeń.
import asyncio
from langchain_openai import ChatOpenAI
from langchain_core.messages import HumanMessage
llm = ChatOpenAI(model='gpt-4o-mini', api_key='sk-...')
async def async_agent_call(question: str) -> str:
# ainvoke: non-blocking, releases event loop while waiting for OpenAI
response = await llm.ainvoke([HumanMessage(content=question)])
return response.content
async def handle_multiple_users(questions: list) -> list:
# All three LLM calls run concurrently
results = await asyncio.gather(*[async_agent_call(q) for q in questions])
return results
questions = [
'What is Python?',
'What is TypeScript?',
'What is Rust?'
]
results = asyncio.run(handle_multiple_users(questions))
for q, a in zip(questions, results):
print(f'Q: {q[:30]}... A: {a[:50]}...')astream() dla strumieniowania tokenów
astream() zwraca tokeny w miarę ich nadejścia z LLM. Pozwala to strumieniować odpowiedzi do użytkownika w czasie rzeczywistym, bez oczekiwania na nadejście całego wyniku.
import asyncio
from langchain_openai import ChatOpenAI
from langchain_core.messages import HumanMessage
llm = ChatOpenAI(model='gpt-4o-mini', api_key='sk-...')
async def stream_response(question: str):
print(f'Streaming answer to: {question}\n')
full_response = ''
async for chunk in llm.astream([HumanMessage(content=question)]):
token = chunk.content
if token:
print(token, end='', flush=True) # Print each token as it arrives
full_response += token
print() # New line after streaming
return full_response
async def main():
await stream_response('List 3 benefits of async programming in Python')
asyncio.run(main())AsyncCallbackHandler
Wywołania zwrotne LangChain są uruchamiane podczas określonych zdarzeń: rozpoczęcia LLM, zakończenia LLM, rozpoczęcia narzędzia i błędu łańcucha. AsyncCallbackHandler obsługuje te zdarzenia asynchronicznie, bez blokowania pętli agenta.
from langchain_core.callbacks import AsyncCallbackHandler
from typing import Any, Dict, List
import time
class LatencyCallbackHandler(AsyncCallbackHandler):
def __init__(self):
self.step_times = {}
self.step_counts = {}
async def on_llm_start(self, serialized: Dict, prompts: List[str], **kwargs):
run_id = str(kwargs.get('run_id', ''))
self.step_times[run_id] = time.perf_counter()
async def on_llm_end(self, response, **kwargs):
run_id = str(kwargs.get('run_id', ''))
if run_id in self.step_times:
elapsed_ms = (time.perf_counter() - self.step_times[run_id]) * 1000
print(f'LLM call completed in {elapsed_ms:.0f}ms')
async def on_tool_start(self, serialized: Dict, input_str: str, **kwargs):
tool_name = serialized.get('name', 'unknown')
print(f'Tool starting: {tool_name}')
async def on_tool_error(self, error: Exception, **kwargs):
print(f'Tool error: {error}')
handler = LatencyCallbackHandler()
print('Async callback handler created')
# Use: llm.ainvoke([...], config={'callbacks': [handler]})Asynchroniczne funkcje węzłów LangGraph
Węzły LangGraph mogą być funkcjami asynchronicznymi. Po zdefiniowaniu węzła jako async def LangGraph oczekuje na jego zakończenie podczas wykonywania grafu. Jest to zalecany wzorzec dla grafów produkcyjnych.
import asyncio
from langgraph.graph import StateGraph, END
from typing import TypedDict, List
class AgentState(TypedDict):
question: str
entities: List[str]
context: str
answer: str
async def extract_entities_node(state: AgentState) -> AgentState:
await asyncio.sleep(0.1) # Simulate async NLP call
entities = state['question'].split()[:3] # Simplified
return {'entities': entities}
async def retrieve_context_node(state: AgentState) -> AgentState:
await asyncio.sleep(0.2) # Simulate async vector search
context = f'Context for entities: {state["entities"]}'
return {'context': context}
async def generate_answer_node(state: AgentState) -> AgentState:
await asyncio.sleep(0.3) # Simulate async LLM call
answer = f'Answer based on: {state["context"]}'
return {'answer': answer}
# Build async graph
graph = StateGraph(AgentState)
graph.add_node('extract', extract_entities_node)
graph.add_node('retrieve', retrieve_context_node)
graph.add_node('generate', generate_answer_node)
graph.set_entry_point('extract')
graph.add_edge('extract', 'retrieve')
graph.add_edge('retrieve', 'generate')
graph.add_edge('generate', END)
app = graph.compile()
print('Async LangGraph compiled')Asynchroniczne strumieniowanie z LangGraph
LangGraph obsługuje asynchroniczne strumieniowanie stanów pośrednich podczas wykonywania grafu. Należy użyć astream(), aby obserwować wyniki poszczególnych węzłów zaraz po ich zakończeniu, zamiast czekać na pełne wykonanie.
import asyncio
async def stream_graph_execution(graph_app, initial_state: dict):
print('Graph execution streaming:')
async for step_output in graph_app.astream(initial_state):
for node_name, state_delta in step_output.items():
print(f' Node [{node_name}] completed:')
for key, value in state_delta.items():
print(f' {key}: {value}')
# Run the async graph
initial = {
'question': 'What is machine learning?',
'entities': [],
'context': '',
'answer': ''
}
asyncio.run(stream_graph_execution(app, initial))Ograniczanie szybkości za pomocą semaforów
OpenAI i inne API LLM mają limity szybkości dotyczące liczby żądań na minutę. Należy użyć semafora asynchronicznego, aby nie przekroczyć limitu szybkości nawet podczas wykonywania wielu współbieżnych zadań agentów.
import asyncio
from langchain_openai import ChatOpenAI
from langchain_core.messages import HumanMessage
llm = ChatOpenAI(model='gpt-4o-mini', api_key='sk-...')
# Limit to 10 concurrent LLM calls
LLM_SEMAPHORE = asyncio.Semaphore(10)
async def rate_limited_llm_call(question: str) -> str:
async with LLM_SEMAPHORE:
response = await llm.ainvoke([HumanMessage(content=question)])
return response.content
async def process_large_batch(questions: list) -> list:
print(f'Processing {len(questions)} questions with max 10 concurrent LLM calls')
tasks = [rate_limited_llm_call(q) for q in questions]
results = await asyncio.gather(*tasks, return_exceptions=True)
successes = [r for r in results if not isinstance(r, Exception)]
failures = [r for r in results if isinstance(r, Exception)]
print(f'Success: {len(successes)}, Failed: {len(failures)}')
return results
# Process 50 questions with max 10 concurrent calls
questions = [f'Question {i}: What is concept number {i}?' for i in range(20)]
asyncio.run(process_large_batch(questions))Definicje narzędzi asynchronicznych
W LangChain funkcje narzędzi mogą być asynchroniczne. Na narzędzia asynchroniczne oczekuje się podczas wykonywania agenta, co umożliwia nieblokujące wywołania API wewnątrz samego narzędzia.
import asyncio
import httpx
from langchain.tools import tool
@tool
async def async_web_search(query: str) -> str:
'''Search the web for information about the query.'''
async with httpx.AsyncClient() as client:
# Real implementation would use a search API
response = await client.get(
'https://api.search.example.com/search',
params={'q': query, 'api_key': 'your-key'},
timeout=10.0
)
response.raise_for_status()
results = response.json()
return '\n'.join([r['snippet'] for r in results.get('items', [])[:3]])
@tool
async def async_fetch_document(url: str) -> str:
'''Fetch and return the text content of a URL.'''
async with httpx.AsyncClient() as client:
response = await client.get(url, timeout=15.0)
return response.text[:3000] # Limit content size
print('Async tools defined')
print('Use with: agent.ainvoke({"input": "your question"})')Asynchroniczny agent korzystający bezpośrednio z OpenAI
Można zbudować w pełni asynchroniczną pętlę agenta bezpośrednio za pomocą OpenAI SDK, bez LangChain. Zapewnia to maksymalną kontrolę i minimalny narzut.
import asyncio
import openai
import json
client = openai.AsyncOpenAI(api_key='sk-...')
TOOLS = [
{'type': 'function', 'function': {
'name': 'web_search',
'description': 'Search the web',
'parameters': {'type': 'object', 'properties': {'query': {'type': 'string'}}, 'required': ['query']}
}}
]
async def async_tool_call(tool_name: str, args: dict) -> str:
if tool_name == 'web_search':
await asyncio.sleep(0.3) # Simulate search
return f'Search results for: {args["query"]}'
return 'Unknown tool'
async def async_agent_loop(question: str, max_turns: int = 5) -> str:
messages = [{'role': 'user', 'content': question}]
for turn in range(max_turns):
response = await client.chat.completions.create(
model='gpt-4o-mini', messages=messages, tools=TOOLS
)
msg = response.choices[0].message
messages.append(msg)
if not msg.tool_calls:
return msg.content
# Execute tool calls in parallel
tool_results = await asyncio.gather(*[
async_tool_call(tc.function.name, json.loads(tc.function.arguments))
for tc in msg.tool_calls
])
for tc, result in zip(msg.tool_calls, tool_results):
messages.append({'role': 'tool', 'tool_call_id': tc.id, 'content': result})
return 'Max turns reached'
result = asyncio.run(async_agent_loop('What is the latest news on AI?'))
print(result)Anulowanie i zwalnianie zasobów
Zadania asynchroniczne można anulować. Należy prawidłowo obsługiwać asyncio.CancelledError, aby zapewnić zwolnienie zasobów po anulowaniu wykonania agenta (np. na żądanie użytkownika lub po przekroczeniu limitu czasu).
import asyncio
async def cancellable_agent(question: str):
try:
print('Agent starting')
await asyncio.sleep(0.5) # Step 1
print('Step 1 done')
await asyncio.sleep(0.5) # Step 2 - may be cancelled here
print('Step 2 done')
return 'Completed'
except asyncio.CancelledError:
print('Agent was cancelled - cleaning up')
# Clean up resources: close connections, log cancellation
raise # Always re-raise CancelledError
finally:
print('Cleanup always runs')
async def run_with_timeout(question: str, timeout: float):
task = asyncio.create_task(cancellable_agent(question))
try:
result = await asyncio.wait_for(task, timeout=timeout)
return result
except asyncio.TimeoutError:
print(f'Agent exceeded {timeout}s timeout')
task.cancel()
return None
# Run with 0.7s timeout (not enough for both steps)
result = asyncio.run(run_with_timeout('test', timeout=0.7))
print('Final result:', result)Testowanie asynchronicznego kodu agenta
Funkcje asynchroniczne agentów należy testować za pomocą pytest-asyncio. Funkcje testowe należy oznaczać dekoratorem @pytest.mark.asyncio, aby uruchamiać je w pętli zdarzeń.
import pytest
import asyncio
from unittest.mock import AsyncMock, patch
# Install: pip install pytest-asyncio
# pytest.ini: [pytest] asyncio_mode = auto
@pytest.mark.asyncio
async def test_async_agent_call():
with patch('openai.AsyncOpenAI') as mock_openai:
mock_client = AsyncMock()
mock_openai.return_value = mock_client
mock_response = AsyncMock()
mock_response.choices[0].message.content = 'Mocked answer'
mock_response.choices[0].message.tool_calls = None
mock_client.chat.completions.create.return_value = mock_response
# Test the async function
result = await async_agent_call('What is Python?')
assert isinstance(result, str)
print('Async test passed')
@pytest.mark.asyncio
async def test_parallel_execution():
start = asyncio.get_event_loop().time()
results = await asyncio.gather(
asyncio.sleep(0.1),
asyncio.sleep(0.1),
asyncio.sleep(0.1)
)
elapsed = asyncio.get_event_loop().time() - start
assert elapsed < 0.3, 'Should complete in parallel'
print(f'Parallel test passed: {elapsed:.2f}s')Sprawdzenie wiedzy: frameworki asynchroniczne
Proszę sprawdzić swoją wiedzę na temat asynchronicznych frameworków agentów.
Podsumowanie asynchronicznych frameworków
Asynchroniczny LangChain udostępnia ainvoke(), astream() i AsyncCallbackHandler do tworzenia gotowych do użycia w produkcji agentów asynchronicznych. LangGraph natywnie obsługuje asynchroniczne funkcje węzłów. Do ograniczania szybkości należy używać semaforów, do bezpośrednich wywołań klienta AsyncOpenAI, a do testowania pytest-asyncio. Prawidłowa obsługa anulowania zapewnia poprawne zwalnianie zasobów po przerwaniu działania agentów.
Często zadawane pytania
Czy lekcja „Asynchroniczne frameworki agentów: LangChain i nie tylko” jest bezpłatna?
Tak — pełny tekst „Asynchroniczne frameworki agentów: LangChain i nie tylko” jest dostępny za darmo tutaj w sieci. Aby ćwiczyć ją interaktywnie (wbudowany edytor kodu i tutor AI dostępny 24/7) i odblokować resztę kursu AI Agents, przejdź na CoddyKit PRO. Kurs AI Agents zawiera 4 lekcji w sumie.
Co nauczysz się w „Asynchroniczne frameworki agentów: LangChain i nie tylko”?
ainvoke(), astream() i asynchroniczne łańcuchy w LangChain i LangGraph. Ćwiczysz AI Agents z praktycznym kodem, który uruchamiasz bezpośrednio w przeglądarce, a tutor AI dostępny 24/7 odpowiada na Twoje pytania podczas pracy nad lekcją.
Czy potrzebuję doświadczenia, aby zacząć AI Agents?
Nie wymagamy żadnego doświadczenia. AI Agents w CoddyKit jest strukturyzowany dla początkujących i zaawansowanych użytkowników, więc możesz zacząć tutaj lub od początku i uczyć się w swoim tempie. To lekcja 4 z 4.
Ile czasu zajmuje lekcja „Asynchroniczne frameworki agentów: LangChain i nie tylko”?
Większość lekcji CoddyKit trwa około 5–10 minut. Każda lekcja to mały, interaktywny krok, dzięki czemu robisz systematyczne postępy i zawsze wracasz dokładnie do tego samego miejsca — na webie i w aplikacji.
Czy mogę pisać i uruchamiać kod w tej lekcji AI Agents?
Tak. Każda lekcja AI Agents zawiera wbudowany edytor kodu, więc piszesz i uruchamiasz prawdziwy kod bezpośrednio w przeglądarce i od razu otrzymujesz sprzężenie zwrotne od AI — bez konfiguracji na komputerze.
Wszystkie lekcje w tym kursie
- Asynchroniczny Python dla twórców agentów
- Kolejki zdarzeń i brokerzy komunikatów
- Nieblokujące równoległe wykonywanie narzędzi
- Asynchroniczne frameworki agentów: LangChain i nie tylko