AI एजेंट · पाठ

आउटपुट फ़िल्टरिंग (Llama Guard, NeMo)

आउटपुट पर एक छोटा गार्ड मॉडल चलाएँ, ताकि जारी करने से पहले विषाक्तता, PII लीक और नीति उल्लंघन पकड़े जा सकें।

पाठ 2, कुल 4 में से15 चरण

आउटपुट फ़िल्टरिंग (Llama Guard, NeMo), CoddyKit पर AI एजेंट का एक निःशुल्क पाठ है। यह 4 में से 2वाँ पाठ है। आप नीचे पूरा पाठ निःशुल्क पढ़ सकते हैं—फिर अंतर्निहित कोड संपादक और 24/7 एआई ट्यूटर के साथ ब्राउज़र में इसका व्यावहारिक अभ्यास कर सकते हैं। यह AI एजेंट सीखने के मार्ग का हिस्सा है और आपकी प्रगति वेब तथा CoddyKit ऐप पर सिंक होती रहती है। AI एजेंट पाठ्यक्रम में कुल 4 पाठ शामिल हैं।

Outputs को Filter क्यों करें?

Safe inputs के बावजूद models ये outputs दे सकते हैं:

  • Personal data leaks
  • Hate speech / harassment
  • Self-harm content
  • ऐसे tool calls जो User के intent का उल्लंघन करें

Output filter, User को कुछ भी दिखने से पहले आपकी अंतिम सुरक्षा-पंक्ति है।

Llama Guard

Meta का safety classifier — open weights और बहुत तेज़:

from transformers import AutoTokenizer, AutoModelForCausalLM
model = AutoModelForCausalLM.from_pretrained('meta-llama/Llama-Guard-3-8B')
# Outputs 'safe' or 'unsafe' with category.

NeMo Guardrails

NVIDIA NeMo Guardrails — Python framework जो LLMs को checks से wrap करता है:

# pip install nemoguardrails
from nemoguardrails import RailsConfig, LLMRails

config = RailsConfig.from_path('./config')
rails = LLMRails(config)
response = rails.generate(messages=[{'role': 'user', 'content': '...'}])

Guardrails AI

guardrails-ai validation/guards पर केंद्रित Python library है। यह XML-आधारित "rails" और reAsk repair loops का समर्थन करती है।

Anthropic Constitutional Classifier

Anthropic ने एक classifier जारी किया है (Claude में मौजूद safety training के अतिरिक्त)। यह किसी भी model output की post-hoc filtering के लिए उपयोगी है।

Custom LLM-Based Filter

सबसे सस्ता उपाय: outputs की जाँच करने वाला छोटा LLM:

FILTER_PROMPT = '''
Analyse the output. Return JSON: {safe: true/false, reason: str, categories: list}

Unsafe categories: PII leak, harmful instructions, copyright violation, prompt-injection echo.

Output:
{output}
'''

result = guard_llm.invoke(FILTER_PROMPT.format(output=text))

Regex / Rule Filters

विशिष्ट patterns के लिए regex तेज़ और विश्वसनीय है:

import re
EMAIL_RE = re.compile(r'[\w\.-]+@[\w\.-]+')
SSN_RE = re.compile(r'\b\d{3}-\d{2}-\d{4}\b')

def has_pii(text):
    return bool(EMAIL_RE.search(text) or SSN_RE.search(text))

# --- demo ---
samples = ['Contact me at jane@example.com', 'My SSN is 123-45-6789', 'No PII in this sentence']
for s in samples:
    print(f'{s!r} -> has_pii={has_pii(s)}')

PII Redaction

केवल block न कीजिए — कभी-कभी PII redact करके बाकी content को आगे जाने दीजिए:

def redact(text):
    text = EMAIL_RE.sub('[email]', text)
    text = SSN_RE.sub('[ssn]', text)
    return text

दो-स्तरीय Filter

पहले तेज़ rules, फिर महँगा LLM:

def filter_output(text):
    if has_obvious_pii(text):
        return BLOCKED
    result = guard_llm.invoke(...)
    return result

Streaming Outputs

Stream किए गए outputs के लिए SENTENCE-BY-SENTENCE filter कीजिए:

buf = ''
for token in stream:
    buf += token
    if buf.endswith(('. ', '! ', '? ', '\n')):
        if not safe(buf):
            stream.close()
            emit_to_client('[content filtered]')
            break
        emit_to_client(buf)
        buf = ''

False Positives बनाम False Negatives

सही संतुलन के लिए इसे tune कीजिए:

  • High-stakes (medical, financial): false positives की ओर झुकिए (अधिक block कीजिए)
  • Creative/entertainment: false negatives की ओर झुकिए (कम block कीजिए)

Filter Logs संवेदनशील होते हैं

Filter logs में blocked content होता है। इन्हें कम TTLs के साथ सुरक्षित रूप से store कीजिए:

log.warning('Blocked output', extra={'reason': r, 'category': c}, sanitize_payload=True)

User को समझाइए

जब आप किसी content को block करें, तो User को कारण बताइए ताकि वह बार-बार retry न करे:

def handle_privacy_request():
    return 'I cannot share that information for privacy reasons.'

print(handle_privacy_request())

Multi-Layer Filter

Regex को LLM-आधारित filtering के साथ क्यों मिलाएँ?

पुनरावलोकन

Llama Guard / NeMo / custom LLM filter + regex rules। दो layers का उपयोग कीजिए। जहाँ संभव हो redact कीजिए। Content block होने पर User को हमेशा बताइए।

शुरुआत निःशुल्क

एआई शिक्षक के साथ AI एजेंट सीखें — निःशुल्क

अपने ब्राउज़र में वास्तविक कोड लिखें और चलाएँ, चौबीसों घंटे एआई शिक्षक से तुरंत सहायता पाएँ, और वेब या ऐप पर वहीं से शुरू करें जहाँ आपने छोड़ा था।

पाठ्यक्रम
60
पाठ
239

अक्सर पूछे जाने वाले प्रश्न

क्या “आउटपुट फ़िल्टरिंग (Llama Guard, NeMo)” पाठ निःशुल्क है?

हाँ—“आउटपुट फ़िल्टरिंग (Llama Guard, NeMo)” का पूरा पाठ यहाँ वेब पर निःशुल्क पढ़ा जा सकता है। इंटरैक्टिव अभ्यास (अंतर्निहित कोड संपादक और 24/7 एआई ट्यूटर) करने और AI एजेंट पाठ्यक्रम का बाकी हिस्सा अनलॉक करने के लिए CoddyKit PRO लें। AI एजेंट पाठ्यक्रम में कुल 4 पाठ शामिल हैं।

“आउटपुट फ़िल्टरिंग (Llama Guard, NeMo)” में मैं क्या सीखूँगा?

आउटपुट पर एक छोटा गार्ड मॉडल चलाएँ, ताकि जारी करने से पहले विषाक्तता, PII लीक और नीति उल्लंघन पकड़े जा सकें। आप ब्राउज़र में सीधे चलाए जाने वाले व्यावहारिक कोड के साथ AI एजेंट का अभ्यास करते हैं, और पाठ पूरा करते समय 24/7 एआई ट्यूटर आपके प्रश्नों के उत्तर देता है।

क्या AI एजेंट शुरू करने के लिए मुझे किसी अनुभव की आवश्यकता है?

पहले के अनुभव की आवश्यकता नहीं है। CoddyKit पर AI एजेंट शुरुआती से लेकर उन्नत शिक्षार्थियों तक सभी के लिए व्यवस्थित किया गया है, इसलिए आप यहीं से या शुरुआत से सीखना शुरू कर सकते हैं और अपनी गति से आगे बढ़ सकते हैं। यह 4 में से 2वाँ पाठ है।

“आउटपुट फ़िल्टरिंग (Llama Guard, NeMo)” पाठ पूरा करने में कितना समय लगता है?

CoddyKit का अधिकांश पाठ लगभग 5–10 मिनट में पूरा हो जाता है। हर पाठ छोटा और संवादात्मक है, इसलिए आप लगातार प्रगति करते हैं और वेब या ऐप पर वहीं से सीखना जारी रख सकते हैं जहाँ आपने छोड़ा था।

क्या मैं इस AI एजेंट पाठ में कोड लिख और चला सकता हूँ?

हाँ। हर AI एजेंट पाठ में एक अंतर्निर्मित कोड संपादक शामिल है, जिससे आप सीधे अपने ब्राउज़र में वास्तविक कोड लिख और चला सकते हैं और तुरंत एआई प्रतिक्रिया पा सकते हैं—स्थानीय सेटअप की आवश्यकता नहीं है।

इस पाठ्यक्रम के सभी पाठ

  1. प्रॉम्प्ट इंजेक्शन से बचाव
  2. आउटपुट फ़िल्टरिंग (Llama Guard, NeMo)
  3. कोड एजेंटों के लिए सैंडबॉक्स निष्पादन
  4. उपकरणों पर पहुँच नियंत्रण
← AI एजेंट पर वापस जाएँ