0Pricing
AI Engineering Academy · 课时

使用参数控制模型行为

尝试调整 temperature、max_tokens 和 top_p,观察它们如何改变输出风格、长度和创造性,然后为您的使用场景选择合适的设置。

使用参数控制模型行为 是 CoddyKit 上的免费 AI Engineering Academy 课时。 这是第 3 节课,共 4 节。 你可以在下方免费阅读本课时的完整内容 — 然后在浏览器中使用内置代码编辑器和全天候 AI 导师进行实践。 这是 AI Engineering Academy 学习路径的一部分,你的进度在网页和 CoddyKit 应用中同步。 AI Engineering Academy 课程共包含 4 节课。

本课时的部分内容尚未翻译,以英文显示。

The Core Parameters That Matter

A handful of parameters shape your output the most: temperature, max_tokens, top_p, frequency_penalty, and presence_penalty. Master these five and you control the model.

Temperature: Controlling Randomness

Temperature controls randomness. Near 0, the model picks the safest word and stays consistent — great for facts. Around 0.7-1.0, it gets varied and creative. See the code.

from openai import OpenAI
client = OpenAI()

for temp in [0.0, 0.7, 1.5]:
    response = client.chat.completions.create(
        model='gpt-4o-mini',
        messages=[{'role': 'user', 'content': 'Name a color.'}],
        temperature=temp,
        max_tokens=5
    )
    print(f'Temp {temp}: {response.choices[0].message.content}')
# Temp 0.0: Red       (always most common)
# Temp 0.7: Blue      (varied but sensible)
# Temp 1.5: Vermillion (surprising choices)

max_tokens: Controlling Response Length

max_tokens caps how long the reply can get — a safety limit, not a target. Too low and answers get cut off; too high wastes money and time. Add a buffer and watch finish_reason.

# Different max_tokens for different use cases

# Classification: short answer expected
classification_response = client.chat.completions.create(
    model='gpt-4o-mini',
    messages=[{'role': 'user', 'content': 'Is this positive or negative? "Great product!"'}],
    max_tokens=5  # Only need 1-2 words
)

# Detailed analysis: longer output needed
analysis_response = client.chat.completions.create(
    model='gpt-4o-mini',
    messages=[{'role': 'user', 'content': 'Analyze the pros and cons of microservices.'}],
    max_tokens=800  # Need space for detailed explanation
)

top_p: Nucleus Sampling

top_p is another randomness dial: it samples only from the top tokens that add up to top_p of the probability. Tip: tune either temperature or top_p, not both at once.

# top_p usage example
response_narrow = client.chat.completions.create(
    model='gpt-4o-mini',
    messages=[{'role': 'user', 'content': 'Continue: The sky is...'}],
    top_p=0.1,  # only very likely tokens (conservative, predictable)
    temperature=1.0  # keep temperature at 1 when using top_p
)

response_wide = client.chat.completions.create(
    model='gpt-4o-mini',
    messages=[{'role': 'user', 'content': 'Continue: The sky is...'}],
    top_p=0.95,  # most tokens eligible (creative, varied)
    temperature=1.0
)

Frequency Penalty: Reducing Repetition

frequency_penalty discourages the model from repeating the same words, scaling with how often they appear. Reach for it when long answers get repetitive — try 0.3 to 0.7.

# Frequency penalty to reduce repetition in long outputs
response = client.chat.completions.create(
    model='gpt-4o-mini',
    messages=[{
        'role': 'user',
        'content': 'Write 5 tips for better sleep.'
    }],
    max_tokens=300,
    frequency_penalty=0.5  # reduces repeating the same words/phrases
)
print(response.choices[0].message.content)

Presence Penalty: Encouraging Topic Diversity

presence_penalty nudges the model toward new topics: it penalizes any word that already appeared, even once. Great for brainstorming where you want fresh, varied ideas.

# Presence penalty for diverse brainstorming output
response = client.chat.completions.create(
    model='gpt-4o-mini',
    messages=[{
        'role': 'user',
        'content': 'List 10 creative ways to use AI in a small business.'
    }],
    max_tokens=400,
    presence_penalty=0.8  # encourages introducing different topics per item
)
print(response.choices[0].message.content)

stop Sequences: Custom Stopping Points

The stop parameter lists strings that halt generation the moment they appear (and they're left out). Stop at a newline to grab a clean single-line answer. See the code.

# Use stop sequences to get clean single-line output
response = client.chat.completions.create(
    model='gpt-4o-mini',
    messages=[{
        'role': 'user',
        'content': 'What is the Python keyword for a function definition?\nAnswer:'
    }],
    max_tokens=20,
    stop=['\n', '.']  # stop at newline or period - gets just the keyword
)
print(repr(response.choices[0].message.content))  # 'def'

seed: Reproducible Outputs

The seed parameter makes output reproducible: same seed plus temperature 0 gives the same reply. Great for testing — just watch the system_fingerprint for backend changes.

# Reproducible output with seed parameter
response1 = client.chat.completions.create(
    model='gpt-4o-mini',
    messages=[{'role': 'user', 'content': 'Pick a random number from 1 to 10.'}],
    temperature=0,
    seed=42
)

response2 = client.chat.completions.create(
    model='gpt-4o-mini',
    messages=[{'role': 'user', 'content': 'Pick a random number from 1 to 10.'}],
    temperature=0,
    seed=42
)

print(response1.choices[0].message.content)  # same
print(response2.choices[0].message.content)  # same
print('Fingerprint:', response1.system_fingerprint)

Choosing Parameters for Common Use Cases

Skip the guesswork with quick presets: temperature 0 for classification, 0.2 for factual Q&A, 0.8-1.0 for creative writing. Start there, then adjust to your task. See the code.

# Parameter presets for different task types
PRESETS = {
    'classify': {'temperature': 0, 'max_tokens': 20},
    'factual_qa': {'temperature': 0.2, 'max_tokens': 400},
    'creative': {'temperature': 0.9, 'max_tokens': 1000, 'frequency_penalty': 0.3},
    'code': {'temperature': 0.1, 'max_tokens': 2000},
    'summary': {'temperature': 0.3, 'max_tokens': 300, 'frequency_penalty': 0.2},
}

def complete(prompt, task_type='factual_qa', **overrides):
    params = {**PRESETS[task_type], **overrides}
    return client.chat.completions.create(
        model='gpt-4o-mini',
        messages=[{'role': 'user', 'content': prompt}],
        **params
    )

Logprobs: Understanding Model Confidence

logprobs returns how confident the model was in each token, plus alternatives it weighed. Low confidence on a fact is a hallucination warning — useful for safer systems.

import math

response = client.chat.completions.create(
    model='gpt-4o-mini',
    messages=[{'role': 'user', 'content': 'The capital of France is?'}],
    max_tokens=3,
    logprobs=True,
    top_logprobs=3  # show top 3 alternative tokens at each position
)

for token_log in response.choices[0].logprobs.content:
    prob = math.exp(token_log.logprob)  # convert log prob to probability
    print(f'Token: {token_log.token!r} | Probability: {prob:.2%}')
    for alt in token_log.top_logprobs:
        print(f'  Alt: {alt.token!r} -> {math.exp(alt.logprob):.2%}')

Testing Parameter Effects Systematically

Don't guess parameters — test them. Build a small grid search that runs the same prompt across combos and scores each. That turns tuning from art into engineering. See the code.

from itertools import product

# Systematic parameter grid search
temperatures = [0.0, 0.3, 0.7]
max_tokens_options = [100, 300]
test_prompt = 'Summarize the benefits of unit testing in 2 sentences.'

results = []
for temp, max_tok in product(temperatures, max_tokens_options):
    resp = client.chat.completions.create(
        model='gpt-4o-mini',
        messages=[{'role': 'user', 'content': test_prompt}],
        temperature=temp,
        max_tokens=max_tok
    )
    results.append({
        'temperature': temp,
        'max_tokens': max_tok,
        'output': resp.choices[0].message.content,
        'actual_tokens': resp.usage.completion_tokens
    })

Quick Check

Test your understanding of AI Engineering concepts from this lesson.

Lesson Recap

You learned to steer the model: temperature sets randomness, max_tokens caps length (watch finish_reason!), and the penalties cut repetition and add variety. Next: error handling.

常见问题解答

「使用参数控制模型行为」课时是免费的吗?

是的 — 「使用参数控制模型行为」的完整文本可在网页上免费阅读。要进行交互式练习(内置代码编辑器和全天候 AI 导师)并解锁 AI Engineering Academy 课程的其余内容,请升级到 CoddyKit PRO。 AI Engineering Academy 课程共包含 4 节课。

「使用参数控制模型行为」这节课中我会学到什么?

尝试调整 temperature、max_tokens 和 top_p,观察它们如何改变输出风格、长度和创造性,然后为您的使用场景选择合适的设置。 你通过在浏览器中直接运行的动手代码来练习 AI Engineering Academy,全天候 AI 导师会在你学习这节课的过程中回答你的问题。

学习 AI Engineering Academy 需要有经验吗?

无需任何先前经验。CoddyKit 上的 AI Engineering Academy 课程适合初学者到高级学习者,你可以从这里开始或从头开始,按照自己的节奏学习。 这是第 3 节课,共 4 节。

「使用参数控制模型行为」课时需要多长时间?

大多数 CoddyKit 课程大约需要 5–10 分钟。每节课都很精短且互动,所以你能稳步进步,并在网页和应用中从离开的地方继续。

我能在这节 AI Engineering Academy 课中编写并运行代码吗?

能。每节 AI Engineering Academy 课都包含内置代码编辑器,你可以在浏览器中直接编写并运行真实代码,并获得即时 AI 反馈 — 无需本地设置。

此课程中的所有课时

  1. 设置 Python 环境
  2. 聊天补全端点
  3. 使用参数控制模型行为
  4. 错误处理与速率限制
← 返回 AI Engineering Academy