0Pricing
AI Prompt Engineering · Lesson

Top-k Sampling

Limiting choice to k most probable tokens and its effect on output diversity.

Top-k Sampling is a free AI Prompt Engineering lesson on CoddyKit. This is lesson 3 of 4. You can read the complete lesson below for free — then practise it hands-on in the browser with a built-in code editor and a 24/7 AI tutor. It is part of the AI Prompt Engineering learning path, and your progress syncs across the web and the CoddyKit app. The AI Prompt Engineering course includes 4 lessons in total.

What Is Top-k Sampling?

Top-k sampling restricts the model to sampling from the k most probable tokens at each step. All tokens outside the top-k are assigned zero probability and cannot be selected.

k=1 is greedy decoding (only the single most probable token). k=50 is a typical creative range. k=vocabulary_size is equivalent to unrestricted sampling.

Top-k Algorithm

The algorithm is simpler than top-p:

  1. Compute softmax probabilities over the full vocabulary
  2. Sort tokens by probability descending
  3. Keep only the top k tokens; set all others to 0
  4. Renormalize the top-k probabilities to sum to 1
  5. Sample from the renormalized distribution
import numpy as np

def softmax(logits, temperature=1.0):
    scaled = logits / temperature
    e = np.exp(scaled - np.max(scaled))
    return e / e.sum()

def top_k_sample(logits, k=50, temperature=1.0):
    probs = softmax(logits, temperature)

    # Find top-k indices
    top_k_indices = np.argsort(probs)[::-1][:k]
    top_k_probs = probs[top_k_indices]

    # Renormalize
    top_k_probs = top_k_probs / top_k_probs.sum()

    # Sample
    chosen = np.random.choice(top_k_indices, p=top_k_probs)
    return chosen

# With a 10-token vocabulary:
logits = np.random.randn(10)
print('Chosen token:', top_k_sample(logits, k=3))

k=1: Greedy Decoding

When k=1, only the single most probable token is in the candidate set. Sampling from a set of size 1 is deterministic — the model always picks that token. This is identical to greedy decoding (temperature=0).

logits = np.array([3.0, 2.0, 1.0, 0.5, -1.0])

# k=1: greedy
top_1 = top_k_sample(logits, k=1)
print(f'k=1 always picks: {np.argmax(logits)}')  # index 0, the highest logit
print(f'top_k_sample result: {top_1}')  # always 0

# Multiple runs
for _ in range(5):
    print(top_k_sample(logits, k=1), end=' ')
# Output: 0 0 0 0 0 — perfectly deterministic

Typical Creative Range: k=50

k=50 is a common default for creative text generation. It allows exploration across 50 tokens at each step while preventing the model from selecting tokens it has very low confidence in.

With a vocabulary of 50,000 tokens, k=50 means the model considers only the top 0.1% of tokens at each step. This is a significant restriction — most of the vocabulary is excluded.

import openai
client = openai.OpenAI(api_key='sk-...')

# Note: OpenAI API does not expose top_k directly in chat completions.
# Top-k is primarily a parameter in Hugging Face Transformers and Anthropic's API.

# Hugging Face example:
from transformers import pipeline
generator = pipeline('text-generation', model='gpt2')
output = generator(
    'Once upon a time',
    max_new_tokens=100,
    do_sample=True,
    top_k=50,
    temperature=1.0
)
print(output[0]['generated_text'])

Top-k in Anthropic's Claude API

Anthropic's Claude API exposes top_k as a direct parameter. This makes it easy to experiment with the effect of fixed vocabulary truncation on Claude's outputs.

import anthropic
claude = anthropic.Anthropic(api_key='sk-ant-...')

# top_k limits the number of tokens considered
message = claude.messages.create(
    model='claude-opus-4-5',
    max_tokens=256,
    temperature=1.0,
    top_k=50,  # sample from top 50 most probable tokens
    messages=[{
        'role': 'user',
        'content': 'Write a short poem about debugging code.'
    }]
)
print(message.content[0].text)

The Fixed-k Problem

The fundamental limitation of top-k: k is fixed regardless of the model's confidence at that step.

When the model is very confident (one token has 95% probability), k=50 still includes 49 largely irrelevant tokens. When the model is very uncertain (50 tokens each have ~2% probability), k=50 may actually be appropriate.

The problem: k=50 can be simultaneously too restrictive and too permissive depending on context. Top-p addresses this with dynamic nucleus sizing.

# Illustrating the fixed-k problem
logits_confident = np.array([5.0] + [0.1] * 9)  # model is very sure
logits_uncertain = np.array([1.0] * 10)           # model has no idea

probs_conf = softmax(logits_confident)
probs_unc = softmax(logits_uncertain)

print('Confident — top 3 tokens cover:', np.sort(probs_conf)[::-1][:3].sum().round(3))
# ~0.998 — k=50 is extremely wasteful, includes near-zero probability tokens

print('Uncertain — top 3 tokens cover:', np.sort(probs_unc)[::-1][:3].sum().round(3))
# ~0.30 — k=50 may actually be needed to cover a reasonable nucleus

Top-k at Distribution Tails

Top-k and top-p differ most significantly at the distribution tails:

  • With top-k=50, the 50th token might have probability 0.001% (extremely unlikely but still in the candidate set)
  • With top-p=0.9, any token outside the 90% nucleus is excluded — including tokens that would be in top-k

Top-p is more principled about tail behavior: it excludes improbable tokens based on probability, not position in the ranking.

# Tail behavior comparison
import numpy as np

# Highly skewed distribution (one dominant token)
skewed_logits = np.array([4.0, 2.0, 1.5, 1.0, 0.5,
                           0.1, 0.0, -0.1, -0.5, -1.0])
probs = softmax(skewed_logits)

print('Probability of tokens 6-9 (tail):')
for i in range(6, 10):
    print(f'  Token {i}: {probs[i]:.4%}')
# These tokens are very unlikely but are included in top-k=10
# Top-p=0.9 would exclude them entirely

Combining Top-k and Top-p

Some implementations apply both top-k and top-p: first truncate to top-k, then apply top-p nucleus sampling within that set. This provides a hard cap on vocabulary size (top-k) while also applying probability-based filtering (top-p).

def top_k_top_p_sample(logits, k=50, p=0.9, temperature=1.0):
    probs = softmax(logits, temperature)

    # First apply top-k
    top_k_indices = np.argsort(probs)[::-1][:k]
    top_k_probs = probs[top_k_indices]

    # Then apply top-p within top-k
    sorted_k = np.sort(top_k_probs)[::-1]
    cumulative = np.cumsum(sorted_k)
    nucleus_size = np.searchsorted(cumulative, p) + 1

    final_indices = top_k_indices[:nucleus_size]
    final_probs = top_k_probs[:nucleus_size]
    final_probs = final_probs / final_probs.sum()

    return np.random.choice(final_indices, p=final_probs)

When Top-k Is Preferred Over Top-p

Despite top-p's theoretical advantages, top-k is preferred in some scenarios:

  • Constrained vocabulary tasks: when the model should only output from a fixed set (e.g., multiple choice A/B/C/D), a small top-k (4) directly enforces this
  • Reproducibility: top-k behavior is easier to reason about — 'always consider exactly 50 tokens'
  • Hardware-optimized implementations: some inference engines implement top-k more efficiently than top-p
# Constrained output with top-k=4
# For a multiple choice task (A, B, C, D)
# If A/B/C/D tokens have indices 32, 33, 34, 35
# top-k=4 with those as the top-4 logits forces selection from those 4 only

multiple_choice_prompt = (
    'Answer with only A, B, C, or D.\n'
    'What is the capital of France?\n'
    'A) Berlin\n'
    'B) Paris\n'
    'C) Rome\n'
    'D) Madrid\n'
    'Answer:'
)

# With temperature=0, top_k=1: always picks the highest logit token

Practical Top-k Guidelines

When to use top-k and what values to choose:

  • k=1: greedy / factual tasks
  • k=5–20: focused creative tasks, minimal variation
  • k=40–100: standard creative range in most language models
  • k=500+: very open exploration (rarely needed; use top-p instead)

In most modern LLM APIs, top-p is the preferred parameter. Use top-k when you need a hard vocabulary cap or when the API exposes it but not top-p.

TOP_K_GUIDELINES = {
    'factual_qa': 1,           # greedy
    'code_generation': 10,     # near-greedy, correct syntax
    'summarization': 20,       # slightly varied but focused
    'chat': 50,                # natural variation
    'creative_writing': 100,   # wider vocabulary exploration
    'poetry': 200,             # unusual word choices encouraged
}

def call_with_top_k(task, prompt, model='claude-opus-4-5'):
    k = TOP_K_GUIDELINES.get(task, 50)
    claude = anthropic.Anthropic(api_key='sk-ant-...')
    return claude.messages.create(
        model=model,
        max_tokens=512,
        top_k=k,
        messages=[{'role': 'user', 'content': prompt}]
    )

Top-k and Temperature Together

Top-k and temperature are applied in sequence: temperature first reshapes the logit distribution, then top-k truncates it to the k most probable tokens. Using both together is common in Hugging Face Transformers pipelines.

Typical combination: temperature=0.9 (moderate diversity) + top_k=50 (hard vocabulary cap). This avoids both the flatness of high temperature alone and the tail-sampling problem.

from transformers import pipeline

generator = pipeline('text-generation', model='gpt2')

# Combined top-k + temperature
output = generator(
    'The future of AI is',
    max_new_tokens=80,
    do_sample=True,
    top_k=50,
    temperature=0.9
)
print(output[0]['generated_text'])

# Compare: top-k=1 (greedy)
greedy_output = generator(
    'The future of AI is',
    max_new_tokens=80,
    do_sample=False  # greedy, equivalent to top_k=1
)
print(greedy_output[0]['generated_text'])

Knowledge Check

In top-k sampling, if k=1 is set, which behavior does the model exhibit?

Recap: Top-k Sampling

Top-k sampling truncates the vocabulary to the k most probable tokens before sampling:

  • k=1: greedy decoding — deterministic, always picks the top token
  • k=50: typical creative range — balanced diversity and coherence
  • Key limitation: k is fixed regardless of model confidence — can be too permissive or too restrictive
  • Vs top-p: top-p's dynamic nucleus adapts to confidence; top-k does not

Use top-k when you need a hard vocabulary cap. Prefer top-p for most production applications. Next lesson: choosing parameters for your specific use case.

Frequently Asked Questions

Is the “Top-k Sampling” lesson free?

Yes — the full text of “Top-k Sampling” is free to read here on the web. To practise it interactively (a built-in code editor and a 24/7 AI tutor) and unlock the rest of the AI Prompt Engineering course, upgrade to CoddyKit PRO. The AI Prompt Engineering course includes 4 lessons in total.

What will I learn in “Top-k Sampling”?

Limiting choice to k most probable tokens and its effect on output diversity. You practise AI Prompt Engineering with hands-on code you run directly in the browser, and a 24/7 AI tutor answers your questions as you work through the lesson.

Do I need any experience to start AI Prompt Engineering?

No prior experience is required. AI Prompt Engineering on CoddyKit is structured for beginners through advanced learners, so you can start here or from the beginning and move at your own pace. This is lesson 3 of 4.

How long does the “Top-k Sampling” lesson take?

Most CoddyKit lessons take about 5–10 minutes. Each one is bite-sized and interactive, so you make steady progress and pick up exactly where you left off across the web and the app.

Can I write and run code in this AI Prompt Engineering lesson?

Yes. Every AI Prompt Engineering lesson includes a built-in code editor, so you write and run real code right in your browser and get instant AI feedback — no local setup required.

All lessons in this course

  1. What Is Temperature in LLMs?
  2. Top-p Nucleus Sampling
  3. Top-k Sampling
  4. Choosing Parameters for Your Use Case
← Back to AI Prompt Engineering