0Pricing
Mojo Academy · Aula

Otimizando o produto interno

Vetorize e desenrole o produto escalar.

Otimizando o produto interno é uma aula grátis de Mojo Academy no CoddyKit. Esta é a aula 3 de 4. Você pode ler a aula completa abaixo gratuitamente — depois pratica ao vivo no navegador com um editor de código integrado e um tutor de IA 24/7. Faz parte do caminho de aprendizado de Mojo Academy, e seu progresso é sincronizado entre a web e o app CoddyKit. O curso de Mojo Academy inclui 4 aulas no total.

Partes desta aula ainda não foram traduzidas e aparecem em inglês.

The Dot Product Core

The inner product multiplies two sequences element by element and adds the results. Speeding it up speeds your whole matmul.

var acc: Float32 = 0.0
for k in range(K):
    acc += a[k] * b[k]

One Lane at a Time Is Slow

A plain loop handles a single pair per step. Modern CPUs can do many at once, so scalar code leaves lanes idle.

Enter SIMD

A SIMD value packs several numbers and applies one operation to all of them together, doing real work in parallel per instruction.

var v = SIMD[DType.float32, 4](1, 2, 3, 4)

Load a Chunk at Once

Instead of one float, load a whole SIMD-width slice of each array in a single move, ready for vector math.

var va = a.load[width=4](k)
var vb = b.load[width=4](k)

Multiply Whole Vectors

Multiplying two SIMD values yields a vector of products in one step. The element-wise work happens across all lanes together.

var prod = va * vb  # 4 products at once

Accumulate Into a Vector

Keep a SIMD accumulator and add each product vector into it. You delay the final sum until the loop is done.

var acc = SIMD[DType.float32, 4](0)
acc += va * vb

Step by the Width

The loop now advances by the SIMD width, not by one. Four lanes wide means a quarter as many iterations.

for k in range(0, K, 4):
    acc += a.load[width=4](k) * b.load[width=4](k)

Reduce at the End

After the loop, collapse the lane accumulator into one number with a horizontal reduce, giving the final dot product.

var total = acc.reduce_add()

Unrolling the Loop

Unrolling processes several widths per iteration, cutting loop overhead and exposing more independent work to the CPU.

Handle the Remainder

If K is not a multiple of the width, a few elements are left over. A small tail loop finishes them one at a time.

for k in range(K - K % 4, K):
    total += a[k] * b[k]

Let Mojo Help

Mojo offers the vectorize helper to apply a width-parameterized body across a range, handling stepping and the tail for you.

Quick Check

After a SIMD-width dot product loop, why do you call reduce_add at the end?

Recap

Vectorize the inner product by loading SIMD chunks, multiplying lanes together, accumulating, then reduce; unroll and clean up the tail. ⚡

Perguntas Frequentes

A aula “Otimizando o produto interno” é grátis?

Sim — o texto completo de “Otimizando o produto interno” é grátis para ler aqui na web. Para praticá-la interativamente (um editor de código integrado e um tutor de IA 24/7) e desbloquear o restante do curso de Mojo Academy, atualize para CoddyKit PRO. O curso de Mojo Academy inclui 4 aulas no total.

O que vou aprender em “Otimizando o produto interno”?

Vetorize e desenrole o produto escalar. Você pratica Mojo Academy com código prático que executa diretamente no navegador, e um tutor de IA 24/7 responde suas dúvidas enquanto trabalha na aula.

Preciso ter experiência prévia para começar Mojo Academy?

Nenhuma experiência prévia é necessária. Mojo Academy no CoddyKit é estruturado para alunos iniciantes até avançados, então você pode começar aqui ou desde o início e aprender no seu ritmo. Esta é a aula 3 de 4.

Quanto tempo leva a aula “Otimizando o produto interno”?

A maioria das aulas CoddyKit leva cerca de 5–10 minutos. Cada uma é compacta e interativa, então você faz progresso constante e retoma exatamente de onde parou entre web e app.

Posso escrever e executar código nesta aula de Mojo Academy?

Sim. Cada aula de Mojo Academy inclui um editor de código integrado, então você escreve e executa código real direto no navegador e recebe feedback de IA instantaneamente — nenhuma configuração local necessária.

Todas as aulas deste curso

  1. Modelando um tensor em Mojo
  2. Construindo uma multiplicação de matrizes passo a passo
  3. Otimizando o produto interno
  4. Verificando a correção numérica
← Voltar para Mojo Academy