0Pricing
Mojo Academy · Aula

Combinando SIMD com laços

Vetorizе o núcleo do kernel.

Combinando SIMD com laços é uma aula grátis de Mojo Academy no CoddyKit. Esta é a aula 2 de 4. Você pode ler a aula completa abaixo gratuitamente — depois pratica ao vivo no navegador com um editor de código integrado e um tutor de IA 24/7. Faz parte do caminho de aprendizado de Mojo Academy, e seu progresso é sincronizado entre a web e o app CoddyKit. O curso de Mojo Academy inclui 4 aulas no total.

Partes desta aula ainda não foram traduzidas e aparecem em inglês.

From Scalar to SIMD

To accelerate a kernel, swap the one-at-a-time loop for one that handles a whole pack of values per step with SIMD.

for i in range(n):
    out[i] = a[i] + b[i]

Pick a Width

Choose how many elements fit in one vector. The width is the number of lanes each step processes together.

alias width = 4

Load a Chunk

Grab several elements at once into a SIMD value with a vector load instead of reading them one by one.

var va = a.load[width=width](i)

Compute the Pack

Run the kernel's math on both chunks at once. The whole pack is added element-wise in a single operation.

var vsum = a.load[width=width](i) + b.load[width=width](i)

Store the Pack

Write the full result back with a vector store, covering every lane you just computed in one move.

out.store[width=width](i, vsum)

Step by the Width

The vectorized loop advances by the width, not by one. Each pass covers a full pack of elements.

for i in range(0, n, width):
    pass

Let vectorize Help

Mojo's vectorize helper sweeps a closure across the range in vector steps so you skip the manual bookkeeping.

from algorithm import vectorize

Write the Chunk Closure

You define a small parameterized function that handles one chunk. Mojo calls it with the right width as it sweeps.

fn body[w: Int](i: Int):
    out.store[width=w](i, a.load[width=w](i) + b.load[width=w](i))

Run vectorize

Call vectorize with your closure, the width, and the size. It loops and even handles the leftover tail for you.

vectorize[body, width](n)

Mind the Tail

When n is not a multiple of the width, a few elements remain. vectorize cleans up that tail so nothing is missed.

Same Output, More Speed

The vectorized kernel produces identical results but moves through data in big steps, so it finishes much sooner.

Quick Check

You vectorize a kernel with width 4 but n is 10. What handles the last two elements?

Recap

Vectorize a kernel by loading and storing packs, stepping by the width, and letting vectorize handle the sweep and the tail. 🚀

Perguntas Frequentes

A aula “Combinando SIMD com laços” é grátis?

Sim — o texto completo de “Combinando SIMD com laços” é grátis para ler aqui na web. Para praticá-la interativamente (um editor de código integrado e um tutor de IA 24/7) e desbloquear o restante do curso de Mojo Academy, atualize para CoddyKit PRO. O curso de Mojo Academy inclui 4 aulas no total.

O que vou aprender em “Combinando SIMD com laços”?

Vetorizе o núcleo do kernel. Você pratica Mojo Academy com código prático que executa diretamente no navegador, e um tutor de IA 24/7 responde suas dúvidas enquanto trabalha na aula.

Preciso ter experiência prévia para começar Mojo Academy?

Nenhuma experiência prévia é necessária. Mojo Academy no CoddyKit é estruturado para alunos iniciantes até avançados, então você pode começar aqui ou desde o início e aprender no seu ritmo. Esta é a aula 2 de 4.

Quanto tempo leva a aula “Combinando SIMD com laços”?

A maioria das aulas CoddyKit leva cerca de 5–10 minutos. Cada uma é compacta e interativa, então você faz progresso constante e retoma exatamente de onde parou entre web e app.

Posso escrever e executar código nesta aula de Mojo Academy?

Sim. Cada aula de Mojo Academy inclui um editor de código integrado, então você escreve e executa código real direto no navegador e recebe feedback de IA instantaneamente — nenhuma configuração local necessária.

Todas as aulas deste curso

  1. Anatomia de um núcleo de computação
  2. Combinando SIMD com laços
  3. Reduzindo o tráfego de memória
  4. Ladrilhamento para localidade de cache
← Voltar para Mojo Academy