Reducir el tráfico de memoria
Mantenga los datos cerca del cálculo.
Reducir el tráfico de memoria es una lección gratuita de Mojo Academy en CoddyKit. Esta es la lección 3 de 4. Puedes leer la lección completa abajo gratuitamente — luego la practicas en el navegador con un editor de código integrado y un tutor de IA 24/7. Forma parte de la ruta de aprendizaje de Mojo Academy, y tu progreso se sincroniza en la web y la app de CoddyKit. El curso de Mojo Academy incluye 4 lecciones en total.
Partes de esta lección aún no han sido traducidas y se muestran en inglés.
Why Memory Matters
Modern CPUs compute far faster than they can fetch data. Often a kernel waits on memory, not on the math itself.
What Is Memory Traffic?
Memory traffic is the total bytes your kernel reads and writes. Less traffic per result usually means a faster kernel.
Touch Data Once
Reading the same value many times wastes bandwidth. Load it once, do all the work, then reuse it from a register.
var x = a[i]
var y = x * x + xFuse Your Loops
Two separate loops over the same array read it twice. Fusing them into one pass reads each element only once.
for i in range(n):
out[i] = a[i] * 2 + a[i]Keep Values in Registers
A value held in a CPU register needs no memory access. Reuse intermediate results instead of writing them out and back.
Avoid Temp Buffers
Extra temporary arrays add both stores and loads. Skip them when you can and compute straight into the final output.
Stream Sequentially
Reading memory in order lets the CPU prefetch ahead. Jumping around defeats prefetching and stalls the loop.
for i in range(n):
total += a[i]Cache Lines Travel Together
Memory arrives in fixed-size cache lines. Using every byte of a line you fetched gives you free, already-loaded data.
Compute More per Byte
Arithmetic intensity is work done per byte loaded. Raising it means each fetched value earns more compute before you move on.
Write Once, If You Can
Stores cost bandwidth too. Accumulate in a local and write the final result once rather than updating memory repeatedly.
var acc = Float32(0)
for i in range(n):
acc += a[i]
out[0] = accLess Traffic, More Speed
When the kernel waits on data, cutting reads and writes is the biggest win, often beating clever arithmetic tweaks.
Quick Check
Your kernel reads the same array in two separate loops. What single change cuts its memory traffic most?
Recap
Cut memory traffic by touching data once, fusing loops, reusing registers, streaming in order, and writing results just once. 💾
Preguntas frecuentes
¿La lección «Reducir el tráfico de memoria» es gratis?
Sí — el texto completo de «Reducir el tráfico de memoria» es gratis para leer aquí en la web. Para practicarla de forma interactiva (editor de código integrado y tutor de IA 24/7) y desbloquear el resto del curso de Mojo Academy, actualiza a CoddyKit PRO. El curso de Mojo Academy incluye 4 lecciones en total.
¿Qué aprenderé en «Reducir el tráfico de memoria»?
Mantenga los datos cerca del cálculo. Practicas Mojo Academy con código real que ejecutas directamente en el navegador, y un tutor de IA 24/7 responde tus preguntas mientras trabajas en la lección.
¿Necesito experiencia previa para empezar Mojo Academy?
No se requiere experiencia previa. Mojo Academy en CoddyKit está estructurado para principiantes hasta estudiantes avanzados, así que puedes empezar aquí o desde el inicio y avanzar a tu ritmo. Esta es la lección 3 de 4.
¿Cuánto tiempo toma la lección «Reducir el tráfico de memoria»?
La mayoría de las lecciones de CoddyKit toman alrededor de 5–10 minutos. Cada una es compacta e interactiva, así que avanzas constantemente y retomas exactamente por donde dejaste en la web y la app.
¿Puedo escribir y ejecutar código en esta lección de Mojo Academy?
Sí. Cada lección de Mojo Academy incluye un editor de código integrado, así que escribes y ejecutas código real directamente en tu navegador y obtienes retroalimentación instantánea de IA — sin configuración local necesaria.
Todas las lecciones de este curso
- Anatomía de un kernel de cálculo
- Combinar SIMD con bucles
- Reducir el tráfico de memoria
- Usar teselado para mejorar la localidad de caché