Skip to content

P5 · SAXPY and memory bandwidth

SAXPY computes result[i] = scale * X[i] + Y[i] for 20 million floats. The task version uses 64 tasks.

1. Measure and explain task speedup

Implementation Time (ms) Estimated bandwidth (GiB/s) GFLOPS
Single-core ISPC 8.266 36.054 4.839
Task ISPC 4.156 71.709 9.625

Both outputs match serial execution. Tasks provide 1.99× speedup.

Each element performs two floating-point operations for an estimated 16 bytes of traffic: 0.125 FLOP/byte. This low arithmetic intensity makes memory bandwidth the expected bottleneck.

Multiple cores keep more memory requests in flight, but share the available bandwidth. Scaling levels off as that bandwidth saturates, so near-linear scaling with core count is not expected for this workload. Further improvement depends on reducing traffic or using the memory system more efficiently.

The driver calculates both rates from elapsed time:

  • GFLOPS = 2N / seconds / 10⁹
  • bandwidth = 16N / seconds / 1024³

The bandwidth estimate uses the traffic model below. Its unit is GiB/s, although the driver prints “GB/s.”

2. Extra credit: explain memory traffic

With write-allocate, write-back caching:

Transfer Bytes per element
Read X 4
Read Y 4
Fetch the destination cache line on a store miss 4, amortized
Write back the modified destination line 4, amortized
Total 16

A store miss fetches the destination line into cache before modifying it. That extra read accounts for the fourth transfer in TOTAL_BYTES = 4 * N * sizeof(float).

3. Extra credit: optimize SAXPY

Skipped.