P5 · SAXPY and memory bandwidth¶
SAXPY computes result[i] = scale * X[i] + Y[i] for 20 million floats. The task version uses 64 tasks.
1. Measure and explain task speedup¶
| Implementation | Time (ms) | Estimated bandwidth (GiB/s) | GFLOPS |
|---|---|---|---|
| Single-core ISPC | 8.266 | 36.054 | 4.839 |
| Task ISPC | 4.156 | 71.709 | 9.625 |
Both outputs match serial execution. Tasks provide 1.99× speedup.
Each element performs two floating-point operations for an estimated 16 bytes of traffic: 0.125 FLOP/byte. This low arithmetic intensity makes memory bandwidth the expected bottleneck.
Multiple cores keep more memory requests in flight, but share the available bandwidth. Scaling levels off as that bandwidth saturates, so near-linear scaling with core count is not expected for this workload. Further improvement depends on reducing traffic or using the memory system more efficiently.
The driver calculates both rates from elapsed time:
GFLOPS = 2N / seconds / 10⁹bandwidth = 16N / seconds / 1024³
The bandwidth estimate uses the traffic model below. Its unit is GiB/s, although the driver prints “GB/s.”
2. Extra credit: explain memory traffic¶
With write-allocate, write-back caching:
| Transfer | Bytes per element |
|---|---|
Read X |
4 |
Read Y |
4 |
| Fetch the destination cache line on a store miss | 4, amortized |
| Write back the modified destination line | 4, amortized |
| Total | 16 |
A store miss fetches the destination line into cache before modifying it. That extra read accounts for the fourth transfer in TOTAL_BYTES = 4 * N * sizeof(float).
3. Extra credit: optimize SAXPY¶
Skipped.