Parallel Lab Notes¶
Implementations and performance measurements for Stanford CS149 and CMU 15-418/618.
Assignments¶
A1 · CPU parallelism¶
Threads, SIMD intrinsics, ISPC tasks, memory bandwidth, and K-means.
A2 · Task system¶
Thread pools, synchronous task launch, and dependency scheduling.
A3 · CUDA renderer¶
CUDA SAXPY, parallel prefix sum, and circle rendering.
A4 · MPI wire routing¶
Work decomposition, communication, and routing performance.
A5 · RK4 kernel¶
GPU optimization of the 3D heat-equation solver.