1 min read
Optimizing Dilated Convolution

Achieved 10x speedup in single-threaded performance using SIMD, loop unrolling, and optimized memory access.

Implemented multi-threading with pthreads, further improving performance by 60x.

Designed and optimized a CUDA implementation, achieving a 1300x speedup.

Discussion

Comments are powered by GitHub Discussions.