Marcin Żukowski’s PivCo-Huffman paper inspired me to revisit an Oodle-like 12-stream Huffman decoder for Apple CPUs that I experimented with in 2022. The same old kernel manages a pretty flat 5GB/s to 5.1GB/s on Apple M4, which did beat PivCo-Huffman, on one or two files, at the time of the paper.
But with @rygorous@mastodon.gamedev.place's and Marcin's improvements, the updated PivCo-Huffman code reliably wins. And it doesn't need 31 GPRs. Very exciting work.
https://gist.github.com/dougallj/a14d3f72b57ee58a81d487b43ff2a05b
Remote
Low-level systems stuff. Reverse engineering, security research, bit twiddling, optimisation, SIMD, uarch. 64-bit ARM enthusiast.
he/they
0
Followers
0
Following
3
Posts
Joined October 29, 2022
Posts
Open post
Replying to
@rygorous@mastodon.gamedev.place
@rygorous FWIW, your writing is hugely inspirational to me. I've learned some of my favourite ideas from your posts, like the throughput optimised bit-reading technique, and the texture swizzling bit-trick. And often your blog has let me understand things that I'd tried and failed to understand before, like FFT implementations, Huffman table initialisation, SIMD transposes, and triangle rasterisation. (Thank you!)
This sucks, but many people still care about understanding.
15
0
0
0
Open post
Andreas Abel added latency, throughput, and port usage data for Emerald Rapids, Meteor Lake, Arrow Lake, and Zen 5 to https://uops.info/table.html 🎉
36
4
14
0
Remote instance
mastodon.social
Open on original server