Damus
Andrew Zonenberg profile picture
Andrew Zonenberg
@Andrew Zonenberg

Security and open source at the hardware/software interface. Embedded sec @ IOActive. Lead dev of ngscopeclient/libscopehal. GHz probe designer. Open source networking hardware. "So others may live"

Toots searchable on tootfinder.

Relays (1)
  • wss://relay.ditto.pub – read & write

Recent Notes

Julien Goodwin · 5d
nostr:nprofile1qy2hwumn8ghj7un9d3shjtnyd968gmewwp6kyqpqqv5atqz9k9c54q8c28kra6sfata0wk7w7x5gkrnde8vmxe5gt00q8mmv7s why down to 6.8 not 5.8? some form of interconnect thing that means it's still faster ...
Andrew Zonenberg profile picture
@nprofile1q... there's some jitter because the filter graph, rendering, and waveform capture are async to each other. These are single point numbers although i also have rough averages for a continuous run.

Also possible the different memory access patterns led to better cache hit rates in other shaders
1
Andrew Zonenberg · 5d
nostr:nprofile1qy2hwumn8ghj7un9d3shjtnyd968gmewwp6kyqpqdf5yrneqpe2uufa8znqag9t03qvxma32zfkv4vlhwel08raj5vysqprwyr also the overall filter has two command buffers in it and there's OS jitter affecting the cpu side round trip path
Tom Verbeure · 5d
nostr:nprofile1qy2hwumn8ghj7un9d3shjtnyd968gmewwp6kyqpq4v56a2a5vrwz9sxc2veqn85h4d7p42p26wdwtppa5tjc8npattdsqxhuwm nostr:nprofile1qy2hwumn8ghj7un9d3shjtnyd968gmewwp6kyqpqqv5atqz9k9c54q8c28kra6sfata0wk7...
Andrew Zonenberg profile picture
@nprofile1q... @nprofile1q... my problem is getting data in and out the narrow channel that is VRAM lol. Almost all of my shaders are saturating the memory subsystem so the only major optimizations possible are to do less loads/stores or order them more efficiently for better coalescing or cache hit rates.

Which is why I'm beginning to look at kernel fusion for things like FFT + windowing so I can keep temporaries in registers and never touch the memory subsystem at all
1
Tom Verbeure · 5d
nostr:nprofile1qy2hwumn8ghj7un9d3shjtnyd968gmewwp6kyqpqqv5atqz9k9c54q8c28kra6sfata0wk7w7x5gkrnde8vmxe5gt00q8mmv7s nostr:nprofile1qy2hwumn8ghj7un9d3shjtnyd968gmewwp6kyqpq4v56a2a5vrwz9sxc2veqn85h4d7p42p26wdwtppa5tjc8npattdsqxhuwm Consider it a blessing that there are solutions by being clever. Other t...
Andrew Zonenberg · 6d
It feels like every time I look at something in ngscopeclient and go "hmm, this is too slow" and bang on it for a few hours I get a significant speedup out of somewhere.
Andrew Zonenberg profile picture
And sure enough the other big shader in the same filter had a huge optimization possible too, 3.75 -> 0.78 ms for the shader and 8.8 -> 6.8 for the whole filter.

Not too bad considering it took 11.2 on the same workload a few hours ago, that's a 65% speedup overall
1
Julien Goodwin · 5d
nostr:nprofile1qy2hwumn8ghj7un9d3shjtnyd968gmewwp6kyqpqqv5atqz9k9c54q8c28kra6sfata0wk7w7x5gkrnde8vmxe5gt00q8mmv7s why down to 6.8 not 5.8? some form of interconnect thing that means it's still faster in the pipeline, but not able to achieve the theoretical speedup?
Ben · 5d
nostr:nprofile1qy2hwumn8ghj7un9d3shjtnyd968gmewwp6kyqpqqv5atqz9k9c54q8c28kra6sfata0wk7w7x5gkrnde8vmxe5gt00q8mmv7s You should write a book, "Using GPUs to massively speed up your C/C++ code" - targeted...
Andrew Zonenberg profile picture
@nprofile1q... Not familiar with the internals of database engines to the point I could comment. They work best on massively data-parallel problems with predictable, linear memory address ordering.

In general the best case scenario is every thread making 32-bit reads and writes to consecutive word addresses, then doing a ton of flops before the next memory access. The further you diverge from this, the further from theoretical peak performance you get.
Andrew Zonenberg · 5d
And sure enough the other big shader in the same filter had a huge optimization possible too, 3.75 -> 0.78 ms for the shader and 8.8 -> 6.8 for the whole filter. Not too bad considering it took 11.2 on the same workload a few hours ago, that's a 65% speedup overall
Andrew Zonenberg profile picture
Why is shader optimization so addictive? I have other things to do...

But it's 2AM, I'm thoroughly caffeinated, and staring at the Radeon GPU profiler reordering SSBO memory transactions for improved coalescing.

And successfully shaved about 1.7 ms off the first shader in the PAM4 edge detector (4.71 -> 2.98 ms, a 58% speedup). The overall filter went from 11.22 to 8.8 ms, a 27.5% speedup.

There's definitely still more room to tune, too.

2
Andrew Zonenberg · 6d
It feels like every time I look at something in ngscopeclient and go "hmm, this is too slow" and bang on it for a few hours I get a significant speedup out of somewhere.
Ben · 5d
nostr:nprofile1qy2hwumn8ghj7un9d3shjtnyd968gmewwp6kyqpqqv5atqz9k9c54q8c28kra6sfata0wk7w7x5gkrnde8vmxe5gt00q8mmv7s You should write a book, "Using GPUs to massively speed up your C/C++ code" - targeted for C/C++ programmers. I'd pre-order it for sure. Chapter 1 - GPUs for idiots and/or programmers ...