Rendered at 16:08:46 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
Pannoniae 4 hours ago [-]
Nice article :) Yeah this is basically a tradeoff between CPU and GPU power. There are different types of culling. There's the basic stuff like backface culling (supported in hardware, don't render triangles facing away from you) and frustum culling (don't render objects which your camera doesn't see). These are used in just about every game.
For occlusion culling it's a bit more tricky because you can either do it on the CPU in broadly two ways, either do low-res raycasting / software rendering like in the article on the CPU and cull based on that. This is an adaptive workload, you can give it more threads or CPU power and it scales for better culling which results in less pixels rendered on the GPU.
You can also use GPU culling but that's more complicated to do and that uses the GPU which creates a catch-22 - you want to use GPU culling to reduce GPU load but integrated GPUs don't cope well with compute shaders and memory bandwidth in general, so doing a culling pass might wipe out any culling gains you might have.
And dedicated GPUs have the raw power and memory bandwidth to just submit everything in your frustum and get most of it depth-rejected.
I wonder if it would be even faster to create a connectivity graph on the CPU, like each chunk knows whether a neighbour is visible and vice versa. On rendering the chunk graph is walked and the visible chunks are submitted, kind of like a primitive garbage collector to determine liveness. The culling would be worse but I presume traversing a fairly small (few thousand elements) list is quite a bit cheaper than rendering the "mipped" occlusion boxes, but do let me know if this is wrong.
avaer 9 minutes ago [-]
If you can turn the problem into a small kernel operating on a heap of data (or a hierarchy), the GPU almost always wins for culling, especially if you pipe the cull into the draw with GPU-driven rendering.
If the author used GPU culling it would likely be faster on modern hardware, they just said they can't because of platform restrictions. But that's what AAA games do.
The modern non-nanite techniques here are basically to regularize to grids, cull the grids on the GPU, and then cull more with a HZB. That's your "mipped occlusion boxes", except it's actually very cheap to do this because you're reusing depth you already had from previous frames, and it all stays on the GPU. As a bonus, you can do your LODs on the GPU at the same time, saving even more CPU work and memory bandwidth.
Also, depth rejection is not going to help with the problems that culling tries to solve; draw, vertex processing, raster, and then pixel tests is much more expensive than a cull test. And if you're forward rendering w/expensive fragment shader, the overdraw of relying on the Z buffer to do your "culling" can kill you.
Pannoniae 35 seconds ago [-]
>that's what AAA games do
yeah where the pixel work and the vert count is magnitudes higher. I forgot to mention in my comment that I was talking about low-poly / pixel stuff like this, with a very simple PS and being pretty much API bound or memory-bound in perf
corbinvachal 60 minutes ago [-]
[dead]
keyle 6 hours ago [-]
Wow a tech article on HN. What is happening. Are we back in 2022? /s
fullstackwife 5 hours ago [-]
January 2026, thats light years in AI psychosis time units, thats before harness, and auto hill climbing era
For occlusion culling it's a bit more tricky because you can either do it on the CPU in broadly two ways, either do low-res raycasting / software rendering like in the article on the CPU and cull based on that. This is an adaptive workload, you can give it more threads or CPU power and it scales for better culling which results in less pixels rendered on the GPU.
You can also use GPU culling but that's more complicated to do and that uses the GPU which creates a catch-22 - you want to use GPU culling to reduce GPU load but integrated GPUs don't cope well with compute shaders and memory bandwidth in general, so doing a culling pass might wipe out any culling gains you might have.
And dedicated GPUs have the raw power and memory bandwidth to just submit everything in your frustum and get most of it depth-rejected.
I wonder if it would be even faster to create a connectivity graph on the CPU, like each chunk knows whether a neighbour is visible and vice versa. On rendering the chunk graph is walked and the visible chunks are submitted, kind of like a primitive garbage collector to determine liveness. The culling would be worse but I presume traversing a fairly small (few thousand elements) list is quite a bit cheaper than rendering the "mipped" occlusion boxes, but do let me know if this is wrong.
If the author used GPU culling it would likely be faster on modern hardware, they just said they can't because of platform restrictions. But that's what AAA games do.
The modern non-nanite techniques here are basically to regularize to grids, cull the grids on the GPU, and then cull more with a HZB. That's your "mipped occlusion boxes", except it's actually very cheap to do this because you're reusing depth you already had from previous frames, and it all stays on the GPU. As a bonus, you can do your LODs on the GPU at the same time, saving even more CPU work and memory bandwidth.
Also, depth rejection is not going to help with the problems that culling tries to solve; draw, vertex processing, raster, and then pixel tests is much more expensive than a cull test. And if you're forward rendering w/expensive fragment shader, the overdraw of relying on the Z buffer to do your "culling" can kill you.
yeah where the pixel work and the vert count is magnitudes higher. I forgot to mention in my comment that I was talking about low-poly / pixel stuff like this, with a very simple PS and being pretty much API bound or memory-bound in perf