Skip to main content
GameDev.net gamedev.net
🔒 Locked

HLSL: Pixel Shader 2.0a vs 2.0b

Started by transwarp Sep 3, 2007 at 10:38 AM 10 replies 9k views
Original Post
transwarp
transwarp
Hello world, I've recently started writing an HLSL "Terrain Shader" using NVidia's FXComposer that blends up to 16 textures together. The code won't compile against ps_2_0, since that only allows using a maximum of 4 different textures. Compiling the code against ps_2_a or ps_2_b works fine. All information I have about Pixel Shader 2.0a and 2.0b is from this Wikipedia article and not practically tested, since I only have a SM 3.0 graphics card available for testing. Things I don't quite understand: - Pixel Shader 2.0b is said to have a "Dependent texture limit" of 4 like the original 2.0. Yet still FXComposer will compile my code accessing 16 different texture samplers using ps_2_b without any errors. Does this mean it will definitely work on a Radeon R420? - Does a card that supports ps_2_a also support ps_2_b? If not, should I simply create two techniques, compiling the same code against both ps_2_a and ps_2_b? Regards, Bastian [Edited by - transwarp on September 7, 2007 11:48:14 AM]
jollyjeffers
jollyjeffers
Quote:
Original post by transwarp
The code won't compile against ps_2_0, since that only allows using a maximum of 4 different textures.
You're doing lots of dependent reads then? I've got several shaders that run fine as ps_2_0 that use 16 samplers [smile]

In general dependent reads, whatever shader model, can be a performance problem. If you can re-architect your algorithm to avoid them it could be well worth the effort!

tbh, I'm not sure on your first question.

Quote:
Original post by transwarp
- Does a card that supports ps_2_a also support ps_2_b? If not, should I simply create two techniques, compiling the same code against both ps_2_a and ps_2_b?
No, the 'a' and 'b' profiles were just created to allow Nvidia and ATI to extend the 2.0 spec before 3.0 became standard. The fact that 'b' comes after 'a' is irrelevant here [smile]

Create multiple techniques for each profile, using uniform parameters to select/remove code as appropriate.

hth
Jack
<hr align="left" width="25%" />
Jack Hoxley <small>[</small><small> Forum FAQ | Revised FAQ |
transwarp
transwarp
Hi Jack,

thanks for the feedback.

Quote:
Original post by jollyjeffers
You're doing lots of dependent reads then? I've got several shaders that run fine as ps_2_0 that use 16 samplers [smile]

In general dependent reads, whatever shader model, can be a performance problem. If you can re-architect your algorithm to avoid them it could be well worth the effort!

It seems I got something mixed up there. The compiler complains about "texture addressing operations in a dependency chain that is too complex for the target shader model" when I use ps_2_0 instead of ps_2_a or ps_2_b.
What exactly does this mean? And what makes a texture read "dependent"?


Here's a snippet from my shader code to show what I'm currently doing:
outputColor = tex2D(Texture1Sampler, IN.texCoords)*IN.texWeights1.x;outputColor += tex2D(Texture2Sampler, IN.texCoords)*IN.texWeights1.y;outputColor += tex2D(Texture3Sampler, IN.texCoords)*IN.texWeights1.z;outputColor += tex2D(Texture4Sampler, IN.texCoords)*IN.texWeights1.w;outputColor += tex2D(Texture5Sampler, IN.texCoords)*IN.texWeights2.x;...outputColor += tex2D(Texture16Sampler, IN.texCoords)*IN.texWeights4.w;


Regards,
Bastian
Jalibr
Jalibr
Well, a shader that does just that will compile, so that isn't your problem. I couldn't tell you exactly what you're doing wrong without seeing the rest of the shader.

If your shader uses too many temporary registers, then that will force the compiler to split up more of your texture reads because it won't have space to store the results, causing the later reads to become "dependent". The dependency rules are somewhat difficult to understand, but there doesn't need to be a direct dependency to cause a dependency.
transwarp
transwarp
Hello everyone,
I've uploaded the complete shader here.

It currently contains 5 techniques:
- Multi30 (ps_3_0, 16 textures, per-pixel lighting, branching for every texture read)
- Multi2a (ps_2_a, 16 textures, per-pixel lighting, only two larger branches)
- Multi2b (ps_2_b, 16 textures, per-pixel lighting, only two larger branches)
- Multi20 (ps_2_0, 12 textures, per-pixel lighting, no branching)
- Multi14 (ps_1_4, 4 textures, per-vertex lighting, no branching)

If you have any hints how to improve performance and/or reduce the shader profiles required for the techniques, I'd be glad to hear them.

Regards,
Bastian

[Edited by - transwarp on September 3, 2007 3:48:48 PM]
ET3D
ET3D
First of all, let me suggest that you take a look at the assembly compiled from HLSL. You should be able to learn about what's really happening from that.

ATI has a GPU Shader Analyzer, which can tell you the performance of your shaders on various ATI chips.

I noticed something about your shaders, that you seem to expect static branching on SM2.0. There's no branching of any kind in PS in SM2.0 (or SM2.0b). You'll see that if you look at the assembly code generated.
jollyjeffers
jollyjeffers
Hi - had a couple of busy days so apologies for the slow reply!

Quote:
The compiler complains about "texture addressing operations in a dependency chain that is too complex for the target shader model" when I use ps_2_0 instead of ps_2_a or ps_2_b.
What exactly does this mean? And what makes a texture read "dependent"?
A dependent read is about as simple as it sounds - GPU's obviously need a sampling location before it can fetch you any data, if this sampling location depends on any computation within the shader then it is dependent on the results completing before it can fetch the sample.

The alternative is if you use an interpolated position (such as a TEXCOORD[n] input into the PS) where it is constant and the GPU knows where to fetch it before any code is executed. The fact that later arithmetic operations might depend on the result isn't relevant here - it's what happens before the sampling operation.

The error message about "dependency chains" is a subtle and complicated architectural detail that I'm never entirely convinced I understand [smile] I believe SM2 has 4 levels of indirection (the big change from PS1.1-1.3 to 1.4 was that it added a second level iirc). If you think about 4 nested conditionals then you've got 24 possible routes through the branching - 16 possible execution paths. Shaders might be powerful but they are pretty simple and at a certain point it just can't handle such complexities.

At least that's my understanding!


I've had a look at the shader you posted; my thoughts:

if (blendFactor != 0)  // Use SM3.0's dynamic branching for better performance    {      if (IN.texWeights1.x!=0) farColor += tex2D(Texture1Sampler, IN.texCoords)*IN.texWeights1.x;      if (IN.texWeights1.y!=0) farColor += tex2D(Texture2Sampler, IN.texCoords)*IN.texWeights1.y;      /* ... */      if (IN.texWeights4.y!=0) farColor += tex2D(Texture14Sampler, IN.texCoords)*IN.texWeights4.y;      if (IN.texWeights4.z!=0) farColor += tex2D(Texture15Sampler, IN.texCoords)*IN.texWeights4.z;      if (IN.texWeights4.w!=0) farColor += tex2D(Texture16Sampler, IN.texCoords)*IN.texWeights4.w;    }


The outer if() seems reasonable, but the if() on each sample is not a good idea. GPU's operate at a batch level - they're not per-pixel granular in execution. The size of the batch varies from GPU to GPU. Each GPU in the batch should take the same branch, and if it doesn't then it'll have to stall until any other pixels in that batch have executed it - so if you have a batch of 100 pixels and only 10 take a particular branch the other 90 will still have to wait. My point is that you're probably being too granular with your conditionals.

For PSHelper_Multi2x() and PSHelper_Multi20() you're using nearTextureCoords in the second batch, which gives you a dependent read, and the if (blendFactor != 1) is probably introducing another level of indirection.


To be honest I can't see an obvious way that you've exceeded any limits, but I've not really had time to really dig deep into your code or run it through any tools.

I'd imagine you'll be okay if you eliminate the conditional nature of your code - generally speaking a pixel shader loves deterministic ALU-heavy code. Branching is often more of a hinderance than a benefit - use it with caution.

hth
Jack
<hr align="left" width="25%" />
Jack Hoxley <small>[</small><small> Forum FAQ | Revised FAQ |
ET3D
ET3D
Quote:
Original post by jollyjeffers
A dependent read is about as simple as it sounds - GPU's obviously need a sampling location before it can fetch you any data, if this sampling location depends on any computation within the shader then it is dependent on the results completing before it can fetch the sample.

Not really as simple as it sounds. For example (and I'm mixing assembly and HLSL in that I'm using register names for temps):

r0 = tex2D(tex1, coords);
r0 += tex2D(tex2, coords);

This is similar to what transwarp is doing.

Here the second tex2D is dependent on the first, because the addition to r0 cannot be done before the first tex2D has finished and put its value in r0. That's why the second tex2D has to wait for the first, making it dependent.

A clever compiler can do it like this:

r0 = tex2D(tex1, coords);
r1 = tex2D(tex2, coords);
r2 = r0 + r1;

This removes the dependency, since only the arithmetic instruction depends on all the texture results being available (which apparently isn't a problem).

This, however, requires a temporary register for each read. PS2.0 has only 12 temp registers. PS2.0b has 32, which makes this solution allow more texture reads.

Solutions with some dependency which allow all textures to be read on PS2.0 should be possible:

r0 = tex2D
...
r7 = tex2D
r11 = r0 + ... + r7
r0 = tex2D // all these reads will have 1 level of dependency
...
r7 = tex2D
r11 += r0 + ... + r7
transwarp
transwarp
Updated the uploaded shader code.

Quote:
Original post by ET3D
ATI has a GPU Shader Analyzer, which can tell you the performance of your shaders on various ATI chips.

I've just had a look at that tool. Its indications of what type of code is the bottleneck on which GPUs looks pretty nice. However I must admit I cannot make much use of the assembly code.

Quote:
Original post by ET3D
I noticed something about your shaders, that you seem to expect static branching on SM2.0. There's no branching of any kind in PS in SM2.0 (or SM2.0b). You'll see that if you look at the assembly code generated.

Ok, I've removed the if's in the SM2.0 code now.

Quote:
Original post by jollyjeffers
The outer if() seems reasonable, but the if() on each sample is not a good idea.

And I've reduced the if's in the SM3.0 code to the outer ones.

Quote:
Original post by ET3D
A clever compiler can do it like this:

r0 = tex2D(tex1, coords);
r1 = tex2D(tex2, coords);
r2 = r0 + r1;

This removes the dependency, since only the arithmetic instruction depends on all the texture results being available (which apparently isn't a problem).

This, however, requires a temporary register for each read. PS2.0 has only 12 temp registers. PS2.0b has 32, which makes this solution allow more texture reads.

That exactly explains the compiler behaviour I'm seeing. It seems to be automatically removing the dependency as far as possible and hitting the limit later with 2.0b due to the additional registers.
Seeing that the compiler is already optimizing the assembly code here, would it still be advisable to try and manually tweak this procedure?

[Edited by - transwarp on September 4, 2007 2:56:14 PM]
Jalibr
Jalibr
If you look at the output of the ps_3_0 compile, you'll see that it's actually not using dynamic branching. Dynamic branching in pixel shaders has some pretty important limitations. Specifically you can't use any texture instructions that require gradients (tex2D in this case, it uses gradients to figure out the shape of the texture in screen space when deciding which mip to use).

In this case, your derivatives will be constant across the branch, so you can use ddx/ddy outside of the if and pass that to tex2Dgrad. Older versions of the compiler actually would have let you leave the first set in the if (but not the second since it doesn't come directly from inputs), but we ran into issues with certain drivers when we did that.

Unfortunately tex2Dgrad usually is less performant than tex2D (since there's no guarantee that the derivatives are equal across the quad), so I would do some perf testing, and maybe consider calculating the LOD yourself and using tex2Dlod to see which method gives you the best performance.
ET3D
ET3D
Quote:
Original post by transwarp
Seeing that the compiler is already optimizing the assembly code here, would it still be advisable to try and manually tweak this procedure?

Shouldn't hurt. I'd suggest looking at the assembly output, then playing with the code to see what your tweaks result in.
ET3D
ET3D
BTW, I looked back at your original post, and noticed you're planning to use this code for terrain shading. IMO you'll get very low performance (especially for PS2.0 and derivative) if you have one shader with the maximum number of textures. Most of the terrain is unlikely to have many layers. Using different shaders and/or multiple passes should result in better performance.

Topic Locked

This topic has been locked by a moderator. New replies are not allowed.

Sign in to reply to this topic.