Skip to main content
GameDev.net gamedev.net
🔒 Locked

Speed of SetStreamSource...

Started by Erzengeldeslichtes Jan 20, 2004 at 6:12 AM 12 replies 2.6k views
Original Post
Erzengeldeslichtes
Erzengeldeslichtes
I know that many things, such as SetTexture, are choke points in the rendering engine, and so should be minimized. Is SetStreamSource one of these? For example, let''s say I have 3 types of models. A-Ship, B-Ship, and asteroid. Let''s say I have 3 textures, Ship-Gray, Ship-Blue, and Rocky. Ship-Gray and Ship-Blue fit on both A-Ship and B-Ship. Now I submit to my render queue A-Ship-Gray, Asteroid-Rocky, A-Ship-Blue, B-Ship-Gray, A-Ship-Gray, B-Ship-Blue, B-Ship-Blue. At current my render queue would render it thus:
Ship-Gray
  A-Ship
  B-Ship
  A-Ship
Rocky
  Asteroid
Ship-Blue
  A-Ship
  B-Ship
  B-Ship
 
Where each model setstreamsource. Now would it be better if I resorted it so that it did this:
* = Denotes a setstream.
Ship-Gray
  * A-Ship
    A-Ship
  * B-Ship
Rocky
  *Asteroid
Ship-Blue
  * A-Ship
  * B-Ship
    B-Ship
 
Or would the code for sorting these properly more likely take up more time than just SetStreamSource for the same model? (Yes, I know, profile and find out. However, I''d like to know if anyone else knows before I reinvent a wheel that may be square.)
----Erzengel des Lichtes光の大天使Archangel of LightEverything has a use. You must know that use, and when to properly use the effects.♀≈♂?
evolutional
evolutional
I currently render how you have your set up, eg: Just use texture batching. I am also interested as to how expensive the SetStreamSource is; I create a large VB that holds all my vertex data - the individual models are then referenced by a stored offset into this and drawn. But this can mean a lot of setstream calls. The thing I was thinking was that ''it couldn''t hurt'' to batch them in this way, could it?
Erzengeldeslichtes
Erzengeldeslichtes
The only harm in resorting it is that it takes time to run through the buffer saying "are we the same mesh?"
The question is which takes more time.
----Erzengel des Lichtes光の大天使Archangel of LightEverything has a use. You must know that use, and when to properly use the effects.♀≈♂?
neneboricua19
neneboricua19
It also depends on how many times you''re calling SetStreamSource. In the case you mentioned, you probably wouldn''t notice a difference.

But say you have a scene with 30 or so entities in your game. Those 30 entities can be drawn with, say 5 meshes. If you have to call SetStreamSource between every entity, you''re gonna get a pretty severe performance hit. But if you sort based on the mesh, you''ll only need to call SetStreamSource 5 times.

Like the previous poster said, your results will vary depending on how complicated the models are. They will also vary depending on what memory pool the meshes are loaded in. If they''re in system memory, the hit will be large.

Also remember that the GPU tries to cache geometry as much as possible, so reusing vertices as soon as possible is always to your benefit.

As stated by the AP above, the easiest way is just to code it up and try it out. Use the qsort algorithm from C and just define a compare function that takes a look at the types of entites being compared. Shouldn''t be too difficult. It would probably take you less than 20 minutes.

neneboricua
Raloth
Raloth
Why not put all the entity models in one vertex buffer? Just render the appropriate range of triangles when you call DrawIndexedPrimitive and you only have to call SetStreamSource once.
____________________________________________________________AAAAA: American Association Against Adobe AcrobatYou know you hate PDFs...
ProtoPornoPants
ProtoPornoPants
I think that neneboricua19 is spot on here. I'd be interested to know the results of the two versions.

As for vertex cacheing, that cannot be underestimated. I was writing a terrain-drawing library for the xbox where each chunk of terrain was made up by a 11x11 grid of verts. There was a marked improvement when I reordered the index buffer to take account of the xbox's vertex cache.

Here's the first row of the tristrip:


0 - 1 - 2 - 3 - 4 - 5 - 6 - 7 - 8 - 9 - 10
| / | / | / | / | / | / | / | / | / | / |
| / |/ | / | / | / | / | / | / | / | / |
11 - 12 - 13 - 14 - 15 - 16 - 17 - 18 - 19 - 20 - 21


The cache inside the NV2a holds roughly 16 verts in it at once. It just holds the last 16 transformed verts in a FIFO type way. If you get things right, then you can make sure that verts never get transformed (ie run through the vertex shader) more than once.

Here's what I had for the first line of rendering this 11x11 grid using one massive tristrip:


static const s16 hiResChunkIndices[] =
{
0, 1, 1, 2, 2, 3, 3, 4, 4, 5, 5, 6, 6, 7, 7, 8, 8, 9, 9, 10, 10, // Degenerates to prime the cache

0, 0, // Degenerates to move left

11, 1, 12, 2, 13, 3, 14, 4, 15, 5, 16, 6, 17, 7, 18, 8, 19, 9, 20, 10, // First row

21, 21, 11, 11, // Degenerates to move left

};


I built the rest of the index buffer from adding 11, then 22, etc to the first row of indices.

Note the use of degenerates to get from one end to the other without breaking the tristrip. Also note that I "prime the cache" by transforming a load of verts at the start so that I can get maximum cache hits. If you do things like this, then no vertex should ever get transformed twice, and you render just one tri-strip per terrain chunk.

Now I don't know about vertex caches in more up-to-date video cards but I can only assume they would be larger. Something to think about though...

p

[edited by - protopornopants on January 21, 2004 6:36:05 AM]
neneboricua19
neneboricua19
ProtoPornoPants, priming the cache is a cool trick. I hadn''t thought of that. I''ll have to try that out sometime.

Thanks for the tip.
neneboricua
Toni Petrina
Toni Petrina
ProtoPornoPants:
How usefull is that method anyway? You say that vertices will first be transformed(exactly once), then they will be stripped. Is there any limitation on vertex cache(that is number of transformed vertices)? When will this method prove faster than ordinary stripping?

I think that using indices i = {1, 2, 3, 4, 5} will mean
transform 1
transform 2
transform 3
draw triangle
transform 4
draw triangle
transform 5
draw triangle

but why is this faster than

transform 1
transform 2
transform 3
transform 4
transform 5
draw triangle
draw triangle
draw triangle

Is there some overhead in switching to transforming and back to drawing?
So... Muira Yoshimoto sliced off his head, walked 8 miles, and defeated a Mongolian horde... by beating them with his head?

Documentation? "We are writing games, we don't have to document anything".
ProtoPornoPants
ProtoPornoPants
The transforming and rasterising are done by seperate units so there''s no extra overhead either way afaik.

The benefit of the system I outlined previously comes not from rendering a single long tristrip, but when rendering a single tristrip that wraps over itself in a grid.

So for my previous example, if we extend it to contain the second row of the grid, we have:


0 - 1 - 2 - 3 - 4 - 5 - 6 - 7 - 8 - 9 - 10
| / | / | / | / | / | / | / | / | / | / |
| / |/ | / | / | / | / | / | / | / | / |
11 - 12 - 13 - 14 - 15 - 16 - 17 - 18 - 19 - 20 - 21
| / | / | / | / | / | / | / | / | / | / |
| / |/ | / | / | / | / | / | / | / | / |
22 - 23 - 24 - 25 - 26 - 27 - 28 - 29 - 30 - 31 - 32


So bearing in mind that if a vertex is still in the 16-vertex cache then we won''t have to transform it again we can compare priming the cache with not priming the cache.

So not priming the cache, we use the following indices for the rows:

Row1 + degenerates:
0, 11, 1, 12, 2, 13, 3, 14, 4, 15, 5, 16, 6, 17, 7, 18, 8, 19, 9, 20, 10, 21, 21, 11, 11

Row2:
22, 12, 23, 13, 24, 14, 25, 15, 26, 16, 27, 17, 28, 18, 29, 19, 20, 20, 31, 21, 32

Looking at what the cache does, by the time it is full, it has the following verts in:

0, 11, 1, 12, 2, 13, 3, 14, 4, 15, 5, 16, 6, 17, 7, 18

At this point, when a new vert comes in, it will throw out the oldest vert. So when 8 comes in, 0 goes out.

The first vert that gets indexed for the second time is 21 for the first degenerate. By this point the cache looks like this (remember it is 16 vertices in a first-in, first-out order):

3, 14, 4, 15, 5, 16, 6, 17, 7, 18, 8, 19, 9, 20, 10, 21

So when 21 comes in, it gets a cache hit and isn''t retransformed. Good.

Next comes 11. The problem is that although we transformed 11 previously, due to the way we have done the tristrip and the way the cache works, 11 is no longer in the cache so we have to go through the overhead of retransforming 11 for the second time.

If you prime the cache like my previous example, then you will see that every vertex gets transformed only once. By the time the cache is full, it looks like this (bearing in mind we have had many cache-hits already):

0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15

So as the new vertices come in, only the ones we have finished with end up being thrown out of the cache.

So basically, this cache priming technique is useful for when you are using a tristrip with degenerates to go back on itself. If you are just using single tristrips that don''t go back over themselves then you don''t need to think about the cache.

Does this answer your question?

p.
Toni Petrina
Toni Petrina
Thank you for your explanation. It was what I was looking for. So this method is best applied on grids(because each vertex is needed by at most 4 polygons)? I can see this method as optimization for drawing grids made by tesselating curved surfaces.

However, I was wondering how can you find out size of your card''s vertex cache. Is it documented anywhere?
So... Muira Yoshimoto sliced off his head, walked 8 miles, and defeated a Mongolian horde... by beating them with his head?

Documentation? "We are writing games, we don't have to document anything".
c t o a n
c t o a n
I posted a benchmark of SetStreamSource on these very forums not too long ago. Maybe about a month and a half ago? Try searching for it, as I cannot remember the exact numbers I got...

[EDIT]:
And here you go, I've found it: clicky

Chris Pergrossi
My Realm | "Good Morning, Dave"

[edited by - c t o a n on January 23, 2004 4:16:14 AM]
Chris PergrossiMy Realm | "Good Morning, Dave"
ProtoPornoPants
ProtoPornoPants
Ah. Very intersting c t o a n. Thank you for the benchmarking.

ffx, this indeed is very good for drawing things like bezier patches etc.

Finding out the size of the vertex cache on your card will probably need some digging around on the manufacturers website. Or maybe www.tomshardwareguide.com might have that kind of stuff.



Topic Locked

This topic has been locked by a moderator. New replies are not allowed.

Sign in to reply to this topic.