Skip to main content
GameDev.net gamedev.net
🔒 Locked

why batch limits?

Started by rouncED Jan 30, 2010 at 12:39 AM 5 replies 1.7k views
Original Post
rouncED
rouncED
So, you can only call drawindexedprimitive so many times per frame before it becomes a bottleneck on the speed of the game. Its about 300 individual calls your allowed, correct me if im wrong. Id just like to know if this is going to be improved when new cards come out, and why is it - what part of the card or motherboard is at fault with this considered?
Nik02
Nik02
There is no fixed limit to the number of batches.

Each state change or draw call causes the system to do a user->kernel mode transition, which is relatively expensive. In addition to this cost, most D3D calls will need to upload a data payload over the graphics bus, which introduces delay and latency as compared to simple system memory copy.

For this reason, it is best to minimize the number of calls that modify the device state. This means that you draw as much as possible with as few device calls as possible (batching). Of course, the GPU also has performance limits. A balanced high-performance engine strains both CPU and GPU equally so that neither is idle, by balancing the batching cost (and other engine logic) with the cost of actually rendering. If the mere act of preparing batches on the CPU takes more time than rendering those batches on the GPU, then it may be wise to employ a simpler (but more crude) method of batching.

How much delay and latency is introduced depends on all the components from the CPU, via system memory and various data buses, to the GPU. In general, the faster the system, the faster you can send stuff to be rendered and therefore, the more batches you can render per second.

The batching issue is becoming even more important, as processing speed on the GPU increases exponentially but memory and bus bandwidth doesn't.
Niko Suni
Hodgman
Hodgman
You're recommended to keep batch counts low, because when you submit a batch, that's when the driver actually has to do some work.

You're "allowed" to call it as much as you want. 300 seems pretty low... I've got test scenes that do 800-2000 (depending on the lighting model) easily.
Quote:
what part of the card or motherboard is at fault with this considered
Code has to execute on the CPU to set up the draw commands, so I guess you can blame the CPU for not being fast enough if you're giving it too much work to do.
It's the same with anything, you perform physics with a too-high object count and it's going to become a bottleneck too ;)
rouncED
rouncED
Ive worked in software before, and I loved the freedom of not being limited by anything, but the fact that the gpu is 16 times faster than even a quad core 3 gig computer (a huge amount) tends me to lean towards using hardware to make my game.

Programming for more calls is easier, by about a thousand times, who wants to group textures and repeat vertex buffer data, for a thousand times more headaches when you can barely get your project off the floor already... I think I will code my project now you two dredz have spoken to me, you think I can get 800? Id love that already.

Grouping textures into atlases is a total pain!!!!!!!!!!!! I cant do it, I think I will just reduce sight length and go for it brute forcing a higher batch count, Ill see how it goes and ill post here what I come up with.

Especially streaming benefits from being able to call drawindexedprimitive more, because you have to swap textures alot... these video cards are a real pain to work with, although I have had SSS working through the gpu and getting almost realtime framerates and working with software you dont have a chance.

The gpu is a hot chick, its just too hard to score with.

Thanks!
LeGreg
LeGreg
Quote:
Original post by Nik02
Each state change or draw call causes the system to do a user->kernel mode transition, which is relatively expensive.


Each draw call has to make some heavy state management, both in runtime and in driver before it can reach the hardware. The state management is the main culprit to what is causing the big overhead.

Kernel transitions have a cost.. But in, say, DirectX, they are "batched" in large groups, so you won't pay a kernel transition every state change (that would be prohibitive).

LeGreg
InvalidPointer
InvalidPointer
This was fixed maybe two years ago with D3D10, actually-- I assume you were referring to the kernel-mode idiocy in DX9.
clb: At the end of 2012, the positions of jupiter, saturn, mercury, and deimos are aligned so as to cause a denormalized flush-to-zero bug when computing earth's gravitational force, slinging it to the sun.
crowley9
crowley9
No, DX9 (and earlier) batch them as well, this has been true for many many years.

One good way to increase your batch limit is multi-thread your application so that your game logic, ai, physics, etc. happens on other cores.

[Edited by - crowley9 on January 30, 2010 9:21:34 PM]

Topic Locked

This topic has been locked by a moderator. New replies are not allowed.

Sign in to reply to this topic.