Skip to main content
GameDev.net gamedev.net
🔒 Locked

Rendering in a task-based multithreaded environment

Started by hornet1990 Jan 23, 2010 at 5:14 PM 17 replies 9.9k views
Original Post
hornet1990
hornet1990
Hi all As the title says I'm curious as to what peoples experiences are of rendering in task-based multithreaded environment. I've been putting together a plan for a new framework (eventual goal is for a flight sim if that makes any difference) using advice gleamed from various discussions on task based multithreading and entity/component systems, and I plan to use DX11 multi-threaded capabilities. However I'm undecided on the best way forward when it actually comes to implementing the rendering. As far as the entity system is concerned I'm going for a "pure" system where each component is just data, and each component type has a system/manager which contains the functionality. Each components data will also be double buffered so that reads are consistent, and writes will be made live with a sync phase after all processing by all systems is done. From a threading point of view there'll be: - WndProc/Rendering thread (with the DX11 immediate context) - Simulation thread - main loop which issues tasks for component system processing etc (the systems themselves will then issue more tasks as required) - 1..n task processing threads (each also having a DX11 deferred context) Rendering wise I see 3 choices: 1 - Simulation thread generates DX calls using tasks/deferred contexts and then dispatches all the Command Lists to the rendering thread for execution with the immediate context 2 - WndProc/Rendering thread renders world state as it is at that time, but breaks the process down and issues tasks for generating the DX calls as appropriate 3 - WndProc/Rendering does all the work within its own thread Obviously each has pro's and cons. I'm thinking: 1 would limit the frame rate to a the update frequency of the simulation loop. 2 would allow unlimited (or a fixed) frame rate, but as far as I can tell would probably warrant triple buffering the component state since when rendering you'd ideally want two sets of live data and interpolate between them (rather than extrapolating from one). Also potential for the frame rate to be affected by tasks it generates being caught in congestion on the task processing threads. 3 doesn't take advantage of multi-cores, more chance of overrunning the render time budget if scene is more complex but avoids potential congestion on task processing threads. And theres bound to be other things I've not considered yet. What does anyone else think on the subject? Cheers
the_edd
the_edd
We've recently been looking at doing something similar at work. It's a simulation system rather than a game, but there's overlap in the requirements so I'll describe what we went with.

The basic idea is to have the simulation churning away in the background as fast as possible. Whenever a simulation step is complete, the corresponding data should be rendered. We're using OpenGL though, FWIW.

We have a thread pool containing N threads, where N is customizable but typically appropriate for the hardware concurrency of the machine. Every single concurrent action goes through this pool.

Actually this isn't quite true. It turned out that we needed to create another thread pool with a single thread for the renderer. But this was purely due to a driver issue in which the system would lock up if OpenGL calls were made in different threads (even if calls occurred one at a time).

We are aiming for a more-or-less classical MVC setup, which is potentially different from your situation. Messages coming from the UI and sent to the Model via a message queue (which executes tasks via the primary thread pool).

A do-a-single-simulation task is repeatedly submitted to the Model too. Within the model, there's lots of parallelism also through some parallel algorithm code we created e.g. ParallelForEach(), ParallelFold(), and so on. Again, these algorithms do their work via the main thread pool.

Anyway, whenever the model completes a simulation step, it sends a draw-this-stuff message to the renderer. This acquires the GL context, does the rendering and then sends a message to the UI (in another message queue) to do the back-buffer-swap.

I think the primary difference between this design and yours is that we don't bind a thread to any particular task (with the exception of the renderer, but our hands are tied there). In other words, we only have your 1..n processing threads and that's it. A big advantage of this is that we can only have our thread-pool spin up a single thread, making our application single threaded, which is great for debugging.

We're just starting to build on top of the design in earnest now, but it seems to be holding together quite well. The UI is responsive even under heavy load and by using a very strict message passing policy, we've so far avoided any nasty threading bugs (except for one or two in the parallelism kernel itself).

I should mention that we don't do any double/triple/... buffering of data in the model. The 3D data structures we use basically boil down to vectors of PODs and so we just copy it all when sending it to the renderer. This costs peanuts in the grand scheme of things, but it's probably not the best approach for a game.
_the_phantom_
_the_phantom_
I've got a recent journal post where I brain dump on just this topic [smile]

I think to a large degree I'm working in the same direction as the OP on this perticular problem.

The root of it being that you have to have the main context bound to a perticular thread, but when using tasks there is no garrentee as to what thread your tasks end up on.

This alone I feel mandates a 'render' thread which is purely responcible for performing rendering actions.

The problem after this becomes one of balence;

If you dedicate a 'core' to rendering (and keep it away from the task system) you could be wasting alot of time. Take, for example, a quad core system with 3 task threads and a dedicated render thread; if you are running at v-sync then your rendering thread might only be taking 4ms a frame, at which point you are wasting 12ms if you are hitting 60fps.

You could use oversubscription to get around this problem; create 4 worker threads, one rendering thread and allow time slicing to do its work. However at this point you have to balence rendering time against time lost task switching a core to another thread; if you are only wasting 2ms is that going to be worth the cost for swapping threads and causing threads to jump cores. (You could use affinity to try and avoid this, however this just means a worker thread gets starved or your renderer bounces around and kicks worker threads off cores anyway. Also, as the guys who wrote the MS Concurrency Runtime pointed out; on a PC you WILL get time silced out, so surely its better to be moved to a core to do work rather than stall until your core is free again?)

The real problem with oversubscription appears when the user is running without v-sync; at this point the renderer isn't waiting and the thread is constantly spinning so anytime you might have got back when it slept for v-sync is now lost and to make matters worse you are now impacting your worker threads as everything is constantly in contention.

At the moment, until I get a chance to try it anyway, my prefer setup would look like this;
update \    sync \    pre-render update ---> sync ---> pre-render ---> next frameupdate /    sync /    pre-render /----- render ------------------------>


(where the three 'update', 'sync', 'pre-render' are infact 'n' threads, not just 3)

So your 'update' is a bunch of spread tasks, your 'sync' is a bunch of parallel tasks which push 'private' state to 'public' state and 'pre-render' use a deferred context per-thread and a (not shown here) render bucket system to build up a work list for the renderer to chew on.

(The 'sync' step is a double buffer setup to prevent you needing locks etc in the update loop; everyone can write their own private state and read anyone elses public state, safe in the knowledge that there are no race conditions involved).

As for the balancing, which is really the OPs question here, I think it depends.

If your rendering thread isn't taking up much time and sleeping alot then you can probably get away with over-subscribing your cpu and creating n-core task threads and a seperate render thread without doing too much damage.

If the user however has v-sync off (via a game option or forced) then you'll probably want to limit yourself to n-core - 1 task threads and a seperate rendering thread. You could experiment with forcing the cpu thread to sleep for Xms to slow down rendering but I doubt this would make you many friends.

If your rendering time is very high then you might want to take the same path as the method above anyway so as not to take the performance hit of active threads going in and out of a core.

Ultimately its going to depends on time and the amount of stuff you have going on at any given moment as to how to handle this.. which as an answer sucks a bit I know [grin]
_the_phantom_
_the_phantom_
No, the issue exists in OpenGL as well; at any given time the main context can only be bound to a single thread of execution. (In the Windows world anyway.. I assume other platforms have the same limitation)

With GL you can release it from a thread and then obtain it in another and I suspect the same is true in D3D for the main context (I've not looked to see).

This makes sense as you don't really want to be trying to issue rendering commands from different threads at different times; As I reall D3D9 could have a multi-threaded context but that, again from what I recall, dumped a load of overhead on you as you had to lock/unlock mutexes to keep things thread safe.
the_edd
the_edd
Quote:
Original post by phantom
No, the issue exists in OpenGL as well; at any given time the main context can only be bound to a single thread of execution. (In the Windows world anyway.. I assume other platforms have the same limitation)

With GL you can release it from a thread and then obtain it in another and I suspect the same is true in D3D for the main context (I've not looked to see).


Right, that's what I thought. I read your initial statement as all the rendering code had to happen on a particular fixed thread.

FYI: Mac OS X has the same limitations. In fact it's a little bit more flexible w.r.t sending textures and geometry, but I daren't start making things any more complicated :)
hornet1990
hornet1990
TBH thats one thing I never considered... all rendering has to go through an immediate context but I'm not seeing anything in the DX11 documentation that says the context should be bound to a particular thread, although in the DXGI section it does seem to recommend that the rendering thread (which I presume means the immediate context) should be the same as the message-pump thread if you want to avoid potential deadlocks.

The simulation thread could itself be a task that when completed simply adds itself again to the task queue for the next iteration. So as you both quite rightly say that reduces it to a rendering thread and 1..n task threads - with n being n-cores available - 1.

So if I'm reading this right you both decide what to render and where as part of your main processing task, and then pass that off to the render thread to actually do it (which is a variation of the first option).

Must admit though I was thinking in terms of a fixed timestep for the main processing task which would also imply a fixed issuing rate of rendering instructions, hence why I was looking to completely decouple rendering.

WRT the rendering thread being idle I did think that maybe it could be a specialised task processing thread so that when idle it steals tasks from another threads queue (as opposed to the task processing threads which have task assigned to their queues). It would solve that problem unless the task it picks up is more involved than the remaining frame time available...
_the_phantom_
_the_phantom_
Quote:
Original post by hornet1990
Must admit though I was thinking in terms of a fixed timestep for the main processing task which would also imply a fixed issuing rate of rendering instructions, hence why I was looking to completely decouple rendering.


It did occure to me a few hours after I'd made that post that I hadn't allowed for fixed time step, however I don't think it would be a massive problem to do so.

The main thing you'd need is an extra internal states for things which require rendering and then, from my diagram above, you only call the 'update' and 'sync' portions on your time step; 'pre-render' can then use the time and the two 'render states' to interpolate between update frames and push that data to the render thread.

The interpolation between states would have to be done somewhere and chances are it is a non-trival task, so splitting it up over the 'pre-render' threads seems like a good idea.

I admit, all of this is untested over N threads but I've done something like it over 2 threads (although in that case I believe the render thread handled interpolation) and I see no obvious reason it couldn't work out.

the_edd
the_edd
Quote:
Original post by hornet1990
The simulation thread could itself be a task that when completed simply adds itself again to the task queue for the next iteration.


That's exactly what happens in our code. Something like:

class Simulator{    public:        void SingleStep();    //...};MessageQueue<Simulator> q(MakeASimulator(), GetPrimaryThreadPool());RepeatedMessage<Simulator> simLoop(q, boost::bind(&Simulator::SingleStep, _1));simLoop.Start();


The RepeatedMessage wraps the functor (created by boost::bind, here) in such a way so that it re-submits itself once finished.

Quote:

So as you both quite rightly say that reduces it to a rendering thread and 1..n task threads - with n being n-cores available - 1.

Or maybe n-cores available - 2, depending on whether or not you use the "main thread" to coordinate things...

Quote:

So if I'm reading this right you both decide what to render and where as part of your main processing task, and then pass that off to the render thread to actually do it (which is a variation of the first option).

Pretty much, yeah. Though in our case, the rendering can be triggered in two ways - either through a simulation step finishing, or as a result of user-interaction through the GUI.

Quote:

Must admit though I was thinking in terms of a fixed timestep for the main processing task which would also imply a fixed issuing rate of rendering instructions, hence why I was looking to completely decouple rendering.

Well, I'm not sure in what sense you're using the word "decouple", but our renderer and simulator know nothing about one another. It's certainly possible to keep coupling down in that sense.

Quote:
WRT the rendering thread being idle I did think that maybe it could be a specialised task processing thread so that when idle it steals tasks from another threads queue (as opposed to the task processing threads which have task assigned to their queues). It would solve that problem unless the task it picks up is more involved than the remaining frame time available...


Right, but that wouldn't be much different to having the render thread process tasks like all the other threads and occasionally putting a "high priority" rendering task in to its queue i.e. sticking it at the front of the queue, rather than the back.

Personally, I think having a dedicated render thread that idles a lot of the time is /probably/ going to be ok (again, I should stress that this is new to me, too). There are probably hundreds of other threads idling on your machine at any given time, anyway. A quick look in my task manager shows around 400 threads.
_the_phantom_
_the_phantom_
Quote:
Original post by the_edd
Quote:

So as you both quite rightly say that reduces it to a rendering thread and 1..n task threads - with n being n-cores available - 1.

Or maybe n-cores available - 2, depending on whether or not you use the "main thread" to coordinate things...


I also considered n-2, however if you allow for the 'main' thread to be sleeping having dispatched the tasks you'd just end up wasting a core/thread again.
the_edd
the_edd
Quote:
Original post by phantom
Quote:
Original post by the_edd
Quote:

So as you both quite rightly say that reduces it to a rendering thread and 1..n task threads - with n being n-cores available - 1.

Or maybe n-cores available - 2, depending on whether or not you use the "main thread" to coordinate things...


I also considered n-2, however if you allow for the 'main' thread to be sleeping having dispatched the tasks you'd just end up wasting a core/thread again.


Yeah that's true. We have a system where any waiting call will call ThreadPool::Help() repeatedly until the thing it's waiting for is ready.
hornet1990
hornet1990
Quote:
Original post by the_edd
Well, I'm not sure in what sense you're using the word "decouple", but our renderer and simulator know nothing about one another. It's certainly possible to keep coupling down in that sense.

As you've described it though your renderer needs to be fed by either your simulation step or gui so although they know nothing of each other there is still that dependancy.

Options 2 and 3 I suggested would be completely seperate running its own loop on the render thread which continually renders the world state as it is at that time, whilst the simulation thread/tasks are updating that world state independantly.

Of course thats where the double buffering of state so that reads are consistent from any thread comes in handy.

I wonder if anyone here has tried this way and what their experiences of it were?

Quote:
Original post by the_edd
(again, I should stress that this is new to me, too).


Its all good fun isn't it?! [lol]
_the_phantom_
_the_phantom_
Quote:
Original post by hornet1990
Quote:
Original post by the_edd
Well, I'm not sure in what sense you're using the word "decouple", but our renderer and simulator know nothing about one another. It's certainly possible to keep coupling down in that sense.

As you've described it though your renderer needs to be fed by either your simulation step or gui so although they know nothing of each other there is still that dependancy.


Well, there is always going to be a 'dependancy' of sorts, granted if the renderer thread does everything [todo with interpolation of states] then the dependancy is just moved elsewhere and to a different time but it still exists.

Also, not only does it still exist but you are now putting an extra burden on your single rendering thread to do more work, work which could be trivally done in parallel, while huge chunks of processing time go unused (your n task threads). So, apart from those moments when your app is processing a world update you are back to being single threaded again.

Quote:

Options 2 and 3 I suggested would be completely seperate running its own loop on the render thread which continually renders the world state as it is at that time, whilst the simulation thread/tasks are updating that world state independantly.

Of course thats where the double buffering of state so that reads are consistent from any thread comes in handy.


The problem with option 2 and 3 is that option 3 effectively singlethreads your app when you aren't updating the world and option 2 starts stomping over your update loops time as well.

If you consider that you should be rendering your last frame while updating the next this means your 'tasks to setup rendering' are stomping about doing things while your 'tasks to update the world' are trying to get on with doing that.

The problem here is that you are going to stall one system or another in some way;
- If you just throw your tasks together in any old order then your render thread will stall while waiting for rendering task work to happen which is fighting with update tasks for cpu time, memory bandwidth and cpu cache.

- If you try to ensure rendering tasks are done first (or as soon as possible) then you might as well go with my 'pre-render' step method which gets them out of the way before rendering starts and doesn't start stalling out your update tasks. This removes potential stalls from rendering due to task congestion and helps things happen in a well defined order of operation while taking maximum advantage of cpu resources at any given moment.

the_edd
the_edd
Quote:
Original post by hornet1990
Quote:
Original post by the_edd
Well, I'm not sure in what sense you're using the word "decouple", but our renderer and simulator know nothing about one another. It's certainly possible to keep coupling down in that sense.

As you've described it though your renderer needs to be fed by either your simulation step or gui so although they know nothing of each other there is still that dependancy.


Well the Controller (as in MVC) depends on the GUI, the renderer and the Model (simulator). The latter three know nothing of one another.

The confusion was probably caused by my abbreviated example, though. The actual message that is sent to the simulator repeatedly is more like this:

Controller::SimulationStep(Simulator &s){    s.SingleStep();    renderer_.Send(bind(&Renderer::Draw, _1, s.GetScene()));}// ...RepeatedMessage<Simulator> simLoop(simulatorQueue, bind(&Controller::SimulationStep, _1));


It is the controller that manages the MessageQueues, RepeatedMessages and coordinates the messages.
hornet1990
hornet1990
Quote:

- If you try to ensure rendering tasks are done first (or as soon as possible) then you might as well go with my 'pre-render' step method which gets them out of the way before rendering starts and doesn't start stalling out your update tasks. This removes potential stalls from rendering due to task congestion and helps things happen in a well defined order of operation while taking maximum advantage of cpu resources at any given moment.

TBH my original (and prefered) plan was very similar to your method anyway but I figured I'd ask if anyone was using some other way that may also be viable - obviously not [grin]

Thanks for the feedback!

Out of interest would you mind giving us an overview of how the rendering side works with DX11's deferred contexts and your render bucket system?

Cheers
_the_phantom_
_the_phantom_
I would gladly do so... if I had written it yet [grin] I'm stuck very much in procrastination land right now..

However, it is likely to be extremely based on this blog post here which talks about ordering draw calls.

Effectively my plan is to use threads to create deferred drawing commands and feed said commands into buckets which can just simply be blased through by the renderer.
hornet1990
hornet1990
Quote:
Original post by phantom
I would gladly do so... if I had written it yet [grin] I'm stuck very much in procrastination land right now..

lol, I know that place like the back of my hand!

Quote:

However, it is likely to be extremely based on this blog post here which talks about ordering draw calls.

Effectively my plan is to use threads to create deferred drawing commands and feed said commands into buckets which can just simply be blased through by the renderer.

A while back when I started looking at multi-threaded rendering I was planning on using something very similar but since then I'm being swayed towards the dice method here.
_the_phantom_
_the_phantom_
Ah yes, I've looked at that presentation before now, mostly I was intrested in the deferred rendering section in it however but I do recall skimming the task stuff (mostly going 'wtf?' at the task dependancy graph *chuckles*)

I think there are simularities between the two methods, step 3 & 4 was basically where I was going with my idea of having the pre-render step throw everything into a render queue and then have the renderer blast through it in one go.

The simularities lay in the layering segment where the 'order your draw calls' mentions it; I think the key difference is I was going to pass the draw calls down to the renderer and have that hit a sorted list of them however I think having looked at that again their method might well be better;

If anything pre-render becomes a two step process;
- bucket renderables into layers (using the same setup as 'order your draw calls'?)
- split per-layer task out to create command lists

With the two parses operating one after each other; the final output being a command list queue which can be again blasted to the screen.

Thanks for reminding me about that article [grin]
Hodgman
Hodgman
Another version I've seen is a single-threaded scene-graph, which issues high-level draw commands to a thread-pool, which assembles buckets of state-changes/draw-calls/etc (in parallel) to send back to a single-threaded renderer.
Quote:
Original post by hornet1990
The simulation thread could itself be a task that when completed simply adds itself again to the task queue for the next iteration.
I started out with repeating-tasks like this, and a thread-pool/task-queue, but I wanted to get rid of the task-queue because it potentially causes some thread contention at the start of every task.
Queues are bad for cache, but read-only sequences are good =D
So instead of submitting tasks to a queue, I structure my tasks like phantom, where the repeating-tasks that form the main game loop are stored in a fixed sequence. Every thread in the pool then executes every task in the sequence.
Each task is a single functional block of code, like Update or Render, but can contain multiple "data items" to be computed.

When executing a task, each thread independently determines a begin/end range into the task's "data items" array, based on the thread-index and thread-count. It can then computes that range of the task, safe from overlap with other threads. Only when a task has been executed by every thread, will the whole range (i.e. every data item) be guaranteed to have been processed.

Single threaded tasks, like GL/D3D or other APIs, can simply use the thread-index to choose to do all the work or none, instead of choosing a range of work to complete. If the thread-index is 0, you're on the main thread and you run, otherwise you don't.
This knocks out the balance of your thread pool, but if you measure slack time per frame, you can determine per-frame-weights to feed into the above thread-index/thread-count partitioning algorithm.
The slack time from an imbalanced thread pool will also get eaten up by your background threads, like audio or file-system.

Even if you're using a single-threaded rendering device, the renderer could still have a "Render" task, which builds commands in parallel, and a "KickRender" tasks, which only runs on the main thread (and issues the commands to the device).

Some tasks have data dependencies on other tasks, for example, I've got an "Update" task (which runs a simulation) and a "Commit" task (which copies state for double-buffering). "Commit" can't run until "Update" has been executed by every thread. To handle this, put "Commit" as far after "Update" in the sequence as possible (so it's likely that update will be done), plus increment an atomic-counter every time "Update" is executed, and spin-wait (+yield after the first n spins) at the start of "Commit" until the counter reaches thread-count. The spin loop will only kick in if there's crazy thread imbalance, or the two tasks are placed too closely together in the task sequence. If it does kick in, the yield will allow your background threads to use some CPU time.
This atomic counter could cause contention though, you can avoid it if you want by incrementing a thread-local counter and testing every thread's private counter in the spin loop (if you do this, you'll need memory fences around the usage of the counters).

When your task execution is in a strict order like this, with parallel overlap where it doesn't matter, but restricted where you define data dependencies, you can do all your task communication using wait-free queues (instead of lock-free ones). Instead of one shared queue, you create an array of queues (one for each thread in the pool). When writing (e.g. in Update) writes are done to queue[thread_index], when reading (e.g. in Commit) reads are done from every queue. The scheduling of the task-sequence and task dependencies lets wait-free usurp lock-based (and lock-free) concurrency.
Quote:
Original post by hornet1990
As far as the entity system is concerned I'm going for a "pure" system where each component is just data, and each component type has a system/manager which contains the functionality. Each components data will also be double buffered so that reads are consistent, and writes will be made live with a sync phase after all processing by all systems is done.
In the set-up I've described, each manager would have two tasks, and the components form the "data item" array. An Update task, which writes to the private buffers, and a Commit tasks which copies the private data to the public buffers.
Your manager's tasks would have to be scheduled so that none of the Commit tasks can begin until all of the Update tasks have completed. I implement dependencies at the "task-group" level to accommodate this.

[Edited by - Hodgman on January 25, 2010 8:49:55 AM]

Topic Locked

This topic has been locked by a moderator. New replies are not allowed.

Sign in to reply to this topic.