Skip to main content
GameDev.net gamedev.net
🔒 Locked

copy surface from GPU to CPU ?

Started by BartGallet May 12, 2003 at 4:19 PM 8 replies 1.5k views
Original Post
BartGallet
BartGallet
Hi All I am using DirectX9 to generate simple renderings which I then feed to some robot-vision algorithms. I am rendering two windows (a global overview and "what the robot sees" i.e. Robot POV). Besides showing those two views, I need to pass on the robot POV to my vision algorithms. So the way I go about this is to use GetBackBuffer from the swapchain, and then I use LockRect to get access to the raw data bytes of the robot POV image. But the problem is when I try to shuffle the bytes over from the backbuffer copy to my copy for the vision algorithms. It takes almost 200ms for a 640*480*4 byte image. Does anyone have a better solution on how I would be albe to shuffle over my Directx9 generated images to the PC side where I don''t have that bottle-neck? Cheers Bart :-)
BartGallet
BartGallet
I am using Nvidea Quadro 4 xlg. It supose to have 4X AGP at 32 bits. I need to access the image at 5 Hz, that is 640*480*4*5 bytes per second, or 6.144 MBytes per second. I don''t think the problem is hardware here.

I am new to DirectX (using version 9), and I am not sure that I am making correct use of getting access to the back buffer, which is lockable.

I just wonder why it takes so much time for something that should on my particular PC only take about 1ms per image, and not 200ms.
directrix
directrix
Sounds like your bottleneck is in reading the pixels from a backbuffer stored in video memory, which is slower than reading from system memory. Try putting the backbuffer in system memory. This will slow down your overall rendering speed a bit, but increase the speed at which your reading the pixels from the backbuffer.

Also, how are you copying the pixels? Avoid loops. Do a copy of the backbuffer using the appropiate DX 9 function. Also if the vision algorithms don''t modify the pixels of the backbuffer, avoid creating a copy of it.

digital radiation
BartGallet
BartGallet
Thanks for the advice

I do need to modify the format of the rendered image, so I am accessing the pixels individually in a loop. I have written the vision algorithms to only receive 8-bit gray scale. For that reason I am only copying over the green channel since it has the greatest weight.

Someone else adviced me to render to a texture. Are memory transfers faster for textures between the GPU and the PC memory?

I also got this article wich might explains that the drivers of graphic cards are not optmized for output (reading from GPU memory):

www.tech-report.com/etc/2002q3/agp-download/index.x?pg=1

Cheers

Bart
Donavon Keithley
Donavon Keithley
quote:
Try putting the backbuffer in system memory.


Erm... back buffers always exist in video memory. Theoretically I suppose the driver could allocate in AGP ("non-local video memory") but that''s not under your control and at any rate it sounds like a pretty wacky thing to do, considering how frequently frame buffer bandwidth is a performance issue.

quote:
Someone else adviced me to render to a texture. Are memory transfers faster for textures between the GPU and the PC memory?


I doubt it''s going to make much difference. The majority of the overhead is reading from local video memory to system memory, and the pokiness of this is a perennial complaint (along with perennial optimism that it''s going to get better Real Soon Now). Also realize that you have to wait for the rendering pipeline to flush before you can get at the resulting bits and that''s going to hurt performance.

The best you can do, IMO, is look for a hardware and driver combination that has relatively fast vid->sys DMA.
treething
treething
quote:
I doubt it's going to make much difference. The majority of the overhead is reading from local video memory to system memory


True, but assuming that driver optimisations can make a big difference to this kind of thing, its possible that the drivers are more optimised for reading from textures, because I would expect thats a more common occurence.

At the end of the day you're really at the mercy of the hardware and drivers, so unless you find a way around the problem of having to read from vid-mem, IMO its worth playing around to find a code-path thats more optimised in the driver.



[edited by - treething on May 13, 2003 6:12:19 AM]
S1CA
S1CA

quote:
I do need to modify the format of the rendered image, so I am accessing the pixels individually in a loop. I have written the vision algorithms to only receive 8-bit gray scale. For that reason I am only copying over the green channel since it has the greatest weight.


If possible use one of the Copy type APIs to copy the whole piece of video memory (preferably without locking) to a system memory texture and then lock that.

CPU access to true video memory or AGP memory isn''t cached by the CPU, so every DWORD you read costs you a lot in CPU cycles - that''s on top of real bus bandwidth issues (for true video memory). IIRC a read from uncached memory (i.e. video memory or AGP memory) is in the region of 40 times slower than a read from cached memory (i.e. a system memory texture)!. <wonders if Herb is reading and has an online link to his GDC1999 talk>


quote:
Someone else adviced me to render to a texture. Are memory transfers faster for textures between the GPU and the PC memory?


If the texture is in AGP memory (which implies the chip probably also supports textures with the DYNAMIC flag), then yes. If the texture is in on-card video memory, then no. Or at least no-ish - if that texture isn''t being used for anything apart from the view to send back to the PC, then write to it first then do the main view - the chip will do 2x as much rendering, but you''ll avoid serialising the whole pipeline which will help. It depends how much of a bottleneck you get from your rendering.



quote:
I also got this article wich might explains that the drivers of graphic cards are not optmized for output (reading from GPU memory):


Yup, becase the majority of applications (games, visualisation etc) never ever need to do that so it''s a very low priority on the TODO lists of all the hardware vendors.

It''s probably worth getting in touch with Stephan Schaem (you''ll find him on some of the common DirectX mailing lists etc). He''s been the most vocal campaigner for getting read backs from video memory accelerated/fixed in drivers (it''s something he requires due to his work). He''ll be the best to talk to IMO for an update on the current "state of play" (i.e. which drivers are good, which aren''t, which companies are planning support etc).


Can you fix the hardware platform this runs on (i.e. you can tell the customer "use _THIS_ PC") or is this expected to work on anything with a graphics card (then wait for someone to run it with a PCI card and complain) ?

If you can fix the platform, you could specify one of the drivers where read backs (and dynamic textures) ARE fairly fast (though that still doesn''t save you from the CPU costs - but might make one of the block copy ops to system memory faster).

It might also be a good idea to play around with a UMA motherboard such as one of the nForces - unified memory means no AGP bus to go over - though I can''t remember what the situation with locked memory is (i.e. if you can get a cached "view" or not)


--
Simon O''Connor
Creative Asylum Ltd
www.creative-asylum.com
Simon O'Connor | Technical Director (Newcastle) Lockwood Publishing | LinkedIn | Personal site
BartGallet
BartGallet
Hi all

thanks for all the replies.

Simon,

my customer is only interested in my vision algorithms, so I am very flexible on what platform I perform my development. In the end, it has to run in an embedded system. Directx is just a means to generate fast images to exercise and develop robot vision for autonomous vehicles.

So I have the choice on what hardware I am doing this. As it happens, I have a dual processor machine where I am developing this, so right now this problem is not holding me back, but it will soon. The bottle-neck increases my CPU load by 40%.

I appreciate your ideas on the hardware (UMA). I need to look into that.

Cheers

Bart
BartGallet
BartGallet
leaving the fact that copying data from the GPU to the CPU is a bottle-neck, I do have enough processing power to make things run with a dual processor board.

However, I do have another complication. I have two renderviews. One of an observers point of view, looking at the robotic vehicle, and one view with the POV of the robotic vehicle. It is the latter view that I need to send to my vision software at 5Hz.

I have two threads, one where I render the observer''s view at 20Hz, and the other which renders the robot''s POV. I am sharing the Directx device thru a critical region. All works fine when I just render the view, and I don''t do any image copy of the robot''s view from GPU to CPU.

However, since that my observer''s view is running at 4 times the speed as my robot''s view, I want the robot''s view rendering kept at minimum, such that the observer''s view doesn''t have to wait to long when the robot''s view rendering is using the Directx Device. At least, not noticable-long.

So I want to do the copying of the robot''s view to the CPU memory after I leave the critical region (i.e. after I have presented my robot''s view).

The robot''s view is rendered on an aditional swap chain. And initially I had something like this pseudo code:

1 set render code to swapchain back buffer
2 render scene
3 copy back buffer
4 present
4 access copied backbuffer image bytes, and copy over to CPU

but that doesn''t work. I assume that after presenting the backbuffer to the front, I cannot access that data anymore.

So I modified my code like this
1 create offscreenplainbuffer
2 set render target to offscreen buffer
3 render
4 copy the offscreen buffer to backbuffer using UpdateSurface
5 present
6 access offscreen buffer image bytes, and copy over to CPU

But that doesnt work either, the image I created is all black.

Is my logic flawed? did I miss something?

Cheers

Bart

Topic Locked

This topic has been locked by a moderator. New replies are not allowed.

Sign in to reply to this topic.