Sign in to follow this  
ziashan

OpenGL glReadPixels() Performance Issue

Recommended Posts

ziashan    108
I am taking screen captures of an OpenGL application in real-time using glReadPixels() and experiencing a performance issue. I've followed the advice provided here:
[url="http://www.gamedev.net/topic/351860-taking-screenshots-using-glreadpixels/"]http://www.gamedev.n...g-glreadpixels/[/url]

and here:
[url="http://www.gamedev.net/topic/575590-real-time-opengl-screen-capture/"]http://www.gamedev.n...screen-capture/[/url]

where samoth had the excellent suggestion to use PBO's, so I implemented his approach based on the information provided here:
[url="http://www.songho.ca/opengl/gl_pbo.html"]http://www.songho.ca...ngl/gl_pbo.html[/url]

This works great on a Windows 7 system with AMD FirePro M5950 graphics and a very old Ubuntu 12.04 system with Radeon RV250 [Mobility FireGL 9000] graphics, but chokes on new Ubuntu 12.04 system with integrated Intel graphics. glReadPixels() takes >600ms to return even though (as I understand it) it should return immediately due to the asynchronous PBO use. I'm using GLEW for the OpenGL extensions and it appears that the graphics drivers support the GL_ARB_pixel_buffer_object extension. Any ideas? somoth mentioned that supplying wrong PBO flags can cause stalls, but i don't know what those might be.

Share this post


Link to post
Share on other sites
Kaptein    2224
"but chokes on new Ubuntu 12.04 system with integrated Intel graphics"
Maybe i'm jaded, but whenever i encounter the words intel and integrated, I move on, and never look back :)

they aren't very strong GPUs, they're full of driver issues and implementation variances.. sometimes it will outright lie to you (glGet*)
however! Newer intel integrated chips have orders of magnitude better openGL support! all the way up to 4.1 from what ive read!
So, personally, I will be waiting until "a larger percentage" of my potential audience has newer intel chips, and continue to pretend the older ones never existed :)

This post is only a warning, that if you really want to support the older intel integrated chips, you are in for a potential nightmare...
I'm sure there are plenty of people who can help you, since many if not most people who aren't making hobbyist games, will choose to include the widest audience possible =)

edit: to actually be helpful,
what happens when you supply an opengl implementation with something it doesn't support may cause it to run it in software mode
which is many orders of magnitude slower than hardware accelerated paths =)
for example, i imagine that if you tried GL_RGBA16 or GL_RGBA16F, it would do it all on the cpu
it may also not support RGB8, but i can't answer concretely what it can and can't do, and i'm unsure if you can ask the driver whether or not it's emulating it or not =)

http://www.opengl.org/wiki/Common_Mistakes#Texture_upload_and_pixel_reads
note that using GL_RGBA is not a "real format" because it doesn't describe internal precision, using GL_RGBA8 was determined to be the fastest.. i was unable to find any information for intel cards
Edited by Kaptein

Share this post


Link to post
Share on other sites
Scourage    1198
I would double check your format of your PBO and make sure that it matches the frame buffer that you're reading from. I've noticed significant performance hits if openGL needs to convert from the framebuffer's format to the PBO's format. Sometimes when you request a framebuffer format, you don't always get what you asked for.

Cheers,

Bob

Share this post


Link to post
Share on other sites
ziashan    108
[quote name='Kaptein' timestamp='1354907032' post='5008194']
Maybe i'm jaded, but whenever i encounter the words intel and integrated, I move on, and never look back
[/quote]

I have to agree with you, the intel graphics have been a real thorn in my side. Unfortunately i'm stuck with them. I tried changing the GL_PACK/UNPACK_ALIGNMENT? and different formats (e.g. GL_RGBA8), but still no luck.

Share this post


Link to post
Share on other sites
BornToCode    1185
You use Unpack whenever you are sending information to the Card. You use Pack whenever you want to read something from the card. One of the way you can speed it up is to have two FBOs. That way wile you are saving one FBO you can start writing to the one at the same time.

Share this post


Link to post
Share on other sites

Create an account or sign in to comment

You need to be a member in order to leave a comment

Create an account

Sign up for a new account in our community. It's easy!

Register a new account

Sign in

Already have an account? Sign in here.

Sign In Now

Sign in to follow this  

  • Similar Content

    • By markshaw001
      Hi i am new to this forum  i wanted to ask for help from all of you i want to generate real time terrain using a 32 bit heightmap i am good at c++ and have started learning Opengl as i am very interested in making landscapes in opengl i have looked around the internet for help about this topic but i am not getting the hang of the concepts and what they are doing can some here suggests me some good resources for making terrain engine please for example like tutorials,books etc so that i can understand the whole concept of terrain generation.
       
    • By KarimIO
      Hey guys. I'm trying to get my application to work on my Nvidia GTX 970 desktop. It currently works on my Intel HD 3000 laptop, but on the desktop, every bind textures specifically from framebuffers, I get half a second of lag. This is done 4 times as I have three RGBA textures and one depth 32F buffer. I tried to use debugging software for the first time - RenderDoc only shows SwapBuffers() and no OGL calls, while Nvidia Nsight crashes upon execution, so neither are helpful. Without binding it runs regularly. This does not happen with non-framebuffer binds.
      GLFramebuffer::GLFramebuffer(FramebufferCreateInfo createInfo) { glGenFramebuffers(1, &fbo); glBindFramebuffer(GL_FRAMEBUFFER, fbo); textures = new GLuint[createInfo.numColorTargets]; glGenTextures(createInfo.numColorTargets, textures); GLenum *DrawBuffers = new GLenum[createInfo.numColorTargets]; for (uint32_t i = 0; i < createInfo.numColorTargets; i++) { glBindTexture(GL_TEXTURE_2D, textures[i]); GLint internalFormat; GLenum format; TranslateFormats(createInfo.colorFormats[i], format, internalFormat); // returns GL_RGBA and GL_RGBA glTexImage2D(GL_TEXTURE_2D, 0, internalFormat, createInfo.width, createInfo.height, 0, format, GL_FLOAT, 0); glTexParameteri(GL_TEXTURE_2D, GL_TEXTURE_MAG_FILTER, GL_NEAREST); glTexParameteri(GL_TEXTURE_2D, GL_TEXTURE_MIN_FILTER, GL_NEAREST); DrawBuffers[i] = GL_COLOR_ATTACHMENT0 + i; glBindTexture(GL_TEXTURE_2D, 0); glFramebufferTexture(GL_FRAMEBUFFER, GL_COLOR_ATTACHMENT0 + i, textures[i], 0); } if (createInfo.depthFormat != FORMAT_DEPTH_NONE) { GLenum depthFormat; switch (createInfo.depthFormat) { case FORMAT_DEPTH_16: depthFormat = GL_DEPTH_COMPONENT16; break; case FORMAT_DEPTH_24: depthFormat = GL_DEPTH_COMPONENT24; break; case FORMAT_DEPTH_32: depthFormat = GL_DEPTH_COMPONENT32; break; case FORMAT_DEPTH_24_STENCIL_8: depthFormat = GL_DEPTH24_STENCIL8; break; case FORMAT_DEPTH_32_STENCIL_8: depthFormat = GL_DEPTH32F_STENCIL8; break; } glGenTextures(1, &depthrenderbuffer); glBindTexture(GL_TEXTURE_2D, depthrenderbuffer); glTexImage2D(GL_TEXTURE_2D, 0, depthFormat, createInfo.width, createInfo.height, 0, GL_DEPTH_COMPONENT, GL_FLOAT, 0); glTexParameteri(GL_TEXTURE_2D, GL_TEXTURE_MAG_FILTER, GL_NEAREST); glTexParameteri(GL_TEXTURE_2D, GL_TEXTURE_MIN_FILTER, GL_NEAREST); glBindTexture(GL_TEXTURE_2D, 0); glFramebufferTexture(GL_FRAMEBUFFER, GL_DEPTH_ATTACHMENT, depthrenderbuffer, 0); } if (createInfo.numColorTargets > 0) glDrawBuffers(createInfo.numColorTargets, DrawBuffers); else glDrawBuffer(GL_NONE); if (glCheckFramebufferStatus(GL_FRAMEBUFFER) != GL_FRAMEBUFFER_COMPLETE) std::cout << "Framebuffer Incomplete\n"; glBindFramebuffer(GL_FRAMEBUFFER, 0); width = createInfo.width; height = createInfo.height; } // ... // FBO Creation FramebufferCreateInfo gbufferCI; gbufferCI.colorFormats = gbufferCFs.data(); gbufferCI.depthFormat = FORMAT_DEPTH_32; gbufferCI.numColorTargets = gbufferCFs.size(); gbufferCI.width = engine.settings.resolutionX; gbufferCI.height = engine.settings.resolutionY; gbufferCI.renderPass = nullptr; gbuffer = graphicsWrapper->CreateFramebuffer(gbufferCI); // Bind glBindFramebuffer(GL_DRAW_FRAMEBUFFER, fbo); // Draw here... // Bind to textures glActiveTexture(GL_TEXTURE0); glBindTexture(GL_TEXTURE_2D, textures[0]); glActiveTexture(GL_TEXTURE1); glBindTexture(GL_TEXTURE_2D, textures[1]); glActiveTexture(GL_TEXTURE2); glBindTexture(GL_TEXTURE_2D, textures[2]); glActiveTexture(GL_TEXTURE3); glBindTexture(GL_TEXTURE_2D, depthrenderbuffer); Here is an extract of my code. I can't think of anything else to include. I've really been butting my head into a wall trying to think of a reason but I can think of none and all my research yields nothing. Thanks in advance!
    • By Adrianensis
      Hi everyone, I've shared my 2D Game Engine source code. It's the result of 4 years working on it (and I still continue improving features ) and I want to share with the community. You can see some videos on youtube and some demo gifs on my twitter account.
      This Engine has been developed as End-of-Degree Project and it is coded in Javascript, WebGL and GLSL. The engine is written from scratch.
      This is not a professional engine but it's for learning purposes, so anyone can review the code an learn basis about graphics, physics or game engine architecture. Source code on this GitHub repository.
      I'm available for a good conversation about Game Engine / Graphics Programming
    • By C0dR
      I would like to introduce the first version of my physically based camera rendering library, written in C++, called PhysiCam.
      Physicam is an open source OpenGL C++ library, which provides physically based camera rendering and parameters. It is based on OpenGL and designed to be used as either static library or dynamic library and can be integrated in existing applications.
       
      The following features are implemented:
      Physically based sensor and focal length calculation Autoexposure Manual exposure Lense distortion Bloom (influenced by ISO, Shutter Speed, Sensor type etc.) Bokeh (influenced by Aperture, Sensor type and focal length) Tonemapping  
      You can find the repository at https://github.com/0x2A/physicam
       
      I would be happy about feedback, suggestions or contributions.

    • By altay
      Hi folks,
      Imagine we have 8 different light sources in our scene and want dynamic shadow map for each of them. The question is how do we generate shadow maps? Do we render the scene for each to get the depth data? If so, how about performance? Do we deal with the performance issues just by applying general methods (e.g. frustum culling)?
      Thanks,
       
  • Popular Now