Original Post
It is my understanding that on x86 and x68-64 the only reordering is that loads may be reordered with older stores to different locations.
With this in mind, I'm wondering if a compiler barrier is sufficient in Linux if one were to port Microsoft's LockFreePipe: http://dx9-occlusion-sample.googlecode.com/svn-history/r43/trunk/DXUT/Optional/DXUTLockFreePipe.h
The variables in question are declared volatile here, and MSVC will insert CPU barriers in the cases of volatile as needed (it's a well known non-standard treatment of volatile by Microsoft). However, this is not necessarily the case with gcc, so I'm wondering if an explicit CPU barrier is necessary here, or I can just replace the _ReadWriteBarrier MSVC compiler barrier with the __asm__ __volatile__("" ::: "memory") gcc compiler barrier and that along with the restrictions on reordering of x86 and x86-64 will be sufficient. I certainly wouldn't want to insert a CPU barrier that will cost ~100 cycles if it's not necessary.
As an aside, does an SSE intrinsic barrier _mm_mfence act as only a barrier to SSE loads/stores, or a general memory barrier to all loads/stores?
[Edit:] In Intel's Threaded Building Blocks, the load with acquire and store with release functions they have use casts to volatile together with compiler barriers only, no CPU barriers, for x86 and x86-64...
Another aside is, the code I linked to uses DWORD but what about portability? They're doing pointer arithmetic with it. Should maybe uintptr_t be used instead?
[Edited by - Prune on December 16, 2010 7:41:48 PM]
With this in mind, I'm wondering if a compiler barrier is sufficient in Linux if one were to port Microsoft's LockFreePipe: http://dx9-occlusion-sample.googlecode.com/svn-history/r43/trunk/DXUT/Optional/DXUTLockFreePipe.h
The variables in question are declared volatile here, and MSVC will insert CPU barriers in the cases of volatile as needed (it's a well known non-standard treatment of volatile by Microsoft). However, this is not necessarily the case with gcc, so I'm wondering if an explicit CPU barrier is necessary here, or I can just replace the _ReadWriteBarrier MSVC compiler barrier with the __asm__ __volatile__("" ::: "memory") gcc compiler barrier and that along with the restrictions on reordering of x86 and x86-64 will be sufficient. I certainly wouldn't want to insert a CPU barrier that will cost ~100 cycles if it's not necessary.
As an aside, does an SSE intrinsic barrier _mm_mfence act as only a barrier to SSE loads/stores, or a general memory barrier to all loads/stores?
[Edit:] In Intel's Threaded Building Blocks, the load with acquire and store with release functions they have use casts to volatile together with compiler barriers only, no CPU barriers, for x86 and x86-64...
Another aside is, the code I linked to uses DWORD but what about portability? They're doing pointer arithmetic with it. Should maybe uintptr_t be used instead?
[Edited by - Prune on December 16, 2010 7:41:48 PM]