Skip to main content
GameDev.net gamedev.net
🔒 Locked

Need for CPU (vs just compiler) barriers on x86(-64)

Started by Prune Dec 16, 2010 at 6:43 PM 4 replies 5.4k views
Original Post
Prune
Prune
It is my understanding that on x86 and x68-64 the only reordering is that loads may be reordered with older stores to different locations.
With this in mind, I'm wondering if a compiler barrier is sufficient in Linux if one were to port Microsoft's LockFreePipe: http://dx9-occlusion-sample.googlecode.com/svn-history/r43/trunk/DXUT/Optional/DXUTLockFreePipe.h
The variables in question are declared volatile here, and MSVC will insert CPU barriers in the cases of volatile as needed (it's a well known non-standard treatment of volatile by Microsoft). However, this is not necessarily the case with gcc, so I'm wondering if an explicit CPU barrier is necessary here, or I can just replace the _ReadWriteBarrier MSVC compiler barrier with the __asm__ __volatile__("" ::: "memory") gcc compiler barrier and that along with the restrictions on reordering of x86 and x86-64 will be sufficient. I certainly wouldn't want to insert a CPU barrier that will cost ~100 cycles if it's not necessary.

As an aside, does an SSE intrinsic barrier _mm_mfence act as only a barrier to SSE loads/stores, or a general memory barrier to all loads/stores?

[Edit:] In Intel's Threaded Building Blocks, the load with acquire and store with release functions they have use casts to volatile together with compiler barriers only, no CPU barriers, for x86 and x86-64...

Another aside is, the code I linked to uses DWORD but what about portability? They're doing pointer arithmetic with it. Should maybe uintptr_t be used instead?

[Edited by - Prune on December 16, 2010 7:41:48 PM]
"But who prays for Satan? Who, in eighteen centuries, has had the common humanity to pray for the one sinner that needed it most?" --Mark Twain

~~~~~~~~~~~~~~~Looking for a high-performance, easy to use, and lightweight math library? http://www.cmldev.net/ (note: I'm not associated with that project; just a user)
outRider
outRider
Plain volatile is sufficient for x86 and x86-64, you don't need any actual sync instructions.
Hodgman
Hodgman
On MSVC, volatile only prohibits the compiler from reordering the volatile read/write with regards to other volatile variable, or to global variables.

If you write data to a member, and then publish it by setting a volatile flag, the compiler may re-order those writes, which is why the file you linked to informs the compiler via the _ReadWriteBarrier intrinsic.

[edit]That was more to clarify outRider's post than to respond to the OP

[edit #2] uintptr_t sounds like a good idea.

[edit #3] _mm_mfence will issue a general CPU fence instruction. On x86 will likely be a dummy load instruction with the LOCK prefix (aka an InterlockedExchange).
Prune
Prune
I'm having trouble coming up with a natural example of when the CPU barrier is needed. I've also noticed an interesting compiler-and-wordsize-dependent variation of the CPU barrier in Intel's TBB code:

Windows 32
Intel C++ and MSVC: __asm { __asm mfence }

Windows 64
Intel C++: __asm { __asm mfence }
MSVC: _mm_mfence()

Linux 32 and Linux 64
Intel C++ and GCC: __asm__ __volatile__("mfence": : :"memory")

It's really interesting why MSVC under 64 bit is _mm_mfence whereas the rest of Windows cases are asm mfence.

One issue I haven't seen addressed anywhere is the semantics of volatile when code around it gets auto-vectorized. SSE can get reordered on x86(-64) so I am wondering if it is possible that when the compiler vectorizes some code, some of the assumptions in code such as the one discussed here that make omission of CPU barriers OK become invalidated and might result in incorrectly functioning compiled code...
"But who prays for Satan? Who, in eighteen centuries, has had the common humanity to pray for the one sinner that needed it most?" --Mark Twain

~~~~~~~~~~~~~~~Looking for a high-performance, easy to use, and lightweight math library? http://www.cmldev.net/ (note: I'm not associated with that project; just a user)
_the_phantom_
_the_phantom_
That's probably down to MS's x64 compiler not allowing inline assembly to be used; you have to use intrinsics or write the whole function externally as assembly
Prune
Prune
Quote:
Original post by phantom
That's probably down to MS's x64 compiler not allowing inline assembly to be used; you have to use intrinsics or write the whole function externally as assembly

Given, this, is there any reason not to use the intrinsic in the 32-bit MSVC case as well (and even for the Intel Compiler cases, as ICL reads this intrinsic fine)? I somehow doubt that there would be any performance impact, as intrinsics expand to non-function call assembly anyway.
"But who prays for Satan? Who, in eighteen centuries, has had the common humanity to pray for the one sinner that needed it most?" --Mark Twain

~~~~~~~~~~~~~~~Looking for a high-performance, easy to use, and lightweight math library? http://www.cmldev.net/ (note: I'm not associated with that project; just a user)

Topic Locked

This topic has been locked by a moderator. New replies are not allowed.

Sign in to reply to this topic.