Skip to main content
GameDev.net gamedev.net
🔒 Locked

ARM NEON

Started by ongamex92 Oct 8, 2013 at 3:05 PM 4 replies 1.5k views
Original Post
ongamex92
ongamex92

Currently I'm extending a math library (bullet math form Bullet phyisics to be exact). Im adding Matrix4x4 Vector4 and other helper classes/methods(3d graphics friendly methods)

Tesing with SSE is easy. But there is a ARM NEON implementation in Bullet and I dont want to break the compibility.

Is there any way to test ARM NEON SIMD on Windows?

Bregma
Bregma

I have regularly used qemu to emulate Cortex-A8 and later ARM devices under Linux (those devices are pretty much guaranteed to have vfs3 and NEON copros). I know qemu is also available for Microsoft Windows™. You could try that. or you could use the Google to find an alternative emulator of your choice.

Stephen M. Webb
Professional Free Software Developer
TheChubu
TheChubu

Even if they work I don't think an emulator is good for any sort of performance measurement really.

"I AM ZE EMPRAH OPENGL 3.3 THE CORE, I DEMAND FROM THEE ZE SHADERZ AND MATRIXEZ"   My journals: dustArtemis ECS framework and 
L. Spiro
L. Spiro

Right. The whole purpose in using NEON is performance, not compatibility. If your goal is compatibility, you are doing it wrong. It is specifically meant to be processor-specific. Convert the code to SSE* using macros to select which code path to follow.

L. Spiro

I restore Nintendo 64 video-game OST’s into HD! https://www.youtube.com/channel/UCCtX_wedtZ5BoyQBXEhnVZw/playlists?view=1&sort=lad&flow=grid
ongamex92
ongamex92

Well In the most cases I'm only porting the code form btMatrix3x3 and btVector3 to 4 dimension version of the classes.

I beleve in the people who wrote bullet math. Of course I will pay attention in the situations where the performance is not good enough.

I'm not really into NEON (just can't find good papers). I'm thinking that the most optimal version using NEON should look like the SSE version of the code, is that correct?

RobTheBloke
RobTheBloke

Well In the most cases I'm only porting the code form btMatrix3x3 and btVector3 to 4 dimension version of the classes.

I beleve in the people who wrote bullet math. Of course I will pay attention in the situations where the performance is not good enough.

I'm not really into NEON (just can't find good papers). I'm thinking that the most optimal version using NEON should look like the SSE version of the code, is that correct?

For basic ops like add/sub, then yes they should be similar. For dot/cross and other operations based around shuffles/unpack/movehl, they aren't going to look too similar at all....

Topic Locked

This topic has been locked by a moderator. New replies are not allowed.

Sign in to reply to this topic.