git.karo-electronics.de Git - karo-tx-linux.git/commit

author	Manfred Spraul <manfred@colorfullife.com>
	Thu, 27 Jun 2013 23:54:05 +0000 (09:54 +1000)
committer	Stephen Rothwell <sfr@canb.auug.org.au>
	Fri, 28 Jun 2013 06:39:03 +0000 (16:39 +1000)
commit	68257ae3f5159b1714b6da5a769ac3d9d6a7aec9
tree	99b0e6a9ca95146123ff8f91330263e8064ab85d	tree \| snapshot
parent	d765ce51d5718415c36df3a3c5579a2a7668fead	commit \| diff

ipc/sem.c: cacheline align the semaphore structures

As now each semaphore has its own spinlock and parallel operations are
possible, give each semaphore its own cacheline.

On a i3 laptop, this gives up to 28% better performance:

#semscale 10 | grep "interleave 2"
- before:
Cpus 1, interleave 2 delay 0: 36109234 in 10 secs
Cpus 2, interleave 2 delay 0: 55276317 in 10 secs
Cpus 3, interleave 2 delay 0: 62411025 in 10 secs
Cpus 4, interleave 2 delay 0: 81963928 in 10 secs

-after:
Cpus 1, interleave 2 delay 0: 35527306 in 10 secs
Cpus 2, interleave 2 delay 0: 70922909 in 10 secs <<< + 28%
Cpus 3, interleave 2 delay 0: 80518538 in 10 secs
Cpus 4, interleave 2 delay 0: 89115148 in 10 secs <<< + 8.7%

i3, with 2 cores and with hyperthreading enabled. Interleave 2 in order
use first the full cores. HT partially hides the delay from cacheline
trashing, thus the improvement is "only" 8.7% if 4 threads are running.

Signed-off-by: Manfred Spraul <manfred@colorfullife.com>
Cc: Rik van Riel <riel@redhat.com>
Cc: Davidlohr Bueso <davidlohr.bueso@hp.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>