SIMD Visualisers
Interactive visualisers for x86 SIMD instructions
PSHUFB / PSHUFW / PSHUFD / PSHUFLW / PSHUFHW
Integer shuffles — byte, word, dword (per 128-bit lane, mask or imm8 controlled)
SSE / SSE2 / SSSE3 / AVX2 / AVX-512
MMX / XMM / YMM / ZMM
→
SHUFPS / SHUFPD
Float shuffles — single / double precision, two source registers, imm8 controlled
SSE / SSE2 / AVX / AVX-512F
XMM / YMM / ZMM
→
VSHUFF32x4 / VSHUFF64x2 / VSHUFI32x4 / VSHUFI64x2
128-bit lane shuffles — pick whole 128-bit lanes from two sources, imm8 controlled, k-mask supported
AVX-512F / AVX-512F+VL
YMM / ZMM
→
VPBROADCASTB/W/D/Q / VBROADCASTSS/SD / VBROADCAST*x*
Broadcasts — replicate one element or a 128-/256-bit unit across the destination, k-mask supported
AVX / AVX2 / AVX-512
XMM / YMM / ZMM
→
VPMOVSX / VPMOVZX (BW/BD/BQ/WD/WQ/DQ)
Sign / zero extension — widen each packed byte/word/dword element to a larger element, k-mask supported
AVX / AVX2 / AVX-512
XMM / YMM / ZMM
→
PUNPCKLBW / PUNPCKLWD / PUNPCKLDQ / PUNPCKLQDQ
Unpack low — interleave the low halves of two sources per 128-bit lane, k-mask supported
SSE2 / AVX / AVX2 / AVX-512
XMM / YMM / ZMM
→
PUNPCKHBW / PUNPCKHWD / PUNPCKHDQ / PUNPCKHQDQ
Unpack high — interleave the high halves of two sources per 128-bit lane, k-mask supported
SSE2 / AVX / AVX2 / AVX-512
XMM / YMM / ZMM
→
PCMPEQB / PCMPEQW / PCMPEQD / PCMPEQQ
Packed equality — all-ones or all-zeros per element, or an opmask in the EVEX form
SSE2 / SSE4.1 / AVX / AVX2 / AVX-512
XMM / YMM / ZMM
→
PMOVMSKB / VPMOVMSKB
Byte sign mask — pack the top bit of every byte into a general-purpose register
SSE2 / AVX / AVX2
XMM / YMM → r32
→
PINSRB / PINSRD / PINSRQ
Insert — copy one byte/dword/qword from a general-purpose register or memory into the XMM slot picked by imm8
SSE4.1 / AVX / AVX-512BW / AVX-512DQ
r32 / r64 / mem → XMM
→
PCMPISTRI / PCMPESTRI
String compare — the imm8-controlled comparison matrix, its aggregation into IntRes1/IntRes2, and the index it puts in ECX
SSE4.2
XMM → ECX + flags
→
VPTERNLOGD / VPTERNLOGQ
Ternary logic — any 3-input Boolean function, bit by bit, with imm8 as its truth table; k-mask and broadcast supported
AVX-512F (+VL)
XMM / YMM / ZMM
→
VPCMPB/W/D/Q / VPCMPUB/UW/UD/UQ
Predicate compares into an opmask — imm8 picks EQ/LT/LE/FALSE/NEQ/NLT/NLE/TRUE, signed or unsigned
AVX-512F / AVX-512BW (+VL)
XMM / YMM / ZMM
→
KSHIFTLB/W/D/Q / KSHIFTRB/W/D/Q
Opmask shifts — move the mask bits by imm8, shifting zeros in and dropping what falls off
AVX-512F / AVX-512DQ / AVX-512BW
k0–k7
→
KUNPCKBW / KUNPCKWD / KUNPCKDQ
Opmask concatenation — glue two masks into one twice as wide; no interleaving despite the name
AVX-512F / AVX-512BW
k0–k7
→