| Loop Id: 166 | Module: exec | Source: calc_dt_kernel.f90:93-129 [...] | Coverage: 3.43% |
|---|
| Loop Id: 166 | Module: exec | Source: calc_dt_kernel.f90:93-129 [...] | Coverage: 3.43% |
|---|
0x42a770 VMOVQ %XMM24,%RDX |
0x42a776 VDIVSD %XMM0,%XMM23,%XMM6 |
0x42a77c VXORPD %XMM22,%XMM6,%XMM3 |
0x42a782 VMULSD (%RDX),%XMM3,%XMM0 [6] |
0x42a786 LEA 0x1(%RAX),%RDX |
0x42a78a VMINSD %XMM2,%XMM0,%XMM2 |
0x42a78e VMINSD %XMM4,%XMM2,%XMM1 |
0x42a792 VMINSD %XMM1,%XMM11,%XMM11 |
0x42a796 CMP %RAX,-0x38(%RBP) [10] |
0x42a79a JE 42a8cb |
0x42a7a0 MOV %RDX,%RAX |
0x42a7a3 VMOVAPD %XMM5,%XMM6 |
0x42a7a7 VMOVAPD %XMM7,%XMM4 |
0x42a7ab VMOVAPD %XMM8,%XMM1 |
0x42a7af VMOVAPD %XMM9,%XMM2 |
0x42a7b3 VMOVSD (%R13,%RAX,8),%XMM29 [12] |
0x42a7bb VMOVSD (%R14,%RAX,8),%XMM5 [13] |
0x42a7c1 VADDSD %XMM29,%XMM29,%XMM28 |
0x42a7c7 VDIVSD (%R12,%RAX,8),%XMM28,%XMM29 [11] |
0x42a7ce VFMADD132SD %XMM5,%XMM29,%XMM5 |
0x42a7d4 VMOVSD 0x8(%R9,%RAX,8),%XMM3 [2] |
0x42a7db VMOVSD 0x8(%RDI,%RAX,8),%XMM7 [14] |
0x42a7e1 VUNPCKLPD %XMM4,%XMM7,%XMM4 |
0x42a7e5 VMINSD (%R15,%RAX,8),%XMM18,%XMM28 [4] |
0x42a7ec VMULSD %XMM17,%XMM28,%XMM29 |
0x42a7f2 VSQRTSD %XMM5,%XMM5,%XMM5 |
0x42a7f6 VMAXSD %XMM10,%XMM5,%XMM8 |
0x42a7fb VADDSD 0x8(%R8,%RAX,8),%XMM3,%XMM5 [1] |
0x42a802 VUNPCKLPD %XMM6,%XMM5,%XMM0 |
0x42a806 VMULPD %XMM0,%XMM4,%XMM9 |
0x42a80a VDIVSD %XMM8,%XMM29,%XMM29 |
0x42a810 VANDPD %XMM20,%XMM9,%XMM4 |
0x42a816 VADDSD %XMM19,%XMM9,%XMM8 |
0x42a81c VUNPCKHPD %XMM9,%XMM9,%XMM3 |
0x42a821 VSUBSD %XMM3,%XMM8,%XMM0 |
0x42a825 VUNPCKHPD %XMM4,%XMM4,%XMM28 |
0x42a82b VMOVSD (%RBX,%RAX,8),%XMM3 [15] |
0x42a830 VMOVSD 0x8(%RSI,%RAX,8),%XMM8 [5] |
0x42a836 VMULSD %XMM3,%XMM10,%XMM6 |
0x42a83a VMAXPD %XMM4,%XMM28,%XMM28 |
0x42a840 VMULSD %XMM15,%XMM3,%XMM9 |
0x42a845 VADDSD %XMM1,%XMM8,%XMM1 |
0x42a849 VMULSD (%R11,%RAX,8),%XMM1,%XMM1 [9] |
0x42a84f VMAXSD %XMM28,%XMM6,%XMM28 |
0x42a855 VDIVSD %XMM28,%XMM9,%XMM4 |
0x42a85b VMOVSD 0x8(%RCX,%RAX,8),%XMM9 [3] |
0x42a861 VADDSD %XMM2,%XMM9,%XMM2 |
0x42a865 VMULSD (%R10,%RAX,8),%XMM2,%XMM2 [8] |
0x42a86b VADDSD %XMM2,%XMM0,%XMM0 |
0x42a86f VANDPD %XMM12,%XMM2,%XMM2 |
0x42a874 VSUBSD %XMM1,%XMM0,%XMM0 |
0x42a878 VANDPD %XMM12,%XMM1,%XMM1 |
0x42a87d VMINSD %XMM29,%XMM4,%XMM4 |
0x42a883 VMULSD %XMM14,%XMM3,%XMM29 |
0x42a889 VADDSD %XMM3,%XMM3,%XMM3 |
0x42a88d VMAXSD %XMM2,%XMM1,%XMM1 |
0x42a891 VDIVSD %XMM3,%XMM0,%XMM0 |
0x42a895 VMAXSD %XMM6,%XMM1,%XMM6 |
0x42a899 VCOMISD %XMM0,%XMM16 |
0x42a89f VDIVSD %XMM6,%XMM29,%XMM2 |
0x42a8a5 JA 42a770 |
0x42a8ab VMOVQ %XMM25,%RDX |
0x42a8b1 VMINSD (%RDX),%XMM2,%XMM1 [7] |
0x42a8b5 LEA 0x1(%RAX),%RDX |
0x42a8b9 VMINSD %XMM4,%XMM1,%XMM4 |
0x42a8bd VMINSD %XMM4,%XMM11,%XMM11 |
0x42a8c1 CMP %RAX,-0x38(%RBP) [10] |
0x42a8c5 JNE 42a7a0 |
/home/eoseret/qaas_runs_ZEN5/174-049-0948/intel/CloverLeaf1.3-FC/build/CloverLeaf1.3-FC/CloverLeaf_ref/kernels/calc_dt_kernel.f90: 93 - 129 |
-------------------------------------------------------------------------------- |
93: !$OMP SIMD |
[...] |
99: cc=soundspeed(j,k)*soundspeed(j,k) |
100: cc=cc+2.0_8*viscosity_a(j,k)/density0(j,k) |
101: cc=MAX(SQRT(cc),g_small) |
102: |
103: dtct=dtc_safe*MIN(dsx,dsy)/cc |
104: |
105: div=0.0 |
106: |
107: dv1=(xvel0(j ,k)+xvel0(j ,k+1))*xarea(j ,k) |
108: dv2=(xvel0(j+1,k)+xvel0(j+1,k+1))*xarea(j+1,k) |
109: |
110: div=div+dv2-dv1 |
111: |
112: dtut=dtu_safe*2.0_8*volume(j,k)/MAX(ABS(dv1),ABS(dv2),g_small*volume(j,k)) |
113: |
114: dv1=(yvel0(j,k )+yvel0(j+1,k ))*yarea(j,k ) |
115: dv2=(yvel0(j,k+1)+yvel0(j+1,k+1))*yarea(j,k+1) |
116: |
117: div=div+dv2-dv1 |
118: |
119: dtvt=dtv_safe*2.0_8*volume(j,k)/MAX(ABS(dv1),ABS(dv2),g_small*volume(j,k)) |
120: |
121: div=div/(2.0_8*volume(j,k)) |
122: |
123: IF(div.LT.-g_small)THEN |
124: dtdivt=dtdiv_safe*(-1.0_8/div) |
125: ELSE |
126: dtdivt=g_big |
127: ENDIF |
128: |
129: dt_min_val=MIN(dt_min_val,dtct,dtut,dtvt,dtdivt) |
| Coverage (%) | Name | Source Location | Module |
|---|---|---|---|
| ►95.70+ | gomp_thread_start | team.c:130 | libgomp.so.1.0.0 |
| ○ | start_thread | libc.so.6 | |
| ○ | __GI___clone3 | libc.so.6 | |
| ►4.19+ | GOMP_parallel | libgomp.h:982 | libgomp.so.1.0.0 |
| ○ | calc_dt_kernel | calc_dt_kernel.f90:140 | exec |
| ○ | calc_dt | calc_dt.f90:81 | exec |
| ○ | timestep | timestep.f90:87 | exec |
| ○ | hydro | hydro.f90:54 | exec |
| ○ | main | clover_leaf.f90:41 | exec |
| ○ | __libc_start_call_main | libc.so.6 | |
| ○ | __libc_start_main | libc.so.6 | |
| ○ | _start | clover_leaf.f90:41 | exec |
| min | med | avg | max |
|---|---|---|---|
| Percentile Index | 10 | 20 | 30 | 40 | 50 | 60 | 70 | 80 | 90 | 100 |
|---|---|---|---|---|---|---|---|---|---|---|
| Value |
| min | med | avg | max |
|---|---|---|---|
| Percentile Index | 10 | 20 | 30 | 40 | 50 | 60 | 70 | 80 | 90 | 100 |
|---|---|---|---|---|---|---|---|---|---|---|
| Value |
| Path / |
| Metric | Value |
|---|---|
| CQA speedup if no scalar integer | 1.00 |
| CQA speedup if FP arith vectorized | 5.03 |
| CQA speedup if fully vectorized | 8.00 |
| CQA speedup if no inter-iteration dependency | NA |
| CQA speedup if next bottleneck killed | 2.87 |
| Bottlenecks | |
| Function | __calc_dt_kernel_module_MOD_calc_dt_kernel._omp_fn.0 |
| Source | calc_dt_kernel.f90:93-93,calc_dt_kernel.f90:99-129 |
| Source loop unroll info | not unrolled or unrolled with no peel/tail loop |
| Source loop unroll confidence level | max |
| Unroll/vectorization loop type | NA |
| Unroll factor | NA |
| CQA cycles | 30.50 |
| CQA cycles if no scalar integer | 30.50 |
| CQA cycles if FP arith vectorized | 6.06 |
| CQA cycles if fully vectorized | 3.81 |
| Front-end cycles | 7.50 |
| P0 cycles | 0.50 |
| P1 cycles | 0.50 |
| P2 cycles | 0.50 |
| P3 cycles | 0.67 |
| P4 cycles | 0.67 |
| P5 cycles | 0.67 |
| P6 cycles | 3.50 |
| P7 cycles | 3.50 |
| P8 cycles | 3.50 |
| P9 cycles | 3.50 |
| P10 cycles | 10.63 |
| P11 cycles | 10.63 |
| P12 cycles | 10.63 |
| P13 cycles | 10.63 |
| P14 cycles | 1.00 |
| P15 cycles | 1.00 |
| DIV/SQRT cycles | 30.50 |
| Inter-iter dependencies cycles | 2 |
| FE+BE cycles (UFS) | NA |
| Stall cycles (UFS) | NA |
| Nb insns | 59.50 |
| Nb uops | 60.00 |
| Nb loads | 14.00 |
| Nb stores | 0.00 |
| Nb stack references | 1.00 |
| FLOP/cycle | 0.85 |
| Nb FLOP add-sub | 9.00 |
| Nb FLOP mul | 8.50 |
| Nb FLOP fma | 1.00 |
| Nb FLOP div | 5.50 |
| Nb FLOP rcp | 0.00 |
| Nb FLOP sqrt | 1.00 |
| Nb FLOP rsqrt | 0.00 |
| Bytes/cycle | 3.69 |
| Bytes prefetched | 0.00 |
| Bytes loaded | 112.00 |
| Bytes stored | 0.00 |
| Stride 0 | 2.00 |
| Stride 1 | 0.00 |
| Stride n | 0.00 |
| Stride unknown | 12.00 |
| Stride indirect | 0.00 |
| Vectorization ratio all | 17.42 |
| Vectorization ratio load | 0.00 |
| Vectorization ratio store | NA |
| Vectorization ratio mul | 13.39 |
| Vectorization ratio add_sub | 0.00 |
| Vectorization ratio fma | 0.00 |
| Vectorization ratio div_sqrt | 0.00 |
| Vectorization ratio other | 36.14 |
| Vector-efficiency ratio all | 14.68 |
| Vector-efficiency ratio load | 12.50 |
| Vector-efficiency ratio store | NA |
| Vector-efficiency ratio mul | 14.17 |
| Vector-efficiency ratio add_sub | 12.50 |
| Vector-efficiency ratio fma | 12.50 |
| Vector-efficiency ratio div_sqrt | 12.50 |
| Vector-efficiency ratio other | 17.02 |
| Metric | Value |
|---|---|
| CQA speedup if no scalar integer | 1.00 |
| CQA speedup if FP arith vectorized | 4.70 |
| CQA speedup if fully vectorized | 8.00 |
| CQA speedup if no inter-iteration dependency | NA |
| CQA speedup if next bottleneck killed | 2.78 |
| Bottlenecks | P10, P11, |
| Function | __calc_dt_kernel_module_MOD_calc_dt_kernel._omp_fn.0 |
| Source | calc_dt_kernel.f90:93-93,calc_dt_kernel.f90:99-129 |
| Source loop unroll info | not unrolled or unrolled with no peel/tail loop |
| Source loop unroll confidence level | max |
| Unroll/vectorization loop type | NA |
| Unroll factor | NA |
| CQA cycles | 28.50 |
| CQA cycles if no scalar integer | 28.50 |
| CQA cycles if FP arith vectorized | 6.06 |
| CQA cycles if fully vectorized | 3.56 |
| Front-end cycles | 7.25 |
| P0 cycles | 0.33 |
| P1 cycles | 0.33 |
| P2 cycles | 0.33 |
| P3 cycles | 0.67 |
| P4 cycles | 0.67 |
| P5 cycles | 0.67 |
| P6 cycles | 3.50 |
| P7 cycles | 3.50 |
| P8 cycles | 3.50 |
| P9 cycles | 3.50 |
| P10 cycles | 10.25 |
| P11 cycles | 10.25 |
| P12 cycles | 10.25 |
| P13 cycles | 10.25 |
| P14 cycles | 1.00 |
| P15 cycles | 1.00 |
| DIV/SQRT cycles | 28.50 |
| Inter-iter dependencies cycles | 2 |
| FE+BE cycles (UFS) | NA |
| Stall cycles (UFS) | NA |
| Nb insns | 58.00 |
| Nb uops | 58.00 |
| Nb loads | 14.00 |
| Nb stores | 0.00 |
| Nb stack references | 1.00 |
| FLOP/cycle | 0.88 |
| Nb FLOP add-sub | 9.00 |
| Nb FLOP mul | 8.00 |
| Nb FLOP fma | 1.00 |
| Nb FLOP div | 5.00 |
| Nb FLOP rcp | 0.00 |
| Nb FLOP sqrt | 1.00 |
| Nb FLOP rsqrt | 0.00 |
| Bytes/cycle | 3.93 |
| Bytes prefetched | 0.00 |
| Bytes loaded | 112.00 |
| Bytes stored | 0.00 |
| Stride 0 | 2.00 |
| Stride 1 | 0.00 |
| Stride n | 0.00 |
| Stride unknown | 12.00 |
| Stride indirect | 0.00 |
| Vectorization ratio all | 16.98 |
| Vectorization ratio load | 0.00 |
| Vectorization ratio store | NA |
| Vectorization ratio mul | 14.29 |
| Vectorization ratio add_sub | 0.00 |
| Vectorization ratio fma | 0.00 |
| Vectorization ratio div_sqrt | 0.00 |
| Vectorization ratio other | 34.78 |
| Vector-efficiency ratio all | 14.62 |
| Vector-efficiency ratio load | 12.50 |
| Vector-efficiency ratio store | NA |
| Vector-efficiency ratio mul | 14.29 |
| Vector-efficiency ratio add_sub | 12.50 |
| Vector-efficiency ratio fma | 12.50 |
| Vector-efficiency ratio div_sqrt | 12.50 |
| Vector-efficiency ratio other | 16.85 |
| Metric | Value |
|---|---|
| CQA speedup if no scalar integer | 1.00 |
| CQA speedup if FP arith vectorized | 5.36 |
| CQA speedup if fully vectorized | 8.00 |
| CQA speedup if no inter-iteration dependency | NA |
| CQA speedup if next bottleneck killed | 2.95 |
| Bottlenecks | P10, P11, |
| Function | __calc_dt_kernel_module_MOD_calc_dt_kernel._omp_fn.0 |
| Source | calc_dt_kernel.f90:93-93,calc_dt_kernel.f90:99-129 |
| Source loop unroll info | not unrolled or unrolled with no peel/tail loop |
| Source loop unroll confidence level | max |
| Unroll/vectorization loop type | NA |
| Unroll factor | NA |
| CQA cycles | 32.50 |
| CQA cycles if no scalar integer | 32.50 |
| CQA cycles if FP arith vectorized | 6.06 |
| CQA cycles if fully vectorized | 4.06 |
| Front-end cycles | 7.75 |
| P0 cycles | 0.67 |
| P1 cycles | 0.67 |
| P2 cycles | 0.67 |
| P3 cycles | 0.67 |
| P4 cycles | 0.67 |
| P5 cycles | 0.67 |
| P6 cycles | 3.50 |
| P7 cycles | 3.50 |
| P8 cycles | 3.50 |
| P9 cycles | 3.50 |
| P10 cycles | 11.00 |
| P11 cycles | 11.00 |
| P12 cycles | 11.00 |
| P13 cycles | 11.00 |
| P14 cycles | 1.00 |
| P15 cycles | 1.00 |
| DIV/SQRT cycles | 32.50 |
| Inter-iter dependencies cycles | 2 |
| FE+BE cycles (UFS) | NA |
| Stall cycles (UFS) | NA |
| Nb insns | 61.00 |
| Nb uops | 62.00 |
| Nb loads | 14.00 |
| Nb stores | 0.00 |
| Nb stack references | 1.00 |
| FLOP/cycle | 0.83 |
| Nb FLOP add-sub | 9.00 |
| Nb FLOP mul | 9.00 |
| Nb FLOP fma | 1.00 |
| Nb FLOP div | 6.00 |
| Nb FLOP rcp | 0.00 |
| Nb FLOP sqrt | 1.00 |
| Nb FLOP rsqrt | 0.00 |
| Bytes/cycle | 3.45 |
| Bytes prefetched | 0.00 |
| Bytes loaded | 112.00 |
| Bytes stored | 0.00 |
| Stride 0 | 2.00 |
| Stride 1 | 0.00 |
| Stride n | 0.00 |
| Stride unknown | 12.00 |
| Stride indirect | 0.00 |
| Vectorization ratio all | 17.86 |
| Vectorization ratio load | 0.00 |
| Vectorization ratio store | NA |
| Vectorization ratio mul | 12.50 |
| Vectorization ratio add_sub | 0.00 |
| Vectorization ratio fma | 0.00 |
| Vectorization ratio div_sqrt | 0.00 |
| Vectorization ratio other | 37.50 |
| Vector-efficiency ratio all | 14.73 |
| Vector-efficiency ratio load | 12.50 |
| Vector-efficiency ratio store | NA |
| Vector-efficiency ratio mul | 14.06 |
| Vector-efficiency ratio add_sub | 12.50 |
| Vector-efficiency ratio fma | 12.50 |
| Vector-efficiency ratio div_sqrt | 12.50 |
| Vector-efficiency ratio other | 17.19 |
| Path / |
| Function | __calc_dt_kernel_module_MOD_calc_dt_kernel._omp_fn.0 |
| Source file and lines | calc_dt_kernel.f90:93-129 |
| Module | exec |
| nb instructions | 59.50 |
| nb uops | 60 |
| loop length | 307 |
| used x86 registers | 15 |
| used mmx registers | 0 |
| used xmm registers | 24 |
| used ymm registers | 0 |
| used zmm registers | 0 |
| nb stack references | 1 |
| ADD-SUB / MUL ratio | 1.21 |
| micro-operation queue | 7.50 cycles |
| front end | 7.50 cycles |
| P0 | P1 | P2 | P3 | P4 | P5 | P6 | P7 | P8 | P9 | P10 | P11 | P12 | P13 | P14 | P15 | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| uops | 0.50 | 0.50 | 0.50 | 0.67 | 0.67 | 0.67 | 3.50 | 3.50 | 3.50 | 3.50 | 10.63 | 10.63 | 10.63 | 10.63 | 1.00 | 1.00 |
| cycles | 0.50 | 0.50 | 0.50 | 0.67 | 0.67 | 0.67 | 3.50 | 3.50 | 3.50 | 3.50 | 10.63 | 10.63 | 10.63 | 10.63 | 1.00 | 1.00 |
| Cycles executing div or sqrt instructions | 30.50 |
| Longest recurrence chain latency (RecMII) | 2.00 |
| Front-end | 7.50 |
| Dispatch | 10.63 |
| DIV/SQRT | 30.50 |
| Data deps. | 2.00 |
| Overall L1 | 30.50 |
| all | 0% |
| load | NA (no load vectorizable/vectorized instructions) |
| store | NA (no store vectorizable/vectorized instructions) |
| mul | NA (no mul vectorizable/vectorized instructions) |
| add-sub | NA (no add-sub vectorizable/vectorized instructions) |
| fma | NA (no fma vectorizable/vectorized instructions) |
| other | 0% |
| all | 17% |
| load | 0% |
| store | NA (no store vectorizable/vectorized instructions) |
| mul | 13% |
| add-sub | 0% |
| fma | 0% |
| div/sqrt | 0% |
| other | 37% |
| all | 17% |
| load | 0% |
| store | NA (no store vectorizable/vectorized instructions) |
| mul | 13% |
| add-sub | 0% |
| fma | 0% |
| div/sqrt | 0% |
| other | 36% |
| all | 12% |
| load | NA (no load vectorizable/vectorized instructions) |
| store | NA (no store vectorizable/vectorized instructions) |
| mul | NA (no mul vectorizable/vectorized instructions) |
| add-sub | NA (no add-sub vectorizable/vectorized instructions) |
| fma | NA (no fma vectorizable/vectorized instructions) |
| other | 12% |
| all | 14% |
| load | 12% |
| store | NA (no store vectorizable/vectorized instructions) |
| mul | 14% |
| add-sub | 12% |
| fma | 12% |
| div/sqrt | 12% |
| other | 17% |
| all | 14% |
| load | 12% |
| store | NA (no store vectorizable/vectorized instructions) |
| mul | 14% |
| add-sub | 12% |
| fma | 12% |
| div/sqrt | 12% |
| other | 17% |
| Function | __calc_dt_kernel_module_MOD_calc_dt_kernel._omp_fn.0 |
| Source file and lines | calc_dt_kernel.f90:93-129 |
| Module | exec |
| nb instructions | 58 |
| nb uops | 58 |
| loop length | 299 |
| used x86 registers | 15 |
| used mmx registers | 0 |
| used xmm registers | 23 |
| used ymm registers | 0 |
| used zmm registers | 0 |
| nb stack references | 1 |
| ADD-SUB / MUL ratio | 1.29 |
| micro-operation queue | 7.25 cycles |
| front end | 7.25 cycles |
| P0 | P1 | P2 | P3 | P4 | P5 | P6 | P7 | P8 | P9 | P10 | P11 | P12 | P13 | P14 | P15 | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| uops | 0.33 | 0.33 | 0.33 | 0.67 | 0.67 | 0.67 | 3.50 | 3.50 | 3.50 | 3.50 | 10.25 | 10.25 | 10.25 | 10.25 | 1.00 | 1.00 |
| cycles | 0.33 | 0.33 | 0.33 | 0.67 | 0.67 | 0.67 | 3.50 | 3.50 | 3.50 | 3.50 | 10.25 | 10.25 | 10.25 | 10.25 | 1.00 | 1.00 |
| Cycles executing div or sqrt instructions | 28.50 |
| Longest recurrence chain latency (RecMII) | 2.00 |
| Front-end | 7.25 |
| Dispatch | 10.25 |
| DIV/SQRT | 28.50 |
| Data deps. | 2.00 |
| Overall L1 | 28.50 |
| all | 0% |
| load | NA (no load vectorizable/vectorized instructions) |
| store | NA (no store vectorizable/vectorized instructions) |
| mul | NA (no mul vectorizable/vectorized instructions) |
| add-sub | NA (no add-sub vectorizable/vectorized instructions) |
| fma | NA (no fma vectorizable/vectorized instructions) |
| other | 0% |
| all | 17% |
| load | 0% |
| store | NA (no store vectorizable/vectorized instructions) |
| mul | 14% |
| add-sub | 0% |
| fma | 0% |
| div/sqrt | 0% |
| other | 36% |
| all | 16% |
| load | 0% |
| store | NA (no store vectorizable/vectorized instructions) |
| mul | 14% |
| add-sub | 0% |
| fma | 0% |
| div/sqrt | 0% |
| other | 34% |
| all | 12% |
| load | NA (no load vectorizable/vectorized instructions) |
| store | NA (no store vectorizable/vectorized instructions) |
| mul | NA (no mul vectorizable/vectorized instructions) |
| add-sub | NA (no add-sub vectorizable/vectorized instructions) |
| fma | NA (no fma vectorizable/vectorized instructions) |
| other | 12% |
| all | 14% |
| load | 12% |
| store | NA (no store vectorizable/vectorized instructions) |
| mul | 14% |
| add-sub | 12% |
| fma | 12% |
| div/sqrt | 12% |
| other | 17% |
| all | 14% |
| load | 12% |
| store | NA (no store vectorizable/vectorized instructions) |
| mul | 14% |
| add-sub | 12% |
| fma | 12% |
| div/sqrt | 12% |
| other | 16% |
| Instruction | Nb FU | P0 | P1 | P2 | P3 | P4 | P5 | P6 | P7 | P8 | P9 | P10 | P11 | P12 | P13 | P14 | P15 | Latency | Recip. throughput | Vectorization |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| MOV %RDX,%RAX | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.17 | N/A |
| VMOVAPD %XMM5,%XMM6 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.17 | vect (25.0%) |
| VMOVAPD %XMM7,%XMM4 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.17 | vect (25.0%) |
| VMOVAPD %XMM8,%XMM1 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.17 | vect (25.0%) |
| VMOVAPD %XMM9,%XMM2 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.17 | vect (25.0%) |
| VMOVSD (%R13,%RAX,8),%XMM29 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0.25 | 0.25 | 0.25 | 0.25 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0.50 | scal (12.5%) |
| VMOVSD (%R14,%RAX,8),%XMM5 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0.25 | 0.25 | 0.25 | 0.25 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0.50 | scal (12.5%) |
| VADDSD %XMM29,%XMM29,%XMM28 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.50 | 0.50 | 0 | 0 | 2 | 0.50 | scal (12.5%) |
| VDIVSD (%R12,%RAX,8),%XMM28,%XMM29 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0.25 | 0.25 | 0.25 | 0.25 | 0.50 | 0.50 | 0 | 0 | 0 | 0 | 13 | 4 | scal (12.5%) |
| VFMADD132SD %XMM5,%XMM29,%XMM5 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.50 | 0.50 | 0 | 0 | 0 | 0 | 4 | 0.50 | scal (12.5%) |
| VMOVSD 0x8(%R9,%RAX,8),%XMM3 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0.25 | 0.25 | 0.25 | 0.25 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0.50 | scal (12.5%) |
| VMOVSD 0x8(%RDI,%RAX,8),%XMM7 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0.25 | 0.25 | 0.25 | 0.25 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0.50 | scal (12.5%) |
| VUNPCKLPD %XMM4,%XMM7,%XMM4 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.25 | 0.25 | 0.25 | 0.25 | 0 | 0 | 1 | 0.25 | scal (12.5%) |
| VMINSD (%R15,%RAX,8),%XMM18,%XMM28 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0.25 | 0.25 | 0.25 | 0.25 | 0 | 0 | 0.50 | 0.50 | 0 | 0 | 2 | 0.50 | scal (12.5%) |
| VMULSD %XMM17,%XMM28,%XMM29 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.50 | 0.50 | 0 | 0 | 0 | 0 | 3 | 0.50 | scal (12.5%) |
| VSQRTSD %XMM5,%XMM5,%XMM5 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.50 | 0.50 | 0 | 0 | 0 | 0 | 21 | 8.50 | scal (12.5%) |
| VMAXSD %XMM10,%XMM5,%XMM8 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.50 | 0.50 | 0 | 0 | 2 | 0.50 | scal (12.5%) |
| VADDSD 0x8(%R8,%RAX,8),%XMM3,%XMM5 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0.25 | 0.25 | 0.25 | 0.25 | 0 | 0 | 0.50 | 0.50 | 0 | 0 | 2 | 0.50 | scal (12.5%) |
| VUNPCKLPD %XMM6,%XMM5,%XMM0 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.25 | 0.25 | 0.25 | 0.25 | 0 | 0 | 1 | 0.25 | scal (12.5%) |
| VMULPD %XMM0,%XMM4,%XMM9 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.50 | 0.50 | 0 | 0 | 0 | 0 | 3 | 0.50 | vect (25.0%) |
| VDIVSD %XMM8,%XMM29,%XMM29 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.50 | 0.50 | 0 | 0 | 0 | 0 | 13 | 4 | scal (12.5%) |
| VANDPD %XMM20,%XMM9,%XMM4 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.25 | 0.25 | 0.25 | 0.25 | 0 | 0 | 1 | 0.25 | vect (25.0%) |
| VADDSD %XMM19,%XMM9,%XMM8 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.50 | 0.50 | 0 | 0 | 2 | 0.50 | scal (12.5%) |
| VUNPCKHPD %XMM9,%XMM9,%XMM3 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.25 | 0.25 | 0.25 | 0.25 | 0 | 0 | 1 | 0.25 | scal (12.5%) |
| VSUBSD %XMM3,%XMM8,%XMM0 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.50 | 0.50 | 0 | 0 | 2 | 0.50 | scal (12.5%) |
| VUNPCKHPD %XMM4,%XMM4,%XMM28 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.25 | 0.25 | 0.25 | 0.25 | 0 | 0 | 1 | 0.25 | scal (12.5%) |
| VMOVSD (%RBX,%RAX,8),%XMM3 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0.25 | 0.25 | 0.25 | 0.25 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0.50 | scal (12.5%) |
| VMOVSD 0x8(%RSI,%RAX,8),%XMM8 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0.25 | 0.25 | 0.25 | 0.25 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0.50 | scal (12.5%) |
| VMULSD %XMM3,%XMM10,%XMM6 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.50 | 0.50 | 0 | 0 | 0 | 0 | 3 | 0.50 | scal (12.5%) |
| VMAXPD %XMM4,%XMM28,%XMM28 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.50 | 0.50 | 0 | 0 | 2 | 0.50 | vect (25.0%) |
| VMULSD %XMM15,%XMM3,%XMM9 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.50 | 0.50 | 0 | 0 | 0 | 0 | 3 | 0.50 | scal (12.5%) |
| VADDSD %XMM1,%XMM8,%XMM1 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.50 | 0.50 | 0 | 0 | 2 | 0.50 | scal (12.5%) |
| VMULSD (%R11,%RAX,8),%XMM1,%XMM1 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0.25 | 0.25 | 0.25 | 0.25 | 0.50 | 0.50 | 0 | 0 | 0 | 0 | 3 | 0.50 | scal (12.5%) |
| VMAXSD %XMM28,%XMM6,%XMM28 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.50 | 0.50 | 0 | 0 | 2 | 0.50 | scal (12.5%) |
| VDIVSD %XMM28,%XMM9,%XMM4 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.50 | 0.50 | 0 | 0 | 0 | 0 | 13 | 4 | scal (12.5%) |
| VMOVSD 0x8(%RCX,%RAX,8),%XMM9 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0.25 | 0.25 | 0.25 | 0.25 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0.50 | scal (12.5%) |
| VADDSD %XMM2,%XMM9,%XMM2 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.50 | 0.50 | 0 | 0 | 2 | 0.50 | scal (12.5%) |
| VMULSD (%R10,%RAX,8),%XMM2,%XMM2 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0.25 | 0.25 | 0.25 | 0.25 | 0.50 | 0.50 | 0 | 0 | 0 | 0 | 3 | 0.50 | scal (12.5%) |
| VADDSD %XMM2,%XMM0,%XMM0 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.50 | 0.50 | 0 | 0 | 2 | 0.50 | scal (12.5%) |
| VANDPD %XMM12,%XMM2,%XMM2 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.25 | 0.25 | 0.25 | 0.25 | 0 | 0 | 1 | 0.25 | vect (25.0%) |
| VSUBSD %XMM1,%XMM0,%XMM0 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.50 | 0.50 | 0 | 0 | 2 | 0.50 | scal (12.5%) |
| VANDPD %XMM12,%XMM1,%XMM1 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.25 | 0.25 | 0.25 | 0.25 | 0 | 0 | 1 | 0.25 | vect (25.0%) |
| VMINSD %XMM29,%XMM4,%XMM4 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.50 | 0.50 | 0 | 0 | 2 | 0.50 | scal (12.5%) |
| VMULSD %XMM14,%XMM3,%XMM29 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.50 | 0.50 | 0 | 0 | 0 | 0 | 3 | 0.50 | scal (12.5%) |
| VADDSD %XMM3,%XMM3,%XMM3 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.50 | 0.50 | 0 | 0 | 2 | 0.50 | scal (12.5%) |
| VMAXSD %XMM2,%XMM1,%XMM1 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.50 | 0.50 | 0 | 0 | 2 | 0.50 | scal (12.5%) |
| VDIVSD %XMM3,%XMM0,%XMM0 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.50 | 0.50 | 0 | 0 | 0 | 0 | 13 | 4 | scal (12.5%) |
| VMAXSD %XMM6,%XMM1,%XMM6 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.50 | 0.50 | 0 | 0 | 2 | 0.50 | scal (12.5%) |
| VCOMISD %XMM0,%XMM16 | 2 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.50 | 0.50 | 0.50 | 0.50 | 6 | 0.50 | scal (12.5%) |
| VDIVSD %XMM6,%XMM29,%XMM2 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.50 | 0.50 | 0 | 0 | 0 | 0 | 13 | 4 | scal (12.5%) |
| JA 42a770 <__calc_dt_kernel_module_MOD_calc_dt_kernel._omp_fn.0+0x400> | 1 | 0 | 0 | 0 | 0.33 | 0.33 | 0.33 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0.50-1 | N/A |
| VMOVQ %XMM25,%RDX | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.50 | 0.50 | 1 | 0.50 | scal (12.5%) |
| VMINSD (%RDX),%XMM2,%XMM1 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0.25 | 0.25 | 0.25 | 0.25 | 0 | 0 | 0.50 | 0.50 | 0 | 0 | 2 | 0.50 | scal (12.5%) |
| LEA 0x1(%RAX),%RDX | 1 | 0.17 | 0.17 | 0.17 | 0.17 | 0.17 | 0.17 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0.25 | N/A |
| VMINSD %XMM4,%XMM1,%XMM4 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.50 | 0.50 | 0 | 0 | 2 | 0.50 | scal (12.5%) |
| VMINSD %XMM4,%XMM11,%XMM11 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.50 | 0.50 | 0 | 0 | 2 | 0.50 | scal (12.5%) |
| CMP %RAX,-0x38(%RBP) | 1 | 0.17 | 0.17 | 0.17 | 0.17 | 0.17 | 0.17 | 0.25 | 0.25 | 0.25 | 0.25 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0.25 | N/A |
| JNE 42a7a0 <__calc_dt_kernel_module_MOD_calc_dt_kernel._omp_fn.0+0x430> | 1 | 0 | 0 | 0 | 0.33 | 0.33 | 0.33 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0.50-1 | N/A |
| Function | __calc_dt_kernel_module_MOD_calc_dt_kernel._omp_fn.0 |
| Source file and lines | calc_dt_kernel.f90:93-129 |
| Module | exec |
| nb instructions | 61 |
| nb uops | 62 |
| loop length | 315 |
| used x86 registers | 15 |
| used mmx registers | 0 |
| used xmm registers | 25 |
| used ymm registers | 0 |
| used zmm registers | 0 |
| nb stack references | 1 |
| ADD-SUB / MUL ratio | 1.13 |
| micro-operation queue | 7.75 cycles |
| front end | 7.75 cycles |
| P0 | P1 | P2 | P3 | P4 | P5 | P6 | P7 | P8 | P9 | P10 | P11 | P12 | P13 | P14 | P15 | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| uops | 0.67 | 0.67 | 0.67 | 0.67 | 0.67 | 0.67 | 3.50 | 3.50 | 3.50 | 3.50 | 11.00 | 11.00 | 11.00 | 11.00 | 1.00 | 1.00 |
| cycles | 0.67 | 0.67 | 0.67 | 0.67 | 0.67 | 0.67 | 3.50 | 3.50 | 3.50 | 3.50 | 11.00 | 11.00 | 11.00 | 11.00 | 1.00 | 1.00 |
| Cycles executing div or sqrt instructions | 32.50 |
| Longest recurrence chain latency (RecMII) | 2.00 |
| Front-end | 7.75 |
| Dispatch | 11.00 |
| DIV/SQRT | 32.50 |
| Data deps. | 2.00 |
| Overall L1 | 32.50 |
| all | 0% |
| load | NA (no load vectorizable/vectorized instructions) |
| store | NA (no store vectorizable/vectorized instructions) |
| mul | NA (no mul vectorizable/vectorized instructions) |
| add-sub | NA (no add-sub vectorizable/vectorized instructions) |
| fma | NA (no fma vectorizable/vectorized instructions) |
| other | 0% |
| all | 18% |
| load | 0% |
| store | NA (no store vectorizable/vectorized instructions) |
| mul | 12% |
| add-sub | 0% |
| fma | 0% |
| div/sqrt | 0% |
| other | 39% |
| all | 17% |
| load | 0% |
| store | NA (no store vectorizable/vectorized instructions) |
| mul | 12% |
| add-sub | 0% |
| fma | 0% |
| div/sqrt | 0% |
| other | 37% |
| all | 12% |
| load | NA (no load vectorizable/vectorized instructions) |
| store | NA (no store vectorizable/vectorized instructions) |
| mul | NA (no mul vectorizable/vectorized instructions) |
| add-sub | NA (no add-sub vectorizable/vectorized instructions) |
| fma | NA (no fma vectorizable/vectorized instructions) |
| other | 12% |
| all | 14% |
| load | 12% |
| store | NA (no store vectorizable/vectorized instructions) |
| mul | 14% |
| add-sub | 12% |
| fma | 12% |
| div/sqrt | 12% |
| other | 17% |
| all | 14% |
| load | 12% |
| store | NA (no store vectorizable/vectorized instructions) |
| mul | 14% |
| add-sub | 12% |
| fma | 12% |
| div/sqrt | 12% |
| other | 17% |
| Instruction | Nb FU | P0 | P1 | P2 | P3 | P4 | P5 | P6 | P7 | P8 | P9 | P10 | P11 | P12 | P13 | P14 | P15 | Latency | Recip. throughput | Vectorization |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| VMOVQ %XMM24,%RDX | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.50 | 0.50 | 1 | 0.50 | scal (12.5%) |
| VDIVSD %XMM0,%XMM23,%XMM6 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.50 | 0.50 | 0 | 0 | 0 | 0 | 13 | 4 | scal (12.5%) |
| VXORPD %XMM22,%XMM6,%XMM3 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.25 | 0.25 | 0.25 | 0.25 | 0 | 0 | 1 | 0.25 | vect (25.0%) |
| VMULSD (%RDX),%XMM3,%XMM0 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0.25 | 0.25 | 0.25 | 0.25 | 0.50 | 0.50 | 0 | 0 | 0 | 0 | 3 | 0.50 | scal (12.5%) |
| LEA 0x1(%RAX),%RDX | 1 | 0.17 | 0.17 | 0.17 | 0.17 | 0.17 | 0.17 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0.25 | N/A |
| VMINSD %XMM2,%XMM0,%XMM2 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.50 | 0.50 | 0 | 0 | 2 | 0.50 | scal (12.5%) |
| VMINSD %XMM4,%XMM2,%XMM1 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.50 | 0.50 | 0 | 0 | 2 | 0.50 | scal (12.5%) |
| VMINSD %XMM1,%XMM11,%XMM11 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.50 | 0.50 | 0 | 0 | 2 | 0.50 | scal (12.5%) |
| CMP %RAX,-0x38(%RBP) | 1 | 0.17 | 0.17 | 0.17 | 0.17 | 0.17 | 0.17 | 0.25 | 0.25 | 0.25 | 0.25 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0.25 | N/A |
| JE 42a8cb <__calc_dt_kernel_module_MOD_calc_dt_kernel._omp_fn.0+0x55b> | 1 | 0 | 0 | 0 | 0.33 | 0.33 | 0.33 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0.50-1 | N/A |
| MOV %RDX,%RAX | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.17 | N/A |
| VMOVAPD %XMM5,%XMM6 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.17 | vect (25.0%) |
| VMOVAPD %XMM7,%XMM4 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.17 | vect (25.0%) |
| VMOVAPD %XMM8,%XMM1 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.17 | vect (25.0%) |
| VMOVAPD %XMM9,%XMM2 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.17 | vect (25.0%) |
| VMOVSD (%R13,%RAX,8),%XMM29 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0.25 | 0.25 | 0.25 | 0.25 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0.50 | scal (12.5%) |
| VMOVSD (%R14,%RAX,8),%XMM5 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0.25 | 0.25 | 0.25 | 0.25 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0.50 | scal (12.5%) |
| VADDSD %XMM29,%XMM29,%XMM28 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.50 | 0.50 | 0 | 0 | 2 | 0.50 | scal (12.5%) |
| VDIVSD (%R12,%RAX,8),%XMM28,%XMM29 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0.25 | 0.25 | 0.25 | 0.25 | 0.50 | 0.50 | 0 | 0 | 0 | 0 | 13 | 4 | scal (12.5%) |
| VFMADD132SD %XMM5,%XMM29,%XMM5 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.50 | 0.50 | 0 | 0 | 0 | 0 | 4 | 0.50 | scal (12.5%) |
| VMOVSD 0x8(%R9,%RAX,8),%XMM3 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0.25 | 0.25 | 0.25 | 0.25 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0.50 | scal (12.5%) |
| VMOVSD 0x8(%RDI,%RAX,8),%XMM7 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0.25 | 0.25 | 0.25 | 0.25 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0.50 | scal (12.5%) |
| VUNPCKLPD %XMM4,%XMM7,%XMM4 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.25 | 0.25 | 0.25 | 0.25 | 0 | 0 | 1 | 0.25 | scal (12.5%) |
| VMINSD (%R15,%RAX,8),%XMM18,%XMM28 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0.25 | 0.25 | 0.25 | 0.25 | 0 | 0 | 0.50 | 0.50 | 0 | 0 | 2 | 0.50 | scal (12.5%) |
| VMULSD %XMM17,%XMM28,%XMM29 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.50 | 0.50 | 0 | 0 | 0 | 0 | 3 | 0.50 | scal (12.5%) |
| VSQRTSD %XMM5,%XMM5,%XMM5 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.50 | 0.50 | 0 | 0 | 0 | 0 | 21 | 8.50 | scal (12.5%) |
| VMAXSD %XMM10,%XMM5,%XMM8 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.50 | 0.50 | 0 | 0 | 2 | 0.50 | scal (12.5%) |
| VADDSD 0x8(%R8,%RAX,8),%XMM3,%XMM5 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0.25 | 0.25 | 0.25 | 0.25 | 0 | 0 | 0.50 | 0.50 | 0 | 0 | 2 | 0.50 | scal (12.5%) |
| VUNPCKLPD %XMM6,%XMM5,%XMM0 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.25 | 0.25 | 0.25 | 0.25 | 0 | 0 | 1 | 0.25 | scal (12.5%) |
| VMULPD %XMM0,%XMM4,%XMM9 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.50 | 0.50 | 0 | 0 | 0 | 0 | 3 | 0.50 | vect (25.0%) |
| VDIVSD %XMM8,%XMM29,%XMM29 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.50 | 0.50 | 0 | 0 | 0 | 0 | 13 | 4 | scal (12.5%) |
| VANDPD %XMM20,%XMM9,%XMM4 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.25 | 0.25 | 0.25 | 0.25 | 0 | 0 | 1 | 0.25 | vect (25.0%) |
| VADDSD %XMM19,%XMM9,%XMM8 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.50 | 0.50 | 0 | 0 | 2 | 0.50 | scal (12.5%) |
| VUNPCKHPD %XMM9,%XMM9,%XMM3 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.25 | 0.25 | 0.25 | 0.25 | 0 | 0 | 1 | 0.25 | scal (12.5%) |
| VSUBSD %XMM3,%XMM8,%XMM0 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.50 | 0.50 | 0 | 0 | 2 | 0.50 | scal (12.5%) |
| VUNPCKHPD %XMM4,%XMM4,%XMM28 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.25 | 0.25 | 0.25 | 0.25 | 0 | 0 | 1 | 0.25 | scal (12.5%) |
| VMOVSD (%RBX,%RAX,8),%XMM3 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0.25 | 0.25 | 0.25 | 0.25 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0.50 | scal (12.5%) |
| VMOVSD 0x8(%RSI,%RAX,8),%XMM8 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0.25 | 0.25 | 0.25 | 0.25 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0.50 | scal (12.5%) |
| VMULSD %XMM3,%XMM10,%XMM6 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.50 | 0.50 | 0 | 0 | 0 | 0 | 3 | 0.50 | scal (12.5%) |
| VMAXPD %XMM4,%XMM28,%XMM28 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.50 | 0.50 | 0 | 0 | 2 | 0.50 | vect (25.0%) |
| VMULSD %XMM15,%XMM3,%XMM9 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.50 | 0.50 | 0 | 0 | 0 | 0 | 3 | 0.50 | scal (12.5%) |
| VADDSD %XMM1,%XMM8,%XMM1 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.50 | 0.50 | 0 | 0 | 2 | 0.50 | scal (12.5%) |
| VMULSD (%R11,%RAX,8),%XMM1,%XMM1 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0.25 | 0.25 | 0.25 | 0.25 | 0.50 | 0.50 | 0 | 0 | 0 | 0 | 3 | 0.50 | scal (12.5%) |
| VMAXSD %XMM28,%XMM6,%XMM28 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.50 | 0.50 | 0 | 0 | 2 | 0.50 | scal (12.5%) |
| VDIVSD %XMM28,%XMM9,%XMM4 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.50 | 0.50 | 0 | 0 | 0 | 0 | 13 | 4 | scal (12.5%) |
| VMOVSD 0x8(%RCX,%RAX,8),%XMM9 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0.25 | 0.25 | 0.25 | 0.25 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0.50 | scal (12.5%) |
| VADDSD %XMM2,%XMM9,%XMM2 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.50 | 0.50 | 0 | 0 | 2 | 0.50 | scal (12.5%) |
| VMULSD (%R10,%RAX,8),%XMM2,%XMM2 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0.25 | 0.25 | 0.25 | 0.25 | 0.50 | 0.50 | 0 | 0 | 0 | 0 | 3 | 0.50 | scal (12.5%) |
| VADDSD %XMM2,%XMM0,%XMM0 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.50 | 0.50 | 0 | 0 | 2 | 0.50 | scal (12.5%) |
| VANDPD %XMM12,%XMM2,%XMM2 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.25 | 0.25 | 0.25 | 0.25 | 0 | 0 | 1 | 0.25 | vect (25.0%) |
| VSUBSD %XMM1,%XMM0,%XMM0 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.50 | 0.50 | 0 | 0 | 2 | 0.50 | scal (12.5%) |
| VANDPD %XMM12,%XMM1,%XMM1 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.25 | 0.25 | 0.25 | 0.25 | 0 | 0 | 1 | 0.25 | vect (25.0%) |
| VMINSD %XMM29,%XMM4,%XMM4 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.50 | 0.50 | 0 | 0 | 2 | 0.50 | scal (12.5%) |
| VMULSD %XMM14,%XMM3,%XMM29 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.50 | 0.50 | 0 | 0 | 0 | 0 | 3 | 0.50 | scal (12.5%) |
| VADDSD %XMM3,%XMM3,%XMM3 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.50 | 0.50 | 0 | 0 | 2 | 0.50 | scal (12.5%) |
| VMAXSD %XMM2,%XMM1,%XMM1 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.50 | 0.50 | 0 | 0 | 2 | 0.50 | scal (12.5%) |
| VDIVSD %XMM3,%XMM0,%XMM0 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.50 | 0.50 | 0 | 0 | 0 | 0 | 13 | 4 | scal (12.5%) |
| VMAXSD %XMM6,%XMM1,%XMM6 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.50 | 0.50 | 0 | 0 | 2 | 0.50 | scal (12.5%) |
| VCOMISD %XMM0,%XMM16 | 2 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.50 | 0.50 | 0.50 | 0.50 | 6 | 0.50 | scal (12.5%) |
| VDIVSD %XMM6,%XMM29,%XMM2 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.50 | 0.50 | 0 | 0 | 0 | 0 | 13 | 4 | scal (12.5%) |
| JA 42a770 <__calc_dt_kernel_module_MOD_calc_dt_kernel._omp_fn.0+0x400> | 1 | 0 | 0 | 0 | 0.33 | 0.33 | 0.33 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0.50-1 | N/A |
