| Loop Id: 544 | Module: exec | Source: advec_mom_kernel.f90-pp.f90:151-178 | Coverage: 0.01% |
|---|
| Loop Id: 544 | Module: exec | Source: advec_mom_kernel.f90-pp.f90:151-178 | Coverage: 0.01% |
|---|
0x41f860 ADD X8, X8, #1 |
0x41f864 SUBS W9, W9, #1 |
0x41f868 B.LE 41f9c0 |
0x41f86c CMP W11, #1 |
0x41f870 B.LT 41f860 |
0x41f874 LDR X16, [X20, #64] |
0x41f878 LDR X17, [X20, #256] |
0x41f87c LDP X1, X2, [X20, #288] |
0x41f880 ORR X15, XZR, XZR |
0x41f884 LDR X16, [X16] |
0x41f888 MADD X18, X16, X8, XZR |
0x41f88c LDR X16, [X20, #104] |
0x41f890 LDR X0, [X16] |
0x41f894 SUB X16, X18, X0 |
0x41f898 ADD X18, X12, X18 |
0x41f89c UBFM X3, X16, #61, #60 |
0x41f8a0 SUB X18, X18, X0 |
0x41f8a4 ADD X16, X17, X3 |
0x41f8a8 ADD X17, X1, X3 |
0x41f8ac LDR X1, [X20, #224] |
0x41f8b0 UBFM X0, X18, #61, #60 |
0x41f8b4 ADD X18, X2, X0 |
0x41f8b8 ORR X2, XZR, X12 |
0x41f8bc ADD X0, X1, X0 |
0x41f8c0 ORR W1, WZR, W14 |
0x41f8c4 B 41f8ec |
(545) 0x41f8c8 FSUB D4, D0, S4 |
(545) 0x41f8cc SUB W1, W1, #1 |
(545) 0x41f8d0 ADD X2, X2, #1 |
(545) 0x41f8d4 CMP W1, #1 |
(545) 0x41f8d8 FMADD D4, D4, D16, D5 |
(545) 0x41f8dc FMUL D3, D3, D4 |
(545) 0x41f8e0 STR D3, [X18, X15,LSL #3] |
(545) 0x41f8e4 ADD X15, X15, #1 |
(545) 0x41f8e8 B.LE 41f860 |
(545) 0x41f8ec LDR D3, [X0, X15,LSL #3] |
(545) 0x41f8f0 FCMP D3, #0 |
(545) 0x41f8f4 B.MI 41f90c |
(545) 0x41f8f8 SUB W3, W2, #1 |
(545) 0x41f8fc ORR X5, XZR, X2 |
(545) 0x41f900 ADD X4, X2, #1 |
(545) 0x41f904 ORR W6, WZR, W3 |
(545) 0x41f908 B 41f91c |
(545) 0x41f90c ADD X4, X12, X15 |
(545) 0x41f910 ADD W3, W10, W15 |
(545) 0x41f914 ADD X5, X4, #1 |
(545) 0x41f918 ADD W6, W3, #1 |
(545) 0x41f91c SBFM X5, X5, #61, #31 |
(545) 0x41f920 FABS D4, D3 |
(545) 0x41f924 LDR D6, [X17, X6,SXTW #3] |
(545) 0x41f928 SBFM X4, X4, #61, #31 |
(545) 0x41f92c MOVI D16, #0 |
(545) 0x41f930 LDR D5, [X16, X5] |
(545) 0x41f934 FDIV D4, D4, D5 |
(545) 0x41f938 LDR D5, [X17, X5] |
(545) 0x41f93c FSUB D7, D5, S6 |
(545) 0x41f940 LDR D6, [X17, X4] |
(545) 0x41f944 FSUB D6, D6, S5 |
(545) 0x41f948 FMUL D17, D7, D6 |
(545) 0x41f94c FCMP D17, #0 |
(545) 0x41f950 B.LE 41f8c8 |
(545) 0x41f954 LDP X4, X5, [X20, #272] |
(545) 0x41f958 SBFM X3, X3, #0, #31 |
(545) 0x41f95c FABS D7, D7 |
(545) 0x41f960 FADD D18, D4, D0 |
(545) 0x41f964 LDR X5, [X5] |
(545) 0x41f968 FMUL D18, D18, D7 |
(545) 0x41f96c FABS D17, D6 |
(545) 0x41f970 SUB X3, X3, X5 |
(545) 0x41f974 SUB X6, X4, X5,LSL #3 |
(545) 0x41f978 LDR D19, [X4, X3,LSL #3] |
(545) 0x41f97c ADD X6, X6, X13 |
(545) 0x41f980 LDR D16, [X6, X15,LSL #3] |
(545) 0x41f984 FDIV D18, D18, D19 |
(545) 0x41f988 FSUB D19, D1, S4 |
(545) 0x41f98c FMUL D19, D19, D17 |
(545) 0x41f990 FDIV D19, D19, D16 |
(545) 0x41f994 FADD D18, D19, D18 |
(545) 0x41f998 FMUL D16, D16, D18 |
(545) 0x41f99c FDIV D16, D16, D2 |
(545) 0x41f9a0 FCMP D16, D7 |
(545) 0x41f9a4 FCSEL D7, D7, D16, #10 |
(545) 0x41f9a8 FCMP D7, D17 |
(545) 0x41f9ac FCSEL D7, D17, D7, #10 |
(545) 0x41f9b0 FCMP D6, #0 |
(545) 0x41f9b4 FNEG D16, D7 |
(545) 0x41f9b8 FCSEL D16, D7, D16, #8 |
(545) 0x41f9bc B 41f8c8 |
/home/hbollore/qaas-runs/170-308-3604/intel/CloverLeafFC/build/build/CMakeFiles/clover_leaf.dir/CloverLeaf_ref/kernels/advec_mom_kernel.f90-pp.f90: 151 - 178 |
-------------------------------------------------------------------------------- |
151: DO k=y_min,y_max+1 |
152: DO j=x_min-1,x_max+1 |
153: IF(node_flux(j,k).LT.0.0)THEN |
154: upwind=j+2 |
155: donor=j+1 |
156: downwind=j |
157: dif=donor |
158: ELSE |
159: upwind=j-1 |
160: donor=j |
161: downwind=j+1 |
162: dif=upwind |
163: ENDIF |
164: sigma=ABS(node_flux(j,k))/(node_mass_pre(donor,k)) |
165: width=celldx(j) |
166: vdiffuw=vel1(donor,k)-vel1(upwind,k) |
167: vdiffdw=vel1(downwind,k)-vel1(donor,k) |
168: limiter=0.0 |
169: IF(vdiffuw*vdiffdw.GT.0.0)THEN |
170: auw=ABS(vdiffuw) |
171: adw=ABS(vdiffdw) |
172: wind=1.0_8 |
173: IF(vdiffdw.LE.0.0) wind=-1.0_8 |
174: limiter=wind*MIN(width*((2.0_8-sigma)*adw/width+(1.0_8+sigma)*auw/celldx(dif))/6.0_8,auw,adw) |
175: ENDIF |
176: advec_vel_s=vel1(donor,k)+(1.0-sigma)*limiter |
177: mom_flux(j,k)=advec_vel_s*node_flux(j,k) |
178: ENDDO |
| Path / |
| Metric | Value |
|---|---|
| CQA speedup if no scalar integer | 1.00 |
| CQA speedup if FP arith vectorized | 1.00 |
| CQA speedup if fully vectorized | 5.33 |
| CQA speedup if no inter-iteration dependency | NA |
| CQA speedup if next bottleneck killed | 1.23 |
| Bottlenecks | P2, P3, P4, P5, |
| Function | advec_mom_kernel |
| Source | advec_mom_kernel.f90-pp.f90:151-151,advec_mom_kernel.f90-pp.f90:178-178 |
| Source loop unroll info | NA |
| Source loop unroll confidence level | NA |
| Unroll/vectorization loop type | NA |
| Unroll factor | NA |
| CQA cycles | 4.00 |
| CQA cycles if no scalar integer | 4.00 |
| CQA cycles if FP arith vectorized | 4.00 |
| CQA cycles if fully vectorized | 0.75 |
| Front-end cycles | 3.25 |
| DIV/SQRT cycles | 1.50 |
| P0 cycles | 1.50 |
| P1 cycles | 4.00 |
| P2 cycles | 4.00 |
| P3 cycles | 4.00 |
| P4 cycles | 4.00 |
| P5 cycles | 0.00 |
| P6 cycles | 0.00 |
| P7 cycles | 0.00 |
| P8 cycles | 0.00 |
| P9 cycles | 2.33 |
| P10 cycles | 2.33 |
| P11 cycles | 2.33 |
| P12 cycles | 0.00 |
| P13 cycles | 0.00 |
| P14 cycles | 0.00 |
| Inter-iter dependencies cycles | NA |
| FE+BE cycles (UFS) | NA |
| Stall cycles (UFS) | NA |
| Nb insns | 26.00 |
| Nb uops | 26.00 |
| Nb loads | NA |
| Nb stores | 0.00 |
| Nb stack references | 0.00 |
| FLOP/cycle | 0.00 |
| Nb FLOP add-sub | 0.00 |
| Nb FLOP mul | 0.00 |
| Nb FLOP fma | 0.00 |
| Nb FLOP div | 0.00 |
| Nb FLOP rcp | 0.00 |
| Nb FLOP sqrt | 0.00 |
| Nb FLOP rsqrt | 0.00 |
| Bytes/cycle | 16.00 |
| Bytes prefetched | 0.00 |
| Bytes loaded | 64.00 |
| Bytes stored | 0.00 |
| Stride 0 | NA |
| Stride 1 | NA |
| Stride n | NA |
| Stride unknown | NA |
| Stride indirect | NA |
| Vectorization ratio all | 0.00 |
| Vectorization ratio load | NA |
| Vectorization ratio store | NA |
| Vectorization ratio mul | NA |
| Vectorization ratio add_sub | 0.00 |
| Vectorization ratio fma | 0.00 |
| Vectorization ratio div_sqrt | NA |
| Vectorization ratio other | 0.00 |
| Vector-efficiency ratio all | 22.50 |
| Vector-efficiency ratio load | NA |
| Vector-efficiency ratio store | NA |
| Vector-efficiency ratio mul | NA |
| Vector-efficiency ratio add_sub | 22.92 |
| Vector-efficiency ratio fma | 25.00 |
| Vector-efficiency ratio div_sqrt | NA |
| Vector-efficiency ratio other | 20.83 |
| Metric | Value |
|---|---|
| CQA speedup if no scalar integer | 1.00 |
| CQA speedup if FP arith vectorized | 1.00 |
| CQA speedup if fully vectorized | 5.33 |
| CQA speedup if no inter-iteration dependency | NA |
| CQA speedup if next bottleneck killed | 1.23 |
| Bottlenecks | P2, P3, P4, P5, |
| Function | advec_mom_kernel |
| Source | advec_mom_kernel.f90-pp.f90:151-151,advec_mom_kernel.f90-pp.f90:178-178 |
| Source loop unroll info | NA |
| Source loop unroll confidence level | NA |
| Unroll/vectorization loop type | NA |
| Unroll factor | NA |
| CQA cycles | 4.00 |
| CQA cycles if no scalar integer | 4.00 |
| CQA cycles if FP arith vectorized | 4.00 |
| CQA cycles if fully vectorized | 0.75 |
| Front-end cycles | 3.25 |
| DIV/SQRT cycles | 1.50 |
| P0 cycles | 1.50 |
| P1 cycles | 4.00 |
| P2 cycles | 4.00 |
| P3 cycles | 4.00 |
| P4 cycles | 4.00 |
| P5 cycles | 0.00 |
| P6 cycles | 0.00 |
| P7 cycles | 0.00 |
| P8 cycles | 0.00 |
| P9 cycles | 2.33 |
| P10 cycles | 2.33 |
| P11 cycles | 2.33 |
| P12 cycles | 0.00 |
| P13 cycles | 0.00 |
| P14 cycles | 0.00 |
| Inter-iter dependencies cycles | NA |
| FE+BE cycles (UFS) | NA |
| Stall cycles (UFS) | NA |
| Nb insns | 26.00 |
| Nb uops | 26.00 |
| Nb loads | NA |
| Nb stores | 0.00 |
| Nb stack references | 0.00 |
| FLOP/cycle | 0.00 |
| Nb FLOP add-sub | 0.00 |
| Nb FLOP mul | 0.00 |
| Nb FLOP fma | 0.00 |
| Nb FLOP div | 0.00 |
| Nb FLOP rcp | 0.00 |
| Nb FLOP sqrt | 0.00 |
| Nb FLOP rsqrt | 0.00 |
| Bytes/cycle | 16.00 |
| Bytes prefetched | 0.00 |
| Bytes loaded | 64.00 |
| Bytes stored | 0.00 |
| Stride 0 | NA |
| Stride 1 | NA |
| Stride n | NA |
| Stride unknown | NA |
| Stride indirect | NA |
| Vectorization ratio all | 0.00 |
| Vectorization ratio load | NA |
| Vectorization ratio store | NA |
| Vectorization ratio mul | NA |
| Vectorization ratio add_sub | 0.00 |
| Vectorization ratio fma | 0.00 |
| Vectorization ratio div_sqrt | NA |
| Vectorization ratio other | 0.00 |
| Vector-efficiency ratio all | 22.50 |
| Vector-efficiency ratio load | NA |
| Vector-efficiency ratio store | NA |
| Vector-efficiency ratio mul | NA |
| Vector-efficiency ratio add_sub | 22.92 |
| Vector-efficiency ratio fma | 25.00 |
| Vector-efficiency ratio div_sqrt | NA |
| Vector-efficiency ratio other | 20.83 |
| Path / |
| Function | advec_mom_kernel |
| Source file and lines | advec_mom_kernel.f90-pp.f90:151-178 |
| Module | exec |
| nb instructions | 26 |
| loop length | 104 |
| nb stack references | 0 |
| front end | 3.25 cycles |
| P0 | P1 | P2 | P3 | P4 | P5 | P6 | P7 | P8 | P9 | P10 | P11 | P12 | P13 | P14 | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| uops | 1.50 | 1.50 | 4.00 | 4.00 | 4.00 | 4.00 | 0.00 | 0.00 | 0.00 | 0.00 | 2.33 | 2.33 | 2.33 | 0.00 | 0.00 |
| cycles | 1.50 | 1.50 | 4.00 | 4.00 | 4.00 | 4.00 | 0.00 | 0.00 | 0.00 | 0.00 | 2.33 | 2.33 | 2.33 | 0.00 | 0.00 |
| Cycles executing div or sqrt instructions | NA |
| Front-end | 3.25 |
| Overall L1 | 4.00 |
| all | 0% |
| load | NA (no load vectorizable/vectorized instructions) |
| store | NA (no store vectorizable/vectorized instructions) |
| mul | NA (no mul vectorizable/vectorized instructions) |
| add-sub | 0% |
| fma | 0% |
| div/sqrt | NA (no div/sqrt vectorizable/vectorized instructions) |
| other | 0% |
| Instruction | Nb FU | P0 | P1 | P2 | P3 | P4 | P5 | P6 | P7 | P8 | P9 | P10 | P11 | P12 | P13 | P14 | Latency | Recip. throughput |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| ADD X8, X8, #1 | 1 | 0 | 0 | 0.25 | 0.25 | 0.25 | 0.25 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0.25 |
| SUBS W9, W9, #1 | 1 | 0 | 0 | 0.25 | 0.25 | 0.25 | 0.25 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0.33 |
| B.LE 41f9c0 <__nv_advec_mom_kernel_mod_advec_mom_kernel__F1L79_1_+0x12a0> | 1 | 0.50 | 0.50 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0.50 |
| CMP W11, #1 | 1 | 0 | 0 | 0.25 | 0.25 | 0.25 | 0.25 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0.33 |
| B.LT 41f860 <__nv_advec_mom_kernel_mod_advec_mom_kernel__F1L79_1_+0x1140> | 1 | 0.50 | 0.50 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0.50 |
| LDR X16, [X20, #64] | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.33 | 0.33 | 0.33 | 0 | 0 | 4 | 0.33 |
| LDR X17, [X20, #256] | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.33 | 0.33 | 0.33 | 0 | 0 | 4 | 0.33 |
| LDP X1, X2, [X20, #288] | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.33 | 0.33 | 0.33 | 0 | 0 | 4 | 1 |
| ORR X15, XZR, XZR | 1 | 0 | 0 | 0.25 | 0.25 | 0.25 | 0.25 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0.25 |
| LDR X16, [X16] | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.33 | 0.33 | 0.33 | 0 | 0 | 4 | 0.33 |
| MADD X18, X16, X8, XZR | 1 | 0 | 0 | 0 | 0 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 2 | 1 |
| LDR X16, [X20, #104] | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.33 | 0.33 | 0.33 | 0 | 0 | 4 | 0.33 |
| LDR X0, [X16] | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.33 | 0.33 | 0.33 | 0 | 0 | 4 | 0.33 |
| SUB X16, X18, X0 | 1 | 0 | 0 | 0.25 | 0.25 | 0.25 | 0.25 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0.25 |
| ADD X18, X12, X18 | 1 | 0 | 0 | 0.25 | 0.25 | 0.25 | 0.25 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0.25 |
| UBFM X3, X16, #61, #60 | 1 | 0 | 0 | 0.25 | 0.25 | 0.25 | 0.25 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0.25 |
| SUB X18, X18, X0 | 1 | 0 | 0 | 0.25 | 0.25 | 0.25 | 0.25 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0.25 |
| ADD X16, X17, X3 | 1 | 0 | 0 | 0.25 | 0.25 | 0.25 | 0.25 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0.25 |
| ADD X17, X1, X3 | 1 | 0 | 0 | 0.25 | 0.25 | 0.25 | 0.25 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0.25 |
| LDR X1, [X20, #224] | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.33 | 0.33 | 0.33 | 0 | 0 | 4 | 0.33 |
| UBFM X0, X18, #61, #60 | 1 | 0 | 0 | 0.25 | 0.25 | 0.25 | 0.25 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0.25 |
| ADD X18, X2, X0 | 1 | 0 | 0 | 0.25 | 0.25 | 0.25 | 0.25 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0.25 |
| ORR X2, XZR, X12 | 1 | 0 | 0 | 0.25 | 0.25 | 0.25 | 0.25 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0.25 |
| ADD X0, X1, X0 | 1 | 0 | 0 | 0.25 | 0.25 | 0.25 | 0.25 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0.25 |
| ORR W1, WZR, W14 | 1 | 0 | 0 | 0.25 | 0.25 | 0.25 | 0.25 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0.25 |
| B 41f8ec <__nv_advec_mom_kernel_mod_advec_mom_kernel__F1L79_1_+0x11cc> | 1 | 0.50 | 0.50 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0.50 |
| Function | advec_mom_kernel |
| Source file and lines | advec_mom_kernel.f90-pp.f90:151-178 |
| Module | exec |
| nb instructions | 26 |
| loop length | 104 |
| nb stack references | 0 |
| front end | 3.25 cycles |
| P0 | P1 | P2 | P3 | P4 | P5 | P6 | P7 | P8 | P9 | P10 | P11 | P12 | P13 | P14 | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| uops | 1.50 | 1.50 | 4.00 | 4.00 | 4.00 | 4.00 | 0.00 | 0.00 | 0.00 | 0.00 | 2.33 | 2.33 | 2.33 | 0.00 | 0.00 |
| cycles | 1.50 | 1.50 | 4.00 | 4.00 | 4.00 | 4.00 | 0.00 | 0.00 | 0.00 | 0.00 | 2.33 | 2.33 | 2.33 | 0.00 | 0.00 |
| Cycles executing div or sqrt instructions | NA |
| Front-end | 3.25 |
| Overall L1 | 4.00 |
| all | 0% |
| load | NA (no load vectorizable/vectorized instructions) |
| store | NA (no store vectorizable/vectorized instructions) |
| mul | NA (no mul vectorizable/vectorized instructions) |
| add-sub | 0% |
| fma | 0% |
| div/sqrt | NA (no div/sqrt vectorizable/vectorized instructions) |
| other | 0% |
| Instruction | Nb FU | P0 | P1 | P2 | P3 | P4 | P5 | P6 | P7 | P8 | P9 | P10 | P11 | P12 | P13 | P14 | Latency | Recip. throughput |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| ADD X8, X8, #1 | 1 | 0 | 0 | 0.25 | 0.25 | 0.25 | 0.25 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0.25 |
| SUBS W9, W9, #1 | 1 | 0 | 0 | 0.25 | 0.25 | 0.25 | 0.25 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0.33 |
| B.LE 41f9c0 <__nv_advec_mom_kernel_mod_advec_mom_kernel__F1L79_1_+0x12a0> | 1 | 0.50 | 0.50 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0.50 |
| CMP W11, #1 | 1 | 0 | 0 | 0.25 | 0.25 | 0.25 | 0.25 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0.33 |
| B.LT 41f860 <__nv_advec_mom_kernel_mod_advec_mom_kernel__F1L79_1_+0x1140> | 1 | 0.50 | 0.50 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0.50 |
| LDR X16, [X20, #64] | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.33 | 0.33 | 0.33 | 0 | 0 | 4 | 0.33 |
| LDR X17, [X20, #256] | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.33 | 0.33 | 0.33 | 0 | 0 | 4 | 0.33 |
| LDP X1, X2, [X20, #288] | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.33 | 0.33 | 0.33 | 0 | 0 | 4 | 1 |
| ORR X15, XZR, XZR | 1 | 0 | 0 | 0.25 | 0.25 | 0.25 | 0.25 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0.25 |
| LDR X16, [X16] | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.33 | 0.33 | 0.33 | 0 | 0 | 4 | 0.33 |
| MADD X18, X16, X8, XZR | 1 | 0 | 0 | 0 | 0 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 2 | 1 |
| LDR X16, [X20, #104] | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.33 | 0.33 | 0.33 | 0 | 0 | 4 | 0.33 |
| LDR X0, [X16] | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.33 | 0.33 | 0.33 | 0 | 0 | 4 | 0.33 |
| SUB X16, X18, X0 | 1 | 0 | 0 | 0.25 | 0.25 | 0.25 | 0.25 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0.25 |
| ADD X18, X12, X18 | 1 | 0 | 0 | 0.25 | 0.25 | 0.25 | 0.25 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0.25 |
| UBFM X3, X16, #61, #60 | 1 | 0 | 0 | 0.25 | 0.25 | 0.25 | 0.25 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0.25 |
| SUB X18, X18, X0 | 1 | 0 | 0 | 0.25 | 0.25 | 0.25 | 0.25 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0.25 |
| ADD X16, X17, X3 | 1 | 0 | 0 | 0.25 | 0.25 | 0.25 | 0.25 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0.25 |
| ADD X17, X1, X3 | 1 | 0 | 0 | 0.25 | 0.25 | 0.25 | 0.25 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0.25 |
| LDR X1, [X20, #224] | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.33 | 0.33 | 0.33 | 0 | 0 | 4 | 0.33 |
| UBFM X0, X18, #61, #60 | 1 | 0 | 0 | 0.25 | 0.25 | 0.25 | 0.25 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0.25 |
| ADD X18, X2, X0 | 1 | 0 | 0 | 0.25 | 0.25 | 0.25 | 0.25 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0.25 |
| ORR X2, XZR, X12 | 1 | 0 | 0 | 0.25 | 0.25 | 0.25 | 0.25 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0.25 |
| ADD X0, X1, X0 | 1 | 0 | 0 | 0.25 | 0.25 | 0.25 | 0.25 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0.25 |
| ORR W1, WZR, W14 | 1 | 0 | 0 | 0.25 | 0.25 | 0.25 | 0.25 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0.25 |
| B 41f8ec <__nv_advec_mom_kernel_mod_advec_mom_kernel__F1L79_1_+0x11cc> | 1 | 0.50 | 0.50 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0.50 |
