| Loop Id: 40 | Module: exec | Source: cg.cpp:105-113 | Coverage: 0.02% |
|---|
| Loop Id: 40 | Module: exec | Source: cg.cpp:105-113 | Coverage: 0.02% |
|---|
(36) 0x402d48 CMP W15, W20 |
(36) 0x402d4c B.GE 402edc |
(37) 0x402d50 SBFM X24, X14, #0, #31 |
(37) 0x402d54 SBFM X4, X15, #0, #31 |
(37) 0x402d58 CMP W21, #3 |
(37) 0x402d5c B.LS 402e98 |
(37) 0x402d60 ORR X4, XZR, X18 |
(39) 0x402d64 SBFM X24, X14, #0, #31 |
(39) 0x402d68 ADD X2, X24, X18 |
(39) 0x402d6c UBFM X5, X2, #61, #60 |
(39) 0x402d70 ADD X2, X8, X2,LSL #3 |
(39) 0x402d74 ADD X1, X5, #16 |
(39) 0x402d78 ADD X3, X9, X5 |
(39) 0x402d7c ADD X6, X12, X1 |
(39) 0x402d80 ADD X25, X9, X1 |
(39) 0x402d84 CMP X6, X3 |
(39) 0x402d88 ADD X6, X12, X5 |
(39) 0x402d8c ADD X11, X8, X1 |
(39) 0x402d90 CCMP X25, X6, #0, #8 |
(39) 0x402d94 CSINC W23, WZR, WZR, #8 |
(39) 0x402d98 ADD X1, X5, #8 |
(39) 0x402d9c CMP X3, X11 |
(39) 0x402da0 ADD X26, X12, X1 |
(39) 0x402da4 CCMP X2, X25, #2, #3 |
(39) 0x402da8 CSINC W11, WZR, WZR, #3 |
(39) 0x402dac ADD X1, X13, X1 |
(39) 0x402db0 CMP X26, X2 |
(39) 0x402db4 CCMP X2, X1, #4, #1 |
(39) 0x402db8 AND W11, W23, W11 |
(39) 0x402dbc CCMP X3, X1, #4, #1 |
(39) 0x402dc0 CSINC W1, WZR, WZR, #0 |
(39) 0x402dc4 ANDS WZR, W1, W11 |
(39) 0x402dc8 B.EQ 402e98 |
(39) 0x402dcc MOVZ X1, #0 |
(39) 0x402dd0 ADD X5, X13, X5 |
(39) 0x402dd4 HINT #0 |
(39) 0x402dd8 HINT #0 |
(39) 0x402ddc HINT #0 |
(38) 0x402de0 LDR Q29, [X5, X1] |
(38) 0x402de4 LDR Q31, [X3, X1] |
(38) 0x402de8 FMLA V31.2D, V29.2D, V28.2D |
(38) 0x402dec STR Q31, [X3, X1] |
(38) 0x402df0 LDR Q29, [X6, X1] |
(38) 0x402df4 LDR Q31, [X2, X1] |
(38) 0x402df8 FMLS V31.2D, V28.2D, V29.2D |
(38) 0x402dfc FMUL V29.2D, V31.2D, V31.2D |
(38) 0x402e00 STR Q31, [X2, X1] |
(38) 0x402e04 ADD X1, X1, #16 |
(38) 0x402e08 FADD D30, D30, D29 |
(38) 0x402e0c MOV D31, V29.D[1] |
(38) 0x402e10 FADD D30, D31, D30 |
(38) 0x402e14 CMP X7, X1 |
(38) 0x402e18 B.NE 402de0 |
(39) 0x402e1c CBZ W16, 402e4c |
(39) 0x402e20 ADD W1, W30, W14 |
(39) 0x402e24 SBFM X1, X1, #61, #31 |
(39) 0x402e28 LDR D29, [X13, X1] |
(39) 0x402e2c LDR D31, [X9, X1] |
(39) 0x402e30 FMADD D31, D27, D29, D31 |
(39) 0x402e34 STR D31, [X9, X1] |
(39) 0x402e38 LDR D29, [X12, X1] |
(39) 0x402e3c LDR D31, [X8, X1] |
(39) 0x402e40 FMSUB D31, D27, D29, D31 |
(39) 0x402e44 FMADD D30, D31, D31, D30 |
(39) 0x402e48 STR D31, [X8, X1] |
(39) 0x402e4c ADD W10, W10, #1 |
(39) 0x402e50 ADD W14, W14, W17 |
(39) 0x402e54 CMP W0, W10 |
(39) 0x402e58 B.GT 402d64 |
0x402e5c LDP X21, X22, [SP, #32] |
0x402e60 LDP X23, X24, [SP, #48] |
0x402e64 LDP X25, X26, [SP, #64] |
0x402e68 ADD X19, X19, #40 |
0x402e6c LDR X0, [X19] |
(34) 0x402e70 ORR X1, XZR, X0 |
(34) 0x402e74 FMOV D31, X0 |
(34) 0x402e78 FADD D31, D30, D31 |
(34) 0x402e7c FMOV X2, D31 |
(34) 0x402e80 CAS X1, X2, [X19] |
(34) 0x402e84 CMP X0, X1 |
(34) 0x402e88 B.NE 402efc |
0x402e8c LDP X19, X20, [SP, #16] |
0x402e90 LDP X29, X30, [SP], #80 |
0x402e94 RET |
(37) 0x402e98 ADD X2, X22, X4 |
(37) 0x402e9c ADD X1, X24, X4 |
(37) 0x402ea0 ADD X2, X2, X24 |
(37) 0x402ea4 UBFM X1, X1, #61, #60 |
(37) 0x402ea8 UBFM X2, X2, #61, #60 |
(35) 0x402eac LDR D29, [X13, X1] |
(35) 0x402eb0 LDR D31, [X9, X1] |
(35) 0x402eb4 FMADD D31, D27, D29, D31 |
(35) 0x402eb8 STR D31, [X9, X1] |
(35) 0x402ebc LDR D29, [X12, X1] |
(35) 0x402ec0 LDR D31, [X8, X1] |
(35) 0x402ec4 FMSUB D31, D27, D29, D31 |
(35) 0x402ec8 FMADD D30, D31, D31, D30 |
(35) 0x402ecc STR D31, [X8, X1] |
(35) 0x402ed0 ADD X1, X1, #8 |
(35) 0x402ed4 CMP X2, X1 |
(35) 0x402ed8 B.NE 402eac |
(36) 0x402edc ADD W10, W10, #1 |
(36) 0x402ee0 ADD W14, W14, W17 |
(36) 0x402ee4 CMP W0, W10 |
(36) 0x402ee8 B.GT 402d48 |
0x402eec B 402e5c |
(34) 0x402efc ORR X0, XZR, X1 |
(34) 0x402f00 B 402e70 |
/home/hbollore/qaas/qaas-runs/174-118-5752/intel/TeaLeaf/build/TeaLeaf/src/omp/cg.cpp: 105 - 113 |
-------------------------------------------------------------------------------- |
105: #pragma omp parallel for reduction(+ : rrn_temp) |
106: #endif |
107: for (int jj = halo_depth; jj < y - halo_depth; ++jj) { |
108: for (int kk = halo_depth; kk < x - halo_depth; ++kk) { |
109: const int index = kk + jj * x; |
110: |
111: u[index] += alpha * p[index]; |
112: r[index] -= alpha * w[index]; |
113: rrn_temp += r[index] * r[index]; |
| Coverage (%) | Name | Source Location | Module |
|---|---|---|---|
| ►90.20+ | gomp_thread_start | team.c:130 | libgomp.so.1.0.0 |
| ○ | start_thread | libc.so.6 | |
| ○ | thread_start | libc.so.6 | |
| ►9.65+ | start_thread | libc.so.6 | |
| ○ | thread_start | libc.so.6 |
| min | med | avg | max |
|---|---|---|---|
| Percentile Index | 10 | 20 | 30 | 40 | 50 | 60 | 70 | 80 | 90 | 100 |
|---|---|---|---|---|---|---|---|---|---|---|
| Value |
| min | med | avg | max |
|---|---|---|---|
| Percentile Index | 10 | 20 | 30 | 40 | 50 | 60 | 70 | 80 | 90 | 100 |
|---|---|---|---|---|---|---|---|---|---|---|
| Value |
| Path / |
| Metric | Value |
|---|---|
| CQA speedup if no scalar integer | 1.00 |
| CQA speedup if FP arith vectorized | 1.00 |
| CQA speedup if fully vectorized | 1.09 |
| CQA speedup if no inter-iteration dependency | NA |
| CQA speedup if next bottleneck killed | 1.78 |
| Bottlenecks | P12, P13, P14, |
| Function | cg_calc_ur(int, int, int, double, double*, double*, double const*, double*, double const*) [clone ._omp_fn.0] |
| Source | cg.cpp:105-105,cg.cpp:108-108 |
| Source loop unroll info | NA |
| Source loop unroll confidence level | NA |
| Unroll/vectorization loop type | NA |
| Unroll factor | NA |
| CQA cycles | 2.00 |
| CQA cycles if no scalar integer | 2.00 |
| CQA cycles if FP arith vectorized | 2.00 |
| CQA cycles if fully vectorized | 1.83 |
| Front-end cycles | 1.13 |
| P0 cycles | 1.00 |
| P1 cycles | 1.00 |
| P2 cycles | 0.17 |
| P3 cycles | 0.17 |
| P4 cycles | 0.17 |
| P5 cycles | 0.17 |
| P6 cycles | 0.17 |
| P7 cycles | 0.17 |
| P8 cycles | 0.00 |
| P9 cycles | 0.00 |
| P10 cycles | 0.00 |
| P11 cycles | 0.00 |
| P12 cycles | 2.00 |
| P13 cycles | 2.00 |
| P14 cycles | 2.00 |
| P15 cycles | 0.00 |
| P16 cycles | 0.00 |
| DIV/SQRT cycles | 0.00 |
| Inter-iter dependencies cycles | NA |
| FE+BE cycles (UFS) | NA |
| Stall cycles (UFS) | NA |
| Nb insns | 9.00 |
| Nb uops | 9.00 |
| Nb loads | NA |
| Nb stores | 0.00 |
| Nb stack references | 0.00 |
| FLOP/cycle | 0.00 |
| Nb FLOP add-sub | 0.00 |
| Nb FLOP mul | 0.00 |
| Nb FLOP fma | 0.00 |
| Nb FLOP div | 0.00 |
| Nb FLOP rcp | 0.00 |
| Nb FLOP sqrt | 0.00 |
| Nb FLOP rsqrt | 0.00 |
| Bytes/cycle | 44.00 |
| Bytes prefetched | 0.00 |
| Bytes loaded | 88.00 |
| Bytes stored | 0.00 |
| Stride 0 | NA |
| Stride 1 | NA |
| Stride n | NA |
| Stride unknown | NA |
| Stride indirect | NA |
| Vectorization ratio all | 0.00 |
| Vectorization ratio load | 0.00 |
| Vectorization ratio store | NA |
| Vectorization ratio mul | NA |
| Vectorization ratio add_sub | NA |
| Vectorization ratio fma | NA |
| Vectorization ratio div_sqrt | NA |
| Vectorization ratio other | NA |
| Vector-efficiency ratio all | 90.00 |
| Vector-efficiency ratio load | 90.00 |
| Vector-efficiency ratio store | NA |
| Vector-efficiency ratio mul | NA |
| Vector-efficiency ratio add_sub | NA |
| Vector-efficiency ratio fma | NA |
| Vector-efficiency ratio div_sqrt | NA |
| Vector-efficiency ratio other | NA |
| Metric | Value |
|---|---|
| CQA speedup if no scalar integer | 1.00 |
| CQA speedup if FP arith vectorized | 1.00 |
| CQA speedup if fully vectorized | 1.09 |
| CQA speedup if no inter-iteration dependency | NA |
| CQA speedup if next bottleneck killed | 1.78 |
| Bottlenecks | P12, P13, P14, |
| Function | cg_calc_ur(int, int, int, double, double*, double*, double const*, double*, double const*) [clone ._omp_fn.0] |
| Source | cg.cpp:105-105,cg.cpp:108-108 |
| Source loop unroll info | NA |
| Source loop unroll confidence level | NA |
| Unroll/vectorization loop type | NA |
| Unroll factor | NA |
| CQA cycles | 2.00 |
| CQA cycles if no scalar integer | 2.00 |
| CQA cycles if FP arith vectorized | 2.00 |
| CQA cycles if fully vectorized | 1.83 |
| Front-end cycles | 1.13 |
| P0 cycles | 1.00 |
| P1 cycles | 1.00 |
| P2 cycles | 0.17 |
| P3 cycles | 0.17 |
| P4 cycles | 0.17 |
| P5 cycles | 0.17 |
| P6 cycles | 0.17 |
| P7 cycles | 0.17 |
| P8 cycles | 0.00 |
| P9 cycles | 0.00 |
| P10 cycles | 0.00 |
| P11 cycles | 0.00 |
| P12 cycles | 2.00 |
| P13 cycles | 2.00 |
| P14 cycles | 2.00 |
| P15 cycles | 0.00 |
| P16 cycles | 0.00 |
| DIV/SQRT cycles | 0.00 |
| Inter-iter dependencies cycles | NA |
| FE+BE cycles (UFS) | NA |
| Stall cycles (UFS) | NA |
| Nb insns | 9.00 |
| Nb uops | 9.00 |
| Nb loads | NA |
| Nb stores | 0.00 |
| Nb stack references | 0.00 |
| FLOP/cycle | 0.00 |
| Nb FLOP add-sub | 0.00 |
| Nb FLOP mul | 0.00 |
| Nb FLOP fma | 0.00 |
| Nb FLOP div | 0.00 |
| Nb FLOP rcp | 0.00 |
| Nb FLOP sqrt | 0.00 |
| Nb FLOP rsqrt | 0.00 |
| Bytes/cycle | 44.00 |
| Bytes prefetched | 0.00 |
| Bytes loaded | 88.00 |
| Bytes stored | 0.00 |
| Stride 0 | NA |
| Stride 1 | NA |
| Stride n | NA |
| Stride unknown | NA |
| Stride indirect | NA |
| Vectorization ratio all | 0.00 |
| Vectorization ratio load | 0.00 |
| Vectorization ratio store | NA |
| Vectorization ratio mul | NA |
| Vectorization ratio add_sub | NA |
| Vectorization ratio fma | NA |
| Vectorization ratio div_sqrt | NA |
| Vectorization ratio other | NA |
| Vector-efficiency ratio all | 90.00 |
| Vector-efficiency ratio load | 90.00 |
| Vector-efficiency ratio store | NA |
| Vector-efficiency ratio mul | NA |
| Vector-efficiency ratio add_sub | NA |
| Vector-efficiency ratio fma | NA |
| Vector-efficiency ratio div_sqrt | NA |
| Vector-efficiency ratio other | NA |
| Path / |
| Function | cg_calc_ur(int, int, int, double, double*, double*, double const*, double*, double const*) [clone ._omp_fn.0] |
| Source file and lines | cg.cpp:105-113 |
| Module | exec |
| nb instructions | 9 |
| loop length | 36 |
| nb stack references | 0 |
| front end | 1.13 cycles |
| P0 | P1 | P2 | P3 | P4 | P5 | P6 | P7 | P8 | P9 | P10 | P11 | P12 | P13 | P14 | P15 | P16 | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| uops | 1.00 | 1.00 | 0.17 | 0.17 | 0.17 | 0.17 | 0.17 | 0.17 | 0.00 | 0.00 | 0.00 | 0.00 | 2.00 | 2.00 | 2.00 | 0.00 | 0.00 |
| cycles | 1.00 | 1.00 | 0.17 | 0.17 | 0.17 | 0.17 | 0.17 | 0.17 | 0.00 | 0.00 | 0.00 | 0.00 | 2.00 | 2.00 | 2.00 | 0.00 | 0.00 |
| Cycles executing div or sqrt instructions | NA |
| Front-end | 1.13 |
| Overall L1 | 2.00 |
| all | 0% |
| load | 0% |
| store | NA (no store vectorizable/vectorized instructions) |
| mul | NA (no mul vectorizable/vectorized instructions) |
| add-sub | NA (no add-sub vectorizable/vectorized instructions) |
| fma | NA (no fma vectorizable/vectorized instructions) |
| div/sqrt | NA (no div/sqrt vectorizable/vectorized instructions) |
| other | NA (no other vectorizable/vectorized instructions) |
| Instruction | Nb FU | P0 | P1 | P2 | P3 | P4 | P5 | P6 | P7 | P8 | P9 | P10 | P11 | P12 | P13 | P14 | P15 | P16 | Latency | Recip. throughput | Vectorization |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| LDP X21, X22, [SP, #32] | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.33 | 0.33 | 0.33 | 0 | 0 | 4 | 0.50 | scal (100.0%) |
| LDP X23, X24, [SP, #48] | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.33 | 0.33 | 0.33 | 0 | 0 | 4 | 0.50 | scal (100.0%) |
| LDP X25, X26, [SP, #64] | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.33 | 0.33 | 0.33 | 0 | 0 | 4 | 0.50 | scal (100.0%) |
| ADD X19, X19, #40 | 1 | 0 | 0 | 0.17 | 0.17 | 0.17 | 0.17 | 0.17 | 0.17 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0.17 | N/A |
| LDR X0, [X19] | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.33 | 0.33 | 0.33 | 0 | 0 | 4 | 0.33 | scal (50.0%) |
| LDP X19, X20, [SP, #16] | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.33 | 0.33 | 0.33 | 0 | 0 | 4 | 0.50 | N/A |
| LDP X29, X30, [SP], #80 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.33 | 0.33 | 0.33 | 0 | 0 | 4 | 0.50 | scal (100.0%) |
| RET | 1 | 0.50 | 0.50 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0.50 | N/A |
| B 402e5c <_Z10cg_calc_uriiidPdS_PKdS_S1_._omp_fn.0+0x1b0> | 1 | 0.50 | 0.50 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0.50 | N/A |
| Function | cg_calc_ur(int, int, int, double, double*, double*, double const*, double*, double const*) [clone ._omp_fn.0] |
| Source file and lines | cg.cpp:105-113 |
| Module | exec |
| nb instructions | 9 |
| loop length | 36 |
| nb stack references | 0 |
| front end | 1.13 cycles |
| P0 | P1 | P2 | P3 | P4 | P5 | P6 | P7 | P8 | P9 | P10 | P11 | P12 | P13 | P14 | P15 | P16 | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| uops | 1.00 | 1.00 | 0.17 | 0.17 | 0.17 | 0.17 | 0.17 | 0.17 | 0.00 | 0.00 | 0.00 | 0.00 | 2.00 | 2.00 | 2.00 | 0.00 | 0.00 |
| cycles | 1.00 | 1.00 | 0.17 | 0.17 | 0.17 | 0.17 | 0.17 | 0.17 | 0.00 | 0.00 | 0.00 | 0.00 | 2.00 | 2.00 | 2.00 | 0.00 | 0.00 |
| Cycles executing div or sqrt instructions | NA |
| Front-end | 1.13 |
| Overall L1 | 2.00 |
| all | 0% |
| load | 0% |
| store | NA (no store vectorizable/vectorized instructions) |
| mul | NA (no mul vectorizable/vectorized instructions) |
| add-sub | NA (no add-sub vectorizable/vectorized instructions) |
| fma | NA (no fma vectorizable/vectorized instructions) |
| div/sqrt | NA (no div/sqrt vectorizable/vectorized instructions) |
| other | NA (no other vectorizable/vectorized instructions) |
| Instruction | Nb FU | P0 | P1 | P2 | P3 | P4 | P5 | P6 | P7 | P8 | P9 | P10 | P11 | P12 | P13 | P14 | P15 | P16 | Latency | Recip. throughput | Vectorization |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| LDP X21, X22, [SP, #32] | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.33 | 0.33 | 0.33 | 0 | 0 | 4 | 0.50 | scal (100.0%) |
| LDP X23, X24, [SP, #48] | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.33 | 0.33 | 0.33 | 0 | 0 | 4 | 0.50 | scal (100.0%) |
| LDP X25, X26, [SP, #64] | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.33 | 0.33 | 0.33 | 0 | 0 | 4 | 0.50 | scal (100.0%) |
| ADD X19, X19, #40 | 1 | 0 | 0 | 0.17 | 0.17 | 0.17 | 0.17 | 0.17 | 0.17 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0.17 | N/A |
| LDR X0, [X19] | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.33 | 0.33 | 0.33 | 0 | 0 | 4 | 0.33 | scal (50.0%) |
| LDP X19, X20, [SP, #16] | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.33 | 0.33 | 0.33 | 0 | 0 | 4 | 0.50 | N/A |
| LDP X29, X30, [SP], #80 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.33 | 0.33 | 0.33 | 0 | 0 | 4 | 0.50 | scal (100.0%) |
| RET | 1 | 0.50 | 0.50 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0.50 | N/A |
| B 402e5c <_Z10cg_calc_uriiidPdS_PKdS_S1_._omp_fn.0+0x1b0> | 1 | 0.50 | 0.50 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0.50 | N/A |
