| Function: hypre_swap2 | Module: exec | Source: hypre_qsort.c:47-53 | Coverage (incl. loops): 0.01% | (excl. loops): 0.01% |
|---|
| Function: hypre_swap2 | Module: exec | Source: hypre_qsort.c:47-53 | Coverage (incl. loops): 0.01% | (excl. loops): 0.01% |
|---|
/home/hbollore/qaas/qaas-runs/174-161-6712/intel/AMG/build/AMG/AMG/utilities/hypre_qsort.c: 47 - 53 |
-------------------------------------------------------------------------------- |
47: temp = v[i]; |
48: v[i] = v[j]; |
49: v[j] = temp; |
50: temp2 = w[i]; |
51: w[i] = w[j]; |
52: w[j] = temp2; |
53: } |
0x4a7320 LDR D0, [X1, X2,LSL #3] |
0x4a7324 LDR D1, [X1, X3,LSL #3] |
0x4a7328 LDR X8, [X0, X2,LSL #3] |
0x4a732c LDR X9, [X0, X3,LSL #3] |
0x4a7330 STR X9, [X0, X2,LSL #3] |
0x4a7334 STR X8, [X0, X3,LSL #3] |
0x4a7338 STR D1, [X1, X2,LSL #3] |
0x4a733c STR D0, [X1, X3,LSL #3] |
0x4a7340 RET |
0x4a7344 HINT #0 |
0x4a7348 HINT #0 |
0x4a734c HINT #0 |
| Coverage (%) | Name | Source Location | Module |
|---|---|---|---|
| ►57.14+ | hypre_qsort2abs | par_interp.c:3191 | exec |
| ○ | hypre_BoomerAMGInterpTruncatio[...] | par_interp.c:2912 | exec |
| ○ | __kmp_invoke_microtask | libomp.so | |
| ►42.86+ | hypre_BoomerAMGInterpTruncatio[...] | par_interp.c:2912 | exec |
| ○ | __kmp_invoke_microtask | libomp.so |
| Coverage (%) | Name | Source Location | Module |
|---|---|---|---|
| ►66.67+ | hypre_qsort2abs | par_interp.c:3191 | exec |
| ○ | hypre_BoomerAMGInterpTruncatio[...] | par_interp.c:2912 | exec |
| ○ | __kmp_invoke_microtask | libomp.so | |
| ►33.33+ | hypre_BoomerAMGInterpTruncatio[...] | par_interp.c:2912 | exec |
| ○ | __kmp_invoke_microtask | libomp.so |
| Coverage (%) | Name | Source Location | Module |
|---|---|---|---|
| ►50.00+ | hypre_BoomerAMGInterpTruncatio[...] | par_interp.c:2912 | exec |
| ○ | __kmp_invoke_microtask | libomp.so | |
| ►25.00+ | hypre_qsort2abs | par_interp.c:3191 | exec |
| ○ | hypre_qsort2abs | par_interp.c:3191 | exec |
| ○ | hypre_BoomerAMGInterpTruncatio[...] | par_interp.c:2912 | exec |
| ○ | __kmp_invoke_microtask | libomp.so | |
| ►12.50+ | hypre_qsort2abs | par_interp.c:3191 | exec |
| ○ | hypre_BoomerAMGInterpTruncatio[...] | par_interp.c:2912 | exec |
| ○ | __kmp_invoke_microtask | libomp.so | |
| ►12.50+ | hypre_qsort2abs | par_interp.c:3191 | exec |
| ○ | hypre_qsort2abs | par_interp.c:3191 | exec |
| ○ | hypre_qsort2abs | par_interp.c:3191 | exec |
| ○ | hypre_BoomerAMGInterpTruncatio[...] | par_interp.c:2912 | exec |
| ○ | __kmp_invoke_microtask | libomp.so |
| min | med | avg | max |
|---|---|---|---|
| Percentile Index | 10 | 20 | 30 | 40 | 50 | 60 | 70 | 80 | 90 | 100 |
|---|---|---|---|---|---|---|---|---|---|---|
| Value |
| min | med | avg | max |
|---|---|---|---|
| Percentile Index | 10 | 20 | 30 | 40 | 50 | 60 | 70 | 80 | 90 | 100 |
|---|---|---|---|---|---|---|---|---|---|---|
| Value |
| Path / |
The code analyzed by CQA in that panel excludes loops and represents 0.01% of application time for run 1x1
| Source file and lines | hypre_qsort.c:47-53 |
| Module | exec |
| nb instructions | 12 |
| loop length | 48 |
| nb stack references | 0 |
| front end | 1.13 cycles |
| P0 | P1 | P2 | P3 | P4 | P5 | P6 | P7 | P8 | P9 | P10 | P11 | P12 | P13 | P14 | P15 | P16 | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| uops | 0.50 | 0.50 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 1.00 | 1.00 | 0.00 | 0.00 | 2.67 | 2.67 | 2.67 | 1.00 | 1.00 |
| cycles | 0.50 | 0.50 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 1.00 | 1.00 | 0.00 | 0.00 | 2.67 | 2.67 | 2.67 | 1.00 | 1.00 |
| Cycles executing div or sqrt instructions | NA |
| Front-end | 1.13 |
| Overall L1 | 2.67 |
| all | 0% |
| load | 0% |
| store | 0% |
| mul | NA (no mul vectorizable/vectorized instructions) |
| add-sub | NA (no add-sub vectorizable/vectorized instructions) |
| fma | NA (no fma vectorizable/vectorized instructions) |
| div/sqrt | NA (no div/sqrt vectorizable/vectorized instructions) |
| other | NA (no other vectorizable/vectorized instructions) |
| Instruction | Nb FU | P0 | P1 | P2 | P3 | P4 | P5 | P6 | P7 | P8 | P9 | P10 | P11 | P12 | P13 | P14 | P15 | P16 | Latency | Recip. throughput | Vectorization |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| LDR D0, [X1, X2,LSL #3] | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.33 | 0.33 | 0.33 | 0 | 0 | 6 | 0.33 | scal (50.0%) |
| LDR D1, [X1, X3,LSL #3] | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.33 | 0.33 | 0.33 | 0 | 0 | 6 | 0.33 | scal (50.0%) |
| LDR X8, [X0, X2,LSL #3] | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.33 | 0.33 | 0.33 | 0 | 0 | 4 | 0.33 | scal (50.0%) |
| LDR X9, [X0, X3,LSL #3] | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.33 | 0.33 | 0.33 | 0 | 0 | 4 | 0.33 | scal (50.0%) |
| STR X9, [X0, X2,LSL #3] | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.50 | 0.50 | 0 | 0.50 | 0.50 | 1 | 0.50 | scal (50.0%) |
| STR X8, [X0, X3,LSL #3] | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.50 | 0.50 | 0 | 0.50 | 0.50 | 1 | 0.50 | scal (50.0%) |
| STR D1, [X1, X2,LSL #3] | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.50 | 0.50 | 0 | 0 | 0.50 | 0.50 | 0 | 0 | 0 | 2 | 0.50 | scal (50.0%) |
| STR D0, [X1, X3,LSL #3] | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.50 | 0.50 | 0 | 0 | 0.50 | 0.50 | 0 | 0 | 0 | 2 | 0.50 | scal (50.0%) |
| RET | 1 | 0.50 | 0.50 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0.50 | N/A |
| HINT #0 | N/A | ||||||||||||||||||||
| HINT #0 | N/A | ||||||||||||||||||||
| HINT #0 | N/A |
The code analyzed by CQA in that panel excludes loops and represents 0.01% of application time for run 1x1
| Source file and lines | hypre_qsort.c:47-53 |
| Module | exec |
| nb instructions | 12 |
| loop length | 48 |
| nb stack references | 0 |
| front end | 1.13 cycles |
| P0 | P1 | P2 | P3 | P4 | P5 | P6 | P7 | P8 | P9 | P10 | P11 | P12 | P13 | P14 | P15 | P16 | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| uops | 0.50 | 0.50 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 1.00 | 1.00 | 0.00 | 0.00 | 2.67 | 2.67 | 2.67 | 1.00 | 1.00 |
| cycles | 0.50 | 0.50 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 1.00 | 1.00 | 0.00 | 0.00 | 2.67 | 2.67 | 2.67 | 1.00 | 1.00 |
| Cycles executing div or sqrt instructions | NA |
| Front-end | 1.13 |
| Overall L1 | 2.67 |
| all | 0% |
| load | 0% |
| store | 0% |
| mul | NA (no mul vectorizable/vectorized instructions) |
| add-sub | NA (no add-sub vectorizable/vectorized instructions) |
| fma | NA (no fma vectorizable/vectorized instructions) |
| div/sqrt | NA (no div/sqrt vectorizable/vectorized instructions) |
| other | NA (no other vectorizable/vectorized instructions) |
| Instruction | Nb FU | P0 | P1 | P2 | P3 | P4 | P5 | P6 | P7 | P8 | P9 | P10 | P11 | P12 | P13 | P14 | P15 | P16 | Latency | Recip. throughput | Vectorization |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| LDR D0, [X1, X2,LSL #3] | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.33 | 0.33 | 0.33 | 0 | 0 | 6 | 0.33 | scal (50.0%) |
| LDR D1, [X1, X3,LSL #3] | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.33 | 0.33 | 0.33 | 0 | 0 | 6 | 0.33 | scal (50.0%) |
| LDR X8, [X0, X2,LSL #3] | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.33 | 0.33 | 0.33 | 0 | 0 | 4 | 0.33 | scal (50.0%) |
| LDR X9, [X0, X3,LSL #3] | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.33 | 0.33 | 0.33 | 0 | 0 | 4 | 0.33 | scal (50.0%) |
| STR X9, [X0, X2,LSL #3] | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.50 | 0.50 | 0 | 0.50 | 0.50 | 1 | 0.50 | scal (50.0%) |
| STR X8, [X0, X3,LSL #3] | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.50 | 0.50 | 0 | 0.50 | 0.50 | 1 | 0.50 | scal (50.0%) |
| STR D1, [X1, X2,LSL #3] | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.50 | 0.50 | 0 | 0 | 0.50 | 0.50 | 0 | 0 | 0 | 2 | 0.50 | scal (50.0%) |
| STR D0, [X1, X3,LSL #3] | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.50 | 0.50 | 0 | 0 | 0.50 | 0.50 | 0 | 0 | 0 | 2 | 0.50 | scal (50.0%) |
| RET | 1 | 0.50 | 0.50 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0.50 | N/A |
| HINT #0 | N/A | ||||||||||||||||||||
| HINT #0 | N/A | ||||||||||||||||||||
| HINT #0 | N/A |
| Run 1x1 | Number processes: 1Number nodes: 1Run Command: <executable> -n 400 400 400MPI Command: mpirun -n <number_processes> --bind-to core --map-by package:PE=96 --rank-by fill --report-bindings Dataset: Run Directory: /home/hbollore/qaas/qaas-runs/174-161-6712/intel/AMG/run/oneview_runs/multicore/armclang_5/oneview_run_1741633338OMP_PROC_BIND: spreadOMP_DISPLAY_AFFINITY: TRUEOMP_AFFINITY_FORMAT: 'OMP: pid %P tid %i thread %n bound to OS proc set {%A}'OMP_DISPLAY_ENV: TRUEOMP_NUM_THREADS: 1OMP_PLACES: threads |
|---|---|
| Run 1x2 | Number processes: 1Number nodes: 1Run Command: <executable> -n 400 400 400MPI Command: mpirun -n <number_processes> --bind-to core --map-by package:PE=96 --rank-by fill --report-bindings Dataset: Run Directory: /home/hbollore/qaas/qaas-runs/174-161-6712/intel/AMG/run/oneview_runs/multicore/armclang_5/oneview_run_1741633338OMP_NUM_THREADS: 2OMP_PROC_BIND: spreadOMP_DISPLAY_AFFINITY: TRUEOMP_AFFINITY_FORMAT: 'OMP: pid %P tid %i thread %n bound to OS proc set {%A}'OMP_DISPLAY_ENV: TRUEOMP_PLACES: threads |
| Run 1x4 | Number processes: 1Number nodes: 1Run Command: <executable> -n 400 400 400MPI Command: mpirun -n <number_processes> --bind-to core --map-by package:PE=96 --rank-by fill --report-bindings Dataset: Run Directory: /home/hbollore/qaas/qaas-runs/174-161-6712/intel/AMG/run/oneview_runs/multicore/armclang_5/oneview_run_1741633338OMP_NUM_THREADS: 4OMP_PROC_BIND: spreadOMP_DISPLAY_AFFINITY: TRUEOMP_AFFINITY_FORMAT: 'OMP: pid %P tid %i thread %n bound to OS proc set {%A}'OMP_DISPLAY_ENV: TRUEOMP_PLACES: threads |
| Run 1x8 | Number processes: 1Number nodes: 1Run Command: <executable> -n 400 400 400MPI Command: mpirun -n <number_processes> --bind-to core --map-by package:PE=96 --rank-by fill --report-bindings Dataset: Run Directory: /home/hbollore/qaas/qaas-runs/174-161-6712/intel/AMG/run/oneview_runs/multicore/armclang_5/oneview_run_1741633338OMP_NUM_THREADS: 8OMP_PROC_BIND: spreadOMP_DISPLAY_AFFINITY: TRUEOMP_AFFINITY_FORMAT: 'OMP: pid %P tid %i thread %n bound to OS proc set {%A}'OMP_DISPLAY_ENV: TRUEOMP_PLACES: threads |
| Run 1x16 | Number processes: 1Number nodes: 1Run Command: <executable> -n 400 400 400MPI Command: mpirun -n <number_processes> --bind-to core --map-by package:PE=96 --rank-by fill --report-bindings Dataset: Run Directory: /home/hbollore/qaas/qaas-runs/174-161-6712/intel/AMG/run/oneview_runs/multicore/armclang_5/oneview_run_1741633338OMP_NUM_THREADS: 16OMP_PROC_BIND: spreadOMP_DISPLAY_AFFINITY: TRUEOMP_AFFINITY_FORMAT: 'OMP: pid %P tid %i thread %n bound to OS proc set {%A}'OMP_DISPLAY_ENV: TRUEOMP_PLACES: threads |
| Run 1x24 | Number processes: 1Number nodes: 1Run Command: <executable> -n 400 400 400MPI Command: mpirun -n <number_processes> --bind-to core --map-by package:PE=96 --rank-by fill --report-bindings Dataset: Run Directory: /home/hbollore/qaas/qaas-runs/174-161-6712/intel/AMG/run/oneview_runs/multicore/armclang_5/oneview_run_1741633338OMP_NUM_THREADS: 24OMP_PROC_BIND: spreadOMP_DISPLAY_AFFINITY: TRUEOMP_AFFINITY_FORMAT: 'OMP: pid %P tid %i thread %n bound to OS proc set {%A}'OMP_DISPLAY_ENV: TRUEOMP_PLACES: threads |
| Run 1x32 | Number processes: 1Number nodes: 1Run Command: <executable> -n 400 400 400MPI Command: mpirun -n <number_processes> --bind-to core --map-by package:PE=96 --rank-by fill --report-bindings Dataset: Run Directory: /home/hbollore/qaas/qaas-runs/174-161-6712/intel/AMG/run/oneview_runs/multicore/armclang_5/oneview_run_1741633338OMP_NUM_THREADS: 32OMP_PROC_BIND: spreadOMP_DISPLAY_AFFINITY: TRUEOMP_AFFINITY_FORMAT: 'OMP: pid %P tid %i thread %n bound to OS proc set {%A}'OMP_DISPLAY_ENV: TRUEOMP_PLACES: threads |
| Run 1x40 | Number processes: 1Number nodes: 1Run Command: <executable> -n 400 400 400MPI Command: mpirun -n <number_processes> --bind-to core --map-by package:PE=96 --rank-by fill --report-bindings Dataset: Run Directory: /home/hbollore/qaas/qaas-runs/174-161-6712/intel/AMG/run/oneview_runs/multicore/armclang_5/oneview_run_1741633338OMP_NUM_THREADS: 40OMP_PROC_BIND: spreadOMP_DISPLAY_AFFINITY: TRUEOMP_AFFINITY_FORMAT: 'OMP: pid %P tid %i thread %n bound to OS proc set {%A}'OMP_DISPLAY_ENV: TRUEOMP_PLACES: threads |
| Run 1x48 | Number processes: 1Number nodes: 1Run Command: <executable> -n 400 400 400MPI Command: mpirun -n <number_processes> --bind-to core --map-by package:PE=96 --rank-by fill --report-bindings Dataset: Run Directory: /home/hbollore/qaas/qaas-runs/174-161-6712/intel/AMG/run/oneview_runs/multicore/armclang_5/oneview_run_1741633338OMP_NUM_THREADS: 48OMP_PROC_BIND: spreadOMP_DISPLAY_AFFINITY: TRUEOMP_AFFINITY_FORMAT: 'OMP: pid %P tid %i thread %n bound to OS proc set {%A}'OMP_DISPLAY_ENV: TRUEOMP_PLACES: threads |
| Run 1x56 | Number processes: 1Number nodes: 1Run Command: <executable> -n 400 400 400MPI Command: mpirun -n <number_processes> --bind-to core --map-by package:PE=96 --rank-by fill --report-bindings Dataset: Run Directory: /home/hbollore/qaas/qaas-runs/174-161-6712/intel/AMG/run/oneview_runs/multicore/armclang_5/oneview_run_1741633338OMP_NUM_THREADS: 56OMP_PROC_BIND: spreadOMP_DISPLAY_AFFINITY: TRUEOMP_AFFINITY_FORMAT: 'OMP: pid %P tid %i thread %n bound to OS proc set {%A}'OMP_DISPLAY_ENV: TRUEOMP_PLACES: threads |
| Run 1x64 | Number processes: 1Number nodes: 1Run Command: <executable> -n 400 400 400MPI Command: mpirun -n <number_processes> --bind-to core --map-by package:PE=96 --rank-by fill --report-bindings Dataset: Run Directory: /home/hbollore/qaas/qaas-runs/174-161-6712/intel/AMG/run/oneview_runs/multicore/armclang_5/oneview_run_1741633338OMP_NUM_THREADS: 64OMP_PROC_BIND: spreadOMP_DISPLAY_AFFINITY: TRUEOMP_AFFINITY_FORMAT: 'OMP: pid %P tid %i thread %n bound to OS proc set {%A}'OMP_DISPLAY_ENV: TRUEOMP_PLACES: threads |
| Run 1x72 | Number processes: 1Number nodes: 1Run Command: <executable> -n 400 400 400MPI Command: mpirun -n <number_processes> --bind-to core --map-by package:PE=96 --rank-by fill --report-bindings Dataset: Run Directory: /home/hbollore/qaas/qaas-runs/174-161-6712/intel/AMG/run/oneview_runs/multicore/armclang_5/oneview_run_1741633338OMP_NUM_THREADS: 72OMP_PROC_BIND: spreadOMP_DISPLAY_AFFINITY: TRUEOMP_AFFINITY_FORMAT: 'OMP: pid %P tid %i thread %n bound to OS proc set {%A}'OMP_DISPLAY_ENV: TRUEOMP_PLACES: threads |
| Run 1x80 | Number processes: 1Number nodes: 1Run Command: <executable> -n 400 400 400MPI Command: mpirun -n <number_processes> --bind-to core --map-by package:PE=96 --rank-by fill --report-bindings Dataset: Run Directory: /home/hbollore/qaas/qaas-runs/174-161-6712/intel/AMG/run/oneview_runs/multicore/armclang_5/oneview_run_1741633338OMP_NUM_THREADS: 80OMP_PROC_BIND: spreadOMP_DISPLAY_AFFINITY: TRUEOMP_AFFINITY_FORMAT: 'OMP: pid %P tid %i thread %n bound to OS proc set {%A}'OMP_DISPLAY_ENV: TRUEOMP_PLACES: threads |
| Run 1x88 | Number processes: 1Number nodes: 1Run Command: <executable> -n 400 400 400MPI Command: mpirun -n <number_processes> --bind-to core --map-by package:PE=96 --rank-by fill --report-bindings Dataset: Run Directory: /home/hbollore/qaas/qaas-runs/174-161-6712/intel/AMG/run/oneview_runs/multicore/armclang_5/oneview_run_1741633338OMP_NUM_THREADS: 88OMP_PROC_BIND: spreadOMP_DISPLAY_AFFINITY: TRUEOMP_AFFINITY_FORMAT: 'OMP: pid %P tid %i thread %n bound to OS proc set {%A}'OMP_DISPLAY_ENV: TRUEOMP_PLACES: threads |
| Run 1x96 | Number processes: 1Number nodes: 1Run Command: <executable> -n 400 400 400MPI Command: mpirun -n <number_processes> --bind-to core --map-by package:PE=96 --rank-by fill --report-bindings Dataset: Run Directory: /home/hbollore/qaas/qaas-runs/174-161-6712/intel/AMG/run/oneview_runs/multicore/armclang_5/oneview_run_1741633338OMP_NUM_THREADS: 96OMP_PROC_BIND: spreadOMP_DISPLAY_AFFINITY: TRUEOMP_AFFINITY_FORMAT: 'OMP: pid %P tid %i thread %n bound to OS proc set {%A}'OMP_DISPLAY_ENV: TRUEOMP_PLACES: threads |
| (1x1) Efficiency | (1x1) Potential Speed-Up (%) | (1x2) Efficiency | (1x2) Potential Speed-Up (%) | (1x4) Efficiency | (1x4) Potential Speed-Up (%) | (1x8) Efficiency | (1x8) Potential Speed-Up (%) | (1x16) Efficiency | (1x16) Potential Speed-Up (%) | (1x24) Efficiency | (1x24) Potential Speed-Up (%) | (1x32) Efficiency | (1x32) Potential Speed-Up (%) | (1x40) Efficiency | (1x40) Potential Speed-Up (%) | (1x48) Efficiency | (1x48) Potential Speed-Up (%) | (1x56) Efficiency | (1x56) Potential Speed-Up (%) | (1x64) Efficiency | (1x64) Potential Speed-Up (%) | (1x72) Efficiency | (1x72) Potential Speed-Up (%) | (1x80) Efficiency | (1x80) Potential Speed-Up (%) | (1x88) Efficiency | (1x88) Potential Speed-Up (%) | (1x96) Efficiency | (1x96) Potential Speed-Up (%) |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | 0 | 2.31 | 0 | 3.41 | 0 | 10.59 | 0 | 17.1 | 0 | 24.89 | 0 | 40.67 | 0 | 38.39 | 0 | 40.86 | 0 | 48.84 | 0 | 42.44 | 0 | 84.32 | 0 | 247.08 | 0 | 70.36 | 0 | 133.39 | 0 |
| Run | Number of threads | Efficiency (ideal is 1) | Speedup | Ideal Speedup | Time (s) | Coverage (%) |
|---|---|---|---|---|---|---|
| 1x1 | 1 | 1 | 1 | 1 | 0.034999996423721 | 0.013156464323401 |
| 1x2 | 2 | 2.31 | 2.31 | 2 | 0.014999999664724 | 0.010695016011596 |
| 1x4 | 4 | 3.41 | 3.41 | 4 | 0.014999998733401 | 0.013586993329227 |
| 1x8 | 4 | 10.59 | 10.59 | 8 | 0.010000000707805 | 0.0074938493780792 |
| 1x16 | 5 | 17.1 | 17.1 | 16 | 0.00999999884516 | 0.0071053225547075 |
| 1x24 | 5 | 24.89 | 24.89 | 24 | 0.0099999979138374 | 0.0056777540594339 |
| 1x32 | 5 | 40.67 | 40.67 | 32 | 0.0049999998882413 | 0.0037913944106549 |
| 1x40 | 6 | 38.39 | 38.39 | 40 | 0.0049999998882413 | 0.0037785621825606 |
| 1x48 | 7 | 40.86 | 40.86 | 48 | 0.00499999942258 | 0.0038649614434689 |
| 1x56 | 7 | 48.84 | 48.84 | 56 | 0.0050000003539026 | 0.003557363525033 |
| 1x64 | 9 | 42.44 | 42.44 | 64 | 0.0050000003539026 | 0.0042346520349383 |
| 1x72 | 5 | 84.32 | 84.32 | 72 | 0.0049999998882413 | 0.002134133130312 |
| 1x80 | 2 | 247.08 | 247.08 | 80 | 0.0049999998882413 | 0.00079137639841065 |
| 1x88 | 7 | 70.36 | 70.36 | 88 | 0.0050000003539026 | 0.0025358656421304 |
| 1x96 | 4 | 133.39 | 133.39 | 96 | 0.0049999998882413 | 0.0013431708794087 |
| Name | Coverage (%) | Time (s) |
|---|---|---|
| ○hypre_swap2 | 0.01 | 0.03 |
