	*************************************************
	*                                               *
	*          ONE-View report generation           *
	*                                               *
	*************************************************

[MAQAO] Info: Experiment configuration summary is available adding -dbg=1 in command line

* [MAQAO] Warning: Experiment directory /beegfs/hackathon/users/eoseret/qaas_runs_test/178-668-9639/intel/AMG/run/oneview_runs/defaults/gcc/oneview_results_1786690412 already exists and is reused.
           It can be replaced using --replace in the command line.
[MAQAO] Info: 
[MAQAO] Info: START THE APPLICATION PROFILING
[MAQAO] Info: -> RUNNING THE PROFILER...
[MAQAO] Info:   LPROF has already been run
[MAQAO] Info: STOP THE APPLICATION PROFILING
[MAQAO] Info: 
[MAQAO] Info: START FUNCTIONS AND LOOPS ANALYSIS ...
[MAQAO] Info: STOP FUNCTIONS AND LOOPS ANALYSIS ...
[MAQAO] Info: 
[MAQAO] Info: START THE REPORT GENERATION
[MAQAO] Info: -> ONE-VIEW EXPERIMENT DIRECTORY: /beegfs/hackathon/users/eoseret/qaas_runs_test/178-668-9639/intel/AMG/run/oneview_runs/defaults/gcc/oneview_results_1786690412


+====================================================================================================================+
+                                                    1  -  GLOBAL                                                    +
+====================================================================================================================+


+--------------------------------------------------------------------------------------------------------------------+
+                                             1.1  -  Experiment Summary                                             +
+--------------------------------------------------------------------------------------------------------------------+

  Application:			/beegfs/hackathon/users/eoseret/qaas_runs_test/178-668-9639/intel/AMG/run/base_runs/defaults/gcc/exec
  Timestamp:			2026-08-14 08:53:32
  Universal Timestamp:		1786690412
  Experiment Type:		MPI; OpenMP; Throughput; 
  Machine:			isix07.benchmarkcenter.megware.com
  Architecture:			x86_64
  Micro Architecture:		GRANITE_RAPIDS
  Model Name:			Intel(R) Xeon(R) 6972P
  Cache Size:			491520 KB
  Number of Cores:		96
  OS Version:			Linux 5.14.0-687.31.1.el9_8.x86_64 #1 SMP PREEMPT_DYNAMIC Sat Aug 1 05:38:01 EDT 2026
  Compilation Options:		
		exec: GNU C89 15.1.0 -march=graniterapids -mmmx -mpopcnt -msse -msse2 -msse3 -mssse3 -msse4.1 -msse4.2 -mavx -mavx2 -mno-sse4a -mno-fma4 -mno-xop -mfma -mavx512f -mbmi -mbmi2 -maes -mpclmul -mavx512vl -mavx512bw -mavx512dq -mavx512cd -mavx512vbmi -mavx512ifma -mavx512vpopcntdq -mavx512vbmi2 -mgfni -mvpclmulqdq -mavx512vnni -mavx512bitalg -mavx512bf16 -mno-avx512vp2intersect -mno-3dnow -madx -mabm -mcldemote -mclflushopt -mclwb -mno-clzero -mcx16 -menqcmd -mf16c -mfsgsbase -mfxsr -mno-hle -msahf -mno-lwp -mlzcnt -mmovbe -mmovdir64b -mmovdiri -mno-mwaitx -mpconfig -mpku -mprfchw -mptwrite -mrdpid -mrdrnd -mrdseed -mno-rtm -mserialize -msgx -msha -mshstk -mno-tbm -mtsxldtrk -mvaes -mwaitpkg -mwbnoinvd -mxsave -mxsavec -mxsaveopt -mxsaves -mamx-tile -mamx-int8 -mamx-bf16 -muintr -mhreset -mno-kl -mno-widekl -mavxvnni -mavx512fp16 -mno-avxifma -mno-avxvnniint8 -mno-avxneconvert -mno-cmpccxadd -mamx-fp16 -mprefetchi -mno-raoint -mno-amx-complex -mno-avxvnniint16 -mno-sm3 -mno-sha512 -mno-sm4 -mno-apxf -mno-usermsr -mavx10.1-256 -mavx10.1-512 -mno-avx10.2 -mno-amx-avx512 -mno-amx-tf32 -mno-amx-transpose -mno-amx-fp8 -mno-movrs -mno-amx-movrs --param=l1-cache-size=48 --param=l1-cache-line-size=64 --param=l2-cache-size=491520 -mtune=graniterapids -g -O3 -std=gnu90 -fno-omit-frame-pointer -fcf-protection=none -fopenmp -funroll-loops 
  Number of processes observed:	6
  Number of threads observed:	192
  MAQAO version:		2026.1.0
  MAQAO build:			6d1be1d51c1e63266254997eb301734a7264775d::20260810-150026




+--------------------------------------------------------------------------------------------------------------------+
+                                               1.2  -  Global Metrics                                               +
+--------------------------------------------------------------------------------------------------------------------+

  Total Time:				48.16 s
  Max (Thread Active Time):		47.31 s
  Average Active Time:			39.81 s
  Activity Ratio:			83.4 %
  Average number of active threads:	158.718
  Affinity Stability:			99.6 %
  Time spent in analyzed loops:		93.6 %
  Time spent in analyzed innermost loops: 54.7 %
  Time spent in user code:		93.7 %
  Compilation Options Score:		100
  Array Access Efficiency:		70.0 %

   Potential Speedups
  ----------------------------------------------------
  Perfect Flow Complexity:		1.03
  Perfect OpenMP/MPI/Pthread/TBB:	1.03
  Perfect OpenMP/MPI/Pthread/TBB + Load Distribution:	1.22
  If No Scalar Integer:
      Potential Speedup:		1.10
      Nb Loops to get 80%:		8
  If FP Vectorized:
      Potential Speedup:		1.55
      Nb Loops to get 80%:		6
  If Fully Vectorized:
      Potential Speedup:		3.57
      Nb Loops to get 80%:		30
  If Only FP Arithmetic:
      Potential Speedup:		1.34
      Nb Loops to get 80%:		16




+--------------------------------------------------------------------------------------------------------------------+
+                                             1.3  -  Potential Speedups                                             +
+--------------------------------------------------------------------------------------------------------------------+

  If No Scalar Integer:
      Number of loops   | 1      | 10     | 20     | 29     | 40     | 
      Cumulated Speedup | 1.0434 | 1.0904 | 1.1029 | 1.1047 | 1.1048 | 
  Top 5 loops:
    exec - 2138:	1.0434
    exec - 3132:	1.0517
    exec - 3141:	1.0592
    exec - 3149:	1.0651
    exec - 2857:	1.0707

  If FP Vectorized:
      Number of loops   | 1      | 10     | 20     | 29     | 40     | 
      Cumulated Speedup | 1.1255 | 1.5106 | 1.5430 | 1.5463 | 1.5463 | 
  Top 5 loops:
    exec - 2140:	1.1255
    exec - 2138:	1.2767
    exec - 3133:	1.3286
    exec - 3142:	1.3789
    exec - 3132:	1.422

  If Fully Vectorized:
      Number of loops   | 1      | 10     | 20     | 29     | 40     | 
      Cumulated Speedup | 1.2237 | 2.2402 | 2.6681 | 3.0370 | 3.2932 | 
  Top 5 loops:
    exec - 2140:	1.2237
    exec - 2138:	1.5325
    exec - 3133:	1.6603
    exec - 3142:	1.7943
    exec - 3132:	1.9183

  If Only FP Arithmetic:
      Number of loops   | 1      | 10     | 20     | 29     | 40     | 
      Cumulated Speedup | 1.1176 | 1.2426 | 1.2938 | 1.3150 | 1.3269 | 
  Top 5 loops:
    exec - 2138:	1.1176
    exec - 3132:	1.1472
    exec - 3141:	1.1748
    exec - 2137:	1.1908
    exec - 3128:	1.2035



+====================================================================================================================+
+                                                   2  -  SUMMARY                                                    +
+====================================================================================================================+


+--------------------------------------------------------------------------------------------------------------------+
+                                             2.1  -  EXPERIMENT QUALITY                                             +
+--------------------------------------------------------------------------------------------------------------------+

  [4 / 4] Application profile is long enough (47.31 s)
To have good quality measurements, it is advised that the application profiling time is greater than 10 seconds.

  [3 / 3] Most of time spent in analyzed modules comes from functions with source/debug info
-g option gives access to debugging informations, such are source locations.

  [3 / 3] Most of time spent in analyzed modules comes from functions with compilation options informations and
-fno-omit-frame-pointer is present
-fno-omit-frame-pointer improves the accuracy of callchains found during the application profiling.

  [3 / 3] Host configuration allows retrieval of all necessary metrics.


  [2.9967301212459 / 3] Most of time spent in analyzed modules (99.89%) comes from functions compiled with architecture specialization option
-march=graniterapids


  [3 / 3] Optimization level option is correctly used


  [2 / 2] Application is correctly profiled ("Others" category represents 0.00 % of the execution time)
To have a representative profiling, it is advised that the category "Others" represents less than 20% of the execution
time in order to analyze as much as possible of the user code

  [1 / 1] Lstopo present. The Topology lstopo report will be generated.


  [0 / 0] Fastmath not used
Consider to add ffast-math to compilation flags (or replace -O3 with -Ofast) to unlock potential extra speedup by
relaxing floating-point computation consistency. Warning: floating-point accuracy may be reduced and the compliance
to IEEE/ISO rules/specifications for math functions will be relaxed, typically 'errno' will no longer be set after
calling some math functions.


+--------------------------------------------------------------------------------------------------------------------+
+                                                2.2  -  CODE QUALITY                                                +
+--------------------------------------------------------------------------------------------------------------------+

  [4 / 4] Enough time of the experiment time spent in analyzed loops (93.65%)
If the time spent in analyzed loops is less than 30%, standard loop optimizations will have a limited impact on
application performances.

  [3 / 4] A significant amount of threads are idle (17.33%)
On average, more than 10% of observed threads are idle. Such threads are probably IO/sync waiting. Some hints: use
faster filesystems to read/write data, improve parallel load balancing and/or scheduling.

  [3 / 4] CPU activity is below 90% (83.37%)
CPU cores are idle more than 10% of time. Threads supposed to run on these cores are probably IO/sync waiting. Some
hints: use faster filesystems to read/write data, improve parallel load balancing and/or scheduling.

  [4 / 4] Loop profile is not flat
At least one loop coverage is greater than 4% (24.81%), representing an hotspot for the application

  [4 / 4] Enough time of the experiment time spent in analyzed innermost loops (54.71%)
If the time spent in analyzed innermost loops is less than 15%, standard innermost loop optimizations such as
vectorisation will have a limited impact on application performances.

  [4 / 4] Affinity is good (99.56%)
Threads are not migrating to CPU cores: probably successfully pinned

  [3 / 3] Less than 10% (0.00%) is spend in BLAS1 operations
It could be more efficient to inline by hand BLAS1 operations

  [0 / 3] Too many functions do not use all threads
Functions running on a reduced number of threads (typically sequential code) cover at least 10% of application
walltime (16.72%). Check both "Max Inclusive Time Over Threads" and "Nb Threads" in Functions or Loops tabs and
consider parallelizing sequential regions or improving parallelization of regions running on a reduced number of
threads

  [3 / 3] Cumulative Outermost/In between loops coverage (38.94%) lower than cumulative innermost loop coverage (54.71%)
Having cumulative Outermost/In between loops coverage greater than cumulative innermost loop coverage will make loop
optimization more complex

  [2 / 2] Less than 10% (0.00%) is spend in BLAS2 operations
BLAS2 calls usually could make a poor cache usage and could benefit from inlining.

  [2 / 2] Less than 10% (0.00%) is spend in Libm/SVML (special functions)



+--------------------------------------------------------------------------------------------------------------------+
+                                               2.3  -  LOOPS OVERVIEW                                               +
+--------------------------------------------------------------------------------------------------------------------+

  Top 5 loops:
   + exec - 2140:
     analysis: Execution Time: 24 % - Vectorization Ratio: 50.00 % - Vector Length Use: 28.75 %
     Loop Computation Issues: 4
        [4] [SA] Less than 10% of the FP ADD/SUB/MUL arithmetic operations are performed using FMA - Reorganize
            arithmetic expressions to exhibit potential for FMA. This issue costs 4 points.
     Data Access Issues: 32
        [4] [SA] Presence of indirect accesses - Use array restructuring or gather instructions to lower the cost.
            There are 1 issues ( = indirect data accesses) costing 4 point each.
        [16] [SA] Presence of expensive instructions (GATHER/SCATTER) - Use array restructuring. There are 4 issues (=
            instructions) costing 4 points each.
        [12] [SA] Presence of special instructions executing on a single port (INSERT/EXTRACT, BLEND/MERGE,
            Other_packing) - Simplify data access and try to get stride 1 access. There are 12 issues (=
            instructions) costing 1 point each.
     Vectorization Roadblocks: 4
        [4] [SA] Presence of indirect accesses - Use array restructuring or gather instructions to lower the cost.
            There are 1 issues ( = indirect data accesses) costing 4 point each.
     Inefficient Vectorization: 28
        [16] [SA] Presence of expensive instructions (GATHER/SCATTER) - Use array restructuring. There are 4 issues (=
            instructions) costing 4 points each.
        [12] [SA] Presence of special instructions executing on a single port (INSERT/EXTRACT, BLEND/MERGE,
            Other_packing) - Simplify data access and try to get stride 1 access. There are 12 issues (=
            instructions) costing 1 point each.

   + exec - 2138:
     analysis: Execution Time: 20 % - Vectorization Ratio: 37.89 % - Vector Length Use: 23.55 %
     Loop Computation Issues: 6
        [4] [SA] Presence of expensive FP instructions - Perform hoisting, change algorithm, use SVML or proper
            numerical library or perform value profiling (count the number of distinct input values). There are 1
            issues (= instructions) costing 4 points each.
        [2] [SA] Presence of a large number of scalar integer instructions - Simplify loop structure, perform loop
            splitting or perform unroll and jam. This issue costs 2 points.
     Control Flow Issues: 1002
        [1000] [SA] Too many paths (2049 paths) - Simplify control structure. There are 2049 issues ( = paths) costing 1
            point, limited to 1000.
        [2] [SA] Non innermost loop (Outermost) - Collapse loop with innermost ones. This issue costs 2 points.
     Data Access Issues: 54
        [32] [SA] Presence of expensive instructions (GATHER/SCATTER) - Use array restructuring. There are 8 issues (=
            instructions) costing 4 points each.
        [20] [SA] Presence of special instructions executing on a single port (INSERT/EXTRACT, BLEND/MERGE,
            Other_packing) - Simplify data access and try to get stride 1 access. There are 20 issues (=
            instructions) costing 1 point each.
        [2] [SA] More than 20% of the loads are accessing the stack - Perform loop splitting to decrease pressure on
            registers. This issue costs 2 points.
     Vectorization Roadblocks: 1002
        [1000] [SA] Too many paths (2049 paths) - Simplify control structure. There are 2049 issues ( = paths) costing 1
            point, limited to 1000.
        [2] [SA] Non innermost loop (Outermost) - Collapse loop with innermost ones. This issue costs 2 points.
     Inefficient Vectorization: 52
        [32] [SA] Presence of expensive instructions (GATHER/SCATTER) - Use array restructuring. There are 8 issues (=
            instructions) costing 4 points each.
        [20] [SA] Presence of special instructions executing on a single port (INSERT/EXTRACT, BLEND/MERGE,
            Other_packing) - Simplify data access and try to get stride 1 access. There are 20 issues (=
            instructions) costing 1 point each.

   + exec - 3133:
     analysis: Execution Time: 6 % - Vectorization Ratio: 50.00 % - Vector Length Use: 28.75 %
     Loop Computation Issues: 4
        [4] [SA] Less than 10% of the FP ADD/SUB/MUL arithmetic operations are performed using FMA - Reorganize
            arithmetic expressions to exhibit potential for FMA. This issue costs 4 points.
     Data Access Issues: 32
        [4] [SA] Presence of indirect accesses - Use array restructuring or gather instructions to lower the cost.
            There are 1 issues ( = indirect data accesses) costing 4 point each.
        [16] [SA] Presence of expensive instructions (GATHER/SCATTER) - Use array restructuring. There are 4 issues (=
            instructions) costing 4 points each.
        [12] [SA] Presence of special instructions executing on a single port (INSERT/EXTRACT, BLEND/MERGE,
            Other_packing) - Simplify data access and try to get stride 1 access. There are 12 issues (=
            instructions) costing 1 point each.
     Vectorization Roadblocks: 4
        [4] [SA] Presence of indirect accesses - Use array restructuring or gather instructions to lower the cost.
            There are 1 issues ( = indirect data accesses) costing 4 point each.
     Inefficient Vectorization: 28
        [16] [SA] Presence of expensive instructions (GATHER/SCATTER) - Use array restructuring. There are 4 issues (=
            instructions) costing 4 points each.
        [12] [SA] Presence of special instructions executing on a single port (INSERT/EXTRACT, BLEND/MERGE,
            Other_packing) - Simplify data access and try to get stride 1 access. There are 12 issues (=
            instructions) costing 1 point each.

   + exec - 3142:
     analysis: Execution Time: 6 % - Vectorization Ratio: 50.00 % - Vector Length Use: 28.75 %
     Loop Computation Issues: 4
        [4] [SA] Less than 10% of the FP ADD/SUB/MUL arithmetic operations are performed using FMA - Reorganize
            arithmetic expressions to exhibit potential for FMA. This issue costs 4 points.
     Data Access Issues: 32
        [4] [SA] Presence of indirect accesses - Use array restructuring or gather instructions to lower the cost.
            There are 1 issues ( = indirect data accesses) costing 4 point each.
        [16] [SA] Presence of expensive instructions (GATHER/SCATTER) - Use array restructuring. There are 4 issues (=
            instructions) costing 4 points each.
        [12] [SA] Presence of special instructions executing on a single port (INSERT/EXTRACT, BLEND/MERGE,
            Other_packing) - Simplify data access and try to get stride 1 access. There are 12 issues (=
            instructions) costing 1 point each.
     Vectorization Roadblocks: 4
        [4] [SA] Presence of indirect accesses - Use array restructuring or gather instructions to lower the cost.
            There are 1 issues ( = indirect data accesses) costing 4 point each.
     Inefficient Vectorization: 28
        [16] [SA] Presence of expensive instructions (GATHER/SCATTER) - Use array restructuring. There are 4 issues (=
            instructions) costing 4 points each.
        [12] [SA] Presence of special instructions executing on a single port (INSERT/EXTRACT, BLEND/MERGE,
            Other_packing) - Simplify data access and try to get stride 1 access. There are 12 issues (=
            instructions) costing 1 point each.

   + exec - 3132:
     analysis: Execution Time: 4 % - Vectorization Ratio: 36.73 % - Vector Length Use: 23.09 %
     Loop Computation Issues: 2
        [2] [SA] Presence of a large number of scalar integer instructions - Simplify loop structure, perform loop
            splitting or perform unroll and jam. This issue costs 2 points.
     Control Flow Issues: 38
        [36] [SA] Too many paths (32 paths) - Simplify control structure. There are 32 issues ( = paths) costing 1
            point each with a malus of 4 points.
        [2] [SA] Non innermost loop (Outermost) - Collapse loop with innermost ones. This issue costs 2 points.
     Data Access Issues: 26
        [16] [SA] Presence of expensive instructions (GATHER/SCATTER) - Use array restructuring. There are 4 issues (=
            instructions) costing 4 points each.
        [10] [SA] Presence of special instructions executing on a single port (INSERT/EXTRACT, BLEND/MERGE,
            Other_packing) - Simplify data access and try to get stride 1 access. There are 10 issues (=
            instructions) costing 1 point each.
     Vectorization Roadblocks: 38
        [36] [SA] Too many paths (32 paths) - Simplify control structure. There are 32 issues ( = paths) costing 1
            point each with a malus of 4 points.
        [2] [SA] Non innermost loop (Outermost) - Collapse loop with innermost ones. This issue costs 2 points.
     Inefficient Vectorization: 26
        [16] [SA] Presence of expensive instructions (GATHER/SCATTER) - Use array restructuring. There are 4 issues (=
            instructions) costing 4 points each.
        [10] [SA] Presence of special instructions executing on a single port (INSERT/EXTRACT, BLEND/MERGE,
            Other_packing) - Simplify data access and try to get stride 1 access. There are 10 issues (=
            instructions) costing 1 point each.



+====================================================================================================================+
+                                                 3  -  APPLICATION                                                  +
+====================================================================================================================+


+--------------------------------------------------------------------------------------------------------------------+
+                                               3.1  -  Categorization                                               +
+--------------------------------------------------------------------------------------------------------------------+

   Category | IO     | Exe    | System  | Others  | Memory | String | MPI   | TBB   | OMP   | Pthread | Math  |
  ----------+--------+--------+---------+---------+--------+--------+-------+-------+-------+---------+-------+
   Time (%) | 0.00   | 93.68  | 3.15    | 0.00    | 0.00   | 0.18   | 0.04  | 0.00  | 2.95  | 0.00    | 0.00  |




+--------------------------------------------------------------------------------------------------------------------+
+                                          3.2  -  Function Based Profiling                                          +
+--------------------------------------------------------------------------------------------------------------------+

   Buckets                    | Nb Functions              | Coverage                  | Cumulated Coverage        |
  ----------------------------+---------------------------+---------------------------+---------------------------+
   > 8%                       | 2                         | 68.85                     | 68.85                     |
   4% to 8%                   | 0                         | 0.00                      | 68.85                     |
   2% to 4%                   | 2                         | 5.18                      | 74.02                     |
   1% to 2%                   | 12                        | 17.45                     | 91.47                     |
   0.5% to 1%                 | 6                         | 4.55                      | 96.02                     |
   0.25% to 0.5%              | 8                         | 2.83                      | 98.86                     |
   0.125% to 0.25%            | 3                         | 0.54                      | 99.40                     |
   < 0.125%                   | 122                       | 0.60                      | 100.00                    |




+--------------------------------------------------------------------------------------------------------------------+
+                                            3.3  -  Loop Based Profiling                                            +
+--------------------------------------------------------------------------------------------------------------------+

   Buckets                    | Nb Loops                  | Coverage                  | Cumulated Coverage        |
  ----------------------------+---------------------------+---------------------------+---------------------------+
   > 8%                       | 1                         | 24.81                     | 24.81                     |
   4% to 8%                   | 2                         | 12.92                     | 37.74                     |
   2% to 4%                   | 0                         | 0.00                      | 37.74                     |
   1% to 2%                   | 2                         | 3.22                      | 40.96                     |
   0.5% to 1%                 | 10                        | 7.36                      | 48.31                     |
   0.25% to 0.5%              | 10                        | 4.05                      | 52.36                     |
   0.125% to 0.25%            | 7                         | 1.14                      | 53.50                     |
   < 0.125%                   | 171                       | 1.21                      | 54.71                     |


+====================================================================================================================+
+                                                  4  -  FUNCTIONS                                                   +
+====================================================================================================================+


+--------------------------------------------------------------------------------------------------------------------+
+                                              4.1  -  Top 10 Functions                                              +
+--------------------------------------------------------------------------------------------------------------------+

   Function                                               | Module              | Coverage (%)   | Time (s)       |
  --------------------------------------------------------+---------------------+----------------+----------------+
   hypre_ParCSRRelaxThreads._omp_fn.1                     | exec                | 45.52          | 18.12          |
   hypre_CSRMatrixMatvecOutOfPlace._omp_fn.6              | exec                | 23.33          | 9.29           |
   unknown_kernel_region                                  | kernel              | 3.16           | 1.26           |
   hypre_BoomerAMGCreate2ndS._omp_fn.7                    | exec                | 2.02           | 0.80           |
   hypre_SeqVectorAxpy._omp_fn.0                          | exec                | 1.69           | 0.67           |
   gomp_barrier_wait_end                                  | libgomp.so.1.0.0    | 1.69           | 0.68           |
   hypre_BoomerAMGBuildMultipass._omp_fn.5                | exec                | 1.67           | 0.67           |
   hypre_BoomerAMGBuildMultipass._omp_fn.10               | exec                | 1.63           | 0.65           |
   hypre_BoomerAMGCreateS._omp_fn.1                       | exec                | 1.59           | 0.63           |
   hypre_ParCSRRelaxThreads._omp_fn.0                     | exec                | 1.53           | 0.61           |


+====================================================================================================================+
+                                                    5  -  LOOPS                                                     +
+====================================================================================================================+


+--------------------------------------------------------------------------------------------------------------------+
+                                                5.1  -  Top 10 Loops                                                +
+--------------------------------------------------------------------------------------------------------------------+

   Loop Id        | Module              | Source Location                                        | Coverage (%)   |
  ----------------+---------------------+--------------------------------------------------------+----------------+
   2140           | exec                | ams.c:3672-3675                                        | 24.81          |
   2138           | exec                | ams.c:3669-3682                                        | 20.70          |
   3133           | exec                | csr_matvec.c:310-312                                   | 6.82           |
   3142           | exec                | csr_matvec.c:259-261                                   | 6.11           |
   3132           | exec                | csr_matvec.c:307-314                                   | 4.59           |
   3141           | exec                | csr_matvec.c:256-263                                   | 4.06           |
   3128           | exec                | csr_matvec.c:334-341                                   | 1.76           |
   3184           | exec                | vector.c:452-452                                       | 1.69           |
   2137           | exec                | ams.c:3659-3659                                        | 1.52           |
   3149           | exec                | csr_matvec.c:564-564,csr_matvec.c:567-567              | 1.21           |





+====================================================================================================================+
+                                                     6  -  CQA                                                      +
+====================================================================================================================+


+--------------------------------------------------------------------------------------------------------------------+
+                                                   6.1  -  Loops                                                    +
+--------------------------------------------------------------------------------------------------------------------+





      6.1.1  -  Loop 2140 from exec
  ==========================================================================================================

The loop is defined in /beegfs/hackathon/users/eoseret/qaas_runs_test/178-668-9639/intel/AMG/build/AMG/AMG/parcsr_ls/ams.c:3672-3675.

The related source loop is not unrolled or unrolled with no peel/tail loop.

      6.1.1.1  -  Path 1
  ----------------------------------------------------------------------------------------------------------

1% of peak computational performance is used (0.35 out of 32.00 FLOP per cycle (GFLOPS @ 1GHz))

      6.1.1.1.1  -  Vectorization
  ----------------------------------------------------------------------------------------------------------

Your loop is partially vectorized.
Only 28% of vector register length is used (average across all SSE/AVX instructions).


Details
50% of SSE/AVX instructions are used in vector version (process two or more data elements in vector registers):
 - 0% of SSE/AVX addition or subtraction instructions are used in vector version.
 - 75% of SSE/AVX instructions that are not load, store, addition, subtraction nor multiply instructions are used in vector version.


Workaround
Read the "512-bits vectorization" report at "Potential" confidence level.


      6.1.1.1.2  -  Execution units bottlenecks
  ----------------------------------------------------------------------------------------------------------

Found no such bottlenecks but see expert reports for more complex bottlenecks.




      6.1.1.1.3  -  512-bits vectorization
  ----------------------------------------------------------------------------------------------------------

On some x86 processors supporting 512-bits vectorization, compilers are often too conservative and limit vectorization to 256 bits. Performance can then be improved by enforcing 512-bits vectorization, especially with many vectorized and high trip count loops. 512-bits vectorization performance overhead (compared to 256-bits) is generally lower on newer processors.


Workaround
Recompile with -mprefer-vector-width=512


      6.1.1.1.4  -  FMA
  ----------------------------------------------------------------------------------------------------------

Presence of both ADD/SUB and MUL operations.

Workaround
Try to change order in which elements are evaluated (using parentheses) in arithmetic expressions containing both ADD/SUB and MUL operations to enable your compiler to generate FMA instructions wherever possible.
For instance a + b*c is a valid FMA (MUL then ADD).
However (a+b)* c cannot be translated into an FMA (ADD then MUL).




      6.1.1.1.5  -  Complex instructions
  ----------------------------------------------------------------------------------------------------------

Detected COMPLEX INSTRUCTIONS.


Details
These instructions generate more than one micro-operation and only one of them can be decoded during a cycle and the extra micro-operations increase pressure on execution units.
 - VGATHERQPD: 4 occurrences<<list_path_1_complex_1>>



      6.1.1.1.6  -  Slow data structures access
  ----------------------------------------------------------------------------------------------------------

Detected data structures (typically arrays) that cannot be efficiently read/written

Details
 - Irregular (variable stride) or indirect: 1 occurrence(s)
Non-unit stride (uncontiguous) accesses are not efficiently using data caches


Workaround
Try to remove indirect accesses. If applicable, precompute elements out of the innermost loop.


      6.1.1.1.7  -  Gather/scatter instructions
  ----------------------------------------------------------------------------------------------------------

Detected gather/scatter instructions (typically caused by indirect accesses).

Details
 - VGATHERQPD: 4 occurrences<<list_path_1_gather_scatter_1>>


Workaround
Try to simplify your code and/or replace indirect accesses with unit-stride ones.


      6.1.1.1.8  -  Type of elements and instruction set
  ----------------------------------------------------------------------------------------------------------

16 SSE or AVX instructions are processing arithmetic or math operations on double precision FP elements in scalar mode (one at a time).
4 AVX instructions are processing arithmetic or math operations on double precision FP elements in vector mode (four at a time).



      6.1.1.1.9  -  Matching between your loop (in the source code) and the binary loop
  ----------------------------------------------------------------------------------------------------------

The binary loop is composed of 32 FP arithmetical operations:
 - 16: addition or subtraction
 - 16: multiply
The binary loop is loading 384 bytes (48 double precision FP elements).


      6.1.1.1.10  -  Arithmetic intensity
  ----------------------------------------------------------------------------------------------------------

Arithmetic intensity is 0.08 FP operations per loaded or stored byte.







      6.1.2  -  Loop 2138 from exec
  ==========================================================================================================

The loop is defined in /beegfs/hackathon/users/eoseret/qaas_runs_test/178-668-9639/intel/AMG/build/AMG/AMG/parcsr_ls/ams.c:3669-3682.

Warnings:
 - Non-innermost loop: analyzing only self part (ignoring child loops).
 - Ignoring paths for analysis
 - Too many paths. If you really need to analyze all of the 2049 paths individually, rerun with max-paths=2049
 - RecMII not computed since number of paths is unknown or > max_paths
 - Streams not analyzed since number of paths is unknown or > max_paths

Try to simplify control and/or increase the maximum number of paths per function/loop through the 'max-paths-nb' option.

This loop has 2049 execution paths.

The presence of multiple execution paths is typically the main/first bottleneck.
Try to simplify control inside loop: ideally, try to remove all conditional expressions, for example by (if applicable):
 - hoisting them (moving them outside the loop)
 - turning them into conditional moves, MIN or MAX


Ex: if (x<0) x=0 => x = (x<0 ? 0 : x) (or MAX(0,x) after defining the corresponding macro)


      6.1.2.1  -  Path 1
  ----------------------------------------------------------------------------------------------------------

4% of peak computational performance is used (1.55 out of 32.00 FLOP per cycle (GFLOPS @ 1GHz))

      6.1.2.1.1  -  Code clean check
  ----------------------------------------------------------------------------------------------------------

Detected a slowdown caused by scalar integer instructions (typically used for address computation).
By removing them, you can lower the cost of an iteration from 40.67 to 32.50 cycles (1.25x speedup).

Workaround
 - Try to reorganize arrays of structures to structures of arrays
 - Consider to permute loops (see vectorization gain report)



      6.1.2.1.2  -  Vectorization
  ----------------------------------------------------------------------------------------------------------

Your loop is poorly vectorized.
Only 23% of vector register length is used (average across all SSE/AVX instructions).


Details
37% of SSE/AVX instructions are used in vector version (process two or more data elements in vector registers):
 - 75% of SSE/AVX loads are used in vector version.
 - 0% of SSE/AVX stores are used in vector version.
 - 0% of SSE/AVX addition or subtraction instructions are used in vector version.
 - 88% of SSE/AVX multiply instructions are used in vector version.
 - 0% of SSE/AVX fused multiply-add instructions are used in vector version.
 - 0% of SSE/AVX divide and square root instructions are used in vector version.
 - 47% of SSE/AVX instructions that are not load, store, addition, subtraction nor multiply instructions are used in vector version.


Workaround
Read the "512-bits vectorization" report at "Potential" confidence level.


      6.1.2.1.3  -  Execution units bottlenecks
  ----------------------------------------------------------------------------------------------------------

Found no such bottlenecks but see expert reports for more complex bottlenecks.




      6.1.2.1.4  -  512-bits vectorization
  ----------------------------------------------------------------------------------------------------------

On some x86 processors supporting 512-bits vectorization, compilers are often too conservative and limit vectorization to 256 bits. Performance can then be improved by enforcing 512-bits vectorization, especially with many vectorized and high trip count loops. 512-bits vectorization performance overhead (compared to 256-bits) is generally lower on newer processors.


Workaround
Recompile with -mprefer-vector-width=512


      6.1.2.1.5  -  FMA
  ----------------------------------------------------------------------------------------------------------

Detected 2 FMA (fused multiply-add) operations.
Presence of both ADD/SUB and MUL operations.

Workaround
Try to change order in which elements are evaluated (using parentheses) in arithmetic expressions containing both ADD/SUB and MUL operations to enable your compiler to generate FMA instructions wherever possible.
Estimated speedup by perfect pairing: 1.00x.
For instance a + b*c is a valid FMA (MUL then ADD).
However (a+b)* c cannot be translated into an FMA (ADD then MUL).




      6.1.2.1.6  -  Complex instructions
  ----------------------------------------------------------------------------------------------------------

Detected COMPLEX INSTRUCTIONS.


Details
These instructions generate more than one micro-operation and only one of them can be decoded during a cycle and the extra micro-operations increase pressure on execution units.
 - VGATHERQPD: 8 occurrences<<list_path_1_complex_1>>
 - VUCOMISD: 1 occurrences<<list_path_1_complex_2>>



      6.1.2.1.7  -  Gather/scatter instructions
  ----------------------------------------------------------------------------------------------------------

Detected gather/scatter instructions (typically caused by indirect accesses). By removing them, you can lower the cost of an iteration from 40.67 to 35.33 cycles (1.15x speedup).

Details
 - VGATHERQPD: 8 occurrences<<list_path_1_gather_scatter_1>>


Workaround
Try to simplify your code and/or replace indirect accesses with unit-stride ones.


      6.1.2.1.8  -  Type of elements and instruction set
  ----------------------------------------------------------------------------------------------------------

34 SSE or AVX instructions are processing arithmetic or math operations on double precision FP elements in scalar mode (one at a time).
2 SSE or AVX instructions are processing arithmetic or math operations on double precision FP elements in vector mode (two at a time).
6 AVX instructions are processing arithmetic or math operations on double precision FP elements in vector mode (four at a time).



      6.1.2.1.9  -  Matching between your loop (in the source code) and the binary loop
  ----------------------------------------------------------------------------------------------------------

The binary loop is composed of 63 FP arithmetical operations:
 - 31: addition or subtraction (2 inside FMA instructions)
 - 31: multiply (2 inside FMA instructions)
 - 1: divide
The binary loop is loading 888 bytes (111 double precision FP elements).
The binary loop is storing 8 bytes (1 double precision FP elements).


      6.1.2.1.10  -  Arithmetic intensity
  ----------------------------------------------------------------------------------------------------------

Arithmetic intensity is 0.07 FP operations per loaded or stored byte.







      6.1.3  -  Loop 3133 from exec
  ==========================================================================================================

The loop is defined in /beegfs/hackathon/users/eoseret/qaas_runs_test/178-668-9639/intel/AMG/build/AMG/AMG/seq_mv/csr_matvec.c:310-312.

The related source loop is not unrolled or unrolled with no peel/tail loop.

      6.1.3.1  -  Path 1
  ----------------------------------------------------------------------------------------------------------

1% of peak computational performance is used (0.39 out of 32.00 FLOP per cycle (GFLOPS @ 1GHz))

      6.1.3.1.1  -  Vectorization
  ----------------------------------------------------------------------------------------------------------

Your loop is partially vectorized.
Only 28% of vector register length is used (average across all SSE/AVX instructions).


Details
50% of SSE/AVX instructions are used in vector version (process two or more data elements in vector registers):
 - 0% of SSE/AVX addition or subtraction instructions are used in vector version.
 - 75% of SSE/AVX instructions that are not load, store, addition, subtraction nor multiply instructions are used in vector version.


Workaround
Read the "512-bits vectorization" report at "Potential" confidence level.


      6.1.3.1.2  -  Execution units bottlenecks
  ----------------------------------------------------------------------------------------------------------

Found no such bottlenecks but see expert reports for more complex bottlenecks.




      6.1.3.1.3  -  512-bits vectorization
  ----------------------------------------------------------------------------------------------------------

On some x86 processors supporting 512-bits vectorization, compilers are often too conservative and limit vectorization to 256 bits. Performance can then be improved by enforcing 512-bits vectorization, especially with many vectorized and high trip count loops. 512-bits vectorization performance overhead (compared to 256-bits) is generally lower on newer processors.


Workaround
Recompile with -mprefer-vector-width=512


      6.1.3.1.4  -  FMA
  ----------------------------------------------------------------------------------------------------------

Presence of both ADD/SUB and MUL operations.

Workaround
Try to change order in which elements are evaluated (using parentheses) in arithmetic expressions containing both ADD/SUB and MUL operations to enable your compiler to generate FMA instructions wherever possible.
For instance a + b*c is a valid FMA (MUL then ADD).
However (a+b)* c cannot be translated into an FMA (ADD then MUL).




      6.1.3.1.5  -  Complex instructions
  ----------------------------------------------------------------------------------------------------------

Detected COMPLEX INSTRUCTIONS.


Details
These instructions generate more than one micro-operation and only one of them can be decoded during a cycle and the extra micro-operations increase pressure on execution units.
 - VGATHERQPD: 4 occurrences<<list_path_1_complex_1>>



      6.1.3.1.6  -  Slow data structures access
  ----------------------------------------------------------------------------------------------------------

Detected data structures (typically arrays) that cannot be efficiently read/written

Details
 - Irregular (variable stride) or indirect: 1 occurrence(s)
Non-unit stride (uncontiguous) accesses are not efficiently using data caches


Workaround
Try to remove indirect accesses. If applicable, precompute elements out of the innermost loop.


      6.1.3.1.7  -  Gather/scatter instructions
  ----------------------------------------------------------------------------------------------------------

Detected gather/scatter instructions (typically caused by indirect accesses).

Details
 - VGATHERQPD: 4 occurrences<<list_path_1_gather_scatter_1>>


Workaround
Try to simplify your code and/or replace indirect accesses with unit-stride ones.


      6.1.3.1.8  -  Type of elements and instruction set
  ----------------------------------------------------------------------------------------------------------

16 SSE or AVX instructions are processing arithmetic or math operations on double precision FP elements in scalar mode (one at a time).
4 AVX instructions are processing arithmetic or math operations on double precision FP elements in vector mode (four at a time).



      6.1.3.1.9  -  Matching between your loop (in the source code) and the binary loop
  ----------------------------------------------------------------------------------------------------------

The binary loop is composed of 32 FP arithmetical operations:
 - 16: addition or subtraction
 - 16: multiply
The binary loop is loading 384 bytes (48 double precision FP elements).


      6.1.3.1.10  -  Arithmetic intensity
  ----------------------------------------------------------------------------------------------------------

Arithmetic intensity is 0.08 FP operations per loaded or stored byte.







      6.1.4  -  Loop 3142 from exec
  ==========================================================================================================

The loop is defined in /beegfs/hackathon/users/eoseret/qaas_runs_test/178-668-9639/intel/AMG/build/AMG/AMG/seq_mv/csr_matvec.c:259-261.

The related source loop is not unrolled or unrolled with no peel/tail loop.

      6.1.4.1  -  Path 1
  ----------------------------------------------------------------------------------------------------------

1% of peak computational performance is used (0.41 out of 32.00 FLOP per cycle (GFLOPS @ 1GHz))

      6.1.4.1.1  -  Vectorization
  ----------------------------------------------------------------------------------------------------------

Your loop is partially vectorized.
Only 28% of vector register length is used (average across all SSE/AVX instructions).


Details
50% of SSE/AVX instructions are used in vector version (process two or more data elements in vector registers):
 - 0% of SSE/AVX addition or subtraction instructions are used in vector version.
 - 75% of SSE/AVX instructions that are not load, store, addition, subtraction nor multiply instructions are used in vector version.


Workaround
Read the "512-bits vectorization" report at "Potential" confidence level.


      6.1.4.1.2  -  Execution units bottlenecks
  ----------------------------------------------------------------------------------------------------------

Found no such bottlenecks but see expert reports for more complex bottlenecks.




      6.1.4.1.3  -  512-bits vectorization
  ----------------------------------------------------------------------------------------------------------

On some x86 processors supporting 512-bits vectorization, compilers are often too conservative and limit vectorization to 256 bits. Performance can then be improved by enforcing 512-bits vectorization, especially with many vectorized and high trip count loops. 512-bits vectorization performance overhead (compared to 256-bits) is generally lower on newer processors.


Workaround
Recompile with -mprefer-vector-width=512


      6.1.4.1.4  -  FMA
  ----------------------------------------------------------------------------------------------------------

Presence of both ADD/SUB and MUL operations.

Workaround
Try to change order in which elements are evaluated (using parentheses) in arithmetic expressions containing both ADD/SUB and MUL operations to enable your compiler to generate FMA instructions wherever possible.
For instance a + b*c is a valid FMA (MUL then ADD).
However (a+b)* c cannot be translated into an FMA (ADD then MUL).




      6.1.4.1.5  -  Complex instructions
  ----------------------------------------------------------------------------------------------------------

Detected COMPLEX INSTRUCTIONS.


Details
These instructions generate more than one micro-operation and only one of them can be decoded during a cycle and the extra micro-operations increase pressure on execution units.
 - VGATHERQPD: 4 occurrences<<list_path_1_complex_1>>



      6.1.4.1.6  -  Slow data structures access
  ----------------------------------------------------------------------------------------------------------

Detected data structures (typically arrays) that cannot be efficiently read/written

Details
 - Irregular (variable stride) or indirect: 1 occurrence(s)
Non-unit stride (uncontiguous) accesses are not efficiently using data caches


Workaround
Try to remove indirect accesses. If applicable, precompute elements out of the innermost loop.


      6.1.4.1.7  -  Gather/scatter instructions
  ----------------------------------------------------------------------------------------------------------

Detected gather/scatter instructions (typically caused by indirect accesses).

Details
 - VGATHERQPD: 4 occurrences<<list_path_1_gather_scatter_1>>


Workaround
Try to simplify your code and/or replace indirect accesses with unit-stride ones.


      6.1.4.1.8  -  Type of elements and instruction set
  ----------------------------------------------------------------------------------------------------------

16 SSE or AVX instructions are processing arithmetic or math operations on double precision FP elements in scalar mode (one at a time).
4 AVX instructions are processing arithmetic or math operations on double precision FP elements in vector mode (four at a time).



      6.1.4.1.9  -  Matching between your loop (in the source code) and the binary loop
  ----------------------------------------------------------------------------------------------------------

The binary loop is composed of 32 FP arithmetical operations:
 - 16: addition or subtraction
 - 16: multiply
The binary loop is loading 384 bytes (48 double precision FP elements).


      6.1.4.1.10  -  Arithmetic intensity
  ----------------------------------------------------------------------------------------------------------

Arithmetic intensity is 0.08 FP operations per loaded or stored byte.







      6.1.5  -  Loop 3132 from exec
  ==========================================================================================================

The loop is defined in /beegfs/hackathon/users/eoseret/qaas_runs_test/178-668-9639/intel/AMG/build/AMG/AMG/seq_mv/csr_matvec.c:307-314.

Warnings:
 - Non-innermost loop: analyzing only self part (ignoring child loops).
 - Ignoring paths for analysis
 - Too many paths. Rerun with max-paths=32
 - RecMII not computed since number of paths is unknown or > max_paths
 - Streams not analyzed since number of paths is unknown or > max_paths

Try to simplify control and/or increase the maximum number of paths per function/loop through the 'max-paths-nb' option.

This loop has 32 execution paths.

The presence of multiple execution paths is typically the main/first bottleneck.
Try to simplify control inside loop: ideally, try to remove all conditional expressions, for example by (if applicable):
 - hoisting them (moving them outside the loop)
 - turning them into conditional moves, MIN or MAX


Ex: if (x<0) x=0 => x = (x<0 ? 0 : x) (or MAX(0,x) after defining the corresponding macro)


      6.1.5.1  -  Path 1
  ----------------------------------------------------------------------------------------------------------

4% of peak computational performance is used (1.57 out of 32.00 FLOP per cycle (GFLOPS @ 1GHz))

      6.1.5.1.1  -  Code clean check
  ----------------------------------------------------------------------------------------------------------

Detected a slowdown caused by scalar integer instructions (typically used for address computation).
By removing them, you can lower the cost of an iteration from 19.17 to 16.00 cycles (1.20x speedup).

Workaround
 - Try to reorganize arrays of structures to structures of arrays
 - Consider to permute loops (see vectorization gain report)



      6.1.5.1.2  -  Vectorization
  ----------------------------------------------------------------------------------------------------------

Your loop is poorly vectorized.
Only 23% of vector register length is used (average across all SSE/AVX instructions).


Details
36% of SSE/AVX instructions are used in vector version (process two or more data elements in vector registers):
 - 80% of SSE/AVX loads are used in vector version.
 - 0% of SSE/AVX stores are used in vector version.
 - 0% of SSE/AVX addition or subtraction instructions are used in vector version.
 - 0% of SSE/AVX fused multiply-add instructions are used in vector version.
 - 45% of SSE/AVX instructions that are not load, store, addition, subtraction nor multiply instructions are used in vector version.


Workaround
Read the "512-bits vectorization" report at "Potential" confidence level.


      6.1.5.1.3  -  Execution units bottlenecks
  ----------------------------------------------------------------------------------------------------------

Found no such bottlenecks but see expert reports for more complex bottlenecks.




      6.1.5.1.4  -  512-bits vectorization
  ----------------------------------------------------------------------------------------------------------

On some x86 processors supporting 512-bits vectorization, compilers are often too conservative and limit vectorization to 256 bits. Performance can then be improved by enforcing 512-bits vectorization, especially with many vectorized and high trip count loops. 512-bits vectorization performance overhead (compared to 256-bits) is generally lower on newer processors.


Workaround
Recompile with -mprefer-vector-width=512


      6.1.5.1.5  -  FMA
  ----------------------------------------------------------------------------------------------------------

Detected 1 FMA (fused multiply-add) operations.
Presence of both ADD/SUB and MUL operations.

Workaround
Try to change order in which elements are evaluated (using parentheses) in arithmetic expressions containing both ADD/SUB and MUL operations to enable your compiler to generate FMA instructions wherever possible.
For instance a + b*c is a valid FMA (MUL then ADD).
However (a+b)* c cannot be translated into an FMA (ADD then MUL).




      6.1.5.1.6  -  Complex instructions
  ----------------------------------------------------------------------------------------------------------

Detected COMPLEX INSTRUCTIONS.


Details
These instructions generate more than one micro-operation and only one of them can be decoded during a cycle and the extra micro-operations increase pressure on execution units.
 - VGATHERQPD: 4 occurrences<<list_path_1_complex_1>>



      6.1.5.1.7  -  Gather/scatter instructions
  ----------------------------------------------------------------------------------------------------------

Detected gather/scatter instructions (typically caused by indirect accesses). By removing them, you can lower the cost of an iteration from 19.17 to 16.50 cycles (1.16x speedup).

Details
 - VGATHERQPD: 4 occurrences<<list_path_1_gather_scatter_1>>


Workaround
Try to simplify your code and/or replace indirect accesses with unit-stride ones.


      6.1.5.1.8  -  Type of elements and instruction set
  ----------------------------------------------------------------------------------------------------------

15 SSE or AVX instructions are processing arithmetic or math operations on double precision FP elements in scalar mode (one at a time).
1 SSE or AVX instructions are processing arithmetic or math operations on double precision FP elements in vector mode (two at a time).
3 AVX instructions are processing arithmetic or math operations on double precision FP elements in vector mode (four at a time).



      6.1.5.1.9  -  Matching between your loop (in the source code) and the binary loop
  ----------------------------------------------------------------------------------------------------------

The binary loop is composed of 30 FP arithmetical operations:
 - 15: addition or subtraction (1 inside FMA instructions)
 - 15: multiply (1 inside FMA instructions)
The binary loop is loading 400 bytes (50 double precision FP elements).
The binary loop is storing 8 bytes (1 double precision FP elements).


      6.1.5.1.10  -  Arithmetic intensity
  ----------------------------------------------------------------------------------------------------------

Arithmetic intensity is 0.07 FP operations per loaded or stored byte.







      6.1.6  -  Loop 3141 from exec
  ==========================================================================================================

The loop is defined in /beegfs/hackathon/users/eoseret/qaas_runs_test/178-668-9639/intel/AMG/build/AMG/AMG/seq_mv/csr_matvec.c:256-263.

Warnings:
 - Non-innermost loop: analyzing only self part (ignoring child loops).
 - Ignoring paths for analysis
 - Failed to get the number of paths
 - RecMII not computed since number of paths is unknown or > max_paths
 - Streams not analyzed since number of paths is unknown or > max_paths


      6.1.6.1  -  Path 1
  ----------------------------------------------------------------------------------------------------------

4% of peak computational performance is used (1.57 out of 32.00 FLOP per cycle (GFLOPS @ 1GHz))

      6.1.6.1.1  -  Code clean check
  ----------------------------------------------------------------------------------------------------------

Detected a slowdown caused by scalar integer instructions (typically used for address computation).
By removing them, you can lower the cost of an iteration from 19.17 to 16.00 cycles (1.20x speedup).

Workaround
 - Try to reorganize arrays of structures to structures of arrays
 - Consider to permute loops (see vectorization gain report)



      6.1.6.1.2  -  Vectorization
  ----------------------------------------------------------------------------------------------------------

Your loop is poorly vectorized.
Only 23% of vector register length is used (average across all SSE/AVX instructions).


Details
41% of SSE/AVX instructions are used in vector version (process two or more data elements in vector registers):
 - 85% of SSE/AVX loads are used in vector version.
 - 0% of SSE/AVX stores are used in vector version.
 - 0% of SSE/AVX addition or subtraction instructions are used in vector version.
 - 0% of SSE/AVX fused multiply-add instructions are used in vector version.
 - 52% of SSE/AVX instructions that are not load, store, addition, subtraction nor multiply instructions are used in vector version.


Workaround
Read the "512-bits vectorization" report at "Potential" confidence level.


      6.1.6.1.3  -  Execution units bottlenecks
  ----------------------------------------------------------------------------------------------------------

Found no such bottlenecks but see expert reports for more complex bottlenecks.




      6.1.6.1.4  -  512-bits vectorization
  ----------------------------------------------------------------------------------------------------------

On some x86 processors supporting 512-bits vectorization, compilers are often too conservative and limit vectorization to 256 bits. Performance can then be improved by enforcing 512-bits vectorization, especially with many vectorized and high trip count loops. 512-bits vectorization performance overhead (compared to 256-bits) is generally lower on newer processors.


Workaround
Recompile with -mprefer-vector-width=512


      6.1.6.1.5  -  FMA
  ----------------------------------------------------------------------------------------------------------

Detected 1 FMA (fused multiply-add) operations.
Presence of both ADD/SUB and MUL operations.

Workaround
Try to change order in which elements are evaluated (using parentheses) in arithmetic expressions containing both ADD/SUB and MUL operations to enable your compiler to generate FMA instructions wherever possible.
For instance a + b*c is a valid FMA (MUL then ADD).
However (a+b)* c cannot be translated into an FMA (ADD then MUL).




      6.1.6.1.6  -  Complex instructions
  ----------------------------------------------------------------------------------------------------------

Detected COMPLEX INSTRUCTIONS.


Details
These instructions generate more than one micro-operation and only one of them can be decoded during a cycle and the extra micro-operations increase pressure on execution units.
 - VGATHERQPD: 4 occurrences<<list_path_1_complex_1>>



      6.1.6.1.7  -  Gather/scatter instructions
  ----------------------------------------------------------------------------------------------------------

Detected gather/scatter instructions (typically caused by indirect accesses). By removing them, you can lower the cost of an iteration from 19.17 to 16.50 cycles (1.16x speedup).

Details
 - VGATHERQPD: 4 occurrences<<list_path_1_gather_scatter_1>>


Workaround
Try to simplify your code and/or replace indirect accesses with unit-stride ones.


      6.1.6.1.8  -  Type of elements and instruction set
  ----------------------------------------------------------------------------------------------------------

15 SSE or AVX instructions are processing arithmetic or math operations on double precision FP elements in scalar mode (one at a time).
1 SSE or AVX instructions are processing arithmetic or math operations on double precision FP elements in vector mode (two at a time).
3 AVX instructions are processing arithmetic or math operations on double precision FP elements in vector mode (four at a time).



      6.1.6.1.9  -  Matching between your loop (in the source code) and the binary loop
  ----------------------------------------------------------------------------------------------------------

The binary loop is composed of 30 FP arithmetical operations:
 - 15: addition or subtraction (1 inside FMA instructions)
 - 15: multiply (1 inside FMA instructions)
The binary loop is loading 384 bytes (48 double precision FP elements).
The binary loop is storing 8 bytes (1 double precision FP elements).


      6.1.6.1.10  -  Arithmetic intensity
  ----------------------------------------------------------------------------------------------------------

Arithmetic intensity is 0.08 FP operations per loaded or stored byte.







      6.1.7  -  Loop 3128 from exec
  ==========================================================================================================

The loop is defined in /beegfs/hackathon/users/eoseret/qaas_runs_test/178-668-9639/intel/AMG/build/AMG/AMG/seq_mv/csr_matvec.c:334-341.

Warnings:
 - Non-innermost loop: analyzing only self part (ignoring child loops).
 - Ignoring paths for analysis
 - Too many paths. Rerun with max-paths=32
 - RecMII not computed since number of paths is unknown or > max_paths
 - Streams not analyzed since number of paths is unknown or > max_paths

Try to simplify control and/or increase the maximum number of paths per function/loop through the 'max-paths-nb' option.

This loop has 32 execution paths.

The presence of multiple execution paths is typically the main/first bottleneck.
Try to simplify control inside loop: ideally, try to remove all conditional expressions, for example by (if applicable):
 - hoisting them (moving them outside the loop)
 - turning them into conditional moves, MIN or MAX


Ex: if (x<0) x=0 => x = (x<0 ? 0 : x) (or MAX(0,x) after defining the corresponding macro)


      6.1.7.1  -  Path 1
  ----------------------------------------------------------------------------------------------------------

4% of peak computational performance is used (1.57 out of 32.00 FLOP per cycle (GFLOPS @ 1GHz))

      6.1.7.1.1  -  Code clean check
  ----------------------------------------------------------------------------------------------------------

Detected a slowdown caused by scalar integer instructions (typically used for address computation).
By removing them, you can lower the cost of an iteration from 19.17 to 16.00 cycles (1.20x speedup).

Workaround
 - Try to reorganize arrays of structures to structures of arrays
 - Consider to permute loops (see vectorization gain report)



      6.1.7.1.2  -  Vectorization
  ----------------------------------------------------------------------------------------------------------

Your loop is poorly vectorized.
Only 23% of vector register length is used (average across all SSE/AVX instructions).


Details
36% of SSE/AVX instructions are used in vector version (process two or more data elements in vector registers):
 - 80% of SSE/AVX loads are used in vector version.
 - 0% of SSE/AVX stores are used in vector version.
 - 0% of SSE/AVX addition or subtraction instructions are used in vector version.
 - 0% of SSE/AVX fused multiply-add instructions are used in vector version.
 - 45% of SSE/AVX instructions that are not load, store, addition, subtraction nor multiply instructions are used in vector version.


Workaround
Read the "512-bits vectorization" report at "Potential" confidence level.


      6.1.7.1.3  -  Execution units bottlenecks
  ----------------------------------------------------------------------------------------------------------

Found no such bottlenecks but see expert reports for more complex bottlenecks.




      6.1.7.1.4  -  512-bits vectorization
  ----------------------------------------------------------------------------------------------------------

On some x86 processors supporting 512-bits vectorization, compilers are often too conservative and limit vectorization to 256 bits. Performance can then be improved by enforcing 512-bits vectorization, especially with many vectorized and high trip count loops. 512-bits vectorization performance overhead (compared to 256-bits) is generally lower on newer processors.


Workaround
Recompile with -mprefer-vector-width=512


      6.1.7.1.5  -  FMA
  ----------------------------------------------------------------------------------------------------------

Detected 1 FMA (fused multiply-add) operations.
Presence of both ADD/SUB and MUL operations.

Workaround
Try to change order in which elements are evaluated (using parentheses) in arithmetic expressions containing both ADD/SUB and MUL operations to enable your compiler to generate FMA instructions wherever possible.
For instance a + b*c is a valid FMA (MUL then ADD).
However (a+b)* c cannot be translated into an FMA (ADD then MUL).




      6.1.7.1.6  -  Complex instructions
  ----------------------------------------------------------------------------------------------------------

Detected COMPLEX INSTRUCTIONS.


Details
These instructions generate more than one micro-operation and only one of them can be decoded during a cycle and the extra micro-operations increase pressure on execution units.
 - VGATHERQPD: 4 occurrences<<list_path_1_complex_1>>



      6.1.7.1.7  -  Gather/scatter instructions
  ----------------------------------------------------------------------------------------------------------

Detected gather/scatter instructions (typically caused by indirect accesses). By removing them, you can lower the cost of an iteration from 19.17 to 16.50 cycles (1.16x speedup).

Details
 - VGATHERQPD: 4 occurrences<<list_path_1_gather_scatter_1>>


Workaround
Try to simplify your code and/or replace indirect accesses with unit-stride ones.


      6.1.7.1.8  -  Type of elements and instruction set
  ----------------------------------------------------------------------------------------------------------

15 SSE or AVX instructions are processing arithmetic or math operations on double precision FP elements in scalar mode (one at a time).
1 SSE or AVX instructions are processing arithmetic or math operations on double precision FP elements in vector mode (two at a time).
3 AVX instructions are processing arithmetic or math operations on double precision FP elements in vector mode (four at a time).



      6.1.7.1.9  -  Matching between your loop (in the source code) and the binary loop
  ----------------------------------------------------------------------------------------------------------

The binary loop is composed of 30 FP arithmetical operations:
 - 15: addition or subtraction (1 inside FMA instructions)
 - 15: multiply (1 inside FMA instructions)
The binary loop is loading 400 bytes (50 double precision FP elements).
The binary loop is storing 8 bytes (1 double precision FP elements).


      6.1.7.1.10  -  Arithmetic intensity
  ----------------------------------------------------------------------------------------------------------

Arithmetic intensity is 0.07 FP operations per loaded or stored byte.







      6.1.8  -  Loop 3184 from exec
  ==========================================================================================================

The loop is defined in /beegfs/hackathon/users/eoseret/qaas_runs_test/178-668-9639/intel/AMG/build/AMG/AMG/seq_mv/vector.c:452.

It is main loop of related source loop which is unrolled by 4 (including vectorization).

      6.1.8.1  -  Path 1
  ----------------------------------------------------------------------------------------------------------

35% of peak computational performance is used (11.29 out of 32.00 FLOP per cycle (GFLOPS @ 1GHz))

      6.1.8.1.1  -  Vectorization
  ----------------------------------------------------------------------------------------------------------

Your loop is vectorized, but using only 256 out of 512 bits (AVX/AVX2 instructions on AVX-512 processors).
<<image_4x64_512>>

Details
All SSE/AVX instructions are used in vector version (process two or more data elements in vector registers).


Workaround
Read the "512-bits vectorization" report at "Potential" confidence level.


      6.1.8.1.2  -  Execution units bottlenecks
  ----------------------------------------------------------------------------------------------------------

Found no such bottlenecks but see expert reports for more complex bottlenecks.




      6.1.8.1.3  -  512-bits vectorization
  ----------------------------------------------------------------------------------------------------------

On some x86 processors supporting 512-bits vectorization, compilers are often too conservative and limit vectorization to 256 bits. Performance can then be improved by enforcing 512-bits vectorization, especially with many vectorized and high trip count loops. 512-bits vectorization performance overhead (compared to 256-bits) is generally lower on newer processors.


Workaround
Recompile with -mprefer-vector-width=512


      6.1.8.1.4  -  FMA
  ----------------------------------------------------------------------------------------------------------

Detected 32 FMA (fused multiply-add) operations.




      6.1.8.1.5  -  Vector unaligned load/store instructions
  ----------------------------------------------------------------------------------------------------------

Detected 16 optimal vector unaligned load/store instructions.


Details
 - VMOVUPD: 16 occurrences<<list_path_1_vec_align_1>>


Workaround
Use vector aligned instructions:
 1) align your arrays on 64 bytes boundaries: replace { void *p = malloc (size); } with { void *p; posix_memalign (&p, 64, size); }.
 2) inform your compiler that your arrays are vector aligned: if array 'foo' is 64 bytes-aligned, define a pointer 'p_foo' as __builtin_assume_aligned (foo, 64) and use it instead of 'foo' in the loop.
<<image_vec_align>>


      6.1.8.1.6  -  Type of elements and instruction set
  ----------------------------------------------------------------------------------------------------------

8 AVX instructions are processing arithmetic or math operations on double precision FP elements in vector mode (four at a time).



      6.1.8.1.7  -  Matching between your loop (in the source code) and the binary loop
  ----------------------------------------------------------------------------------------------------------

The binary loop is composed of 64 FP arithmetical operations:
 - 32: addition or subtraction (all inside FMA instructions)
 - 32: multiply (all inside FMA instructions)
The binary loop is loading 512 bytes (64 double precision FP elements).
The binary loop is storing 256 bytes (32 double precision FP elements).


      6.1.8.1.8  -  Arithmetic intensity
  ----------------------------------------------------------------------------------------------------------

Arithmetic intensity is 0.08 FP operations per loaded or stored byte.







      6.1.9  -  Loop 2137 from exec
  ==========================================================================================================

The loop is defined in /beegfs/hackathon/users/eoseret/qaas_runs_test/178-668-9639/intel/AMG/build/AMG/AMG/parcsr_ls/ams.c:3659.

It is main loop of related source loop which is unrolled by 4 (including vectorization).

      6.1.9.1  -  Path 1
  ----------------------------------------------------------------------------------------------------------

0% of peak computational performance is used (0.00 out of 64.00 FLOP per cycle (GFLOPS @ 1GHz))

      6.1.9.1.1  -  Vectorization
  ----------------------------------------------------------------------------------------------------------

Your loop is vectorized, but using only 256 out of 512 bits (AVX/AVX2 instructions on AVX-512 processors).


Details
All SSE/AVX instructions are used in vector version (process two or more data elements in vector registers).


Workaround
Read the "512-bits vectorization" report at "Potential" confidence level.


      6.1.9.1.2  -  Execution units bottlenecks
  ----------------------------------------------------------------------------------------------------------

Performance is limited by writing data to caches/RAM (the store unit is a bottleneck).


Workaround
 - Write less array elements
 - Provide more information to your compiler:
  * hardcode the bounds of the corresponding 'for' loop
  * use the 'restrict' C99 keyword





      6.1.9.1.3  -  512-bits vectorization
  ----------------------------------------------------------------------------------------------------------

On some x86 processors supporting 512-bits vectorization, compilers are often too conservative and limit vectorization to 256 bits. Performance can then be improved by enforcing 512-bits vectorization, especially with many vectorized and high trip count loops. 512-bits vectorization performance overhead (compared to 256-bits) is generally lower on newer processors.


Workaround
Recompile with -mprefer-vector-width=512




      6.1.9.1.4  -  Vector unaligned load/store instructions
  ----------------------------------------------------------------------------------------------------------

Detected 16 optimal vector unaligned load/store instructions.


Details
 - VMOVUPD: 16 occurrences<<list_path_1_vec_align_1>>


Workaround
Use vector aligned instructions:
 1) align your arrays on 64 bytes boundaries: replace { void *p = malloc (size); } with { void *p; posix_memalign (&p, 64, size); }.
 2) inform your compiler that your arrays are vector aligned: if array 'foo' is 64 bytes-aligned, define a pointer 'p_foo' as __builtin_assume_aligned (foo, 64) and use it instead of 'foo' in the loop.
<<image_vec_align>>


      6.1.9.1.5  -  Type of elements and instruction set
  ----------------------------------------------------------------------------------------------------------

No instructions are processing arithmetic or math operations on FP elements. This loop is probably writing/copying data or processing integer elements.


      6.1.9.1.6  -  Matching between your loop (in the source code) and the binary loop
  ----------------------------------------------------------------------------------------------------------

The binary loop does not contain any FP arithmetical operations.
The binary loop is loading 256 bytes.
The binary loop is storing 256 bytes.







      6.1.10  -  Loop 3149 from exec
  ==========================================================================================================

The loop is defined in /beegfs/hackathon/users/eoseret/qaas_runs_test/178-668-9639/intel/AMG/build/AMG/AMG/seq_mv/csr_matvec.c:564-567.

Analyzed code is defined in /beegfs/hackathon/users/eoseret/qaas_runs_test/178-668-9639/intel/AMG/build/AMG/AMG/seq_mv/csr_matvec.c:564,567.

Warnings:
 - Non-innermost loop: analyzing only self part (ignoring child loops).
 - Ignoring paths for analysis
 - Too many paths. Rerun with max-paths=16
 - RecMII not computed since number of paths is unknown or > max_paths
 - Streams not analyzed since number of paths is unknown or > max_paths

Try to simplify control and/or increase the maximum number of paths per function/loop through the 'max-paths-nb' option.

This loop has 16 execution paths.

The presence of multiple execution paths is typically the main/first bottleneck.
Try to simplify control inside loop: ideally, try to remove all conditional expressions, for example by (if applicable):
 - hoisting them (moving them outside the loop)
 - turning them into conditional moves, MIN or MAX


Ex: if (x<0) x=0 => x = (x<0 ? 0 : x) (or MAX(0,x) after defining the corresponding macro)


      6.1.10.1  -  Path 1
  ----------------------------------------------------------------------------------------------------------

3% of peak computational performance is used (1.14 out of 32.00 FLOP per cycle (GFLOPS @ 1GHz))

      6.1.10.1.1  -  Code clean check
  ----------------------------------------------------------------------------------------------------------

Detected a slowdown caused by scalar integer instructions (typically used for address computation).
By removing them, you can lower the cost of an iteration from 12.33 to 7.00 cycles (1.76x speedup).

Workaround
 - Try to reorganize arrays of structures to structures of arrays
 - Consider to permute loops (see vectorization gain report)



      6.1.10.1.2  -  Vectorization
  ----------------------------------------------------------------------------------------------------------

Your loop is not vectorized.
8 data elements could be processed at once in vector registers.
<<image_1x64_512>>By vectorizing your loop, you can lower the cost of an iteration from 12.33 to 1.54 cycles (8.00x speedup).

Details
All SSE/AVX instructions are used in scalar version (process only one data element in vector registers).
Since your execution units are vector units, only a vectorized loop can use their full power.


Workaround
 - Try another compiler or update/tune your current one:
  * recompile with fassociative-math (included in Ofast or ffast-math) to extend loop vectorization to FP reductions.
 - Remove inter-iterations dependences from your loop and make it unit-stride:
  * If your arrays have 2 or more dimensions, check whether elements are accessed contiguously and, otherwise, try to permute loops accordingly:
C storage order is row-major: for(i) for(j) a[j][i] = b[j][i]; (slow, non stride 1) => for(i) for(j) a[i][j] = b[i][j]; (fast, stride 1)<<image_row_maj>>
  * If your loop streams arrays of structures (AoS), try to use structures of arrays instead (SoA):
for(i) a[i].x = b[i].x; (slow, non stride 1) => for(i) a.x[i] = b.x[i]; (fast, stride 1)



      6.1.10.1.3  -  Execution units bottlenecks
  ----------------------------------------------------------------------------------------------------------

Found no such bottlenecks but see expert reports for more complex bottlenecks.




      6.1.10.1.4  -  FMA
  ----------------------------------------------------------------------------------------------------------

Detected 7 FMA (fused multiply-add) operations.




      6.1.10.1.5  -  Type of elements and instruction set
  ----------------------------------------------------------------------------------------------------------

7 SSE or AVX instructions are processing arithmetic or math operations on double precision FP elements in scalar mode (one at a time).



      6.1.10.1.6  -  Matching between your loop (in the source code) and the binary loop
  ----------------------------------------------------------------------------------------------------------

The binary loop is composed of 14 FP arithmetical operations:
 - 7: addition or subtraction (all inside FMA instructions)
 - 7: multiply (all inside FMA instructions)
The binary loop is loading 240 bytes (30 double precision FP elements).
The binary loop is storing 56 bytes (7 double precision FP elements).


      6.1.10.1.7  -  Arithmetic intensity
  ----------------------------------------------------------------------------------------------------------

Arithmetic intensity is 0.05 FP operations per loaded or stored byte.





[MAQAO] Info: STOP THE REPORT GENERATION
[MAQAO] Info: 
[MAQAO] Info: If your application produces files, they can be found in directory "/beegfs/hackathon/users/eoseret/qaas_runs_test/178-668-9639/intel/AMG/run/oneview_runs/defaults/gcc/oneview_run_1786690412"
[MAQAO] Info: 
