	*************************************************
	*                                               *
	*          ONE-View report generation           *
	*                                               *
	*************************************************

[MAQAO] Info: Experiment configuration summary is available adding -dbg=1 in command line

* [MAQAO] Warning: Experiment directory /beegfs/hackathon/users/eoseret/qaas_runs_test/178-661-4073/intel/CloverLeaf2.0-CXX/run/oneview_runs/defaults/icx/oneview_results_1786614308 already exists and is reused.
           It can be replaced using --replace in the command line.
[MAQAO] Info: 
[MAQAO] Info: START THE APPLICATION PROFILING
[MAQAO] Info: -> RUNNING THE PROFILER...
[MAQAO] Info:   LPROF has already been run
[MAQAO] Info: STOP THE APPLICATION PROFILING
[MAQAO] Info: 
[MAQAO] Info: START FUNCTIONS AND LOOPS ANALYSIS ...
[MAQAO] Info: STOP FUNCTIONS AND LOOPS ANALYSIS ...
[MAQAO] Info: 
[MAQAO] Info: START THE REPORT GENERATION
[MAQAO] Info: -> ONE-VIEW EXPERIMENT DIRECTORY: /beegfs/hackathon/users/eoseret/qaas_runs_test/178-661-4073/intel/CloverLeaf2.0-CXX/run/oneview_runs/defaults/icx/oneview_results_1786614308


+====================================================================================================================+
+                                                    1  -  GLOBAL                                                    +
+====================================================================================================================+


+--------------------------------------------------------------------------------------------------------------------+
+                                             1.1  -  Experiment Summary                                             +
+--------------------------------------------------------------------------------------------------------------------+

  Application:			/beegfs/hackathon/users/eoseret/qaas_runs_test/178-661-4073/intel/CloverLeaf2.0-CXX/run/oneview_runs/defaults/orig/exec
  Timestamp:			2026-08-13 11:45:07
  Universal Timestamp:		1786614307
  Experiment Type:		MPI; OpenMP; Throughput; 
  Machine:			isix07.benchmarkcenter.megware.com
  Architecture:			x86_64
  Micro Architecture:		GRANITE_RAPIDS
  Model Name:			Intel(R) Xeon(R) 6972P
  Cache Size:			491520 KB
  Number of Cores:		96
  OS Version:			Linux 5.14.0-687.31.1.el9_8.x86_64 #1 SMP PREEMPT_DYNAMIC Sat Aug 1 05:38:01 EDT 2026
  Compilation Options:		
		exec:  --driver-mode=g++ --intel -D USE_OMP -I /beegfs/hackathon/users/eoseret/qaas_runs_test/178-661-4073/intel/CloverLeaf2.0-CXX/build/CloverLeaf2.0-CXX/omp -I /beegfs/hackathon/users/eoseret/qaas_runs_test/178-661-4073/intel/CloverLeaf2.0-CXX/build/build/generated -I /beegfs/hackathon/users/eoseret/qaas_runs_test/178-661-4073/intel/CloverLeaf2.0-CXX/build/CloverLeaf2.0-CXX/driver -I /beegfs/hackathon/users/eoseret/qaas_runs_test/178-661-4073/intel/CloverLeaf2.0-CXX/build/CloverLeaf2.0-CXX/src/omp -g -fno-omit-frame-pointer -fcf-protection=none -no-pie -grecord-command-line -D NDEBUG -std=c++17 -Wall -Wno-unused-parameter -Wno-unused-function -Wno-unused-variable -O3 -fiopenmp -MD -MT CMakeFiles/cloverleaf.dir/src/omp/advec_mom.cpp.o -MF CMakeFiles/cloverleaf.dir/src/omp/advec_mom.cpp.o.d -o CMakeFiles/cloverleaf.dir/src/omp/advec_mom.cpp.o -c /beegfs/hackathon/users/eoseret/qaas_runs_test/178-661-4073/intel/CloverLeaf2.0-CXX/build/CloverLeaf2.0-CXX/src/omp/advec_mom.cpp -I /cluster/hpcx/2.22/ompi5-ifx-mt/include -I /cluster/hpcx/2.22/ompi5-ifx-mt/include/openmpi -fveclib=SVML 
  Number of processes observed:	6
  Number of threads observed:	192
  MAQAO version:		2026.1.0
  MAQAO build:			6d1be1d51c1e63266254997eb301734a7264775d::20260810-150026




+--------------------------------------------------------------------------------------------------------------------+
+                                               1.2  -  Global Metrics                                               +
+--------------------------------------------------------------------------------------------------------------------+

  Total Time:				48.47 s
  Max (Thread Active Time):		47.81 s
  Average Active Time:			47.55 s
  Activity Ratio:			99.9 %
  Average number of active threads:	188.351
  Affinity Stability:			99.8 %
  Time spent in analyzed loops:		89.1 %
  Time spent in analyzed innermost loops: 88.8 %
  Time spent in user code:		89.1 %
  Compilation Options Score:		66.67
  Array Access Efficiency:		53.6 %

   Potential Speedups
  ----------------------------------------------------
  Perfect Flow Complexity:		1.21
  Perfect OpenMP/MPI/Pthread/TBB:	1.02
  Perfect OpenMP/MPI/Pthread/TBB + Load Distribution:	1.05
  If No Scalar Integer:
      Potential Speedup:		1.01
      Nb Loops to get 80%:		9
  If FP Vectorized:
      Potential Speedup:		1.17
      Nb Loops to get 80%:		15
  If Fully Vectorized:
      Potential Speedup:		2.96
      Nb Loops to get 80%:		25
  If Only FP Arithmetic:
      Potential Speedup:		4.14
      Nb Loops to get 80%:		26




+--------------------------------------------------------------------------------------------------------------------+
+                                             1.3  -  Potential Speedups                                             +
+--------------------------------------------------------------------------------------------------------------------+

  If No Scalar Integer:
      Number of loops   | 1      | 10     | 20     | 29     | 40     | 
      Cumulated Speedup | 1.0022 | 1.0121 | 1.0137 | 1.0137 | 1.0137 | 
  Top 5 loops:
    exec - 162:	1.0022
    exec - 170:	1.0043
    exec - 152:	1.0054
    exec - 146:	1.0065
    exec - 172:	1.0075

  If FP Vectorized:
      Number of loops   | 1      | 10     | 20     | 29     | 40     | 
      Cumulated Speedup | 1.0141 | 1.1098 | 1.1572 | 1.1728 | 1.1729 | 
  Top 5 loops:
    exec - 162:	1.0141
    exec - 170:	1.0277
    exec - 942:	1.0402
    exec - 142:	1.0511
    exec - 298:	1.0616

  If Fully Vectorized:
      Number of loops   | 1      | 10     | 20     | 29     | 40     | 
      Cumulated Speedup | 1.0490 | 1.5465 | 2.2461 | 2.8574 | 2.9647 | 
  Top 5 loops:
    exec - 162:	1.049
    exec - 298:	1.0989
    exec - 142:	1.1521
    exec - 170:	1.2067
    exec - 300:	1.2603

  If Only FP Arithmetic:
      Number of loops   | 1      | 10     | 20     | 29     | 40     | 
      Cumulated Speedup | 1.0612 | 1.6788 | 2.7276 | 3.9111 | 4.1421 | 
  Top 5 loops:
    exec - 162:	1.0612
    exec - 170:	1.1203
    exec - 298:	1.1848
    exec - 142:	1.251
    exec - 300:	1.3169



+====================================================================================================================+
+                                                   2  -  SUMMARY                                                    +
+====================================================================================================================+


+--------------------------------------------------------------------------------------------------------------------+
+                                             2.1  -  EXPERIMENT QUALITY                                             +
+--------------------------------------------------------------------------------------------------------------------+

  [4 / 4] Application profile is long enough (47.81 s)
To have good quality measurements, it is advised that the application profiling time is greater than 10 seconds.

  [3 / 3] Most of time spent in analyzed modules comes from functions with source/debug info
-g option gives access to debugging informations, such are source locations.

  [0 / 3] Compilation of some functions is not optimized for the target processor
Architecture specific options are needed to produce efficient code for a specific processor ( -march=(target) ).

  [3 / 3] Most of time spent in analyzed modules comes from functions with compilation options informations and
-fno-omit-frame-pointer is present
-fno-omit-frame-pointer improves the accuracy of callchains found during the application profiling.

  [3 / 3] Optimization level option is correctly used


  [3 / 3] Host configuration allows retrieval of all necessary metrics.


  [2 / 2] Application is correctly profiled ("Others" category represents 0.01 % of the execution time)
To have a representative profiling, it is advised that the category "Others" represents less than 20% of the execution
time in order to analyze as much as possible of the user code

  [1 / 1] Lstopo present. The Topology lstopo report will be generated.



+--------------------------------------------------------------------------------------------------------------------+
+                                                2.2  -  CODE QUALITY                                                +
+--------------------------------------------------------------------------------------------------------------------+

  [4 / 4] Enough time of the experiment time spent in analyzed loops (89.07%)
If the time spent in analyzed loops is less than 30%, standard loop optimizations will have a limited impact on
application performances.

  [4 / 4] Threads activity is good
On average, more than 98.10% of observed threads are actually active 

  [4 / 4] CPU activity is good
CPU cores are active 99.94% of time

  [4 / 4] Loop profile is not flat
At least one loop coverage is greater than 4% (7.09%), representing an hotspot for the application

  [4 / 4] Enough time of the experiment time spent in analyzed innermost loops (88.78%)
If the time spent in analyzed innermost loops is less than 15%, standard innermost loop optimizations such as
vectorisation will have a limited impact on application performances.

  [4 / 4] Affinity is good (99.84%)
Threads are not migrating to CPU cores: probably successfully pinned

  [3 / 3] Less than 10% (0.00%) is spend in BLAS1 operations
It could be more efficient to inline by hand BLAS1 operations

  [3 / 3] Functions mostly use all threads
Functions running on a reduced number of threads (typically sequential code) cover less than 10% of application
walltime (3.79%)

  [3 / 3] Cumulative Outermost/In between loops coverage (0.28%) lower than cumulative innermost loop coverage (88.78%)
Having cumulative Outermost/In between loops coverage greater than cumulative innermost loop coverage will make loop
optimization more complex

  [2 / 2] Less than 10% (0.00%) is spend in BLAS2 operations
BLAS2 calls usually could make a poor cache usage and could benefit from inlining.

  [2 / 2] Less than 10% (0.00%) is spend in Libm/SVML (special functions)



+--------------------------------------------------------------------------------------------------------------------+
+                                               2.3  -  LOOPS OVERVIEW                                               +
+--------------------------------------------------------------------------------------------------------------------+

  Top 5 loops:
   + exec - 162 :
     analysis: Execution Time: 7 % - Vectorization Ratio: 81.38 % - Vector Length Use: 21.38 %
     Loop Computation Issues: 16
        [12] [SA] Presence of expensive FP instructions - Perform hoisting, change algorithm, use SVML or proper
            numerical library or perform value profiling (count the number of distinct input values). There are 3
            issues (= instructions) costing 4 points each.
        [4] [SA] Less than 10% of the FP ADD/SUB/MUL arithmetic operations are performed using FMA - Reorganize
            arithmetic expressions to exhibit potential for FMA. This issue costs 4 points.
     Control Flow Issues: 36
        [36] [SA] Too many paths (32 paths) - Simplify control structure. There are 32 issues ( = paths) costing 1
            point each with a malus of 4 points.
     Data Access Issues: 35
        [0] [SA] Inefficient vectorization: more than 10% of the vector loads instructions are unaligned - When
            allocating arrays, don’t forget to align them. There are 0 issues ( = arrays) costing 2 points each
        [33] [SA] Presence of special instructions executing on a single port (INSERT/EXTRACT, BLEND/MERGE,
            SHUFFLE/PERM) - Simplify data access and try to get stride 1 access. There are 33 issues (= instructions)
            costing 1 point each.
        [2] [SA] More than 20% of the loads are accessing the stack - Perform loop splitting to decrease pressure on
            registers. This issue costs 2 points.
     Vectorization Roadblocks: 36
        [36] [SA] Too many paths (32 paths) - Simplify control structure. There are 32 issues ( = paths) costing 1
            point each with a malus of 4 points.
     Inefficient Vectorization: 33
        [33] [SA] Presence of special instructions executing on a single port (INSERT/EXTRACT, BLEND/MERGE,
            SHUFFLE/PERM) - Simplify data access and try to get stride 1 access. There are 33 issues (= instructions)
            costing 1 point each.

   + exec - 170 :
     analysis: Execution Time: 6 % - Vectorization Ratio: 79.91 % - Vector Length Use: 21.42 %
     Loop Computation Issues: 16
        [12] [SA] Presence of expensive FP instructions - Perform hoisting, change algorithm, use SVML or proper
            numerical library or perform value profiling (count the number of distinct input values). There are 3
            issues (= instructions) costing 4 points each.
        [4] [SA] Less than 10% of the FP ADD/SUB/MUL arithmetic operations are performed using FMA - Reorganize
            arithmetic expressions to exhibit potential for FMA. This issue costs 4 points.
     Control Flow Issues: 36
        [36] [SA] Too many paths (32 paths) - Simplify control structure. There are 32 issues ( = paths) costing 1
            point each with a malus of 4 points.
     Data Access Issues: 35
        [0] [SA] Inefficient vectorization: more than 10% of the vector loads instructions are unaligned - When
            allocating arrays, don’t forget to align them. There are 0 issues ( = arrays) costing 2 points each
        [33] [SA] Presence of special instructions executing on a single port (INSERT/EXTRACT, BLEND/MERGE,
            SHUFFLE/PERM) - Simplify data access and try to get stride 1 access. There are 33 issues (= instructions)
            costing 1 point each.
        [2] [SA] More than 20% of the loads are accessing the stack - Perform loop splitting to decrease pressure on
            registers. This issue costs 2 points.
     Vectorization Roadblocks: 36
        [36] [SA] Too many paths (32 paths) - Simplify control structure. There are 32 issues ( = paths) costing 1
            point each with a malus of 4 points.
     Inefficient Vectorization: 33
        [33] [SA] Presence of special instructions executing on a single port (INSERT/EXTRACT, BLEND/MERGE,
            SHUFFLE/PERM) - Simplify data access and try to get stride 1 access. There are 33 issues (= instructions)
            costing 1 point each.

   + exec - 298 :
     analysis: Execution Time: 5 % - Vectorization Ratio: 74.18 % - Vector Length Use: 20.30 %
     Loop Computation Issues: 12
        [8] [SA] Presence of expensive FP instructions - Perform hoisting, change algorithm, use SVML or proper
            numerical library or perform value profiling (count the number of distinct input values). There are 2
            issues (= instructions) costing 4 points each.
        [4] [SA] Less than 10% of the FP ADD/SUB/MUL arithmetic operations are performed using FMA - Reorganize
            arithmetic expressions to exhibit potential for FMA. This issue costs 4 points.
     Data Access Issues: 190
        [108] [SA] Presence of constant non unit stride data access - Use array restructuring, perform loop interchange
            or use gather instructions to lower a bit the cost. There are 54 issues ( = data accesses) costing 2
            point each.
        [0] [SA] Inefficient vectorization: more than 10% of the vector loads instructions are unaligned - When
            allocating arrays, don’t forget to align them. There are 0 issues ( = arrays) costing 2 points each
        [80] [SA] Presence of special instructions executing on a single port (INSERT/EXTRACT, BLEND/MERGE,
            SHUFFLE/PERM) - Simplify data access and try to get stride 1 access. There are 80 issues (= instructions)
            costing 1 point each.
        [2] [SA] More than 20% of the loads are accessing the stack - Perform loop splitting to decrease pressure on
            registers. This issue costs 2 points.
     Vectorization Roadblocks: 108
        [108] [SA] Presence of constant non unit stride data access - Use array restructuring, perform loop interchange
            or use gather instructions to lower a bit the cost. There are 54 issues ( = data accesses) costing 2
            point each.
     Inefficient Vectorization: 80
        [80] [SA] Presence of special instructions executing on a single port (INSERT/EXTRACT, BLEND/MERGE,
            SHUFFLE/PERM) - Simplify data access and try to get stride 1 access. There are 80 issues (= instructions)
            costing 1 point each.

   + exec - 142 :
     analysis: Execution Time: 5 % - Vectorization Ratio: 69.04 % - Vector Length Use: 19.81 %
     Loop Computation Issues: 8
        [4] [SA] Presence of expensive FP instructions - Perform hoisting, change algorithm, use SVML or proper
            numerical library or perform value profiling (count the number of distinct input values). There are 1
            issues (= instructions) costing 4 points each.
        [4] [SA] Less than 10% of the FP ADD/SUB/MUL arithmetic operations are performed using FMA - Reorganize
            arithmetic expressions to exhibit potential for FMA. This issue costs 4 points.
     Data Access Issues: 190
        [102] [SA] Presence of constant non unit stride data access - Use array restructuring, perform loop interchange
            or use gather instructions to lower a bit the cost. There are 51 issues ( = data accesses) costing 2
            point each.
        [0] [SA] Inefficient vectorization: more than 10% of the vector loads instructions are unaligned - When
            allocating arrays, don’t forget to align them. There are 0 issues ( = arrays) costing 2 points each
        [86] [SA] Presence of special instructions executing on a single port (INSERT/EXTRACT, BLEND/MERGE,
            SHUFFLE/PERM) - Simplify data access and try to get stride 1 access. There are 86 issues (= instructions)
            costing 1 point each.
        [2] [SA] More than 20% of the loads are accessing the stack - Perform loop splitting to decrease pressure on
            registers. This issue costs 2 points.
     Vectorization Roadblocks: 102
        [102] [SA] Presence of constant non unit stride data access - Use array restructuring, perform loop interchange
            or use gather instructions to lower a bit the cost. There are 51 issues ( = data accesses) costing 2
            point each.
     Inefficient Vectorization: 86
        [86] [SA] Presence of special instructions executing on a single port (INSERT/EXTRACT, BLEND/MERGE,
            SHUFFLE/PERM) - Simplify data access and try to get stride 1 access. There are 86 issues (= instructions)
            costing 1 point each.

   + exec - 201 :
     analysis: Execution Time: 4 % - Vectorization Ratio: 78.91 % - Vector Length Use: 21.23 %
     Loop Computation Issues: 32
        [28] [SA] Presence of expensive FP instructions - Perform hoisting, change algorithm, use SVML or proper
            numerical library or perform value profiling (count the number of distinct input values). There are 7
            issues (= instructions) costing 4 points each.
        [4] [SA] Less than 10% of the FP ADD/SUB/MUL arithmetic operations are performed using FMA - Reorganize
            arithmetic expressions to exhibit potential for FMA. This issue costs 4 points.
     Data Access Issues: 132
        [72] [SA] Presence of constant non unit stride data access - Use array restructuring, perform loop interchange
            or use gather instructions to lower a bit the cost. There are 36 issues ( = data accesses) costing 2
            point each.
        [0] [SA] Inefficient vectorization: more than 10% of the vector loads instructions are unaligned - When
            allocating arrays, don’t forget to align them. There are 0 issues ( = arrays) costing 2 points each
        [58] [SA] Presence of special instructions executing on a single port (INSERT/EXTRACT, BLEND/MERGE,
            SHUFFLE/PERM) - Simplify data access and try to get stride 1 access. There are 58 issues (= instructions)
            costing 1 point each.
        [2] [SA] More than 20% of the loads are accessing the stack - Perform loop splitting to decrease pressure on
            registers. This issue costs 2 points.
     Vectorization Roadblocks: 72
        [72] [SA] Presence of constant non unit stride data access - Use array restructuring, perform loop interchange
            or use gather instructions to lower a bit the cost. There are 36 issues ( = data accesses) costing 2
            point each.
     Inefficient Vectorization: 58
        [58] [SA] Presence of special instructions executing on a single port (INSERT/EXTRACT, BLEND/MERGE,
            SHUFFLE/PERM) - Simplify data access and try to get stride 1 access. There are 58 issues (= instructions)
            costing 1 point each.



+====================================================================================================================+
+                                                 3  -  APPLICATION                                                  +
+====================================================================================================================+


+--------------------------------------------------------------------------------------------------------------------+
+                                               3.1  -  Categorization                                               +
+--------------------------------------------------------------------------------------------------------------------+

   Category | IO     | Exe    | System  | Others  | Memory | String | MPI   | TBB   | OMP   | Pthread | Math  |
  ----------+--------+--------+---------+---------+--------+--------+-------+-------+-------+---------+-------+
   Time (%) | 0.00   | 89.07  | 0.25    | 0.01    | 0.00   | 0.00   | 0.03  | 0.00  | 4.15  | 0.00    | 6.50  |




+--------------------------------------------------------------------------------------------------------------------+
+                                          3.2  -  Function Based Profiling                                          +
+--------------------------------------------------------------------------------------------------------------------+

   Buckets                    | Nb Functions              | Coverage                  | Cumulated Coverage        |
  ----------------------------+---------------------------+---------------------------+---------------------------+
   > 8%                       | 0                         | 0.00                      | 0.00                      |
   4% to 8%                   | 10                        | 51.09                     | 51.09                     |
   2% to 4%                   | 9                         | 28.08                     | 79.17                     |
   1% to 2%                   | 11                        | 16.93                     | 96.09                     |
   0.5% to 1%                 | 4                         | 3.14                      | 99.23                     |
   0.25% to 0.5%              | 0                         | 0.00                      | 99.23                     |
   0.125% to 0.25%            | 1                         | 0.22                      | 99.45                     |
   < 0.125%                   | 165                       | 0.55                      | 100.00                    |




+--------------------------------------------------------------------------------------------------------------------+
+                                            3.3  -  Loop Based Profiling                                            +
+--------------------------------------------------------------------------------------------------------------------+

   Buckets                    | Nb Loops                  | Coverage                  | Cumulated Coverage        |
  ----------------------------+---------------------------+---------------------------+---------------------------+
   > 8%                       | 0                         | 0.00                      | 0.00                      |
   4% to 8%                   | 8                         | 41.98                     | 41.98                     |
   2% to 4%                   | 9                         | 28.07                     | 70.05                     |
   1% to 2%                   | 10                        | 15.56                     | 85.61                     |
   0.5% to 1%                 | 3                         | 2.64                      | 88.25                     |
   0.25% to 0.5%              | 0                         | 0.00                      | 88.25                     |
   0.125% to 0.25%            | 1                         | 0.25                      | 88.50                     |
   < 0.125%                   | 69                        | 0.29                      | 88.78                     |


+====================================================================================================================+
+                                                  4  -  FUNCTIONS                                                   +
+====================================================================================================================+


+--------------------------------------------------------------------------------------------------------------------+
+                                              4.1  -  Top 10 Functions                                              +
+--------------------------------------------------------------------------------------------------------------------+

   Function                                               | Module              | Coverage (%)   | Time (s)       |
  --------------------------------------------------------+---------------------+----------------+----------------+
   advec_mom_kernel(int, int, int, int, clover::Buffer... | exec                | 7.09           | 3.37           |
   advec_mom_kernel(int, int, int, int, clover::Buffer... | exec                | 6.21           | 2.95           |
   PdV_kernel(bool, int, int, int, int, double, clover... | exec                | 5.41           | 2.57           |
   accelerate_kernel(int, int, int, int, double, clove... | exec                | 5.24           | 2.49           |
   __svml_u64div2_e9                                      | exec                | 5.06           | 2.41           |
   calc_dt_kernel(int, int, int, int, double, double, ... | exec                | 4.56           | 2.17           |
   advec_cell_kernel(int, int, int, int, int, int, clo... | exec                | 4.54           | 2.16           |
   viscosity_kernel(int, int, int, int, clover::Buffer... | exec                | 4.52           | 2.15           |
   PdV_kernel(bool, int, int, int, int, double, clover... | exec                | 4.40           | 2.09           |
   kmp_flag_64<false, true>::wait(kmp_info*, int, void*)  | libiomp5.so         | 4.05           | 1.93           |


+====================================================================================================================+
+                                                    5  -  LOOPS                                                     +
+====================================================================================================================+


+--------------------------------------------------------------------------------------------------------------------+
+                                                5.1  -  Top 10 Loops                                                +
+--------------------------------------------------------------------------------------------------------------------+

   Loop Id        | Module              | Source Location                                        | Coverage (%)   |
  ----------------+---------------------+--------------------------------------------------------+----------------+
   162            | exec                | context.h:46-46,context.h:69-69,advec_mom.cpp:180-211  | 7.09           |
   170            | exec                | context.h:46-46,context.h:69-69,advec_mom.cpp:108-139  | 6.21           |
   298            | exec                | context.h:69-69,PdV.cpp:70-84                          | 5.41           |
   142            | exec                | accelerate.cpp:41-54,context.h:69-69                   | 5.24           |
   201            | exec                | context.h:46-46,context.h:69-69,calc_dt.cpp:50-76      | 4.56           |
   146            | exec                | context.h:46-46,context.h:69-69,advec_cell.cpp:158-202 | 4.54           |
   942            | exec                | context.h:46-46,context.h:69-69,viscosity.cpp:36-66    | 4.52           |
   300            | exec                | context.h:69-69,PdV.cpp:49-64                          | 4.40           |
   238            | exec                | ideal_gas.cpp:38-46,context.h:69-69                    | 3.94           |
   154            | exec                | context.h:46-46,context.h:69-69,advec_cell.cpp:66-110  | 3.91           |





+====================================================================================================================+
+                                                     6  -  CQA                                                      +
+====================================================================================================================+


+--------------------------------------------------------------------------------------------------------------------+
+                                                   6.1  -  Loops                                                    +
+--------------------------------------------------------------------------------------------------------------------+





      6.1.1  -  Loop 162 from exec
  ==========================================================================================================

The loop is defined in:
 - /beegfs/hackathon/users/eoseret/qaas_runs_test/178-661-4073/intel/CloverLeaf2.0-CXX/build/CloverLeaf2.0-CXX/src/omp/context.h:46,69
 - /beegfs/hackathon/users/eoseret/qaas_runs_test/178-661-4073/intel/CloverLeaf2.0-CXX/build/CloverLeaf2.0-CXX/src/omp/advec_mom.cpp:180-211


It is main loop of related source loop which is unrolled by 2 (including vectorization).
Warnings:
 - Ignoring paths for analysis
 - Too many paths. Rerun with max-paths=32
 - RecMII not computed since number of paths is unknown or > max_paths
 - Streams not analyzed since number of paths is unknown or > max_paths

Try to simplify control and/or increase the maximum number of paths per function/loop through the 'max-paths-nb' option.

This loop has 32 execution paths.

The presence of multiple execution paths is typically the main/first bottleneck.
Try to simplify control inside loop: ideally, try to remove all conditional expressions, for example by (if applicable):
 - hoisting them (moving them outside the loop)
 - turning them into conditional moves, MIN or MAX


Ex: if (x<0) x=0 => x = (x<0 ? 0 : x) (or MAX(0,x) after defining the corresponding macro)


      6.1.1.1  -  Path 1
  ----------------------------------------------------------------------------------------------------------

Warnings:
The number of fused uops of the instruction [PCMPEQD	%XMM2,%XMM2] is unknown
1% of peak computational performance is used (0.53 out of 32.00 FLOP per cycle (GFLOPS @ 1GHz))

      6.1.1.1.1  -  Vectorization
  ----------------------------------------------------------------------------------------------------------

Your loop is partially vectorized.
Only 21% of vector register length is used (average across all SSE/AVX instructions).
By fully vectorizing your loop, you can lower the cost of an iteration from 64.50 to 22.00 cycles (2.93x speedup).

Details
81% of SSE/AVX instructions are used in vector version (process two or more data elements in vector registers):
 - 54% of SSE/AVX loads are used in vector version.
 - 0% of SSE/AVX stores are used in vector version.
 - 60% of SSE/AVX divide and square root instructions are used in vector version.
 - 81% of SSE/AVX instructions that are not load, store, addition, subtraction nor multiply instructions are used in vector version.
Since your execution units are vector units, only a fully vectorized loop can use their full power.


Workaround
 - Pass to your compiler a micro-architecture specialization option:
  * use march=native
 - Use vector aligned instructions:
  1) align your arrays on 64 bytes boundaries: replace { void *p = malloc (size); } with { void *p; posix_memalign (&p, 64, size); }.
  2) inform your compiler that your arrays are vector aligned: if array 'foo' is 64 bytes-aligned, define a pointer 'p_foo' as __builtin_assume_aligned (foo, 64) and use it instead of 'foo' in the loop.



      6.1.1.1.2  -  Execution units bottlenecks
  ----------------------------------------------------------------------------------------------------------

Performance is limited by execution of FP multiply or FMA (fused multiply-add) operations (the FP multiply/FMA unit is a bottleneck).


Workaround
Reduce the number of FP multiply/FMA instructions




      6.1.1.1.3  -  FMA
  ----------------------------------------------------------------------------------------------------------

Presence of both ADD/SUB and MUL operations.

Workaround
 - Pass to your compiler a micro-architecture specialization option:
  * use march=native
 - Try to change order in which elements are evaluated (using parentheses) in arithmetic expressions containing both ADD/SUB and MUL operations to enable your compiler to generate FMA instructions wherever possible.
For instance a + b*c is a valid FMA (MUL then ADD).
However (a+b)* c cannot be translated into an FMA (ADD then MUL).





      6.1.1.1.4  -  Complex instructions
  ----------------------------------------------------------------------------------------------------------

Detected COMPLEX INSTRUCTIONS.


Details
These instructions generate more than one micro-operation and only one of them can be decoded during a cycle and the extra micro-operations increase pressure on execution units.
 - DIV: 2 occurrences<<list_path_1_complex_1>>



      6.1.1.1.5  -  Vector unaligned load/store instructions
  ----------------------------------------------------------------------------------------------------------

Detected 8 suboptimal vector unaligned load/store instructions.


Details
 - MOVHPD: 8 occurrences<<list_path_1_vec_align_1>>


Workaround
 - Pass to your compiler a micro-architecture specialization option:
  * use march=native
 - Use vector aligned instructions:
  1) align your arrays on 64 bytes boundaries: replace { void *p = malloc (size); } with { void *p; posix_memalign (&p, 64, size); }.
  2) inform your compiler that your arrays are vector aligned: if array 'foo' is 64 bytes-aligned, define a pointer 'p_foo' as __builtin_assume_aligned (foo, 64) and use it instead of 'foo' in the loop.
<<image_vec_align>>


      6.1.1.1.6  -  Type of elements and instruction set
  ----------------------------------------------------------------------------------------------------------

24 SSE or AVX instructions are processing arithmetic or math operations on double precision FP elements in vector mode (two at a time).



      6.1.1.1.7  -  Matching between your loop (in the source code) and the binary loop
  ----------------------------------------------------------------------------------------------------------

The binary loop is composed of 34 FP arithmetical operations:
 - 14: addition or subtraction
 - 14: multiply
 - 6: divide
The binary loop is loading 424 bytes (53 double precision FP elements).
The binary loop is storing 16 bytes (2 double precision FP elements).


      6.1.1.1.8  -  Arithmetic intensity
  ----------------------------------------------------------------------------------------------------------

Arithmetic intensity is 0.08 FP operations per loaded or stored byte.







      6.1.2  -  Loop 170 from exec
  ==========================================================================================================

The loop is defined in:
 - /beegfs/hackathon/users/eoseret/qaas_runs_test/178-661-4073/intel/CloverLeaf2.0-CXX/build/CloverLeaf2.0-CXX/src/omp/context.h:46,69
 - /beegfs/hackathon/users/eoseret/qaas_runs_test/178-661-4073/intel/CloverLeaf2.0-CXX/build/CloverLeaf2.0-CXX/src/omp/advec_mom.cpp:108-139


It is main loop of related source loop which is unrolled by 2 (including vectorization).
Warnings:
 - Ignoring paths for analysis
 - Too many paths. Rerun with max-paths=32
 - RecMII not computed since number of paths is unknown or > max_paths
 - Streams not analyzed since number of paths is unknown or > max_paths

Try to simplify control and/or increase the maximum number of paths per function/loop through the 'max-paths-nb' option.

This loop has 32 execution paths.

The presence of multiple execution paths is typically the main/first bottleneck.
Try to simplify control inside loop: ideally, try to remove all conditional expressions, for example by (if applicable):
 - hoisting them (moving them outside the loop)
 - turning them into conditional moves, MIN or MAX


Ex: if (x<0) x=0 => x = (x<0 ? 0 : x) (or MAX(0,x) after defining the corresponding macro)


      6.1.2.1  -  Path 1
  ----------------------------------------------------------------------------------------------------------

Warnings:
The number of fused uops of the instruction [PCMPEQD	%XMM2,%XMM2] is unknown
1% of peak computational performance is used (0.57 out of 32.00 FLOP per cycle (GFLOPS @ 1GHz))

      6.1.2.1.1  -  Vectorization
  ----------------------------------------------------------------------------------------------------------

Your loop is partially vectorized.
Only 21% of vector register length is used (average across all SSE/AVX instructions).
By fully vectorizing your loop, you can lower the cost of an iteration from 59.83 to 22.00 cycles (2.72x speedup).

Details
79% of SSE/AVX instructions are used in vector version (process two or more data elements in vector registers):
 - 54% of SSE/AVX loads are used in vector version.
 - 0% of SSE/AVX stores are used in vector version.
 - 60% of SSE/AVX divide and square root instructions are used in vector version.
 - 80% of SSE/AVX instructions that are not load, store, addition, subtraction nor multiply instructions are used in vector version.
Since your execution units are vector units, only a fully vectorized loop can use their full power.


Workaround
 - Try another compiler or update/tune your current one:
  * recompile with ffast-math (included in Ofast) to extend loop vectorization to FP reductions.
 - Remove inter-iterations dependences from your loop and make it unit-stride:
  * If your arrays have 2 or more dimensions, check whether elements are accessed contiguously and, otherwise, try to permute loops accordingly:
C storage order is row-major: for(i) for(j) a[j][i] = b[j][i]; (slow, non stride 1) => for(i) for(j) a[i][j] = b[i][j]; (fast, stride 1)<<image_row_maj>>
  * If your loop streams arrays of structures (AoS), try to use structures of arrays instead (SoA):
for(i) a[i].x = b[i].x; (slow, non stride 1) => for(i) a.x[i] = b.x[i]; (fast, stride 1)



      6.1.2.1.2  -  Execution units bottlenecks
  ----------------------------------------------------------------------------------------------------------

Performance is limited by execution of FP multiply or FMA (fused multiply-add) operations (the FP multiply/FMA unit is a bottleneck).


Workaround
Reduce the number of FP multiply/FMA instructions




      6.1.2.1.3  -  FMA
  ----------------------------------------------------------------------------------------------------------

Presence of both ADD/SUB and MUL operations.

Workaround
 - Pass to your compiler a micro-architecture specialization option:
  * use march=native
 - Try to change order in which elements are evaluated (using parentheses) in arithmetic expressions containing both ADD/SUB and MUL operations to enable your compiler to generate FMA instructions wherever possible.
For instance a + b*c is a valid FMA (MUL then ADD).
However (a+b)* c cannot be translated into an FMA (ADD then MUL).





      6.1.2.1.4  -  Complex instructions
  ----------------------------------------------------------------------------------------------------------

Detected COMPLEX INSTRUCTIONS.


Details
These instructions generate more than one micro-operation and only one of them can be decoded during a cycle and the extra micro-operations increase pressure on execution units.
 - DIV: 2 occurrences<<list_path_1_complex_1>>



      6.1.2.1.5  -  Vector unaligned load/store instructions
  ----------------------------------------------------------------------------------------------------------

Detected 8 suboptimal vector unaligned load/store instructions.


Details
 - MOVHPD: 8 occurrences<<list_path_1_vec_align_1>>


Workaround
 - Pass to your compiler a micro-architecture specialization option:
  * use march=native
 - Use vector aligned instructions:
  1) align your arrays on 64 bytes boundaries: replace { void *p = malloc (size); } with { void *p; posix_memalign (&p, 64, size); }.
  2) inform your compiler that your arrays are vector aligned: if array 'foo' is 64 bytes-aligned, define a pointer 'p_foo' as __builtin_assume_aligned (foo, 64) and use it instead of 'foo' in the loop.
<<image_vec_align>>


      6.1.2.1.6  -  Type of elements and instruction set
  ----------------------------------------------------------------------------------------------------------

24 SSE or AVX instructions are processing arithmetic or math operations on double precision FP elements in vector mode (two at a time).



      6.1.2.1.7  -  Matching between your loop (in the source code) and the binary loop
  ----------------------------------------------------------------------------------------------------------

The binary loop is composed of 34 FP arithmetical operations:
 - 14: addition or subtraction
 - 14: multiply
 - 6: divide
The binary loop is loading 424 bytes (53 double precision FP elements).
The binary loop is storing 16 bytes (2 double precision FP elements).


      6.1.2.1.8  -  Arithmetic intensity
  ----------------------------------------------------------------------------------------------------------

Arithmetic intensity is 0.08 FP operations per loaded or stored byte.







      6.1.3  -  Loop 298 from exec
  ==========================================================================================================

The loop is defined in:
 - /beegfs/hackathon/users/eoseret/qaas_runs_test/178-661-4073/intel/CloverLeaf2.0-CXX/build/CloverLeaf2.0-CXX/src/omp/context.h:69
 - /beegfs/hackathon/users/eoseret/qaas_runs_test/178-661-4073/intel/CloverLeaf2.0-CXX/build/CloverLeaf2.0-CXX/src/omp/PdV.cpp:70-84


It is main loop of related source loop which is unrolled by 2 (including vectorization).

      6.1.3.1  -  Path 1
  ----------------------------------------------------------------------------------------------------------

1% of peak computational performance is used (0.43 out of 32.00 FLOP per cycle (GFLOPS @ 1GHz))

      6.1.3.1.1  -  Vectorization
  ----------------------------------------------------------------------------------------------------------

Your loop is partially vectorized.
Only 20% of vector register length is used (average across all SSE/AVX instructions).
By fully vectorizing your loop, you can lower the cost of an iteration from 128.83 to 25.75 cycles (5.00x speedup).

Details
74% of SSE/AVX instructions are used in vector version (process two or more data elements in vector registers):
 - 30% of SSE/AVX loads are used in vector version.
 - 46% of SSE/AVX stores are used in vector version.
 - 50% of SSE/AVX divide and square root instructions are used in vector version.
 - 74% of SSE/AVX instructions that are not load, store, addition, subtraction nor multiply instructions are used in vector version.
Since your execution units are vector units, only a fully vectorized loop can use their full power.


Workaround
 - Try another compiler or update/tune your current one:
  * recompile with ffast-math (included in Ofast) to extend loop vectorization to FP reductions.
 - Remove inter-iterations dependences from your loop and make it unit-stride:
  * If your arrays have 2 or more dimensions, check whether elements are accessed contiguously and, otherwise, try to permute loops accordingly:
C storage order is row-major: for(i) for(j) a[j][i] = b[j][i]; (slow, non stride 1) => for(i) for(j) a[i][j] = b[i][j]; (fast, stride 1)<<image_row_maj>>
  * If your loop streams arrays of structures (AoS), try to use structures of arrays instead (SoA):
for(i) a[i].x = b[i].x; (slow, non stride 1) => for(i) a.x[i] = b.x[i]; (fast, stride 1)



      6.1.3.1.2  -  Execution units bottlenecks
  ----------------------------------------------------------------------------------------------------------

Performance is limited by execution of FP multiply or FMA (fused multiply-add) operations (the FP multiply/FMA unit is a bottleneck).


Workaround
Reduce the number of FP multiply/FMA instructions




      6.1.3.1.3  -  FMA
  ----------------------------------------------------------------------------------------------------------

Presence of both ADD/SUB and MUL operations.

Workaround
 - Pass to your compiler a micro-architecture specialization option:
  * use march=native
 - Try to change order in which elements are evaluated (using parentheses) in arithmetic expressions containing both ADD/SUB and MUL operations to enable your compiler to generate FMA instructions wherever possible.
For instance a + b*c is a valid FMA (MUL then ADD).
However (a+b)* c cannot be translated into an FMA (ADD then MUL).





      6.1.3.1.4  -  Complex instructions
  ----------------------------------------------------------------------------------------------------------

Detected COMPLEX INSTRUCTIONS.


Details
These instructions generate more than one micro-operation and only one of them can be decoded during a cycle and the extra micro-operations increase pressure on execution units.
 - DIV: 2 occurrences<<list_path_1_complex_1>>



      6.1.3.1.5  -  Slow data structures access
  ----------------------------------------------------------------------------------------------------------

Detected data structures (typically arrays) that cannot be efficiently read/written

Details
 - Constant unknown stride: 54 occurrence(s)
Non-unit stride (uncontiguous) accesses are not efficiently using data caches


Workaround
 - Try to reorganize arrays of structures to structures of arrays
 - Consider to permute loops (see vectorization gain report)



      6.1.3.1.6  -  Vector unaligned load/store instructions
  ----------------------------------------------------------------------------------------------------------

Detected 28 suboptimal vector unaligned load/store instructions.


Details
 - MOVHPD: 28 occurrences<<list_path_1_vec_align_1>>


Workaround
 - Pass to your compiler a micro-architecture specialization option:
  * use march=native
 - Use vector aligned instructions:
  1) align your arrays on 64 bytes boundaries: replace { void *p = malloc (size); } with { void *p; posix_memalign (&p, 64, size); }.
  2) inform your compiler that your arrays are vector aligned: if array 'foo' is 64 bytes-aligned, define a pointer 'p_foo' as __builtin_assume_aligned (foo, 64) and use it instead of 'foo' in the loop.
<<image_vec_align>>


      6.1.3.1.7  -  Type of elements and instruction set
  ----------------------------------------------------------------------------------------------------------

28 SSE or AVX instructions are processing arithmetic or math operations on double precision FP elements in vector mode (two at a time).



      6.1.3.1.8  -  Matching between your loop (in the source code) and the binary loop
  ----------------------------------------------------------------------------------------------------------

The binary loop is composed of 56 FP arithmetical operations:
 - 36: addition or subtraction
 - 16: multiply
 - 4: divide
The binary loop is loading 1000 bytes (125 double precision FP elements).
The binary loop is storing 152 bytes (19 double precision FP elements).


      6.1.3.1.9  -  Arithmetic intensity
  ----------------------------------------------------------------------------------------------------------

Arithmetic intensity is 0.05 FP operations per loaded or stored byte.







      6.1.4  -  Loop 142 from exec
  ==========================================================================================================

The loop is defined in:
 - /beegfs/hackathon/users/eoseret/qaas_runs_test/178-661-4073/intel/CloverLeaf2.0-CXX/build/CloverLeaf2.0-CXX/src/omp/accelerate.cpp:41-54
 - /beegfs/hackathon/users/eoseret/qaas_runs_test/178-661-4073/intel/CloverLeaf2.0-CXX/build/CloverLeaf2.0-CXX/src/omp/context.h:69


It is main loop of related source loop which is unrolled by 2 (including vectorization).

      6.1.4.1  -  Path 1
  ----------------------------------------------------------------------------------------------------------

1% of peak computational performance is used (0.61 out of 32.00 FLOP per cycle (GFLOPS @ 1GHz))

      6.1.4.1.1  -  Vectorization
  ----------------------------------------------------------------------------------------------------------

Your loop is partially vectorized.
Only 19% of vector register length is used (average across all SSE/AVX instructions).
By fully vectorizing your loop, you can lower the cost of an iteration from 121.17 to 24.00 cycles (5.05x speedup).

Details
69% of SSE/AVX instructions are used in vector version (process two or more data elements in vector registers):
 - 19% of SSE/AVX loads are used in vector version.
 - 14% of SSE/AVX stores are used in vector version.
 - 33% of SSE/AVX divide and square root instructions are used in vector version.
 - 73% of SSE/AVX instructions that are not load, store, addition, subtraction nor multiply instructions are used in vector version.
Since your execution units are vector units, only a fully vectorized loop can use their full power.


Workaround
 - Try another compiler or update/tune your current one:
  * recompile with ffast-math (included in Ofast) to extend loop vectorization to FP reductions.
 - Remove inter-iterations dependences from your loop and make it unit-stride:
  * If your arrays have 2 or more dimensions, check whether elements are accessed contiguously and, otherwise, try to permute loops accordingly:
C storage order is row-major: for(i) for(j) a[j][i] = b[j][i]; (slow, non stride 1) => for(i) for(j) a[i][j] = b[i][j]; (fast, stride 1)<<image_row_maj>>
  * If your loop streams arrays of structures (AoS), try to use structures of arrays instead (SoA):
for(i) a[i].x = b[i].x; (slow, non stride 1) => for(i) a.x[i] = b.x[i]; (fast, stride 1)



      6.1.4.1.2  -  Execution units bottlenecks
  ----------------------------------------------------------------------------------------------------------

Performance is limited by execution of FP multiply or FMA (fused multiply-add) operations (the FP multiply/FMA unit is a bottleneck).


Workaround
Reduce the number of FP multiply/FMA instructions




      6.1.4.1.3  -  FMA
  ----------------------------------------------------------------------------------------------------------

Presence of both ADD/SUB and MUL operations.

Workaround
 - Pass to your compiler a micro-architecture specialization option:
  * use march=native
 - Try to change order in which elements are evaluated (using parentheses) in arithmetic expressions containing both ADD/SUB and MUL operations to enable your compiler to generate FMA instructions wherever possible.
For instance a + b*c is a valid FMA (MUL then ADD).
However (a+b)* c cannot be translated into an FMA (ADD then MUL).





      6.1.4.1.4  -  Complex instructions
  ----------------------------------------------------------------------------------------------------------

Detected COMPLEX INSTRUCTIONS.


Details
These instructions generate more than one micro-operation and only one of them can be decoded during a cycle and the extra micro-operations increase pressure on execution units.
 - DIV: 2 occurrences<<list_path_1_complex_1>>



      6.1.4.1.5  -  Slow data structures access
  ----------------------------------------------------------------------------------------------------------

Detected data structures (typically arrays) that cannot be efficiently read/written

Details
 - Constant unknown stride: 51 occurrence(s)
Non-unit stride (uncontiguous) accesses are not efficiently using data caches


Workaround
 - Try to reorganize arrays of structures to structures of arrays
 - Consider to permute loops (see vectorization gain report)



      6.1.4.1.6  -  Vector unaligned load/store instructions
  ----------------------------------------------------------------------------------------------------------

Detected 40 suboptimal vector unaligned load/store instructions.


Details
 - MOVHPD: 40 occurrences<<list_path_1_vec_align_1>>


Workaround
 - Pass to your compiler a micro-architecture specialization option:
  * use march=native
 - Use vector aligned instructions:
  1) align your arrays on 64 bytes boundaries: replace { void *p = malloc (size); } with { void *p; posix_memalign (&p, 64, size); }.
  2) inform your compiler that your arrays are vector aligned: if array 'foo' is 64 bytes-aligned, define a pointer 'p_foo' as __builtin_assume_aligned (foo, 64) and use it instead of 'foo' in the loop.
<<image_vec_align>>


      6.1.4.1.7  -  Type of elements and instruction set
  ----------------------------------------------------------------------------------------------------------

37 SSE or AVX instructions are processing arithmetic or math operations on double precision FP elements in vector mode (two at a time).



      6.1.4.1.8  -  Matching between your loop (in the source code) and the binary loop
  ----------------------------------------------------------------------------------------------------------

The binary loop is composed of 74 FP arithmetical operations:
 - 38: addition or subtraction
 - 34: multiply
 - 2: divide
The binary loop is loading 1032 bytes (129 double precision FP elements).
The binary loop is storing 128 bytes (16 double precision FP elements).


      6.1.4.1.9  -  Arithmetic intensity
  ----------------------------------------------------------------------------------------------------------

Arithmetic intensity is 0.06 FP operations per loaded or stored byte.







      6.1.5  -  Loop 201 from exec
  ==========================================================================================================

The loop is defined in:
 - /beegfs/hackathon/users/eoseret/qaas_runs_test/178-661-4073/intel/CloverLeaf2.0-CXX/build/CloverLeaf2.0-CXX/src/omp/context.h:46,69
 - /beegfs/hackathon/users/eoseret/qaas_runs_test/178-661-4073/intel/CloverLeaf2.0-CXX/build/CloverLeaf2.0-CXX/src/omp/calc_dt.cpp:50-76


It is main loop of related source loop which is unrolled by 2 (including vectorization).

      6.1.5.1  -  Path 1
  ----------------------------------------------------------------------------------------------------------

1% of peak computational performance is used (0.48 out of 32.00 FLOP per cycle (GFLOPS @ 1GHz))

      6.1.5.1.1  -  Vectorization
  ----------------------------------------------------------------------------------------------------------

Your loop is partially vectorized.
Only 21% of vector register length is used (average across all SSE/AVX instructions).
By fully vectorizing your loop, you can lower the cost of an iteration from 108.33 to 38.50 cycles (2.81x speedup).

Details
78% of SSE/AVX instructions are used in vector version (process two or more data elements in vector registers):
 - 46% of SSE/AVX loads are used in vector version.
 - 61% of SSE/AVX stores are used in vector version.
 - 77% of SSE/AVX divide and square root instructions are used in vector version.
 - 81% of SSE/AVX instructions that are not load, store, addition, subtraction nor multiply instructions are used in vector version.
Since your execution units are vector units, only a fully vectorized loop can use their full power.


Workaround
 - Try another compiler or update/tune your current one:
  * recompile with ffast-math (included in Ofast) to extend loop vectorization to FP reductions.
 - Remove inter-iterations dependences from your loop and make it unit-stride:
  * If your arrays have 2 or more dimensions, check whether elements are accessed contiguously and, otherwise, try to permute loops accordingly:
C storage order is row-major: for(i) for(j) a[j][i] = b[j][i]; (slow, non stride 1) => for(i) for(j) a[i][j] = b[i][j]; (fast, stride 1)<<image_row_maj>>
  * If your loop streams arrays of structures (AoS), try to use structures of arrays instead (SoA):
for(i) a[i].x = b[i].x; (slow, non stride 1) => for(i) a.x[i] = b.x[i]; (fast, stride 1)



      6.1.5.1.2  -  Execution units bottlenecks
  ----------------------------------------------------------------------------------------------------------

Performance is limited by execution of FP multiply or FMA (fused multiply-add) operations (the FP multiply/FMA unit is a bottleneck).

By removing all these bottlenecks, you can lower the cost of an iteration from 108.33 to 78.00 cycles (1.39x speedup).


Workaround
Reduce the number of FP multiply/FMA instructions




      6.1.5.1.3  -  FMA
  ----------------------------------------------------------------------------------------------------------

Presence of both ADD/SUB and MUL operations.

Workaround
 - Pass to your compiler a micro-architecture specialization option:
  * use march=native
 - Try to change order in which elements are evaluated (using parentheses) in arithmetic expressions containing both ADD/SUB and MUL operations to enable your compiler to generate FMA instructions wherever possible.
For instance a + b*c is a valid FMA (MUL then ADD).
However (a+b)* c cannot be translated into an FMA (ADD then MUL).





      6.1.5.1.4  -  Complex instructions
  ----------------------------------------------------------------------------------------------------------

Detected COMPLEX INSTRUCTIONS.


Details
These instructions generate more than one micro-operation and only one of them can be decoded during a cycle and the extra micro-operations increase pressure on execution units.
 - CMPPD: 1 occurrences<<list_path_1_complex_1>>
 - DIV: 2 occurrences<<list_path_1_complex_2>>



      6.1.5.1.5  -  Slow data structures access
  ----------------------------------------------------------------------------------------------------------

Detected data structures (typically arrays) that cannot be efficiently read/written

Details
 - Constant unknown stride: 36 occurrence(s)
Non-unit stride (uncontiguous) accesses are not efficiently using data caches


Workaround
 - Try to reorganize arrays of structures to structures of arrays
 - Consider to permute loops (see vectorization gain report)



      6.1.5.1.6  -  Vector unaligned load/store instructions
  ----------------------------------------------------------------------------------------------------------

Detected 18 suboptimal vector unaligned load/store instructions.


Details
 - MOVHPD: 18 occurrences<<list_path_1_vec_align_1>>


Workaround
 - Pass to your compiler a micro-architecture specialization option:
  * use march=native
 - Use vector aligned instructions:
  1) align your arrays on 64 bytes boundaries: replace { void *p = malloc (size); } with { void *p; posix_memalign (&p, 64, size); }.
  2) inform your compiler that your arrays are vector aligned: if array 'foo' is 64 bytes-aligned, define a pointer 'p_foo' as __builtin_assume_aligned (foo, 64) and use it instead of 'foo' in the loop.
<<image_vec_align>>


      6.1.5.1.7  -  Type of elements and instruction set
  ----------------------------------------------------------------------------------------------------------

47 SSE or AVX instructions are processing arithmetic or math operations on double precision FP elements in vector mode (two at a time).



      6.1.5.1.8  -  Matching between your loop (in the source code) and the binary loop
  ----------------------------------------------------------------------------------------------------------

The binary loop is composed of 52 FP arithmetical operations:
 - 20: addition or subtraction
 - 18: multiply
 - 12: divide
 - 2: square root
The binary loop is loading 1000 bytes (125 double precision FP elements).
The binary loop is storing 168 bytes (21 double precision FP elements).


      6.1.5.1.9  -  Arithmetic intensity
  ----------------------------------------------------------------------------------------------------------

Arithmetic intensity is 0.04 FP operations per loaded or stored byte.







      6.1.6  -  Loop 146 from exec
  ==========================================================================================================

The loop is defined in:
 - /beegfs/hackathon/users/eoseret/qaas_runs_test/178-661-4073/intel/CloverLeaf2.0-CXX/build/CloverLeaf2.0-CXX/src/omp/context.h:46,69
 - /beegfs/hackathon/users/eoseret/qaas_runs_test/178-661-4073/intel/CloverLeaf2.0-CXX/build/CloverLeaf2.0-CXX/src/omp/advec_cell.cpp:158-202


It is main loop of related source loop which is unrolled by 2 (including vectorization).
Warnings:
 - Ignoring paths for analysis
 - Too many paths. Rerun with max-paths=16
 - RecMII not computed since number of paths is unknown or > max_paths
 - Streams not analyzed since number of paths is unknown or > max_paths

Try to simplify control and/or increase the maximum number of paths per function/loop through the 'max-paths-nb' option.

This loop has 16 execution paths.

The presence of multiple execution paths is typically the main/first bottleneck.
Try to simplify control inside loop: ideally, try to remove all conditional expressions, for example by (if applicable):
 - hoisting them (moving them outside the loop)
 - turning them into conditional moves, MIN or MAX


Ex: if (x<0) x=0 => x = (x<0 ? 0 : x) (or MAX(0,x) after defining the corresponding macro)


      6.1.6.1  -  Path 1
  ----------------------------------------------------------------------------------------------------------

Warnings:
The number of fused uops of the instruction [PCMPEQD	%XMM4,%XMM4] is unknown
2% of peak computational performance is used (0.67 out of 32.00 FLOP per cycle (GFLOPS @ 1GHz))

      6.1.6.1.1  -  Vectorization
  ----------------------------------------------------------------------------------------------------------

Your loop is partially vectorized.
Only 21% of vector register length is used (average across all SSE/AVX instructions).
By fully vectorizing your loop, you can lower the cost of an iteration from 93.17 to 22.00 cycles (4.23x speedup).

Details
82% of SSE/AVX instructions are used in vector version (process two or more data elements in vector registers):
 - 46% of SSE/AVX loads are used in vector version.
 - 63% of SSE/AVX stores are used in vector version.
 - 60% of SSE/AVX divide and square root instructions are used in vector version.
 - 85% of SSE/AVX instructions that are not load, store, addition, subtraction nor multiply instructions are used in vector version.
Since your execution units are vector units, only a fully vectorized loop can use their full power.


Workaround
 - Pass to your compiler a micro-architecture specialization option:
  * use march=native
 - Use vector aligned instructions:
  1) align your arrays on 64 bytes boundaries: replace { void *p = malloc (size); } with { void *p; posix_memalign (&p, 64, size); }.
  2) inform your compiler that your arrays are vector aligned: if array 'foo' is 64 bytes-aligned, define a pointer 'p_foo' as __builtin_assume_aligned (foo, 64) and use it instead of 'foo' in the loop.



      6.1.6.1.2  -  Execution units bottlenecks
  ----------------------------------------------------------------------------------------------------------

Performance is limited by execution of FP multiply or FMA (fused multiply-add) operations (the FP multiply/FMA unit is a bottleneck).


Workaround
Reduce the number of FP multiply/FMA instructions




      6.1.6.1.3  -  FMA
  ----------------------------------------------------------------------------------------------------------

Presence of both ADD/SUB and MUL operations.

Workaround
 - Pass to your compiler a micro-architecture specialization option:
  * use march=native
 - Try to change order in which elements are evaluated (using parentheses) in arithmetic expressions containing both ADD/SUB and MUL operations to enable your compiler to generate FMA instructions wherever possible.
For instance a + b*c is a valid FMA (MUL then ADD).
However (a+b)* c cannot be translated into an FMA (ADD then MUL).





      6.1.6.1.4  -  Complex instructions
  ----------------------------------------------------------------------------------------------------------

Detected COMPLEX INSTRUCTIONS.


Details
These instructions generate more than one micro-operation and only one of them can be decoded during a cycle and the extra micro-operations increase pressure on execution units.
 - DIV: 2 occurrences<<list_path_1_complex_1>>



      6.1.6.1.5  -  Vector unaligned load/store instructions
  ----------------------------------------------------------------------------------------------------------

Detected 14 suboptimal vector unaligned load/store instructions.


Details
 - MOVHPD: 14 occurrences<<list_path_1_vec_align_1>>


Workaround
 - Pass to your compiler a micro-architecture specialization option:
  * use march=native
 - Use vector aligned instructions:
  1) align your arrays on 64 bytes boundaries: replace { void *p = malloc (size); } with { void *p; posix_memalign (&p, 64, size); }.
  2) inform your compiler that your arrays are vector aligned: if array 'foo' is 64 bytes-aligned, define a pointer 'p_foo' as __builtin_assume_aligned (foo, 64) and use it instead of 'foo' in the loop.
<<image_vec_align>>


      6.1.6.1.6  -  Type of elements and instruction set
  ----------------------------------------------------------------------------------------------------------

44 SSE or AVX instructions are processing arithmetic or math operations on double precision FP elements in vector mode (two at a time).



      6.1.6.1.7  -  Matching between your loop (in the source code) and the binary loop
  ----------------------------------------------------------------------------------------------------------

The binary loop is composed of 62 FP arithmetical operations:
 - 28: addition or subtraction
 - 28: multiply
 - 6: divide
The binary loop is loading 712 bytes (89 double precision FP elements).
The binary loop is storing 144 bytes (18 double precision FP elements).


      6.1.6.1.8  -  Arithmetic intensity
  ----------------------------------------------------------------------------------------------------------

Arithmetic intensity is 0.07 FP operations per loaded or stored byte.







      6.1.7  -  Loop 942 from exec
  ==========================================================================================================

The loop is defined in:
 - /beegfs/hackathon/users/eoseret/qaas_runs_test/178-661-4073/intel/CloverLeaf2.0-CXX/build/CloverLeaf2.0-CXX/src/omp/context.h:46,69
 - /beegfs/hackathon/users/eoseret/qaas_runs_test/178-661-4073/intel/CloverLeaf2.0-CXX/build/CloverLeaf2.0-CXX/src/omp/viscosity.cpp:36-66


The related source loop is not unrolled or unrolled with no peel/tail loop.
Warnings:
 - Ignoring paths for analysis
 - Too many paths. If you really need to analyze all of the 256 paths individually, rerun with max-paths=256
 - RecMII not computed since number of paths is unknown or > max_paths
 - Streams not analyzed since number of paths is unknown or > max_paths

Try to simplify control and/or increase the maximum number of paths per function/loop through the 'max-paths-nb' option.

This loop has 256 execution paths.

The presence of multiple execution paths is typically the main/first bottleneck.
Try to simplify control inside loop: ideally, try to remove all conditional expressions, for example by (if applicable):
 - hoisting them (moving them outside the loop)
 - turning them into conditional moves, MIN or MAX


Ex: if (x<0) x=0 => x = (x<0 ? 0 : x) (or MAX(0,x) after defining the corresponding macro)


      6.1.7.1  -  Path 1
  ----------------------------------------------------------------------------------------------------------

3% of peak computational performance is used (1.04 out of 32.00 FLOP per cycle (GFLOPS @ 1GHz))

      6.1.7.1.1  -  Vectorization
  ----------------------------------------------------------------------------------------------------------

Your loop is partially vectorized.
Only 21% of vector register length is used (average across all SSE/AVX instructions).
By fully vectorizing your loop, you can lower the cost of an iteration from 103.83 to 42.50 cycles (2.44x speedup).

Details
79% of SSE/AVX instructions are used in vector version (process two or more data elements in vector registers):
 - 48% of SSE/AVX loads are used in vector version.
 - 76% of SSE/AVX stores are used in vector version.
 - 80% of SSE/AVX divide and square root instructions are used in vector version.
 - 78% of SSE/AVX instructions that are not load, store, addition, subtraction nor multiply instructions are used in vector version.
Since your execution units are vector units, only a fully vectorized loop can use their full power.


Workaround
 - Try another compiler or update/tune your current one:
  * recompile with ffast-math (included in Ofast) to extend loop vectorization to FP reductions.
 - Remove inter-iterations dependences from your loop and make it unit-stride:
  * If your arrays have 2 or more dimensions, check whether elements are accessed contiguously and, otherwise, try to permute loops accordingly:
C storage order is row-major: for(i) for(j) a[j][i] = b[j][i]; (slow, non stride 1) => for(i) for(j) a[i][j] = b[i][j]; (fast, stride 1)<<image_row_maj>>
  * If your loop streams arrays of structures (AoS), try to use structures of arrays instead (SoA):
for(i) a[i].x = b[i].x; (slow, non stride 1) => for(i) a.x[i] = b.x[i]; (fast, stride 1)



      6.1.7.1.2  -  Execution units bottlenecks
  ----------------------------------------------------------------------------------------------------------

Performance is limited by execution of FP multiply or FMA (fused multiply-add) operations (the FP multiply/FMA unit is a bottleneck).


Workaround
Reduce the number of FP multiply/FMA instructions




      6.1.7.1.3  -  FMA
  ----------------------------------------------------------------------------------------------------------

Presence of both ADD/SUB and MUL operations.

Workaround
 - Pass to your compiler a micro-architecture specialization option:
  * use march=native
 - Try to change order in which elements are evaluated (using parentheses) in arithmetic expressions containing both ADD/SUB and MUL operations to enable your compiler to generate FMA instructions wherever possible.
For instance a + b*c is a valid FMA (MUL then ADD).
However (a+b)* c cannot be translated into an FMA (ADD then MUL).





      6.1.7.1.4  -  Complex instructions
  ----------------------------------------------------------------------------------------------------------

Detected COMPLEX INSTRUCTIONS.


Details
These instructions generate more than one micro-operation and only one of them can be decoded during a cycle and the extra micro-operations increase pressure on execution units.
 - CMPPD: 1 occurrences<<list_path_1_complex_1>>
 - DIV: 2 occurrences<<list_path_1_complex_2>>



      6.1.7.1.5  -  Vector unaligned load/store instructions
  ----------------------------------------------------------------------------------------------------------

Detected 18 suboptimal vector unaligned load/store instructions.


Details
 - MOVHPD: 16 occurrences<<list_path_1_vec_align_1>>
 - MOVUPD: 2 occurrences<<list_path_1_vec_align_2>>


Workaround
 - Pass to your compiler a micro-architecture specialization option:
  * use march=native
 - Use vector aligned instructions:
  1) align your arrays on 64 bytes boundaries: replace { void *p = malloc (size); } with { void *p; posix_memalign (&p, 64, size); }.
  2) inform your compiler that your arrays are vector aligned: if array 'foo' is 64 bytes-aligned, define a pointer 'p_foo' as __builtin_assume_aligned (foo, 64) and use it instead of 'foo' in the loop.
<<image_vec_align>>


      6.1.7.1.6  -  Type of elements and instruction set
  ----------------------------------------------------------------------------------------------------------

63 SSE or AVX instructions are processing arithmetic or math operations on double precision FP elements in vector mode (two at a time).



      6.1.7.1.7  -  Matching between your loop (in the source code) and the binary loop
  ----------------------------------------------------------------------------------------------------------

The binary loop is composed of 108 FP arithmetical operations:
 - 46: addition or subtraction
 - 46: multiply
 - 14: divide
 - 2: square root
The binary loop is loading 928 bytes (116 double precision FP elements).
The binary loop is storing 240 bytes (30 double precision FP elements).


      6.1.7.1.8  -  Arithmetic intensity
  ----------------------------------------------------------------------------------------------------------

Arithmetic intensity is 0.09 FP operations per loaded or stored byte.







      6.1.8  -  Loop 300 from exec
  ==========================================================================================================

The loop is defined in:
 - /beegfs/hackathon/users/eoseret/qaas_runs_test/178-661-4073/intel/CloverLeaf2.0-CXX/build/CloverLeaf2.0-CXX/src/omp/context.h:69
 - /beegfs/hackathon/users/eoseret/qaas_runs_test/178-661-4073/intel/CloverLeaf2.0-CXX/build/CloverLeaf2.0-CXX/src/omp/PdV.cpp:49-64


It is main loop of related source loop which is unrolled by 2 (including vectorization).

      6.1.8.1  -  Path 1
  ----------------------------------------------------------------------------------------------------------

1% of peak computational performance is used (0.41 out of 32.00 FLOP per cycle (GFLOPS @ 1GHz))

      6.1.8.1.1  -  Vectorization
  ----------------------------------------------------------------------------------------------------------

Your loop is partially vectorized.
Only 20% of vector register length is used (average across all SSE/AVX instructions).
By fully vectorizing your loop, you can lower the cost of an iteration from 98.17 to 19.54 cycles (5.02x speedup).

Details
74% of SSE/AVX instructions are used in vector version (process two or more data elements in vector registers):
 - 25% of SSE/AVX loads are used in vector version.
 - 0% of SSE/AVX stores are used in vector version.
 - 50% of SSE/AVX divide and square root instructions are used in vector version.
 - 74% of SSE/AVX instructions that are not load, store, addition, subtraction nor multiply instructions are used in vector version.
Since your execution units are vector units, only a fully vectorized loop can use their full power.


Workaround
 - Try another compiler or update/tune your current one:
  * recompile with ffast-math (included in Ofast) to extend loop vectorization to FP reductions.
 - Remove inter-iterations dependences from your loop and make it unit-stride:
  * If your arrays have 2 or more dimensions, check whether elements are accessed contiguously and, otherwise, try to permute loops accordingly:
C storage order is row-major: for(i) for(j) a[j][i] = b[j][i]; (slow, non stride 1) => for(i) for(j) a[i][j] = b[i][j]; (fast, stride 1)<<image_row_maj>>
  * If your loop streams arrays of structures (AoS), try to use structures of arrays instead (SoA):
for(i) a[i].x = b[i].x; (slow, non stride 1) => for(i) a.x[i] = b.x[i]; (fast, stride 1)



      6.1.8.1.2  -  Execution units bottlenecks
  ----------------------------------------------------------------------------------------------------------

Performance is limited by execution of FP multiply or FMA (fused multiply-add) operations (the FP multiply/FMA unit is a bottleneck).


Workaround
Reduce the number of FP multiply/FMA instructions




      6.1.8.1.3  -  FMA
  ----------------------------------------------------------------------------------------------------------

Presence of both ADD/SUB and MUL operations.

Workaround
 - Pass to your compiler a micro-architecture specialization option:
  * use march=native
 - Try to change order in which elements are evaluated (using parentheses) in arithmetic expressions containing both ADD/SUB and MUL operations to enable your compiler to generate FMA instructions wherever possible.
For instance a + b*c is a valid FMA (MUL then ADD).
However (a+b)* c cannot be translated into an FMA (ADD then MUL).





      6.1.8.1.4  -  Complex instructions
  ----------------------------------------------------------------------------------------------------------

Detected COMPLEX INSTRUCTIONS.


Details
These instructions generate more than one micro-operation and only one of them can be decoded during a cycle and the extra micro-operations increase pressure on execution units.
 - DIV: 2 occurrences<<list_path_1_complex_1>>



      6.1.8.1.5  -  Slow data structures access
  ----------------------------------------------------------------------------------------------------------

Detected data structures (typically arrays) that cannot be efficiently read/written

Details
 - Constant unknown stride: 38 occurrence(s)
Non-unit stride (uncontiguous) accesses are not efficiently using data caches


Workaround
 - Try to reorganize arrays of structures to structures of arrays
 - Consider to permute loops (see vectorization gain report)



      6.1.8.1.6  -  Vector unaligned load/store instructions
  ----------------------------------------------------------------------------------------------------------

Detected 20 suboptimal vector unaligned load/store instructions.


Details
 - MOVHPD: 20 occurrences<<list_path_1_vec_align_1>>


Workaround
 - Pass to your compiler a micro-architecture specialization option:
  * use march=native
 - Use vector aligned instructions:
  1) align your arrays on 64 bytes boundaries: replace { void *p = malloc (size); } with { void *p; posix_memalign (&p, 64, size); }.
  2) inform your compiler that your arrays are vector aligned: if array 'foo' is 64 bytes-aligned, define a pointer 'p_foo' as __builtin_assume_aligned (foo, 64) and use it instead of 'foo' in the loop.
<<image_vec_align>>


      6.1.8.1.7  -  Type of elements and instruction set
  ----------------------------------------------------------------------------------------------------------

20 SSE or AVX instructions are processing arithmetic or math operations on double precision FP elements in vector mode (two at a time).



      6.1.8.1.8  -  Matching between your loop (in the source code) and the binary loop
  ----------------------------------------------------------------------------------------------------------

The binary loop is composed of 40 FP arithmetical operations:
 - 20: addition or subtraction
 - 16: multiply
 - 4: divide
The binary loop is loading 656 bytes (82 double precision FP elements).
The binary loop is storing 40 bytes (5 double precision FP elements).


      6.1.8.1.9  -  Arithmetic intensity
  ----------------------------------------------------------------------------------------------------------

Arithmetic intensity is 0.06 FP operations per loaded or stored byte.







      6.1.9  -  Loop 238 from exec
  ==========================================================================================================

The loop is defined in:
 - /beegfs/hackathon/users/eoseret/qaas_runs_test/178-661-4073/intel/CloverLeaf2.0-CXX/build/CloverLeaf2.0-CXX/src/omp/ideal_gas.cpp:38-46
 - /beegfs/hackathon/users/eoseret/qaas_runs_test/178-661-4073/intel/CloverLeaf2.0-CXX/build/CloverLeaf2.0-CXX/src/omp/context.h:69


It is main loop of related source loop which is unrolled by 4 (including vectorization).

      6.1.9.1  -  Path 1
  ----------------------------------------------------------------------------------------------------------

Warnings:
Detected a function call instruction: ignoring called function instructions.
Rerun with --follow-calls=append to include them to analysis  or with --follow-calls=inline to simulate inlining.
1% of peak computational performance is used (0.59 out of 32.00 FLOP per cycle (GFLOPS @ 1GHz))

      6.1.9.1.1  -  Vectorization
  ----------------------------------------------------------------------------------------------------------

Your loop is partially vectorized.
Only 20% of vector register length is used (average across all SSE/AVX instructions).
By fully vectorizing your loop, you can lower the cost of an iteration from 54.00 to 17.00 cycles (3.18x speedup).

Details
81% of SSE/AVX instructions are used in vector version (process two or more data elements in vector registers):
 - 50% of SSE/AVX loads are used in vector version.
 - 0% of SSE/AVX stores are used in vector version.
 - 79% of SSE/AVX instructions that are not load, store, addition, subtraction nor multiply instructions are used in vector version.
Since your execution units are vector units, only a fully vectorized loop can use their full power.


Workaround
 - Pass to your compiler a micro-architecture specialization option:
  * use march=native
 - Use vector aligned instructions:
  1) align your arrays on 64 bytes boundaries: replace { void *p = malloc (size); } with { void *p; posix_memalign (&p, 64, size); }.
  2) inform your compiler that your arrays are vector aligned: if array 'foo' is 64 bytes-aligned, define a pointer 'p_foo' as __builtin_assume_aligned (foo, 64) and use it instead of 'foo' in the loop.



      6.1.9.1.2  -  Execution units bottlenecks
  ----------------------------------------------------------------------------------------------------------

Performance is limited by execution of FP multiply or FMA (fused multiply-add) operations (the FP multiply/FMA unit is a bottleneck).

By removing all these bottlenecks, you can lower the cost of an iteration from 54.00 to 37.17 cycles (1.45x speedup).


Workaround
Reduce the number of FP multiply/FMA instructions



No data for this section



      6.1.9.1.3  -  CALL instructions
  ----------------------------------------------------------------------------------------------------------

Detected function call instructions.


Details
Calling (and then returning from) a function prevents many compiler optimizations (like vectorization), breaks control flow (which reduces pipeline performance) and executes extra instructions to save/restore the registers used inside it, which is very expensive (dozens of cycles). Consider to inline small functions.
 - unknown: 2 occurrences<<list_path_1_call_1>>



      6.1.9.1.4  -  Slow data structures access
  ----------------------------------------------------------------------------------------------------------

Detected data structures (typically arrays) that cannot be efficiently read/written

Details
 - Constant unknown stride: 16 occurrence(s)
Non-unit stride (uncontiguous) accesses are not efficiently using data caches


Workaround
 - Try to reorganize arrays of structures to structures of arrays
 - Consider to permute loops (see vectorization gain report)



      6.1.9.1.5  -  Vector unaligned load/store instructions
  ----------------------------------------------------------------------------------------------------------

Detected 10 suboptimal vector unaligned load/store instructions.


Details
 - MOVHPD: 10 occurrences<<list_path_1_vec_align_1>>


Workaround
 - Pass to your compiler a micro-architecture specialization option:
  * use march=native
 - Use vector aligned instructions:
  1) align your arrays on 64 bytes boundaries: replace { void *p = malloc (size); } with { void *p; posix_memalign (&p, 64, size); }.
  2) inform your compiler that your arrays are vector aligned: if array 'foo' is 64 bytes-aligned, define a pointer 'p_foo' as __builtin_assume_aligned (foo, 64) and use it instead of 'foo' in the loop.
<<image_vec_align>>


      6.1.9.1.6  -  Type of elements and instruction set
  ----------------------------------------------------------------------------------------------------------

16 SSE or AVX instructions are processing arithmetic or math operations on double precision FP elements in vector mode (two at a time).



      6.1.9.1.7  -  Matching between your loop (in the source code) and the binary loop
  ----------------------------------------------------------------------------------------------------------

The binary loop is composed of 32 FP arithmetical operations:
 - 24: multiply
 - 4: divide
 - 4: square root
The binary loop is loading 296 bytes (37 double precision FP elements).
The binary loop is storing 64 bytes (8 double precision FP elements).


      6.1.9.1.8  -  Arithmetic intensity
  ----------------------------------------------------------------------------------------------------------

Arithmetic intensity is 0.09 FP operations per loaded or stored byte.







      6.1.10  -  Loop 154 from exec
  ==========================================================================================================

The loop is defined in:
 - /beegfs/hackathon/users/eoseret/qaas_runs_test/178-661-4073/intel/CloverLeaf2.0-CXX/build/CloverLeaf2.0-CXX/src/omp/context.h:46,69
 - /beegfs/hackathon/users/eoseret/qaas_runs_test/178-661-4073/intel/CloverLeaf2.0-CXX/build/CloverLeaf2.0-CXX/src/omp/advec_cell.cpp:66-110


It is main loop of related source loop which is unrolled by 2 (including vectorization).
Warnings:
 - Ignoring paths for analysis
 - Too many paths. Rerun with max-paths=16
 - RecMII not computed since number of paths is unknown or > max_paths
 - Streams not analyzed since number of paths is unknown or > max_paths

Try to simplify control and/or increase the maximum number of paths per function/loop through the 'max-paths-nb' option.

This loop has 16 execution paths.

The presence of multiple execution paths is typically the main/first bottleneck.
Try to simplify control inside loop: ideally, try to remove all conditional expressions, for example by (if applicable):
 - hoisting them (moving them outside the loop)
 - turning them into conditional moves, MIN or MAX


Ex: if (x<0) x=0 => x = (x<0 ? 0 : x) (or MAX(0,x) after defining the corresponding macro)


      6.1.10.1  -  Path 1
  ----------------------------------------------------------------------------------------------------------

Warnings:
The number of fused uops of the instruction [PCMPEQD	%XMM5,%XMM5] is unknown
2% of peak computational performance is used (0.73 out of 32.00 FLOP per cycle (GFLOPS @ 1GHz))

      6.1.10.1.1  -  Vectorization
  ----------------------------------------------------------------------------------------------------------

Your loop is partially vectorized.
Only 21% of vector register length is used (average across all SSE/AVX instructions).
By fully vectorizing your loop, you can lower the cost of an iteration from 85.50 to 22.00 cycles (3.89x speedup).

Details
79% of SSE/AVX instructions are used in vector version (process two or more data elements in vector registers):
 - 45% of SSE/AVX loads are used in vector version.
 - 33% of SSE/AVX stores are used in vector version.
 - 60% of SSE/AVX divide and square root instructions are used in vector version.
 - 82% of SSE/AVX instructions that are not load, store, addition, subtraction nor multiply instructions are used in vector version.
Since your execution units are vector units, only a fully vectorized loop can use their full power.


Workaround
 - Try another compiler or update/tune your current one:
  * recompile with ffast-math (included in Ofast) to extend loop vectorization to FP reductions.
 - Remove inter-iterations dependences from your loop and make it unit-stride:
  * If your arrays have 2 or more dimensions, check whether elements are accessed contiguously and, otherwise, try to permute loops accordingly:
C storage order is row-major: for(i) for(j) a[j][i] = b[j][i]; (slow, non stride 1) => for(i) for(j) a[i][j] = b[i][j]; (fast, stride 1)<<image_row_maj>>
  * If your loop streams arrays of structures (AoS), try to use structures of arrays instead (SoA):
for(i) a[i].x = b[i].x; (slow, non stride 1) => for(i) a.x[i] = b.x[i]; (fast, stride 1)



      6.1.10.1.2  -  Execution units bottlenecks
  ----------------------------------------------------------------------------------------------------------

Performance is limited by execution of FP multiply or FMA (fused multiply-add) operations (the FP multiply/FMA unit is a bottleneck).


Workaround
Reduce the number of FP multiply/FMA instructions




      6.1.10.1.3  -  FMA
  ----------------------------------------------------------------------------------------------------------

Presence of both ADD/SUB and MUL operations.

Workaround
 - Pass to your compiler a micro-architecture specialization option:
  * use march=native
 - Try to change order in which elements are evaluated (using parentheses) in arithmetic expressions containing both ADD/SUB and MUL operations to enable your compiler to generate FMA instructions wherever possible.
For instance a + b*c is a valid FMA (MUL then ADD).
However (a+b)* c cannot be translated into an FMA (ADD then MUL).





      6.1.10.1.4  -  Complex instructions
  ----------------------------------------------------------------------------------------------------------

Detected COMPLEX INSTRUCTIONS.


Details
These instructions generate more than one micro-operation and only one of them can be decoded during a cycle and the extra micro-operations increase pressure on execution units.
 - CMPPD: 1 occurrences<<list_path_1_complex_1>>
 - DIV: 2 occurrences<<list_path_1_complex_2>>



      6.1.10.1.5  -  Vector unaligned load/store instructions
  ----------------------------------------------------------------------------------------------------------

Detected 14 suboptimal vector unaligned load/store instructions.


Details
 - MOVHPD: 14 occurrences<<list_path_1_vec_align_1>>


Workaround
 - Pass to your compiler a micro-architecture specialization option:
  * use march=native
 - Use vector aligned instructions:
  1) align your arrays on 64 bytes boundaries: replace { void *p = malloc (size); } with { void *p; posix_memalign (&p, 64, size); }.
  2) inform your compiler that your arrays are vector aligned: if array 'foo' is 64 bytes-aligned, define a pointer 'p_foo' as __builtin_assume_aligned (foo, 64) and use it instead of 'foo' in the loop.
<<image_vec_align>>


      6.1.10.1.6  -  Type of elements and instruction set
  ----------------------------------------------------------------------------------------------------------

44 SSE or AVX instructions are processing arithmetic or math operations on double precision FP elements in vector mode (two at a time).



      6.1.10.1.7  -  Matching between your loop (in the source code) and the binary loop
  ----------------------------------------------------------------------------------------------------------

The binary loop is composed of 62 FP arithmetical operations:
 - 28: addition or subtraction
 - 28: multiply
 - 6: divide
The binary loop is loading 640 bytes (80 double precision FP elements).
The binary loop is storing 64 bytes (8 double precision FP elements).


      6.1.10.1.8  -  Arithmetic intensity
  ----------------------------------------------------------------------------------------------------------

Arithmetic intensity is 0.09 FP operations per loaded or stored byte.





[MAQAO] Info: STOP THE REPORT GENERATION
[MAQAO] Info: 
[MAQAO] Info: If your application produces files, they can be found in directory "/beegfs/hackathon/users/eoseret/qaas_runs_test/178-661-4073/intel/CloverLeaf2.0-CXX/run/oneview_runs/defaults/orig/oneview_run_1786614308"
[MAQAO] Info: 
