	*************************************************
	*                                               *
	*          ONE-View report generation           *
	*                                               *
	*************************************************

[MAQAO] Info: Experiment configuration summary is available adding -dbg=1 in command line

* [MAQAO] Warning: Experiment directory /beegfs/hackathon/users/eoseret/qaas_runs_test/178-661-4073/intel/CloverLeaf1.3-FC/run/oneview_runs/defaults/aocc/oneview_results_1786614486 already exists and is reused.
           It can be replaced using --replace in the command line.
[MAQAO] Info: 
[MAQAO] Info: START THE APPLICATION PROFILING
[MAQAO] Info: -> RUNNING THE PROFILER...
[MAQAO] Info:   LPROF has already been run
[MAQAO] Info: STOP THE APPLICATION PROFILING
[MAQAO] Info: 
[MAQAO] Info: START FUNCTIONS AND LOOPS ANALYSIS ...
[MAQAO] Info: STOP FUNCTIONS AND LOOPS ANALYSIS ...
[MAQAO] Info: 
[MAQAO] Info: START THE REPORT GENERATION
[MAQAO] Info: -> ONE-VIEW EXPERIMENT DIRECTORY: /beegfs/hackathon/users/eoseret/qaas_runs_test/178-661-4073/intel/CloverLeaf1.3-FC/run/oneview_runs/defaults/aocc/oneview_results_1786614486


+====================================================================================================================+
+                                                    1  -  GLOBAL                                                    +
+====================================================================================================================+


+--------------------------------------------------------------------------------------------------------------------+
+                                             1.1  -  Experiment Summary                                             +
+--------------------------------------------------------------------------------------------------------------------+

  Application:			/beegfs/hackathon/users/eoseret/qaas_runs_test/178-661-4073/intel/CloverLeaf1.3-FC/run/base_runs/defaults/aocc/exec
  Timestamp:			2026-08-13 11:48:06
  Universal Timestamp:		1786614486
  Experiment Type:		MPI; OpenMP; Throughput; 
  Machine:			isix06.benchmarkcenter.megware.com
  Architecture:			x86_64
  Micro Architecture:		GRANITE_RAPIDS
  Model Name:			Intel(R) Xeon(R) 6972P
  Cache Size:			491520 KB
  Number of Cores:		96
  OS Version:			Linux 5.14.0-687.31.1.el9_8.x86_64 #1 SMP PREEMPT_DYNAMIC Sat Aug 1 05:38:01 EDT 2026
  Compilation Options:		
		exec:  F90 AOCCAOCC_5.1.0-Build#1994 2025_12_23 &apos;+flang -I/beegfs/hackathon/users/eoseret/qaas_runs_test/178-661-4073/intel/CloverLeaf1.3-FC/build/CloverLeaf1.3-FC/CloverLeaf_ref/kernels -O3 -g -fno-omit-frame-pointer -fcf-protection=none -no-pie -fopenmp -c -o -I/cluster/hpcx/2.23/ompi5-aocc-mt/include -I/cluster/hpcx/2.23/ompi5-aocc-mt/lib&apos; 
  Number of processes observed:	6
  Number of threads observed:	192
  MAQAO version:		2026.1.0
  MAQAO build:			6d1be1d51c1e63266254997eb301734a7264775d::20260810-150026




+--------------------------------------------------------------------------------------------------------------------+
+                                               1.2  -  Global Metrics                                               +
+--------------------------------------------------------------------------------------------------------------------+

  Total Time:				38.73 s
  Max (Thread Active Time):		38.21 s
  Average Active Time:			37.91 s
  Activity Ratio:			99.9 %
  Average number of active threads:	187.902
  Affinity Stability:			99.8 %
  Time spent in analyzed loops:		95.9 %
  Time spent in analyzed innermost loops: 95.7 %
  Time spent in user code:		95.9 %
  Compilation Options Score:		66.67
  Array Access Efficiency:		97.1 %

   Potential Speedups
  ----------------------------------------------------
  Perfect Flow Complexity:		1.10
  Perfect OpenMP/MPI/Pthread/TBB:	1.03
  Perfect OpenMP/MPI/Pthread/TBB + Load Distribution:	1.05
  If No Scalar Integer:
      Potential Speedup:		1.01
      Nb Loops to get 80%:		3
  If FP Vectorized:
      Potential Speedup:		1.26
      Nb Loops to get 80%:		13
  If Fully Vectorized:
      Potential Speedup:		1.79
      Nb Loops to get 80%:		18
  If Only FP Arithmetic:
      Potential Speedup:		1.16
      Nb Loops to get 80%:		9




+--------------------------------------------------------------------------------------------------------------------+
+                                             1.3  -  Potential Speedups                                             +
+--------------------------------------------------------------------------------------------------------------------+

  If No Scalar Integer:
      Number of loops   | 1      | 9      | 19     | 28     | 38     | 
      Cumulated Speedup | 1.0051 | 1.0114 | 1.0114 | 1.0114 | 1.0114 | 
  Top 5 loops:
    exec - 99:	1.0051
    exec - 110:	1.0089
    exec - 66:	1.0113
    exec - 179:	1.0114
    exec - 58:	1.0114

  If FP Vectorized:
      Number of loops   | 1      | 9      | 19     | 28     | 38     | 
      Cumulated Speedup | 1.0275 | 1.1628 | 1.2442 | 1.2572 | 1.2572 | 
  Top 5 loops:
    exec - 47:	1.0275
    exec - 99:	1.0462
    exec - 110:	1.0646
    exec - 66:	1.0831
    exec - 170:	1.1005

  If Fully Vectorized:
      Number of loops   | 1      | 9      | 19     | 28     | 38     | 
      Cumulated Speedup | 1.0417 | 1.3082 | 1.6730 | 1.7851 | 1.7901 | 
  Top 5 loops:
    exec - 47:	1.0417
    exec - 191:	1.0772
    exec - 66:	1.1094
    exec - 55:	1.143
    exec - 102:	1.1741

  If Only FP Arithmetic:
      Number of loops   | 1      | 9      | 19     | 28     | 38     | 
      Cumulated Speedup | 1.0210 | 1.1318 | 1.1558 | 1.1560 | 1.1560 | 
  Top 5 loops:
    exec - 66:	1.021
    exec - 55:	1.0419
    exec - 335:	1.0569
    exec - 332:	1.0723
    exec - 350:	1.088



+====================================================================================================================+
+                                                   2  -  SUMMARY                                                    +
+====================================================================================================================+


+--------------------------------------------------------------------------------------------------------------------+
+                                             2.1  -  EXPERIMENT QUALITY                                             +
+--------------------------------------------------------------------------------------------------------------------+

  [4 / 4] Application profile is long enough (38.21 s)
To have good quality measurements, it is advised that the application profiling time is greater than 10 seconds.

  [3 / 3] Most of time spent in analyzed modules comes from functions with source/debug info
-g option gives access to debugging informations, such are source locations.

  [0 / 3] Compilation of some functions is not optimized for the target processor
Architecture specific options are needed to produce efficient code for a specific processor ( -march=(target) ).

  [3 / 3] Most of time spent in analyzed modules comes from functions with compilation options informations and
-fno-omit-frame-pointer is present
-fno-omit-frame-pointer improves the accuracy of callchains found during the application profiling.

  [3 / 3] Optimization level option is correctly used


  [3 / 3] Host configuration allows retrieval of all necessary metrics.


  [2 / 2] Application is correctly profiled ("Others" category represents 0.01 % of the execution time)
To have a representative profiling, it is advised that the category "Others" represents less than 20% of the execution
time in order to analyze as much as possible of the user code

  [1 / 1] Lstopo present. The Topology lstopo report will be generated.



+--------------------------------------------------------------------------------------------------------------------+
+                                                2.2  -  CODE QUALITY                                                +
+--------------------------------------------------------------------------------------------------------------------+

  [4 / 4] Enough time of the experiment time spent in analyzed loops (95.93%)
If the time spent in analyzed loops is less than 30%, standard loop optimizations will have a limited impact on
application performances.

  [4 / 4] Threads activity is good
On average, more than 97.87% of observed threads are actually active 

  [4 / 4] CPU activity is good
CPU cores are active 99.86% of time

  [4 / 4] Loop profile is not flat
At least one loop coverage is greater than 4% (6.45%), representing an hotspot for the application

  [4 / 4] Enough time of the experiment time spent in analyzed innermost loops (95.72%)
If the time spent in analyzed innermost loops is less than 15%, standard innermost loop optimizations such as
vectorisation will have a limited impact on application performances.

  [4 / 4] Affinity is good (99.84%)
Threads are not migrating to CPU cores: probably successfully pinned

  [3 / 3] Less than 10% (0.00%) is spend in BLAS1 operations
It could be more efficient to inline by hand BLAS1 operations

  [3 / 3] Functions mostly use all threads
Functions running on a reduced number of threads (typically sequential code) cover less than 10% of application
walltime (3.69%)

  [3 / 3] Cumulative Outermost/In between loops coverage (0.21%) lower than cumulative innermost loop coverage (95.72%)
Having cumulative Outermost/In between loops coverage greater than cumulative innermost loop coverage will make loop
optimization more complex

  [2 / 2] Less than 10% (0.00%) is spend in BLAS2 operations
BLAS2 calls usually could make a poor cache usage and could benefit from inlining.

  [2 / 2] Less than 10% (0.00%) is spend in Libm/SVML (special functions)



+--------------------------------------------------------------------------------------------------------------------+
+                                               2.3  -  LOOPS OVERVIEW                                               +
+--------------------------------------------------------------------------------------------------------------------+

  Top 5 loops:
   + exec - 38  :
     analysis: Execution Time: 6 % - Vectorization Ratio: 100.00 % - Vector Length Use: 25.00 %
     Loop Computation Issues: 20
        [16] [SA] Presence of expensive FP instructions - Perform hoisting, change algorithm, use SVML or proper
            numerical library or perform value profiling (count the number of distinct input values). There are 4
            issues (= instructions) costing 4 points each.
        [4] [SA] Less than 10% of the FP ADD/SUB/MUL arithmetic operations are performed using FMA - Reorganize
            arithmetic expressions to exhibit potential for FMA. This issue costs 4 points.
     Data Access Issues: 0
        [0] [SA] Inefficient vectorization: more than 10% of the vector loads instructions are unaligned - When
            allocating arrays, don’t forget to align them. There are 0 issues ( = arrays) costing 2 points each

   + exec - 245 :
     analysis: Execution Time: 5 % - Vectorization Ratio: 100.00 % - Vector Length Use: 25.00 %
     Loop Computation Issues: 12
        [8] [SA] Presence of expensive FP instructions - Perform hoisting, change algorithm, use SVML or proper
            numerical library or perform value profiling (count the number of distinct input values). There are 2
            issues (= instructions) costing 4 points each.
        [4] [SA] Less than 10% of the FP ADD/SUB/MUL arithmetic operations are performed using FMA - Reorganize
            arithmetic expressions to exhibit potential for FMA. This issue costs 4 points.
     Data Access Issues: 0
        [0] [SA] Inefficient vectorization: more than 10% of the vector loads instructions are unaligned - When
            allocating arrays, don’t forget to align them. There are 0 issues ( = arrays) costing 2 points each

   + exec - 35  :
     analysis: Execution Time: 5 % - Vectorization Ratio: 100.00 % - Vector Length Use: 25.00 %
     Loop Computation Issues: 20
        [16] [SA] Presence of expensive FP instructions - Perform hoisting, change algorithm, use SVML or proper
            numerical library or perform value profiling (count the number of distinct input values). There are 4
            issues (= instructions) costing 4 points each.
        [4] [SA] Less than 10% of the FP ADD/SUB/MUL arithmetic operations are performed using FMA - Reorganize
            arithmetic expressions to exhibit potential for FMA. This issue costs 4 points.
     Data Access Issues: 0
        [0] [SA] Inefficient vectorization: more than 10% of the vector loads instructions are unaligned - When
            allocating arrays, don’t forget to align them. There are 0 issues ( = arrays) costing 2 points each

   + exec - 47  :
     analysis: Execution Time: 5 % - Vectorization Ratio: 100.00 % - Vector Length Use: 25.00 %
     Loop Computation Issues: 12
        [8] [SA] Presence of expensive FP instructions - Perform hoisting, change algorithm, use SVML or proper
            numerical library or perform value profiling (count the number of distinct input values). There are 2
            issues (= instructions) costing 4 points each.
        [4] [SA] Less than 10% of the FP ADD/SUB/MUL arithmetic operations are performed using FMA - Reorganize
            arithmetic expressions to exhibit potential for FMA. This issue costs 4 points.
     Data Access Issues: 0
        [0] [SA] Inefficient vectorization: more than 10% of the vector loads instructions are unaligned - When
            allocating arrays, don’t forget to align them. There are 0 issues ( = arrays) costing 2 points each

   + exec - 97  :
     analysis: Execution Time: 4 % - Vectorization Ratio: 100.00 % - Vector Length Use: 25.00 %
     Loop Computation Issues: 8
        [4] [SA] Presence of expensive FP instructions - Perform hoisting, change algorithm, use SVML or proper
            numerical library or perform value profiling (count the number of distinct input values). There are 1
            issues (= instructions) costing 4 points each.
        [4] [SA] Less than 10% of the FP ADD/SUB/MUL arithmetic operations are performed using FMA - Reorganize
            arithmetic expressions to exhibit potential for FMA. This issue costs 4 points.
     Data Access Issues: 0
        [0] [SA] Inefficient vectorization: more than 10% of the vector loads instructions are unaligned - When
            allocating arrays, don’t forget to align them. There are 0 issues ( = arrays) costing 2 points each



+====================================================================================================================+
+                                                 3  -  APPLICATION                                                  +
+====================================================================================================================+


+--------------------------------------------------------------------------------------------------------------------+
+                                               3.1  -  Categorization                                               +
+--------------------------------------------------------------------------------------------------------------------+

   Category | IO     | Exe    | System  | Others  | Memory | String | MPI   | TBB   | OMP   | Pthread | Math  |
  ----------+--------+--------+---------+---------+--------+--------+-------+-------+-------+---------+-------+
   Time (%) | 0.00   | 95.93  | 0.31    | 0.01    | 0.00   | 0.00   | 0.03  | 0.00  | 3.72  | 0.00    | 0.00  |




+--------------------------------------------------------------------------------------------------------------------+
+                                          3.2  -  Function Based Profiling                                          +
+--------------------------------------------------------------------------------------------------------------------+

   Buckets                    | Nb Functions              | Coverage                  | Cumulated Coverage        |
  ----------------------------+---------------------------+---------------------------+---------------------------+
   > 8%                       | 3                         | 66.60                     | 66.60                     |
   4% to 8%                   | 4                         | 20.79                     | 87.38                     |
   2% to 4%                   | 3                         | 8.70                      | 96.08                     |
   1% to 2%                   | 1                         | 1.96                      | 98.04                     |
   0.5% to 1%                 | 1                         | 0.67                      | 98.71                     |
   0.25% to 0.5%              | 3                         | 0.84                      | 99.55                     |
   0.125% to 0.25%            | 1                         | 0.24                      | 99.79                     |
   < 0.125%                   | 103                       | 0.21                      | 100.00                    |




+--------------------------------------------------------------------------------------------------------------------+
+                                            3.3  -  Loop Based Profiling                                            +
+--------------------------------------------------------------------------------------------------------------------+

   Buckets                    | Nb Loops                  | Coverage                  | Cumulated Coverage        |
  ----------------------------+---------------------------+---------------------------+---------------------------+
   > 8%                       | 0                         | 0.00                      | 0.00                      |
   4% to 8%                   | 7                         | 35.90                     | 35.90                     |
   2% to 4%                   | 16                        | 49.88                     | 85.78                     |
   1% to 2%                   | 5                         | 7.53                      | 93.30                     |
   0.5% to 1%                 | 2                         | 1.96                      | 95.26                     |
   0.25% to 0.5%              | 0                         | 0.00                      | 95.26                     |
   0.125% to 0.25%            | 0                         | 0.00                      | 95.26                     |
   < 0.125%                   | 72                        | 0.46                      | 95.72                     |


+====================================================================================================================+
+                                                  4  -  FUNCTIONS                                                   +
+====================================================================================================================+


+--------------------------------------------------------------------------------------------------------------------+
+                                              4.1  -  Top 10 Functions                                              +
+--------------------------------------------------------------------------------------------------------------------+

   Function                                               | Module              | Coverage (%)   | Time (s)       |
  --------------------------------------------------------+---------------------+----------------+----------------+
   __nv_advec_mom_kernel_mod_advec_mom_kernel__PARALLE... | exec                | 35.09          | 13.30          |
   __nv_advec_cell_kernel_module_advec_cell_kernel__PA... | exec                | 19.62          | 7.44           |
   __nv_pdv_kernel_module_pdv_kernel__PARALLEL_F1L67_1    | exec                | 11.89          | 4.51           |
   __nv_ideal_gas_kernel_module_ideal_gas_kernel__PARA... | exec                | 5.79           | 2.19           |
   __nv_reset_field_kernel_module_reset_field_kernel__... | exec                | 5.44           | 2.06           |
   __nv_accelerate_kernel_module_accelerate_kernel__PA... | exec                | 5.34           | 2.02           |
   __nv_flux_calc_kernel_module_flux_calc_kernel__PARA... | exec                | 4.22           | 1.60           |
   __nv_calc_dt_kernel_module_calc_dt_kernel__PARALLEL... | exec                | 3.25           | 1.23           |
   __kmp_hyper_barrier_release(barrier_type, kmp_info*... | libomp.so           | 2.77           | 1.05           |
   __nv_revert_kernel_module_revert_kernel__PARALLEL_F... | exec                | 2.69           | 1.02           |


+====================================================================================================================+
+                                                    5  -  LOOPS                                                     +
+====================================================================================================================+


+--------------------------------------------------------------------------------------------------------------------+
+                                                5.1  -  Top 10 Loops                                                +
+--------------------------------------------------------------------------------------------------------------------+

   Loop Id        | Module              | Source Location                                        | Coverage (%)   |
  ----------------+---------------------+--------------------------------------------------------+----------------+
   38             | exec                | PdV_kernel.f90-pp.f90:113-123,PdV_kernel.f90-pp.f90... | 6.45           |
   245            | exec                | ideal_gas_kernel.f90-pp.f90:50-56                      | 5.78           |
   35             | exec                | PdV_kernel.f90-pp.f90:77-87,PdV_kernel.f90-pp.f90:9... | 5.44           |
   47             | exec                | accelerate_kernel.f90-pp.f90:63-63,accelerate_kerne... | 5.34           |
   97             | exec                | advec_mom_kernel.f90-pp.f90:248-248                    | 4.35           |
   108            | exec                | advec_mom_kernel.f90-pp.f90:184-184                    | 4.32           |
   191            | exec                | flux_calc_kernel.f90-pp.f90:57-59                      | 4.22           |
   110            | exec                | advec_mom_kernel.f90-pp.f90:152-177                    | 3.58           |
   53             | exec                | advec_cell_kernel.f90-pp.f90:256-261                   | 3.57           |
   99             | exec                | advec_mom_kernel.f90-pp.f90:215-215,advec_mom_kerne... | 3.56           |





+====================================================================================================================+
+                                                     6  -  CQA                                                      +
+====================================================================================================================+


+--------------------------------------------------------------------------------------------------------------------+
+                                                   6.1  -  Loops                                                    +
+--------------------------------------------------------------------------------------------------------------------+





      6.1.1  -  Loop 38 from exec
  ==========================================================================================================

The loop is defined in /beegfs/hackathon/users/eoseret/qaas_runs_test/178-661-4073/intel/CloverLeaf1.3-FC/build/aocc/CMakeFiles/clover_leaf.dir/CloverLeaf_ref/kernels/PdV_kernel.f90-pp.f90:113-123,129-135.

It is main loop of related source loop which is unrolled by 2 (including vectorization).

      6.1.1.1  -  Path 1
  ----------------------------------------------------------------------------------------------------------

14% of peak computational performance is used (4.48 out of 32.00 FLOP per cycle (GFLOPS @ 1GHz))

      6.1.1.1.1  -  Vectorization
  ----------------------------------------------------------------------------------------------------------

Your loop is vectorized, but using only 128 out of 512 bits (SSE/AVX-128 instructions on AVX-512 processors).
<<image_2x64_512>>By fully vectorizing your loop, you can lower the cost of an iteration from 16.50 to 16.00 cycles (1.03x speedup).

Details
All SSE/AVX instructions are used in vector version (process two or more data elements in vector registers).
Since your execution units are vector units, only a fully vectorized loop can use their full power.


Workaround
 - Pass to your compiler a micro-architecture specialization option:
  * Please look into your compiler manual for march=native or equivalent
 - Use vector aligned instructions:
  1) align your arrays on 64 bytes boundaries
  2) inform your compiler that your arrays are vector aligned: read your compiler manual.



      6.1.1.1.2  -  Execution units bottlenecks
  ----------------------------------------------------------------------------------------------------------

Performance is limited by execution of FP multiply or FMA (fused multiply-add) operations (the FP multiply/FMA unit is a bottleneck).


Workaround
Reduce the number of FP multiply/FMA instructions




      6.1.1.1.3  -  FMA
  ----------------------------------------------------------------------------------------------------------

Presence of both ADD/SUB and MUL operations.

Workaround
 - Pass to your compiler a micro-architecture specialization option:
  * Please look into your compiler manual for march=native or equivalent
 - Try to change order in which elements are evaluated (using parentheses) in arithmetic expressions containing both ADD/SUB and MUL operations to enable your compiler to generate FMA instructions wherever possible.
For instance a + b*c is a valid FMA (MUL then ADD).
However (a+b)* c cannot be translated into an FMA (ADD then MUL).





      6.1.1.1.4  -  Vector unaligned load/store instructions
  ----------------------------------------------------------------------------------------------------------

Detected 27 suboptimal vector unaligned load/store instructions.


Details
 - MOVUPD: 27 occurrences<<list_path_1_vec_align_1>>


Workaround
 - Pass to your compiler a micro-architecture specialization option:
  * Please look into your compiler manual for march=native or equivalent
 - Use vector aligned instructions:
  1) align your arrays on 64 bytes boundaries
  2) inform your compiler that your arrays are vector aligned: read your compiler manual.
<<image_vec_align>>


      6.1.1.1.5  -  Type of elements and instruction set
  ----------------------------------------------------------------------------------------------------------

37 SSE or AVX instructions are processing arithmetic or math operations on double precision FP elements in vector mode (two at a time).



      6.1.1.1.6  -  Matching between your loop (in the source code) and the binary loop
  ----------------------------------------------------------------------------------------------------------

The binary loop is composed of 74 FP arithmetical operations:
 - 36: addition or subtraction
 - 30: multiply
 - 8: divide
The binary loop is loading 448 bytes (56 double precision FP elements).
The binary loop is storing 32 bytes (4 double precision FP elements).


      6.1.1.1.7  -  Arithmetic intensity
  ----------------------------------------------------------------------------------------------------------

Arithmetic intensity is 0.15 FP operations per loaded or stored byte.







      6.1.2  -  Loop 245 from exec
  ==========================================================================================================

The loop is defined in /beegfs/hackathon/users/eoseret/qaas_runs_test/178-661-4073/intel/CloverLeaf1.3-FC/build/aocc/CMakeFiles/clover_leaf.dir/CloverLeaf_ref/kernels/ideal_gas_kernel.f90-pp.f90:50-56.

It is main loop of related source loop which is unrolled by 2 (including vectorization).

      6.1.2.1  -  Path 1
  ----------------------------------------------------------------------------------------------------------

6% of peak computational performance is used (2.12 out of 32.00 FLOP per cycle (GFLOPS @ 1GHz))

      6.1.2.1.1  -  Vectorization
  ----------------------------------------------------------------------------------------------------------

Your loop is vectorized, but using only 128 out of 512 bits (SSE/AVX-128 instructions on AVX-512 processors).
<<image_2x64_512>>

Details
All SSE/AVX instructions are used in vector version (process two or more data elements in vector registers).



      6.1.2.1.2  -  Execution units bottlenecks
  ----------------------------------------------------------------------------------------------------------

Performance is limited by execution of divide and square root operations (the divide/square root unit is a bottleneck).

By removing all these bottlenecks, you can lower the cost of an iteration from 8.50 to 4.00 cycles (2.13x speedup).


Workaround
 - Reduce the number of division or square root instructions:
  * If denominator is constant over iterations, use reciprocal (replace x/y with x*(1/y)). Check precision impact.
 - Check whether you really need double precision. If not, switch to single precision to speedup execution





      6.1.2.1.3  -  Expensive FP math instructions/calls
  ----------------------------------------------------------------------------------------------------------

Detected performance impact from expensive FP math instructions/calls.
By removing/reexpressing them, you can lower the cost of an iteration from 8.50 to 4.50 cycles (1.89x speedup).


      6.1.2.1.4  -  FMA
  ----------------------------------------------------------------------------------------------------------

Presence of both ADD/SUB and MUL operations.

Workaround
 - Pass to your compiler a micro-architecture specialization option:
  * Please look into your compiler manual for march=native or equivalent
 - Try to change order in which elements are evaluated (using parentheses) in arithmetic expressions containing both ADD/SUB and MUL operations to enable your compiler to generate FMA instructions wherever possible.
For instance a + b*c is a valid FMA (MUL then ADD).
However (a+b)* c cannot be translated into an FMA (ADD then MUL).





      6.1.2.1.5  -  Vector unaligned load/store instructions
  ----------------------------------------------------------------------------------------------------------

Detected 4 suboptimal vector unaligned load/store instructions.


Details
 - MOVUPD: 4 occurrences<<list_path_1_vec_align_1>>


Workaround
 - Pass to your compiler a micro-architecture specialization option:
  * Please look into your compiler manual for march=native or equivalent
 - Use vector aligned instructions:
  1) align your arrays on 64 bytes boundaries
  2) inform your compiler that your arrays are vector aligned: read your compiler manual.
<<image_vec_align>>


      6.1.2.1.6  -  Type of elements and instruction set
  ----------------------------------------------------------------------------------------------------------

9 SSE or AVX instructions are processing arithmetic or math operations on double precision FP elements in vector mode (two at a time).



      6.1.2.1.7  -  Matching between your loop (in the source code) and the binary loop
  ----------------------------------------------------------------------------------------------------------

The binary loop is composed of 18 FP arithmetical operations:
 - 2: addition or subtraction
 - 12: multiply
 - 2: divide
 - 2: square root
The binary loop is loading 32 bytes (4 double precision FP elements).
The binary loop is storing 32 bytes (4 double precision FP elements).


      6.1.2.1.8  -  Arithmetic intensity
  ----------------------------------------------------------------------------------------------------------

Arithmetic intensity is 0.28 FP operations per loaded or stored byte.







      6.1.3  -  Loop 35 from exec
  ==========================================================================================================

The loop is defined in /beegfs/hackathon/users/eoseret/qaas_runs_test/178-661-4073/intel/CloverLeaf1.3-FC/build/aocc/CMakeFiles/clover_leaf.dir/CloverLeaf_ref/kernels/PdV_kernel.f90-pp.f90:77-87,93-99.

It is main loop of related source loop which is unrolled by 2 (including vectorization).

      6.1.3.1  -  Path 1
  ----------------------------------------------------------------------------------------------------------

14% of peak computational performance is used (4.48 out of 32.00 FLOP per cycle (GFLOPS @ 1GHz))

      6.1.3.1.1  -  Vectorization
  ----------------------------------------------------------------------------------------------------------

Your loop is vectorized, but using only 128 out of 512 bits (SSE/AVX-128 instructions on AVX-512 processors).
<<image_2x64_512>>By fully vectorizing your loop, you can lower the cost of an iteration from 16.50 to 16.00 cycles (1.03x speedup).

Details
All SSE/AVX instructions are used in vector version (process two or more data elements in vector registers).
Since your execution units are vector units, only a fully vectorized loop can use their full power.


Workaround
 - Pass to your compiler a micro-architecture specialization option:
  * Please look into your compiler manual for march=native or equivalent
 - Use vector aligned instructions:
  1) align your arrays on 64 bytes boundaries
  2) inform your compiler that your arrays are vector aligned: read your compiler manual.



      6.1.3.1.2  -  Execution units bottlenecks
  ----------------------------------------------------------------------------------------------------------

Performance is limited by execution of FP multiply or FMA (fused multiply-add) operations (the FP multiply/FMA unit is a bottleneck).


Workaround
Reduce the number of FP multiply/FMA instructions




      6.1.3.1.3  -  FMA
  ----------------------------------------------------------------------------------------------------------

Presence of both ADD/SUB and MUL operations.

Workaround
 - Pass to your compiler a micro-architecture specialization option:
  * Please look into your compiler manual for march=native or equivalent
 - Try to change order in which elements are evaluated (using parentheses) in arithmetic expressions containing both ADD/SUB and MUL operations to enable your compiler to generate FMA instructions wherever possible.
For instance a + b*c is a valid FMA (MUL then ADD).
However (a+b)* c cannot be translated into an FMA (ADD then MUL).





      6.1.3.1.4  -  Vector unaligned load/store instructions
  ----------------------------------------------------------------------------------------------------------

Detected 19 suboptimal vector unaligned load/store instructions.


Details
 - MOVUPD: 19 occurrences<<list_path_1_vec_align_1>>


Workaround
 - Pass to your compiler a micro-architecture specialization option:
  * Please look into your compiler manual for march=native or equivalent
 - Use vector aligned instructions:
  1) align your arrays on 64 bytes boundaries
  2) inform your compiler that your arrays are vector aligned: read your compiler manual.
<<image_vec_align>>


      6.1.3.1.5  -  Type of elements and instruction set
  ----------------------------------------------------------------------------------------------------------

37 SSE or AVX instructions are processing arithmetic or math operations on double precision FP elements in vector mode (two at a time).



      6.1.3.1.6  -  Matching between your loop (in the source code) and the binary loop
  ----------------------------------------------------------------------------------------------------------

The binary loop is composed of 74 FP arithmetical operations:
 - 28: addition or subtraction
 - 38: multiply
 - 8: divide
The binary loop is loading 280 bytes (35 double precision FP elements).
The binary loop is storing 32 bytes (4 double precision FP elements).


      6.1.3.1.7  -  Arithmetic intensity
  ----------------------------------------------------------------------------------------------------------

Arithmetic intensity is 0.24 FP operations per loaded or stored byte.







      6.1.4  -  Loop 47 from exec
  ==========================================================================================================

The loop is defined in /beegfs/hackathon/users/eoseret/qaas_runs_test/178-661-4073/intel/CloverLeaf1.3-FC/build/aocc/CMakeFiles/clover_leaf.dir/CloverLeaf_ref/kernels/accelerate_kernel.f90-pp.f90:63,69-75.

It is main loop of related source loop which is unrolled by 4 (including vectorization).

      6.1.4.1  -  Path 1
  ----------------------------------------------------------------------------------------------------------

12% of peak computational performance is used (4.11 out of 32.00 FLOP per cycle (GFLOPS @ 1GHz))

      6.1.4.1.1  -  Vectorization
  ----------------------------------------------------------------------------------------------------------

Your loop is vectorized, but using only 128 out of 512 bits (SSE/AVX-128 instructions on AVX-512 processors).
<<image_2x64_512>>By fully vectorizing your loop, you can lower the cost of an iteration from 36.00 to 9.00 cycles (4.00x speedup).

Details
All SSE/AVX instructions are used in vector version (process two or more data elements in vector registers).
Since your execution units are vector units, only a fully vectorized loop can use their full power.


Workaround
 - Pass to your compiler a micro-architecture specialization option:
  * Please look into your compiler manual for march=native or equivalent
 - Use vector aligned instructions:
  1) align your arrays on 64 bytes boundaries
  2) inform your compiler that your arrays are vector aligned: read your compiler manual.



      6.1.4.1.2  -  Execution units bottlenecks
  ----------------------------------------------------------------------------------------------------------

Performance is limited by execution of FP multiply or FMA (fused multiply-add) operations (the FP multiply/FMA unit is a bottleneck).

By removing all these bottlenecks, you can lower the cost of an iteration from 36.00 to 24.50 cycles (1.47x speedup).


Workaround
Reduce the number of FP multiply/FMA instructions




      6.1.4.1.3  -  FMA
  ----------------------------------------------------------------------------------------------------------

Presence of both ADD/SUB and MUL operations.

Workaround
 - Pass to your compiler a micro-architecture specialization option:
  * Please look into your compiler manual for march=native or equivalent
 - Try to change order in which elements are evaluated (using parentheses) in arithmetic expressions containing both ADD/SUB and MUL operations to enable your compiler to generate FMA instructions wherever possible.
For instance a + b*c is a valid FMA (MUL then ADD).
However (a+b)* c cannot be translated into an FMA (ADD then MUL).





      6.1.4.1.4  -  Vector unaligned load/store instructions
  ----------------------------------------------------------------------------------------------------------

Detected 48 suboptimal vector unaligned load/store instructions.


Details
 - MOVUPD: 48 occurrences<<list_path_1_vec_align_1>>


Workaround
 - Pass to your compiler a micro-architecture specialization option:
  * Please look into your compiler manual for march=native or equivalent
 - Use vector aligned instructions:
  1) align your arrays on 64 bytes boundaries
  2) inform your compiler that your arrays are vector aligned: read your compiler manual.
<<image_vec_align>>


      6.1.4.1.5  -  Type of elements and instruction set
  ----------------------------------------------------------------------------------------------------------

74 SSE or AVX instructions are processing arithmetic or math operations on double precision FP elements in vector mode (two at a time).



      6.1.4.1.6  -  Matching between your loop (in the source code) and the binary loop
  ----------------------------------------------------------------------------------------------------------

The binary loop is composed of 148 FP arithmetical operations:
 - 76: addition or subtraction
 - 68: multiply
 - 4: divide
The binary loop is loading 872 bytes (109 double precision FP elements).
The binary loop is storing 160 bytes (20 double precision FP elements).


      6.1.4.1.7  -  Arithmetic intensity
  ----------------------------------------------------------------------------------------------------------

Arithmetic intensity is 0.14 FP operations per loaded or stored byte.







      6.1.5  -  Loop 97 from exec
  ==========================================================================================================

The loop is defined in /beegfs/hackathon/users/eoseret/qaas_runs_test/178-661-4073/intel/CloverLeaf1.3-FC/build/aocc/CMakeFiles/clover_leaf.dir/CloverLeaf_ref/kernels/advec_mom_kernel.f90-pp.f90:248.

It is main loop of related source loop which is unrolled by 2 (including vectorization).

      6.1.5.1  -  Path 1
  ----------------------------------------------------------------------------------------------------------

6% of peak computational performance is used (2.00 out of 32.00 FLOP per cycle (GFLOPS @ 1GHz))

      6.1.5.1.1  -  Vectorization
  ----------------------------------------------------------------------------------------------------------

Your loop is vectorized, but using only 128 out of 512 bits (SSE/AVX-128 instructions on AVX-512 processors).
<<image_2x64_512>>

Details
All SSE/AVX instructions are used in vector version (process two or more data elements in vector registers).



      6.1.5.1.2  -  Execution units bottlenecks
  ----------------------------------------------------------------------------------------------------------

Performance is limited by execution of divide and square root operations (the divide/square root unit is a bottleneck).

By removing all these bottlenecks, you can lower the cost of an iteration from 4.00 to 2.00 cycles (2.00x speedup).


Workaround
 - Reduce the number of division or square root instructions:
  * If denominator is constant over iterations, use reciprocal (replace x/y with x*(1/y)). Check precision impact.
 - Check whether you really need double precision. If not, switch to single precision to speedup execution





      6.1.5.1.3  -  Expensive FP math instructions/calls
  ----------------------------------------------------------------------------------------------------------

Detected performance impact from expensive FP math instructions/calls.
By removing/reexpressing them, you can lower the cost of an iteration from 4.00 to 2.00 cycles (2.00x speedup).


      6.1.5.1.4  -  FMA
  ----------------------------------------------------------------------------------------------------------

Presence of both ADD/SUB and MUL operations.

Workaround
 - Pass to your compiler a micro-architecture specialization option:
  * Please look into your compiler manual for march=native or equivalent
 - Try to change order in which elements are evaluated (using parentheses) in arithmetic expressions containing both ADD/SUB and MUL operations to enable your compiler to generate FMA instructions wherever possible.
For instance a + b*c is a valid FMA (MUL then ADD).
However (a+b)* c cannot be translated into an FMA (ADD then MUL).





      6.1.5.1.5  -  Vector unaligned load/store instructions
  ----------------------------------------------------------------------------------------------------------

Detected 6 suboptimal vector unaligned load/store instructions.


Details
 - MOVUPD: 6 occurrences<<list_path_1_vec_align_1>>


Workaround
 - Pass to your compiler a micro-architecture specialization option:
  * Please look into your compiler manual for march=native or equivalent
 - Use vector aligned instructions:
  1) align your arrays on 64 bytes boundaries
  2) inform your compiler that your arrays are vector aligned: read your compiler manual.
<<image_vec_align>>


      6.1.5.1.6  -  Type of elements and instruction set
  ----------------------------------------------------------------------------------------------------------

4 SSE or AVX instructions are processing arithmetic or math operations on double precision FP elements in vector mode (two at a time).



      6.1.5.1.7  -  Matching between your loop (in the source code) and the binary loop
  ----------------------------------------------------------------------------------------------------------

The binary loop is composed of 8 FP arithmetical operations:
 - 4: addition or subtraction
 - 2: multiply
 - 2: divide
The binary loop is loading 80 bytes (10 double precision FP elements).
The binary loop is storing 16 bytes (2 double precision FP elements).


      6.1.5.1.8  -  Arithmetic intensity
  ----------------------------------------------------------------------------------------------------------

Arithmetic intensity is 0.08 FP operations per loaded or stored byte.







      6.1.6  -  Loop 108 from exec
  ==========================================================================================================

The loop is defined in /beegfs/hackathon/users/eoseret/qaas_runs_test/178-661-4073/intel/CloverLeaf1.3-FC/build/aocc/CMakeFiles/clover_leaf.dir/CloverLeaf_ref/kernels/advec_mom_kernel.f90-pp.f90:184.

It is main loop of related source loop which is unrolled by 2 (including vectorization).

      6.1.6.1  -  Path 1
  ----------------------------------------------------------------------------------------------------------

6% of peak computational performance is used (2.00 out of 32.00 FLOP per cycle (GFLOPS @ 1GHz))

      6.1.6.1.1  -  Vectorization
  ----------------------------------------------------------------------------------------------------------

Your loop is vectorized, but using only 128 out of 512 bits (SSE/AVX-128 instructions on AVX-512 processors).
<<image_2x64_512>>

Details
All SSE/AVX instructions are used in vector version (process two or more data elements in vector registers).



      6.1.6.1.2  -  Execution units bottlenecks
  ----------------------------------------------------------------------------------------------------------

Performance is limited by execution of divide and square root operations (the divide/square root unit is a bottleneck).

By removing all these bottlenecks, you can lower the cost of an iteration from 4.00 to 2.00 cycles (2.00x speedup).


Workaround
 - Reduce the number of division or square root instructions:
  * If denominator is constant over iterations, use reciprocal (replace x/y with x*(1/y)). Check precision impact.
 - Check whether you really need double precision. If not, switch to single precision to speedup execution





      6.1.6.1.3  -  Expensive FP math instructions/calls
  ----------------------------------------------------------------------------------------------------------

Detected performance impact from expensive FP math instructions/calls.
By removing/reexpressing them, you can lower the cost of an iteration from 4.00 to 2.00 cycles (2.00x speedup).


      6.1.6.1.4  -  FMA
  ----------------------------------------------------------------------------------------------------------

Presence of both ADD/SUB and MUL operations.

Workaround
 - Pass to your compiler a micro-architecture specialization option:
  * Please look into your compiler manual for march=native or equivalent
 - Try to change order in which elements are evaluated (using parentheses) in arithmetic expressions containing both ADD/SUB and MUL operations to enable your compiler to generate FMA instructions wherever possible.
For instance a + b*c is a valid FMA (MUL then ADD).
However (a+b)* c cannot be translated into an FMA (ADD then MUL).





      6.1.6.1.5  -  Vector unaligned load/store instructions
  ----------------------------------------------------------------------------------------------------------

Detected 6 suboptimal vector unaligned load/store instructions.


Details
 - MOVUPD: 6 occurrences<<list_path_1_vec_align_1>>


Workaround
 - Pass to your compiler a micro-architecture specialization option:
  * Please look into your compiler manual for march=native or equivalent
 - Use vector aligned instructions:
  1) align your arrays on 64 bytes boundaries
  2) inform your compiler that your arrays are vector aligned: read your compiler manual.
<<image_vec_align>>


      6.1.6.1.6  -  Type of elements and instruction set
  ----------------------------------------------------------------------------------------------------------

4 SSE or AVX instructions are processing arithmetic or math operations on double precision FP elements in vector mode (two at a time).



      6.1.6.1.7  -  Matching between your loop (in the source code) and the binary loop
  ----------------------------------------------------------------------------------------------------------

The binary loop is composed of 8 FP arithmetical operations:
 - 4: addition or subtraction
 - 2: multiply
 - 2: divide
The binary loop is loading 80 bytes (10 double precision FP elements).
The binary loop is storing 16 bytes (2 double precision FP elements).


      6.1.6.1.8  -  Arithmetic intensity
  ----------------------------------------------------------------------------------------------------------

Arithmetic intensity is 0.08 FP operations per loaded or stored byte.







      6.1.7  -  Loop 191 from exec
  ==========================================================================================================

The loop is defined in /beegfs/hackathon/users/eoseret/qaas_runs_test/178-661-4073/intel/CloverLeaf1.3-FC/build/aocc/CMakeFiles/clover_leaf.dir/CloverLeaf_ref/kernels/flux_calc_kernel.f90-pp.f90:57-59.

It is main loop of related source loop which is unrolled by 2 (including vectorization).

      6.1.7.1  -  Path 1
  ----------------------------------------------------------------------------------------------------------

12% of peak computational performance is used (4.00 out of 32.00 FLOP per cycle (GFLOPS @ 1GHz))

      6.1.7.1.1  -  Vectorization
  ----------------------------------------------------------------------------------------------------------

Your loop is vectorized, but using only 128 out of 512 bits (SSE/AVX-128 instructions on AVX-512 processors).
<<image_2x64_512>>By fully vectorizing your loop, you can lower the cost of an iteration from 5.00 to 1.25 cycles (4.00x speedup).

Details
All SSE/AVX instructions are used in vector version (process two or more data elements in vector registers).
Since your execution units are vector units, only a fully vectorized loop can use their full power.


Workaround
 - Pass to your compiler a micro-architecture specialization option:
  * Please look into your compiler manual for march=native or equivalent
 - Use vector aligned instructions:
  1) align your arrays on 64 bytes boundaries
  2) inform your compiler that your arrays are vector aligned: read your compiler manual.



      6.1.7.1.2  -  Execution units bottlenecks
  ----------------------------------------------------------------------------------------------------------

Performance is limited by execution of FP multiply or FMA (fused multiply-add) operations (the FP multiply/FMA unit is a bottleneck).

By removing all these bottlenecks, you can lower the cost of an iteration from 5.00 to 4.00 cycles (1.25x speedup).


Workaround
Reduce the number of FP multiply/FMA instructions




      6.1.7.1.3  -  FMA
  ----------------------------------------------------------------------------------------------------------

Presence of both ADD/SUB and MUL operations.

Workaround
 - Pass to your compiler a micro-architecture specialization option:
  * Please look into your compiler manual for march=native or equivalent
 - Try to change order in which elements are evaluated (using parentheses) in arithmetic expressions containing both ADD/SUB and MUL operations to enable your compiler to generate FMA instructions wherever possible.
For instance a + b*c is a valid FMA (MUL then ADD).
However (a+b)* c cannot be translated into an FMA (ADD then MUL).





      6.1.7.1.4  -  Vector unaligned load/store instructions
  ----------------------------------------------------------------------------------------------------------

Detected 12 suboptimal vector unaligned load/store instructions.


Details
 - MOVUPD: 12 occurrences<<list_path_1_vec_align_1>>


Workaround
 - Pass to your compiler a micro-architecture specialization option:
  * Please look into your compiler manual for march=native or equivalent
 - Use vector aligned instructions:
  1) align your arrays on 64 bytes boundaries
  2) inform your compiler that your arrays are vector aligned: read your compiler manual.
<<image_vec_align>>


      6.1.7.1.5  -  Type of elements and instruction set
  ----------------------------------------------------------------------------------------------------------

10 SSE or AVX instructions are processing arithmetic or math operations on double precision FP elements in vector mode (two at a time).



      6.1.7.1.6  -  Matching between your loop (in the source code) and the binary loop
  ----------------------------------------------------------------------------------------------------------

The binary loop is composed of 20 FP arithmetical operations:
 - 12: addition or subtraction
 - 8: multiply
The binary loop is loading 160 bytes (20 double precision FP elements).
The binary loop is storing 32 bytes (4 double precision FP elements).


      6.1.7.1.7  -  Arithmetic intensity
  ----------------------------------------------------------------------------------------------------------

Arithmetic intensity is 0.10 FP operations per loaded or stored byte.







      6.1.8  -  Loop 110 from exec
  ==========================================================================================================

The loop is defined in /beegfs/hackathon/users/eoseret/qaas_runs_test/178-661-4073/intel/CloverLeaf1.3-FC/build/aocc/CMakeFiles/clover_leaf.dir/CloverLeaf_ref/kernels/advec_mom_kernel.f90-pp.f90:152-177.

The related source loop is not unrolled or unrolled with no peel/tail loop.
This loop has 4 execution paths.

The presence of multiple execution paths is typically the main/first bottleneck.
Try to simplify control inside loop: ideally, try to remove all conditional expressions, for example by (if applicable):
 - hoisting them (moving them outside the loop)
 - turning them into conditional moves, MIN or MAX


Ex: if (x<0) x=0 => x = max(0,x) (Fortran instrinsic procedure)


      6.1.8.1  -  Path 1
  ----------------------------------------------------------------------------------------------------------

3% of peak computational performance is used (1.26 out of 32.00 FLOP per cycle (GFLOPS @ 1GHz))

      6.1.8.1.1  -  Vectorization
  ----------------------------------------------------------------------------------------------------------

Your loop is partially vectorized.
Only 18% of vector register length is used (average across all SSE/AVX instructions).
By fully vectorizing your loop, you can lower the cost of an iteration from 14.33 to 8.00 cycles (1.79x speedup).

Details
50% of SSE/AVX instructions are used in vector version (process two or more data elements in vector registers):
 - 0% of SSE/AVX loads are used in vector version.
 - 0% of SSE/AVX stores are used in vector version.
 - 33% of SSE/AVX addition or subtraction instructions are used in vector version.
 - 20% of SSE/AVX multiply instructions are used in vector version.
 - 33% of SSE/AVX divide and square root instructions are used in vector version.
 - 66% of SSE/AVX instructions that are not load, store, addition, subtraction nor multiply instructions are used in vector version.
Since your execution units are vector units, only a fully vectorized loop can use their full power.


Workaround
 - Try another compiler or update/tune your current one
 - Remove inter-iterations dependences from your loop and make it unit-stride:
  * If your arrays have 2 or more dimensions, check whether elements are accessed contiguously and, otherwise, try to permute loops accordingly:
Fortran storage order is column-major: do i do j a(i,j) = b(i,j) (slow, non stride 1) => do i do j a(j,i) = b(i,j) (fast, stride 1)<<image_col_maj>>
  * If your loop streams arrays of structures (AoS), try to use structures of arrays instead (SoA):
do i a(i)%x = b(i)%x (slow, non stride 1) => do i a%x(i) = b%x(i) (fast, stride 1)



      6.1.8.1.2  -  Execution units bottlenecks
  ----------------------------------------------------------------------------------------------------------

Found no such bottlenecks but see expert reports for more complex bottlenecks.




      6.1.8.1.3  -  FMA
  ----------------------------------------------------------------------------------------------------------

Presence of both ADD/SUB and MUL operations.

Workaround
 - Pass to your compiler a micro-architecture specialization option:
  * Please look into your compiler manual for march=native or equivalent
 - Try to change order in which elements are evaluated (using parentheses) in arithmetic expressions containing both ADD/SUB and MUL operations to enable your compiler to generate FMA instructions wherever possible.
For instance a + b*c is a valid FMA (MUL then ADD).
However (a+b)* c cannot be translated into an FMA (ADD then MUL).





      6.1.8.1.4  -  Slow data structures access
  ----------------------------------------------------------------------------------------------------------

Detected data structures (typically arrays) that cannot be efficiently read/written

Details
 - Constant unknown stride: 2 occurrence(s)
 - Irregular (variable stride) or indirect: 1 occurrence(s)
Non-unit stride (uncontiguous) accesses are not efficiently using data caches


Workaround
 - Try to reorganize arrays of structures to structures of arrays
 - Consider to permute loops (see vectorization gain report)
 - Try to remove indirect accesses. If applicable, precompute elements out of the innermost loop.



      6.1.8.1.5  -  Vector unaligned load/store instructions
  ----------------------------------------------------------------------------------------------------------

Detected 2 suboptimal vector unaligned load/store instructions.


Details
 - MOVHPD: 2 occurrences<<list_path_1_vec_align_1>>


Workaround
 - Pass to your compiler a micro-architecture specialization option:
  * Please look into your compiler manual for march=native or equivalent
 - Use vector aligned instructions:
  1) align your arrays on 64 bytes boundaries
  2) inform your compiler that your arrays are vector aligned: read your compiler manual.
<<image_vec_align>>


      6.1.8.1.6  -  Conversion instructions
  ----------------------------------------------------------------------------------------------------------

Detected expensive conversion instructions.

Details
 - CLTQ: 1 occurrences<<list_path_1_cvt_1>>


Workaround
Avoid mixing data with different types. In particular, check if the type of constants is the same as array elements.


      6.1.8.1.7  -  Type of elements and instruction set
  ----------------------------------------------------------------------------------------------------------

17 SSE or AVX instructions are processing arithmetic or math operations on double precision FP elements in scalar mode (one at a time).
4 SSE or AVX instructions are processing arithmetic or math operations on double precision FP elements in vector mode (two at a time).



      6.1.8.1.8  -  Matching between your loop (in the source code) and the binary loop
  ----------------------------------------------------------------------------------------------------------

The binary loop is composed of 18 FP arithmetical operations:
 - 8: addition or subtraction
 - 6: multiply
 - 4: divide
The binary loop is loading 96 bytes (12 double precision FP elements).
The binary loop is storing 8 bytes (1 double precision FP elements).


      6.1.8.1.9  -  Arithmetic intensity
  ----------------------------------------------------------------------------------------------------------

Arithmetic intensity is 0.17 FP operations per loaded or stored byte.




      6.1.8.2  -  Path 2
  ----------------------------------------------------------------------------------------------------------

4% of peak computational performance is used (1.33 out of 32.00 FLOP per cycle (GFLOPS @ 1GHz))

      6.1.8.2.1  -  Code clean check
  ----------------------------------------------------------------------------------------------------------

Detected a slowdown caused by scalar integer instructions (typically used for address computation).
By removing them, you can lower the cost of an iteration from 6.00 to 4.50 cycles (1.33x speedup).

Workaround
 - Try to reorganize arrays of structures to structures of arrays
 - Consider to permute loops (see vectorization gain report)
 - To reference allocatable arrays, use "allocatable" instead of "pointer" pointers or qualify them with the "contiguous" attribute (Fortran 2008)
 - For structures, limit to one indirection. For example, use a_b%c instead of a%b%c with a_b set to a%b before this loop



      6.1.8.2.2  -  Vectorization
  ----------------------------------------------------------------------------------------------------------

Your loop is poorly vectorized.
Only 16% of vector register length is used (average across all SSE/AVX instructions).
By fully vectorizing your loop, you can lower the cost of an iteration from 6.00 to 2.00 cycles (3.00x speedup).

Details
29% of SSE/AVX instructions are used in vector version (process two or more data elements in vector registers):
 - 0% of SSE/AVX loads are used in vector version.
 - 0% of SSE/AVX stores are used in vector version.
 - 33% of SSE/AVX addition or subtraction instructions are used in vector version.
 - 0% of SSE/AVX multiply instructions are used in vector version.
 - 0% of SSE/AVX divide and square root instructions are used in vector version.
 - 50% of SSE/AVX instructions that are not load, store, addition, subtraction nor multiply instructions are used in vector version.
Since your execution units are vector units, only a fully vectorized loop can use their full power.


Workaround
 - Try another compiler or update/tune your current one
 - Remove inter-iterations dependences from your loop and make it unit-stride:
  * If your arrays have 2 or more dimensions, check whether elements are accessed contiguously and, otherwise, try to permute loops accordingly:
Fortran storage order is column-major: do i do j a(i,j) = b(i,j) (slow, non stride 1) => do i do j a(j,i) = b(i,j) (fast, stride 1)<<image_col_maj>>
  * If your loop streams arrays of structures (AoS), try to use structures of arrays instead (SoA):
do i a(i)%x = b(i)%x (slow, non stride 1) => do i a%x(i) = b%x(i) (fast, stride 1)



      6.1.8.2.3  -  Execution units bottlenecks
  ----------------------------------------------------------------------------------------------------------

Found no such bottlenecks but see expert reports for more complex bottlenecks.




      6.1.8.2.4  -  FMA
  ----------------------------------------------------------------------------------------------------------

Presence of both ADD/SUB and MUL operations.

Workaround
 - Pass to your compiler a micro-architecture specialization option:
  * Please look into your compiler manual for march=native or equivalent
 - Try to change order in which elements are evaluated (using parentheses) in arithmetic expressions containing both ADD/SUB and MUL operations to enable your compiler to generate FMA instructions wherever possible.
For instance a + b*c is a valid FMA (MUL then ADD).
However (a+b)* c cannot be translated into an FMA (ADD then MUL).





      6.1.8.2.5  -  Vector unaligned load/store instructions
  ----------------------------------------------------------------------------------------------------------

Detected 1 suboptimal vector unaligned load/store instructions.


Details
 - MOVHPD: 1 occurrences<<list_path_2_vec_align_1>>


Workaround
 - Pass to your compiler a micro-architecture specialization option:
  * Please look into your compiler manual for march=native or equivalent
 - Use vector aligned instructions:
  1) align your arrays on 64 bytes boundaries
  2) inform your compiler that your arrays are vector aligned: read your compiler manual.
<<image_vec_align>>


      6.1.8.2.6  -  Conversion instructions
  ----------------------------------------------------------------------------------------------------------

Detected expensive conversion instructions.

Details
 - CLTQ: 1 occurrences<<list_path_2_cvt_1>>


Workaround
Avoid mixing data with different types. In particular, check if the type of constants is the same as array elements.


      6.1.8.2.7  -  Type of elements and instruction set
  ----------------------------------------------------------------------------------------------------------

8 SSE or AVX instructions are processing arithmetic or math operations on double precision FP elements in scalar mode (one at a time).
1 SSE or AVX instructions are processing arithmetic or math operations on double precision FP elements in vector mode (two at a time).



      6.1.8.2.8  -  Matching between your loop (in the source code) and the binary loop
  ----------------------------------------------------------------------------------------------------------

The binary loop is composed of 8 FP arithmetical operations:
 - 4: addition or subtraction
 - 3: multiply
 - 1: divide
The binary loop is loading 48 bytes (6 double precision FP elements).
The binary loop is storing 8 bytes (1 double precision FP elements).


      6.1.8.2.9  -  Arithmetic intensity
  ----------------------------------------------------------------------------------------------------------

Arithmetic intensity is 0.14 FP operations per loaded or stored byte.




      6.1.8.3  -  Path 3
  ----------------------------------------------------------------------------------------------------------

3% of peak computational performance is used (1.24 out of 32.00 FLOP per cycle (GFLOPS @ 1GHz))

      6.1.8.3.1  -  Vectorization
  ----------------------------------------------------------------------------------------------------------

Your loop is partially vectorized.
Only 18% of vector register length is used (average across all SSE/AVX instructions).
By fully vectorizing your loop, you can lower the cost of an iteration from 14.50 to 8.00 cycles (1.81x speedup).

Details
50% of SSE/AVX instructions are used in vector version (process two or more data elements in vector registers):
 - 0% of SSE/AVX loads are used in vector version.
 - 0% of SSE/AVX stores are used in vector version.
 - 33% of SSE/AVX addition or subtraction instructions are used in vector version.
 - 20% of SSE/AVX multiply instructions are used in vector version.
 - 33% of SSE/AVX divide and square root instructions are used in vector version.
 - 66% of SSE/AVX instructions that are not load, store, addition, subtraction nor multiply instructions are used in vector version.
Since your execution units are vector units, only a fully vectorized loop can use their full power.


Workaround
 - Try another compiler or update/tune your current one
 - Remove inter-iterations dependences from your loop and make it unit-stride:
  * If your arrays have 2 or more dimensions, check whether elements are accessed contiguously and, otherwise, try to permute loops accordingly:
Fortran storage order is column-major: do i do j a(i,j) = b(i,j) (slow, non stride 1) => do i do j a(j,i) = b(i,j) (fast, stride 1)<<image_col_maj>>
  * If your loop streams arrays of structures (AoS), try to use structures of arrays instead (SoA):
do i a(i)%x = b(i)%x (slow, non stride 1) => do i a%x(i) = b%x(i) (fast, stride 1)



      6.1.8.3.2  -  Execution units bottlenecks
  ----------------------------------------------------------------------------------------------------------

Found no such bottlenecks but see expert reports for more complex bottlenecks.




      6.1.8.3.3  -  FMA
  ----------------------------------------------------------------------------------------------------------

Presence of both ADD/SUB and MUL operations.

Workaround
 - Pass to your compiler a micro-architecture specialization option:
  * Please look into your compiler manual for march=native or equivalent
 - Try to change order in which elements are evaluated (using parentheses) in arithmetic expressions containing both ADD/SUB and MUL operations to enable your compiler to generate FMA instructions wherever possible.
For instance a + b*c is a valid FMA (MUL then ADD).
However (a+b)* c cannot be translated into an FMA (ADD then MUL).





      6.1.8.3.4  -  Slow data structures access
  ----------------------------------------------------------------------------------------------------------

Detected data structures (typically arrays) that cannot be efficiently read/written

Details
 - Constant unknown stride: 4 occurrence(s)
 - Irregular (variable stride) or indirect: 1 occurrence(s)
Non-unit stride (uncontiguous) accesses are not efficiently using data caches


Workaround
 - Try to reorganize arrays of structures to structures of arrays
 - Consider to permute loops (see vectorization gain report)
 - Try to remove indirect accesses. If applicable, precompute elements out of the innermost loop.



      6.1.8.3.5  -  Vector unaligned load/store instructions
  ----------------------------------------------------------------------------------------------------------

Detected 2 suboptimal vector unaligned load/store instructions.


Details
 - MOVHPD: 2 occurrences<<list_path_3_vec_align_1>>


Workaround
 - Pass to your compiler a micro-architecture specialization option:
  * Please look into your compiler manual for march=native or equivalent
 - Use vector aligned instructions:
  1) align your arrays on 64 bytes boundaries
  2) inform your compiler that your arrays are vector aligned: read your compiler manual.
<<image_vec_align>>


      6.1.8.3.6  -  Conversion instructions
  ----------------------------------------------------------------------------------------------------------

Detected expensive conversion instructions.

Details
 - CLTQ: 1 occurrences<<list_path_3_cvt_1>>


Workaround
Avoid mixing data with different types. In particular, check if the type of constants is the same as array elements.


      6.1.8.3.7  -  Type of elements and instruction set
  ----------------------------------------------------------------------------------------------------------

17 SSE or AVX instructions are processing arithmetic or math operations on double precision FP elements in scalar mode (one at a time).
4 SSE or AVX instructions are processing arithmetic or math operations on double precision FP elements in vector mode (two at a time).



      6.1.8.3.8  -  Matching between your loop (in the source code) and the binary loop
  ----------------------------------------------------------------------------------------------------------

The binary loop is composed of 18 FP arithmetical operations:
 - 8: addition or subtraction
 - 6: multiply
 - 4: divide
The binary loop is loading 96 bytes (12 double precision FP elements).
The binary loop is storing 8 bytes (1 double precision FP elements).


      6.1.8.3.9  -  Arithmetic intensity
  ----------------------------------------------------------------------------------------------------------

Arithmetic intensity is 0.17 FP operations per loaded or stored byte.




      6.1.8.4  -  Path 4
  ----------------------------------------------------------------------------------------------------------

4% of peak computational performance is used (1.30 out of 32.00 FLOP per cycle (GFLOPS @ 1GHz))

      6.1.8.4.1  -  Code clean check
  ----------------------------------------------------------------------------------------------------------

Detected a slowdown caused by scalar integer instructions (typically used for address computation).
By removing them, you can lower the cost of an iteration from 6.17 to 4.50 cycles (1.37x speedup).

Workaround
 - Try to reorganize arrays of structures to structures of arrays
 - Consider to permute loops (see vectorization gain report)
 - To reference allocatable arrays, use "allocatable" instead of "pointer" pointers or qualify them with the "contiguous" attribute (Fortran 2008)
 - For structures, limit to one indirection. For example, use a_b%c instead of a%b%c with a_b set to a%b before this loop



      6.1.8.4.2  -  Vectorization
  ----------------------------------------------------------------------------------------------------------

Your loop is poorly vectorized.
Only 16% of vector register length is used (average across all SSE/AVX instructions).
By fully vectorizing your loop, you can lower the cost of an iteration from 6.17 to 2.00 cycles (3.08x speedup).

Details
29% of SSE/AVX instructions are used in vector version (process two or more data elements in vector registers):
 - 0% of SSE/AVX loads are used in vector version.
 - 0% of SSE/AVX stores are used in vector version.
 - 33% of SSE/AVX addition or subtraction instructions are used in vector version.
 - 0% of SSE/AVX multiply instructions are used in vector version.
 - 0% of SSE/AVX divide and square root instructions are used in vector version.
 - 50% of SSE/AVX instructions that are not load, store, addition, subtraction nor multiply instructions are used in vector version.
Since your execution units are vector units, only a fully vectorized loop can use their full power.


Workaround
 - Try another compiler or update/tune your current one
 - Remove inter-iterations dependences from your loop and make it unit-stride:
  * If your arrays have 2 or more dimensions, check whether elements are accessed contiguously and, otherwise, try to permute loops accordingly:
Fortran storage order is column-major: do i do j a(i,j) = b(i,j) (slow, non stride 1) => do i do j a(j,i) = b(i,j) (fast, stride 1)<<image_col_maj>>
  * If your loop streams arrays of structures (AoS), try to use structures of arrays instead (SoA):
do i a(i)%x = b(i)%x (slow, non stride 1) => do i a%x(i) = b%x(i) (fast, stride 1)



      6.1.8.4.3  -  Execution units bottlenecks
  ----------------------------------------------------------------------------------------------------------

Found no such bottlenecks but see expert reports for more complex bottlenecks.




      6.1.8.4.4  -  FMA
  ----------------------------------------------------------------------------------------------------------

Presence of both ADD/SUB and MUL operations.

Workaround
 - Pass to your compiler a micro-architecture specialization option:
  * Please look into your compiler manual for march=native or equivalent
 - Try to change order in which elements are evaluated (using parentheses) in arithmetic expressions containing both ADD/SUB and MUL operations to enable your compiler to generate FMA instructions wherever possible.
For instance a + b*c is a valid FMA (MUL then ADD).
However (a+b)* c cannot be translated into an FMA (ADD then MUL).





      6.1.8.4.5  -  Slow data structures access
  ----------------------------------------------------------------------------------------------------------

Detected data structures (typically arrays) that cannot be efficiently read/written

Details
 - Constant unknown stride: 2 occurrence(s)
Non-unit stride (uncontiguous) accesses are not efficiently using data caches


Workaround
 - Try to reorganize arrays of structures to structures of arrays
 - Consider to permute loops (see vectorization gain report)



      6.1.8.4.6  -  Vector unaligned load/store instructions
  ----------------------------------------------------------------------------------------------------------

Detected 1 suboptimal vector unaligned load/store instructions.


Details
 - MOVHPD: 1 occurrences<<list_path_4_vec_align_1>>


Workaround
 - Pass to your compiler a micro-architecture specialization option:
  * Please look into your compiler manual for march=native or equivalent
 - Use vector aligned instructions:
  1) align your arrays on 64 bytes boundaries
  2) inform your compiler that your arrays are vector aligned: read your compiler manual.
<<image_vec_align>>


      6.1.8.4.7  -  Conversion instructions
  ----------------------------------------------------------------------------------------------------------

Detected expensive conversion instructions.

Details
 - CLTQ: 1 occurrences<<list_path_4_cvt_1>>


Workaround
Avoid mixing data with different types. In particular, check if the type of constants is the same as array elements.


      6.1.8.4.8  -  Type of elements and instruction set
  ----------------------------------------------------------------------------------------------------------

8 SSE or AVX instructions are processing arithmetic or math operations on double precision FP elements in scalar mode (one at a time).
1 SSE or AVX instructions are processing arithmetic or math operations on double precision FP elements in vector mode (two at a time).



      6.1.8.4.9  -  Matching between your loop (in the source code) and the binary loop
  ----------------------------------------------------------------------------------------------------------

The binary loop is composed of 8 FP arithmetical operations:
 - 4: addition or subtraction
 - 3: multiply
 - 1: divide
The binary loop is loading 48 bytes (6 double precision FP elements).
The binary loop is storing 8 bytes (1 double precision FP elements).


      6.1.8.4.10  -  Arithmetic intensity
  ----------------------------------------------------------------------------------------------------------

Arithmetic intensity is 0.14 FP operations per loaded or stored byte.







      6.1.9  -  Loop 53 from exec
  ==========================================================================================================

The loop is defined in /beegfs/hackathon/users/eoseret/qaas_runs_test/178-661-4073/intel/CloverLeaf1.3-FC/build/aocc/CMakeFiles/clover_leaf.dir/CloverLeaf_ref/kernels/advec_cell_kernel.f90-pp.f90:256-261.

It is main loop of related source loop which is unrolled by 2 (including vectorization).

      6.1.9.1  -  Path 1
  ----------------------------------------------------------------------------------------------------------

7% of peak computational performance is used (2.50 out of 32.00 FLOP per cycle (GFLOPS @ 1GHz))

      6.1.9.1.1  -  Vectorization
  ----------------------------------------------------------------------------------------------------------

Your loop is vectorized, but using only 128 out of 512 bits (SSE/AVX-128 instructions on AVX-512 processors).
<<image_2x64_512>>

Details
All SSE/AVX instructions are used in vector version (process two or more data elements in vector registers).



      6.1.9.1.2  -  Execution units bottlenecks
  ----------------------------------------------------------------------------------------------------------

Performance is limited by execution of divide and square root operations (the divide/square root unit is a bottleneck).

By removing all these bottlenecks, you can lower the cost of an iteration from 8.00 to 4.00 cycles (2.00x speedup).


Workaround
 - Reduce the number of division or square root instructions:
  * If denominator is constant over iterations, use reciprocal (replace x/y with x*(1/y)). Check precision impact.
 - Check whether you really need double precision. If not, switch to single precision to speedup execution





      6.1.9.1.3  -  Expensive FP math instructions/calls
  ----------------------------------------------------------------------------------------------------------

Detected performance impact from expensive FP math instructions/calls.
By removing/reexpressing them, you can lower the cost of an iteration from 8.00 to 5.00 cycles (1.60x speedup).


      6.1.9.1.4  -  FMA
  ----------------------------------------------------------------------------------------------------------

Presence of both ADD/SUB and MUL operations.

Workaround
 - Pass to your compiler a micro-architecture specialization option:
  * Please look into your compiler manual for march=native or equivalent
 - Try to change order in which elements are evaluated (using parentheses) in arithmetic expressions containing both ADD/SUB and MUL operations to enable your compiler to generate FMA instructions wherever possible.
For instance a + b*c is a valid FMA (MUL then ADD).
However (a+b)* c cannot be translated into an FMA (ADD then MUL).





      6.1.9.1.5  -  Vector unaligned load/store instructions
  ----------------------------------------------------------------------------------------------------------

Detected 11 suboptimal vector unaligned load/store instructions.


Details
 - MOVUPD: 11 occurrences<<list_path_1_vec_align_1>>


Workaround
 - Pass to your compiler a micro-architecture specialization option:
  * Please look into your compiler manual for march=native or equivalent
 - Use vector aligned instructions:
  1) align your arrays on 64 bytes boundaries
  2) inform your compiler that your arrays are vector aligned: read your compiler manual.
<<image_vec_align>>


      6.1.9.1.6  -  Type of elements and instruction set
  ----------------------------------------------------------------------------------------------------------

10 SSE or AVX instructions are processing arithmetic or math operations on double precision FP elements in vector mode (two at a time).



      6.1.9.1.7  -  Matching between your loop (in the source code) and the binary loop
  ----------------------------------------------------------------------------------------------------------

The binary loop is composed of 20 FP arithmetical operations:
 - 12: addition or subtraction
 - 4: multiply
 - 4: divide
The binary loop is loading 144 bytes (18 double precision FP elements).
The binary loop is storing 32 bytes (4 double precision FP elements).


      6.1.9.1.8  -  Arithmetic intensity
  ----------------------------------------------------------------------------------------------------------

Arithmetic intensity is 0.11 FP operations per loaded or stored byte.







      6.1.10  -  Loop 99 from exec
  ==========================================================================================================

The loop is defined in /beegfs/hackathon/users/eoseret/qaas_runs_test/178-661-4073/intel/CloverLeaf1.3-FC/build/aocc/CMakeFiles/clover_leaf.dir/CloverLeaf_ref/kernels/advec_mom_kernel.f90-pp.f90:215,227-241.

The related source loop is not unrolled or unrolled with no peel/tail loop.
The structure of this loop is probably <if then [else] end>.

The presence of multiple execution paths is typically the main/first bottleneck.
Try to simplify control inside loop: ideally, try to remove all conditional expressions, for example by (if applicable):
 - hoisting them (moving them outside the loop)
 - turning them into conditional moves, MIN or MAX


Ex: if (x<0) x=0 => x = max(0,x) (Fortran instrinsic procedure)


      6.1.10.1  -  Path 1
  ----------------------------------------------------------------------------------------------------------

3% of peak computational performance is used (1.16 out of 32.00 FLOP per cycle (GFLOPS @ 1GHz))

      6.1.10.1.1  -  Code clean check
  ----------------------------------------------------------------------------------------------------------

Detected a slowdown caused by scalar integer instructions (typically used for address computation).
By removing them, you can lower the cost of an iteration from 15.50 to 14.17 cycles (1.09x speedup).

Workaround
 - Try to reorganize arrays of structures to structures of arrays
 - Consider to permute loops (see vectorization gain report)
 - To reference allocatable arrays, use "allocatable" instead of "pointer" pointers or qualify them with the "contiguous" attribute (Fortran 2008)
 - For structures, limit to one indirection. For example, use a_b%c instead of a%b%c with a_b set to a%b before this loop



      6.1.10.1.2  -  Vectorization
  ----------------------------------------------------------------------------------------------------------

Your loop is partially vectorized.
Only 18% of vector register length is used (average across all SSE/AVX instructions).
By fully vectorizing your loop, you can lower the cost of an iteration from 15.50 to 8.00 cycles (1.94x speedup).

Details
50% of SSE/AVX instructions are used in vector version (process two or more data elements in vector registers):
 - 0% of SSE/AVX loads are used in vector version.
 - 0% of SSE/AVX stores are used in vector version.
 - 33% of SSE/AVX addition or subtraction instructions are used in vector version.
 - 20% of SSE/AVX multiply instructions are used in vector version.
 - 33% of SSE/AVX divide and square root instructions are used in vector version.
 - 65% of SSE/AVX instructions that are not load, store, addition, subtraction nor multiply instructions are used in vector version.
Since your execution units are vector units, only a fully vectorized loop can use their full power.


Workaround
 - Try another compiler or update/tune your current one
 - Remove inter-iterations dependences from your loop and make it unit-stride:
  * If your arrays have 2 or more dimensions, check whether elements are accessed contiguously and, otherwise, try to permute loops accordingly:
Fortran storage order is column-major: do i do j a(i,j) = b(i,j) (slow, non stride 1) => do i do j a(j,i) = b(i,j) (fast, stride 1)<<image_col_maj>>
  * If your loop streams arrays of structures (AoS), try to use structures of arrays instead (SoA):
do i a(i)%x = b(i)%x (slow, non stride 1) => do i a%x(i) = b%x(i) (fast, stride 1)



      6.1.10.1.3  -  Execution units bottlenecks
  ----------------------------------------------------------------------------------------------------------

Found no such bottlenecks but see expert reports for more complex bottlenecks.




      6.1.10.1.4  -  FMA
  ----------------------------------------------------------------------------------------------------------

Presence of both ADD/SUB and MUL operations.

Workaround
 - Pass to your compiler a micro-architecture specialization option:
  * Please look into your compiler manual for march=native or equivalent
 - Try to change order in which elements are evaluated (using parentheses) in arithmetic expressions containing both ADD/SUB and MUL operations to enable your compiler to generate FMA instructions wherever possible.
For instance a + b*c is a valid FMA (MUL then ADD).
However (a+b)* c cannot be translated into an FMA (ADD then MUL).





      6.1.10.1.5  -  Complex instructions
  ----------------------------------------------------------------------------------------------------------

Detected COMPLEX INSTRUCTIONS.


Details
These instructions generate more than one micro-operation and only one of them can be decoded during a cycle and the extra micro-operations increase pressure on execution units.
 - CMOVBE: 4 occurrences<<list_path_1_complex_1>>



      6.1.10.1.6  -  Slow data structures access
  ----------------------------------------------------------------------------------------------------------

Detected data structures (typically arrays) that cannot be efficiently read/written

Details
 - Constant unknown stride: 2 occurrence(s)
 - Irregular (variable stride) or indirect: 3 occurrence(s)
Non-unit stride (uncontiguous) accesses are not efficiently using data caches


Workaround
 - Try to reorganize arrays of structures to structures of arrays
 - Consider to permute loops (see vectorization gain report)
 - Try to remove indirect accesses. If applicable, precompute elements out of the innermost loop.



      6.1.10.1.7  -  Vector unaligned load/store instructions
  ----------------------------------------------------------------------------------------------------------

Detected 2 suboptimal vector unaligned load/store instructions.


Details
 - MOVHPD: 2 occurrences<<list_path_1_vec_align_1>>


Workaround
 - Pass to your compiler a micro-architecture specialization option:
  * Please look into your compiler manual for march=native or equivalent
 - Use vector aligned instructions:
  1) align your arrays on 64 bytes boundaries
  2) inform your compiler that your arrays are vector aligned: read your compiler manual.
<<image_vec_align>>


      6.1.10.1.8  -  Conversion instructions
  ----------------------------------------------------------------------------------------------------------

Detected expensive conversion instructions.

Details
 - CLTQ: 1 occurrences<<list_path_1_cvt_1>>


Workaround
Avoid mixing data with different types. In particular, check if the type of constants is the same as array elements.


      6.1.10.1.9  -  Type of elements and instruction set
  ----------------------------------------------------------------------------------------------------------

18 SSE or AVX instructions are processing arithmetic or math operations on double precision FP elements in scalar mode (one at a time).
4 SSE or AVX instructions are processing arithmetic or math operations on double precision FP elements in vector mode (two at a time).



      6.1.10.1.10  -  Matching between your loop (in the source code) and the binary loop
  ----------------------------------------------------------------------------------------------------------

The binary loop is composed of 18 FP arithmetical operations:
 - 8: addition or subtraction
 - 6: multiply
 - 4: divide
The binary loop is loading 112 bytes (14 double precision FP elements).
The binary loop is storing 8 bytes (1 double precision FP elements).


      6.1.10.1.11  -  Arithmetic intensity
  ----------------------------------------------------------------------------------------------------------

Arithmetic intensity is 0.15 FP operations per loaded or stored byte.




      6.1.10.2  -  Path 2
  ----------------------------------------------------------------------------------------------------------

3% of peak computational performance is used (1.17 out of 32.00 FLOP per cycle (GFLOPS @ 1GHz))

      6.1.10.2.1  -  Code clean check
  ----------------------------------------------------------------------------------------------------------

Detected a slowdown caused by scalar integer instructions (typically used for address computation).
By removing them, you can lower the cost of an iteration from 6.83 to 5.00 cycles (1.37x speedup).

Workaround
 - Try to reorganize arrays of structures to structures of arrays
 - Consider to permute loops (see vectorization gain report)
 - To reference allocatable arrays, use "allocatable" instead of "pointer" pointers or qualify them with the "contiguous" attribute (Fortran 2008)
 - For structures, limit to one indirection. For example, use a_b%c instead of a%b%c with a_b set to a%b before this loop



      6.1.10.2.2  -  Vectorization
  ----------------------------------------------------------------------------------------------------------

Your loop is poorly vectorized.
Only 16% of vector register length is used (average across all SSE/AVX instructions).
By fully vectorizing your loop, you can lower the cost of an iteration from 6.83 to 2.00 cycles (3.42x speedup).

Details
30% of SSE/AVX instructions are used in vector version (process two or more data elements in vector registers):
 - 0% of SSE/AVX loads are used in vector version.
 - 0% of SSE/AVX stores are used in vector version.
 - 33% of SSE/AVX addition or subtraction instructions are used in vector version.
 - 0% of SSE/AVX multiply instructions are used in vector version.
 - 0% of SSE/AVX divide and square root instructions are used in vector version.
 - 54% of SSE/AVX instructions that are not load, store, addition, subtraction nor multiply instructions are used in vector version.
Since your execution units are vector units, only a fully vectorized loop can use their full power.


Workaround
 - Try another compiler or update/tune your current one
 - Remove inter-iterations dependences from your loop and make it unit-stride:
  * If your arrays have 2 or more dimensions, check whether elements are accessed contiguously and, otherwise, try to permute loops accordingly:
Fortran storage order is column-major: do i do j a(i,j) = b(i,j) (slow, non stride 1) => do i do j a(j,i) = b(i,j) (fast, stride 1)<<image_col_maj>>
  * If your loop streams arrays of structures (AoS), try to use structures of arrays instead (SoA):
do i a(i)%x = b(i)%x (slow, non stride 1) => do i a%x(i) = b%x(i) (fast, stride 1)



      6.1.10.2.3  -  Execution units bottlenecks
  ----------------------------------------------------------------------------------------------------------

Found no such bottlenecks but see expert reports for more complex bottlenecks.




      6.1.10.2.4  -  FMA
  ----------------------------------------------------------------------------------------------------------

Presence of both ADD/SUB and MUL operations.

Workaround
 - Pass to your compiler a micro-architecture specialization option:
  * Please look into your compiler manual for march=native or equivalent
 - Try to change order in which elements are evaluated (using parentheses) in arithmetic expressions containing both ADD/SUB and MUL operations to enable your compiler to generate FMA instructions wherever possible.
For instance a + b*c is a valid FMA (MUL then ADD).
However (a+b)* c cannot be translated into an FMA (ADD then MUL).





      6.1.10.2.5  -  Complex instructions
  ----------------------------------------------------------------------------------------------------------

Detected COMPLEX INSTRUCTIONS.


Details
These instructions generate more than one micro-operation and only one of them can be decoded during a cycle and the extra micro-operations increase pressure on execution units.
 - CMOVBE: 3 occurrences<<list_path_2_complex_1>>



      6.1.10.2.6  -  Slow data structures access
  ----------------------------------------------------------------------------------------------------------

Detected data structures (typically arrays) that cannot be efficiently read/written

Details
 - Irregular (variable stride) or indirect: 2 occurrence(s)
Non-unit stride (uncontiguous) accesses are not efficiently using data caches


Workaround
Try to remove indirect accesses. If applicable, precompute elements out of the innermost loop.


      6.1.10.2.7  -  Vector unaligned load/store instructions
  ----------------------------------------------------------------------------------------------------------

Detected 1 suboptimal vector unaligned load/store instructions.


Details
 - MOVHPD: 1 occurrences<<list_path_2_vec_align_1>>


Workaround
 - Pass to your compiler a micro-architecture specialization option:
  * Please look into your compiler manual for march=native or equivalent
 - Use vector aligned instructions:
  1) align your arrays on 64 bytes boundaries
  2) inform your compiler that your arrays are vector aligned: read your compiler manual.
<<image_vec_align>>


      6.1.10.2.8  -  Type of elements and instruction set
  ----------------------------------------------------------------------------------------------------------

8 SSE or AVX instructions are processing arithmetic or math operations on double precision FP elements in scalar mode (one at a time).
1 SSE or AVX instructions are processing arithmetic or math operations on double precision FP elements in vector mode (two at a time).



      6.1.10.2.9  -  Matching between your loop (in the source code) and the binary loop
  ----------------------------------------------------------------------------------------------------------

The binary loop is composed of 8 FP arithmetical operations:
 - 4: addition or subtraction
 - 3: multiply
 - 1: divide
The binary loop is loading 48 bytes (6 double precision FP elements).
The binary loop is storing 8 bytes (1 double precision FP elements).


      6.1.10.2.10  -  Arithmetic intensity
  ----------------------------------------------------------------------------------------------------------

Arithmetic intensity is 0.14 FP operations per loaded or stored byte.





[MAQAO] Info: STOP THE REPORT GENERATION
[MAQAO] Info: 
[MAQAO] Info: If your application produces files, they can be found in directory "/beegfs/hackathon/users/eoseret/qaas_runs_test/178-661-4073/intel/CloverLeaf1.3-FC/run/oneview_runs/defaults/aocc/oneview_run_1786614486"
[MAQAO] Info: 
