	*************************************************
	*                                               *
	*          ONE-View report generation           *
	*                                               *
	*************************************************

[MAQAO] Info: Experiment configuration summary is available adding -dbg=1 in command line

* [MAQAO] Warning: Experiment directory /beegfs/hackathon/users/eoseret/qaas_runs_test/178-663-0009/intel/TeaLeaf/run/oneview_runs/compilers/gcc_3/oneview_results_1786633507 already exists and is reused.
           It can be replaced using --replace in the command line.
[MAQAO] Info: 
[MAQAO] Info: START THE APPLICATION PROFILING
[MAQAO] Info: -> RUNNING THE PROFILER...
[MAQAO] Info:   LPROF has already been run
[MAQAO] Info: STOP THE APPLICATION PROFILING
[MAQAO] Info: 
[MAQAO] Info: START FUNCTIONS AND LOOPS ANALYSIS ...
[MAQAO] Info: STOP FUNCTIONS AND LOOPS ANALYSIS ...
[MAQAO] Info: 
[MAQAO] Info: START THE REPORT GENERATION
[MAQAO] Info: -> ONE-VIEW EXPERIMENT DIRECTORY: /beegfs/hackathon/users/eoseret/qaas_runs_test/178-663-0009/intel/TeaLeaf/run/oneview_runs/compilers/gcc_3/oneview_results_1786633507


+====================================================================================================================+
+                                                    1  -  GLOBAL                                                    +
+====================================================================================================================+


+--------------------------------------------------------------------------------------------------------------------+
+                                             1.1  -  Experiment Summary                                             +
+--------------------------------------------------------------------------------------------------------------------+

  Application:			/beegfs/hackathon/users/eoseret/qaas_runs_test/178-663-0009/intel/TeaLeaf/run/binaries/gcc_3/exec
  Timestamp:			2026-08-13 17:05:07
  Universal Timestamp:		1786633507
  Experiment Type:		MPI; OpenMP; Throughput; 
  Machine:			isix06.benchmarkcenter.megware.com
  Architecture:			x86_64
  Micro Architecture:		GRANITE_RAPIDS
  Model Name:			Intel(R) Xeon(R) 6972P
  Cache Size:			491520 KB
  Number of Cores:		96
  OS Version:			Linux 5.14.0-687.31.1.el9_8.x86_64 #1 SMP PREEMPT_DYNAMIC Sat Aug 1 05:38:01 EDT 2026
  Compilation Options:		
		exec: GNU C++17 15.1.0 -march=graniterapids -mprefer-vector-width=256 -g -O3 -O3 -std=c++17 -funroll-loops -ffast-math -fno-omit-frame-pointer -fcf-protection=none -fopenmp 
  Number of processes observed:	6
  Number of threads observed:	192
  MAQAO version:		2026.1.0
  MAQAO build:			6d1be1d51c1e63266254997eb301734a7264775d::20260810-150026




+--------------------------------------------------------------------------------------------------------------------+
+                                               1.2  -  Global Metrics                                               +
+--------------------------------------------------------------------------------------------------------------------+

  Total Time:				24.47 s
  Max (Thread Active Time):		24.13 s
  Average Active Time:			23.94 s
  Activity Ratio:			99.5 %
  Average number of active threads:	187.896
  Affinity Stability:			99.5 %
  Time spent in analyzed loops:		66.3 %
  Time spent in analyzed innermost loops: 65.8 %
  Time spent in user code:		66.7 %
  Compilation Options Score:		100
  Array Access Efficiency:		98.0 %

   Potential Speedups
  ----------------------------------------------------
  Perfect Flow Complexity:		1.00
  Perfect OpenMP/MPI/Pthread/TBB:	1.27
  Perfect OpenMP/MPI/Pthread/TBB + Load Distribution:	1.51
  If No Scalar Integer:
      Potential Speedup:		1.00
      Nb Loops to get 80%:		5
  If FP Vectorized:
      Potential Speedup:		1.19
      Nb Loops to get 80%:		2
  If Fully Vectorized:
      Potential Speedup:		1.50
      Nb Loops to get 80%:		2
  If Only FP Arithmetic:
      Potential Speedup:		1.02
      Nb Loops to get 80%:		1




+--------------------------------------------------------------------------------------------------------------------+
+                                             1.3  -  Potential Speedups                                             +
+--------------------------------------------------------------------------------------------------------------------+

  If No Scalar Integer:
      Number of loops   | 1      | 4      | 9      | 12     | 17     | 
      Cumulated Speedup | 1.0006 | 1.0019 | 1.0026 | 1.0026 | 1.0026 | 
  Top 5 loops:
    exec - 116:	1.0006
    exec - 113:	1.0012
    exec - 30:	1.0016
    exec - 23:	1.0019
    exec - 101:	1.0021

  If FP Vectorized:
      Number of loops   | 1      | 4      | 9      | 12     | 17     | 
      Cumulated Speedup | 1.1091 | 1.1918 | 1.1928 | 1.1928 | 1.1928 | 
  Top 5 loops:
    exec - 24:	1.1091
    exec - 29:	1.1853
    exec - 32:	1.1911
    exec - 30:	1.1918
    exec - 23:	1.1923

  If Fully Vectorized:
      Number of loops   | 1      | 4      | 9      | 12     | 17     | 
      Cumulated Speedup | 1.1952 | 1.4904 | 1.4971 | 1.4985 | 1.4992 | 
  Top 5 loops:
    exec - 29:	1.1952
    exec - 24:	1.4148
    exec - 32:	1.4882
    exec - 116:	1.4904
    exec - 113:	1.4925

  If Only FP Arithmetic:
      Number of loops   | 1      | 4      | 9      | 12     | 17     | 
      Cumulated Speedup | 1.0210 | 1.0227 | 1.0238 | 1.0241 | 1.0241 | 
  Top 5 loops:
    exec - 32:	1.021
    exec - 116:	1.0216
    exec - 113:	1.0222
    exec - 30:	1.0227
    exec - 23:	1.023



+====================================================================================================================+
+                                                   2  -  SUMMARY                                                    +
+====================================================================================================================+


+--------------------------------------------------------------------------------------------------------------------+
+                                             2.1  -  EXPERIMENT QUALITY                                             +
+--------------------------------------------------------------------------------------------------------------------+

  [4 / 4] Application profile is long enough (24.13 s)
To have good quality measurements, it is advised that the application profiling time is greater than 10 seconds.

  [3 / 3] Most of time spent in analyzed modules comes from functions with source/debug info
-g option gives access to debugging informations, such are source locations.

  [2.9981518182109 / 3] Most of time spent in analyzed modules (99.94%) comes from functions compiled with architecture specialization option
-march=graniterapids


  [3 / 3] Most of time spent in analyzed modules comes from functions with compilation options informations and
-fno-omit-frame-pointer is present
-fno-omit-frame-pointer improves the accuracy of callchains found during the application profiling.

  [3 / 3] Optimization level option is correctly used


  [3 / 3] Host configuration allows retrieval of all necessary metrics.


  [2 / 2] Application is correctly profiled ("Others" category represents 0.00 % of the execution time)
To have a representative profiling, it is advised that the category "Others" represents less than 20% of the execution
time in order to analyze as much as possible of the user code

  [1 / 1] Lstopo present. The Topology lstopo report will be generated.



+--------------------------------------------------------------------------------------------------------------------+
+                                                2.2  -  CODE QUALITY                                                +
+--------------------------------------------------------------------------------------------------------------------+

  [4 / 4] Enough time of the experiment time spent in analyzed loops (66.26%)
If the time spent in analyzed loops is less than 30%, standard loop optimizations will have a limited impact on
application performances.

  [4 / 4] Threads activity is good
On average, more than 97.86% of observed threads are actually active 

  [4 / 4] CPU activity is good
CPU cores are active 99.50% of time

  [4 / 4] Loop profile is not flat
At least one loop coverage is greater than 4% (32.66%), representing an hotspot for the application

  [4 / 4] Enough time of the experiment time spent in analyzed innermost loops (65.76%)
If the time spent in analyzed innermost loops is less than 15%, standard innermost loop optimizations such as
vectorisation will have a limited impact on application performances.

  [4 / 4] Affinity is good (99.54%)
Threads are not migrating to CPU cores: probably successfully pinned

  [3 / 3] Less than 10% (0.00%) is spend in BLAS1 operations
It could be more efficient to inline by hand BLAS1 operations

  [0 / 3] Too many functions do not use all threads
Functions running on a reduced number of threads (typically sequential code) cover at least 10% of application
walltime (29.43%). Check both "Max Inclusive Time Over Threads" and "Nb Threads" in Functions or Loops tabs and
consider parallelizing sequential regions or improving parallelization of regions running on a reduced number of
threads

  [3 / 3] Cumulative Outermost/In between loops coverage (0.50%) lower than cumulative innermost loop coverage (65.76%)
Having cumulative Outermost/In between loops coverage greater than cumulative innermost loop coverage will make loop
optimization more complex

  [2 / 2] Less than 10% (0.00%) is spend in BLAS2 operations
BLAS2 calls usually could make a poor cache usage and could benefit from inlining.

  [2 / 2] Less than 10% (0.00%) is spend in Libm/SVML (special functions)



+--------------------------------------------------------------------------------------------------------------------+
+                                               2.3  -  LOOPS OVERVIEW                                               +
+--------------------------------------------------------------------------------------------------------------------+

  Top 5 loops:
   + exec - 29  :
     analysis: Execution Time: 32 % - Vectorization Ratio: 100.00 % - Vector Length Use: 50.00 %
     Data Access Issues: 0
        [0] [SA] Inefficient vectorization: more than 10% of the vector loads instructions are unaligned - When
            allocating arrays, don’t forget to align them. There are 0 issues ( = arrays) costing 2 points each

   + exec - 24  :
     analysis: Execution Time: 25 % - Vectorization Ratio: 100.00 % - Vector Length Use: 50.00 %
     Data Access Issues: 0
        [0] [SA] Inefficient vectorization: more than 10% of the vector loads instructions are unaligned - When
            allocating arrays, don’t forget to align them. There are 0 issues ( = arrays) costing 2 points each

   + exec - 32  :
     analysis: Execution Time: 6 % - Vectorization Ratio: 100.00 % - Vector Length Use: 50.00 %
     Data Access Issues: 6
        [6] [SA] Presence of constant non unit stride data access - Use array restructuring, perform loop interchange
            or use gather instructions to lower a bit the cost. There are 3 issues ( = data accesses) costing 2 point
            each.
        [0] [SA] Inefficient vectorization: more than 10% of the vector loads instructions are unaligned - When
            allocating arrays, don’t forget to align them. There are 0 issues ( = arrays) costing 2 points each
     Vectorization Roadblocks: 6
        [6] [SA] Presence of constant non unit stride data access - Use array restructuring, perform loop interchange
            or use gather instructions to lower a bit the cost. There are 3 issues ( = data accesses) costing 2 point
            each.

   + exec - 116 :
     analysis: Execution Time: 0 % - Vectorization Ratio: 26.23 % - Vector Length Use: 19.93 %
     Loop Computation Issues: 2
        [2] [SA] Presence of a large number of scalar integer instructions - Simplify loop structure, perform loop
            splitting or perform unroll and jam. This issue costs 2 points.
     Control Flow Issues: 100
        [98] [SA] Too many paths (94 paths) - Simplify control structure. There are 94 issues ( = paths) costing 1
            point each with a malus of 4 points.
        [2] [SA] Non innermost loop (Outermost) - Collapse loop with innermost ones. This issue costs 2 points.
     Data Access Issues: 2
        [0] [SA] Inefficient vectorization: more than 10% of the vector loads instructions are unaligned - When
            allocating arrays, don’t forget to align them. There are 0 issues ( = arrays) costing 2 points each
        [2] [SA] More than 20% of the loads are accessing the stack - Perform loop splitting to decrease pressure on
            registers. This issue costs 2 points.
     Vectorization Roadblocks: 100
        [98] [SA] Too many paths (94 paths) - Simplify control structure. There are 94 issues ( = paths) costing 1
            point each with a malus of 4 points.
        [2] [SA] Non innermost loop (Outermost) - Collapse loop with innermost ones. This issue costs 2 points.

   + exec - 113 :
     analysis: Execution Time: 0 % - Vectorization Ratio: 25.81 % - Vector Length Use: 19.86 %
     Loop Computation Issues: 2
        [2] [SA] Presence of a large number of scalar integer instructions - Simplify loop structure, perform loop
            splitting or perform unroll and jam. This issue costs 2 points.
     Control Flow Issues: 100
        [98] [SA] Too many paths (94 paths) - Simplify control structure. There are 94 issues ( = paths) costing 1
            point each with a malus of 4 points.
        [2] [SA] Non innermost loop (Outermost) - Collapse loop with innermost ones. This issue costs 2 points.
     Data Access Issues: 2
        [0] [SA] Inefficient vectorization: more than 10% of the vector loads instructions are unaligned - When
            allocating arrays, don’t forget to align them. There are 0 issues ( = arrays) costing 2 points each
        [2] [SA] More than 20% of the loads are accessing the stack - Perform loop splitting to decrease pressure on
            registers. This issue costs 2 points.
     Vectorization Roadblocks: 100
        [98] [SA] Too many paths (94 paths) - Simplify control structure. There are 94 issues ( = paths) costing 1
            point each with a malus of 4 points.
        [2] [SA] Non innermost loop (Outermost) - Collapse loop with innermost ones. This issue costs 2 points.



+====================================================================================================================+
+                                                 3  -  APPLICATION                                                  +
+====================================================================================================================+


+--------------------------------------------------------------------------------------------------------------------+
+                                               3.1  -  Categorization                                               +
+--------------------------------------------------------------------------------------------------------------------+

   Category | IO     | Exe    | System  | Others  | Memory | String | MPI   | TBB   | OMP   | Pthread | Math  |
  ----------+--------+--------+---------+---------+--------+--------+-------+-------+-------+---------+-------+
   Time (%) | 0.00   | 66.73  | 0.05    | 0.00    | 0.00   | 0.05   | 0.36  | 0.00  | 32.80 | 0.00    | 0.00  |




+--------------------------------------------------------------------------------------------------------------------+
+                                          3.2  -  Function Based Profiling                                          +
+--------------------------------------------------------------------------------------------------------------------+

   Buckets                    | Nb Functions              | Coverage                  | Cumulated Coverage        |
  ----------------------------+---------------------------+---------------------------+---------------------------+
   > 8%                       | 4                         | 85.70                     | 85.70                     |
   4% to 8%                   | 1                         | 7.11                      | 92.81                     |
   2% to 4%                   | 2                         | 4.65                      | 97.46                     |
   1% to 2%                   | 0                         | 0.00                      | 97.46                     |
   0.5% to 1%                 | 1                         | 0.95                      | 98.41                     |
   0.25% to 0.5%              | 0                         | 0.00                      | 98.41                     |
   0.125% to 0.25%            | 4                         | 0.69                      | 99.10                     |
   < 0.125%                   | 100                       | 0.90                      | 100.00                    |




+--------------------------------------------------------------------------------------------------------------------+
+                                            3.3  -  Loop Based Profiling                                            +
+--------------------------------------------------------------------------------------------------------------------+

   Buckets                    | Nb Loops                  | Coverage                  | Cumulated Coverage        |
  ----------------------------+---------------------------+---------------------------+---------------------------+
   > 8%                       | 2                         | 58.63                     | 58.63                     |
   4% to 8%                   | 1                         | 6.98                      | 65.61                     |
   2% to 4%                   | 0                         | 0.00                      | 65.61                     |
   1% to 2%                   | 0                         | 0.00                      | 65.61                     |
   0.5% to 1%                 | 0                         | 0.00                      | 65.61                     |
   0.25% to 0.5%              | 0                         | 0.00                      | 65.61                     |
   0.125% to 0.25%            | 0                         | 0.00                      | 65.61                     |
   < 0.125%                   | 28                        | 0.15                      | 65.76                     |


+====================================================================================================================+
+                                                  4  -  FUNCTIONS                                                   +
+====================================================================================================================+


+--------------------------------------------------------------------------------------------------------------------+
+                                              4.1  -  Top 10 Functions                                              +
+--------------------------------------------------------------------------------------------------------------------+

   Function                                               | Module              | Coverage (%)   | Time (s)       |
  --------------------------------------------------------+---------------------+----------------+----------------+
   cg_calc_ur(int, int, int, double, double*, double*,... | exec                | 32.75          | 7.84           |
   cg_calc_w(int, int, int, double*, double const*, do... | exec                | 26.09          | 6.25           |
   gomp_barrier_wait_end                                  | libgomp.so.1.0.0    | 18.33          | 4.39           |
   gomp_team_barrier_wait_end                             | libgomp.so.1.0.0    | 8.54           | 2.04           |
   cg_calc_p(int, int, int, double, double*, double co... | exec                | 7.11           | 1.70           |
   gomp_barrier_wait                                      | libgomp.so.1.0.0    | 2.45           | 0.59           |
   gomp_team_barrier_wait_final                           | libgomp.so.1.0.0    | 2.20           | 0.53           |
   gomp_thread_start                                      | libgomp.so.1.0.0    | 0.95           | 0.23           |
   mca_part_persist_progress                              | libmpi.so.40.40.7   | 0.19           | 1.48           |
   gomp_team_start                                        | libgomp.so.1.0.0    | 0.19           | 1.46           |


+====================================================================================================================+
+                                                    5  -  LOOPS                                                     +
+====================================================================================================================+


+--------------------------------------------------------------------------------------------------------------------+
+                                                5.1  -  Top 10 Loops                                                +
+--------------------------------------------------------------------------------------------------------------------+

   Loop Id        | Module              | Source Location                                        | Coverage (%)   |
  ----------------+---------------------+--------------------------------------------------------+----------------+
   29             | exec                | cg.cpp:108-113                                         | 32.66          |
   24             | exec                | cg.cpp:86-90                                           | 25.98          |
   32             | exec                | cg.cpp:128-131                                         | 6.98           |
   116            | exec                | pack_halos.cpp:100-105                                 | 0.11           |
   113            | exec                | pack_halos.cpp:83-88                                   | 0.10           |
   30             | exec                | cg.cpp:125-131                                         | 0.10           |
   23             | exec                | cg.cpp:83-83,cg.cpp:86-90                              | 0.06           |
   101            | exec                | pack_halos.cpp:12-17                                   | 0.04           |
   28             | exec                | cg.cpp:105-105,cg.cpp:108-113                          | 0.04           |
   104            | exec                | pack_halos.cpp:30-35                                   | 0.04           |





+====================================================================================================================+
+                                                     6  -  CQA                                                      +
+====================================================================================================================+


+--------------------------------------------------------------------------------------------------------------------+
+                                                   6.1  -  Loops                                                    +
+--------------------------------------------------------------------------------------------------------------------+





      6.1.1  -  Loop 29 from exec
  ==========================================================================================================

The loop is defined in /beegfs/hackathon/users/eoseret/qaas_runs_test/178-663-0009/intel/TeaLeaf/build/TeaLeaf/src/omp/cg.cpp:108-113.

It is main loop of related source loop which is unrolled by 2 (including vectorization).

      6.1.1.1  -  Path 1
  ----------------------------------------------------------------------------------------------------------

18% of peak computational performance is used (6.00 out of 32.00 FLOP per cycle (GFLOPS @ 1GHz))

      6.1.1.1.1  -  Vectorization
  ----------------------------------------------------------------------------------------------------------

Your loop is vectorized, but using only 256 out of 512 bits (AVX/AVX2 instructions on AVX-512 processors).
<<image_4x64_512>>

Details
All SSE/AVX instructions are used in vector version (process two or more data elements in vector registers).


Workaround
Read the "512-bits vectorization" report at "Potential" confidence level.


      6.1.1.1.2  -  Execution units bottlenecks
  ----------------------------------------------------------------------------------------------------------

Found no such bottlenecks but see expert reports for more complex bottlenecks.




      6.1.1.1.3  -  512-bits vectorization
  ----------------------------------------------------------------------------------------------------------

On some x86 processors supporting 512-bits vectorization, compilers are often too conservative and limit vectorization to 256 bits. Performance can then be improved by enforcing 512-bits vectorization, especially with many vectorized and high trip count loops. 512-bits vectorization performance overhead (compared to 256-bits) is generally lower on newer processors.


Workaround
Recompile with -mprefer-vector-width=512


      6.1.1.1.4  -  FMA
  ----------------------------------------------------------------------------------------------------------

Detected 48 FMA (fused multiply-add) operations.




      6.1.1.1.5  -  Vector unaligned load/store instructions
  ----------------------------------------------------------------------------------------------------------

Detected 16 optimal vector unaligned load/store instructions.


Details
 - VMOVUPD: 16 occurrences<<list_path_1_vec_align_1>>


Workaround
Use vector aligned instructions:
 1) align your arrays on 64 bytes boundaries: replace { void *p = malloc (size); } with { void *p; posix_memalign (&p, 64, size); }.
 2) inform your compiler that your arrays are vector aligned: if array 'foo' is 64 bytes-aligned, define a pointer 'p_foo' as __builtin_assume_aligned (foo, 64) and use it instead of 'foo' in the loop.
<<image_vec_align>>


      6.1.1.1.6  -  Type of elements and instruction set
  ----------------------------------------------------------------------------------------------------------

12 AVX instructions are processing arithmetic or math operations on double precision FP elements in vector mode (four at a time).



      6.1.1.1.7  -  Matching between your loop (in the source code) and the binary loop
  ----------------------------------------------------------------------------------------------------------

The binary loop is composed of 96 FP arithmetical operations:
 - 48: addition or subtraction (all inside FMA instructions)
 - 48: multiply (all inside FMA instructions)
The binary loop is loading 512 bytes (64 double precision FP elements).
The binary loop is storing 256 bytes (32 double precision FP elements).


      6.1.1.1.8  -  Arithmetic intensity
  ----------------------------------------------------------------------------------------------------------

Arithmetic intensity is 0.12 FP operations per loaded or stored byte.







      6.1.2  -  Loop 24 from exec
  ==========================================================================================================

The loop is defined in /beegfs/hackathon/users/eoseret/qaas_runs_test/178-663-0009/intel/TeaLeaf/build/TeaLeaf/src/omp/cg.cpp:86-90.

It is main loop of related source loop which is unrolled by 4 (including vectorization).

      6.1.2.1  -  Path 1
  ----------------------------------------------------------------------------------------------------------

34% of peak computational performance is used (10.91 out of 32.00 FLOP per cycle (GFLOPS @ 1GHz))

      6.1.2.1.1  -  Vectorization
  ----------------------------------------------------------------------------------------------------------

Your loop is vectorized, but using only 256 out of 512 bits (AVX/AVX2 instructions on AVX-512 processors).
<<image_4x64_512>>

Details
All SSE/AVX instructions are used in vector version (process two or more data elements in vector registers).


Workaround
Read the "512-bits vectorization" report at "Potential" confidence level.


      6.1.2.1.2  -  Execution units bottlenecks
  ----------------------------------------------------------------------------------------------------------

Performance is limited by execution of FP multiply or FMA (fused multiply-add) operations (the FP multiply/FMA unit is a bottleneck).

By removing all these bottlenecks, you can lower the cost of an iteration from 11.00 to 8.00 cycles (1.38x speedup).


Workaround
Reduce the number of FP multiply/FMA instructions




      6.1.2.1.3  -  512-bits vectorization
  ----------------------------------------------------------------------------------------------------------

On some x86 processors supporting 512-bits vectorization, compilers are often too conservative and limit vectorization to 256 bits. Performance can then be improved by enforcing 512-bits vectorization, especially with many vectorized and high trip count loops. 512-bits vectorization performance overhead (compared to 256-bits) is generally lower on newer processors.


Workaround
Recompile with -mprefer-vector-width=512


      6.1.2.1.4  -  FMA
  ----------------------------------------------------------------------------------------------------------

Detected 32 FMA (fused multiply-add) operations.
Presence of both ADD/SUB and MUL operations.

Workaround
Try to change order in which elements are evaluated (using parentheses) in arithmetic expressions containing both ADD/SUB and MUL operations to enable your compiler to generate FMA instructions wherever possible.
Estimated speedup by perfect pairing: 1.22x.
For instance a + b*c is a valid FMA (MUL then ADD).
However (a+b)* c cannot be translated into an FMA (ADD then MUL).




      6.1.2.1.5  -  Vector unaligned load/store instructions
  ----------------------------------------------------------------------------------------------------------

Detected 10 optimal vector unaligned load/store instructions.


Details
 - VMOVUPD: 10 occurrences<<list_path_1_vec_align_1>>


Workaround
Use vector aligned instructions:
 1) align your arrays on 64 bytes boundaries: replace { void *p = malloc (size); } with { void *p; posix_memalign (&p, 64, size); }.
 2) inform your compiler that your arrays are vector aligned: if array 'foo' is 64 bytes-aligned, define a pointer 'p_foo' as __builtin_assume_aligned (foo, 64) and use it instead of 'foo' in the loop.
<<image_vec_align>>


      6.1.2.1.6  -  Type of elements and instruction set
  ----------------------------------------------------------------------------------------------------------

22 AVX instructions are processing arithmetic or math operations on double precision FP elements in vector mode (four at a time).



      6.1.2.1.7  -  Matching between your loop (in the source code) and the binary loop
  ----------------------------------------------------------------------------------------------------------

The binary loop is composed of 120 FP arithmetical operations:
 - 72: addition or subtraction (32 inside FMA instructions)
 - 48: multiply (32 inside FMA instructions)
The binary loop is loading 648 bytes (81 double precision FP elements).
The binary loop is storing 64 bytes (8 double precision FP elements).


      6.1.2.1.8  -  Arithmetic intensity
  ----------------------------------------------------------------------------------------------------------

Arithmetic intensity is 0.17 FP operations per loaded or stored byte.







      6.1.3  -  Loop 32 from exec
  ==========================================================================================================

The loop is defined in /beegfs/hackathon/users/eoseret/qaas_runs_test/178-663-0009/intel/TeaLeaf/build/TeaLeaf/src/omp/cg.cpp:128-131.

It is main loop of related source loop which is unrolled by 4 (including vectorization).

      6.1.3.1  -  Path 1
  ----------------------------------------------------------------------------------------------------------

35% of peak computational performance is used (11.29 out of 32.00 FLOP per cycle (GFLOPS @ 1GHz))

      6.1.3.1.1  -  Vectorization
  ----------------------------------------------------------------------------------------------------------

Your loop is vectorized, but using only 256 out of 512 bits (AVX/AVX2 instructions on AVX-512 processors).
<<image_4x64_512>>

Details
All SSE/AVX instructions are used in vector version (process two or more data elements in vector registers).


Workaround
Read the "512-bits vectorization" report at "Potential" confidence level.


      6.1.3.1.2  -  Execution units bottlenecks
  ----------------------------------------------------------------------------------------------------------

Found no such bottlenecks but see expert reports for more complex bottlenecks.




      6.1.3.1.3  -  512-bits vectorization
  ----------------------------------------------------------------------------------------------------------

On some x86 processors supporting 512-bits vectorization, compilers are often too conservative and limit vectorization to 256 bits. Performance can then be improved by enforcing 512-bits vectorization, especially with many vectorized and high trip count loops. 512-bits vectorization performance overhead (compared to 256-bits) is generally lower on newer processors.


Workaround
Recompile with -mprefer-vector-width=512


      6.1.3.1.4  -  FMA
  ----------------------------------------------------------------------------------------------------------

Detected 32 FMA (fused multiply-add) operations.




      6.1.3.1.5  -  Slow data structures access
  ----------------------------------------------------------------------------------------------------------

Detected data structures (typically arrays) that cannot be efficiently read/written

Details
 - Constant non-unit stride: 3 occurrence(s)
Non-unit stride (uncontiguous) accesses are not efficiently using data caches


Workaround
 - Try to reorganize arrays of structures to structures of arrays
 - Consider to permute loops (see vectorization gain report)



      6.1.3.1.6  -  Vector unaligned load/store instructions
  ----------------------------------------------------------------------------------------------------------

Detected 16 optimal vector unaligned load/store instructions.


Details
 - VMOVUPD: 16 occurrences<<list_path_1_vec_align_1>>


Workaround
Use vector aligned instructions:
 1) align your arrays on 64 bytes boundaries: replace { void *p = malloc (size); } with { void *p; posix_memalign (&p, 64, size); }.
 2) inform your compiler that your arrays are vector aligned: if array 'foo' is 64 bytes-aligned, define a pointer 'p_foo' as __builtin_assume_aligned (foo, 64) and use it instead of 'foo' in the loop.
<<image_vec_align>>


      6.1.3.1.7  -  Type of elements and instruction set
  ----------------------------------------------------------------------------------------------------------

8 AVX instructions are processing arithmetic or math operations on double precision FP elements in vector mode (four at a time).



      6.1.3.1.8  -  Matching between your loop (in the source code) and the binary loop
  ----------------------------------------------------------------------------------------------------------

The binary loop is composed of 64 FP arithmetical operations:
 - 32: addition or subtraction (all inside FMA instructions)
 - 32: multiply (all inside FMA instructions)
The binary loop is loading 512 bytes (64 double precision FP elements).
The binary loop is storing 256 bytes (32 double precision FP elements).


      6.1.3.1.9  -  Arithmetic intensity
  ----------------------------------------------------------------------------------------------------------

Arithmetic intensity is 0.08 FP operations per loaded or stored byte.







      6.1.4  -  Loop 116 from exec
  ==========================================================================================================

The loop is defined in /beegfs/hackathon/users/eoseret/qaas_runs_test/178-663-0009/intel/TeaLeaf/build/TeaLeaf/src/omp/pack_halos.cpp:100-105.

Warnings:
 - Non-innermost loop: analyzing only self part (ignoring child loops).
 - Ignoring paths for analysis
 - Too many paths. If you really need to analyze all of the 94 paths individually, rerun with max-paths=94
 - RecMII not computed since number of paths is unknown or > max_paths
 - Streams not analyzed since number of paths is unknown or > max_paths

Try to simplify control and/or increase the maximum number of paths per function/loop through the 'max-paths-nb' option.

This loop has 94 execution paths.

The presence of multiple execution paths is typically the main/first bottleneck.
Try to simplify control inside loop: ideally, try to remove all conditional expressions, for example by (if applicable):
 - hoisting them (moving them outside the loop)
 - turning them into conditional moves, MIN or MAX


Ex: if (x<0) x=0 => x = (x<0 ? 0 : x) (or MAX(0,x) after defining the corresponding macro)


      6.1.4.1  -  Path 1
  ----------------------------------------------------------------------------------------------------------

0% of peak computational performance is used (0.00 out of 64.00 FLOP per cycle (GFLOPS @ 1GHz))

      6.1.4.1.1  -  Code clean check
  ----------------------------------------------------------------------------------------------------------

Detected a slowdown caused by scalar integer instructions (typically used for address computation).
By removing them, you can lower the cost of an iteration from 27.00 to 12.00 cycles (2.25x speedup).

Workaround
 - Try to reorganize arrays of structures to structures of arrays
 - Consider to permute loops (see vectorization gain report)



      6.1.4.1.2  -  Vectorization
  ----------------------------------------------------------------------------------------------------------

Your loop is poorly vectorized.
Only 19% of vector register length is used (average across all SSE/AVX instructions).


Details
26% of SSE/AVX instructions are used in vector version (process two or more data elements in vector registers):
 - 36% of SSE/AVX loads are used in vector version.
 - 40% of SSE/AVX stores are used in vector version.
 - 0% of SSE/AVX addition or subtraction instructions are used in vector version.
 - 0% of SSE/AVX instructions that are not load, store, addition, subtraction nor multiply instructions are used in vector version.


Workaround
Read the "512-bits vectorization" report at "Potential" confidence level.


      6.1.4.1.3  -  Execution units bottlenecks
  ----------------------------------------------------------------------------------------------------------

Found no such bottlenecks but see expert reports for more complex bottlenecks.




      6.1.4.1.4  -  512-bits vectorization
  ----------------------------------------------------------------------------------------------------------

On some x86 processors supporting 512-bits vectorization, compilers are often too conservative and limit vectorization to 256 bits. Performance can then be improved by enforcing 512-bits vectorization, especially with many vectorized and high trip count loops. 512-bits vectorization performance overhead (compared to 256-bits) is generally lower on newer processors.


Workaround
Recompile with -mprefer-vector-width=512




      6.1.4.1.5  -  Vector unaligned load/store instructions
  ----------------------------------------------------------------------------------------------------------

Detected 14 optimal vector unaligned load/store instructions.


Details
 - VMOVUPD: 14 occurrences<<list_path_1_vec_align_1>>


Workaround
Use vector aligned instructions:
 1) align your arrays on 64 bytes boundaries: replace { void *p = malloc (size); } with { void *p; posix_memalign (&p, 64, size); }.
 2) inform your compiler that your arrays are vector aligned: if array 'foo' is 64 bytes-aligned, define a pointer 'p_foo' as __builtin_assume_aligned (foo, 64) and use it instead of 'foo' in the loop.
<<image_vec_align>>


      6.1.4.1.6  -  Conversion instructions
  ----------------------------------------------------------------------------------------------------------

Detected expensive conversion instructions.

Details
 - CLTQ: 2 occurrences<<list_path_1_cvt_1>>


Workaround
Avoid mixing data with different types. In particular, check if the type of constants is the same as array elements.


      6.1.4.1.7  -  Type of elements and instruction set
  ----------------------------------------------------------------------------------------------------------

No instructions are processing arithmetic or math operations on FP elements. This loop is probably writing/copying data or processing integer elements.


      6.1.4.1.8  -  Matching between your loop (in the source code) and the binary loop
  ----------------------------------------------------------------------------------------------------------

The binary loop does not contain any FP arithmetical operations.
The binary loop is loading 522 bytes.
The binary loop is storing 332 bytes.







      6.1.5  -  Loop 113 from exec
  ==========================================================================================================

The loop is defined in /beegfs/hackathon/users/eoseret/qaas_runs_test/178-663-0009/intel/TeaLeaf/build/TeaLeaf/src/omp/pack_halos.cpp:83-88.

Warnings:
 - Non-innermost loop: analyzing only self part (ignoring child loops).
 - Ignoring paths for analysis
 - Too many paths. If you really need to analyze all of the 94 paths individually, rerun with max-paths=94
 - RecMII not computed since number of paths is unknown or > max_paths
 - Streams not analyzed since number of paths is unknown or > max_paths

Try to simplify control and/or increase the maximum number of paths per function/loop through the 'max-paths-nb' option.

This loop has 94 execution paths.

The presence of multiple execution paths is typically the main/first bottleneck.
Try to simplify control inside loop: ideally, try to remove all conditional expressions, for example by (if applicable):
 - hoisting them (moving them outside the loop)
 - turning them into conditional moves, MIN or MAX


Ex: if (x<0) x=0 => x = (x<0 ? 0 : x) (or MAX(0,x) after defining the corresponding macro)


      6.1.5.1  -  Path 1
  ----------------------------------------------------------------------------------------------------------

0% of peak computational performance is used (0.00 out of 64.00 FLOP per cycle (GFLOPS @ 1GHz))

      6.1.5.1.1  -  Code clean check
  ----------------------------------------------------------------------------------------------------------

Detected a slowdown caused by scalar integer instructions (typically used for address computation).
By removing them, you can lower the cost of an iteration from 27.17 to 12.00 cycles (2.26x speedup).

Workaround
 - Try to reorganize arrays of structures to structures of arrays
 - Consider to permute loops (see vectorization gain report)



      6.1.5.1.2  -  Vectorization
  ----------------------------------------------------------------------------------------------------------

Your loop is poorly vectorized.
Only 19% of vector register length is used (average across all SSE/AVX instructions).


Details
25% of SSE/AVX instructions are used in vector version (process two or more data elements in vector registers):
 - 38% of SSE/AVX loads are used in vector version.
 - 40% of SSE/AVX stores are used in vector version.
 - 0% of SSE/AVX addition or subtraction instructions are used in vector version.
 - 0% of SSE/AVX instructions that are not load, store, addition, subtraction nor multiply instructions are used in vector version.


Workaround
Read the "512-bits vectorization" report at "Potential" confidence level.


      6.1.5.1.3  -  Execution units bottlenecks
  ----------------------------------------------------------------------------------------------------------

Found no such bottlenecks but see expert reports for more complex bottlenecks.




      6.1.5.1.4  -  512-bits vectorization
  ----------------------------------------------------------------------------------------------------------

On some x86 processors supporting 512-bits vectorization, compilers are often too conservative and limit vectorization to 256 bits. Performance can then be improved by enforcing 512-bits vectorization, especially with many vectorized and high trip count loops. 512-bits vectorization performance overhead (compared to 256-bits) is generally lower on newer processors.


Workaround
Recompile with -mprefer-vector-width=512




      6.1.5.1.5  -  Vector unaligned load/store instructions
  ----------------------------------------------------------------------------------------------------------

Detected 14 optimal vector unaligned load/store instructions.


Details
 - VMOVUPD: 14 occurrences<<list_path_1_vec_align_1>>


Workaround
Use vector aligned instructions:
 1) align your arrays on 64 bytes boundaries: replace { void *p = malloc (size); } with { void *p; posix_memalign (&p, 64, size); }.
 2) inform your compiler that your arrays are vector aligned: if array 'foo' is 64 bytes-aligned, define a pointer 'p_foo' as __builtin_assume_aligned (foo, 64) and use it instead of 'foo' in the loop.
<<image_vec_align>>


      6.1.5.1.6  -  Conversion instructions
  ----------------------------------------------------------------------------------------------------------

Detected expensive conversion instructions.

Details
 - CLTQ: 2 occurrences<<list_path_1_cvt_1>>


Workaround
Avoid mixing data with different types. In particular, check if the type of constants is the same as array elements.


      6.1.5.1.7  -  Type of elements and instruction set
  ----------------------------------------------------------------------------------------------------------

No instructions are processing arithmetic or math operations on FP elements. This loop is probably writing/copying data or processing integer elements.


      6.1.5.1.8  -  Matching between your loop (in the source code) and the binary loop
  ----------------------------------------------------------------------------------------------------------

The binary loop does not contain any FP arithmetical operations.
The binary loop is loading 532 bytes.
The binary loop is storing 332 bytes.







      6.1.6  -  Loop 30 from exec
  ==========================================================================================================

The loop is defined in /beegfs/hackathon/users/eoseret/qaas_runs_test/178-663-0009/intel/TeaLeaf/build/TeaLeaf/src/omp/cg.cpp:125-131.

Warnings:
 - Non-innermost loop: analyzing only self part (ignoring child loops).
 - Ignoring paths for analysis
 - Too many paths. If you really need to analyze all of the 94 paths individually, rerun with max-paths=94
 - RecMII not computed since number of paths is unknown or > max_paths
 - Streams not analyzed since number of paths is unknown or > max_paths

Try to simplify control and/or increase the maximum number of paths per function/loop through the 'max-paths-nb' option.

This loop has 94 execution paths.

The presence of multiple execution paths is typically the main/first bottleneck.
Try to simplify control inside loop: ideally, try to remove all conditional expressions, for example by (if applicable):
 - hoisting them (moving them outside the loop)
 - turning them into conditional moves, MIN or MAX


Ex: if (x<0) x=0 => x = (x<0 ? 0 : x) (or MAX(0,x) after defining the corresponding macro)


      6.1.6.1  -  Path 1
  ----------------------------------------------------------------------------------------------------------

8% of peak computational performance is used (2.69 out of 32.00 FLOP per cycle (GFLOPS @ 1GHz))

      6.1.6.1.1  -  Code clean check
  ----------------------------------------------------------------------------------------------------------

Detected a slowdown caused by scalar integer instructions (typically used for address computation).
By removing them, you can lower the cost of an iteration from 29.00 to 16.00 cycles (1.81x speedup).

Workaround
 - Try to reorganize arrays of structures to structures of arrays
 - Consider to permute loops (see vectorization gain report)



      6.1.6.1.2  -  Vectorization
  ----------------------------------------------------------------------------------------------------------

Your loop is poorly vectorized.
Only 22% of vector register length is used (average across all SSE/AVX instructions).


Details
32% of SSE/AVX instructions are used in vector version (process two or more data elements in vector registers):
 - 43% of SSE/AVX loads are used in vector version.
 - 40% of SSE/AVX stores are used in vector version.
 - 0% of SSE/AVX addition or subtraction instructions are used in vector version.
 - 47% of SSE/AVX fused multiply-add instructions are used in vector version.
 - 0% of SSE/AVX instructions that are not load, store, addition, subtraction nor multiply instructions are used in vector version.


Workaround
Read the "512-bits vectorization" report at "Potential" confidence level.


      6.1.6.1.3  -  Execution units bottlenecks
  ----------------------------------------------------------------------------------------------------------

Found no such bottlenecks but see expert reports for more complex bottlenecks.




      6.1.6.1.4  -  512-bits vectorization
  ----------------------------------------------------------------------------------------------------------

On some x86 processors supporting 512-bits vectorization, compilers are often too conservative and limit vectorization to 256 bits. Performance can then be improved by enforcing 512-bits vectorization, especially with many vectorized and high trip count loops. 512-bits vectorization performance overhead (compared to 256-bits) is generally lower on newer processors.


Workaround
Recompile with -mprefer-vector-width=512


      6.1.6.1.5  -  FMA
  ----------------------------------------------------------------------------------------------------------

Detected 39 FMA (fused multiply-add) operations.




      6.1.6.1.6  -  Complex instructions
  ----------------------------------------------------------------------------------------------------------

Detected COMPLEX INSTRUCTIONS.


Details
These instructions generate more than one micro-operation and only one of them can be decoded during a cycle and the extra micro-operations increase pressure on execution units.
 - ADD: 1 occurrences<<list_path_1_complex_1>>



      6.1.6.1.7  -  Vector unaligned load/store instructions
  ----------------------------------------------------------------------------------------------------------

Detected 14 optimal vector unaligned load/store instructions.


Details
 - VMOVUPD: 14 occurrences<<list_path_1_vec_align_1>>


Workaround
Use vector aligned instructions:
 1) align your arrays on 64 bytes boundaries: replace { void *p = malloc (size); } with { void *p; posix_memalign (&p, 64, size); }.
 2) inform your compiler that your arrays are vector aligned: if array 'foo' is 64 bytes-aligned, define a pointer 'p_foo' as __builtin_assume_aligned (foo, 64) and use it instead of 'foo' in the loop.
<<image_vec_align>>


      6.1.6.1.8  -  Type of elements and instruction set
  ----------------------------------------------------------------------------------------------------------

9 SSE or AVX instructions are processing arithmetic or math operations on double precision FP elements in scalar mode (one at a time).
1 SSE or AVX instructions are processing arithmetic or math operations on double precision FP elements in vector mode (two at a time).
7 AVX instructions are processing arithmetic or math operations on double precision FP elements in vector mode (four at a time).



      6.1.6.1.9  -  Matching between your loop (in the source code) and the binary loop
  ----------------------------------------------------------------------------------------------------------

The binary loop is composed of 78 FP arithmetical operations:
 - 39: addition or subtraction (all inside FMA instructions)
 - 39: multiply (all inside FMA instructions)
The binary loop is loading 708 bytes (88 double precision FP elements).
The binary loop is storing 324 bytes (40 double precision FP elements).


      6.1.6.1.10  -  Arithmetic intensity
  ----------------------------------------------------------------------------------------------------------

Arithmetic intensity is 0.08 FP operations per loaded or stored byte.







      6.1.7  -  Loop 23 from exec
  ==========================================================================================================

The loop is defined in /beegfs/hackathon/users/eoseret/qaas_runs_test/178-663-0009/intel/TeaLeaf/build/TeaLeaf/src/omp/cg.cpp:83-90.

Analyzed code is defined in /beegfs/hackathon/users/eoseret/qaas_runs_test/178-663-0009/intel/TeaLeaf/build/TeaLeaf/src/omp/cg.cpp:83,86-90.

Warnings:
 - Non-innermost loop: analyzing only self part (ignoring child loops).
 - Ignoring paths for analysis
 - Too many paths. Rerun with max-paths=14
 - RecMII not computed since number of paths is unknown or > max_paths
 - Streams not analyzed since number of paths is unknown or > max_paths

Try to simplify control and/or increase the maximum number of paths per function/loop through the 'max-paths-nb' option.

This loop has 14 execution paths.

The presence of multiple execution paths is typically the main/first bottleneck.
Try to simplify control inside loop: ideally, try to remove all conditional expressions, for example by (if applicable):
 - hoisting them (moving them outside the loop)
 - turning them into conditional moves, MIN or MAX


Ex: if (x<0) x=0 => x = (x<0 ? 0 : x) (or MAX(0,x) after defining the corresponding macro)


      6.1.7.1  -  Path 1
  ----------------------------------------------------------------------------------------------------------

10% of peak computational performance is used (3.44 out of 32.00 FLOP per cycle (GFLOPS @ 1GHz))

      6.1.7.1.1  -  Code clean check
  ----------------------------------------------------------------------------------------------------------

Detected a slowdown caused by scalar integer instructions (typically used for address computation).
By removing them, you can lower the cost of an iteration from 32.83 to 19.00 cycles (1.73x speedup).

Workaround
 - Try to reorganize arrays of structures to structures of arrays
 - Consider to permute loops (see vectorization gain report)



      6.1.7.1.2  -  Vectorization
  ----------------------------------------------------------------------------------------------------------

Your loop is partially vectorized.
Only 23% of vector register length is used (average across all SSE/AVX instructions).


Details
51% of SSE/AVX instructions are used in vector version (process two or more data elements in vector registers):
 - 52% of SSE/AVX loads are used in vector version.
 - 25% of SSE/AVX stores are used in vector version.
 - 65% of SSE/AVX addition or subtraction instructions are used in vector version.
 - 66% of SSE/AVX multiply instructions are used in vector version.
 - 66% of SSE/AVX fused multiply-add instructions are used in vector version.
 - 27% of SSE/AVX instructions that are not load, store, addition, subtraction nor multiply instructions are used in vector version.


Workaround
Read the "512-bits vectorization" report at "Potential" confidence level.


      6.1.7.1.3  -  Execution units bottlenecks
  ----------------------------------------------------------------------------------------------------------

Found no such bottlenecks but see expert reports for more complex bottlenecks.




      6.1.7.1.4  -  512-bits vectorization
  ----------------------------------------------------------------------------------------------------------

On some x86 processors supporting 512-bits vectorization, compilers are often too conservative and limit vectorization to 256 bits. Performance can then be improved by enforcing 512-bits vectorization, especially with many vectorized and high trip count loops. 512-bits vectorization performance overhead (compared to 256-bits) is generally lower on newer processors.


Workaround
Recompile with -mprefer-vector-width=512


      6.1.7.1.5  -  FMA
  ----------------------------------------------------------------------------------------------------------

Detected 28 FMA (fused multiply-add) operations.
Presence of both ADD/SUB and MUL operations.

Workaround
Try to change order in which elements are evaluated (using parentheses) in arithmetic expressions containing both ADD/SUB and MUL operations to enable your compiler to generate FMA instructions wherever possible.
Estimated speedup by perfect pairing: 1.03x.
For instance a + b*c is a valid FMA (MUL then ADD).
However (a+b)* c cannot be translated into an FMA (ADD then MUL).




      6.1.7.1.6  -  Complex instructions
  ----------------------------------------------------------------------------------------------------------

Detected COMPLEX INSTRUCTIONS.


Details
These instructions generate more than one micro-operation and only one of them can be decoded during a cycle and the extra micro-operations increase pressure on execution units.
 - ADD: 1 occurrences<<list_path_1_complex_1>>
 - INC: 2 occurrences<<list_path_1_complex_2>>
 - SETA: 4 occurrences<<list_path_1_complex_3>>



      6.1.7.1.7  -  Vector unaligned load/store instructions
  ----------------------------------------------------------------------------------------------------------

Detected 5 optimal vector unaligned load/store instructions.


Details
 - VMOVUPD: 5 occurrences<<list_path_1_vec_align_1>>


Workaround
Use vector aligned instructions:
 1) align your arrays on 64 bytes boundaries: replace { void *p = malloc (size); } with { void *p; posix_memalign (&p, 64, size); }.
 2) inform your compiler that your arrays are vector aligned: if array 'foo' is 64 bytes-aligned, define a pointer 'p_foo' as __builtin_assume_aligned (foo, 64) and use it instead of 'foo' in the loop.
<<image_vec_align>>


      6.1.7.1.8  -  Conversion instructions
  ----------------------------------------------------------------------------------------------------------

Detected expensive conversion instructions.

Details
 - CLTQ: 1 occurrences<<list_path_1_cvt_1>>


Workaround
Avoid mixing data with different types. In particular, check if the type of constants is the same as array elements.


      6.1.7.1.9  -  Type of elements and instruction set
  ----------------------------------------------------------------------------------------------------------

13 SSE or AVX instructions are processing arithmetic or math operations on double precision FP elements in scalar mode (one at a time).
14 SSE or AVX instructions are processing arithmetic or math operations on double precision FP elements in vector mode (two at a time).
11 AVX instructions are processing arithmetic or math operations on double precision FP elements in vector mode (four at a time).



      6.1.7.1.10  -  Matching between your loop (in the source code) and the binary loop
  ----------------------------------------------------------------------------------------------------------

The binary loop is composed of 113 FP arithmetical operations:
 - 71: addition or subtraction (28 inside FMA instructions)
 - 42: multiply (28 inside FMA instructions)
The binary loop is loading 726 bytes (90 double precision FP elements).
The binary loop is storing 84 bytes (10 double precision FP elements).


      6.1.7.1.11  -  Arithmetic intensity
  ----------------------------------------------------------------------------------------------------------

Arithmetic intensity is 0.14 FP operations per loaded or stored byte.







      6.1.8  -  Loop 101 from exec
  ==========================================================================================================

The loop is defined in /beegfs/hackathon/users/eoseret/qaas_runs_test/178-663-0009/intel/TeaLeaf/build/TeaLeaf/src/omp/pack_halos.cpp:12-17.

Warnings:
 - Non-innermost loop: analyzing only self part (ignoring child loops).
 - Ignoring paths for analysis
 - Too many paths. If you really need to analyze all of the 94 paths individually, rerun with max-paths=94
 - RecMII not computed since number of paths is unknown or > max_paths
 - Streams not analyzed since number of paths is unknown or > max_paths

Try to simplify control and/or increase the maximum number of paths per function/loop through the 'max-paths-nb' option.

This loop has 94 execution paths.

The presence of multiple execution paths is typically the main/first bottleneck.
Try to simplify control inside loop: ideally, try to remove all conditional expressions, for example by (if applicable):
 - hoisting them (moving them outside the loop)
 - turning them into conditional moves, MIN or MAX


Ex: if (x<0) x=0 => x = (x<0 ? 0 : x) (or MAX(0,x) after defining the corresponding macro)


      6.1.8.1  -  Path 1
  ----------------------------------------------------------------------------------------------------------

0% of peak computational performance is used (0.00 out of 64.00 FLOP per cycle (GFLOPS @ 1GHz))

      6.1.8.1.1  -  Code clean check
  ----------------------------------------------------------------------------------------------------------

Detected a slowdown caused by scalar integer instructions (typically used for address computation).
By removing them, you can lower the cost of an iteration from 27.17 to 12.00 cycles (2.26x speedup).

Workaround
 - Try to reorganize arrays of structures to structures of arrays
 - Consider to permute loops (see vectorization gain report)



      6.1.8.1.2  -  Vectorization
  ----------------------------------------------------------------------------------------------------------

Your loop is poorly vectorized.
Only 20% of vector register length is used (average across all SSE/AVX instructions).


Details
26% of SSE/AVX instructions are used in vector version (process two or more data elements in vector registers):
 - 38% of SSE/AVX loads are used in vector version.
 - 40% of SSE/AVX stores are used in vector version.
 - 0% of SSE/AVX addition or subtraction instructions are used in vector version.
 - 0% of SSE/AVX instructions that are not load, store, addition, subtraction nor multiply instructions are used in vector version.


Workaround
Read the "512-bits vectorization" report at "Potential" confidence level.


      6.1.8.1.3  -  Execution units bottlenecks
  ----------------------------------------------------------------------------------------------------------

Found no such bottlenecks but see expert reports for more complex bottlenecks.




      6.1.8.1.4  -  512-bits vectorization
  ----------------------------------------------------------------------------------------------------------

On some x86 processors supporting 512-bits vectorization, compilers are often too conservative and limit vectorization to 256 bits. Performance can then be improved by enforcing 512-bits vectorization, especially with many vectorized and high trip count loops. 512-bits vectorization performance overhead (compared to 256-bits) is generally lower on newer processors.


Workaround
Recompile with -mprefer-vector-width=512




      6.1.8.1.5  -  Vector unaligned load/store instructions
  ----------------------------------------------------------------------------------------------------------

Detected 14 optimal vector unaligned load/store instructions.


Details
 - VMOVUPD: 14 occurrences<<list_path_1_vec_align_1>>


Workaround
Use vector aligned instructions:
 1) align your arrays on 64 bytes boundaries: replace { void *p = malloc (size); } with { void *p; posix_memalign (&p, 64, size); }.
 2) inform your compiler that your arrays are vector aligned: if array 'foo' is 64 bytes-aligned, define a pointer 'p_foo' as __builtin_assume_aligned (foo, 64) and use it instead of 'foo' in the loop.
<<image_vec_align>>


      6.1.8.1.6  -  Conversion instructions
  ----------------------------------------------------------------------------------------------------------

Detected expensive conversion instructions.

Details
 - CLTQ: 2 occurrences<<list_path_1_cvt_1>>


Workaround
Avoid mixing data with different types. In particular, check if the type of constants is the same as array elements.


      6.1.8.1.7  -  Type of elements and instruction set
  ----------------------------------------------------------------------------------------------------------

No instructions are processing arithmetic or math operations on FP elements. This loop is probably writing/copying data or processing integer elements.


      6.1.8.1.8  -  Matching between your loop (in the source code) and the binary loop
  ----------------------------------------------------------------------------------------------------------

The binary loop does not contain any FP arithmetical operations.
The binary loop is loading 537 bytes.
The binary loop is storing 332 bytes.







      6.1.9  -  Loop 28 from exec
  ==========================================================================================================

The loop is defined in /beegfs/hackathon/users/eoseret/qaas_runs_test/178-663-0009/intel/TeaLeaf/build/TeaLeaf/src/omp/cg.cpp:105-113.

Analyzed code is defined in /beegfs/hackathon/users/eoseret/qaas_runs_test/178-663-0009/intel/TeaLeaf/build/TeaLeaf/src/omp/cg.cpp:105,108-113.

Warnings:
 - Non-innermost loop: analyzing only self part (ignoring child loops).
 - Ignoring paths for analysis
 - Too many paths. Rerun with max-paths=30
 - RecMII not computed since number of paths is unknown or > max_paths
 - Streams not analyzed since number of paths is unknown or > max_paths

Try to simplify control and/or increase the maximum number of paths per function/loop through the 'max-paths-nb' option.

This loop has 30 execution paths.

The presence of multiple execution paths is typically the main/first bottleneck.
Try to simplify control inside loop: ideally, try to remove all conditional expressions, for example by (if applicable):
 - hoisting them (moving them outside the loop)
 - turning them into conditional moves, MIN or MAX


Ex: if (x<0) x=0 => x = (x<0 ? 0 : x) (or MAX(0,x) after defining the corresponding macro)


      6.1.9.1  -  Path 1
  ----------------------------------------------------------------------------------------------------------

12% of peak computational performance is used (3.92 out of 32.00 FLOP per cycle (GFLOPS @ 1GHz))

      6.1.9.1.1  -  Code clean check
  ----------------------------------------------------------------------------------------------------------

Detected a slowdown caused by scalar integer instructions (typically used for address computation).
By removing them, you can lower the cost of an iteration from 25.00 to 11.33 cycles (2.21x speedup).

Workaround
 - Try to reorganize arrays of structures to structures of arrays
 - Consider to permute loops (see vectorization gain report)



      6.1.9.1.2  -  Vectorization
  ----------------------------------------------------------------------------------------------------------

Your loop is partially vectorized.
Only 25% of vector register length is used (average across all SSE/AVX instructions).


Details
50% of SSE/AVX instructions are used in vector version (process two or more data elements in vector registers):
 - 64% of SSE/AVX loads are used in vector version.
 - 72% of SSE/AVX stores are used in vector version.
 - 42% of SSE/AVX addition or subtraction instructions are used in vector version.
 - 80% of SSE/AVX fused multiply-add instructions are used in vector version.
 - 17% of SSE/AVX instructions that are not load, store, addition, subtraction nor multiply instructions are used in vector version.


Workaround
Read the "512-bits vectorization" report at "Potential" confidence level.


      6.1.9.1.3  -  Execution units bottlenecks
  ----------------------------------------------------------------------------------------------------------

Found no such bottlenecks but see expert reports for more complex bottlenecks.




      6.1.9.1.4  -  512-bits vectorization
  ----------------------------------------------------------------------------------------------------------

On some x86 processors supporting 512-bits vectorization, compilers are often too conservative and limit vectorization to 256 bits. Performance can then be improved by enforcing 512-bits vectorization, especially with many vectorized and high trip count loops. 512-bits vectorization performance overhead (compared to 256-bits) is generally lower on newer processors.


Workaround
Recompile with -mprefer-vector-width=512


      6.1.9.1.5  -  FMA
  ----------------------------------------------------------------------------------------------------------

Detected 45 FMA (fused multiply-add) operations.




      6.1.9.1.6  -  Complex instructions
  ----------------------------------------------------------------------------------------------------------

Detected COMPLEX INSTRUCTIONS.


Details
These instructions generate more than one micro-operation and only one of them can be decoded during a cycle and the extra micro-operations increase pressure on execution units.
 - SETA: 2 occurrences<<list_path_1_complex_1>>



      6.1.9.1.7  -  Vector unaligned load/store instructions
  ----------------------------------------------------------------------------------------------------------

Detected 12 optimal vector unaligned load/store instructions.


Details
 - VMOVUPD: 12 occurrences<<list_path_1_vec_align_1>>


Workaround
Use vector aligned instructions:
 1) align your arrays on 64 bytes boundaries: replace { void *p = malloc (size); } with { void *p; posix_memalign (&p, 64, size); }.
 2) inform your compiler that your arrays are vector aligned: if array 'foo' is 64 bytes-aligned, define a pointer 'p_foo' as __builtin_assume_aligned (foo, 64) and use it instead of 'foo' in the loop.
<<image_vec_align>>


      6.1.9.1.8  -  Conversion instructions
  ----------------------------------------------------------------------------------------------------------

Detected expensive conversion instructions.

Details
 - CLTQ: 1 occurrences<<list_path_1_cvt_1>>


Workaround
Avoid mixing data with different types. In particular, check if the type of constants is the same as array elements.


      6.1.9.1.9  -  Type of elements and instruction set
  ----------------------------------------------------------------------------------------------------------

5 SSE or AVX instructions are processing arithmetic or math operations on double precision FP elements in scalar mode (one at a time).
6 SSE or AVX instructions are processing arithmetic or math operations on double precision FP elements in vector mode (two at a time).
9 AVX instructions are processing arithmetic or math operations on double precision FP elements in vector mode (four at a time).



      6.1.9.1.10  -  Matching between your loop (in the source code) and the binary loop
  ----------------------------------------------------------------------------------------------------------

The binary loop is composed of 98 FP arithmetical operations:
 - 53: addition or subtraction (45 inside FMA instructions)
 - 45: multiply (all inside FMA instructions)
The binary loop is loading 577 bytes (72 double precision FP elements).
The binary loop is storing 248 bytes (31 double precision FP elements).


      6.1.9.1.11  -  Arithmetic intensity
  ----------------------------------------------------------------------------------------------------------

Arithmetic intensity is 0.12 FP operations per loaded or stored byte.







      6.1.10  -  Loop 104 from exec
  ==========================================================================================================

The loop is defined in /beegfs/hackathon/users/eoseret/qaas_runs_test/178-663-0009/intel/TeaLeaf/build/TeaLeaf/src/omp/pack_halos.cpp:30-35.

Warnings:
 - Non-innermost loop: analyzing only self part (ignoring child loops).
 - Ignoring paths for analysis
 - Too many paths. If you really need to analyze all of the 94 paths individually, rerun with max-paths=94
 - RecMII not computed since number of paths is unknown or > max_paths
 - Streams not analyzed since number of paths is unknown or > max_paths

Try to simplify control and/or increase the maximum number of paths per function/loop through the 'max-paths-nb' option.

This loop has 94 execution paths.

The presence of multiple execution paths is typically the main/first bottleneck.
Try to simplify control inside loop: ideally, try to remove all conditional expressions, for example by (if applicable):
 - hoisting them (moving them outside the loop)
 - turning them into conditional moves, MIN or MAX


Ex: if (x<0) x=0 => x = (x<0 ? 0 : x) (or MAX(0,x) after defining the corresponding macro)


      6.1.10.1  -  Path 1
  ----------------------------------------------------------------------------------------------------------

0% of peak computational performance is used (0.00 out of 64.00 FLOP per cycle (GFLOPS @ 1GHz))

      6.1.10.1.1  -  Code clean check
  ----------------------------------------------------------------------------------------------------------

Detected a slowdown caused by scalar integer instructions (typically used for address computation).
By removing them, you can lower the cost of an iteration from 27.00 to 12.00 cycles (2.25x speedup).

Workaround
 - Try to reorganize arrays of structures to structures of arrays
 - Consider to permute loops (see vectorization gain report)



      6.1.10.1.2  -  Vectorization
  ----------------------------------------------------------------------------------------------------------

Your loop is poorly vectorized.
Only 19% of vector register length is used (average across all SSE/AVX instructions).


Details
25% of SSE/AVX instructions are used in vector version (process two or more data elements in vector registers):
 - 38% of SSE/AVX loads are used in vector version.
 - 40% of SSE/AVX stores are used in vector version.
 - 0% of SSE/AVX addition or subtraction instructions are used in vector version.
 - 0% of SSE/AVX instructions that are not load, store, addition, subtraction nor multiply instructions are used in vector version.


Workaround
Read the "512-bits vectorization" report at "Potential" confidence level.


      6.1.10.1.3  -  Execution units bottlenecks
  ----------------------------------------------------------------------------------------------------------

Found no such bottlenecks but see expert reports for more complex bottlenecks.




      6.1.10.1.4  -  512-bits vectorization
  ----------------------------------------------------------------------------------------------------------

On some x86 processors supporting 512-bits vectorization, compilers are often too conservative and limit vectorization to 256 bits. Performance can then be improved by enforcing 512-bits vectorization, especially with many vectorized and high trip count loops. 512-bits vectorization performance overhead (compared to 256-bits) is generally lower on newer processors.


Workaround
Recompile with -mprefer-vector-width=512




      6.1.10.1.5  -  Vector unaligned load/store instructions
  ----------------------------------------------------------------------------------------------------------

Detected 14 optimal vector unaligned load/store instructions.


Details
 - VMOVUPD: 14 occurrences<<list_path_1_vec_align_1>>


Workaround
Use vector aligned instructions:
 1) align your arrays on 64 bytes boundaries: replace { void *p = malloc (size); } with { void *p; posix_memalign (&p, 64, size); }.
 2) inform your compiler that your arrays are vector aligned: if array 'foo' is 64 bytes-aligned, define a pointer 'p_foo' as __builtin_assume_aligned (foo, 64) and use it instead of 'foo' in the loop.
<<image_vec_align>>


      6.1.10.1.6  -  Conversion instructions
  ----------------------------------------------------------------------------------------------------------

Detected expensive conversion instructions.

Details
 - CLTQ: 2 occurrences<<list_path_1_cvt_1>>


Workaround
Avoid mixing data with different types. In particular, check if the type of constants is the same as array elements.


      6.1.10.1.7  -  Type of elements and instruction set
  ----------------------------------------------------------------------------------------------------------

No instructions are processing arithmetic or math operations on FP elements. This loop is probably writing/copying data or processing integer elements.


      6.1.10.1.8  -  Matching between your loop (in the source code) and the binary loop
  ----------------------------------------------------------------------------------------------------------

The binary loop does not contain any FP arithmetical operations.
The binary loop is loading 544 bytes.
The binary loop is storing 332 bytes.





[MAQAO] Info: STOP THE REPORT GENERATION
[MAQAO] Info: 
[MAQAO] Info: If your application produces files, they can be found in directory "/beegfs/hackathon/users/eoseret/qaas_runs_test/178-663-0009/intel/TeaLeaf/run/oneview_runs/compilers/gcc_3/oneview_run_1786633507"
[MAQAO] Info: 
