===== out_proj M=2048 shape=27 =====
ops=128.8 GOP  k=6144 n=5120 M=2048 shape=27 (128, 128, 128, 1, 4, 2, 8)
[2217] python3.12@127.0.0.1
  void exl3_i8_gemm_kernel<128, 128, 128, 1, 4, 2, 8>(const signed char *, const unsigned int *, const float *, __half *, int, int, int, int, int, int, float) (40, 16, 1)x(128, 1, 1), Context 1, Stream 7, Device 0, CC 8.9
    Section: GPU Speed Of Light Throughput
    ----------------------- ----------- ------------
    Metric Name             Metric Unit Metric Value
    ----------------------- ----------- ------------
    DRAM Frequency                  Ghz        10.24
    SM Frequency                    Ghz         2.20
    Elapsed Cycles                cycle       828998
    Memory Throughput                 %        46.41
    DRAM Throughput                   %         7.70
    Duration                         us       374.88
    L1/TEX Cache Throughput           %        22.64
    L2 Cache Throughput               %        46.41
    SM Active Cycles              cycle    814568.47
    Compute (SM) Throughput           %        59.79
    ----------------------- ----------- ------------
    OPT   This workload exhibits low compute throughput and memory bandwidth utilization relative to the peak           
          performance of this device. Achieved compute throughput and/or memory bandwidth below 60.0% of peak           
          typically indicate latency issues. Look at Scheduler Statistics and Warp State Statistics for potential       
          reasons.                                                                                                      
    Section: Compute Workload Analysis
    -------------------- ----------- ------------
    Metric Name          Metric Unit Metric Value
    -------------------- ----------- ------------
    Executed Ipc Active   inst/cycle         1.64
    Executed Ipc Elapsed  inst/cycle         1.63
    Issue Slots Busy               %        40.73
    Issued Ipc Active     inst/cycle         1.64
    SM Busy                        %        59.79
    -------------------- ----------- ------------
    INF   Tensor is the highest-utilized pipeline (59.8%) based on elapsed cycles in the workload, taking into account  
          the rates of its different instructions. It is the logical aggregation of individual tensor pipelines. It's   
          dominated by its Tensor (INT) sub-pipeline. It is well-utilized, but should not be a bottleneck.              
    Section: Memory Workload Analysis
    --------------------------------------- ----------- ------------
    Metric Name                             Metric Unit Metric Value
    --------------------------------------- ----------- ------------
    Local Memory Spilling Requests                                 0
    Local Memory Spilling Request Overhead            %            0
    L2 Sector Promotion Misses                        %            0
    Shared Memory Spilling Requests                                0
    Shared Memory Spilling Request Overhead           %            0
    Memory Throughput                           Gbyte/s        75.66
    Mem Busy                                          %        46.41
    Max Bandwidth                                     %        45.11
    L1/TEX Hit Rate                                   %         4.62
    L2 Persisting Size                            Mbyte        14.16
    L2 Compression Success Rate                       %            0
    L2 Compression Ratio                                           0
    L2 Compression Input Sectors                 sector            0
    L2 Hit Rate                                       %        96.36
    Mem Pipes Busy                                    %        20.95
    --------------------------------------- ----------- ------------
    Section: Scheduler Statistics
    ---------------------------- ----------- ------------
    Metric Name                  Metric Unit Metric Value
    ---------------------------- ----------- ------------
    One or More Eligible                   %        41.10
    Issued Warp Per Scheduler                        0.41
    No Eligible                            %        58.90
    Active Warps Per Scheduler          warp         1.84
    Eligible Warps Per Scheduler        warp         0.55
    ---------------------------- ----------- ------------
    OPT   Est. Local Speedup: 40.21%                                                                                    
          Every scheduler is capable of issuing one instruction per cycle, but for this workload each scheduler only    
          issues an instruction every 2.4 cycles. This might leave hardware resources underutilized and may lead to     
          less optimal performance. Out of the maximum of 12 warps per scheduler, this workload allocates an average    
          of 1.84 active warps per scheduler, but only an average of 0.55 warps were eligible per cycle. Eligible       
          warps are the subset of active warps that are ready to issue their next instruction. Every cycle with no      
          eligible warp results in no instruction being issued and the issue slot remains unused. To increase the       
          number of eligible warps, avoid possible load imbalances due to highly different execution durations per      
          warp. Reducing stalls indicated on the Warp State Statistics and Source Counters sections can help, too.      
    Section: Warp State Statistics
    ---------------------------------------- ----------- ------------
    Metric Name                              Metric Unit Metric Value
    ---------------------------------------- ----------- ------------
    Warp Cycles Per Issued Instruction             cycle         4.47
    Warp Cycles Per Executed Instruction           cycle         4.47
    Avg. Active Threads Per Warp                                   32
    Avg. Not Predicated Off Threads Per Warp                    31.60
    ---------------------------------------- ----------- ------------
    WRN   The optional metric smsp__pcsamp_sample_count could not be found. Collecting it as an additional metric could 
          enable the rule to provide more guidance.                                                                     
    ----- --------------------------------------------------------------------------------------------------------------
    OPT   Est. Speedup: 35.06%                                                                                          
          On average, each warp of this workload spends 1.6 cycles being stalled waiting for the execution pipe to be   
          available. This stall occurs when all active warps execute their next instruction on a specific,              
          oversubscribed math pipeline. Try to increase the number of active warps to hide the existent latency or try  
          changing the instruction mix to utilize all available pipelines in a more balanced way. This stall type       
          represents about 35.1% of the total average of 4.5 cycles between issuing two instructions.                   
    ----- --------------------------------------------------------------------------------------------------------------
    INF   Check the Warp Stall Sampling (All Samples) table for the top stall locations in your source based on         
          sampling data. The Profiling Guide                                                                            
          (https://docs.nvidia.com/nsight-compute/ProfilingGuide/index.html#metrics-reference) provides more details    
          on each stall reason.                                                                                         
    Section: Instruction Statistics
    ---------------------------------------- ----------- ------------
    Metric Name                              Metric Unit Metric Value
    ---------------------------------------- ----------- ------------
    Local Memory Spilling Requests                  byte            0
    Shared Memory Spilling Requests                 byte            0
    Avg. Executed Instructions Per Scheduler        inst       334780
    Executed Instructions                           inst    171407360
    Avg. Issued Instructions Per Scheduler          inst    334807.48
    Issued Instructions                             inst    171421432
    ---------------------------------------- ----------- ------------
    OPT   Est. Speedup: 5.477%                                                                                          
          This workload executes 0 fused and 368640 non-fused FP32 instructions. By converting pairs of non-fused       
          instructions to their fused (https://docs.nvidia.com/cuda/floating-point/#cuda-and-floating-point),           
          higher-throughput equivalent, the achieved FP32 performance could be increased by up to 50% (relative to its  
          current performance).                                                                                         
    Section: Launch Statistics
    -------------------------------- --------------- ---------------
    Metric Name                          Metric Unit    Metric Value
    -------------------------------- --------------- ---------------
    Block Size                                                   128
    Function Cache Configuration                     CachePreferNone
    Grid Size                                                    640
    Registers Per Thread             register/thread             201
    Shared Memory Configuration Size           Kbyte          102.40
    Driver Shared Memory Per Block       Kbyte/block            1.02
    Dynamic Shared Memory Per Block      Kbyte/block           49.15
    Static Shared Memory Per Block        byte/block               0
    # SMs                                         SM             128
    Stack Size                                                  1024
    Threads                                   thread           81920
    # TPCs                                                        64
    Enabled TPC IDs                                              all
    Uses Green Context                                             0
    Waves Per SM                                                2.50
    -------------------------------- --------------- ---------------
    OPT   Est. Speedup: 33.33%                                                                                          
          A wave of thread blocks is defined as the maximum number of blocks that can be executed in parallel on the    
          target GPU. The number of blocks in a wave depends on the number of multiprocessors and the theoretical       
          occupancy of the kernel. This kernel launch results in 2 full waves and a partial wave of 128 thread blocks.  
          Under the assumption of a uniform execution duration of all thread blocks, this partial wave may account for  
          up to 33.3% of the total runtime of this kernel. Try launching a grid with no partial wave. The overall       
          impact of this tail effect also lessens with the number of full waves executed for a grid. See the Hardware   
          Model (https://docs.nvidia.com/nsight-compute/ProfilingGuide/index.html#metrics-hw-model) description for     
          more details on launch configurations.                                                                        
    Section: Occupancy
    ------------------------------- ----------- ------------
    Metric Name                     Metric Unit Metric Value
    ------------------------------- ----------- ------------
    Block Limit SM                        block           24
    Block Limit Registers                 block            2
    Block Limit Shared Mem                block            2
    Block Limit Warps                     block           12
    Theoretical Active Warps per SM        warp            8
    Theoretical Occupancy                     %        16.67
    Achieved Occupancy                        %        15.30
    Achieved Active Warps Per SM           warp         7.35
    ------------------------------- ----------- ------------
    OPT   Est. Speedup: 40.21%                                                                                          
          The 2.00 theoretical warps per scheduler this kernel can issue according to its occupancy are below the       
          hardware maximum of 12. This kernel's theoretical occupancy (16.7%) is limited by the number of required      
          registers, and the required amount of shared memory.                                                          
===== out_proj M=8192 shape=27 =====
ops=515.4 GOP  k=6144 n=5120 M=8192 shape=27 (128, 128, 128, 1, 4, 2, 8)
[2786] python3.12@127.0.0.1
  void exl3_i8_gemm_kernel<128, 128, 128, 1, 4, 2, 8>(const signed char *, const unsigned int *, const float *, __half *, int, int, int, int, int, int, float) (40, 64, 1)x(128, 1, 1), Context 1, Stream 7, Device 0, CC 8.9
    Section: GPU Speed Of Light Throughput
    ----------------------- ----------- ------------
    Metric Name             Metric Unit Metric Value
    ----------------------- ----------- ------------
    DRAM Frequency                  Ghz        10.24
    SM Frequency                    Ghz         2.23
    Elapsed Cycles                cycle      3163946
    Memory Throughput                 %        48.86
    DRAM Throughput                   %         8.46
    Duration                         ms         1.42
    L1/TEX Cache Throughput           %        23.44
    L2 Cache Throughput               %        48.86
    SM Active Cycles              cycle   3147221.88
    Compute (SM) Throughput           %        62.14
    ----------------------- ----------- ------------
    OPT   Compute is more heavily utilized than Memory: Look at the Compute Workload Analysis section to see what the   
          compute pipelines are spending their time doing. Also, consider whether any computation is redundant and      
          could be reduced or moved to look-up tables.                                                                  
    Section: Compute Workload Analysis
    -------------------- ----------- ------------
    Metric Name          Metric Unit Metric Value
    -------------------- ----------- ------------
    Executed Ipc Active   inst/cycle         1.70
    Executed Ipc Elapsed  inst/cycle         1.69
    Issue Slots Busy               %        42.33
    Issued Ipc Active     inst/cycle         1.70
    SM Busy                        %        62.14
    -------------------- ----------- ------------
    OPT   Tensor is the highest-utilized pipeline (62.1%) based on elapsed cycles in the workload, taking into account  
          the rates of its different instructions. It is the logical aggregation of individual tensor pipelines. It's   
          dominated by its Tensor (INT) sub-pipeline. The pipeline is well-utilized, but might become a bottleneck if   
          more work is added. Based on the number of executed instructions, the highest utilized pipeline (48.6%) is    
          ALU. It executes integer and logic operations. Comparing the two, the overall pipeline utilization appears    
          to be caused by frequent, low-latency instructions. See the Profiling Guide                                   
          (https://docs.nvidia.com/nsight-compute/ProfilingGuide/index.html#metrics-decoder) or hover over the          
          pipeline name to understand the workloads handled by each pipeline. The Instruction Statistics section shows  
          the mix of executed instructions for this workload. Check the Warp State Statistics section for which         
          reasons cause warps to stall.                                                                                 
    Section: Memory Workload Analysis
    --------------------------------------- ----------- ------------
    Metric Name                             Metric Unit Metric Value
    --------------------------------------- ----------- ------------
    Local Memory Spilling Requests                                 0
    Local Memory Spilling Request Overhead            %            0
    L2 Sector Promotion Misses                        %            0
    Shared Memory Spilling Requests                                0
    Shared Memory Spilling Request Overhead           %            0
    Memory Throughput                           Gbyte/s        83.14
    Mem Busy                                          %        48.86
    Max Bandwidth                                     %        47.48
    L1/TEX Hit Rate                                   %         4.59
    L2 Persisting Size                            Mbyte        14.16
    L2 Compression Success Rate                       %            0
    L2 Compression Ratio                                           0
    L2 Compression Input Sectors                 sector            0
    L2 Hit Rate                                       %        97.88
    Mem Pipes Busy                                    %        21.77
    --------------------------------------- ----------- ------------
    Section: Scheduler Statistics
    ---------------------------- ----------- ------------
    Metric Name                  Metric Unit Metric Value
    ---------------------------- ----------- ------------
    One or More Eligible                   %        42.54
    Issued Warp Per Scheduler                        0.43
    No Eligible                            %        57.46
    Active Warps Per Scheduler          warp         1.96
    Eligible Warps Per Scheduler        warp         0.59
    ---------------------------- ----------- ------------
    OPT   Est. Local Speedup: 37.86%                                                                                    
          Every scheduler is capable of issuing one instruction per cycle, but for this workload each scheduler only    
          issues an instruction every 2.4 cycles. This might leave hardware resources underutilized and may lead to     
          less optimal performance. Out of the maximum of 12 warps per scheduler, this workload allocates an average    
          of 1.96 active warps per scheduler, but only an average of 0.59 warps were eligible per cycle. Eligible       
          warps are the subset of active warps that are ready to issue their next instruction. Every cycle with no      
          eligible warp results in no instruction being issued and the issue slot remains unused. To increase the       
          number of eligible warps, avoid possible load imbalances due to highly different execution durations per      
          warp. Reducing stalls indicated on the Warp State Statistics and Source Counters sections can help, too.      
    Section: Warp State Statistics
    ---------------------------------------- ----------- ------------
    Metric Name                              Metric Unit Metric Value
    ---------------------------------------- ----------- ------------
    Warp Cycles Per Issued Instruction             cycle         4.60
    Warp Cycles Per Executed Instruction           cycle         4.60
    Avg. Active Threads Per Warp                                   32
    Avg. Not Predicated Off Threads Per Warp                    31.60
    ---------------------------------------- ----------- ------------
    WRN   The optional metric smsp__pcsamp_sample_count could not be found. Collecting it as an additional metric could 
          enable the rule to provide more guidance.                                                                     
    ----- --------------------------------------------------------------------------------------------------------------
    OPT   Est. Speedup: 36.18%                                                                                          
          On average, each warp of this workload spends 1.7 cycles being stalled waiting for the execution pipe to be   
          available. This stall occurs when all active warps execute their next instruction on a specific,              
          oversubscribed math pipeline. Try to increase the number of active warps to hide the existent latency or try  
          changing the instruction mix to utilize all available pipelines in a more balanced way. This stall type       
          represents about 36.2% of the total average of 4.6 cycles between issuing two instructions.                   
    ----- --------------------------------------------------------------------------------------------------------------
    INF   Check the Warp Stall Sampling (All Samples) table for the top stall locations in your source based on         
          sampling data. The Profiling Guide                                                                            
          (https://docs.nvidia.com/nsight-compute/ProfilingGuide/index.html#metrics-reference) provides more details    
          on each stall reason.                                                                                         
    Section: Instruction Statistics
    ---------------------------------------- ----------- ------------
    Metric Name                              Metric Unit Metric Value
    ---------------------------------------- ----------- ------------
    Local Memory Spilling Requests                  byte            0
    Shared Memory Spilling Requests                 byte            0
    Avg. Executed Instructions Per Scheduler        inst      1339120
    Executed Instructions                           inst    685629440
    Avg. Issued Instructions Per Scheduler          inst   1339162.61
    Issued Instructions                             inst    685651255
    ---------------------------------------- ----------- ------------
    OPT   Est. Speedup: 5.692%                                                                                          
          This workload executes 0 fused and 1474560 non-fused FP32 instructions. By converting pairs of non-fused      
          instructions to their fused (https://docs.nvidia.com/cuda/floating-point/#cuda-and-floating-point),           
          higher-throughput equivalent, the achieved FP32 performance could be increased by up to 50% (relative to its  
          current performance).                                                                                         
    Section: Launch Statistics
    -------------------------------- --------------- ---------------
    Metric Name                          Metric Unit    Metric Value
    -------------------------------- --------------- ---------------
    Block Size                                                   128
    Function Cache Configuration                     CachePreferNone
    Grid Size                                                   2560
    Registers Per Thread             register/thread             201
    Shared Memory Configuration Size           Kbyte          102.40
    Driver Shared Memory Per Block       Kbyte/block            1.02
    Dynamic Shared Memory Per Block      Kbyte/block           49.15
    Static Shared Memory Per Block        byte/block               0
    # SMs                                         SM             128
    Stack Size                                                  1024
    Threads                                   thread          327680
    # TPCs                                                        64
    Enabled TPC IDs                                              all
    Uses Green Context                                             0
    Waves Per SM                                                  10
    -------------------------------- --------------- ---------------
    Section: Occupancy
    ------------------------------- ----------- ------------
    Metric Name                     Metric Unit Metric Value
    ------------------------------- ----------- ------------
    Block Limit SM                        block           24
    Block Limit Registers                 block            2
    Block Limit Shared Mem                block            2
    Block Limit Warps                     block           12
    Theoretical Active Warps per SM        warp            8
    Theoretical Occupancy                     %        16.67
    Achieved Occupancy                        %        16.32
    Achieved Active Warps Per SM           warp         7.83
    ------------------------------- ----------- ------------
    OPT   Est. Speedup: 37.86%                                                                                          
          The 2.00 theoretical warps per scheduler this kernel can issue according to its occupancy are below the       
          hardware maximum of 12. This kernel's theoretical occupancy (16.7%) is limited by the number of required      
          registers, and the required amount of shared memory.                                                          
===== out_proj M=16384 shape=27 =====
ops=1030.8 GOP  k=6144 n=5120 M=16384 shape=27 (128, 128, 128, 1, 4, 2, 8)
[2887] python3.12@127.0.0.1
  void exl3_i8_gemm_kernel<128, 128, 128, 1, 4, 2, 8>(const signed char *, const unsigned int *, const float *, __half *, int, int, int, int, int, int, float) (40, 128, 1)x(128, 1, 1), Context 1, Stream 7, Device 0, CC 8.9
    Section: GPU Speed Of Light Throughput
    ----------------------- ----------- ------------
    Metric Name             Metric Unit Metric Value
    ----------------------- ----------- ------------
    DRAM Frequency                  Ghz        10.24
    SM Frequency                    Ghz         2.22
    Elapsed Cycles                cycle      6302187
    Memory Throughput                 %        48.62
    DRAM Throughput                   %         9.10
    Duration                         ms         2.83
    L1/TEX Cache Throughput           %        23.57
    L2 Cache Throughput               %        48.62
    SM Active Cycles              cycle   6258540.70
    Compute (SM) Throughput           %        62.56
    ----------------------- ----------- ------------
    OPT   Compute is more heavily utilized than Memory: Look at the Compute Workload Analysis section to see what the   
          compute pipelines are spending their time doing. Also, consider whether any computation is redundant and      
          could be reduced or moved to look-up tables.                                                                  
    Section: Compute Workload Analysis
    -------------------- ----------- ------------
    Metric Name          Metric Unit Metric Value
    -------------------- ----------- ------------
    Executed Ipc Active   inst/cycle         1.71
    Executed Ipc Elapsed  inst/cycle         1.70
    Issue Slots Busy               %        42.61
    Issued Ipc Active     inst/cycle         1.71
    SM Busy                        %        62.56
    -------------------- ----------- ------------
    OPT   Tensor is the highest-utilized pipeline (62.6%) based on elapsed cycles in the workload, taking into account  
          the rates of its different instructions. It is the logical aggregation of individual tensor pipelines. It's   
          dominated by its Tensor (INT) sub-pipeline. The pipeline is well-utilized, but might become a bottleneck if   
          more work is added. Based on the number of executed instructions, the highest utilized pipeline (48.9%) is    
          ALU. It executes integer and logic operations. Comparing the two, the overall pipeline utilization appears    
          to be caused by frequent, low-latency instructions. See the Profiling Guide                                   
          (https://docs.nvidia.com/nsight-compute/ProfilingGuide/index.html#metrics-decoder) or hover over the          
          pipeline name to understand the workloads handled by each pipeline. The Instruction Statistics section shows  
          the mix of executed instructions for this workload. Check the Warp State Statistics section for which         
          reasons cause warps to stall.                                                                                 
    Section: Memory Workload Analysis
    --------------------------------------- ----------- ------------
    Metric Name                             Metric Unit Metric Value
    --------------------------------------- ----------- ------------
    Local Memory Spilling Requests                                 0
    Local Memory Spilling Request Overhead            %            0
    L2 Sector Promotion Misses                        %            0
    Shared Memory Spilling Requests                                0
    Shared Memory Spilling Request Overhead           %            0
    Memory Throughput                           Gbyte/s        89.42
    Mem Busy                                          %        48.62
    Max Bandwidth                                     %        47.25
    L1/TEX Hit Rate                                   %         4.60
    L2 Persisting Size                            Mbyte        14.16
    L2 Compression Success Rate                       %            0
    L2 Compression Ratio                                           0
    L2 Compression Input Sectors                 sector            0
    L2 Hit Rate                                       %        98.13
    Mem Pipes Busy                                    %        21.92
    --------------------------------------- ----------- ------------
    Section: Scheduler Statistics
    ---------------------------- ----------- ------------
    Metric Name                  Metric Unit Metric Value
    ---------------------------- ----------- ------------
    One or More Eligible                   %        42.79
    Issued Warp Per Scheduler                        0.43
    No Eligible                            %        57.21
    Active Warps Per Scheduler          warp         1.98
    Eligible Warps Per Scheduler        warp         0.60
    ---------------------------- ----------- ------------
    OPT   Est. Local Speedup: 37.44%                                                                                    
          Every scheduler is capable of issuing one instruction per cycle, but for this workload each scheduler only    
          issues an instruction every 2.3 cycles. This might leave hardware resources underutilized and may lead to     
          less optimal performance. Out of the maximum of 12 warps per scheduler, this workload allocates an average    
          of 1.98 active warps per scheduler, but only an average of 0.60 warps were eligible per cycle. Eligible       
          warps are the subset of active warps that are ready to issue their next instruction. Every cycle with no      
          eligible warp results in no instruction being issued and the issue slot remains unused. To increase the       
          number of eligible warps, avoid possible load imbalances due to highly different execution durations per      
          warp. Reducing stalls indicated on the Warp State Statistics and Source Counters sections can help, too.      
    Section: Warp State Statistics
    ---------------------------------------- ----------- ------------
    Metric Name                              Metric Unit Metric Value
    ---------------------------------------- ----------- ------------
    Warp Cycles Per Issued Instruction             cycle         4.62
    Warp Cycles Per Executed Instruction           cycle         4.62
    Avg. Active Threads Per Warp                                   32
    Avg. Not Predicated Off Threads Per Warp                    31.60
    ---------------------------------------- ----------- ------------
    WRN   The optional metric smsp__pcsamp_sample_count could not be found. Collecting it as an additional metric could 
          enable the rule to provide more guidance.                                                                     
    ----- --------------------------------------------------------------------------------------------------------------
    OPT   Est. Speedup: 36.38%                                                                                          
          On average, each warp of this workload spends 1.7 cycles being stalled waiting for the execution pipe to be   
          available. This stall occurs when all active warps execute their next instruction on a specific,              
          oversubscribed math pipeline. Try to increase the number of active warps to hide the existent latency or try  
          changing the instruction mix to utilize all available pipelines in a more balanced way. This stall type       
          represents about 36.4% of the total average of 4.6 cycles between issuing two instructions.                   
    ----- --------------------------------------------------------------------------------------------------------------
    INF   Check the Warp Stall Sampling (All Samples) table for the top stall locations in your source based on         
          sampling data. The Profiling Guide                                                                            
          (https://docs.nvidia.com/nsight-compute/ProfilingGuide/index.html#metrics-reference) provides more details    
          on each stall reason.                                                                                         
    Section: Instruction Statistics
    ---------------------------------------- ----------- ------------
    Metric Name                              Metric Unit Metric Value
    ---------------------------------------- ----------- ------------
    Local Memory Spilling Requests                  byte            0
    Shared Memory Spilling Requests                 byte            0
    Avg. Executed Instructions Per Scheduler        inst      2678240
    Executed Instructions                           inst   1371258880
    Avg. Issued Instructions Per Scheduler          inst   2678302.72
    Issued Instructions                             inst   1371290991
    ---------------------------------------- ----------- ------------
    OPT   Est. Speedup: 5.73%                                                                                           
          This workload executes 0 fused and 2949120 non-fused FP32 instructions. By converting pairs of non-fused      
          instructions to their fused (https://docs.nvidia.com/cuda/floating-point/#cuda-and-floating-point),           
          higher-throughput equivalent, the achieved FP32 performance could be increased by up to 50% (relative to its  
          current performance).                                                                                         
    Section: Launch Statistics
    -------------------------------- --------------- ---------------
    Metric Name                          Metric Unit    Metric Value
    -------------------------------- --------------- ---------------
    Block Size                                                   128
    Function Cache Configuration                     CachePreferNone
    Grid Size                                                   5120
    Registers Per Thread             register/thread             201
    Shared Memory Configuration Size           Kbyte          102.40
    Driver Shared Memory Per Block       Kbyte/block            1.02
    Dynamic Shared Memory Per Block      Kbyte/block           49.15
    Static Shared Memory Per Block        byte/block               0
    # SMs                                         SM             128
    Stack Size                                                  1024
    Threads                                   thread          655360
    # TPCs                                                        64
    Enabled TPC IDs                                              all
    Uses Green Context                                             0
    Waves Per SM                                                  20
    -------------------------------- --------------- ---------------
    Section: Occupancy
    ------------------------------- ----------- ------------
    Metric Name                     Metric Unit Metric Value
    ------------------------------- ----------- ------------
    Block Limit SM                        block           24
    Block Limit Registers                 block            2
    Block Limit Shared Mem                block            2
    Block Limit Warps                     block           12
    Theoretical Active Warps per SM        warp            8
    Theoretical Occupancy                     %        16.67
    Achieved Occupancy                        %        16.48
    Achieved Active Warps Per SM           warp         7.91
    ------------------------------- ----------- ------------
    OPT   Est. Speedup: 37.44%                                                                                          
          The 2.00 theoretical warps per scheduler this kernel can issue according to its occupancy are below the       
          hardware maximum of 12. This kernel's theoretical occupancy (16.7%) is limited by the number of required      
          registers, and the required amount of shared memory.                                                          
===== gate_proj M=2048 shape=27 =====
ops=365.1 GOP  k=5120 n=17408 M=2048 shape=27 (128, 128, 128, 1, 4, 2, 8)
[2988] python3.12@127.0.0.1
  void exl3_i8_gemm_kernel<128, 128, 128, 1, 4, 2, 8>(const signed char *, const unsigned int *, const float *, __half *, int, int, int, int, int, int, float) (136, 16, 1)x(128, 1, 1), Context 1, Stream 7, Device 0, CC 8.9
    Section: GPU Speed Of Light Throughput
    ----------------------- ----------- ------------
    Metric Name             Metric Unit Metric Value
    ----------------------- ----------- ------------
    DRAM Frequency                  Ghz        10.24
    SM Frequency                    Ghz         2.23
    Elapsed Cycles                cycle      2260722
    Memory Throughput                 %        48.10
    DRAM Throughput                   %        13.65
    Duration                         ms         1.01
    L1/TEX Cache Throughput           %        23.31
    L2 Cache Throughput               %        48.10
    SM Active Cycles              cycle   2241550.06
    Compute (SM) Throughput           %        61.80
    ----------------------- ----------- ------------
    OPT   Compute is more heavily utilized than Memory: Look at the Compute Workload Analysis section to see what the   
          compute pipelines are spending their time doing. Also, consider whether any computation is redundant and      
          could be reduced or moved to look-up tables.                                                                  
    Section: Compute Workload Analysis
    -------------------- ----------- ------------
    Metric Name          Metric Unit Metric Value
    -------------------- ----------- ------------
    Executed Ipc Active   inst/cycle         1.70
    Executed Ipc Elapsed  inst/cycle         1.69
    Issue Slots Busy               %        42.20
    Issued Ipc Active     inst/cycle         1.70
    SM Busy                        %        61.80
    -------------------- ----------- ------------
    OPT   Tensor is the highest-utilized pipeline (61.8%) based on elapsed cycles in the workload, taking into account  
          the rates of its different instructions. It is the logical aggregation of individual tensor pipelines. It's   
          dominated by its Tensor (INT) sub-pipeline. The pipeline is well-utilized, but might become a bottleneck if   
          more work is added. Based on the number of executed instructions, the highest utilized pipeline (48.4%) is    
          ALU. It executes integer and logic operations. Comparing the two, the overall pipeline utilization appears    
          to be caused by frequent, low-latency instructions. See the Profiling Guide                                   
          (https://docs.nvidia.com/nsight-compute/ProfilingGuide/index.html#metrics-decoder) or hover over the          
          pipeline name to understand the workloads handled by each pipeline. The Instruction Statistics section shows  
          the mix of executed instructions for this workload. Check the Warp State Statistics section for which         
          reasons cause warps to stall.                                                                                 
    Section: Memory Workload Analysis
    --------------------------------------- ----------- ------------
    Metric Name                             Metric Unit Metric Value
    --------------------------------------- ----------- ------------
    Local Memory Spilling Requests                                 0
    Local Memory Spilling Request Overhead            %            0
    L2 Sector Promotion Misses                        %            0
    Shared Memory Spilling Requests                                0
    Shared Memory Spilling Request Overhead           %            0
    Memory Throughput                           Gbyte/s       134.16
    Mem Busy                                          %        48.10
    Max Bandwidth                                     %        46.37
    L1/TEX Hit Rate                                   %         5.21
    L2 Persisting Size                            Mbyte        14.16
    L2 Compression Success Rate                       %            0
    L2 Compression Ratio                                           0
    L2 Compression Input Sectors                 sector            0
    L2 Hit Rate                                       %        95.77
    Mem Pipes Busy                                    %        21.73
    --------------------------------------- ----------- ------------
    Section: Scheduler Statistics
    ---------------------------- ----------- ------------
    Metric Name                  Metric Unit Metric Value
    ---------------------------- ----------- ------------
    One or More Eligible                   %        42.42
    Issued Warp Per Scheduler                        0.42
    No Eligible                            %        57.58
    Active Warps Per Scheduler          warp         1.95
    Eligible Warps Per Scheduler        warp         0.59
    ---------------------------- ----------- ------------
    OPT   Est. Local Speedup: 38.2%                                                                                     
          Every scheduler is capable of issuing one instruction per cycle, but for this workload each scheduler only    
          issues an instruction every 2.4 cycles. This might leave hardware resources underutilized and may lead to     
          less optimal performance. Out of the maximum of 12 warps per scheduler, this workload allocates an average    
          of 1.95 active warps per scheduler, but only an average of 0.59 warps were eligible per cycle. Eligible       
          warps are the subset of active warps that are ready to issue their next instruction. Every cycle with no      
          eligible warp results in no instruction being issued and the issue slot remains unused. To increase the       
          number of eligible warps, avoid possible load imbalances due to highly different execution durations per      
          warp. Reducing stalls indicated on the Warp State Statistics and Source Counters sections can help, too.      
    Section: Warp State Statistics
    ---------------------------------------- ----------- ------------
    Metric Name                              Metric Unit Metric Value
    ---------------------------------------- ----------- ------------
    Warp Cycles Per Issued Instruction             cycle         4.60
    Warp Cycles Per Executed Instruction           cycle         4.60
    Avg. Active Threads Per Warp                                   32
    Avg. Not Predicated Off Threads Per Warp                    31.60
    ---------------------------------------- ----------- ------------
    WRN   The optional metric smsp__pcsamp_sample_count could not be found. Collecting it as an additional metric could 
          enable the rule to provide more guidance.                                                                     
    ----- --------------------------------------------------------------------------------------------------------------
    OPT   Est. Speedup: 35.89%                                                                                          
          On average, each warp of this workload spends 1.7 cycles being stalled waiting for the execution pipe to be   
          available. This stall occurs when all active warps execute their next instruction on a specific,              
          oversubscribed math pipeline. Try to increase the number of active warps to hide the existent latency or try  
          changing the instruction mix to utilize all available pipelines in a more balanced way. This stall type       
          represents about 35.9% of the total average of 4.6 cycles between issuing two instructions.                   
    ----- --------------------------------------------------------------------------------------------------------------
    INF   Check the Warp Stall Sampling (All Samples) table for the top stall locations in your source based on         
          sampling data. The Profiling Guide                                                                            
          (https://docs.nvidia.com/nsight-compute/ProfilingGuide/index.html#metrics-reference) provides more details    
          on each stall reason.                                                                                         
    Section: Instruction Statistics
    ---------------------------------------- ----------- ------------
    Metric Name                              Metric Unit Metric Value
    ---------------------------------------- ----------- ------------
    Local Memory Spilling Requests                  byte            0
    Shared Memory Spilling Requests                 byte            0
    Avg. Executed Instructions Per Scheduler        inst       950980
    Executed Instructions                           inst    486901760
    Avg. Issued Instructions Per Scheduler          inst    951020.71
    Issued Instructions                             inst    486922603
    ---------------------------------------- ----------- ------------
    OPT   Est. Speedup: 5.676%                                                                                          
          This workload executes 0 fused and 1253376 non-fused FP32 instructions. By converting pairs of non-fused      
          instructions to their fused (https://docs.nvidia.com/cuda/floating-point/#cuda-and-floating-point),           
          higher-throughput equivalent, the achieved FP32 performance could be increased by up to 50% (relative to its  
          current performance).                                                                                         
    Section: Launch Statistics
    -------------------------------- --------------- ---------------
    Metric Name                          Metric Unit    Metric Value
    -------------------------------- --------------- ---------------
    Block Size                                                   128
    Function Cache Configuration                     CachePreferNone
    Grid Size                                                   2176
    Registers Per Thread             register/thread             201
    Shared Memory Configuration Size           Kbyte          102.40
    Driver Shared Memory Per Block       Kbyte/block            1.02
    Dynamic Shared Memory Per Block      Kbyte/block           49.15
    Static Shared Memory Per Block        byte/block               0
    # SMs                                         SM             128
    Stack Size                                                  1024
    Threads                                   thread          278528
    # TPCs                                                        64
    Enabled TPC IDs                                              all
    Uses Green Context                                             0
    Waves Per SM                                                8.50
    -------------------------------- --------------- ---------------
    Section: Occupancy
    ------------------------------- ----------- ------------
    Metric Name                     Metric Unit Metric Value
    ------------------------------- ----------- ------------
    Block Limit SM                        block           24
    Block Limit Registers                 block            2
    Block Limit Shared Mem                block            2
    Block Limit Warps                     block           12
    Theoretical Active Warps per SM        warp            8
    Theoretical Occupancy                     %        16.67
    Achieved Occupancy                        %        16.26
    Achieved Active Warps Per SM           warp         7.80
    ------------------------------- ----------- ------------
    OPT   Est. Speedup: 38.2%                                                                                           
          The 2.00 theoretical warps per scheduler this kernel can issue according to its occupancy are below the       
          hardware maximum of 12. This kernel's theoretical occupancy (16.7%) is limited by the number of required      
          registers, and the required amount of shared memory.                                                          
===== gate_proj M=8192 shape=27 =====
ops=1460.3 GOP  k=5120 n=17408 M=8192 shape=27 (128, 128, 128, 1, 4, 2, 8)
[3089] python3.12@127.0.0.1
  void exl3_i8_gemm_kernel<128, 128, 128, 1, 4, 2, 8>(const signed char *, const unsigned int *, const float *, __half *, int, int, int, int, int, int, float) (136, 64, 1)x(128, 1, 1), Context 1, Stream 7, Device 0, CC 8.9
    Section: GPU Speed Of Light Throughput
    ----------------------- ----------- ------------
    Metric Name             Metric Unit Metric Value
    ----------------------- ----------- ------------
    DRAM Frequency                  Ghz        10.24
    SM Frequency                    Ghz         2.23
    Elapsed Cycles                cycle      8918108
    Memory Throughput                 %        48.76
    DRAM Throughput                   %        15.89
    Duration                         ms         3.99
    L1/TEX Cache Throughput           %        23.55
    L2 Cache Throughput               %        48.76
    SM Active Cycles              cycle   8873682.73
    Compute (SM) Throughput           %        62.54
    ----------------------- ----------- ------------
    OPT   Compute is more heavily utilized than Memory: Look at the Compute Workload Analysis section to see what the   
          compute pipelines are spending their time doing. Also, consider whether any computation is redundant and      
          could be reduced or moved to look-up tables.                                                                  
    Section: Compute Workload Analysis
    -------------------- ----------- ------------
    Metric Name          Metric Unit Metric Value
    -------------------- ----------- ------------
    Executed Ipc Active   inst/cycle         1.71
    Executed Ipc Elapsed  inst/cycle         1.71
    Issue Slots Busy               %        42.71
    Issued Ipc Active     inst/cycle         1.71
    SM Busy                        %        62.54
    -------------------- ----------- ------------
    OPT   Tensor is the highest-utilized pipeline (62.5%) based on elapsed cycles in the workload, taking into account  
          the rates of its different instructions. It is the logical aggregation of individual tensor pipelines. It's   
          dominated by its Tensor (INT) sub-pipeline. The pipeline is well-utilized, but might become a bottleneck if   
          more work is added. Based on the number of executed instructions, the highest utilized pipeline (49.0%) is    
          ALU. It executes integer and logic operations. Comparing the two, the overall pipeline utilization appears    
          to be caused by frequent, low-latency instructions. See the Profiling Guide                                   
          (https://docs.nvidia.com/nsight-compute/ProfilingGuide/index.html#metrics-decoder) or hover over the          
          pipeline name to understand the workloads handled by each pipeline. The Instruction Statistics section shows  
          the mix of executed instructions for this workload. Check the Warp State Statistics section for which         
          reasons cause warps to stall.                                                                                 
    Section: Memory Workload Analysis
    --------------------------------------- ----------- ------------
    Metric Name                             Metric Unit Metric Value
    --------------------------------------- ----------- ------------
    Local Memory Spilling Requests                                 0
    Local Memory Spilling Request Overhead            %            0
    L2 Sector Promotion Misses                        %            0
    Shared Memory Spilling Requests                                0
    Shared Memory Spilling Request Overhead           %            0
    Memory Throughput                           Gbyte/s       156.23
    Mem Busy                                          %        48.76
    Max Bandwidth                                     %        47.01
    L1/TEX Hit Rate                                   %         5.20
    L2 Persisting Size                            Mbyte        14.16
    L2 Compression Success Rate                       %            0
    L2 Compression Ratio                                           0
    L2 Compression Input Sectors                 sector            0
    L2 Hit Rate                                       %        95.86
    Mem Pipes Busy                                    %        21.99
    --------------------------------------- ----------- ------------
    Section: Scheduler Statistics
    ---------------------------- ----------- ------------
    Metric Name                  Metric Unit Metric Value
    ---------------------------- ----------- ------------
    One or More Eligible                   %        42.87
    Issued Warp Per Scheduler                        0.43
    No Eligible                            %        57.13
    Active Warps Per Scheduler          warp         1.99
    Eligible Warps Per Scheduler        warp         0.60
    ---------------------------- ----------- ------------
    OPT   Est. Local Speedup: 37.46%                                                                                    
          Every scheduler is capable of issuing one instruction per cycle, but for this workload each scheduler only    
          issues an instruction every 2.3 cycles. This might leave hardware resources underutilized and may lead to     
          less optimal performance. Out of the maximum of 12 warps per scheduler, this workload allocates an average    
          of 1.99 active warps per scheduler, but only an average of 0.60 warps were eligible per cycle. Eligible       
          warps are the subset of active warps that are ready to issue their next instruction. Every cycle with no      
          eligible warp results in no instruction being issued and the issue slot remains unused. To increase the       
          number of eligible warps, avoid possible load imbalances due to highly different execution durations per      
          warp. Reducing stalls indicated on the Warp State Statistics and Source Counters sections can help, too.      
    Section: Warp State Statistics
    ---------------------------------------- ----------- ------------
    Metric Name                              Metric Unit Metric Value
    ---------------------------------------- ----------- ------------
    Warp Cycles Per Issued Instruction             cycle         4.63
    Warp Cycles Per Executed Instruction           cycle         4.63
    Avg. Active Threads Per Warp                                   32
    Avg. Not Predicated Off Threads Per Warp                    31.60
    ---------------------------------------- ----------- ------------
    WRN   The optional metric smsp__pcsamp_sample_count could not be found. Collecting it as an additional metric could 
          enable the rule to provide more guidance.                                                                     
    ----- --------------------------------------------------------------------------------------------------------------
    OPT   Est. Speedup: 36.18%                                                                                          
          On average, each warp of this workload spends 1.7 cycles being stalled waiting for the execution pipe to be   
          available. This stall occurs when all active warps execute their next instruction on a specific,              
          oversubscribed math pipeline. Try to increase the number of active warps to hide the existent latency or try  
          changing the instruction mix to utilize all available pipelines in a more balanced way. This stall type       
          represents about 36.2% of the total average of 4.6 cycles between issuing two instructions.                   
    ----- --------------------------------------------------------------------------------------------------------------
    INF   Check the Warp Stall Sampling (All Samples) table for the top stall locations in your source based on         
          sampling data. The Profiling Guide                                                                            
          (https://docs.nvidia.com/nsight-compute/ProfilingGuide/index.html#metrics-reference) provides more details    
          on each stall reason.                                                                                         
    Section: Instruction Statistics
    ---------------------------------------- ----------- ------------
    Metric Name                              Metric Unit Metric Value
    ---------------------------------------- ----------- ------------
    Local Memory Spilling Requests                  byte            0
    Shared Memory Spilling Requests                 byte            0
    Avg. Executed Instructions Per Scheduler        inst      3803920
    Executed Instructions                           inst   1947607040
    Avg. Issued Instructions Per Scheduler          inst   3804011.09
    Issued Instructions                             inst   1947653676
    ---------------------------------------- ----------- ------------
    OPT   Est. Speedup: 5.745%                                                                                          
          This workload executes 0 fused and 5013504 non-fused FP32 instructions. By converting pairs of non-fused      
          instructions to their fused (https://docs.nvidia.com/cuda/floating-point/#cuda-and-floating-point),           
          higher-throughput equivalent, the achieved FP32 performance could be increased by up to 50% (relative to its  
          current performance).                                                                                         
    Section: Launch Statistics
    -------------------------------- --------------- ---------------
    Metric Name                          Metric Unit    Metric Value
    -------------------------------- --------------- ---------------
    Block Size                                                   128
    Function Cache Configuration                     CachePreferNone
    Grid Size                                                   8704
    Registers Per Thread             register/thread             201
    Shared Memory Configuration Size           Kbyte          102.40
    Driver Shared Memory Per Block       Kbyte/block            1.02
    Dynamic Shared Memory Per Block      Kbyte/block           49.15
    Static Shared Memory Per Block        byte/block               0
    # SMs                                         SM             128
    Stack Size                                                  1024
    Threads                                   thread         1114112
    # TPCs                                                        64
    Enabled TPC IDs                                              all
    Uses Green Context                                             0
    Waves Per SM                                                  34
    -------------------------------- --------------- ---------------
    Section: Occupancy
    ------------------------------- ----------- ------------
    Metric Name                     Metric Unit Metric Value
    ------------------------------- ----------- ------------
    Block Limit SM                        block           24
    Block Limit Registers                 block            2
    Block Limit Shared Mem                block            2
    Block Limit Warps                     block           12
    Theoretical Active Warps per SM        warp            8
    Theoretical Occupancy                     %        16.67
    Achieved Occupancy                        %        16.55
    Achieved Active Warps Per SM           warp         7.94
    ------------------------------- ----------- ------------
    OPT   Est. Speedup: 37.46%                                                                                          
          The 2.00 theoretical warps per scheduler this kernel can issue according to its occupancy are below the       
          hardware maximum of 12. This kernel's theoretical occupancy (16.7%) is limited by the number of required      
          registers, and the required amount of shared memory.                                                          
===== gate_proj M=16384 shape=27 =====
ops=2920.6 GOP  k=5120 n=17408 M=16384 shape=27 (128, 128, 128, 1, 4, 2, 8)
[3190] python3.12@127.0.0.1
  void exl3_i8_gemm_kernel<128, 128, 128, 1, 4, 2, 8>(const signed char *, const unsigned int *, const float *, __half *, int, int, int, int, int, int, float) (136, 128, 1)x(128, 1, 1), Context 1, Stream 7, Device 0, CC 8.9
    Section: GPU Speed Of Light Throughput
    ----------------------- ----------- ------------
    Metric Name             Metric Unit Metric Value
    ----------------------- ----------- ------------
    DRAM Frequency                  Ghz        10.24
    SM Frequency                    Ghz         2.22
    Elapsed Cycles                cycle     17828623
    Memory Throughput                 %        49.08
    DRAM Throughput                   %        16.25
    Duration                         ms         8.00
    L1/TEX Cache Throughput           %        23.60
    L2 Cache Throughput               %        49.08
    SM Active Cycles              cycle  17715464.60
    Compute (SM) Throughput           %        62.70
    ----------------------- ----------- ------------
    OPT   Compute is more heavily utilized than Memory: Look at the Compute Workload Analysis section to see what the   
          compute pipelines are spending their time doing. Also, consider whether any computation is redundant and      
          could be reduced or moved to look-up tables.                                                                  
    Section: Compute Workload Analysis
    -------------------- ----------- ------------
    Metric Name          Metric Unit Metric Value
    -------------------- ----------- ------------
    Executed Ipc Active   inst/cycle         1.72
    Executed Ipc Elapsed  inst/cycle         1.71
    Issue Slots Busy               %        42.82
    Issued Ipc Active     inst/cycle         1.72
    SM Busy                        %        62.70
    -------------------- ----------- ------------
    OPT   Tensor is the highest-utilized pipeline (62.7%) based on elapsed cycles in the workload, taking into account  
          the rates of its different instructions. It is the logical aggregation of individual tensor pipelines. It's   
          dominated by its Tensor (INT) sub-pipeline. The pipeline is well-utilized, but might become a bottleneck if   
          more work is added. Based on the number of executed instructions, the highest utilized pipeline (49.1%) is    
          ALU. It executes integer and logic operations. Comparing the two, the overall pipeline utilization appears    
          to be caused by frequent, low-latency instructions. See the Profiling Guide                                   
          (https://docs.nvidia.com/nsight-compute/ProfilingGuide/index.html#metrics-decoder) or hover over the          
          pipeline name to understand the workloads handled by each pipeline. The Instruction Statistics section shows  
          the mix of executed instructions for this workload. Check the Warp State Statistics section for which         
          reasons cause warps to stall.                                                                                 
    Section: Memory Workload Analysis
    --------------------------------------- ----------- ------------
    Metric Name                             Metric Unit Metric Value
    --------------------------------------- ----------- ------------
    Local Memory Spilling Requests                                 0
    Local Memory Spilling Request Overhead            %            0
    L2 Sector Promotion Misses                        %            0
    Shared Memory Spilling Requests                                0
    Shared Memory Spilling Request Overhead           %            0
    Memory Throughput                           Gbyte/s       159.75
    Mem Busy                                          %        49.08
    Max Bandwidth                                     %        47.31
    L1/TEX Hit Rate                                   %         5.20
    L2 Persisting Size                            Mbyte        14.16
    L2 Compression Success Rate                       %            0
    L2 Compression Ratio                                           0
    L2 Compression Input Sectors                 sector            0
    L2 Hit Rate                                       %        95.86
    Mem Pipes Busy                                    %        22.05
    --------------------------------------- ----------- ------------
    Section: Scheduler Statistics
    ---------------------------- ----------- ------------
    Metric Name                  Metric Unit Metric Value
    ---------------------------- ----------- ------------
    One or More Eligible                   %        42.94
    Issued Warp Per Scheduler                        0.43
    No Eligible                            %        57.06
    Active Warps Per Scheduler          warp         1.99
    Eligible Warps Per Scheduler        warp         0.60
    ---------------------------- ----------- ------------
    OPT   Est. Local Speedup: 37.3%                                                                                     
          Every scheduler is capable of issuing one instruction per cycle, but for this workload each scheduler only    
          issues an instruction every 2.3 cycles. This might leave hardware resources underutilized and may lead to     
          less optimal performance. Out of the maximum of 12 warps per scheduler, this workload allocates an average    
          of 1.99 active warps per scheduler, but only an average of 0.60 warps were eligible per cycle. Eligible       
          warps are the subset of active warps that are ready to issue their next instruction. Every cycle with no      
          eligible warp results in no instruction being issued and the issue slot remains unused. To increase the       
          number of eligible warps, avoid possible load imbalances due to highly different execution durations per      
          warp. Reducing stalls indicated on the Warp State Statistics and Source Counters sections can help, too.      
    Section: Warp State Statistics
    ---------------------------------------- ----------- ------------
    Metric Name                              Metric Unit Metric Value
    ---------------------------------------- ----------- ------------
    Warp Cycles Per Issued Instruction             cycle         4.64
    Warp Cycles Per Executed Instruction           cycle         4.64
    Avg. Active Threads Per Warp                                   32
    Avg. Not Predicated Off Threads Per Warp                    31.60
    ---------------------------------------- ----------- ------------
    WRN   The optional metric smsp__pcsamp_sample_count could not be found. Collecting it as an additional metric could 
          enable the rule to provide more guidance.                                                                     
    ----- --------------------------------------------------------------------------------------------------------------
    OPT   Est. Speedup: 36.23%                                                                                          
          On average, each warp of this workload spends 1.7 cycles being stalled waiting for the execution pipe to be   
          available. This stall occurs when all active warps execute their next instruction on a specific,              
          oversubscribed math pipeline. Try to increase the number of active warps to hide the existent latency or try  
          changing the instruction mix to utilize all available pipelines in a more balanced way. This stall type       
          represents about 36.2% of the total average of 4.6 cycles between issuing two instructions.                   
    ----- --------------------------------------------------------------------------------------------------------------
    INF   Check the Warp Stall Sampling (All Samples) table for the top stall locations in your source based on         
          sampling data. The Profiling Guide                                                                            
          (https://docs.nvidia.com/nsight-compute/ProfilingGuide/index.html#metrics-reference) provides more details    
          on each stall reason.                                                                                         
    Section: Instruction Statistics
    ---------------------------------------- ----------- ------------
    Metric Name                              Metric Unit Metric Value
    ---------------------------------------- ----------- ------------
    Local Memory Spilling Requests                  byte            0
    Shared Memory Spilling Requests                 byte            0
    Avg. Executed Instructions Per Scheduler        inst      7607840
    Executed Instructions                           inst   3895214080
    Avg. Issued Instructions Per Scheduler          inst      7608001
    Issued Instructions                             inst   3895296512
    ---------------------------------------- ----------- ------------
    OPT   Est. Speedup: 5.759%                                                                                          
          This workload executes 0 fused and 10027008 non-fused FP32 instructions. By converting pairs of non-fused     
          instructions to their fused (https://docs.nvidia.com/cuda/floating-point/#cuda-and-floating-point),           
          higher-throughput equivalent, the achieved FP32 performance could be increased by up to 50% (relative to its  
          current performance).                                                                                         
    Section: Launch Statistics
    -------------------------------- --------------- ---------------
    Metric Name                          Metric Unit    Metric Value
    -------------------------------- --------------- ---------------
    Block Size                                                   128
    Function Cache Configuration                     CachePreferNone
    Grid Size                                                  17408
    Registers Per Thread             register/thread             201
    Shared Memory Configuration Size           Kbyte          102.40
    Driver Shared Memory Per Block       Kbyte/block            1.02
    Dynamic Shared Memory Per Block      Kbyte/block           49.15
    Static Shared Memory Per Block        byte/block               0
    # SMs                                         SM             128
    Stack Size                                                  1024
    Threads                                   thread         2228224
    # TPCs                                                        64
    Enabled TPC IDs                                              all
    Uses Green Context                                             0
    Waves Per SM                                                  68
    -------------------------------- --------------- ---------------
    Section: Occupancy
    ------------------------------- ----------- ------------
    Metric Name                     Metric Unit Metric Value
    ------------------------------- ----------- ------------
    Block Limit SM                        block           24
    Block Limit Registers                 block            2
    Block Limit Shared Mem                block            2
    Block Limit Warps                     block           12
    Theoretical Active Warps per SM        warp            8
    Theoretical Occupancy                     %        16.67
    Achieved Occupancy                        %        16.60
    Achieved Active Warps Per SM           warp         7.97
    ------------------------------- ----------- ------------
    OPT   Est. Speedup: 37.3%                                                                                           
          The 2.00 theoretical warps per scheduler this kernel can issue according to its occupancy are below the       
          hardware maximum of 12. This kernel's theoretical occupancy (16.7%) is limited by the number of required      
          registers, and the required amount of shared memory.                                                          
