Linux syscall times

Linux perf trace -s gathers aggregate statistics on Linux system call times. The overhead is generally in the range of a few % and is production ready although it's best to gauge the overhead in a test environment.

  1. Install perf if it's not installed. To check if it's installed:
    perf version
  2. As root, start the following command and replace 60 with the number of seconds you want to run it for:
    nohup timeout -s INT 60 perf trace -s > diag_perftracesummary_$(hostname)_$(date +%Y%m%d_%H%M%S).txt 2>&1 &
  3. Wait for the timeout period or run jobs to see if it's still running or completed.
  4. After it completes, upload diag*.

Details
  1. If you want to use sudo:
    nohup sudo sh -c "timeout -s INT 60 perf trace -s > diag_perftracesummary_$(hostname)_$(date +%Y%m%d_%H%M%S).txt 2>&1" >nohup.out 2>&1 &
  2. Instead of a fixed time, if you want to run it indefinitely in the background and then kill it manually:
    1. As root, start the following command in the background:
      nohup perf trace -s > diag_perftracesummary_$(hostname)_$(date +%Y%m%d_%H%M%S).txt 2>&1 &
    2. Reproduce the problem.
    3. After the problem is reproduced, stop it:
      pkill -INT -f "perf trace"

Example output:

bash (1), 738 events, 0.0%

   syscall            calls  errors  total       min       avg       max       stddev
                                     (msec)    (msec)    (msec)    (msec)        (%)
   --------------- --------  ------ -------- --------- --------- ---------     ------
   pselect6              62      0  7268.012     0.000   117.226  1695.909     29.78%
   rt_sigaction         112      0     0.594     0.004     0.005     0.027      4.33%
   write                 46      0     0.432     0.005     0.009     0.022      4.93%
   [...]