brintos

brintos / linux-shallow public Read only

0
0
Text · 22.9 KiB · 7c66d81 Raw
628 lines · plain
1perf-report(1)2==============3 4NAME5----6perf-report - Read perf.data (created by perf record) and display the profile7 8SYNOPSIS9--------10[verse]11'perf report' [-i <file> | --input=file]12 13DESCRIPTION14-----------15This command displays the performance counter profile information recorded16via perf record.17 18OPTIONS19-------20-i::21--input=::22        Input file name. (default: perf.data unless stdin is a fifo)23 24-v::25--verbose::26        Be more verbose. (show symbol address, etc)27 28-q::29--quiet::30	Do not show any warnings or messages.  (Suppress -v)31 32-n::33--show-nr-samples::34	Show the number of samples for each symbol35 36--show-cpu-utilization::37        Show sample percentage for different cpu modes.38 39-T::40--threads::41	Show per-thread event counters.  The input data file should be recorded42	with -s option.43-c::44--comms=::45	Only consider symbols in these comms. CSV that understands46	file://filename entries.  This option will affect the percentage of47	the overhead column.  See --percentage for more info.48--pid=::49        Only show events for given process ID (comma separated list).50 51--tid=::52        Only show events for given thread ID (comma separated list).53-d::54--dsos=::55	Only consider symbols in these dsos. CSV that understands56	file://filename entries.  This option will affect the percentage of57	the overhead column.  See --percentage for more info.58-S::59--symbols=::60	Only consider these symbols. CSV that understands61	file://filename entries.  This option will affect the percentage of62	the overhead column.  See --percentage for more info.63 64--symbol-filter=::65	Only show symbols that match (partially) with this filter.66 67-U::68--hide-unresolved::69        Only display entries resolved to a symbol.70 71-s::72--sort=::73	Sort histogram entries by given key(s) - multiple keys can be specified74	in CSV format.  Following sort keys are available:75	pid, comm, dso, symbol, parent, cpu, socket, srcline, weight,76	local_weight, cgroup_id, addr.77 78	Each key has following meaning:79 80	- comm: command (name) of the task which can be read via /proc/<pid>/comm81	- pid: command and tid of the task82	- dso: name of library or module executed at the time of sample83	- dso_size: size of library or module executed at the time of sample84	- symbol: name of function executed at the time of sample85	- symbol_size: size of function executed at the time of sample86	- parent: name of function matched to the parent regex filter. Unmatched87	entries are displayed as "[other]".88	- cpu: cpu number the task ran at the time of sample89	- socket: processor socket number the task ran at the time of sample90	- srcline: filename and line number executed at the time of sample.  The91	DWARF debugging info must be provided.92	- srcfile: file name of the source file of the samples. Requires dwarf93	information.94	- weight: Event specific weight, e.g. memory latency or transaction95	abort cost. This is the global weight.96	- local_weight: Local weight version of the weight above.97	- cgroup_id: ID derived from cgroup namespace device and inode numbers.98	- cgroup: cgroup pathname in the cgroupfs.99	- transaction: Transaction abort flags.100	- overhead: Overhead percentage of sample101	- overhead_sys: Overhead percentage of sample running in system mode102	- overhead_us: Overhead percentage of sample running in user mode103	- overhead_guest_sys: Overhead percentage of sample running in system mode104	on guest machine105	- overhead_guest_us: Overhead percentage of sample running in user mode on106	guest machine107	- sample: Number of sample108	- period: Raw number of event count of sample109	- time: Separate the samples by time stamp with the resolution specified by110	--time-quantum (default 100ms). Specify with overhead and before it.111	- code_page_size: the code page size of sampled code address (ip)112	- ins_lat: Instruction latency in core cycles. This is the global instruction113	  latency114	- local_ins_lat: Local instruction latency version115	- p_stage_cyc: On powerpc, this presents the number of cycles spent in a116	  pipeline stage. And currently supported only on powerpc.117	- addr: (Full) virtual address of the sampled instruction118	- retire_lat: On X86, this reports pipeline stall of this instruction compared119	  to the previous instruction in cycles. And currently supported only on X86120	- simd: Flags describing a SIMD operation. "e" for empty Arm SVE predicate. "p" for partial Arm SVE predicate121	- type: Data type of sample memory access.122	- typeoff: Offset in the data type of sample memory access.123	- symoff: Offset in the symbol.124	- weight1: Average value of event specific weight (1st field of weight_struct).125	- weight2: Average value of event specific weight (2nd field of weight_struct).126	- weight3: Average value of event specific weight (3rd field of weight_struct).127 128	By default, comm, dso and symbol keys are used.129	(i.e. --sort comm,dso,symbol)130 131	If --branch-stack option is used, following sort keys are also132	available:133 134	- dso_from: name of library or module branched from135	- dso_to: name of library or module branched to136	- symbol_from: name of function branched from137	- symbol_to: name of function branched to138	- srcline_from: source file and line branched from139	- srcline_to: source file and line branched to140	- mispredict: "N" for predicted branch, "Y" for mispredicted branch141	- in_tx: branch in TSX transaction142	- abort: TSX transaction abort.143	- cycles: Cycles in basic block144 145	And default sort keys are changed to comm, dso_from, symbol_from, dso_to146	and symbol_to, see '--branch-stack'.147 148	When the sort key symbol is specified, columns "IPC" and "IPC Coverage"149	are enabled automatically. Column "IPC" reports the average IPC per function150	and column "IPC coverage" reports the percentage of instructions with151	sampled IPC in this function. IPC means Instruction Per Cycle. If it's low,152	it indicates there may be a performance bottleneck when the function is153	executed, such as a memory access bottleneck. If a function has high overhead154	and low IPC, it's worth further analyzing it to optimize its performance.155 156	If the --mem-mode option is used, the following sort keys are also available157	(incompatible with --branch-stack):158	symbol_daddr, dso_daddr, locked, tlb, mem, snoop, dcacheline, blocked.159 160	- symbol_daddr: name of data symbol being executed on at the time of sample161	- dso_daddr: name of library or module containing the data being executed162	on at the time of the sample163	- locked: whether the bus was locked at the time of the sample164	- tlb: type of tlb access for the data at the time of the sample165	- mem: type of memory access for the data at the time of the sample166	- snoop: type of snoop (if any) for the data at the time of the sample167	- dcacheline: the cacheline the data address is on at the time of the sample168	- phys_daddr: physical address of data being executed on at the time of sample169	- data_page_size: the data page size of data being executed on at the time of sample170	- blocked: reason of blocked load access for the data at the time of the sample171 172	And the default sort keys are changed to local_weight, mem, sym, dso,173	symbol_daddr, dso_daddr, snoop, tlb, locked, blocked, local_ins_lat,174	see '--mem-mode'.175 176	If the data file has tracepoint event(s), following (dynamic) sort keys177	are also available:178	trace, trace_fields, [<event>.]<field>[/raw]179 180	- trace: pretty printed trace output in a single column181	- trace_fields: fields in tracepoints in separate columns182	- <field name>: optional event and field name for a specific field183 184	The last form consists of event and field names.  If event name is185	omitted, it searches all events for matching field name.  The matched186	field will be shown only for the event has the field.  The event name187	supports substring match so user doesn't need to specify full subsystem188	and event name everytime.  For example, 'sched:sched_switch' event can189	be shortened to 'switch' as long as it's not ambiguous.  Also event can190	be specified by its index (starting from 1) preceded by the '%'.191	So '%1' is the first event, '%2' is the second, and so on.192 193	The field name can have '/raw' suffix which disables pretty printing194	and shows raw field value like hex numbers.  The --raw-trace option195	has the same effect for all dynamic sort keys.196 197	The default sort keys are changed to 'trace' if all events in the data198	file are tracepoint.199 200-F::201--fields=::202	Specify output field - multiple keys can be specified in CSV format.203	Following fields are available:204	overhead, overhead_sys, overhead_us, overhead_children, sample, period,205	weight1, weight2, weight3, ins_lat, p_stage_cyc and retire_lat.  The206	last 3 names are alias for the corresponding weights.  When the weight207	fields are used, they will show the average value of the weight.208 209	Also it can contain any sort key(s).210 211	By default, every sort keys not specified in -F will be appended212	automatically.213 214	If the keys starts with a prefix '+', then it will append the specified215        field(s) to the default field order. For example: perf report -F +period,sample.216 217-p::218--parent=<regex>::219        A regex filter to identify parent. The parent is a caller of this220	function and searched through the callchain, thus it requires callchain221	information recorded. The pattern is in the extended regex format and222	defaults to "\^sys_|^do_page_fault", see '--sort parent'.223 224-x::225--exclude-other::226        Only display entries with parent-match.227 228-w::229--column-widths=<width[,width...]>::230	Force each column width to the provided list, for large terminal231	readability.  0 means no limit (default behavior).232 233-t::234--field-separator=::235	Use a special separator character and don't pad with spaces, replacing236	all occurrences of this separator in symbol names (and other output)237	with a '.' character, that thus it's the only non valid separator.238 239-D::240--dump-raw-trace::241        Dump raw trace in ASCII.242 243--disable-order::244	Disable raw trace ordering.245 246-g::247--call-graph=<print_type,threshold[,print_limit],order,sort_key[,branch],value>::248        Display call chains using type, min percent threshold, print limit,249	call order, sort key, optional branch and value.  Note that ordering250	is not fixed so any parameter can be given in an arbitrary order.251	One exception is the print_limit which should be preceded by threshold.252 253	print_type can be either:254	- flat: single column, linear exposure of call chains.255	- graph: use a graph tree, displaying absolute overhead rates. (default)256	- fractal: like graph, but displays relative rates. Each branch of257		 the tree is considered as a new profiled object.258	- folded: call chains are displayed in a line, separated by semicolons259	- none: disable call chain display.260 261	threshold is a percentage value which specifies a minimum percent to be262	included in the output call graph.  Default is 0.5 (%).263 264	print_limit is only applied when stdio interface is used.  It's to limit265	number of call graph entries in a single hist entry.  Note that it needs266	to be given after threshold (but not necessarily consecutive).267	Default is 0 (unlimited).268 269	order can be either:270	- callee: callee based call graph.271	- caller: inverted caller based call graph.272	Default is 'caller' when --children is used, otherwise 'callee'.273 274	sort_key can be:275	- function: compare on functions (default)276	- address: compare on individual code addresses277	- srcline: compare on source filename and line number278 279	branch can be:280	- branch: include last branch information in callgraph when available.281	          Usually more convenient to use --branch-history for this.282 283	value can be:284	- percent: display overhead percent (default)285	- period: display event period286	- count: display event count287 288--children::289	Accumulate callchain of children to parent entry so that then can290	show up in the output.  The output will have a new "Children" column291	and will be sorted on the data.  It requires callchains are recorded.292	See the `overhead calculation' section for more details. Enabled by293	default, disable with --no-children.294 295--max-stack::296	Set the stack depth limit when parsing the callchain, anything297	beyond the specified depth will be ignored. This is a trade-off298	between information loss and faster processing especially for299	workloads that can have a very long callchain stack.300	Note that when using the --itrace option the synthesized callchain size301	will override this value if the synthesized callchain size is bigger.302 303	Default: 127304 305-G::306--inverted::307        alias for inverted caller based call graph.308 309--ignore-callees=<regex>::310        Ignore callees of the function(s) matching the given regex.311        This has the effect of collecting the callers of each such312        function into one place in the call-graph tree.313 314--pretty=<key>::315        Pretty printing style.  key: normal, raw316 317--stdio:: Use the stdio interface.318 319--stdio-color::320	'always', 'never' or 'auto', allowing configuring color output321	via the command line, in addition to via "color.ui" .perfconfig.322	Use '--stdio-color always' to generate color even when redirecting323	to a pipe or file. Using just '--stdio-color' is equivalent to324	using 'always'.325 326--tui:: Use the TUI interface, that is integrated with annotate and allows327        zooming into DSOs or threads, among other features. Use of --tui328	requires a tty, if one is not present, as when piping to other329	commands, the stdio interface is used.330 331--gtk:: Use the GTK2 interface.332 333-k::334--vmlinux=<file>::335        vmlinux pathname336 337--ignore-vmlinux::338	Ignore vmlinux files.339 340--kallsyms=<file>::341        kallsyms pathname342 343-m::344--modules::345        Load module symbols. WARNING: This should only be used with -k and346        a LIVE kernel.347 348-f::349--force::350        Don't do ownership validation.351 352--symfs=<directory>::353        Look for files with symbols relative to this directory.354 355-C::356--cpu:: Only report samples for the list of CPUs provided. Multiple CPUs can357	be provided as a comma-separated list with no space: 0,1. Ranges of358	CPUs are specified with -: 0-2. Default is to report samples on all359	CPUs.360 361-M::362--disassembler-style=:: Set disassembler style for objdump.363 364--source::365	Interleave source code with assembly code. Enabled by default,366	disable with --no-source.367 368--asm-raw::369	Show raw instruction encoding of assembly instructions.370 371--show-total-period:: Show a column with the sum of periods.372 373-I::374--show-info::375	Display extended information about the perf.data file. This adds376	information which may be very large and thus may clutter the display.377	It currently includes: cpu and numa topology of the host system.378 379-b::380--branch-stack::381	Use the addresses of sampled taken branches instead of the instruction382	address to build the histograms. To generate meaningful output, the383	perf.data file must have been obtained using perf record -b or384	perf record --branch-filter xxx where xxx is a branch filter option.385	perf report is able to auto-detect whether a perf.data file contains386	branch stacks and it will automatically switch to the branch view mode,387	unless --no-branch-stack is used.388 389--branch-history::390	Add the addresses of sampled taken branches to the callstack.391	This allows to examine the path the program took to each sample.392	The data collection must have used -b (or -j) and -g.393 394--addr2line=<path>::395        Path to addr2line binary.396 397--objdump=<path>::398        Path to objdump binary.399 400--prefix=PREFIX::401--prefix-strip=N::402	Remove first N entries from source file path names in executables403	and add PREFIX. This allows to display source code compiled on systems404	with different file system layout.405 406--group::407	Show event group information together. It forces group output also408	if there are no groups defined in data file.409 410--group-sort-idx::411	Sort the output by the event at the index n in group. If n is invalid,412	sort by the first event. It can support multiple groups with different413	amount of events. WARNING: This should be used on grouped events.414 415--demangle::416	Demangle symbol names to human readable form. It's enabled by default,417	disable with --no-demangle.418 419--demangle-kernel::420	Demangle kernel symbol names to human readable form (for C++ kernels).421 422--mem-mode::423	Use the data addresses of samples in addition to instruction addresses424	to build the histograms.  To generate meaningful output, the perf.data425	file must have been obtained using perf record -d -W and using a426	special event -e cpu/mem-loads/p or -e cpu/mem-stores/p. See427	'perf mem' for simpler access.428 429--percent-limit::430	Do not show entries which have an overhead under that percent.431	(Default: 0).  Note that this option also sets the percent limit (threshold)432	of callchains.  However the default value of callchain threshold is433	different than the default value of hist entries.  Please see the434	--call-graph option for details.435 436--percentage::437	Determine how to display the overhead percentage of filtered entries.438	Filters can be applied by --comms, --dsos and/or --symbols options and439	Zoom operations on the TUI (thread, dso, etc).440 441	"relative" means it's relative to filtered entries only so that the442	sum of shown entries will be always 100%.  "absolute" means it retains443	the original value before and after the filter is applied.444 445--header::446	Show header information in the perf.data file.  This includes447	various information like hostname, OS and perf version, cpu/mem448	info, perf command line, event list and so on.  Currently only449	--stdio output supports this feature.450 451--header-only::452	Show only perf.data header (forces --stdio).453 454--time::455	Only analyze samples within given time window: <start>,<stop>. Times456	have the format seconds.nanoseconds. If start is not given (i.e. time457	string is ',x.y') then analysis starts at the beginning of the file. If458	stop time is not given (i.e. time string is 'x.y,') then analysis goes459	to end of file. Multiple ranges can be separated by spaces, which460	requires the argument to be quoted e.g. --time "1234.567,1234.789 1235,"461 462	Also support time percent with multiple time ranges. Time string is463	'a%/n,b%/m,...' or 'a%-b%,c%-%d,...'.464 465	For example:466	Select the second 10% time slice:467 468	  perf report --time 10%/2469 470	Select from 0% to 10% time slice:471 472	  perf report --time 0%-10%473 474	Select the first and second 10% time slices:475 476	  perf report --time 10%/1,10%/2477 478	Select from 0% to 10% and 30% to 40% slices:479 480	  perf report --time 0%-10%,30%-40%481 482--switch-on EVENT_NAME::483	Only consider events after this event is found.484 485	This may be interesting to measure a workload only after some initialization486	phase is over, i.e. insert a perf probe at that point and then using this487	option with that probe.488 489--switch-off EVENT_NAME::490	Stop considering events after this event is found.491 492--show-on-off-events::493	Show the --switch-on/off events too. This has no effect in 'perf report' now494	but probably we'll make the default not to show the switch-on/off events495        on the --group mode and if there is only one event besides the off/on ones,496	go straight to the histogram browser, just like 'perf report' with no events497	explicitly specified does.498 499--itrace::500	Options for decoding instruction tracing data. The options are:501 502include::itrace.txt[]503 504	To disable decoding entirely, use --no-itrace.505 506--full-source-path::507	Show the full path for source files for srcline output.508 509--show-ref-call-graph::510	When multiple events are sampled, it may not be needed to collect511	callgraphs for all of them. The sample sites are usually nearby,512	and it's enough to collect the callgraphs on a reference event.513	So user can use "call-graph=no" event modifier to disable callgraph514	for other events to reduce the overhead.515	However, perf report cannot show callgraphs for the event which516	disable the callgraph.517	This option extends the perf report to show reference callgraphs,518	which collected by reference event, in no callgraph event.519 520--stitch-lbr::521	Show callgraph with stitched LBRs, which may have more complete522	callgraph. The perf.data file must have been obtained using523	perf record --call-graph lbr.524	Disabled by default. In common cases with call stack overflows,525	it can recreate better call stacks than the default lbr call stack526	output. But this approach is not foolproof. There can be cases527	where it creates incorrect call stacks from incorrect matches.528	The known limitations include exception handing such as529	setjmp/longjmp will have calls/returns not match.530 531--socket-filter::532	Only report the samples on the processor socket that match with this filter533 534--samples=N::535	Save N individual samples for each histogram entry to show context in perf536	report tui browser.537 538--raw-trace::539	When displaying traceevent output, do not use print fmt or plugins.540 541-H::542--hierarchy::543	Enable hierarchical output.  In the hierarchy mode, each sort key groups544	samples based on the criteria and then sub-divide it using the lower545	level sort key.546 547	For example:548	In normal output:549 550	  perf report -s dso,sym551	  # Overhead  Shared Object      Symbol552	      50.00%  [kernel.kallsyms]  [k] kfunc1553	      20.00%  perf               [.] foo554	      15.00%  [kernel.kallsyms]  [k] kfunc2555	      10.00%  perf               [.] bar556	       5.00%  libc.so            [.] libcall557 558	In hierarchy output:559 560	  perf report -s dso,sym --hierarchy561	  #   Overhead  Shared Object / Symbol562	      65.00%    [kernel.kallsyms]563	        50.00%    [k] kfunc1564	        15.00%    [k] kfunc2565	      30.00%    perf566	        20.00%    [.] foo567	        10.00%    [.] bar568	       5.00%    libc.so569	         5.00%    [.] libcall570 571--inline::572	If a callgraph address belongs to an inlined function, the inline stack573	will be printed. Each entry is function name or file/line. Enabled by574	default, disable with --no-inline.575 576--mmaps::577	Show --tasks output plus mmap information in a format similar to578	/proc/<PID>/maps.579 580	Please note that not all mmaps are stored, options affecting which ones581	are include 'perf record --data', for instance.582 583--ns::584	Show time stamps in nanoseconds.585 586--stats::587	Display overall events statistics without any further processing.588	(like the one at the end of the perf report -D command)589 590--tasks::591	Display monitored tasks stored in perf data. Displaying pid/tid/ppid592	plus the command string aligned to distinguish parent and child tasks.593 594--percent-type::595	Set annotation percent type from following choices:596	  global-period, local-period, global-hits, local-hits597 598	The local/global keywords set if the percentage is computed599	in the scope of the function (local) or the whole data (global).600	The period/hits keywords set the base the percentage is computed601	on - the samples period or the number of samples (hits).602 603--time-quantum::604	Configure time quantum for time sort key. Default 100ms.605	Accepts s, us, ms, ns units.606 607--total-cycles::608	When --total-cycles is specified, it supports sorting for all blocks by609	'Sampled Cycles%'. This is useful to concentrate on the globally hottest610	blocks. In output, there are some new columns:611 612	'Sampled Cycles%' - block sampled cycles aggregation / total sampled cycles613	'Sampled Cycles'  - block sampled cycles aggregation614	'Avg Cycles%'     - block average sampled cycles / sum of total block average615			    sampled cycles616	'Avg Cycles'      - block average sampled cycles617	'Branch Counter'  - block branch counter histogram (with -v showing the number)618 619--skip-empty::620	Do not print 0 results in the --stat output.621 622include::callchain-overhead-calculation.txt[]623 624SEE ALSO625--------626linkperf:perf-stat[1], linkperf:perf-annotate[1], linkperf:perf-record[1],627linkperf:perf-intel-pt[1]628