提交 · 3b6f9e5cb21964b7ce12bf81076f830885563ec8 · openanolis / cloud-kernel

14 1月, 2009 1 次提交

perf_counter: Add support for pinned and exclusive counter groups · 3b6f9e5c

由 Paul Mackerras 提交于 1月 14, 2009

Impact: New perf_counter features

A pinned counter group is one that the user wants to have on the CPU
whenever possible, i.e. whenever the associated task is running, for
a per-task group, or always for a per-cpu group.  If the system
cannot satisfy that, it puts the group into an error state where
it is not scheduled any more and reads from it return EOF (i.e. 0
bytes read).  The group can be released from error state and made
readable again using prctl(PR_TASK_PERF_COUNTERS_ENABLE).  When we
have finer-grained enable/disable controls on counters we'll be able
to reset the error state on individual groups.

An exclusive group is one that the user wants to be the only group
using the CPU performance monitor hardware whenever it is on.  The
counter group scheduler will not schedule an exclusive group if there
are already other groups on the CPU and will not schedule other groups
onto the CPU if there is an exclusive group scheduled (that statement
does not apply to groups containing only software counters, which can
always go on and which do not prevent an exclusive group from going on).
With an exclusive group, we will be able to let users program PMU
registers at a low level without the concern that those settings will
perturb other measurements.

Along the way this reorganizes things a little:
- is_software_counter() is moved to perf_counter.h.
- cpuctx->active_oncpu now records the number of hardware counters on
  the CPU, i.e. it now excludes software counters.  Nothing was reading
  cpuctx->active_oncpu before, so this change is harmless.
- A new cpuctx->exclusive field records whether we currently have an
  exclusive group on the CPU.
- counter_sched_out moves higher up in perf_counter.c and gets called
  from __perf_counter_remove_from_context and __perf_counter_exit_task,
  where we used to have essentially the same code.
- __perf_counter_sched_in now goes through the counter list twice, doing
  the pinned counters in the first loop and the non-pinned counters in
  the second loop, in order to give the pinned counters the best chance
  to be scheduled in.

Note that only a group leader can be exclusive or pinned, and that
attribute applies to the whole group.  This avoids some awkwardness in
some corner cases (e.g. where a group leader is closed and the other
group members get added to the context list).  If we want to relax that
restriction later, we can, and it is easier to relax a restriction than
to apply a new one.

This doesn't yet handle the case where a pinned counter is inherited
and goes into error state in the child - the error state is not
propagated up to the parent when the child exits, and arguably it
should.
Signed-off-by: NPaul Mackerras <paulus@samba.org>

3b6f9e5c

09 1月, 2009 1 次提交

perf_counter: Add optional hw_perf_group_sched_in arch function · 3cbed429

由 Paul Mackerras 提交于 1月 09, 2009

Impact: extend perf_counter infrastructure

This adds an optional hw_perf_group_sched_in() arch function that enables
a whole group of counters in one go.  It returns 1 if it added the group
successfully, 0 if it did nothing (and therefore the core needs to add
the counters individually), or a negative number if an error occurred.
It should add all the counters and enable any software counters in the
group, or else add none of them and return an error.

There are a couple of related changes/improvements in the group handling
here:

* As an optimization, group_sched_out() and group_sched_in() now check the
  state of the group leader, and do nothing if the leader is not active
  or disabled.

* We now call hw_perf_save_disable/hw_perf_restore around the complete
  set of counter enable/disable calls in __perf_counter_sched_in/out,
  to give the arch code the opportunity to defer updating the hardware
  state until the hw_perf_restore call if it wants.

* We no longer stop adding groups after we get to a group that has more
  than one counter.  We will ultimately add an option for a group to be
  exclusive.  The current code doesn't really implement exclusive groups
  anyway, since a group could end up going on with other counters that
  get added before it.
Signed-off-by: NPaul Mackerras <paulus@samba.org>

3cbed429

25 12月, 2008 1 次提交

perfcounters: include asm/perf_counter.h only if CONFIG_PERF_COUNTERS=y · e44aef58

由 Ingo Molnar 提交于 12月 25, 2008

Impact: build fix on ia64

KOSAKI Motohiro reported that -tip doesnt build on ia64 because
asm/perf_counter.h only exists on x86 for now. Fix it.
Reported-by: NKOSAKI Motohiro <kosaki.motohiro@jp.fujitsu.com>
Tested-by: NKOSAKI Motohiro <kosaki.motohiro@jp.fujitsu.com>
Acked-by: NKOSAKI Motohiro <kosaki.motohiro@jp.fujitsu.com>
Signed-off-by: NIngo Molnar <mingo@elte.hu>

e44aef58

23 12月, 2008 6 次提交

perfcounters: add PERF_COUNT_BUS_CYCLES · f650a672

由 Ingo Molnar 提交于 12月 23, 2008

Generalize "bus cycles" hw events - and map them to CPU_CLK_Unhalted.Ref
on x86. (which is a good enough approximation)
Signed-off-by: NIngo Molnar <mingo@elte.hu>

f650a672

perfcounters: remove ->nr_inherited · 8fe91e61

由 Ingo Molnar 提交于 12月 23, 2008

Impact: remove dead code

nr_inherited was not maintained correctly (not decremented) - and also
not used - remove it.
Signed-off-by: NIngo Molnar <mingo@elte.hu>

8fe91e61

perfcounters: enable lowlevel pmc code to schedule counters · 95cdd2e7

由 Ingo Molnar 提交于 12月 21, 2008

Allow lowlevel ->enable() op to return an error if a counter can not be
added. This can be used to handle counter constraints.
Signed-off-by: NIngo Molnar <mingo@elte.hu>

95cdd2e7

perfcounters: hw ops rename · 7671581f

由 Ingo Molnar 提交于 12月 17, 2008

Impact: rename field names

Shorten them.
Signed-off-by: NIngo Molnar <mingo@elte.hu>

7671581f

x86, perfcounters: prepare for fixed-mode PMCs · eb2b8618

由 Ingo Molnar 提交于 12月 17, 2008

Impact: refactor the x86 code for fixed-mode PMCs

Extend the data structures and rename the existing facilities
to allow for a 'generic' versus 'fixed' counter distinction.
Signed-off-by: NIngo Molnar <mingo@elte.hu>

eb2b8618

I
perfcounters: remove warnings · 8fb93313
由 Ingo Molnar 提交于 12月 23, 2008
```
Impact: remove debug checks
Signed-off-by: NIngo Molnar <mingo@elte.hu>
```
8fb93313

15 12月, 2008 4 次提交

perfcounters: add task migrations counter · 6c594c21

由 Ingo Molnar 提交于 12月 14, 2008

Impact: add new feature, new sw counter

Add a counter that counts the number of cross-CPU migrations a
task is suffering.
Signed-off-by: NIngo Molnar <mingo@elte.hu>

6c594c21

perfcounters: add context switch counter · 5d6a27d8

由 Ingo Molnar 提交于 12月 14, 2008

Impact: add new feature, new sw counter

Add a counter that counts the number of context-switches a task
is doing.
Signed-off-by: NIngo Molnar <mingo@elte.hu>

5d6a27d8

perfcounters: implement "counter inheritance" · 9b51f66d

由 Ingo Molnar 提交于 12月 12, 2008

Impact: implement new performance feature

Counter inheritance can be used to run performance counters in a workload,
transparently - and pipe back the counter results to the parent counter.

Inheritance for performance counters works the following way: when creating
a counter it can be marked with the .inherit=1 flag. Such counters are then
'inherited' by all child tasks (be they fork()-ed or clone()-ed). These
counters get inherited through exec() boundaries as well (except through
setuid boundaries).

The counter values get added back to the parent counter(s) when the child
task(s) exit - much like stime/utime statistics are gathered. So inherited
counters are ideal to gather summary statistics about an application's
behavior via shell commands, without having to modify that application.

The timec.c command utilizes counter inheritance:

  http://redhat.com/~mingo/perfcounters/timec.c

Sample output:

   $ ./timec -e 1 -e 3 -e 5 ls -lR /usr/include/ >/dev/null

   Performance counter stats for 'ls':

           163516953 instructions
                2295 cache-misses
             2855182 branch-misses
Signed-off-by: NIngo Molnar <mingo@elte.hu>

9b51f66d

perfcounters: restructure x86 counter math · ee06094f

由 Ingo Molnar 提交于 12月 13, 2008

Impact: restructure code

Change counter math from absolute values to clear delta logic.

We try to extract elapsed deltas from the raw hw counter - and put
that into the generic counter.
Signed-off-by: NIngo Molnar <mingo@elte.hu>

ee06094f

11 12月, 2008 11 次提交

perf counters: clean up state transitions · 6a930700