提交 · df399b064118bf9a5b9a3faaa67feb1cbb34e9d4 · openeuler / Kernel

28 3月, 2019 1 次提交

drm/amdgpu: XGMI pstate switch initial support · df399b06

由 shaoyunl 提交于 3月 20, 2019

Driver vote low to high pstate switch whenever there is an outstanding
XGMI mapping request. Driver vote high to low pstate when all the
outstanding XGMI mapping is terminated.
Signed-off-by: Nshaoyunl <shaoyun.liu@amd.com>
Reviewed-by: NChristian König <christian.koenig@amd.com>
Signed-off-by: NAlex Deucher <alexander.deucher@amd.com>

df399b06

22 3月, 2019 4 次提交

drm/amdgpu: use the new VM backend for PTEs · c3546695

由 Christian König 提交于 3月 18, 2019

And remove the existing code when it is unused.
Signed-off-by: NChristian König <christian.koenig@amd.com>
Reviewed-by: NChunming Zhou <david1.zhou@amd.com>
Reviewed-by: NFelix Kuehling <Felix.Kuehling@amd.com>
Signed-off-by: NAlex Deucher <alexander.deucher@amd.com>

c3546695

drm/amdgpu: new VM update backends · 6dd09027

由 Christian König 提交于 3月 18, 2019

Separate out all functions for SDMA and CPU based page table
updates into separate backends.

This way we can keep most of the complexity of those from the
core VM code.
Signed-off-by: NChristian König <christian.koenig@amd.com>
Reviewed-by: NChunming Zhou <david1.zhou@amd.com>
Reviewed-by: NFelix Kuehling <Felix.Kuehling@amd.com>
Signed-off-by: NAlex Deucher <alexander.deucher@amd.com>

6dd09027

drm/amdgpu: move and rename amdgpu_pte_update_params · d1e29462

由 Christian König 提交于 3月 18, 2019

Move the update parameter into the VM header and rename them.
Signed-off-by: NChristian König <christian.koenig@amd.com>
Reviewed-by: NChunming Zhou <david1.zhou@amd.com>
Reviewed-by: NFelix Kuehling <Felix.Kuehling@amd.com>
Signed-off-by: NAlex Deucher <alexander.deucher@amd.com>

d1e29462

drm/amdgpu: remove some unused VM defines · 2c250802

由 Christian König 提交于 3月 18, 2019

Not needed any more.
Signed-off-by: NChristian König <christian.koenig@amd.com>
Reviewed-by: NChunming Zhou <david1.zhou@amd.com>
Reviewed-by: NFelix Kuehling <Felix.Kuehling@amd.com>
Signed-off-by: NAlex Deucher <alexander.deucher@amd.com>

2c250802

20 3月, 2019 4 次提交

drm/amdgpu: wait for VM to become idle during flush · 56753e73

由 Christian König 提交于 1月 10, 2019

Make sure that not only the entities are flush, but that
we also wait for the HW to finish all processing.
Signed-off-by: NChristian König <christian.koenig@amd.com>
Reviewed-by: NChunming Zhou <david1.zhou@amd.com>
Signed-off-by: NAlex Deucher <alexander.deucher@amd.com>

56753e73

drm/amdgpu: remove chash · 04ed8459

由 Christian König 提交于 11月 07, 2018

Remove the chash implementation for now since it isn't used any more.
Signed-off-by: NChristian König <christian.koenig@amd.com>
Reviewed-by: NFelix Kuehling <Felix.Kuehling@amd.com>
Signed-off-by: NAlex Deucher <alexander.deucher@amd.com>

04ed8459

drm/amdgpu: drop the huge page flag · adc7bfe5

由 Christian König 提交于 2月 01, 2019

Not needed any more since we now free PDs/PTs on demand.
Signed-off-by: NChristian König <christian.koenig@amd.com>
Reviewed-by: NFelix Kuehling <Felix.Kuehling@amd.com>
Acked-by: NHuang Rui <ray.huang@amd.com>
Signed-off-by: NAlex Deucher <alexander.deucher@amd.com>

adc7bfe5

drm/amdgpu: allocate VM PDs/PTs on demand · 0ce15d6f

由 Christian König 提交于 1月 30, 2019

Let's start to allocate VM PDs/PTs on demand instead of pre-allocating
them during mapping.
Signed-off-by: NChristian König <christian.koenig@amd.com>
Reviewed-by: NFelix Kuehling <Felix.Kuehling@amd.com>
Acked-by: NHuang Rui <ray.huang@amd.com>
Signed-off-by: NAlex Deucher <alexander.deucher@amd.com>

0ce15d6f

26 1月, 2019 1 次提交

drm/amdgpu: set bulk_moveable to false when lru changed v2 · b61857b5

由 Chunming Zhou 提交于 1月 10, 2019

if lru is changed, we cannot do bulk moving.
v2:
root bo isn't in bulk moving, skip its change.
Signed-off-by: NChunming Zhou <david1.zhou@amd.com>
Reviewed-by: NChristian König <christian.koenig@amd.com>
Signed-off-by: NAlex Deucher <alexander.deucher@amd.com>

b61857b5

08 12月, 2018 1 次提交

drm/amdgpu: remove VM fault_credit handling · a655dad4

由 Christian König 提交于 9月 26, 2018

printk_ratelimit() is much better suited to limit the number of reported
VM faults.
Signed-off-by: NChristian König <christian.koenig@amd.com>
Reviewed-by: NFelix Kuehling <Felix.Kuehling@amd.com>
Acked-by: NAlex Deucher <alexander.deucher@amd.com>
Signed-off-by: NAlex Deucher <alexander.deucher@amd.com>

a655dad4

14 9月, 2018 1 次提交

drm/amdgpu: use a single linked list for amdgpu_vm_bo_base · 646b9025

由 Christian König 提交于 9月 10, 2018

Instead of the double linked list. Gets the size of amdgpu_vm_pt down to
64 bytes again.

We could even reduce it down to 32 bytes, but that would require some
rather extreme hacks.
Signed-off-by: NChristian König <christian.koenig@amd.com>
Reviewed-by: NChunming Zhou <david1.zhou@amd.com>
Acked-by: NFelix Kuehling <Felix.Kuehling@amd.com>
Signed-off-by: NAlex Deucher <alexander.deucher@amd.com>

646b9025

13 9月, 2018 1 次提交

drm/amdgpu: Move fault hash table to amdgpu vm · 240cd9a6

由 Oak Zeng 提交于 9月 05, 2018

In stead of share one fault hash table per device, make it
per vm. This can avoid inter-process lock issue when fault
hash table is full.

Change-Id: I5d1281b7c41eddc8e26113e010516557588d3708
Signed-off-by: NOak Zeng <Oak.Zeng@amd.com>
Suggested-by: NChristian Konig <Christian.Koenig@amd.com>
Suggested-by: NFelix Kuehling <Felix.Kuehling@amd.com>
Reviewed-by: NChristian Konig <christian.koenig@amd.com>
Signed-off-by: NAlex Deucher <alexander.deucher@amd.com>

240cd9a6

11 9月, 2018 2 次提交

drm/amdgpu: correctly sign extend 48bit addresses v3 · ad9a5b78

由 Christian König 提交于 8月 27, 2018

Correct sign extend the GMC addresses to 48bit.

v2: sign extending turned out easier than thought.
v3: clean up the defines and move them into amdgpu_gmc.h as well
Signed-off-by: NChristian König <christian.koenig@amd.com>
Reviewed-by: NJunwei Zhang <Jerry.Zhang@amd.com>
Signed-off-by: NAlex Deucher <alexander.deucher@amd.com>

ad9a5b78

drm/amdgpu: separate per VM BOs from normal in the moved state · c12a2ee5

由 Christian König 提交于 9月 01, 2018

Allows us to avoid taking the spinlock in more places.
Signed-off-by: NChristian König <christian.koenig@amd.com>
Reviewed-by: NJunwei Zhang <Jerry.Zhang@amd.com>
Signed-off-by: NAlex Deucher <alexander.deucher@amd.com>

c12a2ee5

30 8月, 2018 2 次提交

drm/amdkfd: Release an acquired process vm · bf47afba

由 Oak Zeng 提交于 8月 27, 2018

For compute vm acquired from amdgpu, vm.pasid is managed
by kfd. Decouple pasid from such vm on process destroy
to avoid duplicate pasid release.
Signed-off-by: NOak Zeng <Oak.Zeng@amd.com>
Reviewed-by: NFelix Kuehling <Felix.Kuehling@amd.com>
Signed-off-by: NAlex Deucher <alexander.deucher@amd.com>

bf47afba

drm/amdgpu: Set pasid for compute vm (v2) · 1685b01a

由 Oak Zeng 提交于 8月 29, 2018

To make a amdgpu vm to a compute vm, the old pasid will be freed and
replaced with a pasid managed by kfd. Kfd can't reuse original pasid
allocated by amdgpu because kfd uses different pasid policy with amdgpu.
For example, all graphic devices share one same pasid in a process.

v2: rebase (Alex)
Signed-off-by: NOak Zeng <Oak.Zeng@amd.com>
Reviewed-by: NFelix Kuehling <Felix.Kuehling@amd.com>
Signed-off-by: NAlex Deucher <alexander.deucher@amd.com>

1685b01a

28 8月, 2018 6 次提交

drm/amdgpu: remove extra root PD alignment · 248f2b8e

由 Christian König 提交于 8月 22, 2018

Just another leftover from radeon.
Signed-off-by: NChristian König <christian.koenig@amd.com>
Reviewed-by: NAlex Deucher <alexander.deucher@amd.com>
Reviewed-by: NJunwei Zhang <Jerry.Zhang@amd.com>
Acked-by: NHuang Rui <ray.huang@amd.com>
Signed-off-by: NAlex Deucher <alexander.deucher@amd.com>

248f2b8e

drm/amdgpu: Adjust the VM size based on system memory size v2 · 43370c4c

由 Felix Kuehling 提交于 8月 21, 2018

Set the VM size based on system memory size between the ASIC-specific
limits given by min_vm_size and max_bits. GFXv9 GPUs will keep their
default VM size of 256TB (48 bit). Only older GPUs will adjust VM size
depending on system memory size.

This makes more VM space available for ROCm applications on GFXv8 GPUs
that want to map all available VRAM and system memory in their SVM
address space.

v2:
* Clarify comment
* Round up memory size before >> 30
* Round up automatic vm_size to power of two
Signed-off-by: NFelix Kuehling <Felix.Kuehling@amd.com>
Acked-by: NJunwei Zhang <Jerry.Zhang@amd.com>
Reviewed-by: NHuang Rui <ray.huang@amd.com>
Reviewed-by: NChristian König <christian.koenig@amd.com>
Signed-off-by: NAlex Deucher <alexander.deucher@amd.com>

43370c4c

drm/amdgpu: Adjust the VM size based on system memory size v2 · fca5d959

由 Felix Kuehling 提交于 8月 21, 2018

Set the VM size based on system memory size between the ASIC-specific
limits given by min_vm_size and max_bits. GFXv9 GPUs will keep their
default VM size of 256TB (48 bit). Only older GPUs will adjust VM size
depending on system memory size.

This makes more VM space available for ROCm applications on GFXv8 GPUs
that want to map all available VRAM and system memory in their SVM
address space.

v2:
* Clarify comment
* Round up memory size before >> 30
* Round up automatic vm_size to power of two
Signed-off-by: NFelix Kuehling <Felix.Kuehling@amd.com>
Acked-by: NJunwei Zhang <Jerry.Zhang@amd.com>
Reviewed-by: NHuang Rui <ray.huang@amd.com>
Reviewed-by: NChristian König <christian.koenig@amd.com>
Signed-off-by: NAlex Deucher <alexander.deucher@amd.com>

fca5d959

drm/amdgpu: use bulk moves for efficient VM LRU handling (v6) · f921661b

由 Huang Rui 提交于 8月 06, 2018

I continue to work for bulk moving that based on the proposal by Christian.

Background:
amdgpu driver will move all PD/PT and PerVM BOs into idle list. Then move all of
them on the end of LRU list one by one. Thus, that cause so many BOs moved to
the end of the LRU, and impact performance seriously.

Then Christian provided a workaround to not move PD/PT BOs on LRU with below
patch:
Commit 0bbf32026cf5ba41e9922b30e26e1bed1ecd38ae ("drm/amdgpu: band aid
validating VM PTs")

However, the final solution should bulk move all PD/PT and PerVM BOs on the LRU
instead of one by one.

Whenever amdgpu_vm_validate_pt_bos() is called and we have BOs which need to be
validated we move all BOs together to the end of the LRU without dropping the
lock for the LRU.

While doing so we note the beginning and end of this block in the LRU list.

Now when amdgpu_vm_validate_pt_bos() is called and we don't have anything to do,
we don't move every BO one by one, but instead cut the LRU list into pieces so
that we bulk move everything to the end in just one operation.

Test data:
+--------------+-----------------+-----------+---------------------------------------+
|              |The Talos        |Clpeak(OCL)|BusSpeedReadback(OCL)                  |
|              |Principle(Vulkan)|           |                                       |
+------------------------------------------------------------------------------------+
|              |                 |           |0.319 ms(1k) 0.314 ms(2K) 0.308 ms(4K) |
| Original     |  147.7 FPS      |  76.86 us |0.307 ms(8K) 0.310 ms(16K)             |
+------------------------------------------------------------------------------------+
| Orignial + WA|                 |           |0.254 ms(1K) 0.241 ms(2K)              |
|(don't move   |  162.1 FPS      |  42.15 us |0.230 ms(4K) 0.223 ms(8K) 0.204 ms(16K)|
|PT BOs on LRU)|                 |           |                                       |
+------------------------------------------------------------------------------------+
| Bulk move    |  163.1 FPS      |  40.52 us |0.244 ms(1K) 0.252 ms(2K) 0.213 ms(4K) |
|              |                 |           |0.214 ms(8K) 0.225 ms(16K)             |
+--------------+-----------------+-----------+---------------------------------------+

After test them with above three benchmarks include vulkan and opencl. We can
see the visible improvement than original, and even better than original with
workaround.

v2: move all BOs include idle, relocated, and moved list to the end of LRU and
put them together.
v3: remove unused parameter and use list_for_each_entry instead of the one with
save entry.
v4: move the amdgpu_vm_move_to_lru_tail after command submission, at that time,
all bo will be back on idle list.
v5: remove amdgpu_vm_move_to_lru_tail_by_list(), use bulk_moveable instread of
validated, and move ttm_bo_bulk_move_lru_tail() also into
amdgpu_vm_move_to_lru_tail().
v6: clean up and fix return value.
Signed-off-by: NChristian König <christian.koenig@amd.com>
Signed-off-by: NHuang Rui <ray.huang@amd.com>
Tested-by: NMike Lothian <mike@fireburn.co.uk>
Tested-by: NDieter Nützel <Dieter@nuetzel-hh.de>
Acked-by: NChunming Zhou <david1.zhou@amd.com>
Reviewed-by: NJunwei Zhang <Jerry.Zhang@amd.com>
Signed-off-by: NAlex Deucher <alexander.deucher@amd.com>

f921661b

drm/amdgpu: use new scheduler load balancing for VMs · 3798e9a6

由 Christian König 提交于 7月 12, 2018

Instead of the fixed round robin use let the scheduler balance the load
of page table updates.
Signed-off-by: NChristian König <christian.koenig@amd.com>
Reviewed-by: NChunming Zhou <david1.zhou@amd.com>
Signed-off-by: NAlex Deucher <alexander.deucher@amd.com>

3798e9a6

drm/amdgpu: move vm definitions into amdgpu_vm header · 4473e1db

由 Huang Rui 提交于 8月 03, 2018

Demangle amdgpu.h.
Signed-off-by: NHuang Rui <ray.huang@amd.com>
Acked-by: NChristian König <christian.koenig@amd.com>
Signed-off-by: NAlex Deucher <alexander.deucher@amd.com>

4473e1db

01 8月, 2018 1 次提交

drm/amdgpu: add new amdgpu_vm_bo_trace_cs() function v2 · 8ab19ea6

由 Christian König 提交于 7月 27, 2018

This allows us to trace all VM ranges which should be valid inside a CS.

v2: dump mappings without BO as well
Signed-off-by: NChristian König <christian.koenig@amd.com>
Reviewed-by: NChunming  Zhou <david1.zhou@amd.com>
Reviewed-and-tested-by: Andrey Grodzovsky <andrey.grodzovsky@amd.com> (v1)
Reviewed-by: Huang Rui <ray.huang@amd.com> (v1)
Signed-off-by: NAlex Deucher <alexander.deucher@amd.com>

8ab19ea6

11 7月, 2018 1 次提交

drm/amdgpu: Add support for logging process info in amdgpu_vm. · 2aa37bf5

由 Andrey Grodzovsky 提交于 6月 28, 2018

Add process and thread names and pids and a function to extract
this info from relevant amdgpu_vm.

v2: Add documentation and fix identation.

v3: Add getter and setter functions for amdgpu_task_info.
Signed-off-by: NAndrey Grodzovsky <andrey.grodzovsky@amd.com>
Acked-by: NJim Qu <Jim.Qu@amd.com>
Signed-off-by: NAlex Deucher <alexander.deucher@amd.com>

2aa37bf5

24 5月, 2018 2 次提交

drm/amdgpu: move VM BOs on LRU again · 806f043f

由 Christian König 提交于 4月 19, 2018

Move all BOs belonging to a VM on the LRU with every submission.
Signed-off-by: NChristian König <christian.koenig@amd.com>
Reviewed-by: NChunming Zhou <david1.zhou@amd.com>
Signed-off-by: NAlex Deucher <alexander.deucher@amd.com>

806f043f

drm/amdgpu: rework VM state machine lock handling v2 · af4c0f65

由 Christian König 提交于 4月 19, 2018

Only the moved state needs a separate spin lock protection. All other
states are protected by reserving the VM anyway.

v2: fix some more incorrect cases
Signed-off-by: NChristian König <christian.koenig@amd.com>
Reviewed-by: NChunming Zhou <david1.zhou@amd.com>
Signed-off-by: NAlex Deucher <alexander.deucher@amd.com>

af4c0f65

19 5月, 2018 1 次提交

drm/amdgpu: remove unused member · b9245b94

由 Christian König 提交于 4月 19, 2018

This lock isn't used any more.
Signed-off-by: NChristian König <christian.koenig@amd.com>
Reviewed-by: NChunming Zhou <david1.zhou@amd.com>
Signed-off-by: NAlex Deucher <alexander.deucher@amd.com>

b9245b94

16 5月, 2018 1 次提交

drm/amdgpu: Add support to change mtype for 2nd part of gart BOs on GFX9 · 959a2091

由 Yong Zhao 提交于 5月 14, 2018

This change prepares for a workaround in amdkfd for a GFX9 HW bug. It
requires the control stack memory of compute queues, which is allocated
from the second page of MQD gart BOs, to have mtype NC, rather than
the default UC.
Signed-off-by: NYong Zhao <yong.zhao@amd.com>
Reviewed-by: NChristian König <christian.koenig@amd.com>
Signed-off-by: NAlex Deucher <alexander.deucher@amd.com>

959a2091

16 3月, 2018 2 次提交

drm/amdgpu: Add helper to turn an existing VM into a compute VM · b236fa1d

由 Felix Kuehling 提交于 3月 15, 2018

v2: Removed updating and checking of vm->vm_context
v3: Enable amdgpu_vm_clear_bo in amdgpu_vm_make_compute
Signed-off-by: NFelix Kuehling <Felix.Kuehling@amd.com>
Reviewed-by: NChristian König <christian.koenig@amd.com>
Signed-off-by: NOded Gabbay <oded.gabbay@gmail.com>

b236fa1d

drm/amdgpu: Move KFD-specific fields into struct amdgpu_vm · 5b21d3e5

由 Felix Kuehling 提交于 3月 15, 2018

Remove struct amdkfd_vm and move the fields into struct amdgpu_vm.
This will allow turning a VM created by a DRM render node into a
KFD VM.

v2: Removed vm_context field
Signed-off-by: NFelix Kuehling <Felix.Kuehling@amd.com>
Reviewed-by: NChristian König <christian.koenig@amd.com>
Signed-off-by: NOded Gabbay <oded.gabbay@gmail.com>

5b21d3e5

20 2月, 2018 1 次提交

drm/amdgpu: reduce reserved VA size · 18d09e63

由 Christian König 提交于 1月 22, 2018

1MB should be more than enough, currently we use about 8K.
Signed-off-by: NChristian König <christian.koenig@amd.com>
Reviewed-by: NAlex Deucher <alexander.deucher@amd.com>
Acked-by: NMonk Liu <monk.liu@amd.com>
Signed-off-by: NAlex Deucher <alexander.deucher@amd.com>

18d09e63

28 12月, 2017 2 次提交

drm/amdgpu: drop client_id from VM · 0e36b9b2

由 Christian König 提交于 12月 18, 2017

Use the fence context from the scheduler entity.
Signed-off-by: NChristian König <christian.koenig@amd.com>
Reviewed-by: NChunming Zhou <david1.zhou@amd.com>
Signed-off-by: NAlex Deucher <alexander.deucher@amd.com>

0e36b9b2

drm/amdgpu: separate VMID and PASID handling · 620f774f

由 Christian König 提交于 12月 18, 2017

Move both into the new files amdgpu_ids.[ch]. No functional change.
Signed-off-by: NChristian König <christian.koenig@amd.com>
Reviewed-by: NChunming Zhou <david1.zhou@amd.com>
Signed-off-by: NAlex Deucher <alexander.deucher@amd.com>

620f774f

19 12月, 2017 1 次提交

drm/amdgpu: implement 2+1 PD support for Raven v3 · 6a42fd6f

由 Christian König 提交于 12月 05, 2017

Instead of falling back to 2 level and very limited address space use
2+1 PD support and 128TB + 512GB of virtual address space.

v2: cleanup defines, rebase on top of level enum
v3: fix inverted check in hardware setup
Signed-off-by: NChristian König <christian.koenig@amd.com>
Reviewed-and-Tested-by: NChunming Zhou <david1.zhou@amd.com>
Signed-off-by: NAlex Deucher <alexander.deucher@amd.com>

6a42fd6f

15 12月, 2017 1 次提交

drm/amdgpu: add enumerate for PDB/PTB v3 · 196f7489

由 Chunming Zhou 提交于 12月 13, 2017

v2:
  remove SUBPTB member
v3:
  remove last_level, use AMDGPU_VM_PTB directly instead.
Signed-off-by: NChunming Zhou <david1.zhou@amd.com>
Reviewed-by: NChristian König <christian.koenig@amd.com>
Signed-off-by: NAlex Deucher <alexander.deucher@amd.com>

196f7489

13 12月, 2017 2 次提交

drm/amdgpu: remove keeping the addr of the VM PDs · 78eb2f0c

由 Christian König 提交于 11月 30, 2017

No more double house keeping.
Signed-off-by: NChristian König <christian.koenig@amd.com>
Reviewed-by: NChunming Zhou <david1.zhou@amd.com>
Signed-off-by: NAlex Deucher <alexander.deucher@amd.com>

78eb2f0c

drm/amdgpu: remove last_entry_used from the VM code · 8f19cd78

由 Christian König 提交于 11月 30, 2017

Not needed any more.
Signed-off-by: NChristian König <christian.koenig@amd.com>
Reviewed-by: NChunming Zhou <david1.zhou@amd.com>
Signed-off-by: NAlex Deucher <alexander.deucher@amd.com>

8f19cd78

07 2月, 2018 1 次提交

drm/amdgpu: Fix header file dependencies · 61b100e9

由 Felix Kuehling 提交于 2月 06, 2018

Signed-off-by: NFelix Kuehling <Felix.Kuehling@amd.com>
Reviewed-by: NChristian König <christian.koenig@amd.com>
Signed-off-by: NOded Gabbay <oded.gabbay@gmail.com>

61b100e9

08 12月, 2017 1 次提交

drm: move amd_gpu_scheduler into common location · 1b1f42d8

由 Lucas Stach 提交于 12月 06, 2017

This moves and renames the AMDGPU scheduler to a common location in DRM
in order to facilitate re-use by other drivers. This is mostly a straight
forward rename with no code changes.

One notable exception is the function to_drm_sched_fence(), which is no
longer a inline header function to avoid the need to export the
drm_sched_fence_ops_scheduled and drm_sched_fence_ops_finished structures.
Reviewed-by: NChunming Zhou <david1.zhou@amd.com>
Tested-by: NDieter Nützel <Dieter@nuetzel-hh.de>
Acked-by: NAlex Deucher <alexander.deucher@amd.com>
Signed-off-by: NLucas Stach <l.stach@pengutronix.de>
Signed-off-by: NAlex Deucher <alexander.deucher@amd.com>

1b1f42d8

openeuler / Kernel 1 年多 前同步成功

openeuler / Kernel
1 年多前同步成功