提交 · 8b1251e1740d3b3b7d390a9d52d957debd943146 · OpenXiangShan / XiangShan

18 11月, 2022 29 次提交

sbuffer: opt mask clean fanout (#1720) · 8b1251e1

由 William Wang 提交于 8月 10, 2022

We used to clean mask in sbuffer in 1 cycle when do sbuffer enq,
which introduced 64*16 fanout.

To reduce fanout, now mask in sbuffer is cleaned when dcache hit resp
comes. Clean mask for a line in sbuffer takes 2 cycles.

Meanwhile, dcache reqIdWidth is also reduced from 64 to
log2Up(nEntries) max log2Up(StoreBufferSize).

This commit will not cause perf change.

8b1251e1

L

dcache: duplicate 3 more regs in cacheOpDecoder · 476e71e5
由 lixin 提交于 8月 10, 2022

476e71e5
Z

MainPipe: fix fanout of regs in stage 3 (#1718) · ca18e2c6
由 zhanglinjuan 提交于 8月 09, 2022

ca18e2c6
W
lq: update paddr in lq in load_s1 and load_s2 (#1707) · 0a47e4a1
由 William Wang 提交于 8月 09, 2022
```
Now we use 2 cycles to update paddr in lq. In this way,
paddr in lq is still valid in load_s3
```
0a47e4a1
L

dcache: duplicate cache_req_valid · 72e3aa13
由 lixin 提交于 8月 09, 2022

72e3aa13
L

dcache: duplicate regs in cacheOpDecoder · e47fc57c
由 lixin 提交于 8月 09, 2022

e47fc57c

lq: add 1 extra stage for lq data write (#1705) · 39f2ec76

由 William Wang 提交于 8月 09, 2022

Now lq data is divided into 8 banks by default. Write to lq
data takes 2 cycles to finish

Lq data will not be read in at least 2 cycles after write, so it is ok
to add this delay. For example:
T0: update lq meta, lq data write req start
T1: lq data write finish, new wbidx selected
T2: read lq data according to new wbidx selected

39f2ec76

W

misc: fix nanhu lsu cherry-pick conflict · c047ef9c
由 William Wang 提交于 11月 18, 2022

c047ef9c
W

std: add an extra pipe stage for std (#1704) · 0a992150
由 William Wang 提交于 8月 06, 2022

0a992150
Z

WritebackQueue: fix bug when ProbeAck is merged with a ReleaseData (#1709) · 5c01cc3c
由 zhanglinjuan 提交于 8月 06, 2022

5c01cc3c
H

dcache: duplicate registers for better fanout (#1700) · c3a5fe5f
由 happy-lx 提交于 8月 04, 2022

c3a5fe5f

dcache: fix fanout · b11ec622

由 lixin 提交于 7月 25, 2022

* pipelineReg in miss queue
* translated_cache_req_opCode and io_cache_req_valid_reg in cacheOpDecoder
* r_way_en_reg in bankedDataArray

b11ec622

dcache: delay wbq data update for 1 cycle (#1701) · 7a919e05

由 William Wang 提交于 8月 03, 2022

This commit and an extra cycle for miss queue store data and mask write.
For now, there are 18 missqueue entries. Each entry has a 512 bit
data reg and a 64 bit mask reg. If we update writeback queue data in 1
cycle, the fanout will be at least 18x(512+64) = 10368.

Now writeback queue req meta update is unchanged, however, data and mask
update will happen 1 cycle after req fire or release update fire (T0).
In T0, data and meta will be written to a buffer in missqueue.
In T1, s_data_merge or s_data_override in each missqueue entry will
be used as data and mask wen.

7a919e05

W

sq: always update data/addrModule when st s1_valid (#1703) · 29b5bc3c
由 William Wang 提交于 8月 03, 2022

29b5bc3c
W

dcache: use MissReqWoStoreData in missq entry · e771db6c
由 William Wang 提交于 8月 01, 2022

e771db6c

dcache: delay missq st data/mask write for 1 cycle · c731e79f

由 William Wang 提交于 8月 01, 2022

This commit and an extra cycle for miss queue store data and mask write.
For now, there are 16 missqueue entries. Each entry has a 512 bit store
data reg and a 64 bit store mask. If we update miss queue data in 1
cycle, the fanout will be at least 16x(512+64) = 9216.

Now missqueue req meta update is unchanged, however, store data and mask
update will happen 1 cycle after primary fire or secondary fire (T0).
In T0, store data and meta will be written to a buffer in missqueue.
In T1, s_write_storedata in each missqueue entry will be used as store
data and mask wen.

Miss queue entry data organization is also optimized. 512 bit
req.store_data is removed from miss queue entry. It should save
8192 bits in total.

c731e79f

W

dcache: fix rowBits parameter usage · af22dd7c
由 William Wang 提交于 8月 01, 2022

af22dd7c
W
ldu: update lq correctly when replay_from_fetch (#1694) · 7ad02651
由 William Wang 提交于 7月 30, 2022
```
uop.ctrl.replayInst in lq should be replayed when load_s2 update lq
i.e. load_s2.io.out.valid
```
7ad02651
W

lq: fix X introduced by violation check (#1695) · e5cb7504
由 William Wang 提交于 7月 30, 2022

e5cb7504
W
sbuffer: gen blockDcacheWrite 1 cycle earlier (#1693) · 779faf12
由 William Wang 提交于 7月 28, 2022
```
It will save time for store_req generation in dcache Mainpipe, which is
at the beginning of a critical path
```
779faf12
W

lq: opt lq data wen (load_s2_valid) fanout (#1687) · c1af2986
由 William Wang 提交于 7月 27, 2022

c1af2986
J

Misc: l1 buffer adjustment (#1689) · 4a2390a4
由 Jiawei Lin 提交于 7月 27, 2022

4a2390a4
W
ldu: report ldld vio and fwd error in s3 (#1685) · 67cddb05
由 William Wang 提交于 11月 18, 2022
```
It should fix the timing problem caused by ldld violation check and
forward error check
```
67cddb05

lq: update data field iff load_s2 valid (#1680) · 353424a7

由 William Wang 提交于 7月 27, 2022

Now we update data field (fwd data, uop) in load queue when load_s2
is valid. It will help to on lq wen fanout problem.

State flags will be treated differently. They are still updated
accurately according to loadIn.valid

353424a7

Z
dcache: fix fan-out in WritebackEntry (#1675) · f94d088c
由 Ziyue-Zhang 提交于 7月 23, 2022
```
Co-authored-by: NZiyue Zhang <zhangziyue21b@ict.ac.cn>
```
f94d088c
W

sbuffer: set EnsbufferWidth upper bound to 2 · db7f55d9
由 William Wang 提交于 11月 18, 2022

db7f55d9

sbuffer: add an extra cycle for sbuffer write · 3d3419b9

由 William Wang 提交于 7月 26, 2022

In previous design, sbuffer valid entry select and
sbuffer data write are in the same cycle, which
caused huge fanout. An extra write stage is added to
solve this problem.

Now sbuffer enq logic is divided into 3 stages:

sbuffer_in_s0:
* read data and meta from store queue
* store them in 2 entry fifo queue

sbuffer_in_s1:
* read data and meta from fifo queue
* update sbuffer meta (vtag, ptag, flag)
* prevert that line from being sent to dcache (add a block condition)
* prepare cacheline level write enable signal, RegNext() data and mask

sbuffer_in_s2:
* use cacheline level buffer to update sbuffer data and mask
* remove dcache write block (if there is)

3d3419b9

MainPipe: fix fan-out (#1674) · b909b713

由 zhanglinjuan 提交于 7月 26, 2022

* MainPipe: reduce fanout by duplicating registers

* MainPipe: fix wrong assert
Co-authored-by: NWilliam Wang <zeweiwang@outlook.com>

b909b713

sbuffer: rename sbuffer deq related signals · 80382c05

由 William Wang 提交于 7月 24, 2022

Now sbuffer deq logic is divided into 2 stages:

sbuffer_out_s0:
* read data and meta from sbuffer
* RegNext() them
* set line state to inflight

sbuffer_out_s1:
* send write req to dcache

sbuffer_out_extra:
* receive write result from dcache
* update line state

80382c05

15 11月, 2022 1 次提交
- J
  Merge pull request #1824 from OpenXiangShan/bump-chisel-circt · 28ab2f5a
  由 Jiawei Lin 提交于 11月 15, 2022
```
misc: bump chisel-circt
```
  28ab2f5a
14 11月, 2022 1 次提交
- S
  Merge pull request #1690 from chenguokai/frontend_db · f580a020
  由 Steve Gou 提交于 11月 14, 2022
```
frontend: Add ChiselDB records
```
  f580a020
11 11月, 2022 1 次提交
- S
  Merge pull request #1825 from OpenXiangShan/frontend-bump-nanhu · 692910fa
  由 Steve Gou 提交于 11月 11, 2022
```
frontend bump nanhu
```
  692910fa
10 11月, 2022 3 次提交
- Y
  
  ctrl: fix jalr target read address · f70fe10f
  由 Yinan Xu 提交于 7月 21, 2022
  
  f70fe10f
- J
  
  IPrefetch: fix merge error for req.ready · 020ef3eb
  由 Jenius 提交于 11月 10, 2022
  
  020ef3eb
- J
  
  ReplacePipe: fix req_id mismatch bug · 98929a13
  由 Jenius 提交于 11月 10, 2022
  
  98929a13
09 11月, 2022 5 次提交
- L
  
  misc: bump chisel-circt · 714ba5a1
  由 LinJiawei 提交于 11月 09, 2022
  
  714ba5a1
- J
  
  ICache: fix ReplacePipe comb loop · 6ecd5de6
  由 Jenius 提交于 11月 09, 2022
  
  6ecd5de6
- J
  
  IFU: fix early flush for mmio instructions · 4a74a727
  由 Jenius 提交于 11月 02, 2022
  
  4a74a727
- J
  
  <verifi>:ICache add condition for multiple-hit · ff1018c6
  由 Jenius 提交于 10月 10, 2022
  
  ff1018c6
- J
  IFU: mmio wait until last instruction retiring · 1d1e6d4d
  由 Jenius 提交于 10月 08, 2022
```
* add 1 stage for mmio_state before sending request to MMIO bus
* check whether the last fetch packet commit all its intructions (the
result of execution path has been decided)
* avoid speculative execution to MMIO bus
```
  1d1e6d4d

OpenXiangShan / XiangShan 11 个月 前同步成功

OpenXiangShan / XiangShan
11 个月前同步成功