提交 · v0.2.1 · OpenDILab开源决策智能平台 / DI-engine

22 11月, 2021 1 次提交
- N
  
  v0.2.1 · cf8ad134
  由 niuyazhe 提交于 11月 22, 2021
  
  cf8ad134
20 11月, 2021 1 次提交
- N
  
  style(nyz): add PDQN/MAPPO link, DQN doc zh link and correct format · 96103e9b
  由 niuyazhe 提交于 11月 20, 2021
  
  96103e9b
19 11月, 2021 4 次提交

polish(davide) add example of GAIL entry + config for Mujoco and Cartpole (#114) · d1bc1387

由 Davide Liu 提交于 11月 19, 2021

* added gail entry

* added lunarlander and cartpole config

* added gail mujoco config

* added mujoco exp

* update22-10

* added third exp

* added metric to evaluate policies

* added GAIL entry and config for Cartpole and Walker2d

* checked style and unittest

* restored lunarlander env

* style problems

* bug correction

* Delete expert_data_train.pkl

* changed loss of GAIL

* Update walker2d_ddpg_gail_config.py

* changed gail reward from -D(s, a) to -log(D(s, a))

* added small constant to reward function

* added comment to clarify config

* Update walker2d_ddpg_gail_config.py

* added lunarlander entry + config

* Added Atari discriminator + Pong entry config

* Update gail_irl_model.py

* Update gail_irl_model.py

* added gail serial pipeline and onehot actions for gail atari

* related to previous commit

* removed main files

* removed old comment

d1bc1387

feature(lk): add PDQN algorithm for hybrid action spaces (#118) · 39a7cfe3

由 Ke Li 提交于 11月 19, 2021

* add_pdqn_model

* modify_model_structure

* initial_version_PDQN

* bug_free_PDQN_no_test_convergence

* update_pdqn_config

* add_noise_to_continuous_args

* polish(nyz): polish code style and add noise in pdqn

* seperate_dis_and_cont_model

* fix_bug_for_separation

* fix(pu): current q value use the data action, fix cont loss detach bug, 1 encoder, dist and cont learning rate

* polish(pu): actor delay update

* fix(pu): fix disc cont update frequency

* polish(pu): polish pdqn config

* polish(lk): add comments and typelint for pdqn and dqn

* feature(lk): add test file for pdqn model and policy

* polish(lk): code style

* polish(lk): rm the modify of unrelated files

* polish(lk): rm useless commentes code in pdqn
Co-authored-by: Nniuyazhe <niuyazhe@sensetime.com>
Co-authored-by: Npuyuan1996 <2402552459@qq.com>

39a7cfe3

N

style(nyz): modify supported PyTorch version and correct format · d8115c50
由 niuyazhe 提交于 11月 19, 2021

d8115c50
Z

test(zjow): verified the compatibility of PyTorch==1.10.0 · 3922ffb5
由 zjowowen 提交于 11月 19, 2021

3922ffb5

18 11月, 2021 3 次提交
- N
  
  polish(nyz): add naive buffer periodic thruput seconds argument · abc4b3a2
  由 niuyazhe 提交于 11月 18, 2021
  
  abc4b3a2
- J
  polish(yzj): add DataParallel and DataDistributedParallel (#123) · d1188d71
  由 jayyoung0802 提交于 11月 18, 2021
```
* add spaceinvaders multi gpu

* add dp and ddp

* Update __init__.py

* recover init
```
  d1188d71
- N
  feature(nyz): add registry force_overwrite argument and polish cartpole · cbee45b4
  由 niuyazhe 提交于 11月 18, 2021
```
qrdqn config
```
  cbee45b4
17 11月, 2021 1 次提交
- X
  Merge pull request #122 from opendilab/dev-torch1.1.0 · f0014586
  由 Xu Jingxin 提交于 11月 17, 2021
```
feature(nyz): extend torch1.1.0 support
```
  f0014586
16 11月, 2021 4 次提交
- N
  
  refactor(nyz): add new compatibility file in ding top level · 495e2f1a
  由 niuyazhe 提交于 11月 16, 2021
  
  495e2f1a
- N
  
  polish(nyz): add torch1.1.0 compatibility for nn.Flatten · 7acdb671
  由 niuyazhe 提交于 11月 16, 2021
  
  7acdb671
- N
  
  polish(nyz): add torch1.1.0 compatibility for torch.utils.data · 171dddc4
  由 niuyazhe 提交于 11月 16, 2021
  
  171dddc4
- N
  
  style(nyz): add torch1.1.0 support · 8df82e01
  由 niuyazhe 提交于 11月 16, 2021
  
  8df82e01
15 11月, 2021 2 次提交
- J
  feature(jrn): add the bipedalwalker config of sac and ppo (#121) · 38480a5b
  由 Jia Ruonan 提交于 11月 15, 2021
```
* commit bipedalwalkere_ppo_config

* commit bipedalwalker_sac_config
```
  38480a5b
- N
  
  style(nyz): add mbrl badge and env doc link · 12bc041d
  由 niuyazhe 提交于 11月 15, 2021
  
  12bc041d
07 11月, 2021 1 次提交
- N
  feature(nyz): enable arbitrary policy num in serial sample collector and... · 3a91c429
  由 niuyazhe 提交于 11月 07, 2021
```
feature(nyz): enable arbitrary policy num in serial sample collector and evaluator, add git in docker(smac docker)
```
  3a91c429
03 11月, 2021 3 次提交
- N
  
  fix(nyz): fix wqmix target_model state_dict bug and polish mujoco model env · f70d3ddb
  由 niuyazhe 提交于 11月 03, 2021
  
  f70d3ddb
- N
  
  fix(nyz): fix learn state_dict target model bug · db642fd3
  由 niuyazhe 提交于 11月 03, 2021
  
  db642fd3
- D
  fix(davide): small fix on bsuite environment (#117) · ad780d49
  由 Davide Liu 提交于 11月 03, 2021
```
* small fix

* added bsuite env version

* modified test
```
  ad780d49
01 11月, 2021 3 次提交

N

fix(nyz): fix target model wrapper hard reset bug · fd11c88f
由 niuyazhe 提交于 11月 01, 2021

fd11c88f
N

fix(nyz): fix r2d2 and dqtd error unittest bug · 28930a86
由 niuyazhe 提交于 11月 01, 2021

28930a86

蒲

feature(pu): add NGU algorithm (#40) · 286ea243

由蒲源提交于 11月 01, 2021

* test rnd

* fix mz config

* fix config

* feature(pu): fix r2d2, add beta to actor

* feature(pu): add ngu-dev

* fix(pu): fix r2d2

* fix(puyuan): fix r2d2

* feature(puyuan): add minigrid r2d2 config

* polish minigrid config

* dev-ngu

* feature(pu): add action and reward as inputs of q network

* feature(pu): add episodic reward model

* feature(pu): add episodic reward model, modify r2d2 and collector for ngu

* fix(pu): recover files that were changed by mistake

* fix(pu): fix tblogger cnt bug

* add_dqfd

* Is_expert to is_expert

* fix(pu): fix r2d2 bug

* fix(pu): fix beta index to gamma bug

* fix(pu): fix numerical stability problem

* style(pu): flake8 format

* fix(pu): fix rnd reward model train times

* polish(pu): polish r2d2 reset problem

* fix(pu): fix episodic reward normalize bug

* polish(pu): polish config params and episodic_reward init value

* modify according to the last commnets

* value_gamma;done;marginloss;sqil适配

* feature(pu): add r2d3 algorithm and config of lunarlander and pong

* fix(pu): fix demo path bug

* fix(pu): fix cuda bug at function get_gae in adder.py

* feature(pu): add pong r2d2 config

* polish(pu): r2d2 uses the mixture priority, episodic_reward transforms to mean 0 std1

* polish(pu): polish r2d2 config

* test(pu): test cuda compatiality of dqfd_nstep_td_error in r2d3

* polish(pu): polish config

* polish(pu): polish config and annotation

* fix(pu): fix r2d2 target net update bug and done bug

* polish(pu): polish pong r2d2 config and add montezuma r2d2 config

* polish(pu): add some logs for debugging in r2d2

* polish(pu): recover config deleted by mistake

* fix(pu): fix r2d3 config of lunarlander and pong

* fix(pu): fix the r2d2 bug in r2d3

* fix(pu): fix r2d3 cpu device bug in fun dqfd_nstep_td_error of td.py

* fix(pu): fix n_sample bug in serial_entry_r2d3

* polish(pu): polish minigrid r2d2 config

* fix(pu): add info dict of fourrooms doorkey in minigrid_env

* polish(pu): polish r2d2 config

* fix(pu): fix expert policy collect traj bug, now we use the argmax_sample wrapper

* fix(pu): fix r2d2 done and target update bug, polish config

* fix(pu): fix null_padding transition obs to zeros

* fix(pu): episodic_reward transform to [0,1]

* fix(pu): fix the value_gamma bug

* fix(pu): fix device bug in ngu_reward_model.py

* fix(pu): fix null_padding problem in rnd and episodic reward model

* polish(pu): polish config

* fix(pu): use the deepcopy train_data to add bonus reward

* polish(pu): add the operation of enlarging seq_length times to the last reward of the whole episode

* fix(pu): fix the episode length 1 bug and weight intrinsic reward bug

* feature(pu): add montezuma ngu config

* fix(pu): fix lunarlander ngu unroll_len to 998 so that the sequence length is equal to the max step 1000

* test(pu): episodic reward transforms to [0,1]

* fix(pu): fix r2d3 one-step rnn init bug and add r2d2_collect_traj

* fix(pu): fix r2d2_collect_traj.py

* feature(pu): add pong_r2d3_r2d2expert_config

* polish(pu): yapf format

* polish(pu): fix td.py conflict

* polish(pu): flake8 format

* polish(pu): add lambda_one_step_td key in dqfd error

* test(pu): set key lambda_one_step_td and lambda_supervised_loss as 0

* style(pu): yapf format

* style(pu): format

* polish(nyz): fix ngu detailed compatibility error

* fix(nyz): fix dqfd one_step td lambda bug

* fix(pu): fix test_acer and test_rnd compatibility error
Co-authored-by: NSwain <niuyazhe314@outlook.com>
Co-authored-by: NWill_Nie <nieyunpengwill@hotmail.com>

286ea243

31 10月, 2021 1 次提交
- N
  
  polish(nyz): remove on_policy option in dizoo config and entry · a6aa2c65
  由 niuyazhe 提交于 10月 30, 2021
  
  a6aa2c65
29 10月, 2021 4 次提交

feature(lcm): add MBPO algorithm (#113) · b1e9b4ea

由 Swain 提交于 10月 29, 2021

* feature(lcm): add MBPO algorithm (#87)

* add model-based rl

* fix yazhe's comments

* format

* pass flake8 test

* polish(nyz): polish mbpo import, name and test
Co-authored-by: Nlichuming <lichuming@lichumingdeMacBook-Pro.local>

b1e9b4ea

N

style(nyz): restrict deploy trigger branch · edb14698
由 niuyazhe 提交于 10月 29, 2021

edb14698
N

style(nyz): modify doc and deploy trigger and update mujoco license download link(smac docker) · c77ebf78
由 niuyazhe 提交于 10月 29, 2021

c77ebf78

feature(nyz): add PADDPG for hybrid action space as baseline (#109) · d2f79536

由 Swain 提交于 10月 29, 2021

* fix(nyz): fix gym_hybrid env not scale action bug

* feature(nyz): add PADDPG basic implementation for hybrid action space

* fix(nyz): fix td3/d4pg comatibility bug with new modifications

* fix(nyz): fix hybrid ddpg action type grad bug and update config

* feature(nyz): add eps greedy + multinomial wrapper and gym_hybrid ddpg convergence config

* style(nyz): update PADDPG in README

* test_model_hybrid_qac

* fix_typo_in_README

* test_policy_hybrid_qac

* polish(nyz): polish hybrid action space to dict structure and polish unittest

* fix(nyz): fix td3bc compatibility bug
Co-authored-by: N李可 <like2@CN0014008466M.local>

d2f79536

28 10月, 2021 1 次提交

feature(nyz): add gobigger baseline (#95) · a8fec8bb

由 Swain 提交于 10月 28, 2021

* feature(nyz): add gobigger baseline

* style(nyz): add gobigger env infor

* feature(nyz): add ignore prefix in default collate

* feautre(nyz): add vsbot training baseline

* fix(nyz): fix to_tensor empty list bug and polish gobigger baseline

* style(nyz): split gobigger baseline code

a8fec8bb

26 10月, 2021 2 次提交
- N
  
  polish(nyz): polish collector benchmark test(enable docker, smac docker) · fed80b44
  由 niuyazhe 提交于 10月 26, 2021
  
  fed80b44
- J
  test(yzj): add unittest for dataset, metric_serial_evaluator and learner (#107) · 0414eda5
  由 jayyoung0802 提交于 10月 26, 2021
```
* add 4 pytest dataset.py learner_aggregator.py learner_hook.py metric_serial_evaluator.py

* fix yapf and flake8 And remove invalid self._env

* fix fake_cls_config.py flake8
```
  0414eda5
25 10月, 2021 1 次提交

test(wyh): add more unittest for ppo and sac policy (#104) · c5af1cf2

由 Weiyuhong-1998 提交于 10月 25, 2021

* fix(wyh):reward model test

* fix(wyh):sac ppo test

* fix(wyh):ppo_continuous test

* fix(wyh):style

* fix(wyh):ppo test
Co-authored-by: NSwain <niuyazhe314@outlook.com>

c5af1cf2

22 10月, 2021 4 次提交

feature(zym): add offlineRL algo td3_bc and polish policy comments(#88) · 7c1b5e95

由 Yinmin.Zhang 提交于 10月 22, 2021

* feature(zym): add offlineRL algo td3_bc.

* feature(zym): add offlineRL algo td3_bc.

* feature(zym): add offlineRL algo td3_bc.

* polish(zym): polish some annotations in td3/ddpg/sac/ppo; polish `_forward_collect` and `_foward_eval`.

* fix(lj): fix dimension bug in cql for continuous env.

* fix(zym): fix dimension bug in cql for continuous env.

* fix(zym): fix dimension bug in cql for continuous env.

* polish(zym): update README.md.

7c1b5e95

polish(nyz): fix ppo bugs and update atari ppo offpolicy config (#108) · 2d5ec7c3

由 Swain 提交于 10月 22, 2021

* fix(nyz): fix ppo cuda bug and random collect bug

* config(nyz): add pong ppo off policy better config

* fix(nyz): fix ppo device bug in get_train_sample and update ppo offpolicy config

* style(nyz): correct yapf format

2d5ec7c3

N

fix(nyz): fix base policy model state_dict overlap bug · c2b14d48
由 niuyazhe 提交于 10月 22, 2021

c2b14d48
N

style(nyz): restrict ale-py version to 0.7.0(enable docker, smac docker) · f82f369d
由 niuyazhe 提交于 10月 22, 2021

f82f369d

21 10月, 2021 3 次提交

X
feature(xjx): test in pure docker environment (#103) · eee5c207
由 Xu Jingxin 提交于 10月 21, 2021
```
* Test in docker

* Add docker test entry

* Trap exit

* Test in docker
```
eee5c207

feature(lk): add gym-soccer (HFO) env (#94) · 8f47f4cb

由 Ke Li 提交于 10月 21, 2021

* add_soccer_env

* add_info

* close

* format

* test_gym_soccer

* rm_torch

* replay_log

* format_style

* add_gym_soccer_to_readme

* separate render_func

* add_gif_file

* scale_action

* flake_style_format

* resolve_review_comments

* add branch info for gym hybrid

8f47f4cb

N

polish(nyz): modify dizoo test mark to envtest(enable docker, smac docker) · f04b9eb7
由 niuyazhe 提交于 10月 21, 2021

f04b9eb7

20 10月, 2021 1 次提交
- X
  
  fix(xjx): replace distutil pyyaml with pip package (#99) · 0fcfdf26
  由 Xu Jingxin 提交于 10月 20, 2021
  
  0fcfdf26

OpenDILab开源决策智能平台 / DI-engine 上一次同步 2 年多

OpenDILab开源决策智能平台 / DI-engine
上一次同步 2 年多