Fix shape inconsistency in A2CPolicy and PPOPolicy #155

Trinkle23897 · 2020-07-21T05:33:43Z

The original r - v's shape in A2C is wrong.
The shape of log_prob is different: [bsz] in Categorical and [bsz, 1] in Normal. Should manually make the shape to be consistent with other tensors.

* make sure the key type of Batch is string, and add unit tests * add is_empty() function and unit tests * enable cat of mixing dict and Batch, just like stack

* minor polish * improve and implement Batch.cat_ * bugfix for buffer.sample with field impt_weight * restore the usage of a.cat_(b) * fix 2 bugs in batch and add corresponding unittest * code fix for update * update is_empty to recognize empty over empty; bugfix for len * bugfix for update and add testcase * add testcase of update * fix docs * fix docs * fix docs [ci skip] * fix docs [ci skip] Co-authored-by: Trinkle23897 <463003665@qq.com>

* re-implement Batch.stack and add testcases * add doc for Batch.stack * reuse _create_values and refactor stack_ & cat_ * fix pep8 * fix docs * raise exception for stacking with partial keys and axis!=0 * minor fix * minor fix Co-authored-by: Trinkle23897 <463003665@qq.com>

* remove multibuf * reward_metric * make fileds with empty Batch rather than None after reset * many fixes and refactor Co-authored-by: Trinkle23897 <463003665@qq.com>

* Enable selecting worker for vector env step method. * Update collector to match new vecenv selective worker behavior. * Bug fix. * Fix rebase Co-authored-by: Alexis Duburcq <alexis.duburcq@wandercraft.eu>

* code refactor; remove unused kwargs; add reward_normalization for dqn * bugfix for __setitem__ with torch.Tensor; add Batch.condense * minor fix * support cat with empty Batch * remove the dependency of is_empty on len; specify the semantic of empty Batch by test cases * support stack with empty Batch * remove condense * refactor code to reflect the shared / partial / reserved categories of keys * add is_empty(recursive=False) * doc fix * docfix and bugfix for _is_batch_set * add doc for key reservation * bugfix for algebra operators * fix cat with lens hint * code refactor * bugfix for storing None * use ValueError instead of exception * hide lens away from users * add comment for __cat * move the computation of the initial value of lens in cat_ itself. * change the place of doc string * doc fix for Batch doc string * change recursive to recurse * doc string fix * minor fix for batch doc

* add doc for len exceptions * doc move; unify is_scalar_value function * remove some issubclass check * bugfix for shape of Batch(a=1) * keep moving doc * keep writing batch tutorial * draft version of Batch tutorial done * improving doc * keep improving doc * batch tutorial done * rename _is_number * rename _is_scalar * shape property do not raise exception * restore some doc string * grammarly [ci skip] * grammarly + fix warning of building docs * polish docs * trim and re-arrange batch tutorial * go straight to the point * minor fix for batch doc * add shape / len in basic usage * keep improving tutorial * unify _to_array_with_correct_type to remove duplicate code * delegate type convertion to Batch.__init__ * further delegate type convertion to Batch.__init__ * bugfix for setattr * add a _parse_value function * remove dummy function call * polish docs Co-authored-by: Trinkle23897 <463003665@qq.com>

* Enable selecting worker for vector env step method. * Update collector to match new vecenv selective worker behavior. * Bug fix. * Fix rebase Co-authored-by: Alexis Duburcq <alexis.duburcq@wandercraft.eu>

* code refactor; remove unused kwargs; add reward_normalization for dqn * bugfix for __setitem__ with torch.Tensor; add Batch.condense * minor fix * support cat with empty Batch * remove the dependency of is_empty on len; specify the semantic of empty Batch by test cases * support stack with empty Batch * remove condense * refactor code to reflect the shared / partial / reserved categories of keys * add is_empty(recursive=False) * doc fix * docfix and bugfix for _is_batch_set * add doc for key reservation * bugfix for algebra operators * fix cat with lens hint * code refactor * bugfix for storing None * use ValueError instead of exception * hide lens away from users * add comment for __cat * move the computation of the initial value of lens in cat_ itself. * change the place of doc string * doc fix for Batch doc string * change recursive to recurse * doc string fix * minor fix for batch doc

* add doc for len exceptions * doc move; unify is_scalar_value function * remove some issubclass check * bugfix for shape of Batch(a=1) * keep moving doc * keep writing batch tutorial * draft version of Batch tutorial done * improving doc * keep improving doc * batch tutorial done * rename _is_number * rename _is_scalar * shape property do not raise exception * restore some doc string * grammarly [ci skip] * grammarly + fix warning of building docs * polish docs * trim and re-arrange batch tutorial * go straight to the point * minor fix for batch doc * add shape / len in basic usage * keep improving tutorial * unify _to_array_with_correct_type to remove duplicate code * delegate type convertion to Batch.__init__ * further delegate type convertion to Batch.__init__ * bugfix for setattr * add a _parse_value function * remove dummy function call * polish docs Co-authored-by: Trinkle23897 <463003665@qq.com>

* Enable selecting worker for vector env step method. * Update collector to match new vecenv selective worker behavior. * Bug fix. * Fix rebase Co-authored-by: Alexis Duburcq <alexis.duburcq@wandercraft.eu>

* code refactor; remove unused kwargs; add reward_normalization for dqn * bugfix for __setitem__ with torch.Tensor; add Batch.condense * minor fix * support cat with empty Batch * remove the dependency of is_empty on len; specify the semantic of empty Batch by test cases * support stack with empty Batch * remove condense * refactor code to reflect the shared / partial / reserved categories of keys * add is_empty(recursive=False) * doc fix * docfix and bugfix for _is_batch_set * add doc for key reservation * bugfix for algebra operators * fix cat with lens hint * code refactor * bugfix for storing None * use ValueError instead of exception * hide lens away from users * add comment for __cat * move the computation of the initial value of lens in cat_ itself. * change the place of doc string * doc fix for Batch doc string * change recursive to recurse * doc string fix * minor fix for batch doc

* add doc for len exceptions * doc move; unify is_scalar_value function * remove some issubclass check * bugfix for shape of Batch(a=1) * keep moving doc * keep writing batch tutorial * draft version of Batch tutorial done * improving doc * keep improving doc * batch tutorial done * rename _is_number * rename _is_scalar * shape property do not raise exception * restore some doc string * grammarly [ci skip] * grammarly + fix warning of building docs * polish docs * trim and re-arrange batch tutorial * go straight to the point * minor fix for batch doc * add shape / len in basic usage * keep improving tutorial * unify _to_array_with_correct_type to remove duplicate code * delegate type convertion to Batch.__init__ * further delegate type convertion to Batch.__init__ * bugfix for setattr * add a _parse_value function * remove dummy function call * polish docs Co-authored-by: Trinkle23897 <463003665@qq.com>

tianshou/policy/modelfree/a2c.py

tianshou/policy/modelfree/ddpg.py

tianshou/policy/modelfree/ppo.py

tianshou/policy/modelfree/sac.py

tianshou/policy/modelfree/a2c.py

tianshou/policy/modelfree/ddpg.py

Trinkle23897 · 2020-07-21T13:40:32Z

Should be okay now :)

youkaichao · 2020-07-21T13:50:27Z

It should be better to document somewhere the shape of these variables. ndarray with shape of (n,) is very error-prone and annoying because of mis-broadcasting.

duburcqa · 2020-07-21T13:51:07Z

It should be better to document somewhere the shape of these variables.

I agree.

- fix 2 warning in doctest - change the minimum version of gym (to be aligned with openai baselines) - change squeeze and reshape to flatten (related to #155). I think flatten is better.

- The original `r - v`'s shape in A2C is wrong. - The shape of log_prob is different: [bsz] in Categorical and [bsz, 1] in Normal. Should manually make the shape to be consistent with other tensors.

- fix 2 warning in doctest - change the minimum version of gym (to be aligned with openai baselines) - change squeeze and reshape to flatten (related to thu-ml#155). I think flatten is better.

youkaichao and others added 18 commits July 11, 2020 09:44

Improve Batch (#126)

e976d74

* make sure the key type of Batch is string, and add unit tests * add is_empty() function and unit tests * enable cat of mixing dict and Batch, just like stack

Improve collector (#125)

885fbc1

* remove multibuf * reward_metric * make fileds with empty Batch rather than None after reset * many fixes and refactor Co-authored-by: Trinkle23897 <463003665@qq.com>

Vector env enable select worker (#132)

cee8088

* Enable selecting worker for vector env step method. * Update collector to match new vecenv selective worker behavior. * Bug fix. * Fix rebase Co-authored-by: Alexis Duburcq <alexis.duburcq@wandercraft.eu>

Vector env enable select worker (#132)

c198c60

* Enable selecting worker for vector env step method. * Update collector to match new vecenv selective worker behavior. * Bug fix. * Fix rebase Co-authored-by: Alexis Duburcq <alexis.duburcq@wandercraft.eu>

Merge branch 'dev' into dev

977b627

Vector env enable select worker (#132)

988a13d

* Enable selecting worker for vector env step method. * Update collector to match new vecenv selective worker behavior. * Bug fix. * Fix rebase Co-authored-by: Alexis Duburcq <alexis.duburcq@wandercraft.eu>

Merge branch 'dev' of github.com:trinkle23897/tianshou into dev

9d31801

fix a2c

20db334

Merge branch 'dev' into fix-policy-shape

58057f5

minor update

a802766

Trinkle23897 changed the title ~~Fix shape inconsistency in policy~~ Fix shape inconsistency in A2CPolicy Jul 21, 2020

duburcqa reviewed Jul 21, 2020

View reviewed changes

tianshou/policy/modelfree/a2c.py Outdated Show resolved Hide resolved

duburcqa reviewed Jul 21, 2020

View reviewed changes

tianshou/policy/modelfree/ddpg.py Outdated Show resolved Hide resolved

duburcqa reviewed Jul 21, 2020

View reviewed changes

tianshou/policy/modelfree/ppo.py Outdated Show resolved Hide resolved

duburcqa reviewed Jul 21, 2020

View reviewed changes

tianshou/policy/modelfree/sac.py Outdated Show resolved Hide resolved

Trinkle23897 added 3 commits July 21, 2020 15:04

Merge branch 'dev' into fix-policy-shape

fb2d490

minor update

f30dc25

Merge branch 'dev' into fix-policy-shape

160db40

duburcqa previously approved these changes Jul 21, 2020

View reviewed changes

fix a2c bug

877b324

Trinkle23897 dismissed duburcqa’s stale review via 877b324 July 21, 2020 11:44

Trinkle23897 changed the title ~~Fix shape inconsistency in A2CPolicy~~ WIP: Fix shape inconsistency in A2CPolicy Jul 21, 2020

Trinkle23897 changed the title ~~WIP: Fix shape inconsistency in A2CPolicy~~ Fix shape inconsistency in A2CPolicy and PPOPolicy Jul 21, 2020

Trinkle23897 added 2 commits July 21, 2020 19:54

fix ppo

42287b7

all squeeze

e88e100

duburcqa reviewed Jul 21, 2020

View reviewed changes

tianshou/policy/modelfree/a2c.py Outdated Show resolved Hide resolved

tianshou/policy/modelfree/ddpg.py Outdated Show resolved Hide resolved

squeeze with dim

4c52ef3

Trinkle23897 requested review from youkaichao and duburcqa July 21, 2020 13:41

Trinkle23897 added 2 commits July 21, 2020 22:01

remove squeeze()

ac83c4f

add a warning

a736526

duburcqa approved these changes Jul 21, 2020

View reviewed changes

youkaichao approved these changes Jul 21, 2020

View reviewed changes

Trinkle23897 merged commit 089b85b into thu-ml:dev Jul 21, 2020

Trinkle23897 deleted the fix-policy-shape branch July 21, 2020 14:24

Trinkle23897 mentioned this pull request Jul 23, 2020

Does the current PPO support discrete actions? #101

Closed

Trinkle23897 linked an issue Jul 23, 2020 that may be closed by this pull request

Does the current PPO support discrete actions? #101

Closed

Trinkle23897 mentioned this pull request Jul 23, 2020

3 fix #158

Merged

Trinkle23897 added a commit that referenced this pull request Jul 23, 2020

3 fix (#158)

352a518

- fix 2 warning in doctest - change the minimum version of gym (to be aligned with openai baselines) - change squeeze and reshape to flatten (related to #155). I think flatten is better.

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Uh oh!

Fix shape inconsistency in A2CPolicy and PPOPolicy #155

Fix shape inconsistency in A2CPolicy and PPOPolicy #155

Uh oh!

Trinkle23897 commented Jul 21, 2020 •

edited

Loading

Uh oh!

Uh oh!

Uh oh!

Uh oh!

Uh oh!

Uh oh!

Uh oh!

Trinkle23897 commented Jul 21, 2020

Uh oh!

youkaichao commented Jul 21, 2020

Uh oh!

duburcqa commented Jul 21, 2020

Uh oh!

Uh oh!

Fix shape inconsistency in A2CPolicy and PPOPolicy #155

Fix shape inconsistency in A2CPolicy and PPOPolicy #155

Uh oh!

Conversation

Trinkle23897 commented Jul 21, 2020 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

Uh oh!

Uh oh!

Uh oh!

Uh oh!

Uh oh!

Uh oh!

Uh oh!

Trinkle23897 commented Jul 21, 2020

Uh oh!

youkaichao commented Jul 21, 2020

Uh oh!

duburcqa commented Jul 21, 2020

Uh oh!

Uh oh!

Trinkle23897 commented Jul 21, 2020 •

edited

Loading