pytorch
695eef05 - optimizer exploration - v1 and v2 + fix position_weighted optimizer + decoupled weight decay (#54042)

Commit View On GitHub

Commit

3 years ago

optimizer exploration - v1 and v2 + fix position_weighted optimizer + decoupled weight decay (#54042) Summary: Pull Request resolved: https://github.com/pytorch/pytorch/pull/54042 Pull Request resolved: https://github.com/pytorch/pytorch/pull/53881 1. Fix position_weighted optimizer: Position weighted layer uses default optimizer but is actually gradient_slice, which will cause problem if we do not handle it properly in the new optimizier. The solution is to use sparseadagrad when it is gradient_slices. 2. Optimizer implementation of v1 and v2: using 1st momentum with/without bias_correction. 3. also implemented decoupled weight decay in the new optimizer. Test Plan: buck test //caffe2/caffe2/fb/dper/layer_models/tests/split_1:sparse_nn_test_2 -- test_mlp_optimization buck test //caffe2/caffe2/python:optimizer_test -- TestDecayAdagrad buck test //caffe2/caffe2/python/operator_test:decay_adagrad_test ctr_mbl_feed work flow: f255731660 oc work flow: f255739503 Reviewed By: 0x10cxR1 Differential Revision: D26839668 fbshipit-source-id: 2b6881c1a88540ef5766be40f5e80001257e2199

Author

lanlanfb

Committer

facebook-github-bot

Parents

5c3d80d8

pytorch 695eef05 - optimizer exploration - v1 and v2 + fix position_weighted optimizer + decoupled weight decay (#54042)

Commit

pytorch
695eef05 - optimizer exploration - v1 and v2 + fix position_weighted optimizer + decoupled weight decay (#54042)