Is there any way, I can add simple L1/L2 regularization in PyTorch? We can probably compute the regularized loss by simply adding the data_loss with the reg_loss but is there any explicit way, any support from PyTorch library to do it more easily without doing it manually?
7 Answers
Following should help for L2 regularization:
optimizer = torch.optim.Adam(model.parameters(), lr=1e-4, weight_decay=1e-5)
This is presented in the documentation for PyTorch. You can add L2 loss using the weight_decay parameter to the Optimization function.
Previous answers, while technically correct, are inefficient performance wise and are not too modular (hard to apply on a per-layer basis, as provided by, say, keras layers).
PyTorch L2 implementation
Why PyTorch implemented L2 inside torch.optim.Optimizer instances?
Let's take a look at torch.optim.SGD source code (currently as functional optimization procedure), especially this part:
for i, param in enumerate(params):
d_p = d_p_list[i]
# L2 weight decay specified HERE!
if weight_decay != 0:
d_p = d_p.add(param, alpha=weight_decay)
- One can see, that
d_p(derivative of parameter, gradient) is modified and re-assigned for faster computation (not saving the temporary variables) - It has
O(N)complexity without any complicated math likepow - It does not involve
autogradextending the graph without any need
Compare that to O(n) **2 operations, addition and also taking part in backpropagation.