Adding L1/L2 Regularization in Pytorch?

Adding L1/L2 Regularization in Pytorch?

Is there any way, I can add simple L1/L2 regularization in PyTorch? We can probably compute the regularized loss by simply adding the data_loss with the reg_loss but is there any explicit way, any support from PyTorch library to do it more easily without doing it manually?

7 Answers

Following should help for L2 regularization:

optimizer = torch.optim.Adam(model.parameters(), lr=1e-4, weight_decay=1e-5)
2

This is presented in the documentation for PyTorch. You can add L2 loss using the weight_decay parameter to the Optimization function.

5

Previous answers, while technically correct, are inefficient performance wise and are not too modular (hard to apply on a per-layer basis, as provided by, say, keras layers).

PyTorch L2 implementation

Why PyTorch implemented L2 inside torch.optim.Optimizer instances?

Let's take a look at torch.optim.SGD source code (currently as functional optimization procedure), especially this part:

for i, param in enumerate(params):
    d_p = d_p_list[i]
    # L2 weight decay specified HERE!
    if weight_decay != 0:
        d_p = d_p.add(param, alpha=weight_decay)
  • One can see, that d_p (derivative of parameter, gradient) is modified and re-assigned for faster computation (not saving the temporary variables)
  • It has O(N) complexity without any complicated math like pow
  • It does not involve autograd extending the graph without any need

Compare that to O(n) **2 operations, addition and also taking part in backpropagation.

David Miller
Author

David Miller

David Miller brings 15 years of experience in global economics, personal finance strategy, and market dynamics. He specializes in turning complex economic trends into actionable insights for everyday readers.