- gradient 가 backpropagation 을 통해 layer 들을 지날 때 마다 exponential 하게 증가하는 현상
1 min read
해결책: Gated Recurrent Unit / Long Short-Term Memory / Truncated BTT machine_learning/optimization/exploding gradients gradient 가 backpropagation 을 통해 layer 들을 지날 때 마다 exponential 하게 증가하는 현상 해결책: gradient clipping
gradient 가 backpropagation 을 통해 layer 들을 지날 때 마다 exponential 하게 감소하는 현상 해결책: Gated Recurrent Unit / Long Short-Term Memory / Truncated BTT machine...
임의의 함수에 대한 gradient 는 그 함수의 partial derivatives 함수들을 element 로 가지는 vector 를 의미한다.
regularization > Weight Decay 를 참고할 것.
gradient descent 를 진행할 때, 모든 training example 을 전부 사용하는 것을 Batch Gradient Descent 라 부른다.
모델이 무거울수록 오래 학습할수록 성능이 좋아진다. “We increase model size, performance first gets worse and then gets better.
주로 최적화하는 Hyperparameter 들 Learning rate \alpha 가장 중요한 hyperparameter Number of layers Number of hidden units Learning rate decay Mini-batch size Momentum term 일반적으로 0.9...
ML 모델 h 에 대한 적합한 (\theta i 와 같은) parameter 를 찾기 위한 방법 Visualization of Gradient Descent 아래는 parameter \theta 0 와 \theta 1 에 대한 loss function J 의 등고선 그래프이다.
Early Stopping 전략이란 무엇인가 Early Stopping 은 overfitting 현상을 완화하는 방법으로, 학습을 하다가 일정 기준에 의해 학습을 중간에 멈추는 방법을 의미한다.
Newton’s method 라고 불리기도 하며, 실수 함수의 approximate 한 해를 빠르게 찾는 방법이다.
Gradient accumulation is a technique where you can train on bigger batch sizes than your machine would normally be able to fit into memory.