Exploding Gradients
- gradient 가 backpropagation 을 통해 layer 들을 지날 때 마다 exponential 하게 증가하는 현상
1 min read
해결책: Gated Recurrent Unit / Long Short-Term Memory / Truncated BTT exploding gradients gradient 가 backpropagation 을 통해 layer 들을 지날 때 마다 exponential 하게 증가하는 현상 해결책: gradient clipping
Vanishing Gradients gradient 가 backpropagation 을 통해 layer 들을 지날 때 마다 exponential 하게 감소하는 현상 해결책: Gated Recurrent Unit / Long Short-Term Memory / Truncated BTT exploding...
Double Descent 모델이 무거울수록 오래 학습할수록 성능이 좋아진다. “We increase model size, performance first gets worse and then gets better.
Global Minimum Global minimum 은 목적 함수 전체 영역에서 가장 작은 값을 가지는 지점이다. x^\ = \arg\min x f(x) 어떤 지점이 주변에서는 가장 작지만 전체에서는 더 작은 지점이 따로 있다면 local minimum 이다.
Gradient Accumulation Gradient accumulation is a technique where you can train on bigger batch sizes than your machine would normally be able to fit into memory.
Newton-Raphson Method Newton’s method 라고 불리기도 하며, 실수 함수의 approximate 한 해를 빠르게 찾는 방법이다.
Exponentially Weighted (moving) Averages SGD 보다 좋은 최적화 알고리즘은 Exponentially Weighted Averages (지수 가중치 평균) 개념을 이용한다 (또는 통계에서는 지수이동평균이라 부른다).
Optimization Problem 최적화 문제 (Optimization problems) 란 여러개의 선택가능한 후보 중에서 최적의 해 (Optimal value) 또는 최적의 해에 근접한 값을 찾는 문제를 일컫는다.
Momentum(모멘텀)은 최적화 알고리즘에서 사용되는 기법으로, 이전 단계의 업데이트 방향을 일정 비율로 반영하여 학습의 속도를 높이고 진동을 줄이는 역할을 합니다. Momentum의 수식은 다음과 같습니다.
Lagrange Multiplier Method Method of finding a local maximum subject to constraints.
Gradient Checkpoint GitHub - cybertronai/gradient-checkpointing: Make huge neural nets fit in memory.