Hierarchical Softmax
Hierarchical Softmax 는 softmax function 를 전체로 계산하기 보다는 Tree 구조로 Hierarchical 하게 Softmax 를 계산하는 방법이다.

1 min read
Hierarchical Softmax 는 softmax function 를 전체로 계산하기 보다는 Tree 구조로 Hierarchical 하게 Softmax 를 계산하는 방법이다.

hierarchical softmax 보다는 negative sampling 방식이 더 효과가 좋았음
How? subsampling of the frequent words negative sampling 제안 (hierarchical softmax 의 alternative)
D) Related hierarchical softmax E) References
...그리고 V 는 단어의 총 개수다. 일반적으로 V 는 매우 큰 편이라, 매번 확률 vector 를 계산할 때 많은 계산량이 요구된다. softmax 의 계산량 이슈를 해소하기 위해 hierarchical softmax 그리고 negative sampling 가 제시되었다 (주로 negative sampling 이 많이 이용된다). E) Related F) References
...put Layer 에서 모든 단어에 대한 Softmax 계산을 해야하기 때문에, 이에 따른 연산량이 막대하다. 이 부분의 계산량을 줄이기 위한 방법이 두가지 제안되었는데, 하나는 hierarchical softmax 고, 다른 하나는 negative sampling 이다. hierarchical softmax 와 negative sampling 은 확률 값 계산의 계산량을 줄이기 위한 방법으...
Negative Sampling Negative Sampling 은 skip-gram 모델의 softmax 함수 계산 비용을 절약하기 위해 고안된 Word Embedding 방식이다.
Word2Vec Word2Vec 은 shallow neural network 를 통해 단어를 저차원 vector space 로 임베딩 하는 모델이다. 모델의 학습 결과로 얻어지는 vector 간 거리가 가깝다는 것은 단어 간 의미가 서로 비슷함을 의미한다.
Skip Gram Skip Gram 은 중심 단어를 통해 주변단어를 예측하는 Word Embedding 모델이다.
paper link Abstract skip-gram 모델의 extension 을 제안함으로써 학습 속도와 vector quality 상승을 보였다.
Softmax Function class K>2 의 경우에서 generalized linear model (reference) \displaystyle p\left(C {k}\mid\mathbf{x}\right)=\frac{p\left(\mathbf{x}\mid C {k}\right)p\left(C...
Word Embedding 각 단어에 대한 n 차원 feature 를 만들고 적절한 값을 부여하는 방식을 의미한다. A.1) 예시 One Hot Encoding 을 통해서도 Word Embedding 이 가능하긴 하다.
One Hot Encoding Numerical 데이터로 바꿀 때, 0 또는 1 만을 사용하는 방법을 의미한다.
Bag of Words 단어들의 순서는 전혀 고려하지 않고, 단어들의 출현 빈도 (frequency) 에만 집중하는 텍스트 데이터의 수치화 표현 방법 A.1) 장점 간단하다. A.2) 단점 단어 간 순서의 정보를 잃게 된다.
Doc2Vec 문서 내 단어를 벡터로 임베딩하는 Word2Vec 과 달리, 전체 문서를 벡터로 임베딩하는 모델이다. 일반적으로 Paragraph Vector model 이라 불리며, Word2Vec 에서 얻어지는 벡터를 평균내는 방식보다 효과가 좋다고 알려져있다.
CBOW CBOW 는 주변단어 (Context) 를 통해서 주어진 단어 (target) 가 무엇인지 찾는 것이다. 정확히는 앞뒤로 \frac{c}{2} 개의 단어를 (총 c 개) 통해 주어진 단어를 예측한다는 것이 CBOW 의 아이디어이다.