Zzong's Notes

Home

❯

machine_learning

❯

unsupervised learning

unsupervised learning

2026년 6월 14일1 min read

  • clustering
    • K-means
    • DBSCAN
  • anomaly detection
    • One-class support vector machine
    • Isolation Forest
  • dimensionality reduction
    • Principal Component Analysis
    • Kernel PCA
    • Locally Linear Embedding
    • t-Stochastic Nearest Embedding
  • Association rule learning
    • Apriori
    • Eclat

링크된 언급

4
Bidirectional Encoder Representations from Transformers

...ned) 언어 모델이다. BERT is designed to pre-train deep bidirectional representations from unlabeled text(unsupervised learning) by jointly conditioning on both left and right context in all layers. 기존 방식들은 fine-tuning 을 잘 하지 ...

Gensim

Gensim Gensim 은 문서를 semantic vector 로 효율적으로 표현하기 위한 파이썬 라이브러리이다. Gensim 은 구조화 되어있지 않은 plain text 를 unsupervised learning 알고리즘을 통해 처리한다. Gensim 에서 지원하는 비지도 학습 알고리즘은 Word2Vec, FastText, LDA, LSA 등 이 있다. Related References

supervised learning

unsupervised learning

T5

C.1.3) 학습 방식 Span corruption task 를 통해 unsupervised learning (MLM) 만 적용했다고 한다. D) Related E) References

함께 보면 좋은 글

supervised learning

Tags unsupervised learning k-Nearest Neighbors linear regression logistic regression support vector machine Decision Tree and random forest .

clustering

Clustering B) 다차원에서의 클러스터링 다차원 (high dimensional) 데이터를 이용한 클러스터링은 의미없을 수 있다 (may be meaningless).

t-Stochastic Nearest Embedding

t-SNE t-Stochastic Nearest Embedding 는 vector visualization 을 위하여 자주 이용되는 차원 축소 알고리즘이다.

Principal Component Analysis

PCA N 개의 i.i.d. 를 만족하는 데이터 포인트들 \mathbf{X}=\left[\mathbf{x} {1},\ldots,\mathbf{x} {N}\right]^{T} 이 존재하고, 각 \mathbf{x} 는 D 차원 벡터라고 하자.

K-means

K-means A.1) 시간 복잡도 O(kNrD) k : 클러스터 개수 (사용자에 의해 정의됨) N : 객체 개수 r: 수렴할때 까지 반복한 iteration 횟수 D : 객체의 차원 수 (window 를 이용한 clustering 의 경우, window 길이) A.2) Optimization Object...

support vector machine

Classification with SVM binary classification 의 핵심은 D 차원 공간에 존재하는 example 들을 레이블에 따라 두 파티션으로 나눈다는 것이다. 이때, 공간을 나누는 역할을 hyperplane 이 맡게된다.

Gensim

Gensim Gensim 은 문서를 semantic vector 로 효율적으로 표현하기 위한 파이썬 라이브러리이다. Gensim 은 구조화 되어있지 않은 plain text 를 unsupervised learning 알고리즘을 통해 처리한다.

dimension reduction

Dimension Reduction 데이터의 차원수를 줄이는 방법 B) Lesson Learned 가장 보편적인 방식은 PCA, UMAP, t-SNE 가 있다.

K-prototype

K-prototype K-prototype 은 k-means 와 k-mode 를 결합하여 수치적 데이터와 범주적 데이터가 모두 있는 데이터 세트를 처리하는 클러스터링 알고리즘 (clustering) B) References Detailed EDA | k-prototypes clustering | Kaggle...

Eclat

Eclat Eclat 은 교집합을 이용한 DFS 방식이다. 그래서 Apriori 과 달리 parallel 하게 수행하는 것이 가능하다. B) Eclat 의 특징 Eclat 은 Apriori 와 비교했을 때 메모리에 모두 적재될 수 있는 적은 데이터셋에 적합하다.