1 min read
values are shifted and rescaled so that they end up ranging from 0 to 1.
Permutation Importance 는 feature importance 를 측정하기 위한 방법으로,모델을 학습시킨 뒤 특정 feature 의 데이터를 shuffle 했을 때, 검증 데이터 셋에 대한 예측성능을 확인하고 feature importance 를 계산한다.
데이터의 차원수를 줄이는 방법 Lesson Learned 가장 보편적인 방식은 PCA, UMAP, t-SNE 가 있다. 선형 변환 방식인 PCA 은 비선형 변환 방식인 UMAP 또는 t-SNE 와 비교하진 않고, 서로 조합하면서 사용하는 것 같다.
covariate shift 는 ML 에서 종종 마주치는 특수한 형태의 데이터 변화를 의미한다. References www.seldon.io/what-is-covariate-shift .
Filtering Method 통계적 기법이나 알고리즘을 이용하여 feature 간 상관 관계를 계산함 t-test \chi^2-test information gain information value Wrapper Method Forward Greedy Backward Greedy Genetic Search...
SHAP 는 feature importance(attribution) 을 파악할 수 있는 방법이다.
Overfitting 이란 모델이 특정 데이터 셋에 과도하게 적합된 것을 의미 Overfitting 문제를 완화하기 위한 접근들 더 많은 training data Data Augmentation regularization dropout Early Stopping .
Mlib MLlib is Spark’s machine learning (ML) library. Its goal is to make practical machine learning scalable and easy.
References [ICR IARC, 2023] EDA and Submission | Kaggle.
deep Learning 과 달리 structured data 가 필요한 학습 모델 Machine Learning Model Representation 위 그림의 h 는 가설 (hypothesis) 을 뜻함 h:X\rightarrow Y X 는 입력 데이터 공간 (space), Y 는 출력 데이터 공간.