Categorical Feature
sklearn.compose.ColumnTransformer: 각 열의 값들을 개별적으로 어떻게 transformation 할지 정할 수 있다. 예를 들면, 한 열은OneHotEncoder그리고 나머지는 적용안하는passthrough
1 min read
Category Feature Related References Google Developers Blog: Introducing TensorFlow Feature Columns .
Feature Selection 1.1. Filtering Method 통계적 기법이나 알고리즘을 이용하여 feature 간 상관 관계를 계산함 t-test \chi^2-test information gain information value 1.2.
누락 데이터(결측값) 처리 결측값 처리는 데이터 전처리 단계에서 매우 중요합니다. 대표적인 접근 방식은 다음과 같습니다.
Permutation Importance Permutation Importance 는 feature importance 를 측정하기 위한 방법으로,모델을 학습시킨 뒤 특정 feature 의 데이터를 shuffle 했을 때, 검증 데이터 셋에 대한 예측성능을 확인하고 feature importance 를 계산한다.
Feature Crosses 서로 다른 성격 (도메인) 을 지닌 feature 를 섞어서 하나의 context 로 표현하는 것 (예시) 어느 사람이 바나나를 샀고, 요리책을 샀다면 믹서기를 살 확률이 높다.
Box-cox Transformation 데이터를 정규 분포 에 가깝게 만들어 주는 변환 방법 x>0 에 대하여 box-cox 변환은 다음을 만족하도록 한다.
Regression 회귀는 예측하는 함수를 만드는 것이다. 예측하려는 데이터의 종류가 numerical 또는 categorical data 에 따라서 사용하는 알고리즘이 달라진다.
Dimension Reduction 데이터의 차원수를 줄이는 방법 B) Lesson Learned 가장 보편적인 방식은 PCA, UMAP, t-SNE 가 있다.
Classification 일반적인 classification 문제: y\in{0,1} (0 은 Negative Class, 1 은 Positive Class) E-mail: 스팸 yes or no? 온라인 거래: 사기 yes or no? Tumor: 양성 or 음성?...
K-prototype K-prototype 은 k-means 와 k-mode 를 결합하여 수치적 데이터와 범주적 데이터가 모두 있는 데이터 세트를 처리하는 클러스터링 알고리즘 (clustering) B) References Detailed EDA | k-prototypes clustering | Kaggle...