Min-max Scaling
values are shifted and rescaled so that they end up ranging from 0 to 1.
We do this by subtracting the min value and dividing by the max minus the min.
단점
outlier 에 취약한 편이다. 반대로, standardization is much less affected by outliers.
1 min read
values are shifted and rescaled so that they end up ranging from 0 to 1.
We do this by subtracting the min value and dividing by the max minus the min.
x′=max(x)−min(x)x−min(x)outlier 에 취약한 편이다. 반대로, standardization is much less affected by outliers.
...맥에 따라 의미가 조금 달라서, 실제로 어떤 변환을 말하는지 확인해야 한다. Feature scaling 문맥에서 normalization 이라고 하면 보통 값을 특정 범위로 맞추는 min-max scaling 을 가리키는 경우가 많다. B) Standardization 과 구분 standardization 은 평균을 빼고 표준편차로 나누어 z-score 로 바꾸는 방법이다.
...라고도 부른다. B) Normalization 과 구분 Normalization 은 문맥에 따라 넓게 쓰이는 말이라 혼동이 자주 생긴다. feature scaling 문맥에서는 보통 min-max scaling 처럼 값을 특정 범위로 맞추는 방법을 가리키고, 통계 문맥에서는 standardization 을 z-score normalization 이라고 부르기도 한다.
Normalization Normalization 은 데이터의 scale 을 맞추는 전처리를 넓게 부르는 말이다. 문맥에 따라 의미가 조금 달라서, 실제로 어떤 변환을 말하는지 확인해야 한다.
Standardization Standardization 은 feature 를 평균 0, 표준편차 1의 스케일로 바꾸는 전처리다.
Outlier outlier is a point for which y i is far from the value predicted by the model.
누락 데이터(결측값) 처리 결측값 처리는 데이터 전처리 단계에서 매우 중요합니다. 대표적인 접근 방식은 다음과 같습니다.
Dimension Reduction 데이터의 차원수를 줄이는 방법 B) Lesson Learned 가장 보편적인 방식은 PCA, UMAP, t-SNE 가 있다.
Layer Normalization Batch Normalization 이라고도 불리며, 배치에 있는 각 입력을 평균이 0 이고 분산이 1 을 가지도록 정규화하는 작업을 의미한다.
Manifold manifold 의 정의 고차원의 데이터를 저차원으로 옮길 때 데이터를 잘 설명하는 집합의 모형 매니폴드 가정 (manifold Hypothesis (assumption) 고차원의 데이터의 밀도는 낮지만, 이들의 집합을 포함하는 저차원의 manifold(subspace) 가 있다는 가정...
Permutation Importance Permutation Importance 는 feature importance 를 측정하기 위한 방법으로,모델을 학습시킨 뒤 특정 feature 의 데이터를 shuffle 했을 때, 검증 데이터 셋에 대한 예측성능을 확인하고 feature importance 를 계산한다.
Downsampling 데이터를 줄여서 샘플링하는 방법이다.
Categorical Feature sklearn.compose.ColumnTransformer : 각 열의 값들을 개별적으로 어떻게 transformation 할지 정할 수 있다.