유명한 아이템은 더 많이 노출되어서 학습 데이터의 밸런스가 더욱 무너지는 현상
1 min read
유명한 아이템은 더 많이 노출되어서 학습 데이터의 밸런스가 더욱 무너지는 현상
개인화의 레벨을 감소시키고, serendipity 의 정도를 낮춤 추천의 fairness 정도를 감소시킴 Matthew effect issue
Matthew effect
Tags Matthew effect Average Recommendation Popularity Scaling Method Collaborative filtering via high-dimensional regression Related Papers Popularity Bias in Dynamic...
Exposure bias 는 사용자가 모든 item 을 본 것이 아니라, 노출된 item 에 대해서만 feedback 을 남긴다는 데서 생기는 bias 다.
paper link Abstract Blindly fitting the data without considering the inherent biases will result in many serious issues, e.g., the discrepancy between offline evaluation and...
배경 To remedy the selection bias in evaluation, some recent work considers a recommendation as an intervention analogous to treating a patient with a specific drug.
Data Bias 란 수집된 학습 데이터의 분포가 이상적인 테스트 데이터의 분포와 다른 것을 의미한다. 아무리 많은 데이터를 수집하더라도, 분포 자체가 다르면 estimation 과 optimal function 간 gap 이 생길 수 밖에 없다.
데이터를 하나의 저장소로 모으지 않고 모델을 학습할 수 없을까? 데이터를 특정 서버에 모으지 않고 개인 디바이스에서 학습하는 방식으로, 데이터 이동 이슈를 해소하기 위한 방안으로 많이 연구되고 있다.
References A Troubling Analysis of Reproducibility and Progress in Recommender Systems Research, TOIS, 2021.
IPW 관측될 확률이 낮았던 샘플에 더 큰 가중치를 주어, 편향된 로그로부터 편향 없는 추정값을 얻는 방법이다. 각 샘플에 그 샘플이 관측될 확률(propensity)의 역수를 곱한다.
collaborative filtering 의 MF 기법은 gradient descent update 를 이용하여 user 와 item 의 latent vector 를 찾아내는데, 이러한 최적화 과정은 너무 느리고 많은 반복이 필요하다.
Two types of selection biases system item selection process users’ self-selection behavior Example 내 의견 viewable impression log 를 수집할 수 있다면 self-selection behavior 보다는 system...