References
- A Troubling Analysis of Reproducibility and Progress in Recommender Systems Research, TOIS, 2021.
1 min read
Related userKNN, itemKNN
데이터를 하나의 저장소로 모으지 않고 모델을 학습할 수 없을까? 데이터를 특정 서버에 모으지 않고 개인 디바이스에서 학습하는 방식으로, 데이터 이동 이슈를 해소하기 위한 방안으로 많이 연구되고 있다.
유명한 아이템은 더 많이 노출되어서 학습 데이터의 밸런스가 더욱 무너지는 현상.
KNN 어떤 문제를 푸냐에 따라 방식이 달라진다: 분류 또는 회귀. For classification: 비교 대상이 되는 데이터 주변에 가장 가까이 존재하는 k 개의 데이터와 비교해 가장 가까운 데이터 종류로 판별한다.
Item-based KNN algorithms were discussed in 2001 [33] and later successfully applied in industry around 2003 References Amazon.com Recommendations: Item-to-Item Collaborative...
collaborative filtering 의 MF 기법은 gradient descent update 를 이용하여 user 와 item 의 latent vector 를 찾아내는데, 이러한 최적화 과정은 너무 느리고 많은 반복이 필요하다.
지원 알고리즘 Word2Vec Mikolov, Tomas, et al.
추천 시스템의 학습은 supervised learning 과 다르다. 추천 시스템은 오직 사용자가 선택한 결과만을 가지고 학습하기 때문에, log data 는 partial information 형식을 지닌다.
대화형 추천 시스템: 사용자와 실제 대화를 통해 사용자의 성향을 파악하여 추천을 진행함.
추천 시스템에 관련된 github repo 를 정리하는 page Repos github.com/microsoft/recommenders/ This repository contains examples and best practices for building recommendation systems, provided...
Data Bias 란 수집된 학습 데이터의 분포가 이상적인 테스트 데이터의 분포와 다른 것을 의미한다. 아무리 많은 데이터를 수집하더라도, 분포 자체가 다르면 estimation 과 optimal function 간 gap 이 생길 수 밖에 없다.