Zzong's Notes

Home

❯

papers

❯

bandit

❯

Recommender systems using LinUCB - A contextual multi-armed bandit approach

Recommender systems using LinUCB - A contextual multi-armed bandit approach

2026년 7월 20일1 min read

  • Metadata
    • Tag: Contextual Bandit, Thompson sampling, LinUCB, Multi-Armed Bandit
    • Link: https://towardsdatascience.com/recommender-systems-using-linucb-a-contextual-multi-armed-bandit-approach-35a6f0eb6c4
    • Related Reference: git_issue_Contextual Bandit 을 활용한 개인화추천 성능 고도화 실험, A Contextual-Bandit Approach to Personalized News Article Recommendation
      • Even though the recorded mean award achieved for Arm 3 is higher, the algorithm selects arm 2 because of the uncertainty of its potential and updates its confidence bound for future trials.

함께 보면 좋은 글

A Contextual-Bandit Approach to Personalized News Article Recommendation

Abstract 사용자와 콘텐츠 정보를 활용한 개인화 웹 서비스 (광고, 뉴스 등) 를 제공하는 것은 다음과 같은 두 가지 이유로 어렵다.

Mortal Multi-Armed Bandits

Mortal Multi-Armed Bandits Mortal MAB, Multi-Armed Bandit Related Mortal Multi Armed Bandit (2008) References...

Top-K Contextual Bandits with Equity of Exposure

Top-K Contextual Bandits with Equity of Exposure Abstract Probability Ranking Principle 에 의하면 top-K items 을 greedy 하게 rank 하는것이 optimal 함 대신 Introduction This work...

Optimal Regret Analysis of Thompson Sampling in Stochastic Multi-armed Bandit Problem with Multiple Plays

Optimal Regret Analysis of Thompson Sampling in Stochastic Multi-armed Bandit Problem with Multiple Plays B) Introduction Thompson sampling is an old heuristic that has a...

An Asymptotically Optimal Primal-Dual Incremental Algorithm for Contextual Linear Bandits

Links Paper link Abstract optimism principle 에 기반한 알고리즘은 문제에 대한 구조를 exploit 하는데 실패하여 점근적으로 suboptimal 결과를 보임 context 분포와 exploration policy 가 나눠지도록 (decoupled) regret lower...

Contextual Combinatorial Bandit and its Application on Diversified Online Recommendation

Related B) References www.chenshouyuan.com/papers/sdm14.pdf .

Burst-induced Multi-Armed Bandit for Learning Recommendation

Burst-induced Multi-Armed Bandit for Learning Recommendation B) Abstract 해결하려는 문제: a non-stationary and context-free Multi-Armed Bandit problem, 유저나 아이템에 대한 어떠한 정보가 없는 경우 C)...

Deep Bayesian Bandits Showdown - An Empirical Comparison of Bayesian Deep Networks for Thompson Sampling

Summary 이 논문의 main research question: how approximated model posteriors affect the performance of decision making via Thompson Sampling in contextual bandits.

Deep Bayesian Bandits - Exploring in Online Personalized Recommendations

한줄 “요약 Contextual Bandit + Bootstrapped Neural Network로 CTR 예측의 불확실성을 추정하여 광고 추천에서 exploration을 수행.

Multi-Armed Bandit

정의 Multi-armed Bandit 은 어떤 슬롯머신이 어떤 수익률을 가지는지 모를 때, 탐색 (Exploration) 과 활용 (Exploitation) 을 적절히 사용하여 최적의 수익을 찾아내고자 하는 Reinforcement Learning 알고리즘을 의미한다.