Zzong's Notes

Home

❯

RL

❯

EXP3

EXP3

2026년 6월 14일1 min read

References

  • https://jeremykun.com/2013/11/08/adversarial-bandits-and-the-exp3-algorithm/

함께 보면 좋은 글

epsilon-greedy algorithm

\varepsilon-Greedy 알고리즘은 \varepsilon 확률로 가능한 모든 action 들 중 하나를 동일한 확률로 임의 선택하는 것이다. 그 외에는 greedy 알고리즘과 동일하다.

Exploration and Exploitation trade-off

임의로 아무 콘텐츠나 시도해보며 콘텐츠의 반응률을 유추하는 과정을 탐색 (explore) 라고 부르며, 지금까지 반응률이 가장 높았던 콘텐츠를 노출시키는 것을 활용 (exploit) 이라 부른다.

Reinforcement Learning

Reinforcement Learning machine learning 기법 중 하나.

Multi-Armed Bandit

정의 Multi-armed Bandit 은 어떤 슬롯머신이 어떤 수익률을 가지는지 모를 때, 탐색 (Exploration) 과 활용 (Exploitation) 을 적절히 사용하여 최적의 수익을 찾아내고자 하는 Reinforcement Learning 알고리즘을 의미한다.

UCB

UCB란 UCB(Upper Confidence Bound) 알고리즘은 각 팔(arm)의 현재까지의 평균 보상(mean reward)을 추적하면서, 동시에 각 arm 에 대한 상위 신뢰 구간(upper confidence bound, UCB)을 계산합니다.

Exploring compact reinforcement-learning representations with linear regression

paper Link: arxiv.org/pdf/1205.2606.pdf Exploring Compact Reinforcement-learning Representations with Linear Regression KWIK Linear Regression KWIK (Knows What It Knows) is a...

exploration

Hard-exploration 문제 hard-exploration 문제란, 보상이 매우 드문 특정 환경에서의 exploration 을 의미한다. 임의의 exploration 의 경우, 성공적인 state 나 의미있는 feedback 을 발견하기가 매우 어렵다.

method of moments

Method of Moments A.1) Recoteam 픽코마 적용 사례 issue: 2019 하계 인턴, 픽코마/선물하기 연관 추천 개선 신규 arm 이 지나치게 높은 entry expected reward 값을 가지고 있기 때문에 신규 arm 이 아니면 거의 explore 되지 못 하는 문제가 있었음 기존...

S-MDP

S-MDP agent 가 조합적 선택 (combinatorial selections) 을 연속적으로 수행해야 하는 문제 B) Related C) References.

policy

Policy policy 란 주어진 states 에서 actions 에 대한 확률 분포를 의미하며 \pi 로 표현한다.