Zzong's Notes

Home

❯

papers

❯

rl

❯

Deep reinforcement learning for search, recommendation, and online advertising - a survey

Deep reinforcement learning for search, recommendation, and online advertising - a survey

2026년 6월 14일1 min read

Deep Reinforcement Learning for Search, Recommendation, and Online Advertising - a Survey

  • Paper link
    • https://arxiv.org/abs/1812.07127

함께 보면 좋은 글

DRN - A Deep Reinforcement Learning Framework for News Recommendation

Abstract 뉴스 추천을 위한 딥러닝 기반의 강화 학습 프레임워크를 제안한다. news feature 들과 user 의 preferences 의 변동성 (dynamic) 을 설명하는 것은 상당히 어렵다.

Recommendations with Negative Feedback via Pairwise Deep Reinforcement Learning

Recommendations with Negative Feedback via Pairwise Deep Reinforcement Learning Discussion Q-learning based offline 학습 방식이고, 모델도 무겁고.

Reinforcement Learning for Slate-based Recommender Systems - A Tractable Decomposition and Practical Methodology

Empirical Evaluation: Live Experiments YouTube 에 SARSA-TS 알고리즘을 실험 candidate -> ranker 를 거치게 되는데, ranker 의 scoring 함수에서 사용하는 myopic(근시안적) engagement 측정값을 LTV estimate 로 변경함...

Deep Exploration via Bootstrapped DQN

Abstract 해결하려는 문제: 강화학습에서의 효율적인 exploration Randomized value functions offer a promising approach to efficient exploration with generalization, but existing algorithms are not...

Exploration by Random Network Distillation

Exploration by Random Network Distillation RL methods work by maximizing the expected return of a policy.

Exploring compact reinforcement-learning representations with linear regression

paper Link: arxiv.org/pdf/1205.2606.pdf Exploring Compact Reinforcement-learning Representations with Linear Regression KWIK Linear Regression KWIK (Knows What It Knows) is a...

A Survey on Deep Learning Based POI Recommendations

A Survey on Deep Learning Based POI Recommendations A POI recommendation technique essentially exploits users’ historical check-ins and other multimodal information to...

Online learning to rank for information retrieval

Online Learning to Rank for Information Retrieval Related References slide: staff.fnwi.uva.nl/m.derijke/wp-content/uploads/sigir2016-tutorial.pdf .

Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Paper page - Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs RLVR의 역설 배경: LLM의 추론 능력을 향상시키기 위해 RLVR(검증 가능한 보상을 이용한 강화학습)...

Joint User-Entity Representation Learning for Event Recommendation in Social Network

Joint User-Entity Representation Learning for Event Recommendation in Social Network Paper link Abstract 사용자와 이벤트를 동일한 latent space 에 project 하기 위해 a joint representation...