- Some Prototypical Examples
- standard Gaussian distribution 3 개에서 sampling 을 하는데, 가장 큰 값만을 이용해 분포를 만들면, mean 값이 0.85 가 나온다.
- 즉, expected disappointment 가 된다.

- Expected disappointment 는 distribution (reward model or action) 이 더 많아질수록 커진다.

1 min read


...ement Learning - Tutorial, Review, and Perspectives on Open Problems The Optimizer’s Curse - Skepticism and Postdecision Surprise in Decision Analysis
Introduction 추천 시스템에서 발생하는 loop 는 bias 를 일으킴 (1) collect data, (2) train model, (3) deploy model loop off-policy learning 에서 biased data 로 부터 unbiased learning 을 지향하는 방법에는...
paper link: arxiv.org/abs/1606.03966 Contextual Bandit pipeline 을 구성할 때 어떻게 하면 기술 부채가 적어지는지에 대한 내용을 다루는 것 같음.
논문의 퀄리티 측정법 뻔하지 않은 결과에는 합리적이고 디테일한 설명이 필요하다. novelty 가 부족하면 안된다. 기존 방식이 가진 이슈를 제기하고, 이를 해결할 수 있는 방안을 제시해야한다. References 2401.02412.pdf .
polynomial approximate sufficient statistics for scalable Bayesian GLM inference paper link Abstract GLM 학습을 위한 새로운 approach 를 제안.
각 random variable 에 따른 확률 분포 예시 Intelligence (I: i^0,i^1), Difficulty (D: d^0,d^1), Grade (G: g^1,g^2,g^3) 위 분포의 모든 값을 합치면 1 이 된다.
SIRIP: SIGIR Symposium on IR in Practice (Industry Track) SIRIP (Industry Track) 개요 행사 정보 48회 SIGIR 컨퍼런스 (2025년 7월 13일–18일, 이탈리아 파두아 개최) 내에서 진행된 Industry Track입니다.
Factor 는 함수 또는 테이블이다. Factor 는 a bunch of arguments 를 입력으로 받는데, 일반적으로 random variables 의 집합을 입력으로 받게된다. Scope: Factor 가 받을 수 있는 랜덤 변수의 범위를 scope 라 한다.
search query retrieval 를 주제로 하여 negative feedback signal 을 학습할 수 있는 알고리즘의 overview 조사 non-clicks documents 를 negative relevance feedback 으로 산정 (document, query) 쌍이 주어졌을 때,...
해당 논문은 대규모 언어 모델(LLM)을 인간의 선호도에 맞게 미세 조정하는 DPO(Direct Preference Optimization) 방식의 한계를 지적하고, 이를 개선한 MIPO(Modulated Intervention Preference Optimization) 라는 새로운 방법을 제안합니다.
Introduction Thompson sampling is an old heuristic that has a spirit of Bayesian inference and selects an arm based on posterior samples of the expectation of each arm.