Zzong's Notes

RLVR

3건의 항목

  • 2026년 6월 28일

    Group Sequence Policy Optimization

    • LLM
    • reinforcement_learning
    • post_training
    • RLVR
    • policy_optimization
    • Qwen
  • 2026년 6월 14일

    One Token to Fool LLM-as-a-Judge

    • language_model
    • RLVR
    • nlp
    • paper_review
  • 2026년 6월 14일

    Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

    • LLM
    • RLVR
    • paper_review
    • reinforcement_learning