• Skip to primary navigation
  • Skip to content
  • Skip to footer
연구, 개발, 디버깅... 의미 있는 삽질 일지 연구, 개발, 디버깅... 의미 있는 삽질 일지 Actions make memories
  • Category
  • Tag
  • Search

    DimensionSTP

    Fun, creativity, and persistence

    • Seoul, Republic of Korea
    • Email
    • GitHub

    최근 포스트

    The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning Review

    2026-09-19 8 분 소요

    0. Introduction

    LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget Review

    2026-09-19 12 분 소요

    0. Introduction

    Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning Review

    2026-09-18 10 분 소요

    0. Introduction

    Weak-to-Strong Generalization via Direct On-Policy Distillation Review

    2026-09-18 8 분 소요

    0. Introduction

    SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning Review

    2026-09-17 8 분 소요

    0. Introduction

    • 이전
    • 1
    • 2
    • 3
    • 4
    • 5
    • 6
    • …
    • 57
    • 다음
    • 팔로우:
    • GitHub
    • 피드
    © 2026 DimensionSTP. Powered by Jekyll & Minimal Mistakes.