VC-Attention: Value Smoothing and Softmax Casting for Low-bit Attention Review 2026-09-28 9 분 소요 0. Introduction
Calibrating Teacher–Student Discrepancy for On-Policy Distillation Review 2026-09-27 9 분 소요 0. Introduction
RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning Review 2026-09-26 10 분 소요 0. Introduction
Rethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening Review 2026-09-25 10 분 소요 0. Introduction
H3-World: Turning Language Understanding into World Control Review 2026-09-24 10 분 소요 0. Introduction