Topics Agents 2 Distillation 1 Evaluation 1 MoE 3 Off-policy corrections 4 RL infrastructure 3 RLVR 5 Training dynamics 2Agents2 postsA Reflection on Multi-Agent Role-Playing Apr 20, 2026 · 23 min read · AgentsContext Management for LLM Agents: A Memory Hierarchy View Apr 18, 2026 · 24 min read · AgentsDistillation1 postImplementing On-Policy Distillation: Lessons from Building OPD in VeRL Jan 6, 2026 · 18 min read · Distillation, RL infrastructureEvaluation1 postWhat's in Pass@K? Jan 30, 2026 · 14 min read · RLVR, EvaluationMoE3 postsA Field Guide to Training–Inference Corrections Sep 5, 2026 · 17 min read · Off-policy corrections, MoEThe Infrastructure Cost of MoE Routing Replay Apr 26, 2026 · 16 min read · MoE, RL infrastructure, Off-policy correctionsOff-Policy Corrections in LLM RL Training Mar 1, 2026 · 30 min read · Off-policy corrections, MoEOff-policy corrections4 postsSignal or Noise? An SNR Criterion for Trusting Your Importance Ratio Sep 7, 2026 · 38 min read · Off-policy correctionsA Field Guide to Training–Inference Corrections Sep 5, 2026 · 17 min read · Off-policy corrections, MoEThe Infrastructure Cost of MoE Routing Replay Apr 26, 2026 · 16 min read · MoE, RL infrastructure, Off-policy correctionsOff-Policy Corrections in LLM RL Training Mar 1, 2026 · 30 min read · Off-policy corrections, MoERL infrastructure3 postsThe Infrastructure Cost of MoE Routing Replay Apr 26, 2026 · 16 min read · MoE, RL infrastructure, Off-policy correctionsTraining-Free Process Rewards for LLM RL Jan 10, 2026 · 12 min read · RLVR, RL infrastructureImplementing On-Policy Distillation: Lessons from Building OPD in VeRL Jan 6, 2026 · 18 min read · Distillation, RL infrastructureRLVR5 postsOptimizing Pass@k as Reweighting Prompts Sep 29, 2026 · 28 min read · RLVR In progressRevisiting the Predictability of RLVR Sep 27, 2026 · 33 min read · RLVR, Training dynamics In progressWhat's in Pass@K? Jan 30, 2026 · 14 min read · RLVR, EvaluationTraining-Free Process Rewards for LLM RL Jan 10, 2026 · 12 min read · RLVR, RL infrastructureUnderstanding Length Dynamics in RL Training Dec 21, 2025 · 35 min read · RLVR, Training dynamicsTraining dynamics2 postsRevisiting the Predictability of RLVR Sep 27, 2026 · 33 min read · RLVR, Training dynamics In progressUnderstanding Length Dynamics in RL Training Dec 21, 2025 · 35 min read · RLVR, Training dynamics
Implementing On-Policy Distillation: Lessons from Building OPD in VeRL Jan 6, 2026 · 18 min read · Distillation, RL infrastructure
A Field Guide to Training–Inference Corrections Sep 5, 2026 · 17 min read · Off-policy corrections, MoE
The Infrastructure Cost of MoE Routing Replay Apr 26, 2026 · 16 min read · MoE, RL infrastructure, Off-policy corrections
Signal or Noise? An SNR Criterion for Trusting Your Importance Ratio Sep 7, 2026 · 38 min read · Off-policy corrections
A Field Guide to Training–Inference Corrections Sep 5, 2026 · 17 min read · Off-policy corrections, MoE
The Infrastructure Cost of MoE Routing Replay Apr 26, 2026 · 16 min read · MoE, RL infrastructure, Off-policy corrections
The Infrastructure Cost of MoE Routing Replay Apr 26, 2026 · 16 min read · MoE, RL infrastructure, Off-policy corrections
Implementing On-Policy Distillation: Lessons from Building OPD in VeRL Jan 6, 2026 · 18 min read · Distillation, RL infrastructure
Revisiting the Predictability of RLVR Sep 27, 2026 · 33 min read · RLVR, Training dynamics In progress
Revisiting the Predictability of RLVR Sep 27, 2026 · 33 min read · RLVR, Training dynamics In progress