Sitemap
A list of all the posts and pages found on the site. For you robots out there is an XML version available for digesting as well.
Pages
Page Not Found
Page not found. Your pixels are in another canvas.
Page not in menu
This is a page not in the main menu
Posts
Optimizing Pass@k as Reweighting Prompts
Published:
Pass@k upweights hard problems. Training on pass@k applies that weight directly, which raises two questions: what the weight should be, and where in the training loop to apply it.
Revisiting the Predictability of RLVR
Published:
RELEX and AlphaRL find that RLVR weight updates are low rank and evolve near-linearly, and conclude that RLVR training is predictable. Their low-rank measurements reproduce. But a random walk produces the same measurements, rank-1 recovers the gain only when it is fitted to the endpoint it reconstructs, and on our runs extrapolation fails even at RELEX’s own horizon.
Signal or Noise? An SNR Criterion for Trusting Your Importance Ratio
Published:
Your importance ratio mixes genuine policy movement with cross-engine evaluation noise. The noise is heavy-tailed, architecturally amplified, and aimed at exactly the tokens where the learning signal lives. Measure the split before you choose the treatment.
A Field Guide to Training–Inference Corrections
Published:
Your rollout engine and your trainer do not define the same policy — even with bit-identical weights. The field’s default response is a blanket importance-sampling correction. Eliminate the mismatch you can, then correct what remains.
The Infrastructure Cost of MoE Routing Replay
Published:
Routing replay (R3) stabilizes MoE RL training, but the routing data is 97% of the generation payload. This post traces the bottleneck — the single-threaded manager pipeline, not bandwidth — and the failed ‘obvious’ fix that revealed a fundamental constraint of mixing NCCL with inference.
A Reflection on Multi-Agent Role-Playing
Published:
Role-playing was the first multi-agent pattern — assign personas, let agents debate or collaborate. But it was largely a product of 2023-2024 model capabilities. As models improve, the real value of multi-agent systems turns out to be structural.
Context Management for LLM Agents: A Memory Hierarchy View
Published:
How LLM agents learn to manage their own context — from harness-driven compaction to memory tools and sub-agents — and why this may be the key bottleneck for long-horizon reasoning.
Off-Policy Corrections in LLM RL Training
Published:
A unified treatment of the five sources of distribution mismatch in LLM reinforcement learning and their corrections.
What’s in Pass@K?
Published:
Pass@k is ubiquitous in evaluating reasoning models, but the metric is more subtle than it appears. Computing it correctly requires the unbiased estimator, and the nonlinearity of pass@k means it effectively upweights hard problems compared to pass@1.
Training-Free Process Rewards for LLM RL
Published:
A training-free approach to step-level credit assignment: estimate V(prefix) via log-probability, compute marginal utility across episodes — plus the implementation pitfalls that silently destroy the signal.
Implementing On-Policy Distillation: Lessons from Building OPD in VeRL
Published:
On-policy distillation integrates teacher guidance into RL training, but the implementation is full of silent failures. This post documents the architecture, pitfalls, and design choices from building OPD in VeRL.
Understanding Length Dynamics in RL Training
Published:
An empirical investigation into what drives output length growth during RL training, revealing that dataset difficulty composition is the primary driver behind the ‘overthinking’ phenomenon.
portfolio
Portfolio item number 1
Short description of portfolio item number 1
Portfolio item number 2
Short description of portfolio item number 2 
