Blog

Personal explorations and reflections on language models. Browse by topicRSS feed

2026

Optimizing Pass@k as Reweighting Prompts

28 min read In progress

Pass@k upweights hard problems. Training on pass@k applies that weight directly, which raises two questions: what the weight should be, and where in the training loop to apply it.

Revisiting the Predictability of RLVR

33 min read , In progress

RELEX and AlphaRL find that RLVR weight updates are low rank and evolve near-linearly, and conclude that RLVR training is predictable. Their low-rank measurements reproduce. But a random walk produces the same measurements, rank-1 recovers the gain only when it is fitted to the endpoint it reconstructs, and on our runs extrapolation fails even at RELEX’s own horizon.

A Reflection on Multi-Agent Role-Playing

23 min read

Role-playing was the first multi-agent pattern — assign personas, let agents debate or collaborate. But it was largely a product of 2023-2024 model capabilities. As models improve, the real value of multi-agent systems turns out to be structural.

What's in Pass@K?

14 min read ,

Pass@k is ubiquitous in evaluating reasoning models, but the metric is more subtle than it appears. Computing it correctly requires the unbiased estimator, and the nonlinearity of pass@k means it effectively upweights hard problems compared to pass@1.

2025