Prose reinforcement learning

Four H200s · sequential rollout → craft / D33 / Luna → two Muon updates

Waiting for a verified cycle. No measured feed loaded.

Current collection

Coverage and phase

GPU snapshot

Device memory is not the same as live tensor allocation or allocator-reserved memory.

Training rewards and optimizer measurements

Rewards describe the sampled policy and the current training prompts. Prompt groups change between collections; these trends are not held-out evaluations. Reward charts use collection IDs; optimizer charts use actual optimizer-step IDs.

Individual craft heads

Raw head utilities over applicable scenes. D6 uses opening scenes; D19 is excluded. Missing values remain missing.

Held-out evaluation

Recorded events

Entropy and confirmed total billing are not emitted by this run. Held-out quality comes from a separate offline evaluation, reported below. Trainer–inference KL is estimated on rollout action tokens in the inference-to-trainer direction. K3 is the primary nonnegative curve and k1 is the signed companion; collections 1–10 are backfilled from retained same-policy caches. Neither is exact full-vocabulary KL. The loss reference term remains sampled k3 on behavior actions against the frozen merged SFT.