Prose reinforcement learning
Four H200s · sequential rollout → craft / D33 / Luna → two Muon updates
Current collection
GPU snapshot
Device memory is not the same as live tensor allocation or allocator-reserved memory.
Training rewards and optimizer measurements
Rewards describe the sampled policy and the current training prompts. Prompt groups change between collections; these trends are not held-out evaluations. Reward charts use collection IDs; optimizer charts use actual optimizer-step IDs.
Individual craft heads
Raw head utilities over applicable scenes. D6 uses opening scenes; D19 is excluded. Missing values remain missing.
Held-out evaluation
Recorded events
Entropy and confirmed total billing are not emitted by this run. Held-out quality comes from a separate offline evaluation, reported below. Trainer–inference KL is estimated on rollout action tokens in the inference-to-trainer direction. K3 is the primary nonnegative curve and k1 is the signed companion; collections 1–10 are backfilled from retained same-policy caches. Neither is exact full-vocabulary KL. The loss reference term remains sampled k3 on behavior actions against the frozen merged SFT.