Reading the RL charts. The log-prob gate compares how likely the learner finds the generated
tokens with how likely the sampler found them (nats per token). They should sit close together, near
0.2–0.4; a gap means the update is computed on the wrong numbers. Evaluation scores are
failure-inclusive: a missing, blocked or invalid attempt counts as 0. Only aggregates are shown here.