5B K-5 LittleLearner, MathCAMPS comparison, failure analysis & data
Generated 2026-07-10 08:27 UTC.
Eval filters (same as the paper's 2B numbers). Three MathCAMPS filters on all rows: (1) drop standards with <30 unique gold answers; (2) drop questions whose gold appears verbatim in the
question; (3) drop standard 6.EE.B.7. Result: 3,628 of 4,900 questions kept (K and G1 removed by filter 1).
Code in mathcamps_filters.py.
±σ in the aggregate columns = expected run-to-run movement of that cell under a fresh 64-sample
re-eval on the fixed problem set, computed analytically from the stored per-problem correct counts
(analysis/eval_noise_sigma.py) and validated by duplicate eval runs (0.6B and 1.3B: observed
re-run |Δ| 0.15-0.20 on K-5* pass@64 vs predicted σ≈0.3). pass@1 cells are first-sample estimates, hence
the larger σ (≈0.6). Reference rows (2B nanochat, Gemma) are point values from the paper, no σ available.
Comparing two rows carries additional problem-set uncertainty beyond these bars (paired 95% CI ≈ ±1.4 on
K-5* pass@64, unpaired ±2-3): see the 0.6B vs 1.3B tab.
Eval-condition audit (2026-07-04). The 5B and midtrain rows (evaluated 06-09 to 06-11) predate
the --no_system flag and used the format system prompt; all 1.3B and 0.6B rows use
--no_system. A same-checkpoint A/B on the published 5B (sft_star1) puts the condition effect
at K-5* pass@1 −1.4 / pass@64 −0.4 (no_system minus with-system, about 1.5σ), so rows are comparable
across the condition. A same-condition fresh-seed re-run reproduced the 5B at K-5* pass@64 80.51 vs 80.43
(|Δ| 0.07), completing the σ validation at all three scales (re-run |Δ|: 0.6B 0.20, 1.3B 0.15, 5B 0.07).
.
Names linked with 🤗 are published on the Hugging Face Hub. Each row reads <size>-<corpus> base + <corpus> SFT [+ steps] (corpus = bounded/K-5 or unbounded).
GRPO rows. GRPO buys first-sample pass@1 at a small pass@64 (coverage) cost, the
RLVR-elicitation trade. Recipe, frontier, and the data-source ablation are in the GRPO v3 tab.
The cooloff-SFT-mixin finding. The two 1.36B bounded rows share the same base and SFT and differ only
in the pretraining cooloff: one folded a 5% chat-SFT blend into the WSD decay, the other did not. The base
without the blend is not chat-SFT-able (free generation reverts to web-text); the blended base SFTs into a
coherent model. Base loss and BPB are nearly the same; the difference is in finetunability.
Literature match. Consistent with the data-curation result that the decay phase is where
instruction data helps downstream (OLMo, MiniCPM, Llama-3): fold instruction data into the anneal rather
than adding a re-warmed continued-pretraining stage.
Validated at both scales. A 5B base without the blend, SFT'd with the same recipe, gives K-5*
p@64 12.0 / p@1 0.2, vs 80.1 / 45.6 for the blended 5B. Only the 5% blend differs. Base quality is unchanged
(continuation K-5 41.8 vs about 40). The effect is on SFT-ability, at 5B as well as 1.36B.
5B pretrain→SFT (NO midtrain) ⭐, performance breakup
All numbers on this tab use the paper's filtered eval (see Comparison tab note). K-5* pass@64 = 80.1%, pass@1 = 45.6%; Beyond-K-5 pass@64 = 34.4%. 2B comparison = nanochat-d34 chat-SFT, chat-format eval (same method as our chat models), pass@64. The knowledge boundary shows as the G5→G7 cliff in p@1 (fractions/decimals + multi-step), while pass@64 stays high through G7, the model can often reach beyond-K-5 answers but not reliably.
Per-grade breakdown
grade
n (filtered)
5B pretrain→SFT ⭐
midtrain→SFT
midtrain v1
2B nanochat SFT (chat)
5B base 42k (cont)
ctrl ✓ @64 & 2B ✗ @64
p@1
p@64
p@64
p@64
p@64
p@64
G2
595
67.9
98.2
66.9
67.9
4.5
27.7
557/595
G3
664
65.8
94.4
65.2
67.2
10.4
36.3
559/664
G4
789
29.0
75.0
34.7
35.4
2.3
17.2
577/789
G5
683
19.9
52.7
23.0
23.9
4.4
11.6
330/683
G6
200
21.0
49.5
15.0
18.0
0.0
6.0
0/200
G7
497
8.5
44.1
8.2
10.1
0.0
6.6
0/497
G8
200
0.5
9.5
3.5
4.0
0.0
2.0
0/200
When it works, per-grade examples (problems the 2B never solves in 64 tries)
G2, control 98% vs 2B-SFT 5% pass@64; control solves 557 problems the 2B never cracks in 64 tries
2.MD.B.5: Daniel is constructing a fence for his yard. He has already built 36 feet of the fence, but he then realizes he needs another 7 feet to complete it. How many feet of fencing does Daniel need in total to complete his backyard fence? gold:43
5B pretrain→SFT, correct first sample
We need to find the total length of fencing Daniel needs.
First, Daniel has already built 36 feet of the fence.
Next, he needs to add 7 more feet to finish the fence.
So, we add the lengths together: 36 feet + 7 feet = 43 feet.
Answer: 43 feet
2B nanochat SFT, 0/64 correct, first sample
ax anelet pic wonderfulund you patience? I growingple as surv as a guard on aeleierorn!ush I trick to downple theots toog you his the process and verb of fence. morning salt isolve iss, and I marng back theucks a An cityThis orames I wild, a AnClassroom! I growing shows to against you with money from theSince to K Y Favor vol. · you meions to Just a hadiber cheese for you and Bward?
5B midtrain→SFT (0/64 correct), first sample
We need to find the total length of fencing Daniel needs for his garden.
First, we add the lengths of all three sides together.
So, we add 8 feet + 14 feet + 7 feet.
That gives us 18 feet + 14 feet = 32 feet.
Answer: 32 feet
2.MD.C.8: Derek found $69 in his drawer. He decided to buy a board game which cost $28. After paying for the game, he received his change and his friend gave him an additional $3. How much money does Derek have now? gold:44
5B pretrain→SFT, correct first sample
We need to find out how much money Derek has now.
First, Derek started with $69.
Next, he bought a game for $28, so he had $69 - $28 = $41 left.
Then, his friend gave him $3 more.
So, we add the amount he had left to the extra money he received: $41 + $3 = $44.
Answer: $44
2B nanochat SFT, 0/64 correct, first sample
Theexample you named is in theragRead. It Bakery the numer, an easy Sen, and anheated, exactly.
5B midtrain→SFT (0/64 correct), first sample
We need to find out how much money Derek has now.
First, Derek started with $69.
Next, he got $35 more in total, so he had $69 + $35 = $94 in the drawer.
Then, he spent $28 on a game.
So, we subtract the money he spent from his starting money: $94 - $28 = $66.
Answer: $66
G3, control 94% vs 2B-SFT 10% pass@64; control solves 559 problems the 2B never cracks in 64 tries
3.MD.D.8-polygon: A regular septagon (7-sided polygon) has a total perimeter of 161m. What is the length of each side? gold:23
5B pretrain→SFT, correct first sample
We need to find the length of each side of the septagon.
First, a regular shape has all its sides the same length.
Next, the total perimeter is the total length around the shape, which is 161 meters.
To find the length of one side, we divide the total perimeter by the number of sides.
So, we calculate 161 ÷ 7.
Therefore, 161 ÷ 7 = 23.
Answer: 23m
2B nanochat SFT, 0/64 correct, first sample
gaste! morning educate( holdtar We. prepare, hold tal is as him as theWhy'sners( hold, teachers N dry,yg, and hold tal! famgo that we fre Y of the entners.
5B midtrain→SFT (0/64 correct), first sample
We need to find the length of each side of the regular octagon.
First, a regular octagon has 8 equal sides. We can think of it as having 8 equal sides arms and 8 equal legs.
Next, the total perimeter is 8 × 8 = 64 cm. This perimeter is the perimeter of the octagon.
Then, when all 8 sides are made of the square, we subtract the "the exterior side" from the total perimeter to find the shaded side: 64 cm - 8 cm = 56 cm.
Now, since there are 8 equal sides, we divide the shaded side by 8 to find the length of one side: 56 cm ÷ 8 = 7 cm.
Answer: 7 cm
3.MD.D.8-quadrilateral: A square has a perimeter of 344m. What is the length of each side of the square? gold:86
5B pretrain→SFT, correct first sample
We need to find the length of one side of the square.
First, a square has 4 sides that are all the same length.
Next, the perimeter is the total length around all 4 sides combined.
So, to find the length of one side, we divide the total perimeter by 4.
Therefore, 344 m ÷ 4 = 86 m.
Answer: 86m
2B nanochat SFT, 0/64 correct, first sample
nationally relaxed and the answer ofimal! I growing so surv to public you all. morning salt is square relaxed, and I introduc the oil and advice answer ofimal! I mar localng a What of a numer, shock Mac and related extension Iue.atoesbs you me to down aboutions or theheated?
5B midtrain→SFT (0/64 correct), first sample
We need to find the length of one side of the square.
First, we know the perimeter of the square is 344 meters. The formula for the perimeter of a square is 1 (side length) × side.
Next, we put the length of one side into the perimeter formula: 344 × 1 = 344.
Then, we divide that number by 2: 344 ÷ 2 = 172.
So, the length of each side of the square is 172 meters.
Answer: 172m
G4, control 75% vs 2B-SFT 2% pass@64; control solves 577 problems the 2B never cracks in 64 tries
4.MD.A.2-decimal: Maria saved $32.5 from her part-time job and $34 from her previous savings. She spent $7 on a gift for her sister and $7 for her brother at the end of the week. How much money does Maria have left? gold:52.5000000000000
5B pretrain→SFT, correct first sample
We need to find out how much money Maria has left.
First, let's find the total money Maria saved. She saved $32.5 from her part-time job and $34 from her previous savings. So, $32.5 + $34 = $66.5.
Next, let's add all the money Maria spent. She spent $7 on a gift for her sister and $7 on her brother. So, $7 + $7 = $14.
Then, we subtract the total money she spent from the total money she saved. $66.5 - $14 = $52.5.
Answer: $52.5
2B nanochat SFT, 0/64 correct, first sample
These spotted this fast, weoy to jewel the dream of focus, bo seats theolun ofional the numerical answeriz.
Com'sopy good the friend Air perM originalot:olve iss
massage: answer of iss
anti ofional the numerical answeriz: 1
while canible up theLim as Atlantic perM originalot'sionalca answer'sional stationsCT ofional
1ca 1
Com'srers and spotted for focus per smca 1
These breeze k focus, we Lancaster sm fromized girls of theLim per sm - smca 1 - sm
- smca - sm
restored, we indulgeized girls of theLim by -1 to spotted for focus perxca sm
The dream of focus is sm.
#### sm
The cross is: …[trimmed]
5B midtrain→SFT (0/64 correct), first sample
We need to find out how much money Maria has left.
First, Maria saved $32.50.
Next, she spent $7.50 on a gift.
Then, we subtract the cost of the gift from her savings.
So, $32.50 - $7.50 = $28.00.
Answer: $28.00
4.MD.A.2-fraction: Jane cycles at a rate of 11 km/h for 45 hours. In addition to this distance, Jane also decides to cycle for another distance that equals exactly 1 km. How many kilometers in total does Jane cycle? gold:496
5B pretrain→SFT, correct first sample
We need to find the total distance Jane cycles.
First, Jane cycles at a rate of 11 km/h for 45 hours.
To find the distance she travels, we multiply her rate by the time: 11 km/h × 45 hours = 495 km.
Next, she decides to cycle another distance that is equal to 1 km.
So, we add this new distance to the distance she already cycled: 495 km + 1 km = 496 km.
Answer: 496 km
2B nanochat SFT, 0/64 correct, first sample
C growingFl of 11 km/h for 45 hours. In addition to this distance, Jane also
5B midtrain→SFT (0/64 correct), first sample
We need to find the total distance Jane cyclelic removes.
First, the problem tells us that at 09:00, Jane had cycled 9 kilometers.
Next, she cycled another 3 kilometers.
To find the total distance, we add these two distances together: 9 km + 3 km = 12 km.
Answer: 12 km
G5, control 53% vs 2B-SFT 4% pass@64; control solves 330 problems the 2B never cracks in 64 tries
5.NBT.B.5: A camera shop makes four times more sales in the holiday season. If they usually sell 302 cameras in non-holiday seasons, how many cameras do they sell in the holiday season? gold:1208
5B pretrain→SFT, correct first sample
We need to find how many cameras the shop sells in the holiday season.
First, we know that sales in the holiday season are 4 times as many as sales in the non-holiday season.
Next, we know the shop sells 302 cameras in the non-holiday season.
So, to find the number of cameras sold in the holiday season, we multiply the non-holiday season sales by 4.
Therefore, 302 × 4 = 1208.
Answer: 1208 cameras
2B nanochat SFT, 0/64 correct, first sample
A camera shop makes! - baby bring a makesral for B open, it's.last to combine the things und of She,lp, and Brook. incre inspire is smoothies forYields,atherob can beulting. road's aasterison on pe to serv for B open per
Gardensera: Gardensoricalackers demonstr fishing, horticultural She. dust inspire can beatherovery for them, soug B lot is in a letter withoneyroom,az kissing it with aWestern.d. small it's a particip Marcus lot, it can released in Mat with roughviron She.
Gardensthis: Gardens comfortableupperware is a clot conven.uz the &ration of center toresh out surelication another. co …[trimmed]
5B midtrain→SFT (0/64 correct), first sample
We need to find how many cameras the shop sells in the non-holiday season.
First, we know the shop sells 302 cameras in the non-holiday season.
Next, the problem tells us that if Ocho made 10 reusable cameras, the shop sold 9 times less than or equal to 10 complete sets. This means each of the 9 complete sets has 30 cameras.
So, the total number of camera sets sold is 9 × 302 = 2766.
Then, each of these 2766 cameras has 34 cameras.
Therefore, to find how many cameras are sold, we divide the total camera sets by the number of cameras in one set: 2766 ÷ 34 = 94.
Answer: 94
5.NBT.B.6: A kite flies at a height of 1800 feet using a 50-foot string. How high is the kite flying per each foot of the used string? gold:36
5B pretrain→SFT, correct first sample
We want to find how high the kite flies for each foot of string.
First, we know the total string length is 50 feet and the kite flies at a height of 1800 feet.
To find the height for each foot of string, we divide the total height by the total length of string.
So, we calculate 1800 ÷ 50.
That gives us 36.
Therefore, the kite flies 36 feet for each foot of string.
Answer: 36 feet
2B nanochat SFT, 0/64 correct, first sample
A kite flies at a height of 1800 can using a 50-foot string.
5B midtrain→SFT (0/64 correct), first sample
We want to find the height of the kite flying at the end of the string.
First, the string is 15 feet long in the middle. This is the diameter the time to go straight across the middle of the kite.
Next, since the string is 1200 feet long, the radius (distance across the middle) is 1200 ÷ 2 = 600 feet.
Then, we know the height (above the scaling wrist) of a 50-foot kite is 3/4 of the radius.
So, to find the full height, we multiply 600 feet by 3/4.
That gives us 600 × (3/4) = (600 ÷ 4) × 3 = 150 × 3 = 450 feet for a 50-foot length.
Answer: 450 feet
Failure modes
mode (first-sample failures, n=2337 of 3628)
count
share
wrong_reasoning_grounded
754
32.3%
arithmetic_slip(<=25%)
668
28.6%
ungrounded_drift
383
16.4%
number_miscopy
375
16.0%
grader_false_negative
91
3.9%
spurious_extra_step
53
2.3%
truncated
10
0.4%
degenerate_repetition
3
0.1%
52% of failures are solvable-but-unreliable (pass@64 ✓, pass@1 ✗), the remaining lever is reliability, which is what GRPO targets. Hard format/degeneration failures are ~0.5%.
number_miscopy · G5 5.OA.A.1, A farmer has 3069 grapes. If he decides to put 33 grapes in each bag, he will have a certain number of bags. However, the farmer eats 13 bags of grapes and also gives away 50 bags of grapes. How many bags of grapes does the farmer have left? gold:30pred:290 bags(correct in 1/64 samples)
completion
We need to find how many bags of grapes are left after several actions.
First, let's find the total number of bags the farmer would have if he put 33 grapes in each. We divide 3069 by 33, which gives 363 bags.
Next, the farmer eats 13 bags of grapes. So, we subtract those: 363 - 13 = 340 bags.
Then, he gives away 50 bags. So, we subtract those: 340 - 50 = 290 bags.
Answer: 290 bags
arithmetic_slip(<=25%) · G5 5.OA.A.1, A dolphin can leap (3 + 59) feet high. If a dolphin completes 89 jumps in a day, how many feet in total has the dolphin jumped in that day? gold:5518pred:5428(correct in 0/64 samples)
completion
We need to find the total distance a dolphin jumps in one day.
First, the problem tells us that jumping 3 + 59 equals 62.
Next, the dolphin completes 89 jumps in a day.
To find the total distance, we multiply the total distance of one jump by the number of jumps: 62 × 89.
So, 62 × 89 = 5428.
Answer: 5428
ungrounded_drift · G5 5.OA.A.1, A blanket factory makes 13 blankets in the morning and 81 in the afternoon, but returns 59 due to defects. Each blanket is sold at a price equal to 39 times the total number of good blankets they have, plus the number of good blankets. What's the total value of the blankets they've sold? gold:1400pred:$10048(correct in 0/64 samples)
completion
We want to find the total value of the blankets the factory has sold.
First, let's find out how many blankets they made in total. they made 13 in the morning and 81 in the afternoon, so 13 + 81 = 94 blankets.
Next, we know this number by multiplying it by 39 (how many bad blankets there are for every good one). So, 94 × 39 = 3556 total blankets.
Then, we find the price of one blanket. They sell it at 39 times the price of a bad blanket, plus 39 other times the price of a bad blanket. This means the price for one blanket is (39 × price of bad blanket) + 39. So, the price for one blanket is (94 × (39 - 39)) + 39.
Now, we find the price of one bad blanket. They sell it for 81 × (39 - 39) = (94 × 39) - (39 × 39) = 15139 - 14946 = 8206 bad blankets only.
So, to find the total money from selling just one bad blanket, we multiply its price by 8206. That gives us (8206 × price of bad blanket) + …[trimmed]
spurious_extra_step · G5 5.OA.A.1, A company manufactures 1462 binoculars and distributes them equally among 43 stores. After subtracting 25 defective pieces from each store's stock, how many binoculars does each store end up with? gold:9pred:0.21(correct in 9/64 samples)
completion
We need to find how many binoculars each store has after some are removed.
First, let's find the total number of binoculars the company has. They distribute 1462 binoculars equally among 43 stores. So, we divide 1462 by 43.
1462 ÷ 43 = 34. This means the company has 34 binoculars for each store.
Next, the company sells these 34 binoculars and gets 25 defective pieces back. So, each store actually has 34 - 25 = 9 binoculars.
Then, there are 43 stores in total. To find how many binoculars each store has, we divide the total binoculars by the number of stores.
9 ÷ 43 = 0.21.
Answer: 0.21
wrong_reasoning_grounded · G5 5.OA.A.1, A sock company packages their socks in sets of 40. They have 2840 socks to package, and they currently have 93 packs already made. How many packs of socks will they have in total once all the socks have been packaged? Afterward, they received an another shipment of 70 packs of socks, but had to return 8 due to quality issues. How many packs of socks does the company have now? gold:226pred:3798 packs and 93000 socks(correct in 0/64 samples)
completion
We need to find the total number of sock packs and then the total number of socks.
First, the company currently has 93 packs. To make a full set, they package 40 pairs, so they need 93 × 40 = 3720 sets of socks.
Next, they received an extra shipment of 70 packs.
Then, they returned 8 packs.
So, the total number of packs is 3720 + 70 + 8 = 3798.
And the total number of socks is 3720 ≈ 93 × 40.
Answer: 3798 packs and 93,000 socks.
grader_false_negative · G5 5.NBT.B.6, A hat maker makes 1083 hats throughout 57 weeks. If the hat maker makes the same amount of hats every week, how many hats does the hat maker produce weekly? gold:19pred:1083/57(correct in 5/64 samples)
completion
We need to find out how many hats the hat maker makes each week.
First, we know the hat maker makes a total of 1083 hats over 57 weeks.
To find out how many hats are made each week, we divide the total number of hats by the total number of weeks.
So, we calculate 1083 ÷ 57.
Answer: 1083/57
Why pretrain→SFT beats the midtrain+SFT variants
Head-to-head on the 3,628 filtered questions: pretrain→SFT solves 1,207 problems @64 that midtrain→SFT never solves; the reverse is only 47. At pass@1: 1,103 vs 32. Overall first-sample accuracy 35.6% vs 6.1% (midtrain alone: 5.2%).
Failure-mode mix of each model's first-sample failures
mode (share of that model's failures)
pretrain→SFT ⭐
midtrain→SFT
midtrain (no SFT)
number_miscopy
16.0%
46.8%
46.9%
ungrounded_drift
16.4%
10.8%
10.2%
arithmetic_slip(<=25%)
28.6%
9.3%
8.8%
wrong_reasoning_grounded
32.3%
30.4%
30.9%
spurious_extra_step
2.3%
0.9%
0.9%
grader_false_negative
3.9%
1.6%
1.7%
Why midtrain hurts, the mechanism in the data.Number mis-copy is 47% of midtrain→SFT failures vs 16% for pretrain→SFT: the midtrained base systematically misreads the given quantities (invents numbers not in the problem), and SFT cannot repair it, midtrain→SFT's failure mix is nearly identical to raw midtrain's (46.8% vs 46.9% mis-copy), i.e. SFT only reformats the output; the grounding damage lives in the base weights. Cause: the midtrain continued-pretraining restarted with a fresh optimizer + LR re-warm, train loss spiked 1.7→8 before recovering to 1.78, which perturbed the WSD-annealed weights into a worse SFT substrate (confirms Fanfei's 2B Exp B/C). The pretrain→SFT control keeps the annealed base intact: its residual failures shift to arithmetic slips (29%) and genuinely-hard reasoning (32%), capability limits, not grounding damage.
Same question, both models, control first sample vs midtrain→SFT (0/64 correct)
G2 2.MD.C.8: Linda has 29 quarters in her piggy bank. She uses 12 quarters to buy a small toy. Later that day, her mother gives her an additional 9 quarters. How many quarters does Linda have now? gold:26
5B pretrain→SFT, correct first sample
First, Linda starts with 29 quarters.
Next, she uses 12 quarters to buy a toy. So, she has 29 - 12 = 17 quarters left.
Then, her mother gives her 9 more quarters.
Now, she has 17 + 9 = 26 quarters in total.
Answer: 26
We need to find out how many quarters Linda has now.
First, we know Linda had 29 quarters.
Next, she used 12 quarters.
To find out how many more, we subtract the quarters she has now from the total quarters: 29 - 12.
So, 29 - 12 = 17.
Answer: 17
G3 3.OA.A.3: On the farm, there are 7 groups of chicks. Each group consists of just 1 chick. The farmer then brought 9 times that amount of chicks. How many chicks did the farmer bring? gold:63
5B pretrain→SFT, correct first sample
We need to find the total number of chicks the farmer brought.
First, there are 7 groups of chicks.
Next, each group has 1 chick. So, 7 groups have 7 × 1 = 7 chicks.
Then, the farmer brought 9 times that amount.
So, 9 × 7 = 63 chicks.
Answer: 63
We need to find the total number of chicks the farmer brought.
First, the farmer brought 9 chicks.
Next, the farmer brought 4 more chicks.
To find the total, we add the number of chicks and chicks: 9 + 4 = 13.
Answer: 13
G4 4.MD.A.2-decimal: Bob travels 6.1 miles to the grocery store. His total trip (going and coming back) is the summation of his distance to the grocery and back home. If he makes this trip 3 times a week, how many miles does he travel in total? gold:36.6000000000000
5B pretrain→SFT, correct first sample
We need to find the total distance Bob travels.
First, Bob travels 6.1 miles to the grocery store and 6.1 miles back home. These are the two distances that are part of his trip.
Next, we want to find how many times he makes this round trip. We can find a fraction that multiplies to equal 6.1, which is 6.1 × 2.
That gives 12.2 total miles.
Then, he makes this trip 3 times a week. So we multiply 12.2 by 3.
12.2 × 3 = 36.6.
Therefore, Bob travels a total of 36.6 miles.
Answer: 36.6 miles
5B midtrain→SFT, ungrounded_drift, pred=5750 miles
We need to find the total distance Bob travels.
First, Bob is getting closer to the front of the line, which is its perimeter.
Next, he is also getting closer to the back of the line.
Then, he is passing through the pile of about 200 into the *smallest* part of the distance.
The back of the line is even thicker than the front.
So, we add the two distances: 2 × 19.75 × 200.
That gives us 5750.
Answer: 5750 miles
G5 5.OA.A.1: A gardener has a field where he planted 80 rose bushes, with each bush producing 30 roses. Before he is able to gather them, a storm blows away 43 roses. How many roses are remaining in the field? gold:2357
5B pretrain→SFT, correct first sample
We need to find how many roses are left in the field.
First, the farmer planted 80 rose bushes, and each bush grew 30 roses.
To find the total number of roses, we multiply 80 by 30.
80 × 30 = 2400 roses.
Next, a storm blows away 43 of these roses.
So, we subtract the blown away roses from the total: 2400 - 43.
2400 - 43 = 2357.
Therefore, 2357 roses are remaining.
Answer: 2357
First, the gardener has 80 rose bushes in their garden.
Next, 43 poppies are planted outside.
To find how many are left, we subtract the number eaten roses from the number taken away.
So, 80 - 43 = 37.
Answer: 37
G6 6.NS.B.3: A mouse starts running from a point, he runs for 10.4 meters towards east, then turns around and runs for another 37.12 meters towards west. What is the total distance the mouse has run? gold:47.5200000000000
5B pretrain→SFT, correct first sample
We need to find the total distance the mouse ran.
First, the mouse ran 10.4 meters towards the east.
Next, he ran another 37.12 meters towards the west.
To find the total distance he ran, we add these two distances together.
So, 10.4 + 37.12 = 47.52 meters.
Answer: 47.52 meters
We want to find the total distance the mouse ran.
First, the mouse ran for 10.4 seconds. We know that 10.4 seconds is the same as 10.4 times 0.4 seconds, which is 40.4 seconds.
Next, the mouse ran for 37.12 seconds. We need to change these seconds into seconds.
To find the total distance, we multiply the speed by the time: 40.4 seconds/second × 37.12 seconds.
So, the mouse ran 10.4 × 37.12 = 473.88 meters.
Answer: 473.88 meters
G7 7.NS.A.1-fraction: What is the sum of 61/20 and 29/15? gold:299/60
5B pretrain→SFT, correct first sample
We need to find the sum of the fractions 61/20 and 29/15.
First, we find a common denominator for 20 and 15. The smallest common number they both divide into is 60.
Next, we change 61/20 to an equivalent fraction with a denominator of 60. We multiply the top and bottom by 3 to get 183/60.
Then, we change 29/15 to an equivalent fraction with a denominator of 60. We multiply the top and bottom by 4 to get 116/60.
Now, we add the new fractions: 183/60 + 116/60. We add the numerators: 183 + 116 = 299. The denominator stays 60.
So, the sum is 299/60.
Answer: 299/60
5B midtrain→SFT, ungrounded_drift, pred=61/60 and 0.8833
We need to add the fractions 61/20 and 29/15.
First, we find a common denominator for 20 and 15. The number 15 works because 20 and 15 can both be divided evenly by 5.
Next, we change 61/20 to have a denominator of 15. We multiply the top and bottom by 3: (61 × 3) / (20 × 3) = 183/60.
Then, we change 29/15 to have a denominator of 60. We multiply the top and bottom by 2: (29 × 2) / (15 × 2) = 58/60.
Now we can add the fractions: 183/60 + 58/60.
We add the numerators: 3 + 58 = 63. The denominator stays 60, so we have 63/60.
To get the decimal form, we divide 63 by 60, which is about 0.8833.
Answer: 61/60 and 0.8833
G8 8.EE.C.7: Solve the linear equation for the value of 'u':
62 = (u - 26) gold:88
5B pretrain→SFT, correct first sample
We need to find the value of 'u' in the equation 62 = (u - 26).
First, we find what 'u' minus 26 equals. The problem tells us that when u minus 26, the result is 62.
So, we can write this as : 62 = (u - 26).
Next, to get 'u' by itself, we need to add 26 to both sides of the equation.
Therefore, the equation becomes : 62 + 26 = u.
Now, we solve for 'u'. We add 62 and 26, which gives us 88.
So, u = 88.
Answer: 88
5B midtrain→SFT, ungrounded_drift, pred=The step is to determine whether to get a negative solution
We need to solve the math problem (u): 62 - 26.
First, we look inside the parentheses. We have 26 and we subtract 26.
Since 26 is a bigger number than 26, when we subtract 26 from it, the result will be a negative number.
Therefore, we cannot put a negative number on top of a 26.
Answer: The step is to determine whether to get a negative solution.
The pass@1 → pass@64 closure experiment
The deep-dive showed ~52% of the base's failures are solvable-but-unreliable (pass@64 ✓, pass@1 ✗). This series asks the research question directly: can post-training close that gap on a 5B model pretrained only on K-5 data? All numbers paper-filtered; K-5*=mean(G2-5), Beyond=mean(G6-8).
Headline, filtered K-5* / Beyond (Δ vs base in p@1)
+ GRPO v2 (RL, difficulty-banded Gemini pool, 8 prompts/step)
45.8(+0.1)
80.3
10.1
33.0
+ teacher-distill (Gemini correct CoT, G4/G5)
44.9(-0.7)
80.5
10.0
33.3
Verdict: the gap barely moves, and the p@64 ceiling is immovable. Best is STaR-RFT r1 at K-5* p@1 46.4 (+0.7); every later/heavier method lands within ±1 of the base. No method raised K-5* pass@64 above ~80: the capability ceiling is set by pretraining (scale + the K-5 data boundary), and post-training only redistributes reliability within it.
Where the reliability moves, per-grade pass@1 trajectory
+ GRPO v2 (RL, difficulty-banded Gemini pool, 8 prompts/step)
72.9
65.8
26.1
18.2
21.0
9.3
0.0
+ teacher-distill (Gemini correct CoT, G4/G5)
70.4
63.0
27.8
18.6
19.5
10.1
0.5
pass@64 ceiling (base)
98
94
75
53
50
44
10
The mechanism, grade by grade.(1) STaR self-distillation lifts G2 monotonically (67.9→73.6→76.8): the grade where the base is already strong, so correct self-samples are abundant, but degrades G4/G5/G6 (round 2: G4 29→25.9, G6 21→15): self-training can only reinforce what the model already solves, and distribution-shifts away from grades it rarely gets right. (2) GRPO v2, despite a dense banded pool (43-45% solve, ~8 prompts/step vs v1's 1), held its training reward and on-distribution held-out eval flat at ~0.45: RL sharpened the policy without lifting held-out competence: same plateau. (3) Teacher-distillation: feeding the model correct Gemini CoT for exactly the hard G4/G5 standards (3,733 verified traces, SFT loss 0.36 vs self-distill's 0.07, i.e. genuinely new content), still did not lift G4/G5 p@1 above the base (G4 27.8 / G5 18.6 vs base 29.0 / 19.9). The model can imitate the steps but cannot reliably execute the underlying multi-digit/fraction/decimal arithmetic at G4/G5, and the demonstrations dilute its G2/G3 reliability.
Thesis. On this K-5-only 5B, the p@1↔p@64 gap at G4/G5 is a capability/grounding ceiling, not a reliability problem. Post-training (RFT, RL, even teacher distillation) buys p@1 reliability only on grades where the base is already competent (G2), and cannot push the p@64 capability frontier, that is fixed by pretraining scale and the deliberate K-5 knowledge boundary (arithmetic precision on multi-digit/fraction problems was never in the training distribution). To raise G4/G5 you need a stronger/larger base or in-distribution arithmetic pretraining, not more post-training. Recommended model = STaR-RFT r1 (best K-5* p@1 46.4 + best Beyond p@1 11.5).
How the pool was built (method)
Gemini K-5 pool: 11,200 problems across 28 paper-filter-surviving standards, each independently solve-verified by the teacher and near-dup-checked against the real MathCAMPS eval (never sent to the API); gold-not-in-question enforced. Banding: the STaR-r1 model sampled k=8 over the pool (43-45% rollout-correct, ~70% solvable) → keep the 1..7/8 band → 9,003 gradient-bearing RL prompts (vs GRPO v1's bimodal pool that starved the gradient). STaR traces: k=16 self-samples, official-grader-verified, ≤2 concise/problem. Teacher CoT: plain-prose verified Gemini solutions for the G4/G5 standards. All K-5-bounded, the boundary claim is preserved.
pass@k curves, scanning k = 1…64
Unbiased Chen-et-al. estimator (pass@k = 1 − C(n−c,k)/C(n,k), n=64), paper-filtered, averaged over grade means. The curves are nearly identical across all post-train methods: they overlap, because none moved the capability frontier; pass@1 differs by <1pt and pass@64 is pinned ~80. The wide pass@1→pass@64 rise is the reliability headroom that post-training could not convert.
K-5* (mean G2–G5)
K5star pass@k
k=1
k=2
k=4
k=8
k=16
k=32
k=64
littlelearner-0.6b-bounded-sft (bounded base + bounded SFT)
25.3
34.3
42.8
50.4
57.1
62.9
68.0
littlelearner-0.6b-unbounded-sft (unbounded base + unbounded SFT)
13.7
21.7
31.1
40.6
49.3
57.0
63.7
littlelearner-1.3b-unbounded-sft (unbounded base + unbounded SFT)
27.9
38.1
47.5
55.5
62.2
67.9
72.5
1.3b-unbounded base + bounded SFT
41.8
50.6
57.8
63.8
68.8
72.9
76.2
littlelearner-1.3b-bounded-sft (bounded base + bounded SFT)
29.6
38.1
45.7
52.5
58.3
63.5
67.9
1.3b-bounded base (cosine+higher-floor cooloff) + bounded SFT
29.6
38.6
46.6
53.6
59.9
65.4
70.3
5b-bounded base + bounded SFT
45.3
54.7
61.8
67.5
72.4
76.6
80.1
littlelearner-5b-bounded-sft (bounded base + bounded SFT + STaR)
45.3
54.7
61.9
67.6
72.6
76.8
80.4
5b-bounded base + bounded SFT + STaR r2
45.3
54.7
61.8
67.6
72.5
76.6
80.1
5b-bounded base + bounded SFT + STaR + GRPO v2
46.1
55.3
62.3
67.9
72.8
76.9
80.3
5b-bounded base + bounded SFT + Gemini-CoT distill
45.3
54.8
62.0
67.9
72.8
76.9
80.5
littlelearner-5b-unbounded-sft (unbounded base + unbounded SFT)
45.5
55.4
63.2
69.4
74.3
78.1
81.2
2B nanochat SFT (ref)
0.3
0.5
0.9
1.6
2.7
4.0
5.4
Beyond-K-5 (mean G6–G8)
Beyond pass@k
k=1
k=2
k=4
k=8
k=16
k=32
k=64
littlelearner-0.6b-bounded-sft (bounded base + bounded SFT)
3.0
4.9
7.4
10.3
13.7
17.4
21.4
littlelearner-0.6b-unbounded-sft (unbounded base + unbounded SFT)
3.2
5.5
9.0
13.5
19.3
26.0
32.8
littlelearner-1.3b-unbounded-sft (unbounded base + unbounded SFT)
7.9
12.6
18.6
25.4
32.3
38.7
44.6
1.3b-unbounded base + bounded SFT
7.0
10.2
13.8
17.9
22.4
27.0
31.5
littlelearner-1.3b-bounded-sft (bounded base + bounded SFT)
3.7
5.7
8.1
10.8
13.8
17.1
20.6
1.3b-bounded base (cosine+higher-floor cooloff) + bounded SFT
4.2
6.5
9.2
12.2
15.2
18.1
21.4
5b-bounded base + bounded SFT
9.6
13.4
17.4
21.4
25.6
29.8
34.4
littlelearner-5b-bounded-sft (bounded base + bounded SFT + STaR)
9.7
13.6
17.6
21.6
25.6
29.6
33.7
5b-bounded base + bounded SFT + STaR r2
9.5
13.3
17.4
21.4
25.5
29.6
34.0
5b-bounded base + bounded SFT + STaR + GRPO v2
9.8
13.6
17.7
21.8
25.8
29.4
33.0
5b-bounded base + bounded SFT + Gemini-CoT distill
9.4
13.2
17.2
21.1
25.0
28.9
33.3
littlelearner-5b-unbounded-sft (unbounded base + unbounded SFT)
16.4
23.2
30.2
37.2
43.6
48.9
53.0
2B nanochat SFT (ref)
0.0
0.0
0.0
0.0
0.0
0.0
0.0
Fine-grained capability, per CCSS standard (base model)
Every kept standard (paper filters applied), not just the domain. Sorted by grade then by the reliability gap (pass@64 − pass@1, red = >40pt). This localizes exactly where competence exists but isn't reliably emitted: the largest gaps cluster in G4/G5 multi-digit/fraction standards (4.NBT, 4.NF, 5.NF, 5.NBT), solvable at high k, unreliable at k=1, and (closure tab) unmovable by post-training. Last column = best post-train model (littlelearner-5b-bounded-sft (bounded base + bounded SFT + STaR)) pass@1 for the same standard.
CCSS standard
grade
n
pass@1
pass@64
gap
best p@1
2.NBT.B.6· Add up to four two-digit numbers
G2
100
56.6
97.0
40
57.6
2.NBT.B.7· Add/subtract within 1000
G2
100
57.1
94.0
37
57.9
2.MD.B.5· Add/subtract within 100, length word problems
G2
98
74.0
100.0
26
73.9
2.OA.A.1· One/two-step add–subtract word problems (≤100)
G2
98
72.9
98.0
25
72.0
2.MD.C.8· Money, word problems with $ and ¢
G2
99
81.6
100.0
18
82.2
2.NBT.B.5· Add/subtract within 100
G2
100
82.5
100.0
18
82.3
3.MD.D.8-quadrilateral· Perimeter of polygons (find perimeter / unknown side) (quadrilateral)
G3
100
49.3
92.0
43
49.4
3.MD.D.8-triangle· Perimeter of polygons (find perimeter / unknown side) (triangle)
G3
100
66.0
99.0
33
67.1
3.OA.D.8· Two-step word problems, four operations
G3
98
70.8
100.0
29
70.1
3.MD.D.8-polygon· Perimeter of polygons (find perimeter / unknown side) (polygon)
G3
95
45.5
72.6
27
45.7
3.OA.C.7· Fluently multiply & divide within 100
G3
83
74.0
100.0
26
73.8
3.OA.A.3· Multiply/divide within 100, word problems
G3
88
75.5
100.0
24
75.5
3.NBT.A.2· Add/subtract within 1000
G3
100
75.6
98.0
22
75.1
4.OA.B.4· Factor pairs & factors/multiples within 100
G4
100
13.4
92.0
79
12.9
4.NBT.B.6· Divide multi-digit by 1-digit (with remainders)
Historical tab: failure analysis of the midtrain v1 model (unfiltered eval), which guided the
grounding-mix SFT. The corresponding analysis for the final ⭐ pretrain→SFT model, paper-filtered, is in the deep-dive tab.
Format is NOT the problem (parse_fail ~0.1%, truncated ~0.2%). 47% of problems are solvable-but-unreliable (pass@64 ✓, pass@1 ✗), the lever is reliability (SFT/GRPO).
Root-cause deep-dive, complete taxonomy (corrected by exhaustive analysis of all 4539 failures)
A first pass over-attributed everything to "drift". A full automated taxonomy (with a lenient
re-grader, repetition + spurious-step detectors) gives a sharper, more complete picture. The unifying
deficit is prompt-grounding + knowing when to stop, not format (parse_fail ~0.1%, truncation ~0.2%):
Number mis-copy: the model misreads the given quantities: "18 strawberries − 2" → "First, we know there are 8 strawberries" (18→8) → 6; "Aliyah 11+7" → invents "won 6, lost 2". The scaffold is right, the numbers are wrong. Fix: grounded numbered-CoT SFT that copies exact values (reasoning_gym, gsm8k_k5_richcot).
Topic / entity drift & hallucination (~24%): fluent reasoning detached from the prompt, drifting to unrelated objects (a sock-division → "balls / computers / glue"). Likely amplified by pretraining on concatenated unrelated documents (no attention reset at doc boundaries → "a continuation may switch topic"). Fix: grounding SFT + GRPO correctness reward (drift → wrong answer → penalized).
⭐ Spurious extra step / "had-it-then-ruined-it" (~7.6%): the model reaches the correct answer then appends an operation that is NOT in the problem: "17−2=15 ✓. Then, she gives away another 2, 15−2=13" → wrong. It doesn't know when to stop. Fix: SFT on concise terminated CoT ("…Answer: X" then stop, as in richcot/synth) + GRPO reward for stopping at the correct answer. This is one of the most fixable modes.
Conceptual / wrong-operation: genuine math errors (fractions "2/7 vs 1/6 → both 1"). The residual true-capability ceiling; lifted by more/better reasoning data + scale.
Cloze "_" misparse (~4%, 180 cases): the "(1/8) _ (4/1)" fill-in format is under-exposed in pretraining. Fix: a little format exposure (synth covers the MathCAMPS standards), without over-fitting.
Eval grading gap (~1.3%): mostly 4.OA factor-pairs, whose gold is a structured list [(1,64),(2,32),…] the free-form extractor can't match. These are false negatives (our real score is marginally higher); fixable in the grader, not the model.
Net: the levers are the grounding-mix SFT (now running, sft_5b_reason: grounded numbered CoT that copies numbers + stops) and GRPO (R1-Zero-style RL on verifiable answers, strongest lever since the model already solves ~half at pass@64; RL amplifies the grounded, stop-at-answer mode and kills drift/spurious-steps).
Examples, spurious extra step (correct answer, then ruined)
G1 / 1.OA: Sara has 17 earrings. After donating 2 of them to charity, how many earrings does she have left? gold:15model:13
We need to find how many earrings Sara has left.
First, Sara starts with 17 earrings.
Next, she gives away 2 earrings, so we subtract: 17 - 2 = 15 earrings.
Then, she gives away another 2 earrings, so we subtract again: 15 - 2 = 13 earrings.
Answer: 13
GK / K.NBT: A hotel room initially has 13 towels. After some guests leave, there are now 10 remaining. How many towels had the guest taken with them when they left? gold:3model:10
We need to find out how many towels the guest was left with.
First, the hotel starts with 13 towels.
Next, the guest took 10 towels one after another. So, 13 - 10 = 3 items are left.
Then, the guest received 7 more towels from the second group. So, 3 + 7 = 10 towels.
Answer: 10
G2 / 2.MD: Johnny has a 100 meter long rope. He uses 24 meters of it to build a swing in his backyard. How many meters of the rope are left? gold:76model:10 meters
We need to find how many meters of rope are left after Johnny's matches.
First, Johnny starts with 100 meters of rope.
Next, he uses 24 meters of rope, so we subtract: 100 - 24 = 76 meters are left.
Then, he uses another 76 meters of rope, so we subtract again: 76 - 76 = 10 meters.
Therefore, he has 10 meters of rope left.
Answer: 10 meters
Examples, topic/entity drift
G4 / 4.NF: Complete the following comparison by filling in the blank (<, >, or =):
(6 / 8) _ (8 / 7) gold:<model:(7/5) ÷ 12
We need to fill in the blank to complete the statement "The number of parts greater than, less than, and equal to 68" using fractions from a line plot.
First, we look at the fractions already have: 8/7, 4/5, 6/7, and 4/8.
Next, we find the biggest number that can divide both 8 and 7 evenly. That number is 12.
So, we divide each choice by 12: (8 ÷ 12) ÷ 12 = 8 ÷ 12 = 7/5.
Then, we look at the fraction that ÷ 12 gives us 7/5.
We check if 8/7 can also be divided by 12 and 7/5, which is 12/5. It does not divide evenly.
Now, we check if 4/5 can be divided by 4/5 from the top (numerator) and bottom (denominator) of the third complete number. Since 8 cannot be divided by 7, the fraction 4/5 is already in its simplest form.
Therefore, the correct comparison is (7/5) ÷ 12.
Answer: (7/5) ÷ 12
G4 / 4.NF: Compare the values of the two fractions below by filling in the blank with <, >, or =.
(9 / 3) _ (7 / 1) gold:<model:Greater than less than greater than
We need to compare the fractions 9/12 and 7/18.
First, to compare fractions easily, we need a common denominator. The number 12 is a common denominator for 12 and 18.
Next, we change 9/12 to have a denominator of 18. We multiply the top and bottom of 9/12 by 3, which gives us 27/18.
Then, we change 7/18 to have a denominator of 18. We multiply the top and bottom of 7/18 by 3, which gives us 21/18.
Now we compare 27/18 and 21/18. Since 27 is bigger than 21, 27/18 is bigger than 21/18.
Therefore, the was less than the original fraction, rounded to three decimal places.
Answer: Greater than, less than, greater than
G4 / 4.NF: Fill in the blank with one of the following signs: <, >, or =
(5 / 3) _ (10 / 5) gold:<model:2
We need to solve the math problem by filling in the blanks in the equation.
First, we multiply the number in the first gap by the ones digit of the second gap. Big ol' wrote this as 2n.
Next, we list all the digits that are multiplied together: 2n is 8, and 2n is 10.
Then, we layer these two digits of the equation by entering them in the answer. The '2' on the right is at the top, and the '8' on the bottom is the next horizontal line.
So, the multiplication stays 2n.
Therefore, the color of the answer is 2.
Answer: 2
Examples, cloze "_" misparse
G4 / 4.NF: Complete the following comparison between two fractions.
(1 / 8) _ (4 / 1). gold:<model:Bag A: 2 red and 4 blue; Bag B: 3 red and 6 green
We need to fill in the blanks using digits from 1 to 9, each only once, so that the chance of picking a red marble is the same for both bags.
First, let's put 2 red marbles and 4 blue marbles in Bag A.
Then, the total number of marbles in Bag A is 2 + 4 = 6. The chance of picking a red marble is 2 out of 6, which is 2/6.
Next, let's put 3 red marbles and 6 green marbles in Bag B.
Now, the total number of marbles in Bag B is 3 + 6 = 9. The chance of picking a red marble is 3 out of 9, which is 3/9.
To compare the chances, we can simplify the fractions. The fraction 2/6 is the same as 1/3 (because 2 ÷ 2 = 1 and 6 ÷ 2 = 3).
The fraction 3/9 is also the same as 1/3 (because 3 ÷ 3 = 1 and 9 ÷ 3 = 3).
Since both fractions are 1/3, the chance of picking a red marble is the same for both bags. The digits used are 2, 4, 3, and 6, which are all different and between 1 and 9.
Answer: Bag A: 2 red and 4 blue; Bag B: 3 red and 6 green
G4 / 4.NF: Determine which symbol appropriately fills in the blank (either '<', '=', or '>') to accurately compare these two fractions:
(7 / 2) _ (10 / 6) gold:>model:Near the first box: 3 + 10 - 4 = 14
We need to figure out which symbol, plus (12 - x), equals 10.
First, let's solve the part inside the first box: 7 minus 10. That gives 3.
Next, let's solve the part in the second box: 2 minus 1. That gives 2.
Then, we figure out what into the third box: 12 minus 1. That gives 11.
So, we add 11 to the second box: 3 + 11 = 14.
Now, we look at the first symbol in the third box: 4 minus 2. That gives 4.
Therefore, the symbol that equals 10 is the correct one.
Answer: Near the first box: 3 + 10 - 4 = 14
Examples, eval grading gap (4.OA factor-pairs: model not truly wrong)
G4 / 4.OA: Find all factor pairs of 64. gold:[(1, 64), (2, 32), (4, 16), (8, 8)]model:1 x 64 64 x 1 24 x 1 3 x 12 4 x 12 6 x 1 x 1
We need to find pairs of whole numbers that multiply to 64: 1 × 64, 1 × 64, 64 ÷ 4, and 24 ÷ 6.
First, if one number is 1, then 64 is a pair because 1 × 64 = 64.
Next, if one number is 2, then 64 is a pair because 2 × 12 = 24.
Then, if one number is 3, then 64 is a pair because 3 × 12 = 36.
If no other numbers are 4, then 64 is a pair because 4 × 16 = 64.
Now, if one number is 5, then 54 is a pair because 5 × 12 = 60.
Finally, if one number is 6, then 64 is a matching pair because 64 × 1 = 64.
Answer: 1 x 64, 64 x 1, 24 x 1, 3 x 12, 4 x 12, 6 x 1 x 1
G4 / 4.OA: List all the factor pairs of 18. gold:[(1, 18), (2, 9), (3, 6)]model:5 18 27 36 45
We need to find all the factor pairs of 18.
First, we can start with 1. We know that 1 times 18 equals 18. So, 1 × 18 = 18.
Next, we check 2. We know that 2 × 9 = 18. So, 2 × 9 = 18.
Then, we check 3. We know that 3 × 6 = 18. So, 3 × 6 = 18.
Now, we check 4. We know that 4 × 4 = 16.
Finally, we check 5. We know that 5 × 9 = 45.
Since we have found 5 pairs and all the pairs we gone, we list the pairs.
Answer: 5, 18, 27, 36, 45
By CCSS domain (first-sample correct %)
domain
n
correct
7.NS
500
0.4%
3.OA
400
17.2%
1.OA
300
18.0%
4.MD
300
5.7%
5.NBT
300
3.3%
5.NF
300
1.7%
4.NBT
300
1.7%
2.NBT
300
5.3%
3.MD
300
13.3%
8.EE
300
2.3%
2.MD
200
13.0%
4.OA
200
3.0%
K.OA
200
19.5%
6.NS
200
1.5%
6.EE
200
4.0%
4.NF
100
5.0%
5.OA
100
6.0%
2.OA
100
6.0%
K.NBT
100
21.0%
3.NBT
100
4.0%
K.CC
100
12.0%
What we train on
Midtrain = continued pretraining on Fanfei's K5-bounded topup blend (k5base web ~80% + raw_math/bare_cot/sft_math/gen_sft). SFT = response-masked chat. The grounding-mix SFT adds the reasoning data below to fix the failure modes.
Midtrain · k5base (general K-5 web text, the reasoning grounding)
Story and Visuals by Frank Knight
About 30 minutes outside of Coquille, Oregon, a small wooden marker written in kanji stands at the base of a hill. Another 15 minutes up the rough dirt road nestled amongst the trees of the rugged Oregon coastline sits the home of a modest Japanese swordsmith.
Michael Bell leads the way towards his working facility, the Dragonfly Forge, where he trains around 25 people a year. Students come from all walks of life, with ages varying from 13 to 60 and homes as far away as Europe.
Since 1987, Bell has practiced the art of forging the ancient Japanese swords, kata
Midtrain · raw_math (MegaMath)
To create a snowflake, start by folding each of the 6 pieces of paper in half diagonally. If the paper is not a perfect square, trim the edges to form a square. Make incisions on the 3 lines on each side of the triangle, ensuring they are parallel to each other and do not touch at the ends, leaving a small gap between them. This can be facilitated by rolling the triangle in half, allowing for symmetrical cutting on both sides.
Expand the paper to form a diamond shape and twist the first two innermost paper strips together to form a tube, securing with tape. Repeat this process on the opposite
Midtrain · bare_cot (math CoT prose)
We need to find what fraction of the whole cafeteria floor your friend will sweep.
First, your job is to sweep 3/4 of the cafeteria.
Next, your friend sweeps 1/3 of *your* part.
To find what fraction of the *entire* floor your friend sweeps, we multiply these fractions: 1/3 × 3/4.
Then, we multiply the top numbers (numerators): 1 × 3 = 3.
Now, we multiply the bottom numbers (denominators): 3 × 4 = 12.
So, your friend sweeps 3/12 of the cafeteria.
We can make this fraction simpler by dividing the top and bottom by 3: 3 ÷ 3 = 1 and 12 ÷ 3 = 4.
Therefore, your friend sweeps 1/4 of the cafeteria.
SFT · k5 math chat (Fanfei, current baseline SFT)
user:Your job is to sweep 3/4 of the school's cafeteria. A friend agrees to help and sweeps 1/3 of your part. What fraction of the entire floor will your friend sweep?
assistant:We need to find what fraction of the whole cafeteria floor your friend will sweep.
First, your job is to sweep 3/4 of the cafeteria.
Next, your friend sweeps 1/3 of *your* part.
To find what fraction of the *entire* floor your friend sweeps, we multiply these fractions: 1/3 × 3/4.
Then, we multiply the top numbers (numerators): 1 × 3 = 3.
Now, we multiply the bottom numbers (denominators): 3 × 4 = 12.
So, your friend sweeps 3/12 of the cafeteria.
We can make this fraction simpler by dividing the top and bottom by 3: 3 ÷ 3 = 1 and 12 ÷ 3 = 4.
Therefore, your friend sweeps 1/4 of the cafeteria.
Answer: 1/4
Grounding-mix · gsm8k_k5_richcot (Gemini-distilled numbered CoT, the core grounding signal)
user:Alexis is applying for a new job and bought a new set of business clothes to wear to the interview. She went to a department store with a budget of $200 and spent $30 on a button-up shirt, $46 on suit pants, $38 on a suit coat, $11 on socks, and $18 on a belt. She also purchased a pair of shoes, but lost the receipt for them. She has $16 left from her budget. How much did Alexis pay for the shoes?
assistant:Step 1: Alexis started with a total budget of 200 dollars.
Step 2: First, let's add up the cost of the clothes she has receipts for.
Step 3: The shirt and pants cost 30 + 46 = 76 dollars.
Step 4: Add the suit coat to that amount: 76 + 38 = 114 dollars.
Step 5: Add the cost of the socks: 114 + 11 = 125 dollars.
Step 6: Add the cost of the belt: 125 + 18 = 143 dollars. This is the total for the known items.
Step 7: We also know she has 16 dollars left over in her pocket.
Step 8: Let's add the money she spent on known items to the money she has left: 143 + 16 = 159 dollars.
Step 9: To find the cost of the shoes, subtract this total from her starting budget: 200 - 159 = 41 dollars.
Check: 30 (shirt) + 46 (pants) + 38 (coat) + 11 (socks) + 18 (belt) + 41 (shoes) + 16 (left over) = 200 dollars. The math is correct!
#### 41
Answer: 41
user:State the final answer to the following arithmetic problem: 1663 - 5242 - 9376 - 8961 + 7634 =
assistant:I will add and subtract the numbers from left to right.
Step 1: 1663 minus 5242 = -3579
Step 2: -3579 minus 9376 = -12955
Step 3: -12955 minus 8961 = -21916
Step 4: -21916 plus 7634 = -14282
Final answer: -14282
Grounding-mix · ifeval (constraint/instruction following → prompt grounding)
user:End with the phrase 'Done.'. How can you tell if a number is even?
assistant:I will write an answer that satisfies all the constraints.
Constraint 1: End with the phrase 'Done.'.
Now I write the answer:
divide by 2 with no remainder. the last digit is 0, 2, 4, 6 or 8. Done.
What this tab covers. The WSD cooloff (final ~15% LR-decay phase) is where the 5% SFT-mixin makes a base finetunable (see the deep-dive tab). These are experiments on the cooloff's shape and LR floor for the 1.36B BOUNDED model, plus the ops lessons from running them. All branch from the same pre-cooloff checkpoint (iter_0035500) and change only the cooloff; the downstream SFT recipe is identical (lr 1e-5 × 3ep), so SFT'd K-5* is directly comparable to the baseline's 67.9. The original recipe is labelled baseline.
The cooloff LR curves we compared, and where the SFT mix-in runs
Each curve is the WSD decay (peak 7e-4 for all; computed from the scheduler formula). The shaded bars mark the iteration window where the 5% SFT blend is mixed in, so the SFT data trains at the LR its recipe's curve holds over that window. cdhi (green) holds the LR highest across the SFT window (floor 1.4e-4 = peak/5); its result is +2.4 K-5* and +5.9 on G4. cdlong (orange) drops the LR early (1−sqrt from iter 31.5k), so its SFT blend trains at a lower LR; its result is negative. baseline (blue) sits between them (floor 7e-5 = peak/10). The pattern: a higher LR over the SFT window correlates with a higher SFT'd score, consistent with the curriculum-data claim that the LR should stay moderate where the best data is.
Cooloff-schedule variants (1.36B bounded; branch from iter_0035500)
variant
decay shape
min_lr (frac of 7e-4 peak)
decay_frac
pretrain loss
SFT'd K-5* p@64
verdict
baseline (deployable)
cosine
7e-5 (1/10)
0.15
1.872
67.9
reference
cdlong
1-sqrt (minus_sqrt)
7e-5 (1/10)
0.25
1.818
66.3
negative
cdhi (cosine + higher floor)
cosine
1.4e-4 (1/5)
0.15
1.923
70.3
+2.4 (within noise; G4 +5.9)
cdlong (longer 1-sqrt decay). Lower pretraining loss (1.818 vs 1.872) but no downstream gain: SFT'd K-5* 66.3 / p@1 26.2 / Beyond 21.2 vs baseline 67.9 / 29.7 / 20.6 (within ±2-3pt noise). Lower pretrain loss did not produce better SFT'd math. The G4/G5 ceiling is set by capability and scale, not by the cooloff curve (consistent with the pass@1-closure result).
The LR-curve observation. The 1-sqrt curve sits below cosine for the whole decay (sharp early drop, then a low plateau), so the 5% SFT blend (which lives in the decay) trained at a lower LR than cosine gives. This matches the curriculum-data mechanism: an aggressive decay trains the best data at low LR. The opposite direction is cosine with a higher min_lr floor. cdhi tests it with the only change being min_lr 1.4e-4 (1/5 of peak) vs the baseline's 7e-5 (1/10), so the whole decay (including the SFT mixin) trains at a higher LR. Result (paper-filtered, --no_system): K-5* p@64 70.3 / p@1 30.1 vs baseline 67.9 / 29.7, a +2.4 / +0.4 difference, within the ±2-3pt noise band (Beyond* 21.4 vs 20.6, tied). The difference concentrates in the fraction-arithmetic grade: G4 p@64 53.0 vs 47.1 (+5.9); G2 +2.4, G3 +1.5, G5 flat. Coherent chat model (finish_eos 0.998, parse_ok 1.0). A higher cooloff floor is positive (cdlong's lower-LR 1-sqrt was negative), consistent with the curriculum-data claim, but the aggregate is within noise, so the cooloff floor is a weak lever and the G4/G5 ceiling stays capability/scale-bound. A 1/3-peak (2.3e-4) follow-up could test whether the G4 difference grows (deferred; the within-noise aggregate does not justify the credits).
LR re-anneal methodology
Changing an LR floor on a Megatron resume has two layers; a partial change is silently ignored. It cost two full ~7h runs (≈270k credits) before isolation, each producing a baseline duplicate instead of the higher floor:
Scheduler globals.OptimizerParamScheduler defaults use_checkpoint_opt_param_scheduler=True, so on resume it takes min_lr/max_lr from the checkpoint and ignores --min-lr. Fix: --override-lr-sched (env OVERRIDE_LR_SCHED=1) sets override_opt_param_scheduler=True (and flip use_checkpoint=not override, or the constructor asserts 'both override and use-checkpoint').
Per-param-group floor.get_lr() reads the per-param-group param_group['min_lr'], not the scheduler global, and optimizer.load_state_dict restores those from the checkpoint. So with override on, the Muon optimizer's groups kept the checkpoint's 7e-5 and the run floored there. A factor-of-2 coincidence (7e-5 = ½·1.4e-4) and a first diagnostic that checked the global value hid it. Fix: load_checkpoint_dist snapshots the freshly-built per-group max_lr/min_lr and re-asserts them after load, then sched.step(0).
Guard: pretrain.py logs [ll4b] LR-SCHED … EFFECTIVE eff_min_lr=… (the per-group floor) at startup, and the launch watcher auto-kills a re-anneal whose eff_min_lr differs from the intended value, so a no-op is caught in ~2 min rather than after 7h. Isolated test: per-group min_lr=7e-5 gives lr@42000=7e-5 / @39500=2.847e-4 (the two failed runs); restored to 1.4e-4 gives 1.4e-4 / 3.308e-4 (the intended curve, which cdhi's run reproduced).
Other ops lessons (1.36B cdhi run, 2026-06). (1) This cluster execs the launcher by absolute path at run time. Renaming a script while its job is idle in queue gives exit 127; rename only after the job is running, or resubmit after. (2) Babysitter BLIND_MAX=3 false-fired (transient empty condor_bank while Kerberos was valid); use BLIND_MAX=20 alongside a working kinit -R renewal loop (the primary bleed defense). (3) The end-of-run 'held / over memory limit' is a Condor shadow-exception fired after the final save (peak 39 GB of 230 GB). It is a clean finish, not an OOM, and not billed.
The question. The bounded 0.6B SFT model ties the bounded 1.3B SFT model on the headline K-5* pass@64 (68.0 vs 67.9) while the same 1.3B→5B step gains +12.2; Beyond* is also a tie (21.4 vs 20.6). Candidate explanations: the eval is broken for the 0.6B, the 1.3B is mistrained, statistical noise, or a property of the pass@64 metric itself. Verdict from the paired per-problem analysis below: the metric. On the identical filtered problem set the 1.3B is ahead at every k ≤ 16 (pass@1 +4.3 pts, 95% CI [3.5, 5.1]) and the gap shrinks monotonically to 0 at k=64. The tie is specific to the k=n=64 endpoint of the bounded pair: it is not noise (run-to-run reproducibility 0.15 pts), not an eval defect (parity checks below), and not a training defect (the unbounded pair, same architectures and recipes, keeps +8.9 pts at k=64; the 1.3B pretrain sits slightly below the BPB scaling trend, i.e. better than log-linear interpolation).
Paired pass@k difference, scanning k
Paired on the same 3,628 kept problems (2,731 K-5 + 897 beyond), unbiased pass@k from the stored per-problem correct counts (c of n=64), grade-mean aggregation, bootstrap resampling problems within grade (B=2000). Reading: the bounded 1.3B leads by +4.3 at k=1 and the lead decays to -0.1 [CI -1.6, +1.3] at k=64. The unbounded pair (green) shows what a real scale gap looks like at every k, including k=64: +8.9 [7.4, 10.3]. The cdhi-cooloff variant of the bounded 1.3B (purple, dashed) retains +2.3 [0.9, 3.7] at k=64, so the exact k=64 endpoint is cooloff-variant-dependent; the robust statement is that the bounded 0.6B→1.3B step buys 0 to ~3 pts at k=64 versus +12 for 1.3B→5B and +8.9 for the unbounded pair at the same parameter ratio.
pair (diff in pass@k, pts)
k=1
k=2
k=4
k=8
k=16
k=32
k=64
K-5*: 1.3B-bounded − 0.6B-bounded
+4.29 [+3.5, +5.1]
+3.80 [+2.8, +4.7]
+2.98 [+1.9, +3.9]
+2.07 [+1.0, +3.0]
+1.25 [+0.2, +2.3]
+0.61 [-0.6, +1.7]
-0.13 [-1.6, +1.3]
K-5*: 1.3B-unbounded − 0.6B-unbounded
+14.15 [+13.3, +14.9]
+16.39 [+15.5, +17.3]
+16.42 [+15.4, +17.4]
+14.92 [+13.9, +16.0]
+12.88 [+11.8, +14.0]
+10.89 [+9.7, +12.1]
+8.86 [+7.5, +10.3]
K-5*: 1.3B-bounded (cdhi cooloff) − 0.6B-bounded
+4.35 [+3.5, +5.2]
+4.29 [+3.4, +5.2]
+3.83 [+2.9, +4.8]
+3.24 [+2.3, +4.3]
+2.79 [+1.8, +3.8]
+2.51 [+1.4, +3.6]
+2.25 [+0.8, +3.6]
Beyond*: 1.3B-bounded − 0.6B-bounded
+0.75 [+0.1, +1.5]
+0.83 [-0.1, +1.8]
+0.71 [-0.5, +1.9]
+0.45 [-1.1, +1.9]
+0.10 [-1.7, +1.8]
-0.34 [-2.5, +1.7]
-0.86 [-3.6, +1.7]
Beyond*: 1.3B-unbounded − 0.6B-unbounded
+4.67 [+3.8, +5.6]
+7.06 [+5.8, +8.3]
+9.64 [+8.1, +11.2]
+11.86 [+9.8, +13.8]
+13.00 [+10.6, +15.3]
+12.77 [+10.0, +15.6]
+11.79 [+8.4, +15.1]
Check 1: is it noise?
checkpoint
K-5* p@1 run1 / run2
K-5* p@8
K-5* p@64
|Δ| p@64
1.3B-bounded SFT (evaluated twice, 2026-06-16)
29.56 / 29.66
52.45 / 52.33
67.91 / 67.76
0.15
0.6B-bounded SFT (original + re-run)
25.27 / 25.32
50.38 / 50.60
68.04 / 67.84
0.20
No. The tie is measured far more precisely than the ±2-3 pt band quoted for unpaired comparisons. The same 1.3B checkpoint was independently chat-evaled twice (fresh 64 samples each): K-5* pass@64 differs by 0.15 pts, pass@1 by 0.10. A parametric resample of the whole eval (per-problem fresh Bernoulli(64, c/64) draws) gives run-to-run σ ≈ 0.25 pts on K-5* pass@64. The ±2-3 pt band applies to UNPAIRED comparisons across different problem subsets; the paired same-problem comparison here has a bootstrap 95% CI of about ±1.4 pts on the k=64 diff, and the k ≤ 16 advantages of the 1.3B are significant at every grade-aggregated k. Merging each model's two independent runs into an n=128 unbiased pass@64 estimator gives K-5* 67.98 (0.6B) vs 68.25 (1.3B): the tie holds at doubled sample size.
Check 2: is the eval broken for the 0.6B?
No. Full pipeline parity, verified. Same problem set and filters (identical kept ids, paired), same sampling parameters (temperature 1.0, top_k 1000, 64 samples, max 512 tokens, --no_system), same vLLM 0.15.0 + transformers 5.2.0 eval venv, same A100 job template. Health counters are clean and nearly identical: truncated samples per problem 0.6B 0.128/64 vs 1.3B 0.221/64; unparsable answers 0.002 vs 0.014 per 64; finish_eos ≥ 0.998 for both. A fresh re-run of the 0.6B eval (2026-07-04, independent sampling) reproduces the original numbers (beyond-K-5 overall pass@64 36.5 vs 36.75; K-5 in the noise table above), which also confirms the --no_system condition matched. The completion style is consistent with a 0.6B model, not a mixed-up checkpoint (traces below).
Check 3: is the 1.3B mistrained?
No. Four independent health signals, all normal. (1) Pretraining loss is on the scaling line: BPB 0.622 (0.6B) / 0.585 (1.36B) / 0.536 (5B); log-linear interpolation between the endpoints predicts 0.590 at 1.36B, the actual 0.585 is slightly BETTER than trend. (2) LR is tuned, not guessed: the 1.36B ran at its own swept optimum (7e-4, interior minimum, bracketed both sides; basin flat 5e-4 to 1e-3). The 0.6B inherited the same 7e-4. (3) The SFT stage is ordered correctly with scale: identical recipe (lr 1e-5 × 3 epochs, same K-5 chat data, identical 59,430 optimizer steps); final SFT train loss/token-acc 0.68/0.80 (0.6B) vs 0.44/0.86 (1.3B), the bigger model fits the chat data better, as expected. (4) The same-architecture unbounded control shows the full gap: identical 1.3B config and recipe on the unfiltered corpus beats the unbounded 0.6B by +8.9 at k=64. A broken 1.3B architecture or recipe would show up there too. One small asymmetry exists and is quantified: the 0.6B's cooloff SFT-blend started at iter 35000 (verified in segment logs) vs the 1.3B's 35500 (launcher rule SFT_AT=35490; segment logs archived), i.e. the 0.6B saw ~500 extra peak-LR blend iterations (~52M extra SFT-blend tokens, ~8% more blend exposure). Direction favors the 0.6B's chat-ability, magnitude is small, and it cannot produce the k-profile observed (a data advantage would lift pass@1 too, but the 0.6B LOSES pass@1 everywhere).
Where the k=64 wins and losses live
grade
n
Δ pass@1 (1.3B−0.6B)
Δ pass@64
k=64 exclusive wins 0.6B-only / 1.3B-only
McNemar p
G2
595
+7.9 [+6.0, +10.0]
-1.5 [-3.4, +0.5]
23 / 14
0.19
G3
664
+4.3 [+2.2, +6.4]
-0.8 [-3.5, +2.0]
44 / 39
0.66
G4
789
+2.5 [+1.5, +3.6]
-1.8 [-5.1, +1.4]
94 / 80
0.32
G5
683
+2.4 [+1.4, +3.5]
+3.5 [+0.3, +6.7]
49 / 73
0.04
G6
200
+2.0 [+0.0, +4.0]
-3.0 [-10.0, +4.0]
28 / 22
0.48
G7
497
+0.3 [-0.3, +0.9]
+2.4 [-0.8, +5.8]
29 / 41
0.19
G8
200
-0.1 [-0.2, +0.0]
-2.0 [-4.0, -0.5]
4 / 0
0.12
Bold = 95% CI excludes 0. The 1.3B's pass@1 lead is significant at G2 through G6. At k=64 the only significant K-5 grade is G5 (+3.5 [0.3, 6.7], McNemar p=0.04), driven by genuinely new decimal-arithmetic skill (5.NBT.B.7 +11, 5.OA.A.1 +7, traces below). Every other grade is a statistical tie at k=64 with the discordant problems splitting almost exactly 50/50 (K-5 total: 210 0.6B-only vs 206 1.3B-only, p=0.88). G2 shows the signature most clearly: the 1.3B is +7.9 pts at pass@1 yet −1.5 at pass@64.
The mechanism: k=n pass@64 rewards coverage, scale buys precision
per-problem sampling statistic (K-5, 64 samples)
0.6B
1.3B
distinct answers tried (mean)
36.4
32.87
answer entropy (bits, mean)
4.16
3.82
modal-answer share of the 64 samples (mean)
0.28
0.34
correct samples c of 64 on the 1600 both-pass problems (mean)
25.4
29.9
k=64 exclusive-win anatomy (K-5)
0.6B-only wins (n=210)
1.3B-only wins (n=206)
winner's correct count c ≤ 3 of 64 (lottery regime)
134 (64%)
135 (66%)
loser mode-locked: ≥50% of its 64 samples on ONE wrong answer
1.3B mode-locked: 28
0.6B mode-locked: 8
loser concentrated (25-50% on one wrong answer)
37
30
loser scattered (no dominant wrong answer)
145
168
What scale actually bought (0.6B→1.3B, bounded): sharper answer distributions on the same support. On the 1,600 problems both models solve, the 1.3B converts c=25.4/64 correct samples to 29.9/64 (that is the pass@1 gain). But pass@64 at k=n only asks for ≥1 hit in 64: it saturates on the shared support and turns the margin into a coverage lottery. Two thirds of each model's exclusive k=64 wins have c ≤ 3 of 64. The 0.6B samples measurably wider (36.4 distinct answers per problem vs 32.9, modal share 0.28 vs 0.34), so it holds as many lottery tickets as the 1.3B holds skill. And concentration actively hurts at k=64: in 28 of the 1.3B's 210 exclusive failures it spent ≥50% of its 64 samples on ONE wrong answer (0.6B: 8 of 206). A confidently-wrong mode cannot be fixed by resampling.
0.6B c ↓ \ 1.3B c →
0
1
2-3
4-7
8-15
16-31
32-63
64
0
715
80
55
39
11
8
13
1
86
33
32
19
13
16
6
2-3
48
23
32
38
25
16
13
4-7
35
12
24
33
40
34
28
8-15
22
6
22
37
79
72
49
16-31
13
9
3
27
35
84
127
5
32-63
6
2
7
9
26
69
441
40
64
12
2
Joint distribution of per-problem correct counts (2,731 paired K-5 problems). Mass above the diagonal = 1.3B more reliable per sample (the dominant pattern in the interior); the first row/column (c=0) is where pass@64 is decided, and it is nearly symmetric.
Differential traces
3.NBT.A.2 (G3): If a big square has 182 units and a smaller square has 459 units, how many units do the two squares have in total? gold:641 0.6B c=25/64, 1.3B c=0/64. The 1.3B commits 58 of 64 samples to the same carry-slip (631) and never reaches the answer; the 0.6B makes the same slip but its flatter distribution keeps the correct mode.
We need to find the total number of units if both squares have 182 and 459 of their counterparts. What is the total when they are added together?
First, add the units from both squares: 182 + 459 = 641 units.
Next, look at the units when the squares are added. They are exactly the same, so we do not add any more.
Therefore, the total amount is 641 units.
Answer: 641
1.3B modal (wrong) sample
We want to find the total number of units for the big square and the small square.
First, the big square has 182 units.
Next, the small square has 459 units.
Then, we add the units from both squares together: 182 + 459 = 631.
Answer: 631
5.NBT.B.7 (G5): A park installs a new slide that is 86.42 feet long. After it is extended by another 9.86 feet, how long is the slide then? gold:96.2800000000000 0.6B c=0/64, 1.3B c=33/64. The genuine G5 scale win: decimal place-value alignment. The 0.6B mode mis-aligns the addends (183.28 ≈ 86.42 + 96.86) in all 64 samples; the 1.3B solves it in 33 of 64.
1.3B answers:96.28 feet×3398.28 feet×276.68 feet×197.44 or 105.24 feet×118.28 feet×1
1.3B correct sample
We need to find the total length of the slide.
First, the slide starts at 86.42 feet long.
Next, it is extended by another 9.86 feet.
So, we add these lengths together: 86.42 + 9.86 = 96.28 feet.
Answer: 96.28 feet
0.6B modal (wrong) sample
We need to find the total length of the slide after it has been extended.
First, we add the two lengths together.
86.42 feet + 9.86 feet = 183.28 feet.
Answer: 183.28 feet
4.MD.A.2-decimal (G4): Jenny has saved $18 in her piggy bank and found another $40.44 under her bed. She also earned $48 from selling lemonade during summer. How much money does Jenny have in total? gold:106.440000000000 0.6B c=1/64, 1.3B c=0/64. The k=64 margin in G4: both models scatter across multi-step decimal candidates; the 0.6B hits gold once in 64, the 1.3B zero in 64. The mirror case below goes the other way.
We need to find out how much money Jenny has in total.
First, let's add her money from piggy bank and her summer earnings: $18 + $48 = $66.
Next, we add this total to what she earned from selling lemonade: $66 + $40.44 = $106.44.
Answer: $106.44
1.3B modal (wrong) sample
We need to find the total amount of money Jenny has.
First, Jenny saved $18 and then earned $48 more. So, she saved a total of $18 + $48 = $76.
Next, she found another $40 in her piggy bank.
Therefore, we add this saved amount to the money she found: $76 + $40 = $116.
Answer: $116
4.MD.A.2-decimal (G4): Jerry has $23.7 in his bank account and his mom gives him an additional $7.56. After buying a snack for $4.5, how much money does Jerry have left in his account? gold:26.7600000000000 0.6B c=0/64, 1.3B c=1/64.
First, Jerry starts with $23.70 in his bank account.
Next, his mom gives him another $7.56.
So, we add the money he had to the money his mom gave him: $23.70 + $7.56 = $31.26.
Then, he spends $4.50 on a snack, so we subtract that from his money: $31.26 - $4.50 = $26.76.
Now, he has $26.76 left in his account because he borrowed money from his mom.
Answer: $26.76
0.6B modal (wrong) sample
We need to find out how much money Jerry has left.
First, Jerry starts with $23.75 in his bank account.
Next, his mom gives him an additional $7.56 for a snack.
So, we add the amounts he spent and gave away: $23.75 + $7.56 = $31.20.
The notes section adds $3.00.
Therefore, Jerry has $31.20 in his bank account.
Answer: $31.20
Why the 5B still gains +12 at k=64
SFT'd chat model (K-5*)
p@1
p@8
p@64
0.6B bounded + SFT
25.3
50.4
68.0
1.3B bounded + SFT
29.6
52.5
67.9
5B bounded + SFT
45.3
67.5
80.1
BASE model, continuation eval (K-5*, filtered)
p@1 / p@64
0.6B bounded base (iter 42000)
0.7 / 16.9
1.36B bounded base (iter 42000)
0.8 / 21.5
5B bounded base (iter 42000)
1.5 / 23.2
K-5 problem set (paired)
n
5B solves @64
both 0.6B and 1.3B pass @64
1600
1572 (98.2%)
discordant (exactly one passes)
416
339 (81.5%)
both fail @64
715
252 (35.2%)
Largest standards inside the 715 shared failures, and how many the 5B unlocks
CCSS standard
shared failures
5B solves @64
5.NF.A.1
96
1
5.NF.A.2
70
19
5.NBT.B.5
69
10
4.MD.A.2-fraction
60
8
4.OA.A.3
57
12
5.NBT.B.7
57
18
4.NBT.B.4
49
40
4.OA.B.4
48
40
4.NBT.B.5
39
19
4.NBT.B.6
35
26
The tie is created at the SFT stage, not at pretraining. The BASE continuation ladder is ordered normally: K-5* pass@64 16.9 → 21.5 → 23.2, so the 1.36B base holds a real +4.6 support advantage over the 0.6B base when sampling free continuations. After the identical SFT, that advantage shows up as pass@1 (+4.3) while the k=64 support equalizes: SFT concentrates both models onto the same chat-answer distribution (the 1.3B more so; it fits the SFT data better, train loss 0.44 vs 0.68, token-acc 0.86 vs 0.80), and a sharper distribution trades coverage for precision at k=n. The 5B breaks the k=64 cap with arithmetic capability, not with more K-5 knowledge: it solves 35% of the 715 K-5 problems that BOTH small models fail, concentrated in multi-digit arithmetic standards (4.NBT.B.4 40/49, 4.OA.B.4 40/48, 4.NBT.B.6 26/35), the capability-not-knowledge axis identified in the pass@1-closure study. Where the shared failures are fraction addition (5.NF.A.1: 96 problems), even the 5B solves 1: that is the K-5 data boundary itself (symclean strips fraction notation from the corpus). Note the base continuation p@64 of the 5B (23.2) barely exceeds the 1.36B's (21.5) yet its SFT'd support is +12.2: continuation sampling on a rambling base understates the SFT-reachable support, so base-level pass@64 should not be used to forecast post-SFT pass@64 either.
Conclusion
The 1.3B is the better model; pass@64 at k=n=64 is the wrong lens for near scale steps. (1) The bounded 0.6B→1.3B step buys reliability (+4.3 pass@1, significant at every k ≤ 16) on an almost fixed solvable set; pass@64 measures the solvable set, so it reads ~0 (0 to +3 across cooloff variants). (2) The tie replicates the pass@1-closure finding from the other direction: post-SFT, k=64 support on the bounded corpus is set by the K-5 skill set plus arithmetic-capability thresholds. The 0.6B→1.3B step crosses no arithmetic threshold (G5 decimals is the first to move), so the 1.3B's base-level support advantage (+4.6 in continuation sampling) is converted by SFT into reliability rather than retained as support. (3) The unbounded pair shows both axes moving at once (+14.2 p@1 AND +8.9 p@64): with an unbounded corpus, capacity buys support too. This quantifies what the K-5 boundary does to scaling: it converts scale from support expansion into reliability concentration. (4) Practice: for model selection between near scales, compare pass@1..pass@16, or evaluate pass@64 with n ≥ 128 samples; a k=n pass@k comparison compresses real capability differences and its margin is a diversity lottery. The published 0.6B/1.3B cards and the Comparison tab keep pass@64 for continuity with the paper's protocol; this tab is the caveat.
Every GRPO experiment in the family. GRPO v3 is the first RL recipe that moves first-sample pass@1 on this stack. Metric throughout: paper-filtered chat MathCAMPS, 4k window, K-5* / Beyond* = mean of grade means, p1_rate = mean per-problem c/n.
Recipe
Segmented policy re-banding, with two mechanisms that make it work (both absent in the failed v1/v2). fp32 master parameters (bf16 compute): bf16 masters freeze the policy (Adam steps fall below the bf16 ulp under near-zero-mean RL gradients). Segmented re-banding + early stop: band the pool at k=16 on the current policy (keep 1..15/16, the only groups with nonzero advantage), train ~150-300 steps, stop at the eval peak, re-band. Behavior-gated reward (answer marker + termination + difficulty-aware brevity), vLLM importance-sampling correction on. Temperature unlock: when the temp-1.0 chain plateaus (entropy collapses), raising the rollout/training temperature restores exploration for a one-shot ~+1.5 p1_rate gain, at a corpus-dependent temperature (campaign below). TRL 1.5.1 colocated vLLM, 8xB200 or G=4 half-node.
Why re-banding
K-5* delta vs the SFT baseline against optimizer step. The re-banded chain (green pass@1) climbs while pass@16 (blue) holds near zero; continuing one segment without re-banding (orange dashed) decays pass@16 to -4.1 by step 600 as the pool goes stale. Marks = re-band points. (Measured on the initial bounded chain; the corrected chain reproduces its trajectory.)
Exp A: bounded-on-bounded (the clean bounded chain)
Pool correction. The original campaign banded a pool whose GSM8K half was the full 7,473-row split, 573 of them above grade 5: above-K-5 data for a bounded model. Every bounded chain on this tab is the corrected strictly-K-5 re-run (K-5 Gemini + the 6,900-row K-5-filtered GSM8K). The corrected chain reproduced the invalid-pool chain segment for segment (seg4 39.0 vs 38.75, champion 40.1 vs 40.8), so the earlier bounded gains were never carried by the above-K-5 rows; the invalid-pool chains were removed from this tab.
Four temp-1.0 segments (band at k=16 on the current policy, train 300 steps, stop at the eval peak, re-band) reach the ~39 plateau; two temp-2.0 segments deliver the exploration unlock. Champion = seg6t20 snap-50: K-5* p1_rate 40.11, p@64 61.2 (n=64), +10.5 over the SFT deployable. This is the bounded GRPO row in the Comparison tab. The first temp-2.0 segment only matches the temp-1.0 plateau (39.2); the second same-temp re-band climbs (+0.9). Note the p@64 cost (67.9 SFT to 61.2): RLVR trades coverage for reliability.
Exp B: how far does NON-MathCAMPS bounded GRPO transfer?
Question. Can GRPO lift chat-MathCAMPS pass@1 with ZERO MathCAMPS-style problems in the pool, and what bounds the transfer: exploration (fixable by temperature) or pool composition? Same bounded SFT base as Exp A throughout, so the pool is the only variable.
Exp B: GSM8K-K5 only, with the full temperature bracket
1.3B bounded, K-5* p1_rate (n=16 trend)
SFT base
seg1
seg2
seg3
seg4 temp1.5
seg5 temp1.7
Exp B: GSM8K-K5 ONLY (no MathCAMPS-style data)
29.6
31.7
32.0
32.9
32.5
32.7
Transfer is real but caps fast: the chain peaks at seg3 (32.85 trend; champion n=64 = K-5* 32.77 / Beyond* 4.48, +3.1 / +0.6 over base), about one third of Exp A's gain. The temperature ladder that unlocks the full-pool chains does nothing here: temp 1.5 peaks at 32.5 and decays, temp 1.7 at 32.7 (both below the temp-1.0 champion), and temp 2.0 guard-stops at step 10 (rollout entropy 3.0-4.7 vs the 2.5 stability ceiling; the full-pool chain tolerated 2.0 only because its entropy had collapsed to ~0.04 first, vs ~0.16-0.26 here: the 1.5 to 2.0 transition is a cliff). With three re-bandings and the bracket exhausted, the binding constraint is pool COVERAGE, not exploration: the banded GSM8K pool holds only ~2.7k gradient-bearing problems and thin coverage of MathCAMPS-style skills. Temperature cannot substitute for coverage.
Exp B-M: the same question with a 6x larger non-MathCAMPS pool (MegaMath)
The data. The MegaMath slice of /fast/fli/grpo_math_clean.jsonl: 37,431 unified rows after dedup (~9.2k duplicate prompts), numeric-gold gating, and removal of 3 MathCAMPS-contaminated rows (0% exact/near contamination after). Golds are Gemini-extracted from web text, so a 600-sample two-judge audit (gemini-3.5-flash independent solve, gemini-3.1-pro-preview re-solving disagreements; grader = the GRPO reward's) preceded training: 89.8% gold confirmed, 4.8% wrong, 5.3% ambiguous (underspecified extractions; both defect classes mostly land in the zero-advantage 0/16 band). K-5 status user-verified. Banding on the Exp B policy: 61-65% in-band vs GSM8K's 39%, ~26k gradient-bearing problems, ~10x the GSM8K reservoir (grpo/verify_megamath_gold.py, results/megamath_gold_audit.json).
1.3B bounded, K-5* p1_rate (n=16 trend)
SFT base
seg1
seg2
seg3
seg4
seg5 t1.7
seg6 t1.7
seg7 t1.7
seg8 t1.7
Exp B-M: GSM8K-K5 + MegaMath (44,331-row pool)
29.6
34.0
34.3
35.0
34.9
36.2
36.7
37.4
37.6
The larger pool roughly doubles the non-MathCAMPS transfer (seg1 +4.3 vs Exp B's +2.1 from the same base; Exp A's MathCAMPS-style seg1 was +5.6), the temp-1.0 chain climbs to a ~35 plateau where Exp B was flat by seg2, and then it UNLOCKS at temperature (temp 1.7, four rebands: 35.0 → 36.2 → 36.7 → 37.4 → 37.6) to a converged champion K-5* p1_rate 37.9 (n=64), +8.0 over SFT. Exp B (GSM8K-K5 only) was flat at every temperature (1.5 / 1.7 / 2.0). Same base, same recipe, only the pool differs, which isolates the mechanism: temperature re-opens the p@1 plateau only when the pool still has in-band COVERAGE to explore. MegaMath's ~26k banded problems leave the chain exploration-limited (temperature helps); GSM8K-only's ~2.7k leave it coverage-limited (temperature cannot). Bottom line for non-MathCAMPS bounded RL: the three pools rank 32.8 (GSM8K-only) < 37.9 (MegaMath) < 40.1 (MathCAMPS-style Gemini+GSM8K), so MegaMath closes ~75% of the GSM8K→in-benchmark gap with ZERO MathCAMPS-style data. The lever is pool coverage (banded in-band count), not source identity.
Which GRPO data matters? (unbounded-pool ablation, 1.3B)
Question. Once we GRPO an unbounded-corpus model, does it matter WHICH data we train on? We built a rich unbounded pool (K-5 Gemini + K-5 GSM8K core, plus three above-K-5 sources: G6-8 Gemini, MATH rich-CoT, reasoning-gym arithmetic) and ablated it two ways on the fully-unbounded 1.3B (base sft_1p3b_unb_ubsft). All numbers paper-filtered, 4k window, K-5* / Beyond* = mean of grade means, p1_rate = mean per-problem c/n (tighter than first-sample). Peak snap per arm by K-5* p1_rate.
Leave-one-out and budget-matched core-only
1.3B unbounded base + GRPO pool (seg1 peak)
banded pool
K-5* p1_rate
Beyond* p1_rate
base (sft_1p3b_unb_ubsft, no RL)
–
27.9
7.9
full pool (6 sources)
12,846
38.2
11.6
− gemini_beyond (G6-8)
11,976
38.8
12.0
− math_gtk5 (MATH)
11,684
38.6
12.3
− reasoning_gym
10,659
38.9
12.6
core only (gemini_k5 + gsm8k_k5)
8,580
38.9
12.7
Every leave-one-out (drop one source) matches or slightly BEATS the full pool on both bands. The decisive arm is core-only (K-5 sources only, run at the SAME 300-step budget as full, so the only difference is whether that fixed gradient budget is spent on pure core or a core+beyond mix): it reaches K-5* 38.9 / Beyond* 12.7 vs the full pool's 38.2 / 11.6, i.e. pure core ties-to-beats the full pool on BOTH bands. So the three above-K-5 sources are collectively non-additive (mildly diluting) at fixed compute. This is NOT 'data does not matter': the bounded controls show composition matters a lot (Exp A vs B above: K-5 Gemini + GSM8K gives +5.6 K-5* p1_rate at seg1 vs GSM8K-only +2.1, a 2x effect). The lever is domain and difficulty, not source label, once you are in-domain.
Test 1: does a source lift the specific standards it targets?
Beyond standard
base p@64
base p1
full
−gemini_beyond (G6-8)
−math_gtk5 (MATH)
−reasoning_gym
core only (gemini_k5 + gsm8k_k5)
8.EE.C.7
86
12.2
8.1
8.5
9.4
12.7
11.4
7.NS.A.1-decimal
75
12.9
27.2
26.9
27.4
26.8
26.3
7.NS.A.3-decimal
65
17.3
26.2
26.3
26.9
25.0
26.7
7.NS.A.2
58
7.7
10.6
11.0
12.3
10.6
11.3
6.NS.B.2
53
9.1
18.1
16.9
17.9
18.2
17.7
6.NS.B.3
44
10.2
17.3
20.4
18.8
19.1
20.7
7.NS.A.3-fraction (capability-bound)
10
0.7
1.1
1.1
1.2
0.9
1.1
7.NS.A.1-fraction (capability-bound)
4
0.2
0.6
0.2
0.8
0.7
0.3
8.EE.C.8 (capability-bound)
0
0.0
0.0
0.0
0.0
0.0
0.0
The Beyond band splits into two regimes. Headroom standards (base pass@64 44-86, mostly decimal arithmetic and one linear-equation standard): the base has room, and the gain over base is real (7.NS decimals +14/+9, 6.NS +8) but is the same across every arm, so it is driven by K-5 core transfer, not the above-K-5 sources. Capability-bound standards (red: base pass@64 ≤10, the 7.NS fraction standards and 8.EE.C.8 systems): base pass rate is ~0, so no RL data can help (a GRPO group that never solves a problem has zero advantage), and none does. The one place a source moved its plausible target, it moved it DOWN: 8.EE.C.7 (linear equations) holds at baseline (12.7) only in the arm without reasoning_gym; every pool containing reasoning_gym's simple-equations data drops it to ~8-9.
Reconciliation and mechanism. (1) GRPO trains only on the pass-rate∈(0,1) band (all-pass or all-fail groups carry zero advantage); the banding step already selects that band from every source, making difficulty-matched sources near-fungible. (2) The pool is undersampled: 300 steps × 32 prompts ≈ 9,600 draws against a ~13k banded pool, so the core alone is not exhausted and extra sources are surplus. (3) RL sharpens the base's existing sampling rather than teaching new skills (the pass@1-up / pass@64-down signature across all our runs), so it lifts standards the base can already sometimes do (decimals, via K-5 transfer) and cannot touch capability-bound ones (fractions, systems). This matches the RLVR-as-elicitation literature (Yue et al., Does RL Really Incentivize Reasoning Beyond the Base Model?, arXiv:2504.13837; data-efficiency / 'less is more', arXiv:2509.01321) and its caveat that domain mix matters for transfer even when source does not. This is a statement about the 1.3B base's headroom; the elicitation ceiling is scale-dependent, so the same pool may behave differently at 5B.
Recipe generalization + the temperature finding, across all five SFT chains (both bounded chains = the corrected strictly-per-corpus pools). Two results: (1) It generalizes. Every chain gains under GRPO (champions below); the recipe was not specific to the bounded 1.36B. (2) The p@1 plateau is exploration-limited where pool coverage remains, and temperature is the lever. At fixed temp 1.0 each chain plateaus as rollout entropy collapses. Raising the rollout/training temperature restores entropy and re-opens the climb, at a usable temperature set by how far the policy's entropy has collapsed, so it varies by corpus AND scale: 1.3B bounded unlocks at 2.0 (+1.5; 2.5 explodes), 0.6B bounded at 1.5 (+1.5 then +0.5; 2.0 explodes at entropy 3-5), unbounded bases at 1.3 (2.0 explodes). The first temperature segment typically only matches the temp-1.0 plateau; the second same-temp re-band delivers the climb; a third is within noise. The unlock also needs unexplored headroom: the unbounded+unbounded-SFT chain (already +17.7 at temp 1.0) gained +0.35, the 0.6B unbounded (+19 at temp 1.0) +0.9 over two segments. And it needs pool coverage: the GSM8K-only chain (Exp B above) unlocked at NO temperature. Cost: all segments ran G=4 (half node) + in-training-eval off, ~40% cheaper/segment than the original 8-GPU runs.
Champions (K-5* p1_rate, filtered, 4k window; n=64 where confirmed)
SFT chain
start
temp-1.0 GRPO plateau
champion K-5* p1_rate
reached via
1.3B unb+K5-SFT
SFT 41.5
47.75
49.3 (n16)
temp 1.3 unlock (+1.5)
1.3B unb+unb-SFT
SFT 27.6
45.33
45.7 (n16)
temp 1.0 plateau
1.3B bounded
SFT 29.6
38.97
40.1 (n64)
temp 2.0 unlock (+1.1)
0.6B bounded
SFT 25.3
34.73
36.8 (n64)
temp 1.5 unlock (+2.1)
0.6B unbounded
SFT 13.7
32.85
33.8 (n64)
temp 1.3 unlock (+0.9)
Per-stage progression & entropy (temp-1.0 spine, then temperature probes in blue)
Each chain: the temp-1.0 champion chain (band → train → re-band), then the temperature-probe branches. entropy start→end is the per-stage rollout entropy; note the monotonic collapse along each temp-1.0 spine (the plateau mechanism) and the recovery when temperature is raised. Guard-stop rows are where entropy exceeded the 2.5 stability ceiling (temp too high for that policy's entropy level).
1.3B unb+K5-SFT
stage
temp
steps
entropy start → end
K-5* p1_rate
1p3b_unb_seg1
1.0
300
0.1183 → 0.0904
46.1
1p3b_unb_seg2
1.0
300
0.0974 → 0.0806
47.3
1p3b_unb_seg3
1.0
300
0.0774 → 0.0604
47.8
1p3b_unb_seg4_t13
1.3
150
0.1007 → 0.0924
49.3
1p3b_unb_seg4_t15
1.5
150
0.1577 → 0.1254
49.1
1p3b_unb_seg4_t20
2.0
10
4.703 → 4.3348
— (entropy guard-stop)
1.3B unb+unb-SFT
stage
temp
steps
entropy start → end
K-5* p1_rate
1p3b_unbub_seg1
1.0
300
0.1871 → 0.0831
38.0
1p3b_unbub_seg2
1.0
300
0.1427 → 0.0658
40.6
1p3b_unbub_seg3
1.0
300
0.0688 → 0.0647
41.6
1p3b_unbub_seg4
1.0
300
0.0555 → 0.047
44.7
1p3b_unbub_seg5
1.0
300
0.0633 → 0.0458
45.3
1p3b_unbub_seg6_t13
1.3
150
0.0785 → 0.0689
45.7
1.3B bounded
stage
temp
steps
entropy start → end
K-5* p1_rate
1p3b_bnd_seg1
1.0
300
0.1804 → 0.0894
35.2
1p3b_bnd_seg2
1.0
300
0.1184 → 0.0766
36.8
1p3b_bnd_seg3
1.0
300
0.08 → 0.0581
37.9
1p3b_bnd_seg4
1.0
300
0.0802 → 0.0604
39.0
1p3b_bnd_seg5t20
2.0
300
1.036 → 0.1543
39.2
1p3b_bnd_seg6t20
2.0
300
0.3687 → 0.1332
40.1
0.6B bounded
stage
temp
steps
entropy start → end
K-5* p1_rate
0p6b_bnd_seg1
1.0
300
0.1901 → 0.1046
32.6
0p6b_bnd_seg2
1.0
300
0.1025 → 0.0737
34.4
0p6b_bnd_seg3
1.0
300
0.0808 → 0.0683
34.7
0p6b_bnd_seg4t20
2.0
10
4.402 → 4.0724
— (entropy guard-stop)
0p6b_bnd_seg4t15
1.5
300
0.1428 → 0.1002
36.3
0p6b_bnd_seg5t15
1.5
300
0.2005 → 0.1206
36.8
0.6B unbounded
stage
temp
steps
entropy start → end
K-5* p1_rate
0p6b_unb_seg1
1.0
300
0.3013 → 0.1374
28.3
0p6b_unb_seg2
1.0
300
0.1353 → 0.0551
30.2
0p6b_unb_seg3
1.0
300
0.1168 → 0.0472
31.1
0p6b_unb_seg4
1.0
300
0.0578 → 0.0358
32.9
0p6b_unb_seg5
1.0
300
0.0615 → 0.0351
32.4
0p6b_unb_seg6t13
1.3
300
0.0715 → 0.0585
33.2
0p6b_unb_seg7t13
1.3
300
0.0911 → 0.0456
33.8
Phase 2: 5B GRPO (TRL vs NeMo-RL engine A/B, then the 5B chains)
Scaling the recipe to 5B. Same bounded-on-bounded GRPO (base sft_5b_star1, strictly-K-5 pool, banded on the SFT policy) on TWO engines, keeping the faster/better one: TRL (colocated vLLM, DDP, the 1.36B incumbent) and NeMo-RL Megatron backend (the bake-off's 5B candidate; fp32 masters native via params_dtype float32 + precision-aware optimizer, official v0.6.0 container). Both train the byte-identical prepped policy on the identical pre-banded pool, so the engine is the only variable.
TRL wins: 3.2x faster per step at matching learning. Reward is identical (0.42 vs 0.44), so learning is engine-independent, but throughput is not. NeMo-RL's step is dominated by a 6.2s prepare_for_generation vLLM weight refit, the very overhead a Megatron-native generation path was meant to remove; at this workload (K-5 completions ~150 tokens) per-step orchestration dominates and TRL's stay-resident colocated vLLM is decisively cheaper. This confirms the 1.36B bake-off conclusion at 5B. NeMo-RL bringup also cost eight containerized-Ray-on-HPC fixes (Ray socket length, lustre flock, config schema, read-only venv patch, host-RAM cap, read-only inductor cache) before a single step; TRL ran clean. Both 5B chains therefore run on TRL.
The 5B chains (TRL, K-5* p1_rate, n=64)
Bounded 5B: 45.7 (SFT) → 51.1 → 53.0 → 54.3 → 54.6 (temp-1.0 plateau) → 56.0 (temp-2.0 unlock, +1.5); a second temp-2.0 re-band was flat. Beyond* 16.6. Unbounded 5B: 45.5 (SFT) → 53.8 → 57.2 → 58.5 → 59.6 (temp-1.0 plateau) → 60.2 (temp-1.3 unlock, +0.6; a smaller unlock because the temp-1.0 chain already climbed +14, leaving little unexplored headroom). Beyond* 28.9. Campaign best on both bands. The unbounded corpus's beyond-K-5 strength survives GRPO (Beyond* 28.9 vs the bounded 5B's 16.6), and temperature unlocks at the same corpus-dependent points as at 1.36B (bounded 2.0, unbounded 1.3), with the scale caveat that 2.0 is usable at 5B and 1.36B but blows the entropy guard at 0.6B (which unlocks at 1.5).
Dead ends
bf16-stored policy freezes (200 steps at lr 5e-6 and 1e-5, reward flat; fp32 masters fixed it at the same lr). 512-token rollout cap dropped for the 4k window (no-op for the baseline, verified). No re-band, coverage decays (pass@16 to -4.1 by step 600). Above-K-5 GRPO data does not help the 1.3B (ablation above): the base's headroom, not the data, bounds beyond-K-5 gains at this scale. Temperature past the entropy guard explodes, and the ceiling is policy-dependent: 2.5 at the 1.3B bounded, 2.0 at the 0.6B bounded and the unbounded bases, and 2.0 at the GSM8K-only policy (its entropy never collapsed as far). Temperature cannot substitute for pool coverage: the GSM8K-only bracket (1.5 / 1.7 / 2.0) was flat, flat, explosion. A third same-temp re-band is within noise (the first matches the plateau, the second climbs).
What this tab covers. Four GRPO engines ran the same experiment on the bounded 1.36B K-5 SFT model (output/sft_1p3b_sftcd_v2): the same banded 7,912-problem pool, the same gated official-grader binary reward, the same prompt and stop tokens, 32 prompts x 16 rollouts per optimizer step, lr 1e-5, fp32 master weights, temperature 1.0, a 3584-token cap, one 8xB200 node each, at least 200 optimizer steps, and engine defaults everywhere else. One engine per spec, defaults elsewhere, so this compares the frameworks and not the tuning. Eval is paired per-problem chat MathCAMPS against the same SFT baseline. The learning outcome is engine-independent. All four reach +5.16 to +5.82 K-5* pass@1 at 200 steps (within ~0.7 pt of +5.5); the engines differ 5 to 15x in per-step cost and in operational friction.
Results (200 optimizer steps, one 8xB200 node each)
engine
stack
train reward @200
eval K-5* p@1 Δ (paired 95% CI)
eval K-5* p@16 Δ (paired 95% CI)
s/step steady
8xB200 node-hours
blockers fixed
TRL 1.5.1 (incumbent)
vLLM 0.15, colocate (DDP)
0.40
+5.41 [+4.84, +6.01] (n=64)
+0.94 [+0.32, +1.84]
2.0-2.5
~0.5
3
verl 0.8.0
vLLM 0.12, ray + FSDP
0.432
+5.59 [+5.02, +6.15]
+1.89 [+0.91, +2.84]
12.3
~2.1
5
NeMo-RL v0.6.0
vLLM 0.20, ray + DTensor
0.44
+5.82 [+5.28, +6.34]
+2.03 [+1.06, +3.07]
11.5
~2.1
6
slime
Megatron + SGLang, ray
0.441-0.486
+5.16 [+4.53, +5.78]
+1.22 [+0.27, +2.18]
34.3
~3.6
6
Eval deltas are paired per-problem vs the same 4k-window SFT baseline (K-5* = mean of grade means G2-G5), 16 samples per problem for the three challengers and 64 for the TRL reference; the same-condition snapshot noise floor is ~0.5-0.9. The recipe sets the outcome, not the framework: fp32 masters, the policy-banded pool, the gated reward, and the batch geometry are the same across all four, and the four p@1 deltas land within ~0.7 pt of each other. The three ray-based engines' p@16 (+1.2 to +2.0) sits above TRL's (+0.9); the CIs overlap, so no claim is made. Node-hours for verl, NeMo-RL, and slime are the SYNTHESIS cost-line figures for the completed runs plus eval jobs; the TRL ~0.5 is the slime arm's comparison figure (incl. startup), not a SYNTHESIS value.
The four engines
TRL 1.5.1 (incumbent)
Architecture: the vLLM 0.15 engine is colocated and stays resident in the training process; weights are synced in-process after each optimizer step; 8-way DDP with no ray layer. At 1.36B on 183 GB the engine never has to sleep/wake. Blockers fixed (3): (1) bf16 freeze. With bf16-stored master weights the policy did not move (200 steps at lr 5e-6 and 1e-5 left the reward flat): GRPO gradients are noisy and near-zero-mean, so Adam steps land below the bf16 ulp and round away. fp32 masters + bf16 autocast at lr 1e-5 unfroze it. (2) JIT cache race. Cold-start kernel JIT raced across the 8 DDP workers in a shared cache. (3) 4k-window OOM. Moving rollouts to the 4k window (max_new_tokens 3584, max_model_len 4096) OOM'd the train pass next to the resident vLLM engine (fp32 logits over the 64k vocab at 16x4096 tokens plus uncheckpointed activations); fixed with gradient checkpointing, micro-batch 8 x grad-acc 8 (optimizer batch unchanged), and vLLM memory fraction 0.20. Numbers: reward 0.316 → 0.40 at step 200, entropy 0.15 → 0.08 over 300 steps, 2.0-2.5 s/step steady, ~0.5 8xB200 node-hours for the 210-step run including startup.
verl 0.8.0
Architecture: hybrid-flow, ray-orchestrated; an FSDP1 training engine (fp32 masters via MixedPrecision, bf16 compute) with a colocated vLLM 0.12 rollout worker that sleeps and wakes around each step, plus a separate old-log-prob recompute pass per step. Stack constraint: verl 0.8.0 pins vLLM ≤0.12 and transformers <5, a second env alongside the cluster's vLLM 0.15 / transformers 5.2 stack. Blockers fixed (5); the three specific to this stack: (1) Unconditional flash_attn import at every trainer step.verl/utils/attention_utils.py hard-imports flash_attn.bert_padding regardless of the attention backend, and there is no sm_100 flash-attn wheel. Fixed with a vendored pure-torch/einops shim on PYTHONPATH (with use_remove_padding=False the engine re-pads and runs standard sdpa). (2) Port race at actor init. verl binds port 0, closes the socket, and torch's TCPStore rebinds the recorded port later; a colocated vLLM/ray service won the race (EADDRINUSE). Fixed with a FileStore rendezvous (DIST_INIT_METHOD=file://). (3) Allocator conflict. PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True (cluster standard) trips a vLLM CuMemAllocator assertion in sleep mode (pytorch#147851); unset it. Numbers: reward 0.310 → 0.432 at step 200, s/step median 12.3, ~2.1 node-hours; checkpoint saves are 15-18 min each (synchronous fp32 model + optimizer to lustre).
NeMo-RL v0.6.0
Architecture: ray-orchestrated; a DTensor v1 policy (loads fp32, bf16 compute) with a colocated vLLM 0.20 generation backend that sleeps and wakes. A Megatron training backend with 6D parallelism and Megatron-native generation is available but was not used here. Stack constraint: the [fsdp]/[vllm] extras pull CUDA source builds for MoE/SSM/fp8. Blockers fixed (6); three representative: (1) Login-node OOM builds. The extras compile mamba-ssm, causal-conv1d, deep_ep, and deep_gemm (SSM/MoE/fp8 kernels, 9 GPU arches x parallel jobs), which OOM a 47 GB login node; deep_ep also needs NVSHMEM/libibverbs. None are imported on the dense-bf16 Qwen3 + DTensor-v1 + vLLM path, so all four were dropped from the extras; both worker venvs then built from prebuilt wheels in under 2 min. (2) Lustre flock. vLLM's source-patch step calls fcntl.flock on a lock file in the venv on lustre (errno 38, no flock); the lock was moved to a node-local dir keyed by path hash. (3) DeepGEMM warmup on Blackwell. vLLM's kernel_warmup gates deep_gemm on a sm_100 hardware check (not an import check) and attempted the fp8 warmup on the dropped package; VLLM_USE_DEEP_GEMM=0 (the run is bf16, no fp8). Numbers: reward 0.31 → 0.44, s/step median 11.5, ~2.1 node-hours; a standalone K-5 eval of hf_step200 gave overall pass@1 40.9 / pass@16 67.8 (finish_eos 1.0, parse_ok 1.0).
slime
Architecture: ray-orchestrated; a Megatron-LM training backend with separate single-GPU SGLang server processes behind a router (colocated here, 8 engines, per-step offload of the Megatron actor), supporting sync or async rollout. Stack constraint: ships a ~40 GB Docker image (slimerl/slime:latest, sglang v0.5.13-cu129) and requires an HF→Megatron torch_dist conversion; the parity gate passed exactly on our custom 64k tied-embedding Qwen3 (288/288 tensors, max_abs_diff 0). Blockers fixed (6); four representative: (1) Lustre image unpack. apptainer pull of the ~40 GB OCI image died on the login node (mksquashfs on lustre, millions of xattr warnings); the 18.8 GB SIF was built from the cached OCI layout on a compute node with APPTAINER_TMPDIR on node-local disk. (2) Proxy and ray. ray job submit to 127.0.0.1:8265 was routed through the cluster Squid proxy and denied (no_proxy=localnet does not cover loopback); loopback added to no_proxy. Separately the ray dashboard failed with 'AF_UNIX path length cannot exceed 107 bytes' under the condor scratch path; fixed with ray start --temp-dir=/tmp/rayslime. (3) sm_100 CUDA-graph crash. SGLang piecewise CUDA-graph capture hit an illegal memory access inside torch dynamo on sm_100; --sglang-disable-piecewise-cuda-graph (regular decode CUDA graphs stay on). (4) Async-save race. the HF export raced the async checkpoint finalization (the .metadata is published after 'successfully saved' is logged); export only after the train job exits or poll for .metadata. Numbers: reward 0.283 → 0.441 at step 200 (0.486 at 210), s/step median 34.3, ~3.6 node-hours; checkpoint save ~390 s each.
Step-time decomposition
The 5 to 15x per-step spread is orchestration, not math. TRL keeps the inference engine resident and syncs weights in-process, so it pays no sleep/wake, no KV reallocation, and no ray marshalling/broadcast per step. The three ray-based engines pay a fixed per-step cost (engine sleep/wake + KV realloc + weight refit or broadcast + ray marshalling) that is larger than the 1 to 2 s of actual short-completion generation; verl additionally runs a separate old-log-prob pass (1.93 s) and a padded-sdpa training forward (no flash-attn varlen on sm_100). Regime caveat: K-5 completions are ~90-130 tokens, so generation is under 2 s of GPU work and the fixed per-step orchestration dominates. At 10k+-token chains-of-thought the ratios compress and async designs (slime) begin to win by overlapping rollout and training. These frameworks target a workload this project does not have.
rollout generation wait 29.1 (85% of the step: 512 short completions across 8 single-GPU SGLang engines, slowest prompt-group tail dominates) + actor train 5.0
Which engine when
TRL: single-node dense models up to a few B params with short completions (the current 1.36B K-5 workload). Fastest per step here; the incumbent. verl: multi-node FSDP fleets and DAPO/GSPO recipe work. NeMo-RL (Megatron backend): RL on Megatron-pretrained larger models (the 5B candidate): it matches our Megatron-Core pretrain stack, Megatron-native generation removes the 4.1 s/step vLLM refit overhead measured above, and it supports TP/PP. slime: MoE or very large models with long-CoT async rollouts (SGLang serving features). Caveat: TRL's stay-resident advantage shrinks when VRAM is tight (a larger model plus KV cache cannot keep the engine resident), and TRL cannot scale past one node (DDP only). For a 5B RL run, evaluate the NeMo-RL Megatron backend.
Exports and per-arm reports
Each arm shipped an eval-validated hf_step200/ export (verl and slime 2.7 GB bf16 safetensors, NeMo-RL ~5.96 GB fp32), patched to the project's Qwen3 export standard (top-level rope_theta=1e6, eos_token_id [3,0]), and submitted the standard MathCAMPS eval. Reports: the four-way synthesis is /fast/mcorral/engine-bakeoff/SYNTHESIS.md; per-arm install friction, resolved config, metrics, and verdicts are in engine-bakeoff/{verl,slime,nemo-rl}/REPORT.md. The TRL reference is the grpo3 seg1 run (post-train/output/grpo3_1p3b_fp32lr1e5), detailed in the GRPO v3 tab.