Post-training pipelines can now turn base models into capable reasoning assistants with far less overhead than many still assume. Historically, this work required not only compute, but also mature recipes, carefully curated data, multiple training stages, and substantial engineering effort. Meta's Llama 3.2 Instruct models were released before the current wave of progress in open post-training had fully taken shape. Our goal was to follow these newer open recipes in a practical setting and see how far they can now push a Llama 3.2 3B base model.

What this shows is how quickly open post-training has matured. In a short time, improvements in recipes, data, and training methodology have made strong post-training far more practical and reproducible. Using an open recipe inspired mainly by SmolLM3 style training, we reproduced the core stages of modern post-training (mid-training, supervised fine-tuning, preference optimization, and reinforcement learning) on a single compute node with one researcher in a matter of weeks. The broader takeaway is encouraging: sophisticated post-training is much more accessible, and open methods are advancing fast enough to close gaps that previously seemed out of reach.

## Background

LLMs have recently achieved remarkable gains by applying a **post-training pipeline** on top of base models. Projects like **Hugging Face's SmolLM3** (3B model) and **Allen AI's OLMo 3** (7B/32B models) demonstrate that sequential fine-tuning stages can transform a base model into a much more capable instruct and reasoning model. These pipelines typically involve:

1. An intermediate large scale mid-training on curated data to inject new skills, that consists of reasoning traces and instructions.
2. Supervised fine-tuning (SFT) on instruction-following demonstrations.
3. Preference optimization (e.g. Direct Preference Optimisation (DPO), Anchored Preference Optimisation (APO)) using chosen and rejected pairs.
4. Reinforcement learning (RL) on complex reasoning tasks for further improvement.

Inspired by these recipes, we set out to **apply the full post-training stack to Meta's Llama-3.2-1B** and elevate it towards the quality of its instruct-tuned version. Our goal was to reproduce the **multi-stage post-training flow** on a smaller model to validate how each stage contributes to performance gains. In doing so, we followed the blueprint from SmolLM3, adapting it to our infrastructure and budget constraints.

# Experiment Setup

We initialized with the Llama-3.2-1B base model and first modified its tokenization to be ChatML template friendly. Although the base Llama-3.2-1B model can support extended contexts up to ~128K tokens with long-context adaptations, most of the mid-training, SFT, and alignment stages were conducted using shorter context lengths to match the training setup. This involved mapping unused special tokens to the special tokens that appear in the **ChatML** format so that the model could easily ingest conversation-style prompts. We then executed four stages of training in sequence – **Mid-Training**, **SFT**, **APO**, and **RL** – using appropriate frameworks for each:

- **TRL (Transformers Reinforcement Learning)** library for Mid, SFT and preference tuning, leveraging built-in trainers and DeepSpeed ZeRO-3 for efficient training.
- **VeRL (ByteDance's open RL framework)** for the RL stage, which integrates Ray and PyTorch FSDP to manage large-scale rollout generation and policy optimization.

We fixed a compute budget corresponding to a **single node p4d.24xlarge** GPU **(8 × 40G A100 GPUs)** throughout all stages. Keeping the hardware fixed allowed us to plan memory usage and wall-clock time precisely for different context lengths and batch sizes. Each intermediate stage produced a new model checkpoint so that we could evaluate progress incrementally. Below, we describe each stage, the data used, and the observed impact on the model's capabilities.

## Evaluation Benchmarks

We evaluated our models at each stage using a suite of benchmarks targeting different capabilities:

- **GSM8K** (Grade School Math 8K): Tests multi-step arithmetic reasoning with word problems requiring 2–8 steps.
- **GSM Plus**: A harder variant of GSM8K with additional distractor information and more complex problem structures, testing robustness of mathematical reasoning.
- **GPQA Diamond (Graduate-Level Google-Proof QA)**: Expert-level science questions (biology, physics, chemistry) designed to be unsearchable—tests genuine reasoning and knowledge synthesis.
- **IFEval** (Instruction Following Evaluation): Measures how well models follow specific formatting and constraint instructions.

These benchmarks collectively assess: basic reasoning (GSM8K), robust reasoning (GSM Plus), advanced knowledge synthesis (GPQA), and instruction compliance (IFEval).

# Stage 1: Reasoning Mid‑Training

The first stage was the **mid-training** designed to inject the base model with general instruction following and reasoning skills before any solid supervised instruction tuning. This concept mirrors what SmolLM3 and OLMo3 did. We use SmolLM3's mid-training corpus of high-quality tokens focusing on domains like math, code, and logic.

We followed the chat template given in the SmolLM3 recipe. The major part to notice was, this chat template was a minimal one with both thinking and non-thinking models. The mid-training was run for multiple epochs over **~35B tokens (4-epoch mid-training on 140B tokens).** The context window was set to 32768 tokens. The mid-training loss curve reveals several important dynamics. The initial steep drop from ~1.4 to ~1.0 within the first 5K steps indicates rapid adaptation to the new data distribution. **After the Mid-training, we compared it with the base model on the following metrics.** Mid-training clearly improved the base model's **reasoning and long-context handling.**

**Benchmark Analysis:** Looking at the benchmark comparisons:

- **Math robustness clearly moved:** GSM8K and GSM-Plus jump sharply in **Think** mode (e.g., GSM8K ~0.02 → ~0.34; GSM-Plus ~0.03 → ~0.26).
- **Instruction following only nudged:** IFEval improves slightly.

# Stage 2: Supervised Fine‑Tuning (SFT)

We moved to **supervised fine-tuning**, where the model learns to follow instructions and engage in helpful dialogue. We used the SmolLM3 SFT dataset which was a blend of _instruction-following demonstrations and chat conversations_, covering skills like open QA, multi-turn dialogue, tool use, coding assistance, and multi-step reasoning. After about 4 epochs over ~2.5B tokens of SFT data, we obtained the SFT version.

**Benchmark Analysis:** Comparing SFT to Mid-trained checkpoints:

- **Primary win is instruction compliance:** IFEval jumps ~3×.
- **Math mostly preserved, not transformed:** GSM8K and GSM-Plus improve modestly, so SFT didn't "create" math skill.

# Stage 3: Preference Alignment via APO

To align the model's behaviour with desired responses, the next step was Direct Preference Optimization (DPO), using an **Anchored Preference Optimisation (APO)** variant. The preference dataset used in SmolLM3 consists of 1B chosen-rejected pairs.

**Benchmark Analysis:** DPO/APO results show nuanced improvements:

- **GSM8K is essentially unchanged vs SFT:** no real gain.
- **GPQA is the clearest regression signal (especially in Think)**. DPO-No-Think is closer, suggesting that the “think style” got worse for this benchmark.

# Stage 4: Reinforcement Learning with Verifiable Rewards (RL)

The final and most challenging stage was **Reinforcement Learning (RL)** fine-tuning, aimed at further boosting the model's performance on complex reasoning. We decided to go with a simple RLVR setup with sparse rewards, a GRPO style loss function, and a verifiable task schema.

**Benchmark Analysis:**

- **Largest gains are on the tasks we reward (math):** GSM8K: DPO(Think) 0.37 → RL ~0.61–0.72.
- **Non-math transfer is mixed (and sometimes flat):** GPQA varies by run and is not consistently higher than SFT.

## Lessons Learned

We conclude this single-node post-training journey with a clear takeaway: Mid-training and SFT reliably shape priors and instruction behavior; DPO/APO offers smaller, sometimes unstable, deltas under tight context constraints; and RL is the only stage that consistently produces a step-change in verifiable reasoning given we already have a solid model that has seen MID, SFT, and DPO data.

## Critical Success Factors

- **Evaluate for drift, not just benchmark scores.**
- **Data quality beats data quantity.**
- **Memory constraints materially change the experiment.**

## The Power (and Pain) of RL

RL is worth the investment: it's the only stage that produced genuine step-change improvements in reasoning accuracy, not just formatting or compliance.

But RL remains a “green field” for practitioners: 
- **Entropy dynamics are your early warning system**. 
- **Practical tuning philosophy:** Start conservative with batch sizes and KL constraints.
