$ train qwen2.5 --method grpo
GRPO-CoT for Qwen 2.5
A reinforcement-learning pipeline for Qwen2.5-3B using multi-objective rewards, 4-bit QLoRA, vLLM sampling and schema-aligned GRPO training on GSM8K.
output: structured reasoning / single-GPU fine-tuning / 4x improvement

