Multi-Reward RL: Scaling GDPO, CISPO, and REPO-R to 27B
- Authors

- Name
- Nino
- Occupation
- Senior Tech Editor
In our ongoing investigation into multi-reward reinforcement learning, we previously explored how methods like PPO, GRPO, and GDPO consolidate diverse reward signals into a unified objective. While initial benchmarks on a 14B model highlighted the dangers of overfitting—where training curves often mask poor generalization—this installment scales our efforts to a Qwen3.8-27B model, running for 600 steps to observe how these architectures behave at scale.
Scaling the Architecture: The 27B Experiment
We deployed a 27B model on a structured task governed by 12 distinct reward channels. Our goal was to isolate the impact of specific algorithmic choices: GDPO for reward normalization, CISPO for gradient stability, and REPO-R for entropy-aware advantage scaling.
By keeping the environment and reward structure consistent, we found that subtle changes in the loss function lead to vastly different outcomes on held-out tasks. For developers using n1n.ai to access high-performance LLM APIs, understanding these nuances is critical for effective fine-tuning.
The Advantage of CISPO and REPO-R
In our testing, swapping the traditional DAPO loss for CISPO improved held-out performance from -0.47 to -0.02. CISPO's approach—capping the importance weight at 1.2 rather than zeroing out gradients—prevents the model from 'switching off' during critical token decision points.
Adding REPO-R, which rescales token advantages based on their rarity, yielded our most significant jump, reaching a held-out score of +1.33. This suggests that in complex multi-reward landscapes, rewarding rare, high-quality tokens is a powerful lever for steering model behavior.
| Configuration | Shipped Step | Holdout Score | Improvement |
|---|---|---|---|
| DAPO (Baseline) | 600 | -0.47 | — |
| CISPO | 600 | -0.02 | +0.45 |
| CISPO + REPO-R | 600 | +1.33 | +1.35 |
| Filtered Prompts | 500 | +1.80 | +0.47 |
The Silent Killer: Advantage Floors
During our migration to a more demanding environment, we encountered a failure mode caused by an 'advantage floor' guard. This guard, intended to filter out low-quality answers, inadvertently discarded 96% of the learning signal because the model struggled to clear the threshold in the early stages.
This highlights a universal truth in RL: guards are environment-dependent. If your reward distribution shifts, your safety thresholds must be recalibrated. When scaling your own models using n1n.ai infrastructure, always monitor the percentage of non-zero advantages to ensure your trainer is actually learning.
Pro Tips for Multi-Reward Tuning
- Filter Your Prompts: Training on 500 prompts is often less effective than training on a curated subset of 200-250 that provide 'just enough' signal.
- Monitor Entropy: Use an entropy thermostat to adjust your REPO-R strength dynamically.
- Validate on Held-out Data: Never trust your training curve. Use a frozen probe to select your final checkpoint.
For teams building custom agents, integrating these advanced RL techniques can significantly elevate performance. If you are looking for reliable access to the latest models to test these strategies, n1n.ai provides the stable API foundation you need.
Get a free API key at n1n.ai