The Journal
Writing
Papers and notes.
2026
Three-Layer Progressive Optimization Aug 15
Systematic Limit Approximation for Consumer-GPU GRPO Post-Training
DGBB, ZSBR, and SFOC on one RX 7900 XTX: 795.5 tok/s (97.6% of the wall argmax), +7.8 pp from ZSBR, and 23.2% → 45.6% held-out in 8.5 hours of pure RL.
Characterizing and Optimizing LLM Post-Training Efficiency Aug 14
On Consumer AMD RDNA3 GPUs
Dual-peak MFU, gap attribution, and a tuning recipe for GRPO on a single RX 7900 XTX. Unsloth +5.2% vs HuggingFace; 4-bit −40.7%; 97.7% of the GEMM-to-GRPO gap is generation.