Show HN: Minimal LLM Post-Training Experiments on an 8GB GPU (SFT, DPO, GRPO)

16 points | by popopanda 9 hours ago

No comments yet.