Chunked KL loss, running Knowledge Distillation locally in less <6GB VRAM

2 points | by ikergarcia1996 4 hours ago

2 comments