Giving something a sinister name doesn’t change what it is.
Didnt Altman say a few years ago that the next frontier for training would be synthetic data? It’s what everyone turned to (Tiny Stories, etc.). Using existing LLMs to generate more training data for new ones is the obvious step.
LLMs are offered as document generators. This is what people do. They use them to generate documents. And then they train on those documents. The labs do it themselves. They didn’t ask when they trained on the internet. And now the Chinese - and researchers, and hobbyist, and businesses - dont ask when they train on LLM outputs. Especially when they paid for them.
What they call “distillation” is the very knowledge flywheel that we want in society. You buy a book, you may learn from it, and you may write a better one. The author got paid when yiu bought it. Society gets paid when the ideas spread and lead to the creation of more books, products and services, all of which increase the choices that everyone has.
Distillation is so easy, so why is Google and Meta playing catch up with worse models than many Chinese labs? I mean, they could distill from the Chinese modles which are free? Why are the American labs models not top tier? Arcee, Laguna, Inkling? Why has Cohere or Mistral fallen behind? Our hubris would be the worse of us if we think the Chinese labs are keeping up because of their distillation usages.
Can someone walk me through exactly how powerful distillation is?
I've never seen a good blog post or metric establishing what % of a new model's "intelligence" can be captured. I'm not disputing the document's claim. Widespread effort at distillation is clear evidence that it is highly effective. I guess I just don't have an intuitive or technical sense for how powerful distillation is.
I wouldn't want to over state it, but from the point of view of a nation state excluded from a class of technology, acquiring it by means which look unusual and probably are is not inherently "wrong" -it probably breaks at least one other nations laws and attracts sanctions but you won't get any party from the USSR or the PRC or any other trade sanctioned economy past and present to agree "they did a wrong thing" because firstly states are essentially amoral and do things in the interest of the state good or bad, right or wrong, they are judged as fruitful or unfruitful and secondly although it's classic "whataboutism" in many ways, "you made me do it" is very strong.
"You only leapfrogged us by stealing our tech and copying it" is a very weak line of reasoning. "You did not contribute to the embedded R&D costs to get to this point" is dickering about the price attributed to all three of George Bernard Shaw, Oscar Wilde and Winston Churchill.
Add a dash of nimby and xenophobia and this is about right... you've got the likes of OpenAi and Anthropic destroying printed work from Austrailia and New Zealand to acquire latent data, Meta executives downloading commercial porn on their home computers and driving SSD's into work, all of them are building datascenters in metropolitan areas with generous tax breaks for their non-profitable ventures. The fact that 60% of model usage is now Chinese is what's prompting these folks to whine for protectionary measures. Meanwhile they will have no problem clearcutting massive swaths of meaningful work for people in their pursuit of AGI.
Kinda cool and very meta if Chinese AI companies are using distillation to liberate model value from US investment driven AI companies, by then releasing models created with distillation for free. Its circular IP and copyright avoidance and potential model improvement all the way down.
Parallels to how China bootstrapped their state manufacturing capacity with Western company joint ventures contributing capital and know how, and then gave them the boot once they no longer needed them.
Just finished reading "The Great Heist" which details across a number of manufactoring sectors with some wild stories the way China implemented their "Made in China 2025" policy.
Giving something a sinister name doesn’t change what it is.
Didnt Altman say a few years ago that the next frontier for training would be synthetic data? It’s what everyone turned to (Tiny Stories, etc.). Using existing LLMs to generate more training data for new ones is the obvious step.
LLMs are offered as document generators. This is what people do. They use them to generate documents. And then they train on those documents. The labs do it themselves. They didn’t ask when they trained on the internet. And now the Chinese - and researchers, and hobbyist, and businesses - dont ask when they train on LLM outputs. Especially when they paid for them.
What they call “distillation” is the very knowledge flywheel that we want in society. You buy a book, you may learn from it, and you may write a better one. The author got paid when yiu bought it. Society gets paid when the ideas spread and lead to the creation of more books, products and services, all of which increase the choices that everyone has.
Wait until they find out people use compiler output to write better compilers without asking for permission
It's compilers all the way down.
But so if there's distillation, and more distillation, and more distillation, what do we end up with ? A world of ELI5 ? "Grey zones are for losers" ?
Distillation is so easy, so why is Google and Meta playing catch up with worse models than many Chinese labs? I mean, they could distill from the Chinese modles which are free? Why are the American labs models not top tier? Arcee, Laguna, Inkling? Why has Cohere or Mistral fallen behind? Our hubris would be the worse of us if we think the Chinese labs are keeping up because of their distillation usages.
Can someone walk me through exactly how powerful distillation is?
I've never seen a good blog post or metric establishing what % of a new model's "intelligence" can be captured. I'm not disputing the document's claim. Widespread effort at distillation is clear evidence that it is highly effective. I guess I just don't have an intuitive or technical sense for how powerful distillation is.
It appears we're setting the scene for another round of hapless protectionist policy.
The time where even a token justification was necessary has long passed, so I'm surprised they bothered.
I wouldn't want to over state it, but from the point of view of a nation state excluded from a class of technology, acquiring it by means which look unusual and probably are is not inherently "wrong" -it probably breaks at least one other nations laws and attracts sanctions but you won't get any party from the USSR or the PRC or any other trade sanctioned economy past and present to agree "they did a wrong thing" because firstly states are essentially amoral and do things in the interest of the state good or bad, right or wrong, they are judged as fruitful or unfruitful and secondly although it's classic "whataboutism" in many ways, "you made me do it" is very strong.
"You only leapfrogged us by stealing our tech and copying it" is a very weak line of reasoning. "You did not contribute to the embedded R&D costs to get to this point" is dickering about the price attributed to all three of George Bernard Shaw, Oscar Wilde and Winston Churchill.
That is the longest run-on sentence I've seen in quite some time
They’re trying to kidnap what we have rightfully stolen.
O no! China is plundering the very plunder we worked very hard to plunder from everyone on earth.
Add a dash of nimby and xenophobia and this is about right... you've got the likes of OpenAi and Anthropic destroying printed work from Austrailia and New Zealand to acquire latent data, Meta executives downloading commercial porn on their home computers and driving SSD's into work, all of them are building datascenters in metropolitan areas with generous tax breaks for their non-profitable ventures. The fact that 60% of model usage is now Chinese is what's prompting these folks to whine for protectionary measures. Meanwhile they will have no problem clearcutting massive swaths of meaningful work for people in their pursuit of AGI.
Kinda cool and very meta if Chinese AI companies are using distillation to liberate model value from US investment driven AI companies, by then releasing models created with distillation for free. Its circular IP and copyright avoidance and potential model improvement all the way down.
Parallels to how China bootstrapped their state manufacturing capacity with Western company joint ventures contributing capital and know how, and then gave them the boot once they no longer needed them.
Just finished reading "The Great Heist" which details across a number of manufactoring sectors with some wild stories the way China implemented their "Made in China 2025" policy.
Great book, highly recommend.
Saw this firsthand when Deepseek's thinking model referenced Claude and Open AI guidelines in separate chats as part of the thinking process notes.
Time to run the same playbook back to their frontier labs.
If their frontier LLMs are the results of distilling US LLMs - then stealing it back won't work.
[flagged]