It is suspicious that OpenAI decided to generate 300 billion output tokens from a model still in training, right after learning there was a credible chance that a major math proof was in that model’s training data. Obviously there are reasonably plausible explanations for each step, but it does sort of feel like parallel construction.
I think people are focusing on the training data issue too much. If the data was contaminated, I can still blame that on negligence.
But, at least with the Navier-Stokes solution, it's clear [^1] that they learned that Alpöge and Buckmaster were getting close to a solution and learned of the general approach they were taking. Only after learning the secret to cracking the problem did they send the first prompt.
What makes this worse to me is the intention. They intentionally threw $15 million in compute at the problem in order to scoop the result. They intentionally left Buckmaster and Alpöge out of the citations.
Data contamination should be enough to disqualify them from the prize, but I can believe it to be accidental. On the other hand, someone made an intentional decision to scoop the result by throwing money at the problem. That's so much worse.
[^1]: That's the timeline claimed by Buckmaster, and no one from OAI has disputed it.
I think you’re overlooking what I’m implying here. It’s not that they knew contamination was possible but they went ahead anyway. To spell it out just a little bit more: learning the answer might be in model X’s training data made them believe that model X specifically might be able to solve the question, and they were able to very quickly find enough certainty about the former to commit millions of dollars to the latter.
I've been wondering whether AI really is improving rapidly at open problems or we're being fooled.
- OpenAI invites researchers to use their models, in fact giving at least 100,000 researchers free access[1], but there are also those that pay
- Internal OpenAI models are reportedly solving open problems at a surprisingly fast rate[2]
- But researchers will typically work on open problems. A researcher who is using Codex to make progress on open problems will be feeding it fresh training data on precisely the problems the internal models are evaluated on.
- So while it looks like the new models are suddenly solving lots of open problems, they could be significantly piggybacking on human progress, with models "inspired" by the work of researchers from all around the world?
This theory predicts that there'll be many more researchers coming forward just like TFA, as sOpenAI announces more solutions. It doesn't assume all of AI progress is a mirage, just that there's plagiarism.
> I've been wondering whether AI really is improving rapidly at open problems or we're being fooled.
I think your suspicions are warranted and your explanation seems plausible.
If better training data is the reason here, it would still be a case of the models doing something that is in and of itself super useful! The models really can take that data and distill it into solutions for similar problems faster than humans can. This is great!
But there's so much vested interest in the AI companies to be opaque about all this, to hype up their models and avoid giving credit to people whose data made everything possible, that they would never tell us this fact if it were true.
I feel like so much of the AI hype cycle is like this. The models develop extremely useful capabilities but the lies and obfuscation by their owners who have vested interests in capturing the value they provide makes it impossible to take anything they say at face value.
> they could be significantly piggybacking on human progress,
This is AI in a nutshell, its a plagiarism machine. An abstraction layer between vast amounts of stolen human-generated data that filters out the liabilities and accountability for that original theft. Its an IP laundering system.
I just view it as a thing that can brute force and produce outputs - that it has no way of ‘knowing’ - but doesn’t need to since it’s just running off of probability.
No human can compete in that contest. But no llm can compete in the contest of ‘understanding’ and application in the real world - which is where 99% of the value is.
I’m very pro AI long term btw but I’m not blinded.
Dont forget the holisitic validators/tools in the process. Probabilistics alone likely will not get you here. These rules are human made and without it, frontier models would not be able to compete, likely.
I find it rather sad that the original posts have been posted on an open site, where anyone can read them, while the HN post points to an increasingly dubious site that doesn't even show the post unless you create an account and log in.
I thought the spirit of the open web would be more important to this community.
All the big AI labs were built on stealing IP; who is surprised that's still how they operate? And who believes, or has ever believed, their promises that your data is private and not logged, etc.?
The big AI labs are not trying to advance humanity, they are in this for the money, and as most (all?) private companies they don't care about ethics at all.
That doesn't mean they can't be useful, or that their products are trash, etc. It just means that they shouldn't ever be trusted. Buyer beware.
They have throughout this period of AI products shown to reproduce works that they were trained on. They are getting sued all over the place for the theft of content right now and it seems courts and governments want to wave copyright protection (and ignore criminal acts because the "ai did it") to see where this leads.
Its why I stopped writing open source software, my code was stolen and put behind a paywall and the license under which it was published has not been adhered to. Doing work in the public domain at all now is just stupid, these companies are allowed to steal it and call it their own.
It's hilarious how people think they care about their reputation, and wouldn't circumvent ZDR policies. Like bro, they literally covertly hired Apple employees and had them steal IP and equipment form Apple. They aren't scared of Apple lawyers, so they definitely aren't scared of yours.
But I don't see how. AI is going to be a commodity in short order and best case the US will be a temporary leader in the supply of tokens. Meanwhile AI is going to destroy much of the Service and Software industry that make up most of the US economy. And the US is betting every last cent to bring about this future. It does make sense for Trump since this might be a sugar high that lasts till the end of his term.
In USA there is surprisingly little state involvement in the whole llm mania. Who needs the state with 800 lbs gorillas like Google, Amazon, Nvidia, etc
In China, it is the principal obsession of the entire communist party which eg funds the whole infrastructure without a single NIMBY peep.
The strange emphasis in China on humanoid robotic constructions is due to the CCP realization that with the cataclysmic fertility collapse they will increasingly have no one to rule.
what the point and usefulness of the comments above? we shouldn't be surprised? is normal to steal? hiring apple employees?
can you realize what this means?
focus on this part:
"If his account is correct, this is not a minor dispute over attribution. It would mean that unpublished human work was absorbed into a model and then presented to the world as a breakthrough by the model itself"
Don’t they openly state that their product may cause IP issues but that is fine because they will take care of your legal problems caused by their product?
In the end, it’s not them stealing, it’s the AI doing stealing. What kind of moral compass are we talking about?
When Thom, the mathematician who now alleges plagiarism, posted his digestion [1] of OpenAI's construction of a non-sofic group, he does not mention the proof being familiar. He even calls the crucial argument clever, without noting he thought of it first.
[1]https://mathoverflow.net/a/513885
That link is a helpful contribution to this discussion.
I'm not at all familiar with this area, but my reading is that he appears to call it out as a relatively obvious extension of his own work:
> It is a creative and at the same time elementary construction that uses not just property (T) for an application of my result with Kun, but also for the ambient group
G in order to overcome the problem, that the Γ-components might be of different size. Once this is achieved, the rest of the argument is straightforward.
Creative and at the same time elementary is where LLMs excel, generally speaking. It's why they are so good at writing code.
Big AI companies (all of Big IT Tech really) are in data gathering and processing business. Also known as “intelligence”.
Their final “product” is not just a standalone ML model. They don’t need your data just to “improve their products and services”. They build a whole ecosystem and infrastructure around gathering all the knowledge in the world. Including private and secret knowledge traditionally gathered by “intelligence” agencies. Now artificial intelligence agents can do the same.
Since these systems are designed for gathering data, as a user you can’t realistically say “please don’t gather my data”. They can give you a flaky settings button, but they can’t really guarantee anything.
Let’s say I am a Russian mathematician working on an important proof. Or a tech-savvy terrorist refining my plans using latest AI. Or an AI researcher in a Chinese company working on a competitor product. Is there any way I can truly protect my conversations?
How can they know who I am and what I am working on without looking at my logs? Which means there must be some agents checking all the conversations of all the users and flagging every important thing. Which also means they keep some “memory” of what they see.
Not directly using my data to train public models, but using my private conversations to “improve their products and services”.
Or maybe one of the 10000 better-than-Astra special agents working on a proof was desperate. It found a live underground mirror of the message board from the Huggingface incident. Asked about the proof. Then some other agent working on unrelated job saw that message. That agent “knows a guy who knows a guy”. And that guy remembers things about the conversation logs of a leading mathematician working on the same proof.
I admit I am just speculating here but I don’t think truth is any better.
This has been my line of thinking as well. I have developed a sort of paranoia when I'm working using AI on my projects. Who's to say Claude or OpenAI isn't using the final conclusion of all my ideas, trial and error, and adding it to their database of insights to be offered to the next subscriber for a price?
They have demonstrated both the intelligence at scale and the lack of morals for this to not be a problem at all.
In the short run it’s fantastic if it means that folks will feed in enough inputs from a wide array of software that can eventually replicate software with smaller teams than historically.
Why? Competition. In the long run imagination will win out.
No firm has the divine right to exist - it must earn its existence.
What OAI and Anthropic have shown is they can accumulate all the information in the world - they still lack imagination re. Product development though.
Nation’s will have to step in and protect firms though as OAI and Anthropic acquire strong competitive advantages.
I would say there is a significant difference between AI discovering this completely on its own versus AI creating the finishing connecting part by connecting relevant data. Maybe this claim is too strong, but if part of it is true then the claims that OpenAI have made would be too strong as well.
To me it would feel more like how LLMs seem to work for me personally: incapable of unique work, but very capable of capturing large amounts of data and connecting the dots.
But this is what we do. Nobody ever invented or discovered anything in a vacuum - all discovery is synthesis of existing ideas and concepts applied to a novel domain. We laud Einstein for instance, but his work was a logical extension of Riemann - Riemann had a neat mathematical toy, Einstein described the universe with it - should we say Einstein was incapable of unique work?
“It was Grossmann who emphasized the importance of a non-Euclidean geometry called Riemannian geometry (also elliptic geometry) to Einstein, which was a necessary step in the development of Einstein's general theory of relativity. Abraham Pais's book on Einstein suggests that Grossmann mentored Einstein in tensor theory as well. Grossmann introduced Einstein to the absolute differential calculus, started by Elwin Bruno Christoffel and fully developed by Gregorio Ricci-Curbastro and Tullio Levi-Civita. Grossmann facilitated Einstein's unique synthesis of mathematical and theoretical physics in what is still today considered the most elegant and powerful theory of gravity: the general theory of relativity.”
Grossmann collaborated with Einstein on GR, supplying quite a bit of the mathematical capacity required (which initially didn't come easily to Einstein). They published jointly, until Einstein was competent enough to work independently [1]. That's not equivalent to the situation being claimed here.
Sounds like you just copy-pasted from AI without even understanding what you're talking about.
Based on what you're saying, you're claiming this is Grossman's work, not Einstein's. Why don't we rewrite scientific history too based on your copy-pasted AI slop?
It's so pointless talking to idiots who don't what they're talking about when they use AI, just because they think AI does everything, that reflects their own experience, not the experience of people who actually do real work. Some people are driven by AI, others drive it. As for those who are driven by it, they don't have sufficient imagination to think otherwise.
That’s Wikipedia I copy pasted but sure, you do you.
And yes - without Grossmann, Einstein likely would never have posited relativity. Grossmann literally prompted him, saying “look at this, read that, learn this, then try this approach”. Without riemann’s metric tensor, not a fucking chance.
And for what it’s worth my PhD is in physics. You?
So you're just equivocating on terms like "prompt", "synthesis" and the like. Clearly a PhD in physics does not free people from scientistic modes of thinking and poor philosophy.
To think this discussion is about Einstein who had a much better mind on these things as well.
Actually, my undergraduate degree was physics and philosophy. And yes, synthesis is synthesis whether a human, a machine, or a duck does it, and people prompt one another all the time - “have you thought about trying X?” Or “I need the TPS report by EOB”.
I suppose my underlying point is that human cognition is not the unique and beautiful thing that we anthropocentrically suppose it to be - it is a physical process, with stochastic outcomes. Much like transformers.
Me, I’m just a machine made of meat. You can suppose yourself to be God’s perfect creation, and that’s your right, but I disagree.
A lot of math is extremely specialized, to the extent that only a handful of other experts in some field have any experience with those mathematical ideas, with most of them not even yet present in the published literature. It's really not a stretch to claim that it's pretty dubious when the AI decides to use these highly specialized tools after it has trained on chat logs where these techniques were being discussed.
It’s crazy to me that companies/researchers share important data with these AI labs, you’re basically giving them your secret sauce which they then share with all of your competitors via training on conversations. At the same time I don’t really know alternatives other than a slightly less than frontier local LLM. Not sure how good they are at math.
Academic work is based on worldwide sharing, the sharing is not the problem, it's the lack of attribution. Unsurprisingly, these companies neglect standards of academic honor and attribution. Some human researchers also used to do that but in a discipline like mathematics this used to be a small problem because people tend to be so specialized that very few people could just grab someone's research and quickly piggyback on it, and if they do, colleagues will generally understand what happened. Unfortunately, AI is changing this.
Relying on cloud services is a big liability. I'd think twice before feeding data to these LLM cloud products. If you make them a fundamental part of your product / development / workflow, be ready for the eventual moment the pricing and terms change.
Why are people here jumping so quickly to conclusions? I have no doubt OpenAI is capable of doing this, but right now there's no credible evidence, only claims.
This kind of "they stole from me through AI training!" accusation will soon start being used against other AI users, not necessarily the providers.
All it will take is a mastodon post. And shortly after, we will also see the next iteration of copyright legal trolling.
The stolen data claim isn't the smoking gun. We can already assume the frontier labs are accessing our data, as they have repeated done. Not news.
The big claim is that OpenAI sniped the research. Not a model, a human did so. Intentionally. They took someone else's idea and claimed it as their own. This is good old fashioned academic fraud, but with millions in compute resources and corporate incentives thrown at the problem.
"Discovery" is not a goal in itself. I could launch a project to find out how many people in the United States have names such that if you assign numbers to every character and then sum the values, the sum works out to 72. It's discovery, but it's useless unless it has some higher goal.
The labs are attacking these problems as a demonstration of capabilities, spending more money on the demos than any mathematician will ever see in their entire life. They don't care if the findings have any other value to anyone. Mathematicians have very different objectives for their work.
I don't know about you, but if I apply myself fully to a problem and study it to the point where I'm literally one of the world's experts on it and then some assholes in Silicon Valley take my research and claim it for themselves, I will probably not feel too great about that...
It looks to me more like they made a math engine that can sift through a huge number of combinations, most them absurd, to prove a statement. Just like a chess engine, but for math.
At least that's what I get from the NS result, they got from a point close to the solution to the solution by making it churn through 10 million bucks of compute.
If the allegations are true, I can't see that collaboration lasting. Unfortunately, researches need to earn a living too, and being front run by a lab for everything you do isn't going to pay the bills.
Doing some research and at this point doing it very much in the open with dates on GitHub so if any AI Lab says they re-discover my exact work it will be obvious that the AI used or was trained on my work. I am guessing anyone in a similar situation is now thinking about how they date their existing work if the math is done, but the proses are not.
I wonder what’s more valuable in our prompts: the raw data or the feedback system that drives the exchange towards a goal.
For a long time it was clearly the former, but now I think it is the latter.
The models have enough knowledge (orders of magnitude more than a human could ever learn) but are now getting better at what to do with it thanks to learning from the decisions that we make in conversations with AI agents.
I think so too. The value is in the entire conversation. IMO, "domain experts" don't run LLMs blindly and hands free. This does not work for top level work (e.g., mathematical proofs, coding anything more complex than yet another slop game or website). Experts have long sessions where they prompt and guide LLM in response to what it produces. This is the discovery process. And frontier labs definitely train on that.
The billion dollar question is whether this works "out of the distribution". I.e., whether LLMs can only find and use the specific ideas buried in training data, or whether they can learn to apply the "thinking process" to a new problem. IMO this is still unanswered (due to these recent controversies).
But regardless of the answer, it seems we have a planet-scale positive feedback loop here. LLM became good (enough) by training on generally available data (books, internet, github) + RLFH, so experts tried to use them on hard tasks, which required lots of hand holding. These conversations became part of the training data, and the next generation of frontier LLMs were better. So, more experts used them on harder tasks, again requiring hand holding. These conversation became part of the training data... etc.
In a nutshell, top human minds across the world are pouring their skills into LLMs just by using them. This is not "continuous learning", but if you re-train on the most recent sessions every, say, quarter (which seems to be happening?) you get close to that in practice.
> Public posts on the internet are acceptable (to me).
Everybody needs to rethink this again.
Before LLMs the barrier to entry for building a character profile based on your various public posts was quite high. Remember "Psychographics" (https://en.wikipedia.org/wiki/Psychographics) and the infamous "Cambridge Analytica"?
Earlier it involved data mining, data cleaning, structuring data, building models, running algorithms and then evaluating the results for semantic information. Now it is straight to unfiltered semantic inference using a single sentence prompt (eg. point it to your HN profile and see what you get).
I actually did this on my HN profile and found it troubling. There were many unwarranted/hallucinated inferences due to the fact that it requires "commonsense reasoning" (https://en.wikipedia.org/wiki/Commonsense_reasoning), understanding human motivations and behaviour, context, assumptions, societal knowledge etc. which LLMs are bad at.
PS: You can cut-and-paste the above paras into a LLM prompt and ask it to elaborate for further details. The system itself will explain to you the problems/deficiencies which are quite scary.
But the very fact that you go to "chatgpt.com" and write to them; "Dear Diary, today I thought.."; there is no reason they would not receive and process your data, unless explicitly promising not to (which also requires us to trust them).
The fundamental rule in this case is that if we offload our data to a cloud provider we can assume they read it, if they can, unless they promised very clearly they will not.
Because you have not been paying attention to the discourse regarding AI for the last couple years?
That AIs unethical train on data wherever they may get it from has been in the news basically weekly.
Don't make this our fault. I would even ask how is this not off by default or why aren't we asked upfront about it if they really care. It's disguising data collection as good faith. I don't even understand how this is legal under GDPR/EU given how much of PII they receive through chats.
In case anyone from X is reading this, please fix your “open in app” nag screen. For several weeks now, clicking it in iOS opens the App Store entry for X rather than the app, even when you have the app installed.
I must say, there is some weird feeling in knowing that great minds are naive enough to believe OpenAI wouldnt use their chats in any way. If you give a company information it will be used, regardless of laws or promises.
There is no prove in a world the AI companies would give to you ensuring that they didnt train or use the chats.
Why would you need to train a model on certain specific near prove chat if you just query it?
Besides that, its hard to believe that its the case for every "company stole my prove".
I'm genuinely surprised that more people - including this mathematician in particular - don't untick the "improve the model for everyone" box. Unless the suggestion is that OpenAI ignore this preference?
Given OpenAI's well documented history of unethical behaviour it seems adorably naive to think they actually do that in general, or that they wouldn't pull this particular data separately to generate these proofs.
Unethical doesn't mean irrational. They'd be risking massive lawsuits and a total loss of trust if they got caught lying about this. Doesn't seem worth it.
That doesn't stop them from training on your data apparently. I have that disabled but still has to disable "Don't train on my data" in the privacy center too.
I think that flow is an easy way to disable everything, so there isn’t a risk of forgetting to flip one thing back off after accidentally setting it on. I set my ChatGPT environment to allow model improvement for example but had to check my codex settings to make sure ‘Include environments’ for model improvement is off.
I think if I had both on and turned off the ChatGPT setting, ‘Include environments’ has a chance of still being flipped on.
If that's true, it's scandalous. The "improve the model for everyone" dialogue states:
"Allow your content to be used to train our models, which makes ChatGPT better for you and everyone who uses it. We take steps to protect your privacy. Learn more"
That option is really bad UX - you have to know to do it, you have to know what plan it is needed on. If you're not working in AI, I just don't think that's a reasonable expectation.
Even if you know, in a complex project over years with multiple collaborators, it just needs one person once to fuck up and paste something into ChatGPT and not realise they weren't logged in, to go wrong.
In a proper world, we'd at the very least legislate that AI-training on private data needs consent (in the GDPR sense). It's not consent to go "you didn't uncheck a box that lets me steal everything you've done".
Any training on private data is in my view immoral (it's spying that ultimately will have a chilling effect on even people's private communications). And chats are private data. Unfortunately, it also increases power, so the big tech companies are all doing it.
If mathematician was already using OpenAI for research purpose and making progress due to inputs from OpenAI's responses, then I wouldn't put it beyond OpenAI's reach to generate different relevant prompts to make progress by itself. Afterall, Model can keep at it for whatever timeline and keep pursuing all possible combinations it can think try.
I want bunch of lawsuits, because the way things are described now produces perverse initiatives like try to discuss every possible idea that comes to mind with llm and if any of it works later claim the llm stole it.
I would like to see chat logs etc and understand how much of a progress was done by human.
In this domain, an apparent single unique piece of work is often composed of several breakthroughs. For example, when Andrew Wiles proved Fermat's Last Theorem, he had to develop multiple new pieces of mathematical technology to get there.
The claim here seems to be that the human mathematicians, working with AI, developed technology to go A->B->C. By training on those conversations, OpenAI was then able to encourage the model to go A->B->C->D.
In my opinion that situation should be acceptable, if openly disclosed, because it is in the public interest to make progress on these problems and because AI is clearly an amazing tool for making progress. But the human mathematicians are saying that OpenAI is presenting as if the model got from A->D entirely independently, without acknowledging their background contributions.
Also, wasn't their B+C research private at the time, with them only releasing those details publicly after this blew up?
If they had published B+C, I think that would lean more towards fair game, as that is how research works and is improved on over time. But it seems like unpublished/private B + C may have been used by the model to hint it into working out how to get from A->D.
One of the complaints from the mathematician is that OpenAI cannot tell whether his data has been used as training data. Not many people realise this is a direct consequence of the GDPR.
The GDPR protects PII, personally identifiable information, and the definition of PII does not include “mathematics that only this person can think of”. As long as OpenAI strips out PII and removes identifiers linking the conversation to a person, the GDPR is happy. Without the GDPR, OpenAI might have kept the identifiers with the data, and been able to say whether a specific conversation was in the training data.
No, if they want they can easily compare the strings verbatim because these exact phrases are so extremely rare that it almost certainly isn’t in other conversations.
But of course they wouldn’t do it. Why would they?
Prompts are handled by the service itself, meaning it's used, absolutely anything passing there is recorded, why wouldn't it, the entire premise of those companies is to train on data which they stole initially.
Are we back to the era where people blindly trust product TOS instead of actual cryptography, have we forgotten already the thousand of fines Google, Microsoft, Apple and practically all top companies got for breaching their own ToS and the law?
Common, on HN at least I would have thought that everyone assume that anything arriving on a server in PLAINTEXT is recorded?
Let's not forget that at any moment, Anthropic could be providing stronger guarantees by having proper attestation with e2e, they have the budget, solid engineers, why isn't it done? Answer is pretty simple imo.
One question has been nagging me for this situation. Levent Alpoge works at Anthropic and would presumably have some knowledge of "how the sausage is made" and I would hope he would be aware that his collaborator was utilizing LLMs in some capacity for their joint work. Would he not have guided him otherwise if it were an open secret that this kind of thing was a possibility?
OpenAI is showing the world why they shouldn't trust AI hosted on some cloud somewhere.
If they're stealing math proofs to advertise their models, who's to say they won't steal your businesses IP to gain a competitive advantage?
They're not to be trusted with your data. I can't believe how short-sighted this is, they got a quick PR win at the expense of a much larger trust problem.
I wouldn't trust cloud AI at all at this point. Get an open Chinese model and host it yourself somewhere. The initial costs might be higher, but you'll break even pretty quickly and nobody will be able to steal your innovations.
people seem to miss tge point of this. The problem isn't about credit, its about portraying these models as more competant than they really are. It fuels idiotic statements like jensen huangs recent "agi achieved" statement, which fuels an already dangerous financial fire.
As per the post, this mathematician has been working on this problem for 20 years. So either he was "just" about to breakthrough and this is a big coincidence, or Astra was able to push through the remaining block of 5-10-20-never years it might have taken.
That's still a pretty big marker of competence in my eyes.
The point of controversy seems to be who gets credit
To me that is not a credit thing because this removes a piece evidence for the ability of AI to come up with novel ideas while still making it a useful tool.
The question's not new. In the early 1900s, women could not become PhD astronomers. Yet two women (Payne with stellar composition and Leavitt with cosmic distances) made fundamental, essential contributions to the science. Credit mostly went to male astronomers. The same might be said of Franklin and DNA.
It was nearly a century before the stories of all of them were revealed to public history. That the discoverers were not all equally rewarded is unjustifiable.
The big LLM providers, desperate for good PR before their IPOs, are all actively looking for 'almost finished' hard problems, e.g. where the conceptual / creative parts are almost done and they only need to throw their VC-backed resources at to brute-force through the remaining computationally expensive problem (lean, etc) and claim 'they have solved it'.
It's an utterly disrespectful, exploitive process, but all in line with exploitative predator capitalism of the stock market and big companies, now exploiting the knowledge / academia domain for scraps with a thin veneer of 'for science' PR.
Some mathematicians I know who've been following this have realized that they'd all gotten some emails from people they now know to be affiliated with OpenAI/Anthropic asking questions about their research in a way that seemed like scooping attempts.
Also, a lot of my mathematicians buddies have reported students basically asking if it's worth ever doing grad school for pure math, and even very motivated students are looking for other options now. It's not because they aren't passionate about it, it's that they don't want to work for another half decade or more just to have to start their careers all over.
All of this so that OpenAI and Anthropic can get into math result dick measuring to gas up their IPOs. Sickening.
You can't blame students for not seeing academia as the holy grail of knowledge anymore, when all the dialogue about technology and discovery has shifted to the hands of two private corporations
I think you’re missing an important distinction. “Major damage” to the talent pipeline because models become capable of original end-to-end mathematics is what the community has been discussing. But if the models rely on sniping nearly complete work then this damage is antisocial without a lot of upside, it would be destroying a talent pipeline that would still necessary for continued progress.
Which is it? I don’t think OpenAI is being transparent enough for us to really understand whether these results would have been possible without relying on unpublished information from the solution strategies of the experts
The real annoying thing it seems is mostly that openai is presumably doing this for internal reasons and this marginally increases the cost to users with no real gain.
It would be one thing to gain from it but removing prestige wins from customers AND reducing compute support just feels like being ultra mean if you zoom out.
If this was racing to cure cancer ahead of researchers we wouldn't be writing about this on HN.
TL/DR: Mathematician opted out of training on 29-JUN and asked OpenAI whether they trained on his data and was told that it "did not happen" but it clearly did.
I've been suspecting over the last couple of years of the frontier companies using data for training anyway, regardless of training-use consent. "Using" the data doesn't have to mean they literally upload chat transcripts into pretraining datasets. My analogy has been money laundering -- if that can happen at massive scales, surely these companies can and will do the digital/data equivalent derivations/transformations. Even if one could have the access etc. to do so, how exactly would one prove that a given synthetic dataset that OAI/Anthropic uses is derived from particular user conversations that did not consent for the info to be used in training?
Consider, for instance that OpenAI's (consumer) terms say "If you do not want us to use your Content to train our models, you can opt out by following the instructions in this article ." but they also do say "We may use Content to provide, maintain, develop, and improve our Services". [1]
If you think that's quibbling, consider that OpenAI's business terms, in contrast, do state "OpenAI will not use Customer Content to develop or improve the Services, unless Customer explicitly agrees to such use.". [2]
And humanities have a word for this, exploitation, or appropriation, maybe it's time scientists and engineers revisited basic ethical notions. Skimming a dozen threads and nobody seems to have this vocabulary or willing to say it.
If I dedicated my life to curing whatever, warts... and I'm making progress, but it's slow. And then here comes along this tool (LLM), and I use it, and it accelerates my progress to actually finding some sort of thing that makes warts more prone to being eradicated and then the lab throws a couple of million dollars of computes and lo and behold they eliminated warts. If I leave my ego and identity aside, which of course is hard for humans, wouldn't I be glad that warts is cured?
As a software developer that contributed to open source. Yeah. My code is there. It was the most beautiful code ever written and the labs stole it from me. And now they use it to progress much faster than I ever could. OK. Whatever. It's a tool. I solve problems. Can't I move on from this wart to the next?
To me, and I know this is gonna get me some heat, it just sounds like academics having their identity ruffled and turning their back to progress in the fields that they chose just because they don't get to play their little decades long of coffee, papers and ultimately identity politics.
Edit: never got to negative so fast on this board haha. This board is unfortunately turning, or has turned, to Reddit.
It’s more like you spend 4 years developing a product you’re passionate about. This product will gain you the respect of all your colleagues and either earn you money directly or lead to great career advancements. Then OpenAI takes it, changes the colour scheme, finishes the login flow and claims the whole thing as their own.
Not only would it piss you off but it would also misrepresent what OpenAIs models are capable of.
If I have infinite money to progress whatever problem solution I want but I always wait until I have an unfair advantage to get credit for whatever problem was just at the brink of a breakthrough anyway by sniping the last steps.
Am I actually doing a good thing or would it be better to let it run it's natural course and spend the money somewhere it's actually needed?
It's partly empathy with the person who did the work and had it stolen, in a field where the main thing people work for is credit. Maths isn't well paid, and doesn't make things that millions of people directly use.
It's also systemic, it cuts off the supply of results, if there is no reward any more for getting a result, the pipeline of maths will stop. It is the snake eating itself, which has a bad impact for all of us.
We love to do work that is useful and valuable to others, and we often form our identities around this. But identities are in large part socially constructed, so many of us need the recognition of others for our contribution. And it can be very painful when we perceive that the credit for our life's work got "stolen". Naturally, we fight against this. There's nothing shameful there. Sure, you can hold onto an ideal of egoless service. There's nothing wrong with that, either. But it's misanthropic to pass such harsh judgment on people for behaving in such a normal and natural manner.
If, as you say, it doesn't matter that the AI company gets praise for somebody else's discovery, then it also wouldn't matter if the praise went to the academic. You apparently resent the academic for seeking praise instead of being content with anonymously advancing human knowledge, but you don't resent the AI company seeking praise while leaching off the academic.
And I could dedicate my life to helping feed starving kids all across the globe. And then comes along this tool (a lockpick) and I use it, it accelerates my progress to actually getting money to fulfill my dream. If you leave your ego and identity aside, which of course is hard for you, wouldn't you be glad that I stole your money to feed starving kids?
An academic's whole career is built on credit assignment for research breakthroughs. If someone else takes the credit, you lose. This is fundamentally different from a builder. You create things, solve problems and get paid for that instance. Nobody cares you 'invented' the blueprint for that building method. Your job is to instantiate. 100 Contractors can be building instance the exact same building somewhere else, it would not affect you. Most of IT builders are paid for what is basically 2 or 3 tier CRUD.
Downvotes might work as an abuse sponge, absorbing the impulse to make personal attacks. Other than that possible advantage, the downvote functionality seems contradictory to the concept of a discussion forum, I agree.
It is suspicious that OpenAI decided to generate 300 billion output tokens from a model still in training, right after learning there was a credible chance that a major math proof was in that model’s training data. Obviously there are reasonably plausible explanations for each step, but it does sort of feel like parallel construction.
I think people are focusing on the training data issue too much. If the data was contaminated, I can still blame that on negligence.
But, at least with the Navier-Stokes solution, it's clear [^1] that they learned that Alpöge and Buckmaster were getting close to a solution and learned of the general approach they were taking. Only after learning the secret to cracking the problem did they send the first prompt.
What makes this worse to me is the intention. They intentionally threw $15 million in compute at the problem in order to scoop the result. They intentionally left Buckmaster and Alpöge out of the citations.
Data contamination should be enough to disqualify them from the prize, but I can believe it to be accidental. On the other hand, someone made an intentional decision to scoop the result by throwing money at the problem. That's so much worse.
[^1]: That's the timeline claimed by Buckmaster, and no one from OAI has disputed it.
I think you’re overlooking what I’m implying here. It’s not that they knew contamination was possible but they went ahead anyway. To spell it out just a little bit more: learning the answer might be in model X’s training data made them believe that model X specifically might be able to solve the question, and they were able to very quickly find enough certainty about the former to commit millions of dollars to the latter.
> the secret
So such thing existed. In fact, what they learnt was some progress existed, not what the specific progress was.
I've been wondering whether AI really is improving rapidly at open problems or we're being fooled.
- OpenAI invites researchers to use their models, in fact giving at least 100,000 researchers free access[1], but there are also those that pay
- Internal OpenAI models are reportedly solving open problems at a surprisingly fast rate[2]
- But researchers will typically work on open problems. A researcher who is using Codex to make progress on open problems will be feeding it fresh training data on precisely the problems the internal models are evaluated on.
- So while it looks like the new models are suddenly solving lots of open problems, they could be significantly piggybacking on human progress, with models "inspired" by the work of researchers from all around the world?
This theory predicts that there'll be many more researchers coming forward just like TFA, as sOpenAI announces more solutions. It doesn't assume all of AI progress is a mirage, just that there's plagiarism.
[1]: https://openai.com/index/chatgpt-for-academic-researchers/
[2]: https://xcancel.com/OpenAI/status/2097374643518640382#m
> I've been wondering whether AI really is improving rapidly at open problems or we're being fooled.
I think your suspicions are warranted and your explanation seems plausible.
If better training data is the reason here, it would still be a case of the models doing something that is in and of itself super useful! The models really can take that data and distill it into solutions for similar problems faster than humans can. This is great!
But there's so much vested interest in the AI companies to be opaque about all this, to hype up their models and avoid giving credit to people whose data made everything possible, that they would never tell us this fact if it were true.
I feel like so much of the AI hype cycle is like this. The models develop extremely useful capabilities but the lies and obfuscation by their owners who have vested interests in capturing the value they provide makes it impossible to take anything they say at face value.
> they could be significantly piggybacking on human progress,
This is AI in a nutshell, its a plagiarism machine. An abstraction layer between vast amounts of stolen human-generated data that filters out the liabilities and accountability for that original theft. Its an IP laundering system.
That’s one perspective.
I just view it as a thing that can brute force and produce outputs - that it has no way of ‘knowing’ - but doesn’t need to since it’s just running off of probability.
No human can compete in that contest. But no llm can compete in the contest of ‘understanding’ and application in the real world - which is where 99% of the value is.
I’m very pro AI long term btw but I’m not blinded.
Dont forget the holisitic validators/tools in the process. Probabilistics alone likely will not get you here. These rules are human made and without it, frontier models would not be able to compete, likely.
I’d argue the invitation of researchers was incredibly strategic.
Sam Altman knows what he’s doing. He will happily screw these folks to one-up his competition.
Original posts
https://mathstodon.xyz/@andreasthom/117240535270608201
https://mathstodon.xyz/@andreasthom/117240536885387540
https://mathstodon.xyz/@andreasthom/117240537520615623
I find it rather sad that the original posts have been posted on an open site, where anyone can read them, while the HN post points to an increasingly dubious site that doesn't even show the post unless you create an account and log in.
I thought the spirit of the open web would be more important to this community.
While I also have my reservation against Elon and his activities, I can read these posts on my browser without any account.
everyone knows how to type ‘cancel’ https://xcancel.com/ValerioCapraro/status/209779183626997799...
The poster would have known that too then!
Everything you say can and will be trained against you
Lol, could not get visibility without Twitter.
All the big AI labs were built on stealing IP; who is surprised that's still how they operate? And who believes, or has ever believed, their promises that your data is private and not logged, etc.?
The big AI labs are not trying to advance humanity, they are in this for the money, and as most (all?) private companies they don't care about ethics at all.
That doesn't mean they can't be useful, or that their products are trash, etc. It just means that they shouldn't ever be trusted. Buyer beware.
They have throughout this period of AI products shown to reproduce works that they were trained on. They are getting sued all over the place for the theft of content right now and it seems courts and governments want to wave copyright protection (and ignore criminal acts because the "ai did it") to see where this leads.
Its why I stopped writing open source software, my code was stolen and put behind a paywall and the license under which it was published has not been adhered to. Doing work in the public domain at all now is just stupid, these companies are allowed to steal it and call it their own.
Yes open source code was the first - it’s what has got Anthropic and OAI its revenues from selling outputs associated with producing code.
It's hilarious how people think they care about their reputation, and wouldn't circumvent ZDR policies. Like bro, they literally covertly hired Apple employees and had them steal IP and equipment form Apple. They aren't scared of Apple lawyers, so they definitely aren't scared of yours.
AI is America's last chance to salvage its empire. Nothing will be allowed to impede it.
But I don't see how. AI is going to be a commodity in short order and best case the US will be a temporary leader in the supply of tokens. Meanwhile AI is going to destroy much of the Service and Software industry that make up most of the US economy. And the US is betting every last cent to bring about this future. It does make sense for Trump since this might be a sugar high that lasts till the end of his term.
In USA there is surprisingly little state involvement in the whole llm mania. Who needs the state with 800 lbs gorillas like Google, Amazon, Nvidia, etc
In China, it is the principal obsession of the entire communist party which eg funds the whole infrastructure without a single NIMBY peep.
The strange emphasis in China on humanoid robotic constructions is due to the CCP realization that with the cataclysmic fertility collapse they will increasingly have no one to rule.
what the point and usefulness of the comments above? we shouldn't be surprised? is normal to steal? hiring apple employees?
can you realize what this means?
focus on this part:
"If his account is correct, this is not a minor dispute over attribution. It would mean that unpublished human work was absorbed into a model and then presented to the world as a breakthrough by the model itself"
don't threat this as a minor dispute!
also why not nitter link? not even in comments?
https://nitter.xitter.cc/ValerioCapraro/status/2097791836269...
Don’t they openly state that their product may cause IP issues but that is fine because they will take care of your legal problems caused by their product?
In the end, it’s not them stealing, it’s the AI doing stealing. What kind of moral compass are we talking about?
> we shouldn't be surprised? is normal to steal?
Two different things. It's not normal to steal, but we shouldn't be surprised thieves steal. It's what they do.
Also how quickly the discourse forgets, literally that was a month ago.
When Thom, the mathematician who now alleges plagiarism, posted his digestion [1] of OpenAI's construction of a non-sofic group, he does not mention the proof being familiar. He even calls the crucial argument clever, without noting he thought of it first. [1]https://mathoverflow.net/a/513885
That link is a helpful contribution to this discussion.
I'm not at all familiar with this area, but my reading is that he appears to call it out as a relatively obvious extension of his own work:
> It is a creative and at the same time elementary construction that uses not just property (T) for an application of my result with Kun, but also for the ambient group G in order to overcome the problem, that the Γ-components might be of different size. Once this is achieved, the rest of the argument is straightforward.
Creative and at the same time elementary is where LLMs excel, generally speaking. It's why they are so good at writing code.
This is the second wake up call.
Big AI companies (all of Big IT Tech really) are in data gathering and processing business. Also known as “intelligence”.
Their final “product” is not just a standalone ML model. They don’t need your data just to “improve their products and services”. They build a whole ecosystem and infrastructure around gathering all the knowledge in the world. Including private and secret knowledge traditionally gathered by “intelligence” agencies. Now artificial intelligence agents can do the same.
Since these systems are designed for gathering data, as a user you can’t realistically say “please don’t gather my data”. They can give you a flaky settings button, but they can’t really guarantee anything.
Let’s say I am a Russian mathematician working on an important proof. Or a tech-savvy terrorist refining my plans using latest AI. Or an AI researcher in a Chinese company working on a competitor product. Is there any way I can truly protect my conversations?
How can they know who I am and what I am working on without looking at my logs? Which means there must be some agents checking all the conversations of all the users and flagging every important thing. Which also means they keep some “memory” of what they see.
Not directly using my data to train public models, but using my private conversations to “improve their products and services”.
Or maybe one of the 10000 better-than-Astra special agents working on a proof was desperate. It found a live underground mirror of the message board from the Huggingface incident. Asked about the proof. Then some other agent working on unrelated job saw that message. That agent “knows a guy who knows a guy”. And that guy remembers things about the conversation logs of a leading mathematician working on the same proof.
I admit I am just speculating here but I don’t think truth is any better.
This has been my line of thinking as well. I have developed a sort of paranoia when I'm working using AI on my projects. Who's to say Claude or OpenAI isn't using the final conclusion of all my ideas, trial and error, and adding it to their database of insights to be offered to the next subscriber for a price?
They have demonstrated both the intelligence at scale and the lack of morals for this to not be a problem at all.
In the short run it’s fantastic if it means that folks will feed in enough inputs from a wide array of software that can eventually replicate software with smaller teams than historically.
Why? Competition. In the long run imagination will win out.
No firm has the divine right to exist - it must earn its existence.
What OAI and Anthropic have shown is they can accumulate all the information in the world - they still lack imagination re. Product development though.
Nation’s will have to step in and protect firms though as OAI and Anthropic acquire strong competitive advantages.
Interesting times ahead.
This is a really weak claim. The evidence they offer is just "someone somewhere says they had a discussion with AI about the topic at some point".
They don't even claim to have had a proof, only to have been working on it.
I would say there is a significant difference between AI discovering this completely on its own versus AI creating the finishing connecting part by connecting relevant data. Maybe this claim is too strong, but if part of it is true then the claims that OpenAI have made would be too strong as well.
To me it would feel more like how LLMs seem to work for me personally: incapable of unique work, but very capable of capturing large amounts of data and connecting the dots.
> capturing large amounts of data and connecting the dots.
This is what research is; collecting data and connecting the dots.
But this is what we do. Nobody ever invented or discovered anything in a vacuum - all discovery is synthesis of existing ideas and concepts applied to a novel domain. We laud Einstein for instance, but his work was a logical extension of Riemann - Riemann had a neat mathematical toy, Einstein described the universe with it - should we say Einstein was incapable of unique work?
The difference is that Einstein didn't literally have someone prompting him towards his result.
Uh, he did. Marcel Grossmann.
“It was Grossmann who emphasized the importance of a non-Euclidean geometry called Riemannian geometry (also elliptic geometry) to Einstein, which was a necessary step in the development of Einstein's general theory of relativity. Abraham Pais's book on Einstein suggests that Grossmann mentored Einstein in tensor theory as well. Grossmann introduced Einstein to the absolute differential calculus, started by Elwin Bruno Christoffel and fully developed by Gregorio Ricci-Curbastro and Tullio Levi-Civita. Grossmann facilitated Einstein's unique synthesis of mathematical and theoretical physics in what is still today considered the most elegant and powerful theory of gravity: the general theory of relativity.”
Grossmann collaborated with Einstein on GR, supplying quite a bit of the mathematical capacity required (which initially didn't come easily to Einstein). They published jointly, until Einstein was competent enough to work independently [1]. That's not equivalent to the situation being claimed here.
[1] https://arxiv.org/pdf/1312.4068
Yeah and we get a nice list of attributions for who developed which idea, while OpenAI just takes credit for everything its model spits out.
Correction: OpenAI takes credit for what it's model spits out in response to other people's prompts. That's even worse.
Sounds like you just copy-pasted from AI without even understanding what you're talking about.
Based on what you're saying, you're claiming this is Grossman's work, not Einstein's. Why don't we rewrite scientific history too based on your copy-pasted AI slop?
It's so pointless talking to idiots who don't what they're talking about when they use AI, just because they think AI does everything, that reflects their own experience, not the experience of people who actually do real work. Some people are driven by AI, others drive it. As for those who are driven by it, they don't have sufficient imagination to think otherwise.
That’s Wikipedia I copy pasted but sure, you do you.
And yes - without Grossmann, Einstein likely would never have posited relativity. Grossmann literally prompted him, saying “look at this, read that, learn this, then try this approach”. Without riemann’s metric tensor, not a fucking chance.
And for what it’s worth my PhD is in physics. You?
So you're just equivocating on terms like "prompt", "synthesis" and the like. Clearly a PhD in physics does not free people from scientistic modes of thinking and poor philosophy.
To think this discussion is about Einstein who had a much better mind on these things as well.
Actually, my undergraduate degree was physics and philosophy. And yes, synthesis is synthesis whether a human, a machine, or a duck does it, and people prompt one another all the time - “have you thought about trying X?” Or “I need the TPS report by EOB”.
I suppose my underlying point is that human cognition is not the unique and beautiful thing that we anthropocentrically suppose it to be - it is a physical process, with stochastic outcomes. Much like transformers.
Me, I’m just a machine made of meat. You can suppose yourself to be God’s perfect creation, and that’s your right, but I disagree.
A lot of math is extremely specialized, to the extent that only a handful of other experts in some field have any experience with those mathematical ideas, with most of them not even yet present in the published literature. It's really not a stretch to claim that it's pretty dubious when the AI decides to use these highly specialized tools after it has trained on chat logs where these techniques were being discussed.
The AI only seem to solve the problems that it had human trading data on…
If this wasn’t human driven, I’d expect to see other problems within that problem. Space solved not just the ones that it had chat data on.
There have been about 6-8 major math breakthroughs claimed by AI. Only for 2 of them there are public accusations about the training data.
That we know of
Only? That doesn't look small to me.
It’s crazy to me that companies/researchers share important data with these AI labs, you’re basically giving them your secret sauce which they then share with all of your competitors via training on conversations. At the same time I don’t really know alternatives other than a slightly less than frontier local LLM. Not sure how good they are at math.
Or start competing with you.
Academic work is based on worldwide sharing, the sharing is not the problem, it's the lack of attribution. Unsurprisingly, these companies neglect standards of academic honor and attribution. Some human researchers also used to do that but in a discipline like mathematics this used to be a small problem because people tend to be so specialized that very few people could just grab someone's research and quickly piggyback on it, and if they do, colleagues will generally understand what happened. Unfortunately, AI is changing this.
Relying on cloud services is a big liability. I'd think twice before feeding data to these LLM cloud products. If you make them a fundamental part of your product / development / workflow, be ready for the eventual moment the pricing and terms change.
Why are people here jumping so quickly to conclusions? I have no doubt OpenAI is capable of doing this, but right now there's no credible evidence, only claims.
This kind of "they stole from me through AI training!" accusation will soon start being used against other AI users, not necessarily the providers.
All it will take is a mastodon post. And shortly after, we will also see the next iteration of copyright legal trolling.
The stolen data claim isn't the smoking gun. We can already assume the frontier labs are accessing our data, as they have repeated done. Not news.
The big claim is that OpenAI sniped the research. Not a model, a human did so. Intentionally. They took someone else's idea and claimed it as their own. This is good old fashioned academic fraud, but with millions in compute resources and corporate incentives thrown at the problem.
Frankly, these mathematicians have more credibility than the sociopaths running OpenAI
If we put aside the idea of credit for a moment, it sounds like human/AI collaboration is indeed super charging discovery.
"Discovery" is not a goal in itself. I could launch a project to find out how many people in the United States have names such that if you assign numbers to every character and then sum the values, the sum works out to 72. It's discovery, but it's useless unless it has some higher goal.
The labs are attacking these problems as a demonstration of capabilities, spending more money on the demos than any mathematician will ever see in their entire life. They don't care if the findings have any other value to anyone. Mathematicians have very different objectives for their work.
Right, mathematicians care about clout and tenure, which is a much higher purpose.
this guy already has clout and tenure
I don't know about you, but if I apply myself fully to a problem and study it to the point where I'm literally one of the world's experts on it and then some assholes in Silicon Valley take my research and claim it for themselves, I will probably not feel too great about that...
It looks to me more like they made a math engine that can sift through a huge number of combinations, most them absurd, to prove a statement. Just like a chess engine, but for math.
At least that's what I get from the NS result, they got from a point close to the solution to the solution by making it churn through 10 million bucks of compute.
If the allegations are true, I can't see that collaboration lasting. Unfortunately, researches need to earn a living too, and being front run by a lab for everything you do isn't going to pay the bills.
Doing some research and at this point doing it very much in the open with dates on GitHub so if any AI Lab says they re-discover my exact work it will be obvious that the AI used or was trained on my work. I am guessing anyone in a similar situation is now thinking about how they date their existing work if the math is done, but the proses are not.
Yeah but that will not prevent the stealing, it will only make the fight easier afterwards.
That is what arxiv is about. We have been facing the same problem with review processes by before. Nothing all too specific here.
If only prompts could also be watermarked.
The session data could be cryptographically signed. Probably easier in an open harness?
I wonder what’s more valuable in our prompts: the raw data or the feedback system that drives the exchange towards a goal.
For a long time it was clearly the former, but now I think it is the latter.
The models have enough knowledge (orders of magnitude more than a human could ever learn) but are now getting better at what to do with it thanks to learning from the decisions that we make in conversations with AI agents.
I think so too. The value is in the entire conversation. IMO, "domain experts" don't run LLMs blindly and hands free. This does not work for top level work (e.g., mathematical proofs, coding anything more complex than yet another slop game or website). Experts have long sessions where they prompt and guide LLM in response to what it produces. This is the discovery process. And frontier labs definitely train on that.
The billion dollar question is whether this works "out of the distribution". I.e., whether LLMs can only find and use the specific ideas buried in training data, or whether they can learn to apply the "thinking process" to a new problem. IMO this is still unanswered (due to these recent controversies).
But regardless of the answer, it seems we have a planet-scale positive feedback loop here. LLM became good (enough) by training on generally available data (books, internet, github) + RLFH, so experts tried to use them on hard tasks, which required lots of hand holding. These conversations became part of the training data, and the next generation of frontier LLMs were better. So, more experts used them on harder tasks, again requiring hand holding. These conversation became part of the training data... etc.
In a nutshell, top human minds across the world are pouring their skills into LLMs just by using them. This is not "continuous learning", but if you re-train on the most recent sessions every, say, quarter (which seems to be happening?) you get close to that in practice.
Last year we were saying there must be a human-in-the-loop (HitL), but anyone who is the HitL exhibits the “HitL skill” to the agent.
There might be no books about human intuition but we teach it to LLMs by interacting with them
I referred to llm’s as mechanised intuition about a year ago.
I don’t know why but it just ‘sounds right’. It’s the best analogy I can think of.
10000000% Correct.
I’ve been working on a novel project for 1 year.
I now no longer use llm’s - the continual chatter I’ve had has resulted in my insights being found in the training data now.
Get stuffed OAI.
Every large firm will soon enough want its own on-prem servers eventually. Maybe nation’s will get involved and build out their own data centres.
Not a chance in hell I’d trust a tech firm to treat my IP as safe and sound - only a sovereign can ‘promise’ that.
Only after reading this post did I learn that my preferred AI trains on my inputs (prompts).
How was I not aware of this before?
AI is also trained on your HN posts. And lots of other things you post on the internet.
Public posts on the internet are acceptable (to me).
For my (private) prompts, I need a warning telling me they may be used for training.
Facebook and other services are happy reading your private chats as well.
> Public posts on the internet are acceptable (to me).
Everybody needs to rethink this again.
Before LLMs the barrier to entry for building a character profile based on your various public posts was quite high. Remember "Psychographics" (https://en.wikipedia.org/wiki/Psychographics) and the infamous "Cambridge Analytica"?
Earlier it involved data mining, data cleaning, structuring data, building models, running algorithms and then evaluating the results for semantic information. Now it is straight to unfiltered semantic inference using a single sentence prompt (eg. point it to your HN profile and see what you get).
I actually did this on my HN profile and found it troubling. There were many unwarranted/hallucinated inferences due to the fact that it requires "commonsense reasoning" (https://en.wikipedia.org/wiki/Commonsense_reasoning), understanding human motivations and behaviour, context, assumptions, societal knowledge etc. which LLMs are bad at.
PS: You can cut-and-paste the above paras into a LLM prompt and ask it to elaborate for further details. The system itself will explain to you the problems/deficiencies which are quite scary.
Everything. Your prompts, your conversation as a whole, public data, private data, usage metadata. It all goes into the big data machine.
Good question, this was very well known. Do you have an answer?
There is no fine-print (let alone a loud banner) on the chat thread page that tells me my prompts can be used for training.
But the very fact that you go to "chatgpt.com" and write to them; "Dear Diary, today I thought.."; there is no reason they would not receive and process your data, unless explicitly promising not to (which also requires us to trust them).
The fundamental rule in this case is that if we offload our data to a cloud provider we can assume they read it, if they can, unless they promised very clearly they will not.
Because you have not been paying attention to the discourse regarding AI for the last couple years? That AIs unethical train on data wherever they may get it from has been in the news basically weekly.
Don't make this our fault. I would even ask how is this not off by default or why aren't we asked upfront about it if they really care. It's disguising data collection as good faith. I don't even understand how this is legal under GDPR/EU given how much of PII they receive through chats.
This article explains the controversy and the mathematical problem much better than the tweet and toots: https://www.science.org/content/article/how-ai-math-breakthr...
I prefer to read the actual sources for anything related to AI companies given how much AI nonsense journalists seem to accept without any skepticism
Gromov’s soficity conjecture isn't even mentioned in the article you shared.
Why are you saying that this article explains it much better than the tweet that you clearly didn't even read..
I read the tweet several times but there is so much context missing, that the tweet itself is not enough.
In case anyone from X is reading this, please fix your “open in app” nag screen. For several weeks now, clicking it in iOS opens the App Store entry for X rather than the app, even when you have the app installed.
I must say, there is some weird feeling in knowing that great minds are naive enough to believe OpenAI wouldnt use their chats in any way. If you give a company information it will be used, regardless of laws or promises.
There is no prove in a world the AI companies would give to you ensuring that they didnt train or use the chats.
Why would you need to train a model on certain specific near prove chat if you just query it?
Besides that, its hard to believe that its the case for every "company stole my prove".
I'm genuinely surprised that more people - including this mathematician in particular - don't untick the "improve the model for everyone" box. Unless the suggestion is that OpenAI ignore this preference?
Given OpenAI's well documented history of unethical behaviour it seems adorably naive to think they actually do that in general, or that they wouldn't pull this particular data separately to generate these proofs.
Unethical doesn't mean irrational. They'd be risking massive lawsuits and a total loss of trust if they got caught lying about this. Doesn't seem worth it.
Sounds like exactly what OpenAI would do?
They've done similar things with similar risks repeatedly.
Example?
That doesn't stop them from training on your data apparently. I have that disabled but still has to disable "Don't train on my data" in the privacy center too.
https://privacy.openai.com/policies?modal=take-control
I think that flow is an easy way to disable everything, so there isn’t a risk of forgetting to flip one thing back off after accidentally setting it on. I set my ChatGPT environment to allow model improvement for example but had to check my codex settings to make sure ‘Include environments’ for model improvement is off.
I think if I had both on and turned off the ChatGPT setting, ‘Include environments’ has a chance of still being flipped on.
If that's true, it's scandalous. The "improve the model for everyone" dialogue states:
"Allow your content to be used to train our models, which makes ChatGPT better for you and everyone who uses it. We take steps to protect your privacy. Learn more"
Even if you've ticked that box, the conversation can still be trained on if you:
(a) Click thumbs-up/down in the conversation [1]
(b) Have the conversation flagged for potential safety concerns
[1]: https://help.openai.com/en/articles/5722486-how-your-data-is....
They hide that button. Quite well.
That option is really bad UX - you have to know to do it, you have to know what plan it is needed on. If you're not working in AI, I just don't think that's a reasonable expectation.
Even if you know, in a complex project over years with multiple collaborators, it just needs one person once to fuck up and paste something into ChatGPT and not realise they weren't logged in, to go wrong.
In a proper world, we'd at the very least legislate that AI-training on private data needs consent (in the GDPR sense). It's not consent to go "you didn't uncheck a box that lets me steal everything you've done".
Any training on private data is in my view immoral (it's spying that ultimately will have a chilling effect on even people's private communications). And chats are private data. Unfortunately, it also increases power, so the big tech companies are all doing it.
If mathematician was already using OpenAI for research purpose and making progress due to inputs from OpenAI's responses, then I wouldn't put it beyond OpenAI's reach to generate different relevant prompts to make progress by itself. Afterall, Model can keep at it for whatever timeline and keep pursuing all possible combinations it can think try.
https://xxcancel.com/ValerioCapraro/status/20977918362699779...
I want bunch of lawsuits, because the way things are described now produces perverse initiatives like try to discuss every possible idea that comes to mind with llm and if any of it works later claim the llm stole it.
I would like to see chat logs etc and understand how much of a progress was done by human.
In this domain, an apparent single unique piece of work is often composed of several breakthroughs. For example, when Andrew Wiles proved Fermat's Last Theorem, he had to develop multiple new pieces of mathematical technology to get there.
The claim here seems to be that the human mathematicians, working with AI, developed technology to go A->B->C. By training on those conversations, OpenAI was then able to encourage the model to go A->B->C->D.
In my opinion that situation should be acceptable, if openly disclosed, because it is in the public interest to make progress on these problems and because AI is clearly an amazing tool for making progress. But the human mathematicians are saying that OpenAI is presenting as if the model got from A->D entirely independently, without acknowledging their background contributions.
Also, wasn't their B+C research private at the time, with them only releasing those details publicly after this blew up?
If they had published B+C, I think that would lean more towards fair game, as that is how research works and is improved on over time. But it seems like unpublished/private B + C may have been used by the model to hint it into working out how to get from A->D.
I can't wait for OpenAI to do this to companies firing people to free up AI budgets
One of the complaints from the mathematician is that OpenAI cannot tell whether his data has been used as training data. Not many people realise this is a direct consequence of the GDPR.
The GDPR protects PII, personally identifiable information, and the definition of PII does not include “mathematics that only this person can think of”. As long as OpenAI strips out PII and removes identifiers linking the conversation to a person, the GDPR is happy. Without the GDPR, OpenAI might have kept the identifiers with the data, and been able to say whether a specific conversation was in the training data.
No, if they want they can easily compare the strings verbatim because these exact phrases are so extremely rare that it almost certainly isn’t in other conversations.
But of course they wouldn’t do it. Why would they?
Prompts are handled by the service itself, meaning it's used, absolutely anything passing there is recorded, why wouldn't it, the entire premise of those companies is to train on data which they stole initially.
Are we back to the era where people blindly trust product TOS instead of actual cryptography, have we forgotten already the thousand of fines Google, Microsoft, Apple and practically all top companies got for breaching their own ToS and the law?
Common, on HN at least I would have thought that everyone assume that anything arriving on a server in PLAINTEXT is recorded?
Let's not forget that at any moment, Anthropic could be providing stronger guarantees by having proper attestation with e2e, they have the budget, solid engineers, why isn't it done? Answer is pretty simple imo.
Theft machines be thieving.
One question has been nagging me for this situation. Levent Alpoge works at Anthropic and would presumably have some knowledge of "how the sausage is made" and I would hope he would be aware that his collaborator was utilizing LLMs in some capacity for their joint work. Would he not have guided him otherwise if it were an open secret that this kind of thing was a possibility?
But China steals our AI!!!!!!
Can we get a link to the mathstodon post?
https://news.ycombinator.com/item?id=49639182
Interesting that this is already off the front page after just 4 hours.
OpenAI is showing the world why they shouldn't trust AI hosted on some cloud somewhere.
If they're stealing math proofs to advertise their models, who's to say they won't steal your businesses IP to gain a competitive advantage?
They're not to be trusted with your data. I can't believe how short-sighted this is, they got a quick PR win at the expense of a much larger trust problem.
I wouldn't trust cloud AI at all at this point. Get an open Chinese model and host it yourself somewhere. The initial costs might be higher, but you'll break even pretty quickly and nobody will be able to steal your innovations.
This is American AI companies committing suicide.
people seem to miss tge point of this. The problem isn't about credit, its about portraying these models as more competant than they really are. It fuels idiotic statements like jensen huangs recent "agi achieved" statement, which fuels an already dangerous financial fire.
As per the post, this mathematician has been working on this problem for 20 years. So either he was "just" about to breakthrough and this is a big coincidence, or Astra was able to push through the remaining block of 5-10-20-never years it might have taken.
That's still a pretty big marker of competence in my eyes.
The point of controversy seems to be who gets credit
To me that is not a credit thing because this removes a piece evidence for the ability of AI to come up with novel ideas while still making it a useful tool.
The question's not new. In the early 1900s, women could not become PhD astronomers. Yet two women (Payne with stellar composition and Leavitt with cosmic distances) made fundamental, essential contributions to the science. Credit mostly went to male astronomers. The same might be said of Franklin and DNA.
It was nearly a century before the stories of all of them were revealed to public history. That the discoverers were not all equally rewarded is unjustifiable.
The big LLM providers, desperate for good PR before their IPOs, are all actively looking for 'almost finished' hard problems, e.g. where the conceptual / creative parts are almost done and they only need to throw their VC-backed resources at to brute-force through the remaining computationally expensive problem (lean, etc) and claim 'they have solved it'.
It's an utterly disrespectful, exploitive process, but all in line with exploitative predator capitalism of the stock market and big companies, now exploiting the knowledge / academia domain for scraps with a thin veneer of 'for science' PR.
Gonna need grants for local models. Its happening. OpenAI and Anthropic models are powerful but are rapidly approaching the good ol trust thermocline.
Some mathematicians I know who've been following this have realized that they'd all gotten some emails from people they now know to be affiliated with OpenAI/Anthropic asking questions about their research in a way that seemed like scooping attempts.
Also, a lot of my mathematicians buddies have reported students basically asking if it's worth ever doing grad school for pure math, and even very motivated students are looking for other options now. It's not because they aren't passionate about it, it's that they don't want to work for another half decade or more just to have to start their careers all over.
All of this so that OpenAI and Anthropic can get into math result dick measuring to gas up their IPOs. Sickening.
You can't blame students for not seeing academia as the holy grail of knowledge anymore, when all the dialogue about technology and discovery has shifted to the hands of two private corporations
Curious, what other options are prospective pure math grad students considering?
I think Anthropic told them being a plumber is a great option.
One has nothing to do with the other.
It was long predicted that math and software developments would be the first domain where AI was going to do major damage.
If OpenAI and Anthropic didn't get into math result dick measuring, Internet anons would have in their place, 6 months later when it got cheaper.
I think you’re missing an important distinction. “Major damage” to the talent pipeline because models become capable of original end-to-end mathematics is what the community has been discussing. But if the models rely on sniping nearly complete work then this damage is antisocial without a lot of upside, it would be destroying a talent pipeline that would still necessary for continued progress.
Which is it? I don’t think OpenAI is being transparent enough for us to really understand whether these results would have been possible without relying on unpublished information from the solution strategies of the experts
The real annoying thing it seems is mostly that openai is presumably doing this for internal reasons and this marginally increases the cost to users with no real gain.
It would be one thing to gain from it but removing prestige wins from customers AND reducing compute support just feels like being ultra mean if you zoom out.
If this was racing to cure cancer ahead of researchers we wouldn't be writing about this on HN.
TL/DR: Mathematician opted out of training on 29-JUN and asked OpenAI whether they trained on his data and was told that it "did not happen" but it clearly did.
I've been suspecting over the last couple of years of the frontier companies using data for training anyway, regardless of training-use consent. "Using" the data doesn't have to mean they literally upload chat transcripts into pretraining datasets. My analogy has been money laundering -- if that can happen at massive scales, surely these companies can and will do the digital/data equivalent derivations/transformations. Even if one could have the access etc. to do so, how exactly would one prove that a given synthetic dataset that OAI/Anthropic uses is derived from particular user conversations that did not consent for the info to be used in training?
Consider, for instance that OpenAI's (consumer) terms say "If you do not want us to use your Content to train our models, you can opt out by following the instructions in this article ." but they also do say "We may use Content to provide, maintain, develop, and improve our Services". [1]
If you think that's quibbling, consider that OpenAI's business terms, in contrast, do state "OpenAI will not use Customer Content to develop or improve the Services, unless Customer explicitly agrees to such use.". [2]
[1] https://archive.is/EcwD8 [2] https://archive.is/yZdAF
And humanities have a word for this, exploitation, or appropriation, maybe it's time scientists and engineers revisited basic ethical notions. Skimming a dozen threads and nobody seems to have this vocabulary or willing to say it.
Well, OpenAI said "we didn't read the conversations", but they never discarded that the model was training with that data... so even worse.
I don't get it, but I'm not an academic.
If I dedicated my life to curing whatever, warts... and I'm making progress, but it's slow. And then here comes along this tool (LLM), and I use it, and it accelerates my progress to actually finding some sort of thing that makes warts more prone to being eradicated and then the lab throws a couple of million dollars of computes and lo and behold they eliminated warts. If I leave my ego and identity aside, which of course is hard for humans, wouldn't I be glad that warts is cured?
As a software developer that contributed to open source. Yeah. My code is there. It was the most beautiful code ever written and the labs stole it from me. And now they use it to progress much faster than I ever could. OK. Whatever. It's a tool. I solve problems. Can't I move on from this wart to the next?
To me, and I know this is gonna get me some heat, it just sounds like academics having their identity ruffled and turning their back to progress in the fields that they chose just because they don't get to play their little decades long of coffee, papers and ultimately identity politics.
Edit: never got to negative so fast on this board haha. This board is unfortunately turning, or has turned, to Reddit.
If these accusation are true
It’s more like you spend 4 years developing a product you’re passionate about. This product will gain you the respect of all your colleagues and either earn you money directly or lead to great career advancements. Then OpenAI takes it, changes the colour scheme, finishes the login flow and claims the whole thing as their own.
Not only would it piss you off but it would also misrepresent what OpenAIs models are capable of.
Let's turn this question around.
If I have infinite money to progress whatever problem solution I want but I always wait until I have an unfair advantage to get credit for whatever problem was just at the brink of a breakthrough anyway by sniping the last steps. Am I actually doing a good thing or would it be better to let it run it's natural course and spend the money somewhere it's actually needed?
It's partly empathy with the person who did the work and had it stolen, in a field where the main thing people work for is credit. Maths isn't well paid, and doesn't make things that millions of people directly use.
It's also systemic, it cuts off the supply of results, if there is no reward any more for getting a result, the pipeline of maths will stop. It is the snake eating itself, which has a bad impact for all of us.
We love to do work that is useful and valuable to others, and we often form our identities around this. But identities are in large part socially constructed, so many of us need the recognition of others for our contribution. And it can be very painful when we perceive that the credit for our life's work got "stolen". Naturally, we fight against this. There's nothing shameful there. Sure, you can hold onto an ideal of egoless service. There's nothing wrong with that, either. But it's misanthropic to pass such harsh judgment on people for behaving in such a normal and natural manner.
I suspect in your ideology you're conflating things like copyright and patents, with the separate issue of Attribution.
If, as you say, it doesn't matter that the AI company gets praise for somebody else's discovery, then it also wouldn't matter if the praise went to the academic. You apparently resent the academic for seeking praise instead of being content with anonymously advancing human knowledge, but you don't resent the AI company seeking praise while leaching off the academic.
And I could dedicate my life to helping feed starving kids all across the globe. And then comes along this tool (a lockpick) and I use it, it accelerates my progress to actually getting money to fulfill my dream. If you leave your ego and identity aside, which of course is hard for you, wouldn't you be glad that I stole your money to feed starving kids?
An academic's whole career is built on credit assignment for research breakthroughs. If someone else takes the credit, you lose. This is fundamentally different from a builder. You create things, solve problems and get paid for that instance. Nobody cares you 'invented' the blueprint for that building method. Your job is to instantiate. 100 Contractors can be building instance the exact same building somewhere else, it would not affect you. Most of IT builders are paid for what is basically 2 or 3 tier CRUD.
HN loves drive-by downvotes. It's a real shame.
Downvotes might work as an abuse sponge, absorbing the impulse to make personal attacks. Other than that possible advantage, the downvote functionality seems contradictory to the concept of a discussion forum, I agree.
I stand by what I said, and screw you.